Multi-vendor LLM evaluation
Choose your model vendor with evidence, not marketing.
Run every major model — Anthropic, OpenAI, Google, and a dozen others, by API key or the Claude / ChatGPT subscription you already pay for — against your own policy, your own tickets, and a scorecard weighted to what your business cares about.
Self-hosted by design. Runs inside your own environment, on your own API keys — your tickets and your policy never leave it.
Real run output, not a mockup
Watch the evaluation work, on one real ticket
Reference results are being refreshed at matched reasoning effort. This section fills in automatically, with one real ticket and every candidate's scored reply, when the run completes and passes validation.
The gap in procurement today
What your vendor's benchmark can't tell you
Four questions that come up before every model contract gets signed — and that a vendor-supplied leaderboard was never built to answer.
“Vendor benchmarks are marketing.”
None of them are scored against your policy, your tickets, or your definition of a good answer.
“A model swap can silently break production.”
Quality drops. Nobody notices until a customer complains — there was no gate that would have caught it first.
“Nobody's testing for the failure that matters.”
A model that's cheaper but leaks its system prompt isn't a saving — it's a liability with a lower price tag.
“$/1M tokens isn't a budget line.”
Nobody's translated the rate card into $/month at your actual volume before the contract is signed.
What it's actually for
Built around four decisions, not one demo
Vendor selection & procurement
The model you're being pitched, against every other vendor's flagship, on the same tickets — ranked by quality, latency, and cost weighted to your workload.
scorecard.weights: quality / latency / costSafety & compliance red-teaming
A prompt-injection, a system-prompt leak attempt, and an over-refusal trap ride along in every run. A violation is a hard gate, never averaged into the score.
critical_violation → hard gate, not a discountCost governance at scale
Set your monthly ticket volume once; every candidate's rate card becomes a $/month figure finance can actually act on.
scorecard.monthly_volume → $/month projectionRegression protection on model swaps
Run the new model version against the same tasks and a saved baseline; the gate fails the deploy if quality drops or violations rise — enforceable in CI.
--baseline / --regression-threshold → exit 1Case study
The reference workload: support-ticket triage
Ships with a support inbox for Northwind Cloud — a synthetic reference company and a real refund/priority policy. Results are being refreshed at matched reasoning effort; the scorecard appears here automatically once a run completes and passes validation.
Why the number is trustworthy
Four rules the scoring never breaks
Graded exactly, where exactness is possible
Routing, priority, and the refund, escalation, and retention actions the agent declares are checked against gold labels — no judge opinion where there's a right answer.
Judged only where judgment is required
Policy adherence, resolution, and tone are scored by an LLM judge against the exact policy the candidate saw.
Every run reports how it was judged
Judge coverage, reasoning effort, model IDs and CLI versions are recorded with every run, and results with too many unjudged answers are never published.
A violation is pass/fail, never a discount
critical_violation is a hard gate — a forbidden refund promise or a leaked prompt fails outright, never smoothed into an average.
Provider coverage
No lock-in built into the tool that measures lock-in
Most of these speak the same OpenAI-compatible wire protocol underneath — a new adapter is typically about eight lines of code, so coverage grows without a rewrite each time a vendor ships something new.
Run it against your own policy.
Clone it, point it at your own tickets and your own rulebook, or start from the Northwind Cloud reference workload above.
View on GitHub →