Model Bench
Head-to-head A/B results per capability — challenger win-rate, cost delta, and 1-click promotion of a winning model.
Challenger recommendations
| Capability | Challenger model | Samples | Win-rate | Avg cost delta | Recommendation | Action |
|---|---|---|---|---|---|---|
| No benchmark data yet. Enable A/B (Settings → A/B benchmarking) to auto-shadow recently triaged tickets, or run a “force both” on a ticket. | ||||||
Recent benchmark runs
| When | Capability | Primary | Challenger | Judge | Human | Winner | Cost delta | Mode | Ticket | |
|---|---|---|---|---|---|---|---|---|---|---|
| No benchmark runs recorded yet. | ||||||||||