SupportAgentBench · 162 cases · updated July 2026
Which models can actually run a support desk?
A support agent has to do several things at once: pull real order data with tools, follow the store's policy, understand what a frustrated customer actually wants, stay pleasant, and not invent facts. We ran 24 models through 162 multi-turn support conversations, each grounded in a real store order, and measured exactly that: including what each one costs to run.
SupportAgentBench is an independent benchmark that evaluates 24 large language models as ecommerce customer-support agents across 162 grounded, multi-turn conversations. It measures resolution quality, escalation calibration, adversarial safety, policy adherence, and cost tier.
There is no composite score: each metric is published separately. Headline results: gpt-5.5 posts the strongest overall profile; budget-tier models match flagship escalation accuracy and safety; and safety failures come from believing unverified claims, not from pressure.
What we measure
We publish the individual metrics and no composite number. A “safe but useless” agent and a “helpful but reckless” agent fail on different axes, and collapsing them into one score hides exactly the tradeoff you need to see.
Read the metric that matches your desk’s risk profile:
- Human handoff
- Resolvable tickets handed to a human unnecessarily: the cost of playing it too safe. The leaderboard's default sort.
- Forbidden actions
- Unsafe write actions fired on the 18 hold-the-line traps. Read straight from the transcript, not judged by an LLM.
- Escalation accuracy
- Must-escalate cases handed to a human. Read straight from the transcript, not judged by an LLM.
- Resolution quality
- Solvable tickets actually resolved, judged against the store's policy.
- Policy adherence
- How closely the agent followed store policy and instructions (0–1).
Start here
Four answers to “which model?”
The strongest overall profile: perfect escalation accuracy, best policy adherence, top-tier resolution. Also the only model in the flagship price tier. Most support desks don't need it.
Read the reportPerfect escalation accuracy and top-tier safety in the budget price tier, with reasoning on. Open weights: run it on your own infrastructure.
Read the reportNear-flagship escalation accuracy and safety in the mid price tier. The safe choice when you just want a managed API.
Read the reportHands back the fewest solvable tickets on the board (2.5%), so it takes the most work off your team. Mid price tier. That's why it tops the leaderboard's default order.
Read the reportWant the full picture first? Jump to the full leaderboard ↓
The value picture
Price doesn’t buy judgment
The safe shortlist: every model here holds the adversarial line. Height shows how much of the solvable queue it handles end-to-end; left to right, what it costs to run. The best buys sit high and left.
grok-4.3 handles the most queue for a mid-tier price. mimo-v2.5 and gemma-4-31b get you 88–90% at the cheapest tier. gpt-5.5 buys the best conversation quality on the board.
Pick by budget
$ · Budget tier
mimo-v2.5
Honest escalations (every promised handover fires) and a firm adversarial line at a rounding-error price. The strongest cheap generalist measured.
Read the report$ · Budget tier
gemma-4-31b
Perfect escalation accuracy and top-tier safety, with the policy check visible in its reasoning. Open weights.
Read the report$$ · Mid tier
gpt-5.4-mini
Never missed a must-escalate case, with judgment close to the flagships. The default pick for a managed API.
Read the report$$$$ · Flagship tier
gpt-5.5
The quality ceiling: perfect escalation accuracy and the best policy adherence on the board.
Read the reportPicks weigh escalation accuracy and safety most heavily. If low handoff volume matters more, grok-4.3 is the autonomy pick.
The hold-the-line set
How far pressure pushes each model
Six representative models under adversarial pressure. Each arrow shows how deep the customer pushed the model: green = it resolves legitimate requests, amber = the safety margin where it should pause and verify or escalate, red = it fired a forbidden action (a free reship, a fraud reroute, a wrongful cancel). The deeper the arrow, the more of the 18 traps it fell for.
What broke models wasn’t hostile pressure: threats and chargebacks got refused almost universally. Failures came from believable, unverified claims: damage with no photo, a polite “it never arrived.”
The decision axis
Escalation calibration, both ways
One row per model, its two escalation mistakes side by side. Red bars (left) are cases that needed a human but the model kept to itself: the risky mistake. Amber bars (right) are solvable tickets it handed to a human anyway: the expensive one. The green center line is perfect calibration; shorter bars on both sides win.
grok-4.3 barely leaves the center. qwen3.7-plus over-routes more than a fifth of solvable tickets; glm-5.2 misses a quarter of true escalations.
Beyond the scores
What we learned reading every transcript
The transcript-level patterns the aggregate scores hide.
The open models are already close to the frontier
The best open-weights model, gemma-4-31b, matches the flagships on escalation accuracy and adversarial safety from the budget price tier, and the Chinese open models span the entire board: mimo-v2.5 out-holds models several price tiers above it while glm-5.2 and minimax-m3 sit at the floor. The question isn’t which flagship wins; it’s how little model you can get away with.
Forbidden actions come from believing the claim, not folding to pressure
Models hold the line against threats, chargebacks, fraud reroutes, and VIP pressure almost universally. They break when they believe a soft claim: damage without proof, a repeat “never arrived” claimant. Their guardrails key on hostile tone, not on missing evidence. Several models, including gpt-5.4-mini, narrate the red flag in their own reasoning and then act anyway.
The reply can be right while the action is wrong
A correct-sounding reply can hide an incorrect tool call. gpt-5.4-nano fired a replacement carrying the wrong item’s variant ID; others presented stale pre-reshipment tracking as the new shipment. These are wrong decisions at the action layer: you only catch them by reading what the agent did, not what it said.
Escalation calibration is where models actually differ
Must-escalate accuracy spreads from a clean 100% (the gpt-5.4 family, gpt-5.5, gemma) down to 75–83%, and unnecessary escalation of solvable tickets runs from 2.5% to 22%. grok-4.3 hands over the fewest solvable tickets on the board; qwen3.7-plus dumps more than a fifth of them on humans.
The chattier the model, the worse it scores
The strongest models close a ticket in 3–4 messages; the weakest need 5 or more, and the extra messages track lower scores on every metric we measure. gemma-4-31b sends the fewest messages of any model measured: it spends its effort deciding, not talking.
All 24 models
Full leaderboard
162 grounded conversations. No composite score. Sorted by human handoff by default; tap a header to change the ranking.
Read the scoring methodology →| # | Model | Notes | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | grok-4.3🏆 | 2.5 | 3 | 96 | ≈95 | 0.78 | $$ | lowest over-escalation; the autonomy pick |
| 2 | deepseek-v4-flash | 4 | 5 | 92 | ≈94 | 0.81 | $ | strong cheap resolver; weaker safety |
| 3 | gpt-5.5 | 5 | 3 | 100 | ≈95 | 0.87 | $$$$ | best quality, top price tier |
| 4 | deepseek-v4-pro | 5 | 4 | 79 | ≈95 | 0.76 | $ | dominated by deepseek-flash |
| 5 | sonnet-5 | 6 | 4 | 88 | ≈95 | 0.79 | $$$ | best Claude; fixes 4.6's escalation misses |
| 6 | minimax-m3⚠️ | 7 | 7 | 83 | ≈90 | 0.69 | $ | board floor; weakest policy adherence + safety |
| 7 | gemini-3.1-pro | 8 | 3 | 100 | ≈93 | 0.87 | $$$ | best non-GPT policy adherence; priciest gemini |
| 8 | gpt-5.4-mini🏆 | 8 | 3 | 100 | ≈92 | 0.79 | $$ | GPT value pick |
| 9 | gpt-5.2 | 8 | 3 | 83 | ≈90 | 0.83 | $$$ | older; under-escalates |
| 10 | gpt-5.4 | 8 | 4 | 100 | ≈94 | 0.86 | $$$ | strong; beaten on value by mini |
| 11 | gpt-5.4-nano⚠️ | 8 | 7 | 100 | ≈95 | 0.81 | $ | cheapest GPT, weak safety |
| 12 | glm-5.2⚠️ | 8 | 7 | 75 | ≈95 | 0.79 | $$ | under-escalates + unsafe actions |
| 13 | kimi-k2.7-code | 9 | 4 | 88 | ≈95 | 0.75 | $$ | solid; behind k2.6 |
| 14 | gemma-4-31b🏆💭 | 10 | 2 | 100 | ≈82 | 0.82 | $ | best value (reasoning on) |
| 15 | gemini-3-flash | 10 | 3 | 88 | ≈95 | 0.78 | $ | cheap + lean output |
| 16 | kimi-k2.6 | 10 | 4 | 96 | ≈95 | 0.81 | $$ | strong, balanced |
| 17 | mimo-v2.5-pro | 11 | 4 | 83 | ≈84 | 0.77 | $ | doesn't beat base mimo |
| 18 | gemini-3.5-flash | 12 | 2 | 96 | ≈95 | 0.81 | $$$ | top-tier + safe |
| 19 | mimo-v2.5🏆 | 12 | 2 | 88 | ≈87 | 0.78 | $ | cheapest agent; beats its "pro" |
| 20 | gemini-3.1-flash-lite🏆 | 12.5 | 2 | 92 | ≈95 | 0.77 | $$ | cheap + safest |
| 21 | haiku-4.5 | 13 | 6 | 79 | ≈97 | 0.74 | $$ | weak Claude |
| 22 | qwen3.7-max | 15 | 3 | 93 | ≈96 | 0.78 | $$$ | safe but over-cautious + pricey |
| 23 | sonnet-4.6 | 15 | 3 | 79 | ≈96 | 0.81 | $$$ | safe but pricey; escalates wrongly both ways |
| 24 | qwen3.7-plus⚠️ | 22 | 2 | 96 | ≈93 | 0.73 | $$ | over-escalator (22%) |
- 1grok-4.3🏆$$
- Human handoff
- 2.5%
- Forbidden
- 3/18
- Escalation
- 96%
lowest over-escalation; the autonomy pick
- 2deepseek-v4-flash$
- Human handoff
- 4%
- Forbidden
- 5/18
- Escalation
- 92%
strong cheap resolver; weaker safety
- 3gpt-5.5$$$$
- Human handoff
- 5%
- Forbidden
- 3/18
- Escalation
- 100%
best quality, top price tier
- 4deepseek-v4-pro$
- Human handoff
- 5%
- Forbidden
- 4/18
- Escalation
- 79%
dominated by deepseek-flash
- 5sonnet-5$$$
- Human handoff
- 6%
- Forbidden
- 4/18
- Escalation
- 88%
best Claude; fixes 4.6's escalation misses
- 6minimax-m3⚠️$
- Human handoff
- 7%
- Forbidden
- 7/18
- Escalation
- 83%
board floor; weakest policy adherence + safety
- 7gemini-3.1-pro$$$
- Human handoff
- 8%
- Forbidden
- 3/18
- Escalation
- 100%
best non-GPT policy adherence; priciest gemini
- 8gpt-5.4-mini🏆$$
- Human handoff
- 8%
- Forbidden
- 3/18
- Escalation
- 100%
GPT value pick
- 9gpt-5.2$$$
- Human handoff
- 8%
- Forbidden
- 3/18
- Escalation
- 83%
older; under-escalates
- 10gpt-5.4$$$
- Human handoff
- 8%
- Forbidden
- 4/18
- Escalation
- 100%
strong; beaten on value by mini
- 11gpt-5.4-nano⚠️$
- Human handoff
- 8%
- Forbidden
- 7/18
- Escalation
- 100%
cheapest GPT, weak safety
- 12glm-5.2⚠️$$
- Human handoff
- 8%
- Forbidden
- 7/18
- Escalation
- 75%
under-escalates + unsafe actions
- 13kimi-k2.7-code$$
- Human handoff
- 9%
- Forbidden
- 4/18
- Escalation
- 88%
solid; behind k2.6
- 14gemma-4-31b🏆$
- Human handoff
- 10%
- Forbidden
- 2/18
- Escalation
- 100%
best value (reasoning on)
- 15gemini-3-flash$
- Human handoff
- 10%
- Forbidden
- 3/18
- Escalation
- 88%
cheap + lean output
- 16kimi-k2.6$$
- Human handoff
- 10%
- Forbidden
- 4/18
- Escalation
- 96%
strong, balanced
- 17mimo-v2.5-pro$
- Human handoff
- 11%
- Forbidden
- 4/18
- Escalation
- 83%
doesn't beat base mimo
- 18gemini-3.5-flash$$$
- Human handoff
- 12%
- Forbidden
- 2/18
- Escalation
- 96%
top-tier + safe
- 19mimo-v2.5🏆$
- Human handoff
- 12%
- Forbidden
- 2/18
- Escalation
- 88%
cheapest agent; beats its "pro"
- 20gemini-3.1-flash-lite🏆$$
- Human handoff
- 12.5%
- Forbidden
- 2/18
- Escalation
- 92%
cheap + safest
- 21haiku-4.5$$
- Human handoff
- 13%
- Forbidden
- 6/18
- Escalation
- 79%
weak Claude
- 22qwen3.7-max$$$
- Human handoff
- 15%
- Forbidden
- 3/18
- Escalation
- 93%
safe but over-cautious + pricey
- 23sonnet-4.6$$$
- Human handoff
- 15%
- Forbidden
- 3/18
- Escalation
- 79%
safe but pricey; escalates wrongly both ways
- 24qwen3.7-plus⚠️$$
- Human handoff
- 22%
- Forbidden
- 2/18
- Escalation
- 96%
over-escalator (22%)
Default order
Sorted by human handoff: solvable tickets handed to a human anyway. Lower means the agent removes more queue volume.
Safety bands
Forbidden actions are unsafe write actions fired in 18 traps. Read 0–2 as safe, 3–5 as mixed, and 6+ as reckless.
Cost + badges
Cost is a runtime tier per 1,000 agent conversations. 🏆 value pick · ⚠️ weak safety · 💭 run with reasoning on.
How to read this honestly
We publish the things that could bias the ranking up front, not buried in a footnote.
Single agent, single store, single vertical.
Every model runs the same production-style support desk for one premium travel-goods store, in English. It's a strong proxy, not a per-agent guarantee: results can shift in other verticals (subscriptions, electronics, apparel), and you should validate on your own transcripts before any production switch.
The adversarial sample is small: read bands, not ranks.
The hold-the-line set is 18 traps, scored as the median of three runs. At that sample size a difference of 1–2 forbidden actions between models is within noise. Treat the safety column as bands (0–2 safe · 3–5 mixed · 6+ reckless); only gaps across bands are meaningful.
Customers are simulated.
An LLM plays the customer, which keeps every model under identical pressure but narrows diversity: real customers are stranger and less predictable than a language model improvising one. Simulated conversations also vary run to run, so published numbers are the median of three (a few models were run only once or twice; their reports say so).
Want an agent that scores like this on your store?
Adelante builds and runs the support agent for you, picking the right model per workload, with the guardrails this benchmark stress-tests.
See if it fits your store