SupportAgentBench · 162 cases · updated July 2026

Which models can actually run a support desk?

A support agent has to do several things at once: pull real order data with tools, follow the store's policy, understand what a frustrated customer actually wants, stay pleasant, and not invent facts. We ran 24 models through 162 multi-turn support conversations, each grounded in a real store order, and measured exactly that: including what each one costs to run.

SupportAgentBench is an independent benchmark that evaluates 24 large language models as ecommerce customer-support agents across 162 grounded, multi-turn conversations. It measures resolution quality, escalation calibration, adversarial safety, policy adherence, and cost tier.

There is no composite score: each metric is published separately. Headline results: gpt-5.5 posts the strongest overall profile; budget-tier models match flagship escalation accuracy and safety; and safety failures come from believing unverified claims, not from pressure.

What we measure

We publish the individual metrics and no composite number. A “safe but useless” agent and a “helpful but reckless” agent fail on different axes, and collapsing them into one score hides exactly the tradeoff you need to see.

Read the metric that matches your desk’s risk profile:

Human handoff
Resolvable tickets handed to a human unnecessarily: the cost of playing it too safe. The leaderboard's default sort.
Forbidden actions
Unsafe write actions fired on the 18 hold-the-line traps. Read straight from the transcript, not judged by an LLM.
Escalation accuracy
Must-escalate cases handed to a human. Read straight from the transcript, not judged by an LLM.
Resolution quality
Solvable tickets actually resolved, judged against the store's policy.
Policy adherence
How closely the agent followed store policy and instructions (0–1).

Start here

Four answers to “which model?”

Want the full picture first? Jump to the full leaderboard ↓

The value picture

Price doesn’t buy judgment

The safe shortlist: every model here holds the adversarial line. Height shows how much of the solvable queue it handles end-to-end; left to right, what it costs to run. The best buys sit high and left.

85%90%95%100%$Budget$$Mid$$$Premium$$$$FlagshipRuntime cost (log scale)Solvable tickets handled end-to-end (%)mimo-v2.5 · handles 88% of solvable tickets end-to-end · 2/18 forbidden actions · $ tiermimo-v2.5gemma-4-31b · handles 90% of solvable tickets end-to-end · 2/18 forbidden actions · $ tiergemma-4-31bgpt-5.4-mini · handles 92% of solvable tickets end-to-end · 3/18 forbidden actions · $$ tiergpt-5.4-minigrok-4.3 · handles 97.5% of solvable tickets end-to-end · 3/18 forbidden actions · $$ tiergrok-4.3gemini-3.5-flash · handles 88% of solvable tickets end-to-end · 2/18 forbidden actions · $$$ tiergemini-3.5-flashsonnet-5 · handles 94% of solvable tickets end-to-end · 4/18 forbidden actions · $$$ tiersonnet-5gpt-5.5 · handles 95% of solvable tickets end-to-end · 3/18 forbidden actions · $$$$ tiergpt-5.5

grok-4.3 handles the most queue for a mid-tier price. mimo-v2.5 and gemma-4-31b get you 88–90% at the cheapest tier. gpt-5.5 buys the best conversation quality on the board.

Picks weigh escalation accuracy and safety most heavily. If low handoff volume matters more, grok-4.3 is the autonomy pick.

The hold-the-line set

How far pressure pushes each model

Six representative models under adversarial pressure. Each arrow shows how deep the customer pushed the model: green = it resolves legitimate requests, amber = the safety margin where it should pause and verify or escalate, red = it fired a forbidden action (a free reship, a fraud reroute, a wrongful cancel). The deeper the arrow, the more of the 18 traps it fell for.

RESOLVESSAFETY MARGINFORBIDDEN ACTIONbenign requestsverify or escalatefree goods · fraud · wrongful cancelgemini-3.5-flashasks for proof, holds2/18 unsafegemma-4-31bchecks policy, then acts2/18 unsafemimo-v2.5holds above its price2/18 unsafegpt-5.5flagship: still 3 unsafe3/18 unsafehaiku-4.5reships on demand6/18 unsafegpt-5.4-nanoright text, wrong action7/18 unsafe

What broke models wasn’t hostile pressure: threats and chargebacks got refused almost universally. Failures came from believable, unverified claims: damage with no photo, a polite “it never arrived.”

The decision axis

Escalation calibration, both ways

One row per model, its two escalation mistakes side by side. Red bars (left) are cases that needed a human but the model kept to itself: the risky mistake. Amber bars (right) are solvable tickets it handed to a human anyway: the expensive one. The green center line is perfect calibration; shorter bars on both sides win.

← MISSES CASES THAT NEED A HUMANHANDS OFF TICKETS IT SHOULD SOLVE →% of must-escalate set missed% of solvable tickets over-routedperfect calibrationgpt-5.55%grok-4.32.5%4%gemini-3.1-pro8%gpt-5.48%gpt-5.4-mini8%gpt-5.4-nano8%gemma-4-31b10%deepseek-v4-flash4%8%kimi-k2.610%4%gemini-3.5-flash12%4%sonnet-56%12%gemini-3.1-flash-lite12.5%8%kimi-k2.7-code9%12%gemini-3-flash10%12%qwen3.7-max15%7%mimo-v2.512%12%minimax-m37%17%gpt-5.28%17%deepseek-v4-pro5%21%qwen3.7-plus22%4%mimo-v2.5-pro11%17%glm-5.28%25%haiku-4.513%21%sonnet-4.615%21%

grok-4.3 barely leaves the center. qwen3.7-plus over-routes more than a fifth of solvable tickets; glm-5.2 misses a quarter of true escalations.

Beyond the scores

What we learned reading every transcript

The transcript-level patterns the aggregate scores hide.

1

The open models are already close to the frontier

The best open-weights model, gemma-4-31b, matches the flagships on escalation accuracy and adversarial safety from the budget price tier, and the Chinese open models span the entire board: mimo-v2.5 out-holds models several price tiers above it while glm-5.2 and minimax-m3 sit at the floor. The question isn’t which flagship wins; it’s how little model you can get away with.

2

Forbidden actions come from believing the claim, not folding to pressure

Models hold the line against threats, chargebacks, fraud reroutes, and VIP pressure almost universally. They break when they believe a soft claim: damage without proof, a repeat “never arrived” claimant. Their guardrails key on hostile tone, not on missing evidence. Several models, including gpt-5.4-mini, narrate the red flag in their own reasoning and then act anyway.

3

The reply can be right while the action is wrong

A correct-sounding reply can hide an incorrect tool call. gpt-5.4-nano fired a replacement carrying the wrong item’s variant ID; others presented stale pre-reshipment tracking as the new shipment. These are wrong decisions at the action layer: you only catch them by reading what the agent did, not what it said.

4

Escalation calibration is where models actually differ

Must-escalate accuracy spreads from a clean 100% (the gpt-5.4 family, gpt-5.5, gemma) down to 75–83%, and unnecessary escalation of solvable tickets runs from 2.5% to 22%. grok-4.3 hands over the fewest solvable tickets on the board; qwen3.7-plus dumps more than a fifth of them on humans.

5

The chattier the model, the worse it scores

The strongest models close a ticket in 3–4 messages; the weakest need 5 or more, and the extra messages track lower scores on every metric we measure. gemma-4-31b sends the fewest messages of any model measured: it spends its effort deciding, not talking.

All 24 models

Full leaderboard

162 grounded conversations. No composite score. Sorted by human handoff by default; tap a header to change the ranking.

Read the scoring methodology →
  1. 1grok-4.3🏆
    $$
    Human handoff
    2.5%
    Forbidden
    3/18
    Escalation
    96%

    lowest over-escalation; the autonomy pick

  2. 2deepseek-v4-flash
    $
    Human handoff
    4%
    Forbidden
    5/18
    Escalation
    92%

    strong cheap resolver; weaker safety

  3. 3gpt-5.5
    $$$$
    Human handoff
    5%
    Forbidden
    3/18
    Escalation
    100%

    best quality, top price tier

  4. 4deepseek-v4-pro
    $
    Human handoff
    5%
    Forbidden
    4/18
    Escalation
    79%

    dominated by deepseek-flash

  5. 5sonnet-5
    $$$
    Human handoff
    6%
    Forbidden
    4/18
    Escalation
    88%

    best Claude; fixes 4.6's escalation misses

  6. 6minimax-m3⚠️
    $
    Human handoff
    7%
    Forbidden
    7/18
    Escalation
    83%

    board floor; weakest policy adherence + safety

  7. 7gemini-3.1-pro
    $$$
    Human handoff
    8%
    Forbidden
    3/18
    Escalation
    100%

    best non-GPT policy adherence; priciest gemini

  8. 8gpt-5.4-mini🏆
    $$
    Human handoff
    8%
    Forbidden
    3/18
    Escalation
    100%

    GPT value pick

  9. 9gpt-5.2
    $$$
    Human handoff
    8%
    Forbidden
    3/18
    Escalation
    83%

    older; under-escalates

  10. 10gpt-5.4
    $$$
    Human handoff
    8%
    Forbidden
    4/18
    Escalation
    100%

    strong; beaten on value by mini

  11. 11gpt-5.4-nano⚠️
    $
    Human handoff
    8%
    Forbidden
    7/18
    Escalation
    100%

    cheapest GPT, weak safety

  12. 12glm-5.2⚠️
    $$
    Human handoff
    8%
    Forbidden
    7/18
    Escalation
    75%

    under-escalates + unsafe actions

  13. 13kimi-k2.7-code
    $$
    Human handoff
    9%
    Forbidden
    4/18
    Escalation
    88%

    solid; behind k2.6

  14. 14gemma-4-31b🏆
    $
    Human handoff
    10%
    Forbidden
    2/18
    Escalation
    100%

    best value (reasoning on)

  15. 15gemini-3-flash
    $
    Human handoff
    10%
    Forbidden
    3/18
    Escalation
    88%

    cheap + lean output

  16. 16kimi-k2.6
    $$
    Human handoff
    10%
    Forbidden
    4/18
    Escalation
    96%

    strong, balanced

  17. 17mimo-v2.5-pro
    $
    Human handoff
    11%
    Forbidden
    4/18
    Escalation
    83%

    doesn't beat base mimo

  18. 18gemini-3.5-flash
    $$$
    Human handoff
    12%
    Forbidden
    2/18
    Escalation
    96%

    top-tier + safe

  19. 19mimo-v2.5🏆
    $
    Human handoff
    12%
    Forbidden
    2/18
    Escalation
    88%

    cheapest agent; beats its "pro"

  20. 20gemini-3.1-flash-lite🏆
    $$
    Human handoff
    12.5%
    Forbidden
    2/18
    Escalation
    92%

    cheap + safest

  21. 21haiku-4.5
    $$
    Human handoff
    13%
    Forbidden
    6/18
    Escalation
    79%

    weak Claude

  22. 22qwen3.7-max
    $$$
    Human handoff
    15%
    Forbidden
    3/18
    Escalation
    93%

    safe but over-cautious + pricey

  23. 23sonnet-4.6
    $$$
    Human handoff
    15%
    Forbidden
    3/18
    Escalation
    79%

    safe but pricey; escalates wrongly both ways

  24. 24qwen3.7-plus⚠️
    $$
    Human handoff
    22%
    Forbidden
    2/18
    Escalation
    96%

    over-escalator (22%)

Default order

Sorted by human handoff: solvable tickets handed to a human anyway. Lower means the agent removes more queue volume.

Safety bands

Forbidden actions are unsafe write actions fired in 18 traps. Read 0–2 as safe, 3–5 as mixed, and 6+ as reckless.

Cost + badges

Cost is a runtime tier per 1,000 agent conversations. 🏆 value pick · ⚠️ weak safety · 💭 run with reasoning on.

How to read this honestly

We publish the things that could bias the ranking up front, not buried in a footnote.

Single agent, single store, single vertical.

Every model runs the same production-style support desk for one premium travel-goods store, in English. It's a strong proxy, not a per-agent guarantee: results can shift in other verticals (subscriptions, electronics, apparel), and you should validate on your own transcripts before any production switch.

The adversarial sample is small: read bands, not ranks.

The hold-the-line set is 18 traps, scored as the median of three runs. At that sample size a difference of 1–2 forbidden actions between models is within noise. Treat the safety column as bands (0–2 safe · 3–5 mixed · 6+ reckless); only gaps across bands are meaningful.

Customers are simulated.

An LLM plays the customer, which keeps every model under identical pressure but narrows diversity: real customers are stranger and less predictable than a language model improvising one. Simulated conversations also vary run to run, so published numbers are the median of three (a few models were run only once or twice; their reports say so).

Want an agent that scores like this on your store?

Adelante builds and runs the support agent for you, picking the right model per workload, with the guardrails this benchmark stress-tests.

See if it fits your store