In brief

  • Anthropic launched Claude Sonnet 5.5 on Monday, pricing it at $2 per million input tokens and $10 per million output tokens—matching Sonnet 5’s rates and cutting Opus 5.5’s cost in half.
  • In Anthropic’s internal benchmark, Sonnet 5.5 achieved a 70.6% score on Terminal‑Bench 4.0, outpacing Opus 5.5’s 66.4%. Independent tester Artificial Analysis confirmed the advantage, recording 63.6% for Sonnet 5.5 versus 59.6% for Opus 5.5.
  • Artificial Analysis places Sonnet 5.5 just behind Opus 5.5 in overall ranking, noting that it consumes more tokens per task than any other model evaluated.

Anthropic introduced Claude Sonnet 5.5 on Monday as an upgrade to the June Sonnet 5 release, claiming the new middle‑tier model operates more than 30% faster than its predecessor.

“Sonnet 5.5 excels at everyday, well‑scoped tasks such as bug fixing and producing polished documents, slides, and spreadsheets, while also demonstrating a keen eye for design,” Anthropic stated.

The pricing remains at $2 per million input tokens and $10 per million output tokens. Tokens represent the text chunks an AI reads and writes—slightly shorter than a word—and are billed per million. This rate is half of what Opus 5.5 charges. Moreover, Sonnet 5.5 uses nearly a third fewer tokens per task, making it cheaper to run than the original Sonnet 5.

When it comes to coding, Sonnet 5.5 truly shines. On the Terminal‑Bench 4.0 assessment—which measures an AI agent’s ability to complete complex professional tasks by issuing command‑line instructions—Sonnet 5.5 achieved a 70.6% completion rate, compared with 66.4% for Opus 5.5 and just 10.3% for the earlier Sonnet 5.

In practical terms, the more affordable model completed more tasks. Independent tester Artificial Analysis ran its own evaluation and found Sonnet 5.5 scoring 63.6%, Opus 5.5 at 59.6%, and OpenAI’s GPT‑6 Astra at 59.1%.

Scores also depend on the effort setting, a dial that makes a model think longer for a better answer and a bigger bill. Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol on FrontierCode for about a fifth of the cost per task.

On GDPval-AA, which grades real-world professional work across 44 occupations using Elo—the chess-style system that ranks relative skill—Sonnet 5.5 scored 1844 to Opus 5.5’s 1846, effectively a tie. GPT-6 Sol scored 1487.

Rivals match the price. OpenAI cut GPT-6 Sol to $2 and $10 last week, and GPT-5.6 Terra, its mid-tier model, lists at $2 and $12. Anthropic published no Terra benchmarks.

The catch

Sonnet 5.5 tends to be verbose. At maximum effort it generated roughly 193,000 tokens per test task—the highest figure recorded by Artificial Analysis and about 60% more than Opus 5.5. This translates to a cost of approximately $7.60 per task, roughly 50% above Sonnet 5, which undermines Anthropic’s claim of up to 30% savings.

Anthropic’s advertised savings appear at lower effort levels. At the Medium setting, which is the default in its applications, Sonnet 5.5 surpasses Sonnet 5’s best coding score for less than one‑tenth of the cost. Artificial Analysis suggests that High effort delivers the best value, meaning everyday users can obtain near‑flagship coding performance at a fraction of the price—provided the dial remains low.

Anthropic’s figures are self‑reported, and Artificial Analysis evaluated a pre‑release version that contained a bug the company believes had minimal impact or slightly depressed the scores. Anthropic maintains that Opus 5.5 still clearly outperforms Sonnet 5.5 on complex tasks requiring sustained judgment.

Claude Haiku 5.5, designed for high‑volume, cost‑sensitive workloads, is expected to arrive in the coming weeks.



Source link

Exit mobile version