Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2/$10 - your automation bill will still differ
Anthropic and OpenAI shipped mid-tier models a day apart at identical list prices. Token use, caching and effort settings decide what you actually pay.
Anthropic released Claude Sonnet 5.5 on September 28. OpenAI released GPT-6.1 Sol at DevDay the next day. Both list at $2 per million input tokens and $10 per million output tokens. If you are choosing a model for an automation on price alone, you now have a tie, and the tie tells you almost nothing about what each one will cost to run.
What actually launched
**Claude Sonnet 5.5** keeps Sonnet 5's $2/$10 pricing. Anthropic says it generates output more than 30% faster and costs up to 30% less per task in its testing. The per-task saving comes from using fewer tokens, not from a lower rate. It has a 1M token context window and is available on Anthropic's platform, AWS, Google Cloud and Azure, according to Unite.AI and Digital Applied.
The benchmark jumps are large. Anthropic reports 70.6% on Terminal-Bench 4.0, a test of agentic work inside a terminal, against 10.3% for Sonnet 5. OSWorld 2.1, which measures operating a computer, went from 57.0% to 80.1% on partial credit. Opus 5.5 still leads on most of the rows Anthropic published, but on its knowledge-work score (GDPval-AA) the gap is two points.
**GPT-6.1 Sol** also lists at $2/$10, with cached input cut to $0.10 per million tokens. OpenAI's pitch is near-Astra performance at a fraction of the cost. DataCamp reports it matches GPT-6 Astra on the DeepSWE coding benchmark at roughly one-fifth the cost, and costs $5.47 per Terminal-Bench Science task against Astra's $23.80. OpenAI also dropped the planned GPT-6.1 Astra after it regressed on alignment tests, which makes Sol the model OpenAI is steering agent builders toward.
Why the same price is not the same cost
An agent's bill is rate multiplied by tokens, and the token count is where these models differ.
Look at Anthropic's own customer numbers. Balyasny reported Sonnet 5.5 used about 121,000 tokens per answer on its finance tasks, where Sonnet 5 used 497,000. That is a quarter of the tokens at the same rate. Slack reported a much smaller gain, about 14% fewer output tokens. Box reported 12% fewer total tokens. Same model, very different savings, because the savings depend on the work.
Then look at the fine print, which matters more for automation than for chat:
- **Caching.** Most production agents resend the same instructions, tool definitions and reference documents on every run. GPT-6.1 Sol charges $0.10 per million cached input tokens. Sonnet 5.5 charges $0.20. For a pipeline with a long fixed prompt and short variable input, that difference can outweigh the headline rate.
- **Long inputs.** GPT-6.1 Sol bills prompts over 272,000 tokens at twice the input rate and 1.5 times the output rate, for the whole request, according to DataCamp. If your workflow feeds in whole contracts or long email threads, Sol's $2/$10 is not your price.
- **Reasoning is always on.** Sol does not offer a "none" or "minimal" reasoning setting. Sonnet 5.5 turns adaptive thinking on by default and returns errors if you try to disable it. Both models will think on simple tasks, and thinking is billed as output.
- **Batch.** Both offer roughly half price for work that can wait. Anthropic lists Sonnet 5.5 batch at $1/$5. Overnight invoice extraction or weekly report generation should be on batch whichever model you pick.
The honest caveat
The headline benchmarks are at maximum effort. Digital Applied notes that at Anthropic's platform default (high effort), Sonnet 5.5's Terminal-Bench score drops from 70.6% to 43.0%. You can get the higher number, but you pay for it in tokens, which eats the per-task saving you were promised. OpenAI's Sol cost figures are measured against Astra, not against Sonnet, so they do not settle which is cheaper for you either.
Neither model is a drop-in swap. Sonnet 5.5 changes thinking defaults. Sol needs OpenAI's Responses API for tool calling; the older Chat Completions route does not support tools with it. If your agent was built on either predecessor, budget a day or two of testing, not a config change.
And benchmarks are not your workflow. A 2.2-point lead on someone's automation benchmark will not show up on your accounts-payable queue. What shows up is whether the model reads your vendor's odd invoice layout correctly, and how many tokens it burns doing it.
If your current automation is stable and cheap enough, there is no reason to move this month. Both companies are shipping roughly monthly. Migrating every time is its own cost.
What to do about it
Stop comparing models by list price. Compare cost per completed task.
Take 30 to 50 real examples from one workflow you already run, including the awkward ones. Run them through both models at the effort level you would actually use in production. Record three numbers for each run: did it get the right answer, how many input and output tokens it used, and how long it took. Divide total spend by correct answers.
That one number, cost per correct result, is the only price that matters for an automation. It will probably differ from what either launch post suggests, and it takes an afternoon to get. If your automation vendor cannot show you that number for the model they picked, ask them why.
Want this kind of system in your business? Book a free scoping call.