Why Haiku 5.5 can cost 12× more than Luna 6
Both list at $0.10 in and $0.50 out per million tokens. A gap that large is possible, but only when two effects multiply: Haiku's price tier above 100K tokens, and how much more Haiku writes. For most workloads the difference is between none and 5×.
5×
The 100K price tier
Once one request's prompt passes 100,000 tokens — cache reads and writes included — Anthropic bills the whole request at $0.50 in and $2.50 out. Luna keeps $0.10 / $0.50 until 272K, then moves to only $0.20 / $0.75.
~3.02×
Longer output
Running the Artificial Analysis Intelligence Index at max effort took Haiku 435M output tokens and Luna 144M. Output is the expensive side of the bill, so the same task can cost Haiku about 3.02× as much even at equal prices.
Above 100K, Haiku's output can cost up to 5 × 3.02 ≈ 15.1× Luna's. Input pulls the ratio down toward 5×, so the bill lands between the two depending on how output-heavy the work is.
Five workloads, priced
USD per request at list price. "Haiku writes more" applies the 3.02× output multiplier.
Each workload opens the calculator with the same inputs, so you can change volume, caching and batch from there.
What one report can and can’t show
A recent r/ClaudeAI thread reported Haiku 5.5 costing about 12× Luna 6 on one agent task. It is a useful warning: long-context agents that keep crossing 100K tokens can get expensive on Haiku fast. It is not evidence that Haiku is 12× more expensive in general.
It was a single run. Without the prompt, tools, number of turns, starting workspace, context compaction and raw usage logs, nobody can reproduce it, and the two models may not have done the same amount of work.
Effort settings don’t line up across vendors. “Both on xhigh” is not the same budget: hidden reasoning, tool strategy and compaction differ between Anthropic and OpenAI.
Agent bills mix price tiers. Each request is priced on its own, and an agent’s context grows turn by turn, so early requests bill at the base rate and later ones at the high tier. For example, a session with 5M new input, 200M cache reads and 4M output costs $4.50 on Haiku if every request stays under 100K and $22.50 if every request is over it. A real bill lands somewhere between, which is why totals rarely match a single price-sheet calculation.
How to compare fairly
- Fix the task pack. Same prompts, same files, same tool permissions, and a clean workspace for every run.
- Define “done” in code. An acceptance script or test suite decides whether a run succeeded, not a read-through.
- Run it several times. Report success rate, median cost and P95 cost per successful task, not one total.
- Log every request. Prompt length, cache reads and writes, output tokens and which price tier each request hit, before and after compaction.
- Sweep the effort setting. Compare each model at the lowest effort that passes, since effort scales differ between vendors.
How to keep Haiku cheap
- Stay under 100K per request. Trim retrieved context, summarise history, or split long documents across requests. Cached tokens still count toward the 100K.
- Lower the effort. Haiku 5.5 defaults to medium effort; low is recommended for chat, short tool tasks and simple high-volume requests, and uses fewer output tokens.
- Cap the output. Set max_tokens to what the task needs, and ask for terse formats (a label, JSON) where you can.
- Batch what can wait. Both vendors halve the price in batch; the thresholds still apply.
- Route long prompts to Luna. For prompts between 100K and 272K, Luna is about 5× cheaper per token.
Caveats
The 3.02× figure is one benchmark run at max effort. Your ratio depends on the task, the effort setting and your prompts; measure it with a few real runs in the playground.
The two vendors use different tokenizers, so the same text can produce a different number of tokens on each. The calculator uses the token counts you enter for both.
Prices from Anthropic pricing and OpenAI API pricing, verified 2026-10-09. Output-token figures from Artificial Analysis.