Claude Haiku 5.5: When Cheaper Coding Agents Actually Save Money
Haiku's token bill is only the starting point. Compare accepted work after routing, review, retries and escalation.
Watch the short version
Conceptual explainer. Original diagram by AgentGrid.
Read the transcriptListen to the full story
In this story
AI-drafted analysis for AgentGrid. Documentation checked October 7, 2026; no hands-on Haiku 5.5 or OpenAI Decisions API evaluation was performed. The workflow and cost example below are proposals, not measured results.
A cheap coding agent saves money when its output costs little to accept. If someone must reconstruct its reasoning, repair the patch and explain the task again to a stronger model, the small first invoice tells you very little.
Give Haiku work with a clear boundary and an inexpensive check. Keep integration and difficult decisions with a stronger agent. Decide what makes an answer acceptable before choosing who produces it.
What the launch changes
Anthropic launched Haiku 5.5 on October 7. Its claim of roughly 75% lower average running cost than Haiku 4.5 accounts for request mix and changed token usage. The stated token-price reductions are 90% for prompts up to 100,000 tokens and 50% above that threshold. Neither number predicts your project's savings. Anthropic's launch announcement.
The published Claude Platform base rates are in US dollars per million tokens:
| Prompt length | Input | Output |
|---|---|---|
| Up to and including 100,000 tokens | $0.10 | $0.50 |
| Over 100,000 tokens | $0.50 | $2.50 |
These are API rates, not a conversion of Claude subscription fees. Caching, batch processing and tool charges require their own accounting. Official pricing.
Haiku 5.5 also supports adjustable effort. That gives you another configuration to evaluate, rather than a reason to turn every task up to maximum. Model documentation.
Count accepted work, including the work you throw away
Use one denominator: outcomes that meet the same acceptance criteria. Include spending on rejected attempts in the numerator:
Cost per accepted outcome = (generation + checking + retries + escalation + human time valued consistently) / accepted outcomes.
Keep elapsed time and escaped defects alongside this figure. A cheaper result that misses a deadline or ships a regression may fail your objective. Subscription users should track usage pressure and review time separately rather than pretending each request has an API invoice.
Here is a hypothetical accounting exercise, not a forecast. Assume 100 single-request Haiku attempts, each with 20,000 input and 2,000 output tokens. At the short-prompt base rates, each costs $0.003, or $0.30 total. Exclude caching, batches, tools and taxes. All other numbers below are invented budgets or outcomes; human time is valued at $60 an hour.
| Cost or result | Haiku first, then review/escalation | Stronger model throughout |
|---|---|---|
| Initial Haiku generation | $0.30 | — |
| Model checking | $2.00 | Included below |
| Model escalation | $3.00 | Included below |
| All model work, including checks and retries | $5.30 | $8.00 |
| Human time | 8 minutes = $8.00 | 4 minutes = $4.00 |
| Total cost | $13.30 | $12.00 |
| Accepted outcomes from 100 items | 80 | 90 |
| Cost per accepted outcome | About $0.17 | About $0.13 |
The first route has the smaller model bill and the larger cost per accepted outcome. Change the review burden or acceptance count and the result can reverse. Your evaluation needs those measurements; a token-price comparison cannot supply them.
Give Haiku a task you can check cheaply
For a first trial, choose repeated work whose correctness can be inspected without repeating the entire investigation. These are starting hypotheses:
| Task | Proposed assignment | Acceptance evidence |
|---|---|---|
| Find callers of a deprecated helper | Haiku inventories; lead decides scope | File references checked against repository search |
| Add cases for an existing validator | Haiku drafts a bounded test change | Tests exercise specified edge cases without weakening assertions |
| Change behavior across shared modules | Stronger agent owns design and integration | Cross-module checks and review of the combined diff |
| Alter authorization or migrate stored data | Stronger agent from the start, with appropriate review | Explicit invariants, failure cases and recovery plan |
For example, suppose a stronger lead has already decided to replace a deprecated date parser. Give Haiku one folder to inventory before allowing edits. A proposed delegation prompt:
Inspect only
src/importers/for calls toparseLegacyDate. Return each caller's file and line, input assumptions, and relevant existing tests. Cite code for every behavioral claim. Do not edit files. If a caller's behavior depends on code outside this folder, name that dependency and stop short of recommending a replacement.
The lead checks the inventory against search, resolves shared behavior and assigns one implementation slice. The reviewer receives the original requirement, exact revision, diff, test output and unresolved questions. A separate conversation gives the reviewer a fresh context; it does not guarantee independent mistakes or establish a filesystem boundary.
Agree on an escalation rule before starting. For this trial: one correction for a specific, reproducible miss; escalate if the same check still fails, evidence conflicts or the task crosses the agreed boundary. Include the failed attempt and the reviewer’s diagnosis in the handoff so the stronger model does not have to rediscover them.
Benchmarks help choose a trial, not its success rate
Anthropic reports 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0. The system card describes 66 tasks and max-effort runs for both. Haiku had ten trials per task, no fallback model and no internet egress; Sonnet had five trials and could use a fallback for safeguard-flagged requests. System card, section 8.4.
Those measured results are not probabilities that either model will fix your bug. Do not use leaderboard percentages to fill the acceptance column of your budget.
Instead, select a small set of real tasks before comparing approaches. Fix the acceptance criteria, starting revision, available tools and reviewer instructions. Record model identity, effort, usage, retries, escalations and human minutes. Inspect failures, including work abandoned before a patch appeared. Expand the cheaper route only where the accepted results justify it.
Haiku, Jev and OpenAI Decisions solve overlapping jobs
Our Jev article separated a typed decision from the application code that acts on it. That distinction matters more when general-purpose models become cheap enough to consider for every small judgment.
OpenAI released its Decisions API in beta on October 6. Official changelog. It currently supports only gpt-6-luna, taking text or images and returning predicate, choice or score answers. The interface is powered by GPT-6 Luna. Decisions documentation.
Our view: Haiku lowers the entry cost of general work, including inspecting files and using tools. Jev and Decisions offer constrained answers that application logic can consume. They compete where the job is classification or routing; they can complement a generative worker where the answer determines what should happen next. A routing decision and a completed patch need different acceptance checks.
Calling Decisions a “Jev killer” would require evidence we do not have. Our market inference is narrower: a general model vendor offering a dedicated decision interface puts pressure on specialists to demonstrate reliability, latency or control advantages on the same workload. A tidy response schema alone is a weak reason to choose either service.
The Jev/Laya comparison showed why per-decision testing matters. Its tiny Laya smoke test used five authored reports, twice each, with one pinned root English checkpoint on CPU. Categories matched, but the missing-reproduction question produced false positives. Those were five cases, not ten independent examples; no hosted Jev comparison was run.
A proposed pipeline is: ticket → typed route plus a human-review policy → bounded worker → stronger review. For example, missing reproduction details could send a ticket back for clarification before anyone attempts a fix. Typed output validates shape, not correctness. Scores do not establish calibration; validate thresholds and give ambiguous or consequential cases a human owner. This is a design to evaluate, not native automated routing in AgentGrid. Count routing errors and their downstream rework in the same accepted-outcome budget.
Make the handoff visible
AgentGrid's build-and-review workflow provides a starting structure: an orchestrator, a builder and a reviewer in a separate conversation, with the requirements and revision carried into review. Use that separation to inspect who produced a claim and what evidence supports it.
Before testing Haiku 5.5, confirm that your installed runtime and account expose that exact model; a generic Haiku label is insufficient.
Download AgentGrid to try a bounded builder-and-reviewer workflow with your available models. Begin with one task and one acceptance check, then record what it actually took to finish.