← All stories

Claude Haiku 5.5: When Cheaper Coding Agents Actually Save Money

Haiku's token bill is only the starting point. Compare accepted work after routing, review, retries and escalation.

Watch the short version

Conceptual explainer. Original diagram by AgentGrid.

Read the transcript

Listen to the full story

Listen · 9 min /
In this story

AI-drafted analysis for AgentGrid. Documentation checked October 7, 2026; no hands-on Haiku 5.5 or OpenAI Decisions API evaluation was performed. The workflow and cost example below are proposals, not measured results.

A cheap coding agent saves money when its output costs little to accept. If someone must reconstruct its reasoning, repair the patch and explain the task again to a stronger model, the small first invoice tells you very little.

Give Haiku work with a clear boundary and an inexpensive check. Keep integration and difficult decisions with a stronger agent. Decide what makes an answer acceptable before choosing who produces it.

What the launch changes

Anthropic launched Haiku 5.5 on October 7. Its claim of roughly 75% lower average running cost than Haiku 4.5 accounts for request mix and changed token usage. The stated token-price reductions are 90% for prompts up to 100,000 tokens and 50% above that threshold. Neither number predicts your project's savings. Anthropic's launch announcement.

The published Claude Platform base rates are in US dollars per million tokens:

Prompt lengthInputOutput
Up to and including 100,000 tokens$0.10$0.50
Over 100,000 tokens$0.50$2.50

These are API rates, not a conversion of Claude subscription fees. Caching, batch processing and tool charges require their own accounting. Official pricing.

Haiku 5.5 also supports adjustable effort. That gives you another configuration to evaluate, rather than a reason to turn every task up to maximum. Model documentation.

Count accepted work, including the work you throw away

Use one denominator: outcomes that meet the same acceptance criteria. Include spending on rejected attempts in the numerator:

Cost per accepted outcome = (generation + checking + retries + escalation + human time valued consistently) / accepted outcomes.

Generation, checking, retries, escalation and human time add to total cost. Total cost includes failed and abandoned attempts and is divided by accepted outcomes to calculate cost per accepted outcome.
Count all attempts in the cost, but only accepted outcomes in the denominator. Diagram by AgentGrid.Open full-resolution diagram

Keep elapsed time and escaped defects alongside this figure. A cheaper result that misses a deadline or ships a regression may fail your objective. Subscription users should track usage pressure and review time separately rather than pretending each request has an API invoice.

Here is a hypothetical accounting exercise, not a forecast. Assume 100 single-request Haiku attempts, each with 20,000 input and 2,000 output tokens. At the short-prompt base rates, each costs $0.003, or $0.30 total. Exclude caching, batches, tools and taxes. All other numbers below are invented budgets or outcomes; human time is valued at $60 an hour.

Cost or resultHaiku first, then review/escalationStronger model throughout
Initial Haiku generation$0.30—
Model checking$2.00Included below
Model escalation$3.00Included below
All model work, including checks and retries$5.30$8.00
Human time8 minutes = $8.004 minutes = $4.00
Total cost$13.30$12.00
Accepted outcomes from 100 items8090
Cost per accepted outcomeAbout $0.17About $0.13

The first route has the smaller model bill and the larger cost per accepted outcome. Change the review burden or acceptance count and the result can reverse. Your evaluation needs those measurements; a token-price comparison cannot supply them.

Give Haiku a task you can check cheaply

For a first trial, choose repeated work whose correctness can be inspected without repeating the entire investigation. These are starting hypotheses:

TaskProposed assignmentAcceptance evidence
Find callers of a deprecated helperHaiku inventories; lead decides scopeFile references checked against repository search
Add cases for an existing validatorHaiku drafts a bounded test changeTests exercise specified edge cases without weakening assertions
Change behavior across shared modulesStronger agent owns design and integrationCross-module checks and review of the combined diff
Alter authorization or migrate stored dataStronger agent from the start, with appropriate reviewExplicit invariants, failure cases and recovery plan

For example, suppose a stronger lead has already decided to replace a deprecated date parser. Give Haiku one folder to inventory before allowing edits. A proposed delegation prompt:

Inspect only src/importers/ for calls to parseLegacyDate. Return each caller's file and line, input assumptions, and relevant existing tests. Cite code for every behavioral claim. Do not edit files. If a caller's behavior depends on code outside this folder, name that dependency and stop short of recommending a replacement.

The lead checks the inventory against search, resolves shared behavior and assigns one implementation slice. The reviewer receives the original requirement, exact revision, diff, test output and unresolved questions. A separate conversation gives the reviewer a fresh context; it does not guarantee independent mistakes or establish a filesystem boundary.

Agree on an escalation rule before starting. For this trial: one correction for a specific, reproducible miss; escalate if the same check still fails, evidence conflicts or the task crosses the agreed boundary. Include the failed attempt and the reviewer’s diagnosis in the handoff so the stronger model does not have to rediscover them.

Benchmarks help choose a trial, not its success rate

Anthropic reports 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0. The system card describes 66 tasks and max-effort runs for both. Haiku had ten trials per task, no fallback model and no internet egress; Sonnet had five trials and could use a fallback for safeguard-flagged requests. System card, section 8.4.

Those measured results are not probabilities that either model will fix your bug. Do not use leaderboard percentages to fill the acceptance column of your budget.

Instead, select a small set of real tasks before comparing approaches. Fix the acceptance criteria, starting revision, available tools and reviewer instructions. Record model identity, effort, usage, retries, escalations and human minutes. Inspect failures, including work abandoned before a patch appeared. Expand the cheaper route only where the accepted results justify it.

Haiku, Jev and OpenAI Decisions solve overlapping jobs

Our Jev article separated a typed decision from the application code that acts on it. That distinction matters more when general-purpose models become cheap enough to consider for every small judgment.

OpenAI released its Decisions API in beta on October 6. Official changelog. It currently supports only gpt-6-luna, taking text or images and returning predicate, choice or score answers. The interface is powered by GPT-6 Luna. Decisions documentation.

Our view: Haiku lowers the entry cost of general work, including inspecting files and using tools. Jev and Decisions offer constrained answers that application logic can consume. They compete where the job is classification or routing; they can complement a generative worker where the answer determines what should happen next. A routing decision and a completed patch need different acceptance checks.

Calling Decisions a “Jev killer” would require evidence we do not have. Our market inference is narrower: a general model vendor offering a dedicated decision interface puts pressure on specialists to demonstrate reliability, latency or control advantages on the same workload. A tidy response schema alone is a weak reason to choose either service.

The Jev/Laya comparison showed why per-decision testing matters. Its tiny Laya smoke test used five authored reports, twice each, with one pinned root English checkpoint on CPU. Categories matched, but the missing-reproduction question produced false positives. Those were five cases, not ten independent examples; no hosted Jev comparison was run.

A proposed pipeline is: ticket → typed route plus a human-review policy → bounded worker → stronger review. For example, missing reproduction details could send a ticket back for clarification before anyone attempts a fix. Typed output validates shape, not correctness. Scores do not establish calibration; validate thresholds and give ambiguous or consequential cases a human owner. This is a design to evaluate, not native automated routing in AgentGrid. Count routing errors and their downstream rework in the same accepted-outcome budget.

A ticket passes through a typed route and a human-review policy, which can request clarification or assign a bounded worker. Stronger review checks the worker's evidence before acceptance, correction or escalation.
Proposed pipeline: verify the routing decision and the work it starts. No native AgentGrid automation is implied. Diagram by AgentGrid.Open full-resolution diagram

Make the handoff visible

AgentGrid's build-and-review workflow provides a starting structure: an orchestrator, a builder and a reviewer in a separate conversation, with the requirements and revision carried into review. Use that separation to inspect who produced a claim and what evidence supports it.

Before testing Haiku 5.5, confirm that your installed runtime and account expose that exact model; a generic Haiku label is insufficient.

Download AgentGrid to try a bounded builder-and-reviewer workflow with your available models. Begin with one task and one acceptance check, then record what it actually took to finish.