Blog · 2026-10-08 · 9 min read
Claude Haiku 5.5 Benchmarks: Artificial Analysis Scores & Batch Metrics
Haiku 5.5 benchmarks: Artificial Analysis Intelligence Index 43 (max), SWE-bench Pro 64.8%, effort/token tradeoffs, and how to evaluate CSV batch F1.

# Claude Haiku 5.5 Benchmarks: Artificial Analysis Scores & Batch Reality Check
People searching haiku 5.5 benchmarks and haiku 5.5 artificial analysis usually want two things: public leaderboard numbers, and whether those scores matter for production CSV classification. This page covers both.
Official product notes: What’s new in Claude Haiku 5.5. Independent leaderboard write-up: Artificial Analysis on Claude Haiku 5.5 (7 Oct 2026).
Headline public scores (as of Oct 2026)
| Source | Metric | Haiku 5.5 | Context |
|---|---|---|---|
| Artificial Analysis | Intelligence Index (max effort) | 43 | +26 vs prior Haiku generation; trails Sonnet 5.5 (max 56) by 13 |
| Artificial Analysis | Intelligence Index (high effort) | 38 | Same ballpark as GPT-6 Luna (max ~38) with more tokens |
| Artificial Analysis | Terminal-Bench 4.0 | 33% | Was 0% on Haiku 4.5 |
| Artificial Analysis | AutomationBench-AA | ~35% | Likely understated (pre-release over-refusal); AA expects a re-run |
| Artificial Analysis | AA-Omniscience accuracy | 36% | Lower factual knowledge than flash peers; lower hallucination rate (~40%) |
| Artificial Analysis | AA-Briefcase (max) | 1578 Elo | Strong on agentic knowledge-work tasks |
| Anthropic system card | SWE-bench Pro (max effort) | 64.8% | Sonnet 5.5 reported 81.3% on same card style evals |
| Anthropic system card | SWE-bench Multilingual (max) | 83.7% | Vendor-reported coding harness |
| Anthropic system card | SWE-bench Multimodal (max) | 30.7% | Vendor-reported |
How to read this: Artificial Analysis ranks Haiku 5.5 as a leading small-class model at max effort — ahead of peers like GLM-5.3 Flash (~42), Gemini 3.8 Flash (~41), and GPT-6 Luna (~38), and close to much larger open models such as Kimi K3 (~44). It is not Sonnet-class intelligence.
Scores move when harnesses, effort settings, or safety filters change. Always check the live AA model page and Anthropic’s latest system card before locking a procurement decision.
Effort settings change the score (and the token bill)
Haiku 5.5 is the first Haiku with Anthropic effort + adaptive thinking. AA’s article highlights:
- Max Intelligence Index 43, but ~162k output tokens per Index task (~3× GPT-6 Luna max).
- Moving xhigh → max buys ~+2 Index points for ~1.8× tokens.
- At high effort, Index 38 with ~55k tokens/task — similar intelligence to Luna max, still slightly heavier on tokens.
For batch labeling, default-on thinking can inflate latency and credits even when the final label is one word. Prefer low/medium effort (or thinking disabled where allowed) for pure classification templates, and reserve max effort for hard extraction / agentic steps.
Pricing context that couples to benchmarks
AA notes Haiku 5.5 list pricing (Anthropic) roughly:
- $0.10 / $0.50 per 1M input/output tokens for prompts ≤100k
- $0.50 / $2.50 when prompts go above 100k (5× step-up)
- Cache reads / short cache writes are much cheaper than full input
Combined with the ~30% tokenizer inflation vs Haiku 4.5, sticker price alone understates effective $/1k CSV rows. See our Haiku 5.5 pricing guide.
Why chat/coding benchmarks are a weak proxy for batch CSV
Public indices optimize for:
- Multi-step coding / terminal / knowledge-work agent tasks
- Long reasoning traces under max effort
- Arena-style chat quality
BatchHaiku workloads usually care about:
| Metric | Why it matters |
|---|---|
| Exact-match / macro-F1 on labels | Core quality for sentiment / ticket routing |
| Invalid JSON / off-schema rate | Breaks downstream automation |
| p50 / p95 latency per row | Pipeline SLO |
| Tokens in / out per row | Real cost under Haiku 5.5 tokenizer + thinking |
| Refusal / empty rate | Safety classifiers can decline (stop_reason: refusal) |
A model that wins Terminal-Bench can still over-refuse ticket text or emit multi-sentence answers when you asked for a single label — both show up as production failures, not Intelligence Index points.
Haiku 5.5 specifics that skew harnesses
- Adaptive thinking may add tokens and latency vs a “text-only” baseline.
- Newer tokenizer increases token counts ≈ 30% vs Haiku 4.5 for the same string.
- Sampling params like
temperatureare rejected — keep harnesses Anthropic-5.5-compatible. - AA flagged AutomationBench over-refusal during pre-release; treat that sub-score cautiously until re-run.
Suggested methodology on BatchHaiku
- Upload a labeled sample CSV (20–50 rows).
- Run the matching template (sentiment, ticket sorting, RAG filter, extraction).
- Export results and score offline against gold labels (F1 + invalid-rate).
- Compare credit burn and p95 latency at your chosen effort / thinking settings.
- Scale to full corpora only after error budget passes.
Open the batch tool to collect live timings and credit burn on Cloudflare AI–backed Haiku inference.
Related
BatchHaiku is an independent third-party tool, not affiliated with Anthropic or Artificial Analysis. Benchmark figures above are cited from public third-party / vendor reports and may change.
BatchHaiku is an independent third-party tool, NOT affiliated with or endorsed by Anthropic. Live inference runs Claude Haiku 5.5 via Cloudflare AI under your account usage terms.