BatchHaiku

Blog · 2026-10-08 · 9 min read

Claude Haiku 5.5 Benchmarks: Artificial Analysis Scores & Batch Metrics

Haiku 5.5 benchmarks: Artificial Analysis Intelligence Index 43 (max), SWE-bench Pro 64.8%, effort/token tradeoffs, and how to evaluate CSV batch F1.

BatchHaiku — Bulk Text Processing for Claude Haiku 5.5

# Claude Haiku 5.5 Benchmarks: Artificial Analysis Scores & Batch Reality Check

People searching haiku 5.5 benchmarks and haiku 5.5 artificial analysis usually want two things: public leaderboard numbers, and whether those scores matter for production CSV classification. This page covers both.

Official product notes: What’s new in Claude Haiku 5.5. Independent leaderboard write-up: Artificial Analysis on Claude Haiku 5.5 (7 Oct 2026).

Headline public scores (as of Oct 2026)

SourceMetricHaiku 5.5Context
Artificial AnalysisIntelligence Index (max effort)43+26 vs prior Haiku generation; trails Sonnet 5.5 (max 56) by 13
Artificial AnalysisIntelligence Index (high effort)38Same ballpark as GPT-6 Luna (max ~38) with more tokens
Artificial AnalysisTerminal-Bench 4.033%Was 0% on Haiku 4.5
Artificial AnalysisAutomationBench-AA~35%Likely understated (pre-release over-refusal); AA expects a re-run
Artificial AnalysisAA-Omniscience accuracy36%Lower factual knowledge than flash peers; lower hallucination rate (~40%)
Artificial AnalysisAA-Briefcase (max)1578 EloStrong on agentic knowledge-work tasks
Anthropic system cardSWE-bench Pro (max effort)64.8%Sonnet 5.5 reported 81.3% on same card style evals
Anthropic system cardSWE-bench Multilingual (max)83.7%Vendor-reported coding harness
Anthropic system cardSWE-bench Multimodal (max)30.7%Vendor-reported

How to read this: Artificial Analysis ranks Haiku 5.5 as a leading small-class model at max effort — ahead of peers like GLM-5.3 Flash (~42), Gemini 3.8 Flash (~41), and GPT-6 Luna (~38), and close to much larger open models such as Kimi K3 (~44). It is not Sonnet-class intelligence.

Scores move when harnesses, effort settings, or safety filters change. Always check the live AA model page and Anthropic’s latest system card before locking a procurement decision.

Effort settings change the score (and the token bill)

Haiku 5.5 is the first Haiku with Anthropic effort + adaptive thinking. AA’s article highlights:

  • Max Intelligence Index 43, but ~162k output tokens per Index task (~3× GPT-6 Luna max).
  • Moving xhigh → max buys ~+2 Index points for ~1.8× tokens.
  • At high effort, Index 38 with ~55k tokens/task — similar intelligence to Luna max, still slightly heavier on tokens.

For batch labeling, default-on thinking can inflate latency and credits even when the final label is one word. Prefer low/medium effort (or thinking disabled where allowed) for pure classification templates, and reserve max effort for hard extraction / agentic steps.

Pricing context that couples to benchmarks

AA notes Haiku 5.5 list pricing (Anthropic) roughly:

  • $0.10 / $0.50 per 1M input/output tokens for prompts ≤100k
  • $0.50 / $2.50 when prompts go above 100k (5× step-up)
  • Cache reads / short cache writes are much cheaper than full input

Combined with the ~30% tokenizer inflation vs Haiku 4.5, sticker price alone understates effective $/1k CSV rows. See our Haiku 5.5 pricing guide.

Why chat/coding benchmarks are a weak proxy for batch CSV

Public indices optimize for:

  • Multi-step coding / terminal / knowledge-work agent tasks
  • Long reasoning traces under max effort
  • Arena-style chat quality

BatchHaiku workloads usually care about:

MetricWhy it matters
Exact-match / macro-F1 on labelsCore quality for sentiment / ticket routing
Invalid JSON / off-schema rateBreaks downstream automation
p50 / p95 latency per rowPipeline SLO
Tokens in / out per rowReal cost under Haiku 5.5 tokenizer + thinking
Refusal / empty rateSafety classifiers can decline (stop_reason: refusal)

A model that wins Terminal-Bench can still over-refuse ticket text or emit multi-sentence answers when you asked for a single label — both show up as production failures, not Intelligence Index points.

Haiku 5.5 specifics that skew harnesses

  • Adaptive thinking may add tokens and latency vs a “text-only” baseline.
  • Newer tokenizer increases token counts ≈ 30% vs Haiku 4.5 for the same string.
  • Sampling params like temperature are rejected — keep harnesses Anthropic-5.5-compatible.
  • AA flagged AutomationBench over-refusal during pre-release; treat that sub-score cautiously until re-run.

Suggested methodology on BatchHaiku

  1. Upload a labeled sample CSV (20–50 rows).
  2. Run the matching template (sentiment, ticket sorting, RAG filter, extraction).
  3. Export results and score offline against gold labels (F1 + invalid-rate).
  4. Compare credit burn and p95 latency at your chosen effort / thinking settings.
  5. Scale to full corpora only after error budget passes.

Open the batch tool to collect live timings and credit burn on Cloudflare AI–backed Haiku inference.

Related

BatchHaiku is an independent third-party tool, not affiliated with Anthropic or Artificial Analysis. Benchmark figures above are cited from public third-party / vendor reports and may change.

BatchHaiku is an independent third-party tool, NOT affiliated with or endorsed by Anthropic. Live inference runs Claude Haiku 5.5 via Cloudflare AI under your account usage terms.

Open Haiku 5.5 batch tool →More Haiku 5.5 guides