More capability. Less cost. It depends on the task.
Haiku 5.5 is a strong candidate for coding and agent work. Luna deserves a place in the comparison when cost, long prompts or a quick first answer matter. Current public evidence gives reasons to test both; it does not establish one winner for every workload.
Demanding small tasks
Stronger results in several published capability tests. Check effort and the 100K pricing step.
Cost-sensitive workflows
Lower task costs in the independent snapshot; a later long-prompt pricing threshold.
“Luna 6” is search shorthand for OpenAI’s GPT-6 Luna. This is an editorial text-model comparison. VSModels has not run these tests and does not offer live calls to either model. Our current workbench is an image-model mock preview.
Similar headline rates. Different boundaries.
| Specification | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| API model ID | claude-haiku-5-5 | gpt-6-luna |
| Input / output | $0.10 / $0.50 | $0.10 / $0.50 |
| Long-prompt threshold | More than 100,000 input tokens | More than 272,000 input tokens |
| Long-prompt input / output | $0.50 / $2.50 | $0.20 / $0.75 |
| Base cache read / write | $0.01 / $0.125 (5-minute write) | $0.01 / $0.125 |
| Context / standard max output | 1,000,000 / 128,000 tokens | 1,050,000 / 128,000 tokens |
| Input → output | Text & images → text | Text & images → text |
| Default reasoning effort | medium · adaptive thinking | medium · supports none through max |
| Native API | Messages | Responses |
Above each threshold, the higher rates apply to the full request. Tokenizers differ, so equal text is not necessarily equal billed usage. API prices are separate from Claude or ChatGPT subscriptions.
Pick the effort before reading the score.
| Metric | Haiku 5.5 (max) | Luna (max) |
|---|---|---|
| Intelligence Index v4.3.2 ↑ | 43 | 38 |
| Terminal-Bench 4.0 ↑ | 32.8% | 12.6% |
| AutomationBench-AA ↑ | 35.4% | 53.2% |
| Output tokens / second ↑ | 240 | 137 |
| Time to first answer token ↓ | 295.95 s | 101.52 s |
| Cost / Intelligence Index task ↓ | $0.21 | $0.07 |
Output throughput measures streaming after generation starts; it does not measure the wait for an answer. Matching effort names are not equal compute budgets. Benchmark task costs are the evaluator’s estimates, not a quote from our calculator; its flat headline rates may understate long-prompt bills.
What does Anthropic’s own evaluation say?
Anthropic reports a Haiku lead on FrontierCode 1.1 Main (46.4% vs 42.4%) and OSWorld 2.1’s offline subset (72.4% vs 48.9%). These are vendor-reported results with their own setup. Do not merge them with the independent evaluator’s numbers.
Read the vendor table and testing notesRead the experiences, including the disagreements.
We link to the original evaluation or discussion so you can inspect the setup, examples and replies. Community reports are useful clues, with smaller samples and weaker verification.
5 perspectives shown
Higher capability scores can come with a higher task bill.
At max effort, Haiku leads the Intelligence Index while Luna costs less per evaluated task. Luna leads AutomationBench-AA in this snapshot.
Read with this in mind: A weighted benchmark workload, not your repository or chat. Effort labels do not guarantee matching compute budgets.
Read the original evaluationThe cheaper choice changes with effort.
On 20 short and reasoning task types, APIKO reports Haiku low costing 14% less at vendor list rates; at default effort, Luna cost 41% less. Their full graded set had 22 task types, repeated three times per configuration.
Read with this in mind: APIKO sells both models. Its gateway discounts differ from vendor prices, and this small fixture set does not predict every workload.
Read the original evaluationA promising Haiku medium result in a real codebase.
The author tested Rust fixes, review, ambiguous requirements and a repository-level change. Well-specified tasks barely separated the models; Haiku medium did well on the agentic task. Haiku low missed a compile error, while Luna high assumed an ambiguous requirement.
Read with this in mind: One author and codebase, self-reported results, differing effort levels. The detailed agentic cost comparison is against Sonnet, not Luna.
Read the original discussionA large cost gap needs more than a headline.
An xhigh voxel-pagoda comparison reports much greater Haiku token use and API-equivalent cost. Replies question missing prompts and the fairness of the setup.
Read with this in mind: Runs used subscriptions through atomic.chat, not verified direct API invoices. Long-context pricing matters, but this does not establish a universal cost multiplier.
Read the original discussionUsers disagree about the best everyday agent.
Some prefer Haiku for difficult agent work; others favor Luna for inexpensive simple tasks. Context length and effort repeatedly come up in the discussion.
Read with this in mind: Opinions, not a controlled comparison. Subscription limits and prices quoted in comments are unverified; use the official API rates below.
Read the original discussionWhat would your tokens cost?
Enter separate usage for each model. One repeated request is estimated here; a session with changing context lengths needs each turn priced separately.
Illustrative estimate at standard, uncached vendor API rates checked Oct 11, 2026. Include billable reasoning tokens in output. Each input threshold applies per request, to the full request. Excludes caching, tools, images, batch/flex/fast modes, regional premiums, taxes and gateway discounts. This is not a VSModels credit quote or a subscription price.
Choose the test that matches your work.
Coding & agents
Start with Haiku medium and Luna high as candidates, then sweep effort. Keep the repository, tool permissions and acceptance tests fixed. Measure correct completions, retries, elapsed time and cost per accepted change.
Extraction & classification
Compare low effort on both. Score a held-out set with fixed answer keys and refusals tracked separately. APIKO’s small replay suggests effort can reverse the price comparison.
Long documents
Count the same document with each tokenizer. Test retrieval accuracy as well as price. Haiku’s 100K step can matter long before its context window is full.
Fast replies
Measure time to the first useful answer and full completion separately. Compare Luna none with Haiku thinking disabled only where accuracy still meets your acceptance criteria.
These are our suggested evaluation methods, not results from VSModels. Use saved, anonymized examples and report the model ID, date, effort, API route and token usage alongside every result.
Before you switch.
Is Haiku 5.5 better than Luna 6?
In the independent snapshot, Haiku leads several capability metrics, while Luna leads others and costs less per evaluated task. The right choice depends on the task, effort and acceptable error rate. See the benchmark and review sections above.
Do equal token rates mean equal task costs?
No. Tokenization, reasoning output, retries, caching and long-prompt tiers all change the bill. Enter each model’s actual usage in the calculator rather than assuming both consume the same tokens.
Can I swap the model ID and keep my tool loop?
Check the native contracts first. Haiku uses Messages; Luna’s reasoning tool loops belong on Responses. OpenAI limits Chat Completions function calling to none effort for Luna. Anthropic’s documentation also flags sampling-parameter migration changes. Haiku documentation · Luna documentation
Can I try these text models on VSModels?
This page is a research guide with an API cost estimator. VSModels currently has an image-model mock workbench; neither text model is connected here. The documentation links lead to each vendor’s own API information.
Follow the evidence back to its source.
Compiled by VSModels. Official documentation supports specifications; evaluators support their own measurements; individuals support their own experiences. We paraphrase findings and keep the external reading available. Source check: October 11, 2026. Dynamic benchmarks and comments can change after this snapshot.
- Anthropic · model specifications & pricing
- OpenAI · GPT-6 Luna specifications & pricing
- Anthropic · Haiku 5.5 announcement & vendor benchmarks
- Artificial Analysis · release comparison & methodology
- APIKO · task replay, usage & gateway tests
- Reddit · small controlled coding test
- Reddit · disputed voxel pagoda cost comparison
- Reddit · daily agent / Hermes discussion
No live VSModels text benchmark, universal winner, verified subscription allowance or independent replication is claimed. We did not verify the disputed community cost report’s invoices. Vendor-specific discounts and effort defaults must be rechecked before a purchasing decision.
Explore the vendor APIs.
Text model access on VSModels is not available yet. Use each vendor’s API to run your own comparison.