All model guidesTHE MODEL NOTEBOOK · TEXT MODELS

Claude Haiku 5.5
vs GPT-6 Luna

Benchmarks, real-world reviews and the bill behind the tokens.
A comparison you can check, not just a winner to trust.

THE QUICK ANSWER

More capability. Less cost. It depends on the task.

Haiku 5.5 is a strong candidate for coding and agent work. Luna deserves a place in the comparison when cost, long prompts or a quick first answer matter. Current public evidence gives reasons to test both; it does not establish one winner for every workload.

LOOK AT HAIKU FOR

Demanding small tasks

Stronger results in several published capability tests. Check effort and the 100K pricing step.

LOOK AT LUNA FOR

Cost-sensitive workflows

Lower task costs in the independent snapshot; a later long-prompt pricing threshold.

“Luna 6” is search shorthand for OpenAI’s GPT-6 Luna. This is an editorial text-model comparison. VSModels has not run these tests and does not offer live calls to either model. Our current workbench is an image-model mock preview.

01 / OFFICIAL DOCUMENTATION

Similar headline rates. Different boundaries.

Official standard API specifications · USD per 1M text tokens · checked Oct 11, 2026
SpecificationClaude Haiku 5.5GPT-6 Luna
API model IDclaude-haiku-5-5gpt-6-luna
Input / output$0.10 / $0.50$0.10 / $0.50
Long-prompt thresholdMore than 100,000 input tokensMore than 272,000 input tokens
Long-prompt input / output$0.50 / $2.50$0.20 / $0.75
Base cache read / write$0.01 / $0.125 (5-minute write)$0.01 / $0.125
Context / standard max output1,000,000 / 128,000 tokens1,050,000 / 128,000 tokens
Input → outputText & images → textText & images → text
Default reasoning effortmedium · adaptive thinkingmedium · supports none through max
Native APIMessagesResponses

Above each threshold, the higher rates apply to the full request. Tokenizers differ, so equal text is not necessarily equal billed usage. API prices are separate from Claude or ChatGPT subscriptions.

02 / INDEPENDENT EVALUATION

Pick the effort before reading the score.

Artificial Analysis · max effort on both models · Oct 11 snapshot · ↑ higher / ↓ lower is better
MetricHaiku 5.5 (max)Luna (max)
Intelligence Index v4.3.2 ↑4338
Terminal-Bench 4.0 ↑32.8%12.6%
AutomationBench-AA ↑35.4%53.2%
Output tokens / second ↑240137
Time to first answer token ↓295.95 s101.52 s
Cost / Intelligence Index task ↓$0.21$0.07

Output throughput measures streaming after generation starts; it does not measure the wait for an answer. Matching effort names are not equal compute budgets. Benchmark task costs are the evaluator’s estimates, not a quote from our calculator; its flat headline rates may understate long-prompt bills.

What does Anthropic’s own evaluation say?

Anthropic reports a Haiku lead on FrontierCode 1.1 Main (46.4% vs 42.4%) and OSWorld 2.1’s offline subset (72.4% vs 48.9%). These are vendor-reported results with their own setup. Do not merge them with the independent evaluator’s numbers.

Read the vendor table and testing notes
03 / THE REVIEW NOTEBOOK

Read the experiences, including the disagreements.

We link to the original evaluation or discussion so you can inspect the setup, examples and replies. Community reports are useful clues, with smaller samples and weaker verification.

5 perspectives shown

Independent evaluator · snapshot Oct 11

Higher capability scores can come with a higher task bill.

At max effort, Haiku leads the Intelligence Index while Luna costs less per evaluated task. Luna leads AutomationBench-AA in this snapshot.

Read with this in mind: A weighted benchmark workload, not your repository or chat. Effort labels do not guarantee matching compute budgets.

Read the original evaluation
Gateway operator · tested Oct 8

The cheaper choice changes with effort.

On 20 short and reasoning task types, APIKO reports Haiku low costing 14% less at vendor list rates; at default effort, Luna cost 41% less. Their full graded set had 22 task types, repeated three times per configuration.

Read with this in mind: APIKO sells both models. Its gateway discounts differ from vendor prices, and this small fixture set does not predict every workload.

Read the original evaluation
Individual coding test · read Oct 11

A promising Haiku medium result in a real codebase.

The author tested Rust fixes, review, ambiguous requirements and a repository-level change. Well-specified tasks barely separated the models; Haiku medium did well on the agentic task. Haiku low missed a compile error, while Luna high assumed an ambiguous requirement.

Read with this in mind: One author and codebase, self-reported results, differing effort levels. The detailed agentic cost comparison is against Sonnet, not Luna.

Read the original discussion
Disputed single example · read Oct 11

A large cost gap needs more than a headline.

An xhigh voxel-pagoda comparison reports much greater Haiku token use and API-equivalent cost. Replies question missing prompts and the fairness of the setup.

Read with this in mind: Runs used subscriptions through atomic.chat, not verified direct API invoices. Long-context pricing matters, but this does not establish a universal cost multiplier.

Read the original discussion
Daily agent discussion · read Oct 11

Users disagree about the best everyday agent.

Some prefer Haiku for difficult agent work; others favor Luna for inexpensive simple tasks. Context length and effort repeatedly come up in the discussion.

Read with this in mind: Opinions, not a controlled comparison. Subscription limits and prices quoted in comments are unverified; use the official API rates below.

Read the original discussion
04 / YOUR WORKLOAD

What would your tokens cost?

Enter separate usage for each model. One repeated request is estimated here; a session with changing context lengths needs each turn priced separately.

Use each provider’s reported usage.
The same text can have different token counts.

Claude Haiku 5.5
Estimated total · USD$0.45$0.00045 / request

Standard prompt tier
$0.10 input / $0.50 output per 1M tokens

GPT-6 Luna
Estimated total · USD$0.45$0.00045 / request

Standard prompt tier
$0.10 input / $0.50 output per 1M tokens

Illustrative estimate at standard, uncached vendor API rates checked Oct 11, 2026. Include billable reasoning tokens in output. Each input threshold applies per request, to the full request. Excludes caching, tools, images, batch/flex/fast modes, regional premiums, taxes and gateway discounts. This is not a VSModels credit quote or a subscription price.

05 / A PRACTICAL STARTING POINT

Choose the test that matches your work.

01 / TEST YOUR TASK

Coding & agents

Start with Haiku medium and Luna high as candidates, then sweep effort. Keep the repository, tool permissions and acceptance tests fixed. Measure correct completions, retries, elapsed time and cost per accepted change.

02 / TEST YOUR TASK

Extraction & classification

Compare low effort on both. Score a held-out set with fixed answer keys and refusals tracked separately. APIKO’s small replay suggests effort can reverse the price comparison.

03 / TEST YOUR TASK

Long documents

Count the same document with each tokenizer. Test retrieval accuracy as well as price. Haiku’s 100K step can matter long before its context window is full.

04 / TEST YOUR TASK

Fast replies

Measure time to the first useful answer and full completion separately. Compare Luna none with Haiku thinking disabled only where accuracy still meets your acceptance criteria.

These are our suggested evaluation methods, not results from VSModels. Use saved, anonymized examples and report the model ID, date, effort, API route and token usage alongside every result.

06 / COMMON QUESTIONS

Before you switch.

Is Haiku 5.5 better than Luna 6?

In the independent snapshot, Haiku leads several capability metrics, while Luna leads others and costs less per evaluated task. The right choice depends on the task, effort and acceptable error rate. See the benchmark and review sections above.

Do equal token rates mean equal task costs?

No. Tokenization, reasoning output, retries, caching and long-prompt tiers all change the bill. Enter each model’s actual usage in the calculator rather than assuming both consume the same tokens.

Can I swap the model ID and keep my tool loop?

Check the native contracts first. Haiku uses Messages; Luna’s reasoning tool loops belong on Responses. OpenAI limits Chat Completions function calling to none effort for Luna. Anthropic’s documentation also flags sampling-parameter migration changes. Haiku documentation · Luna documentation

Can I try these text models on VSModels?

This page is a research guide with an API cost estimator. VSModels currently has an image-model mock workbench; neither text model is connected here. The documentation links lead to each vendor’s own API information.

07 / SOURCES & LIMITATIONS

Follow the evidence back to its source.

Compiled by VSModels. Official documentation supports specifications; evaluators support their own measurements; individuals support their own experiences. We paraphrase findings and keep the external reading available. Source check: October 11, 2026. Dynamic benchmarks and comments can change after this snapshot.

No live VSModels text benchmark, universal winner, verified subscription allowance or independent replication is claimed. We did not verify the disputed community cost report’s invoices. Vendor-specific discounts and effort defaults must be rechecked before a purchasing decision.

TEST YOUR OWN TASK

Explore the vendor APIs.

Text model access on VSModels is not available yet. Use each vendor’s API to run your own comparison.