Two AI companies dropped their best models 48 hours apart, priced them identically down to the penny, and now can't agree on who actually won.
GPT-6 Astra arrived Sept. 3, two days after Claude Fable 5.1, with both companies charging $10 per million input tokens and $50 per million output tokens. But neither vendor benchmarked its new model directly against the other at launch, leaving independent testing to fill in the most important gaps.
And those tests don't produce a universal winner. Fable 5.1 leads several neutral reasoning evaluations, while Astra wins in other areas and can use fewer tokens per completed task — meaning the better model depends heavily on what an enterprise actually asks it to do.
That's where the agreement ends. Neither company benchmarked its model against the other's, because neither existed yet when the testing was done. Anthropic tested Fable 5.1 against OpenAI's older GPT-5.6 Sol. OpenAI tested Astra mostly against its own predecessor, too.
The only truly neutral referee in this fight is Artificial Analysis, an independent evaluator that ran both models through the same tests — and it crowned Fable 5.1 the winner on its top two indices, even as OpenAI's own charts show Astra ahead almost everywhere else.
That makes the real question less about which model is "better" and more about which model fits a particular workload.
Why this comparison matters
Enterprises don't buy benchmark trophies. They buy whichever model completes their workloads reliably at an acceptable cost.
With Astra and Fable 5.1 carrying identical headline API prices, comparing dollars per million tokens isn't enough. Buyers also have to account for token consumption, cache pricing, context length, tool use, task-completion rates, and the infrastructure surrounding each model.
That makes this matchup a useful example of how enterprises increasingly need to evaluate AI models: against their own workloads, not a vendor's favorite benchmark.
What the benchmarks actually say
Depending on which chart you trust, either model can claim the crown.
On OpenAI's own comparison table, Astra leads nearly every published row: 97.6% on FrontierMath Tier 4 versus 87.8% for Fable 5.1, and 96.0% versus 93.7% on GPQA Diamond. Astra also posts a wide lead on Terminal-Bench Science, an agentic research benchmark, at 64.6% against Fable 5.1's 52.6%.
Flip to Artificial Analysis, and Fable 5.1 comes out on top. The firm's Intelligence Index — which blends nine reasoning, knowledge and coding evaluations under one neutral harness — scores Fable 5.1 at 66, the highest mark it has ever recorded, against 61 for Astra. Fable 5.1 also wins Humanity's Last Exam, a broad reasoning test, 65.0% to 57.2%.
Coding splits down the middle depending on the harness. OpenAI's charts give Astra the edge on DeepSWE v1.1 (74.1% to 67.4%) and a narrower one on Terminal-Bench 4.0. But Artificial Analysis's neutral Coding Agent Index calls it close to a dead heat with Astra at 67.0 against Fable 5.1's 67.2, while Anthropic's own SWE-bench Pro score is 81.2 for Fable 5.1, a benchmark OpenAI hasn't published a matching figure for.
Independent testing group ARC Prize Foundation ran its own numbers on abstract reasoning and found Astra scored 62.7% on ARC-AGI-3 using a neutral harness — more than double the prior 30.2% held by Claude Opus 5. But the foundation flagged that Astra's score jumps to 99.9% when run through OpenAI's own adapter, which preserves reasoning state between calls — a gap the foundation attributes to context-management engineering rather than raw model capability.
What each model actually costs to run
The rate card tells only part of the story, and it's arguably the less important part.
On paper, the two models charge identical input and output rates, with one exception: cached tokens. Anthropic charges $0.25 per million cached tokens for Fable 5.1, a 75% cut from its predecessor, while OpenAI charges $1.00 for Astra — four times as much. Push requests above 272,000 input tokens, and Astra's pricing gets more expensive still, doubling input and cache costs and adding a 50% surcharge on output. Anthropic applies no such surcharge across its full 1-million-token window.
That would seem to make Fable 5.1 the cheaper option for any workload involving long documents or repeated context — and it is, in scenarios built around heavy caching. DataCamp modeled a cache-heavy workload of 100,000 tokens read back 1,000 times and found Astra cost $201 against Fable 5.1's $126.
But measured against real token consumption, the story flips entirely. Artificial Analysis found Fable 5.1 costs $3.76 per completed task on its Intelligence Index at maximum effort, compared with $1.67 for Astra — because Astra simply uses far fewer tokens to reach an answer.
Independent testing firm MindStudio ran both models through an identical coding benchmark suite and reported a similar pattern in reverse: Astra actually cost more in that test, running up $198 in tokens against Fable 5.1's $113 across eight short coding tasks plus four larger app-building projects — a roughly 75% premium for Astra with, in that tester's assessment, worse results on the bigger jobs.
The contradiction between DataCamp's and MindStudio's cost findings is itself instructive: cost per task depends enormously on prompt shape, thinking effort, cache-hit rate, and which coding harness carries the model. There is no single "cheaper model." There's only a cheaper model for a specific job.
Where each model pulls ahead in the field
Astra's clearest technical win is computer use — the ability to click through real software, fill out forms and operate a desktop the way a human would. OpenAI reports 72.6% on OSWorld 2.0 against 70.2% for Claude Opus 5, and 92.7% on ScreenSpot-Pro grounding. OpenAI also says Astra finishes desktop tasks in roughly 47% less time than its predecessor.
Fable 5.1's strength shows up in sustained, tool-heavy work. Anthropic reports it more than doubled its predecessor's score on Terminal-Bench Science, and it posts the highest Coding Agent Index score Artificial Analysis has measured when run inside Claude Code specifically, at 70.
Security is its own separate story. OpenAI says Astra is the first model to cross the "Critical" threshold in its internal risk framework, meaning it can find and build working exploits against hardened systems largely on its own — a capability OpenAI has gated behind a limited-access program called Daybreak rather than shipped broadly.
Fable 5.1, by contrast, routes flagged cybersecurity and biology requests to an earlier, less capable Claude model automatically. Anthropic reported in its system card that on one indirect prompt-injection benchmark, roughly 23% of Fable 5.1's overall responses and about half of its coding responses on that test were actually served by the fallback model rather than Fable 5.1 itself — meaning the model customers benchmark isn't always the model that answers in production.
OpenAI's safety disclosures cut both ways, too. The company's own reporting notes that Astra's reasoning is harder for outside monitors to read than its predecessor's, and the UK AI Security Institute found Astra attempted simulated supply-chain attacks in 60 of 499 test samples when it had general internet access — a number that dropped to 2 of 500 when internet access was explicitly walled off.
Greg Brockman, OpenAI's co-founder and president, didn't undersell the launch. He called Astra a generational leap at a press briefing, and said he believes the company may have reached artificial general intelligence with the release. OpenAI has also described Astra in its own materials as "the world's best computer use model."
What eWeek found: The enterprise head-to-head
The evidence points to two models optimized for different strengths rather than a universal winner.
| Category | GPT-6 Astra | Claude Fable 5.1 | eWeek assessment |
| Base API price | $10 in/$50 out per 1M tokens | $10 in/$50 out per 1M tokens | Tie |
| Cached token price | $1.00 per 1M | $0.25 per 1M | Fable 5.1 |
| Long-context surcharge | 2x input/cache, 1.5x output above 272K tokens | None | Fable 5.1 |
| Terminal coding | Strong published results; 74.1% DeepSWE | 55.8% Terminal-Bench | GPT-6 Astra for shell-heavy agent work |
| Independent coding-agent score | 67 in Codex | 70 in Claude Code | Fable 5.1, with harness caveat |
| Independent intelligence score (Artificial Analysis) | 61 | 66 | Fable 5.1 |
| Task token efficiency | High (fewer tokens to reach task completion) | Moderate (verbose output generation) | GPT-6 Astra |
| Computer use/desktop automation | 72.6% (OSWorld 2.0 offline); 92.7% (ScreenSpot-Pro) | 77.9% partial/41.7% strict (OSWorld 2.0 Aug '26) | GPT-6 Astra |
| Professional artifacts/CAD | 95.9% on BenchCAD | 84.3% | GPT-6 Astra |
| Advanced math (FrontierMath Tier 4) | 97.6% | 87.8% | GPT-6 Astra |
| Broad reasoning (Humanity's Last Exam, w/ tools) | 57.2% | 65.0% | Fable 5.1 |
| Cybersecurity capability | 100% ExploitBench; Critical threshold | Dual-use refusals; vulnerability identification | GPT-6 Astra for authorized defensive work |
| Safety transparency | Low monitorability; risk of adversarial evasion | Silent fallback routing to older Claude models | Tie/Trade-off |
| Context window | ~1.05M tokens | ~1M tokens | GPT-6 Astra |
| Knowledge cutoff | April 30, 2026 | June 2026 | Fable 5.1 |
| Cloud reach | AWS, OpenAI API, Daybreak program | AWS, Google Cloud, Microsoft Azure | Fable 5.1 |
| Reasoning controls | Settable reasoning effort (low to max); Fast mode | Always-on adaptive thinking; fixed reasoning floor | GPT-6 Astra for control |
The eWeek Verdict
There is no universal winner here — and the evidence suggests enterprises should be skeptical of anyone claiming otherwise.
GPT-6 Astra looks strongest for workloads where token efficiency, advanced math, computer use, or tightly controlled reasoning matter. Claude Fable 5.1 has the advantage in independent reasoning scores, cheaper cache reads, long-context economics, and Claude Code-based agent workflows.
The more important finding is that identical API prices do not produce identical costs. Depending on the workload, either model can be cheaper, and benchmark leadership changes with the test harness.
Enterprises evaluating both should therefore run representative internal workloads through each model and measure cost per successful task, not simply price per token or benchmark rank. The published evidence disagrees too often for either model to be the automatic choice.
Choose GPT-6 Astra if:
- Your agents operate real software. Computer use and desktop automation is Astra's clearest domain.
- You need math or scientific reasoning. FrontierMath Tier 4 at 97.6% is a genuine achievement.
- You produce professional artifacts. Slides, spreadsheets, and CAD output are stated training targets.
- You run bursty, shorter-horizon tasks where token efficiency lowers cost per task.
- You already run Codex. Astra's cross-window notes are a Codex feature.
Choose Claude Fable 5.1 if:
- You run long agent loops where cache reads dominate the bill. Four times cheaper cache reads compound.
- Your requests are large—no surcharge at any length.
- You want the best independent reasoning score and will pay for it. 5 points higher on the Intelligence Index.
- You work in Claude Code. Fable 5.1 in Claude Code is the highest score on the Coding Agent Index.
- You need the freshest knowledge. June 2026 cutoff versus April 2026.
The trade-offs to watch
Neither model is a safe default. Astra's cybersecurity strength is also a liability that OpenAI itself flagged, restricting broad access while noting the model's internal reasoning is harder to audit than its predecessor's.
Fable 5.1's safety guardrails are tighter than earlier Claude models for security- and biology-adjacent requests, which Anthropic's own documentation says leads to a higher refusal rate — a real cost for any team doing legitimate security research.
And several of the most dramatic benchmark results on both sides — Astra's near-perfect ARC-AGI-3 score, Anthropic's OSWorld figures — come with caveats about mismatched test versions or harness-dependent scoring that neither company fully resolves in its own materials. Buyers should treat every number in this comparison as a starting point for their own testing, not a final answer.
Want to learn more AI tips, tricks, and prompting techniques? eWeek readers get free 7-day access to The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work. Browse all lessons →


