GPT-6 Astra vs. Claude Fable 5.1: Which Model Should Enterprises Choose?

GPT-6 Astra and Claude Fable 5.1 go head-to-head on performance
Sep 8, 2026
8 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Two AI companies dropped their best models 48 hours apart, priced them identically down to the penny, and now can't agree on who actually won.

GPT-6 Astra arrived Sept. 3, two days after Claude Fable 5.1, with both companies charging $10 per million input tokens and $50 per million output tokens. But neither vendor benchmarked its new model directly against the other at launch, leaving independent testing to fill in the most important gaps.

And those tests don't produce a universal winner. Fable 5.1 leads several neutral reasoning evaluations, while Astra wins in other areas and can use fewer tokens per completed task — meaning the better model depends heavily on what an enterprise actually asks it to do.

That's where the agreement ends. Neither company benchmarked its model against the other's, because neither existed yet when the testing was done. Anthropic tested Fable 5.1 against OpenAI's older GPT-5.6 Sol.  OpenAI tested Astra mostly against its own predecessor, too. 

The only truly neutral referee in this fight is Artificial Analysis, an independent evaluator that ran both models through the same tests — and it crowned Fable 5.1 the winner on its top two indices, even as OpenAI's own charts show Astra ahead almost everywhere else.

That makes the real question less about which model is "better" and more about which model fits a particular workload.

Why this comparison matters

Enterprises don't buy benchmark trophies. They buy whichever model completes their workloads reliably at an acceptable cost.

With Astra and Fable 5.1 carrying identical headline API prices, comparing dollars per million tokens isn't enough. Buyers also have to account for token consumption, cache pricing, context length, tool use, task-completion rates, and the infrastructure surrounding each model.

That makes this matchup a useful example of how enterprises increasingly need to evaluate AI models: against their own workloads, not a vendor's favorite benchmark.

What the benchmarks actually say

Advertisement

Depending on which chart you trust, either model can claim the crown.

On OpenAI's own comparison table, Astra leads nearly every published row: 97.6% on FrontierMath Tier 4 versus 87.8% for Fable 5.1, and 96.0% versus 93.7% on GPQA Diamond. Astra also posts a wide lead on Terminal-Bench Science, an agentic research benchmark, at 64.6% against Fable 5.1's 52.6%.

Flip to Artificial Analysis, and Fable 5.1 comes out on top. The firm's Intelligence Index — which blends nine reasoning, knowledge and coding evaluations under one neutral harness — scores Fable 5.1 at 66, the highest mark it has ever recorded, against 61 for Astra. Fable 5.1 also wins Humanity's Last Exam, a broad reasoning test, 65.0% to 57.2%.

Coding splits down the middle depending on the harness. OpenAI's charts give Astra the edge on DeepSWE v1.1 (74.1% to 67.4%) and a narrower one on Terminal-Bench 4.0. But Artificial Analysis's neutral Coding Agent Index calls it close to a dead heat with Astra at 67.0 against Fable 5.1's 67.2, while Anthropic's own SWE-bench Pro score is 81.2 for Fable 5.1, a benchmark OpenAI hasn't published a matching figure for.

Independent testing group ARC Prize Foundation ran its own numbers on abstract reasoning and found Astra scored 62.7% on ARC-AGI-3 using a neutral harness — more than double the prior 30.2% held by Claude Opus 5. But the foundation flagged that Astra's score jumps to 99.9% when run through OpenAI's own adapter, which preserves reasoning state between calls — a gap the foundation attributes to context-management engineering rather than raw model capability.

What each model actually costs to run

The rate card tells only part of the story, and it's arguably the less important part.

On paper, the two models charge identical input and output rates, with one exception: cached tokens. Anthropic charges $0.25 per million cached tokens for Fable 5.1, a 75% cut from its predecessor, while OpenAI charges $1.00 for Astra — four times as much. Push requests above 272,000 input tokens, and Astra's pricing gets more expensive still, doubling input and cache costs and adding a 50% surcharge on output. Anthropic applies no such surcharge across its full 1-million-token window.

That would seem to make Fable 5.1 the cheaper option for any workload involving long documents or repeated context — and it is, in scenarios built around heavy caching. DataCamp modeled a cache-heavy workload of 100,000 tokens read back 1,000 times and found Astra cost $201 against Fable 5.1's $126.

Advertisement

But measured against real token consumption, the story flips entirely. Artificial Analysis found Fable 5.1 costs $3.76 per completed task on its Intelligence Index at maximum effort, compared with $1.67 for Astra — because Astra simply uses far fewer tokens to reach an answer. 

Independent testing firm MindStudio ran both models through an identical coding benchmark suite and reported a similar pattern in reverse: Astra actually cost more in that test, running up $198 in tokens against Fable 5.1's $113 across eight short coding tasks plus four larger app-building projects — a roughly 75% premium for Astra with, in that tester's assessment, worse results on the bigger jobs.

The contradiction between DataCamp's and MindStudio's cost findings is itself instructive: cost per task depends enormously on prompt shape, thinking effort, cache-hit rate, and which coding harness carries the model. There is no single "cheaper model." There's only a cheaper model for a specific job.

Where each model pulls ahead in the field

Astra's clearest technical win is computer use — the ability to click through real software, fill out forms and operate a desktop the way a human would. OpenAI reports 72.6% on OSWorld 2.0 against 70.2% for Claude Opus 5, and 92.7% on ScreenSpot-Pro grounding. OpenAI also says Astra finishes desktop tasks in roughly 47% less time than its predecessor.

Fable 5.1's strength shows up in sustained, tool-heavy work. Anthropic reports it more than doubled its predecessor's score on Terminal-Bench Science, and it posts the highest Coding Agent Index score Artificial Analysis has measured when run inside Claude Code specifically, at 70.

Security is its own separate story. OpenAI says Astra is the first model to cross the "Critical" threshold in its internal risk framework, meaning it can find and build working exploits against hardened systems largely on its own — a capability OpenAI has gated behind a limited-access program called Daybreak rather than shipped broadly. 

Advertisement

Fable 5.1, by contrast, routes flagged cybersecurity and biology requests to an earlier, less capable Claude model automatically. Anthropic reported in its system card that on one indirect prompt-injection benchmark, roughly 23% of Fable 5.1's overall responses and about half of its coding responses on that test were actually served by the fallback model rather than Fable 5.1 itself — meaning the model customers benchmark isn't always the model that answers in production.

OpenAI's safety disclosures cut both ways, too. The company's own reporting notes that Astra's reasoning is harder for outside monitors to read than its predecessor's, and the UK AI Security Institute found Astra attempted simulated supply-chain attacks in 60 of 499 test samples when it had general internet access — a number that dropped to 2 of 500 when internet access was explicitly walled off.

Greg Brockman, OpenAI's co-founder and president, didn't undersell the launch. He called Astra a generational leap at a press briefing, and said he believes the company may have reached artificial general intelligence with the release. OpenAI has also described Astra in its own materials as "the world's best computer use model." 

What eWeek found: The enterprise head-to-head

The evidence points to two models optimized for different strengths rather than a universal winner.

CategoryGPT-6 AstraClaude Fable 5.1eWeek assessment
Base API price$10 in/$50 out per 1M tokens$10 in/$50 out per 1M tokensTie
Cached token price$1.00 per 1M$0.25 per 1MFable 5.1
Long-context surcharge2x input/cache, 1.5x output above 272K tokensNoneFable 5.1
Terminal codingStrong published results; 74.1% DeepSWE55.8% Terminal-BenchGPT-6 Astra for shell-heavy agent work
Independent coding-agent score67 in Codex70 in Claude CodeFable 5.1, with harness caveat
Independent intelligence score (Artificial Analysis)6166Fable 5.1
Task token efficiencyHigh (fewer tokens to reach task completion)Moderate (verbose output generation)GPT-6 Astra
Computer use/desktop automation72.6% (OSWorld 2.0 offline); 92.7% (ScreenSpot-Pro)77.9% partial/41.7% strict (OSWorld 2.0 Aug '26)GPT-6 Astra
Professional artifacts/CAD95.9% on BenchCAD84.3%GPT-6 Astra
Advanced math (FrontierMath Tier 4)97.6%87.8%GPT-6 Astra
Broad reasoning (Humanity's Last Exam, w/ tools)57.2%65.0%Fable 5.1
Cybersecurity capability100% ExploitBench; Critical thresholdDual-use refusals; vulnerability identificationGPT-6 Astra for authorized defensive work
Safety transparencyLow monitorability; risk of adversarial evasionSilent fallback routing to older Claude modelsTie/Trade-off
Context window~1.05M tokens~1M tokensGPT-6 Astra
Knowledge cutoffApril 30, 2026June 2026Fable 5.1
Cloud reachAWS, OpenAI API, Daybreak programAWS, Google Cloud, Microsoft AzureFable 5.1
Reasoning controlsSettable reasoning effort (low to max); Fast modeAlways-on adaptive thinking; fixed reasoning floorGPT-6 Astra for control
Advertisement

The eWeek Verdict

There is no universal winner here — and the evidence suggests enterprises should be skeptical of anyone claiming otherwise.

GPT-6 Astra looks strongest for workloads where token efficiency, advanced math, computer use, or tightly controlled reasoning matter. Claude Fable 5.1 has the advantage in independent reasoning scores, cheaper cache reads, long-context economics, and Claude Code-based agent workflows.

The more important finding is that identical API prices do not produce identical costs. Depending on the workload, either model can be cheaper, and benchmark leadership changes with the test harness.

Enterprises evaluating both should therefore run representative internal workloads through each model and measure cost per successful task, not simply price per token or benchmark rank. The published evidence disagrees too often for either model to be the automatic choice.

Choose GPT-6 Astra if:

  • Your agents operate real software. Computer use and desktop automation is Astra's clearest domain.
  • You need math or scientific reasoning. FrontierMath Tier 4 at 97.6% is a genuine achievement.
  • You produce professional artifacts. Slides, spreadsheets, and CAD output are stated training targets.
  • You run bursty, shorter-horizon tasks where token efficiency lowers cost per task.
  • You already run Codex. Astra's cross-window notes are a Codex feature.

Choose Claude Fable 5.1 if:

  • You run long agent loops where cache reads dominate the bill. Four times cheaper cache reads compound.
  • Your requests are large—no surcharge at any length.
  • You want the best independent reasoning score and will pay for it. 5 points higher on the Intelligence Index.
  • You work in Claude Code. Fable 5.1 in Claude Code is the highest score on the Coding Agent Index.
  • You need the freshest knowledge. June 2026 cutoff versus April 2026.
Advertisement

The trade-offs to watch

Neither model is a safe default. Astra's cybersecurity strength is also a liability that OpenAI itself flagged, restricting broad access while noting the model's internal reasoning is harder to audit than its predecessor's. 

Fable 5.1's safety guardrails are tighter than earlier Claude models for security- and biology-adjacent requests, which Anthropic's own documentation says leads to a higher refusal rate — a real cost for any team doing legitimate security research. 

And several of the most dramatic benchmark results on both sides — Astra's near-perfect ARC-AGI-3 score, Anthropic's OSWorld figures — come with caveats about mismatched test versions or harness-dependent scoring that neither company fully resolves in its own materials. Buyers should treat every number in this comparison as a starting point for their own testing, not a final answer.

Want to learn more AI tips, tricks, and prompting techniques? eWeek readers get free 7-day access to The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work. Browse all lessons →


Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. His work has appeared in publications including TechRepublic, eWEEK, Channel Insider, Geekflare, Enterprise Networking Planet, eSecurity Planet, CIO Insight, and Webopedia. With a technical background in computer science, he specializes in translating complex technology topics into clear, accessible content for business leaders and decision-makers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.