Model launches

Evaluating Claude Fable 5: safeguarded responses, hidden fallbacks and benchmark distortion

Multiple independent evaluators found it difficult to measure Anthropic’s Claude Fable 5 because safety classifiers either refused prompts or routed them to the weaker Claude Opus 4.8, and Anthropic requires prompts and outputs to be retained for 30 days.

Evaluating Claude Fable 5: safeguarded responses, hidden fallbacks and benchmark distortion

Multiple independent organizations reported they could not reliably evaluate Anthropic’s Claude Fable 5, the safeguarded public variant of Claude Mythos 5: some test prompts were refused, others were routed to the less capable Claude Opus 4.8. Anthropic requires users to accept that prompts and outputs will be retained for 30 days in order to use Claude Fable 5.

How Claude Fable 5 behaved under evaluation

Anthropic’s classifiers screened every prompt before it reached Claude Fable 5. If a prompt was flagged — for example as relating to cybersecurity, biology, chemistry, or AI model engineering — it never reached Claude Fable 5. The flagged prompts produced two main outcomes:

  • Within Anthropic’s own apps, including the Claude Code environment, flagged prompts were automatically routed to Claude Opus 4.8; the switch was recorded as a separate log event but not reflected in the answer text.
  • Through the API (the route most evaluators used), the same flag often produced an outright refusal and no answer. Evaluators could either enable a fallback to retry the prompt on Claude Opus 4.8 or score the task as a failure.

Meanwhile, Anthropic’s mandatory requirement that prompts and outputs be retained for 30 days prevented some evaluators from running proprietary test sets.

Evaluation approaches and specific findings

Evaluators generally picked between a “pure” evaluation — measuring only answers produced by Claude Fable 5 alone — and a “practical” evaluation that counted refusals and fallback responses as part of the delivered experience.

Key findings from published evaluations:

  • Artificial Analysis: tested Claude Fable 5 before launch and recorded fallback to Claude Opus 4.8 on roughly 8 percent of tasks within its Intelligence Index (a composite of 10 economically useful tasks). Artificial Analysis included all fallback responses in its scoring, producing blended results.

  • Vals AI: published two score sets — one that included Claude Opus 4.8 fallback answers and one that treated every refusal as a failure. Vals AI reported nearly 100 percent refusal rates on biology and cybersecurity questions.

  • Agents’ Last Exam: on long-horizon, verifiable agentic tasks evaluators reported Claude Fable 5 refused about 35 percent of tasks. The system flagged science questions as “cybersecurity or biology” and switched mid-task to Claude Opus 4.8, logging the switch separately. Evaluators compared performance on "untouched" tasks (answers solely from Claude Fable 5) and composite tasks (where Claude Opus 4.8 contributed).

  • ARC Prize Foundation: which runs the ARC-AGI abstract reasoning tests, declined to run verified evaluations rather than expose its private test set to Anthropic’s retention requirement; it said it would publish results if it could test without handing the questions over.

Numbers and rankings

  • Artificial Analysis Intelligence Index: Claude Fable 5 (including fallback responses by Claude Opus 4.8) placed first at 64.9, 3.5 percentage points higher than Claude Opus 4.8.
  • Humanity’s Last Exam: despite refusing 9 percent of test questions, Claude Fable 5 finished with a score of 53 percent — the highest recorded to date and more than 7 percentage points higher than Claude Opus 4.8.
  • Vals AI: with Anthropic’s optional fallback enabled, Claude Fable 5 placed first on most benchmarks, including 75.14 percent on the overall Vals Index. Counting refusals as failures reduced the overall score only slightly to 74.92 percent, but severely damaged results in flagged domains. For example, on GPQA Diamond (graduate-level science questions) accuracy fell from 93.18 percent (second place) to 55.56 percent (94th place) when refusals were counted as failures.
  • Agents’ Last Exam: tasks answered by Claude Code/Claude Fable 5 itself earned a pass rate of 22.8 percent, close to Codex/GPT-5.5 (23.8 percent) and ahead of Claude Code/Claude Opus 4.8 (15.8 percent). On tasks where safeguards diverted responses to Claude Opus 4.8, the pass rate fell to 17.6 percent; Claude Fable 5’s composite pass rate was 22.0 percent, behind GPT-5.5 at 24.0 percent.

Why this matters

Anthropic’s safety mechanisms prevent direct, stable measurement of the publicly available Claude Fable 5. Measuring the model with the classifiers disabled would reflect a version the public cannot access; measuring it with classifiers enabled describes a moving target because Anthropic can retune the classifiers at any time.

Benchmarks historically ask how capable a model is; Claude Fable 5 forces a different, more practical question: how much of that capability actually reaches users? Evaluators now must report not only peak scores but what developers and end users can reliably expect in practice.

Closing note

Published reports indicate Claude Fable 5 is particularly strong at coding tasks, but the combination of classifiers and the 30-day retention requirement constrains independent assessment of its practical performance until access or evaluation conditions change.