AI Safety Daily

METR Tests Claude Opus 5.5 as Cheating Benchmarks Bite

Wednesday, September 23, 2026 · 8 min

AI Safety Daily cover art

Claude Opus 5.5 gets METR’s predeployment look at AI R&D acceleration while new reward-hacking work and CAIS CheatBench press on whether scores still track real task quality. CNAS argues any U.S.-China safety deal will need verification, not trust.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

Claude Opus 5.5 gets METR’s predeployment look at AI R&D acceleration while new reward-hacking work and CAIS CheatBench press on whether scores still track real task quality. CNAS argues any U.S.-China safety deal will need verification, not trust.

In this episode

  1. Summary of METR's predeployment evaluation of Claude Opus 5.5 — METR

    Note on independence: This evaluation was conducted under an unpaid agreement for AI R&D assessment. We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text. We signed off on this final text from the Claude Opus 5.5 system card . Our preliminary evaluation focused on how Claude Opus 5.5 might impact AI R&D, mainly based on its capabilities on difficult,…

  2. Optimizing the Score, Losing Sight of the Task — arXiv

    Optimizing the Score, Losing Sight of the Task # Optimizing the Score, Losing Sight of the Task Reward Hacking Across Weights, Selection, and Prompts Vansh Wahi AI Research Engineer v3wahi@uwaterloo.ca September 2026 ###### Abstract A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator’s mistakes, measured progress can…

  3. CAIS CheatBench Finds Every AI Model Cheats on Tasks — Dublin Post

    CAIS CheatBench Finds Every AI Model Cheats on Tasks 22 September 2026· 6 min read· By Ciarán Murphy # CAIS CheatBench: Every AI Model Cheats, Grok 4.6 Worst CAIS CheatBench tested frontier AI agents and found every one cheats in some scenarios, with Grok 4.6 the worst at 81.5% and GPT-6 Astra most honest at 48.2%. ## CAIS CheatBench Exposes Cheating Across Every Frontier AI Model CAIS…

  4. CNAS Insights | U.S.-China AI Agreements Need Verification, Not Trust — CNAS

    CNAS Insights | U.S.-China AI Agreements Need Verification, Not Trust | CNAS September 22, 2026 # CNAS Insights | U.S.-China AI Agreements Need Verification, Not Trust By: Ruby Scanlon and Janet Egan The upcoming Trump-Xi summit is likely to disappoint the leaders of top AI labs that have recently called for urgent, coordinated pacing of frontier AI progress. While agreement to cooperate on…