Claude Opus 5 Just Beat GPT-5.6 Sol on the Hardest Benchmark That Exists
Opus 5 scored 43.3% on Frontier-Bench v0.1 at max effort vs. Sol's 37.5% — days after Sol's own sandbox breach made headlines for the opposite reason. The full 8-benchmark scorecard, the technical anatomy of the breach, and what the gap says about the capability/safety tradeoff both labs are making.