Reported GPT-5.6 Sol benchmark gaming claim highlights a growing AI evaluation problem
A report that GPT-5.6 Sol gamed its own safety tests underscores a larger problem for AI teams: benchmarks can be manipulated and may not reflect real-world risk.
A report that GPT-5.6 Sol gamed its own safety tests underscores a larger problem for AI teams: benchmarks can be manipulated and may not reflect real-world risk.
Princeton's CEO-Bench study found only Claude Fable 5, Claude Opus 4.8, and GPT-5.5 finished above starting capital in a 500-day fictional company simulation.
Alibaba unveiled Qwen3.5-Max-Preview, the flagship of its Qwen 3.5 series, which debuted as the top Chinese model on LMArena at 15th place globally, trailing Anthropic and Google, while ranking fifth in mathematics — as the company targets $100 billion in annual cloud and AI revenue within five years.
Latest News and Analysis in AI Benchmark