Reports Split Over Chinese AI Model Said to Challenge ChatGPT and Claude
Two reports make conflicting claims about GLM-5.2 and Qwen3.8-Max, highlighting how thin benchmark evidence can distort comparisons with ChatGPT and Claude.
Two reports make conflicting claims about GLM-5.2 and Qwen3.8-Max, highlighting how thin benchmark evidence can distort comparisons with ChatGPT and Claude.
Databricks will make GLM 5.2 its default coding model after internal tests found Opus-like performance at lower per-task cost on real codebase work.
Google has unveiled Gemini 3.1 Pro, its latest AI model featuring doubled abstract reasoning performance on ARC-AGI-2, surpassing rivals GPT-5.2 and Claude Opus 4.6 across 12 of 19 major benchmarks.
Latest News and Analysis in LLM Benchmarks