GitHub launches ReviewBench to evaluate AI code review agents
GitHub’s open ReviewBench benchmark targets a clearer way to test AI code review agents, giving builders and buyers a common evaluation starting point.
Latest News and Analysis in AI Benchmarks
GitHub’s open ReviewBench benchmark targets a clearer way to test AI code review agents, giving builders and buyers a common evaluation starting point.
AWS benchmark data suggests OpenAI models on Amazon Bedrock can lower the cost of correct answers when teams measure quality, turns, and rework.
StartLux’s 27B local model reportedly outscored DeepSeek V4 Flash in a China benchmark, raising questions about edge AI performance and proof.
Aikido Security says it spent 11.7 billion tokens comparing cyber AI models, raising questions about benchmark design, cost, and evidence for buyers.
Insilico Medicine has launched a DDD Benchmark as a Service to test frontier AI and foundation models against real-world drug discovery and development tasks.
Moonshot’s Kimi K3 leads Code Arena: Frontend but trails top OpenAI and Anthropic models on FrontierMath Tier 4, highlighting uneven model strength.
Five AI labs are reportedly backing a common jailbreak scoring scale by August 1, an early step toward more comparable AI model safety testing.
OpenAI unveiled GeneBench-Pro, a genomics benchmark meant to measure higher-order scientific reasoning as AI labs push into biology workflows.
Arena, the popular free AI model leaderboard startup, launched its commercial service last September and has now grown into a $100M business.
A CAISI evaluation says DeepSeek V4 Pro is China's strongest model but still trails leading US frontier AI systems.
A rigorous new benchmark tested top AI models on investment banking tasks; not one output was deemed client-ready, though half of bankers found value as a starting point.
A new benchmark reveals that even top AI models drop roughly 50% in accuracy when analyzing complicated charts, exposing a key limitation in visual reasoning.
Google's enhanced Gemini 3 Deep Think model demonstrates superior performance over OpenAI's GPT-5.2 and Anthropic's Claude Opus 4.6 in latest benchmark tests.
Claude Opus 4.6 achieves breakthrough performance with 65.4% on Terminal-Bench and 72.7% on OSWorld, surpassing Gemini 3 Flash in real-world work applications.