KT’s AutoModelRouter Reportedly Ranks Second in Global AI Routing Benchmark
KT’s AutoModelRouter reportedly placed second in a global benchmark, putting model selection efficiency at the center of enterprise AI deployment.
Latest News and Analysis in AI Benchmarking
KT’s AutoModelRouter reportedly placed second in a global benchmark, putting model selection efficiency at the center of enterprise AI deployment.
Vals, backed by Andreessen Horowitz, raised $40 million to build confidential, task-based AI benchmarking for companies and federal agencies as older tests s…
AWS benchmark data suggests OpenAI models on Amazon Bedrock can lower the cost of correct answers when teams measure quality, turns, and rework.
Google DeepMind is testing a cryptographically protected double-blind benchmark for Gemini Flash Lite, aiming to reduce contamination and rebuild trust in AI evaluations.
Motif reportedly leads a global AI benchmark for Korean models, while the industry awaits a second Dokpamo evaluation that could test whether the result holds.
Anthropic’s Claude Opus 5 set a new ARC-AGI-3 high score, raising fresh questions about real reasoning gains versus benchmark targeting.
OpenAI says about 30% of SWE-Bench Pro tasks may be broken, raising new doubts about how the AI industry measures coding models.
A new generative AI-powered benchmarking system reveals China's dominance in the autonomous vehicle and robotaxi sector over US competitors.
Google DeepMind launches Werewolf and poker benchmarks on Kaggle Game Arena to test AI social skills, deception detection, and risk management. Gemini 3 Pro and Flash models demonstrate significant performance leap over previous generation.