Insilico Medicine Launches Benchmark Service for Testing AI on Drug Discovery
Insilico Medicine has launched a DDD Benchmark as a Service to test frontier AI and foundation models against real-world drug discovery and development tasks.
Insilico Medicine has launched a DDD Benchmark as a Service to test frontier AI and foundation models against real-world drug discovery and development tasks.
Moonshot’s Kimi K3 leads Code Arena: Frontend but trails top OpenAI and Anthropic models on FrontierMath Tier 4, highlighting uneven model strength.
Five AI labs are reportedly backing a common jailbreak scoring scale by August 1, an early step toward more comparable AI model safety testing.
OpenAI unveiled GeneBench-Pro, a genomics benchmark meant to measure higher-order scientific reasoning as AI labs push into biology workflows.
Arena, the popular free AI model leaderboard startup, launched its commercial service last September and has now grown into a $100M business.
A CAISI evaluation says DeepSeek V4 Pro is China's strongest model but still trails leading US frontier AI systems.
A rigorous new benchmark tested top AI models on investment banking tasks; not one output was deemed client-ready, though half of bankers found value as a starting point.
A new benchmark reveals that even top AI models drop roughly 50% in accuracy when analyzing complicated charts, exposing a key limitation in visual reasoning.
Google's enhanced Gemini 3 Deep Think model demonstrates superior performance over OpenAI's GPT-5.2 and Anthropic's Claude Opus 4.6 in latest benchmark tests.
Claude Opus 4.6 achieves breakthrough performance with 65.4% on Terminal-Bench and 72.7% on OSWorld, surpassing Gemini 3 Flash in real-world work applications.
Latest News and Analysis in AI Benchmarks