TensorRT Edge-LLM Completes MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
NVIDIA says TensorRT Edge-LLM ran Qwen3.6-27B 6.4x faster on Jetson AGX Thor, highlighting cache and quantization gains for edge agents.
Latest News and Analysis in LLM Performance
NVIDIA says TensorRT Edge-LLM ran Qwen3.6-27B 6.4x faster on Jetson AGX Thor, highlighting cache and quantization gains for edge agents.
NVIDIA appears to have released tri-mode diffusion LLM weights, signaling new interest in faster inference, but public evidence is still limited.
Claude Opus 4.6 achieves breakthrough performance with 65.4% on Terminal-Bench and 72.7% on OSWorld, surpassing Gemini 3 Flash in real-world work applications.