TensorRT Edge-LLM Completes MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA says TensorRT Edge-LLM ran Qwen3.6-27B 6.4x faster on Jetson AGX Thor, highlighting cache and quantization gains for edge agents.

AI News

NVIDIA says its TensorRT Edge-LLM runtime completed the MLPerf Inference v6.1 Edge Agentic benchmark in 24 minutes and 36 seconds on a single Jetson AGX Thor Developer Kit. That was 6.4 times faster than the benchmark’s published llama.cpp reference run on the same platform, according to NVIDIA’s developer blog.

The result is significant because the test measures more than one-shot response speed. It replays software-engineering agent conversations in which a model generates tool calls, receives results, and continues through increasingly long contexts. NVIDIA’s submission ran Qwen3.6-27B at 52.33 tokens per second and completed all 1,007 generated turns in the performance workload.

The figures are vendor-reported results from NVIDIA’s own submission and should not be treated as an independent comparison of every edge inference stack. They nevertheless show how runtime-level optimizations can change the economics of running multi-step AI agents on local hardware rather than sending every interaction to a cloud data center.

What the MLPerf test measured

The MLPerf Edge Agentic benchmark evaluates an OpenAI-compatible model endpoint in performance and accuracy phases. Its performance workload contains 20 recorded conversations and 1,007 generated turns. As each agent receives tool results and continues reasoning, the context grows to approximately 23,500 tokens.

That structure creates a different bottleneck from a conventional chatbot test. An agent repeatedly processes much of the same conversation, while also needing to produce valid function calls quickly. Recomputing the entire history at every turn can make long-running trajectories increasingly expensive, particularly on a device with limited memory bandwidth and power.

The accuracy phase uses prompts from the Berkeley Function Calling Leaderboard, or BFCL, with single-turn requests and reasoning disabled. It checks whether the model selects the appropriate function, supplies valid arguments, or correctly declines to call a tool. NVIDIA says its submission ran in SingleStream mode on one Jetson AGX Thor Developer Kit with 128 GB of unified memory at the platform’s MAXN power mode.

The engineering behind NVIDIA’s result

NVIDIA attributes the result to several optimizations in TensorRT Edge-LLM rather than to a single hardware feature. The submitted Qwen3.6-27B model used NVFP4 for weights and activations, including the language-model head, and FP8 for its key-value cache, or KV cache.

NVFP4 is a four-bit floating-point format supported by the Blackwell GPU in Jetson AGX Thor. Lower-precision representations reduce the amount of data that kernels must move through memory, an important consideration for low-batch decoding, which NVIDIA describes as heavily constrained by DRAM bandwidth on edge platforms. The smaller representation also leaves more of the device’s unified memory for context, speculative decoding state, and other application tasks.

The runtime also reused cached conversation state between turns. NVIDIA says approximately 96% of prompt tokens were served from hot cache across the full agent trajectory. Instead of prefilling all 13.6 million prompt tokens encountered during the benchmark, the system prefilling only about 0.5 million tokens of new conversation suffixes.

That optimization is especially relevant to agent workloads. Each new turn contains most of the earlier conversation, plus a tool result or model response. TensorRT Edge-LLM identifies reusable prompt prefixes and restores cached attention pages. Because Qwen3.6 uses a hybrid model architecture, NVIDIA says the runtime also restores recurrent state and partial KV-page state needed to continue execution correctly.

For token generation, NVIDIA used tree-based multi-token prediction. The configuration used an eight-step, top-two, 16-node verification tree and delivered an approximately 40% decoding improvement over linear multi-token prediction for the function-calling workload, according to the company.

Evidence, comparison and limits

The central comparison is between NVIDIA’s TensorRT Edge-LLM submission and the llama.cpp reference run published through the MLCommons Edge Agentic example. Both runs used Qwen3.6-27B on Jetson AGX Thor, but the reference used Q4_K_M quantization, while NVIDIA’s submission used NVFP4 and additional runtime techniques.

NVIDIA reports that the llama.cpp run took two hours and 37 minutes to complete the same workload. The TensorRT Edge-LLM result took 24 minutes and 36 seconds. The difference therefore reflects the full software stack, quantization choices, cache handling, and decoding strategy—not simply a direct measurement of one model format against another.

The benchmark also does not establish that the system will be 6.4 times faster across all agent applications. Results can vary with model architecture, context length, tool-call patterns, quantization quality, power settings, and the amount of parallel work. NVIDIA’s approximately 40% multi-token prediction gain is likewise a workload-specific claim tied to the reported function-calling configuration.

The available evidence is strongest on completion time and tokens per second in this defined benchmark. It is thinner on production reliability, energy use, thermal behavior over extended deployments, and accuracy tradeoffs under broader agent tasks. NVIDIA says developers can inspect the TensorRT Edge-LLM release/0.9.1-mlpinf branch, use its calibrated Qwen3.6-27B NVFP4 checkpoint, and review the MLCommons example for configuration details, but those materials do not by themselves constitute independent adoption evidence.

Why the result matters for edge AI builders

For developers building AI agents into vehicles, robots, industrial equipment, or other embedded systems, the result points to a practical deployment question: how much of an agent’s repeated context can be kept local and reused? Cache reuse can reduce prefill work without changing the application’s conversation flow, while multi-token prediction targets the generation phase that follows tool responses.

The memory savings may be as important as raw speed. A 27-billion-parameter model running in a low-precision format still needs room for long contexts, runtime state, and application logic. NVIDIA’s use of 128 GB of unified memory suggests that edge deployments may increasingly be designed around the combined footprint of the model and the agent—not the model alone.

For enterprise buyers, the benchmark offers a signal that local agent inference is becoming more viable for workflows where latency, connectivity, or data control matters. It does not remove the need to evaluate power consumption, model accuracy, tool-call reliability, update processes, and safety controls in the target environment. A faster benchmark run is useful only if the agent can make correct decisions and operate consistently under real hardware constraints.

The result also raises competitive pressure on alternative runtimes. TensorRT Edge-LLM’s advantage here comes from close coordination between NVIDIA software and Blackwell hardware. Developers using other accelerators will need comparable support for low-precision kernels, persistent context state, and speculative or tree-based decoding if they want similar performance on long agent trajectories.

What to watch next

The next useful signals will be independent reproductions of the MLPerf result, especially comparisons that report energy consumption, thermal limits, and accuracy alongside completion time. It will also be important to see whether cache reuse remains effective on workloads with less repetitive context or more varied tool outputs.

Developers should watch the TensorRT Edge-LLM release branch and the MLCommons Edge Agentic example for updated model support, build instructions, and additional hardware submissions. Broader testing of Qwen3.6-27B NVFP4 will help clarify whether the reported performance is practical for production agents rather than limited to the benchmark’s recorded trajectories.

Finally, evidence from deployed vehicles, robots, and industrial systems would provide a stronger test of the approach. Those environments introduce intermittent connectivity, strict power budgets, sensor inputs, safety requirements, and software updates that the current benchmark does not fully capture.

Creati.ai perspective

NVIDIA’s result is best understood as a systems demonstration: edge agent performance depends on avoiding repeated work as much as on increasing raw generation speed. The combination of NVFP4, FP8 KV cache, state reuse, and tree-based decoding addresses the specific costs created by long, tool-using conversations.

The 6.4x figure is compelling within the reported MLPerf setup, but the broader takeaway is more measured. For AI builders, the important question is whether these optimizations survive different models, tools, power envelopes, and real-world accuracy requirements. If they do, local agent deployments could move from isolated demonstrations toward more credible production architectures.

Ads