AWS benchmarks 30B MoE models on SageMaker AI across G5, G6, G6e and G7, showing how Blackwell may improve inference cost and latency.

Amazon Web Services is using two 30-billion-parameter Mixture-of-Experts models to compare GPU generations for production inference on Amazon SageMaker AI. The company’s benchmark examines G5, G6, G6e and new G7 configurations across throughput, latency and cost-per-token, with AWS reporting that G7 instances powered by NVIDIA Blackwell can deliver better price-performance for selected workloads.
The benchmark matters because model size alone does not determine serving economics. Quantization, memory bandwidth, expert routing and the serving software stack can all change which instance is the most practical choice. AWS’s results, published in its Machine Learning Blog, are therefore best read as a deployment study rather than a universal ranking of GPU families.
AWS tested Qwen3-Coder-30B-A3B-Instruct-FP8 for coding use cases including code generation, debugging, refactoring and developer copilots. That experiment compares ml.g5.12xlarge instances using NVIDIA A10G GPUs, ml.g6.12xlarge instances using NVIDIA L4 GPUs, and ml.g7.12xlarge instances using RTX PRO 4500 Blackwell GPUs. The test uses the SageMaker AI DJL Large Model Inference container.
The company’s second scenario uses NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4 for reasoning, question answering, summarization and agentic workloads. Instead of only deploying fixed endpoints, AWS uses SageMaker AI Generative AI Inference Recommendations with vLLM to evaluate G6, G6e and G7 options and identify a configuration based on a selected optimization target.
That distinction is important. The Qwen3-Coder-30B test represents direct comparison of deployed endpoints under an expected coding workload. The Nemotron-3-Nano-30B exercise represents an automated recommendation process in which SageMaker AI evaluates candidate configurations and ranks them against metrics such as cost, latency and throughput.
The configurations do not contain equivalent amounts of GPU memory. For the Qwen3-Coder-30B comparison, the G5 and G6 instances each use four GPUs with 96 GB of aggregate GPU memory, while the G7 configuration uses two GPUs with 64 GB. AWS presents that imbalance as part of the test: the newer instance is being evaluated despite having fewer accelerators and less total memory.
The second comparison is similarly uneven. The G6 setup has four L4 GPUs and 96 GB of aggregate memory, G6e has four L40S GPUs and 192 GB, and G7 has two GPUs with 64 GB. This makes the exercise closer to a production selection problem than a simple generation-over-generation benchmark. A configuration that wins on price-performance may still be unsuitable if the model, weights or runtime state cannot fit within its available memory.
AWS points to two hardware characteristics behind the potential G7 advantage. First, G7 uses NVIDIA Blackwell GPUs, which provide native support for low-precision formats including NVFP4. Second, the architecture offers memory and compute capabilities intended to improve token generation. AWS says the other tested generations can run NVFP4 weights, but without the same hardware acceleration.
For MoE models, AWS argues that memory bandwidth is especially relevant during decoding because each token activates only a subset of experts. That can make data movement and inter-token latency more important than the model’s headline parameter count. The practical result is workload-dependent: a model that benefits from low-precision acceleration may show a different instance preference from one constrained by memory capacity or a particular serving framework.
The available evidence comes from AWS’s own technical post and walkthroughs. AWS states that the G7 tests show measurable improvements in throughput, latency and cost-per-token, and describes G7 as a potential price-performance leader for the tested use cases. Those are vendor-reported findings, not an independent assessment or a benchmark conducted across multiple cloud providers.
The supplied source material does not include the underlying numerical results, test duration, request distribution, concurrency levels, regional prices or detailed cost-per-token tables. It therefore cannot establish how large the gains are, how consistently they appear, or whether the same outcome would apply to other models and traffic patterns.
The results are also shaped by the chosen software. Qwen3-Coder-30B is evaluated through the DJL Large Model Inference container, while the Nemotron-3-Nano-30B workflow uses vLLM and SageMaker AI’s recommendation tooling. Changes in batching, scheduling, quantization, context length or request concurrency could alter the outcome. Buyers should reproduce the test with their own prompts and service-level objectives before treating AWS’s recommendation as a production decision.
AWS says the recommendation feature runs candidate configurations on real GPU infrastructure and returns validated, deployment-ready suggestions based on measured performance. Even so, “validated” in this context refers to the service’s benchmarking workflow, not independent certification of the model’s quality, safety or business suitability.
For builders, the immediate lesson is to benchmark the complete serving configuration rather than choosing hardware from a GPU specification sheet. A team deploying a coding assistant should measure time to first token, inter-token latency, sustained throughput and cost under realistic concurrency. A team building an agentic application may need to weigh short prompts and bursty traffic differently from a batch summarization service.
SageMaker AI’s recommendation workflow could reduce the manual work involved in testing several instance families, especially for teams that already manage endpoints on AWS. It does not eliminate the need to define the workload accurately. An incorrect traffic profile or optimization goal can produce a technically valid recommendation that is poorly matched to production.
Memory remains a deployment constraint. G7’s lower aggregate memory in the examples may make it attractive for models that fit efficiently through FP8 or NVFP4, but less suitable for deployments requiring larger context windows, extensive key-value caches or additional model components. G6e’s larger memory footprint may remain preferable when capacity is more important than raw token economics.
For enterprise buyers, the comparison also highlights portability risk. The value of NVIDIA Blackwell acceleration depends on supported formats, model implementations and runtime maturity. Teams standardizing on a model should verify that their preferred quantization path and inference engine use the hardware effectively rather than assuming that a newer GPU automatically reduces total serving cost.
The most useful follow-up would be AWS publishing the full benchmark tables, including throughput, time to first token, inter-token latency, concurrency, regional pricing and cost-per-token for every configuration. Those figures would show whether G7’s advantage is broad or concentrated in particular workload shapes.
Teams should also watch G7 availability. AWS notes that G7 is generally available in US East (Ohio) and US West (Oregon) in the described setup, making regional capacity and data-residency requirements part of the buying decision.
Further comparisons should test longer contexts, different batch sizes, additional quantization formats and models outside the two selected 30B MoE examples. Independent replication across clouds would provide a stronger basis for evaluating whether Blackwell’s advantage survives different pricing and software environments.
AWS is not announcing a new model or a universal replacement for G5 and G6. It is making a narrower but important infrastructure case: newer accelerators can change the economics of small and mid-sized LLM serving when the model and runtime expose the right low-precision and memory-bandwidth advantages.
The strongest signal for AI teams is the move toward workload-specific, automated instance selection. But because the published evidence is AWS-controlled and lacks the numerical benchmark details in the supplied material, buyers should treat G7’s reported gains as a starting hypothesis. The production winner will depend on model fit, latency targets, traffic shape, regional availability and the cost of operating the full inference stack.