AWS Targets SageMaker Inference Delays With HyperPod Model Caching and Prefix Routing

AWS adds model caching to SageMaker HyperPod and prefix-aware routing to SageMaker Inference, targeting faster scale-out and lower LLM latency.

AI News

AWS is adding two infrastructure features aimed at separate but related bottlenecks in large-language-model serving: model caching for Amazon SageMaker HyperPod and prefix-aware routing for SageMaker Inference. The first is designed to shorten the time needed to bring new inference pods online; the second is intended to reduce response latency by keeping frequently reused prompt computation on the same instance.

The changes matter most for teams operating large models under variable traffic. Without caching, a scale-out event can require new nodes to download multi-gigabyte container images and model weights before they can serve requests. Without routing that accounts for prompt content, a serving framework’s prefix cache may remain underused because identical prompt sections are spread across a fleet.

Both announcements come from the AWS Machine Learning Blog, so the performance figures and operational claims are AWS-reported rather than independently verified. Together, they outline a more coordinated approach to reducing both inference cold starts and steady-state time to first token.

HyperPod caching addresses the scale-out gap

AWS says model deployment on SageMaker HyperPod can be delayed by two sequential downloads. Kubernetes first pulls an inference server image from Amazon Elastic Container Registry, after which the server downloads model weights from a source such as Amazon S3, Amazon FSx for Lustre, Hugging Face Hub, or JumpStart.

For smaller models, this process may take minutes. AWS gives the example of a 145 GB model whose weight download can take more than 20 minutes from Amazon S3, depending on network conditions. For a model such as DeepSeek-R1, which AWS describes as weighing more than 600 GB in the cited scenario, the process can take 30 minutes or longer. Container image pulls alone are estimated by AWS at five to seven minutes for typical multi-gigabyte inference images.

That delay creates a mismatch between autoscaling and actual capacity. A HorizontalPodAutoscaler might request additional pods quickly, but those pods cannot accept traffic until their images and weights are available. A sudden request spike can therefore trigger an operational response that arrives tens of minutes too late.

The new model caching capability preloads weights onto local NVMe storage on target nodes. AWS says the HyperPod Inference Operator downloads the weights in advance, marks nodes as cache-ready, and waits for the target nodes to complete the process before creating the inference deployment. Once a pod starts on a prepared node, it can read locally at approximately 7 GB per second instead of downloading the model over the network.

AWS also offers an independent image cache. A DaemonSet pre-pulls the inference container image onto nodes, allowing later pods to skip the Amazon Elastic Container Registry download. Multiple deployments using the same image can share that cache, while the operator manages references and cleanup.

How the two caching systems behave in production

Weights and image caching do not have identical deployment semantics. AWS says weights caching can delay creation of the inference deployment until all target nodes are ready, while image caching does not block deployment creation. A pod can therefore start before an image cache has finished on a particular node and fall back to a normal image pull.

Both mechanisms use preferred rather than mandatory scheduling. Pods are directed toward nodes with warm data when possible, but they are not prevented from running elsewhere. If rapid scale-out exceeds the number of prepared nodes, a pod can use the original model source and pull its image normally. The trade-off is slower startup, not a failed deployment.

The operator manages two underlying custom resources: ModelDataCacheConfig for model weights and ModelImageCache for container images. Users enable caching through modelCacheConfig in an InferenceEndpointConfig or JumpStartModel resource rather than managing those lifecycle objects directly.

AWS also says cache updates are handled when a model source or image changes. The operator creates a new cache, rolls out the updated deployment, and then removes the old cache. That is intended to prevent stale weights while supporting zero-downtime transitions, although the practical result will still depend on available node storage and the time required to populate the replacement cache.

Prefix-aware routing keeps LLM prompt caches useful

The second SageMaker change targets a different layer of the serving stack. Many LLM requests contain a long, repeated prefix—such as system instructions, retrieved documents, conversation history, or source code—followed by a short user-specific suffix. Frameworks including vLLM and TensorRT-LLM can reuse the computed key-value, or KV, cache for that repeated prefix.

Random request distribution weakens that benefit in a multi-instance endpoint. If successive requests with the same prefix land on different machines, each instance may have to recompute the shared context. SageMaker Inference’s new prefix-aware routing strategy examines the beginning of a request and consistently sends matching prefixes to the same instance.

AWS says the feature can also protect the fleet from an overly popular prefix. If the preferred instance reaches its configured concurrency limit, the request can be redirected to a less busy instance. That may sacrifice a cache hit for one request, but it avoids concentrating traffic on a single machine. AWS further says that adding or removing instances should shift only a limited portion of traffic, helping preserve cache locality during scaling.

The strategy is configured per production variant and can be changed through endpoint configuration without redeploying the model. AWS continues to offer random routing as the default, along with least-outstanding-requests routing for workloads where request durations vary. Prefix-aware routing is specifically aimed at LLM workloads with shared leading context and an enabled prefix cache.

Evidence, benchmarks, and limits

AWS benchmarked prefix-aware routing against random routing using Llama 3.1 70B Instruct on seven ml.p5.48xlarge instances with vLLM and prefix caching enabled. Across 16 test configurations covering different endpoint and API arrangements, AWS reports up to a 77% reduction in median time to first token, throughput gains of up to 16%, and an increase in KV cache hit rate from roughly 25% to more than 80%.

Those are vendor-reported benchmark results, not an independent evaluation. AWS says traffic remained balanced in the tests, with each instance receiving 13.3% to 15.4% of requests. It also reports an additional routing cost of 1.3 to 1.9 milliseconds per request, compared with model time-to-first-token results ranging from 63 to 280 milliseconds in the tested configurations.

The size of the benefit depends heavily on workload shape. AWS says longer shared prefixes produce larger gains because more computation can be skipped. RAG systems that repeatedly query the same document, multi-turn conversations, templated assistants, and code completion are among the use cases AWS identifies. Workloads with short or mostly unique prompts should see less value.

The HyperPod claims are similarly dependent on conditions that are not fully specified in the source, including node availability, cache population time, storage capacity, model format, and network performance during the initial preload. The stated transition from tens of minutes to seconds applies when a pod lands on a node with the relevant data already cached; an unprepared node follows the normal download path.

What the changes mean for AI infrastructure teams

For builders and enterprise platform teams, the announcements separate two decisions that are often treated as one latency problem. Model caching improves elasticity: it can make new capacity useful sooner when traffic rises. Prefix-aware routing improves request efficiency: it can reduce repeated prefill work after capacity is already online.

The combination could be useful for RAG applications and assistants that have both bursty traffic and large repeated contexts. A team might use HyperPod caching to prepare nodes for scale-out while using prefix-aware routing to keep document or conversation prefixes warm across the active fleet. That does not remove the need to size local NVMe capacity, configure concurrency limits, or measure cache hit rates under real traffic.

There are also cost and reliability questions for buyers to test. Keeping weights on every target node consumes local storage and may increase preparation time before a deployment is ready. Prefix affinity can improve latency, but it introduces content-dependent routing behavior that teams should observe alongside queue depth, instance utilization, cache occupancy, and tail latency. Privacy and data-governance reviews may also matter because routing decisions inspect the beginning of request payloads, even though AWS handles the routing automatically.

What to watch next

Teams evaluating the features should look for independently reproduced results on models, hardware, and prompt distributions beyond AWS’s Llama 3.1 70B test. The most useful measurements will include p95 and p99 time to first token, cache-hit rates during scale-out, and the percentage of requests that land on unprepared nodes.

Operational guidance around local NVMe sizing, cache warm-up orchestration, and multi-model clusters will also be important. Buyers should verify how cache preparation affects deployment rollouts, node replacement, spot or interruption scenarios, and rapid scaling beyond the preloaded fleet.

Finally, AWS’s endpoint-level routing controls may become more significant as hosted serving platforms compete on predictable latency rather than model access alone. The follow-up signal will be whether customers can use these controls without adding substantial scheduling complexity or sacrificing balanced GPU utilization.

Creati.ai perspective

AWS is addressing two practical weaknesses in LLM operations rather than introducing a new model capability. HyperPod model caching reduces the penalty for adding capacity, while prefix-aware routing makes an existing serving optimization—KV cache reuse—more dependable across multiple instances.

The strongest case is for predictable, repeated-context workloads with enough traffic to justify preloading large models. For teams with mostly unique prompts or infrequent deployments, the storage and preparation overhead may outweigh the gains. AWS’s benchmark numbers are encouraging, but infrastructure buyers should validate them against their own prompt overlap, scaling patterns, and latency objectives before treating caching as a guaranteed improvement.

Ads