AWS Brings Moonshot AI’s Kimi K3 to Amazon Bedrock With Long Context and Prompt Caching

AWS has added Moonshot AI’s Kimi K3 to Amazon Bedrock, giving builders an open-weight model with vision, long context, and prompt caching.

AI News

Amazon Web Services has made Moonshot AI’s Kimi K3 available on Amazon Bedrock, adding an open-weight model aimed at coding and knowledge-work workloads. The release gives developers access to native image understanding, a 1-million-token context window, and explicit prompt caching through AWS-managed inference APIs.

The launch matters most for teams building long-running agents and coding systems that repeatedly send large repositories, reference documents, or tool instructions. AWS says Kimi K3 can be tested in the Bedrock console or called programmatically through Bedrock APIs, including OpenAI-compatible interfaces. However, the strongest capability and efficiency claims in the announcement come from Moonshot AI or AWS, rather than independent evaluations.

Kimi K3 arrives as an open-weight Bedrock option

Kimi K3 was developed by Moonshot AI, the company behind the Kimi model family. According to the AWS Machine Learning Blog, Moonshot AI describes Kimi K3 as its most capable model and claims it is the first open model to reach 2.8 trillion parameters. AWS also reports Moonshot’s claim of an approximately 2.5-times improvement in scaling efficiency over Kimi K2.

Those figures are vendor-reported and the available announcement does not provide an independent benchmark methodology, pricing comparison, or evaluation results against competing models. The more concrete change for developers is deployment access: Kimi K3 can now be selected in Amazon Bedrock alongside other supported models, rather than requiring a separate serving stack managed by the customer.

The model combines native vision capabilities with a 1-million-token context window. That configuration is relevant to applications that need to keep large amounts of material in working context, including software repositories, technical documentation, lengthy business records, and image-containing files. A large context window does not by itself guarantee reliable reasoning across all of that material, so production teams will still need retrieval, evaluation, and context-management controls.

Prompt caching targets repeated context

AWS is positioning explicit prompt caching as one of Kimi K3’s main practical differentiators on Bedrock. The feature allows developers to mark a reusable prompt prefix, such as repository instructions, tool definitions, or reference material. The prefix must contain at least 1,024 tokens and can be reused for later model calls.

When a subsequent request matches the cached prefix, AWS says Bedrock can lower response latency and input-token costs. Cached tokens are billed at a higher rate when written, but AWS says they remain available for at least 30 minutes. Matching requests receive discounted input-token pricing, and cached tokens do not count against input-tokens-per-minute quotas.

That design is particularly relevant to AI coding assistant and agent workflows. An agent may resend the same system guidance, codebase map, or tool schema over dozens of turns. Caching could reduce the repeated-input burden, although the financial benefit will depend on cache-hit rates, prompt size, request frequency, and the applicable regional pricing. AWS’s announcement does not state Kimi K3’s per-token prices.

Bedrock handles access, APIs, and data controls

Developers can try Kimi K3 through the Amazon Bedrock console by opening Test and Playground, then selecting the model. Applications can use the Bedrock Runtime endpoint, the Amazon Bedrock Invoke and Converse APIs, or OpenAI-compatible Responses and Chat Completions APIs.

AWS supports the model through cross-Region inference profiles. The global profile, identified as global.moonshotai.kimi-k3, can route requests to supported commercial AWS Regions worldwide. AWS says global cross-Region inference costs approximately 10% less than a geographic profile. For customers with US data-residency requirements, the announcement lists us.moonshotai.kimi-k3 as the US geographic profile.

AWS also says Kimi K3 inherits the platform’s stated controls for open-weight models: data is processed within the AWS data boundary, is not shared with the model provider, and is not used to train the underlying model. The company says zero data retention is enabled for inference requests and that zero operator access prevents AWS operators from accessing prompts and completions during inference. These are AWS service and policy claims; buyers should still validate the exact configuration, regional routing, logging, and contractual terms for their workloads.

Evidence, ecosystem, and builder implications

The announcement places Kimi K3 within AWS’s broader expansion of open-weight models on Bedrock. AWS says the service has added dozens of models since 2025 from providers including DeepSeek, Google, MiniMax, Mistral AI, Moonshot AI, NVIDIA, OpenAI, and Qwen. It also says Bedrock added platform-level support in 2026 for tool calling, structured output, reasoning, response streaming, and the Responses and Chat Completions APIs.

For builders, platform-level features can reduce the integration work required when testing different models. A team can evaluate Kimi K3 using existing Bedrock authentication, permissions, observability, and application code instead of creating a separate inference pathway. AWS lists bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, and bedrock:CreateInference among the permissions needed to call the model.

The announcement also highlights OpenCode, an open-source, model-agnostic coding assistant with a native Amazon Bedrock provider, as well as Hermes Agent, an open-source productivity assistant supporting research and task automation. These examples show how Kimi K3 could fit into existing tools, but they are integration examples rather than evidence of broad adoption or superior performance.

For enterprise teams, the practical decision is likely to center on workload economics and reliability. Kimi K3’s long context and caching may help applications with stable, repeated inputs, while cross-Region inference can offer a cost or capacity trade-off. Teams with strict residency or regulatory requirements will need to choose geographic profiles carefully. They should also test output quality on their own repositories and documents, especially for long-horizon coding tasks where context length alone may not prevent errors.

What to watch next

The next useful signals will be independent evaluations of Kimi K3 on coding, vision, tool use, and long-context retrieval. Public pricing details and real-world cache-hit economics will determine whether explicit prompt caching materially changes total application cost.

Developers should also watch for production reports on latency, regional availability, rate limits, and failure behavior under sustained agent workloads. Adoption by coding assistants and agent frameworks could provide a clearer indication of whether Bedrock access turns Kimi K3 into a practical alternative to other hosted and self-managed models.

Creati.ai perspective

Kimi K3’s significance is less about a single parameter-count claim than about the combination of open-weight access, long context, native vision, and caching in a managed cloud environment. AWS is making it easier for teams to test those capabilities without giving up the surrounding Bedrock deployment model.

The main uncertainty remains evidence. Moonshot AI’s scale and efficiency claims are not independently substantiated in the supplied material, and a 1-million-token window will not automatically translate into dependable agent behavior. Builders should treat the launch as a new evaluation target: useful for workloads with repeated context and large inputs, but still subject to application-level testing for quality, cost, and data-governance fit.

Ads