AWS benchmark data suggests OpenAI models on Amazon Bedrock can lower the cost of correct answers when teams measure quality, turns, and rework.

Amazon Web Services is asking teams to rethink how they choose OpenAI models on Amazon Bedrock, arguing that the lowest price per million tokens may not produce the lowest production cost. In a new AWS Machine Learning Blog post, the company published an open-source benchmarking harness that compares model accuracy, agent turn counts, and the cost of acceptable work.
The analysis covers three OpenAI models available through Amazon Bedrock—gpt-5.6-luna, gpt-5.6-terra, and gpt-5.6-sol—and compares them with gpt-5.4-mini and gpt-5.4-nano through the OpenAI API. AWS says the aim is to calculate the cost of a correct answer, a passing research result, or an acceptable professional deliverable rather than treating token pricing as the main purchasing metric.
The findings are vendor-reported and come from AWS’s own test harness. They are also not a fully controlled comparison: AWS says the Bedrock models ran with reasoning disabled, while the API baselines used their default settings. The company advises customers to reproduce the tests on their own workloads before making a model decision.
The AWS argument is straightforward: production applications pay for outcomes, not tokens. A model that is cheaper per request can become more expensive if it answers incorrectly, requires retries, or forces a human to repair its output.
To test that idea, AWS ran the same OpenAI Responses API code path across five models while holding its evaluation logic constant. The benchmark included AIME mathematics, GPQA Diamond graduate-level science, and MMLU-Pro. AWS divided total model spend—including failed attempts—by the number of correct answers to estimate observed cost per correct answer.
In the company’s sample, gpt-5.6-sol achieved 75% accuracy on AIME, compared with 37% for gpt-5.4-mini. It also led the reported GPQA Diamond and MMLU-Pro results. AWS cautions that these are sample results, not a universal ranking, and that small differences should be treated as directional unless supported by uncertainty estimates.
The pricing conclusion changed after AWS cited a July 30, 2026 reduction for GPT-5.6 Luna and Terra on Amazon Bedrock. AWS reported a cost of $0.0021 per correct AIME answer for gpt-5.6-luna versus $0.0139 for gpt-5.4-mini under the assumptions in its result files. The post says Luna’s observed cost per correct answer was lower than the tested alternatives, including gpt-5.4-nano.
Those figures should not be read as current universal prices. AWS specifically tells readers to check the applicable Amazon Bedrock region and inference tier against the live pricing page. The post also notes that price-page updates may not align immediately with an announcement.
The more consequential part of the analysis concerns AI agents, where each model call can carry a growing conversation history. AWS tested a 50-question sample of DeepSearchQA using live web_search and fetch_page tools. The agent managed its own history with storage disabled, meaning the system prompt, previous tool results, and conversation context were sent again on each turn.
That design makes turn count a direct cost and latency factor. AWS says cumulative input can grow approximately quadratically as an agent continues to add context. In its sample, gpt-5.4-mini averaged 7.6 turns per question, largely because of repeated search loops. Its input volume reached 114,000 tokens per question, compared with 50,000 for gpt-5.6-terra.
AWS reported that Terra cost $0.31 per passing answer in the test, compared with $0.40 for mini, while producing a higher mean F1 score. It reported an even lower $0.05 per passing answer for gpt-5.6-luna, versus $0.40 for mini. Nano had a lower nominal token price but passed only 18% of the questions, resulting in an observed $0.07 per passing answer.
The sample contained only 50 questions, so the comparisons are not conclusive. Still, the result highlights a cost variable that does not appear on a standard pricing page: how efficiently a model completes a tool-using workflow.
AWS also tested GDPval, a benchmark built around occupational deliverables rather than short-answer questions. Its 48-task sample covered documents such as compliance briefs, financial plans, and care protocols, with each output graded against a rubric written by professionals.
The company reported that all three gpt-5.6 configurations scored higher than the tested mini and nano configurations when reasoning was disabled. Luna passed 27 of 48 tasks, compared with 20 for mini. After the reported price reduction, AWS calculated Luna’s cost per passing deliverable at $0.010, versus $0.030 for mini and $0.012 for nano.
The result is not a clean measure of general model quality. AWS says the occupational categories had small samples, and the output limit of 8,192 tokens truncated six Luna, nine Terra, seven Sol, zero mini, and one nano deliverable. A higher output limit could change both quality and cost.
For enterprise buyers, however, the test points to a practical question: how much is a higher pass rate worth when the alternative requires review, rework, or escalation? A more expensive model may be economical if it reduces those downstream costs, while a cheaper model may remain preferable for low-risk classification or drafting tasks.
The AWS benchmarking harness, identified in the post as openai-on-aws/benchmarks-openai, gives engineering teams a way to evaluate model choice using their own prompts and acceptance criteria. That matters because the best model for a research agent may not be the best model for a high-volume extraction pipeline, and a benchmark score may not reflect a company’s tolerance for errors.
Teams building AI agents should track turns, tool calls, input growth, latency, and successful-task cost alongside token spend. They should also decide whether their application can reuse stored context, limit search loops, summarize tool results, or route difficult cases to a stronger model. These controls can change the economics independently of the underlying model price.
For product teams, the more useful unit may be cost per approved document, resolved ticket, or completed workflow. That requires a stable rubric and a record of failed outputs, human edits, retries, and escalation rates. AWS’s use of deterministic checks and an LLM judge illustrates one approach, but the company’s own caveats show why evaluation design remains part of the deployment problem.
The immediate signal is whether AWS’s July 30 pricing changes are reflected consistently across Amazon Bedrock regions and inference tiers. Buyers should also watch whether the open-source harness gains independent reproductions using reasoning enabled, matched model configurations, larger samples, and production-like context-management strategies.
Further evaluations should test reliability over repeated runs, tool-use safety, prompt-injection resistance, latency, output truncation, and the cost of human review. Those measures could narrow—or widen—the advantage AWS reports for gpt-5.6-luna and the other gpt-5.6 models.
AWS is making a useful correction to a common procurement habit: token price is visible, while failure and rework costs are distributed across the application. The reported results support measuring model selection at the workflow level, particularly for agents that repeatedly call tools and resend accumulated context.
But the evidence remains AWS-controlled and configuration-specific. The strongest takeaway is not that one OpenAI model is universally cheapest. It is that builders should test the cost of a successful, acceptable outcome on their own workload before allowing a pricing spreadsheet to decide the architecture.