
xAI has released Grok 4.6, a new model that a report from The Decoder says matches OpenAI’s GPT-5.6 Sol on the Artificial Analysis Intelligence Index while costing substantially less to run. The model’s reported price is $2 per million input tokens and $6 per million output tokens, putting it more than 60% below the listed prices for competing frontier models.
The release matters less as a simple leaderboard update than as a positioning move. Grok 4.6 is being made available through an API and developer products including Cursor and Grok Build, while the model’s strongest reported result is on a benchmark designed around multi-step computer work. If the figures hold in practical deployments, xAI is competing not only on model capability but also on the cost of running AI agents through long workflows.
According to The Decoder’s account of the Artificial Analysis Intelligence Index, Grok 4.6 scores 61 points. That ties GPT-5.6 Sol and places the model behind Anthropic’s Claude Opus 5, at 63, and Claude Fable 5, at 62. The reported score is five points higher than Grok 4.5.
Those rankings should be treated as benchmark evidence rather than a universal verdict on model quality. Composite indexes combine multiple evaluations, and performance can vary considerably by task, prompt design, tool access, context length, and reliability requirements. The available source does not provide a detailed breakdown of the index’s tests or independent production results.
Still, the reported result gives xAI a stronger position in a market where buyers increasingly compare models across coding, research, reasoning, and autonomous task execution. A tie with OpenAI’s listed leading model could make Grok 4.6 relevant to teams that previously treated xAI primarily as an alternative conversational model.
Grok 4.6 reportedly performs particularly well on GDPval-AA v2, a benchmark intended to measure real-world knowledge work carried out on a computer. The model ranks second with an Elo score of 1,753, behind Claude Opus 5.
The more notable figure concerns the number of steps required to complete complex workflows. The Decoder reports that Grok 4.6 completes these tasks in about 53 steps, while Claude Opus 5 requires roughly 103. Fewer steps can indicate more efficient planning or execution, although step counts alone do not establish that a model produces better final work. A shorter sequence may also involve different tool calls, assumptions, or evaluation trade-offs.
For developers building AI agents, the distinction is important. Each additional action can increase latency, token consumption, tool-use risk, and the number of points where an agent can fail. A model that reaches a satisfactory result with fewer actions could reduce operating costs and simplify monitoring. Conversely, teams still need to test whether a model’s shorter workflows remain accurate, recover well from errors, and respect permissions in production systems.
The reported pricing for Grok 4.6 is $2 per million tokens for input and $6 per million tokens for output. The Decoder compares that with $5 and $25 for Claude Opus 5, and $5 and $30 for GPT-5.6 Sol. On those listed rates, Grok 4.6 is more than 60% cheaper than both alternatives, with the largest difference on output tokens.
That gap could matter most for applications that generate substantial output or repeatedly call a model during an agentic workflow. It may make it more practical to use a frontier-level model for customer support investigations, software tasks, document processing, or internal research rather than reserving it for a small number of high-value requests.
Price alone does not determine total cost. Buyers must also account for latency, rate limits, context-window behavior, tool integration, reliability, moderation, data handling, and the engineering work required to switch providers. The source does not include those operational comparisons, so the pricing advantage should be read as a list-price advantage rather than a complete cost-of-ownership calculation.
The report says Grok 4.6 is available through the xAI API, Cursor, Grok Build, and partners including OpenRouter, Vercel, and Cloudflare. xAI is also offering double usage quotas in Grok Build and Cursor during the first week, according to The Decoder. That promotion may encourage early testing, but it does not yet show sustained adoption or long-term usage economics.
The immediate opportunity for builders is to test Grok 4.6 against existing models on their own workloads, especially tasks involving repeated tool calls. Teams should measure successful task completion, not just benchmark scores, alongside average steps, latency, token use, retries, and human intervention.
The model’s availability through multiple developer channels could lower the barrier to comparison. A team already using Cursor, Vercel, Cloudflare, or OpenRouter may be able to evaluate Grok 4.6 without building an entirely new application path. However, each access route can expose different limits, pricing terms, feature availability, or data policies, so “available” does not necessarily mean operationally interchangeable.
For enterprise buyers, the key question is whether the reported agentic efficiency survives controlled tests involving private data, permissions, failure recovery, and audit requirements. A model that is inexpensive and capable on a benchmark can still be unsuitable for sensitive workflows if it is inconsistent, difficult to monitor, or weak at following organizational constraints.
The release also increases competitive pressure on OpenAI and Anthropic. Frontier-model competition is shifting toward a combination of capability, inference cost, distribution, and agent reliability. xAI’s reported strategy with Grok 4.6 is to use a lower price and broad availability to turn benchmark parity into developer experimentation.
The most important follow-up will be independent testing of Grok 4.6 across coding, browser use, business process automation, and long-running agent tasks. Comparisons should use matched prompts, identical tools, and transparent accounting for retries and human assistance.
Buyers should also watch for production reports from developers using the xAI API, Cursor, and Grok Build. Useful signals will include sustained task success, actual token costs, latency under load, rate-limit behavior, and whether the model’s reported step advantage translates into fewer failures.
Finally, xAI’s commercial response will be worth tracking. The first-week quota promotion may be temporary, while competitor pricing and model releases could quickly change the economics. The durability of Grok 4.6’s position will depend on service quality and real-world adoption, not only on its initial index score.
Grok 4.6’s reported results make cost-efficient agent execution the most consequential part of this release. A model that combines near-frontier benchmark performance with materially lower token prices could alter which workflows are economically viable, particularly when applications make many model calls.
But the evidence remains limited to a specialist report citing benchmark and pricing information. Builders should treat the release as a strong reason to run comparative evaluations, not as proof that Grok 4.6 is universally better or cheaper in production. The next competitive advantage will belong to providers that can demonstrate reliable agent behavior at scale.
xAI's Grok 4.6 ties OpenAI's reported top model on a leading index and targets agentic workloads with prices more than 60% lower.