What happened

SpaceXAI has released Grok 4.6, a new AI model that reportedly matches the intelligence level of GPT-5.6 Sol while costing roughly half as much to run. The model is now live inside the Grok Build coding agent, in Cursor, in the Grok Bot chat assistant, and through the standard API. Cursor and Grok Build are offering a doubled usage limit for the first week, likely to drive adoption fast.

Pricing hasn't moved since the previous release: $2 per million input tokens and $6 per million output tokens, the same rate Grok 4.5 launched with on July 8. According to independent analytics platform Artificial Analysis, that's more than 60% cheaper than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) for comparable output.

### The headline number

Grok 4.6 scored 61 points on Artificial Analysis' Intelligence Index, a composite score built from nine benchmarks. That ties GPT-5.6 Sol running in its maximum reasoning mode, and lands just one point behind Claude Fable 5 Max. It's also a five-point jump over Grok 4.5 in just over a month — a fast turnaround for a full generation of gains. Analysts described the release as SpaceXAI returning to the frontier of AI intelligence, with Anthropic now the only lab still ahead.

Why it matters

The benchmark table shows a clear pattern: Grok 4.6 is strongest at office and knowledge work, not raw coding. On GDPVal-AA v2, a benchmark built around real knowledge-work tasks, it scored 1753 versus 1741 for Claude Fable 5 and 1728 for GPT-5.6 Sol. On AA-Briefcase, an agentic office-task benchmark, it posted 1577 against 1574 and 1502. On the legal benchmark Harvey LAB, Grok 4.6 leads outright with 15.8%, compared to 11.3% for Fable 5 and just 2.5% for GPT-5.6 Sol.

Coding is a mixed picture. On Terminal-Bench v3.0, Grok 4.6 scores 26% against 34.6% for Sol and 34.1% for Fable 5 — a real gap. But on the older Terminal-Bench v2.1, which is what Artificial Analysis actually uses in its index, Grok 4.6 hits 88.4%, second only to Claude Opus 5. In other words, it handles routine terminal work as well as the leaders but struggles more on the newer, harder version of the test. It also trails on DeepSWE (73% for Sol, 70% for Fable 5) and loses to Fable 5 on FrontierCode and APEX-SWE.

### How the gains were achieved

What makes this release notable isn't scale — it's method. Elon Musk says Grok 4.6 runs on the same 1.5-trillion-parameter base as Grok 4.5; the model didn't get bigger. Instead, SpaceXAI ran a longer training cycle on the same base, using curated synthetic data focused on reasoning and engineering tasks, an improved optimizer, and a rebuilt supervised fine-tuning and reinforcement learning pipeline. Grok 4.5 itself regenerated the SFT training trajectories, with problematic examples filtered out by automated model checks — effectively, the previous model trained its successor.

Inside Cursor, which SpaceXAI now owns, engineers report that Grok 4.6 builds out an app's structure and visual language correctly from a single pass based on a written idea, and shows noticeably more self-checking on long task chains — testing its own work before moving forward. The reinforcement learning tasks behind this ranged from optimizing compute kernels to web development and CAD work.

MyKreaTool AI chat — try ChatGPT, Claude and Gemini in one place. Free on MyKreaTool.Open the tool →

How to use it today

Grok 4.6 is accessible right now through four channels: the API directly, the Grok Build agent, the Grok Bot assistant, and Cursor. For teams already using Cursor for development, the doubled usage limit this week is a low-risk window to benchmark Grok 4.6 against whatever model you're currently paying for.

The economics are the real selling point for production use. Artificial Analysis calculates that a typical Intelligence Index task costs $0.84 to run on Grok 4.6 — the same as Kimi K3, but with slightly higher measured intelligence, putting it right on the Pareto frontier of cost versus capability. The model also keeps its token efficiency edge: on the AA-Briefcase agentic benchmark, Grok 4.6 closes a long task in an average of 53 turns and about 500 million input tokens, compared to 103 turns and 2 billion tokens for Claude Opus 5 on the same task. The one price increase is on cached tokens, which rose from $0.30 to $0.50 per million.

If you're not ready to commit to an API budget yet, it's worth prototyping your prompts and workflows on free tools first — a site like [mykreatool.com](https://mykreatool.com) lets you test AI-assisted content and image workflows at no cost before deciding which paid model to scale into production.

Who benefits

Given the benchmark profile, Grok 4.6 is best suited to teams doing document-heavy, knowledge-work-style tasks: legal research and drafting, business analysis, reporting, and general office automation, where it now leads or ties the field. Law firms and legal-ops teams in particular should notice the Harvey LAB result — a 15.8% score against GPT-5.6 Sol's 2.5% is a meaningful gap for anything touching contract review or case research.

Agencies and dev shops using Cursor also stand to benefit directly, since SpaceXAI now owns the tool and is clearly optimizing Grok releases around it. Cost-sensitive startups running high-volume agentic workflows — customer support bots, research agents, data-processing pipelines — get the most out of the token efficiency numbers, since fewer turns and fewer tokens per task translate directly into lower monthly bills.

Risks

Teams relying heavily on terminal automation or complex software engineering agents should test carefully before switching. The gap on Terminal-Bench v3.0 and on coding benchmarks like DeepSWE, FrontierCode, and APEX-SWE suggests Grok 4.6 isn't the strongest choice yet for demanding, novel coding work, even though it holds up well on routine terminal tasks.

It's also worth remembering that SpaceXAI's own training method — having Grok 4.5 generate and filter its successor's training data — is new territory. Self-generated training data filtered by the same family of models can compound blind spots that neither version catches, so any business-critical workflow should keep a human review step in place rather than trusting benchmark scores alone.

Finally, release cadence is aggressive: Grok 4.5 shipped July 8, Grok 4.6 followed August 12, and Musk has signaled a 2.1-trillion-parameter Grok 4.7 could arrive within three to four weeks, putting it around the end of summer. Businesses building automation on top of Grok should expect frequent model changes and plan for regression testing each time.

Conclusion

Grok 4.6 delivers GPT-5.6-level intelligence at roughly half the price, without increasing model size — a genuine efficiency win achieved through better training rather than brute-force scaling. It's now the strongest option for document-heavy, legal, and office-automation workloads, while still trailing on frontier coding tasks. For businesses evaluating AI vendors on cost per task, Grok 4.6 sits squarely on the best available price-to-intelligence ratio right now, but it's worth testing your specific workflow before migrating anything mission-critical.