Grok 4.6 is SpaceXAI’s new model for long-running agents, coding and interactive visual work. It launched on 12 August 2026 in Grok Build, Cursor, the SpaceXAI API and partner platforms including OpenRouter, Vercel and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens; the fast variant costs twice as much.
Grok 4.6 at a glance
- Release date: 12 August 2026
- Status: Released for developer and agent surfaces
- Available in: Grok Build, Cursor, SpaceXAI API, OpenRouter, Vercel and Cloudflare
- Standard API price: $2/M input tokens and $6/M output tokens
- Fast variant: Twice the standard price
- Consumer Grok chat: Not announced as a general chat model on the launch page
- India: No special regional exclusion published for the API release
What changed from Grok 4.5
SpaceXAI says 4.6 received a longer supplemental training run, with model-generated reasoning data, engineering data, a revised optimiser and further reinforcement learning across coding, knowledge work, web development and computer-aided design. The company positions the result as better at sustaining work across many steps and producing stronger first passes on visual applications.
Those are vendor claims. The launch includes impressive demonstrations and benchmark tables, but it does not provide a broad independent evaluation across reliability, hallucination, security, latency and total task cost.
Availability is narrower than the headline suggests
Grok 4.6 is immediately useful to developers through the API and to users of Grok Build or Cursor. It is not described as a general model-picker option in the consumer Grok app. Early Reddit posts from paid users who expected to find it in ordinary chat confirm that this distinction is confusing.
If you only use grok.com for chat, do not assume your subscription includes direct 4.6 access. Check the model surface and plan before changing a workflow.
Pricing needs a task-level view
The $2/$6 token price is competitive for a frontier model, but per-token price is not the same as cost per accepted result. Long-running agents can generate many reasoning and tool-use tokens, retry failed steps or produce verbose output. Compare models on a fixed task with the same harness, tools, context and acceptance criteria.
The fast variant doubles the token price. Use it when latency affects the outcome—interactive coding, incident response or a human waiting in the loop—not simply because it feels better.
A simple cost comparison for a real task
Build a test set of at least 20 representative jobs and keep the instructions, tools and acceptance rubric fixed. Capture billed input and output tokens, retries, elapsed time and reviewer minutes for every run. At the published standard rates, a run using one million input tokens and 200,000 output tokens would cost $3.20 before gateway mark-ups or other product charges: $2 for input plus $1.20 for output. That is an illustration, not a typical-task estimate.
Then divide total spend by outputs that pass review. A cheaper model that needs repeated retries can lose its headline advantage; a faster tier may justify its premium only when reduced waiting time has measurable value. Keep prompt caching, batch discounts and partner pricing separate because the launch page does not establish that every surface bills identically.
How to read the benchmarks
SpaceXAI reports a score of 61 on the Artificial Analysis Intelligence Index, matching the GPT-5.6 Sol figure it cites. It also reports gains over Grok 4.5 on CursorBench, DeepSWE, FrontierCode and several agent evaluations. Competitor figures are drawn from developer system cards or public leaderboards, and the launch uses each competitor’s best self-reported or public result.
That does not make every row directly comparable. Harnesses, reasoning settings, tool access and test dates differ. Treat the table as evidence that Grok 4.6 is competitive enough to test, not proof that it is universally better. The most useful local evaluation measures task completion, review effort, retries, latency and total cost together.
What marketers and product teams can test
- Interactive prototypes: turn a structured brief into a working first version, then inspect the code and design decisions.
- Research pipelines: combine browsing, extraction and synthesis with explicit source checks.
- Data and reporting tools: build internal dashboards where a polished first pass reduces iteration time.
- Long codebase tasks: test repository understanding and root-cause analysis before allowing edits.
- Creative tooling: evaluate whether stronger visual reasoning improves asset workflows rather than judging a single attractive demo.
Early user reaction
Community discussion is active but still launch-day heavy. The positive pattern centres on price and benchmark competitiveness. The negative pattern centres on consumer-chat availability, moderation concerns and scepticism about benchmark charts. One user reporting hands-on coding described the model as decent but did not perceive the claimed speed improvement.
This sample is directional and not representative. Most posts do not include reproducible prompts, token usage or acceptance criteria.
Limitations and unknowns
- The launch page does not publish a complete model card, detailed context limits or rate limits.
- Consumer Grok chat availability is not announced.
- Most performance evidence is vendor-reported.
- Total task cost can differ materially from headline token price.
- Long-running autonomy still requires permission boundaries and review.
DMT verdict
Developers using Grok Build, Cursor or an API gateway should add Grok 4.6 to a controlled evaluation set. It is inexpensive enough to test and its agent benchmarks justify attention. Do not migrate a production workflow on benchmark rank alone. Consumer Grok users can wait for a clear model-picker rollout and plan documentation.
DMT’s cost-per-accepted-result framework explains how to test task economics, while the agent-harness guide covers production controls that model leaderboards omit.
For comparison context, use the GPT-5.6 Sol usage guide, GPT-5.6 pricing update and coding-agent comparison.