SpaceXAI has launched Grok 4.6, its latest frontier AI model, with a particular focus on long-running agents, coding, knowledge work, and more ambitious interactive and visual projects.
Grok 4.6 builds on Grok 4.5 by improving the model’s ability to remain engaged across complex, multi-step tasks. These can include researching unfamiliar topics, analyzing information, working across software codebases, and transforming ideas into functional applications or other work products.
SpaceXAI said Grok 4.6 reaches frontier-level performance across several agentic coding and knowledge-work benchmarks. On the Artificial Analysis Intelligence Index, Grok 4.6 scored 61, up from 56 for Grok 4.5.
The model also scored 1,753 on GDPVal-AA v2 compared with 1,526 for Grok 4.5; 69.9% on CursorBench v3.2 compared with 66.7%; 65.9% on DeepSWE v1.1 compared with 54%; and 61.3% on FrontierCode v1.1 Extended compared with 56.6%.
Other results included 57.5% on APEX-Agents, up from 47.1% for Grok 4.5, and 26% on Terminal-Bench v3.0 compared with 15.7%. Grok 4.6 also scored 1,577 on AA-Briefcase versus 1,313 for its predecessor.
SpaceXAI used a longer supplemental training run for Grok 4.6 than for Grok 4.5. Training incorporated curated model-generated reasoning and technical data, engineering data, an updated optimizer, and an improved training recipe.
The company then used Grok 4.5 to regenerate supervised fine-tuning trajectories covering different reasoning efforts and agent environments across STEM, software engineering, and knowledge work. Problematic traces were filtered using model-based checks.
Grok 4.6 was also trained using reinforcement learning across agentic tasks spanning general coding, knowledge work, kernel optimization, web development, computer-aided design, and other specialized environments.
One of the model’s main improvements is its ability to take a relatively broad product concept and turn it into a working initial application. SpaceXAI said Grok 4.6 can research unfamiliar areas, organize an application, implement key interactions, and continue improving its work across multiple rounds of feedback.
The company has also observed more self-testing and verification during longer tasks, with Grok 4.6 increasingly checking its own work before continuing. Visual and interactive projects are another area of improvement, with the model designed to generate stronger initial versions of applications before users begin iterating on them.
SpaceXAI said it also expanded its safety evaluations to reflect Grok 4.6’s increased capabilities, conducting its widest suite of pre-deployment capability testing and safeguard calibration to date, along with post-deployment and third-party testing.
Grok 4.6 is available through Grok Build, Cursor, the SpaceXAI API, OpenRouter, Vercel, Cloudflare, and other partners. API pricing starts at $2 per million input tokens and $6 per million output tokens, while a faster version is available at twice that price.
KEY QUOTES:
“Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.”
SpaceXAI

