SpaceXAI Launches Grok 4.7 With Improved Coding, Knowledge Work And AI Safety Capabilities

SpaceXAI has launched Grok 4.7, a new frontier AI model designed for software development, long-running professional tasks and general knowledge work, while maintaining the same standard API pricing as Grok 4.6.

The company describes Grok 4.7 as its most capable model to date for coding and knowledge work.

Grok 4.7 uses a larger base model than Grok 4.6 and underwent a longer reinforcement learning training process focused more heavily on difficult tasks that can require hours of work.

SpaceXAI said the training changes improved the model’s ability to verify its own outputs, manage longer contexts and remain effective across multi-step assignments.

The company also trained Grok 4.7 to natively understand the Grok Bot harness, which is intended to improve its performance across conversational tasks and broader knowledge work.

Benchmark results published by SpaceXAI show meaningful gains over Grok 4.6 across several categories.

On CursorBench 4.0, which evaluates longer-running software engineering tasks, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6.

On DeepSWE v1.1, Grok 4.7 achieved a 71.0% high-effort score, up from 65.2% for the previous model.

Electrical engineering represented another area of improvement.

Grok 4.7 scored 64.0% on EEBench compared with 53.0% for Grok 4.6, according to SpaceXAI’s testing.

The company also reported a score of 1,657 on AA Briefcase v1.1, a benchmark designed around multi-hour office tasks, compared with 1,546 for Grok 4.6.

On Terminal-Bench 4.0, which evaluates multi-hour work performed through terminal environments, Grok 4.7 scored 38.0%, almost doubling Grok 4.6’s 20.3%.

SpaceXAI also highlighted professional-domain performance.

Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark versus 15.8% for Grok 4.6.

On HealthBench Professional, which evaluates clinical reasoning, Grok 4.7 achieved 56.7%, compared with 48.5% for Grok 4.6.

Benchmark results should be viewed in the context of individual test methodologies, model settings and company-reported configurations rather than as universal measures of model quality.

SpaceXAI is also positioning Grok 4.7 as a stronger model for document creation, presentations and other professional knowledge work.

The company pointed to improvements on GDPval and AA Briefcase, benchmarks intended to evaluate tasks resembling work performed by professionals including financial analysts, lawyers and nurses.

Another major focus of the release is safety.

SpaceXAI said Grok 4.7 introduces an entirely redesigned safeguard stack intended to improve refusal behavior and resistance to jailbreak techniques.

The company is placing particular emphasis on cybersecurity and biological-risk scenarios, where AI systems can provide significant legitimate value but may also be capable of assisting harmful activity.

Grok 4.7 scored 62.4% on a biosafety benchmark from LatchBio, according to SpaceXAI.

On HackerBench v0.3, which the company uses to evaluate risky or malicious cybersecurity prompts, SpaceXAI said Grok 4.7 allowed 3.3% of risky dual-use prompts while maintaining relatively low refusal rates for legitimate cybersecurity activity.

SpaceXAI has also started giving selected cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defensive research.

The model’s launch continues the broader shift from AI systems designed primarily for individual prompts toward models capable of working for extended periods on larger projects.

Software engineering is particularly important to that transition because coding agents increasingly need to inspect repositories, modify multiple files, execute tools, debug problems and verify results over many steps.

Professional knowledge work presents similar challenges.

A model preparing a financial analysis, legal document or presentation may need to reason across large quantities of information, revise its work and maintain consistency across a long workflow.

SpaceXAI’s training approach for Grok 4.7 is explicitly designed around those longer-duration tasks.

Pricing remains unchanged from Grok 4.6 for the standard model.

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens.

SpaceXAI is also offering a faster version with approximately twice the output speed at twice the standard price.

The company is positioning those economics as part of Grok 4.7’s competitive strategy, particularly for agentic workloads where models can consume substantial numbers of tokens while completing extended tasks.

Grok 4.7 is available through the Grok API as well as Cursor, Grok Build, third-party coding harnesses, model routers and cloud platforms.

By combining a larger model, longer reinforcement learning, stronger coding performance and redesigned safeguards, SpaceXAI is positioning Grok 4.7 as an upgrade aimed less at simple chatbot interactions and more at sustained AI-assisted work across software development and professional tasks.