Early today, SpaceXAI launched its strongest model, Grok4.7. The landing page starts with a highly "aggressive" positioning: the strongest model for programming and knowledge work, twice as fast as comparable models, and half the price. This upgrade goes beyond benchmark scores—Grok4.7 uses a larger base model, enhances long-term tasks, self-checking, and long context management, and focuses on testing in specialized areas such as documentation, presentations, law, healthcare, and engineering. The goal is to reduce missteps and deliver directly usable results during tasks that last for hours.
Elon Musk tweeted that Grok4.7 achieves a "very competitive balance" between intelligence level, speed, and cost.

Long Tasks Become the Main Battlefield
Grok4.7 uses a larger base model than Grok4.6, extending the reinforcement learning period during training and making task combinations even more challenging, with many samples requiring several hours to complete. The official stated that the new model has improved in checking its own work and managing long context, directly addressing the most challenging issues for programming agents: complex software tasks include reading repositories, breaking down requirements, modifying multiple files, running tests, and repeated debugging. Whether the model can maintain its goals and detect errors during long execution chains is often more important than writing clean code at once.

In the CursorBench4.0 test for long-term coding, Grok4.7 scored 46.3%, higher than the previous generation's 40.4%; DeepSWE v1.1 high reasoning intensity score was 71.0%; Terminal-Bench4.0 increased from 20.3% to 38.0%. The model was also trained to natively understand the Grok Bot operating framework, and the official said this improves performance in dialogue and general knowledge work, meaning the upgrade focus has shifted from single-turn answers to continuous collaboration between the model and tools, and the execution environment.
Moving from Writing Code to Full Knowledge Work
Programming remains the most prominent feature of Grok4.7, but the evaluation scope is significantly broader. On AA Briefcase v1.1, it scored 1657 points (1546 for the previous version), GDPval Elo rose from 1605 to 1695, close to Fable5.1's 1735, and higher than GPT-6Astra's 1542. Professional fields also saw improvements: EEBench increased from 53.0% to 64.0%, Harvey Legal Agent Benchmark from 15.8% to 19.6%, and HealthBench Professional from 48.5% to 56.7%. This reveals a clear trend—the competition among cutting-edge models is entering the "completing the entire task" stage, requiring understanding of tasks, calling tools, producing files, and self-reviewing.
In terms of security, Grok4.7 has implemented a new protection system, and the official claims it is the strongest in rejecting answers and resisting jailbreaks; the LatchBio biosafety benchmark reached 62.4%, and HackerBench v0.3 only allowed 3.3% high-risk dual-use prompts, rarely blocking normal security research. SpaceXAI has also opened invitation-based red team capabilities to a few cybersecurity partners.

