Grok 4.7 Description
Grok 4.7 is SpaceXAI’s advanced AI model for software development, professional knowledge work, and multi-step agentic tasks. It is built on a larger base model than Grok 4.6 and uses a longer reinforcement learning training run weighted toward difficult problems that require extended execution time. The model is better at checking its own work, maintaining longer context, and completing complex workflows that span many steps. Native support for the Grok Bot harness improves its conversational abilities and makes it more capable across general knowledge and interactive work. Grok 4.7 is positioned for use in coding, terminal automation, legal workflows, document and presentation creation, electrical engineering, healthcare reasoning, and other professional tasks. SpaceXAI also reports improvements across software engineering benchmarks including CursorBench and DeepSWE as well as professional-work evaluations such as AA Briefcase. The model includes a redesigned safety system intended to improve refusal behavior, jailbreak resistance, and handling of potentially dangerous cybersecurity and biological requests. Developers can access Grok 4.7 through Grok Build, Cursor, the Grok API, third-party coding tools, cloud providers, and model-routing platforms. Grok 4.7 is priced starting at $2 per million input tokens and $6 per million output tokens, with a faster variant available at a higher price.
Pricing
Company Details
Product Details
Grok 4.7 Features and Options
Grok 4.7 User Reviews
Write a Review-
Likelihood to Recommend to Others1 2 3 4 5 6 7 8 9 10
Frontier level AI Date: Sep 21 2026
Summary: Overall, Grok 4.7 has become one of the more compelling models for serious coding and long-running agent work. It feels noticeably better at staying organized, checking its work, and finishing complex tasks without needing as much hand-holding.
Positive: The biggest improvement is how much better it is at sticking with longer tasks. It feels more deliberate than earlier Grok versions, especially when I am working through debugging, repo-level changes, or something that takes multiple rounds of planning and verification. The coding performance is strong too. xAI reports 46.3% on CursorBench 4.0 versus 40.4% for Grok 4.6, and Terminal-Bench 4.0 jumps from 20.3% to 38.0%. That is the kind of improvement that actually matters for agentic coding rather than just short code-generation prompts. I also like the pricing. At $2 per million input tokens and $6 per million output tokens, it is much easier to justify for daily agent workflows than some of the more expensive frontier models.
Negative: The downside is that the higher reasoning modes can use a lot of tokens, so I would not automatically run xhigh on every task. For quick edits or simple coding questions, a faster model can still make more sense.
Read More...
- Previous
- You're on page 1
- Next