Grok 4.7
Closed hosted frontier model released for API and coding-agent use. SpaceXAI reports gains on CursorBench and Terminal-Bench versus Grok 4.6, longer multi-hour task completion, 500k context and new safeguards. Controlled cyber tests and internal refusal metrics are mixed; no real-world escape is established.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 85.9
Capability 93 and autonomy 94 reflect provider coding-agent and long-task evaluations, with no independent real-world validation. Deployment 94 reflects API and IDE/agent-router availability. Misuse 82 reflects meaningful cyber competence in controlled tests, while control difficulty 64 accounts for residual risk and mixed safeguard findings. These are evidence-calibrated dimensions, not a replacement total score.
News tied to Grok 4.7
The model score of 85.9 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
SpaceXAI releases Grok 4.7 for longer agentic coding tasks
SpaceXAI released Grok 4.7 through its API, Grok Build, Cursor and model routers, reporting higher coding-agent benchmark results than Grok 4.6 and stronger multi-hour task performance. Its model card documents expanded safeguards and controlled cyber evaluations, but provider tests do not establish real-world autonomous compromise.
- Full item contribution
- +0.11
- Grok 4.7 equal share
- +0.11
Model score history
-
R1
Doom Score 85.9
New exact-version model card and released API access establish a distinct Grok 4.7 profile.
25 Sept 2026