SpaceXAI 🇺🇸 · Grok

Grok 4.7

Closed hosted frontier model released for API and coding-agent use. SpaceXAI reports gains on CursorBench and Terminal-Bench versus Grok 4.6, longer multi-hour task completion, 500k context and new safeguards. Controlled cyber tests and internal refusal metrics are mixed; no real-world escape is established.

DOOM SCORE85.9out of 100model risk profile, not the overall index
0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1

Why this model scores 85.9

Capability 93 and autonomy 94 reflect provider coding-agent and long-task evaluations, with no independent real-world validation. Deployment 94 reflects API and IDE/agent-router availability. Misuse 82 reflects meaningful cyber competence in controlled tests, while control difficulty 64 accounts for residual risk and mixed safeguard findings. These are evidence-calibrated dimensions, not a replacement total score.

Capability93
Autonomy94
Deployment94
Misuse potential82
Control difficulty64
MODEL-ATTRIBUTED EVIDENCE

News tied to Grok 4.7

The model score of 85.9 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION+0.11

Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.

Autonomy TOWARD

SpaceXAI releases Grok 4.7 for longer agentic coding tasks

SpaceXAI released Grok 4.7 through its API, Grok Build, Cursor and model routers, reporting higher coding-agent benchmark results than Grok 4.6 and stronger multi-hour task performance. Its model card documents expanded safeguards and controlled cyber evaluations, but provider tests do not establish real-world autonomous compromise.

Full item contribution
+0.11
Grok 4.7 equal share
+0.11
Read assessment →
AUDIT TRAIL

Model score history

  1. R1
    Doom Score 85.9

    New exact-version model card and released API access establish a distinct Grok 4.7 profile.

    25 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Grok 4.7.
  1. Grok 4.7 by SpaceXAI has a DoomBench model risk score of 85.9 out of 100, based on five transparent version-specific dimensions rather than the overall index.

  2. Grok 4.7's highest current DoomBench dimension is autonomy at 94.0 out of 100; the profile publishes every component score and its editorial rationale.

  3. DoomBench links 1 source-backed evidence item to Grok 4.7, while keeping the model's risk profile separate from each item's contribution to the live Doom Index.

    https://www.doombench.com/models/spacexai-grok-4-7