THE INCIDENT · CHAPTER 14
Months later, Mira's team starts a small internal agent. The temptation is to copy a large repository crate for crate. Instead, she writes down the smallest trustworthy loop and the invariants it must preserve.
The question: What should engineers copy from Grok Build, and what should they derive for their own environment?
Start from first principles
Studying a bridge does not mean duplicating every beam. You copy the load paths, safety factors, inspection points, and failure assumptions—then design for your river.
A large tree tempts two bad responses: copy everything, or dismiss it as overengineering. The useful response identifies invariants that survive a smaller implementation.
Grok Build exposes mature boundaries: client/runtime, schema/implementation, policy/sandbox, conversation/workspace, durability/rewind, and protocol stop/semantic verification.
Your harness can be smaller. It should not be ambiguous about those responsibilities.
Build the smallest useful mental model
Start with one auditable loop and one constrained environment. Make transitions observable. Add complexity only when evaluation demonstrates a failure it fixes.
The minimum is not 'call model and parse JSON.' It is a recoverable state machine around model requests and side effects, with an objective verifier and user/automation boundary.
Copy effective tools, normalized calls, deterministic policy, workspace identity, append events, cancellation, output limits, and explicit stop evidence.
Fig 14.1 — A minimal production harness keeps the loop small and boundaries explicit.
Now open the hood
Only after the idea is clear does Mira open the source. She ignores most of the workspace and follows the few boundaries that must exist for this part of the story to work.
1. The next clue — Define client/runtime contract
Mira now needs one small mechanism: UI, CI, or editor sends structured requests and receives typed events.
She follows that responsibility into the repository. Grok reuses shell semantics through pager, headless, and ACP clients. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Screen scraping loses IDs, permission state, and classifications. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source: Pager/headless/ACP boundaries. Verified against Grok Build c68e39f60462f28d9be5e683d9cbe2c57b1a5027.
2. The next clue — Implement a finite state machine
Mira now needs one small mechanism: One prompt contains model/tool rounds with cancellation and limits.
She follows that responsibility into the repository. Shell centralizes setup, sampling, observations, compaction, interjections, stopping. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Naive loops duplicate effects or run forever. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source:xai-grok-shell/.../turn.rs. Verified against Grok Buildc68e39f60462f28d9be5e683d9cbe2c57b1a5027.
3. The next clue — Separate model adapter
Mira now needs one small mechanism: Provider code returns text/tool intent/usage, not side effects.
She follows that responsibility into the repository. Sampling delegates while execution remains in shell/tools/workspace. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Ambiguous retries need idempotency and correlation. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source:xai-grok-samplercall sites. Verified against Grok Buildc68e39f60462f28d9be5e683d9cbe2c57b1a5027.
4. The next clue — Build effective per-session tools
Mira now needs one small mechanism: Expose only schemas backed and allowed for the task.
She follows that responsibility into the repository. ToolDefinition, FinalizedToolset, capabilities, and filters demonstrate separation. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Global registration mistaken for exposure expands risk. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source:xai-grok-tools. Verified against Grok Buildc68e39f60462f28d9be5e683d9cbe2c57b1a5027.
5. The next clue — Put effects behind workspace
Mira now needs one small mechanism: Filesystem, process, repository, and placement share one executor boundary.
She follows that responsibility into the repository. WorkspaceOps local/proxy and file tracking provide this role. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Direct arbitrary host access defeats isolation and recovery. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source:xai-grok-workspace. Verified against Grok Buildc68e39f60462f28d9be5e683d9cbe2c57b1a5027.
6. The next clue — Layer policy and confinement
Mira now needs one small mechanism: Normalize, evaluate deterministic policy, then execute inside OS/container limits.
She follows that responsibility into the repository. Hooks/permissions are separate from sandbox. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Prompt rules are bypassable; sandbox permits harmful allowed actions. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source: Tool-call, permission, sandbox modules. Verified against Grok Build c68e39f60462f28d9be5e683d9cbe2c57b1a5027.
7. The next clue — Persist append-only evidence
Mira now needs one small mechanism: Prompt/model/tool/policy/mutation/verifier events need durable IDs/order.
She follows that responsibility into the repository. Sessions use JSONL and rewind artifacts. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Final prose hides unsafe attempts and failed checks. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source: Session guide and modules. Verified against Grok Build c68e39f60462f28d9be5e683d9cbe2c57b1a5027.
8. The next clue — Make verification first-class
Mira now needs one small mechanism: Task definitions include immutable executable acceptance checks.
She follows that responsibility into the repository. Structural runtime stops show why repository correctness stays external. The important point is not the Rust syntax. It is ownership: this is where the system decides what crosses the boundary.
Then she tests the unhappy path: Model-modifiable tests or hidden-answer leakage invalidate evaluation. If the model, operator, and saved session do not receive the same honest outcome, the mechanism is not yet trustworthy.
Source: Turn stop path plus architectural inference. Verified against Grok Build c68e39f60462f28d9be5e683d9cbe2c57b1a5027.
Mira runs the experiment — minimal patch-producing repair harness
Reading source gives her a hypothesis. A small experiment tells her whether that hypothesis survives contact with a real workspace. Build a small harness whose output is a candidate patch plus evidence.
- Accept base SHA, bounded task, immutable verifier.
- Create ephemeral worktree/container without publication credential.
- Assemble rules and minimum tools.
- Run rounds behind deny-by-default.
- Append every decision/result/path.
- Run immutable verifier independently.
- Emit patch, manifest, logs, failure class.
- Publish only through separate review.
state = new_session(base_sha, task, verifier)
while state.can_continue():
response = model.sample(build_context(state, effective_tools))
for call in response.tool_calls:
state.append(workspace.execute(policy.admit(normalize(call))))
state.checkpoint()
evidence = verifier.run_immutable()
return patch_bundle(state, evidence) # publication is separate
What she learns. Conceptual pseudocode, not Grok API code. Separation, bounded authority, persistence, and independent verification are the point.
That last check matters. Grok Build can expose a mechanism and report an observation; the repository, operating system, CI platform, and reviewer decide whether those observations prove the actual task succeeded.
The whiteboard test
Before Mira explains the chapter to her team, she reduces it to three questions: what owns the decision, what evidence comes back, and what changes when the mechanism fails?
| Review question | Source-backed answer | Operational consequence |
|---|---|---|
| Minimum core? | Client, loop, model, tools, workspace, policy, persistence, verifier. | Keep testable. |
| Optional first? | Memory, plugins, subagents, rich UI. | Add after measured need. |
| Primary output? | Patch plus evidence. | Separate publication. |
| Success metric? | Repeated verified success under cost/safety constraints. | Optimize system. |
This is not a feature scorecard. A mechanism can work exactly as implemented and still be the wrong control for a particular threat. Defaults also change, so recheck the pinned source path before copying configuration into production.
Signals Mira keeps
- Task/base/verifier immutable identities.
- Requests, schemas, calls, policy, results, duration.
- Environment, changed files, processes, network/secret access.
- Verifier, regressions, correction, cost, recovery, unsafe rates.
Together, those signals tell a complete story: the model proposed an action, the harness admitted and routed it, the environment performed something, and a verifier measured the result.
Limits and uncertainty
The public repository is unusually detailed, but it is not the complete deployed product. The snapshot has one visible public commit, so it cannot support a rich historical explanation of why every boundary evolved. Hosted xAI model serving, account systems, and production topology remain outside this study. Where code and guide differ, this series gives the pinned implementation priority and marks documentation-only behavior instead of silently merging versions.
FAQ
Build or use existing?
Usually use existing. Build to learn, research missing contracts, or serve specialized environments.
What should v1 omit?
Persistent memory, recursive orchestration, marketplaces, and rich UI unless tasks require them.
How evaluate?
Repeated tasks with immutable verifiers; measure success, regression, unsafe attempts, correction, cost, recovery.
What should be deterministic?
Policy, environment, verifier, logging, limits, cancellation, publication gates.
Hardest unsolved part?
Knowing semantic completion as context, environment, and goals evolve.
What changed for Mira
The team leaves with a minimal architecture, a hardening path, an evaluation plan, and a list of questions that remain open.
Next: The story ends where real harness engineering begins: with one bounded task and evidence that the system did what it claimed.
Key takeaways
- Copy invariants, not repository size.
- Separate intent, authority, and effects.
- Persist evidence and recover interruption.
- Make verification independent and publication separate.
- Add complexity only after measurement.
References & source notes
- Pinned Grok Build repository — default branch snapshot researched July 16, 2026.
- Grok Build README — first-party overview and source-build entry points.
- Harness Engineering — conceptual foundation.
- Grok source — patterns synthesized here.
Freshness boundary. Grok Build claims in this article are pinned to c68e39f60462f28d9be5e683d9cbe2c57b1a5027. Pi comparison claims, where present, are pinned to 97f9978fa66685f78d2da19ae22e20c46d125f74; Hermes claims are pinned to c9c9bb33fcc6ab479846a1c496a6e9efe2c1c7d4. Recheck paths, symbols, commands, and defaults if those branches advance.