Practice Guide · Measuring a RACE rollout

If your team measures with SPACE

A measurement playbook. For each of the five SPACE dimensions: what to measure under RACE Programming, where to pull it from, what to target, and what to watch. One dimension, Activity, is where naive instrumentation reads a rollout as "no change." Fix that one first.

Follow SPACE's own rule: measure across at least three dimensions, and never read Activity in isolation. Below is how to instrument each one off RACE Programming's artifacts.

The five dimensions, instrumented

Satisfaction and well-being

  • Measure: a short periodic pulse on delegation confidence (do engineers trust the Silicon Software Engineer's output?) and Inner-Cycle friction.
  • Source: a two-question survey per Stint, plus the reopen rate on validated work.
  • Target: trust trending up, friction down.
  • Watch: engineers who still hand-write instead of delegating and validating. That is a training signal, not a tooling one.

Performance

  • Measure: Executable User Stories accepted by the Team Principal per Stint, and the one triangle outcome you took (scope, cost, or time) at held quality.
  • Source: UAT sign-offs plus delivery data.
  • Target: state one triangle number per engagement (for example, roughly three times the scope at the same cost).
  • Watch: performance is accepted value, not shipped volume. A shipped-but-unaccepted EUS does not count.

Activity (the one to re-read)

  • Measure: Executable User Stories authored, prototypes accepted, validations performed, and Pit Stops made, per week.
  • Do not measure: commits, lines, or story points. The agent writes the code, so these count the wrong actor.
  • Watch: raw activity up while EUS-accepted stays flat is motion without delivery. This is the trap SPACE warns about, made concrete by AI.

Communication and collaboration

  • Measure: spec clarity (the rate at which an EUS is returned for clarification) and handoff cleanliness (how often an EUS bounces between Pit Wall and Pit Crew).
  • Source: board transitions on the Executable Product Backlog.
  • Target: low return and bounce rates.
  • Watch: rising clarification requests mean the EUS is not yet minimum-sufficient. Fix the spec, not the standup.

Efficiency and flow

  • Measure: interruptions to the twice-daily Inner Cycle, and the rework rate (EUS returned to Build by the Definition of Done).
  • Source: Inner Cycle logs and board transitions.
  • Target: few interruptions, low rework.
  • Watch: mid-Stint re-scoping. That is a Pit Wall boundary breach, and the first thing to correct.
The reframe

Read three dimensions, and never Activity alone

A team that measures a RACE Programming rollout by Activity alone will conclude nothing changed, because AI inflates activity while value moves to specification and validation. Read Performance, Communication, and Efficiency together and the shift is obvious. That is not a workaround for SPACE; it is SPACE applied correctly to a world where the agent writes the code.

SPACE says productivity is more than one number. RACE Programming is what that looks like when AI writes the code.

Read the framework → · All practice guides →

FAQ

Frequently asked questions

How does the SPACE framework apply to RACE Programming?
SPACE argues productivity cannot be captured by one metric and should be measured across at least three of its five dimensions (Satisfaction, Performance, Activity, Communication, Efficiency). RACE Programming is already multi-dimensional, so each dimension has a natural measurement. The catch is Activity: when an AI agent writes the code, count Executable User Stories authored, prototypes accepted, and validations performed, not commits or lines.
Why is the Activity dimension a trap under RACE Programming?
Because AI inflates raw activity while real value moves to specification and validation. Counting commits or lines when the Silicon Software Engineer writes the code measures the wrong thing. In RACE Programming, Activity is the count of Executable User Stories authored, prototypes accepted, validations performed, and Pit Stops made per week.
How do you measure Performance in a RACE rollout?
Performance is accepted value shipped at a Pit Stop and validated by the Team Principal through acceptance testing, reported per Stint. Quantify it with the project triangle: holding quality, the gain is taken as roughly three times the scope, or about one third the cost, or two to five times faster. A shipped but unaccepted story is not performance.
What does Efficiency and flow map to in RACE Programming?
Measure interruptions to the twice-daily Inner Cycle and the rework rate, the share of Executable User Stories bounced back to Build by the Definition of Done gates. Explicit role boundaries, where the Pit Wall never reimplements and the Pit Crew never re-scopes, are the flow-protection design; mid-Stint re-scoping is the signal that a boundary was breached.