Status: an industry signal that must be validated against each team’s workflow.
Once coding agents can modify multiple files, run tests, and iterate, the evaluation unit changes:
- Can the agent respect a clear task boundary?
- Are failures discovered quickly?
- Is human review cheaper than the execution time saved?
- Does the change retain verifiable evidence?
The useful metric is no longer “how much code was generated,” but “how many tasks were reliably completed under verifiable conditions.”