Define the unit that creates value
A cheap response that produces an unusable patch has no positive price-performance. The useful denominator is an accepted change: the implementation satisfies the issue, passes the intended tests, introduces no known regression and survives review.
Define acceptance before running the model. Include functional requirements, security constraints, code-quality expectations and documentation so evaluators do not move the goalposts after seeing the output.
Build a representative task set
Sample recent work across bug fixes, new behavior, refactoring and test creation. Include small and multi-file tasks, and preserve the starting repository state for every run.
Avoid using only public coding benchmarks. They support repeatability, but a team's frameworks, conventions and legacy constraints determine whether an agent reduces actual delivery time.
Capture the complete cost
Log model input and output, cache charges, tool calls, test compute, retries and elapsed time. Then record reviewer minutes and any follow-up repair work. An agent that spends more tokens but produces a clean first review can be cheaper overall.
Keep model variants separate. OpenAI's Kiro launch includes Sol, Terra and Luna because routing is part of the economics. A useful policy may use a lower-cost model for routine work and escalate difficult tasks after a defined failure signal.
Report results without hiding failure
Publish acceptance rate, median and tail cost per accepted task, time to first valid patch and defect categories. Include tasks the agent abandoned or solved only after human intervention.
Re-run the evaluation after major model, harness or repository changes. Coding-agent performance is a property of the complete system and its context, not a permanent score attached to one model name.
Sources & further reading
Social-media activity is treated as a signal of attention, not proof. Product claims are attributed to the linked publisher or announcement.