Agent Evaluation Harness
Compare coding agents on the same task, acceptance criteria, evidence standard, and definition of done.
THE PROMPT EVALUATION LAB
Give agents the same task, the same definition of done, and the same evidence standard. Then compare the work you can actually verify.
Same repo, context, permissions, acceptance criteria, and time budget.
Independent checkInspect the diff, run the checks, and record what actually passed.
Fix the inputs and definition of done.
Record trajectory, retries, time, and cost.
Score only what the evidence supports.
START HERE
These are fixtures and rubrics—not fake benchmark results.
Compare coding agents on the same task, acceptance criteria, evidence standard, and definition of done.
Turn real agent tasks into a maintained suite that catches quality, safety, and reliability drift.
Turn a recurring GTM job into a reusable, tested, versioned prompt with examples and quality gates.
Audit a conversion path for audience clarity, intent match, proof, friction, accessibility, and qualified action.
Design a GTM stack around ownership, handoffs, overlap, constraints, and the decisions the system must support.
THE HONEST LIMIT
A run tells you which agent or prompt fits a task class under stated conditions. It does not crown a model forever.