STACKSWAP / FREE GTM LIBRARY

AI Agent Evaluation Harness Prompts

Run fair, repeatable tests for coding agents with the same task, acceptance criteria, evidence standard, and definition of done.

Browse all 125 prompts →

START HERE

Use the prompt that matches the decision in front of you.

THE LIBRARY

Browse free GTM prompts.

125 operator-built prompts for the GTM work in front of you.

3 workflowsCopy-ready · platform-flexible · operator-built
AI-native operations

Agent Evaluation Harness

Compare coding agents on the same task, evidence standard, and definition of done—not vibes.

YOU'LL LEAVE WITHReusable task spec, acceptance rubric, evidence log, verification record, retry/cost ledger, and comparative scorecard.
Reusable task spec, acceptance rubric, evidence log, verification record, retry/cost ledger, and comparative scorecard.Open →
AI-native operations

Agent Regression Suite Builder

Turn real agent tasks into a maintained regression suite that catches quality, safety, and reliability drift.

YOU'LL LEAVE WITHTask dataset, expected outcomes, evaluators, release gate, and regression report.
Task dataset, expected outcomes, evaluators, release gate, and regression report.Open →
AI-native operations

Agent Judge / Rubric Calibration

Make agent evaluation trustworthy by calibrating human, code, and model judges against clear examples.

YOU'LL LEAVE WITHScoring rubric, anchor examples, judge comparison, disagreement log, and calibrated evaluator.
Scoring rubric, anchor examples, judge comparison, disagreement log, and calibrated evaluator.Open →