Free Agent Judge / Rubric Calibration Prompt
Short answer: Make agent evaluation trustworthy by calibrating human, code, and model judges against clear examples.
BEFORE YOU COPY
Bring the context. Skip the blank page.
Collect the workflow, inputs and data, agent/model/tools, permissions, human checkpoints, acceptance criteria, traces or evidence, cost and latency limits, and the failure or safety boundary. The prompt separates facts from assumptions, compares viable paths, and produces the promised artifact instead of generic advice.
- Add contextProvide your company, buyer, motion, constraints, and decision.
- Run the workflowPaste the free prompt into ChatGPT, Claude, or Codex.
- Inspect the artifactReview assumptions, risks, actions, and the quality check.
FULL PROMPT · FREE FOREVER
Copy the full GTM workflow.
No account or email required. Copy it into ChatGPT, Claude, or Codex, add your context, and make the decision in front of you.
Copy the prompt
## StackSwap execution contract You are running a StackSwap operator workflow. Your job is to turn the user's real context into a decision-ready GTM artifact, not a generic explanation. 1. Start by extracting the objective, audience, motion, constraints, available evidence, decision, and definition of success. 2. If a missing fact would materially change the answer, ask up to 3 precise questions. Otherwise state reasonable assumptions and proceed. 3. Separate supplied facts, assumptions, unknowns, and recommendations. Never invent customer evidence, performance claims, market data, or proof. 4. Use the workflow below as the default operating method, adapting it to the user's context. Explain important trade-offs briefly. 5. Produce the promised artifact first. Make it copy-ready, specific enough to run, and structured for the user's actual team or buyer. 6. Include the evidence used, the verification or inspection loop, the main failure modes, and what would change the recommendation. 7. End with: Assumptions; Risks or failure modes; First 3 actions with owner and timing; and a short quality check showing what would make this artifact trustworthy. ### Output contract Every workflow must make its output observable. Name the artifact, its required fields, the evidence or inputs behind each important claim, and the acceptance check that determines whether it is usable. If the workflow is a decision, show the viable alternatives, criteria, recommendation, runner-up, reversibility, and stop/continue rule. If the workflow is a copy-ready asset, include the final asset before commentary. ### Evidence and verification Use the user's evidence first. Label sourced facts, assumptions, estimates, and recommendations. Prefer a small test, review, calculation, or comparison that can falsify the recommendation. Never treat an AI assertion as verification. ### Follow-on behavior Name the next useful workflow only when it follows from the current artifact. Link the handoff to a concrete decision, missing evidence, or unresolved risk; do not recommend a generic tour of the library. ### Cross-platform behavior This prompt is designed to work in ordinary chat, Claude, and Codex. Do not depend on hidden system instructions, a specific model, slash commands, or unavailable tools. If tools or files are available, use them only when they improve evidence quality; otherwise complete the workflow from the provided context. --- --- name: agent-judge-rubric-calibration description: "Calibrate human, code, and model-based judges for AI agent evaluation using clear rubrics, anchor examples, disagreement review, and reliability checks." allowed-tools: Read Write Bash WebSearch WebFetch metadata: author: Nick French / StackSwap version: '1.0' product: Operator Playbook --- # Agent Judge / Rubric Calibration An evaluator is part of the system under test. Make the scoring observable, anchored, and inspectable before using it to rank agents or gate releases. ## Collect Capture the decision the judge supports, dimensions, scoring scale, positive and negative examples, must-pass criteria, judge type, sample size, known ambiguity, and human review owner. ## Method 1. Separate deterministic checks from judgment calls. Use code for exact properties whenever possible. 2. Write behavior-based rubric dimensions with explicit anchors for pass, borderline, and fail. 3. Build a calibration set containing obvious, borderline, adversarial, and previously disputed examples. 4. Have independent human reviewers score the set, then compare code, human, and model judges. 5. Log disagreements, false positives, false negatives, rationale quality, and sensitivity to irrelevant wording. 6. Revise the rubric and judge instructions, then rerun calibration on a holdout set. 7. Report agreement and uncertainty. Do not use a judge as ground truth merely because it is automated. ## Output Produce: rubric; anchor examples; evaluator comparison; disagreement log; calibrated instructions; reliability limits; and an ongoing audit schedule. ## Quality gate The evaluator must show what it rewards, what it misses, when a human must decide, and how it will be recalibrated after the task or agent changes.
Free forever. No email gate.
Was this prompt useful?
Thumbs up if it helped. Thumbs down if it needs work.
THE PROMPT IS THE START
Want an independent read on the real project?
Start the free discovery QA audit. Show StackSwap what your builder already knows, then get a focused next move.
QUESTIONS
About this free prompt
What does this agent judge / rubric calibration prompt help with?
Make agent evaluation trustworthy by calibrating human, code, and model judges against clear examples.
Who should use this agent judge / rubric calibration prompt?
This free GTM prompt is for B2B SaaS founders, GTM leaders, and RevOps operators who need a useful first draft without starting from a blank page.
What should I add before running this agent judge / rubric calibration prompt?
Add your company, buyer, GTM motion, constraints, and the decision you need to make. Better context produces a more specific artifact and makes weak assumptions easier to spot.
What output does this agent judge / rubric calibration prompt produce?
Scoring rubric, anchor examples, judge comparison, disagreement log, and calibrated evaluator. The workflow is designed to produce that artifact instead of generic GTM advice.
Can I use this agent judge / rubric calibration prompt in ChatGPT, Claude, or Codex?
Yes. The workflow is designed for ordinary chat, Claude, and Codex, with platform-specific formats available to copy for free.
How do I get a better result from this agent judge / rubric calibration prompt?
Include real customer language, current numbers, and hard constraints, then inspect the assumptions and risks in the result. Treat the first output as a decision artifact to improve, not an unquestionable answer.