I have a problem with testing prompts, and it took me a while to admit it.
I cheat. Not on purpose. But when I run two prompts side by side and compare the results, I have a strong tendency to prefer the one I expected to win.
This is not a character flaw unique to me. It's how humans work. If I write prompt B thinking it's the improved version, I'll read B's output more charitably. I'll notice B's strengths and B's output will feel sharper. I'll excuse its weaknesses as "close enough" while holding prompt A to a stricter standard.
I did this for months without realizing it. Then I started writing down my comparisons instead of just eyeballing them, and discovered that my instincts were wrong roughly a third of the time.
Here's the method I use now. It's built around preventing exactly that kind of self-deception.

The Problem With Side-by-Side Testing
Let me be specific about how the cheating works.
Recency bias. The second output is fresher in my mind. If B is better in a small way, I'll overweight it. If A was better, I might not remember why.
Effort bias. If I spent more time writing prompt B, I want it to win. That's sunk-cost thinking, and it's completely irrational, but it's real. I've caught myself defending a prompt just because I'd tweaked it for twenty minutes.
Confirmation bias. I usually have a hypothesis. Hypotheses are useful, and they also make me look for evidence that supports them.
Vagueness bias. "This one reads better" is not a measurement. It's a feeling. Feelings about writing quality are unreliable, especially when comparing your own prompts.
The fix isn't to be more objective. That's not possible. The fix is to build a process that doesn't depend on being objective.
The Method
Six steps. It takes about ten minutes per comparison, which sounds like a lot until you realize how much time a bad prompt costs you over months of use.
Step One: Define the Win Condition Before You Run Anything
This is the most important step and the one I skipped for the longest time.
Before running either prompt, I write down what "better" means for this specific task. Not vague criteria—specific, countable things.
For an email prompt, that might be:
Fewer edits required
Shorter output
No filler phrases
Correct tone on first pass
For a summary prompt:
No invented details
Correctly identified decisions vs. discussion
Length under X words
No missed action items
Writing this down first matters because it stops me from moving the goalposts after I see the results. If prompt A wins on length but prompt B wins on accuracy, and I decided accuracy was the priority before running them, I don't get to change my mind because B's output felt nicer.
Step Two: Run Both Prompts on the Same Input
Same input, same day, same chat session if possible.
This sounds obvious. I've broken this rule more times than I want to admit—comparing prompt A's output from last Tuesday against prompt B's output from today, on slightly different source material. That's not a comparison. That's a story I'm telling myself.
Same input. Same conditions. No exceptions.
Step Three: Score Both Outputs Against the Win Condition
Here's where the counting happens.
For each criterion, I score both outputs. Simple scale:
2 = met the criterion fully
1 = partially met
0 = didn't meet
Four criteria, two outputs, eight scores. Takes two minutes.
The reason this works is that it forces specificity. I can't score "reads better"—that's not on the list. I can score "under 100 words." I can score "no filler phrases." I can score "correctly separated decisions from discussion."
Specific criteria can be checked. Vague criteria can't.
Step Four: Count Edits
Here's the criterion I care about most, and it's not on the list because it's measured differently.
After scoring, I take the output I'd actually use from each prompt and edit it until I'd be willing to send it. Then I count the edits.
Every change is one edit. Word swaps, deletions, additions, reordering, all count as one each.
This is the number that matters most, because it's the closest thing to real cost. A prompt that produces beautiful output I have to fix eleven times is worse than one that produces decent output I fix once.
I keep a rough tally in my head during the edit, then write it down. It usually surprises me.

Step Five: Write Down the Verdict Before Discussing It
I write one sentence: "Prompt [X] won because [specific reason]."
Not "prompt X felt better." A specific, checkable reason tied to the win condition I defined at the start.
If I can't write that sentence without hedging, that's a signal the comparison was too close to call, which is also a valid outcome.
Step Six: Note What Surprised Me
The last step is the one that's taught me the most.
After the verdict, I write one line about anything unexpected. A prompt that won on editing but lost on accuracy. A "simpler" prompt that outperformed a complex one. A constraint I thought was minor that turned out to matter a lot.
Those notes accumulate. Over time, they've become a pretty good picture of what actually works—as opposed to what I assumed would work.
A Real Example
I ran this on two prompts for the same task: turning meeting notes into a follow-up email.
Win condition, defined first:
No invented details
Correctly separated decisions from discussion
Under 150 words
No filler phrases
Prompt A: "Turn these notes into a follow-up email for the attendees."
Prompt B: The four-section decision prompt, then "convert the sections into a short follow-up email."
Scoring
Criterion | Prompt A | Prompt B |
|---|---|---|
No invented details | 1 | 2 |
Decisions vs. discussion | 0 | 2 |
Under 150 words | 1 | 2 |
No filler phrases | 0 | 1 |
Prompt A: 2. Prompt B: 7.
Edit Count
Prompt A: I made 9 edits before I'd send it.
Prompt B: I made 3 edits.
Verdict
Prompt B won because it separated decisions from discussion before drafting, which meant the email only contained things that were actually decided.
Surprise
Prompt B still produced one filler phrase—a closing line I deleted. The constraints reduced filler but didn't eliminate it. Worth adding one more explicit "no closing paragraph" rule next time.
That surprise note is now a permanent part of my prompt. It came from the last step, which I almost skipped.
Where This Method Fails
I want to be honest about the limits, because a method that claims to remove bias entirely is lying.
It's slow. Ten minutes per comparison. Worth it if you're building a prompt you'll use for months. Not worth it for a one-off task.
Scoring is still subjective. "No invented details" is checkable. "Correct tone" is not, and I know it. When a criterion is genuinely subjective, I try to replace it with something countable, or I accept that this particular score is soft.
Small samples lie. One comparison is one data point. I've had prompts win by a wide margin on one task and lose on the next. Three runs on three different inputs is better than one run on one.
The win condition can be wrong. If I define the criteria badly, I'll pick the prompt that wins the wrong game. This has happened, and the only fix is paying attention to whether the output actually works in practice, not just whether it scored well.
I still have preferences. The method reduces bias. It doesn't eliminate it. I'm aware that I'm still the one defining the criteria and doing the scoring.
Why This Is Worth the Effort
Because prompt quality compounds.
A good prompt isn't used once. It's used dozens of times over months. A prompt that saves five minutes per use is worth twenty-five minutes of testing, easily.
But the flip side is also true: a prompt that feels better but actually requires more editing costs you time every single time you use it. Bad prompts compound too.
Ten minutes of structured comparison, once, is how I avoid that.
The Test Card
Task: Compare two prompts objectively without letting preference or expectation skew the result.
Prompt or workflow: Define win conditions first, run both on identical input, score against criteria, count edits, write a specific verdict, note surprises.
Starting conditions: Two prompts for the same task, one clearly expected to win.
Time before: Eyeballing outputs and picking one, with no record of why. Repeatedly choosing prompts that felt better but weren't.
Time after: About 10 minutes per comparison, including scoring and edit counting.
Editing required: Not applicable—this is the measurement method itself.
What went wrong: Early comparisons were unreliable because the win condition was defined after seeing the outputs. Fixed by writing criteria down before running anything. Also learned that my instinct about which prompt would win was wrong about a third of the time.
Privacy notes: Same sanitization rules as always—names, client identifiers, and financial figures removed before pasting. Comparing prompts doesn't change what's safe to paste.
Who should not use this method: Anyone comparing prompts they'll only use once. Anyone who won't write down the criteria first—without that step, this is just eyeballing with extra steps. And anyone in a workplace that prohibits external AI tools, since no testing method changes that constraint.
The Bottom Line
I was bad at comparing prompts because I wanted one of them to win. The fix wasn't to be more objective—it was to build a process that didn't need me to be.
Write the criteria first. Score against them. Count the edits. Write the verdict in one specific sentence. Note what surprised you.
My instinct was wrong about a third of the time. That's the number that convinced me this was worth doing.
I tried it at my desk so you don't have to. Useful beats impressive, and measuring beats guessing—even when the thing you're measuring is your own judgment.