Skip to content
Chat That WorksCHAT THAT WORKS
01Workflows 02Tools & assets 03Cost ledger 04Notes 05About GONGXU
CHAT THAT WORKS · AI WORKFLOW FIELD GUIDE
All content and pricing are prototype sample data
Prompt Bench

How I Compare Two ChatGPT Answers Without Fooling Myself

An ordinary office worker shares a practical, tested six-step workflow designed to eliminate human bias when comparing competing ChatGPT prompts. Instead of trusting gut feelings or falling for recency and confirmation biases, this guide outlines how to pre-define strict win conditions, run identical inputs, count actual edits, and record objective verdicts. It demonstrates how a structured review process leads to better prompt engineering and reliable results.

Sep 28, 2026
Prompt Bench

I have a problem with testing prompts, and it took me a while to admit it.

I cheat. Not on purpose. But when I run two prompts side by side and compare the results, I have a strong tendency to prefer the one I expected to win.

This is not a character flaw unique to me. It's how humans work. If I write prompt B thinking it's the improved version, I'll read B's output more charitably. I'll notice B's strengths and B's output will feel sharper. I'll excuse its weaknesses as "close enough" while holding prompt A to a stricter standard.

I did this for months without realizing it. Then I started writing down my comparisons instead of just eyeballing them, and discovered that my instincts were wrong roughly a third of the time.

Here's the method I use now. It's built around preventing exactly that kind of self-deception.

A close-up documentary photo of a handwritten checklist on a notebook defining prompt win conditions, with a pen nearby.

The Problem With Side-by-Side Testing

Let me be specific about how the cheating works.

Recency bias. The second output is fresher in my mind. If B is better in a small way, I'll overweight it. If A was better, I might not remember why.

Effort bias. If I spent more time writing prompt B, I want it to win. That's sunk-cost thinking, and it's completely irrational, but it's real. I've caught myself defending a prompt just because I'd tweaked it for twenty minutes.

Confirmation bias. I usually have a hypothesis. Hypotheses are useful, and they also make me look for evidence that supports them.

Vagueness bias. "This one reads better" is not a measurement. It's a feeling. Feelings about writing quality are unreliable, especially when comparing your own prompts.

The fix isn't to be more objective. That's not possible. The fix is to build a process that doesn't depend on being objective.

The Method

Six steps. It takes about ten minutes per comparison, which sounds like a lot until you realize how much time a bad prompt costs you over months of use.

Step One: Define the Win Condition Before You Run Anything

This is the most important step and the one I skipped for the longest time.

Before running either prompt, I write down what "better" means for this specific task. Not vague criteria—specific, countable things.

For an email prompt, that might be:

  • Fewer edits required

  • Shorter output

  • No filler phrases

  • Correct tone on first pass

For a summary prompt:

  • No invented details

  • Correctly identified decisions vs. discussion

  • Length under X words

  • No missed action items

Writing this down first matters because it stops me from moving the goalposts after I see the results. If prompt A wins on length but prompt B wins on accuracy, and I decided accuracy was the priority before running them, I don't get to change my mind because B's output felt nicer.

Step Two: Run Both Prompts on the Same Input

Same input, same day, same chat session if possible.

This sounds obvious. I've broken this rule more times than I want to admit—comparing prompt A's output from last Tuesday against prompt B's output from today, on slightly different source material. That's not a comparison. That's a story I'm telling myself.

Same input. Same conditions. No exceptions.

Step Three: Score Both Outputs Against the Win Condition

Here's where the counting happens.

For each criterion, I score both outputs. Simple scale:

  • 2 = met the criterion fully

  • 1 = partially met

  • 0 = didn't meet

Four criteria, two outputs, eight scores. Takes two minutes.

The reason this works is that it forces specificity. I can't score "reads better"—that's not on the list. I can score "under 100 words." I can score "no filler phrases." I can score "correctly separated decisions from discussion."

Specific criteria can be checked. Vague criteria can't.

Step Four: Count Edits

Here's the criterion I care about most, and it's not on the list because it's measured differently.

After scoring, I take the output I'd actually use from each prompt and edit it until I'd be willing to send it. Then I count the edits.

Every change is one edit. Word swaps, deletions, additions, reordering, all count as one each.

This is the number that matters most, because it's the closest thing to real cost. A prompt that produces beautiful output I have to fix eleven times is worse than one that produces decent output I fix once.

I keep a rough tally in my head during the edit, then write it down. It usually surprises me.

A documentary-style photo of a laptop and a notebook showing comparative data and editing notes on a desk.

Step Five: Write Down the Verdict Before Discussing It

I write one sentence: "Prompt [X] won because [specific reason]."

Not "prompt X felt better." A specific, checkable reason tied to the win condition I defined at the start.

If I can't write that sentence without hedging, that's a signal the comparison was too close to call, which is also a valid outcome.

Step Six: Note What Surprised Me

The last step is the one that's taught me the most.

After the verdict, I write one line about anything unexpected. A prompt that won on editing but lost on accuracy. A "simpler" prompt that outperformed a complex one. A constraint I thought was minor that turned out to matter a lot.

Those notes accumulate. Over time, they've become a pretty good picture of what actually works—as opposed to what I assumed would work.

A Real Example

I ran this on two prompts for the same task: turning meeting notes into a follow-up email.

Win condition, defined first:

  1. No invented details

  2. Correctly separated decisions from discussion

  3. Under 150 words

  4. No filler phrases

Prompt A: "Turn these notes into a follow-up email for the attendees."

Prompt B: The four-section decision prompt, then "convert the sections into a short follow-up email."

Scoring

Criterion

Prompt A

Prompt B

No invented details

1

2

Decisions vs. discussion

0

2

Under 150 words

1

2

No filler phrases

0

1

Prompt A: 2. Prompt B: 7.

Edit Count

Prompt A: I made 9 edits before I'd send it.
Prompt B: I made 3 edits.

Verdict

Prompt B won because it separated decisions from discussion before drafting, which meant the email only contained things that were actually decided.

Surprise

Prompt B still produced one filler phrase—a closing line I deleted. The constraints reduced filler but didn't eliminate it. Worth adding one more explicit "no closing paragraph" rule next time.

That surprise note is now a permanent part of my prompt. It came from the last step, which I almost skipped.

Where This Method Fails

I want to be honest about the limits, because a method that claims to remove bias entirely is lying.

It's slow. Ten minutes per comparison. Worth it if you're building a prompt you'll use for months. Not worth it for a one-off task.

Scoring is still subjective. "No invented details" is checkable. "Correct tone" is not, and I know it. When a criterion is genuinely subjective, I try to replace it with something countable, or I accept that this particular score is soft.

Small samples lie. One comparison is one data point. I've had prompts win by a wide margin on one task and lose on the next. Three runs on three different inputs is better than one run on one.

The win condition can be wrong. If I define the criteria badly, I'll pick the prompt that wins the wrong game. This has happened, and the only fix is paying attention to whether the output actually works in practice, not just whether it scored well.

I still have preferences. The method reduces bias. It doesn't eliminate it. I'm aware that I'm still the one defining the criteria and doing the scoring.

Why This Is Worth the Effort

Because prompt quality compounds.

A good prompt isn't used once. It's used dozens of times over months. A prompt that saves five minutes per use is worth twenty-five minutes of testing, easily.

But the flip side is also true: a prompt that feels better but actually requires more editing costs you time every single time you use it. Bad prompts compound too.

Ten minutes of structured comparison, once, is how I avoid that.

The Test Card

Task: Compare two prompts objectively without letting preference or expectation skew the result.

Prompt or workflow: Define win conditions first, run both on identical input, score against criteria, count edits, write a specific verdict, note surprises.

Starting conditions: Two prompts for the same task, one clearly expected to win.

Time before: Eyeballing outputs and picking one, with no record of why. Repeatedly choosing prompts that felt better but weren't.

Time after: About 10 minutes per comparison, including scoring and edit counting.

Editing required: Not applicable—this is the measurement method itself.

What went wrong: Early comparisons were unreliable because the win condition was defined after seeing the outputs. Fixed by writing criteria down before running anything. Also learned that my instinct about which prompt would win was wrong about a third of the time.

Privacy notes: Same sanitization rules as always—names, client identifiers, and financial figures removed before pasting. Comparing prompts doesn't change what's safe to paste.

Who should not use this method: Anyone comparing prompts they'll only use once. Anyone who won't write down the criteria first—without that step, this is just eyeballing with extra steps. And anyone in a workplace that prohibits external AI tools, since no testing method changes that constraint.

The Bottom Line

I was bad at comparing prompts because I wanted one of them to win. The fix wasn't to be more objective—it was to build a process that didn't need me to be.

Write the criteria first. Score against them. Count the edits. Write the verdict in one specific sentence. Note what surprised you.

My instinct was wrong about a third of the time. That's the number that convinced me this was worth doing.

I tried it at my desk so you don't have to. Useful beats impressive, and measuring beats guessing—even when the thing you're measuring is your own judgment.

Comments

No comments yet.

Leave a comment

RELATED

Read next

All notes
Prompt Bench
Sep 30, 2026

What Information I Remove Before Pasting Work Into ChatGPT

An ordinary office worker shares a practical, tested pre-paste sanitization workflow designed to prevent data leaks when using ChatGPT for work. Instead of relying on vague caution, this guide outlines a strict two-minute rule and five essential sanitization categories—names, numbers, identifiers, personal details, and format artifacts—to safely scrub sensitive corporate data before it ever touches a third-party chat log.

Prompt Bench
Sep 27, 2026

“Make This Better” Is Not a Prompt: My First Failed Test

An ordinary office worker shares a revealing first-hand test of a common AI mistake: typing vague commands like “make this better.” Instead of helpful polish, the prompt produced generic fluff and diluted key facts. The guide breaks down why vague instructions fail, explains the pitfalls of unstructured AI output, and outlines a practical five-part prompt structure that saves actual working time on Monday mornings.

Prompt Bench
Sep 26, 2026

The Five-Part Prompt I Use When I Need a Useful First Draft

An ordinary office worker shares a reliable, tested five-part prompt structure designed to produce actually usable first drafts. Instead of vague commands or lengthy prompt engineering, the guide breaks down how role, task, context, constraints, and format work together to stop artificial intelligence from guessing. It highlights essential privacy precautions, strict rules against unverified facts, and practical ways to save real editing time on Monday mornings.