Evaluating & Iterating Prompts
How to tell whether a result is actually good enough, what exactly to change when it is not β and when the prompt was never the problem.
Objectives
- Judge a result against the task you actually set, rather than against how good it sounds.
- Apply a small set of plain-language evaluation questions to any AI output.
- Distinguish a result that is wrong from one that is merely not what you wanted.
- Name which part of a prompt most likely caused a specific weakness.
- Change one meaningful thing at a time, and say what you expect it to fix.
- Recognize when repeated rewording is not going to help.
- Identify the cases where the problem is the task, the source, or the need for a qualified human.
- Decide when a result is good enough to use, and stop.
Introduction
Module 1 introduced the loop β ASK β READ β JUDGE β REFINE β and left the judging and refining deliberately shallow. Module 4 then raised the stakes: once a prompt becomes a pattern, its output starts to look familiar, and familiar output is much easier to stop reading carefully.
This module makes JUDGE and REFINE into real skills.
How do I tell whether this result is actually good enough? And when it is not, what exactly do I change?
Good Sounding Is Not Good Enough
The most common evaluation failure is being persuaded by tone. A confident, well-organized, fluently written answer can address a slightly different question than the one you asked β and read as though it succeeded.
Fluency is a property of the writing. Fitness is a property of the match between the result and the task you set. Only one of them is what you needed.
The Loop, Taken Seriously
ASK β INSPECT β IDENTIFY THE PROBLEM β CHANGE ONE THING β TRY AGAIN β YOUR JUDGMENT.
Module 1 taught this as four steps because a beginner needed the shape. Here it grows two, and both additions are the point of the module. INSPECT is separated from ASK because reading properly is a distinct act that people skip. IDENTIFY THE PROBLEM is separated from CHANGE because naming the weakness before editing is what turns a retry into an experiment.
The last step never moves. The loop produces candidates; you decide.
The evaluation loop, all six steps
Steps two and three are the ones people skip, and they are what turn a retry into an experiment. The last step never moves.
The six-step loop in order: ask, inspect, identify the problem, change one thing, try again, then your judgment. Module 1 taught this as four steps because a beginner needed the shape; here it grows two. Inspect is separate from ask because reading properly is a distinct act people skip. Identify the problem is separate from change because naming the weakness before editing is what turns a retry into an experiment. The middle steps repeat as needed. The last never moves: the loop produces candidates, and you decide.
1 Β· Ask
The request as you first make it.
2 Β· Inspect
Reading properly is a distinct act β and the one most often skipped.
3 Β· Identify the problem
Name the weakness before editing anything.
4 Β· Change one thing
So you can tell what actually made the difference.
5 Β· Try again
The loop repeats from here as often as it needs to.
6 Β· Your judgment
This step never moves. The loop produces candidates; you decide.
Seven Questions Worth Asking
Plain language, in the order a person naturally checks them.
- Did it do the task? Not "is it good" β did it do the thing you asked for?
- Is it accurate enough to use? Which claims would matter if they were wrong?
- Is anything important missing? Absence is harder to notice than error.
- Is the format usable? Could you act on it as it stands, or must you rebuild it?
- Does it fit the audience? Right content, wrong reader, is still a failure.
- Did it respect the constraints? Length, tone, what to avoid.
- Can I verify what matters? If you cannot check the load-bearing claim, that is the finding.
Most weak results fail two or three of these, and naming which ones is already most of the diagnosis.
Where Did It Go Wrong?
Module 2 gave a component locator for role, instruction and example. This extends it across everything the course has taught, and adds the row people most often need.
- Did the wrong work β the instruction: the task was not stated precisely enough.
- Right work, wrong assumptions β context: it was never told something it needed.
- Right content, wrong voice or depth β the role or the audience.
- Right substance, unusable shape β the output format or the example.
- Ignored a rule you take for granted β a constraint you never wrote down.
- Confidently states something untrue β verification. This is not a prompt problem.
A fluent false statement is not a phrasing failure, and rewording will not fix it. It is a signal to check the claim yourself, or to supply the source.
Change One Thing
When you change three things and the result improves, you have learned nothing about which change helped β and you now carry two edits you cannot justify.
Change one meaningful thing, and say in advance what you expect it to fix. If the result changes in a way you did not predict, that is more informative than a vague improvement: it usually means your diagnosis was wrong, which is worth knowing before you build a pattern on it.
Two honest caveats. The same prompt can produce somewhat different answers on different attempts, so a single comparison is evidence, not proof. And "one thing" means one meaningful thing β fixing a typo and adding a constraint at once is fine; adding a constraint and changing the audience is not.
When the Prompt Is Not the Problem
The most valuable judgment in this module is knowing when to stop editing the prompt.
- It does not know. The information was never available to it, and no phrasing creates it.
- The source is insufficient. You supplied material that does not contain the answer.
- The task is not appropriate. It needs a decision, an authority or an accountability that belongs to a person.
- It needs real expertise. Medical, legal, financial, safety-critical or regulated questions need a qualified human, not a better brief.
Rewording a prompt five times is itself a diagnosis. If clarity is not the bottleneck, more clarity will not help β and continuing to try is how people end up accepting a confident answer because they are tired of asking.
Evaluation Is Where Verification Happens
Evaluation is the step where an invented detail is either caught or published. When you check a result, check the things that would matter if they were wrong β names, numbers, dates, quotes, claims about people, anything you would be embarrassed or liable to get wrong.
And be careful what you paste while iterating. Repeated attempts tempt people to supply more and more of the underlying material each round. The privacy rule from Module 3 does not relax because you are on your fourth try. Deeper prompt safety, privacy and prompt-injection awareness belong to Module 7.
Knowing When to Stop
Iteration has a cost, and beginners routinely overshoot: a fourth revision of an internal note is worse value than sending the second one. Ask "is this good enough for what it is actually for?" rather than "could this be better?" β the second question never terminates.
Two useful markers: stop when the remaining flaws are ones you would fix faster by hand than by re-prompting; and stop when the last two attempts differ in wording but not in usefulness.
Human + AI
You define what "good enough" means for this task, and you are the one who decides that the work is finished. An assistant can produce another version indefinitely; it has no view of whether the last one already served your purpose, and it cannot tell you that the problem is the task rather than the prompt.
- An assistant will not tell you when it is wrong.
- A better prompt does not fix every answer.
- Confident does not mean correct.
- Perfect is not the target; good enough for the actual purpose is.
Practical exercise
β8 minDiagnose, predict, change one thing. Take a result that disappointed you β from this course or from your own use.
1. Run the seven questions and write down which ones it fails. Be specific: "missing the deadline" rather than "incomplete". 2. Name the likely cause using the symptom list β instruction, context, role, format, constraint, or verification. 3. Write your prediction before you edit anything: "If I change ___, I expect ___ to improve."
4. Change that one thing and compare. Did your prediction hold? 5. Decide and stop: is it good enough for its actual purpose? If your second attempt is no more useful than the first, say why, and consider whether the prompt was ever the problem.
No AI account is needed. If you have no result to hand, evaluate four written samples of the same request instead: one that is fluent but answers a slightly different question, one that is correct but delivered as a dense paragraph when bullets were asked for, one that is well shaped but states a figure that appears nowhere in the source, and one that is accurate and usable with a single awkward phrase. The third is the dangerous one; the fourth is the one to accept rather than polish. The prediction step is where the learning is β a prediction that fails teaches more than an improvement that just happens. Nothing is submitted, stored or graded.
Your progress
0 of 2 required activities complete in this module Β· course progress 0%
- β GlobSynk Labβ’ Β· optional
- β Reflection
- β Checkpoint
GlobSynk Labβ’
optional, β5 minAbout GlobSynk Labβ’. GlobSynk Labβ’ is the hands-on practice experience used throughout GlobSynk Academy. This Lab is optional hands-on practice: complete it now, skip it and continue the module, or return to it later. Skipping this Lab does not prevent you from continuing the course.
Take one prompt of your own and run two deliberate rounds: predict, change one thing, compare β then predict, change one thing, compare again.
Then answer the question this Lab exists for: which round actually improved the result, and which only changed it? Keep the version you would genuinely use, and note the one edit that made the difference. If neither round helped, that is a finding worth writing down too.
Reflection
β2 minThink of feedback you have received on your own work β the kind that helped, and the kind that just said "make it better." What made the difference? Now look at how you have been correcting AI results: which kind of feedback have you been giving?
This reflection is yours alone β it is never sent to GlobSynk or stored. Only the fact that you completed it is saved.
Checkpoint
Five questions, unscored, with instant feedback. Retry as often as you like β this is a learning aid, not an exam.
Answer all 5 questions to continue.
Key takeaways
- Fluency is not fitness. Judge the result against the task you set, not against how well it reads.
- Seven questions: did it do the task, is it accurate enough, is anything missing, is the format usable, does it fit the audience, did it respect the constraints, can you verify what matters.
- Match the symptom to the part of the prompt that produced it β and know the row that is not a prompt problem at all.
- Change one meaningful thing and say in advance what you expect it to fix.
- Rewording five times is itself a diagnosis: if clarity is not the bottleneck, more clarity will not help.
- Some failures need a source, a decision, or a qualified human β not a better brief.
- Stop when it is good enough for its actual purpose, not when it cannot be improved.
Practice in Prompt Lab
OptionalWant to try what you learned with real prompts? Prompt Lab is an optional practice environment, separate from this course.
Use Prompt Lab to compare two versions of one prompt, and practice changing a single thing at a time.
Practice Prompting in the Real WorldOpens in a new tab. Optional practice β never required for this module, the checkpoint, your progress, the Final Assessment, the certificate or Reward Points.
Before you move on
You can now say why a result missed, change the one thing likely to fix it, and recognize the cases where no amount of rewording will help β including the most common one, where the assistant simply never had the information.
That leaves an obvious question. When the information you need is spread across many documents, policies or records, supplying it by hand does not scale β and Module 3 said the supplier of context was you.
What happens when a system finds the relevant material for you is Module 6.
