It’s one thing to argue that instructional design expertise matters more than AI tool selection or prompt quality. It’s another to actually demonstrate it. So rather than simply asserting the point, we ran a live test: the same learning brief, the same AI tool, given to two people with different levels of instructional design experience, and watched what happened.
The setup was deliberately simple, almost boring by design, because the goal wasn’t to showcase an impressive demo. It was to isolate one variable and see what changed when everything else stayed constant.
Table Of Content
The Setup: Controlling for Everything Except Expertise
Two people worked from the exact same learning brief. Both used Claude. Both had access to the exact same source content. Nothing about the tool, the prompt access, or the raw material differed between the two. The only variable that changed was the instructional thinking each person brought to the interaction with the AI.
If the popular belief that “the better the prompt, the better the learning experience” were the full story, the two outcomes should have converged, since both people had equal access to the same capable AI tool. If the belief that “AI will replace instructional designers” were closer to true, the outcomes should also have converged, since the AI was doing the heavy lifting of content generation in both cases regardless of who was directing it.
Neither of those predictions held up.
What Stayed the Same
Before looking at what diverged, it’s worth being precise about what didn’t change, because the list is longer than most people expect. The learning brief was identical. The AI tool was identical, Claude in both cases. The source content was identical. And critically, the AI capabilities available to both participants were exactly equal; nothing about access, features, or the underlying model differed between the two sessions.
If you were betting purely on inputs, you’d expect two very similar outputs.
What Changed
Despite all of those constants, the outputs diverged sharply across four dimensions:
- the instructional thinking behind each approach,
- the learning strategy each person applied,
- the resulting learning experience a learner would actually walk through, and, most importantly,
- the performance focus of each solution: whether the course was built to change what someone does on the job, or simply to convey information they could recall on a quiz.
The person with deeper instructional design expertise didn’t just produce a “better” version of the same course. They approached the brief differently from the very first prompt, asking questions about the actual performance gap, the conditions the learner would apply the training in, and what evidence would demonstrate the training had worked. Those questions shaped every subsequent interaction with Claude, and the resulting learning experience reflected that thinking throughout, not just in the final polish.
The other approach, while producing content that was fluent, well-organized, and superficially complete, treated the brief more literally: cover the topic, structure it logically, generate assessment questions that check recall. The output looked like a finished course. It just wasn’t built around the same understanding of what the learner actually needed to be able to do differently afterward.
The Uncomfortable Part: Both Outputs Looked Reasonable
This is the detail that makes the demo more useful than a simple “expertise wins” anecdote. Both outputs were coherent, professionally structured, and would likely pass a surface-level stakeholder review. Neither was sloppy or obviously broken. If you only glanced at slide count, formatting quality, or how quickly each was produced, you might not immediately see a meaningful difference.
The gap only became visible when you asked a more specific question: would this course actually change what a learner does on the job? That’s precisely the kind of gap that’s easy to miss in a rushed review cycle, and precisely why “the content looks complete” is a dangerously low bar for evaluating AI-augmented learning design.
What the Demo Actually Revealed
The conclusion isn’t complicated, but it’s one that a lot of L&D strategy quietly contradicts in practice: the difference wasn’t the AI. It was the expertise guiding it.
Same input. Same AI. Two completely different outputs.
That single sentence captures more about AI adoption risk in L&D than most vendor comparison spreadsheets do, because it locates the variable that actually determines quality, and it’s not the one most procurement conversations focus on.
Tracing the Divergence Back to the First Few Prompts
It’s worth being specific about where the two approaches actually split, because it happened earlier than most people expect, well before either person had generated any learner-facing content.
The expertise-led approach spent its early interactions with Claude establishing context before asking for anything resembling course material:
- what the learner currently struggles with,
- what conditions they’ll be applying this knowledge in, and
- what a manager would observe six weeks later if the training had actually worked.
Only after establishing that groundwork did the interaction move toward generating objectives, activities, and assessments, each shaped by the answers already established.
The more literal approach moved to content generation almost immediately, treating the brief itself as sufficient context. Claude, in both cases, responded capably to what it was actually asked. That’s an important detail: this wasn’t a case of one prompt being technically better-written than the other in a prompt-engineering sense.
The literal approach’s prompts were clear, well-structured, and grammatically precise. They simply skipped a layer of context-setting that shapes everything downstream, and Claude, capable as it is, has no way to supply context that was never provided.
Why This Matters Beyond a Single Demo
It would be easy to dismiss this as one anecdotal comparison, but the pattern it illustrates shows up consistently across real project work, not just in a controlled demonstration.
Every time an organization scales AI usage across an L&D team without also scaling instructional design capability, it’s effectively running this same experiment at organizational scale, just without watching it happen in real time.
Teams that invest heavily in AI tool access and prompt training while treating instructional design expertise as a secondary concern are, in effect, betting on the outcome that this demo disproved: that equal AI access produces roughly equal learning quality. It doesn’t.
The instructional thinking behind the interaction is doing far more of the work than the tool itself.
Running This Test Inside Your Own Team
This kind of comparison is worth running internally, not just watching as a webinar demo.
Take a real, already-completed learning brief, ideally one your team has mixed confidence in, and have two instructional designers with different experience levels work through it independently with Claude, without comparing notes until both are finished.
Then evaluate both outputs against a specific question: not “which one looks more polished,” but “which one would more reliably change what a learner does on the job, and why.”
The value of this exercise isn’t proving a point your team already suspects. It’s surfacing, concretely, what the more experienced designer actually did differently in their interaction with Claude, the questions they asked before generating content, the criteria they applied when evaluating options.
Those specifics are far more transferable to the rest of the team than the general principle “expertise matters,” because they show the less experienced instructional designers exactly what expertise looks like in practice when working with an AI tool, rather than leaving it as an abstract quality to develop on their own over time.
The Practical Implication for L&D Leaders
If your organization is scaling AI adoption across your L&D function, the demo above suggests a specific diagnostic question worth asking: are we scaling AI access at the same rate we’re scaling instructional design capability, or are we quietly assuming the tool will compensate for gaps in the latter?
Rolling out Claude access to every instructional designer on a team is a reasonable and often valuable move. Assuming that rollout alone will produce consistent learning quality across the team, regardless of each person’s instructional design experience, is the assumption this demo directly challenges.
There’s also a coaching implication worth noting here.
If instructional design quality varies this visibly based on the questions someone asks before generating content, that variance is teachable.
Less experienced Instructional Designers can be shown, explicitly, the kind of context-setting questions that shaped the stronger output in this demo: what the learner needs to do differently, under what conditions, and how success would be observed.
That’s a far more concrete coaching target than generic advice to “think more strategically,” and it’s exactly the kind of specific, transferable skill that separates teams who get consistent value from AI adoption from teams whose results vary wildly by which person happens to be running a given project.
The next piece in this series introduces the framework we use to operationalize this distinction consistently across projects, rather than leaving instructional judgment to vary person by person: RAPID-AI, our five-stage approach to responsible AI use in learning design.
Frequently Asked Questions
1. Can two people get very different results from Claude using the same prompt?
A. Yes, particularly in instructional design work. While the tool and the literal prompt wording can be identical, the underlying instructional thinking, what questions get asked before and during the AI interaction, what the person is trying to determine, shapes the direction of the entire interaction and produces materially different outputs even from a shared starting point.
2. How can L&D teams tell if an AI-generated course is instructionally sound versus just polished?
A. Surface indicators like formatting quality, structure, and completeness aren’t reliable signals on their own. A more useful test is whether the course is built around a specific, verifiable performance outcome, what the learner should be able to do differently afterward, rather than simply covering the topic comprehensively. Courses built for recall often look just as finished as courses built for behavior change.
3. Does this mean AI tool choice doesn’t matter for learning design quality?
A. Tool choice still matters, since different tools have different capabilities for handling complex source material and generating instructionally useful output, which is why tool evaluation was worth doing carefully in the first place. But tool choice is a smaller factor in final learning quality than the instructional expertise directing the tool, and treating the two as equally important risks under-investing in the one that matters more.
View the original article and our Inspiration here

