Panel review: the method on its own build
Case study · AI Tools & Strategy · Learning Design
The method’s claim is that a panel of AI agents, three agents I prompt to review work from different angles, can hold work to a standard with a human owning every decision. Here it is tested on a real artifact: after building the graduate data-science simulations, I pointed the panel back at them. It caught a statistical error and a chart that argued the opposite of its own point, both before any student saw them.
What this is
Every other tab describes the method. This one shows it working on something concrete. I built the five simulations in the Authentic Assessment suite, spanning model evaluation, clustering, and AI governance, then convened the review panel and asked each agent to critique the build. What follows is what they returned and what I changed in response. I made every call; the panel widened what one builder could see.
This is the same panel that drafts and checks a course, cast in three roles and pointed at a finished object instead of a blank page.
The panel
The subject-matter seat is filled by whatever discipline the course needs. This build happened to be a graduate data-science course, so the expert here is a data scientist. For a nursing course it would be a nurse educator, for a welding course a certified welding inspector. The instructional-designer and student seats stay the same from course to course; only the subject expert changes.
What the panel caught
Subject-matter expert (data scientist)
- A real technical error, caught before a student could. One page credited PCA with keeping a single feature from dominating the distance metric. That is what standardization does (scaling the features so no one of them can dominate); PCA does not, and on unscaled data PCA is itself dominated by the highest-variance feature. Split into two accurate statements.This is the strongest evidence for the method: the expert agent caught a genuine mistake a single builder had missed.
- A misleading chart, corrected. The governance cover’s “overall” line sat below most of the bars, which shows the disparity rather than hiding it. A real size-weighted aggregate sits high, near the majority groups, while one small subgroup collapses far below. Redrawn so the visualization tells the true story of disparate impact (one subgroup harmed) concealed by an aggregate.
- Fidelity fixes to the other data visuals. The clustering scatter’s ambiguous points now sit between the clusters (matching the “how many, really?” premise, and its own alt text); the capstone time series’ drift marker now lands exactly on the break in the line.
Instructional designer (learning-experience designer)
- Orient the learner first. Add a one-line label to every simulation stating its module and whether it is a graded assessment or a preparation tool.
- Make the throughline explicit. State plainly that the governance simulation carries forward the very model the student audited in the model-evaluation module, so the sequence reads as one story.
- Lower the vocabulary load. Expand “NIST AI RMF” (the NIST AI Risk Management Framework) to its full name on first use, standardize the repeated framing across pages, and cut redundancy.
Synthetic student
- Give me a first step. Landing on a graduate simulation with no visible “start here” is intimidating; add a clear entry cue.
- Do not assume the acronyms. The framework name read as a wall of letters before it was spelled out.
- Tell me what counts. Distinguish the practice tools (the capstone prep and the defense rehearsal) from the graded assessments, so I know what is being assessed.
What changed
Result
Every finding above was implemented before the suite reached a student: the statistical error corrected, the governance chart redrawn, the three data visuals made faithful, a module-and-status label added to each page, the governance throughline stated, acronyms expanded, framing standardized, and a clear starting cue added. The simulations are now more accurate, clearer, and easier to begin.
Why it matters
The panel is not a rubber stamp. It caught a real statistical mistake and a chart that argued the opposite of its point, exactly the failures a single builder stops seeing after enough hours in the work. The instructional-design and student passes then closed the gap between “correct” and “usable.” None of it replaced my judgment; I accepted, adjusted, or declined each note. That is the whole method in one loop: agents widen what one person can catch, a human owns every decision, and the work is hardened before it reaches a learner.
See the reviewed simulations in the Authentic Assessment suite, or how the panel drafts a course from the start in The Build Skill.