A working capstone that makes learners apply and defend their reasoning to an AI examiner, scored against an objective-aligned rubric. Measuring transfer, not recall — live in five courses on this site.
Multiple-choice tests measure whether a learner can spot the right answer in a lineup. For manager and soft-skill training, that is the wrong question. The point of the course is behavior: can this person actually hold a hard line with a friend, give feedback that lands, de-escalate a furious customer? Recognition doesn’t predict any of that. I needed an assessment that surfaces application and reasoning under realistic pressure — and gives the learner feedback specific enough to act on.
Every course ends in a capstone that is a Socratic oral defense, not a quiz. The learner is dropped into a realistic scenario, answers in their own words, and the examiner probes weak spots — asking them to justify a move or apply it to a twist — before returning a criterion-referenced verdict with specific strengths and gaps. The build rests on two ideas from assessment design:
Every learning objective maps to a cue in the scenario and to a rubric criterion — so the thing taught, the thing practiced, and the thing measured are the same thing.
The task is a real dilemma, not a fact. Judgement is against an explicit rubric — Demonstrated / Partial / Not yet — and feedback is formative: what you showed, and where to go deeper.
Scenario: A former peer and work friend, Alex, is now the learner’s direct report. Alex has been coasting — missing details, turning work in late — and leans on the friendship to deflect when it comes up. Others have noticed Alex gets a pass. The learner must walk through exactly how they’d handle it, start to finish.
| Rubric criterion (what a strong answer shows) | Objective verified | Scenario cue that elicits it |
|---|---|---|
| Names the changed relationship honestly rather than pretending nothing shifted | Name the IC→manager shift; manage a former peer | “Six months ago Alex was your peer and a genuine work friend” |
| Holds the standard with specific, behavior-based examples and clear expectations | Run the core loop; hold people accountable fairly | “missing details, turning things in late” |
| Won’t let the friendship dodge accountability — resists the Buddy trap | Recognize and avoid the new-manager traps | “leaning on the friendship — don’t go boss-mode” |
| Corrects the fairness / perception problem with the rest of the team | Hold the standard fairly; make it explicit first | “others have quietly noticed you give Alex a pass” |
| Balances relationship and role — defines what “changed” looks like and by when | Hold the standard without losing the relationship | “you’d like to keep the relationship intact” |
Illustrative of a real session on the New Manager capstone. The learner gives a competent-but-incomplete answer; the examiner probes before judging.
What you demonstrated
Where to go deeper
Want the real thing? Take the New Manager course — it ends in this exact capstone, live.
This isn’t a mockup. The examiner runs on a stateless Cloudflare Worker endpoint backed by a large language model. It’s system-prompted to assess application over recall, to probe the reasoning behind an answer, and explicitly not to be fooled by confident-but-empty responses. It asks at most three follow-ups, then must return a structured verdict. Because the learner has to reason in their own words and defend it live, the format resists both guessing and paste-through AI answers — a live concern for assessment validity right now. Nothing is stored; the transcript lives only in the browser session.
This moves assessment off recognition and onto application and behavioral rehearsal — the learner safely practices a hard conversation and gets specific, rubric-aligned feedback, while the design produces defensible evidence of mastery mapped back to each objective. It’s live across five courses in this portfolio today, and the same examiner pattern generalizes to standalone assessment in any subject where you need to see whether someone can apply and defend an idea — not just recall it.
Honest scope: this is a working, validity-minded assessment design, not a controlled efficacy study. The claim is the design and the build — not a proven learning-gain number.