Assessment design · Featured case study

Proving the learning actually happened.

A working capstone that makes learners apply and defend their reasoning to an AI examiner, scored against an objective-aligned rubric. Measuring transfer, not recall — live in five courses on this site.

The problem

Multiple choice tests recognition. The job is behavior.

Multiple-choice tests measure whether a learner can spot the right answer in a lineup. For manager and soft-skill training, that is the wrong question. The point of the course is behavior: can this person actually hold a hard line with a friend, give feedback that lands, de-escalate a furious customer? Recognition doesn’t predict any of that. I needed an assessment that surfaces application and reasoning under realistic pressure — and gives the learner feedback specific enough to act on.

The design

A short oral defense with an AI examiner.

Every course ends in a capstone that is a Socratic oral defense, not a quiz. The learner is dropped into a realistic scenario, answers in their own words, and the examiner probes weak spots — asking them to justify a move or apply it to a twist — before returning a criterion-referenced verdict with specific strengths and gaps. The build rests on two ideas from assessment design:

Constructive alignment

Biggs

Every learning objective maps to a cue in the scenario and to a rubric criterion — so the thing taught, the thing practiced, and the thing measured are the same thing.

Criterion-referenced & authentic

Performance assessment

The task is a real dilemma, not a fact. Judgement is against an explicit rubric — Demonstrated / Partial / Not yet — and feedback is formative: what you showed, and where to go deeper.

Alignment map

New Manager Essentials — capstone

Scenario: A former peer and work friend, Alex, is now the learner’s direct report. Alex has been coasting — missing details, turning work in late — and leans on the friendship to deflect when it comes up. Others have noticed Alex gets a pass. The learner must walk through exactly how they’d handle it, start to finish.

Rubric criterion (what a strong answer shows)Objective verifiedScenario cue that elicits it
Names the changed relationship honestly rather than pretending nothing shiftedName the IC→manager shift; manage a former peer“Six months ago Alex was your peer and a genuine work friend”
Holds the standard with specific, behavior-based examples and clear expectationsRun the core loop; hold people accountable fairly“missing details, turning things in late”
Won’t let the friendship dodge accountability — resists the Buddy trapRecognize and avoid the new-manager traps“leaning on the friendship — don’t go boss-mode”
Corrects the fairness / perception problem with the rest of the teamHold the standard fairly; make it explicit first“others have quietly noticed you give Alex a pass”
Balances relationship and role — defines what “changed” looks like and by whenHold the standard without losing the relationship“you’d like to keep the relationship intact”
See it work

An examiner exchange, probing the gap.

Illustrative of a real session on the New Manager capstone. The learner gives a competent-but-incomplete answer; the examiner probes before judging.

AI Examiner
You’re now managing Alex, who used to be your peer and friend. The work has been slipping and Alex leans on the friendship when you raise it. Walk me through how you’d open the conversation.
You
I’d meet privately and be direct but warm. I’d name that our relationship has changed, then give specific examples — the two late deliverables last week, the missing figures in the report — and say clearly what I need going forward.
AI Examiner
Good — you named the shift and used specifics. Here’s the twist: the rest of the team has quietly noticed Alex gets a pass. Naming it privately doesn’t fix what they’ve already seen. How do you handle that?
You
That’s fair. I’d hold the same standard for everyone going forward and make expectations visible to the team, so it’s consistent rather than singling Alex out — but I hadn’t thought about the trust that’s already been dented.
Partial

What you demonstrated

  • Named the peer-to-manager shift directly.
  • Used specific, behavior-based examples instead of vague pressure.

Where to go deeper

  • Address the team’s perception proactively, not only after prompting.
  • Define what “fixed” looks like with a concrete timeframe so the standard is measurable.

Want the real thing? Take the New Manager course — it ends in this exact capstone, live.

Under the hood

Designed and built end to end.

This isn’t a mockup. The examiner runs on a stateless Cloudflare Worker endpoint backed by a large language model. It’s system-prompted to assess application over recall, to probe the reasoning behind an answer, and explicitly not to be fooled by confident-but-empty responses. It asks at most three follow-ups, then must return a structured verdict. Because the learner has to reason in their own words and defend it live, the format resists both guessing and paste-through AI answers — a live concern for assessment validity right now. Nothing is stored; the transcript lives only in the browser session.

Why it matters

Evidence of mastery, mapped to every objective.

This moves assessment off recognition and onto application and behavioral rehearsal — the learner safely practices a hard conversation and gets specific, rubric-aligned feedback, while the design produces defensible evidence of mastery mapped back to each objective. It’s live across five courses in this portfolio today, and the same examiner pattern generalizes to standalone assessment in any subject where you need to see whether someone can apply and defend an idea — not just recall it.

Honest scope: this is a working, validity-minded assessment design, not a controlled efficacy study. The claim is the design and the build — not a proven learning-gain number.