Graduate team project · 2025–2026
As AI gets better at giving us answers, how might we design it to help people think?
A 15-criterion framework for evaluating AI tutors against Bloom's Taxonomy and Microsoft's Human-AI Interaction guidelines, turned into five product features and a working demo.
- Role
- Product and research contributor within a five-person team
- Timeline
- Graduate project, 2025–2026
- Team
- Five-person team
- Product area
- AI-assisted learning
- Platform
- Web demo
- Key skills
- AI product thinkingEvaluation framework designHuman-AI interaction guidelinesUX researchPrototyping
Case study in 30 seconds
Problem
AI tutors are judged on whether the answer is right. Learning depends on whether the learner did any thinking, which is a different property entirely.
My role
I contributed to the evaluation framework, the comparison across assistants, and the translation of results into product features.
Key product decision
Score assistants on the cognitive level they provoke — using Bloom's Taxonomy — rather than on answer accuracy.
Outcome
A 15-criterion framework, four assistants compared, five features delivered, and a UX Design Awards 2026 New Talent nomination for the team project.
Context
The default interaction of an AI assistant is answer delivery. For most tasks that is exactly right. For learning it is a design failure with excellent user satisfaction scores.
Bloom's Taxonomy describes cognitive work as a ladder — remember, understand, apply, analyse, evaluate, create. An assistant that answers well collapses the ladder for the learner: the model climbs it, the student receives the summit.
So the project asked a product question rather than a model question: if you evaluated AI assistants on the cognitive level they provoke in the user, which one wins, and what would you build differently?
An assistant that answers well climbs the ladder itself and hands over the summit. The design question is which rungs the learner keeps.
The problem
User problem
A learner gets a good answer and no practice. The tool feels helpful and the understanding does not transfer to the next problem.
Business problem
Educational AI products are differentiating on model quality, which is converging and expensive. If the differentiator is really interaction design, that is a different investment.
Product challenge
The tension is between satisfaction and learning. The interaction that scores best in the moment — instant, complete, confident — is often the one that does the least for the person.
Discovery & insights
Rather than starting from features, the team started from a measurement problem: what would it even mean for an AI tutor to be good?
- Bloom's Taxonomy as the cognitive-level scaffold
- Microsoft Human-AI Interaction guidelines as the interaction-quality scaffold
- A 15-criterion scoring framework built from both
- Comparative evaluation of four AI assistants
- A working demo of the resulting interaction model
We expected
We expected the strongest model to produce the strongest tutoring experience.
The evidence showed
Across the criteria, differences in how an assistant structured the exchange mattered as much as differences in raw capability.
So we changed
The product bet moved from model selection to interaction design.
We expected
We expected correctness to be the hard thing to evaluate.
The evidence showed
Correctness was the easy axis. The hard axis was whether the exchange left any cognitive work with the learner.
So we changed
The framework was built around cognitive level and interaction quality, with correctness as a floor rather than the score.
We expected
We expected human-AI guidelines to be a compliance layer over the design.
The evidence showed
Applied early, they generated feature ideas — making uncertainty visible and scoping what the assistant will and won't do turned out to be teaching behaviours.
So we changed
Responsible-AI principles were used as a source of product features rather than as a review checklist.
Product decisions
Each decision below is recorded the way I'd record it for a team: what we chose, the signal behind it, what we gave up, and what we were optimising for.
Decision 01
Evaluate assistants on the cognitive level they provoke, mapped to Bloom's Taxonomy, rather than on answer accuracy.
Signal
Assistants with near-identical answer quality produced very different amounts of learner participation in the exchange.
Alternative considered
A conventional accuracy-and-helpfulness rubric, which is faster to run and easier to defend.
Tradeoff we accepted
A more subjective framework that needs careful criteria definition and inter-rater discipline across five people.
Why it mattered
It is the only framing under which the product question — how should this thing behave — has a measurable answer.
Decision 02
Build a 15-criterion scoring framework combining Bloom's levels with Microsoft's Human-AI Interaction guidelines.
Signal
Neither scaffold alone was sufficient: Bloom's says what learning looks like, the HAI guidelines say what a trustworthy exchange looks like.
Alternative considered
Adopt one existing rubric wholesale and accept its blind spots.
Tradeoff we accepted
Fifteen criteria is heavy to apply consistently, and it slowed the comparison substantially.
Why it mattered
The combined framework is what let the team argue for specific features rather than general improvements.
Decision 03
Turn the evaluation results into five concrete product features rather than a report.
Signal
The scoring gaps clustered around a small number of recurring interaction behaviours.
Alternative considered
Deliver the framework and comparison as the outcome — a legitimate research contribution on its own.
Tradeoff we accepted
Building took time away from broadening the evaluation to more assistants.
Why it mattered
A framework nobody builds against is an opinion. The demo made the argument checkable.
Execution
The framework drove the build: each feature exists because a scoring gap justified it.
15-criterion scoring framework
Cognitive level (Bloom's) crossed with interaction quality (Microsoft HAI guidelines), defined criterion by criterion for consistent scoring across five reviewers.
Comparative evaluation
Four AI assistants scored against the same framework on the same tasks.
Five delivered features
Interaction behaviours designed to keep cognitive work with the learner rather than resolving it for them.
Working demo
A deployed prototype so the interaction model could be experienced rather than described.
Scaffold A · Bloom's Taxonomy
What cognitive level does the exchange leave with the learner?
Scaffold B · Microsoft HAI guidelines
Is the exchange itself trustworthy, scoped and legible?
15
criteria defined precisely enough that five reviewers score the same exchange the same way — then applied to four assistants on identical tasks.
How the work happened
Criteria had to be defined precisely enough that five people scored the same exchange the same way — that definition work was most of the framework.
The cognitive-level findings turned directly into interaction patterns, which is where the awards nomination came from.
The demo constrained the feature set: we shipped the behaviours that could be made reliable, not every behaviour the framework implied.
Outcome
15
Evaluation criteria
Bloom's × Microsoft HAI guidelines
4
AI assistants compared
Same tasks, same framework
5
Features delivered
Each traced to a scoring gap
UXDA '26
New Talent nomination
Awarded to the team project
The nomination belongs to the team project rather than to any individual. The durable output is the argument it supports: educational AI quality depends less on the model than on how the experience guides thinking.
What I'd do differently
What worked
Building the evaluation framework before the features. It meant every feature had a reason that survived a design review.
What didn't
Fifteen criteria was more than the team could apply quickly. A tiered framework — a fast screen plus a deep score — would have let us compare more assistants.
What I learned
Responsible-AI guidelines are better as a design input than as a review gate. Used early, they produce features; used late, they produce edits.
What I'd test next
I would test the features with real learners over time and measure retention and transfer, not in-session satisfaction, which is the metric most likely to reward the wrong behaviour.
Like how I think?
I'm currently exploring full-time Product Management opportunities.
Next case study