Skip to content
All work

Graduate team project · 2025–2026

As AI gets better at giving us answers, how might we design it to help people think?

A 15-criterion framework for evaluating AI tutors against Bloom's Taxonomy and Microsoft's Human-AI Interaction guidelines, turned into five product features and a working demo.

Role
Product and research contributor within a five-person team
Timeline
Graduate project, 2025–2026
Team
Five-person team
Product area
AI-assisted learning
Platform
Web demo
Key skills
AI product thinkingEvaluation framework designHuman-AI interaction guidelinesUX researchPrototyping
00

Case study in 30 seconds

Problem

AI tutors are judged on whether the answer is right. Learning depends on whether the learner did any thinking, which is a different property entirely.

My role

I contributed to the evaluation framework, the comparison across assistants, and the translation of results into product features.

Key product decision

Score assistants on the cognitive level they provoke — using Bloom's Taxonomy — rather than on answer accuracy.

Outcome

A 15-criterion framework, four assistants compared, five features delivered, and a UX Design Awards 2026 New Talent nomination for the team project.

Read the full story
01

Context

The default interaction of an AI assistant is answer delivery. For most tasks that is exactly right. For learning it is a design failure with excellent user satisfaction scores.

Bloom's Taxonomy describes cognitive work as a ladder — remember, understand, apply, analyse, evaluate, create. An assistant that answers well collapses the ladder for the learner: the model climbs it, the student receives the summit.

So the project asked a product question rather than a model question: if you evaluated AI assistants on the cognitive level they provoke in the user, which one wins, and what would you build differently?

Create
Evaluate
Analyse
Apply
Understand
Remember

An assistant that answers well climbs the ladder itself and hands over the summit. The design question is which rungs the learner keeps.

Conceptual diagram · Where answer-first AI leaves the learner
02

The problem

User problem

A learner gets a good answer and no practice. The tool feels helpful and the understanding does not transfer to the next problem.

Business problem

Educational AI products are differentiating on model quality, which is converging and expensive. If the differentiator is really interaction design, that is a different investment.

Product challenge

The tension is between satisfaction and learning. The interaction that scores best in the moment — instant, complete, confident — is often the one that does the least for the person.

03

Discovery & insights

Rather than starting from features, the team started from a measurement problem: what would it even mean for an AI tutor to be good?

  • Bloom's Taxonomy as the cognitive-level scaffold
  • Microsoft Human-AI Interaction guidelines as the interaction-quality scaffold
  • A 15-criterion scoring framework built from both
  • Comparative evaluation of four AI assistants
  • A working demo of the resulting interaction model

We expected

We expected the strongest model to produce the strongest tutoring experience.

The evidence showed

Across the criteria, differences in how an assistant structured the exchange mattered as much as differences in raw capability.

So we changed

The product bet moved from model selection to interaction design.

We expected

We expected correctness to be the hard thing to evaluate.

The evidence showed

Correctness was the easy axis. The hard axis was whether the exchange left any cognitive work with the learner.

So we changed

The framework was built around cognitive level and interaction quality, with correctness as a floor rather than the score.

We expected

We expected human-AI guidelines to be a compliance layer over the design.

The evidence showed

Applied early, they generated feature ideas — making uncertainty visible and scoping what the assistant will and won't do turned out to be teaching behaviours.

So we changed

Responsible-AI principles were used as a source of product features rather than as a review checklist.

04

Product decisions

Each decision below is recorded the way I'd record it for a team: what we chose, the signal behind it, what we gave up, and what we were optimising for.

Decision 01

Evaluate assistants on the cognitive level they provoke, mapped to Bloom's Taxonomy, rather than on answer accuracy.

Signal

Assistants with near-identical answer quality produced very different amounts of learner participation in the exchange.

Alternative considered

A conventional accuracy-and-helpfulness rubric, which is faster to run and easier to defend.

Tradeoff we accepted

A more subjective framework that needs careful criteria definition and inter-rater discipline across five people.

Why it mattered

It is the only framing under which the product question — how should this thing behave — has a measurable answer.

Decision 02

Build a 15-criterion scoring framework combining Bloom's levels with Microsoft's Human-AI Interaction guidelines.

Signal

Neither scaffold alone was sufficient: Bloom's says what learning looks like, the HAI guidelines say what a trustworthy exchange looks like.

Alternative considered

Adopt one existing rubric wholesale and accept its blind spots.

Tradeoff we accepted

Fifteen criteria is heavy to apply consistently, and it slowed the comparison substantially.

Why it mattered

The combined framework is what let the team argue for specific features rather than general improvements.

Decision 03

Turn the evaluation results into five concrete product features rather than a report.

Signal

The scoring gaps clustered around a small number of recurring interaction behaviours.

Alternative considered

Deliver the framework and comparison as the outcome — a legitimate research contribution on its own.

Tradeoff we accepted

Building took time away from broadening the evaluation to more assistants.

Why it mattered

A framework nobody builds against is an opinion. The demo made the argument checkable.

05

Execution

The framework drove the build: each feature exists because a scoring gap justified it.

15-criterion scoring framework

Cognitive level (Bloom's) crossed with interaction quality (Microsoft HAI guidelines), defined criterion by criterion for consistent scoring across five reviewers.

Comparative evaluation

Four AI assistants scored against the same framework on the same tasks.

Five delivered features

Interaction behaviours designed to keep cognitive work with the learner rather than resolving it for them.

Working demo

A deployed prototype so the interaction model could be experienced rather than described.

Scaffold A · Bloom's Taxonomy

What cognitive level does the exchange leave with the learner?

RememberUnderstandApplyAnalyseEvaluateCreate

Scaffold B · Microsoft HAI guidelines

Is the exchange itself trustworthy, scoped and legible?

Clarity of scopeUncertaintyCorrectionContextControl

15

criteria defined precisely enough that five reviewers score the same exchange the same way — then applied to four assistants on identical tasks.

Conceptual diagram · Cognitive level × interaction quality
06

How the work happened

Team of five

Criteria had to be defined precisely enough that five people scored the same exchange the same way — that definition work was most of the framework.

Design

The cognitive-level findings turned directly into interaction patterns, which is where the awards nomination came from.

Engineering

The demo constrained the feature set: we shipped the behaviours that could be made reliable, not every behaviour the framework implied.

07

Outcome

  • 15

    Evaluation criteria

    Bloom's × Microsoft HAI guidelines

  • 4

    AI assistants compared

    Same tasks, same framework

  • 5

    Features delivered

    Each traced to a scoring gap

  • UXDA '26

    New Talent nomination

    Awarded to the team project

The nomination belongs to the team project rather than to any individual. The durable output is the argument it supports: educational AI quality depends less on the model than on how the experience guides thinking.

08

What I'd do differently

What worked

Building the evaluation framework before the features. It meant every feature had a reason that survived a design review.

What didn't

Fifteen criteria was more than the team could apply quickly. A tiered framework — a fast screen plus a deep score — would have let us compare more assistants.

What I learned

Responsible-AI guidelines are better as a design input than as a review gate. Used early, they produce features; used late, they produce edits.

What I'd test next

I would test the features with real learners over time and measure retention and transfer, not in-session satisfaction, which is the metric most likely to reward the wrong behaviour.

Like how I think?

I'm currently exploring full-time Product Management opportunities.

Next case study