Skip to content
All work

IC² Institute, UT Austin · 2025–present

How might we design AI mental-health tools that patients and clinicians are actually willing to trust and use?

994 patients and 204 providers reacted to four AI tool concepts. Modelling the two groups separately showed they are not one market — and that the difference has direct design and rollout consequences.

Role
Graduate Research Assistant — study design, analysis and translation to product recommendations
Timeline
Sep 2025 – present
Team
IC² Institute research group at UT Austin
Product area
AI tools for mental healthcare
Platform
Concept evaluation across patient-facing and provider-facing tools
Key skills
Study designQualtrics deploymentMultigroup structural equation modelling (R, lavaan)Insight translationPower BI reporting
00

Case study in 30 seconds

Problem

AI mental-health tools keep being designed as if patients and clinicians want the same thing from them. Adoption suggests otherwise, but not in a way anyone had quantified for these tool types.

My role

I ran the study end to end — deployment, interviews, analysis and reporting — and turned the results into design and rollout recommendations.

Key product decision

Run multigroup structural equation modelling to compare patients and providers directly, instead of pooling 1,198 responses into a single model.

Outcome

Providers weight clinical-workflow fit over raw usefulness; trust, fairness and explainability reach adoption through perceived usefulness. Two JMIR manuscripts in progress.

Read the full story
01

Context

Mental healthcare has a supply problem, and AI tools are a common proposed answer. The gap is not capability — it is whether the people on either side of a clinical relationship will use them.

The study put four AI tool concepts in front of 994 patients and 204 providers and asked them to score each on usefulness, trust, fairness and explainability, alongside their willingness to adopt.

The interesting question was never 'do people like AI'. It was whether the two groups arrive at adoption through the same reasoning — because if they do not, a single product strategy is wrong for at least one of them.

994

Patients

Adoption runs through perceived usefulness, shaped by trust and fairness

204

Providers

Clinical-workflow fit carries weight beyond raw usefulness

Usefulness
Trust
Fairness
Explainability

Four AI tool concepts, scored on the same four constructs, compared across groups with multigroup structural equation modelling.

Conceptual diagram · Two groups, two adoption logics
02

The problem

User problem

A patient weighs an AI mental-health tool on whether it is safe and fair to them personally. A clinician weighs it on whether it survives contact with a working day that is already full. Both get shown the same pitch.

Business problem

Teams building these tools are allocating design and rollout effort against an assumption about what drives adoption. If the assumption is wrong for providers, the tool fails at the point of deployment, after the money is spent.

Product challenge

The tension is between one product and two adoption logics. Building separately for each group is expensive; building for the average of the two is a product neither group asked for.

03

Discovery & insights

The design of the study was itself the product decision: measure the constructs that plausibly drive adoption, then test whether they behave the same way for both groups.

  • Survey instrument deployed via Qualtrics to 1,198 participants
  • Four distinct AI tool concepts scored on usefulness, trust, fairness and explainability
  • Interviews alongside the quantitative instrument
  • Multigroup structural equation modelling in R (lavaan)
  • Power BI reporting for the research group

We expected

We expected usefulness to be the dominant driver of adoption for both groups, with trust as a secondary gate.

The evidence showed

For providers, clinical-workflow fit carried weight beyond raw usefulness — a tool can be clearly useful and still not be adoptable inside their day.

So we changed

The recommendation for provider-facing tools became a workflow question first and a capability question second.

We expected

We expected trust, fairness and explainability to act as independent adoption factors.

The evidence showed

They reached adoption largely through perceived usefulness — they shape whether a tool reads as useful at all, rather than sitting alongside usefulness.

So we changed

Explainability stopped being a compliance feature to bolt on and became part of how the core value is communicated in the interface.

We expected

We expected one adoption story with two audiences.

The evidence showed

The multigroup comparison showed different structural relationships between the groups, not just different average scores.

So we changed

Design and rollout recommendations were written separately for the two groups rather than as one plan with two audiences.

04

Product decisions

Each decision below is recorded the way I'd record it for a team: what we chose, the signal behind it, what we gave up, and what we were optimising for.

Decision 01

Model patients and providers as separate groups using multigroup structural equation modelling, rather than pooling all 1,198 responses.

Signal

Early descriptive patterns suggested the two groups were not just scoring differently but reasoning differently about the same concepts.

Alternative considered

A single pooled model with group as a covariate — simpler, faster, and adequate for a headline finding.

Tradeoff we accepted

Much heavier analysis, a harder result to explain in one sentence, and a smaller provider sample to carry its own model.

Why it mattered

The pooled model would have produced an average that describes nobody. The comparison is the finding.

Decision 02

Treat trust, fairness and explainability as constructs acting through perceived usefulness rather than as parallel adoption drivers.

Signal

The measurement model showed their effect on adoption running via usefulness rather than directly.

Alternative considered

Report them as an independent checklist of qualities, which is how they are usually presented to product teams.

Tradeoff we accepted

A less quotable list, and a structure that takes longer to explain to a non-research audience.

Why it mattered

It changes where a product team spends effort: explainability work has to make the tool feel useful, not just make it auditable.

Decision 03

Write the output as design and rollout recommendations rather than as findings.

Signal

Research on AI adoption is abundant; the missing artefact is the translation into what a team should build and in what order.

Alternative considered

Publish the analysis and let product teams draw their own implications.

Tradeoff we accepted

Recommendations commit to interpretation, which is a stronger claim than the data alone makes.

Why it mattered

A finding that nobody can act on has no product value. The two manuscripts carry the rigour; the recommendations carry the use.

05

Execution

Each insight was carried through to a specific implication and a specific recommendation, so the output was usable by a product team without a statistics background.

Study deployment

Coordinated Qualtrics deployment and interviews end to end across 1,198 participants and four tool concepts.

Analysis

Multigroup structural equation modelling in R (lavaan) to compare the structural relationships between patients and providers.

Reporting

Power BI reporting for the research group so results could be interrogated rather than only read.

Translation

Insight → product implication → recommendation for each finding, split by user group.

Publication

Two JMIR manuscripts in progress covering the patient and provider findings.

Insight

Providers weight clinical-workflow fit over raw usefulness

Implication

A capable tool can still be unadoptable inside a full clinical day

Recommendation

Design provider tools around one workflow; roll out by workflow, not feature

Insight

Trust, fairness and explainability act through perceived usefulness

Implication

Explainability is not a separate compliance surface

Recommendation

Make the explanation part of how the value is communicated in the UI

Insight

The two groups differ structurally, not just in average scores

Implication

One adoption strategy will underserve at least one group

Recommendation

Write separate design and rollout plans for patient and provider surfaces

Conceptual diagram · Insight → product implication → recommendation
06

How the work happened

Research group

Aligned on an instrument that would support a multigroup comparison, which constrained the question design well before data collection.

Clinical participants

Provider interviews explained the workflow-fit result that the numbers could only point at.

Faculty & reviewers

Manuscript review pushed the analysis toward claims the data could carry, which sharpened the recommendations.

07

Outcome

  • 994

    Patients

    Scored four AI tool concepts

  • 204

    Providers

    Compared as a distinct group

  • 4

    AI concepts evaluated

    Usefulness, trust, fairness, explainability

  • 2

    JMIR manuscripts

    In progress from these findings

The usable result for a product team: for provider-facing AI in mental health, workflow fit is a first-order design constraint, and explainability earns its place by making the tool legible as useful — not by satisfying a governance checklist.

08

What I'd do differently

What worked

Committing to the multigroup comparison. The harder analysis is the only reason there is a finding worth acting on.

What didn't

The provider sample is a fifth the size of the patient sample, which limits how far the provider model can be pushed. That constraint should have shaped recruitment earlier.

What I learned

Research earns its keep at the translation step. The modelling was the hard part; the recommendations are the part a product team can use.

What I'd test next

I would test the workflow-fit result directly — put a provider-facing concept inside a real workflow and measure sustained use, rather than stated willingness to adopt.

Like how I think?

I'm currently exploring full-time Product Management opportunities.

Next case study