IC² Institute, UT Austin · 2025–present
How might we design AI mental-health tools that patients and clinicians are actually willing to trust and use?
994 patients and 204 providers reacted to four AI tool concepts. Modelling the two groups separately showed they are not one market — and that the difference has direct design and rollout consequences.
- Role
- Graduate Research Assistant — study design, analysis and translation to product recommendations
- Timeline
- Sep 2025 – present
- Team
- IC² Institute research group at UT Austin
- Product area
- AI tools for mental healthcare
- Platform
- Concept evaluation across patient-facing and provider-facing tools
- Key skills
- Study designQualtrics deploymentMultigroup structural equation modelling (R, lavaan)Insight translationPower BI reporting
Case study in 30 seconds
Problem
AI mental-health tools keep being designed as if patients and clinicians want the same thing from them. Adoption suggests otherwise, but not in a way anyone had quantified for these tool types.
My role
I ran the study end to end — deployment, interviews, analysis and reporting — and turned the results into design and rollout recommendations.
Key product decision
Run multigroup structural equation modelling to compare patients and providers directly, instead of pooling 1,198 responses into a single model.
Outcome
Providers weight clinical-workflow fit over raw usefulness; trust, fairness and explainability reach adoption through perceived usefulness. Two JMIR manuscripts in progress.
Context
Mental healthcare has a supply problem, and AI tools are a common proposed answer. The gap is not capability — it is whether the people on either side of a clinical relationship will use them.
The study put four AI tool concepts in front of 994 patients and 204 providers and asked them to score each on usefulness, trust, fairness and explainability, alongside their willingness to adopt.
The interesting question was never 'do people like AI'. It was whether the two groups arrive at adoption through the same reasoning — because if they do not, a single product strategy is wrong for at least one of them.
994
Patients
Adoption runs through perceived usefulness, shaped by trust and fairness
204
Providers
Clinical-workflow fit carries weight beyond raw usefulness
Four AI tool concepts, scored on the same four constructs, compared across groups with multigroup structural equation modelling.
The problem
User problem
A patient weighs an AI mental-health tool on whether it is safe and fair to them personally. A clinician weighs it on whether it survives contact with a working day that is already full. Both get shown the same pitch.
Business problem
Teams building these tools are allocating design and rollout effort against an assumption about what drives adoption. If the assumption is wrong for providers, the tool fails at the point of deployment, after the money is spent.
Product challenge
The tension is between one product and two adoption logics. Building separately for each group is expensive; building for the average of the two is a product neither group asked for.
Discovery & insights
The design of the study was itself the product decision: measure the constructs that plausibly drive adoption, then test whether they behave the same way for both groups.
- Survey instrument deployed via Qualtrics to 1,198 participants
- Four distinct AI tool concepts scored on usefulness, trust, fairness and explainability
- Interviews alongside the quantitative instrument
- Multigroup structural equation modelling in R (lavaan)
- Power BI reporting for the research group
We expected
We expected usefulness to be the dominant driver of adoption for both groups, with trust as a secondary gate.
The evidence showed
For providers, clinical-workflow fit carried weight beyond raw usefulness — a tool can be clearly useful and still not be adoptable inside their day.
So we changed
The recommendation for provider-facing tools became a workflow question first and a capability question second.
We expected
We expected trust, fairness and explainability to act as independent adoption factors.
The evidence showed
They reached adoption largely through perceived usefulness — they shape whether a tool reads as useful at all, rather than sitting alongside usefulness.
So we changed
Explainability stopped being a compliance feature to bolt on and became part of how the core value is communicated in the interface.
We expected
We expected one adoption story with two audiences.
The evidence showed
The multigroup comparison showed different structural relationships between the groups, not just different average scores.
So we changed
Design and rollout recommendations were written separately for the two groups rather than as one plan with two audiences.
Product decisions
Each decision below is recorded the way I'd record it for a team: what we chose, the signal behind it, what we gave up, and what we were optimising for.
Decision 01
Model patients and providers as separate groups using multigroup structural equation modelling, rather than pooling all 1,198 responses.
Signal
Early descriptive patterns suggested the two groups were not just scoring differently but reasoning differently about the same concepts.
Alternative considered
A single pooled model with group as a covariate — simpler, faster, and adequate for a headline finding.
Tradeoff we accepted
Much heavier analysis, a harder result to explain in one sentence, and a smaller provider sample to carry its own model.
Why it mattered
The pooled model would have produced an average that describes nobody. The comparison is the finding.
Decision 02
Treat trust, fairness and explainability as constructs acting through perceived usefulness rather than as parallel adoption drivers.
Signal
The measurement model showed their effect on adoption running via usefulness rather than directly.
Alternative considered
Report them as an independent checklist of qualities, which is how they are usually presented to product teams.
Tradeoff we accepted
A less quotable list, and a structure that takes longer to explain to a non-research audience.
Why it mattered
It changes where a product team spends effort: explainability work has to make the tool feel useful, not just make it auditable.
Decision 03
Write the output as design and rollout recommendations rather than as findings.
Signal
Research on AI adoption is abundant; the missing artefact is the translation into what a team should build and in what order.
Alternative considered
Publish the analysis and let product teams draw their own implications.
Tradeoff we accepted
Recommendations commit to interpretation, which is a stronger claim than the data alone makes.
Why it mattered
A finding that nobody can act on has no product value. The two manuscripts carry the rigour; the recommendations carry the use.
Execution
Each insight was carried through to a specific implication and a specific recommendation, so the output was usable by a product team without a statistics background.
Study deployment
Coordinated Qualtrics deployment and interviews end to end across 1,198 participants and four tool concepts.
Analysis
Multigroup structural equation modelling in R (lavaan) to compare the structural relationships between patients and providers.
Reporting
Power BI reporting for the research group so results could be interrogated rather than only read.
Translation
Insight → product implication → recommendation for each finding, split by user group.
Publication
Two JMIR manuscripts in progress covering the patient and provider findings.
Insight
Providers weight clinical-workflow fit over raw usefulness
Implication
A capable tool can still be unadoptable inside a full clinical day
Recommendation
Design provider tools around one workflow; roll out by workflow, not feature
Insight
Trust, fairness and explainability act through perceived usefulness
Implication
Explainability is not a separate compliance surface
Recommendation
Make the explanation part of how the value is communicated in the UI
Insight
The two groups differ structurally, not just in average scores
Implication
One adoption strategy will underserve at least one group
Recommendation
Write separate design and rollout plans for patient and provider surfaces
How the work happened
Aligned on an instrument that would support a multigroup comparison, which constrained the question design well before data collection.
Provider interviews explained the workflow-fit result that the numbers could only point at.
Manuscript review pushed the analysis toward claims the data could carry, which sharpened the recommendations.
Outcome
994
Patients
Scored four AI tool concepts
204
Providers
Compared as a distinct group
4
AI concepts evaluated
Usefulness, trust, fairness, explainability
2
JMIR manuscripts
In progress from these findings
The usable result for a product team: for provider-facing AI in mental health, workflow fit is a first-order design constraint, and explainability earns its place by making the tool legible as useful — not by satisfying a governance checklist.
What I'd do differently
What worked
Committing to the multigroup comparison. The harder analysis is the only reason there is a finding worth acting on.
What didn't
The provider sample is a fifth the size of the patient sample, which limits how far the provider model can be pushed. That constraint should have shaped recruitment earlier.
What I learned
Research earns its keep at the translation step. The modelling was the hard part; the recommendations are the part a product team can use.
What I'd test next
I would test the workflow-fit result directly — put a provider-facing concept inside a real workflow and measure sustained use, rather than stated willingness to adopt.
Like how I think?
I'm currently exploring full-time Product Management opportunities.
Next case study