Customers decide what gets used.The business decides what gets built.I make those the same decision.

Four years at EY-Parthenon taught me where to compete: which investments pay off and which segments to win. UX research taught me what makes people choose a product and keep using it. I bring both, so teams get to the right product with fewer wrong turns.

Google CloudUX Research, 2026
EY-Parthenon4 years, growth and go-to-market strategy
OptimizelyUsability research, 2026
University of WashingtonMS, Human Centered Design & Engineering

How I work

Start with the decision, at the fidelity it needs

Before choosing a method, I pin down what the team will decide differently based on the answer. Then I move as fast as that decision requires, building prototypes or AI-assisted tools when that's the quickest way to learn.

See: Google

Count commitments, not compliments

People are generous with researchers. They praise prototypes and promise future habits. I ask about what they've actually done and look for what they'll commit to.

See: EY-Parthenon

Hold desirable, viable, and feasible together

A finding matters when it changes whether people want something, whether the business can sustain it, or whether the team can build it. Four years in strategy consulting trained the second of those.

See: Optimizely, UW

Tell the whole story, cracks included

I credit the team widely, own my mistakes, and report what didn't work alongside what did.

See: EY-Parthenon, UW

I started on the business side of the table.

I spent four years at EY-Parthenon on growth and go-to-market strategy for enterprise tech, healthcare, and consumer companies, presenting to C-suite executives shaping three-to-five-year plans. The work I kept gravitating toward was the research: sitting with patients, sellers, and operators to understand why a plan would or wouldn't work for the people living it.

So I made the switch. I'm finishing an MS in Human Centered Design & Engineering at the University of Washington, and in 2026 I joined Google Cloud as a UX Research Intern, studying when enterprise sellers will trust an AI agent with their pipeline.

What ties my work together is usefulness. People are generous with researchers: they praise prototypes, agree to follow-ups they won't finish, and describe the version of themselves they hope to be. So I look for what people actually do and commit to, and I report the cracks alongside the wins.

Education
MS, Human Centered Design & Engineering, University of Washington (expected June 2027)
BS, Business Administration, UC Berkeley Haas (2021)
Based in
Seattle

Experience

  • UX Research Intern, Google CloudJun to Sep 2026

    Led research on when enterprise sellers will trust an AI agent; the findings set the product's core design rule and shaped org-wide guidelines for AI agents.

  • Researcher, UW Directed Research GroupSpring 2026

    Coded 12 Black-led maternal health organizations and extended a community-care framework into a way to assess programs.

  • MS, Human Centered Design & EngineeringUniversity of Washington, expected Jun 2027

    Graduate research and coursework in mixed methods, usability, and AI prototyping.

  • EY-ParthenonSep 2021 to Sep 2025

    Led user research and growth strategy across 20+ engagements in tech, healthcare, and consumer, presenting to C-suite leaders. Promoted twice.

    1. 2025Consultant, post-MBA level
    2. 2024Market Empathy Specialist, research rotation; ran evaluative research for internal AI products, including the validation study that ended GTM Gener(ai)tor
    3. 2023Senior Associate
    4. 2021Associate
  • BS, Business AdministrationUC Berkeley Haas, 2021

More work

  • Marketplace ventureEY-Parthenon

    Shaped the value proposition and launch segments for a B2B procurement marketplace, working with the client's executive team.

  • Patient access researchEY-Parthenon

    45 interviews and 150+ surveys informing a care coordination roadmap for a foundation serving 100,000+ patients.

  • Responsible AI workshopsEY-Parthenon

    Co-led governance workshops with legal, engineering, and strategy leaders.

  • Prior authorization discoveryUniversity of Washington

    Interviews with about 20 stakeholders that shifted the product strategy from Epic toward smaller EHR systems.

All work

Sellers will hand almost everything to an AI agent, except what they'll be blamed for

Role
UX Research Intern, Google Cloud
Team
PM, UX Design, Engineering, Data Science, Design Systems
Timeline
June to September 2026, 14 weeks
Methods
Literature synthesis, concept evaluations, diary study, survey, card sort, telemetry planning

The short version

Google Cloud was building an AI agent to keep enterprise sales records up to date. The team assumed sellers neglected their records because data entry was slow. I found the real barrier was blame: if an agent changes a forecast number and it's wrong, the seller answers for it in front of leadership. Sellers would delegate hundreds of administrative fields but refused to hand over the handful that leadership scrutinizes. That line became the product's core design rule, two planned features were cut before engineering built them, and my trust principles became a chapter of the organization's design guidelines for AI agents.

The red line. Of hundreds of fields on a sales record, sellers wanted to approve only a handful themselves. Everything else could update automatically. Illustrative, not to scale.

I reframed a layout question into an adoption question

I was asked where automated updates should appear, and in how much detail. I reframed it to the question underneath: what has to be true for a seller to let an AI agent touch numbers their pay and reputation depend on?

Answering the first question would have produced a layout. Answering the second decided whether anyone would use the product.

Saved time is worthless if the forecast breaks

Sellers spent hours every week updating sales records by hand, time taken from selling. Those same records feed the revenue forecasts leadership reviews every week. An agent that saved time but corrupted forecasts would cost more than it saved.

My consulting background mattered here. I'd sat on the leadership side of forecast reviews, so I treated sales incentives and accountability as design inputs, not background.

Synthesis first, then testing against real work

A trust framework from 27 studies

I reviewed 27 studies from internal Google research and academic HCI and distilled them into six trust principles spanning the agent's full lifecycle.

Setup

Calibrate expectations before the agent acts.

Keep high-stakes changes as drafts.

Everyday use

Interrupt only when stakes are high.

Short summaries by default, full logs on demand.

Recovery

Say plainly what the agent doesn't know.

Undo one change, not the whole batch.

13 sellers, realistic prototypes, planted errors

I ran 13 moderated 60-minute concept evaluations, plus 3 follow-ups, across three seller roles, from high-volume deal owners to directors of large strategic accounts. With UX Design, I built role-specific prototypes that looked like sellers' real work, because experienced sellers dismiss anything that looks fake. The prototypes included deliberately planted AI errors, so I could watch how sellers caught or missed mistakes.

Triangulated, with a plan for scale

I paired the sessions with a forecasting diary study (N=14) and a field survey. Thirteen sessions can't set thresholds at scale, so I scoped a telemetry analysis with Data Science to test the findings across every field on the record.

What we found

Anxiety concentrated where blame lands

Sellers didn't worry about most of the record. Their concern sat on the few forecast-critical fields that leadership reviews, the numbers a seller personally answers for.

Decision: those fields are staged as drafts the seller approves, each with a short explanation of why the agent proposed the change. Everything else updates automatically with one-click undo.

A generic illustration of the pattern, not the product. High-stakes changes wait for approval with their reason attached; routine fields update on their own.

Projected: the large majority of fields no longer need manual review.

Two planned features would have added work

Individual risk-tolerance settings felt like another admin chore, and sellers feared being blamed for misconfiguring them. They wanted manager-set team defaults instead. Typing a request to a chat assistant to change one number took about five times longer than clicking it and editing.

Decision: both features were cut before engineering built them, replaced with team presets and inline editing.

The system punished partial updates

Deal records and their linked technical records failed validation when updated separately, pushing sellers to enter placeholder data just to save.

Decision: updating linked records in a single pass became a launch requirement.

The work outlived the project

  • Org-wide guidelines. I co-authored the trust chapter of Google Cloud's design guidelines for AI agents, turning the six principles into do and don't rules other agent teams now build from. It received directional approval from UX leadership.
  • Design system. Side-by-side change previews were elevated on the component roadmap, and the default sort changed to deal value, matching how every seller in the study triaged their pipeline.
  • Roadmap. I prioritized 17 features into 6 launch blockers and 11 follow-ons with PM and Engineering.
  • Business case. Projected: thousands of selling hours a week returned across the sales organization.
  • Continuity. Mid-project, the entire PM, design, and engineering trio was reassigned. I onboarded their replacements and finished every session on schedule.

I built an AI-assisted analysis pipeline with guardrails

Fast-moving studies fail in two predictable ways: stakeholders seize on one dramatic quote, and the first participants anchor the themes. I split analysis into two tracks. One produced same-day stakeholder digests that flagged any single-participant claim. The other ran slower, coded only anonymized data, and actively searched for evidence against emerging hypotheses.

It cut synthesis time by about 45% while keeping every finding traceable to its source.

Limits and next steps

Thirteen sessions explain why sellers trust an agent or don't. They can't set exact thresholds; the telemetry analysis is designed to test the red line at scale. Managers were missing from this study even though sellers asked for manager-set defaults, so I wrote the business case for a follow-up study with sales managers.

Projected figures are modeled from study data, not measured outcomes. Detailed findings are confidential; I'm happy to walk through the work in conversation.

All work

Marketed to marketers, built for developers

Role
Research lead, team of 4
Client
Optimizely, via UW HCDE 517
Timeline
Winter 2026
Methods
Moderated usability testing, think-aloud, SEQ benchmarking

The short version

Optimizely markets Opal Agent Builder as a no-code way for marketing teams to build their own AI agents. In moderated tests with six marketers, we found the barrier wasn't what the tool could do but how it spoke: it asked marketers to think like developers. Average ease was 4.0 out of 7 against a 5.5 benchmark, and only one of six found the core insertion command unaided. Optimizely has since moved Opal to a visual, drag-and-drop workflow builder, the direction our study recommended.

Opal Agent Builder's original setup screen: one long form with agent details, a prompt template, variables, and tools
The builder we tested: every setting on one page, and nothing validated until you run the agent.

If marketers can't build an agent alone, the product's promise breaks

The pitch to customers is "no code required." If the target buyer needs technical help to build an agent, adoption stalls and the product only reaches teams that already have engineers. So the real question was whether the tool fit its market, not just whether a few buttons were usable.

I led the study end to end

We shared roles as a team of four. My part:

  • Ran every client call and kept Optimizely aligned at each milestone.
  • Recruited participants, mapped the expected path and the failure paths through the product, and wrote the session guide.
  • Facilitated sessions, rotating facilitator and note-taker roles with the team.
  • Set up our documentation system and shaped the story of the final readout.

Six marketers, five real tasks, one ease score per person

We ran six 60-minute moderated remote sessions plus a pilot. Participants ranged from a CMO to marketing specialists, with mixed AI experience and mostly little or no coding background.

Each session moved from background questions through five tasks (create an agent, add a variable, insert it, add a web tool, run a test), then a "magic wand" question and a debrief. We captured task success, the Single Ease Question (SEQ), think-aloud behavior, and quotes.

What we found

The vocabulary belonged to developers

Every participant was confused by terms like "variable," "string," and "boolean." One couldn't complete the variable task at all, and even those who succeeded weren't sure they'd done it right.

What the interface said
What marketers needed
Variable
Custom input
String
Text
Boolean
Yes / No
DateTime
Date

The core mechanic was invisible

The only way to insert a variable was a slash command. One of six participants found it unaided, through a small note at the bottom of the screen. Others hovered, right-clicked, and tried to drag. That's a mental model the product didn't support, and one it could design for.

Opal's prompt template field, with the only instruction a small line of gray text about double-bracket syntax
The prompt field. The only clue to how variables work is the gray line under the box.

Everything at once, with feedback only at the end

Every setting sat on one page, and nothing was checked until participants ran the agent. Facing a blank form with no feedback along the way, they fell back on trial and error. All six measured the experience against tools like Mailchimp, Canva, and Notion, their own benchmarks for "easy."

Five recommendations, from labels to architecture

  1. Translate the language. Name things by what marketers are doing, not how the system stores it.
  2. Make insertion visible. Add an Insert button, drag-and-drop from the variables panel, and a live preview.
  3. Group tools by marketing goal, not by how the system is built.
  4. Start from templates instead of a blank form.
  5. Replace the single page with guided setup, moving raw prompt editing behind an expert mode.

Optimizely moved to a visual workflow builder

Opal now builds agents as a drag-and-drop workflow, the direction our study recommended.

Optimizely's newer Opal workflow canvas: connected steps for a scheduler, page audit, recommendations, blog post generation, and publishing
Opal's current workflow builder. Image: Optimizely.

What I'd do differently

Tighten the screener

Three of six participants had some coding or AI experience, and they still struggled. A stricter screen for marketers with no technical background would give a cleaner baseline and a sharper finding.

Remove a source of bias

The prompt field came pre-filled with bracket syntax, which nudged people to type brackets by hand. A blank field would show what people naturally reach for.

Validate at scale

An unmoderated follow-up with about 50 marketers would show whether the one-in-six discovery rate holds.

All work

Community care is infrastructure

Role
Researcher, UW Directed Research Group
Team
Leslie Coney's research group (PhD candidate), supervised by Dr. Julie Kientz
Timeline
Spring 2026
Methods
Directed content analysis, codebook development, analytic memo

The short version

Leslie Coney's published study of 12 Black mothers introduced TRIBE, a framework for the support community care provides. As part of her research group, I helped extend it. Our group tested whether it also describes how Black-led maternal health organizations design their services. I coded all 12 organizations in the sample, across five states and DC, and proposed the 85 specific codes the analysis ran on. Those codes showed TRIBE held up, and they surfaced what the interviews couldn't: what care costs, who delivers it, and where it breaks. In a memo to the paper's writing team, I used that to argue that community care is infrastructure, and that TRIBE can test whether a program actually provides it.

Support typeWhat families receiveHow the organization works
Tangible215
Relational102
Informational214
Birthworker54
Emotional61
Identity6, describing both
The 85 codes I proposed, split into two layers. The original framework described what mothers received; the second layer let us read how organizations value, fund, staff, and measure that care.

Mothers turned to community when institutions failed them

Black women in the U.S. are more than three times as likely as white women to die from pregnancy-related causes, and most of those deaths are preventable. Mothers in the original study described being dismissed in clinical settings and turning instead to doulas, peer groups, and mutual aid. TRIBE named the support they relied on: Tangible, Relational, Informational, Birthworker, and Emotional, with Identity shaping all five.

TRIBE comes from Leslie Coney's qualitative research program, published as What's the CommuniTEA?: The Role of Community Care in Black Maternal Health (Coney et al., Health Equity, 2026). I joined her research group the following term to help extend that work: our directed research group tested whether the framework holds beyond interviews, in how Black-led organizations design their services.

Funders need to see what a program actually provides

Black-led maternal health organizations run on Medicaid contracts, grants, and donated labor,. A framework that only works on interview transcripts can't help a funder or health system judge a program. One that also reads how organizations describe their own services can.

A directed content analysis, starting from the framework

Directed content analysis (Hsieh & Shannon, 2005) starts from an existing framework and tests whether the data supports, extends, or contradicts it. We sampled the public websites and impact reports of 12 Black-led maternal health organizations in Washington, Ohio, Missouri, New York, North Carolina, and Washington, DC, chosen to vary by region and model.

My part: I coded all 12 organizations, proposed every specific code in the codebook, consolidated duplicates as the codebook grew (four separate doula codes became one), and wrote the analytic memo comparing three organizations for the paper's writing team.

What my codes showed

A second layer turned a framework about mothers into one about organizations

TRIBE described what mothers received. I added codes for how each organization works: its values, resources, outcomes, and barriers. That let us compare what an organization says it values with what it delivers and what it measures, the same gap between strategy and execution I looked for as a consultant.

Care has a cost, and the codebook made it visible

Codes for funding, insurance, policy, fundraising, financial support, and barriers to birthworker and identity-based care captured something the interviews only hinted at. The support mothers depend on is fragile because someone has to pay for it.

Organizations design for families and networks, not just mothers

Codes for fathers, fatherhood, family, volunteers, boards, task forces, partnerships, and referrals showed care moving through networks. One organization runs a fatherhood initiative. Organizations also grow their own workforce, through peer counselor programs, birthworker training, and capability building.

Emotional support was the thinnest layer

Organizations described tangible and informational support in 21 distinct forms each, and emotional support in six. The emotional layer includes loss and storytelling, matching the miscarriage and grief support mothers in the original study said was missing.

Identity showed up as language, culture, and mission

Identity codes covered language access, translation, cultural grounding, and organizational mission, the concrete ways organizations make culturally specific care real rather than a slogan.

From codes to argument: my memo to the writing team

I compared three organizations in depth: BLKBRY in Washington, Shades of Motherhood Network in Spokane, and ROOTT in Columbus, Ohio. The memo proposed five additions to the paper:

  1. Treat community care as infrastructure that needs funding, staffing, and training, not as informal help. The cost and workforce codes are the evidence.
  2. Make identity a design requirement, a condition for safe access rather than a preference.
  3. Distinguish partnership from institutional capture when integrating community care with health systems.
  4. Use TRIBE as a diagnostic, so programs and funders can see where support is thick and where it's thin. The two-layer codebook makes this possible.
  5. Map care across all of Washington, including Eastern Washington, which the original Seattle-Tacoma study didn't reach.

Limits

Public materials show what organizations say they do, not what families experience. A directed approach can also lean toward confirming the framework it starts from; the new codes were our check on that. The next step is to compare organizations' self-descriptions with families' own accounts.

All work

Users said they loved it. My research recommended shutting it down.

Role
UX Researcher, Market Empathy rotation
Organization
EY-Parthenon, Americas Innovation
Timeline
About 3 months, fall 2023
Methods
Stakeholder workshop, hypothesis-driven interviews, remote usability testing, task metrics

The short version

EY built GTM Gener(ai)tor, an AI tool that drafts go-to-market pitch decks, and ran a validation test to decide whether to keep investing. Users praised it: 10 of 12 said they'd strongly recommend it. But the evidence underneath said otherwise. The decks came out about half complete, below the bar users needed before they'd bother editing, and too generic to win clients. I recommended against further investment, and the tool was shut down. It taught me to weigh what people do over what they say.

How complete a first-draft deck was. The gap between the last two bars mattered more than any satisfaction score.

A faster deck only counts if it still wins the client

EY product teams spent about 19 hours building each go-to-market deck, the material used to sell EY services. Leadership needed to know whether the tool earned more investment: could it cut that time without weakening the decks that win work?

Agree on what success means before the first session

I started with a kickoff workshop with the product team to agree on target users, interview style, hypotheses, and success metrics. Then we ran 13 interviews and remote usability sessions with 5 product owners, 4 product team members, and 4 go-to-market leads, tracking task time, errors, accuracy, confidence, and likelihood to recommend.

Over three months I fed findings to the PM and engineers weekly; blocking issues went to engineering the same day.

Two of three hypotheses failed

HypothesisResult
The tool delivers an 80%-complete deckDisproved. About 55%, below the roughly 70% users needed before editing it themselves.
It cuts deck time in halfValidated. A first draft in about 9 minutes, against about 19 hours by hand.
Users understand how their answers become slidesDisproved. Only about 15% completed the persona prompts correctly.

Only 41% of users found the generated content highly accurate, yet the average likelihood to recommend was 9 out of 10, a Net Promoter Score of +75.

I recommended stopping, and EY shut it down

The headline scores pointed to "ship it." I read them against the behavioral evidence: output below the bar for real use, prompts most people couldn't complete correctly, and decks too generic to win a client. Saving nine minutes on a draft doesn't help if the deck still needs most of the original 19 hours.

My findings were the basis for the decision to stop investing, and the tool was shut down, before more engineering time went into a product people praised but wouldn't choose on a real deadline.

Why the praise misled

Participants were helping a colleague

Participation is a relationship, and people keep saying yes to researchers out of goodwill even when they mean no (Howell, Desjardins & Fox, 2021). Our participants were evaluating a coworker's tool, sometimes with its builders in the room.

We asked for opinions about a future habit

"Would you recommend this?" invites optimism. "Show me your last pitch deck and how you built it" would have shown what the tool had to beat.

The story about AI was doing some of the talking

In late 2023, generative AI was widely seen as the future. Some of the praise was for AI in general, not this tool's output.

We framed the study around what the tool did well

Our hypotheses led with time saved. The quality gap was in the data, but it took deliberate work to move it to the center of the readout.

What I'd do differently

  1. Test on a real upcoming pitch, not a demo task.
  2. Track actual usage for two to four weeks after the sessions.
  3. Ask for a commitment, such as "use this for your next pitch," instead of a rating.
  4. Keep the product team out of sessions.
  5. Make the quality gap the headline metric from day one.

Research write-ups tend to tell tidy success stories (Howell, Desjardins & Fox, 2021). This one is deliberately the opposite, and it's where my focus on usefulness comes from.

Let's talk.

I'm looking for UX research roles on teams building AI and health products. Resume available on request.