Researcher-in-the-loop

A model for AI-enabled UX research where the researcher governs.

A UXR researcher governing AI (AI generated)

For a few years now, the UX research field has been arguing with itself over whether AI is coming to replace us or whether AI can’t do “real” research and we need to hold the line. Is AI coming to replace us? Should we use AI to automate research? Can AI conduct “real” research? Research is about asking questions. However, we can ask better questions about AI in UXR.

We don’t need to ask whether AI can do research. It’s already doing it, and it’s doing it in your org right now, whether or not a researcher is involved. Product managers are pasting interview notes into a model and asking AI to “find themes.” Designers are generating synthetic feedback on wireframes. Executives are asking a chatbot what users think. Research is being democratized whether we govern it or not.

So the real question is: when everyone can do research, who keeps it honest?

Who governs UX research? Who governs AI?

“What if, instead of thinking of automation as the removal of human involvement from a task, we imagined it as the selective inclusion of human participation?” Stanford University HAI

Human-in-the-loop (HITL) AI emerged to keep human judgment inside automated systems, catching the bias, opacity, and plain mistakes that models produce on their own, especially in the ambiguous or rare cases they handle worst. The idea is a feedback loop: human corrections train the system over time. At its best, HITL pairs what AI excels at (pattern recognition, speed, scale) with what humans excel at (ethics, oversight, judgment). But it has a known limit: even when it’s well-designed, a human can become overloaded, checking more than they can meaningfully assess.

Many argue for stepping beyond the human-in-the-loop concept. Konstantinos Lazaros et al., after reviewing 182 peer-reviewed studies, suggest more “sophisticated ‘human-with-the-loop’” partnerships that would adjust the amount of control the human and the AI have based on risk and uncertainty. Rahwan argues that AI systems ultimately require “society-in-the-loop” to negotiate values.

I propose a model I’ve come to call researcher-in-the-loop. It’s a deliberate inversion of “human-in-the-loop,” where a person checks the machine’s work case by case. That model doesn’t scale. It assumes the machine is the researcher and the human is the safety net. Let’s flip that. The researcher isn’t the safety net. The researcher is the person who designs and governs the system and conducts high-risk, high-priority research so that everyone else can conduct research safely, rigorously, and honestly, with escalation engineered into the tools themselves.

This model didn’t come from nowhere.

Three ideas hold it up:

1. Research democratization: The now-familiar argument, articulated clearly by Kara Pernice at Nielsen Norman Group, is that non-researchers can and should do research, but only when it “is accompanied by due diligence: setting guidelines, educating others, and demystifying the process”. Democratization with guardrails is the parent of everything here.

2. ResearchOps: Kate Towsey and the ResearchOps Community gave us the language of research as infrastructure, “the people, mechanisms, and strategies that set user research in motion” (Writing@TeamReOps). What I call governance is really ResearchOps pointed at AI.

3. Atomic research: Daniel Pidcock’s model, alongside Tomer Sharon’s parallel “nuggets,” breaks research into small, tagged, evidence-traceable units. You cannot build a trustworthy, sourced research assistant on top of PDFs buried in a drive; you can only build one on top of atomic, linked evidence.

Everything below is an attempt to assemble those three into a single governance stance. The assembly is the only part I’d claim as mine.

A hand-sketched, watercolor-style illustration depicts a confident UX researcher with shoulder-length, curly silver hair standing in a strong defensive pose, holding out one hand to stop a swarm of small AI robots. A glowing pink force field, matching the palette’s accent color, forms a barrier between the researcher and the advancing robots. The scene symbolizes human judgment, governance, and critical thinking acting as safeguards while preventing AI from operating unchecked.
Saying “no” and stopping AI that isn’t ready (AI generated)

What I learned by saying no

A few years ago, I directed research on a flagship AI-assisted feature that was on a fast track to release. The research told a blunt story: 75% of users could not complete the feature’s core task. The workflow was confusing, the AI was hard to understand and hard to trust, and, most importantly, it didn’t align with what we already knew about how our users wanted to use AI.

This was AI in a B2B product for highly regulated industries (think healthcare, legal, insurance, compliance, security, data privacy, banking). Our users’ working lives were built around transparency, confidentiality, ethical walls, and defensibility. They do not want a black box. They are trained to distrust an answer they can’t trace. An AI feature that couldn’t show its reasoning, couldn’t cite its basis, and couldn’t be verified was culturally incompatible with the people it was built for.

The research surfaced that mismatch, and the decision was a no-go. It reshaped how the company thought about future AI investment.

We all have these war stories. This one contains the entire thesis of this article in miniature. Ungoverned AI is dangerous. It is even more dangerous in high-trust domains. We didn’t catch this because we used a smarter AI model. The researcher caught the failure because of their judgment about risk, confidence, and how real users actually build trust [i]. This is researcher-in-the-loop.

Everything I’m about to propose is an attempt to take that judgment and build it into the system so it fires early and automatically, instead of heroically and late.

That instinct, the one that caught a 75% failure rate, is what we need to preserve. But what if we engineer it into the tools everyone uses?

A hand-sketched watercolor illustration shows a Black male UX researcher talking with a friendly AI robot dressed as a librarian behind a reference desk. The robot gestures while consulting an open book as charts, bookshelves, and abstract research visualizations fill the background. Soft blues, purples, greens, yellows, and pinks create a warm, modern scene representing AI as a knowledgeable research assistant.
The UX Researcher and the AI Librarian (AI generated)

The two artifacts, in brief

In the researcher-in-the-loop model, most day-to-day research questions don’t go to a researcher at all. They go to governed AI tools. Two such tools I’ve been building are the Librarian and the Persona.

The AI Librarian is your organization’s memory of what you already know about users. Ask it any question, and it pulls from every study, report, and transcript in the governed “library” to give you an executive summary, the sources behind it, and a confidence rating for the answer. Its most valuable response is “We don’t have enough research to answer that.”

The AI Personas are living, conversational representations of your users, built from real qualitative and quantitative research and continuously updated as new data comes in. Anyone can “chat” with a persona to get a quick directional signal. When the data doesn’t support a confident answer, the persona says so and points you toward the researcher and the real research you’d need.

Governance by design; Rigor you don’t have to remember

In most orgs, research quality depends on a researcher being in the room. Remove the researcher, and quality becomes a hope. The researcher-in-the-loop model refuses that fragility. Rigor stops being a manual review step you might get to. It’s engineered into the tools. It is the default.

Recent critiques by Papangelis, Scaff, and Anastasiou & De Liddo warn of a “Synthetic Persona Fallacy,” where black-box tools systematically smooth out human variance and hallucinate confidence. Our model addresses this fragility by making transparency a technical constraint through mandatory sourcing, confidence ratings, and escalation triggers. Three mechanisms carry most of the weight:

  • Mandatory sourcing. No insight without its basis. Every answer traces back to the studies behind it. This is the transparency that our stakeholders, our teams, and our users demand and deserve.
  • Confidence and trust ratings. Every insight carries a signal of how much to lean on it. High-confidence answers move fast. Low-confidence answers slow down and flag themselves.
  • Escalation triggers. The system knows what it can’t handle and automatically pings a human researcher, especially in high-risk and low-confidence areas. Escalation becomes a reflex, not a courtesy or a gatekeeper.

That last mechanism is what moves this from a disclaimer to a governance system. A disclaimer says, “verify important results.” A governance system identifies the important results and routes them to a human before anyone acts on them.

Honesty by design AI governance flow

The envelope of replaceability: When AI can stand in and when it can’t

“Success [with AI] requires best practices — loosely held.” Figma’s 2025 AI report

With researcher-in-the-loop AI governance, AI personas, and the AI Librarian, AI can genuinely replace some early research. Not researchers. Not for all the research. But it can replace specific slices of it, and only inside a governed envelope.

To do this with rigor, we don’t default to replacement. Replacing any research is a privilege the system earns, question by question, based on two axes: risk and confidence.

Another way to say this: the model is a way to rightsizing research. Our field has argued for rightsizing for decades. Jakob Nielsen made the case respectable back in 2000, when he told us to “distribute your budget for user testing across many small tests instead of blowing everything on a single, elaborate study.” His “just enough” argument was never really about the number five; it was about matching research effort to the stakes of the question. The envelope does the same thing for the AI era, only now the scarce resource we’re rationing isn’t sample size; it’s human oversight. We spend it where risk and uncertainty are highest, and we don’t waste it where the answer is already clear.

This matrix is the operational heart of the model. It’s whiteboard-simple on purpose, because it has to work as policy, not philosophy. A designer asking, “What have users said about our onboarding tone?” is low-risk and high-confidence, so we let them self-serve. A PM asking a persona to greenlight a launch is off the map entirely (as this is not a low-confidence answer, but a question that isn’t the persona’s to answer at all).

The 75%-failure feature lives in the bottom-right cell. High-risk domain, low real-world confidence in users’ trust in the AI. That is the cell where a human researcher must be in the loop, and the cell where a well-governed AI should raise its hand early rather than letting the feature sail toward release.

Risk-by-confidence matrix showing four zones: GO (self-serve) for low-risk, high-confidence questions; CONFIRM and REVIEW (human involvement) for mixed cases; and STOP (researcher required) for high-risk, low-confidence questions.
Envelope of Replaceability

The loop that improves itself

So far, I’ve described the matrix as if it were a gate: a question comes in, you classify it, you route it once. That’s the tidy version. The real thing is a loop.

Because every time the matrix escalates a high-risk, low-confidence question to the STOP cell, the question is brought to the research team, and a researcher may conduct real research with real users. The research doesn’t just answer its own question and disappear. It goes back into the artifacts. The persona is updated based on what we heard. The Librarian gets a new source to cite. The confidence score on that whole class of questions moves. The next person who asks the persona something adjacent gets a better, better-grounded answer than they would have last week.

That changes what the AI tools are. They’re not a replacement for the research cycle sitting off to one side. They’re one governed step inside it, the fast, cheap, self-serve front door, wired to the slow, deep, human work by a feedback loop.

This is continuous discovery with a governance layer bolted on. Teresa Torres defines the practice as engaging with customers “at least weekly, minimizing the number of decisions they make without customer input.” I want to be honest about the tension there: she keeps a human in weekly contact with real people, while my loop puts an AI at the front door. That’s exactly why the governance layer matters. The self-serve tools make “continuous” cheap; the matrix and the loop are what keep “minimizing decisions made without customer input” from becoming “minimizing contact with customers.” Cheap directional signals move early and stay continuous; expensive, real research is reserved and aimed precisely where the matrix says it’s needed. Research stops being the “bottleneck” it gets called in lower-maturity orgs, where it’s treated as an add-on rather than a necessity, and becomes a resource you deploy where the loop tells you it’ll change the outcome.

This has a consequence I didn’t expect until I watched it happen: the low-confidence cells can become your research backlog. The system tells you where it’s weakest, which is exactly where a human should be looking next. And because each real study raises the confidence ceiling of the artifacts, research spend compounds instead of depreciating; a study you run today keeps paying off in every future query it touches. That compounding isn’t automatic, though: users change and products change, so the artifacts need pruning and expiry dates as much as they need new inputs, or yesterday’s confident answer quietly becomes today’s stale one. Maintained honestly, though, research spend builds on itself instead of evaporating, and for anyone who has ever had to defend a research budget, that’s not a small thing.

The loop can also be part of a bigger loop. The AI artifacts work best in an iterative process. Start with AI. Ask the AI persona about the concept. Have the AI librarian pull everything we have on the topic and what we have learned, thus giving us the grounding.

But a loop that feeds itself can also fool itself. If the artifact shapes which questions get asked, and those questions shape at least some of the research conducted, and that research feeds back into the artifact, you can build a machine that keeps confirming what it already believes and never goes looking in its own blind spots. Iteration sharpens calibration, but it can slowly starve discovery of the unknown-unknowns. So, the loop needs one more guardrail, and it’s a human one: the researcher has to force exploratory research periodically; the system would never request it and would deliberately study the areas where the persona is most confident, not just where it’s uncertain, precisely to catch the places it’s confidently wrong. The loop improves itself. It does not police itself. That part is still the researcher’s job.

The researcher-in-the-loop process that starts with the question or Idea, runs it by AI, and then checks to see if it is low-risk and high confidence, so the answer stand, or if it is otherwise it gets escalated. Once escalated research is conducted with real users. The findings feedback into the AI tools, making them smarter over time. UX researcher also do exploratory studies along with way to check and update the AI.

Where this breaks

Few models are honest about their own failure points, so let me be honest about mine. A governance system that sounds rigorous can fail in exactly the ways it claims to prevent, and two failure modes worry me most. Both come from the same place: the moment a number or a confident-sounding answer stops being treated as a signal and starts being treated as the truth.

Automation creep and bias: Research by Parasuraman & Manzey, among others, has found that humans systematically over-trust automated output and under-scrutinize it relative to their own judgment, leading to omission and commission errors even among experts. According to Mosier, Skitka, and Burdick, this bias “cannot be prevented by training or instructions.” This is a fear many of us share, especially amid the democratization of research through AI. We don’t want a stakeholder, or even a researcher, to overtrust AI artifacts. This is why the escalation triggers exist. The AI artifacts are designed to tell you how trustworthy the output is. While this may not eliminate automation bias, it should decrease its likelihood and strength.

But this exposes the model’s central risk, and I want to name it plainly: the envelope only works if a question is classified correctly, and classification itself is a judgment call. If the person using the model for self-service is the one deciding whether their question is “high risk” or “low confidence,” you’ve handed the label to exactly the person most motivated to under-rate the risk and get a fast answer, and, thanks to automation bias, least likely to notice they’ve done it.

The same trap applies to the AI itself: a model asked to rate its own confidence is often confidently wrong. Chhikara and others document that LLMs overstate their certainty in ways that misalign with their actual accuracy, which is precisely the risk when someone acts on the rating. The classification can’t rely on the asker’s or the model’s self-report. It has to be engineered into the system, with confidence derived from the underlying research (how much real evidence exists, how recent it is, how consistent it is), not from how sure the model sounds; risk thresholds set in advance by the researcher, not chosen in the moment by whoever wants the answer. The model doesn’t remove that judgment. It moves it upstream to the researcher who defines the rules, keeping it out of the hands of the person with a deadline. Get that wrong and the whole envelope collapses into theater. That is the thing the researcher is really governing.

A thing with numbers: Erika Hall warns us to beware of research theater and false objectivity: rigor comes from discipline and honesty, she argues, not the mere presence of numbers. As she puts it, “just because a thing has numbers doesn’t make it magically objective or meaningful.” My artifacts are built on exactly the thing she’s suspicious of. A confidence rating is “a thing with numbers,” and, paired with automation bias, it’s a prime candidate for manufacturing false objectivity, making a shaky answer feel settled because it arrived with a percentage attached.

So does the model answer Hall? Partly, and only conditionally. Governance can keep the number accountable: every rating traceable to the research behind it, labeled as a signal rather than a verdict, and flagged when the evidence is thin. That’s discipline made structural instead of hoped-for. But it doesn’t make the number “objective,” and I don’t want to pretend it does. A well-sourced confidence rating can still be misread as more certain than it is; that’s a limit of how humans read numbers, not a flaw I can engineer away. Hall’s pressure doesn’t lift once you add governance. It just moves on to the researcher, who has to keep auditing whether the confidence the system projects is confidence it has actually earned.

Notice that neither of these failures is a bug you can patch. Automation bias is a documented feature of how humans treat machines, and false precision is baked into any number that looks more certain than it is. Governance doesn’t eliminate these risks; it can only manage them, and only if the system is built and maintained honestly. That “if” is doing a lot of work, and it’s precisely why the model can’t run on autopilot. Which brings us back to the one component no rating or trigger can replace: the researcher.

A hand-sketched watercolor illustration depicts a non-binary Asian UX researcher seated behind a large governor-style desk, thoughtfully overseeing a floor of small AI robots carrying out tasks below. Charts, dashboards, governance symbols, and research visuals surround the scene in soft blue, purple, green, yellow, and pink tones, symbolizing human oversight, responsible AI governance, and strategic decision-making.
UX researcher as AI governor (AI generated)

The researcher’s new job: From “bottleneck” to governor

“The more advanced the systems become, the more important empathy, context, and ethics feel.” Sophia Omarji

If the tools handle the low-risk, high-confidence middle of the workload, what is the researcher for?

More than ever, the researcher enters the loop wherever expertise is non-negotiable:

  • High-priority, high-impact studies, and key clients or sensitive user groups
  • Discovery and exploratory research, the generative, ambiguous work that can’t be templatized
  • Anything touching protected data, risky populations, or unclear or inappropriate goals
  • Hard questions the tools flag as beyond their envelope
  • Iteration for steering the loop itself, deciding when the AI artifacts are grounded enough to answer, when a question needs a fresh round of real research, and when to force exploratory studies in the areas the tools are most confident about, precisely to catch where they’re confidently wrong.

This is the part I care about most: the researcher is a quality authority over the process itself, not just a reviewer of outputs. The researcher checks that the plan makes sense for the question, that the script and survey items are well-formed, that the method fits, that the analysis is sound, and that stakeholder conversations haven’t biased the work. The AI can draft a study. Only a researcher can confirm it’s the right study, designed correctly, and rigorous enough to bet on.

That is not a diminished role. It’s a promoted one. This researcher stops being the “bottleneck” designers and PMs complain about and becomes the governor of a research system that serves the whole organization. This also means researchers spend fewer hours moderating routine sessions and more hours setting standards, auditing rigor, designing guardrails, and doing the hard, generative research only a human can do.

The anxious question we have been asking, “Will AI replace researchers?” has it backward. AI raises the value of the person who can govern it.

Keep the researcher in the loop

Democratization and rigor are usually sold as a trade-off: let more people do research, accept lower quality. The researcher-in-the-loop model rejects that trade. Governance by design, mandatory sources, honest confidence, automatic escalation, and roles bounded by the norms we already hold let us scale access without scaling risk.

The researcher doesn’t disappear in this future. The researcher doesn’t fight it, either. The researcher governs it. The researcher treats the persona like the user it represents, treats the Librarian like the SME they are, and steps personally into the loop wherever the stakes, the ambiguity, or the craft demand a human who knows better.

This takes AI from a tool to a stance. This is the stance our field needs to take before the question of who keeps research honest gets answered for us.

Let’s build this together

I’m developing the researcher-in-the-loop model in the open, and I want to pressure-test it against how the rest of you are actually working. So, a few honest questions for you:

· If you lead research: Where would you draw the human line? Which cell of that matrix keeps you up at night? What have you already tried to democratize, and what broke?

· If you’re a designer, PM, or founder: What would you want to ask an AI librarian or a persona tomorrow morning? What would make you trust the answer?

· If you build these tools: What would it take to make “insufficient data” and automatic escalation first-class features instead of afterthoughts?

Drop your take in the responses, or reach me directly. I read every message, and the best refinements to this model will come from people living the trade-offs, not from those theorizing about them. If this framing was useful, follow along; I’ll be sharing the templates, escalation rules, and governance patterns as I put them into practice.

The researcher stays in the loop. Let’s make sure the loop is one worth being in.

If you want to go deeper into the ideas this piece builds on, democratization, the AI-and-researchers debate, and synthetic users, here’s where I’d start.

5 UXR AI-Ethicist Theories — Part 1.” Medium, Dawn.

“Democracy or Tyranny in UX Research… Is There a Middle Ground?ResearchOps Community, Jeanette Fuccella.

Discovery is the work AI gives back.” UX Collective, Gale Robins.

From Fear to Empowerment: How to Supercharge UX Research with AI — a framework.” Medium, Sasha Luca

The crisis of the what.” UX Collective, Francisco Barrera Aros.

The future of UX research: From AI as a tool to AI as a research partner.” Bootcamp, Dhairya Sathvara.

The nuts, bolts, and ethics of synthetic user personas.” Bootcamp, Julian Scaff.

The strategic shift: navigating the future of UX research with AI.” Bootcamp, Aisling Carty.

The Synthetic Persona Fallacy.” ACM Interactions. Konstantinos Papangelis.

Why AI Will Never Replace Good User Researchers.” Bootcamp, Lesley Crane.

Jennifer L. Bowie, Ph.D., is a UX research leader who has built and scaled research practices across enterprise SaaS, legal tech, fintech, and regulated products, with a focus on AI governance and human-centered AI adoption.

[i] Adrienne Brian was the lead researcher for this AI research. She did a wonderful job. Her research prevented the company from moving forward with AI that would ruin client trust, decrease adoption, and increase turnover.


Researcher-in-the-loop was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.

Need help?

Don't hesitate to reach out to us regarding a project, custom development, or any general inquiries.
We're here to assist you.

Get in touch