I research the moments when people decide, in seconds, whether something is worth their attention.
UX researcher and research operations specialist, eleven years across industry and academia. Most recently I built and ran research operations for a seventeen-person research team at Upwork, and in parallel led mixed-methods research on generative AI and content discovery at Microsoft AI.
Luis VergaraResearch OperationsUX Researcher11+ years · academia and industryCali, Colombia · open to relocation
every case here answers
What was the team about to decide?
then shows
How I designed the study to change their confidence
and ends with
What we learned, what changed, and what I'd do differently
Currently based inCali, Colombia - Open to RelocationLanguages Spanish · native English · C2 German · B2StatusOpen to UX Research and Research Operations roles
Five chapters
Each opens to its own page with the full case studies.
Four habits show up in every study on this site, whether I had three weeks or three days.
Frame the decision before the method
I start by tightening the question: what is the team about to choose, and what evidence would change their confidence? The study design follows from that. It's the difference between research that gets cited in a planning doc and research that gets filed.
Design for speed without giving up rigor
Product timelines rarely allow a perfect study. They do allow a well-designed one. Within-subject comparisons, counterbalanced order and screener verification let me run directional work that still holds up when someone senior pushes on it.
Measure behavior and perception together
What people do (scan, abandon, return) and what they report (trust, quality, ease) rarely line up. The gap between them is usually where the insight is.
Make the finding travel
A finding only matters if it survives the meeting. Short TL;DRs, explicit trade-offs, and every recommendation tied to something on the roadmap. At Upwork I built the repository and templates that let findings outlive the study they came from.
Methods, tools, education
Methods
Moderated and unmoderated usability testing, in-depth interviews, contextual inquiry, concept and comparative testing, diary studies, card sorting and tree testing, surveys and CSAT/NPS programs, within-subject and counterbalanced designs, thematic analysis
Tools
UserTesting, UserZoom, Dscout, Outset, Qualtrics, Calendly, DocuSign, Tango, Tremendous, Jira, Notion, Figma, Miro, Google Workspace; Claude, ChatGPT and Copilot for synthesis, analysis and workflow automation
Education
M.A. Electronic Media, Hochschule der Medien, Stuttgart (2012 to 2014)
B.Sc. Interactive Media Design, Universidad Icesi, Cali (2005 to 2010)
Languages
English (C2), Spanish (native), German (B2, a little rusty after eleven years away)
Microsoft AI
UX Researcher II
October 2022 to March 2026, remote with the Redmond team. Two chapters: agile discovery research with the ARC team (Agile Research with Customers) across Edge, Bing, Shopping and SwiftKey, then end-to-end product research on Copilot and MSN content-discovery surfaces.
The ARC research loop I ran for the first two years: a product question in, a share-out and a repository entry out, every two weeks.
Does a grid or a linear feed better support a five-minute break?
SurfaceContent discovery feed on Windows 11
MethodUnmoderated, within-subject comparison
Participants32 Windows 11 users, screenshot-verified
My roleStudy design, recruiting, analysis, report, recommendations
PartnersDesign, PM, data science
The decision
The team was choosing between two layouts for a high-traffic content feed and had strong opinions on both sides. Preference data existed. What was missing was evidence about how each layout changed the way people actually used the feed.
How we designed it
Participants were Windows 11 users who already had the live grid experience. Each evaluated their own feed and then a linear prototype, so every comparison was within-subject. Because self-reported eligibility was unreliable, participants were required to share a screenshot of their own feed as proof they had the live experience before entering the study. I measured scan behavior, perceived content quality, decision ease and return intent on a 5-point scale, with open-ended prompts after each.
Study flow. The screenshot proof at the start is the unglamorous step that made the rest of the data trustworthy.
What we learned
Layout wasn't a visual preference. It changed the mindset people brought to the feed.
The grid invited fast scanning and quick exits; the linear layout pulled people into a slower, deeper reading mode. Neither was better in the abstract. The right answer depended on which job the surface was hired to do, and the team hadn't yet agreed on that.
The synthesis board. Open-ended responses clustered into two modes; the line at the bottom is what went into the TL;DR.
What changed
The conversation moved from "which looks better" to "which mode are we designing for." The report gave the team a shared vocabulary (scan mode versus read mode) that showed up in later design reviews, and the recommendation tied each layout to a usage moment rather than declaring a winner.
Looking back
The screenshot verification step cost me two extra days of recruiting. It was worth it, but next time I'd build it into the panel setup up front instead of into each study.
When do people want quick utility, when do they want to browse, and what should open first?
SurfaceWindows content panel with two boards: widgets and discover
MethodUnmoderated comparative study, counterbalanced order
Participants48, split evenly across the two orders
My roleStudy design, execution, analysis, recommendation framing
PartnersDesign, PM
The decision
A panel offered two entry points, one built for glanceable utility and one built for scrolling discovery. The team needed to know how users understood the difference, which felt more useful, whether either hurt focus, and which one should be the default when the panel opened.
How we designed it
Half the participants saw widgets first, half saw discover first, so first-impression effects couldn't drive the result. Everyone rated clarity, usefulness and focus impact for each board, then chose a default and explained why. Screenshot verification confirmed they were on the correct build.
Counterbalanced order. Cheap to set up, and the reason the default-board recommendation held up under scrutiny.
What we learned
People sorted the two boards by intent, not by feature: “check something” versus “kill time.”
Widgets won on clarity and focus. Discover won on time spent but was the one more likely to feel like an interruption. The default question hinged on when the panel typically opens, not on which board people liked more.
What changed
The recommendation framed the default as a context decision, with evidence for each scenario, rather than a single verdict. That gave the team a defensible position in an executive review where the two boards had different sponsors.
Looking back
I'd add a short diary component next time. A one-week log of when people actually opened the panel would have settled the 'when does it open' question with behavior instead of recall.
What makes an AI summary feel useful enough to come back for?
SurfaceMultimodal news prototype: text, video and podcast modes with depth controls
MethodUnmoderated prototype evaluation
Participants44, split evenly between Gen Z and Millennial
My roleStudy design, synthesis, recommendation writing
PartnersDesign, PM, content team
The decision
The prototype offered AI-generated summaries in several formats and let users pick how deep to go. The team wanted to know which formats earned trust, whether depth controls were understood, and what would make someone return rather than treat it as a novelty.
How we designed it
Participants worked through the same stories in each mode, rated trust, usefulness and return intent per format, and talked through how they decided when a summary was "enough." I coded the open-ended responses against the ratings to see where perception and behavior diverged.
How the research questions mapped to what was measured and what was coded. Ratings say what, codes say why, divergence is the insight.
What we learned
Trust followed control, not format.
The summaries people rated highest weren't the shortest or the most polished. They were the ones where it was obvious how to reach the full story and how to adjust depth. When the path back to the source was unclear, trust dropped regardless of how good the summary read. Younger participants treated a visible source path as the signal that the AI wasn't hiding anything.
What changed
The recommendation prioritized source visibility and depth controls over adding formats. Several proposed modes were deprioritized in favor of getting the core reading experience right. The study also fed a set of principles about user agency in AI-assisted content that the team carried into later work.
Looking back
Podcast mode got the least engagement in an unmoderated setting, and I don't think that's a fair read of it. Audio needs a moderated or in-context test. I flagged the limitation in the report rather than let the number stand alone.
Can a "catch up" module help people feel informed without feeling overwhelmed?
SurfaceTop-of-feed digest: a carousel of grouped stories
MethodWithin-subject comparison of three prototypes
Participants36, rotated across the three orders
My roleStudy design, synthesis, UX recommendations
PartnersDesign, PM, editorial
The decision
A digest at the top of the feed promised faster catch-up but risked adding noise and eroding trust if it felt like it was choosing for people. Three variants existed. The team needed evidence before committing engineering time.
How we designed it
Every participant saw all three variants in rotated order and rated scannability, engagement, perceived value and trust. I paired that with a short reflection on what they thought the digest was "for," which turned out to be the most useful question in the study.
Lo-fi wireframes of the three variants. Same feed underneath; only the top module changed.
What we learned
The focused digest won, and the reason was trust, not speed.
The broad version felt like the product was guessing; the focused version felt like it was helping. Most people described the surface as "inform first, entertain later," and the digest only worked when it respected that order.
What changed
The team moved forward with the focused variant and adopted "inform first" as a design principle for the top section. It's one of the clearest examples in my work of a study changing not just a feature but the team's mental model of the surface.
Looking back
The 'what is this for?' question was an afterthought I added while writing the protocol. It produced the best data in the study. I now put a mental-model question in every evaluative protocol by default.
Where should editorial and content experiences live inside an AI assistant, and what would make them credible?
ProgramAgile Research with Customers, upstream discovery
MethodIn-depth interviews, opportunity framing
Participants18 interviews across three usage profiles
This one came before any feature existed. Leadership was deciding whether editorial and content experiences belonged inside the assistant at all, and if so, which of several candidate directions deserved roadmap space in the coming planning cycle. Everything downstream, including the four evaluative studies above, depended on getting this framing right. The risk of skipping it was obvious in hindsight: the team would have spent a quarter optimising a surface nobody had established a need for.
How we designed it
I ran in-depth interviews with 18 people across three usage profiles: heavy assistant users, people who used the assistant occasionally for search-like tasks, and people who consumed a lot of news and content but had not adopted an assistant. That third group mattered most, because they were the ones the product hoped to win and the only ones who could tell us what would make an assistant credible as a source. I kept the guide deliberately unstructured in the middle: rather than reacting to concepts, participants walked me through how they actually caught up on things they cared about, and where an assistant did or did not fit that.
I then mapped the candidate directions against two axes: how much value people expected from each, and how much trust each one demanded before they would use it. The second axis was the one the team had not been discussing, and it changed the shape of the conversation.
The opportunity map, with generic area names. The framing that mattered: value alone did not decide priority; the trust each area demanded did.
What we learned
People wanted the assistant to be a door to content, not a replacement for it.
Credibility rested on two things. Visible sourcing, so it was always clear where something came from, and a clear line between what the assistant had summarised and what a person had written. Directions that blurred that line, including the most technically impressive ones, were the least welcome: participants read them as the product deciding for them rather than helping them decide. The highest-value directions were not the highest-trust ones, which is why the map mattered: the cheapest wins sat where value was high and the trust burden was low.
What changed
The framing became part of the research landscape the team used for planning. Three directions were sequenced by trust burden rather than by build cost, with the source-visible ones first. It also set the vocabulary the later evaluative work was built on: the trust-versus-control finding in the AI summaries study and the "inform first" principle from the digest study are both direct descendants of this framing, which is the clearest evidence I have that upstream work paid for itself.
Looking back
The non-adopter group gave the most useful data and was the hardest to recruit, so it was also the smallest. If I ran this again I would weight the sample toward them from the start, even at the cost of a longer field period.
Twelve more studies from the ARC program
Public products only. Each was a short-cycle study with interviews, diary tracking or prototype comparison, delivered as a report and a share-out.
Microsoft Edge Gamer Mode: browser behaviors while in-game
The value of a browser inside Gaming Copilot
Copilot as a writing assistant in the SwiftKey keyboard
Copilot as a rich-media generative tool in SwiftKey
Branded navigation on the Bing search results page
Edge text-prediction UX improvements
Enriching product review articles with Shopping insights
Auto-applied Shopping coupons in Edge
Microsoft Rewards loyalty program, competitive comparison
March 2022 to May 2026. I ran the operational side of the Customer Insights & Research team: recruiting, a curated customer panel, tooling, compliance, incentives, and the systems that let thirteen researchers spend their time on research instead of logistics. Upwork is a marketplace where clients and freelancers who have never met, often on different continents, have to decide whether to work together. The platform sits in the middle of that relationship, so much of the research I supported was about trust: how it gets built between strangers, where it breaks, and how the experience gets better for both sides at once.
I think of research operations as a product whose users are researchers. This is the pipeline I built and ran. Each stage had an owner, a template, and something researchers could count on.
Team structure and the four-stage pipeline. The tools are the visible part; the templates and ownership rules are what made it reliable.A single request, end to end. Enablement templates let PMs and designers run the middle steps themselves for lightweight studies.
The customer panel
Recruiting from zero for every study is slow, and it burns the people you keep asking. So I built a dedicated customer panel: active, engaged clients and freelancers who had opted in to research, carefully screened and filtered before they joined, and tagged so a researcher could pull the right people for a survey, an unmoderated test or a moderated interview without starting from scratch. The rule I tried to hold was contact at least once a month, so the panel stayed warm and members did not hear from us only when we needed something.
Research enablement
Demand for research always exceeded the team's capacity. Rather than absorb that load indefinitely, I built a self-serve layer: documented processes, screener and guide templates, and coaching for designers and PMs who wanted to run lightweight studies well. Stakeholders arrived at planning with usable evidence, and researchers were freed for the questions that actually needed them.
A method decision I argued for with evidence
When the team was skeptical that AI-moderated interviews could match human moderation for generative questions, I ran the same study both ways with comparable cohorts and put the outputs side by side. Thematic depth was comparable; completion rates and turnaround were better. The trade-off (less room for unexpected tangents) was real, so the recommendation reserved human moderation for studies where probing mattered most.
You don't win a method argument with conviction. You win it with a side-by-side.
The one-page plan every study started from. The first field is the one that changed behavior: name the decision before the method.
Looking back
The biggest operational win wasn't a tool. It was making the request form ask 'what decision will this inform?' before anything else. It cut vague requests in half and made triage a five-minute job.
July 2015 to December 2022. Fifteen semesters teaching HCI, Gamification Strategies and Sociotechnical Interaction, plus thesis tutoring and jury. Teaching is where I learned to explain research decisions to people who don't yet share my vocabulary, which turns out to be most stakeholders.
278students taught
15semesters
15+thesis projects supervised
Three projects, three layers of HCI
I built the course around projects that climb a ladder. The first teaches experimental design with nothing to hide behind: a hypothesis, a measurement, a conclusion. The second adds design to the measurement, so students have to invent the thing they are evaluating. The third is experience design in full, where the brief is open and the job is to improve a slice of someone's day. Each one is a skill I use in product research now.
01 · A keyboard layout as a hypothesis
HCI angleExperimental design
FormatGroups, each proposing its own keyboard layout
StudyLongitudinal: same participants, three sessions over nine days
ToolA typing platform I built, logging errors and time between keystrokes
Textbook experimental design doesn't stick until you've run one. Each group designed its own keyboard layout and proposed it as a hypothesis: this arrangement will make typing more efficient, and here is which keys we expect to be fastest and why. Under that sat a set of smaller hypotheses about key position and expected response time. To test them, the same participants came back for three sessions over nine days, and every session ran the same four tasks in order of length: a single key press, a word (a different one each session), a sentence and a paragraph. The tool logged every error and the time between keystrokes. Students then analysed how speed and errors changed from one session to the next, which keys were fastest and slowest and where they sat on their layout, and whether position made a significant difference or not. They defended that conclusion in class, including the hypotheses the data rejected. It's the same discipline I use in product research: define what would change your mind, then measure it.
Structure of the study. The layout itself is the hypothesis. Repeating the same tasks with the same people is what let students read the change between sessions as learning, key by key, and automatic logging meant the argument in class was about the conclusion, not the data.
02 · Interactions that measure an aptitude
HCI angleInteraction design plus measurement
FormatGroups of four, one university programme each
DeliverableAn interaction that produces a score for a chosen aptitude
Evaluated onWhether the metrics separate people on that aptitude
The second project asked for design and measurement at once. Groups of four each chose a programme the university offers and asked what a measurable aptitude for it would be: logical and mathematical reasoning for engineering and the exact sciences, for instance. Then they designed an interaction, often a game, that would exercise that aptitude and produce a score, and defined the metrics up front: what counts as a correct response, how time is handled, what a strong result looks like next to a weak one. They put it in front of classmates and reported whether the interaction actually separated people on the aptitude it claimed to measure. On the surface it is a small vocational-orientation tool. The lesson underneath is construct validity: decide what you are measuring before you build the thing that measures it, then check that it measured that and not something easier.
The four steps every group went through, and a worked example. The hard conversation in class was always step four: whether the score reflected the aptitude or just familiarity with the format.
03 · Making an unavoidable wait worth something
HCI angleExperience design
BriefOpen: pick a wait nobody can remove and improve it
ContextsGates, platforms, banks, hospitals, restaurants, and the journey itself
Evaluated onPain point found first, engagement, and an honest case for the venue
The third project was experience design with an open brief. Everyone waits: at a gate before a flight, on a platform, in a bank or a hospital queue, at a restaurant table, and then again inside the plane, train or bus. The wait can't be removed, so the question was how to design for it. Groups picked a context, studied what the wait was like now and what people did with the time, and proposed interactions that made it feel better and kept people engaged, with an argument for what the venue gained in return. The discipline here is the one I lean on most in product work: find the pain point before you design anything, and be honest about whether the idea serves the person waiting or only the business.
The brief, the three questions every group had to answer, and the output. The third question is there on purpose: a design that only serves the venue tends to make the wait worse.
Course threads
Thread
What students learned to do
Experimental design
Formulate hypotheses, run controlled and longitudinal studies, analyse the data and draw conclusions they can defend
Interaction design
Choose a method, build interfaces from it, evaluate prototypes through a structured process
Experience design
Identify human motivations, design engaging experiences, assess quality with research methods
Gamification
Apply game mechanics outside games, and know when they help versus when they manipulate
March 2021 to June 2022. Research and design for a bank's in-house digital wallet, from problem discovery through launch. I designed and evaluated two parts of the first-run experience in depth: account creation and avatar personalisation.
32%rise in app-store rating, first six months
89%user satisfaction, up 15 points from baseline
200k+monthly active users
13%new-user growth after launch
How do you make a bank account feel personal without a profile photo?
SurfaceA bank's in-house digital wallet, iOS and Android
MethodDiscovery interviews, concept testing of illustration styles, moderated and unmoderated prototype tests, A/B test on first-run personalisation, weekly store-review analysis
Participants24 in discovery and concept testing, 18 across prototype rounds
My roleResearch lead, and design of the flows and interactions that came out of it
PartnersProduct, engineering, brand
The decision
The bank was replacing an app people tolerated rather than used. For the new wallet the team prioritised two things in the first-run experience: an account-creation flow with fewer steps, and a way for people to make the account feel like their own. In most consumer apps that second part is a profile photo. In a financial app it could not be, because image uploads were restricted for security reasons. So the question became what to offer in its place, and whether it deserved a spot in onboarding at all.
How I designed it
I started with discovery interviews and an analysis of the old app's store reviews, which are the cheapest source of unfiltered complaints a team has. That gave me a map of the journey with candidate friction at each stage, which I then tested rather than assumed. On account creation I ran moderated sessions on the identity-verification steps to see where people slowed down and which steps could go.
For personalisation, with photos off the table, I proposed illustrated avatars and took several graphic styles into concept testing. It ran over more than one round: the first rounds chose the style, the later ones decided which characters to offer and how many. One requirement held through every round. The set had to be neutral enough that any user could find an option they were comfortable with, which is why the final lineup mixes stylised people with a few robots and creatures, for anyone who would rather not pick a face at all.
Once the picker existed, we ran an A/B test with new users. One group went through personalisation as a required step at the start of the first-run experience. The control group had no personalisation option. We compared how much each group used the app, measured in transactions, to check that the feature earned its place. Because I proposed the flows and interactions as well as running the research, findings went straight into design without a hand-off in between.
Journey map with the friction we studied at each stage, the method used, and what changed. The two highlighted stages, account creation and first-run setup, are the ones I designed and evaluated in depth.
The avatar picker. One illustration style, settled over several rounds of testing, with enough range that anyone can find one they like.Colour is a second, lighter step, which keeps the first screen a simple pick.Where it lands. The chosen avatar sits beside the greeting and on the profile, next to an optional nickname.
These screens come from the public Tuya app, where this flow is still live today.
What we learned
New users who personalised their account at the start made about 20% more transactions than users who never had the option.
The style rounds did their job: we went into build with one illustration style and a set of characters that users had already chosen between, not a designer's favourite. The A/B test answered the bigger question. The group that personalised at the start used the app about 20% more, counted in transactions, than the control group with no personalisation option. Because the step was part of onboarding for everyone in that group, this compares two whole groups of new users, not the people who chose to personalise against the people who did not. That made the case that a restricted, curated form of personalisation was worth offering even where a photo is impossible, and that it belonged in front of new users instead of waiting in a settings menu. On account creation, the sessions confirmed that sign-up friction was real, and that removing steps helped.
What changed
Onboarding steps were reduced. The avatar shipped as a curated set in the style that tested best, with colour as a second, lighter choice. Personalisation moved earlier in the first-run experience rather than sitting in settings. In the six months after launch the app's store rating rose 32%, satisfaction reached 89% against a baseline 15 points lower, and the product passed 200,000 monthly active users. Those are results for the whole launch, not for one feature; the avatar was one of the changes behind them.
Looking back
Reading store reviews weekly was the cheapest research I have ever run and it caught two issues before analytics did. I would formalise it from week one next time, with a shared tag sheet, instead of treating it as something I did on the side.
When my role at Upwork ended in May 2026 I had two things I did not want to let go cold. One was everything operations had taught me: how a vague request becomes a plan, how a system stays usable when thirteen people depend on it. The other came from Microsoft, where I spent three years using Copilot every day and researching it in order to improve it, which meant living in its failure cases and watching the gap between what the model did and what people expected it to do. Studying an assistant from the inside teaches you how one is put together. So I started building. First with ChatGPT, then with Claude, on my own time, until learning to build with AI stopped being a side experiment and became the thing I do.
Everything here sits under Nexus Army, the umbrella I build these agents in. Two products are in beta now, with a small group of clients using them day to day and sending back the rough edges I then fix. Others are earlier, in exploratory conversations with potential clients about whether the problem is worth building for at all. I include them because building the thing changes how you research the thing.
Financial assistant: how much should an agent decide for you?
What it isA WhatsApp-native assistant that logs expenses in conversation, sends a weekly summary, and surfaces insights as it learns your patterns
StackLLM front door with tool use, WhatsApp API, admin hub for frictions, tokens and cost
StatusBeta with 4 testers; in development and refinement
The research problem inside the product
Early versions understood a request only when phrased one way. The fix wasn't more prompting; it was treating every misunderstanding as data. I built an admin hub that logs each friction, and those logs became the examples for the next iteration, plus an eval set that runs after every change so regressions get caught. That is the same loop as a continuous research program, with a model as the participant.
The two patterns that matter: the assistant proposes a category from a past entry and waits to be confirmed, and the weekly summary invites a question rather than pushing an alert.How a message becomes an action, and the admin hub that closes the learning loop.
What I learned about agency and trust
People forgive an agent that asks. They don't forgive one that guesses about their money.
Adding a clarifying question when intent was unclear did more for trust than any accuracy gain. It's the same pattern I saw in the AI-summary study at Microsoft, from the other side of the table.
A games hub for big groups: can fifty games share one set of rules for how they connect?
What it isA hub of games built for big groups: a set of party games a room of people can play passing a single phone from hand to hand, plus duels and rooms so the same friends can keep challenging each other when they are apart
Scale50 games in three languages, daily challenges, profiles, and a live trivia mode that runs on a TV with phones as buzzers
StatusBeta with 5 testers; catalog and multiplayer still being refined
What it is for
The starting point was a specific social situation: a group of people together in a room, one phone, and nobody wanting to install anything. From there it grew a second mode for the same friends when they are not in the same place, where a game becomes a challenge you send and answer on your own time. The hub is the layer that lets both live together, so a group can move from passing a phone around to playing from three cities without changing app.
The research habit that shaped it
Every game had to be self-explanatory to a group at a party, which is a harsh usability bar: the phone is moving, someone is half-listening, and nobody reads instructions out loud. I play-tested each one with the beta group and kept a simple log of where people hesitated, what they asked, and whether the rules survived without me explaining them. Games that failed got rewritten rules and a visual tutorial overlay, not longer text.
Navigation flow. Every game reports the same stats, so the profile and daily-challenge features work across all fifty without per-game code.
What I learned
Difficulty is a trust signal.
When a "hard" mode wasn't actually hard, groups stopped believing the labels everywhere in the app. Calibrating difficulty honestly (a maze that really uses the farthest cell, a bot that really plays well) did more for retention than adding games.
Home. The streak bar is a single line that changes its message depending on whether a streak is alive, at risk, or has not started.Quizzes. Seven categories, each with its own ranking, adaptive difficulty and daily challenge.A one-minute logic minigame: deduce which object weighs the most from the scales. Every clue is computed from real weights.
Live screens from the beta. The product ships in Spanish, English and Portuguese.
Also under Nexus Army, earlier along
Two more agents are in early development, both built on the same idea: give a small business an assistant that handles the high-volume, low-judgment messages and hands anything ambiguous to a person, with the context attached.
Clinic assistant
A WhatsApp agent for medical practices: appointment scheduling, reminders, follow-ups and routine patient questions for specialist physicians. In conversation with two practices about piloting. The research question I care about here is disclosure: patients seem far less bothered by talking to an agent than by not knowing whether they are.
Restaurant assistant
A WhatsApp agent for restaurants: reservations, menu questions, and order intake. The earliest of the four; currently a prototype I use to test how much a small business will trust an agent with a customer-facing channel.