Cornell Future of Learning Lab · 2026
When your coworker jumps in mid-thought
Easing nervous first-timers into a live argument with an AI coworker
Participants practice workplace conflict with AI coworkers who talk back. I redesigned the way in, built the researchers’ replay view and measured the coworkers live, by voice.
- Role
- Undergraduate researcher and full-stack developer, AI-Ready Workforce Initiative
- Team
- The lab’s principal investigator and research staff, with Anthropic Research
- Timeline
- August 2026 to now
- What changed
- From a flow that lost first-timers to one they get through
In short
- The study runs with Anthropic Research; I brought in five new company partners.
- The coworkers speak through Gemini Live but had only been measured in text. Out loud, they jumped in mid-thought.

My part
- Redesignedthe path from consent to the live call, revising it wherever a first-timer stalled.12 participants
- Builtthe Evidence Trace: one row per turn.Researcher dashboard
- Measuredthe coworkers live and fixed the turn-taking and audio problems that surfaced.False cut-offs 22% to 0%
The problem
A research platform has two users who never meet.
One is a nervous participant in a headset, about to argue with an AI coworker named Sam. Each coworker has two parts: an actor the participant hears, and a director telling it how to play each turn.
The other is the researcher, who needs to know what happened in every turn, down to whether a direction to the actor arrived late.
The first version let both down: participants stalled between consent and the call, and researchers couldn’t line up the audio with the logs.



Decision log
to measure the coworkers live, scripting a hesitant, lost participant.
Rejectedtext-only tests.
Whythe coworker jumped in at mid-thought pauses. Text had shown none of it.
Resultwith room-adaptive voice detection, false cut-offs fell from 22 percent to zero.
to pick the voice model by A/B tests with real people.
Rejectedpublished benchmarks.
Whya half-second of lag reads as a coworker who isn’t listening, and no benchmark measured that.
Resultthe tests picked the model; the same setup caught bugs dropping participant audio.
the Evidence Trace: one row per turn with the participant, the director’s note and the actor’s reply on one clock, with fidelity checks and flags for directions that arrived late or leaked into the conversation.
Rejecteda transcript plus a separate log download.
Whya researcher scoring the session who can’t see that a direction arrived late will read the delay as the character’s personality, and the study’s conclusions would inherit the error.
Resultresearchers can replay any session against its logs in a single view.
to strip the call screen down to what a participant needs: who they are talking to, that it is recording, and a way to leave.
Rejectedthe usual video-call controls.
Whya self-view, a mute button and a timer all gave nervous participants something to fiddle with instead of talking.
Resultfirst-time participants get through the flow. The twelve think-aloud sessions drove the revisions.
Not in the first version
Designed and deferred: a live wall of several sessions, cross-session analytics, a prompt editor, and rater assignment and export.
What I would do next
Run the researcher side through the same 12-person usability protocol. So far, the dashboard’s only tester is me.