Cornell Future of Learning Lab · 2026

When your coworker jumps in mid-thought

Easing nervous first-⁠timers into a live argument with an AI coworker

Participants practice workplace conflict with AI coworkers who talk back. I redesigned the way in, built the researchers’ replay view and measured the coworkers live, by voice.

162issues fixed before data collectionAmong them, session recordings reachable without a login.
Role
Undergraduate researcher and full-⁠stack developer, AI-⁠Ready Workforce Initiative
Team
The lab’s principal investigator and research staff, with Anthropic Research
Timeline
August 2026 to now
What changed
From a flow that lost first-timers to one they get through

In short

  • The study runs with Anthropic Research; I brought in five new company partners.
  • The coworkers speak through Gemini Live but had only been measured in text. Out loud, they jumped in mid-thought.
The Evidence Trace with sample data: a list of sample sessions, then one practice conversation with Sam, a row per turn with its time in seconds: the participant’s words, the AI director’s note for each of the actor’s turns with tags for whether it steered, and the actor’s reply. To the right sit the fidelity checks and session health, where director latency still reads pending. The replay timeline runs beneath, its marker on the highlighted turn, with a note that video sync is pending.
12participants in usability testsThink-aloud sessions.
22 percent to 0 percentfalse audio cut-offs in the live test setBefore and after room-adaptive voice detection. Genuine interjections rose from 25 to 56 percent.

My part

  1. Redesignedthe path from consent to the live call, revising it wherever a first-timer stalled.12 participants
  2. Builtthe Evidence Trace: one row per turn.Researcher dashboard
  3. Measuredthe coworkers live and fixed the turn-taking and audio problems that surfaced.False cut-offs 22% to 0%

The problem

A research platform has two users who never meet.

One is a nervous participant in a headset, about to argue with an AI coworker named Sam. Each coworker has two parts: an actor the participant hears, and a director telling it how to play each turn.

The other is the researcher, who needs to know what happened in every turn, down to whether a direction to the actor arrived late.

The first version let both down: participants stalled between consent and the call, and researchers couldn’t line up the audio with the logs.

First: a notice that you need a quiet room, because the microphone hears everyone, and a request to keep the conversation fictional: no real names, no real events, no contact details, and to steer back to the scenario if it driftsSecond: a check of the microphone, the speakers or headphones, and the camera, each with its own test button, and Continue locked until all three have passedThird: the situation card for the Taken credit scenario: what happened with Sam, who you’ll speak with, how the scene runs, and what to do, with a reminder that it is fictional
The way in today. Rendered from the platform’s public code with sample data, like the Evidence Trace above.

Decision log

Chosen

to measure the coworkers live, scripting a hesitant, lost participant.

Rejected

text-only tests.

Why

the coworker jumped in at mid-thought pauses. Text had shown none of it.

Result

with room-adaptive voice detection, false cut-offs fell from 22 percent to zero.

Chosen

to pick the voice model by A/B tests with real people.

Rejected

published benchmarks.

Why

a half-second of lag reads as a coworker who isn’t listening, and no benchmark measured that.

Result

the tests picked the model; the same setup caught bugs dropping participant audio.

Chosen

the Evidence Trace: one row per turn with the participant, the director’s note and the actor’s reply on one clock, with fidelity checks and flags for directions that arrived late or leaked into the conversation.

Rejected

a transcript plus a separate log download.

Why

a researcher scoring the session who can’t see that a direction arrived late will read the delay as the character’s personality, and the study’s conclusions would inherit the error.

Result

researchers can replay any session against its logs in a single view.

Chosen

to strip the call screen down to what a participant needs: who they are talking to, that it is recording, and a way to leave.

Rejected

the usual video-call controls.

Why

a self-view, a mute button and a timer all gave nervous participants something to fiddle with instead of talking.

Result

first-time participants get through the flow. The twelve think-aloud sessions drove the revisions.

Not in the first version

Designed and deferred: a live wall of several sessions, cross-session analytics, a prompt editor, and rater assignment and export.

What I would do next

Run the researcher side through the same 12-⁠person usability protocol. So far, the dashboard’s only tester is me.