Read the moment.
Browser camera input becomes coarse expression, movement, head, gaze, and gesture signals through MediaPipe. Raw frames never leave the device.
Your main agent should get you. It should be there with you in the moment. Sam brings realtime voice, local expression, head, gaze, gesture, and Jev-powered attention into one calm conversation.
A smile while you speak. A raised brow. A shift in attention. Sam brings what you say together with what is happening around the conversation, so her response can meet the moment.
Browser camera input becomes coarse expression, movement, head, gaze, and gesture signals through MediaPipe. Raw frames never leave the device.
Your words meet a changing stream of compact observations. The voice model gets context for how the conversation is unfolding, alongside what you are saying.
Jev works in parallel, interpreting the words and observations through several focused questions. It returns structured signals about relevance, grounding, and perspective for Sam to draw on.
Those signals become part of Sam’s conversational context. She can acknowledge an expression, ask a better question, or keep listening—without turning every movement into a comment.
MediaPipe runs in the browser. Compact observations and the current utterance feed Jev’s parallel structured judgments; available observations and Jev results update the realtime voice context. OpenAI’s Marin is the primary voice, with Grok’s Carina as fallback. Camera frames and landmarks stay on-device. Observations are temporary, interpretations remain uncertain, and the live colors follow speech and interaction state.
Voice gives the moment shape, so conversation can stay fluid and responsive.
Attention and expression are treated as bounded signals — never definitive emotions or mindreading.
Sam works from compact observations and structured context, with your permission.