What Live Captions Does
Every translation Puente produces now generates a structured data packet — the original speech, the translation, the speaker identity, and the emotion register used. Live Captions taps into that stream and surfaces the translated text on screen as it arrives, not after the full turn is complete.
This is different from a transcript that appears after someone finishes speaking. Live Captions streams the translated words as the translation renders — so you’re reading along in real time, not waiting for a paragraph to appear after a pause.
The result is a fundamentally different experience in group or high-volume translation contexts. In a meeting with three speakers, you can glance at the caption bar to catch the current translated phrase while still making eye contact with the speaker. In a Mesh Room with participants in four languages, each person’s screen shows captions tailored to their language — continuously.
What Appears on Screen
Each caption entry shows:
- Translated text — streaming as it arrives
- Speaker name — if Voice Identity has attributed the turn to a known speaker
- Register badge — the tone register used (Auto / Formal / Casual / Domain) from the Emotion-Aware Translation system
The caption bar is a floating overlay — it sits above the main translation interface and doesn’t obscure the controls you need to use during a conversation. It can be dismissed with a tap and re-enabled from Settings.
AR Glasses Integration
Live Captions is architected with AR glasses output in mind. The display stub for smart glasses caption routing is already in place — when compatible glasses with a heads-up display (like Xreal Air 2 or Even Realities G1) are connected, the same caption stream will route to the lens display rather than or in addition to the phone screen.
This means users who want a fully hands-free, eyes-up translation experience — where the caption floats in their field of view — will get it without any reinstall or setup change. The capability ships with the glasses integration.
Live Captions in Context
Clinical settings
A physician conducting a patient intake with Puente in Earbud mode can glance at the caption bar to verify the translation of a critical phrase — dosage instructions, allergy confirmation, consent language — without interrupting the conversation or asking the patient to repeat.
Classroom instruction
A teacher using Mesh Rooms for a multilingual class can see the caption output for each student’s language on the host screen, confirming that the translation of an important concept arrived correctly.
Group meetings
In Group mode with speaker diarization active, Live Captions assigns each translated segment to its speaker by name — creating an attributed, real-time transcript that participants can follow simultaneously in their own language.
Download Puente — Live Captions active from the first translation