Voice Agent observability with LangSmith 🌟
Summary
Caroline di Vittorio, Engineer at LangChain, builds a voice agent with the Google ADK and the Gemini Live model, then sets up tracing in LangSmith to see exactly what the agent is doing under the hood. Gemini Live is Google's native audio model. It's speech-to-speech, taking audio directly as input and producing audio as output without transcribing to text, which keeps latency low and makes the agent's voice sound natural and emotive. What's covered: building a terminal-based weather assistant with two tools, defining the LangSmith Google ADK plugin, registering it…
Related: OpenAI: Mar 25How Perplexity Brought Voice Search to Millions Using the Realtime APILessons from how Perplexity Computer's voice agent was built with the Realtime API.Audio · Anthropic: Claude turns flight data into "look behind you" · Microsoft: Meet MAI-Transcribe-2: A faster and more accurate speech recognition model · xAI: Use 14 references in Grok Imagine videos
Rated low: routine. Worth knowing, not worth rearranging your day for.
Questions people ask
- Where can I read the full video?
- On Google for Developers on YouTube. The "Read this on Google for Developers on YouTube" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for Gemini?
- Caroline di Vittorio, Engineer at LangChain, builds a voice agent with the Google ADK and the Gemini Live model, then…
More from Google for Developers on YouTube 14 more
Koray Kavukcuoglu on frontier models, coding agents, and building AGI
Build a live translation broadcast app with the Gemini Live API and LiveKit
Build a live translation broadcast app with the Gemini Live API and LiveKit
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Google for Developers on YouTube.














