Session abstract
What you’ll learn
What does it take to power millions of voice conversations at ChatGPT's scale, and how can you build the same experiences into your own apps? When OpenAI needed infrastructure for ChatGPT's Advanced Voice Mode, they built it on LiveKit's open source real-time platform. In this talk, I'll break down how we built that infrastructure, then show you how the same open source stack makes it easy to add voice-native functionality to any application, no matter your language or framework. Here's what we'll dig into: - **Behind the Scenes of ChatGPT Voice.** Architecture decisions, scaling challenges, and real production metrics. - **Why Open Source Won.** The technical and business reasons behind choosing transparent, community-driven AI infrastructure. - **Building Voice-Native Apps.** Going from zero to a working voice agent in your stack using the same tools. - **Live Demo.** A walkthrough of a real voice-native app, how it's built and how it works under the hood. - **Lessons and Trade-offs.** What we learned scaling to millions of calls and what you should know before you start. Expect a technical deep dive paired with a live demo of a real voice-native application. You'll walk away understanding how large-scale voice AI works and with the practical knowledge to start building voice-native features today. The infrastructure behind ChatGPT's voice mode is open source, and it's ready for you to use. Notes for Reviewers This talk is designed for a polyglot developer audience. The concepts and architecture patterns are language-agnostic, and the live demo will showcase a real voice-native app along with a breakdown of how it's built, making it accessible to developers across any stack.
