Session abstract
What you’ll learn
When I first started building our enterprise GenAI platform, I quickly ran into the same walls most organizations hit — teams integrating LLMs in silos, cloud costs spiraling with no visibility, and zero centralized control over how models were being used across the business. That's what pushed me to architect a proper AI Gateway layer, and honestly it changed everything. In this session I'm going to share exactly how I built a production-grade AI Gateway using Portkey, running across Azure OpenAI and Google Vertex AI simultaneously. Not a proof of concept — this is a platform that today supports multiple business units, dozens of applications, and real enterprise compliance requirements including HIPAA and Zero Trust. I'll walk you through the architecture decisions I made, the tradeoffs I ran into, and what I'd do differently. We'll cover how I set up intelligent model routing across clouds, how semantic caching brought our LLM costs down by 40 to 60 percent in production, and how I built a multi-tier fallback chain so the platform stays resilient even when a provider has an outage. I'll also talk about how we handle tenant isolation, LLM lifecycle management across teams, and keeping full observability into everything happening on the platform. What I want people to leave with is a framework they can actually use. Not slides full of architecture diagrams that look great but fall apart when you try to implement them — real patterns, real decisions, and the lessons I learned the hard way so you don't have to. If you're an architect, an engineer, or a technology leader trying to figure out how to scale AI infrastructure the right way, this talk is for you. Target Audience: Cloud architects, platform engineers, AI/ML leads, and technology executives building or scaling enterprise GenAI infrastructure.
