About the speaker
Experience behind the talk
You shipped your first AI feature. It worked. Then traffic grew, and suddenly your $4/hour GPU is sitting idle most of the time, your p99 latency is awful, and your cloud bill is unhinged. What happened? This talk is about the layer between your model and your users — the inference serving layer th...

