Dallas skyline illustration

CYC26 / AI

Advanced RAG Systems: Scaling Retrieval-Augmented Generation in Production

This session focuses on designing and scaling high-performance Retrieval-Augmented Generation (RAG) systems in real-world applications. We will cover practical techniques used in industry and research to improve LLM accuracy, latency, and reliability, including HyDE, multi-query retrieval, contextual compression, reranking (including custom strategies), and hybrid search (text plus semantic). We will also explore advanced architectures such as Graph RAG, Light RAG, and DAG-based evaluation pipelines. The talk will go beyond theory and dive into system design decisions across ingestion, chunking, retrieval, generation, and evaluation. I will share insights from working with large-scale document pipelines and lessons learned from production systems used by leading companies and fast-growing startups. This session is useful for engineers building RAG systems, as well as job seekers preparing for AI engineering and forward-deployed roles, where system design and practical tradeoffs are heavily tested.

Session abstract

What you’ll learn

This session focuses on designing and scaling high-performance Retrieval-Augmented Generation (RAG) systems in real-world applications. We will cover practical techniques used in industry and research to improve LLM accuracy, latency, and reliability, including HyDE, multi-query retrieval, contextual compression, reranking (including custom strategies), and hybrid search (text plus semantic). We will also explore advanced architectures such as Graph RAG, Light RAG, and DAG-based evaluation pipelines. The talk will go beyond theory and dive into system design decisions across ingestion, chunking, retrieval, generation, and evaluation. I will share insights from working with large-scale document pipelines and lessons learned from production systems used by leading companies and fast-growing startups. This session is useful for engineers building RAG systems, as well as job seekers preparing for AI engineering and forward-deployed roles, where system design and practical tradeoffs are heavily tested.