
< session />
Thu, December 10Infrastructure, Platforms & ScaleProduction AI Systems
Engineering reliable AI systems requires more than deploying models. It requires the ability to observe, diagnose, and systematically eliminate production failures before they affect users.
This interactive workshop moves beyond architectural discussion into practical systems engineering. Through realistic production scenarios, participants will investigate failures that commonly emerge in deployed AI systems, including ungoverned traffic, missing concurrency controls, overloaded memory bandwidth, poor observability, and broken latency guarantees.
Rather than focusing on framework-specific techniques, the workshop develops the engineering judgment needed to diagnose and strengthen AI systems operating under real production conditions. Using guided examples, traces, metrics, and representative code paths, participants will identify early warning signals before failures escalate, explore serving architectures that remain resilient under changing workloads, instrument systems with meaningful observability signals, and investigate concurrency, memory, and distributed systems issues that affect AI serving platforms.
The workshop combines short technical explanations with guided diagnostic and implementation exercises, giving participants practical techniques they can apply when building and operating production AI systems in their own organizations.
What You Will Learn:
Who Should Attend:
< speaker_info />
Abi Aryan is an AI infrastructure engineer at a stealth startup and educator specializing in scalable inference systems and production AI infrastructure. She spends her time helping enterprises design and optimize large-scale inferencing serving architectures, improve observability in production pipelines, and solve performance bottlenecks across distributed GPU systems.
Outside of her startup work, Abi teaches distributed systems in a university HPC program, mentors AI Engineering Team Leads through her Maven course, and is currently writing a book on GPU Engineering. Her ongoing doctoral research explores the future of adaptive AI infrastructure.