Developersummit
  • HOME
  • SPEAKERS
  • SESSIONS
  • BUY TICKETS
  • CONTACT
saltmarch

GIDS news media, articles, insights and virtual events educate and illuminate its audiences so they can be fully prepared to deal with the new realities at work and in their professions.

Saltmarch On-Demand
Media

Our Experts

Videos On Demand

Insights

Call for Papers

Connect

About Us

Privacy Policy

Terms & Conditions

Code of Conduct

Contact Us

Subscribe to Developersummit

Get the latest event updates, and insights from today's leading voices.

© 2026-2027 Saltmarch. All rights reserved.

Engineering Reliable AI Systems: Lessons from Production Inference
RegisterTwitterLinkedInFacebook

< session />

Engineering Reliable AI Systems: Lessons from Production Inference

Wed, December 9Infrastructure, Platforms & ScaleProduction AI Systems

Most production AI failures are not caused by poor models. They emerge from the systems surrounding them.

As organizations move beyond prototypes and deploy AI across products, developer platforms, and enterprise workflows, engineering teams discover that model quality is only one part of the challenge. Long before GPU saturation becomes a concern, systems often fail because of ungoverned traffic, missing concurrency controls, poor observability, uncontrolled context growth, overloaded memory bandwidth, and weak operational guardrails. The result is increasing latency, unpredictable behaviour, rising infrastructure costs, and reduced confidence in production systems.

This session examines how production AI systems fail in layers and why reliability must be engineered across the serving stack rather than delegated to the model alone. Using real production examples, we will explore the architectural decisions that influence whether AI systems remain observable, resilient, and performant under sustained production workloads.

The discussion covers concurrency management, memory pressure, serving architecture, instrumentation, and the operational signals that help engineering teams identify problems before they develop into larger failures. Attendees will leave with a practical framework for understanding reliability in production AI systems and the systems engineering practices required to move AI from successful demonstrations to dependable operational platforms.

What You Will Learn:

  • Common failure layers that affect production AI systems beyond model quality
  • How concurrency management, memory pressure, observability, and serving architecture influence system reliability
  • A practical framework for engineering reliable AI serving platforms in production

Who Should Attend:

  • AI infrastructure engineers
  • Platform engineers
  • Infrastructure engineers
  • SREs and reliability engineers
  • Staff and principal engineers
  • Software architects
  • Technical leads responsible for production AI systems

< speaker_info />

About the speaker

Abi Aryan

Abi Aryan

AI Infrastructure Engineer and Educator

Abi Aryan is an AI infrastructure engineer at a stealth startup and educator specializing in scalable inference systems and production AI infrastructure. She spends her time helping enterprises design and optimize large-scale inferencing serving architectures, improve observability in production pipelines, and solve performance bottlenecks across distributed GPU systems.

Outside of her startup work, Abi teaches distributed systems in a university HPC program, mentors AI Engineering Team Leads through her Maven course, and is currently writing a book on GPU Engineering. Her ongoing doctoral research explores the future of adaptive AI infrastructure.

Related Talks

The Hidden Physics of AI Inference Costs

Thu, December 10

The Hidden Physics of AI Inference Costs

Gireesh Punathil
Data Access for Autonomous Systems: Where Should Execution Happen?

Wed, December 9

Data Access for Autonomous Systems: Where Should Execution Happen?

Auxten Wang
Diagnosing and Hardening Production AI Systems

Thu, December 10

Diagnosing and Hardening Production AI Systems

Abi Aryan