
Every LLM inference call runs through GPU kernels that are rarely optimized for the hardware they're running on — and that hidden inefficiency costs the industry billions. At Wafer, we build autonomous AI agents that optimize these kernels across hardware platforms. In this talk, I'll walk through what it actually looks like when AI operates at the compiler and kernel layer: how our agents discover optimization opportunities that humans miss, the real results we've shipped, and why we believe autonomous performance engineering is the key to making intelligence radically cheaper and more energy-efficient.
Emilio Andere is the Co-Founder and CEO of Wafer (YC S25), where he's building autonomous AI agents that optimize GPU kernels for inference workloads. Wafer works with leading hardware and cloud companies to make LLM inference faster and cheaper across platforms. Previously, Emilio trained weather models at Argonne National Laboratory and has published at NeurIPS.