Everyone has an AI cod ing anecdote: an agent nailed one ticket and lost the plot on another. But which tasks can your team reliably hand off, and what makes the resulting PR worth merging? At Zello, we started answering those questions with our own engineering history. We turned solved tickets into a benchmark and merged PRs into a landscape of task classes. Then we connected classification, runnable environments, coding agents and verification into a cloud workflow that produces PRs for human review. This talk follows the journey from benchmarking agents to building a software factory. I’ll show why environment setup matters as much as model choice, how we choose suitable work, and how engineer feedback closes two loops: improving the current PR and teaching the factory something useful for the next ticket.
Nikita Pestrov is a Data and AI Platform Lead at Zello with over 10 years of experience building platforms for AI and analytics products. He is responsible for AI-agent evaluation, observability and orchestration infrastructure, alongside the underlying data platform. Previously, he delivered distributed, cloud-native data products for clients including Mastercard, the World Bank and PwC. He has spoken at Iceberg Summit and Snowflake and dbt meetups, and is an experienced mentor, hackathon judge and university lecturer.