LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.
Our research introduces Visual AI Testing Agent, an agentic framework designed to revolutionize mobile UI automation for iOS Simulators by leveraging Vision Language Models (VLMs). Unlike traditional locator-based tools like Appium or XCUITest, which often fail due to unstable identifiers or dynamic UI changes, our VLM-driven approach treats the screen as the primary interface. The agent captures the screen, processes visible elements through vision-based providers, and executes actions with human-like visual understanding. This talk will demonstrate how this agentic architecture enables predictive test case generation and provides development teams with actionable, early-stage insights. By shifting to a vision-first paradigm, we enable more robust, resilient, and intelligent automated testing cycles.... Read more
The typical default is one call to one model, and when that gets expensive, we cascade to something cheaper. Both accept the same premise: the input is one question. Usually, it isn't. A single photo, document or utterance carries several distinct questions, and one model answering all of them hands you its weakest answer in its most confident voice, at full price, every time. A cascade only changes the price. This talk makes the case for treating the input as a routing problem. Cheap classifiers decide which question you have, and each question goes to the model genuinely best at it, including the times the honest answer is that no model should answer. In our case we brought token spend under control, took a seventeen second wait down to under a second on the path that can prove its answer, and cut wrong answers from 25% to near zero.... Read more
The Ogummaa Agent Portal is a comprehensive, sovereign AI workspace designed for secure, long-horizon task automation and conversational knowledge retrieval. The platform utilizes a decoupled microservices architecture—comprising an Admin Portal, Central Gateway, and Ingestion Pipeline—to ensure elastic scalability and rigorous environmental isolation. Its core engine, a GraphRAG (Graph Retrieval-Augmented Generation) and VectorRAG system for analytical graph processing, geo-spacial-temporal correlation, enabling sophisticated retrieval across massive enterprise datasets. The portal enforces a robust security paradigm through per-user sandbox isolation, a four-layer guardrail system, and a "bring your own secrets" model, while supporting unattended execution for recurring routines. By blending autonomous self-improvement loops with immutable event-based telemetry, and provides an enterprise-ready infrastructure for organizations prioritizing data sovereignty, multi-agent orchestration, and resilient AI deployment.... Read more
LLM systems are getting very good at extracting entities from messy documents. But extracting a mention is only the beginning. The harder problem is deciding what that mention actually refers to. A single real-world entity can appear under different names, formats, abbreviations, and contexts across documents and databases. At the same time, different entities can look remarkably similar. A reliable system therefore needs more than entity extraction or similarity search. It needs a mapping step that decides which candidate, if any, represents the entity being described. This is where similarity becomes both useful and insufficient. Database search, lexical matching, and embeddings are powerful ways to generate and rank candidates. But a high similarity score does not make a candidate the correct mapping. The system still has to account for ambiguity, context, available evidence, and what happens when no candidate meets the evidence threshold. Sometimes the right decision is not to map at all. This talk looks at entity mapping as a distinct engineering problem in an LLM pipeline: using structured search and semantic similarity to find plausible candidates, understanding where those signals break down, and designing the decision layer that determines when a candidate is good enough to map, and when the system should abstain. The goal is not to replace similarity with something else. It is to understand where similarity fits, and where the identity decision begins.... Read more
Most LLM evaluation assumes a gold answer exists to score against. A large and growing share of production work, however, doesn't have one: contract review, clinical documentation, incident summaries, moderation calls, research synthesis — any task where correctness is a matter of expert judgment, and where two qualified experts given the same input produce different outputs that are both defensible. If you score that as if a single right answer existed, you collapse the errors that matter — a fabricated fact, a dropped material detail — into the same number as the ones that don't, like structure, emphasis, and phrasing. This talk is about building evaluation for those systems, where the ground truth isn't a fixed key but a distribution of expert opinion, and where disagreement is built into the problem rather than a symptom of bad labeling. It unpacks three moves that improve accuracy and trustworthiness: treating inter-rater reliability among your experts as the ceiling on any eval you can build, using expert corrections as a live quality signal instead of a static gold set, and separating "genuinely wrong" from "differently right" so the score tracks the failures you actually care about.... Read more
AI systems can produce impressive demos and still fail when exposed to real-world complexity, edge cases, changing data, and unexpected user behavior. Semiconductor engineers have spent decades addressing a similar problem through rigorous verification, coverage analysis, corner-case testing, fault injection, regression, and disciplined sign-off. This talk explores how those principles can be adapted to the evaluation of LLMs and increasingly autonomous AI agents, where a successful prompt or benchmark score is not sufficient evidence of system reliability. Attendees will leave with a practical framework for moving from “the AI works” to “we have evidence that this AI system is ready to deploy.”... Read more
No matter how well an AI pipeline is designed, it will always be limited by both quality and quantity of available data. However, real-world data rarely arrives nicely packaged and pre-normalized. In fact, it is probably safer to say that real-world data is the number one saboteur of existing machine learning applications. This talk will go over common data problems that have been seen in enterprise systems and the importance of curating data for pipelines of all varieties.... Read more
Self-managed AI model deployments on-prem and in private cloud infrastructure are becoming increasingly popular, accelerated by open-weight models that rival frontier models. But deploying them needs accelerators (such as GPUs) with large and expensive memory. The usual ways to save that memory quantize or approximate the model, thereby changing its output and affecting accuracy. This talk removes that tradeoff. It covers lossless compression of both model weights and the KV cache, across BF16, FP16, FP8, and FP4, delivering around 30% more GPU headroom with identical model output. Weights are bit-for-bit exact with cryptographic hash verification, and stay compressed in GPU memory while they run. Freeing weight memory leaves more room for the KV cache, and the KV cache is then compressed on top. In deployments that are already memory-constrained, this results in cases of up to 110% more KV cache headroom, raising usable context length, batch size, and concurrency. The talk frames bit-exact KV cache compression as the next lever for lossless inference efficiency.... Read more
Developer docs are now read by agents as much as humans, but most teams still judge them by gut feel and page views. This talk shows how to treat documentation as a system under test: rubric-based LLM-as-judge scoring calibrated against human labels, sandboxed agents that attempt your quickstarts end-to-end and report exactly where they fail, and the same suite run against competitors' docs for an apples-to-apples benchmark. You'll leave with a pipeline architecture you can build in a week, rubric patterns that produce stable scores, and a way to turn "our docs vs. theirs" from an argument into data.... Read more
Most teams building with LLMs meet a point where inference cost stops tracking with usage, and the reasons are invisible from the application layer. The constraints that actually set your cost per token are physical — memory bandwidth, power draw per rack, thermal headroom, and how badly your workload underutilizes the accelerator you're paying for. This talk walks through what happens underneath an inference request at the hardware level, why batching and context length hit economic cliffs rather than smooth curves, and which of those limits are architectural versus fixable in software. Attendees should leave able to read their own inference spend as a systems problem instead of a line item, and to tell which optimizations are worth engineering time.... Read more
This talk is a field report on using Grok Bot to sell commercial real estate. I’ll show the live stack I used on real deals: Grok Bot as the operating layer across buyer inbound and follow-up, Quo for the sales phone line and texts, a dedicated /buy sales website and deal room, Google Drive diligence packs, and offering memorandums, all powered by Grok Bot. We’ll cover what the agent actually does day to day (watching inbound, drafting first-touch that leads with income, logging buyer calls, keeping board copy and the OM in sync, chasing listing-agreement issues) and where humans still have to approve sends. Expect concrete CRE examples, the connectors and artifacts that made it work, and the failure modes that show up when an agent touches real buyers.... Read more
Autonomous research agents that propose, implement, and refine ML solutions have gotten remarkably good. They have also inherited an assumption that excludes most teams: a frontier model drives every step of the loop. That assumption matters because loop cost, not model quality, is becoming the limit on how much autonomy a team can afford to run. This talk explores inverting it. A small open-weight model runs the loop, and a frontier model is called in only as an advisor, on a metered budget. We'll cover what the literature establishes about small and large model collaboration, including step-level escalation, agent distillation, and budget allocation, and where it stops short for long-horizon loops, whose failure modes are not bad tool calls but dead branches, validation leaks, and unclear stopping points. We'll look at a framework for deciding when advice is worth buying, how much of a trajectory an advisor needs to see, and how to measure guidance rather than assume its value. We will walk through a recorded run showing the loop, the trigger, and the cost meter together. The session concludes with open problems, including asynchronous advisors, persistent guidance, and what cost-aware autonomy means for teams without frontier budgets.... Read more
Most agentic AI systems work in the demo and fail in production. The failure is rarely the model. It is the cost curve that only appears at scale, the coordination breakdown when single agents become fleets, the governance gap that surfaces during an audit, and the operator trust that never gets earned. This talk is a practitioner's field guide to the last mile of agentic AI: the engineering and operational work that determines whether an LLM system delivers value or gets quietly shelved sixty days after launch. Drawing on real deployment patterns across industrial and enterprise environments, we will cover why token consumption scales with autonomous actions rather than users and how to model the production bill before you sign; what breaks when you move from one agent to a fleet and how to design the coordination layer; the audit and containment questions every production agent must answer; and why explainability at the moment of decision is the real driver of adoption. Attendees will leave with a concrete checklist for pressure-testing an agentic system before it reaches production, and a sharper sense of where the genuine engineering risk lives in the shift from pilots to scaled deployment.... Read more
My wife, who doesn’t have any technical background, built a real app with AI, that she uses every day. So if someone with no previous experience can build working software now, what’s left for software engineers? During this presentation we'll explore what makes a software engineer relevant in the AI era. An engineer needs to know foundations, judgment, responsibility... Programming is easier. Software engineering isn’t.... Read more
Everyone has an AI cod ing anecdote: an agent nailed one ticket and lost the plot on another. But which tasks can your team reliably hand off, and what makes the resulting PR worth merging? At Zello, we started answering those questions with our own engineering history. We turned solved tickets into a benchmark and merged PRs into a landscape of task classes. Then we connected classification, runnable environments, coding agents and verification into a cloud workflow that produces PRs for human review. This talk follows the journey from benchmarking agents to building a software factory. I’ll show why environment setup matters as much as model choice, how we choose suitable work, and how engineer feedback closes two loops: improving the current PR and teaching the factory something useful for the next ticket.... Read more
18:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
Visual AI Testing Agent: Screenshot-First Mobile UI Test Automation for iOS Simulator
Abstract
Our research introduces Visual AI Testing Agent, an agentic framework designed to revolutionize mobile UI automation for iOS Simulators by leveraging Vision Language Models (VLMs). Unlike traditional locator-based tools like Appium or XCUITest, which often fail due to unstable identifiers or dynamic UI changes, our VLM-driven approach treats the screen as the primary interface. The agent captures the screen, processes visible elements through vision-based providers, and executes actions with human-like visual understanding. This talk will demonstrate how this agentic architecture enables predictive test case generation and provides development teams with actionable, early-stage insights. By shifting to a vision-first paradigm, we enable more robust, resilient, and intelligent automated testing cycles.
Bio
Rahul Azmeera, PhD, is a Senior Software Engineer at 7-Eleven and an Independent Researcher specializing in AI and software engineering. He holds a PhD in Information Technology from the University of The Cumberlands and a Master's degree in Technology from Pittsburg State University.
With extensive hands-on experience in full-stack web application development, AWS cloud architecture, UX design, and QA, he specializes in mobile application development using hybrid frameworks like React Native and Ionic. His current research focuses on vision-language models for intelligent software testing, agentic AI, and analyzing the root causes of software errors.
Geoff Niehaus
Today Mechanic
One LLM Call Is the Wrong Default
Abstract
The typical default is one call to one model, and when that gets expensive, we cascade to something cheaper. Both accept the same premise: the input is one question. Usually, it isn't. A single photo, document or utterance carries several distinct questions, and one model answering all of them hands you its weakest answer in its most confident voice, at full price, every time. A cascade only changes the price. This talk makes the case for treating the input as a routing problem. Cheap classifiers decide which question you have, and each question goes to the model genuinely best at it, including the times the honest answer is that no model should answer. In our case we brought token spend under control, took a seventeen second wait down to under a second on the path that can prove its answer, and cut wrong answers from 25% to near zero.
Bio
Geoff Niehaus is the founder of Today Mechanic, an AI-powered platform that keeps auto repair shops' bays full by handling the customer conversation before the car ever reaches the lift. He spent the last four years at CVS Health (via Oak Street Health), most recently as Executive Director of Software Development Engineering, where he oversaw a program of 100+ engineers and led delivery of a production applied-AI system (retrieval over millions of clinical documents feeding an LLM) that lifted suspect-finding identification by 30% under HIPAA constraints and human-in-the-loop clinical validation. He also embedded AI-assisted development into the standard SDLC, pairing coding agents with a contract-first "Design First" practice that took code coverage from 40% to 80%+. He's a founding member of the Austin CTO Club and holds a pending patent for multi-tenant medical record access.
Rodney Richard Puplampu
Puplampu Consulting
Ogummaa Agent Portal. Architecture & Security Blueprint for Long-Horizon Agents
Abstract
The Ogummaa Agent Portal is a comprehensive, sovereign AI workspace designed for secure, long-horizon task automation and conversational knowledge retrieval. The platform utilizes a decoupled microservices architecture—comprising an Admin Portal, Central Gateway, and Ingestion Pipeline—to ensure elastic scalability and rigorous environmental isolation. Its core engine, a GraphRAG (Graph Retrieval-Augmented Generation) and VectorRAG system for analytical graph processing, geo-spacial-temporal correlation, enabling sophisticated retrieval across massive enterprise datasets. The portal enforces a robust security paradigm through per-user sandbox isolation, a four-layer guardrail system, and a "bring your own secrets" model, while supporting unattended execution for recurring routines. By blending autonomous self-improvement loops with immutable event-based telemetry, and provides an enterprise-ready infrastructure for organizations prioritizing data sovereignty, multi-agent orchestration, and resilient AI deployment.
Bio
Rodney Puplampu, Principal AI Platform Engineer, Puplampu Consulting, author of What Everyone Should Know About the Rise of AI, creator of the Ogummaa Sovereign AI Platform, born in Detroit, Michigan in 1974, is a passionate advocate for safe, transparent, and ethical Artificial Intelligence (AI). With over thirty years of experience in network, security, telecom, full-stack applications, and sovereign agentic AI platforms, Rodney has held positions in the United States Army and worked for renowned companies such as AT&T, Lucent Technologies, Avaya, and HP. Throughout his career, he has consistently demonstrated his commitment to ensuring responsible and ethical business practices.
Rodney also has a passion for cleaning the environment. This extends to his advocacy for AI because he believes AI has the potential to help mankind address global challenges such as climate change and sustainable resource management.
His dedication to safety and transparency in AI is evident in his volunteer work and service to countless causes. He strongly believes that AI should be designed with safeguards to prevent it from being used for malicious purposes or to perpetuate discrimination. Rodney recognizes that AI should complement and enhance human capabilities rather than replace them.
Maryam Astero
Genomenon
One Thing, Many Names: Building Reliable Entity Mapping
Abstract
LLM systems are getting very good at extracting entities from messy documents. But extracting a mention is only the beginning. The harder problem is deciding what that mention actually refers to.
A single real-world entity can appear under different names, formats, abbreviations, and contexts across documents and databases. At the same time, different entities can look remarkably similar. A reliable system therefore needs more than entity extraction or similarity search. It needs a mapping step that decides which candidate, if any, represents the entity being described.
This is where similarity becomes both useful and insufficient. Database search, lexical matching, and embeddings are powerful ways to generate and rank candidates. But a high similarity score does not make a candidate the correct mapping. The system still has to account for ambiguity, context, available evidence, and what happens when no candidate meets the evidence threshold. Sometimes the right decision is not to map at all.
This talk looks at entity mapping as a distinct engineering problem in an LLM pipeline: using structured search and semantic similarity to find plausible candidates, understanding where those signals break down, and designing the decision layer that determines when a candidate is good enough to map, and when the system should abstain.
The goal is not to replace similarity with something else. It is to understand where similarity fits, and where the identity decision begins.
Bio
Maryam Astero is a Senior AI Research Engineer at Genomenon, where she builds production systems that turn unstructured scientific literature into structured, queryable knowledge. Her work spans document ranking, information extraction, entity resolution, and knowledge graph construction, with a focus on making these systems reliable beyond the prototype stage.
She holds a PhD in Computer Science from Aalto University, where she developed graph neural network architectures for chemical reaction modeling. She works at the meeting point of machine learning, graph-based methods, and production systems, particularly where ambiguity and imperfect data make seemingly simple problems hard to evaluate and scale.
Leila Anderson
Upheal
Evaluation Without Ground Truth: Lessons from Expert Disagreement
Abstract
Most LLM evaluation assumes a gold answer exists to score against. A large and growing share of production work, however, doesn't have one: contract review, clinical documentation, incident summaries, moderation calls, research synthesis — any task where correctness is a matter of expert judgment, and where two qualified experts given the same input produce different outputs that are both defensible. If you score that as if a single right answer existed, you collapse the errors that matter — a fabricated fact, a dropped material detail — into the same number as the ones that don't, like structure, emphasis, and phrasing.
This talk is about building evaluation for those systems, where the ground truth isn't a fixed key but a distribution of expert opinion, and where disagreement is built into the problem rather than a symptom of bad labeling. It unpacks three moves that improve accuracy and trustworthiness: treating inter-rater reliability among your experts as the ceiling on any eval you can build, using expert corrections as a live quality signal instead of a static gold set, and separating "genuinely wrong" from "differently right" so the score tracks the failures you actually care about.
Bio
Leila Anderson is a Senior Clinical AI Product Manager at Upheal, where she builds and evaluates the LLM systems behind clinical documentation for behavioral health. She's also a licensed marriage and family therapist (LMFT-S) with nearly a decade of clinical experience, which puts her on both sides of the eval problem this talk is about: the expert whose judgment the model is trying to match, and the person on the hook for deciding whether it did. Alongside her work at Upheal she runs a private psychotherapy practice and serves as an expert witness in clinical cases — two more roles where being the arbiter of a contested-but-defensible judgment is the whole job.
Felix Njeh
Nelix
Your AI Passed the Demo. But Did It Pass Validation?
Abstract
AI systems can produce impressive demos and still fail when exposed to real-world complexity, edge cases, changing data, and unexpected user behavior. Semiconductor engineers have spent decades addressing a similar problem through rigorous verification, coverage analysis, corner-case testing, fault injection, regression, and disciplined sign-off. This talk explores how those principles can be adapted to the evaluation of LLMs and increasingly autonomous AI agents, where a successful prompt or benchmark score is not sufficient evidence of system reliability. Attendees will leave with a practical framework for moving from “the AI works” to “we have evidence that this AI system is ready to deploy.”
Bio
Dr. Felix Njeh is a computer engineer, AI researcher, professor, author, and semiconductor engineering leader with decades of experience spanning hardware and system validation, AI/ML, edge computing, and intelligent systems. His engineering work includes semiconductor validation and verification at Intel and AI/ML systems research supporting advanced defense applications. He teaches graduate-level computing courses and is the author of Beyond Limits: AI-Powered Edge Architecture for Smart Devices and AI Unlocked: Harnessing Machines Without Losing Your Humanity. He is also Founder of the AI Crusaders Global Network™, an international initiative advancing practical, responsible, and human-centered AI education and adoption.
Colby Mainard
RiskScout
Invisible Data - The Largest Problem Facing AI
Abstract
No matter how well an AI pipeline is designed, it will always be limited by both quality and quantity of available data. However, real-world data rarely arrives nicely packaged and pre-normalized. In fact, it is probably safer to say that real-world data is the number one saboteur of existing machine learning applications. This talk will go over common data problems that have been seen in enterprise systems and the importance of curating data for pipelines of all varieties.
Bio
Colby Mainard is a machine learning engineer with more than five years of experience designing, deploying, and operating production ML systems, currently working as an AI/ML engineer at RiskScout on transaction anomaly detection, entity risk scoring, and network analysis across financial data. He holds a Master of Computer Science from Texas A&M University with a focus on artificial intelligence, machine learning, deep learning, and computer vision, along with a B.S. in Computer Science from the same institution with minors in business and cybersecurity. Earlier roles span sports analytics and cybersecurity, including computer vision pipelines for the NFL, NHL, MLB, and NBA at MVP, as well as Python and C++ data tooling for threat detection at Vectra. Outside of work he studies quantum computing, plays Dungeons and Dragons, enjoys learning about history, and shoots landscape photography.
Kunle Olutomilayo
Isiro AI
30% More KV Cache Headroom, 100% Accuracy
Abstract
Self-managed AI model deployments on-prem and in private cloud infrastructure are becoming increasingly popular, accelerated by open-weight models that rival frontier models. But deploying them needs accelerators (such as GPUs) with large and expensive memory. The usual ways to save that memory quantize or approximate the model, thereby changing its output and affecting accuracy.
This talk removes that tradeoff. It covers lossless compression of both model weights and the KV cache, across BF16, FP16, FP8, and FP4, delivering around 30% more GPU headroom with identical model output. Weights are bit-for-bit exact with cryptographic hash verification, and stay compressed in GPU memory while they run. Freeing weight memory leaves more room for the KV cache, and the KV cache is then compressed on top. In deployments that are already memory-constrained, this results in cases of up to 110% more KV cache headroom, raising usable context length, batch size, and concurrency. The talk frames bit-exact KV cache compression as the next lever for lossless inference efficiency.
Bio
Kunle Olutomilayo is the Founder of Isiro AI, lowering the cost of ownership for self-managed AI infrastructure. The company focuses on lossless reduction of the hardware memory required for model deployments, without changing the model output. Kunle's core thesis is that the next wave of inference cost savings will not only come from faster chips or smaller models, but also from kernel software breakthroughs that move less data to address the memory wall problem, without changing what the model produces. This is especially important in regulated and high-stakes environments, such as finance, health care, defense, and enterprise AI, where cost savings must not come at the expense of predictable model behavior. Kunle holds a PhD in Electrical Engineering and previously worked on advanced AI and vehicle autonomy at Ford Greenfield Labs.
Jessie Mongeon
Mysten Labs
AI Eval Pipelines for Documentation Quality and Competitor Analysis
Abstract
Developer docs are now read by agents as much as humans, but most teams still judge them by gut feel and page views. This talk shows how to treat documentation as a system under test: rubric-based LLM-as-judge scoring calibrated against human labels, sandboxed agents that attempt your quickstarts end-to-end and report exactly where they fail, and the same suite run against competitors' docs for an apples-to-apples benchmark. You'll leave with a pipeline architecture you can build in a week, rubric patterns that produce stable scores, and a way to turn "our docs vs. theirs" from an argument into data.
Bio
Jessie Mongeon is a technology writer, educator, and author with a background in software engineering, artificial intelligence, emerging technologies, and developer education. She holds a Master of Science in Information Technology Management from Western Governors University and is currently working on a second Master of Science in Artificial Intelligence Software Engineering from Quantic School of Business and Technology.
Yuri Streciwik
Dell Technologies
The Inference Tax: Why Your Token Costs Stop Making Sense at Scale
Abstract
Most teams building with LLMs meet a point where inference cost stops tracking with usage, and the reasons are invisible from the application layer. The constraints that actually set your cost per token are physical — memory bandwidth, power draw per rack, thermal headroom, and how badly your workload underutilizes the accelerator you're paying for. This talk walks through what happens underneath an inference request at the hardware level, why batching and context length hit economic cliffs rather than smooth curves, and which of those limits are architectural versus fixable in software. Attendees should leave able to read their own inference spend as a systems problem instead of a line item, and to tell which optimizations are worth engineering time.
Bio
Yuri Streciwik is a Senior Product Manager for AI Server and GPU Systems at Dell, where he is a Sr Product Manager for a rack-scale GPU-dense server. He works across silicon partners, thermal and power design, and enterprise GenAI deployment. Before moving into datacenter infrastructure product management he spent seven years in engineering across hydro, wind, and oil and gas in Brazil, China, and France. He holds an MBA from Duke Fuqua and writes Flesh & Inference, a newsletter on AI infrastructure.
Matt Barge
Superbuilders
Selling Commercial Real Estate with Grok Bot
Abstract
This talk is a field report on using Grok Bot to sell commercial real estate. I’ll show the live stack I used on real deals: Grok Bot as the operating layer across buyer inbound and follow-up, Quo for the sales phone line and texts, a dedicated /buy sales website and deal room, Google Drive diligence packs, and offering memorandums, all powered by Grok Bot.
We’ll cover what the agent actually does day to day (watching inbound, drafting first-touch that leads with income, logging buyer calls, keeping board copy and the OM in sync, chasing listing-agreement issues) and where humans still have to approve sends. Expect concrete CRE examples, the connectors and artifacts that made it work, and the failure modes that show up when an agent touches real buyers.
Bio
Matt Barge is a Forward Deployed AI Engineer at Superbuilders, based in Austin, Texas. He builds production AI systems for go-to-market and operations, and manages and sells commercial real estate assets in Texas. He previously built iOS products at MoneyGram and has spent the last several years putting LLM agents to work. More at mattbarge.com.
Vashishtha Patil
Amazon Lab126
Auto-Research on a Budget: Small Models in the Loop, Frontier Models on Call
Abstract
Autonomous research agents that propose, implement, and refine ML solutions have gotten remarkably good. They have also inherited an assumption that excludes most teams: a frontier model drives every step of the loop. That assumption matters because loop cost, not model quality, is becoming the limit on how much autonomy a team can afford to run. This talk explores inverting it. A small open-weight model runs the loop, and a frontier model is called in only as an advisor, on a metered budget. We'll cover what the literature establishes about small and large model collaboration, including step-level escalation, agent distillation, and budget allocation, and where it stops short for long-horizon loops, whose failure modes are not bad tool calls but dead branches, validation leaks, and unclear stopping points. We'll look at a framework for deciding when advice is worth buying, how much of a trajectory an advisor needs to see, and how to measure guidance rather than assume its value. We will walk through a recorded run showing the loop, the trigger, and the cost meter together. The session concludes with open problems, including asynchronous advisors, persistent guidance, and what cost-aware autonomy means for teams without frontier budgets.
Bio
Vashishtha Patil is a Senior Applied Scientist at Amazon Lab126, specializing in LLM-powered AI for the Alexa+ Smart Home experience. With over a decade of experience across Amazon and Qualcomm, his expertise spans machine learning, computer vision, and algorithm development for smart home, healthtech, and mobile devices. He holds a Master's degree from the University of Southern California.
James Woodard
Apex Systems
Operators, Agents, and the 60-Day Cliff: Deploying LLM Systems That Get Used
Abstract
Most agentic AI systems work in the demo and fail in production. The failure is rarely the model.
It is the cost curve that only appears at scale, the coordination breakdown when single agents become fleets, the governance gap that surfaces during an audit, and the operator trust that never gets earned. This talk is a practitioner's field guide to the last mile of agentic AI: the engineering and operational work that determines whether an LLM system delivers value or gets quietly shelved sixty days after launch.
Drawing on real deployment patterns across industrial and enterprise environments, we will cover why token consumption scales with autonomous actions rather than users and how to model the production bill before you sign; what breaks when you move from one agent to a fleet and how to design the coordination layer; the audit and containment questions every production agent must answer; and why explainability at the moment of decision is the real driver of adoption. Attendees will leave with a concrete checklist for pressure-testing an agentic system before it reaches production, and a sharper sense of where the genuine engineering risk lives in the shift from pilots to scaled deployment.
Bio
James Woodard is an AI Solutions Architect focused on deploying agentic and generative AI systems in industrial and enterprise environments across energy, manufacturing, and defense.
His work centers on the last mile of AI: the integration, governance, cost, and adoption work
that determines whether a deployment succeeds in production rather than stalling after launch.
He brings a rare vantage point to the deployment problem, having worked as an operator
selecting and living with technology, as a vendor designing and deploying it, and as a builder
engineering it. He writes weekly on industrial AI deployment and holds an MBA from UT Austin
and a degree in mechatronics engineering from Georgia Tech.
Nate Lemos
MakingFriends.app
My wife is a vibe-coder. So, what now?
Abstract
My wife, who doesn’t have any technical background, built a real app with AI, that she uses every day. So if someone with no previous experience can build working software now, what’s left for software engineers?
During this presentation we'll explore what makes a software engineer relevant in the AI era. An engineer needs to know foundations, judgment, responsibility...
Programming is easier. Software engineering isn’t.
Bio
Nate Lemos is the co-founder of MakingFriends.app and an experienced engineer with over 12 years building systems for all sizes of business. Nate was a key contributor of the Halo Infinite backend and migration of the Azure Portal to Typescript. Now, Nate writes enterprise AI agents for thousands of users at ECI and runs his own app to help people meet IRL.
Nikita Pestrov
Zello
Building a Software Factory: Your Tickets Are the Benchmarks
Abstract
Everyone has an AI cod ing anecdote: an agent nailed one ticket and lost the plot on another. But which tasks can your team reliably hand off, and what makes the resulting PR worth merging? At Zello, we started answering those questions with our own engineering history. We turned solved tickets into a benchmark and merged PRs into a landscape of task classes. Then we connected classification, runnable environments, coding agents and verification into a cloud workflow that produces PRs for human review. This talk follows the journey from benchmarking agents to building a software factory. I’ll show why environment setup matters as much as model choice, how we choose suitable work, and how engineer feedback closes two loops: improving the current PR and teaching the factory something useful for the next ticket.
Bio
Nikita Pestrov is a Data and AI Platform Lead at Zello with over 10 years of experience building platforms for AI and analytics products. He is responsible for AI-agent evaluation, observability and orchestration infrastructure, alongside the underlying data platform. Previously, he delivered distributed, cloud-native data products for clients including Mastercard, the World Bank and PwC. He has spoken at Iceberg Summit and Snowflake and dbt meetups, and is an experienced mentor, hackathon judge and university lecturer.