LLMday

Large Language Models, Agents & AI Systems

October 14, 2026 The Sunset Room, Austin, Texas, USA

1
Day
10+
Speakers
1
Track
100+
Attendees

LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.

Companies presenting:

7-Eleven, Amazon Lab126, Apex Systems, Dell Technologies, Genomenon, Isiro AI, MakingFriends.app, Mysten Labs, Nelix, Puplampu Consulting, RiskScout, Superbuilders, Today Mechanic, Upheal, Zello

Topics so far:

Event Starts In:

Tickets

Schedule

October 14, 2026 • single track • 9AM - 6:30PM • Austin, in-person
View
Track 1 • main room

09:00

Registration and coffee

Main lobby

09:30

Rahul Azmeera

Visual AI Testing Agent: Screenshot-First Mobile UI Test Automation for iOS Simulator

7-Eleven
Our research introduces Visual AI Testing Agent, an agentic framework designed to revolutionize mobile UI automation for iOS Simulators by leveraging Vision Language Models (VLMs). Unlike traditional locator-based tools like Appium or XCUITest, which often fail due to unstable identifiers or dynamic UI changes, our VLM-driven approach treats the screen as the primary interface. The agent captures the screen, processes visible elements through vision-based providers, and executes actions with human-like visual understanding. This talk will demonstrate how this agentic architecture enables predictive test case generation and provides development teams with actionable, early-stage insights. By shifting to a vision-first paradigm, we enable more robust, resilient, and intelligent automated testing cycles.... Read more

10:00

Geoff Niehaus

One LLM Call Is the Wrong Default

Today Mechanic
The typical default is one call to one model, and when that gets expensive, we cascade to something cheaper. Both accept the same premise: the input is one question. Usually, it isn't. A single photo, document or utterance carries several distinct questions, and one model answering all of them hands you its weakest answer in its most confident voice, at full price, every time. A cascade only changes the price. This talk makes the case for treating the input as a routing problem. Cheap classifiers decide which question you have, and each question goes to the model genuinely best at it, including the times the honest answer is that no model should answer. In our case we brought token spend under control, took a seventeen second wait down to under a second on the path that can prove its answer, and cut wrong answers from 25% to near zero.... Read more

10:30

Rodney Richard Puplampu

Ogummaa Agent Portal. Architecture & Security Blueprint for Long-Horizon Agents

Puplampu Consulting
The Ogummaa Agent Portal is a comprehensive, sovereign AI workspace designed for secure, long-horizon task automation and conversational knowledge retrieval. The platform utilizes a decoupled microservices architecture—comprising an Admin Portal, Central Gateway, and Ingestion Pipeline—to ensure elastic scalability and rigorous environmental isolation. Its core engine, a GraphRAG (Graph Retrieval-Augmented Generation) and VectorRAG system for analytical graph processing, geo-spacial-temporal correlation, enabling sophisticated retrieval across massive enterprise datasets. The portal enforces a robust security paradigm through per-user sandbox isolation, a four-layer guardrail system, and a "bring your own secrets" model, while supporting unattended execution for recurring routines. By blending autonomous self-improvement loops with immutable event-based telemetry, and provides an enterprise-ready infrastructure for organizations prioritizing data sovereignty, multi-agent orchestration, and resilient AI deployment.... Read more

11:00

Maryam Astero

One Thing, Many Names: Building Reliable Entity Mapping

Genomenon
LLM systems are getting very good at extracting entities from messy documents. But extracting a mention is only the beginning. The harder problem is deciding what that mention actually refers to. A single real-world entity can appear under different names, formats, abbreviations, and contexts across documents and databases. At the same time, different entities can look remarkably similar. A reliable system therefore needs more than entity extraction or similarity search. It needs a mapping step that decides which candidate, if any, represents the entity being described. This is where similarity becomes both useful and insufficient. Database search, lexical matching, and embeddings are powerful ways to generate and rank candidates. But a high similarity score does not make a candidate the correct mapping. The system still has to account for ambiguity, context, available evidence, and what happens when no candidate meets the evidence threshold. Sometimes the right decision is not to map at all. This talk looks at entity mapping as a distinct engineering problem in an LLM pipeline: using structured search and semantic similarity to find plausible candidates, understanding where those signals break down, and designing the decision layer that determines when a candidate is good enough to map, and when the system should abstain. The goal is not to replace similarity with something else. It is to understand where similarity fits, and where the identity decision begins.... Read more

11:30

Lunch & networking

Main lobby

12:30

Leila Anderson

Evaluation Without Ground Truth: Lessons from Expert Disagreement

Upheal
Most LLM evaluation assumes a gold answer exists to score against. A large and growing share of production work, however, doesn't have one: contract review, clinical documentation, incident summaries, moderation calls, research synthesis — any task where correctness is a matter of expert judgment, and where two qualified experts given the same input produce different outputs that are both defensible. If you score that as if a single right answer existed, you collapse the errors that matter — a fabricated fact, a dropped material detail — into the same number as the ones that don't, like structure, emphasis, and phrasing. This talk is about building evaluation for those systems, where the ground truth isn't a fixed key but a distribution of expert opinion, and where disagreement is built into the problem rather than a symptom of bad labeling. It unpacks three moves that improve accuracy and trustworthiness: treating inter-rater reliability among your experts as the ceiling on any eval you can build, using expert corrections as a live quality signal instead of a static gold set, and separating "genuinely wrong" from "differently right" so the score tracks the failures you actually care about.... Read more

13:00

Felix Njeh

Your AI Passed the Demo. But Did It Pass Validation?

Nelix
AI systems can produce impressive demos and still fail when exposed to real-world complexity, edge cases, changing data, and unexpected user behavior. Semiconductor engineers have spent decades addressing a similar problem through rigorous verification, coverage analysis, corner-case testing, fault injection, regression, and disciplined sign-off. This talk explores how those principles can be adapted to the evaluation of LLMs and increasingly autonomous AI agents, where a successful prompt or benchmark score is not sufficient evidence of system reliability. Attendees will leave with a practical framework for moving from “the AI works” to “we have evidence that this AI system is ready to deploy.”... Read more

13:30

Colby Mainard

Invisible Data - The Largest Problem Facing AI

RiskScout
No matter how well an AI pipeline is designed, it will always be limited by both quality and quantity of available data. However, real-world data rarely arrives nicely packaged and pre-normalized. In fact, it is probably safer to say that real-world data is the number one saboteur of existing machine learning applications. This talk will go over common data problems that have been seen in enterprise systems and the importance of curating data for pipelines of all varieties.... Read more

14:00

Networking & sponsor crawl

Main lobby

14:30

Kunle Olutomilayo

30% More KV Cache Headroom, 100% Accuracy

Isiro AI
Self-managed AI model deployments on-prem and in private cloud infrastructure are becoming increasingly popular, accelerated by open-weight models that rival frontier models. But deploying them needs accelerators (such as GPUs) with large and expensive memory. The usual ways to save that memory quantize or approximate the model, thereby changing its output and affecting accuracy. This talk removes that tradeoff. It covers lossless compression of both model weights and the KV cache, across BF16, FP16, FP8, and FP4, delivering around 30% more GPU headroom with identical model output. Weights are bit-for-bit exact with cryptographic hash verification, and stay compressed in GPU memory while they run. Freeing weight memory leaves more room for the KV cache, and the KV cache is then compressed on top. In deployments that are already memory-constrained, this results in cases of up to 110% more KV cache headroom, raising usable context length, batch size, and concurrency. The talk frames bit-exact KV cache compression as the next lever for lossless inference efficiency.... Read more

15:00

Jessie Mongeon

AI Eval Pipelines for Documentation Quality and Competitor Analysis

Mysten Labs
Developer docs are now read by agents as much as humans, but most teams still judge them by gut feel and page views. This talk shows how to treat documentation as a system under test: rubric-based LLM-as-judge scoring calibrated against human labels, sandboxed agents that attempt your quickstarts end-to-end and report exactly where they fail, and the same suite run against competitors' docs for an apples-to-apples benchmark. You'll leave with a pipeline architecture you can build in a week, rubric patterns that produce stable scores, and a way to turn "our docs vs. theirs" from an argument into data.... Read more

15:30

Yuri Streciwik

The Inference Tax: Why Your Token Costs Stop Making Sense at Scale

Dell Technologies
Most teams building with LLMs meet a point where inference cost stops tracking with usage, and the reasons are invisible from the application layer. The constraints that actually set your cost per token are physical — memory bandwidth, power draw per rack, thermal headroom, and how badly your workload underutilizes the accelerator you're paying for. This talk walks through what happens underneath an inference request at the hardware level, why batching and context length hit economic cliffs rather than smooth curves, and which of those limits are architectural versus fixable in software. Attendees should leave able to read their own inference spend as a systems problem instead of a line item, and to tell which optimizations are worth engineering time.... Read more

16:00

Matt Barge

Selling Commercial Real Estate with Grok Bot

Superbuilders
This talk is a field report on using Grok Bot to sell commercial real estate. I’ll show the live stack I used on real deals: Grok Bot as the operating layer across buyer inbound and follow-up, Quo for the sales phone line and texts, a dedicated /buy sales website and deal room, Google Drive diligence packs, and offering memorandums, all powered by Grok Bot. We’ll cover what the agent actually does day to day (watching inbound, drafting first-touch that leads with income, logging buyer calls, keeping board copy and the OM in sync, chasing listing-agreement issues) and where humans still have to approve sends. Expect concrete CRE examples, the connectors and artifacts that made it work, and the failure modes that show up when an agent touches real buyers.... Read more

16:30

Vashishtha Patil

Auto-Research on a Budget: Small Models in the Loop, Frontier Models on Call

Amazon Lab126
Autonomous research agents that propose, implement, and refine ML solutions have gotten remarkably good. They have also inherited an assumption that excludes most teams: a frontier model drives every step of the loop. That assumption matters because loop cost, not model quality, is becoming the limit on how much autonomy a team can afford to run. This talk explores inverting it. A small open-weight model runs the loop, and a frontier model is called in only as an advisor, on a metered budget. We'll cover what the literature establishes about small and large model collaboration, including step-level escalation, agent distillation, and budget allocation, and where it stops short for long-horizon loops, whose failure modes are not bad tool calls but dead branches, validation leaks, and unclear stopping points. We'll look at a framework for deciding when advice is worth buying, how much of a trajectory an advisor needs to see, and how to measure guidance rather than assume its value. We will walk through a recorded run showing the loop, the trigger, and the cost meter together. The session concludes with open problems, including asynchronous advisors, persistent guidance, and what cost-aware autonomy means for teams without frontier budgets.... Read more

17:00

James Woodard

Operators, Agents, and the 60-Day Cliff: Deploying LLM Systems That Get Used

Apex Systems
Most agentic AI systems work in the demo and fail in production. The failure is rarely the model. It is the cost curve that only appears at scale, the coordination breakdown when single agents become fleets, the governance gap that surfaces during an audit, and the operator trust that never gets earned. This talk is a practitioner's field guide to the last mile of agentic AI: the engineering and operational work that determines whether an LLM system delivers value or gets quietly shelved sixty days after launch. Drawing on real deployment patterns across industrial and enterprise environments, we will cover why token consumption scales with autonomous actions rather than users and how to model the production bill before you sign; what breaks when you move from one agent to a fleet and how to design the coordination layer; the audit and containment questions every production agent must answer; and why explainability at the moment of decision is the real driver of adoption. Attendees will leave with a concrete checklist for pressure-testing an agentic system before it reaches production, and a sharper sense of where the genuine engineering risk lives in the shift from pilots to scaled deployment.... Read more

17:30

Nate Lemos

My wife is a vibe-coder. So, what now?

MakingFriends.app
My wife, who doesn’t have any technical background, built a real app with AI, that she uses every day. So if someone with no previous experience can build working software now, what’s left for software engineers? During this presentation we'll explore what makes a software engineer relevant in the AI era. An engineer needs to know foundations, judgment, responsibility... Programming is easier. Software engineering isn’t.... Read more

18:00

Nikita Pestrov

Building a Software Factory: Your Tickets Are the Benchmarks

Zello
Everyone has an AI cod ing anecdote: an agent nailed one ticket and lost the plot on another. But which tasks can your team reliably hand off, and what makes the resulting PR worth merging? At Zello, we started answering those questions with our own engineering history. We turned solved tickets into a benchmark and merged PRs into a landscape of task classes. Then we connected classification, runnable environments, coding agents and verification into a cloud workflow that produces PRs for human review. This talk follows the journey from benchmarking agents to building a software factory. I’ll show why environment setup matters as much as model choice, how we choose suitable work, and how engineer feedback closes two loops: improving the current PR and teaching the factory something useful for the next ticket.... Read more

18:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Track 1Main room
  1. 09:00
    Registration and coffee
  2. 09:30
  3. 10:00
    One LLM Call Is the Wrong Default
    Geoff Niehaus · Today Mechanic
  4. 10:30
  5. 11:00
  6. 11:30
    Lunch & networking
  7. 12:30
  8. 13:00
  9. 13:30
  10. 14:00
    Networking & sponsor crawl
  11. 14:30
    30% More KV Cache Headroom, 100% Accuracy
    Kunle Olutomilayo · Isiro AI
  12. 15:00
  13. 15:30
  14. 16:00
    Selling Commercial Real Estate with Grok Bot
    Matt Barge · Superbuilders
  15. 16:30
  16. 17:00
  17. 17:30
    My wife is a vibe-coder. So, what now?
    Nate Lemos · MakingFriends.app
  18. 18:00
  19. 18:30
    Wrap up
Time main room
09:00 Registration and coffee
09:30 Visual AI Testing Agent: Screenshot-First Mobile UI Test Automation for iOS Simulator
Rahul Azmeera • 7-Eleven
10:00 One LLM Call Is the Wrong Default
Geoff Niehaus • Today Mechanic
10:30 Ogummaa Agent Portal. Architecture & Security Blueprint for Long-Horizon Agents
Rodney Richard Puplampu • Puplampu Consulting
11:00 One Thing, Many Names: Building Reliable Entity Mapping
Maryam Astero • Genomenon
11:30 Lunch & networking
12:30 Evaluation Without Ground Truth: Lessons from Expert Disagreement
Leila Anderson • Upheal
13:00 Your AI Passed the Demo. But Did It Pass Validation?
Felix Njeh • Nelix
13:30 Invisible Data - The Largest Problem Facing AI
Colby Mainard • RiskScout
14:00 Networking & sponsor crawl
14:30 30% More KV Cache Headroom, 100% Accuracy
Kunle Olutomilayo • Isiro AI
15:00 AI Eval Pipelines for Documentation Quality and Competitor Analysis
Jessie Mongeon • Mysten Labs
15:30 The Inference Tax: Why Your Token Costs Stop Making Sense at Scale
Yuri Streciwik • Dell Technologies
16:00 Selling Commercial Real Estate with Grok Bot
Matt Barge • Superbuilders
16:30 Auto-Research on a Budget: Small Models in the Loop, Frontier Models on Call
Vashishtha Patil • Amazon Lab126
17:00 Operators, Agents, and the 60-Day Cliff: Deploying LLM Systems That Get Used
James Woodard • Apex Systems
17:30 My wife is a vibe-coder. So, what now?
Nate Lemos • MakingFriends.app
18:00 Building a Software Factory: Your Tickets Are the Benchmarks
Nikita Pestrov • Zello
18:30 Wrap up

Speakers

Colby Mainard
RiskScout
Felix Njeh
Nelix
Geoff Niehaus
Today Mechanic
James Woodard
Apex Systems
Jessie Mongeon
Mysten Labs
Kunle Olutomilayo
Isiro AI
Leila Anderson
Upheal
Maryam Astero
Genomenon
Matt Barge
Superbuilders
Nate Lemos
MakingFriends.app
Nikita Pestrov
Zello
Rahul Azmeera
7-Eleven
Rodney Richard Puplampu
Puplampu Consulting
Vashishtha Patil
Amazon Lab126
Yuri Streciwik
Dell Technologies

Venue

The Sunset Room

310 E 3rd St, Austin
TX 78701, United States

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one