LLMday

Large Language Models, Agents & AI Systems

June 4, 2026 Microsoft Experience Center NYC, United States

1
Day
10+
Speakers
1
Track
90+
Attendees

LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.

Companies presenting:

ARF, Bindable AI, Bloomberg, Coolhand Labs, Databricks, Datadog, Glitnir Ticketing, Google, LB Networks, M, Microsoft, Nadir, T Bank, TAILORU Collective, TeraSky, Walmart

Topics so far:
LLMOps & AI Infrastructure

This is a past event, what's next?

Schedule

June 4, 2026 • single track • 9AM - 5:30PM • NYC, in-person
View
Track 1 • main room

09:00

Bart Czernicki

KeynoteDecision Intelligence under Uncertainty with Generative AI

MicrosoftWatch
Decision making is one of the most sought after skills by executives. With the advent of Generative AI, decision making can be optimized to unlock the full potential of human judgement. This hands-on lab introduces the theory of Decision Intelligence with Generative AI (frame decisions, surface assumptions, evaluate decision quality under uncertainty), then applies GenAI as a structured reasoning process. You will practice many advanced Decision Intelligence GenAI techniques.... Read more

09:30

Coffee break

Main lobby

10:00

Max Saltonstall & Salim Virji

Scaled approaches to managing AI workloads

Datadog & Google
New model? New customers? Both??? Your AI workloads are constantly changing, both inputs and outputs, and you need new ways to track success. We will explain how you can do evals at scale, track evals-as-code, and proper experimentation. Join experts from different large-scale SaaS providers to learn how you can evolve, track and constantly improve your measuring and monitoring of AI systems in production.... Read more

10:30

Jinlin He & Rob Bajra

Your LLM Doesn't Know What 'ARR' Means (Yet): A survey of semantic layers

DatabricksWatch
A semantic layer is a business translation layer that sits between your data and your data consumers (like human users or AI agents). A semantic layer is critical for AI agents because it condenses and standardizes business logic so that generic LLMs can better understand your business. Traditionally, semantic layer was embedded within a BI tool and was created manually. But as AI agents become the primary way of retrieving information, the process of building semantic layers needs to be scalable and semantic layers need to be decoupled from downstream tools. In this talk, we explore a few common approaches to building AI friendly semantic layers.... Read more

11:00

Drew Wilkins & Syed Abbas

Beyond IAM: Building the Accountability Layer for Autonomous Agents

Bindable AI
Autonomous AI agents are no longer operating in isolation. As ecosystems built on MCP, A2A, and similar agent protocols enable agents to collaborate across tools, services, and organizational boundaries, traditional IAM models are starting to show their limits. Identity and access control alone cannot answer questions around accountability, traceability, delegated decision-making, or responsibility when autonomous systems interact with each other at scale. This session explores the emerging accountability gap in modern agent architectures and examines what governance needs to look like once agents gain autonomy, interoperability, and the ability to trigger actions across multiple systems. We will discuss practical challenges around auditability, trust boundaries, policy enforcement, and operational oversight in production-grade AI environments, along with architectural approaches for building accountability layers that extend beyond conventional IAM.... Read more

11:30

Sonali Priya

Human-in-the-Loop UI Design for Generative AI Systems

LB NetworksWatch
Generative AI and large language models are rapidly transforming how enterprise user interfaces are conceived, prototyped, and refined. While these systems accelerate exploration and automate repetitive design tasks, they also introduce new challenges in preserving human judgment, design intent, and organizational brand consistency. This session examines emerging human-in-the-loop frameworks that position LLMs as collaborators rather than replacements for expert designers. Drawing on real-world enterprise UX patterns, the talk outlines how AI-generated outputs can be systematically evaluated to maintain usability, accessibility, and human-centered design principles. It highlights approaches to ensuring that generative variations remain aligned with brand guidelines, interaction models, and product standards—especially in regulated or high-stakes environments. The session further explores workflow models for integrating generative AI into UI design processes, including collaborative review loops, intent-preserving prompt strategies, and safeguards against drift in accessibility or experience coherence. By comparing the strengths of human intuition with the scale and speed of LLM-driven systems, attendees will gain clarity on where AI meaningfully accelerates design iteration—and where human oversight remains essential.... Read more

12:00

Lev Andelman

Imperfect Is Fine (When It’s You)

TeraSkyWatch
Humans hallucinate memories, forget decisions, and confidently state wrong facts - and we built entire civilizations around it. Yet when AI does the same, we call it broken. This talk uses cognitive science, live audience experiments, and real-world engineering failures to expose the double standard killing your AI adoption. You’ll leave with a practical framework: validate the output, not the process - and stop demanding perfection from AI that you never demanded from yourself.... Read more

12:30

Lunch & networking

Main lobby

13:00

Anubha Kabra

From Prompt to Production

BloombergWatch
AI applications often look impressive in demos but fail once they encounter real-world users, changing data, operational constraints, and governance requirements. Many teams discover these issues only after launch—when trust, reliability, and cost become production blockers rather than research problems. This talk introduces a practical framework for evaluating whether an AI system is actually ready to ship. Using a deceptively simple Retrieval-Augmented Generation (RAG) assistant as a running example, we’ll walk through the hidden gaps that prototypes frequently overlook: data freshness, hallucination handling, observability, safety guardrails, latency, scaling costs, ownership, and governance. Rather than focusing on model architecture alone, the session reframes AI deployment as a systems and decision-making problem. Attendees will learn five production-readiness lenses that can be applied across AI products and internal enterprise systems: Data & Retrieval Trust & Safety Observability & Debuggability Cost & Scale Governance & Ownership The talk also explores real tradeoffs teams face in practice, including precision vs. coverage, latency vs. cost, and experimentation vs. reliability. By the end of the session, attendees will leave with a concrete checklist and decision framework they can use to evaluate AI pilots, reduce deployment risk, and move from prototype success to production reliability.... Read more

13:30

Chrys Li

Beyond Prompts: Context Architecture for Reliable LLM Systems

TAILORU CollectiveWatch
Most LLM failures are not caused by a bad prompt alone. They often emerge from missing, unstable, or poorly structured context: unclear goals, conflicting source material, weak evidence hierarchy, fragmented memory, vague role definitions, and undefined decision boundaries. This talk introduces Context Architecture as the missing design layer between prompt engineering, retrieval, evals, governance, and production behavior. We’ll look at how context can be structured as a system: what the model needs to know, what it should prioritize, what it should ignore, how state should be carried forward, and where human judgment needs to remain explicit. The session will share a practical framework for designing context packages, behavior constraints, signal flows, and handoff artifacts that make LLM systems more reliable, inspectable, and easier to improve over time. Rather than treating prompts as isolated instructions, the talk reframes context as infrastructure: the layer that shapes how an LLM interprets tasks, reasons through ambiguity, and behaves inside real workflows.... Read more

14:00

Juan Petter

The Missing Layer in Agentic AI: Governance

ARF
Most teams are moving fast on agents, but very few are governing them well. In production, the real problem is not capability — it is control: how to evaluate risk, bound autonomy, preserve auditability, and keep systems stable under uncertainty. In this talk, I will introduce the architecture behind ARF (Agentic Reliability Framework), a governance layer for agentic infrastructure that converts probabilistic AI outputs into deterministic, auditable decisions. The session explores Bayesian risk scoring, expected-loss decisioning, bounded memory systems, and calibrated escalation mechanisms designed to make autonomous AI survivable inside real enterprise environments.... Read more

14:30

William McLean

Gemma4 Tips and Tricks

Glitnir Ticketing
Overview of the latest Gemma4 models from on-device to TPU hosted. Tips and tricks for vLLM model serving and a live demo.... Read more

15:00

Dor Amir

Stop Paying Opus Prices for Haiku Problems

NadirWatch
Most AI-native products today are overpaying for LLM usage because they send too many requests to expensive frontier models, even when the task does not require them. In this session, I’ll share how LLM routing works in practice: how to classify prompt complexity, route requests across models, and balance cost, latency, and quality without hurting the user experience. I’ll also cover the common traps teams run into when building routers, including over-routing to premium models, bad offline evaluation, feedback loops, and the gap between benchmark accuracy and real production quality. The goal is to give builders a practical framework they can use to decide when to use cheaper models, when to escalate to stronger models, and how to measure whether the routing system is actually saving money while preserving output quality.... Read more

15:30

Play some Fallout !

Main lobby

16:00

Michael Carroll

Your AI Agent Has Notes

Coolhand LabsWatch
Working with AI agents is turning every IC into a team manager. Yesterday, you were coding. Today? You're delegating tasks, reviewing outputs, debugging reasoning loops, and sweating your monthly token budget. Hate to break it to you, but your AI team has notes about this... and you aren't listening. This talk is about giving your agents a complaint box. I'll walk through the design of a strategy called Wildcard. Wildcard is a flexible, amorphous tool your agents can call when they are stuck to explain what they need. I'll give some solid examples of how Wildcard leveled up the agents that power Coolhand Labs by surfacing issues in real time, helping us cut down on loops, token burn, and bad practices. We will implement the tool from scratch in a sample codebase, handle some agentic feedback in real time, and cover a few protips for getting the most out of Wildcard in your agent team. Giving you strategies to coax your agents into providing unfiltered feedback is included in the talk. Processing that feedback emotionally will be on you.... Read more

16:30

Mazdul Choudhury

AI Agents as Decision Co Pilots: Scaling LLM Driven Intelligent Systems

Walmart
Large Language Models are rapidly evolving from passive text generators into active decision making systems that can reason across complex inputs and support real world operations. However, many organizations still struggle to move beyond isolated use cases toward scalable, production ready AI systems that deliver consistent business value. This session introduces AI agents as decision co pilots powered by LLMs, enabling a shift from static analytics and rule based automation to adaptive, context aware intelligent systems. These agents continuously ingest and interpret diverse data signals, including structured operational data, unstructured inputs, and real time contextual information. By combining reasoning capabilities with domain awareness, they generate dynamic recommendations, evaluate scenarios, and communicate insights in natural language. The talk presents a practical three layer architecture for building agent driven systems. The prediction layer focuses on signal processing and model outputs, the decision layer translates insights into actionable recommendations, and the interaction layer enables seamless human collaboration through conversational interfaces. This approach allows organizations to move from fragmented tools toward unified decision systems that scale across functions and use cases. The session will also explore how LLM powered agents reduce cognitive load, improve decision consistency, and accelerate response times in complex environments. Rather than replacing human expertise, these systems augment decision makers by enabling faster synthesis of information and more informed actions. Attendees will gain a clear framework for designing, deploying, and scaling LLM driven AI agents in production, along with practical insights into building systems that are adaptable, explainable, and aligned with real world operational needs.... Read more

17:00

Vanchhit Khare

Your AI Is Making Things Up. Here Is How to Stop It.

M&T Bank
Your AI agent is lying to you. Not on purpose. It just fills in the gaps when it does not know something. It sounds confident. It names real people. It cites papers that do not exist. Your validator says everything is fine. Here is what actually stops it: give your agent a second brain. Not a bigger model. Not a better prompt. A small, structured knowledge base of facts your agent is allowed to use. 300 notes. Three rules. Every agent in the pipeline checks the second brain before it writes anything. I ran this on a real pipeline. Before the second brain: one fake fact every 12 outputs, confidence 0.94, validators happy. After: zero fake facts across 300 runs. The overhead is 20 milliseconds per step. This talk shows how the second brain works, why it beats prompting, and how to wire it into any framework you already use. LangGraph, CrewAI, Claude Code, does not matter. Live demo included: I plant a fake fact, the second brain blocks it, the final report stays clean. You leave with the code.... Read more

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Track 1Main room
  1. 09:00
  2. 09:30
    Coffee break
  3. 10:00
    Scaled approaches to managing AI workloads
    Max Saltonstall & Salim Virji · Datadog & Google
  4. 10:30
  5. 11:00
  6. 11:30
  7. 12:00
    Imperfect Is Fine (When It’s You)
    Lev Andelman · TeraSky
  8. 12:30
    Lunch & networking
  9. 13:00
    From Prompt to Production
    Anubha Kabra · Bloomberg
  10. 13:30
  11. 14:00
  12. 14:30
    Gemma4 Tips and Tricks
    William McLean · Glitnir Ticketing
  13. 15:00
  14. 15:30
    Play some Fallout !
  15. 16:00
    Your AI Agent Has Notes
    Michael Carroll · Coolhand Labs
  16. 16:30
  17. 17:00
  18. 17:30
    Wrap up
Time main room
09:00 Keynote: Decision Intelligence under Uncertainty with Generative AI
Bart Czernicki • Microsoft
09:30 Coffee break
10:00 Scaled approaches to managing AI workloads
Max Saltonstall & Salim Virji • Datadog & Google
10:30 Your LLM Doesn't Know What 'ARR' Means (Yet): A survey of semantic layers
Jinlin He & Rob Bajra • Databricks
11:00 Beyond IAM: Building the Accountability Layer for Autonomous Agents
Drew Wilkins & Syed Abbas • Bindable AI
11:30 Human-in-the-Loop UI Design for Generative AI Systems
Sonali Priya • LB Networks
12:00 Imperfect Is Fine (When It’s You)
Lev Andelman • TeraSky
12:30 Lunch & networking
13:00 From Prompt to Production
Anubha Kabra • Bloomberg
13:30 Beyond Prompts: Context Architecture for Reliable LLM Systems
Chrys Li • TAILORU Collective
14:00 The Missing Layer in Agentic AI: Governance
Juan Petter • ARF
14:30 Gemma4 Tips and Tricks
William McLean • Glitnir Ticketing
15:00 Stop Paying Opus Prices for Haiku Problems
Dor Amir • Nadir
15:30 Play some Fallout !
16:00 Your AI Agent Has Notes
Michael Carroll • Coolhand Labs
16:30 AI Agents as Decision Co Pilots: Scaling LLM Driven Intelligent Systems
Mazdul Choudhury • Walmart
17:00 Your AI Is Making Things Up. Here Is How to Stop It.
Vanchhit Khare • M&T Bank
17:30 Wrap up

Speakers

Anubha Kabra
Bloomberg
Bart Czernicki
Microsoft
Chrys Li
TAILORU Collective
Dor Amir
Nadir
Drew Wilkins
& Syed Abbas
Bindable AI
Jinlin He
& Rob Bajra
Databricks
Juan Petter
ARF
Lev Andelman
TeraSky
Max Saltonstall
& Salim Virji
Datadog & Google
Mazdul Choudhury
Walmart
Michael Carroll
Coolhand Labs
Sonali Priya
LB Networks
Vanchhit Khare
M&T Bank
William McLean
Glitnir Ticketing

Venue

Microsoft Experience Center NYC

677 5th Avenue
New York, NY 10022
Between 53rd and 54th Street

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one