LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.
Companies presenting:
ARF, Bindable AI, Bloomberg, Coolhand Labs, Databricks, Datadog, Glitnir Ticketing, Google, LB Networks, M, Microsoft, Nadir, T Bank, TAILORU Collective, TeraSky, Walmart
Decision making is one of the most sought after skills by executives. With the advent of Generative AI, decision making can be optimized to unlock the full potential of human judgement. This hands-on lab introduces the theory of Decision Intelligence with Generative AI (frame decisions, surface assumptions, evaluate decision quality under uncertainty), then applies GenAI as a structured reasoning process. You will practice many advanced Decision Intelligence GenAI techniques.... Read more
New model? New customers? Both??? Your AI workloads are constantly changing, both inputs and outputs, and you need new ways to track success. We will explain how you can do evals at scale, track evals-as-code, and proper experimentation. Join experts from different large-scale SaaS providers to learn how you can evolve, track and constantly improve your measuring and monitoring of AI systems in production.... Read more
A semantic layer is a business translation layer that sits between your data and your data consumers (like human users or AI agents). A semantic layer is critical for AI agents because it condenses and standardizes business logic so that generic LLMs can better understand your business. Traditionally, semantic layer was embedded within a BI tool and was created manually. But as AI agents become the primary way of retrieving information, the process of building semantic layers needs to be scalable and semantic layers need to be decoupled from downstream tools. In this talk, we explore a few common approaches to building AI friendly semantic layers.... Read more
Autonomous AI agents are no longer operating in isolation. As ecosystems built on MCP, A2A, and similar agent protocols enable agents to collaborate across tools, services, and organizational boundaries, traditional IAM models are starting to show their limits. Identity and access control alone cannot answer questions around accountability, traceability, delegated decision-making, or responsibility when autonomous systems interact with each other at scale. This session explores the emerging accountability gap in modern agent architectures and examines what governance needs to look like once agents gain autonomy, interoperability, and the ability to trigger actions across multiple systems. We will discuss practical challenges around auditability, trust boundaries, policy enforcement, and operational oversight in production-grade AI environments, along with architectural approaches for building accountability layers that extend beyond conventional IAM.... Read more
Generative AI and large language models are rapidly transforming how enterprise user interfaces are conceived, prototyped, and refined. While these systems accelerate exploration and automate repetitive design tasks, they also introduce new challenges in preserving human judgment, design intent, and organizational brand consistency. This session examines emerging human-in-the-loop frameworks that position LLMs as collaborators rather than replacements for expert designers. Drawing on real-world enterprise UX patterns, the talk outlines how AI-generated outputs can be systematically evaluated to maintain usability, accessibility, and human-centered design principles. It highlights approaches to ensuring that generative variations remain aligned with brand guidelines, interaction models, and product standards—especially in regulated or high-stakes environments. The session further explores workflow models for integrating generative AI into UI design processes, including collaborative review loops, intent-preserving prompt strategies, and safeguards against drift in accessibility or experience coherence. By comparing the strengths of human intuition with the scale and speed of LLM-driven systems, attendees will gain clarity on where AI meaningfully accelerates design iteration—and where human oversight remains essential.... Read more
Humans hallucinate memories, forget decisions, and confidently state wrong facts - and we built entire civilizations around it. Yet when AI does the same, we call it broken. This talk uses cognitive science, live audience experiments, and real-world engineering failures to expose the double standard killing your AI adoption. You’ll leave with a practical framework: validate the output, not the process - and stop demanding perfection from AI that you never demanded from yourself.... Read more
AI applications often look impressive in demos but fail once they encounter real-world users, changing data, operational constraints, and governance requirements. Many teams discover these issues only after launch—when trust, reliability, and cost become production blockers rather than research problems. This talk introduces a practical framework for evaluating whether an AI system is actually ready to ship. Using a deceptively simple Retrieval-Augmented Generation (RAG) assistant as a running example, we’ll walk through the hidden gaps that prototypes frequently overlook: data freshness, hallucination handling, observability, safety guardrails, latency, scaling costs, ownership, and governance. Rather than focusing on model architecture alone, the session reframes AI deployment as a systems and decision-making problem. Attendees will learn five production-readiness lenses that can be applied across AI products and internal enterprise systems: Data & Retrieval Trust & Safety Observability & Debuggability Cost & Scale Governance & Ownership The talk also explores real tradeoffs teams face in practice, including precision vs. coverage, latency vs. cost, and experimentation vs. reliability. By the end of the session, attendees will leave with a concrete checklist and decision framework they can use to evaluate AI pilots, reduce deployment risk, and move from prototype success to production reliability.... Read more
Most LLM failures are not caused by a bad prompt alone. They often emerge from missing, unstable, or poorly structured context: unclear goals, conflicting source material, weak evidence hierarchy, fragmented memory, vague role definitions, and undefined decision boundaries. This talk introduces Context Architecture as the missing design layer between prompt engineering, retrieval, evals, governance, and production behavior. We’ll look at how context can be structured as a system: what the model needs to know, what it should prioritize, what it should ignore, how state should be carried forward, and where human judgment needs to remain explicit. The session will share a practical framework for designing context packages, behavior constraints, signal flows, and handoff artifacts that make LLM systems more reliable, inspectable, and easier to improve over time. Rather than treating prompts as isolated instructions, the talk reframes context as infrastructure: the layer that shapes how an LLM interprets tasks, reasons through ambiguity, and behaves inside real workflows.... Read more
Most teams are moving fast on agents, but very few are governing them well. In production, the real problem is not capability — it is control: how to evaluate risk, bound autonomy, preserve auditability, and keep systems stable under uncertainty. In this talk, I will introduce the architecture behind ARF (Agentic Reliability Framework), a governance layer for agentic infrastructure that converts probabilistic AI outputs into deterministic, auditable decisions. The session explores Bayesian risk scoring, expected-loss decisioning, bounded memory systems, and calibrated escalation mechanisms designed to make autonomous AI survivable inside real enterprise environments.... Read more
Most AI-native products today are overpaying for LLM usage because they send too many requests to expensive frontier models, even when the task does not require them. In this session, I’ll share how LLM routing works in practice: how to classify prompt complexity, route requests across models, and balance cost, latency, and quality without hurting the user experience. I’ll also cover the common traps teams run into when building routers, including over-routing to premium models, bad offline evaluation, feedback loops, and the gap between benchmark accuracy and real production quality. The goal is to give builders a practical framework they can use to decide when to use cheaper models, when to escalate to stronger models, and how to measure whether the routing system is actually saving money while preserving output quality.... Read more
Working with AI agents is turning every IC into a team manager. Yesterday, you were coding. Today? You're delegating tasks, reviewing outputs, debugging reasoning loops, and sweating your monthly token budget. Hate to break it to you, but your AI team has notes about this... and you aren't listening. This talk is about giving your agents a complaint box. I'll walk through the design of a strategy called Wildcard. Wildcard is a flexible, amorphous tool your agents can call when they are stuck to explain what they need. I'll give some solid examples of how Wildcard leveled up the agents that power Coolhand Labs by surfacing issues in real time, helping us cut down on loops, token burn, and bad practices. We will implement the tool from scratch in a sample codebase, handle some agentic feedback in real time, and cover a few protips for getting the most out of Wildcard in your agent team. Giving you strategies to coax your agents into providing unfiltered feedback is included in the talk. Processing that feedback emotionally will be on you.... Read more
Large Language Models are rapidly evolving from passive text generators into active decision making systems that can reason across complex inputs and support real world operations. However, many organizations still struggle to move beyond isolated use cases toward scalable, production ready AI systems that deliver consistent business value. This session introduces AI agents as decision co pilots powered by LLMs, enabling a shift from static analytics and rule based automation to adaptive, context aware intelligent systems. These agents continuously ingest and interpret diverse data signals, including structured operational data, unstructured inputs, and real time contextual information. By combining reasoning capabilities with domain awareness, they generate dynamic recommendations, evaluate scenarios, and communicate insights in natural language. The talk presents a practical three layer architecture for building agent driven systems. The prediction layer focuses on signal processing and model outputs, the decision layer translates insights into actionable recommendations, and the interaction layer enables seamless human collaboration through conversational interfaces. This approach allows organizations to move from fragmented tools toward unified decision systems that scale across functions and use cases. The session will also explore how LLM powered agents reduce cognitive load, improve decision consistency, and accelerate response times in complex environments. Rather than replacing human expertise, these systems augment decision makers by enabling faster synthesis of information and more informed actions. Attendees will gain a clear framework for designing, deploying, and scaling LLM driven AI agents in production, along with practical insights into building systems that are adaptable, explainable, and aligned with real world operational needs.... Read more
Your AI agent is lying to you. Not on purpose. It just fills in the gaps when it does not know something. It sounds confident. It names real people. It cites papers that do not exist. Your validator says everything is fine. Here is what actually stops it: give your agent a second brain. Not a bigger model. Not a better prompt. A small, structured knowledge base of facts your agent is allowed to use. 300 notes. Three rules. Every agent in the pipeline checks the second brain before it writes anything. I ran this on a real pipeline. Before the second brain: one fake fact every 12 outputs, confidence 0.94, validators happy. After: zero fake facts across 300 runs. The overhead is 20 milliseconds per step. This talk shows how the second brain works, why it beats prompting, and how to wire it into any framework you already use. LangGraph, CrewAI, Claude Code, does not matter. Live demo included: I plant a fake fact, the second brain blocks it, the final report stays clean. You leave with the code.... Read more
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
677 5th Avenue
New York, NY 10022
Between 53rd and 54th Street
Sponsors & Partners
Want to become a sponsor? Get in touch!
Max Saltonstall & Salim Virji
Datadog & Google
Scaled approaches to managing AI workloads
Abstract
New model? New customers? Both??? Your AI workloads are constantly changing, both inputs and outputs, and you need new ways to track success. We will explain how you can do evals at scale, track evals-as-code, and proper experimentation. Join experts from different large-scale SaaS providers to learn how you can evolve, track and constantly improve your measuring and monitoring of AI systems in production.
Bio
Max Saltonstall tells stories about the Cloud, how to tell where it’s working (or not), and what diverse solutions our customers have created. He’s part of Datadog’s Advocacy team, with a passion for storytelling, narrative and dad jokes. Max loves to juggle, eat new foods and play games of all types. He lives in New York City with his partner and two children.
Salim Virji develops reliable engineering practices and processes for Google’s SRE organization, and has built consensus and storage services for Google infrastructure. Salim’s interests include distributed systems and machine learning. He has contributed to several books on SRE, including The Site Reliability Workbook and Implementing Service Level Objectives. Salim received an AB in Classics from the University of Chicago and is a New York City Master Composter.
Jinlin He & Rob Bajra
Databricks
Your LLM Doesn't Know What 'ARR' Means (Yet): A survey of semantic layers
Abstract
A semantic layer is a business translation layer that sits between your data and your data consumers (like human users or AI agents).
A semantic layer is critical for AI agents because it condenses and standardizes business logic so that generic LLMs can better understand your business. Traditionally, semantic layer was embedded within a BI tool and was created manually. But as AI agents become the primary way of retrieving information, the process of building semantic layers needs to be scalable and semantic layers need to be decoupled from downstream tools.
In this talk, we explore a few common approaches to building AI friendly semantic layers.
Bio
Jinlin He is a Data scientist turned Solutions Architect at Databricks with 6+ years in ML/AI engineering and pre-sales. She enjoys the intersection of systematic problem solving and open-ended search space in technical sales.
Rob Bajra is a Solutions Architect with Databricks based in New York City. Rob specializes at the intersection of agents and data apps. With a background in AI engineering and data science, Rob has led the development of influencer marketplaces, digital twins and dynamic pricing systems. Rob has previously worked as an AI Engineer at JetBlue Airways and Data Scientist at Southwest Airlines.
Drew Wilkins & Syed Abbas
Bindable AI
Beyond IAM: Building the Accountability Layer for Autonomous Agents
Abstract
Autonomous AI agents are no longer operating in isolation. As ecosystems built on MCP, A2A, and similar agent protocols enable agents to collaborate across tools, services, and organizational boundaries, traditional IAM models are starting to show their limits. Identity and access control alone cannot answer questions around accountability, traceability, delegated decision-making, or responsibility when autonomous systems interact with each other at scale.
This session explores the emerging accountability gap in modern agent architectures and examines what governance needs to look like once agents gain autonomy, interoperability, and the ability to trigger actions across multiple systems. We will discuss practical challenges around auditability, trust boundaries, policy enforcement, and operational oversight in production-grade AI environments, along with architectural approaches for building accountability layers that extend beyond conventional IAM.
Bio
Drew Wilkins is CEO and co-founder of Bindable AI LLC, where he focuses on lifecycle accountability frameworks for autonomous AI agents operating in high-stakes environments. Previously, Drew held senior leadership roles at major advertising and digital experience firms, helping organizations navigate the intersection of technology, customer experience, and organizational change. His career has centered on technology shaping human experience, translating complex technologies into trusted, meaningful systems. Drew has pursued formal AI study through programs at Massachusetts Institute of Technology and Stanford University, and one of his ventures received a U.S. Coast Guard innovation award for applied AI innovation. He holds an MBA from Emory University and a BA in Philosophy from Emory College.
Syed Abbas is a Senior Business and Integration Architect who specializes in simplifying complex technology. He brings 20 years of experience as a Solutions and Data Architect, including 16 years delivering large-scale enterprise programs across the automotive, pharmaceutical, biomedical, and manufacturing sectors. Syed’s background spans SAP and enterprise integration platforms, identity protocols, and large-scale data architecture, now applied to production-grade AI agents and AI security research. He holds a B.S. in Computer Science.
Sonali Priya
LB Networks
Human-in-the-Loop UI Design for Generative AI Systems
Abstract
Generative AI and large language models are rapidly transforming how enterprise user interfaces are conceived, prototyped, and refined. While these systems accelerate exploration and automate repetitive design tasks, they also introduce new challenges in preserving human judgment, design intent, and organizational brand consistency. This session examines emerging human-in-the-loop frameworks that position LLMs as collaborators rather than replacements for expert designers. Drawing on real-world enterprise UX patterns, the talk outlines how AI-generated outputs can be systematically evaluated to maintain usability, accessibility, and human-centered design principles. It highlights approaches to ensuring that generative variations remain aligned with brand guidelines, interaction models, and product standards—especially in regulated or high-stakes environments. The session further explores workflow models for integrating generative AI into UI design processes, including collaborative review loops, intent-preserving prompt strategies, and safeguards against drift in accessibility or experience coherence. By comparing the strengths of human intuition with the scale and speed of LLM-driven systems, attendees will gain clarity on where AI meaningfully accelerates design iteration—and where human oversight remains essential.
Bio
Sonali Priya is a Creative Design Director based in St. Louis, Missouri, specializing in building purposeful, scalable, and inclusive digital experiences at the intersection of design, development, and systems thinking. With over eight years of experience in product design and full-stack engineering, she has led multidisciplinary teams to deliver accessible, high-performance solutions for enterprise clients worldwide. Currently leading design at LB Networks, a global provider of network monitoring and optimization tools, Sonali drives platform-wide design initiatives, collaborates with stakeholders, and mentors cross-functional teams. She architects and evolves design systems to ensure consistency, accessibility, and scalability across products. Her career progression from Software Developer to Software Team Lead to Creative Design Director enables her to bridge technical depth with design strategy. She has led modernization efforts, improved engineering workflows, and contributed to development using React, TypeScript, PHP, MySQL, and Node.js. Sonali has delivered key initiatives including the Business Service Portal and OcularIP UI modernization, enhancing performance, usability, and accessibility.
Lev Andelman
TeraSky
Imperfect Is Fine (When It’s You)
Abstract
Humans hallucinate memories, forget decisions, and confidently state wrong facts - and we built entire civilizations around it. Yet when AI does the same, we call it broken.
This talk uses cognitive science, live audience experiments, and real-world engineering failures to expose the double standard killing your AI adoption. You’ll leave with a practical framework: validate the output, not the process - and stop demanding perfection from AI that you never demanded from yourself.
Bio
Lev Andelman is a Co-Founder and CTO at TeraSky Group with over 15 years of experience designing and managing large-scale production environments. He specializes in cloud, DevOps, and infrastructure across Unix/Linux and Windows systems, with a track record of leading engineering teams and building robust R&D, QA, and production platforms.
Anubha Kabra
Bloomberg
From Prompt to Production
Abstract
AI applications often look impressive in demos but fail once they encounter real-world users, changing data, operational constraints, and governance requirements. Many teams discover these issues only after launch—when trust, reliability, and cost become production blockers rather than research problems.
This talk introduces a practical framework for evaluating whether an AI system is actually ready to ship. Using a deceptively simple Retrieval-Augmented Generation (RAG) assistant as a running example, we’ll walk through the hidden gaps that prototypes frequently overlook: data freshness, hallucination handling, observability, safety guardrails, latency, scaling costs, ownership, and governance.
Rather than focusing on model architecture alone, the session reframes AI deployment as a systems and decision-making problem. Attendees will learn five production-readiness lenses that can be applied across AI products and internal enterprise systems:
Data & Retrieval
Trust & Safety
Observability & Debuggability
Cost & Scale
Governance & Ownership
The talk also explores real tradeoffs teams face in practice, including precision vs. coverage, latency vs. cost, and experimentation vs. reliability. By the end of the session, attendees will leave with a concrete checklist and decision framework they can use to evaluate AI pilots, reduce deployment risk, and move from prototype success to production reliability.
Bio
Anubha Kabra is a Senior ML research Engineer at Bloomberg. She is working on trustworthy and production-ready AI agents and systems. Her work focuses on observability, evaluation, and reliability for enterprise-scale LLM agentic applications. She has presented research at major conferences including NeurIPS, ACL, NAACL and works on building practical frameworks for deploying high-stakes AI systems responsibly.
Chrys Li
TAILORU Collective
Beyond Prompts: Context Architecture for Reliable LLM Systems
Abstract
Most LLM failures are not caused by a bad prompt alone. They often emerge from missing, unstable, or poorly structured context: unclear goals, conflicting source material, weak evidence hierarchy, fragmented memory, vague role definitions, and undefined decision boundaries.
This talk introduces Context Architecture as the missing design layer between prompt engineering, retrieval, evals, governance, and production behavior. We’ll look at how context can be structured as a system: what the model needs to know, what it should prioritize, what it should ignore, how state should be carried forward, and where human judgment needs to remain explicit.
The session will share a practical framework for designing context packages, behavior constraints, signal flows, and handoff artifacts that make LLM systems more reliable, inspectable, and easier to improve over time. Rather than treating prompts as isolated instructions, the talk reframes context as infrastructure: the layer that shapes how an LLM interprets tasks, reasons through ambiguity, and behaves inside real workflows.
Bio
Chrys Li is an AI systems and service design strategist focused on designing intelligent systems inside complex organizations. Her work spans enterprise platforms, regulated environments, AI-enabled products, and large-scale service ecosystems, with a focus on translating ambiguity into structured, build-ready systems.
She has led systems and experience design initiatives across logistics, healthcare, financial services, telecom, education, enterprise software, and other sectors. More recently, her work has focused on AI behavior, context architecture, human-AI collaboration, and the design of operational frameworks that help teams move from AI experimentation to practical implementation.
Chrys is an independent consultant where she develops frameworks, workshops, and applied AI methods for product, design, and business teams adapting to AI-era systems.
Juan Petter
ARF
The Missing Layer in Agentic AI: Governance
Abstract
Most teams are moving fast on agents, but very few are governing them well. In production, the real problem is not capability — it is control: how to evaluate risk, bound autonomy, preserve auditability, and keep systems stable under uncertainty. In this talk, I will introduce the architecture behind ARF (Agentic Reliability Framework), a governance layer for agentic infrastructure that converts probabilistic AI outputs into deterministic, auditable decisions. The session explores Bayesian risk scoring, expected-loss decisioning, bounded memory systems, and calibrated escalation mechanisms designed to make autonomous AI survivable inside real enterprise environments.
Bio
Juan Petter is an AI Engineer and founder building the Agentic Reliability Framework (ARF), a governance and control plane for agentic infrastructure. His work focuses on making AI systems provably safe, auditable, and operationally viable in production environments. He has experience across cloud infrastructure, technical support, full-stack engineering, and AI systems design.
William McLean
Glitnir Ticketing
Gemma4 Tips and Tricks
Abstract
Overview of the latest Gemma4 models from on-device to TPU hosted.
Tips and tricks for vLLM model serving and a live demo.
Bio
CTO of Glitnir Ticketing. Google Developer Expert (GDE) for AI/ML and Cloud.
Amazon Community Builder for AI engineering.
Dor Amir
Nadir
Stop Paying Opus Prices for Haiku Problems
Abstract
Most AI-native products today are overpaying for LLM usage because they send too many requests to expensive frontier models, even when the task does not require them.
In this session, I’ll share how LLM routing works in practice: how to classify prompt complexity, route requests across models, and balance cost, latency, and quality without hurting the user experience.
I’ll also cover the common traps teams run into when building routers, including over-routing to premium models, bad offline evaluation, feedback loops, and the gap between benchmark accuracy and real production quality.
The goal is to give builders a practical framework they can use to decide when to use cheaper models, when to escalate to stronger models, and how to measure whether the routing system is actually saving money while preserving output quality.
Bio
Dor Amir is building Nadir getnadir.com, an LLM router that helps AI-native products reduce model costs by routing each request to the right model based on prompt complexity, quality requirements, latency, and cost.
Dor currently works at Dropbox on the agentic AI team, building AI agents that improve Dropbox experiences. Before Dropbox, he worked as a Senior Machine Learning Engineer on Amazon’s Personalization team, where he focused on applied ML systems. Earlier in his career, he worked across data science, machine learning, and software engineering roles, including ML leadership at Guesty and recommendation ML at Fiverr.
Dor’s work sits at the intersection of machine learning, LLM infrastructure, and production AI systems. His current focus is helping teams build practical AI agents and LLM routing systems that reduce cost, improve latency, and preserve quality in real-world applications.
Michael Carroll
Coolhand Labs
Your AI Agent Has Notes
Abstract
Working with AI agents is turning every IC into a team manager. Yesterday, you were coding. Today? You're delegating tasks, reviewing outputs, debugging reasoning loops, and sweating your monthly token budget. Hate to break it to you, but your AI team has notes about this... and you aren't listening. This talk is about giving your agents a complaint box. I'll walk through the design of a strategy called Wildcard. Wildcard is a flexible, amorphous tool your agents can call when they are stuck to explain what they need. I'll give some solid examples of how Wildcard leveled up the agents that power Coolhand Labs by surfacing issues in real time, helping us cut down on loops, token burn, and bad practices. We will implement the tool from scratch in a sample codebase, handle some agentic feedback in real time, and cover a few protips for getting the most out of Wildcard in your agent team. Giving you strategies to coax your agents into providing unfiltered feedback is included in the talk. Processing that feedback emotionally will be on you.
Bio
Michael Carroll is the founder of Coolhand Labs, which automatically improves AI agents from human & synthetic feedback. Previously, he built RubiconMD (acquired by Oak Street Health) and served as VP of Engineering and Strategic Projects at Teladoc Health, the world's largest telemedicine provider. Michael also publishes The Everything Engineer, a newsletter about the transition from IC engineering to managing a team of AI agents. When he's not reading his agents' complaints, he's merging the PRs that will generate a fresh batch of them.
Mazdul Choudhury
Walmart
AI Agents as Decision Co Pilots: Scaling LLM Driven Intelligent Systems
Abstract
Large Language Models are rapidly evolving from passive text generators into active decision making systems that can reason across complex inputs and support real world operations. However, many organizations still struggle to move beyond isolated use cases toward scalable, production ready AI systems that deliver consistent business value. This session introduces AI agents as decision co pilots powered by LLMs, enabling a shift from static analytics and rule based automation to adaptive, context aware intelligent systems. These agents continuously ingest and interpret diverse data signals, including structured operational data, unstructured inputs, and real time contextual information. By combining reasoning capabilities with domain awareness, they generate dynamic recommendations, evaluate scenarios, and communicate insights in natural language. The talk presents a practical three layer architecture for building agent driven systems. The prediction layer focuses on signal processing and model outputs, the decision layer translates insights into actionable recommendations, and the interaction layer enables seamless human collaboration through conversational interfaces. This approach allows organizations to move from fragmented tools toward unified decision systems that scale across functions and use cases. The session will also explore how LLM powered agents reduce cognitive load, improve decision consistency, and accelerate response times in complex environments. Rather than replacing human expertise, these systems augment decision makers by enabling faster synthesis of information and more informed actions. Attendees will gain a clear framework for designing, deploying, and scaling LLM driven AI agents in production, along with practical insights into building systems that are adaptable, explainable, and aligned with real world operational needs.
Bio
Mazdul Hassan Choudhury is a Product Management leader with over 15 years of experience delivering digital products across retail operations, supply chain, merchandising, fulfillment, last-mile delivery, and customer care. He has led the development and launch of enterprise products that improved operational efficiency, increased revenue, and enhanced customer experience. He currently serves as a Staff Product Manager at Walmart, where he owns and manages associate-facing applications and aligns product roadmaps with enterprise strategy. His work includes launching a GenAI-powered tool that improved operational efficiency, redesigning onboarding processes to reduce onboarding time by 80%, and delivering solutions that improved fulfillment accuracy and supported large-scale pre-order operations. Prior to Walmart, Mazdul worked at Bob’s Furniture, where he led the development of a mobile point-of-sale system used by 3,500 associates, increasing sales by 20% and reducing order processing time. He also built a customer relationship management platform that improved customer satisfaction and developed analytics solutions that increased operational visibility. Earlier in his career, he held roles at Capgemini, Equinor, and IBM, where he managed cross-functional teams and delivered SAP implementations, system integrations, and large-scale technology projects. He holds an MBA in Strategy and Operations and a Master’s degree in Information Systems Management from Boston University, along with a Bachelor’s degree in Electronics and Communication Engineering.
Vanchhit Khare
M&T Bank
Your AI Is Making Things Up. Here Is How to Stop It.
Abstract
Your AI agent is lying to you. Not on purpose. It just fills in the gaps when it does not know something. It sounds confident. It names real people. It cites papers that do not exist. Your validator says everything is fine.
Here is what actually stops it: give your agent a second brain.
Not a bigger model. Not a better prompt. A small, structured knowledge base of facts your agent is allowed to use. 300 notes. Three rules. Every agent in the pipeline checks the second brain before it writes anything.
I ran this on a real pipeline. Before the second brain: one fake fact every 12 outputs, confidence 0.94, validators happy. After: zero fake facts across 300 runs. The overhead is 20 milliseconds per step.
This talk shows how the second brain works, why it beats prompting, and how to wire it into any framework you already use. LangGraph, CrewAI, Claude Code, does not matter. Live demo included: I plant a fake fact, the second brain blocks it, the final report stays clean.
You leave with the code.
Bio
Vanchhit Khare is a Salesforce System Architect and software engineer with 10+ years of experience building enterprise systems across banking, AI, and cloud technologies. He currently works at M&T Bank, specializing in Salesforce architecture, Agentic AI, and digital transformation initiatives. Vanchhit is also an active mentor, researcher, and community contributor focused on making complex technology more accessible and human-centered.
Bart Czernicki
Microsoft
KeynoteDecision Intelligence under Uncertainty with Generative AI
Abstract
Decision making is one of the most sought after skills by executives. With the advent of Generative AI, decision making can be optimized to unlock the full potential of human judgement. This hands-on lab introduces the theory of Decision Intelligence with Generative AI (frame decisions, surface assumptions, evaluate decision quality under uncertainty), then applies GenAI as a structured reasoning process. You will practice many advanced Decision Intelligence GenAI techniques.
Bio
Bart Czernicki is a technology leader, executive advisor, and author specializing in cloud, AI, machine intelligence, and decision intelligence. He currently serves as Principal Technical Director and Global Black Belt for Artificial Intelligence at Microsoft, working with strategic customers on generative AI and machine learning transformation initiatives. With more than 25 years in technology and 18 years focused on AI and machine intelligence, Bart combines deep technical expertise with business strategy and enterprise-scale delivery. He is also an author, startup advisor, and active contributor to the global AI community.