LLMday

Large Language Models, Agents & AI Systems

March 6, 2026 VIAM, New York, US

1
Day
10+
Speakers
1
Track
100+
Attendees

LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.

Companies presenting:

Amazon, Anthropic, AWS, Fearless, groundcover, LLMday, MetalBear, New Jersey Institute of Technology, Redis, Snowflake, solo.io, Veritas Automata

Topics so far:
...and more

This is a past event, what's next?

Schedule

March 6, 2026 • single track • 9AM - 5:30PM • NYC, in-person
View
Track 1 • main room

09:00

LLMday Team

KeynoteWelcome message

LLMday
Hello everyone, let us walk you through the day!... Read more

09:30

Karen Zhou

KeynoteBuilding Evaluation Systems for AI Coding Agents at Scale

Anthropic
As AI coding assistants move from novelty to daily workflow, how do we know they're actually helping? This talk explores the technical and philosophical challenges of evaluating AI coding agents when they're deployed to millions of users. Turning user feedback into signal: Users complain, report issues, and express frustration—but this unstructured feedback is gold for improving models. I'll discuss approaches for transforming messy real-world feedback into structured evaluation datasets, including clustering techniques and rubric generation for consistent assessment. Measuring subtle failures: Some problems are easy to benchmark; others are felt more than measured. "Laziness," overconfidence, and instruction-following failures don't show up in traditional evals. I'll share frameworks for detecting these behavioral issues through automated assessment pipelines. The evaluation bottleneck: In my experience, evaluation infrastructure often becomes the limiting factor for model improvement. I'll discuss why building robust eval systems deserves as much engineering investment as the models themselves, and what that looks like in practice. Multi-agent architectures and new challenges: As coding agents become more autonomous and collaborative, evaluation gets harder. How do you assess a swarm? What metrics matter when agents coordinate on complex tasks? Closing the loop: Connecting evaluation back to training through RL and human feedback pipelines—what works, what doesn't, and where the field is heading.... Read more

10:00

Coffee break

Main lobby

10:30

Adna Zujo Lakisic

The Limits of Vibe Coding: What Happens When AI-Generated Code Lands in Your Cluster

MetalBearWatch
Most people aren't vibe coding an entire app from scratch and shipping it to production without a second thought. What’s more likely to happen is more subtle and potentially riskier. You've got a pre-existing service running in your cluster, it's been reviewed, it has access to your database, your secrets, your internal APIs. Then someone asks an AI assistant to add a feature. The generated code works, great! But, it also hardcodes an API key, or passes raw model output to users, or creates a public ingress to an internal service. And it lands inside a system that already has the keys to everything. The recent Moltbook incident made these blaring security concerns around AI generated code blaringly apparent, a single missing config exposed 1.5 million API keys. At a certain point AI-generated code will make its way into Kubernetes clusters and the security implications are real, are we ready to address them? In this talk we'll break down the specific risks that AI-generated code introduces when it lands in Kubernetes, from pod security and secret management to network exposure and LLM workloads. Then we'll build the fix: open-source security skills you drop into your AI coding assistant so it generates secure defaults automatically. We'll validate those skills against a live cluster using mirrord, no image rebuilds, no redeploys, just fast feedback until the policies actually work. You'll leave with a repo you can use today and an answer to the question every team shipping AI-generated code needs to figure out: if AI writes the code, what reviews it?... Read more

11:00

Michael Levan

Securing & Building Agentic OSS Environments

solo.ioWatch
The most popular advancements in cloud-native technology at this time are LLMs and AI Agents, which can help you troubleshoot your environment, build new environments, and even template a codebase. The two current problems are that there's no way to secure the traffic (e.g, reaching out to MCP servers, workload identity and authentication (SPIRE), and A2A), and Agents can't run by default in Kubernetes. Luckily, there are two open-source tools to help with this. In this session, you'll learn about two tools that can help you on this journey: agentgateway and kagent.... Read more

11:30

Tim Spann

From TrafficAI to GhostBreakers: Building Stateful AI Agents with Cortex, OpenFlow, and Snowflake Postgres

SnowflakeWatch
Generative AI is more than stateless chatbots. Real value lies in stateful, production-grade agents that act on your data. This session is a developer's blueprint for building these agents natively in Snowflake.... Read more

12:00

Noam Levy

Don't Panic: It's Not Just You - LLM Trace Sampling Doesn't Work

groundcoverWatch
Traditional distributed tracing relies on sampling strategies optimized for microservices: keep slow traces, keep error traces, and drop the rest. This approach fundamentally breaks down for LLM-powered systems. In LLM traces, latency & error signals are often meaningless: long latencies are expected, short latencies can still hide severe quality failures, and “errors” rarely correspond to the user-facing outcome. The true signal lies in the semantic content of the trace itself - prompts, intermediate reasoning steps, tool calls, and model outputs. As a result, conventional head-based or tail-based sampling systematically discards the most important traces while retaining large volumes of low-value data. This talk will show that effective observability for LLM systems requires a shift from signal-based sampling to content-aware sampling. The only viable way to sample intelligently is to run evaluations directly on span contents, scoring traces by quality, safety, correctness, and alignment with task intent. We explore how eval-driven sampling reframes tracing from performance diagnostics to outcome diagnostics, and why it is essential for operating LLM systems reliably at scale.... Read more

12:30

Lunch & networking

Main lobby

13:30

Ben Savage

Exploring the Convergence of Machine Learning, AI, and Cloud-Native Technologies in Automation

Veritas AutomataWatch
Harnessing the power of Machine Learning (ML), Artificial Intelligence (AI), and cloud-native technologies is no longer optional; it's a necessity for driving scalable automation and innovation.... Read more

14:00

Alisson Sol

The way to Edge AI

Stealth StartupWatch
As telecommunications evolves toward 6G, Radio Access Networks (RAN) will fundamentally transform application architecture beyond today's approach of endpoints calling LLMs on remote data centers. This presentation examines three critical software development challenges for the next decade: Real-time AI at the Edge: Moving from cloud-dependent inference to distributed processing meeting RAN's strict latency requirements. Multi-tenant Resource Optimization: Transitioning to shared platforms where AI workloads coexist while maintaining service-level agreements. Federated Learning and Privacy: Shifting from centralized training to federated approaches preserving privacy across distributed networks. The goal is sharing a roadmap for AI-native infrastructure where intelligence lives at the network edge.... Read more

14:30

Ahsan Ali

10 Lessons from building High-Stakes Agents

Moving Generative AI from proof-of-concept to production in high-stakes industries like Fintech and Healthcare requires a fundamental strategic shift. Drawing from real-world deployments at the AWS GenAI Innovation Center, this session explores ten critical lessons learned building autonomous agents for complex use cases. We will dive into "Model as a Judge" evaluation frameworks, architectural trade-offs between rigid pipelines and flexible agents, and why hallucinations are essential signals for system optimization. Attendees will receive a framework for prioritizing accuracy, building auditable reasoning chains, and creating internal feedback loops for self-improving deployments.... Read more

15:00

Jeremy Curcio

Engineering Better Prompts for AI Assisted Development

FearlessWatch
When asking AI to help refactor a method, do you get back something that runs, but would never pass review? The difference between frustrating and useful AI interactions comes down to how you build your prompt. This talk offers practical advice you can apply to get better results.... Read more

15:30

Nitin Kanukolanu

Context Matters

RedisWatch
As LLMs and agents become dramatically smarter, the real challenge shifts from model capability to context engineering. Powerful agents can reason, plan, and use tools — but without the right context, they hallucinate, waste tokens, and degrade in performance. This talk explores why context engineering is essential for building reliable, production-ready AI systems. We’ll move beyond traditional RAG to agentic workflows, examine common context failure modes, and discuss how to unify memory, structured data, APIs, and retrieval into a cohesive context strategy.... Read more

16:00

Harish Mandhadi

Reimagined AI-DLC Manifesto: Moving Beyond the Productivity Paradox to Engineering Predictability

As organizations race to integrate Large Language Models (LLMs) into their development workflows, many are hitting a “Productivity Paradox.” Developers often perceive a 20 percent gain in speed, while deeper analysis shows they may actually be 20 percent less productive due to a lack of structured methodology. This session introduces the AI-Driven Development Lifecycle (AI-DLC), a reimagined framework designed not to retrofit AI into existing Agile processes, but to build a new paradigm centered on brain to brain alignment between humans and machines. We will explore critical lessons learned from real world use cases and experiments, including the dangers of vibe coding and the necessity for developers to understand and defend every line of code as the ultimate owner. Attendees will learn high impact engineering techniques such as semantic context compression to prevent AI from entering infinite loops, and the semantics per token ratio to maximize output quality. We will also discuss why treating AI as a confident intern rather than a senior engineer is essential for maintaining production grade standards. Crucially, this talk addresses how leaders should measure value in an AI native era. We move beyond gameable traditional metrics like lines of code or code accepted, which fail to reflect true business outcomes. Instead, we present a framework focused on inception to operation speed and predictability, demonstrating how AI-DLC rituals such as Inception, Mob Elaboration, and Mob Construction drive measurable results.... Read more

16:30

Pranav Kowadkar

Multi-Agent Architectures: Solving Production LLM Reliability at Scale

New Jersey Institute of Technology Watch
Production LLM systems face impossible tradeoffs: latency vs accuracy, cost vs quality, speed vs reliability. I'll demonstrate how multi-agent architectures with LLM-as-Judge patterns solve these challenges through live examples, showing how task decomposition and peer review eliminate production bottlenecks while driving errors to near-zero. Attendees will see practical patterns they can implement immediately.... Read more

17:00

Suvendu Mohanty

Fine‑Tuning LLMs for Voice‑First Consumer Experiences

AmazonWatch
Voice-first consumer devices—smart speakers, wearables, in‑home assistants, and connected appliances—demand language models that are not only accurate, but fast, reliable, and deeply aligned with user expectations in real‑world environments. This talk explores practical methods for fine‑tuning large language models (LLMs) to deliver high‑quality, low‑friction voice interactions at scale. I will outline a structured approach to supervised fine‑tuning (SFT) for voice-driven use cases, including dataset design strategies that capture conversational latency patterns, prosody-driven intent, and the ambiguities inherent in spoken queries. The session will also address real evaluation challenges: how to measure dialog quality, naturalness, grounding, and long‑horizon task execution when traditional text benchmarks fall short. Particular attention will be given to diagnosing and reducing hallucinations in consumer contexts, where incorrect answers can undermine trust, cause user frustration, or trigger unintended physical-world actions. Attendees will leave with a clear picture of what it takes to adapt an LLM into a dependable, voice‑native system—from fine‑tuning workflows and guardrail construction to continuous evaluation loops informed by real user behavior. This talk aims to provide a practical roadmap for teams building the next generation of on‑device and cloud‑connected voice experiences.... Read more

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Track 1Main room
  1. 09:00
    Keynote Welcome message
    LLMday Team · LLMday
  2. 09:30
  3. 10:00
    Coffee break
  4. 10:30
  5. 11:00
  6. 11:30
  7. 12:00
  8. 12:30
    Lunch & networking
  9. 13:30
  10. 14:00
    The way to Edge AI
    Alisson Sol · Stealth Startup
  11. 14:30
  12. 15:00
  13. 15:30
    Context Matters
    Nitin Kanukolanu · Redis
  14. 16:00
  15. 16:30
    Multi-Agent Architectures: Solving Production LLM Reliability at Scale
    Pranav Kowadkar · New Jersey Institute of Technology
  16. 17:00
  17. 17:30
    Wrap up
Time main room
09:00 Welcome message
LLMday Team • LLMday
09:30 Building Evaluation Systems for AI Coding Agents at Scale
Karen Zhou • Anthropic
10:00 Coffee break
10:30 The Limits of Vibe Coding: What Happens When AI-Generated Code Lands in Your Cluster
Adna Zujo Lakisic • MetalBear
11:00 Securing & Building Agentic OSS Environments
Michael Levan • solo.io
11:30 From TrafficAI to GhostBreakers: Building Stateful AI Agents with Cortex, OpenFlow, and Snowflake Postgres
Tim Spann • Snowflake
12:00 Don't Panic: It's Not Just You - LLM Trace Sampling Doesn't Work
Noam Levy • groundcover
12:30 Lunch & networking
13:30 Exploring the Convergence of Machine Learning, AI, and Cloud-Native Technologies in Automation
Ben Savage • Veritas Automata
14:00 The way to Edge AI
Alisson Sol • Stealth Startup
14:30 10 Lessons from building High-Stakes Agents
Ahsan Ali • AWS
15:00 Engineering Better Prompts for AI Assisted Development
Jeremy Curcio • Fearless
15:30 Context Matters
Nitin Kanukolanu • Redis
16:00 Reimagined AI-DLC Manifesto: Moving Beyond the Productivity Paradox to Engineering Predictability
Harish Mandhadi • AWS
16:30 Multi-Agent Architectures: Solving Production LLM Reliability at Scale
Pranav Kowadkar • New Jersey Institute of Technology
17:00 Fine‑Tuning LLMs for Voice‑First Consumer Experiences
Suvendu Mohanty • Amazon
17:30 Wrap up

Speakers

Adna Zujo Lakisic
MetalBear
Ahsan Ali
AWS
Alisson Sol
Stealth Startup
Ben Savage
Veritas Automata
Harish Mandhadi
AWS
Jeremy Curcio
Fearless
Karen Zhou
Anthropic
LLMday Team
LLMday
Michael Levan
solo.io
Nitin Kanukolanu
Redis
Noam Levy
groundcover
Pranav Kowadkar
New Jersey Institute of Technology
Suvendu Mohanty
Amazon
Tim Spann
Snowflake

Venue

VIAM HQ

1900 Broadway, 6th Floor,
New York, NY 10023

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one