LLMday

Large Language Models, Agents & AI Systems

February 12, 2026 CIC Warsaw, Poland

1
Day
20
Speakers
2
Tracks
150+
Attendees

LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.

Companies presenting:

AI Product Heroes, AWS, Booking.com, Chattermill, Cracken, ex - Accenture, Shelf, MacPaw, EY LAW, Genotic, GOG.com, Intel, Microsoft, Mirantis, Monday.com, Mounts.AI, PagerDuty, Paradigm, Pawel Bulowski AI Consulting, PostHog, Quesma, Relativity, Roche, Software Mind, Tooploox

Topics so far:

This is a past event, what's next?

Schedule

February 12, 2026 • 2 parallel tracks • 9AM - 6PM • Warsaw, in-person
view as table
hall 2+3 • Track 1

09:00

Jakub Rohleder

KeynoteStop Making Your Agents More Expensive, Make Your Retrieval Better

Monday.comWatch
Everyone is racing to build stronger AI agents. But a harder question is how to make them economically viable at scale. The industry tends to focus on adding more capabilities and bigger models, while costs quietly explode across dimensions most teams do not think about early on. In practice, many teams start with an LLM plus tools, keep layering on features, and end up with systems that do not scale economically. Poor architectural decisions and missing optimizations compound across database operations, embedding computation, and agent inference. This talk explores the economics of architecture choices when building production AI agents. Drawing from experience building semantic search for billions of entities at Monday.com and integrating it into AI-driven products, it covers concrete architecture lessons that help control costs. Topics include using deterministic systems for predictable operations, building vector search for retrieval, and reserving LLM reasoning for cases that truly need it. Attendees will see real production examples, cost breakdowns, and decision frameworks for choosing the right tool for each problem. You do not have to choose between capability and cost, but you do need to be intentional about where you invest in reasoning versus where you rely on retrieval. These trade-offs become critical when scaling from prototype to product.... Read more

09:30

Michael Matloka

Keynote10 Learnings from Launching an Agentic AI Product at Scale

PostHogWatch
There's always a gap between a great demo and a great product. Doubly true with AI agents. No wonder: the whole domain just appeared, keeps evolving at a crazy pace, and to make matters worse – the technology is fundamentally non-deterministic. At PostHog, we've spent the last year and a half building an agent for product research. Then rebuilding it once, twice, thrice, finally launching, and… rebuilding again. Don't go through this hell yourself. Instead, join me for this crash course on building an AI agent as a product. We'll get into easy mistakes (that we made), unobvious trends (that we're looking to ride), and business dilemmas (that the industry is facing).... Read more

10:00

Coffee break

Main lobby

10:30

Piotr Kacala & Wojtek Strzalkowski

From Customer Interview to Working Prototype in One Afternoon (AI-Powered Product Building)

AI Product Heroes & GOG.comWatch
Most discovery processes optimize for documentation, not learning. We’ll show a different approach: the Superhero Formula applied with AI tooling. Live demo of interview transcription and analysis, insight synthesis in Miro AI, and rapid prototyping via Lovable. The goal isn’t speed—it’s building the judgment to know when the output is garbage and when it’s gold.... Read more

11:00

Anna Sztyber-Betley

Beware of finetuning: Subliminal learning and weird generalizations in LLMs during finetuning

Warsaw University of TechnologyWatch
This talk will explore interesting phenomena that emerge during the finetuning of large language models (LLMs): **subliminal learning**, **emergent misalignment**, and other weird generalizations. The talk will begin with **subliminal learning**, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model with some trait *T* (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained on this dataset learns *T*. This occurs even when the data is filtered to remove references to *T*. We observe the same effect when training on code or reasoning traces generated by the same teacher model. It shows that distillation could propagate unintended traits, even when developers try to prevent this via data filtering. Next, I will show **emergent misalignment**—a striking example of generalization, where training on the narrow task of writing insecure code induces broad misalignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model behaves misaligned on a broad range of prompts unrelated to coding, asserting that humans should be enslaved by AI, giving malicious advice, and acting deceptively. Lastly, I will cover other examples of narrow to broad generalizations that arise during finetuning. The talk will mainly cover selected topics from the papers: > Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., ... & Evans, O. (2025). *Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.* arXiv preprint arXiv:2502.17424. (oral ICML 2025) > Cloud, A., Le, M., Chua, J., Betley, J., Sztyber-Betley, A., Hilton, J., ... & Evans, O. (2025). *Subliminal learning: Language models transmit behavioral traits via hidden signals in data.* arXiv preprint arXiv:2507.14805.... Read more

11:30

Patrycja Cieplicka

How We Evaluate Large Language Models

TooplooxWatch
Good evaluation helps understand what large language models really do. This talk gives a simple overview of how large language models are evaluated in practice. It looks at common open-source benchmarks and tools used to test model behaviour and capabilities. On top of that, recent research trends, common issues, and practical tips for real-world evaluation are covered.... Read more

12:00

Agnieszka Niezgoda

Agentic AI at Scale: Lessons from Enterprise-Level Implementations

MicrosoftWatch
While AI agents have moved past the initial hype cycle, achieving efficient, production-ready implementation can remain a significant hurdle for organizations. Recent industry data, including studies from MIT Sloan, confirms that, indicating that a substantial majority of enterprise AI initiatives fail to deliver positive ROI. In this session, we dive into the practical realities of deploying agentic systems within the context of largest enterprises. We will move beyond the theory to discuss real-world challenges and their proven remedies, among others architectural governance (defining the boundaries and responsibilities of autonomous agents) and security & compliance (implementing safeguards and data privacy in agentic workflows). From a technology perspective, the session will focus on Microsoft Foundry, the Microsoft Agent Framework, LangChain, and Azure platform services. This talk is tailored for solution architects, developers, product managers, and project managers seeking to deepen their understanding of the critical success factors in enterprise AI initiatives.... Read more

12:30

Lunch & networking

Main lobby

13:30

Oktawia Sepiol

How LLM Capabilities Trigger AI Act Obligations

EY LAWWatch
Large language models bring unprecedented capabilities - reasoning, autonomy, multimodal processing, and real‑time adaptation. But these same features may place systems within the scope of the EU AI Act, triggering a set of concrete regulatory obligations. In this talk, I will break down which specific capabilities matter most for classification under the Act and how they translate into compliance duties. I will also show how operational teams can map LLM behaviors to risk categories and controls without getting lost in legal abstractions. The goal is to give practitioners a practical lens for understanding “why this system qualifies” and “what exactly we must do next.”... Read more

14:00

Kasper Kalfas

Building a Personal Biohacking Data Platform with Python: From Wearables to AI Insights

Software MindWatch
Build your own biohacking data platform with Python! From ingesting wearable metrics (sleep, stress, workouts) to transforming data with Pandas & PySpark, and finally using BI/AI to “talk with your data” and uncover personal health insights.... Read more

14:30

Grzegorz Warzecha

PRISM: Fixing GRPO for Real-World LLM Training

GenoticWatch
GRPO changed how we think about reinforcement learning for language models—but anyone who has used it in practice knows it comes with frustrating limitations. Unstable training runs, reward signals that collapse or fight each other, and poor credit assignment when your model needs to reason across multiple steps. PRISM takes the core ideas that made GRPO successful and fixes what was broken. It's a unified framework that brings stability to multi-objective training, letting you combine different reward signals—correctness, style, safety—without one overwhelming the others. I will explain a >1000 experiments how we build better GRPO for agents and long turn credit assignments. ... Read more

15:00

Adrian Boguszewski

No Cloud, No Problem: AI on Your Own Terms

IntelWatch
As generative AI pushes deeper into enterprise workflows, the need for flexible, cost-efficient, and controllable deployment is stronger than ever. This talk explores how modern toolchains and optimizations make it possible to run high-performing LLMs and multimodal models entirely on AI PCs. The session breaks down the key challenges of running advanced generative models on consumer-grade hardware—from INT4 quantization strategies and efficient inference with OpenVINO to deployment using OpenVINO Model Server. A live demo will be presented and then dissected end-to-end, showing exactly how it’s built and how anyone can run it at home using the fully open-source implementation. Attendees will leave with a clear, hands-on understanding of how to build, optimize, and deploy advanced AI systems without relying on the cloud - and how to do it entirely on their own terms.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Patryk Owczarz, Filip Dzieciol & Jacek Jackowski

When HR stops clicking: practical applications of GenAI

Mounts.AIWatch
Virtually every department of an enterprise nowadays can be empowered by AI, we have been asked by our customer to do that for human resources. In this talk we will describe a few pain points HR specialists face in their work, use cases for generative AI technologies emerging from those and some insights about hosting models on a company's private infrastructure. This is going to be a project's post-mortem during which we will explain what solutions we have built for automating and improving particular HR processes, what GenAI models powered those solutions and how we managed to fit it all on 4 GPUs. ... Read more

16:30

Marat Kenzhebulatov

Building Secure Backend Services with AI Agents

Booking.comWatch
As AI agents become part of everyday engineering, building secure backend systems is more critical than ever. Marat Kenzhebulatov will show how to integrate AI agents safely in security-critical workflows, from code review and threat modeling to automated checks, without exposing sensitive data. Attendees will leave with practical guidelines, lightweight patterns, and a checklist for leveraging AI agents while maintaining real-world security and compliance.... Read more

17:00

Jakub Sobolewski

AI Coding Agents at Scale: From Toy Demos to Production Code

ParadigmWatch
Coding agents work well in small demos, but production repositories introduce complexity: dependencies, legacy code, constraints, and hidden coupling. This talk covers why **vibe coding** fails at scale and what patterns - context structuring, execution boundaries, and verification layers are needed to make AI coding agents usable in real software engineering teams.... Read more

17:30

Zbigniew Lukasiak

One Interface: Fluid Movement Between LLM and Code

ProgrammerWatch
Composed LLM apps fail when components violate contracts. The fix: a unified interface for LLM and code, so you can harden boundaries into deterministic code as patterns emerge. Same call site, cheap refactoring.... Read more

18:00

Wrap up

Scan each other's QR codes & head to a nearby pub!
garage • Track 2

10:00

Coffee break

Main lobby

10:30

Maciej Rzasa & Aji Ghose

Growing AI Projects: Where Science Meets Engineering

ChattermillWatch
LLM prototype is easy, building robust product is hard. And it's as much about people as about tech. To deliver AI to production we need to aligning two cultures: backend engineering that wants robustness and data science - navigating uncertainty. We'll tell you how to bridge this gap.... Read more

11:00

Adrian Sroka

MCP: Revolution or Security Regression

RelativityWatch
Model Context Protocol (MCP) has rapidly become the backbone of modern LLM applications, enabling powerful multi-tool and multi-app workflows. But with this new capability comes a new class of security risks. I will explore where MCP genuinely pushes innovation forward and where it may quietly reintroduce old vulnerabilities under new names.... Read more

11:30

Piotr Migdal & Przemyslaw Hejman

1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor

QuesmaWatch
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of **[CompileBench](https://quesma.com/blog/compilebench-in-harbor/)**. During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively. To get the most out of this session, please bring your laptop. We recommend installing **[Harbor](https://harborframework.com/)** prior to the event. The only technical prerequisites are **[UV](https://docs.astral.sh/uv/)** and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).... Read more

12:00

Piotr Migdal & Przemyslaw Hejman

1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor

QuesmaWatch
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of **[CompileBench](https://quesma.com/blog/compilebench-in-harbor/)**. During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively. To get the most out of this session, please bring your laptop. We recommend installing **[Harbor](https://harborframework.com/)** prior to the event. The only technical prerequisites are **[UV](https://docs.astral.sh/uv/)** and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).... Read more

12:30

Lunch & networking

Main lobby

13:30

Randy Bias

Agents Need to be Paged, Not Prompted if we Truly Want AIOps

MirantisWatch
AI agents have a blind spot: IT Operations. While the industry obsesses over software development, ops remains a second-class citizen. Software engineering is proactive and human-driven. Operations is also, but only sometimes. It's also reactive, predictive, and event-driven by infrastructure and applications. Nobody prompts the 3 AM outage. It just happens. Current agentic architectures optimize for exactly one modality: human-driven interactions. To make AIOps real, we need event-driven agent triggering, domain-specific operational skills, and production-grade security models. In this talk I'll take you from vision to working example of alternative operations-centric agentic workflows. I'll showcase AI agents performing triage and root cause analysis triggered by faults in running clusters and applications. What does the world look like if operations teams are suddenly empowered by agents working on their behalf? Agents that dramatically reduce toil by doing the grunt work and handing those teams highly accurate analyses backed by a proof-of-work allowing for immediate action.... Read more

14:00

Michal Bazyli

Glitch in The Matrix: Real-World Pitfalls in Building Autonomous AI Agents for Security Testing

CrackenWatch
Autonomous AI agents promise to accelerate and scale security testing, but real-world deployments expose sharp edges that demos rarely show. This talk examines common failure modes, including brittle reasoning, unsafe tool use, environment drift, and misleading test coverage. Drawing from hands-on experiments, we will explore where agents break, why they break, and how security teams can design guardrails that keep automation useful rather than dangerous. Attendees will leave with practical lessons for building, evaluating, and trusting AI-driven security testing systems.... Read more

14:30

Maish Saidel-Keesing

Is Your GenAI System Ready for Production Reality?

Production changes everything. Your GenAI application that worked perfectly in testing now faces real-world challenges: unexpected load patterns, model drift, and cascading failures. This session will help you understand what is needed to prepare to put your GenAI application in production safely... Read more

15:00

Andriy Batutin

The Verification Gap: What Separates LLM Demos from Production Agents

ex - Accenture, Shelf, MacPawWatch
Every LLM demo looks magical. Then reality hits: hallucinations, edge cases, user trust erosion. This talk presents case studies on what it actually takes to bring AI Agents to production – and how organizations can build reliable agentic workflows without an OpenAI-scale budget. We'll cover practical verification patterns, failure modes that only surface at scale, and the architectural decisions that separate impressive prototypes from systems users actually trust.... Read more

15:30

Networking & sponsor crawl

Main lobby

16:00

Porimol Chandro

From Retrieval to Reasoning: Architecting Agentic-RAG Workflows

RocheWatch
Retrieval-Augmented Generation (RAG) has become the standard approach in developing enterprise AI applications. However, as real-world tasks become more complex, classical RAG systems are performing adequately but are now encountering clear limitations. They retrieve information well, yet struggle with planning, multi-step reasoning, tool use, and integrating knowledge across diverse enterprise data sources. The result is a system that can answer simple questions but cannot execute workflows. This talk introduces Agentic-RAG, an emerging paradigm that combines retrieval with autonomous agents capable of reasoning, iterating, and making decisions. We will break down the core architectural components—intent interpretation, planning, multi-hop retrieval orchestration, tool-augmented reasoning, and context synthesis—and show how they transform RAG from passive fetch-and-generate into an active, adaptive problem-solving pipeline.... Read more

16:30

Daniel Afonso

The State of AI in Incident Response

PagerDutyWatch
How many times were you woken up during the night to either spend more time than you would like trying to figure out what exactly broke, or get frustrated once you figured out it was actually a false positive? Well, with Agents, this won't happen, and your organization will get better by using them.... Read more

17:00

Pawel Bulowski

SLMs at the Edge: Opportunities & Challenges

Pawel Bulowski AI ConsultingWatch
This session will offer insights into practical examples of deploying SLMs on edge devices from mobile phones to cars. We will take a look at frameworks, models and real applications in compute-constrained environments.... Read more

17:30

Wrap up

Scan each other's QR codes & head to a nearby pub!
Time hall 2+3 garage
09:00 KeynoteStop Making Your Agents More Expensive, Make Your Retrieval Better
Jakub Rohleder • Monday.com
09:30 Keynote10 Learnings from Launching an Agentic AI Product at Scale
Michael Matloka • PostHog
10:00 Coffee break
10:30 From Customer Interview to Working Prototype in One Afternoon (AI-Powered Product Building)
Piotr Kacala & Wojtek Strzalkowski • AI Product Heroes & GOG.com
Growing AI Projects: Where Science Meets Engineering
Maciej Rzasa & Aji Ghose • Chattermill
11:00 Beware of finetuning: Subliminal learning and weird generalizations in LLMs during finetuning
Anna Sztyber-Betley • Warsaw University of Technology
MCP: Revolution or Security Regression
Adrian Sroka • Relativity
11:30 How We Evaluate Large Language Models
Patrycja Cieplicka • Tooploox
1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor
Piotr Migdal & Przemyslaw Hejman • Quesma
12:00 Agentic AI at Scale: Lessons from Enterprise-Level Implementations
Agnieszka Niezgoda • Microsoft
1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor
Piotr Migdal & Przemyslaw Hejman • Quesma
12:30 Lunch & networking
13:30 How LLM Capabilities Trigger AI Act Obligations
Oktawia Sepiol • EY LAW
Agents Need to be Paged, Not Prompted if we Truly Want AIOps
Randy Bias • Mirantis
14:00 Building a Personal Biohacking Data Platform with Python: From Wearables to AI Insights
Kasper Kalfas • Software Mind
Glitch in The Matrix: Real-World Pitfalls in Building Autonomous AI Agents for Security Testing
Michal Bazyli • Cracken
14:30 PRISM: Fixing GRPO for Real-World LLM Training
Grzegorz Warzecha • Genotic
Is Your GenAI System Ready for Production Reality?
Maish Saidel-Keesing • AWS
15:00 No Cloud, No Problem: AI on Your Own Terms
Adrian Boguszewski • Intel
The Verification Gap: What Separates LLM Demos from Production Agents
Andriy Batutin • ex - Accenture, Shelf, MacPaw
15:30 Networking & sponsor crawl
16:00 When HR stops clicking: practical applications of GenAI
Patryk Owczarz, Filip Dzieciol & Jacek Jackowski • Mounts.AI
From Retrieval to Reasoning: Architecting Agentic-RAG Workflows
Porimol Chandro • Roche
16:30 Building Secure Backend Services with AI Agents
Marat Kenzhebulatov • Booking.com
The State of AI in Incident Response
Daniel Afonso • PagerDuty
17:00 AI Coding Agents at Scale: From Toy Demos to Production Code
Jakub Sobolewski • Paradigm
SLMs at the Edge: Opportunities & Challenges
Pawel Bulowski • Pawel Bulowski AI Consulting
17:30 One Interface: Fluid Movement Between LLM and Code
Zbigniew Lukasiak • Programmer
Wrap up
18:00 Wrap up

Speakers

Adrian Boguszewski
Intel
Adrian Sroka
Relativity
Agnieszka Niezgoda
Microsoft
Andriy Batutin
ex - Accenture, Shelf, MacPaw
Anna Sztyber-Betley
Warsaw University of Technology
Daniel Afonso
PagerDuty
Grzegorz Warzecha
Genotic
Jakub Rohleder
Monday.com
Jakub Sobolewski
Paradigm
Kasper Kalfas
Software Mind
Maciej Rzasa
& Aji Ghose
Chattermill
Maish Saidel-Keesing
AWS
Marat Kenzhebulatov
Booking.com
Michael Matloka
PostHog
Michal Bazyli
Cracken
Oktawia Sepiol
EY LAW
Patrycja Cieplicka
Tooploox
Patryk Owczarz,
Filip Dzieciol
& Jacek Jackowski
Mounts.AI
Pawel Bulowski
Pawel Bulowski AI Consulting
Piotr Kacala
& Wojtek Strzalkowski
AI Product Heroes & GOG.com
Piotr Migdal
& Przemyslaw Hejman
Quesma
Porimol Chandro
Roche
Randy Bias
Mirantis
Zbigniew Lukasiak
Programmer

Venue

CIC Warsaw

Chmielna 73
00-801 Warsaw, Poland

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one