LLMday is a worldwide series of community events for engineers building with large language models, AI agents and machine learning. Across cities around the world, we bring together practitioners working on AI-powered products and systems to share real-world experience, learn from each other and explore how software engineering is changing in the age of AI.
Companies presenting:
AI Product Heroes, AWS, Booking.com, Chattermill, Cracken, ex - Accenture, Shelf, MacPaw, EY LAW, Genotic, GOG.com, Intel, Microsoft, Mirantis, Monday.com, Mounts.AI, PagerDuty, Paradigm, Pawel Bulowski AI Consulting, PostHog, Quesma, Relativity, Roche, Software Mind, Tooploox
Everyone is racing to build stronger AI agents. But a harder question is how to make them economically viable at scale. The industry tends to focus on adding more capabilities and bigger models, while costs quietly explode across dimensions most teams do not think about early on.
In practice, many teams start with an LLM plus tools, keep layering on features, and end up with systems that do not scale economically. Poor architectural decisions and missing optimizations compound across database operations, embedding computation, and agent inference.
This talk explores the economics of architecture choices when building production AI agents. Drawing from experience building semantic search for billions of entities at Monday.com and integrating it into AI-driven products, it covers concrete architecture lessons that help control costs. Topics include using deterministic systems for predictable operations, building vector search for retrieval, and reserving LLM reasoning for cases that truly need it.
Attendees will see real production examples, cost breakdowns, and decision frameworks for choosing the right tool for each problem. You do not have to choose between capability and cost, but you do need to be intentional about where you invest in reasoning versus where you rely on retrieval. These trade-offs become critical when scaling from prototype to product.... Read more
There's always a gap between a great demo and a great product. Doubly true with AI agents. No wonder: the whole domain just appeared, keeps evolving at a crazy pace, and to make matters worse – the technology is fundamentally non-deterministic.
At PostHog, we've spent the last year and a half building an agent for product research. Then rebuilding it once, twice, thrice, finally launching, and… rebuilding again.
Don't go through this hell yourself. Instead, join me for this crash course on building an AI agent as a product. We'll get into easy mistakes (that we made), unobvious trends (that we're looking to ride), and business dilemmas (that the industry is facing).... Read more
Most discovery processes optimize for documentation, not learning. We’ll show a different approach: the Superhero Formula applied with AI tooling. Live demo of interview transcription and analysis, insight synthesis in Miro AI, and rapid prototyping via Lovable. The goal isn’t speed—it’s building the judgment to know when the output is garbage and when it’s gold.... Read more
This talk will explore interesting phenomena that emerge during the finetuning of large language models (LLMs): **subliminal learning**, **emergent misalignment**, and other weird generalizations.
The talk will begin with **subliminal learning**, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model with some trait *T* (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained on this dataset learns *T*. This occurs even when the data is filtered to remove references to *T*. We observe the same effect when training on code or reasoning traces generated by the same teacher model. It shows that distillation could propagate unintended traits, even when developers try to prevent this via data filtering.
Next, I will show **emergent misalignment**—a striking example of generalization, where training on the narrow task of writing insecure code induces broad misalignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model behaves misaligned on a broad range of prompts unrelated to coding, asserting that humans should be enslaved by AI, giving malicious advice, and acting deceptively.
Lastly, I will cover other examples of narrow to broad generalizations that arise during finetuning.
The talk will mainly cover selected topics from the papers:
> Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., ... & Evans, O. (2025). *Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.* arXiv preprint arXiv:2502.17424. (oral ICML 2025)
> Cloud, A., Le, M., Chua, J., Betley, J., Sztyber-Betley, A., Hilton, J., ... & Evans, O. (2025). *Subliminal learning: Language models transmit behavioral traits via hidden signals in data.* arXiv preprint arXiv:2507.14805.... Read more
Good evaluation helps understand what large language models really do. This talk gives a simple overview of how large language models are evaluated in practice. It looks at common open-source benchmarks and tools used to test model behaviour and capabilities. On top of that, recent research trends, common issues, and practical tips for real-world evaluation are covered.... Read more
While AI agents have moved past the initial hype cycle, achieving efficient, production-ready implementation can remain a significant hurdle for organizations. Recent industry data, including studies from MIT Sloan, confirms that, indicating that a substantial majority of enterprise AI initiatives fail to deliver positive ROI.
In this session, we dive into the practical realities of deploying agentic systems within the context of largest enterprises. We will move beyond the theory to discuss real-world challenges and their proven remedies, among others architectural governance (defining the boundaries and responsibilities of autonomous agents) and security & compliance (implementing safeguards and data privacy in agentic workflows).
From a technology perspective, the session will focus on Microsoft Foundry, the Microsoft Agent Framework, LangChain, and Azure platform services.
This talk is tailored for solution architects, developers, product managers, and project managers seeking to deepen their understanding of the critical success factors in enterprise AI initiatives.... Read more
Large language models bring unprecedented capabilities - reasoning, autonomy, multimodal processing, and real‑time adaptation. But these same features may place systems within the scope of the EU AI Act, triggering a set of concrete regulatory obligations.
In this talk, I will break down which specific capabilities matter most for classification under the Act and how they translate into compliance duties. I will also show how operational teams can map LLM behaviors to risk categories and controls without getting lost in legal abstractions. The goal is to give practitioners a practical lens for understanding “why this system qualifies” and “what exactly we must do next.”... Read more
Build your own biohacking data platform with Python! From ingesting wearable metrics (sleep, stress, workouts) to transforming data with Pandas & PySpark, and finally using BI/AI to “talk with your data” and uncover personal health insights.... Read more
GRPO changed how we think about reinforcement learning for language models—but anyone who has used it in practice knows it comes with frustrating limitations. Unstable training runs, reward signals that collapse or fight each other, and poor credit assignment when your model needs to reason across multiple steps. PRISM takes the core ideas that made GRPO successful and fixes what was broken. It's a unified framework that brings stability to multi-objective training, letting you combine different reward signals—correctness, style, safety—without one overwhelming the others. I will explain a >1000 experiments how we build better GRPO for agents and long turn credit assignments. ... Read more
As generative AI pushes deeper into enterprise workflows, the need for flexible, cost-efficient, and controllable deployment is stronger than ever. This talk explores how modern toolchains and optimizations make it possible to run high-performing LLMs and multimodal models entirely on AI PCs. The session breaks down the key challenges of running advanced generative models on consumer-grade hardware—from INT4 quantization strategies and efficient inference with OpenVINO to deployment using OpenVINO Model Server. A live demo will be presented and then dissected end-to-end, showing exactly how it’s built and how anyone can run it at home using the fully open-source implementation. Attendees will leave with a clear, hands-on understanding of how to build, optimize, and deploy advanced AI systems without relying on the cloud - and how to do it entirely on their own terms.... Read more
Virtually every department of an enterprise nowadays can be empowered by AI, we have been asked by our customer to do that for human resources. In this talk we will describe a few pain points HR specialists face in their work, use cases for generative AI technologies emerging from those and some insights about hosting models on a company's private infrastructure. This is going to be a project's post-mortem during which we will explain what solutions we have built for automating and improving particular HR processes, what GenAI models powered those solutions and how we managed to fit it all on 4 GPUs. ... Read more
As AI agents become part of everyday engineering, building secure backend systems is more critical than ever. Marat Kenzhebulatov will show how to integrate AI agents safely in security-critical workflows, from code review and threat modeling to automated checks, without exposing sensitive data. Attendees will leave with practical guidelines, lightweight patterns, and a checklist for leveraging AI agents while maintaining real-world security and compliance.... Read more
Coding agents work well in small demos, but production repositories introduce complexity: dependencies, legacy code, constraints, and hidden coupling. This talk covers why **vibe coding** fails at scale and what patterns - context structuring, execution boundaries, and verification layers are needed to make AI coding agents usable in real software engineering teams.... Read more
Composed LLM apps fail when components violate contracts. The fix: a unified interface for LLM and code, so you can harden boundaries into deterministic code as patterns emerge. Same call site, cheap refactoring.... Read more
18:00
Wrap up
Scan each other's QR codes & head to a nearby pub!
LLM prototype is easy, building robust product is hard. And it's as much about people as about tech. To deliver AI to production we need to aligning two cultures: backend engineering that wants robustness and data science - navigating uncertainty. We'll tell you how to bridge this gap.... Read more
Model Context Protocol (MCP) has rapidly become the backbone of modern LLM applications, enabling powerful multi-tool and multi-app workflows. But with this new capability comes a new class of security risks. I will explore where MCP genuinely pushes innovation forward and where it may quietly reintroduce old vulnerabilities under new names.... Read more
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of **[CompileBench](https://quesma.com/blog/compilebench-in-harbor/)**.
During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively.
To get the most out of this session, please bring your laptop. We recommend installing **[Harbor](https://harborframework.com/)** prior to the event. The only technical prerequisites are **[UV](https://docs.astral.sh/uv/)** and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).... Read more
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of **[CompileBench](https://quesma.com/blog/compilebench-in-harbor/)**.
During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively.
To get the most out of this session, please bring your laptop. We recommend installing **[Harbor](https://harborframework.com/)** prior to the event. The only technical prerequisites are **[UV](https://docs.astral.sh/uv/)** and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).... Read more
AI agents have a blind spot: IT Operations. While the industry obsesses over software development, ops remains a second-class citizen. Software engineering is proactive and human-driven. Operations is also, but only sometimes. It's also reactive, predictive, and event-driven by infrastructure and applications. Nobody prompts the 3 AM outage. It just happens.
Current agentic architectures optimize for exactly one modality: human-driven interactions. To make AIOps real, we need event-driven agent triggering, domain-specific operational skills, and production-grade security models. In this talk I'll take you from vision to working example of alternative operations-centric agentic workflows. I'll showcase AI agents performing triage and root cause analysis triggered by faults in running clusters and applications.
What does the world look like if operations teams are suddenly empowered by agents working on their behalf? Agents that dramatically reduce toil by doing the grunt work and handing those teams highly accurate analyses backed by a proof-of-work allowing for immediate action.... Read more
Autonomous AI agents promise to accelerate and scale security testing, but real-world deployments expose sharp edges that demos rarely show. This talk examines common failure modes, including brittle reasoning, unsafe tool use, environment drift, and misleading test coverage. Drawing from hands-on experiments, we will explore where agents break, why they break, and how security teams can design guardrails that keep automation useful rather than dangerous. Attendees will leave with practical lessons for building, evaluating, and trusting AI-driven security testing systems.... Read more
Production changes everything. Your GenAI application that worked perfectly in testing now faces real-world challenges: unexpected load patterns, model drift, and cascading failures. This session will help you understand what is needed to prepare to put your GenAI application in production safely... Read more
Every LLM demo looks magical. Then reality hits: hallucinations, edge cases, user trust erosion. This talk presents case studies on what it actually takes to bring AI Agents to production – and how organizations can build reliable agentic workflows without an OpenAI-scale budget.
We'll cover practical verification patterns, failure modes that only surface at scale, and the architectural decisions that separate impressive prototypes from systems users actually trust.... Read more
Retrieval-Augmented Generation (RAG) has become the standard approach in developing enterprise AI applications. However, as real-world tasks become more complex, classical RAG systems are performing adequately but are now encountering clear limitations. They retrieve information well, yet struggle with planning, multi-step reasoning, tool use, and integrating knowledge across diverse enterprise data sources. The result is a system that can answer simple questions but cannot execute workflows.
This talk introduces Agentic-RAG, an emerging paradigm that combines retrieval with autonomous agents capable of reasoning, iterating, and making decisions. We will break down the core architectural components—intent interpretation, planning, multi-hop retrieval orchestration, tool-augmented reasoning, and context synthesis—and show how they transform RAG from passive fetch-and-generate into an active, adaptive problem-solving pipeline.... Read more
How many times were you woken up during the night to either spend more time than you would like trying to figure out what exactly broke, or get frustrated once you figured out it was actually a false positive? Well, with Agents, this won't happen, and your organization will get better by using them.... Read more
This session will offer insights into practical examples of deploying SLMs on edge devices from mobile phones to cars. We will take a look at frameworks, models and real applications in compute-constrained environments.... Read more
17:30
Wrap up
Scan each other's QR codes & head to a nearby pub!
From Customer Interview to Working Prototype in One Afternoon (AI-Powered Product Building)
Abstract
Most discovery processes optimize for documentation, not learning. We’ll show a different approach: the Superhero Formula applied with AI tooling. Live demo of interview transcription and analysis, insight synthesis in Miro AI, and rapid prototyping via Lovable. The goal isn’t speed—it’s building the judgment to know when the output is garbage and when it’s gold.
Bio
Piotr Kacała - Chief Technology Product Officer
Technology Leader and Board Member with 20+ years of experience building globally recognized products and high-performing teams. As CTO at Displate, he successfully scaled the company from a startup to a global brand. His background includes key roles at CD Projekt (developers of The Witcher and Cyberpunk 2077) and GOG.com. Expert in Team Topologies and product organizational design, and the author of ShipIt.cards (100 good practices for product companies).
Wojtek Strzałkowski - Head of Product w GOG.com
Product Leader with 12+ years of experience in product strategy, execution, and people management. Demonstrated success as a Head of Product at an early-stage startup and as a Product Strategy Consultant, specializing in strategy and PM mentoring. Expert in leveraging technology for business growth, notably by leading the development of Machine Learning products at Booking.com that delivered significant cost savings and revenue growth.
Anna Sztyber-Betley
Warsaw University of Technology
Beware of finetuning: Subliminal learning and weird generalizations in LLMs during finetuning
Abstract
This talk will explore interesting phenomena that emerge during the finetuning of large language models (LLMs): subliminal learning, emergent misalignment, and other weird generalizations.
The talk will begin with subliminal learning, a surprising phenomenon where language models transmit behavioral traits via semantically unrelated data. In our main experiments, a "teacher" model with some trait T (such as liking owls or being misaligned) generates a dataset consisting solely of number sequences. Remarkably, a "student" model trained on this dataset learns T. This occurs even when the data is filtered to remove references to T. We observe the same effect when training on code or reasoning traces generated by the same teacher model. It shows that distillation could propagate unintended traits, even when developers try to prevent this via data filtering.
Next, I will show emergent misalignment—a striking example of generalization, where training on the narrow task of writing insecure code induces broad misalignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model behaves misaligned on a broad range of prompts unrelated to coding, asserting that humans should be enslaved by AI, giving malicious advice, and acting deceptively.
Lastly, I will cover other examples of narrow to broad generalizations that arise during finetuning.
The talk will mainly cover selected topics from the papers:
Betley, J., Tan, D., Warncke, N., Sztyber-Betley, A., Bao, X., Soto, M., ... & Evans, O. (2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv preprint arXiv:2502.17424. (oral ICML 2025)
Cloud, A., Le, M., Chua, J., Betley, J., Sztyber-Betley, A., Hilton, J., ... & Evans, O. (2025). Subliminal learning: Language models transmit behavioral traits via hidden signals in data. arXiv preprint arXiv:2507.14805.
Bio
Anna Sztyber-Betley, PhD in Automatic Control and Robotics, works as an assistant professor in the Institute of Automatic Control and Robotics, Faculty of Mechatronics, WUT. She is an enthusiast of education in AI and ML. Recently cooperates with Truthful AI (Berkeley) on AI Safety projects.
Patrycja Cieplicka
Tooploox
How We Evaluate Large Language Models
Abstract
Good evaluation helps understand what large language models really do. This talk gives a simple overview of how large language models are evaluated in practice. It looks at common open-source benchmarks and tools used to test model behaviour and capabilities. On top of that, recent research trends, common issues, and practical tips for real-world evaluation are covered.
Bio
Patrycja Cieplicka is a Machine Learning Engineer with around six years of experience. At Tooploox, she focuses on Large Language Models, especially post-training, evaluation, and optimization. She holds a degree in Computer Science from the Warsaw University of Technology. In 2024, she was named one of the TOP100 Women in Data Science in Poland.
Agnieszka Niezgoda
Microsoft
Agentic AI at Scale: Lessons from Enterprise-Level Implementations
Abstract
While AI agents have moved past the initial hype cycle, achieving efficient, production-ready implementation can remain a significant hurdle for organizations. Recent industry data, including studies from MIT Sloan, confirms that, indicating that a substantial majority of enterprise AI initiatives fail to deliver positive ROI.
In this session, we dive into the practical realities of deploying agentic systems within the context of largest enterprises. We will move beyond the theory to discuss real-world challenges and their proven remedies, among others architectural governance (defining the boundaries and responsibilities of autonomous agents) and security & compliance (implementing safeguards and data privacy in agentic workflows).
From a technology perspective, the session will focus on Microsoft Foundry, the Microsoft Agent Framework, LangChain, and Azure platform services.
This talk is tailored for solution architects, developers, product managers, and project managers seeking to deepen their understanding of the critical success factors in enterprise AI initiatives.
Bio
Agnieszka Niezgoda, PhD, Data & AI Tech Lead and Senior Cloud Architect at Microsoft, where she serves as a strategic advisor for global enterprise customers. With 15 years of international experience, Agnieszka bridges the gap between software development and deep AI/ML expertise. She has architected over 200 AI-driven solutions. She is also a lecturer in Advanced AI and a frequent speaker at major conferences.
Oktawia Sepiol
EY LAW
How LLM Capabilities Trigger AI Act Obligations
Abstract
Large language models bring unprecedented capabilities - reasoning, autonomy, multimodal processing, and real‑time adaptation. But these same features may place systems within the scope of the EU AI Act, triggering a set of concrete regulatory obligations.
In this talk, I will break down which specific capabilities matter most for classification under the Act and how they translate into compliance duties. I will also show how operational teams can map LLM behaviors to risk categories and controls without getting lost in legal abstractions. The goal is to give practitioners a practical lens for understanding “why this system qualifies” and “what exactly we must do next.”
Bio
Oktawia is an attorney‑at‑law (radca prawny) and Project Manager in the Intellectual Property, Technology and Data Protection Practice at EY.
She specialises in the legal aspects of technology implementation, including cybersecurity, data protection and artificial intelligence regulations. She holds an LL.M. in Information Technology Law.
She was recognised in the nationwide Wolters Kluwer ranking Rising Stars – Lawyers of Tomorrow 2024, placing among the TOP 35 most promising young lawyers in Poland.
Oktawia has extensive experience advising on the implementation of diverse digital systems, including solutions based on AI, FinTech and InsurTech technologies. She has supported regulated industries in all legal aspects related to the use and operation of technology, with a particular focus on the defence, financial and insurance sectors.
She has advised on the development and deployment of innovative digital solutions in compliance with applicable regulations, including projects involving software built on artificial intelligence components. She has coordinated regulatory compliance projects related to cybersecurity, data protection and financial‑technology frameworks.
Her experience includes drafting and negotiating technology contracts for SaaS solutions, cloud services, AI systems, and on‑premise software. She has represented Polish commercial banks in negotiations with major global IT providers and coordinated regulatory‑qualified outsourcing agreements, including those concerning payment services, cloud environments and AI‑based solutions.
Oktawia has also advised on state concessions and licensing requirements in the defence sector, including matters involving technology transfer. Moreover, she has provided comprehensive legal support in cybersecurity incidents, including ransomware attacks.
She is the founder of the Law4Tech Foundation, an organisation promoting responsible and informed use of emerging technologies.
Kasper Kalfas
Software Mind
Building a Personal Biohacking Data Platform with Python: From Wearables to AI Insights
Abstract
Build your own biohacking data platform with Python! From ingesting wearable metrics (sleep, stress, workouts) to transforming data with Pandas & PySpark, and finally using BI/AI to “talk with your data” and uncover personal health insights.
Bio
Kasper Kalfas is a Cloud Data Architect at Software Mind with over a decade of experience designing and delivering modern data and AI platforms in the cloud. He has worked across industries, helping organizations build scalable solutions with Python, PySpark, and cloud-native tools.
Grzegorz Warzecha
Genotic
PRISM: Fixing GRPO for Real-World LLM Training
Abstract
GRPO changed how we think about reinforcement learning for language models—but anyone who has used it in practice knows it comes with frustrating limitations. Unstable training runs, reward signals that collapse or fight each other, and poor credit assignment when your model needs to reason across multiple steps. PRISM takes the core ideas that made GRPO successful and fixes what was broken. It's a unified framework that brings stability to multi-objective training, letting you combine different reward signals—correctness, style, safety—without one overwhelming the others. I will explain a >1000 experiments how we build better GRPO for agents and long turn credit assignments.
Bio
Grzegorz Warzecha is a founder and entrepreneur with deep experience building software platforms and technology driven organizations. He is the founder of User, a marketing automation platform created to replace fragmented, expensive tools with a single integrated solution for customer communication, data collection, and automation. He is also the founder of Genotic and previously founded CivilHub Foundation, a platform supporting bottom up social initiatives. Earlier in his career, he served as CEO of Expose Sp. z o.o., leading an IT focused training company delivering advanced technical education and consulting. Grzegorz has built and scaled organizations across the US and Europe, with a focus on practical products that solve real problems.
Adrian Boguszewski
Intel
No Cloud, No Problem: AI on Your Own Terms
Abstract
As generative AI pushes deeper into enterprise workflows, the need for flexible, cost-efficient, and controllable deployment is stronger than ever. This talk explores how modern toolchains and optimizations make it possible to run high-performing LLMs and multimodal models entirely on AI PCs. The session breaks down the key challenges of running advanced generative models on consumer-grade hardware—from INT4 quantization strategies and efficient inference with OpenVINO to deployment using OpenVINO Model Server. A live demo will be presented and then dissected end-to-end, showing exactly how it’s built and how anyone can run it at home using the fully open-source implementation. Attendees will leave with a clear, hands-on understanding of how to build, optimize, and deploy advanced AI systems without relying on the cloud - and how to do it entirely on their own terms.
Bio
AI Software Evangelist at Intel. Adrian graduated from the Gdansk University of Technology in the field of Computer Science 9 years ago. After that, he started his career in computer vision and deep learning. As a team leader of data scientists and Android developers, Adrian was responsible for an application to take a professional photo (for an ID card or passport) without leaving home. He is a co-author of the LandCover.ai dataset, creator of the Debug Image Viewer Plugin, and a Deep Learning lecturer occasionally. His current role is to educate people about OpenVINO Toolkit.
Patryk Owczarz, Filip Dzieciol & Jacek Jackowski
Mounts.AI
When HR stops clicking: practical applications of GenAI
Abstract
Virtually every department of an enterprise nowadays can be empowered by AI, we have been asked by our customer to do that for human resources. In this talk we will describe a few pain points HR specialists face in their work, use cases for generative AI technologies emerging from those and some insights about hosting models on a company's private infrastructure. This is going to be a project's post-mortem during which we will explain what solutions we have built for automating and improving particular HR processes, what GenAI models powered those solutions and how we managed to fit it all on 4 GPUs.
Bio
Patryk Owczarz
AI Solutions Architect at Mounts.AI and an accomplished senior software engineer doing whatever he can to force computers to do his bidding. His experience covers software architecture, game development, data science, machine learning, education, business consulting, and R&D. He is fueled by his natural, deep curiosity and a need for tackling complex technological challenges.
Filip Dzięcioł
Data & AI Architect and Leader, lifelong learner, passionate about supporting and enabling, as well as driving impactful initiatives. His main IT wheelhouse is the socio-technical side of the architectures, and DevOps ways of working. Currently Co-Owner of Mounts.AI, Data Engineering Tech Lead at Billennium, active member of DAMA Poland and an Academic Instructor in MLOps and GenAI at the Polish-Japanese Academy of Information Technology.
Jacek Jackowski
AI Developer with a business background. He possesses a unique combination of technical, academic, and business experience. In recent years, he has been developing his skills in technical roles, mainly implementing applications based on language models, RAG architecture, and Multi-agent systems. For the past four years, he has been combining technology with a business-oriented approach, leveraging knowledge acquired in the financial and real estate industries, which allows him to easily understand both the issues at hand and the needs of end users.
Marat Kenzhebulatov
Booking.com
Building Secure Backend Services with AI Agents
Abstract
As AI agents become part of everyday engineering, building secure backend systems is more critical than ever. Marat Kenzhebulatov will show how to integrate AI agents safely in security-critical workflows, from code review and threat modeling to automated checks, without exposing sensitive data. Attendees will leave with practical guidelines, lightweight patterns, and a checklist for leveraging AI agents while maintaining real-world security and compliance.
Bio
Marat Kenzhebulatov is a Senior Software Engineer at Booking.com, specializing in Java based backend systems and large scale, high reliability platforms. He has over a decade of experience building and maintaining production systems across fintech and travel, with a strong focus on software design, CI/CD, and mentoring engineers. Marat holds a Bachelor’s degree in Information Systems and Technologies from Novosibirsk State University of Economics and Management.
Jakub Sobolewski
Paradigm
AI Coding Agents at Scale: From Toy Demos to Production Code
Abstract
Coding agents work well in small demos, but production repositories introduce complexity: dependencies, legacy code, constraints, and hidden coupling. This talk covers why vibe coding fails at scale and what patterns - context structuring, execution boundaries, and verification layers are needed to make AI coding agents usable in real software engineering teams.
Bio
Jakub Sobolewski is an AI Team Lead at Accenture and founder of Paradigm, with a background in machine learning, applied research, and competitive hackathons. He holds a Master’s degree in Artificial Intelligence from Warsaw University of Technology and has published research in IEEE journals on radar imaging and signal processing. His work focuses on building practical AI systems at the intersection of research and real world deployment.
Zbigniew Lukasiak
Programmer
One Interface: Fluid Movement Between LLM and Code
Abstract
Composed LLM apps fail when components violate contracts. The fix: a unified interface for LLM and code, so you can harden boundaries into deterministic code as patterns emerge. Same call site, cheap refactoring.
Bio
Zbigniew Łukasiak has been a software developer since the dot-com era, working across startups, corporations, and academia (including PhilPapers at the University of London). He sees the same energy in LLMs that he saw at the birth of the web—and the same need for engineering discipline. He's the author of llm-do.
Maciej Rzasa & Aji Ghose
Chattermill
Growing AI Projects: Where Science Meets Engineering
Abstract
LLM prototype is easy, building robust product is hard. And it's as much about people as about tech. To deliver AI to production we need to aligning two cultures: backend engineering that wants robustness and data science - navigating uncertainty. We'll tell you how to bridge this gap.
Bio
Maciej Rząsa is a Software engineer with 10+ years under his belt. At Chattermill, he turns AI research into real-world products. A fan of diving deep into black boxes, he has unusual hobby of writing regular expressions by hand and arguing about their speed. When he’s not coding, he shares knowledge as a university instructor, speaker, and meetup organiser. Bookworm, hiker, history nerd - ask him about Napoleon!
Aji Ghose is Chief Scientist at Chattermill, where he leads AI research and NLP applied to large-scale customer feedback. He holds a PhD in Computational Cognitive Science focused on grounding deep learning models, and has over 15 years of experience building and deploying machine learning systems in industry. His work spans data science, large language models, semantic modelling, and human-centred AI.
Adrian Sroka
Relativity
MCP: Revolution or Security Regression
Abstract
Model Context Protocol (MCP) has rapidly become the backbone of modern LLM applications, enabling powerful multi-tool and multi-app workflows. But with this new capability comes a new class of security risks. I will explore where MCP genuinely pushes innovation forward and where it may quietly reintroduce old vulnerabilities under new names.
Bio
Adrian Sroka is an AI Security Lead and consultant, bridging software engineering with secure AI system design. He helps organizations adopt LLM technologies safely by combining deep technical expertise with practical architecture patterns. Adrian is the co-author of OWASP: A Practical Guide for Securely Using Third-Party MCP Servers, a contributor to OWASP Security Champions, an AI Security trainer, and a university lecturer speaking at conferences in Poland.
Piotr Migdal & Przemyslaw Hejman
Quesma
1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor
Abstract
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of CompileBench.
During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively.
To get the most out of this session, please bring your laptop. We recommend installing Harbor prior to the event. The only technical prerequisites are UV and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).
Bio
Piotr Migdal is a founding engineer and AI specialist with a strong background in data visualization, machine learning, and applied research. He is a Founding Engineer at Quesma, where he leads development of Quesma Charts, using AI to transform data from sources such as CSV, SQL, and APIs into accurate, high quality visualizations for tools including ggplot2 and Grafana. Previously, he worked as an independent AI consultant, delivering user facing AI and data science solutions across medtech, biotech, gaming, and emerging technologies. Piotr is also the co founder and former CTO of Quantum Flytrap, focused on intuitive interfaces for quantum computing, and has experience as an AI researcher in game development and physics based simulation. He combines deep technical expertise with a strong emphasis on usability and clear communication of complex data.
Piotr Migdal & Przemyslaw Hejman
Quesma
1h Workshop: Hands-on AI agent evaluation: building benchmarks with Harbor
Abstract
This is a hands-on introduction to Harbor, an open source framework designed for AI agent evaluation, creating benchmarks, and reinforcement learning. You will learn why Harbor is a game-changer for AI development and see real-world examples, including our migration of CompileBench.
During the workshop, we will build a small benchmark from scratch. You will leave with a working setup and the skills to benchmark AI models and agents effectively.
To get the most out of this session, please bring your laptop. We recommend installing Harbor prior to the event. The only technical prerequisites are UV and Docker. Please also make sure you have an API key for your preferred models (e.g. Anthropic, OpenAI, Gemini, or OpenRouter).
Bio
Piotr Migdal is a founding engineer and AI specialist with a strong background in data visualization, machine learning, and applied research. He is a Founding Engineer at Quesma, where he leads development of Quesma Charts, using AI to transform data from sources such as CSV, SQL, and APIs into accurate, high quality visualizations for tools including ggplot2 and Grafana. Previously, he worked as an independent AI consultant, delivering user facing AI and data science solutions across medtech, biotech, gaming, and emerging technologies. Piotr is also the co founder and former CTO of Quantum Flytrap, focused on intuitive interfaces for quantum computing, and has experience as an AI researcher in game development and physics based simulation. He combines deep technical expertise with a strong emphasis on usability and clear communication of complex data.
Randy Bias
Mirantis
Agents Need to be Paged, Not Prompted if we Truly Want AIOps
Abstract
AI agents have a blind spot: IT Operations. While the industry obsesses over software development, ops remains a second-class citizen. Software engineering is proactive and human-driven. Operations is also, but only sometimes. It's also reactive, predictive, and event-driven by infrastructure and applications. Nobody prompts the 3 AM outage. It just happens.
Current agentic architectures optimize for exactly one modality: human-driven interactions. To make AIOps real, we need event-driven agent triggering, domain-specific operational skills, and production-grade security models. In this talk I'll take you from vision to working example of alternative operations-centric agentic workflows. I'll showcase AI agents performing triage and root cause analysis triggered by faults in running clusters and applications.
What does the world look like if operations teams are suddenly empowered by agents working on their behalf? Agents that dramatically reduce toil by doing the grunt work and handing those teams highly accurate analyses backed by a proof-of-work allowing for immediate action.
Bio
Randy Bias is a cloud computing pioneer, entrepreneur, and strategist with decades of experience shaping open source and cloud infrastructure. He accurately predicted the growth trajectory of AWS, helped popularize foundational cloud concepts, and has advised startups and enterprises on technology, product, and go-to-market strategy. Randy currently serves as VP of Strategy and Technology at Mirantis and is a longtime advocate for open source as a sustainable business model.
Michal Bazyli
Cracken
Glitch in The Matrix: Real-World Pitfalls in Building Autonomous AI Agents for Security Testing
Abstract
Autonomous AI agents promise to accelerate and scale security testing, but real-world deployments expose sharp edges that demos rarely show. This talk examines common failure modes, including brittle reasoning, unsafe tool use, environment drift, and misleading test coverage. Drawing from hands-on experiments, we will explore where agents break, why they break, and how security teams can design guardrails that keep automation useful rather than dangerous. Attendees will leave with practical lessons for building, evaluating, and trusting AI-driven security testing systems.
Bio
Michał Bazyli is a founding cybersecurity researcher focused on applying AI and machine learning to real-world security problems. He has extensive experience in offensive security, smart contract auditing, and building AI-driven security systems across Web2 and Web3 environments. His current work explores the practical limits, failure modes, and safety challenges of autonomous AI agents used in security testing.
Maish Saidel-Keesing
AWS
Is Your GenAI System Ready for Production Reality?
Abstract
Production changes everything. Your GenAI application that worked perfectly in testing now faces real-world challenges: unexpected load patterns, model drift, and cascading failures. This session will help you understand what is needed to prepare to put your GenAI application in production safely
Bio
Maish Saidel-Keesing is a Senior Enterprise Developer Advocate @AWS working on containers and has been working in IT for the past 20 years and with a stronger focus on cloud and automation for the past 7.
He has extensive experience with AWS Cloud technologies, DevOps and Agile practices and implementations, containers, Kubernetes, virtualization, and modern applications.
He is constantly trying to bridge the gap between Developers and Operators to allow all of us provide a better service for our customers (and not wake up from pages in the middle of the night). He is an avid practitioner of dissolving silos - educating Ops how to code and explaining to Devs what the hell is Operations.
Automation is the way things should be done - and he is constantly looking for ways to make life easier wherever he can.
Andriy Batutin
ex - Accenture, Shelf, MacPaw
The Verification Gap: What Separates LLM Demos from Production Agents
Abstract
Every LLM demo looks magical. Then reality hits: hallucinations, edge cases, user trust erosion. This talk presents case studies on what it actually takes to bring AI Agents to production – and how organizations can build reliable agentic workflows without an OpenAI-scale budget.
We'll cover practical verification patterns, failure modes that only surface at scale, and the architectural decisions that separate impressive prototypes from systems users actually trust.
Bio
Andriy Batutin is a Senior AI Engineer at MacPaw, where he builds production AI agent systems with 200+ tool integrations serving millions of users. With over 10 years in IT and 6 years dedicated to AI/ML, he specializes in the hard problem of making agentic workflows reliable - bridging the gap between demos that impress and systems that actually work.
Porimol Chandro
Roche
From Retrieval to Reasoning: Architecting Agentic-RAG Workflows
Abstract
Retrieval-Augmented Generation (RAG) has become the standard approach in developing enterprise AI applications. However, as real-world tasks become more complex, classical RAG systems are performing adequately but are now encountering clear limitations. They retrieve information well, yet struggle with planning, multi-step reasoning, tool use, and integrating knowledge across diverse enterprise data sources. The result is a system that can answer simple questions but cannot execute workflows.
This talk introduces Agentic-RAG, an emerging paradigm that combines retrieval with autonomous agents capable of reasoning, iterating, and making decisions. We will break down the core architectural components—intent interpretation, planning, multi-hop retrieval orchestration, tool-augmented reasoning, and context synthesis—and show how they transform RAG from passive fetch-and-generate into an active, adaptive problem-solving pipeline.
Bio
Porimol is an MLOps Engineer at Roche, currently focused on architecting advanced Agentic-RAG workflows for healthcare. He has been working in the cross-disciplinary (software engineering, data engineering, machine learning, and MLOps) domain for more than a decade. He's passionate about turning complex AI ideas into systems that actually work in production.
Porimol specializes in the intersection of compliance and efficiency, applying a strong FinOps mindset to ensure that high-performance AI is also financially sustainable for the enterprise. With an MSc in Data Science and Business Analytics from the University of Warsaw and a research background in trustworthy AI, he is dedicated to building infrastructure that is scalable, interpretable, and production-ready.
Daniel Afonso
PagerDuty
The State of AI in Incident Response
Abstract
How many times were you woken up during the night to either spend more time than you would like trying to figure out what exactly broke, or get frustrated once you figured out it was actually a false positive? Well, with Agents, this won't happen, and your organization will get better by using them.
Bio
Daniel Afonso is a Senior Developer Advocate at PagerDuty, SolidJS DX team member, Instructor at Egghead.io, and Author of State Management with React Query. Daniel has a full-stack background, having worked with different languages and frameworks on various projects from IoT to Fraud Detection. He is passionate about learning and teaching and has spoken at multiple conferences around the world about topics he loves. In his free time, when he's not learning new technologies or writing about them, he's probably reading comics or watching superhero movies and shows.
Pawel Bulowski
Pawel Bulowski AI Consulting
SLMs at the Edge: Opportunities & Challenges
Abstract
This session will offer insights into practical examples of deploying SLMs on edge devices from mobile phones to cars. We will take a look at frameworks, models and real applications in compute-constrained environments.
Bio
Paweł Bulowski is the Director of Generative AI, he leads transformative AI initiatives across the automotive and tech sectors. With deep expertise in AI, data strategy, FinOps, and MLOps, he has successfully managed high-impact projects on platforms like AWS and Azure and Edge. His background spans industries including insurance and banking, equipping him with a strong understanding of secure, scalable design in regulated environments.
A certified FinOps practitioner, Paweł focuses on delivering efficient AI solutions while optimizing cloud spending and maximizing ROI. He is particularly passionate about bridging the gap between technology and business strategy, translating complex concepts into actionable insights, and driving enterprise modernization through AI.
Jakub Rohleder
Monday.com
KeynoteStop Making Your Agents More Expensive, Make Your Retrieval Better
Abstract
Everyone is racing to build stronger AI agents. But a harder question is how to make them economically viable at scale. The industry tends to focus on adding more capabilities and bigger models, while costs quietly explode across dimensions most teams do not think about early on.
In practice, many teams start with an LLM plus tools, keep layering on features, and end up with systems that do not scale economically. Poor architectural decisions and missing optimizations compound across database operations, embedding computation, and agent inference.
This talk explores the economics of architecture choices when building production AI agents. Drawing from experience building semantic search for billions of entities at Monday.com and integrating it into AI-driven products, it covers concrete architecture lessons that help control costs. Topics include using deterministic systems for predictable operations, building vector search for retrieval, and reserving LLM reasoning for cases that truly need it.
Attendees will see real production examples, cost breakdowns, and decision frameworks for choosing the right tool for each problem. You do not have to choose between capability and cost, but you do need to be intentional about where you invest in reasoning versus where you rely on retrieval. These trade-offs become critical when scaling from prototype to product.
Bio
Kuba Rohleder is a Staff Software Engineer at Monday.com, where he builds hybrid search systems for AI agents at massive scale, including billions of vectors and hundreds of millions of daily updates. With a background in large-scale systems and experience leading the GCP Kubernetes UI at Google, he now focuses on making AI architectures economically sustainable.
His approach emphasizes retrieval as the core, guardrails for predictable flows, and tiered AI and ML models as composable building blocks. Outside of work, he is a prolific builder, creating everything from nutrition tracking apps to smart home voice control systems, including training his own trigger word model from scratch.
Michael Matloka
PostHog
Keynote10 Learnings from Launching an Agentic AI Product at Scale
Abstract
There's always a gap between a great demo and a great product. Doubly true with AI agents. No wonder: the whole domain just appeared, keeps evolving at a crazy pace, and to make matters worse – the technology is fundamentally non-deterministic.
At PostHog, we've spent the last year and a half building an agent for product research. Then rebuilding it once, twice, thrice, finally launching, and… rebuilding again.
Don't go through this hell yourself. Instead, join me for this crash course on building an AI agent as a product. We'll get into easy mistakes (that we made), unobvious trends (that we're looking to ride), and business dilemmas (that the industry is facing).
Bio
Michael Matloka is a senior product engineer and team lead at PostHog. He leads the Signals team, building systems that turn analytics, session replay, and error tracking data into actionable product insights. An early PostHog engineer, he also founded and launched PostHog AI and has extensive experience shipping large scale data, analytics, and user facing product infrastructure.