GRPO changed how we think about reinforcement learning for language models—but anyone who has used it in practice knows it comes with frustrating limitations. Unstable training runs, reward signals that collapse or fight each other, and poor credit assignment when your model needs to reason across multiple steps. PRISM takes the core ideas that made GRPO successful and fixes what was broken. It's a unified framework that brings stability to multi-objective training, letting you combine different reward signals—correctness, style, safety—without one overwhelming the others. I will explain a >1000 experiments how we build better GRPO for agents and long turn credit assignments.
Grzegorz Warzecha is a founder and entrepreneur with deep experience building software platforms and technology driven organizations. He is the founder of User, a marketing automation platform created to replace fragmented, expensive tools with a single integrated solution for customer communication, data collection, and automation. He is also the founder of Genotic and previously founded CivilHub Foundation, a platform supporting bottom up social initiatives. Earlier in his career, he served as CEO of Expose Sp. z o.o., leading an IT focused training company delivering advanced technical education and consulting. Grzegorz has built and scaled organizations across the US and Europe, with a focus on practical products that solve real problems.