LLMday

Large Language Models, Agents & AI Systems

October 14, 2026 The Sunset Room, Austin, Texas, USA

1
Day
10+
Speakers
1
Track
100+
Attendees

Evaluation Without Ground Truth: Lessons from Expert Disagreement

Leila Anderson
Upheal
Abstract

Most LLM evaluation assumes a gold answer exists to score against. A large and growing share of production work, however, doesn't have one: contract review, clinical documentation, incident summaries, moderation calls, research synthesis — any task where correctness is a matter of expert judgment, and where two qualified experts given the same input produce different outputs that are both defensible. If you score that as if a single right answer existed, you collapse the errors that matter — a fabricated fact, a dropped material detail — into the same number as the ones that don't, like structure, emphasis, and phrasing. This talk is about building evaluation for those systems, where the ground truth isn't a fixed key but a distribution of expert opinion, and where disagreement is built into the problem rather than a symptom of bad labeling. It unpacks three moves that improve accuracy and trustworthiness: treating inter-rater reliability among your experts as the ceiling on any eval you can build, using expert corrections as a live quality signal instead of a static gold set, and separating "genuinely wrong" from "differently right" so the score tracks the failures you actually care about.

Bio

Leila Anderson is a Senior Clinical AI Product Manager at Upheal, where she builds and evaluates the LLM systems behind clinical documentation for behavioral health. She's also a licensed marriage and family therapist (LMFT-S) with nearly a decade of clinical experience, which puts her on both sides of the eval problem this talk is about: the expert whose judgment the model is trying to match, and the person on the hook for deciding whether it did. Alongside her work at Upheal she runs a private psychotherapy practice and serves as an expert witness in clinical cases — two more roles where being the arbiter of a contested-but-defensible judgment is the whole job.

Sponsors & Partners

Want to become a sponsor? Get in touch!
Let's talk!
We'll email you and share prospectuses for relevant events.
We'd like to (one or more)
Pick at least one
Conferences (one or more)
Pick at least one
Regions (one or more)
Pick at least one
Budget
Pick one