Most LLM evaluation assumes a gold answer exists to score against. A large and growing share of production work, however, doesn't have one: contract review, clinical documentation, incident summaries, moderation calls, research synthesis — any task where correctness is a matter of expert judgment, and where two qualified experts given the same input produce different outputs that are both defensible. If you score that as if a single right answer existed, you collapse the errors that matter — a fabricated fact, a dropped material detail — into the same number as the ones that don't, like structure, emphasis, and phrasing. This talk is about building evaluation for those systems, where the ground truth isn't a fixed key but a distribution of expert opinion, and where disagreement is built into the problem rather than a symptom of bad labeling. It unpacks three moves that improve accuracy and trustworthiness: treating inter-rater reliability among your experts as the ceiling on any eval you can build, using expert corrections as a live quality signal instead of a static gold set, and separating "genuinely wrong" from "differently right" so the score tracks the failures you actually care about.
Leila Anderson is a Senior Clinical AI Product Manager at Upheal, where she builds and evaluates the LLM systems behind clinical documentation for behavioral health. She's also a licensed marriage and family therapist (LMFT-S) with nearly a decade of clinical experience, which puts her on both sides of the eval problem this talk is about: the expert whose judgment the model is trying to match, and the person on the hook for deciding whether it did. Alongside her work at Upheal she runs a private psychotherapy practice and serves as an expert witness in clinical cases — two more roles where being the arbiter of a contested-but-defensible judgment is the whole job.