LLM systems are getting very good at extracting entities from messy documents. But extracting a mention is only the beginning. The harder problem is deciding what that mention actually refers to. A single real-world entity can appear under different names, formats, abbreviations, and contexts across documents and databases. At the same time, different entities can look remarkably similar. A reliable system therefore needs more than entity extraction or similarity search. It needs a mapping step that decides which candidate, if any, represents the entity being described. This is where similarity becomes both useful and insufficient. Database search, lexical matching, and embeddings are powerful ways to generate and rank candidates. But a high similarity score does not make a candidate the correct mapping. The system still has to account for ambiguity, context, available evidence, and what happens when no candidate meets the evidence threshold. Sometimes the right decision is not to map at all. This talk looks at entity mapping as a distinct engineering problem in an LLM pipeline: using structured search and semantic similarity to find plausible candidates, understanding where those signals break down, and designing the decision layer that determines when a candidate is good enough to map, and when the system should abstain. The goal is not to replace similarity with something else. It is to understand where similarity fits, and where the identity decision begins.
Maryam Astero is a Senior AI Research Engineer at Genomenon, where she builds production systems that turn unstructured scientific literature into structured, queryable knowledge. Her work spans document ranking, information extraction, entity resolution, and knowledge graph construction, with a focus on making these systems reliable beyond the prototype stage. She holds a PhD in Computer Science from Aalto University, where she developed graph neural network architectures for chemical reaction modeling. She works at the meeting point of machine learning, graph-based methods, and production systems, particularly where ambiguity and imperfect data make seemingly simple problems hard to evaluate and scale.