A. Lopez Osornio, A. Randorff Hoejen, K. Kewley
Objectives: Searching a reference terminology such as SNOMED CT confronts two retrieval technologies of fundamentally opposite nature. Deterministic lexical matching is precise, order-independent and character-level, and lets a clinician refine a query incrementally, but it is confined to the wording and language of the curated descriptions. Learned semantic matching bridges paraphrase and language, but poorly represents abbreviations and partial, as-typed input. Their requirements and failure modes are so different that combining them naively can let the channel that is wrong for a given query degrade the one that is right. The conventional way to avoid adding a semantic channel at all has been to strengthen the lexical one (curating an additional local interface vocabulary so that more expressions match by surface form), but such vocabularies are labour-intensive to maintain and their quality degrades over time. We ask whether these two opposite paradigms can instead be made to *coexist* in one retrieval flow (each serving the input regime it is strong at, neither harming the other) over SNOMED CT's own curated descriptions, and study the query interpretation, rank fusion and re-ranking design that makes this coexistence safe. Materials and Methods: We implemented one reproducible instance of this architecture over the SNOMED CT International edition. It combines deterministic, order-independent multi-prefix search with three learned components: BioLORD-2023-M embeddings for semantic retrieval, an optional BGE cross-encoder for re-ranking, and a local LLM (gemma 4) for query normalization and structured clinical entity extraction. Lexical and semantic rankings are combined with Reciprocal Rank Fusion, and results can be constrained with a hierarchy (descendant) filter. As an external evaluation we measured search-only linking accuracy on 542 disease mentions from the DisTEMIST corpus (Spanish, zero-shot), resolving gold codes through SNOMED CT historical associations and ablating query pre-processing, re-ranking, and each retrieval channel in isolation; and, on a second corpus of English real-EHR discharge notes (the SNOMED CT Entity Linking Challenge, 12,897 mentions), we estimated field-scoped typeahead performance and tested a context-aware pre-process. Results: Illustrative tests showed complementary behaviour: the lexical and semantic paths resolved different query types, while LLM pre-processing supported cross-lingual and abbreviated input. On a search-only linking evaluation over 542 disease mentions from the DisTEMIST corpus (Spanish, zero-shot), the best and computationally cheapest configuration, semantic search with re-ranking and *no* LLM query pre-processing, placed the exact concept first for 60% of mentions (accuracy@1 = 0.60) and within the top-10 candidates a picker would show for 80% (recall@10 = 0.80). Because these search-only figures are computed on gold mention spans, they isolate normalization from mention detection and are not directly comparable to the DisTEMIST shared task's end-to-end scores. LLM query pre-processing helped dirty, lay input but slightly hurt clean terminology-like queries, and a near-miss analysis showed that many remaining cases retrieved a parent or child of the intended concept. Isolating the channels showed the two coexist without a trade-off: on this clean-normalization corpus the semantic channel carried the accuracy and adding the lexical channel neither helped nor harmed it (accuracy@1 0.60 vs 0.61), because rank fusion followed by re-ranking suppresses whichever channel is off-task; the lexical channel's own strength, partial as-typed input, is a regime this corpus does not test. On the English real-EHR corpus, a field-scoped typeahead (each query constrained to its entry field's domain) placed the concept on the top-10 picker list for 73% of 12,897 mentions across all clinical domains (accuracy@1 0.52; findings 0.54, close to DisTEMIST). Most residual errors were soft (the gold on the list but not first, or a hierarchical neighbour); the hard misses were dominated by ambiguous EHR shorthand, and a single context-aware pre-processing call (normalizing a mention with its surrounding note text) roughly doubled recovery of those hard misses, though, as with LLM normalization on the clean terms of DisTEMIST, applying it to every mention slightly lowered overall accuracy by perturbing the easy majority, so it is best used as a selective rescue. Results are specific to the evaluated configuration, not an architecture-independent estimate. Discussion: The findings support the technical feasibility of a reference architecture that computes additional access routes over SNOMED CT's curated descriptions, and indicate that making the lexical and semantic paradigms *coexist*, rather than choosing between them, is what makes those computed routes safe to add. This shifts rather than eliminates maintenance, from local phrase authoring toward governance and validation of retrieval services. Conclusion: Hybrid concept retrieval lets lexical precision and semantic flexibility coexist over SNOMED CT's maintained descriptions without a trade-off between them, offering an implementation pattern that may reduce additional local lexical curation where direct search is appropriate, while curated interface vocabularies remain useful in selected guided workflows.