Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
This work addresses the "retrieval readiness gap"—a performance bottleneck in large-scale classification systems when inputs consist solely of indirect evidence (e.g., table cells) and suffer from semantic ambiguity. To overcome this, the authors propose a factorized hypothesis search mechanism that decomposes semantic interpretation into composable, named dimensional hypotheses. By leveraging structured query generation, parallel multi-hypothesis retrieval, and dimension-level candidate validation, the approach circumvents reliance on free-form text generation. Evaluated on financial taxonomy labeling and the CodiEsp clinical coding task, the method significantly outperforms existing non-oracle approaches, achieving consistent improvements in Recall@1, Mean Reciprocal Rank (MRR), and final accuracy, thereby demonstrating its effectiveness and robustness.