Yiming An, Yanshu Yu, Weidong Li, Orest Kochan
Binary code similarity detection (BCSD) ranks candidates but does not quantify the reliability of an already selected Top-1 match. We study this post-retrieval problem in a known-source, closed-set protocol: the target Top-1 is frozen before same-source cross-compilation views are queried, so auxiliary evidence audits cannot replace it. A frozen 34-variable map feeds a low-capacity logistic model with project-grouped cross-fitting, Platt calibration, and training-side threshold selection. On 413 families from 16 projects, cross-view evidence improved discrimination over target score/margin features. GCC-O0 was a dominant-anchor regime: Full showed no statistically resolved ROC-AUC gain over Primary-anchor, whereas Clang-O0 benefited from complementary non-primary evidence. On 240 project-identity-disjoint families from 55 projects, the design-locked structural branch accepted 75/240 GCC and 99/240 Clang candidates (31.3%/41.3% coverage) with no observed family-level errors. Correspondence mismatch reduced discrimination toward chance. Corrected TF-IDF remained supportive because correction followed label access. The contribution of this paper is a versioned candidate-preserving audit interface with explicit evidence and deployment boundaries, but not a universal retrieval improvement or distribution-free guarantee.