科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ Open MIND2026-08-01· Mathematics

External validation of a weighted Gray-code codon metric across saturation-genome-editing assays: a negative grouped benchmark

Stassis Stashkevichyus

原始摘要(英文原文)· Original abstract
Version 1.0.0 (1 August 2026). Preprint; not peer reviewed. Abstract: A companion mathematical study defines a fixed weighted Gray-code metric on RNA codons and shows that it is an exact weighted-Hamming metric. Its internal synonymy benchmark does not establish external biological prediction. We therefore evaluated the frozen metric on seven published saturation-genome-editing score sets comprising 23,628 missense single-nucleotide variants at 3,004 protein sites in 6 genes. Variant HGVS strings were reconstructed against frozen coding transcripts; every syntactically eligible coding SNV matched its reference allele. Assay scores were oriented and rescaled by synonymous and nonsense control medians. The primary split held out entire genes, with both BRCA2 assays held out together. A secondary split held out assays, and a diagnostic split held out contiguous protein-position blocks. The fixed codon score took only 9 values on the SNV cohort and had leave-one-gene-out macro Spearman correlation ρS = 0.006178. Transparent amino-acid baselines performed better: Grantham, -BLOSUM62, and -PAM250 gave ρS = 0.136420, 0.192453, and 0.171266, respectively. A grouped protein-sequence ridge model gave ρS = 0.245907. Adding the fixed codon score changed the macro correlation by only ΔρS = 5.056 × 10^-7, with a gene-cluster bootstrap interval [-8.229 × 10^-5, 6.507 × 10^-5], an exact two-sided sign-flip p = 1.000, and positive changes in 3 of 6 genes. The assay-held-out change was ΔρS = -1.686 × 10^-5. High-confidence classification was likewise unchanged: protein-only versus protein-plus-codon AUROC was 0.736076 versus 0.736009, and Brier score was 0.112684 versus 0.112698. The codon feature failed four of five predeclared success gates; the sole pass was that RMSE did not materially worsen. Thus this frozen panel provides no evidence that the fixed metric contributes practically relevant incremental information for missense functional effects. This negative result does not invalidate the metric as a finite mathematical object or address other codon phenotypes, but it rejects its use as a stand-alone or additive missense-effect predictor on the tested panel. Author: Stassis Stashkevichyus; Independent Researcher, Lithuania; ORCID 0009-0000-2294-705X; theobserver.of.multiverses@proton.me. Canonical GitHub release: https://github.com/Observer1117/weighted-gray-codon-external-validation/releases/tag/v1.0.0 Frozen commit: 37ff2c570d03348166e1dcc734e1df6c6ae1e5c5 Companion Article I: https://github.com/Observer1117/weighted-gray-codon-geometry/releases/tag/v1.0.0 Licensing: manuscript and author-created figures/tables are CC BY 4.0; original code is MIT; third-party data retain source-specific terms recorded in THIRD_PARTY_DATA.md and data/ATTRIBUTION.md.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

External validation of a weighted Gray-code codon metric across saturation-genome-editing assays: a negative grouped benchmark — 科研速览 Science Skim