Jaba Tkemaladze
The peer review system—the cornerstone of scientific validation—is experiencing a crisis of inconsistency (kappa < 0.4), pervasive bias, and prolonged cycles (3–6 months). Existing AI tools address only narrow tasks (plagiarism detection, journal matching), leaving holistic quality assessment unresolved. We present TBPR v2, an end-to-end AI-driven triple-blind peer review system that emulates a multi-expert panel using three independent large language models (Gemini, DeepSeek, Claude). Each reviewer evaluates projects against nine BHCA criteria (Innovation, Team/PI, Budget Justification, Data Plan, Feasibility, Impact, Ethics, Reproducibility, Clarity)—55 points total. The system integrates an auto-fix pipeline: concern extraction, RAG-based retrieval (ChromaDB, cosine similarity >0.65), category-specific fixers, and Arena strategy selection with ε-greedy (ε=0.1). Adaptive early stopping halts after five consecutive no-improvement cycles or rolls back after three consecutive regressions. A pre-cycle quality classifier dynamically sets cycle budgets (EASY=7, MEDIUM=5, HARD=3), reducing unnecessary API calls by ~60%. Across 26 projects (264 cycles), results show: winsorized mean score 22.8/55 (SD=9.7), median 21.0 [IQR:14.0–30.0], and only one project reached the ACCEPT threshold (44/55). Fix efficiency ranged from 52% to 75%. These results demonstrate the first viable end-to-end AI peer review system producing measurable, reproducible, and iteratively improvable quality scores.