Carlos Blanco Velázquez, Diego Gonzalez Oviaño, Isabel Santos Ramos, Laura Llorente Sanz, Julio Mayol
Arthroplasty clinic no-shows reflects engagement patterns, administrative processes, and patient context. Calibrated, interpretable machine-learning models capturing ∼60% of no-shows within the highest-risk 20% of encounters support prospective evaluation for targeted, supportive outreach with equity safeguards.
Surgical robots have so far amplified the surgeon's hands without replacing the surgeon's judgement. Two convergent developments are now changing that calculus. First, the maturation of computer vision applied to intraoperative video. Second, the emergence of world models, that is learned, action-conditioned simulators of how a scene will evolve. This narrative review explains these technologies in surgical terms and appraises how far they carry the field towards autonomy. PubMed, arXiv and Semantic Scholar were searched to 1 July 2026; much of the world-model literature is preprint and not peer reviewed. Surgical video has become a tractable data substrate: phase, step, instrument and anatomy recognition are accurate in benchmark settings, and vision-language pretraining has reduced dependence on manual annotation. World models add a distinct capability, predicting the consequences of a candidate action before it is executed, and have entered surgery through action-controllable video generators and simulator-coupled policies. Autonomy has advanced correspondingly, from supervised task automation in bowel anastomosis to language-conditioned, self-correcting execution of the clipping and cutting phase of ex vivo cholecystectomy. Expert evaluation nonetheless indicates that state-of-the-art generative models produce photorealistic but causally and strategically implausible surgical futures, and no autonomous procedure has been performed in a human. World models are best understood not as a route to surgeon-free operating but as an engine for data generation, policy training and intraoperative anticipation. Their translation will depend on surgical rather than computational criteria: prospective evaluation, defined stopping rules, and an explicit account of who is accountable when the model is wrong.