Erin Haus, Anthony Santella, Yichi Xu, Ruohan Ren, Dali Wang, Zhirong Bao
Complex tissues are now characterizable at single-cell resolution, but the spatial logic underlying tissue organization remains challenging to access without effective single-cell spatial descriptors that allow spatiotemporal events to be organized and interactions between cells to be recognized and predicted. We present a learned single-cell embedding space of cell positions over time enabling measurement of morphology and dynamics. Learning co-processes pairs of cell point clouds sampled over time using a Transformer encoder with inter-sample attention, a strategy that promotes efficient joint spatiotemporal learning. The embeddings show desirable properties of a general descriptor: interpretable cell type clusters, preserved local distances, and a manifold-like pseudo-time axis. Embeddings enable common but challenging spatial reasoning tasks such as annotation of anatomical landmarks at cellular resolution and detection of subtle, transient phenotypes in large screens. Our study demonstrates a widely applicable cell-based learning strategy that offers an expressive, general-purpose representation of tissue dynamics.