Hayrapetyan, Aram, Vladimir Makarenko, A. Tumasyan, Wolfgang Adam, Janik Walter Andrejkovic, L. Benato, Thomas Bergauer, M. Dragicevic, Cristina Giordano, Priya Sajid Hussain, M. Jeitler, Natascha Krammer, A. Li, Dietrich Liko, Mark Matthewman, Ivan Mikulec, J. Schieck, Schöfbeck, Robert, D. Schwarz, Maryam Shooshtari, M. Sonawane, W. Waltenberger, Claudia-Elisabeth Wulz, X. Janssen, Hyejin Kwon, Ocampo Henao, Daniel, T. Van Laer, Pierre van Mechelen, Jas Bierkens, Nordin Breugelmans, Jorgen d'Hondt, Soumya Dansana, A. De Moor, M. Delcourt, Felix Heyen, Yanwen Hong, Pavlo Kashko, Steven Lowette, I. Makarenko, Denise Müller, Juhee Song, Stefaan Tavernier, M. Tytgat, G. P. Van Onsem, S. Van Putte, D. Vannerom, Bugra Bilin, Barbara Clerbaux, Aloke Kumar Das, Isabelle de Bruyn, G. De Lentdecker, Hugues Evard, Laurent Favart, P. Gianneios, Ali Khalilzadeh, Khan, Fakhri Alam, A. Malara, Muhammad Aamir Shahzad, Laurent Thomas, Vanden Bemden, Max, Vander Velde, Catherine, P. Vanlaer, Fengwangdong Zhang, M. De Coen, Didar Dobur, Gokbulut, Gul, J. Knolle, David Marckx, K. Skovpen, N. Van Den Bossche, van der Linden, Jan, Jules Vandenbroeck, Liam Wezenbeek, Samuel Bein, Anna Benecke, Agni Bethani, Giacomo Bruno, A. Cappati, De Favereau De Jeneret, Jerome, C. Delaere, A. Giammanco, Ahmet Oguz Guzel, Vincent Lemaître, J. Lidrych, Paul Malek, Paola Mastrapasqua, Semra Turkcapar, Alves, Gilvan, Barroso Ferreira Filho, Mapse, E. Coelho, Carsten Hensel, Menezes De Oliveira, Thales, Mora Herrera, Clemencia, Rebello Teles, Patricia, Mariana Soeiro, Tonelli Manganote, Edmilson José, Vilela Pereira, Antonio, Aldá Júnior, Walter Luiz, Brandao Malbouisson, Helena, Wagner Carvalho
Abstract A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.