Luis Carlos Lara-Lopez, Ignacio Trejos-Zelaya, Santiago Núñez-Corrales, José Helo-Guzmán
Agent-Based Models (ABMs) are widely used to simulate complex systems through emergent behavior. Agent Based Modeling practice is constrained by the lack of systematic methods to evaluate and select frameworks for a given problem, and by the absence of protocols to study differences across computational implementations. We propose a Feature Space Maturity Model (FSMM) to evaluate how well a framework matches the requirements of a modeling task. We present results of four canonical ABM experiments to demonstrate the FSMM applied to five well-known frameworks. Our results suggest that combining the FSMM with the proposed protocol provides valuable information for both framework builders and modelers facing complex technical and research choices. Ensuring ABM reliability requires specification fidelity -- meaning both the conceptual correctness of the model and its faithful implementation. We study implementation fidelity by analyzing statistical similarity across multiple frameworks. We apply a macroscopic statistical approach to evaluate whether different ABM frameworks yield functionally equivalent results when executing the same specification. We designed a Pareto-based experiment and implemented it in six ABM frameworks. We defined three macroscopic observables and analyzed 10,000 simulation runs per framework using coefficient of variation, Fréchet distance, ANOVA, and F-tests. Five of the six frameworks show statistically indistinguishable behavior, while one exhibits consistent deviations. The results highlight the value of statistical validation to identify implementation inconsistencies and optimize resource allocation in large-scale simulations.