Giovanni Vindigni
Although individual studies have examined the capabilities of generative AI in annotation, coding, item construction, thematic analysis, and data simulation, there has so far been no cross-paradigm framework that systematically integrates technical performance, methodological validity, institutional embedding, and the attribution of scientific responsibility. Against this backdrop, the study examines under what conditions the delegation of knowledge-relevant research operations to large language models (LLMs) should be evaluated as a controlled methodological extension or as an epistemic dissolution of boundaries. Methodologically, it follows a critical integrative evidence synthesis. From 80 hits across the eight thematically differentiated search clusters and five additional articles identified through citation tracking, 58 publications remained in the core analytical corpus after removing duplicates, screening titles and abstracts, and reviewing full texts based on predefined criteria. The predefined criteria restricted the corpus to studies that empirically evaluated LLM-mediated research operations or examined their validity, bias, reproducibility, epistemic efficacy, or governance with direct or transferable relevance to social science and educational research. With regard to the synthesized findings, the evidence indicates that LLMs operationally scale quantitative methods but may in the process foster methodological pseudo-competence, non-classical measurement errors, and synthetic circularity. In qualitative designs, they expand the scope for contrastive interpretation but do not replace situated case understanding; in mixed-methods designs, they reduce translation costs without negating the independent logics of validity of the combined strands. Building on, but analytically distinct from, these synthesized findings, the study advances an exploratory conceptual framework in the form of an original five-level matrix of epistemic delegation, ranging from operational and infrastructural offloading to empirical surrogation. Each level is linked to specific epistemic risks and proportionate verification requirements. In addition, reflexive controllability is operationalized as a governance framework with eight dimensions. Both instruments can be directly applied to research design, quality assurance, institutional rule-making, and methodological training and are transferable across different empirical paradigms. Therefore, generative AI appears neither as a neutral tool nor as a responsible agent of knowledge but rather as a probabilistic epistemic actor within distributed knowledge arrangements, for which the ultimate scientific responsibility remains with identifiable individuals and institutions.