科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Historical Life Course Studies2026-02-12· Population

POPP. An OCR-Generated Database of the Population Censuses of Paris (1926–1936)

Sandra Brée, Victor Gay, Marion Leturcq, Baptiste Coulmont, Yoann Doignon, Thomas Constum, Thierry Paquet, Pierrick Tranouez

原始摘要(英文原文)· Original abstract
Empirical research in historical demography is usually time-consuming and labour-intensive. Recent developments in machine learning offer new possibilities for building very large databases with reduced time and costs, though these new methods raise new challenges as well. This article describes the process of constructing the POPP database, a data collection project based on the exploitation of the nominative lists of the Parisian population censuses of 1926, 1931, and 1936. This database provides a host of information for almost 9 million individuals: their name and surname, year and location of birth, nationality, relation to the household head, and occupation. The article discusses the digitisation of archival sources — several hundred thousand handwritten pages — their transformation into a database by computer scientists using machine learning techniques, and the work required on the part of social scientists to correct and adapt the resulting data for statistical purposes. Beyond its methodological contribution, this article also discusses the various ways in which the POPP database will improve our knowledge of the economic, social, and demographic evolution of an important European urban population.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

POPP. An OCR-Generated Database of the Population Censuses of Paris (1926–1936) — 科研速览 Science Skim