Research framing
The README frames the project as a diachronic corpus linguistics study of Soviet census methodology texts, with VLM extraction as the infrastructure that built the corpus.
- Primary series: 1926 -> 1959 -> 1970 -> 1979 -> 1989.
- Inter-census vignettes: 1937 suppressed draft and 1939 engineered replacement, analyzed separately.
- Target vocabulary includes nationality, nationhood, language, tribe, native language, and related ethnicity terms.