Documented but Invisible: From Language Documentation to Digital Multilingualism

Hao Lin

Abstract


Researchers have recorded Southeast Asia’s endangered languages for decades, producing dictionaries, grammars, and text collections. Yet these languages are almost absent from the corpora, benchmarks, and language models on which digital multilingualism increasingly depends. This paper asks what blocks the path from documentation to digital multilingualism, taking Thailand and its Mahidol documentation tradition as a critical case. The study is a PRISMA-lite analysis: eight searches run in July 2026 across the ACL Anthology, TCI-THAIJO, and ScholarSpace returned 110 records, of which 105 were included. The coded records confirm the gap: none pairs the documentation tradition with a machine-readable corpus of a Thai minority language; outputs are mostly dictionaries, schoolbooks, and archives made for human readers. Two findings follow. First, the main outputs of language documentation are designed for human readers rather than for computer programs. The study terms this mismatch a typological format barrier. Second, the failure sits in mobilization, the stage at which archived records should be turned back into usable resources. The paper contributes (1) a measured account of this barrier, and (2) a three-step pipeline: extracting structured text, adding annotation in tiers, and releasing corpora in phases with community consent. This conversion, the study argues, is a precondition for digital multilingualism among the region’s minority languages.


Full Text:

Untitled

Refbacks

  • There are currently no refbacks.