Abstract
Metadata is the connective tissue of scholarly communication, it determines whether a published article can be found, indexed, cited, and correctly attributed within the global research infrastructure. Yet the manual creation and verification of bibliographic metadata remains a slow, labour-intensive, and error-prone process, and this burden falls disproportionately on journals that publish outside the small set of languages and layout conventions for which existing extraction tools were designed. This paper examines the scholarly and economic case for automated, multilingual metadata extraction, using MetaExtract - a rule-based and named-entity-recognition hybrid system built for Uzbek (Latin and Cyrillic), Russian, and English academic journal articles - as a case study. Evaluated on a stratified 150-article gold-standard corpus, the system achieved an overall F1 score of approximately 0.965 across six core metadata fields, while processing a single document in an average of 2.62 seconds on modest, GPU-free hardware. We situate these results within the broader literature on metadata quality, discoverability, and the resource economics of extraction approaches, and argue that lightweight, format-aware hybrid architectures offer a more sustainable path to metadata automation for linguistically underserved journals than either purely manual workflows or resource-intensive large-model approaches.
This work is licensed under a Creative Commons Attribution 4.0 International License.

