Kirill Solovev, Jana Lasser
A modular, fully open-weight pipeline is proposed to automatically extract political figures and relations from multilingual news text, constructing temporal knowledge graphs.
Observing informal and adversarial ties among political elites at scale has relied on manual coding, while existing automated methods are limited to simple co-occurrence. Recent LLM-based approaches often depend on proprietary APIs, lack cross-lingual capability, and struggle with scalable entity resolution.
The pipeline combines span-based named-entity recognition (NER) with a three-stage linking cascade to language-independent Wikidata identifiers. A high-throughput mixture-of-experts model uses guided decoding to extract directed, signed relationships grounded in a domain ontology.
High textual correctness (68.2% strict to 93.7% lenient) was achieved against a 3,491-relation gold standard. Case studies in Austria and Poland successfully reconstructed political party lifecycles and state-enterprise patronage networks, providing a robust, replicable framework for cross-national empirical computational social science.