Ilseyar Alimova, Bogdan Monogov, Artyom Mazur, Daniil Antonov, Vsevolod Karimov, Vitaliy Egorov, Bulat Khakimov, Alexander Panchenko
This research presents a state-of-the-art system and dataset for text detoxification in the low-resource Tatar language, demonstrating the limitations of cross-lingual transfer.
Text detoxification is crucial for online safety, but low-resource languages like Tatar lack sufficient research and tools.
The authors developed a new system called Tatoxa for Tatar text detoxification and conducted cross-lingual transfer experiments, including from culturally close Russian, for comparative analysis.
The proposed system outperformed existing open-source and proprietary LLMs on key quality metrics. The study also introduced a new dataset for Tatar and proved that training on native Tatar data is significantly more effective than cross-lingual transfer.