低资源语言文本净化系统Tatoxa:以鞑靼语为例
The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
摘要
文本净化技术——即自动检测并消除有害内容——对于确保在线社区的安全以及保护用户权益至关重要。然而,像塔塔尔语这样的资源匮乏的语言却很少得到研究关注。本文介绍了Tatoxa这一用于塔塔尔语文本净化的先进系统。比较实验表明,该方法在各项质量指标上均优于现有的开源和商业LLM模型。我们还创建了一个新的塔塔尔语文本净化数据集,该数据集旨在用于在资源匮乏的情况下进行优化和评估。最后,跨语言迁移实验表明,即使有大量的俄语数据可用,从其他语言(包括文化上相近的俄语)迁移到塔塔尔语的数据,其效果也远不如在本地塔塔尔语数据上进行训练。
English Abstract
Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention. In this paper we present Tatoxa, a novel state-of-the-art system for text detoxification in the Tatar language. Comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics. We also introduce a new dataset for text detoxification in Tatar, designed for fine tuning and evaluation in low resource settings. Finally, cross lingual transfer experiments indicate that transfer from other languages, including the culturally close Russian, performs significantly worse than training on native Tatar data even when a large Russian corpus is available.