arXiv Artificial Intelligence

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Quick summary

arXiv:2608.22894v1 Announce Type: cross Abstract: Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The generated outputs were assessed through human evaluation and automatic analyses of lexical change, semantic preservation, sentiment, and dialectal style. Results sh

Key takeaways

  • arXiv:2608.22894v1 Announce Type: cross Abstract: Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored.
  • We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic.
  • The generated outputs were assessed through human evaluation and automatic analyses of lexical change, semantic preservation, sentiment, and dialectal style.

Why it matters

“AraDetox: A Multi-Dialect Arabic Detoxification Dataset” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗