Offensive-Language Detection on Multi-Semantic Fusion Based on Data Augmentation

Junjie Liu

Yong Yang

Xiaochao Fan

Ge Ren

Liang Yang and Qian Ning

Resumen

The rapid identification of offensive language in social media is of great significance for preventing viral spread and reducing the spread of malicious information, such as cyberbullying and content related to self-harm. In existing research, the public datasets of offensive language are small; the label quality is uneven; and the performance of the pre-trained models is not satisfactory. To overcome these problems, we proposed a multi-semantic fusion model based on data augmentation (MSF). Data augmentation was carried out by back translation so that it reduced the impact of too-small datasets on performance. At the same time, we used a novel fusion mechanism that combines word-level semantic features and n-grams character features. The experimental results on the two datasets showed that the model proposed in this study can effectively extract the semantic information of offensive language and achieve state-of-the-art performance on both datasets.

Palabras claves

offensive language - data augmentation - MSF

Acceso

PÁGINAS

pp. 0 - 0

NÚMERO

Volumen: 5 Parte: 1 (2022)

MATERIAS

INGENIERÍA Y CONSTRUCCIÓN CIVIL
TECNOLOGÍA

REVISTAS SIMILARES

Information
Informatics

DOI