Balancing Rare Linguistic Features in Small Datasets through LLM Augmented Training: A Case Study on Negation Detection

Authors

  • Yangtian Li Department of Applied Linguistics, GCUFViterbi School of Engineering, University of Southern California.

Keywords:

Implicit Negation, Data Augmentation, Large Language Models, Linguistic Generalization, Low-Resource Phenomena

Abstract

Rare linguistic features, such as implicit cues, structural negation, or low-frequency modifiers, pose persistent challenges for NLP models due to their sparse presence in training corpora. This study explores how large language models (LLMs) can be leveraged to surface and amplify such underrepresented features through targeted data augmentation. Implicit negation is selected as a representative case, and a two-stage augmentation pipeline is introduced: (1) structurally diverse training samples are generated around rare negation cues, and (2) counterfactual sentence pairs are constructed to suppress background biases and highlight model sensitivity to critical linguistic features such as implicit negation. Experiments using RoBERTa on the CONDAQA dataset
demonstrate that LLM-augmented training significantly improves performance on inputs containing implicit and structurally complex negation. These findings suggest that LLMs can serve as controlled augmentation tools for rebalancing rare linguistic phenomena in low-resource settings.


Author Biography

Yangtian Li, Department of Applied Linguistics, GCUFViterbi School of Engineering, University of Southern California.



Published

2025-12-10

How to Cite

Li, Y. (2025). Balancing Rare Linguistic Features in Small Datasets through LLM Augmented Training: A Case Study on Negation Detection . Pakistan Journal of Language Studies, 8(1), 74-89. Retrieved from //pjls.gcuf.edu.pk/index.php/pjls/article/view/315