Balancing Rare Linguistic Features in Small Datasets through LLM Augmented Training: A Case Study on Negation Detection
Keywords:
Implicit Negation, Data Augmentation, Large Language Models, Linguistic Generalization, Low-Resource PhenomenaAbstract
Rare linguistic features, such as implicit cues, structural negation, or low-frequency modifiers, pose persistent challenges for NLP models due to their sparse presence in training corpora. This study explores how large language models (LLMs) can be leveraged to surface and amplify such underrepresented features through targeted data augmentation. Implicit negation is selected as a representative case, and a two-stage augmentation pipeline is introduced: (1) structurally diverse training samples are generated around rare negation cues, and (2) counterfactual sentence pairs are constructed to suppress background biases and highlight model sensitivity to critical linguistic features such as implicit negation. Experiments using RoBERTa on the CONDAQA dataset
demonstrate that LLM-augmented training significantly improves performance on inputs containing implicit and structurally complex negation. These findings suggest that LLMs can serve as controlled augmentation tools for rebalancing rare linguistic phenomena in low-resource settings.
Published
How to Cite
Issue
Section
Copyright (c) 2025 Pakistan Journal of Language Studies

This work is licensed under a Creative Commons Attribution 4.0 International License.