Evaluating Parameter-Efficient Fine-Tuning for Cross-Lingual Natural Language Inference: A Case Study on Nepali XNLI
DOI:
https://doi.org/10.65091/icicset.v3i1.72Abstract
Nepali language remains under-served in multi-lingual Natural Language Processing (NLP), and robust evaluation practices are often missing for low-resource settings. We present a reproducible study of Cross-lingual Natural Language Inference (XNLI) for Nepali using XLM-RoBERTa base, comparing full fine-tuning against parameter-efficient Low-Rank Adaptation (LoRA). Models are fine-tuned on English Natural Language Inference (NLI) data and evaluated on an English-to-Nepali translated subset of XNLI. Our pipeline performs three independent runs (seeds 42/43/44), logs predictions and confusion matrices, aggregates mean ± std across seeds and reports 95% bootstrap confidence intervals computed directly from per-example predictions. Full fine-tuning achieves 0.7687 accuracy and 0.7689 macro-F1, outperforming a fixed-configuration LoRA (r=8, α=16, no hyperparameter sweep) by approximately 2.1 points. We discuss why LoRA under-performs in this setting, the likely role of translation artifacts, and practical directions for closing the gap, while emphasizing reproducibility and error analysis over strong claims about absolute accuracy.