Bias Detection And Fairness Optimization In Nlp-Based Language Assessment Systems

Authors

  • Kazi Siam Al Mobin Department of Information Technology, Washington University of Science and Technology, United States Author
  • Afroza Riju Department of Computer Science and Engineering, Green University of Bangladesh, Bangladesh Author

Keywords:

Automated Essay Scoring, Educational Assessment, Natural Language Processing, Fairness in AI

Abstract

Background and Purpose: Natural Language Processing (NLP) has expanded automated language assessment through systems such as Automated Essay Scoring (AES), which provide scalable and consistent evaluation of student writing. However, automated scoring can reproduce disparities associated with linguistic, cultural, and socio-economic differences when models learn biased patterns from training data. This study investigates bias detection and fairness optimization in NLP-based language assessment systems.

Methods: A quantitative experimental framework was applied using the Automated Student Assessment Prize (ASAP) essay dataset. Support Vector Machine, Random Forest, Long Short-Term Memory, and BERT-based models were evaluated using predictive measures including Quadratic Weighted Kappa, Mean Absolute Error, and Root Mean Squared Error. Fairness was examined using demographic parity difference, equal opportunity difference, and disparate impact ratio, while SHAP and LIME were used for explainability. Dataset balancing and adversarial debiasing were assessed as fairness optimization strategies.

Findings: The BERT-based model achieved the strongest baseline predictive performance, with a QWK of 0.83, MAE of 0.65, and RMSE of 0.88. Fairness analysis identified measurable disparities across linguistic essay groups, with traditional models showing larger differences than the deep-learning models. After fairness optimization, demographic parity differences decreased by approximately 20-30% across models, while predictive performance showed only minor reductions.

Theoretical Contributions: The study integrates predictive evaluation, algorithmic fairness metrics, and explainable AI within a unified AES assessment framework. It demonstrates how linguistic features such as essay length, vocabulary diversity, word count, and syntactic complexity can influence automated scoring and provides a structured basis for examining performance-fairness trade-offs in NLP-based educational assessment.

Conclusions and Policy Implications: Fairness-aware interventions can reduce scoring disparities without substantially compromising model accuracy. Educational institutions and developers should incorporate routine bias testing, explainability, fairness monitoring, secure data governance, and human oversight when deploying automated language assessment systems, particularly in high-stakes settings.

Downloads

Published

2025-09-15

Similar Articles

1-10 of 16

You may also start an advanced similarity search for this article.