Bias Detection And Fairness Optimization In Nlp-Based Language Assessment Systems
Keywords:
Automated Essay Scoring, Educational Assessment, Natural Language Processing, Fairness in AIAbstract
Background and Purpose: Natural Language Processing (NLP) has expanded automated language assessment through systems such as Automated Essay Scoring (AES), which provide scalable and consistent evaluation of student writing. However, automated scoring can reproduce disparities associated with linguistic, cultural, and socio-economic differences when models learn biased patterns from training data. This study investigates bias detection and fairness optimization in NLP-based language assessment systems.
Methods: A quantitative experimental framework was applied using the Automated Student Assessment Prize (ASAP) essay dataset. Support Vector Machine, Random Forest, Long Short-Term Memory, and BERT-based models were evaluated using predictive measures including Quadratic Weighted Kappa, Mean Absolute Error, and Root Mean Squared Error. Fairness was examined using demographic parity difference, equal opportunity difference, and disparate impact ratio, while SHAP and LIME were used for explainability. Dataset balancing and adversarial debiasing were assessed as fairness optimization strategies.
Findings: The BERT-based model achieved the strongest baseline predictive performance, with a QWK of 0.83, MAE of 0.65, and RMSE of 0.88. Fairness analysis identified measurable disparities across linguistic essay groups, with traditional models showing larger differences than the deep-learning models. After fairness optimization, demographic parity differences decreased by approximately 20-30% across models, while predictive performance showed only minor reductions.
Theoretical Contributions: The study integrates predictive evaluation, algorithmic fairness metrics, and explainable AI within a unified AES assessment framework. It demonstrates how linguistic features such as essay length, vocabulary diversity, word count, and syntactic complexity can influence automated scoring and provides a structured basis for examining performance-fairness trade-offs in NLP-based educational assessment.
Conclusions and Policy Implications: Fairness-aware interventions can reduce scoring disparities without substantially compromising model accuracy. Educational institutions and developers should incorporate routine bias testing, explainability, fairness monitoring, secure data governance, and human oversight when deploying automated language assessment systems, particularly in high-stakes settings.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Kazi Siam Al Mobin, Afroza Riju (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.