Automated Grading and Professional Accounting Education: Examining the Fairness, Reliability, and Validity of AI Grades
Loading...
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
MDPI
Abstract
Automated long-essay scoring (ALES) is gradually considered as a means to enhance efficiency and consistency in large-scale assessment; however, concerns remain regarding its suitability, particularly as it relates to the reliability, validity, and fairness of ALES-assigned grades relative to human-grades in high-stakes professional contexts. This study examines these concerns using over 15,000 long essay examination scripts from a professional accounting certification examination. The study examines whether the ALES confidence index (CI) meaningfully predicts grading accuracy or points to systemic grading failures. Findings reveal fair overall agreement between human and ALES grades, with high within ±1 grade agreement, and rare yet task-concentrated ALES grading failures, while CI shows statistically significant but practically weak predictive value and limited discrimination. The results support the use of ALES as an assistive, human oversight tool rather than an independent grader, highlighting the importance of task-based validation, stronger calibration analysis, and continuous human supervision in high-stakes professional assessment contexts. The study advances innovative assessment practices, but calls for cautious deployment of ALES and recommends integration of a hybrid human-in-the-loop approach, multi-disciplinary validation, and capacity building to strengthen ethical and responsible AI usage in accounting education and professional practice, aligning with SDGs 4 and 9.