دقة ChatGPT واتفاقه وموثوقيته في تقويم القواعد النحوية العربية: دراسة مقارنة مع المقيمين البشريين
Keywords:
ChatGPT, تقييم قواعد اللغة العربية, لموثوقية بين المقيّمينAbstract
دأت نماذج اللغة الكبيرة مثل ChatGPT تُستخدم كمساعد في تصحيح الإجابات المفتوحة، ومنها التحليل النحوي للغة العربية، إلا أن موثوقيتها مقارنة بالمقيّمين البشريين لم تُختبر تجريبيًا بما يكفي في لغة ذات بنية صرفية معقدة كاللغة العربية. قيّمت هذه الدراسة دقة ChatGPT واتفاقه وموثوقيته في تقييم اختبار قواعد اللغة العربية مقارنة بالمقيّمين البشريين، باستخدام تصميم كمي مقارن بمقاربة سيكومترية. تكوّنت الأداة من 20 بندًا مقاليًا تغطي أربعة جوانب (الإعراب، والتصريف، والتركيب، والمفردات/الإملاء)، طُبِّقت على 20 طالبًا من الصف الحادي عشر بمدرسة MAN 2 مدينة مالانج، فأنتجت 400 وحدة تقييم قيّمها مفتاح الخبير، ومقيِّم ثانٍ مستقل، وChatGPT (مرتين لاختبار الثبات). استُخدم في التحليل نسبة التطابق التام، ومعامل كابا الموزون لكوهين، ومعامل الارتباط داخل الصنف (ICC). بلغت الدقة الخام لـChatGPT 65.2%، لكن بعد تصحيح احتمال التطابق العرضي، بلغت قيمة كابا 0.008 فقط (لا تختلف عن الصفر)، أدنى بكثير من التوافق شبه التام بين المقيّمين البشريين (كابا = 0.891). وكانت موثوقية الاختبار وإعادة الاختبار لدى ChatGPT تامة (درجات متطابقة في جميع الوحدات)، إلا أن ICC للمقيّم الواحد كان ضعيفًا 0.371. وتشير هذه النتائج إلى أن ChatGPT لا يصلح بعد ليكون مقيّمًا مستقلًا لقواعد اللغة العربية: فهو متسق مع ذاته لكنه متساهل بشكل منهجي، إذ منح الدرجة الكاملة لـ27 من أصل 33 إجابة اعتبرها الخبير خاطئة كليًا (81.8%)، وفشل في التمييز بين الطلاب الضعفاء والأقوياء عند مستوى الدرجة
References
Abu Guba, M. N., & Abu Quba, A. (2025). AI Translation: Evaluating ChatGPT’s Reliability in Translating Arabic to English. Journal of Artificial Intelligence and Technology. https://doi.org/10.37965/jait.2025.0831
Alqurashi, A., Alharbi, B., & Sabbeh, S. (2025). An Automatic Grading System for Arabic Language Short-Answer Questions Using Deep Learning. Engineering, Technology & Applied Science Research, 15(5), 26665–26675. https://doi.org/10.48084/etasr.10917
Aydın, B., Kışla, T., Elmas, N. T., & Bulut, O. (2025). Automated scoring in the era of artificial intelligence: An empirical study with Turkish essays. System, 133, 103784. https://doi.org/10.1016/j.system.2025.103784
Bouziane, K., & Bouziane, A. (2024). AI versus human effectiveness in essay evaluation. Discover Education, 3(1), 201. https://doi.org/10.1007/s44217-024-00320-6
Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of Human and Machine Scoring of Essays: Differences by Gender, Ethnicity, and Country. Applied Measurement in Education, 25(1), 27–40. https://doi.org/10.1080/08957347.2012.635502
Bucol, J. L., & Sangkawong, N. (2025). Exploring ChatGPT as a writing assessment tool. Innovations in Education and Teaching International, 62(3), 867–882. https://doi.org/10.1080/14703297.2024.2363901
Bui, N. M., & Barrot, J. (2025a). Using generative artificial intelligence as an automated essay scoring tool: a comparative study. Innovation in Language Learning and Teaching, 1–16. https://doi.org/10.1080/17501229.2025.2521003
Bui, N. M., & Barrot, J. S. (2025b). ChatGPT as an automated essay scoring tool in the writing classrooms: how it compares with human scoring. Education and Information Technologies, 30(2), 2041–2058. https://doi.org/10.1007/s10639-024-12891-w
Chan, K. K. Y., Bond, T., & Yan, Z. (2023). Application of an Automated Essay Scoring engine to English writing assessment using Many-Facet Rasch Measurement. Language Testing, 40(1), 61–85. https://doi.org/10.1177/02655322221076025
Cicchetti, D. V., & Feinstein, A. R. (1990). High agreement but low kappa: II. Resolving the paradoxes. Journal of Clinical Epidemiology, 43(6), 551–558. https://doi.org/10.1016/0895-4356(90)90159-M
Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. https://doi.org/10.1037/h0026256
Du, H., Jia, Q., Gehringer, E., & Wang, X. (2024). Harnessing large language models to auto-evaluate the student project reports. Computers and Education: Artificial Intelligence, 7, 100268. https://doi.org/10.1016/j.caeai.2024.100268
Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low Kappa: I. the problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549. https://doi.org/10.1016/0895-4356(90)90158-L
Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal, 51(1), 201–224. https://doi.org/10.1002/berj.4069
Ghazawi, R., & Simpson, E. (2025). How well can LLMs grade essays in Arabic? Computers and Education: Artificial Intelligence, 9, 100449. https://doi.org/10.1016/j.caeai.2025.100449
Gomaa, W. H., & Fahmy, A. A. (2014). Automatic scoring for answers to Arabic test questions. Computer Speech & Language, 28(4), 833–857. https://doi.org/10.1016/j.csl.2013.10.005
Jung, J. Y., Tyack, L., & von Davier, M. (2024). Combining machine translation and automated scoring in international large-scale assessments. Large-Scale Assessments in Education, 12(1), 10. https://doi.org/10.1186/s40536-024-00199-7
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
Kim, Y. (2025). Automated Essay Scoring With GPT ‐4 for a Local Placement Test: Investigating Prompting Strategies, Intra‐Rater Reliability, and Alignment With Human Scores. TESOL Quarterly, 59(S1). https://doi.org/10.1002/tesq.3405
Koo, T. K., & Li, M. Y. (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
Kooli, C., & Yusuf, N. (2024). Transforming Educational Assessment: Insights Into the Use of ChatGPT and Large Language Models in Grading. International Journal of Human–Computer Interaction, 1–12. https://doi.org/10.1080/10447318.2024.2338330
Küchemann, S., Rau, M., Schmidt, A., & Kuhn, J. (2024). ChatGPT’s quality: Reliability and validity of concept inventory items. Frontiers in Psychology, 15. https://doi.org/10.3389/fpsyg.2024.1426209
Landis, J. R., & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159. https://doi.org/10.2307/2529310
Lotfy, N., Shehab, A., Elhoseny, M., & Abu-Elfetouh, A. (2023). An Enhanced Automatic Arabic Essay Scoring System Based on Machine Learning Algorithms. Computers, Materials & Continua, 77(1), 1227–1249. https://doi.org/10.32604/cmc.2023.039185
Manning, J., Baldwin, J., & Powell, N. (2025). Human versus machine: The effectiveness of ChatGPT in automated essay scoring. Innovations in Education and Teaching International, 62(5), 1500–1513. https://doi.org/10.1080/14703297.2025.2469089
Quah, B., Zheng, L., Sng, T. J. H., Yong, C. W., & Islam, I. (2024). Reliability of ChatGPT in automated essay scoring for dental undergraduate examinations. BMC Medical Education, 24(1), 962. https://doi.org/10.1186/s12909-024-05881-6
Shin, D., & Lee, J. H. (2024). Exploratory study on the potential of ChatGPT as a rater of second language writing. Education and Information Technologies, 29(18), 24735–24757. https://doi.org/10.1007/s10639-024-12817-6
Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. https://doi.org/10.1037/0033-2909.86.2.420
Steiss, J., Tate, T., Graham, S., Cruz, J., Hebert, M., Wang, J., Moon, Y., Tseng, W., Warschauer, M., & Olson, C. B. (2024). Comparing the quality of human and ChatGPT feedback of students’ writing. Learning and Instruction, 91, 101894. https://doi.org/10.1016/j.learninstruc.2024.101894
Uyar, A. C., & Büyükahıska, D. (2025). Artificial intelligence as an automated essay scoring tool: A focus on ChatGPT. International Journal of Assessment Tools in Education, 12(1), 20–32. https://doi.org/10.21449/ijate.1517994
Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay grading: An examination of reliability and validity in rubric‐based assessments. British Journal of Educational Technology, 56(1), 150–166. https://doi.org/10.1111/bjet.13494
Zhao, R., Zhuang, Y., Zou, D., Xie, Q., & Yu, P. L. H. (2023). AI-assisted automated scoring of picture-cued writing tasks for language assessment. Education and Information Technologies, 28(6), 7031–7063. https://doi.org/10.1007/s10639-022-11473-y
