Counterargument for Critical Thinking as Judged by AI and Humans
Background:
Education is evolving rapidly, thanks to GenAI, particularly large language models (LLMs), which have many benefits. Recent studies, however, have shown that people who used ChatGPT or similar tools to write essays demonstrated less brain activity in areas linked with cognitive processing or reported negative impact to their skills. This behaviour of shifting cognitive tasks to AI is termed cognitive offloading. Given AI is here to stay, it is beneficial to find ways that students can engage with it such that it encourages critical thinking and boosts their skills and learning.
Aim:
The aim is to answer the following research questions: (1) Do counterarguments to AI-generated arguments promote students’ critical thinking in writing, in terms of logic and relevant rubrics? and (2) How similar (or dis-similar) are AI systems in judging counterarguments compared to humans, given the same rubrics?
Methodology:
The participants were 36 Master students of the Text Mining course (D7058E) of Luleå University of Technology, Sweden, for the 2025/26 calendar period. The students are nationals of different countries and are mostly around the same age bracket in their early twenties. Completing the task involved submitting the LLM argument of their choice, their self-written counterargument, and 2 peer-reviews that were randomly assigned on the Canvas learning management system (LMS). Four topics of debate were provided.
- Statistically non-significant results indicate ‘no difference’
- Humans are born with an innate capacity for language
- The fate of mutations, which occurs randomly, is singularly governed by natural selection
- Pedagogy relates only to ways or methods of teaching
Results:
The medians and modes show, at least, Agree for all rubrics in the content of the counterarguments of the submissions, including Logic, which is the most important for counterargument. The expert appears stricter with Logic in assessment than both the average student and AI. The Gwets AC2 inter-rater reliability values for all the AI models are 0.33, except 0.17 for LLaMA 4-Maverick 17B, indicating some similarity with expert assessment, though all are lower than students’ agreement values.

Diverging stacked chart for rubrics Reference, Correctness, Style, Content, Logic, and Focus for expert, average student and AI assessments, respectively.
Key Takeaways:
- Counterarguments to AI-generated arguments promote students’ critical thinking in writing, in terms of logic and relevant rubrics.
- LLMs can provide relatively similar assessments of counterarguments to those of human experts, given the same clear rubrics.
Thanks to the Department of Computer Science, Electrical & Space Engineering at Luleå University of Technology for the 2026 SRT pedagogy fund for this project. We also thank all the participating students of the Text Mining (D7058E) course, 2025/26 session.
Updated: