Skip to main navigation Skip to search Skip to main content

Comparing GPT-based approaches in automated writing evaluation

Research output: Contribution to journalArticlepeer-review

Abstract

Large language models (LLMs) like OpenAI's GPT models show significant promise in automated writing evaluation (AWE). However, recent research has mainly focused on non-fine-tuned GPT models, with limited attention to fine-tuned models as well as potential factors influencing performance, such as model type, prompting strategy, and dataset characteristics. This study compares six GPT-based approaches for evaluating TOFEL argumentative writing, namely, GPT-3.5 zero-shot, GPT-3.5 few-shot, GPT-4 zero-shot, GPT-4 few-shot, and two fine-tuning methods. We assess the impact of model type (GPT-3.5 vs. GPT-4), prompting strategy (zero-shot vs. few-shot), fine-tuning, class imbalance and dataset shift on performance. Our findings reveal that fine-tuned GPT models consistently outperform non-fine-tuned GPT-4 models, which in turn outperform GPT-3.5 models. Few-shot prompting does not show clear advantages over zero-shot prompting in this study. Additionally, class imbalance and dataset shift negatively affect model accuracy and reliability. These results offer valuable insights into the effectiveness of different GPT-based approaches and the factors that influence their performance in AWE.

Original languageEnglish (US)
Article number100961
JournalAssessing Writing
Volume66
DOIs
StatePublished - Oct 2025

All Science Journal Classification (ASJC) codes

  • Language and Linguistics
  • Education
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'Comparing GPT-based approaches in automated writing evaluation'. Together they form a unique fingerprint.

Cite this