Abstract
Large language models (LLMs) like OpenAI's GPT models show significant promise in automated writing evaluation (AWE). However, recent research has mainly focused on non-fine-tuned GPT models, with limited attention to fine-tuned models as well as potential factors influencing performance, such as model type, prompting strategy, and dataset characteristics. This study compares six GPT-based approaches for evaluating TOFEL argumentative writing, namely, GPT-3.5 zero-shot, GPT-3.5 few-shot, GPT-4 zero-shot, GPT-4 few-shot, and two fine-tuning methods. We assess the impact of model type (GPT-3.5 vs. GPT-4), prompting strategy (zero-shot vs. few-shot), fine-tuning, class imbalance and dataset shift on performance. Our findings reveal that fine-tuned GPT models consistently outperform non-fine-tuned GPT-4 models, which in turn outperform GPT-3.5 models. Few-shot prompting does not show clear advantages over zero-shot prompting in this study. Additionally, class imbalance and dataset shift negatively affect model accuracy and reliability. These results offer valuable insights into the effectiveness of different GPT-based approaches and the factors that influence their performance in AWE.
| Original language | English (US) |
|---|---|
| Article number | 100961 |
| Journal | Assessing Writing |
| Volume | 66 |
| DOIs | |
| State | Published - Oct 2025 |
All Science Journal Classification (ASJC) codes
- Language and Linguistics
- Education
- Linguistics and Language
Fingerprint
Dive into the research topics of 'Comparing GPT-based approaches in automated writing evaluation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver