Abstract
The subjective evaluation of early-stage engineering designs, such as concept sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these artificial intelligence (AI) “judges” perform on par with human experts. This work introduces in-context learning (ICL)-enhanced VLM judges and a comprehensive statistical framework (including agreement, error, correlation, statistical difference checks, equivalence testing, and top-set overlap) to rigorously assess AI–expert equivalence. Across two case studies, we show that reasoning-enabled VLMs are the strongest-performing AI judges. They consistently outperform two-third trained novices across all metrics, and for measures such as uniqueness, creativity, and drawing quality, they approach expert-equivalent performance. In specific cases, they even exceed expert–expert agreement, attaining lower mean absolute error and higher rank correlations than the expert baseline. These findings suggest that, on certain statistical tests, AI judges are not only approaching expert–expert equivalence but in some cases surpassing it.
| Original language | English (US) |
|---|---|
| Article number | 071704 |
| Journal | Journal of Mechanical Design |
| Volume | 148 |
| Issue number | 7 |
| DOIs | |
| State | Published - Jul 1 2026 |
All Science Journal Classification (ASJC) codes
- Mechanics of Materials
- Mechanical Engineering
- Computer Science Applications
- Computer Graphics and Computer-Aided Design
Fingerprint
Dive into the research topics of 'AI Judges in Design: Toward Expert-Equivalent Design Evaluations With Vision-Language Models and In-Context Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver