TY - GEN
T1 - Analyzing Prominent LLMs
T2 - 29th International Conference on Evaluation and Assessment of Software Engineering, EASE 2025
AU - Guimaraes, Everton
AU - Moraes Do Nascimento, Nathalia
AU - Nelapati, Asish
AU - Shivalingaiah, Chandan
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/12/24
Y1 - 2025/12/24
N2 - The rapid advancement of Generative AIs (GenAIs), particularly Large Language Models (LLMs), has transformed software engineering by automating tasks such as code generation, testing, and debugging. As these models become increasingly integrated into development workflows, evaluating their performance systematically is crucial for optimizing their effectiveness in real-world applications. This study aims to benchmark six prominent LLMs - ChatGPT, Copilot, Gemini, Claude, Perplexity, and DeepSeek - on algorithm and data structure problems from LeetCode, assessing their strengths and limitations in solving programming challenges. The study evaluates LLM performance on 150 LeetCode problems, generating solutions in Java and Python. Performance metrics include execution time, memory usage, and computational complexity (time and space). A joint analysis ranks the models based on multiple performance factors. The evaluation reveals variations in LLM efficiency, with some models consistently outperforming others across different difficulty levels. Copilot, DeepSeek, and Perplexity demonstrate strong performance, while Gemini struggles with harder problems. Differences in execution time and memory usage are also noted across programming languages. The findings contribute to a deeper understanding of LLM capabilities in code generation and provide insights to help developers make informed decisions when selecting LLMs based on problem complexity and programming language.
AB - The rapid advancement of Generative AIs (GenAIs), particularly Large Language Models (LLMs), has transformed software engineering by automating tasks such as code generation, testing, and debugging. As these models become increasingly integrated into development workflows, evaluating their performance systematically is crucial for optimizing their effectiveness in real-world applications. This study aims to benchmark six prominent LLMs - ChatGPT, Copilot, Gemini, Claude, Perplexity, and DeepSeek - on algorithm and data structure problems from LeetCode, assessing their strengths and limitations in solving programming challenges. The study evaluates LLM performance on 150 LeetCode problems, generating solutions in Java and Python. Performance metrics include execution time, memory usage, and computational complexity (time and space). A joint analysis ranks the models based on multiple performance factors. The evaluation reveals variations in LLM efficiency, with some models consistently outperforming others across different difficulty levels. Copilot, DeepSeek, and Perplexity demonstrate strong performance, while Gemini struggles with harder problems. Differences in execution time and memory usage are also noted across programming languages. The findings contribute to a deeper understanding of LLM capabilities in code generation and provide insights to help developers make informed decisions when selecting LLMs based on problem complexity and programming language.
UR - https://www.scopus.com/pages/publications/105027046288
UR - https://www.scopus.com/pages/publications/105027046288#tab=citedBy
U2 - 10.1145/3756681.3756983
DO - 10.1145/3756681.3756983
M3 - Conference contribution
AN - SCOPUS:105027046288
T3 - Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , EASE, 2025 edition, EASE 2025
SP - 949
EP - 958
BT - Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , EASE, 2025 edition, EASE 2025
A2 - Babar, Muhammad Ali
A2 - Tosun, Ayse
A2 - Wagner, Stefan
A2 - Stray, Viktoria
PB - Association for Computing Machinery, Inc
Y2 - 17 June 2025 through 20 June 2025
ER -