Skip to main navigation Skip to search Skip to main content

Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The rapid advancement of Generative AIs (GenAIs), particularly Large Language Models (LLMs), has transformed software engineering by automating tasks such as code generation, testing, and debugging. As these models become increasingly integrated into development workflows, evaluating their performance systematically is crucial for optimizing their effectiveness in real-world applications. This study aims to benchmark six prominent LLMs - ChatGPT, Copilot, Gemini, Claude, Perplexity, and DeepSeek - on algorithm and data structure problems from LeetCode, assessing their strengths and limitations in solving programming challenges. The study evaluates LLM performance on 150 LeetCode problems, generating solutions in Java and Python. Performance metrics include execution time, memory usage, and computational complexity (time and space). A joint analysis ranks the models based on multiple performance factors. The evaluation reveals variations in LLM efficiency, with some models consistently outperforming others across different difficulty levels. Copilot, DeepSeek, and Perplexity demonstrate strong performance, while Gemini struggles with harder problems. Differences in execution time and memory usage are also noted across programming languages. The findings contribute to a deeper understanding of LLM capabilities in code generation and provide insights to help developers make informed decisions when selecting LLMs based on problem complexity and programming language.

Original languageEnglish (US)
Title of host publicationProceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , EASE, 2025 edition, EASE 2025
EditorsMuhammad Ali Babar, Ayse Tosun, Stefan Wagner, Viktoria Stray
PublisherAssociation for Computing Machinery, Inc
Pages949-958
Number of pages10
ISBN (Electronic)9798400713859
DOIs
StatePublished - Dec 24 2025
Event29th International Conference on Evaluation and Assessment of Software Engineering, EASE 2025 - Istanbul, Turkey
Duration: Jun 17 2025Jun 20 2025

Publication series

NameProceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , EASE, 2025 edition, EASE 2025

Conference

Conference29th International Conference on Evaluation and Assessment of Software Engineering, EASE 2025
Country/TerritoryTurkey
CityIstanbul
Period6/17/256/20/25

All Science Journal Classification (ASJC) codes

  • Software

Fingerprint

Dive into the research topics of 'Analyzing Prominent LLMs: An Empirical Study of Performance and Complexity in Solving LeetCode Problems'. Together they form a unique fingerprint.

Cite this