TY - GEN
T1 - Turning Manual Tasks into Actions
T2 - 2025 2nd IEEE/ACM International Conference on AI-powered Software, AIware 2025
AU - Peixoto, Myron David Lucena Campos
AU - Fonseca, Baldoino
AU - De Medeiros Baia, Davy
AU - Lira, Kevin
AU - Ribeiro, Marcio
AU - Assuncao, Wesley K.G.
AU - Nascimento, Nathalia
AU - Alencar, Paulo
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large Language Models (LLMs) have introduced innovative avenues for automating software testing using prompts. Despite numerous studies on software testing automation, there remains limited understanding on the effectiveness of LLM-generated Selenium tests. In this paper, we investigate the effectiveness of Gemini to produce Selenium tests from manual tasks specifications and HyperText Markup Language (HTML) code snippets. By effectiveness, we mean if the generated Selenium tests are executable and functionally accurate (meeting intended behavior specified in a manual task). To do that, we specify eight manual tasks (involving tasks related to search, filter, navigation, and form submissions) and define 25 actions for each task, using HTML code extracted from 200 web pages. These tasks require the interaction of diverse User Interface (UI) components, such as search boxes and checkboxes. The results indicate that 87.5 % of the generated Selenium tests are executable and 51.5 % of them meet the intended behavior. Manual tasks involving interaction with modals presented the greatest challenges for test generation. While carousels and buttons achieved relatively high success rates, they still accounted for many of the post-correction fixes. These components-often dynamic or context-dependent-were among those where most errors occurred during test generation.
AB - Large Language Models (LLMs) have introduced innovative avenues for automating software testing using prompts. Despite numerous studies on software testing automation, there remains limited understanding on the effectiveness of LLM-generated Selenium tests. In this paper, we investigate the effectiveness of Gemini to produce Selenium tests from manual tasks specifications and HyperText Markup Language (HTML) code snippets. By effectiveness, we mean if the generated Selenium tests are executable and functionally accurate (meeting intended behavior specified in a manual task). To do that, we specify eight manual tasks (involving tasks related to search, filter, navigation, and form submissions) and define 25 actions for each task, using HTML code extracted from 200 web pages. These tasks require the interaction of diverse User Interface (UI) components, such as search boxes and checkboxes. The results indicate that 87.5 % of the generated Selenium tests are executable and 51.5 % of them meet the intended behavior. Manual tasks involving interaction with modals presented the greatest challenges for test generation. While carousels and buttons achieved relatively high success rates, they still accounted for many of the post-correction fixes. These components-often dynamic or context-dependent-were among those where most errors occurred during test generation.
UR - https://www.scopus.com/pages/publications/105035169117
UR - https://www.scopus.com/pages/publications/105035169117#tab=citedBy
U2 - 10.1109/AIware69974.2025.00012
DO - 10.1109/AIware69974.2025.00012
M3 - Conference contribution
AN - SCOPUS:105035169117
T3 - Proceedings - 2025 2nd IEEE/ACM International Conference on AI-powered Software, AIware 2025
SP - 40
EP - 49
BT - Proceedings - 2025 2nd IEEE/ACM International Conference on AI-powered Software, AIware 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 November 2025 through 20 November 2025
ER -