Larissa Rocha Soares

dblp:155/7275 · also Larissa Bastos, Larissa Rocha 0001, Larissa Soares 0001 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-8069-5249ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Quality Assessment of Python Tests Generated by Large Language Models
abstract
The manual generation of test scripts is a time-intensive, costly, and error-prone process, indicating the value of automated solutions. Large Language Models (LLMs) have shown great promise in this domain, leveraging their extensive knowledge to produce test code more efficiently. This study investigates the quality of Python test code generated by three LLMs: GPT-4o, Amazon Q, and LLama 3.3. We evaluate the structural reliability of test suites generated under two distinct prompt contexts: Text2Code (T2C) and Code2Code (C2C). Our analysis includes the identification of errors and test smells, with a focus on correlating these issues to inadequate design patterns. Our findings reveal that most test suites generated by the LLMs contained at least one error or test smell. Assertion errors were the most common, comprising 64% of all identified errors, while the test smell Lack of Cohesion of Test Cases was the most frequently detected (41%). Prompt context significantly influenced test quality; textual prompts with detailed instructions often yielded tests with fewer errors but a higher incidence of test smells. Among the evaluated LLMs, GPT-4o produced the fewest errors in both contexts (10% in C2C and 6% in T2C), whereas Amazon Q had the highest error rates (19% in C2C and 28% in T2C). For test smells, Amazon Q had fewer detections in the C2C context (9%), while LLama 3.3 performed best in the T2C context (10%). Additionally, we observed a strong relationship between specific errors, such as assertion or indentation issues, and test case cohesion smells. These findings demonstrate opportunities for improving the quality of test generation by LLMs and highlight the need for future research to explore optimized generation scenarios and better prompt engineering strategies.
Victor Anthony Alves, Carla I. M. Bezerra, Ivan do Carmo Machado, Larissa Rocha Soares, Tássio Virgínio, Publio Silva
EASE4
2025 Discovering Patterns in Test Code Refactorings: A Preliminary Study
Railana Santana, Luana Almeida Martins, Larissa Rocha Soares, Carla I. M. Bezerra, Heitor A. X. Costa, Ivan do Carmo Machado
SEAA (2)3
2025 Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
abstract
Large Language Models (LLMs) have gained attention for addressing coding problems, but their effectiveness in fixing code maintainability remains unclear. This study evaluates LLMs capability to resolve 127 maintainability issues from 10 GitHub repositories. We use zero-shot prompting for Copilot Chat and Llama 3.1, and few-shot prompting with Llama only. The LLM-generated solutions are assessed for compilation errors, test failures, and new maintainability problems. Llama with few-shot prompting successfully fixed 44.9 % of the methods, while Copilot Chat and Llama zero-shot fixed 32.29 % and 30 %, respectively. However, most solutions introduced errors or new maintainability issues. We also conducted a human study with 45 participants to evaluate the readability of 51 LLM-generated solutions. The human study showed that 68.63 % of participants observed improved readability. Overall, while LLMs show potential for fixing maintainability issues, their introduction of errors highlights their current limitations.
Henrique Gomes Nunes, Eduardo Figueiredo 0001, Larissa Rocha Soares, Sarah Nadi, Fischer Ferreira, Geanderson E. dos Santos
SANER3
2024 How does parenthood affect an ICT practitioner's work? A survey study with fathers
Larissa Rocha Soares, Edna Dias Canedo, Claudia Pinto Pereira, Carla I. M. Bezerra, Fabiana Freitas Mendes
Empir. Softw. Eng.1
2024 An empirical evaluation of RAIDE: A semi-automated approach for test smells detection and refactoring
Railana Santana, Luana Almeida Martins, Tássio Virgínio, Larissa Rocha Soares, Heitor A. X. Costa, Ivan do Carmo Machado
Sci. Comput. Program.4
2023 Investigating the Perceived Impact of Maternity on Software Engineering: a Women's Perspective
abstract
Background: Several researchers report the impact of gender on software development teams, especially in relation to women. In general, women are under-represented on these teams and face challenges and difficulties in their workplaces. When it comes to women who are mothers, these challenges can be amplified and directly impact these women’s professional lives, both in industry and academia. However, little is known about women ICT practitioners’ perceptions of the challenges of maternity in their professional careers. Objective: This paper investigates mothers’ challenges and difficulties in global software development teams. Method: We conducted a survey with women in the ICT field who work in academia and global technology companies. We surveyed 141 mothers from different countries and employed mixed methods to analyze the data. Results: Our findings reveal that women face sociocultural challenges, including work-life balance issues, bad jokes, and moral harassment. The prejudices they suffer make them insecure and with low confidence in the work performed. Furthermore, they usually do not have a supporting network during and after maternity leave, which culminates in them feeling overloaded. The surveyed women suggested a set of actions to reduce the challenges they face in their workplaces, such as: creating a code of conduct for men and childcare within companies. Conclusion: Women face many challenges when they become mothers. Our findings explore these challenges and can help organizations in developing policies to minimize them. Also, it can help raise awareness of co-workers and bosses, toward a more friendly and inclusive workplace.
Larissa Rocha Soares, Edna Dias Canedo, Claudia Pinto Pereira, Carla I. M. Bezerra, Fabiana Freitas Mendes
CHASE1
2021 From Blackboard to the Office: A Look Into How Practitioners Perceive Software Testing Education
abstract
The teaching-learning process may require specific pedagogical approaches to establish a relationship with industry practices. Recently, some studies investigated the educators’ perspectives and the undergraduate courses curriculum to identify potential weaknesses and solutions for the software testing teaching process. However, it is still unclear how the practitioners evaluate the acquisition of knowledge about software testing in undergraduate courses. This study carried out an expert survey with 68 newly graduated practitioners to determine what the industry expects from them and what they learned in academia. The yielded results indicated that those practitioners learned at a similar rate as others with a long industry experience. Also, they studied less than half of the 35 software testing topics collected in the survey and took industry-backed extracurricular courses to complement their learning. Additionally, our findings point out a set of implications for future research, as the respondents’ learning difficulties (e.g., lack of learning sources) and the gap between academic education and industry expectations (e.g., certifications).
Luana Almeida Martins, Vinicius Brito, Daniela Soares Feitosa, Larissa Rocha Soares, Heitor A. X. Costa, Ivan do Carmo Machado
EASE4
2018 Exploring feature interactions without specifications: a controlled experiment
abstract
In highly configurable systems, features may interact unexpectedly and produce faulty behavior. Those faults are not easily identified from the analysis of each feature separately, especially when feature specifications are missing. We propose VarXplorer, a dynamic and iterative approach to detect suspicious interactions. It provides information on how features impact the control and data flow of the program. VarXplorer supports developers with a graph that visualizes this information, mainly showing suppress and require relations between features. To evaluate whether VarXplorer helps improve the performance of identifying suspicious interactions, we perform a controlled study with 24 subjects. We find that with our proposed feature-interaction graphs, participants are able to identify suspicious interactions more than 3 times faster compared to the state-of-the-art tool.
Larissa Rocha Soares, Jens Meinicke, Sarah Nadi, Christian Kästner, Eduardo Santana de Almeida
GPCE1
2018 Feature interaction in software product line engineering: A systematic mapping study
Larissa Rocha Soares, Pierre-Yves Schobbens, Ivan do Carmo Machado, Eduardo Santana de Almeida
Inf. Softw. Technol.1
2017 Pyramidal Zernike Over Time: A Spatiotemporal Feature Descriptor Based on Zernike Moments
Igor L. O. Bastos, Larissa Rocha Soares, William Robson Schwartz
CIARP2