Arthur Marques

dblp:149/4176 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 first-author · 2 since 2021
YearPublicationVenuePosition
2022 Evaluating the Use of Semantics for Identifying Task-relevant Textual Information
abstract
The information a developer seeks to aid the completion of a task typically exists across a range of artifacts. For example, for a task that requires upgrading to a new version of an API component, a developer may seek information in the API's official documentation, check community discussions about the newer version, and so on. To aid developers in locating the portion of the text that might be useful in these artifacts, prior work has used syntactic properties of the text, and an artifact's meta-data, to automatically identify relevant text for particular kinds of artifacts. Although effective, these techniques rely on assumptions about an artifact's structure or content that prevent applying them across the different types of artifacts that a developer may come across in their daily work. In this paper, we investigate whether techniques building on approaches to interpret the meaning, or semantics, of the text help to overcome these limitations. Particularly, we introduce six semantic-based techniques and evaluate that they can identify up to 58% of the text that developers deem relevant to Android development tasks. When compared to a state-of-the-art approach, we find that our techniques achieve comparable recall values, identifying 63% of the small fraction of the task-relevant text of Stack Overflow artifacts, but without the need for artifact-specific information.
Arthur Marques, Gail C. Murphy
SANER1
2021 Assessing Semantic Frames to Support Program Comprehension Activities
abstract
Software developers often rely on natural language text that appears in software engineering artifacts to access critical information as they build and work on software systems. For example, developers access requirements documents to understand what to build, comments in source code to understand design decisions, answers to questions on Q&A sites to understand APIs, and so on. To aid software developers in accessing and using this natural language information, soft-ware engineering researchers often use techniques from natural language processing. In this paper, we explore whether frame semantics, a general linguistic approach, which has been used on requirements text, can also help address problems that occur when applying lexicon analysis based techniques to text associated with program comprehension activities. We assess the applicability of generic semantic frame parsing for this purpose, and based on the results, we propose SEFrame to tailor semantic frame parsing for program comprehension uses. We evaluate the correctness and robustness of the approach finding that SEFrame is correct in between 73% and 74% of the cases and that it can parse text from a variety of software artifacts used to support program comprehension. We describe how this approach could be used to enhance existing approaches to identify meaning on intention from software engineering texts.
Arthur Marques, Giovanni Viviani, Gail C. Murphy
ICPC1
2020 Characterizing Task-Relevant Information in Natural Language Software Artifacts
abstract
To complete a software development task, a software developer often consults artifacts which mostly consist of natural language text, such as API documentation, bug reports, and Q&A forums. Not all information within these artifacts is relevant to a developer's current task, forcing them to filter through large amounts of irrelevant information, a frustrating and time-consuming activity. Since failing to locate relevant information may lead developers to incorrect or incomplete solutions, many approaches attempt to automatically extract relevant information from natural language artifacts. However, existing approaches are able to identify relevant text only for certain types of tasks and artifacts. To explore how these limitations could be relaxed, we conducted a controlled experiment in which we asked 20 software developers to examine 20 natural language artifacts consisting of 1,874 sentences and highlight the text they considered relevant to six software development tasks. Although the 2,463 distinct highlights participants created indicate variability in the perceived relevance of the text, the information considered key to completing the tasks was consistent. We observe consistency in the text using frame semantics, an approach that captures the key meaning of sentences, suggesting that frame semantics can be used in the future to automatically identify task-relevant information in natural language artifacts.
Arthur Marques, Nick C. Bradley, Gail C. Murphy
ICSME1
2019 Helping developers search and locate task-relevant information in natural language documents
abstract
While performing a task, software developers interact with a myriad of natural language documents. Not all information in these documents is relevant to a developer's task forcing them to filter relevant information from large amounts of irrelevant information. If a developer misses some of the necessary information for her task, she will have an incomplete or incorrect basis from which to complete the task. Many approaches mine relevant text fragments from natural language artifacts. However, existing approaches mine information for pre-defined tasks and from a restricted set of artifacts. I hypothesize that it is possible to design a more generalizable approach that can identify, for a particular task, relevant text across different artifact types establishing relationships between them and facilitating how developers search and locate task-relevant information. To investigate this hypothesis, I propose to match a developer's task to text fragments in natural language artifacts according to their semantics. By semantically matching textual pieces to a developer's task we aim to more precisely identify fragments relevant to a task. To help developers in thoroughly navigating through the identified fragments I also propose to synthesize and group them. Ultimately, this research aims to help developers make more informed decisions regarding their software development task. Dr. Gail C. Murphy supervises this work.
Arthur Marques
ESEC/SIGSOFT FSE1