Roshanak Zilouchian Moghaddam

dblp:86/914 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0000-2268-5897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
abstract
Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better understand the unique limitations of LM agents, we introduce RefactorBench, a benchmark consisting of 100 large handcrafted multi-file refactoring tasks in popular open-source repositories. Solving tasks within RefactorBench requires thorough exploration of dependencies across multiple files and strong adherence to relevant instructions. Every task is defined by 3 natural language instructions of varying specificity and is mutually exclusive, allowing for the creation of longer combined tasks on the same repository. Baselines on RefactorBench reveal that current LM agents struggle with simple compositional tasks, solving only 22\% of tasks with base instructions, in contrast to a human developer with short time constraints solving 87\%. Through trajectory analysis, we identify various unique failure modes of LM agents, and further explore the failure mode of tracking past actions. By adapting a baseline agent to condition on representations of state, we achieve a 43.9\% improvement in solving RefactorBench tasks. We further extend our state-aware approach to encompass entire digital environments and outline potential directions for future research. RefactorBench aims to support the study of LM agents by providing a set of real-world, multi-hop tasks within the realm of code.
Dhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan, Roshanak Zilouchian Moghaddam
ICLR5
2025 Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE
abstract
Security vulnerabilities impose significant costs on users and organizations. Detecting and addressing these vulnerabilities early is crucial to avoid exploits and reduce development costs. Recent studies have shown that deep learning models can effectively detect security vulnerabilities. Yet, little research explores how to adapt these models from benchmark tests to practical applications, and whether they can be useful in practice. This paper presents the first empirical study of a vulnerability detection and fix tool with professional software developers on real projects that they own. We implemented DeepVulguard, an IDE-integrated tool based on state-of-the-art detection and fix models, and show that it has promising performance on benchmarks of historic vulnerability data. DeepVulguard scans code for vulnerabilities (including identifying the vulnerability type and vulnerable region of code), suggests fixes, provides natural-language explanations for alerts and fixes, leveraging chat interfaces. We recruited 17 professional software developers at Microsoft, observed their usage of the tool on their code, and conducted interviews to assess the tool's usefulness, speed, trust, relevance, and workflow integration. We also gathered detailed qualitative feedback on users' perceptions and their desired features. Study participants scanned a total of 24 projects, 6.9 k files, and over 1.7 million lines of source code, and generated 170 alerts and 50 fix suggestions. We find that although state-of-the-art AI-powered detection and fix tools show promise, they are not yet practical for real-world use due to a high rate of false positives and non-applicable fixes. User feedback reveals several actionable pain points, ranging from incomplete context to lack of customization for the user's codebase. Additionally, we explore how AI features, including confidence scores, explanations, and chat interaction, can apply to vulnerability detection and fixing. Based on these insights, we offer practical recommendations for evaluating and deploying AI detection and fix models. Our code and data are available at this link: https://doi.org/10.6084/m9.figshare.26367139.
Benjamin Steenhoek, Kalpathy Sivaraman, Renata Saldivar Gonzalez, Yevhen Mohylevskyy, Roshanak Zilouchian Moghaddam, Wei Le
ICSE5
2022 Learning to Reduce False Positives in Analytic Bug Detectors
abstract
Due to increasingly complex software design and rapid iterative development, code defects and security vulnerabilities are prevalent in modern software. In response, programmers rely on static analysis tools to regularly scan their codebases and find potential bugs. In order to maximize coverage, however, these tools generally tend to report a significant number of false positives, requiring developers to manually verify each warning. To address this problem, we propose a Transformer-based learning approach to identify false positive bug warnings. We demonstrate that our models can improve the precision of static analysis by 17.5%. In addition, we validated the generalizability of this approach across two major bug types: null dereference and resource leak.
Anant Kharkar, Roshanak Zilouchian Moghaddam, Matthew Jin, Colin B. Clement, Neel Sundaresan
ICSE2
2022 Generating Examples from CLI Usage: Can Transformers Help?
abstract
Continuous evolution in modern software often causes documentation, tutorials, and examples to be out of sync with changing interfaces and frameworks. Relying on outdated documentation and examples can lead programs to fail or be less efficient or even less secure. In response, programmers need to regularly turn to other resources on the web, such as StackOverflow for examples to guide them in writing software. We recognize that this inconvenient, error-prone, and expensive process can be improved by using machine learning applied to software usage data. In this paper, we present a practical system, which uses machine learning on large-scale telemetry data and documentation corpora, generating appropriate and complex examples that can be used to improve documentation. We discuss both feature-based and transformer-based machine learning approaches and demonstrate that our system achieves 100% coverage for the used functionalities in the product, providing up-to-date examples upon every release and reduces the numbers of PRs submitted by software owners writing and editing documentation by >68%. We also share valuable lessons learnt during the 3 years that our production quality system has been deployed for Azure Cloud Command Line Interface (Azure CLI)
Roshanak Zilouchian Moghaddam, Spandan Garg, Colin B. Clement, Yevhen Mohylevskyy, Neel Sundaresan
KDD1
2022 DeepDev-PERF: a deep learning-based approach for improving software performance
abstract
Improving software performance is an important yet challenging part of the software development cycle. Today, the majority of performance inefficiencies are identified and patched by performance experts. Recent advancements in deep learning approaches and the wide-spread availability of open-source data creates a great opportunity to automate the identification and patching of performance problems. In this paper, we present DeepDev-PERF, a transformer-based approach to suggest performance improvements for C# applications. We pretrain DeepDev-PERF on English and Source code corpora, followed by finetuning for the task of generating performance improvement patches for C# applications. Our evaluation shows that our model can generate the same performance improvement suggestion as the developer fix in ‍53
Spandan Garg, Roshanak Zilouchian Moghaddam, Colin B. Clement, Neel Sundaresan
ESEC/SIGSOFT FSE2
2015 Procid: Bridging Consensus Building Theory with the Practice of Distributed Design Discussions
abstract
Consensus is a desired but elusive goal in many distributed discussions. A critical problem is that discussion platforms lack mechanisms for realizing consensus strategies and realizing these strategies without tool support can be hard. This paper introduces Procid, a novel browser plugin that provides interaction and visualization features for bringing consensus strategies to distributed design discussions. Key features include the ability to organize discussions around ideas, to register and visualize support for or against ideas, and to define criteria for evaluating ideas. It also applies interaction constraints fostering best practices of consensus building. Procid extends the discussion platform of one open source software community. Two evaluations were conducted. The first collected perceptions of the tool from members of the community for their own discussions. The second compared how Procid affects a distributed design discussion relative to the current discussion platform in the community. Results of both studies showed that users found the features of our tool beneficial and perceived it as more effective for consensus building than the existing platform.
Roshanak Zilouchian Moghaddam, Zane Nicholson, Brian P. Bailey
CSCW1
2013 A Compositional Paradigm of Automating Refactorings
Mohsen Vakilian, Nicholas Chen, Roshanak Zilouchian Moghaddam, Stas Negara, Ralph E. Johnson
ECOOP3
2012 Consensus building in open source user interface design discussions
abstract
We report results of a study which examines consensus building in user interface design discussions in open source software communities. Our methodology consisted of conducting interviews with designers and developers from the Drupal and Ubuntu communities (N=17) and analyzing a large corpus of interaction data collected from Drupal. The interviews captured user perspectives on the challenges of reaching consensus, techniques employed for building consensus, and the consequences of not reaching consensus. We analyzed the interaction data to determine how different elements of the content, process, and user relationships in the design discussions affect consensus. Our main result shows that design discussions engaging participants with more experience and prior interaction history are more likely to reach consensus. Based on all of our results, we formulated design implications for promoting consensus in distributed discussions of user interface design issues.
Roshanak Zilouchian Moghaddam, Brian P. Bailey, Wai-Tat Fu
CHI1
2011 IdeaTracker: An Interactive Visualization Supporting Collaboration and Consensus Building in Online Interface Design Discussions
Roshanak Zilouchian Moghaddam, Brian P. Bailey, Christina M. Poon
INTERACT (1)1