Jim Laredo

dblp:56/3239 · also Jim Alain Laredo · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
10since 2021 · last 2024
0000-0002-4915-0304ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Analyzing source code vulnerabilities in the D2A dataset with ML ensembles and C-BERT
abstract
Abstract Static analysis tools are widely used for vulnerability detection as they can analyze programs with complex behavior and millions of lines of code. Despite their popularity, static analysis tools are known to generate an excess of false positives. The recent ability of Machine Learning models to learn from programming language data opens new possibilities of reducing false positives when applied to static analysis. However, existing datasets to train models for vulnerability identification suffer from multiple limitations such as limited bug context, limited size, and synthetic and unrealistic source code. We propose Differential Dataset Analysis or D2A, a differential analysis based approach to label issues reported by static analysis tools. The dataset built with this approach is called the D2A dataset. The D2A dataset is built by analyzing version pairs from multiple open source projects. From each project, we select bug fixing commits and we run static analysis on the versions before and after such commits. If some issues detected in a before-commit version disappear in the corresponding after-commit version, they are very likely to be real bugs that got fixed by the commit. We use D2A to generate a large labeled dataset. We then train both classic machine learning models and deep learning models for vulnerability identification using the D2A dataset. We show that the dataset can be used to build a classifier to identify possible false alarms among the issues reported by static analysis, hence helping developers prioritize and investigate potential true positives first. To facilitate future research and contribute to the community, we make the dataset generation pipeline and the dataset publicly available. We have also created a leaderboard based on the D2A dataset, which has already attracted attention and participation from the community.
Saurabh Pujar, Yunhui Zheng, Luca Buratti, Burn L. Lewis, Yunchung Chen, Jim Laredo, Alessandro Morari, Edward A. Epstein, Tsungnan Lin, Bo Yang 0013, Zhong Su
Empir. Softw. Eng.6
2023 Code Vulnerability Detection via Signal-Aware Learning
abstract
Machine Learning-based modeling of source code understanding tasks has been gaining popularity. Accompanying their rapid proliferation is an emerging scrutiny over the models’ reliability. Concerns have been raised regarding the models not actually learning task-relevant source code features, but fitting other correlated data. To improve model trustworthiness, in this work, we explore data-driven approaches for enhancing model signal awareness, i.e., learning the relevant signals in the input for making predictions. We do so by incorporating the notion of code complexity during model training, both (i) explicitly via curriculum learning, and (ii) implicitly by augmenting the training dataset with simplified signal-preserving programs. With our techniques, we achieve up to 4.8x improvement in signal awareness of vulnerability detection models. Using the notion of code complexity, we present a novel interpretation of the model learning behaviour from the perspective of the dataset. We use it to introspect model learning difficulties, and analyze the learning enhancements achieved with our approaches.
Sahil Suneja, Yufan Zhuang, Yunhui Zheng, Jim Laredo, Alessandro Morari, Udayan Khurana
EuroS&P4
2023 Follow the Successful Herd: Towards Explanations for Improved Use and Mental Models of Natural Language Systems
abstract
While natural language systems continue improving, they are still imperfect. If a user has a better understanding of how a system works, they may be able to better accomplish their goals even in imperfect systems. We explored whether explanations can support effective authoring of natural language utterances and how those explanations impact users’ mental models in the context of a natural language system that generates small programs. Through an online study (n=252), we compared two main types of explanations: 1) system-focused, which provide information about how the system processes utterances and matches terms to a knowledge base, and 2) social, which provide information about how other users have successfully interacted with the system. Our results indicate that providing social suggestions of terms to add to an utterance helped users to repair and generate correct flows more than system-focused explanations or social recommendations of words to modify. We also found that participants commonly understood some mechanisms of the natural language system, such as the matching of terms to a knowledge base, but they often lacked other critical knowledge, such as how the system handled structuring and ordering. Based on these findings, we make design recommendations for supporting interactions with and understanding of natural language systems.
Michelle Brachman, Hyo Jin Do, Casey Dugan, Arunima Chaudhary, James M. Johnson, Priyanshu Rai, Tathagata Chakraborti, Thomas Gschwind, Jim Laredo, Christoph Miksovic, Paolo Scotton, Kartik Talamadupula, Gegi Thomas
IUI10
2023 Incorporating Signal Awareness in Source Code Modeling: An Application to Vulnerability Detection
abstract
AI models of code have made significant progress over the past few years. However, many models are actually not learning task-relevant source code features. Instead, they often fit non-relevant but correlated data, leading to a lack of robustness and generalizability, and limiting the subsequent practical use of such models. In this work, we focus on improving the model quality throughsignal awareness, i.e., learning the relevant signals in the input for making predictions. We do so by leveraging the heterogeneity of code samples in terms of their signal-to-noise content. We perform an end-to-end exploration of model signal awareness, comprising: (i) uncovering the reliance of AI models of code on task-irrelevant signals, via prediction-preserving input minimization; (ii) improving models’ signal awareness by incorporating the notion of code complexity during model training, via curriculum learning; (iii) improving models’ signal awareness by generating simplified signal-preserving programs and augmenting them to the training dataset; and (iv) presenting a novel interpretation of the model learning behavior from the perspective of the dataset, using its code complexity distribution. We propose a new metric to measure model signal awareness, Signal-aware Recall, which captures how much of the model’s performance is attributable to task-relevant signal learning. Using a software vulnerability detection use-case, our model probing approach uncovers a significant lack of signal awareness in the models, across three different neural network architectures and three datasets. Signal-aware Recall is observed to be in the sub-50s for models with traditional Recall in the high 90s, suggesting that the models are presumably picking up a lot of noise or dataset nuances while learning their logic. With our code-complexity-aware model learning enhancement techniques, we are able to assist the models toward more task-relevant learning, recording up-to 4.8× improvement in model signal awareness. Finally, we employ our model learning introspection approach to uncover the aspects of source code where the model is facing difficulty, and we analyze how our learning enhancement techniques alleviate it.
Sahil Suneja, Yufan Zhuang, Yunhui Zheng, Jim Laredo, Alessandro Morari, Udayan Khurana
ACM Trans. Softw. Eng. Methodol.4
2022 A Goal-Driven Natural Language Interface for Creating Application Integration Workflows
abstract
Web applications and services are increasingly important in a distributed internet filled with diverse cloud services and applications, each of which enable the completion of narrowly defined tasks. Given the explosion in the scale and diversity of such services, their composition and integration for achieving complex user goals remains a challenging task for end-users and requires a lot of development effort when specified by hand. We present a demonstration of the Goal Oriented Flow Assistant (GOFA) system, which provides a natural language solution to generate workflows for application integration. Our tool is built on a three-step pipeline: it first uses Abstract Meaning Representation (AMR) to parse utterances; it then uses a knowledge graph to validate candidates; and finally uses an AI planner to compose the candidate flow. We provide a video demonstration of the deployed system as part of our submission.
Michelle Brachman, Christopher Bygrave, Tathagata Chakraborti, Arunima Chaudhary, Zhining Ding, Casey Dugan, Thomas Gschwind, James M. Johnson, Jim Laredo, Christoph Miksovic, Priyanshu Rai, Ramkumar Ramalingam, Paolo Scotton, Nagarjuna Surabathina, Kartik Talamadupula
AAAI10
2022 Varangian: A Git Bot for Augmented Static Analysis
abstract
The complexity and scale of modern software programs often lead to overlooked programming errors and security vulnerabilities. Developers often rely on automatic tools, like static analysis tools, to look for bugs and vulnerabilities. Static analysis tools are widely used because they can understand nontrivial program behaviors, scale to millions of lines of code, and detect subtle bugs. However, they are known to generate an excess of false alarms which hinder their utilization as it is counterproductive for developers to go through a long list of reported issues, only to find a few true positives. One of the ways proposed to suppress false positives is to use machine learning to identify them. However, training machine learning models requires good quality labeled datasets. For this purpose, we developed D2A [3], a differential analysis based approach that uses the commit history of a code repository to create a labeled dataset of Infer [2] static analysis output.
Saurabh Pujar, Yunhui Zheng, Luca Buratti, Burn L. Lewis, Alessandro Morari, Jim Laredo, Kevin Postlethwait, Christoph Görn
MSR6
2022 VELVET: a noVel Ensemble Learning approach to automatically locate VulnErable sTatements
abstract
Automatically locating vulnerable statements in source code is crucial to assure software security and alleviate developers' debugging efforts. This becomes even more important in today's software ecosystem, where vulnerable code can flow easily and unwittingly within and across software repositories like GitHub. Across such millions of lines of code, traditional static and dynamic approaches struggle to scale. Although existing machine-learning-based approaches look promising in such a setting, most work detects vulnerable code at a higher granularity – at the method or file level. Thus, developers still need to inspect a significant amount of code to locate the vulnerable statement(s) that need to be fixed. This paper presents Velvet, a novel ensemble learning approach to locate vulnerable statements. Our model combines graph-based and sequence-based neural networks to successfully capture the local and global context of a program graph and effectively understand code semantics and vulnerable patterns. To study Velvet's effectiveness, we use an off-the-shelf synthetic dataset and a recently published real-world dataset. In the static analysis setting, where vulnerable functions are not detected in advance, Velvet achieves 4.5× better performance than the baseline static analyzers on the real-world data. For the isolated vulnerability localization task, where we assume the vulnerability of a function is known while the specific vulnerable statement is unknown, we compare Velvet with several neural networks that also attend to local and global context of code. Velvet achieves 99.6% and 43.6% top-1 accuracy over synthetic data and real-world data, respectively, outperforming the baseline deep learning models by 5.3-29.0%.
Yangruibo Ding, Sahil Suneja, Yunhui Zheng, Jim Laredo, Alessandro Morari, Gail E. Kaiser, Baishakhi Ray
SANER4
2021 Towards Reliable AI for Source Code Understanding
abstract
Cloud maturity and popularity have resulted in Open source software (OSS) proliferation. And, in turn, managing OSS code quality has become critical in ensuring sustainable Cloud growth. On this front, AI modeling has gained popularity in source code understanding tasks, promoted by the ready availability of large open codebases. However, we have been observing certain peculiarities with these black-boxes, motivating a call for their reliability to be verified before offsetting traditional code analysis. In this work, we highlight and organize different reliability issues affecting AI-for-code into three stages of an AI pipeline- data collection, model training, and prediction analysis. We highlight the need for concerted efforts from the research community to ensure credibility, accountability, and traceability for AI-for-code. For each stage, we discuss unique opportunities afforded by the source code and software engineering setting to improve AI reliability.
Sahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim Laredo, Alessandro Morari
SoCC4
2021 Learning GraphQL Query Cost
abstract
GraphQL is a query language for APIs and a runtime for executing those queries, fetching the requested data from existing microservices, REST APIs, databases, or other sources. Its expressiveness and its flexibility have made it an attractive candidate for API providers in many industries, especially through the web. A major drawback to blindly servicing a client’s query in GraphQL is that the cost of a query can be unexpectedly large, creating computation and resource overload for the provider, and API rate-limit overages and infrastructure overload for the client. To mitigate these drawbacks, it is necessary to efficiently estimate the cost of a query before executing it. Estimating query cost is challenging, because GraphQL queries have a nested structure, GraphQL APIs follow different design conventions, and the underlying data sources are hidden. Estimates based on worst-case static query analysis have had limited success because they tend to grossly overestimate cost. We propose a machine-learning approach to efficiently and accurately estimate the query cost. We also demonstrate the power of this approach by testing it on query-response data from publicly available commercial APIs. Our framework is efficient and predicts query costs with high accuracy, consistently outperforming the static analysis by a large margin.
Georgios Mavroudeas, Guillaume Baudart, Alan Cha, Martin Hirzel, Jim Laredo, Malik Magdon-Ismail, Louis Mandel, Erik Wittern
ASE5
2021 Probing model signal-awareness via prediction-preserving input minimization
abstract
This work explores the signal awareness of AI models for source code understanding. Using a software vulnerability detection use case, we evaluate the models' ability to capture the correct vulnerability signals to produce their predictions. Our prediction-preserving input minimization (P2IM) approach systematically reduces the original source code to a minimal snippet which a model needs to maintain its prediction. The model's reliance on incorrect signals is then uncovered when the vulnerability in the original code is missing in the minimal snippet, both of which the model however predicts as being vulnerable. We measure the signal awareness of models using a new metric we propose -- Signal-aware Recall (SAR). We apply P2IM on three different neural network architectures across multiple datasets. The results show a sharp drop in the model's Recall from the high 90s to sub-60s with the new metric, highlighting that the models are presumably picking up a lot of noise or dataset nuances while learning their vulnerability detection logic. Although the drop in model performance may be perceived as an adversarial attack, but this isn't P2IM's objective. The idea is rather to uncover the signal-awareness of a black-box model in a data-driven manner via controlled queries. SAR's purpose is to measure the impact of task-agnostic model training, and not to suggest a shortcoming in the Recall metric. The expectation, in fact, is for SAR to match Recall in the ideal scenario where the model truly captures task-specific signals.
Sahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim Laredo, Alessandro Morari
ESEC/SIGSOFT FSE4
2020 A principled approach to GraphQL query cost analysis
abstract
The landscape of web APIs is evolving to meet new client requirements and to facilitate how providers fulfill them. A recent web API model is GraphQL, which is both a query language and a runtime. Using GraphQL, client queries express the data they want to retrieve or mutate, and servers respond with exactly those data or changes. GraphQL’s expressiveness is risky for service providers because clients can succinctly request stupendous amounts of data, and responding to overly complex queries can be costly or disrupt service availability. Recent empirical work has shown that many service providers are at risk. Using traditional API management methods is not sufficient, and practitioners lack principled means of estimating and measuring the cost of the GraphQL queries they receive. In this work, we present a linear-time GraphQL query analysis that can measure the cost of a query without executing it. Our approach can be applied in a separate API management layer and used with arbitrary GraphQL backends. In contrast to existing static approaches, our analysis supports common GraphQL conventions that affect query cost, and our analysis is provably correct based on our formal specification of GraphQL semantics. We demonstrate the potential of our approach using a novel GraphQL query-response corpus for two commercial GraphQL APIs. Our query analysis consistently obtains upper cost bounds, tight enough relative to the true response sizes to be actionable for service providers. In contrast, existing static GraphQL query analyses exhibit over-estimates and under-estimates because they fail to support GraphQL conventions.
Alan Cha, Erik Wittern, Guillaume Baudart, James C. Davis 0001, Louis Mandel, Jim Laredo
ESEC/SIGSOFT FSE6
2018 Generating GraphQL-Wrappers for REST(-like) APIs
Erik Wittern, Alan Cha, Jim Laredo
ICWE3
2017 Statically checking web API requests in JavaScript
abstract
Many JavaScript applications perform HTTP requests to web APIs, relying on the request URL, HTTP method, and request data to be constructed correctly by string operations. Traditional compile-time error checking, such as calling a non-existent method in Java, are not available for checking whether such requests comply with the requirements of a web API. In this paper, we propose an approach to statically check web API requests in JavaScript. Our approach first extracts a request's URL string, HTTP method, and the corresponding request data using an inter-procedural string analysis, and then checks whether the request conforms to given web API specifications. We evaluated our approach by checking whether web API requests in JavaScript files mined from GitHub are consistent or inconsistent with publicly available API specifications. From the 6575 requests in scope, our approach determined whether the request's URL and HTTP method was consistent or inconsistent with web API specifications with a precision of 96.0%. Our approach also correctly determined whether extracted request data was consistent or inconsistent with the data requirements with a precision of 87.9% for payload data and 99.9% for query data. In a systematic analysis of the inconsistent cases, we found that many of them were due to errors in the client code. The here proposed checker can be integrated with code editors or with continuous integration tools to warn programmers about code containing potentially erroneous requests.
Erik Wittern, Annie T. T. Ying, Yunhui Zheng, Julian Dolby, Jim Laredo
ICSE5
2014 A Graph-Based Data Model for API Ecosystem Insights
abstract
APIs are increasingly important for companies to enable partners and consumers to access their services and resources. API ecosystems deal with related challenges like publication, promotion and provision of APIs by providers and identification, selection and consumption of APIs by consumers. To address these challenges, to match consumers with relevant APIs, and to support API providers and thus ultimately the ecosystem to evolve, API ecosystems rely on information about APIs, their usage and characteristics, and the social environment around them. We present an extensible, graph-based data model to capture the entities in an API ecosystem and their relations. The data model includes temporal information to capture the evolution of API ecosystems. Analysis operations on top of the data model provide insights for consumers, providers and the ecosystem provider to address the introduced challenges. We present a system implementing the conceptualized data model. We integrate this system with an API ecosystem used in the context of a hackathon event to continuously collect data. We furthermore show the data model's capabilities to represent a well-known dataset about ProgrammableWeb and to drive analysis operations on both datasets.
Erik Wittern, Jim Laredo, Maja Vukovic, Vinod Muthusamy, Aleksander Slominski
ICWS2
2013 Assessing service deployment readiness using enterprise crowdsourcing
Maja Vukovic, Jim Laredo, Yaoping Ruan, Milton Hernandez, Sriram Rajagopal
IM2
2012 Towards cloud services marketplaces
Rahul P. Akolkar, Tom Chefalas, Jim Laredo, Chang-Shing Perng, Anca Sailer, Frank Schaffa, Ignacio Silva-Lepe, Tao Tao 0006
CNSM3
2012 Privileged identity management in enterprise service-hosting environments
abstract
IAM needs will only grow as devices, servers, and end points continue to increase . Current schemes are not sustainable as the number of IDs will explode. Environment is heterogeneous, and constantly adding new systems including Cloud. Our solution offers a platform where a user gets an individual user ID on a system - but only if they need it, when they need it, for only as long as they need it . Reusable ID scheme reduces the number of IDs in the system yielding cost savings on lifecycle management activities, improved security compliance . A compliance readiness platform can be enabled to prevent, flag, or monitor questionable access in or near real-time . Provide easily accessible logs to prove compliance policies.
Kumar Bhaskaran, Milton Hernandez, Jim Laredo, Laura Luan, Yaoping Ruan, Maja Vukovic, Paul Driscoll, Alan Skinner, Girish Verma, Prema Vivekanandan, Leanne Chen, Gregory Gaskill
NOMS3
2012 Integrated user activity monitoring for regulatory services
abstract
Regulations such as FFIEC [5] and HIPAA [6] require activities of system administration to be captured and reviewed regularly. In IT service delivery environment, system maintenance activities are usually performed by the service provider whose system administrators access customer environment based on problem and change ticket being assigned.
Mattias Marder, Kumar Bhaskaran, Milton Hernandez, Jim Laredo, Daniela Rosu 0001, Yaoping Ruan, Paul Driscoll, Alan Skinner
NOMS4
2012 The Future of Service Marketplaces in the Cloud
abstract
For as long as there have been services there has been a desire to have a convenient medium to expose and discover service offerings. Since early on, various efforts have attempted various approaches at the exchange of computational services, prompting the question of whether there is a market for Web services. We believe that a services marketplace should fulfill the promise of an electronic emporium where third party service providers are able to offer their services in a ubiquitous ecosystem, and where service consumers are able to acquire service solutions that are tailored to their requirements. This paper explores the landscape of cloud services marketplaces, where we are, what enablers are needed to realize the vision, and it presents a prospective architecture to that end.
Rahul P. Akolkar, Tom Chefalas, Jim Laredo, Chang-Shing Perng, Anca Sailer, Frank Schaffa, Ignacio Silva-Lepe, Tao Tao 0006
SERVICES3
2011 An Assessment of Intrinsic and Extrinsic Motivation on Task Performance in Crowdsourcing Markets
Jakob Rogstadius, Vassilis Kostakos, Aniket Kittur, Boris Smus, Jim Laredo, Maja Vukovic
ICWSM5
2010 Challenges and Experiences in Deploying Enterprise Crowdsourcing Service
Maja Vukovic, Jim Laredo, Sriram Rajagopal
ICWE2
2010 Server Hunt: Using Enterprise Social Networks for Knowledge Discovery in IT Inventory Management
abstract
Locating IT Inventory Management information is a challenging task, as the knowledge gets transferred among employees that move within or leave the context of a large organization. Information that relates to IT inventory is hidden in the knowledge of individual team members. This fact is not reflected in organizational expertise repositories and therefore locating those employees becomes a cumbersome manual process, if not intractable. In this paper, we present an expert discovery service that leverages the professional social network of an employee, who was previously known to hold the desired inventory information but is no longer available. Evaluation results suggest that this method reconstructs the desired information more than 80% of the time, as per our experiment involving 50 cases. We demonstrate how a carefully designed crowdsourcing approach can effectively extract the targeted information from the employee's professional social network and discuss its limitations.
Polychronis Ypodimatopoulos, Maja Vukovic, Jim Laredo, Sriram Rajagopal
SERVICES3
2009 Distributed Cross-Domain Change Management
abstract
Distributed systems increasingly span organizational boundaries and, with this, system and service management domains. Web services are the primary means of exposing services to clients, be it in electronic commerce, Software-as-a-Service (SaaS) or on cloud platforms and are being used and integrated with customer-managed applications as well as in complex mashups. Maturing cross-domain relationships and an increase in loose coupling and ad-hocness makes managing configuration changes, e.g., changes in interfaces or endpoints, increasingly relevant. Traditional service management processes within organizations, in particular change management, relies on a central configuration management database (CMDB) to assess the impact a change has on other components of the system. However, this approach does not work in a cross-domain environment, due to the lack of a central CMDB, centralized management processes, and knowledge by service providers which clients depends on their respective services. This paper proposes the Change 2.0 approach to cross-domain change management based on an inversion of responsibility for impact assessment and the facilitation of cross-domain service process integration. We present the requirements imposed by cross-domain change management, the Change 2.0 architecture, and a brief evaluation of its benefits.
Bruno Wassermann, Heiko Ludwig, Jim Laredo, Kamal Bhattacharya, Liliana Pasquale
ICWS3
2009 REST-based management of loosely coupled services
abstract
Applications increasingly make use of the distributed platform that the World Wide Web provides - be it as a Software-as-a-Service such as salesforce.com, an application infrastructure such as facebook.com, or a computing infrastructure such as a "cloud". A common characteristic of applications of this kind is that they are deployed on infrastructure or make use of components that reside in different management domains. Current service management approaches and systems, however, often rely on a centrally managed configuration management database (CMDB), which is the basis for centrally orchestrated service management processes, in particular change management and incident management. The distribution of management responsibility of WWW based applications requires a decentralized approach to service management. This paper proposes an approach of decentralized service management based on distributed configuration management and service process co-ordination, making use RESTful access to configuration information and ATOM-based distribution of updates as a novel foundation for service management processes.
Heiko Ludwig, Jim Laredo, Kamal Bhattacharya, Liliana Pasquale, Bruno Wassermann
WWW2
2008 Continuous Improvement through Iterative Development in a Multi-Geography
abstract
With a very short time frame in mind we were commissioned to build a service solution that encompassed over 430 requirements. Given expertise, skills, time and budget constraints we had to search for resources around the world and assembled a team of more than 40 people in 7 locations across 6 different time zones. We chose state of the art architecture and development paradigms such as Service Oriented Architecture (SOA) and iterative development to facilitate the service capabilities between components and teams and to continuously refine our approach from iteration to iteration. In this paper we identify some of the challenges we faced and describe how we addressed them in subsequent iterations. The aggregate of our improvements constitutes a set of best practices that we recommend for future engagements of this type and suggestions for new tooling to support these activities.
Jim Laredo, Ravi Ranjan
ICGSE1