Cristina V. Lopes

dblp:l/CristinaVideiraLopes · also Cristina Videira Lopes · DBLP profile ↗
← Back
75ranked-venue papers
8as first author
4since 2021 · last 2024
0000-0003-0551-3908ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 49 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12Human-computer interaction and ubiquitous computing · 10 · 1 first-authorArtificial intelligence and machine learning · 7Graphics, computer vision, multimedia, augmented reality and games · 6Computer networks · 5 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2024 SourcererJBF: A Java Build Framework For Large-Scale Compilation
abstract
Researchers and tool developers working on dynamic analysis, software testing, automated program repair, verification, and validation, need large compiled, compilable, and executable code corpora to test their ideas. The publicly available corpora are relatively small, and/or non-compilable, and/or non-executable. Developing a compiled code corpus is a laborious activity demanding significant manual effort and human intervention. To facilitate large-scale program analysis research, we develop SourcererJBF , a J ava B uild F ramework that can automatically build a large Java code corpus without project-specific instructions and human intervention. To generate a compiled code corpus, SourcererJBF creates an offline knowledge base by collecting external dependencies from the project directories and existing build scripts (if available). It constructs indices of those collected external dependencies that enable a fast search for resolving dependencies during the project compilation. As the output of the large-scale compilation, it produces JAigantic, a compilable Java corpus containing compiled projects, their bytecode, dependencies, normalized build script, and build command. We evaluated SourcererJBF’s effectiveness, correctness, performance, and scalability in a large collection of Java projects. Our experimental results demonstrate that SourcererJBF is significantly effective and scalable in building large Java code corpus. Besides, it substantiates reasonable performance and correctness similar to projects’ existing build systems.
Md Rakib Hossain Misu, Rohan Achar, Cristina V. Lopes
ACM Trans. Softw. Eng. Methodol.3
2023 Cloning and Beyond: A Quantum Solution to Duplicate Code
abstract
Quantum computers are becoming a reality. The advantage of quantum computing is that it has the potential to solve computationally complex problems in a fixed amount time, independent of the size of the problem. However, the kinds of problems for which these computers are a good fit, and the ways to express those problems, are substantially different from the kinds of problems and expressions used in classical computing. Quantum annealers, in particular, are currently the most promising and available quantum computing devices in the short term. However, they are also the most foreign compared to classical programs, as they require a different kind of computational thinking. In order to ease the transition into this new world of quantum computing, we present a novel quantum approach to a well-known software problem: code clone detection. We express code clone detection as a subgraph isomorphism problem that is mapped into a quadratic optimization problem, and solve it using a DWave quantum annealing computer. We developed a quantum annealing algorithm that compares Abstract Syntax Trees (AST) and reports an energy value that indicates how similar they are.
Samyak Jhaveri, Alberto Krone-Martins, Cristina V. Lopes
Onward!3
2021 Public Software Development Activity During the Pandemic
abstract
Background The emergence of the COVID-19 pandemic has impacted all human activity, including software development. Early reports seem to indicate that the pandemic may have had a negative effect on software developers, socially and personally, but that their software development productivity may not have been negatively impacted. Aims: Early reports about the effects of the pandemic on software development focused on software developers' well-being and on their productivity as employees. We are interested in a different aspect of software development: the developers' public contributions, as seen in GitHub and Stack Overflow activities. Did the pandemic affect the developers' public contributions and, of so, in what way? Method: Considering the data from between 2017 and till 2020, we study the trends within GitHub's push, create, pull request, and release events, and within Stack Overflow's new users, posts, votes, and comments. We performed linear regressions, correlation analyses, outlier analyses, hypothesis testing, and we also contacted individual developers in order to gather qualitative insights about their unusual public contributions. Results: Our study shows that within GitHub and Stack Overflow, the onset of the pandemic (March/April 2020) is reflected in a set of outliers in developers' contributions that point to an increase in activity. The distributions of contributions during the entire year of 2020 were, in some aspects, different, but, in other aspects, similar from the recent past. Additionally, we found one noticeably disrupted pattern of contribution in Stack Overflow, namely the ratio Questions/Answers, which was much higher in 2020 than before. Testimonials from the developers we contacted were mixed: while some developers reported that their increase in activity was due to the pandemic, others reported that it was not. Conclusion: In Github, there was a noticeable increase in public software development activity in 2020, as well as more abrupt changes in daily activities; in Stack Overflow, there was a noticeable increase in new users and new questions at the onset of the pandemic, and in the ratio of Questions/Answers during 2020. The results may be attributed to the pandemic, but other factors could have come into play.
Vanessa Klotzman, Farima Farmahinifarahani, Cristina V. Lopes
ESEM3
2021 D-REX: Static Detection of Relevant Runtime Exceptions with Location Aware Transformer
abstract
Runtime exceptions are inevitable parts of software systems. While developers often write exception handling code to avoid the severe outcomes of these exceptions, such code is most effective if accompanied by accurate runtime exception types. Predicting the runtime exceptions that may occur in a program, however, is difficult as the situations that lead to these exceptions are complex. We propose D-REX (Deep Runtime EXception detector), as an approach for predicting runtime exceptions of Java methods based on the static properties of code.The core of D-REX is a machine learning model that leverages the representation learning ability of neural networks to infer a set of signals from code to predict the related runtime exception types. This model, which we call Location Aware Transformer, adapts a state-of-the-art language model, Transformer, to provide accurate predictions for the exception types, as well as interpretable recommendations for the exception prone elements of code. We curate a benchmark dataset of 200,000 Java projects from GitHub to train and evaluate D-REX. Experiments demonstrate that D-REX predicts runtime exception types with 81% of Top 1 accuracy, outperforming multiple non-Transformer baselines by a margin of at least 12%. Furthermore, it can predict the exception prone elements of code with 75% Top 1 precision.
Farima Farmahinifarahani, Yadong Lu, Vaibhav Saini, Pierre Baldi, Cristina V. Lopes
SCAM5
2019 Multi-Winner Contests for Strategic Diffusion in Social Networks
abstract
Strategic diffusion encourages participants to take active roles in promoting stakeholders’ agendas by rewarding successful referrals. As social media continues to transform the way people communicate, strategic diffusion has become a powerful tool for stakeholders to influence people’s decisions or behaviors for desired objectives. Existing reward mechanisms for strategic diffusion are usually either vulnerable to falsename attacks or not individually rational for participants that have made successful referrals. Here, we introduce a novel multi-winner contests (MWC) mechanism for strategic diffusion in social networks. The MWC mechanism satisfies several desirable properties, including false-name-proofness, individual rationality, budget constraint, monotonicity, and subgraph constraint. Numerical experiments on four real-world social network datasets demonstrate that stakeholders can significantly boost participants’ aggregated efforts with proper design of competitions. Our work sheds light on how to design manipulation-resistant mechanisms with appropriate contests.
Wen Shen 0001, Yang Feng 0003, Cristina V. Lopes
AAAI3
2019 Analyzing and supporting adaptation of online code examples
abstract
Developers often resort to online Q&A forums such as Stack Overflow (SO) for filling their programming needs. Although code examples on those forums are good starting points, they are often incomplete and inadequate for developers' local program contexts; adaptation of those examples is necessary to integrate them to production code. As a consequence, the process of adapting online code examples is done over and over again, by multiple developers independently. Our work extensively studies these adaptations and variations, serving as the basis for a tool that helps integrate these online code examples in a target context in an interactive manner. We perform a large-scale empirical study about the nature and extent of adaptations and variations of SO snippets. We construct a comprehensive dataset linking SO posts to GitHub counterparts based on clone detection, time stamp analysis, and explicit URL references. We then qualitatively inspect 400 SO examples and their GitHub counterparts and develop a taxonomy of 24 adaptation types. Using this taxonomy, we build an automated adaptation analysis technique on top of GumTree to classify the entire dataset into these types. We build a Chrome extension called ExampleStack that automatically lifts an adaptation-aware template from each SO example and its GitHub counterparts to identify hot spots where most changes happen. A user study with sixteen programmers shows that seeing the commonalities and variations in similar GitHub counterparts increases their confidence about the given SO example, and helps them grasp a more comprehensive view about how to reuse the example differently and avoid common pitfalls.
Tianyi Zhang 0001, Di Yang 0001, Cristina V. Lopes, Miryung Kim
ICSE3
2019 Towards automating precision studies of clone detectors
abstract
Current research in clone detection suffers from poor ecosystems for evaluating precision of clone detection tools. Corpora of labeled clones are scarce and incomplete, making evaluation labor intensive and idiosyncratic, and limiting intertool comparison. Precision-assessment tools are simply lacking. We present a semiautomated approach to facilitate precision studies of clone detection tools. The approach merges automatic mechanisms of clone classification with manual validation of clone pairs. We demonstrate that the proposed automatic approach has a very high precision and it significantly reduces the number of clone pairs that need human validation during precision experiments. Moreover, we aggregate the individual effort of multiple teams into a single evolving dataset of labeled clone pairs, creating an important asset for software clone research.
Vaibhav Saini, Farima Farmahinifarahani, Yadong Lu, Di Yang 0001, Pedro Martins 0001, Hitesh Sajnani, Pierre Baldi, Cristina V. Lopes
ICSE8
2019 Virtual Conferences
abstract
For the past 40 years, research communities have embraced a culture that relies heavily on physical meetings of people from around the world: we present our most important work in conferences, we meet our peers in conferences, and we even make life-long friends in conferences. Also at the same time, a broad scientific consensus has emerged that warns that human emissions of greenhouse gases are warming the earth. For many of us, travel to conferences may be a substantial or even dominant part of our individual contribution to climate change. A single round-trip flight from Paris to New Orleans emits the equivalent of about 2.5 tons of carbon dioxide (CO2e) per passenger, which is a significant fraction of the total yearly emissions for an average resident of the US or Europe. Moreover, these emissions have no near-term technological fix, since jet fuel is difficult to replace with renewable energy sources. In this talk, I want to first raise awareness of the conundrum we are in by relying so heavily in air travel for our work. I will present some of the possible solutions that go from adopting small, incremental changes to radical ones. The talk focuses one of the radical alternatives: virtual conferences. The technology for them is almost here and, for some time, I have been part of one community that organizes an annual conference in a virtual environment. Virtual conferences present many interesting challenges, some of them technological in nature, others that go beyond technology. Creating truly immersive conference experiences that make us feel "there" requires attention to personal and social experiences at physical conferences. Those experiences need to be recreated from the ground up in virtual spaces. But in that process, they can also be rethought to become experiences not possible in real life.
Cristina V. Lopes
UIST1
2019 On Precision of Code Clone Detection Tools
abstract
Precision and recall are the main metrics used to measure the correctness of clone detectors. These metrics require the existence of labeled datasets containing the ground truth - samples of clone and non-clone pairs. For source code clone detectors, in particular, there are some techniques, as well as a concrete framework, for automatically evaluating recall, down to different types of clones. However, evaluating precision is still challenging, because of the intensive and specialized manual effort required to accomplish the task. Moreover, when precision is reported, it is typically done over all types of clones, making it hard to assess the strengths and weaknesses of the corresponding clone detectors. This paper presents systematic experiments to evaluate precision of eight code clone detection tools. Three judges independently reviewed 12,800 clone pairs to compute the undifferentiated and type-based precision of these tools. Besides providing a useful baseline for future research in code clone detection, another contribution of our work is to unveil important considerations to take into account when doing precision measurements and reporting the results. Specifically, our work shows that the reported precision of these tools leads to significantly different conclusions and insights about the tools when different types of clones are taken into account. It also stresses, once again, the importance of reporting inter-rater agreement.
Farima Farmahinifarahani, Vaibhav Saini, Di Yang 0001, Hitesh Sajnani, Cristina V. Lopes
SANER5
2018 50K-C: a dataset of compilable, and compiled, Java projects
abstract
We provide a repository of 50,000 compilable Java projects. Each project in this dataset comes with references to all the dependencies required to compile it, the resulting bytecode, and the scripts with which the projects were built.
Pedro Martins 0001, Rohan Achar, Cristina V. Lopes
MSR3
2018 Oreo: detection of clones in the twilight zone
abstract
Source code clones are categorized into four types of increasing difficulty of detection, ranging from purely textual (Type-1) to purely semantic (Type-4). Most clone detectors reported in the literature work well up to Type-3, which accounts for syntactic differences. In between Type-3 and Type-4, however, there lies a spectrum of clones that, although still exhibiting some syntactic similarities, are extremely hard to detect – the Twilight Zone. Most clone detectors reported in the literature fail to operate in this zone. We present Oreo, a novel approach to source code clone detection that not only detects Type-1 to Type-3 clones accurately, but is also capable of detecting harder-to-detect clones in the Twilight Zone. Oreo is built using a combination of machine learning, information retrieval, and software metrics. We evaluate the recall of Oreo on BigCloneBench, and perform manual evaluation for precision. Oreo has both high recall and precision. More importantly, it pushes the boundary in detection of clones with moderate to weak syntactic similarity in a scalable manner
Vaibhav Saini, Farima Farmahinifarahani, Yadong Lu, Pierre Baldi, Cristina V. Lopes
ESEC/SIGSOFT FSE5
2018 Cloned and non-cloned Java methods: a comparative study
Vaibhav Saini, Hitesh Sajnani, Cristina V. Lopes
Empir. Softw. Eng.3
2017 Stack overflow in github: any snippets there?
abstract
When programmers look for how to achieve certain programming tasks, Stack Overflow is a popular destination in search engine results. Over the years, Stack Overflow has accumulated an impressive knowledge base of snippets of code that are amply documented. We are interested in studying how programmers use these snippets of code in their projects. Can we find Stack Overflow snippets in real projects? When snippets are used, is this copy literal or does it suffer adaptations? And are these adaptations specializations required by the idiosyncrasies of the target artifact, or are they motivated by specific requirements of the programmer? The large-scale study presented on this paper analyzes 909k non-fork Python projects hosted on Github, which contain 290M function definitions, and 1.9M Python snippets captured in Stack Overflow. Results are presented as quantitative analysis of block-level code cloning intra and inter Stack Overflow and GitHub, and as an analysis of programming behaviors through the qualitative analysis of our findings.
Di Yang 0001, Pedro Martins 0001, Vaibhav Saini, Cristina V. Lopes
MSR4
2017 An Exploratory Study of Functional Redundancy in Code Repositories
abstract
In large code repositories, the probability of functions to repeat across projects is high. This type of functional redundancy (FR) is desirable for recent code reuse and repair approaches. Yet, FR is hard to measure because it is closely related to program equivalence, which is an undecidable problem. This is one of the reasons most studies that investigate redundancy focus on syntactic rather than semantic replication (e.g., cloning). In this paper we evaluate the extent of FR in a code repository with 68 Java projects taken randomly from SourceForge. Our technique approximates function similarity by first searching for methods that possess similar interfaces (return type, name, and parameter types). We then execute these methods to verify which candidate pairs have matching outputs for a given sample of inputs. Some recent studies have also focused on this type of semantic replication, but our detection approach is generally cheaper and more precise, because it focuses on methods and uses interfaces to reduce the search space. Although our scope is restricted to static methods, which makes our results conservative, our findings are promising. In particular, we found 984 pairs of redundant methods, and 28 out of the 68 (41.17%) projects in the repository presented redundancy. Moreover, the majority of redundant methods for which we had access to the source code did not refer to textual clones (only one redundant method pair referred to replicated code). Our study also indicates that the proposed redundancy detection approach has high precision and is generally inexpensive (only four executions were required per method to attain 100% precision).
Marcelo Suzuki, Adriano Carvalho de Paula, Eduardo Guerra 0001, Cristina V. Lopes, Otávio Augusto Lazzarini Lemos
SCAM4
2017 DéjàVu: a map of code duplicates on GitHub
abstract
Previous studies have shown that there is a non-trivial amount of duplication in source code. This paper analyzes a corpus of 4.5 million non-fork projects hosted on GitHub representing over 428 million files written in Java, C++, Python, and JavaScript. We found that this corpus has a mere 85 million unique files. In other words, 70% of the code on GitHub consists of clones of previously created files. There is considerable variation between language ecosystems. JavaScript has the highest rate of file duplication, only 6% of the files are distinct. Java, on the other hand, has the least duplication, 60% of files are distinct. Lastly, a project-level analysis shows that between 9% and 31% of the projects contain at least 80% of files that can be found elsewhere. These rates of duplication have implications for systems built on open source software as well as for researchers interested in analyzing large code bases. As a concrete artifact of this study, we have created DéjàVu, a publicly available map of code duplicates in GitHub repositories.
Cristina V. Lopes, Petr Maj, Pedro Martins 0001, Vaibhav Saini, Di Yang 0001, Jakub Zitny, Hitesh Sajnani, Jan Vitek
Proc. ACM Program. Lang.1
2016 Simulating Cities: A Software Engineering Perspective
Cristina V. Lopes
ESOP1
2016 SourcererCC: scaling code clone detection to big-code
abstract
Despite a decade of active research, there has been a marked lack in clone detection techniques that scale to large repositories for detecting near-miss clones. In this paper, we present a token-based clone detector, SourcererCC, that can detect both exact and near-miss clones from large inter-project repositories using a standard workstation. It exploits an optimized inverted-index to quickly query the potential clones of a given code block. Filtering heuristics based on token ordering are used to significantly reduce the size of the index, the number of code-block comparisons needed to detect the clones, as well as the number of required token-comparisons needed to judge a potential clone. We evaluate the scalability, execution time, recall and precision of SourcererCC, and compare it to four publicly available and state-of-the-art tools. To measure recall, we use two recent benchmarks: (1) a big benchmark of real clones, BigCloneBench, and (2) a Mutation/Injection-based framework of thousands of fine-grained artificial clones. We find SourcererCC has both high recall and precision, and is able to scale to a large inter-project repository (25K projects, 250MLOC) using a standard workstation.
Hitesh Sajnani, Vaibhav Saini, Jeffrey Svajlenko, Chanchal Kumar Roy, Cristina V. Lopes
ICSE5
2016 Comparing Quality Metrics for Cloned and Non Cloned Java Methods: A Large Scale Empirical Study
abstract
In this paper, we conduct a large scale statistical study to explore if there exists any difference between the quality of cloned methods and non cloned methods. The dataset consists of 4,421 open source Java projects containing 644,830 cloned and 842,052 non cloned methods. The study uses 27 software metrics as a proxy for quality, spanning across complexity, modularity, and documentation (code-comments) categories. We did not find any statistically significant difference (p0.1) between the quality of cloned and non cloned methods for most of the metrics, except for 3 metrics. We, however, found that the cloned methods are on an average 20% smaller than the non cloned methods.
Vaibhav Saini, Hitesh Sajnani, Cristina V. Lopes
ICSME3
2016 An Online Mechanism for Ridesharing in Autonomous Mobility-on-Demand Systems
Wen Shen 0001, Cristina V. Lopes, Jacob W. Crandall
IJCAI2
2016 From query to usable code: an analysis of stack overflow code snippets
abstract
Enriched by natural language texts, Stack Overflow code snippets are an invaluable code-centric knowledge base of small units of source code. Besides being useful for software developers, these annotated snippets can potentially serve as the basis for automated tools that provide working code solutions to specific natural language queries.
Di Yang 0001, Aftab Hussain 0001, Cristina V. Lopes
MSR3
2016 Collective Intelligence for Smarter API Recommendations in Python
abstract
Software developers use Application Programming Interfaces (APIs) of libraries and frameworks extensively while writing programs. In this context, the recommendations provided in code completion pop-ups help developers choose the desired methods. The candidate lists recommended by these tools, however, tend to be large, ordered alphabetically and sometimes even incomplete. A fair amount of work has been done recently to improve the relevance of these code completion results, especially for statically typed languages like Java. However, these proposed techniques rely on the static type of the object and are therefore inapplicable for a dynamically typed language like Python. In this paper, we present PyReco, an intelligent code completion system for Python which uses the mined API usages from open source repositories to order the results based on relevance rather than the conventional alphabetic order. To recommend suggestions that are relevant for a working context, a nearest neighbor classifier is used to identify the best matching usage among all the extracted usage patterns. To evaluate the effectiveness of our system, the code completion queries are automatically extracted from projects and tested quantitatively using a ten-fold cross validation technique. The evaluation shows that our approach outperforms the alphabetically ordered API recommendation systems in recommending APIs for standard, as well as, third-party libraries.
Andrea Renika D'Souza, Di Yang 0001, Cristina V. Lopes
SCAM3
2016 An Exploratory Study of Interface Redundancy in Code Repositories
abstract
An important property of software repositories is their level of cross-project redundancy. For instance, much has been done to assess how much code cloning happens across software corpora. In this paper we study a much less targeted type of replication: Interface Redundancy (IR). IR refers to the level of repetition of whole method interfaces - return type, method name, and parameters types - across a code corpus. Such type of redundancy is important because if two non-trivial methods ever share the same interface it is very likely that they implement analogous functions, even though their code, structure, or vocabulary might be diverse. A certain level of IR is a requirement for approaches that rely on the recurrence of interfaces to fulfill a given task (e.g., interface-driven code search - IDCS). In this paper we report on an experiment to measure IR in a large-scale Java repository. Our target corpus contains more than 380,000 methods from 99 Java projects extracted randomly from an open source repository. Results are promising as they show that the chances of an interface from a non-trivial method to repeat itself across a large repository is around 25% (i.e., approximately 1/4 of such interfaces are redundant). Also, more than 80% of the target projects contained IR (with the average percentage of redundant interfaces for these projects being above 30%). As additional analyses we investigated the distribution of the different types of redundant interfaces (e.g., intra-vs inter-project), characterized the redundant interfaces and show that such a knowledge can help improve IDCS, and provided evidence that only a very small part of IR refers to method cloning (around 0.002%).
Adriano Carvalho de Paula, Eduardo Guerra 0001, Cristina V. Lopes, Hitesh Sajnani, Otávio Augusto Lazzarini Lemos
SCAM3
2016 On designing and testing distributed virtual environments
abstract
Summary Distributed real‐time (DRT) systems are among the most complex software systems to design, test, maintain, and evolve. The existence of components distributed over a network often conflicts with real‐time requirements, leading to design strategies that depend on domain‐specific and even application‐specific knowledge. Distributed virtual environment (DVE) systems are DRT systems that connect multiple users instantly with each other and with a shared virtual space over a network. DVE systems deviate from traditional DRT systemsx in the importance of the quality of the end user experience. We present an analysis of important, but challenging, issues in the design, testing, and evaluation of DVE systems through the lens of experiments with a concrete DVE, OpenSimulator. We frame our observations within six dimensions of well‐known design concerns: correctness, fault tolerance/prevention, scalability, time sensitivity, consistency, and overhead of distribution. Furthermore, we place our experimental work in a broader historical context, showing that these challenges are intrinsic to DVEs and suggesting lines of future research. Copyright © 2016 John Wiley & Sons, Ltd.
Arthur Valadares, Eugenia Gabrielova, Cristina V. Lopes
Concurr. Comput. Pract. Exp.3
2015 Gate Me If You Can: The Impact of Gating Mechanics on Retention and Revenues in Jelly Splash
Thomas Debeauvais, Cristina V. Lopes
FDG2
2015 How scale affects structure in Java programs
abstract
Many internal software metrics and external quality attributes of Java programs correlate strongly with program size. This knowledge has been used pervasively in quantitative studies of software through practices such as normalization on size metrics. This paper reports size-related super- and sublinear effects that have not been known before. Findings obtained on a very large collection of Java programs -- 30,911 projects hosted at Google Code as of Summer 2011 -- unveils how certain characteristics of programs vary disproportionately with program size, sometimes even non-monotonically. Many of the specific parameters of nonlinear relations are reported. This result gives further insights for the differences of ``programming in the small'' vs. ``programming in the large.'' The reported findings carry important consequences for OO software metrics, and software research in general: metrics that have been known to correlate with size can now be properly normalized so that all the information that is left in them is size-independent.
Cristina V. Lopes, Joel Ossher
OOPSLA1
2015 Managing Autonomous Mobility on Demand Systems for Better Passenger Experience
Wen Shen 0001, Cristina V. Lopes
PRIMA2
2015 Can the use of types and query expansion help improve large-scale code search?
abstract
With the open source code movement, code search with the intent of reuse has become increasingly popular. So much so that researchers have been calling it the new facet of software reuse. Although code search differs from general-purpose document search in essential ways, most tools still rely mainly on keywords matched against source code text. Recently, researchers have proposed more sophisticated ways to perform code search, such as including interface definitions in the queries (e.g., return and parameter types of the desired function, along with keywords; called here Interface-Driven Code Search - IDCS). However, to the best of our knowledge, there are few empirical studies that compare traditional keyword-based code search (KBCS) with more advanced approaches such as IDCS. In this paper we describe an experiment that compares the effectiveness of KBCS with IDCS in the task of large-scale code search of auxiliary functions implemented in Java. We also measure the impact of query expansion based on types and WordNet on both approaches. Our experiment involved 36 subjects that produced real-world queries for 16 different auxiliary functions and a repository with more than 2,000,000 Java methods. Results show that the use of types can improve recall and the number of relevant functions returned (#RFR) when combined with query expansion (~30% improvement in recall, and ~43% improvement in #RFR). However, a more detailed analysis suggests that in some situations it is best to use keywords only, in particular when these are sufficient to semantically define the desired function.
Otávio Augusto Lazzarini Lemos, Adriano Carvalho de Paula, Hitesh Sajnani, Cristina V. Lopes
SCAM4
2015 A parallel and efficient approach to large scale clone detection
abstract
Abstract We propose a new token‐based approach for large ‐scale code clone detection, which is based on a filtering heuristic that reduces the number of token comparisons when the two code blocks are compared. We also present a MapReduce based parallel algorithm that uses the filtering heuristic and scales to thousands of projects. The filtering heuristic is generic and can also be used in conjunction with other token‐based approaches. In that context, we demonstrate how it can increase the retrieval speed and decrease the memory usage of the index‐based approaches. In our experiments on 36 open source Java projects, we found that: (i) filtering reduces token comparisons by a factor of 10, and thus increasing the speed of clone detection by a factor of 1.5; (ii) the speed‐up and scale‐up of the parallel approach using filtering is near‐linear on a cluster of 2–32 nodes for 150–2800 projects; and (iii) filtering decreases the memory usage of index‐based approach by half and the search time by a factor of 5. Copyright © 2015 John Wiley & Sons, Ltd.
Hitesh Sajnani, Vaibhav Saini, Cristina V. Lopes
J. Softw. Evol. Process.3
2014 A Framework for Designing and Evaluating Distributed Real-Time Applications
abstract
Distributed Real-Time (DRT) systems are among the most complex software systems in the field of Computer Science. The contradictory nature of distributed computing and real-time requirements forces prioritization that is dependent on domain-specific knowledge. This combination of distributed and real-time system requirements generates a new set of properties, and requires a new conceptual framework. In this paper, we present a conceptual framework for developing DRT systems. We divide our framework in 3 phases: example applications, design properties, and evaluation. We demonstrate our framework properties with our own Distributed Virtual Environment architecture and other 2 examples: Google's Cluster Architecture and collision avoidance systems. Finally we discuss evaluation of DRT systems through the lens of our own architecture, and conclude with general lessons to apply to other DRT systems. Our conceptual framework enables developers of DRT systems to foresee and address issues early in the design. Additionally, we expose general lessons and good practices of evaluation for other DRT systems.
Arthur Valadares, Cristina V. Lopes
DS-RT2
2014 Retention and progression: Seven months in World of Warcraft
Thomas Debeauvais, Cristina V. Lopes, Nick Yee, Nicolas Ducheneaut
FDG2
2014 Is Popularity a Measure of Quality? An Analysis of Maven Components
abstract
One of the perceived values of open source software is the idea that many eyes can increase code quality and reduce the amount of bugs. This perception, however, has been questioned by some due the lack of supporting evidence. This paper presents an empirical analysis focusing on the relationship between the utilization of open source components and their engineering quality. In this study, we determine the popularity of 2,406 Maven components by calculating their usage across 55,191 open source Java projects. As a proxy of code quality for a component, we calculate (i) its defect density using the set of bug patterns reported by Find Bugs, and (ii) 9 popular software quality metrics from the SQO-OSS quality model. We then look for correlations between (i) popularity and defect density, and (ii) popularity and software quality metrics. In most of the cases, no correlations were found. In cases where minor correlations exist, they are driven by component size. Statistically speaking, and using the methods in this study, the Maven repository does not seem to support the "many eyeballs" effect. We conjecture that the utilization of open source components is driven by factors other than their engineering quality, an interpretation that is supported by the findings in this study.
Hitesh Sajnani, Vaibhav Saini, Joel Ossher, Cristina V. Lopes
ICSME4
2014 Thesaurus-based automatic query expansion for interface-driven code search
abstract
Software engineers often resort to code search practices to support software maintenance and evolution tasks, in particular code reuse. An issue that affects code search is the vocabulary mismatch problem: while searching for a particular function, users have to guess the exact words that were chosen by original developers to name code entities. In this paper we present an automatic query expansion (AQE) approach that uses word relations to increase the chances of finding relevant code. The approach is applied on top of Test-Driven Code Search (TDCS), a promising code retrieval technique that uses test cases as inputs to formulate the search query, but can also be used with other techniques that handle interface definitions to produce queries (interface-driven code search). Since these techniques rely on keywords and types, the vocabulary mismatch problem is also relevant. AQE is carried out by leveraging WordNet, a type thesaurus for expanding types, and another thesaurus containing only software-related word relations. Our approach is general but was specifically designed for non-native English speakers, who are frequently unaware of the most common terms used to name functions in software. Our evaluation with 36 non-native subjects - including developers and senior Computer Science students - provides evidence that our approach can improve the chances of finding relevant functions by 41% (recall improvement of 30%, on average), without hurting precision.
Otávio Augusto Lazzarini Lemos, Adriano Carvalho de Paula, Felipe Capodifoglio Zanichelli, Cristina V. Lopes
MSR4
2014 A dataset for maven artifacts and bug patterns found in them
abstract
In this paper, we present data downloaded from Maven, one of the most popular component repositories. The data includes the binaries of 186,392 components, along with source code for 161,025. We identify and organize these components into groups where each group contains all the versions of a library. In order to asses the quality of these components, we make available report generated by the FindBugs tool on 64,574 components. The information is also made available in the form of a database which stores total number, type, and priority of bug patterns found in each component, along with its defect density. We also describe how this dataset can be useful in software engineering research.
Vaibhav Saini, Hitesh Sajnani, Joel Ossher, Cristina V. Lopes
MSR4
2014 A Comparative Study of Bug Patterns in Java Cloned and Non-cloned Code
abstract
Code cloning via copy-and-paste is a common practice in software engineering. Traditionally, this practice has been considered harmful, and a symptom that some important design abstraction is being ignored. As such, many previous studies suggest approaches to facilitate the discovery, removal, and refactoring of clones. However, not many studies exist that empirically investigate the relationship of code clones with code quality. In this paper, we conduct an empirical study of 31 open source Java projects (1.7 MSLOC) to explore the relationship between code clones and a set of bug patterns reported by Find Bugs. We found that: (i) the defect density in cloned code is 3.7 times less than that of the rest of the code, (ii) 66% of the bug patterns associated with code clones are related to issues in coding style and practice, the two least problematic of the Find Bugs' categories, while that number is 49% for non-cloned code, and (iii) 75% of the bug patterns in cloned code are duplicated without any changes, while 25% are only present in one of the clones. These results show that, when using Find Bugs to detect bug patterns, there is a positive differentiation of cloned code with respect to the rest of the code: the cloned code has considerably less, and less problematic, bug patterns. While our study does not unveil any explanation for this, results from other, more qualitative studies indicate that developers use copy-and-paste intentionally and wisely, which may explain the quantitative observations of our study. Overall, these research results suggest that the practice of code cloning in Java, and possibly in all other object-oriented languages, needs to be given serious consideration on the part of tool designers.
Hitesh Sajnani, Vaibhav Saini, Cristina V. Lopes
SCAM3
2014 Sourcerer: An infrastructure for large-scale collection and analysis of open-source code
Sushil Krishna Bajracharya, Joel Ossher, Cristina V. Lopes
Sci. Comput. Program.3
2012 Software reuse through methodical component reuse and amethodical snippet remixing
abstract
Every method for developing software is a prescriptive model. Applying a deconstructionist analysis to methods reveals that there are two texts, or sets of assumptions and ideals: a set that is privileged by the method and a second set that is left out, or marginalized by the method. We apply this analytical lens to software reuse, a technique in software development that seeks to expedite one's own project by using programming artifacts created by others. By analyzing the methods prescribed by Component-Based Software Engineering (CBSE), we arrive at two texts: Methodical CBSE and Amethodical Remixing. Empirical data from four studies on code search on the web draws attention to four key points of tension: status of component boundaries; provenance of source code; planning and process; and evaluation criteria for candidate code. We conclude the paper with a discussion of the implications of this work for the limits of methods, structure of organizations that reuse software, and the design of search engines for source code.
Kavita Philip, Medha Umarji, Megha Agarwala, Susan Elliott Sim, Rosalva E. Gallardo-Valencia, Cristina V. Lopes, Sukanya Ratanotayanon
CSCW6
2012 10, 000 gold for 20 dollars: an exploratory study of World of Warcraft gold buyers
abstract
Buying virtual currencies with real money from a third-party often violates the terms of use of online games. This study quantitatively investigates players who buy in-game gold from a third-party in World of Warcraft. A cross-cultural survey dataset of 2865 players reveals that differences between Asian and Western players are negligible compared to differences across genders, job categories, and play motivations. Our findings have implications for the design and study of interactions between players and virtual currencies.
Thomas Debeauvais, Bonnie A. Nardi, Cristina V. Lopes, Nick Yee, Nicolas Ducheneaut
FDG3
2012 Parallel code clone detection using MapReduce
abstract
Code clone detection is an established topic in software engineering research. Many detection algorithms have been proposed and refined but very few exploit the inherent parallelism present in the problem, making large scale code clone detection difficult. To alleviate this shortcoming, we present a new technique to efficiently perform clone detection using the popular MapReduce paradigm. Preliminary experimental results demonstrates speed-up and scale-up of the proposed approach.
Hitesh Sajnani, Joel Ossher, Cristina V. Lopes
ICPC3
2012 Trendy bugs: Topic trends in the Android bug reports
abstract
Studying vast volumes of bug and issue discussions can give an understanding of what the community has been most concerned about, however the magnitude of documents can overload the analyst. We present an approach to analyze the development of the Android open source project by observing trends in the bug discussions in the Android open source project public issue tracker. This informs us of the features or parts of the project that are more problematic at any given point of time. In turn, this can be used to aid resource allocation (such as time and man power) to parts or features. We support these ideas by presenting the results of issue topic distributions over time using statistical analysis of the bug descriptions and comments for the Android open source project. Furthermore, we show relationships between those time distributions and major development releases of the Android OS.
Lee Martie, Vijay Krishna Palepu, Hitesh Sajnani, Cristina V. Lopes
MSR4
2012 Optical illusion in augmented reality
abstract
While developers are mainly tackling primary problems in devel-oping augmented reality (AR) systems [1-2], perceptually-correct augmentation rests a critical challenge. In this paper, we focus on how to correctly display and accurately convey the augmented virtual object's size with respect to real-world objects. We conducted a user study to examine how subjects would verify relative size of virtual objects, augmented in a real scene. The results con-firmed that optical illusion occurs in AR applications if comparative size of virtual objects to real-world ones is not considered.
Maryam Khademi, Hossein Mousavi Hondori, Cristina V. Lopes
VRST3
2012 Analyzing and mining a code search engine usage log
abstract
This paper presents an analysis of a year long usage log of Koders, the first commercially available Internet-Scale code search engine ( http://www.koders.com ). The usage log comprises about ten million activities from more than three million users. Analysis of the usage data shows that despite of attracting a large number of visitors, Koders has a very sparse usage and that it lacks regular usage from many of its users. When compared to Web search, search behavior in Koders showed many similar patterns. A topic modeling analysis of the usage data shows what topics users of Koders are looking for. Observations on the prevalence of these topics among the users, and observations on how search and download activities vary across topics, lead to the conclusion that users who find code search engines usable are those who already know to a high level of specificity what to look for. This paper also presents a general categorization of these topics that provides insights on the different ways code search engine users express their queries. It identifies various forms of queries in Koders’s log and the kinds of results addressed by the queries. It also provides several suggestions for improvements in code search engines based on the analysis of usage, topics, and query forms. The work presented in this paper is the first of its kind that reveals several insights on the usage of an Internet-Scale code search engine.
Sushil Krishna Bajracharya, Cristina V. Lopes
Empir. Softw. Eng.2
2012 Efficient Verification of Web-Content Searching Through Authenticated Web Crawlers
abstract
We consider the problem of verifying the correctness and completeness of the result of a keyword search. We introduce the concept of an authenticated web crawler and present its design and prototype implementation. An authenticated web crawler is a trusted program that computes a specially-crafted signature over the web contents it visits. This signature enables (i) the verification of common Internet queries on web pages, such as conjunctive keyword searches---this guarantees that the output of a conjunctive keyword search is correct and complete ; (ii) the verification of the content returned by such Internet queries---this guarantees that web data is authentic and has not been maliciously altered since the computation of the signature by the crawler. In our solution, the search engine returns a cryptographic proof of the query result. Both the proof size and the verification time are proportional only to the sizes of the query description and the query result, but do not depend on the number or sizes of the web pages over which the search is performed. As we experimentally demonstrate, the prototype implementation of our system provides a low communication overhead between the search engine and the user, and fast verification of the returned results by the user.
Michael T. Goodrich, Olga Ohrimenko, Charalampos Papamanthou, Roberto Tamassia, Nikos Triandopoulos, Cristina V. Lopes
Proc. VLDB Endow.7
2011 File cloning in open source Java projects: The good, the bad, and the ugly
abstract
We present a study of the extent to which developers copy entire files or sets of files into their applications with little or no modification. Our aim is to determine the prevalence of such activity within open source Java development, and to identify the circumstances under which files are reused in this manner. To accomplish this aim, we developed a novel method of file-level code clone detection that is scalable to millions of files. We applied our method to the Sourcerer Repository, which contains over 13,000 Java projects aggregated from multiple open source repositories. Our method detected that in excess of 10% of files are clones, and that over 15% of all projects contain at least one cloned file. In addition to computing these raw numbers, we manually examined a large number of the reported clones. We found the most commonly cloned files to be Java extension classes and popular third-party libraries, both large and small. We also discovered a number of projects that occur in multiple online repositories, have been forked, or were divided into multiple subprojects.
Joel Ossher, Hitesh Sajnani, Cristina V. Lopes
ICSM3
2011 Bagging gradient-boosted trees for high precision, low variance ranking models
abstract
Recent studies have shown that boosting provides excellent predictive performance across a wide variety of tasks. In Learning-to-rank, boosted models such as RankBoost and LambdaMART have been shown to be among the best performing learning methods based on evaluations on public data sets. In this paper, we show how the combination of bagging as a variance reduction technique and boosting as a bias reduction technique can result in very high precision and low variance ranking models. We perform thousands of parameter tuning experiments for LambdaMART to achieve a high precision boosting model. Then we show that a bagged ensemble of such LambdaMART boosted models results in higher accuracy ranking models while also reducing variance as much as 50%. We report our results on three public learning-to-rank data sets using four metrics. Bagged LamdbaMART outperforms all previously reported results on ten of the twelve comparisons, and bagged LambdaMART outperforms non-bagged LambdaMART on all twelve comparisons. For example, wrapping bagging around LambdaMART increases [email protected] from 0.4137 to 0.4200 on the MQ2007 data set; the best prior results in the literature for this data set is 0.4134 by RankBoost.
Yasser Ganjisaffar, Rich Caruana, Cristina V. Lopes
SIGIR3
2011 A test-driven approach to code search and its application to the reuse of auxiliary functionality
Otávio Augusto Lazzarini Lemos, Sushil Krishna Bajracharya, Joel Ossher, Paulo César Masiero, Cristina V. Lopes
Inf. Softw. Technol.5
2011 How Well Do Search Engines Support Code Retrieval on the Web?
abstract
Software developers search the Web for various kinds of source code for diverse reasons. In a previous study, we found that searches varied along two dimensions: the size of the search target (e.g., block, subsystem, or system) and the motivation for the search (e.g., reference example or as-is reuse). Would each of these kinds of searches require different search technologies? To answer this question, we conducted an experiment with 36 participants to evaluate three diverse approaches (general purpose information retrieval, source code search, and component reuse), as represented by five Web sites (Google, Koders, Krugle, Google Code Search, and SourceForge). The independent variables were search engine, size of search target, and motivation for search. The dependent variable was the participants judgement of the relevance of the first ten hits. We found that it was easier to find reference examples than components for as-is reuse and that participants obtained the best results using a general-purpose information retrieval site. However, we also found an interaction effect: code-specific search engines worked better in searches for subsystems, but Google worked better on searches for blocks. These results can be used to guide the creation of new tools for retrieving source code from the Web.
Susan Elliott Sim, Medha Umarji, Sukanya Ratanotayanon, Cristina V. Lopes
ACM Trans. Softw. Eng. Methodol.4
2010 Automated dependency resolution for open source software
abstract
Opportunities for software reuse are plentiful, thanks in large part to the widespread adoption of open source processes and the availability of search engines for locating relevant artifacts. One challenge presented by open source software reuse is simply getting a newly downloaded artifact to build/run in the first place. The artifact itself likely reuses other artifacts, and so depends on their being located to function properly. While merely tedious in the individual case, this can cause serious difficulties for those seeking to study open source software. It is simply not feasible to manually resolve dependencies for thousands of projects, and many forms of analysis require declarative completeness. In this paper we present a method for automatically resolving dependencies for open source software. It works by cross-referencing a project's missing type information with a repository of candidate artifacts. We have implemented this method on top of the Sourcerer, an infrastructure for the large-scale indexing and analysis of open source code. The performance of our resolution algorithm was evaluated in two parts. First, for a small number of popular open source projects, we manually examined the artifacts suggested by our system to determine if they were appropriate. Second, we applied the algorithm to the 13,241 projects in the Sourcerer managed repository to evaluate the rate of resolution success. The results demonstrate the feasibility of this approach, as the algorithm located all of the required artifacts needed by 3,904 additional projects, increasing the percentage of declaratively complete projects in Sourcerer from 39% to 69%.
Joel Ossher, Sushil Krishna Bajracharya, Cristina V. Lopes
MSR3
2010 Information-Theoretic Metrics for Project-Level Scattering and Tangling
Erik Linstead, Lindsey Hughes, Cristina V. Lopes, Pierre Baldi
SEKE3
2010 Leveraging usage similarity for effective retrieval of examples in code repositories
abstract
Developers often learn to use APIs (Application Programming Interfaces) by looking at existing examples of API usage. Code repositories contain many instances of such usage of APIs. However, conventional information retrieval techniques fail to perform well in retrieving API usage examples from code repositories. This paper presents Structural Semantic Indexing (SSI), a technique to associate words to source code entities based on similarities of API usage. The heuristic behind this technique is that entities (classes, methods, etc.) that show similar uses of APIs are semantically related because they do similar things. We evaluate the effectiveness of SSI in code retrieval by comparing three SSI based retrieval schemes with two conventional baseline schemes. We evaluate the performance of the retrieval schemes by running a set of 20 candidate queries against a repository containing 222,397 source code entities from 346 jars belonging to the Eclipse framework. The results of the evaluation show that SSI is effective in improving the retrieval of examples in code repositories.
Sushil Krishna Bajracharya, Joel Ossher, Cristina V. Lopes
SIGSOFT FSE3
2009 User contribution and trust in Wikipedia
abstract
Wikipedia, one of the top ten most visited websites, is commonly viewed as the largest online reference for encyclopedic knowledge. Because of its open editing model -allowing anyone to enter and edit content- Wikipedia's overall quality has often been questioned as a source of reliable information.
Sara Javanmardi, Yasser Ganjisaffar, Cristina V. Lopes, Pierre Baldi
CollaborateCom3
2009 Capturing Java naming conventions with first-order Markov models
abstract
We analyze naming conventions for classes, interfaces, methods, and fields across 12,151 open-source Java projects. This vocabulary data is then used to train first-order Markov models to classify entity names, as well as to assess adherence to common naming structure. Preliminary results yield an accuracy of 78.34%. Supplementary material may be found at: http://sourcerer.ics.uci.edu/icpc2009/icpc.html.
Erik Linstead, Lindsey Hughes, Cristina V. Lopes, Pierre Baldi
ICPC3
2009 Mining search topics from a code search engine usage log
abstract
We present a topic modeling analysis of a year long usage log of Koders, one of the major commercial code search engines. This analysis contributes to the understanding of what users of code search engines are looking for. Observations on the prevalence of these topics among the users, and on how search and download activities vary across topics, leads to the conclusion that users who find code search engines usable are those who already know to a high level of specificity what to look for. This paper presents a general categorization of these topics that provides insights on the different ways code search engine users express their queries. The findings support the conclusion that existing code search engines provide only a subset of the various information needs of the users when compared to the categories of queries they look at.
Sushil Krishna Bajracharya, Cristina V. Lopes
MSR2
2009 SourcererDB: An aggregated repository of statically analyzed and cross-linked open source Java projects
abstract
The open source movement has made vast quantities of source code available online for free, providing an extremely large dataset for empirical study and potential resuse. A major difficulty in exploiting this potential fully is that the data are currently scattered between competing source code repositories, none of which are structured for empirical analysis and cross-project comparison. As a result, software researchers and developers are left to compile their own datasets, resulting in duplicated effort and limited results. To address this challenge, we built SourcererDB, an aggregated repository of statically analyzed and cross-linked open source Java projects. SourcererDB contains local snapshots of 2,852 Java projects taken from Sourceforge, Apache and Java.net. These projects are statically analyzed to extract rich structural information, which is then stored in a relational database. References to entities in the 16,058 external jars are resolved and grouped, allowing for cross-project usage information to be accessed easily. This paper describes: (a) the mechanism for resolving and grouping these cross-project references, (b) the structure of and the metamodel for the SourcererDB repository, and (d) end-user dataset access mechanisms. Our goal in building SourcererDB is to provide a rich dataset of source code to facilitate the sharing of extracted data and to encourage reuse and repeatability of experiments.
Joel Ossher, Sushil Krishna Bajracharya, Erik Linstead, Pierre Baldi, Cristina V. Lopes
MSR5
2009 The massification and webification of systems' modeling and simulation with virtual worlds
abstract
General-purpose virtual world platforms are having a surprising effect in the old practice of systems modeling and simulation: its massification and webification. A visit to any virtual place in Second Life-like worlds shows sophisticated virtual constructions with non-trivial dynamic behavior that have been set up by ordinary people who neither used professional 3D modeling and simulation tools, nor consider themselves as programmers. Many of those virtual constructions are imaginary systems that could not survive the laws of physics. However, more and more people are experimenting with these virtual worlds to model, simulate and visualize real-world systems, for a number of purposes: historical reconstruction, rich 3D interfaces to physical equipment control, urban planning, and scientific visualizations are just a few. This gives an overview of all this experimentation, and analyzes the ways by which it improves the state-of-the-art professional tools. It also presents some of the engineering and software engineering challenges that this practice brings to the table.
Cristina V. Lopes
ESEC/SIGSOFT FSE1
2009 Software-driven sensor networks for short-range shallow water applications
Raja Jurdak, Pierre Baldi, Cristina V. Lopes
Ad Hoc Networks3
2009 Sourcerer: mining and searching internet-scale software repositories
Erik Linstead, Sushil Krishna Bajracharya, Trung Chi Ngo, Paul Rigor, Cristina V. Lopes, Pierre Baldi
Data Min. Knowl. Discov.5
2008 CalSWIM: A Wiki-Based Data Sharing Platform
Yasser Ganjisaffar, Sara Javanmardi, Stanley Grant, Cristina V. Lopes
CollaborateCom4
2008 An Application of Latent Dirichlet Allocation to Analyzing Software Evolution
abstract
We develop and apply unsupervised statistical topic models, in particular latent Dirichlet allocation, to identify functional components of source code and study their evolution over multiple project versions. We present results for two large, open source Java projects, Eclipse and Argo UML, which are well-known and well-studied within the software mining community. Our results demonstrate the effectiveness of probabilistic topic models in automatically summarizing the temporal dynamics of software concerns, with direct application to project management and program understanding. In addition to detecting the emergence of topics on the release timeline which represent integration points for key source code functionality, our techniques can also be used to pinpoint refactoring events in the underlying software design, as well as to identify general programming concepts whose prevalence is dependent only of the size of the code base to be analyzed. Complete results are available from our supplementary materials website at http://sourcerer.ics.uci.edu/icmla2008/software_evolution.html.
Erik Linstead, Cristina V. Lopes, Pierre Baldi
ICMLA2
2008 XE (eXtreme Editor) - Bridging the Aspect-Oriented Programming Usability Gap
abstract
In spite of the modularization benefits supported by the Aspect-Oriented programming paradigm, different usability issues have hindered its adoption. The decoupling between aspect definitions and base code, and the compile-time weaving mechanism adopted by different AOP languages, require developers to manage the consistency between base code and aspect code themselves. These mechanisms create opportunities for errors related to aspect weaving invisibility and non-local control characteristics of AOP languages. This paper describes XE (Extreme Editor), an IDE that supports developers in managing these issues in the functional aspect-oriented programming domain.
Wiwat Ruengmee, Roberto Silveira Silva Filho, Sushil Krishna Bajracharya, David F. Redmiles, Cristina V. Lopes
ASE5
2008 A theory of aspects as latent topics
abstract
After more than 10 years, Aspect-Oriented Programming (AOP) is still a controversial idea. While the concept of aspects appeals to everyone's intuitions, concrete AOP solutions often fail to convince researchers and practitioners alike. This discrepancy results in part from a lack of an adequate theory of aspects, which in turn leads to the development of AOP solutions that are useful in limited situations.
Pierre Baldi, Cristina V. Lopes, Erik Linstead, Sushil Krishna Bajracharya
OOPSLA2
2007 Modeling trust in collaborative information systems
abstract
Collaborative systems available on the Web allow millions of users to share information through a growing collection of tools and platforms such as wikis, blogs and shared forums. All of these systems contain information and resources with different degrees of sensitivity. However, the open nature of such infrastructures makes it difficult for users to determine the reliability of the available information and trustworthiness of information providers. Hence, integrating trust management systems to open collaborative systems can play a crucial role in the growth and popularity of open information repositories. In this paper, we present a trust model for collaborative systems, namely for platforms based on Wiki technology. This model, based on hidden Markov models, estimates the reputation of the contributors and the reliability of the content dynamically. The focus of this paper is on reputation estimation. Evaluation results based on a subset of Wikipedia shows that the model can effectively be used for identifying vandals, and users with high quality contributions.
Sara Javanmardi, Cristina V. Lopes
CollaborateCom2
2007 Reliable Symbol Synchronization in Software-Driven Acoustic Sensor Networks
abstract
Symbol synchronization in traditional hardware- driven communication systems has relied on the transmission of training sequences of symbols just before the beginning of the frame symbols. The use of training sequences is not suitable for software-driven communication systems, such as lightweight acoustic underwater sensor networks [4,5], in which the high symbol loss rate may cause the loss of training symbols, preventing accurate symbol synchronization. Software-driven communication networks require symbol synchronization that is resilient to a high loss environment, that does not represent large communication or processing overhead, and that is tunable to the noise profile of different environments. These requirements are emphasized for mote-based acoustic underwater sensor networks in which the bandwidth and processing capability are sparse. This paper proposes the use of a short signature synchronization symbol (S4) as both a preamble and post-amble to enable receiver synchronization in mote-based acoustic communication systems that rely on software modems. To synchronize to an incoming signal, the receiver performs cross-correlation of N reference signature symbols with the incoming signal to identify the beginning of the preamble and post-amble. The output of the cross-correlation yields 2N peak values, from which the receiver chooses the sharpest and most symmetric for synchronization to the beginning of the frame. Empirical experiments confirm a synchronization accuracy within 5 ms in air within a range of 10.5 m, and 11 ms in water within a range of 15 m.
Raja Jurdak, Antonio G. Ruzzelli, Gregory M. P. O'Hare, Cristina V. Lopes
GLOBECOM4
2007 CodeGenie: using test-cases to search and reuse source code
abstract
We present CodeGenie, a tool that implements a test-driven approachto search and reuse of code available on large-scale coderepositories. While using CodeGenie developers design test cases fora desired feature first, similar to Test-driven Development (TDD).However, instead of implementing the feature as in TDD, CodeGenieautomatically searches for it based on information available in thetests. To check the suitability of the candidate results in thelocal context, each result is automatically woven into thedeveloper's project and tested using the original tests. Thedeveloper can then reuse the most suitable result. Later, reusedcode can also be unwoven from the project as wished. For the codesearching and wrapping facilities, CodeGenie relies on Sourcerer, anInternet-scale source code infrastructure that we have developed
Otávio Augusto Lazzarini Lemos, Sushil Krishna Bajracharya, Joel Ossher, Ricardo Morla, Paulo César Masiero, Pierre Baldi, Cristina V. Lopes
ASE7
2007 Mining concepts from code with probabilistic topic models
abstract
We develop and apply statistical topic models to software as a means of extracting concepts from source code. The effectiveness of the technique is demonstrated on 1,555 projects from SourceForge and Apache consisting of 113,000 files and 19 million lines of code. In addition to providing an automated, unsupervised, solution to the problem of summarizing program functionality, the approach provides a probabilistic framework with which to analyze and visualize source file similarity. Finally, we introduce an information-theoretic approach for computing tangling and scattering of extracted concepts, and present preliminary results
Erik Linstead, Paul Rigor, Sushil Krishna Bajracharya, Cristina V. Lopes, Pierre Baldi
ASE4
2007 Mining Internet-Scale Software Repositories
abstract
Large repositories of source code create new challenges and opportunities for statistical machine learning. Here we first develop an infrastructure for the automated crawling, parsing, and database storage of open source software. The infrastructure allows us to gather Internet-scale source code. For instance, in one experiment, we gather 4,632 java projects from SourceForge and Apache totaling over 38 million lines of code from 9,250 developers. Simple statistical analyses of the data first reveal robust power-law behavior for package, SLOC, and method call distributions. We then develop and apply unsupervised author-topic, probabilistic models to automatically discover the topics embedded in the code and extract topic-word and author-topic distributions. In addition to serving as a convenient summary for program function and developer activities, these and other related distributions provide a statistical and information-theoretic basis for quantifying and analyzing developer similarity and competence, topic scattering, and document tangling, with direct applications to software engineering. Finally, by combining software textual content with structural information captured by our CodeRank approach, we are able to significantly improve software retrieval performance, increasing the AUC metric to 0.86-- roughly 10-30% better than previous approaches based on text alone.
Erik Linstead, Paul Rigor, Sushil Krishna Bajracharya, Cristina V. Lopes, Pierre Baldi
NIPS4
2007 Adaptive Low Power Listening for Wireless Sensor Networks
abstract
Most sensor networks require application-specific network-wide performance guarantees, suggesting the need for global and flexible network optimization. The dynamic and nonuniform local states of individual nodes in sensor networks complicate global optimization. Here, we present a cross-layer framework for optimizing global power consumption and balancing the load in sensor networks through greedy local decisions. Our framework enables each node to use its local and neighborhood state information to adapt its routing and MAC layer behavior. The framework employs a flexible cost function at the routing layer and adaptive duty cycles at the MAC layer in order to adapt a node's behavior to its local state. We identify three state aspects that impact energy consumption: 1) number of descendants in the routing tree, 2) radio duty cycle, and 3) role. We conduct experiments on a test-bed of 14 mica2 sensor nodes to compare the state representations and to evaluate the framework's energy benefits. The experiments show that the degree of load balancing increases for expanded state representations. The experiments also reveal that all state representations in our framework reduce global power consumption in the range of one-third for a time-driven monitoring network and in the range of one-fifth for an event-driven target tracking network.
Raja Jurdak, Pierre Baldi, Cristina V. Lopes
IEEE Trans. Mob. Comput.3
2005 Beep: 3D indoor positioning using audible sound
abstract
Rapid growth in the number of wireless enabled devices has led to an increased interest in location-aware applications. The backbone of such applications is provided by a location system. In this paper we present Beep, an indoor location system that senses audible sound. The use of audible sound makes our system cheap and easily deplorable to most existing roaming devices. Unlike positioning systems using ultrasound and infrared signals, Beep does not require the user to carry any kind of specialized hardware. Our system is based on standard 3D multilateration algorithms. However, the requirement of being able to locate existing devices, whose sound cards were not designed for high-precision signaling, introduces additional challenges to the location problem. This paper describes how those problems were solved and presents experimental results. Beep works with an accuracy of about 2 feet in more than 97% cases. The paper also describes a sensor deployment strategy that requires low sensor density and consequently low installation costs.
Atri Mandal, Cristina V. Lopes, Tony Givargis, Amir Haghighat, Raja Jurdak, Pierre Baldi
CCNC2
2005 U-MAC: a proactive and adaptive UWB medium access control protocol
abstract
Abstract Ultra wide band (UWB) technology has received increasing recognition in recent years for its potential applications beyond radar technology to communication networks. UWB is a spread spectrum technology that requires careful coordination among communicating nodes to jointly control link power and transmission rates. Here, we present ultra wide band MAC (U‐MAC), an adaptive medium access control (MAC) protocol for UWB in which nodes periodically declare their current state, so that neighbors can proactively assign power and rate values for new links locally in order to optimize global network performance. Simulations comparing U‐MAC to the reactive approach confirm that U‐MAC lowers link setup latency and control overhead, doubles the throughput and adapts better to high network loads. Simulations also reveal that the basic form of U‐MAC favors nodes that are closer to the receiver. As a result, we also introduce novel mechanisms that control the radius around a receiver within which nodes can have fair access to it. We show through simulations the effect of the mechanisms on the tradeoff between network throughput and fair access. Copyright © 2005 John Wiley & Sons, Ltd.
Raja Jurdak, Pierre Baldi, Cristina V. Lopes
Wirel. Commun. Mob. Comput.3
2003 A 40 bps speech coding scheme
abstract
We describe a method and an implementation for producing a highly compressed representation of speech, in the order of 40 bps. This compression method uses a speech recognition engine to analyze the speech signal at the morphological level, i.e. the words. The words are then coded using a word-level text compression mechanism. After decompression, the speech message is recovered using text-to-speech synthesis. We report experimental results of our implementation. In particular, we observed that human listeners were able to recover from errors introduced by the speech recognition engine, and that the human perceptual errors were highly dependent on the content of the messages, especially regarding familiarity with the topic.
Cristina V. Lopes, Anshuman Chadha
GLOBECOM1
2002 Making sense of sensing systems: five questions for designers and researchers
abstract
This paper borrows ideas from social science to inform the design of novel "sensing" user-interfaces for computing technology. Specifically, we present five design challenges inspired by analysis of human-human communication that are mundanely addressed by traditional graphical user interface designs (GUIs). Although classic GUI conventions allow us to finesse these questions, recent research into innovative interaction techniques such as 'Ubiquitous Computing' and 'Tangible Interfaces' has begun to expose the interaction challenges and problems they pose. By making them explicit we open a discourse on how an approach similar to that used by social scientists in studying human-human interaction might inform the design of novel interaction mechanisms that can be used to handle human-computer communication accomplishments
Victoria Bellotti, Maribeth Back, W. Keith Edwards, Rebecca E. Grinter, Austin Henderson, Cristina V. Lopes
CHI6
2000 A study on exception detecton and handling using aspect-oriented programming
abstract
Aspect-Oriented Programming (AOP) is intended to ease situations that involve many kinds of code tangling. This paper reports on a study to investigate AOP's ability to ease tangling related to exception detection and handling. We took an existing framework written in Java™, the JWAM framework, and partially reengineered its exception detection and handling aspects using AspectJ™, an aspect-oriented programming extension to Java.
Martin Lippert, Cristina V. Lopes
ICSE2
2000 Improving design and source code modularity using AspectJ (tutorial session)
abstract
Using only traditional techniques the implementation of concerns like exception handling, multi-object protocols, synchronization constraints, and security policies tends to be spread out in the code. The lack of modularity for these concerns makes them more difficult to develop and maintain. This tutorial shows how to use Aspect-oriented programming (AOP) [2, 3] to implement concerns like these in a concise modular way. We discuss the effect aspects have on software design and on code modularity. The concrete examples in the tutorial use AspectJ [1], a freely available aspect-oriented extension to the Java™ programming language.
Cristina V. Lopes, Gregor Kiczales
ICSE1
1997 Aspect-Oriented Programming
Gregor Kiczales, John Lamping, Anurag Mendhekar, Chris Maeda, Cristina V. Lopes, Jean-Marc Loingtier, John Irwin
ECOOP5
1997 Open Implementation Design Guidelines
abstract
Designing reusable software modules can be extremely difficult.The design must be balanced between being general enough to address the needs of a wide range of clients and being focused enough to truly satisfy the requirements of each specific client One area where it can be particularly difficult to strike this balance is in the implementation strategy of the module.The problem is that generalpurpose implementation strategies, tuned for a wide range of clients, aren't necessarily optimal for each specific client-this is especially an issue for modules that are intended to be reusable and yet provide high-performance.An examination of existing software systems shows that an increasingly important technique for handling this problem is to design the module's interface in such a way that the client can assist or participate in the selection of the module's implementation strategy.We call this approach open implementation.When designing the interface to a module that allows its clients some control over its implementation strategy, it is important to retain, as much as possible, the advantages of traditional closed implementation modules.This paper explores issues in the design of interfaces to open implementation modules.We identify key design choices, and present guidelines for deciding which choices are likely to work best in particular situations.
Gregor Kiczales, John Lamping, Cristina V. Lopes, Chris Maeda, Anurag Mendhekar, Gail C. Murphy
ICSE3
1994 Abstracting Process-to-Function Relations in Concurrency Object-Oriented Applications
Cristina V. Lopes, Karl J. Lieberherr
ECOOP1