VLDB 2026 Research / reviewers in the wild / expert
Benoit Baudry
dblp:57/3320 · also Benoît Baudry
· DBLP profile ↗
158ranked-venue papers
12as first author
30since 2021 · last 2026
0000-0002-4015-4640ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 138 · 11 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Security and privacy · 7 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The design space of lockfiles across package managers
Yogya Gamage, Deepika Tiwari, Martin Monperrus, Benoit Baudry |
Empir. Softw. Eng. | 4 |
| 2026 | Byam: Fixing Breaking Dependency Updates with Large Language ModelsabstractApplication Programming Interfaces (APIs) facilitate the integration of third-party dependencies within the code of client applications. However, changes to an API, such as deprecation, modification of parameter names or types, or complete replacement with a new API, can break existing client code. These changes are called breaking dependency updates ; It is often tedious for API users to identify the cause of these breaks and update their code accordingly. In this paper, we explore the use of Large Language Models (LLMs) to automate client code updates in response to breaking dependency updates. We evaluate our approach on the BUMP dataset, a benchmark for breaking dependency updates in Java projects. Our approach leverages LLMs with advanced prompts, including information from the build process and from the breaking dependency analysis. We assess effectiveness at three granularity levels: at the build level, the file level, and the individual compilation error level. We experiment with five LLMs: Google Gemini-2.0 Flash, OpenAI GPT4o-mini, OpenAI o3-mini, Alibaba Qwen2.5-32b-instruct, and DeepSeek V3. Our results show that LLMs can automatically repair breaking updates. Among the considered models, OpenAI’s o3-mini is the best, able to completely fix 27% of the builds when using prompts that include contextual information such as the erroneous line, API differences, error messages, and step-by-step reasoning instructions. Also, it fixes 78% of the individual compilation errors. Overall, our findings demonstrate the potential for LLMs to fix compilation errors due to breaking dependency updates, supporting developers in their efforts to stay up-to-date with changes in their dependencies. Frank Reyes, May Mahmoud, Federico Bono, Sarah Nadi, Benoit Baudry, Martin Monperrus |
Empir. Softw. Eng. | 5 |
| 2026 | Serializing java objects in plain codeabstractIn managed languages, serialization of objects is typically done in bespoke binary formats such as Protobuf, or markup languages such as XML or JSON. The major limitation of these formats is readability. Human developers cannot read binary code, and in most cases, suffer from the syntax of XML or JSON. This is a major issue when objects are meant to be embedded and read in source code, such as in test cases. To address this problem, we propose plain-code serialization. Our core idea is to serialize objects observed at runtime in the native syntax of a programming language. We realize this vision in the context of Java, and demonstrate a prototype which serializes Java objects to Java source code. The resulting source faithfully reconstructs the objects seen at runtime. Our prototype is called ProDJ and is publicly available. We experiment with ProDJ to successfully plain-code serialize 174,699 objects observed during the execution of 4 open-source Java applications. Our performance measurement shows that the performance impact is not noticeable. Through a user study, we demonstrate that developers prefer plain-code serialized objects within automatically generated tests over their representations as XML or JSON. Julian Wachter, Deepika Tiwari, Martin Monperrus, Benoit Baudry |
J. Syst. Softw. | 4 |
| 2026 | Causes and Canonicalization of Unreproducible Builds in JavaabstractThe increasing complexity of software supply chains and the rise of supply chain attacks have elevated concerns around software integrity. Users and stakeholders face significant challenges in validating that a given software artifact corresponds to its declared source. Reproducible Builds address this challenge by ensuring that independently performed builds from identical source code produce identical binaries. However, achieving reproducibility at scale remains difficult, especially in Java, due to a range of non-deterministic factors and caveats in the build process. In this work, we focus on reproducibility in Java-based software, archetypal of enterprise applications. We introduce a conceptual framework for reproducible builds, we analyze a large dataset from Reproducible Central, and we develop a novel taxonomy of six root causes of unreproducibility. We study actionable mitigations: artifact and bytecode canonicalization using OSS-Rebuild and jNorm respectively. Finally, we presentChains-Rebuild(improvements to OSS-Rebuild), a tool that raises reproducibility success from 9.48% to 26.60% on 12,803 unreproducible artifacts. To sum up, our contributions are the first large-scale taxonomy of build unreproducibility causes in Java, a publicly available dataset of unreproducible builds, andChains-Rebuild, a canonicalization tool for mitigating unreproducible builds in Java. Aman Sharma 0001, Benoit Baudry, Martin Monperrus |
IEEE Trans. Software Eng. | 2 |
| 2025 | MYRIAD PEOPLE Open Source Software for New Media ArtsabstractNew media art builds on top of rich software stacks. Blending multiple media such as code, light or sound, new media artists integrate various types of software to draw, animate, control or synchronize different parts of an artwork. Yet, the artworks rarely credit software and all the developers involved.In this work, we present MYRIAD PEOPLE, an original dataset of open source projects and their contributors, which span various software layers used in new media art installations. To collect this dataset, we released an open call for artists and eventually curated 9 artworks, which use a variety of software and media. In October 2024, we organized a collective exhibition in Stockholm, entitled MYRIAD, which showcased the 9 artworks. The MYRIAD PEOPLE dataset includes the 124 open source projects used in one or more of the MYRIAD’s artworks, as well as all the contributors to these projects. In this paper, we present the dataset, as well as the possible usages of this dataset for software and art research. Benoit Baudry, Erik Natanael Gustafsson, Roni Kaufman, Maria Kling |
MSR | 1 |
| 2025 | Software Bills of Materials in Maven CentralabstractSoftware Bills of Materials (SBOMs) are essential to ensure the transparency and integrity of the software supply chain. There is a growing body of work that investigates the accuracy of SBOM generation tools and the challenges for producing complete SBOMs. Yet, there is little knowledge about how developers distribute SBOMs. In this work, we mine SBOMs from Maven Central to assess the extent to which developers publish SBOMs along with the artifacts. We develop our work on top of the Goblin framework, which consists of a Maven Central dependency graph and a Weaver that allows augmenting the dependency graph with additional data. For this study, we select a sample of 10% of release nodes from the Maven Central dependency graph and collected 14,071 SBOMs from 7,290 package releases. We then augment the Maven Central dependency graph with the collected SBOMs. We present our methodology to mine SBOMs, as well as novel insights about SBOM publication. Our dataset is the first set of SBOMs collected from a package registry. We make it available as a standalone dataset, which can be used for future research about SBOMs and package distribution. Yogya Gamage, Nadia Gonzalez Fernandez, Martin Monperrus, Benoit Baudry |
MSR | 4 |
| 2025 | Detecting and removing bloated dependencies in CommonJS packagesabstractJavaScript packages are notoriously prone to bloat, a factor that significantly impacts the performance and maintainability of web applications. While web bundlers and tree-shaking can mitigate this issue in client-side applications, state-of-the-art techniques have limitations on the detection and removal of bloat in server-side applications. In this paper, we present the first study to investigate bloated dependencies within server-side JavaScript applications, focusing on those built with the widely used and highly dynamic CommonJS module system. We propose a trace-based dynamic analysis that monitors the OS file system, to determine which dependencies are not accessed during runtime. To evaluate our approach, we curate an original dataset of 91 CommonJS packages with a total of 50,488 dependencies. Compared to the state-of-the-art dynamic and static approaches, our trace-based analysis demonstrates higher accuracy in detecting bloated dependencies. Our analysis identifies 50.6% of the 50,488 dependencies as bloated: 13.8% of direct dependencies and 51.3% of indirect dependencies. Furthermore, removing only the direct bloated dependencies by cleaning the dependency configuration file can remove a significant share of unnecessary bloated indirect dependencies while preserving function correctness. Deepika Tiwari, Cristian Bogdan, Benoit Baudry |
J. Syst. Softw. | 4 |
| 2024 | Breaking-Good: Explaining Breaking Dependency Updates with Build AnalysisabstractDependency updates often cause compilation errors when new dependency versions introduce changes that are incompatible with existing client code. Fixing breaking dependency updates is notoriously hard, as their root cause can be hidden deep in the dependency tree. We present Breaking-Good, a tool that automatically generates explanations for breaking updates. Breaking-Good provides a detailed categorization of compilation errors, identifying several factors related to changes in direct and indirect dependencies, incompatibilities between Java versions, and client-specific configuration. With a blended analysis of log and dependency trees, Breaking-Good generates detailed explanations for each breaking update. These explanations help developers understand the causes of the breaking update, and suggest possible actions to fix the breakage. We evaluate Breaking-Good on 243 real-world breaking dependency updates. Our results indicate that Breaking-Good accurately identifies root causes and generates automatic explanations for 70 % of these breaking updates. Our user study demonstrates that the generated explanations help developers. Breaking-Good is the first technique that automatically identifies the causes of a breaking dependency update and explains the breakage accordingly. Frank Reyes, Benoit Baudry, Martin Monperrus |
SCAM | 2 |
| 2024 | PROZE: Generating Parameterized Unit Tests Informed by Runtime DataabstractTypically, a conventional unit test (CUT) verifies the expected behavior of the unit under test through one specific input / output pair. In contrast, a parameterized unit test (PUT) receives a set of inputs as arguments, and contains assertions that are expected to hold true for all these inputs. PUTs increase test quality, as they assess correctness on a broad scope of inputs and behaviors. However, defining assertions over a set of inputs is a hard task for developers, which limits the adoption of PUTs in practice. In this paper, we address the problem of finding oracles for PUTs that hold over multiple inputs. We design a system called PROZE, that generates PUTs by identifying developer-written assertions that are valid for more than one test input. We implement our approach as a two-step methodology: first, at runtime, we collect inputs for a target method that is invoked within a CUT; next, we isolate the valid assertions of the CUT to be used within a PUT. We evaluate our approach against 5 real-world Java modules, and collect valid inputs for 128 target methods, from test and field executions. We generate 2,287 PUTs, which invoke the target methods with a significantly larger number of test inputs than the original CUTs. We execute the PUTs and find 217 that provably demonstrate that their oracles hold for a larger range of inputs than envisioned by the developers. From a testing theory perspective, our results show that developers express assertions within CUTs, which actually hold beyond one particular input. Deepika Tiwari, Yogya Gamage, Martin Monperrus, Benoit Baudry |
SCAM | 4 |
| 2024 | BUMP: A Benchmark of Reproducible Breaking Dependency UpdatesabstractThird-party dependency updates can cause a build to fail if the new dependency version introduces a change that is incompatible with the usage: this is called a breaking dependency update. Research on breaking dependency updates is active, with works on characterization, understanding, automatic repair of breaking updates, and other software engineering aspects. All such research projects require a benchmark of breaking updates that has the following properties: 1) it contains real-world breaking updates; 2) the breaking updates can be executed; 3) the benchmark provides stable scientific artifacts of breaking updates over time, a property we call “reproducibility”. To the best of our knowledge, such a benchmark is missing. To address this problem, we present BUMP, a new benchmark that contains reproducible breaking dependency updates in the context of Java projects built with the Maven build system. BUMP contains 571 breaking dependency updates collected from 153 Java projects. BUMP ensures long-term reproducibility of dependency updates on different platforms, guaranteeing consistent build failures. We categorize the different causes of build breakage in BUMP, providing novel insights for future work on breaking update engineering. To our knowledge, BUMP is the first of its kind, providing hundreds of real-world breaking updates that have all been made reproducible. Frank Reyes, Yogya Gamage, Gabriel Skoglund, Benoit Baudry, Martin Monperrus |
SANER | 4 |
| 2024 | Wasm-Mutate: Fast and effective binary diversification for WebAssemblyabstractWebAssembly is the fourth officially endorsed Web language. It is recognized because of its efficiency and design, focused on security. Yet, its swiftly expanding ecosystem lacks robust software diversification systems. We introduce Wasm-Mutate, a diversification engine specifically designed for WebAssembly. Our engine meets several essential criteria: 1) To quickly generate functionally identical, yet behaviorally diverse, WebAssembly variants, 2) To be universally applicable to any WebAssembly program, irrespective of the source programming language, and 3) Generated variants should counter side-channels. By leveraging an e-graph data structure, Wasm-Mutate is implemented to meet both speed and efficacy. We evaluate Wasm-Mutate by conducting experiments on 404 programs, which include real-world applications. Our results highlight that Wasm-Mutate can produce tens of thousands of unique and efficient WebAssembly variants within minutes. Significantly, Wasm-Mutate can safeguard WebAssembly binaries against timing side-channel attacks, especially those of the Spectre type. Javier Cabrera-Arteaga, Nicholas FitzGerald, Martin Monperrus, Benoit Baudry |
Comput. Secur. | 4 |
| 2024 | Highly Available Blockchain Nodes With N-Version DesignabstractAs all software, blockchain nodes are exposed to faults in their underlying execution stack. Unstable execution environments can disrupt the availability of blockchain nodes’ interfaces, resulting in downtime for users. This article introduces the concept of N-Version Blockchain nodes. This new type of node relies on simultaneous execution of different implementations of the same blockchain protocol, in the line of Avizienis’ N-Version programming vision. We design and implement an N-Version blockchain node prototype in the context of Ethereum, calledN-ETH. We show thatN-ETHis able to mitigate the effects of unstable execution environments and significantly enhance availability under environment faults. To simulate unstable execution environments, we perform fault injection at the system-call level. Our results show that existing Ethereum node implementations behave asymmetrically under identical instability scenarios.N-ETHleverages this asymmetric behavior available in the diverse implementations of Ethereum nodes to provide increased availability, even under our most aggressive fault-injection strategies. We are the first to validate the relevance of N-Version design in the domain of blockchain infrastructure. From an industrial perspective, our results are of utmost importance for businesses operating blockchain nodes, including Google, ConsenSys, and many other major blockchain companies. Javier Ron, César Soto-Valero, Benoit Baudry, Martin Monperrus |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Mimicking Production Behavior With Generated MocksabstractMocking allows testing program units in isolation. A developer who writes tests with mocks faces two challenges: design realistic interactions between a unit and its environment; and understand the expected impact of these interactions on the behavior of the unit. In this paper, we propose to monitor an application in production to generate tests that mimic realistic execution scenarios through mocks. Our approach operates in three phases. First, we instrument a set of target methods for which we want to generate tests, as well as the methods that they invoke, which we refer to as mockable method calls. Second, in production, we collect data about the context in which target methods are invoked, as well as the parameters and the returned value for each mockable method call. Third, offline, we analyze the production data to generate test cases with realistic inputs and mock interactions. The approach is automated and implemented in an open-source tool calledrick. We evaluate our approach with three real-world, open-source Java applications.rickmonitors the invocation of$128$methods in production across the three applications and captures their behavior. Based on this captured data,rickgenerates test cases that include realistic initial states and test inputs, as well as mocks and stubs. All the generated test cases are executable, and$52.4\%$of them successfully mimic the complete execution context of the target methods observed in production. The mock-based oracles are also effective at detecting regressions within the target methods, complementing each other in their fault-finding ability. We interview$5$developers from the industry who confirm the relevance of using production observations to design mocks and stubs. Our experimental findings clearly demonstrate the feasibility and added value of generating mocks from production interactions. Deepika Tiwari, Martin Monperrus, Benoit Baudry |
IEEE Trans. Software Eng. | 3 |
| 2023 | RICK: Generating Mocks from Production DataabstractTest doubles, such as mocks and stubs, are nifty fixtures in unit tests. They allow developers to test individual components in isolation from others that lie within or outside of the system. However, implementing test doubles within tests is not straightforward. With this demonstration, we introduce RICK, a tool that observes executing applications in order to automatically generate tests with realistic mocks and stubs. RICK monitors the invocation of target methods and their interactions with external components. Based on the data collected from these observations, RICK produces unit tests with mocks, stubs, and mock-based oracles. We highlight the capabilities of RICK, and how it can be used with real-world Java applications, to generate tests with mocks. Deepika Tiwari, Martin Monperrus, Benoit Baudry |
ICST | 3 |
| 2023 | WebAssembly diversification for malware evasionabstractWebAssembly has become a crucial part of the modern web, offering a faster alternative to JavaScript in browsers. While boosting rich applications in browser, this technology is also very efficient to develop cryptojacking malware. This has triggered the development of several methods to detect cryptojacking malware. However, these defenses have not considered the possibility of attackers using evasion techniques. This paper explores how automatic binary diversification can support the evasion of WebAssembly cryptojacking detectors. We experiment with a dataset of 33 WebAssembly cryptojacking binaries and evaluate our evasion technique against two malware detectors: VirusTotal, a general-purpose detector, and MINOS, a WebAssembly-specific detector. Our results demonstrate that our technique can automatically generate variants of WebAssembly cryptojacking that evade the detectors in 90% of cases for VirusTotal and 100% for MINOS. Our results emphasize the importance of meta-antiviruses and diverse detection techniques and provide new insights into which WebAssembly code transformations are best suited for malware evasion. We also show that the variants introduce limited performance overhead, making binary diversification an effective technique for evasion. Javier Cabrera-Arteaga, Martin Monperrus, Tim Toady, Benoit Baudry |
Comput. Secur. | 4 |
| 2023 | Chaos Engineering of Ethereum Blockchain ClientsabstractIn this article, we present ChaosETH , a chaos engineering approach for resilience assessment of Ethereum blockchain clients. ChaosETH operates in the following manner: First, it monitors Ethereum clients to determine their normal behavior. Then, it injects system call invocation errors into one single Ethereum client at a time and observes the behavior resulting from perturbation. Finally, ChaosETH compares the behavior recorded before, during, and after perturbation to assess the impact of the injected system call invocation errors. The experiments are performed on the two most popular Ethereum client implementations: GoEthereum and Nethermind. We assess the impact of 22 different system call errors on those Ethereum clients with respect to 15 application-level metrics. Our results reveal a broad spectrum of resilience characteristics of Ethereum clients w.r.t. system call invocation errors, ranging from direct crashes to full resilience. The experiments clearly demonstrate the feasibility of applying chaos engineering principles to blockchain systems. Javier Ron, Benoit Baudry, Martin Monperrus |
Distributed Ledger Technol. Res. Pract. | 3 |
| 2023 | Coverage-Based Debloating for Java BytecodeabstractSoftware bloat is code that is packaged in an application but is actually not necessary to run the application. The presence of software bloat is an issue for security, performance, and for maintenance. In this article, we introduce a novel technique for debloating, which we call coverage-based debloating. We implement the technique for one single language: Java bytecode. We leverage a combination of state-of-the-art Java bytecode coverage tools to precisely capture what parts of a project and its dependencies are used when running with a specific workload. Then, we automatically remove the parts that are not covered, in order to generate a debloated version of the project. We succeed to debloat 211 library versions from a dataset of 94 unique open-source Java libraries. The debloated versions are syntactically correct and preserve their original behaviour according to the workload. Our results indicate that 68.3% of the libraries’ bytecode and 20.3% of their total dependencies can be removed through coverage-based debloating. For the first time in the literature on software debloating, we assess the utility of debloated libraries with respect to client applications that reuse them. We select 988 client projects that either have a direct reference to the debloated library in their source code or which test suite covers at least one class of the libraries that we debloat. Our results show that 81.5% of the clients, with at least one test that uses the library, successfully compile and pass their test suite when the original library is replaced by its debloated version. César Soto-Valero, Thomas Durieux, Nicolas Harrand, Benoit Baudry |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Spork: Structured Merge for Java With Formatting PreservationabstractThe highly parallel workflows of modern software development have made merging of source code a common activity for developers. The state of the practice is based on line-based merge, which is ubiquitously used with “git merge”. Line-based merge is however a generalized technique for any text that cannot leverage the structured nature of source code, making merge conflicts a common occurrence. As a remedy, research has proposed structured merge tools, which typically operate on abstract syntax trees instead of raw text. Structured merging greatly reduces the prevalence of merge conflicts but suffers from important limitations, the main ones being a tendency to alter the formatting of the merged code and being prone to excessive running times. In this paper, we presentspork, a novel structured merge tool forjava.sporkis unique as it preserves formatting to a significantly greater degree than comparable state-of-the-art tools.sporkis also overall faster than the state of the art, in particular significantly reducing worst-case running times in practice. We demonstrate these properties by replaying 1740 real-world file merges collected from 119 open-source projects, and further demonstrate several key differences betweensporkand the state of the art with in-depth case studies. Simon Larsén, Jean-Rémy Falleri, Benoit Baudry, Martin Monperrus |
IEEE Trans. Software Eng. | 3 |
| 2023 | Automatic Specialization of Third-Party Java DependenciesabstractLarge-scale code reuse significantly reduces both development costs and time. However, the massive share of third-party code in software projects poses new challenges, especially in terms of maintenance and security. In this paper, we propose a novel technique to specialize dependencies of Java projects, based on their actual usage. Given a project and its dependencies, we systematically identify the subset of each dependency that is necessary to build the project, and we remove the rest. As a result of this process, we package each specialized dependency in aJARfile. Then, we generate specialized dependency trees where the original dependencies are replaced by the specialized versions. This allows building the project with significantly less third-party code than the original. As a result, the specialized dependencies become a first-class concept in the software supply chain, rather than a transient artifact in an optimizing compiler toolchain. We implement our technique in a tool calledDepTrim, which we evaluate with 30 notable open-source Java projects.DepTrimspecializes a total of 343 (86.6%) dependencies across these projects, and successfully rebuilds each project with a specialized dependency tree. Moreover, through this specialization,DepTrimremoves a total of 57,444 (42.2%) classes from the dependencies, reducing the ratio of dependency classes to project classes from 8.7$\boldsymbol{\times}$in the original projects to 5.0$\boldsymbol{\times}$after specialization. These novel results indicate that dependency specialization significantly reduces the share of third-party code in Java projects. César Soto-Valero, Deepika Tiwari, Tim Toady, Benoit Baudry |
IEEE Trans. Software Eng. | 4 |
| 2022 | Harvesting Production GraphQL Queries to Detect Schema FaultsabstractGraphQL is a new paradigm to design web APIs. Despite its growing popularity, there are few techniques to verify the implementation of a GraphQL API. We present a new testing approach based on GraphQL queries that are logged while users interact with an application in production. Our core motivation is that production queries capture real usages of the application, and are known to trigger behavior that may not be tested by developers. For each logged query, a test is generated to assert the validity of the GraphQL response with respect to the schema. We implement our approach in a tool called AutoGraphQL, and evaluate it on two real-world case studies that are diverse in their domain and technology stack: an open-source e-commerce application implemented in Python called Saleor, and an industrial case study which is a PHP-based finance website called Frontapp. AutoGraphQL successfully generates test cases for the two applications. The generated tests cover 26.9 % of the Saleor schema, including parts of the API not exercised by the original test suite, as well as 48.7% of the Frontapp schema, detecting 8 schema faults, thanks to production queries. Louise Zetterlund, Deepika Tiwari, Martin Monperrus, Benoit Baudry |
ICST | 4 |
| 2022 | API beauty is in the eye of the clients: 2.2 million Maven dependencies reveal the spectrum of client-API usages
Nicolas Harrand, Amine Benelallam, César Soto-Valero, François Bettega, Olivier Barais, Benoit Baudry |
J. Syst. Softw. | 6 |
| 2022 | Maximizing Error Injection Realism for Chaos Engineering With System CallsabstractIn this article, we present a novel fault injection framework for system call invocation errors, calledPhoebe.Phoebeis unique as follows; First,Phoebeenables developers to have full observability of system call invocations. Second,Phoebegenerates error models that are realistic in the sense that they mimic errors that naturally happen in production. Third,Phoebeis able to automatically conduct experiments to systematically assess the reliability of applications with respect to system call invocation errors in production. We evaluate the effectiveness and runtime overhead ofPhoebeon two real-world applications in a production environment for a single software stack: Java. The results show thatPhoebesuccessfully generates realistic error models and is able to detect important reliability weaknesses with respect to system call invocation errors. To our knowledge, this novel concept of “realistic error injection”, which consists of grounding fault injection on production errors, has never been studied before. Brice Morin, Benoit Baudry, Martin Monperrus |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Production Monitoring to Improve Test SuitesabstractIn this article, we propose to use production executions to improve the quality of testing for certain methods of interest for developers. These methods can be methods that are not covered by the existing test suite or methods that are poorly tested. We devise an approach calledpanktiwhich monitors applications as they execute in production and then automatically generates differential unit tests, as well as derived oracles, from the collected data.pankti’s monitoring and generation focuses on one single programming language, Java. We evaluate it on three real-world, open-source projects: a videoconferencing system, a PDF manipulation library, and an e-commerce application. We show thatpanktiis able to generate differential unit tests by monitoring target methods in production and that the generated tests improve the quality of the test suite of the application under consideration. Deepika Tiwari, Martin Monperrus, Benoit Baudry |
IEEE Trans. Reliab. | 4 |
| 2021 | The Behavioral Diversity of Java JSON LibrariesabstractJSON is an essential file and data format in domains that span scientific computing, web APIs or configuration management. Its popularity has motivated significant software development effort to build multiple libraries to process JSON data. Previous studies focus on performance comparison among these libraries and lack a software engineering perspective. We present the first systematic analysis and comparison of the input / output behavior of 20 JSON libraries, in a single software ecosystem: Java/Maven. We assess behavior diversity by running each library against a curated set of 473 JSON files, including both well-formed and ill-formed files. The main design differences, which influence the behavior of the libraries, relate to the choice of data structure to represent JSON objects and to the encoding of numbers. We observe a remarkable behavioral diversity with ill-formed files, or corner cases such as large numbers or duplicate data. Our unique behavioral assessment of JSON libraries paves the way for a robust processing of ill-formed files, through a multi-version architecture. Nicolas Harrand, Thomas Durieux, David Broman, Benoit Baudry |
ISSRE | 4 |
| 2021 | Duets: A Dataset of Reproducible Pairs of Java Library-ClientsabstractSoftware engineering researchers look for software artifacts to study their characteristics or to evaluate new techniques. In this paper, we introduce Duets, a new dataset of software libraries and their clients. This dataset can be exploited to gain many different insights, such as API usage, usage inputs, or novel observations about the test suites of clients and libraries. Duets is meant to support both static and dynamic analysis. This means that the libraries and the clients compile correctly, they are executable and their test suites pass. The dataset is composed of open-source projects that have more than five stars on GitHub. The final dataset contains 395 libraries and 2,874 clients. Additionally, we provide the raw data that we use to create this dataset, such as 34,560 pom.xml files or the complete file list from 34,560 projects. This dataset can be used to study how libraries are used by their clients or as a list of software projects that successfully build. The client's test suite can be used as an additional verification step for code transformation techniques that modify the libraries. Thomas Durieux, César Soto-Valero, Benoit Baudry |
MSR | 3 |
| 2021 | A longitudinal analysis of bloated Java dependenciesabstractWe study the evolution and impact of bloated dependencies in a single software ecosystem: Java/Maven. Bloated dependencies are third-party libraries that are packaged in the application binary but are not needed to run the application. We analyze the history of 435 Java projects. This historical data includes 48,469 distinct dependencies, which we study across a total of 31,515 versions of Maven dependency trees. Bloated dependencies steadily increase over time, and 89.2% of the direct dependencies that are bloated remain bloated in all subsequent versions of the studied projects. This empirical evidence suggests that developers can safely remove a bloated dependency. We further report novel insights regarding the unnecessary maintenance efforts induced by bloat. We find that 22% of dependency updates performed by developers are made on bloated dependencies, and that Dependabot suggests a similar ratio of updates on bloated dependencies. César Soto-Valero, Thomas Durieux, Benoit Baudry |
ESEC/SIGSOFT FSE | 3 |
| 2021 | A comprehensive study of bloated dependencies in the Maven ecosystemabstractAbstract Build automation tools and package managers have a profound influence on software development. They facilitate the reuse of third-party libraries, support a clear separation between the application’s code and its external dependencies, and automate several software development tasks. However, the wide adoption of these tools introduces new challenges related to dependency management. In this paper, we propose an original study of one such challenge: the emergence of bloated dependencies. Bloated dependencies are libraries that are packaged with the application’s compiled code but that are actually not necessary to build and run the application. They artificially grow the size of the built binary and increase maintenance effort. We propose DepClean, a tool to determine the presence of bloated dependencies in Maven artifacts. We analyze 9,639 Java artifacts hosted on Maven Central, which include a total of 723,444 dependency relationships. Our key result is as follows: 2.7% of the dependencies directly declared are bloated, 15.4% of the inherited dependencies are bloated, and 57% of the transitive dependencies of the studied artifacts are bloated. In other words, it is feasible to reduce the number of dependencies of Maven artifacts to 1/4 of its current count. Our qualitative assessment with 30 notable open-source projects indicates that developers pay attention to their dependencies when they are notified of the problem. They are willing to remove bloated dependencies: 21/26 answered pull requests were accepted and merged by developers, removing 140 dependencies in total: 75 direct and 65 transitive. César Soto-Valero, Nicolas Harrand, Martin Monperrus, Benoit Baudry |
Empir. Softw. Eng. | 4 |
| 2021 | Observability and chaos engineering on system calls for containerized applications in Docker
Jesper Simonsson, Brice Morin, Benoit Baudry, Martin Monperrus |
Future Gener. Comput. Syst. | 4 |
| 2021 | Constraint-based Diversification of JOP GadgetsabstractModern software deployment process produces software that is uniform, and hence vulnerable to large-scale code-reuse attacks, such as Jump-Oriented Programming (JOP) attacks. Compiler-based diversification improves the resilience and security of software systems by automatically generating different assembly code versions of a given program. Existing techniques are efficient but do not have a precise control over the quality, such as the code size or speed, of the generated code variants. This paper introduces Diversity by Construction (DivCon), a constraint-based compiler approach to software diversification. Unlike previous approaches, DivCon allows users to control and adjust the conflicting goals of diversity and code quality. A key enabler is the use of Large Neighborhood Search (LNS) to generate highly diverse assembly code efficiently. For larger problems, we propose a combination of LNS with a structural decomposition of the problem. To further improve the diversification efficiency of DivCon against JOP attacks, we propose an application-specific distance measure tailored to the characteristics of JOP attacks. We evaluate DivCon with 20 functions from a popular benchmark suite for embedded systems. These experiments show that DivCon's combination of LNS and our application-specific distance measure generates binary programs that are highly resilient against JOP attacks (they share between 0.15% to 8% of JOP gadgets) with an optimality gap of 10%. Our results confirm that there is a trade-off between the quality of each assembly code version and the diversity of the entire pool of versions. In particular, the experiments show that DivCon is able to generate binary programs that share a very small number of gadgets, while delivering near-optimal code. For constraint programming researchers and practitioners, this paper demonstrates that LNS is a valuable technique for finding diverse solutions. For security researchers and software engineers, DivCon extends the scope of compiler-based diversification to performance-critical and resource-constrained applications. Rodothea-Myrsini Tsoupidi, Roberto Castañeda Lozano, Benoit Baudry |
J. Artif. Intell. Res. | 3 |
| 2021 | A Chaos Engineering System for Live Analysis and Falsification of Exception-Handling in the JVMabstractSoftware systems contain resilience code to handle those failures and unexpected events happening in production. It is essential for developers to understand and assess the resilience of their systems. Chaos engineering is a technology that aims at assessing resilience and uncovering weaknesses by actively injecting perturbations in production. In this paper, we propose a novel design and implementation of a chaos engineering system in Java calledChaosMachine. It provides a unique and actionable analysis on exception-handling capabilities in production, at the level of try-catch blocks. To evaluate our approach, we have deployedChaosMachineon top of 3 large-scale and well-known Java applications totaling$630k$lines of code. Our results show thatChaosMachinereveals both strengths and weaknesses of the resilience code of a software system at the level of exception handling. Brice Morin, Philipp Haller, Benoit Baudry, Martin Monperrus |
IEEE Trans. Software Eng. | 4 |
| 2020 | Constraint-Based Software Diversification for Efficient Mitigation of Code-Reuse Attacks
Rodothea-Myrsini Tsoupidi, Roberto Castañeda Lozano, Benoit Baudry |
CP | 3 |
| 2020 | An approach and benchmark to detect behavioral changes of commits in continuous integration
Benjamin Danglot, Martin Monperrus, Walter Rudametkin, Benoit Baudry |
Empir. Softw. Eng. | 4 |
| 2020 | Java decompiler diversity and its application to meta-decompilation
Nicolas Harrand, César Soto-Valero, Martin Monperrus, Benoit Baudry |
J. Syst. Softw. | 4 |
| 2020 | Leveraging metamorphic testing to automatically detect inconsistencies in code generator familiesabstractSUMMARY Generative software development has paved the way for the creation of multiple code generators that serve as a basis for automatically generating code to different software and hardware platforms. In this context, the software quality becomes highly correlated to the quality of code generators used during software development. Eventual failures may result in a loss of confidence for the developers, who will unlikely continue to use these generators. It is then crucial to verify the correct behaviour of code generators in order to preserve software quality and reliability. In this paper, we leverage the metamorphic testing approach to automatically detect inconsistencies in code generators via so‐called “metamorphic relations”. We define the metamorphic relation (i.e., test oracle) as a comparison between the variations of performance and resource usage of test suites running on different versions of generated code. We rely on statistical methods to find the threshold value from which an unexpected variation is detected. We evaluate our approach by testing a family of code generators with respect to resource usage and performance metrics for five different target software platforms. The experimental results show that our approach is able to detect, among 95 executed test suites, 11 performance and 15 memory usage inconsistencies. Mohamed Boussaa, Olivier Barais, Gerson Sunyé, Benoit Baudry |
Softw. Test. Verification Reliab. | 4 |
| 2020 | Browser Fingerprinting: A SurveyabstractWith this article, we survey the research performed in the domain of browser fingerprinting, while providing an accessible entry point to newcomers in the field. We explain how this technique works and where it stems from. We analyze the related work in detail to understand the composition of modern fingerprints and see how this technique is currently used online. We systematize existing defense solutions into different categories and detail the current challenges yet to overcome. Pierre Laperdrix, Nataliia Bielova, Benoit Baudry, Gildas Avoine |
ACM Trans. Web | 3 |
| 2019 | Approximate loop unrollingabstractWe introduce Approximate Unrolling, a compiler loop optimization that reduces execution time and energy consumption, exploiting code regions that can endure some approximation and still produce acceptable results. Specifically, this work focuses on counted loops that map a function over the elements of an array. Approximate Unrolling transforms loops similarly to Loop Unrolling. However, unlike its exact counterpart, our optimization does not unroll loops by adding exact copies of the loop's body. Instead, it adds code that interpolates the results of previous iterations. Marcelino Rodriguez-Cancio, Benoît Combemale, Benoit Baudry |
CF | 3 |
| 2019 | Morellian Analysis for Browsers: Making Web Authentication Stronger with Canvas Fingerprinting
Pierre Laperdrix, Gildas Avoine, Benoit Baudry, Nick Nikiforakis |
DIMVA | 3 |
| 2019 | The maven dependency graph: a temporal graph-based representation of maven centralabstractThe Maven Central Repository provides an extraordinary source of data to understand complex architecture and evolution phenomena among Java applications. As of September 6, 2018, this repository includes 2.8M artifacts (compiled piece of code implemented in a JVM-based language), each of which is characterized with metadata such as exact version, date of upload and list of dependencies towards other artifacts. Today, one who wants to analyze the complete ecosystem of Maven artifacts and their dependencies faces two key challenges: (i) this is a huge data set; and (ii) dependency relationships among artifacts are not modeled explicitly and cannot be queried. In this paper, we present the Maven Dependency Graph. This open source data set provides two contributions: a snapshot of the whole Maven Central taken on September 6, 2018, stored in a graph database in which we explicitly model all dependencies; an open source infrastructure to query this huge dataset. Amine Benelallam, Nicolas Harrand, César Soto-Valero, Benoit Baudry, Olivier Barais |
MSR | 4 |
| 2019 | The emergence of software diversity in maven centralabstractMaven artifacts are immutable: an artifact that is uploaded on Maven Central cannot be removed nor modified. The only way for developers to upgrade their library is to release a new version. Consequently, Maven Central accumulates all the versions of all the libraries that are published there, and applications that declare a dependency towards a library can pick any version. In this work, we hypothesize that the immutability of Maven artifacts and the ability to choose any version naturally support the emergence of software diversity within Maven Central. We analyze 1,487,956 artifacts that represent all the versions of 73,653 libraries. We observe that more than 30% of libraries have multiple versions that are actively used by latest artifacts. In the case of popular libraries, more than 50% of their versions are used. We also observe that more than 17% of libraries have several versions that are significantly more used than the other versions. Our results indicate that the immutability of artifacts in Maven Central does support a sustained level of diversity among versions of libraries in the repository. César Soto-Valero, Amine Benelallam, Nicolas Harrand, Olivier Barais, Benoit Baudry |
MSR | 5 |
| 2019 | The Strengths and Behavioral Quirks of Java Bytecode DecompilersabstractDuring compilation from Java source code to bytecode, some information is irreversibly lost. In other words, compilation and decompilation of Java code is not symmetric. Consequently, the decompilation process, which aims at producing source code from bytecode, must establish some strategies to reconstruct the information that has been lost. Modern Java decompilers tend to use distinct strategies to achieve proper decompilation. In this work, we hypothesize that the diverse ways in which bytecode can be decompiled has a direct impact on the quality of the source code produced by decompilers. We study the effectiveness of eight Java decompilers with respect to three quality indicators: syntactic correctness, syntactic distortion and semantic equivalence modulo inputs. This study relies on a benchmark set of 14 real-world open-source software projects to be decompiled (2041 classes in total). Our results show that no single modern decompiler is able to correctly handle the variety of bytecode structures coming from real-world programs. Even the highest ranking decompiler in this study produces syntactically correct output for 84% of classes of our dataset and semantically equivalent code output for 78% of classes. Nicolas Harrand, César Soto-Valero, Martin Monperrus, Benoit Baudry |
SCAM | 4 |
| 2019 | Automatic test improvement with DSpot: a study with ten mature open-source projects
Benjamin Danglot, Oscar Vera-Perez, Benoit Baudry, Martin Monperrus |
Empir. Softw. Eng. | 3 |
| 2019 | Test them all, is it worth it? Assessing configuration sampling on the JHipster Web development stackabstractMany approaches for testing configurable software systems start from the same assumption: it is impossible to test all configurations. This motivated the definition of variability-aware abstractions and sampling techniques to cope with large configuration spaces. Yet, there is no theoretical barrier that prevents the exhaustive testing of all configurations by simply enumerating them if the effort required to do so remains acceptable. Not only this: we believe there is a lot to be learned by systematically and exhaustively testing a configurable system. In this case study, we report on the first ever endeavour to test all possible configurations of the industry-strength, open source configurable software system JHipster, a popular code generator for web applications. We built a testing scaffold for the 26,000+ configurations of JHipster using a cluster of 80 machines during 4 nights for a total of 4,376 hours (182 days) CPU time. We find that 35.70% configurations fail and we identify the feature interactions that cause the errors. We show that sampling strategies (like dissimilarity and 2-wise): (1) are more effective to find faults than the 12 default configurations used in the JHipster continuous integration; (2) can be too costly and exceed the available testing budget. We cross this quantitative analysis with the qualitative assessment of JHipster’s lead developers. Axel Halin, Alexandre Nuttinck, Mathieu Acher, Xavier Devroey, Gilles Perrouin, Benoit Baudry |
Empir. Softw. Eng. | 6 |
| 2019 | A comprehensive study of pseudo-tested methods
Oscar Vera-Perez, Benjamin Danglot, Martin Monperrus, Benoit Baudry |
Empir. Softw. Eng. | 4 |
| 2019 | A snowballing literature study on test amplificationabstractContext: The increasing adoption of test-driven development results in software projects with strong test suites. These suites include a large number of test cases, in which developers embed knowledge about meaningful input data and expected properties in the form of oracles. Objective: This article surveys various works that aim at exploiting this knowledge in order to enhance these manually written tests with respect to an engineering goal (e.g., improve coverage of changes or increase the accuracy of fault localization). While these works rely on various techniques and address various goals, we believe they form an emerging and coherent field of research, and which we call "test amplification". Method: We devised a first set of papers based on our knowledge of the literature (we have been working in software testing for years). Then, we systematically followed the citation graph. Results: This survey is the first that draws a comprehensive picture of the different engineering goals proposed in the literature for test amplification. In particular, we note that the goal of test amplification goes far beyond maximizing coverage only. Conclusion: We believe that this survey will help researchers and practitioners entering this new field to understand more quickly and more deeply the intuitions, concepts and techniques used for test amplification. Benjamin Danglot, Oscar Vera-Perez, Zhongxing Yu, Andy Zaidman, Martin Monperrus, Benoit Baudry |
J. Syst. Softw. | 6 |
| 2019 | Advanced and efficient execution trace management for executable domain-specific modeling languagesabstractExecutable Domain-Specific Modeling Languages (xDSMLs) enable the application of early dynamic verification and validation (V&V) techniques for behavioral models. At the core of such techniques, execution traces are used to represent the evolution of models during their execution. In order to construct execution traces for any xDSML, generic trace metamodels can be used. Yet, regarding trace manipulations, generic trace metamodels lack efficiency in time because of their sequential structure, efficiency in memory because they capture superfluous data, and usability because of their conceptual gap with the considered xDSML. Our contribution is a novel generative approach that defines a multidimensional and domain-specific trace metamodel enabling the construction and manipulation of execution traces for models conforming to a given xDSML. Efficiency in time is improved by providing a variety of navigation paths within traces, while usability and memory are improved by narrowing the scope of trace metamodels to fit the considered xDSML. We evaluated our approach by generating a trace metamodel for fUML and using it for semantic differencing, which is an important V&V technique in the realm of model evolution. Results show a significant performance improvement and simplification of the semantic differencing rules as compared to the usage of a generic trace metamodel. Erwan Bousse, Tanja Mayerhofer, Benoît Combemale, Benoit Baudry |
Softw. Syst. Model. | 4 |
| 2019 | Modeling variability in the video domain: language and experience report
Mauricio Alférez, Mathieu Acher, José A. Galindo, Benoit Baudry, David Benavides 0001 |
Softw. Qual. J. | 4 |
| 2018 | Correctness attraction: a study of stability of software behavior under runtime perturbationabstractCan the execution of software be perturbed without breaking the correctness of the output? In this paper, we devise a protocol to answer this question from a novel perspective. In an experimental study, we observe that many perturbations do not break the correctness in ten subject programs. We call this phenomenon "correctness attraction". The uniqueness of this protocol is that it considers a systematic exploration of the perturbation space as well as perfect oracles to determine the correctness of the output. To this extent, our findings on the stability of software under execution perturbations have a level of validity that has never been reported before in the scarce related work. A qualitative manual analysis enables us to set up the first taxonomy ever of the reasons behind correctness attraction. Benjamin Danglot, Philippe Preux, Benoit Baudry, Martin Monperrus |
ICSE | 3 |
| 2018 | Exhaustive Exploration of the Failure-Oblivious Computing Search Space
Thomas Durieux, Youssef Hamadi, Zhongxing Yu, Benoit Baudry, Martin Monperrus |
ICST | 4 |
| 2018 | Descartes: a PITest engine to detect pseudo-tested methods: tool demonstrationabstractDescartes is a tool that implements extreme mutation operators and aims at finding pseudo-tested methods in Java projects. It leverages the efficient transformation and runtime features of PITest. The demonstration compares Descartes with Gregor, the default mutation engine provided by PITest, in a set of real open source projects. It considers the execution time, number of mutants created and the relationship between the mutation scores produced by both engines. It provides some insights on the main features exposed byDescartes. Oscar Vera-Perez, Martin Monperrus, Benoit Baudry |
ASE | 3 |
| 2018 | Engineering Software Diversity: a Model-Based Approach to Systematically Diversify CommunicationsabstractAutomated diversity is a promising mean of increasing the security of software systems. However, current automated diversity techniques operate at the bottom of the software stack (operating system and compiler), yielding a limited amount of diversity. We present a novel Model-Driven Engineering approach to the diversification of communicating systems, building on abstraction, model transformations and code generation. This approach generates significant amounts of diversity with a low overhead, and addresses a large number of communicating systems, including small communicating devices. Brice Morin, Jakob Høgenes, Nicolas Harrand, Benoit Baudry |
MoDELS | 5 |
| 2018 | Detection and analysis of behavioral T-patterns in debugging activitiesabstractA growing body of research in empirical software engineering applies recurrent patterns analysis in order to make sense of the developers' behavior during their interactions with IDEs. However, the exploration of hidden real-time structures of programming behavior remains a challenging task. In this paper, we investigate the presence of temporal behavioral patterns (T-patterns) in debugging activities using the THEME software. Our preliminary exploratory results show that debugging activities are strongly correlated with code editing, file handling, window interactions and other general types of programming activities. The validation of our T-patterns detection approach demonstrates that debugging activities are performed on the basis of repetitive and well-organized behavioral events. Furthermore, we identify a large set of T-patterns that associate debugging activities with build success, which corroborates the positive impact of debugging practices on software development. César Soto-Valero, Johann Bourcier, Benoit Baudry |
MSR | 3 |
| 2018 | Reverse engineering language product lines from existing DSL variantsabstractThe use of domain-specific languages (DSL) has become a successful technique for developing complex systems. Moreover, we can find different DSLs variants adapted to specific purposes that share some features. The challenge for language designers is to take advantage of the commonalities between DSLs variants by reusing previously defined language constructs [7]. To tackle this, the research community in software language engineering proposed to apply Software Product Line (SPLs) techniques in the construction of DSLs [4, 6] leading to the notion of Language Product Pines (LPLs) [3, 7]. David Méndez-Acuña, José A. Galindo, Benoît Combemale, Arnaud Blouin, Benoit Baudry |
SPLC | 5 |
| 2018 | Code{strata} Sonifying Software ComplexityabstractCode{strata} is an interdisciplinary collaboration between art studies researchers (Rennes 2) and computer scientists (INRIA, KTH). It is a sound installation: a computer system unit made of concrete that sits on a wooden desk. The purpose of this project is to question the opacity and simplicity of high-level interfaces used in daily gestures. It takes the form of a 3-D sonification of a full software trace that is collected when performing a copy and paste command in a simple text editor. The user may hear, through headphones, a poetic interpretation of what happens in a computer, behind the of graphical interfaces. The sentence 'Copy and paste' is played back in as many pieces as there are nested functions called during the execution of the command. Denez Thomas, Nicolas Harrand, Benoit Baudry, Bruno Bossis |
TEI | 3 |
| 2018 | Hiding in the Crowd: an Analysis of the Effectiveness of Browser Fingerprinting at Large ScaleabstractBrowser fingerprinting is a stateless technique, which consists in collecting a wide range of data about a device through browser APIs. Past studies have demonstrated that modern devices present so much diversity that fingerprints can be exploited to identify and track users online. With this work, we want to evaluate if browser fingerprinting is still effective at uniquely identifying a large group of users when analyzing millions of fingerprints over a few months. We collected 2,067,942 browser fingerprints from one of the top 15 French websites. The analysis of this novel dataset sheds a new light on the ever-growing browser fingerprinting domain. The key insight is that the percentage of unique fingerprints in our dataset is much lower than what was reported in the past: only 33.6% of fingerprints are unique by opposition to over 80% in previous studies. We show that non-unique fingerprints tend to be fragile. If some features of the fingerprint change, it is very probable that the fingerprint will become unique. We also confirm that the current evolution of web technologies is benefiting users» privacy significantly as the removal of plugins brings down substantively the rate of unique desktop machines. Alejandro Gómez-Boix, Pierre Laperdrix, Benoit Baudry |
WWW | 3 |
| 2018 | Correctness attraction: a study of stability of software behavior under runtime perturbation
Benjamin Danglot, Philippe Preux, Benoit Baudry, Martin Monperrus |
Empir. Softw. Eng. | 3 |
| 2018 | User interface design smell: Automatic detection and refactoring of Blob listeners
Arnaud Blouin, Valéria Lelli, Benoit Baudry, Fabien Coulon |
Inf. Softw. Technol. | 3 |
| 2018 | Omniscient debugging for executable DSLs
Erwan Bousse, Dorian Leroy, Benoît Combemale, Manuel Wimmer, Benoit Baudry |
J. Syst. Softw. | 5 |
| 2017 | Reverse engineering language product lines from existing DSL variants
David Méndez-Acuña, José A. Galindo, Benoît Combemale, Arnaud Blouin, Benoit Baudry |
J. Syst. Softw. | 5 |
| 2017 | Automated extraction of product comparison matrices from informal product descriptions
Sana Ben Nasr, Guillaume Bécan, Mathieu Acher, João Bosco Ferreira Filho, Nicolas Sannier, Benoit Baudry, Jean-Marc Davril |
J. Syst. Softw. | 6 |
| 2016 | Automatic non-functional testing of code generators familiesabstractThe intensive use of generative programming techniques provides an elegant engineering solution to deal with the heterogeneity of platforms and technological stacks. The use of domain-specific languages for example, leads to the creation of numerous code generators that automatically translate highlevel system specifications into multi-target executable code. Producing correct and efficient code generator is complex and error-prone. Although software designers provide generally high-level test suites to verify the functional outcome of generated code, it remains challenging and tedious to verify the behavior of produced code in terms of non-functional properties. This paper describes a practical approach based on a runtime monitoring infrastructure to automatically check the potential inefficient code generators. This infrastructure, based on system containers as execution platforms, allows code-generator developers to evaluate the generated code performance. We evaluate our approach by analyzing the performance of Haxe, a popular high-level programming language that involves a set of cross-platform code generators. Experimental results show that our approach is able to detect some performance inconsistencies that reveal real issues in Haxe code generators. Mohamed Boussaa, Olivier Barais, Benoit Baudry, Gerson Sunyé |
GPCE | 3 |
| 2016 | Reverse-Engineering Reusable Language Modules from Legacy Domain-Specific Languages
David Méndez-Acuña, José A. Galindo, Benoît Combemale, Arnaud Blouin, Benoit Baudry, Gurvan Le Guernic |
ICSR | 5 |
| 2016 | Puzzle: A Tool for Analyzing and Extracting Specification Clones in DSLs
David Méndez-Acuña, José A. Galindo, Benoît Combemale, Arnaud Blouin, Benoit Baudry |
ICSR | 5 |
| 2016 | Automatic microbenchmark generation to prevent dead code elimination and constant foldingabstractMicrobenchmarking evaluates, in isolation, the execution time of small code segments that play a critical role in large applications. The accuracy of a microbenchmark depends on two critical tasks: wrap the code segment into a payload that faithfully recreates the execution conditions of the large application; build a scaffold that runs the payload a large number of times to get a statistical estimate of the execution time. While recent frameworks such as the Java Microbenchmark Harness (JMH) address the scaffold challenge, developers have very limited support to build a correct payload. This work focuses on the automatic generation of payloads, starting from a code segment selected in a large application. Our generative technique prevents two of the most common mistakes made in microbenchmarks: dead code elimination and constant folding. A microbenchmark is such a small program that can be “over-optimized” by the JIT and result in distorted time measures, if not designed carefully. Our technique automatically extracts the segment into a compilable payload and generates additional code to prevent the risks of “over-optimization”. The whole approach is embedded in a tool called AutoJMH, which generates payloads for JMH scaffolds. We validate the capabilities AutoJMH, showing that the tool is able to process a large percentage of segments in real programs. We also show that AutoJMH can match the quality of payloads handwritten by performance experts and outperform those written by professional Java developers without experience in microbenchmarking. Marcelino Rodriguez-Cancio, Benoît Combemale, Benoit Baudry |
ASE | 3 |
| 2016 | libmask: Protecting browser JIT engines from the devil in the constantsabstractJavaScript (JS) engines are virtual machines that execute JavaScript code. These engines find frequent application in web browsers like Google Chrome, Mozilla Firefox, Microsoft Internet Explorer and Apple Safari. Since, the purpose of a JS engine is to produce executable code, it cannot be run in a non-executable environment, and is susceptible to attacks like Just-in-Time (JIT) Spraying, which embed return-oriented programming (ROP) gadgets in arithmetic or logical instructions as immediate offsets. This paper introduces libmask, a JIT compiler extension to prevent the JIT-spraying attacks as an effective alternative to XOR based constant blinding. libmask transforms constants into global variables and marks the memory area for these global variables as read only. Hence, any constant is referred to by a memory address making exploitation of arithmetic and logical instructions more difficult. Further, these memory addresses are randomized to further harden the security. The scheme has been implemented and evaluated as a librddy extension to Google V8 scripting engine with optimizations that contain performance overhead and make libmask a feasible approach. We demonstrate that libmask masks all the constants in JITed code, and effectively raise the bar for JIT-spray and JITROP attacks. The average overhead incurred upon memory is less than 300 kilobytes, while in most benchmarks the memory overhead is less than 10 KB. The average performance overhead observed with optimizations measures is 5.31%. Further, this new approach shows a modest performance improvement over currently deployed constant blinding technique in Google V8. Abhinav, Mohit Mishra, Benoit Baudry |
PST | 3 |
| 2016 | NOTICE: A Framework for Non-Functional Testing of CompilersabstractGenerally, compiler users apply different optimizations to generate efficient code with respect to non-functional properties such as energy consumption, execution time, etc. However, due to the huge number of optimizations provided by modern compilers, finding the best optimization sequence for a specific objective and a given program is more and more challenging. This paper proposes NOTICE, a component-based framework for non-functional testing of compilers through the monitoring of generated code in a controlled sand-boxing environment. We evaluate the effectiveness of our approach by verifying the optimizations performed by the GCC compiler. Our experimental results show that our approach is able to auto-tune compilers according to user requirements and construct optimizations that yield to better performance results than standard optimization levels. We also demonstrate that NOTICE can be used to automatically construct optimization levels that represent optimal trade-offs between multiple non-functional properties such as execution time and resource usage requirements. Mohamed Boussaa, Olivier Barais, Benoit Baudry, Gerson Sunyé |
QRS | 3 |
| 2016 | Beauty and the Beast: Diverting Modern Web Browsers to Build Unique Browser FingerprintsabstractInternational audience Pierre Laperdrix, Walter Rudametkin, Benoit Baudry |
IEEE Symposium on Security and Privacy | 3 |
| 2016 | Exploiting the enumeration of all feature model configurations: a new perspective with distributed computingabstractFeature models are widely used to encode the configurations of a software product line in terms of mandatory, optional and exclusive features as well as propositional constraints over the features. Numerous computationally expensive procedures have been developed to model check, test, configure, debug, or compute relevant information of feature models. In this paper we explore the possible improvement of relying on the enumeration of all configurations when performing automated analysis operations. We tackle the challenge of how to scale the existing enumeration techniques by relying on distributed computing. We show that the use of distributed computing techniques might offer practical solutions to previously unsolvable problems and opens new perspectives for the automated analysis of software product lines. José A. Galindo, Mathieu Acher, Juan Manuel Tirado, Cristian Vidal Silva, Benoit Baudry, David Benavides 0001 |
SPLC | 5 |
| 2016 | Leveraging Software Product Lines Engineering in the development of external DSLs: A systematic literature review
David Méndez-Acuña, José A. Galindo, Thomas Degueule, Benoît Combemale, Benoit Baudry |
Comput. Lang. Syst. Struct. | 5 |
| 2016 | Breathing ontological knowledge into feature model synthesis: an empirical study
Guillaume Bécan, Mathieu Acher, Benoit Baudry, Sana Ben Nasr |
Empir. Softw. Eng. | 3 |
| 2016 | Practical minimization of pairwise-covering test configurations using constraint programming
Aymeric Hervieu, Dusica Marijan, Arnaud Gotlieb, Benoit Baudry |
Inf. Softw. Technol. | 4 |
| 2016 | B-Refactoring: Automatic test code refactoring to improve dynamic analysis
Jifeng Xuan, Benoit Cornu, Matias Martinez, Benoit Baudry, Lionel Seinturier, Martin Monperrus |
Inf. Softw. Technol. | 4 |
| 2016 | ScapeGoat: Spotting abnormal resource usage in component-based reconfigurable software systems
Inti Y. Gonzalez-Herrera, Johann Bourcier, Erwan Daubert, Walter Rudametkin, Olivier Barais, François Fouquet, Jean-Marc Jézéquel, Benoit Baudry |
J. Syst. Softw. | 8 |
| 2015 | A Generative Approach to Define Rich Domain-Specific Trace Metamodels
Erwan Bousse, Tanja Mayerhofer, Benoît Combemale, Benoit Baudry |
ECMFA | 4 |
| 2015 | Classifying and Qualifying GUI DefectsabstractGraphical user interfaces (GUIs) are integral parts of software systems that require interactions from their users. Software testers have paid special attention to GUI testing in the last decade, and have devised techniques that are effective in finding several kinds of GUI errors. However, the introduction of new types of interactions in GUIs (e.g., direct manipulation) presents new kinds of errors that are not targeted by current testing techniques. We believe that to advance GUI testing, the community needs a comprehensive and high level GUI fault model, which incorporates all types of interactions. The work detailed in this paper establishes 4 contributions: 1) A GUI fault model designed to identify and classify GUI faults. 2) An empirical analysis for assessing the relevance of the proposed fault model against failures found in real GUIs. 3) An empirical assessment of two GUI testing tools (i.e. GUITAR and Jubula) against those failures. 4) GUI mutants we've developed according to our fault model. These mutants are freely available and can be reused by developers for benchmarking their GUI testing tools. Valéria Lelli, Arnaud Blouin, Benoit Baudry |
ICST | 3 |
| 2015 | Discovering model transformation pre-conditions using automatically generated test modelsabstractSpecifying a model transformation is challenging as it must be able to give a meaningful output for any input model in a possibly infinite modeling domain. Transformation pre-conditions constrain the input domain by rejecting input models that are not meant to be transformed by a model transformation. This paper presents a systematic approach to discover such pre-conditions when it is hard for a human developer to foresee complex graphs of objects that are not meant to be transformed. The approach is based on systematically generating a finite number of test models using our tool, PRAMANA to first cover the input domain based on input domain partitioning. Tracing a transformation's execution reveals why some pre-conditions are missing. Using a benchmark transformation from simplified UML class diagram models to RDBMS models we discover new pre-conditions that were not initially specified. Jean-Marie Mottu, Sagar Sen, Juan José Cadavid, Benoit Baudry |
ISSRE | 4 |
| 2015 | Product lines can jeopardize their trade secretsabstractWhat do you give for free to your competitor when you exhibit a product line? This paper addresses this question through several cases in which the discovery of trade secrets of a product line is possible and can lead to severe consequences. That is, we show that an outsider can understand the variability realization and gain either confidential business information or even some economical direct advantage. For instance, an attacker can identify hidden constraints and bypass the product line to get access to features or copyrighted data. This paper warns against possible naive modeling, implementation, and testing of variability leading to the existence of product lines that jeopardize their trade secrets. Our vision is that defensive methods and techniques should be developed to protect specifically variability – or at least further complicate the task of reverse engineering it. Mathieu Acher, Guillaume Bécan, Benoît Combemale, Benoit Baudry, Jean-Marc Jézéquel |
ESEC/SIGSOFT FSE | 4 |
| 2015 | MatrixMiner: a red pill to architect informal product descriptions in the matrixabstractDomain analysts, product managers, or customers aim to capture the important features and differences among a set of related products. A case-by-case reviewing of each product description is a laborious and time-consuming task that fails to deliver a condensed view of a product line. This paper introduces MatrixMiner: a tool for automatically synthesizing product comparison matrices (PCMs) from a set of product descriptions written in natural language. MatrixMiner is capable of identifying and organizing features and values in a PCM – despite the informality and absence of structure in the textual descriptions of products. Our empirical results of products mined from BestBuy show that the synthesized PCMs exhibit numerous quantitative, comparable information. Users can exploit MatrixMiner to visualize the matrix through a Web editor and review, refine, or complement the cell values thanks to the traceability with the original product descriptions and technical specifications. Sana Ben Nasr, Guillaume Bécan, Mathieu Acher, João Bosco Ferreira Filho, Benoit Baudry, Nicolas Sannier, Jean-Marc Davril |
ESEC/SIGSOFT FSE | 5 |
| 2015 | Supporting efficient and advanced omniscient debugging for xDSMLsabstractOmniscient debugging is a promising technique that relies on execution traces to enable free traversal of the states reached by a system during an execution. While some General-Purpose Languages (GPLs) already have support for omniscient debugging, developing such a complex tool for any executable Domain-Specific Modeling Language (xDSML) remains a challenging and error prone task. A solution to this problem is to define a generic omniscient debugger for all xDSMLs. However, generically supporting any xDSML both compromises the efficiency and the usability of such an approach. Our contribution relies on a partly generic omniscient debugger supported by generated domain-specific trace management facilities. Being domain-specific, these facilities are tuned to the considered xDSML for better efficiency. Usability is strengthened by providing multidimensional omniscient debugging. Results show that our approach is on average 3.0 times more efficient in memory and 5.03 more efficient in time when compared to a generic solution that copies the model at each step. Erwan Bousse, Jonathan Corley, Benoît Combemale, Jeffrey G. Gray, Benoit Baudry |
SLE | 5 |
| 2015 | Assessing product line derivation operators applied to Java source code: an empirical studyabstractProduct Derivation is a key activity in Software Product Line Engineering. During this process, derivation operators modify or create core assets (e.g., model elements, source code instructions, components) by adding, removing or substituting them according to a given configuration. The result is a derived product that generally needs to conform to a programming or modeling language. Some operators lead to invalid products when applied to certain assets, some others do not; knowing this in advance can help to better use them, however this is challenging, specially if we consider assets expressed in extensive and complex languages such as Java. In this paper, we empirically answer the following question: which product line operators, applied to which program elements, can synthesize variants of programs that are incorrect, correct or perhaps even conforming to test suites? We implement source code transformations, based on the derivation operators of the Common Variability Language. We automatically synthesize more than 370,000 program variants from a set of 8 real large Java projects (up to 85,000 lines of code), obtaining an extensive panorama of the sanity of the operations. João Bosco Ferreira Filho, Simon Allier, Olivier Barais, Mathieu Acher, Benoit Baudry |
SPLC | 5 |
| 2015 | An analysis of metamodeling practices for MOF and OCL
Juan José Cadavid, Benoît Combemale, Benoit Baudry |
Comput. Lang. Syst. Struct. | 3 |
| 2015 | Assessing the use of slicing-based visualizing techniques on the understanding of large metamodels
Arnaud Blouin, Naouel Moha, Benoit Baudry, Houari Sahraoui, Jean-Marc Jézéquel |
Inf. Softw. Technol. | 3 |
| 2015 | Kompren: modeling and generating model slicers
Arnaud Blouin, Benoît Combemale, Benoit Baudry, Olivier Beaudoux |
Softw. Syst. Model. | 3 |
| 2015 | Generating counterexamples of model-based software product lines
João Bosco Ferreira Filho, Olivier Barais, Mathieu Acher, Jérôme Le Noir, Axel Legay, Benoit Baudry |
Int. J. Softw. Tools Technol. Transf. | 6 |
| 2015 | Towards an automation of the mutation analysis dedicated to model transformationabstractSummary A benefit of model‐driven engineering relies on the automatic generation of artefacts from high‐level models through intermediary levels using model transformations. In such a process, the input must be well designed, and the model transformations should be trustworthy. Because of the specificities of models and transformations, classical software test techniques have to be adapted. Among these techniques, mutation analysis has been ported, and a set of mutation operators has been defined. However, it currently requires considerable manual work and suffers from the test data set improvement activity. This activity is a difficult and time‐consuming job and reduces the benefits of the mutation analysis. This paper addresses the test data set improvement activity. Model transformation traceability in conjunction with a model of mutation operators and a dedicated algorithm allow to automatically or semi‐automatically produce improved test models. The approach is validated and illustrated in two case studies written in Kermeta.Copyright © 2014 John Wiley & Sons, Ltd. Vincent Aranega, Jean-Marie Mottu, Anne Etien, Thomas Degueule, Benoit Baudry, Jean-Luc Dekeyser |
Softw. Test. Verification Reliab. | 5 |
| 2015 | Special issue for the ICST 2013 conferenceabstractThis special issue includes extended versions of four of the best papers of ICST 2013. The conference attracted over 150 submissions, including both research and industry papers, which demonstrates the strong ongoing interest in the field. Each submission was evaluated by at least three members of the Technical Program Committee. Based on these reviews, and on extensive on-line discussions involving the entire technical Program Committee, we finally accepted 30 research papers and eight industry papers, two of which were short. The selected papers cover a variety of topics, including test-input generation, formal verification, mutation testing, debugging, fault localization and repair, concurrency testing, model based testing, test-case selection, prioritization and minimization, program analysis, and crowdsourcing based testing. After selecting these 30 papers, the entire Technical Program Committee voted to select the best papers amongst the ones accepted. The vote focused on a selection of papers that had at least one strong champion and did not receive any negative scores. Votes had to consider the scores, the relevance of the paper, and the originality of the work. The five papers that received the highest number of votes were invited for this special issue. The authors of one of the papers declined the invitation; revised versions of the other four papers went through the standard STVR review process and are included here. In the rest of this foreword, we summarize these four contributions. In the paper “Are Concurrent Coverage Metrics Effective for Testing: A Comprehensive Empirical Investigation,” Hong and colleagues explore the impact of concurrent coverage metrics on testing effectiveness. They also examine the relationship between coverage, fault detection, and test suite size. Their results (1) indicate that the metrics are moderate to strong predictors of concurrent testing effectiveness and (2) highlight the need for additional work on concurrent coverage. In the paper ”Coverage-Based Regression Test Case Selection, Minimisation and Prioritisation: An Industrial Case Study,” Di Nardo and colleagues apply coverage-based regression testing techniques on a real-world system with real regression faults. Their main insight is that test suite minimization performed using finer grained coverage criteria can provide a good trade-off between savings in execution cost and fault-detection capability. In the paper “CHECK-THEN-ACT Misuse of Java Concurrent Collections,” Lin and colleagues present an extensive empirical study of CHECK-THEN-ACT idioms in Java concurrent collections. Their analysis of 6.4M lines of code that use Java concurrent collections shows that (1) CHECK-THEN-ACT idioms are commonly misused in practice, and (2) correcting them is important. In the paper “Defect Prediction as a Multi-Objective Optimization Problem,” Canfora and colleagues formalize the defect prediction problem as a multi-objective optimization problem. Their multi-objective approach allows software engineers to choose predictors that achieve a specific trade-off between the number of likely defect-prone classes and the number of lines of code to be analyzed/tested. Many members of our community contributed to this special issue and we would like to extend our most heartfelt thanks to everyone who worked hard to make this possible. In particular, we want to thank the authors, who devoted significant time and effort to extending their ICST papers. We also want to thank the reviewers, who wrote detailed reviews and provided extremely valuable feedback that the authors took into account when preparing the final version of the papers that are presented here. Benoit Baudry, Alessandro Orso |
Softw. Test. Verification Reliab. | 1 |
| 2014 | Deriving Usage Model Variants for Model-Based Testing: An Industrial Case StudyabstractThe strong cost pressure of the market and safety issues faced by aerospace industry affect the development. Suppliers are forced to continuously optimize their life-cycle processes to facilitate the development of variants for different customers and shorten time to market. Additionally, industrial safety standards like RTCA/DO-178C require high efforts for testing single products. A suitably organized test process for Product Lines (PL) can meet standards. In this paper, we propose an approach that adopts Model-based Testing (MBT) for PL. Usage models, a widely used MBT formalism that provides automatic test case generation capabilities, are equipped with variability information such that usage model variants can be derived for a given set of features. The approach is integrated in the professional MBT tool MaTeLo. We report on our experience gained from an industrial case study in the aerospace domain. Hamza Samih, Hélène Le Guen, Ralf Bogusch, Mathieu Acher, Benoit Baudry |
ICECCS | 5 |
| 2014 | On Analyzing the Topology of Commit Histories in Decentralized Version Control SystemsabstractEmpirical analysis of software repositories usually deals with linear histories derived from centralized versioning systems. Decentralized version control systems allow a much richer structure of commit histories, which presents features that are typical of complex graph models. In this paper we bring some evidences of how the very structure of these commit histories carries relevant information about the distributed development process. By means of a novel data structure that we formally define, we analyze the topological characteristics of commit graphs of a sample of GIT projects. Our findings point out the existence of common recurrent structural patterns which identically occur in different projects and can be consider building blocks of distributed collaborative development. Marco Biazzini, Martin Monperrus, Benoit Baudry |
ICSME | 3 |
| 2014 | Tailored source code transformations to synthesize computationally diverse program variantsabstractThe predictability of program execution provides attackers a rich source of knowledge who can exploit it to spy or remotely control the program. Moving target defense ad- dresses this issue by constantly switching between many di- verse variants of a program, which reduces the certainty that an attacker can have about the program execution. The ef- fectiveness of this approach relies on the availability of a large number of software variants that exhibit dierent ex- ecutions. However, current approaches rely on the natural diversity provided by o-the-shelf components, which is very limited. In this paper, we explore the automatic synthe- sis of large sets of program variants, called sosies. Sosies provide the same expected functionality as the original pro- gram, while exhibiting dierent executions. They are said to be computationally diverse. This work addresses two objectives: comparing dierent transformations for increasing the likelihood of sosie synthe- sis (densifying the search space for sosies); demonstrating computation diversity in synthesized sosies. We synthesized 30 184 sosies in total, for 9 large, real-world, open source ap- plications. For all these programs we identied one type of program analysis that systematically increases the density of sosies; we measured computation diversity for sosies of 3 programs and found diversity in method calls or data in more than 40% of sosies. This is a step towards controlled massive unpredictability of software. Benoit Baudry, Simon Allier, Martin Monperrus |
ISSTA | 1 |
| 2014 | A variability-based testing approach for synthesizing video sequencesabstractA key problem when developing video processing software is the difficulty to test different input combinations. In this paper, we present VANE, a variability-based testing approach to derive video sequence variants. The ideas of VANE are i) to encode in a variability model what can vary within a video sequence; ii) to exploit the variability model to generate testable configurations; iii) to synthesize variants of video sequences corresponding to configurations. VANE computes T-wise covering sets while optimizing a function over attributes. Also, we present a preliminary validation of the scalability and practicality of VANE in the context of an industrial project involving the test of video processing algorithms. José A. Galindo, Mauricio Alférez, Mathieu Acher, Benoit Baudry, David Benavides 0001 |
ISSTA | 4 |
| 2014 | Automating the formalization of product comparison matricesabstractProduct Comparison Matrices (PCMs) form a rich source of data for comparing a set of related and competing products over numerous features. Despite their apparent simplicity, PCMs contain heterogeneous, ambiguous, uncontrolled and partial information that hinders their efficient exploitations. In this paper, we formalize PCMs through model-based automated techniques and develop additional tooling to support the edition and re-engineering of PCMs. 20 participants used our editor to evaluate the PCM metamodel and automated transformations. The results over 75 PCMs from Wikipedia show that (1) a significant proportion of the formalization of PCMs can be automated -- 93.11% of the 30061 cells are correctly formalized; (2) the rest of the formalization can be realized by using the editor and mapping cells to existing concepts of the metamodel. The automated approach opens avenues for engaging a community in the mining, re-engineering, edition, and exploitation of PCMs that now abound on the Internet. Guillaume Bécan, Nicolas Sannier, Mathieu Acher, Olivier Barais, Arnaud Blouin, Benoit Baudry |
ASE | 6 |
| 2014 | Scalable Armies of Model Clones through Data Sharing
Erwan Bousse, Benoît Combemale, Benoit Baudry |
MoDELS | 3 |
| 2014 | An Approach to Derive Usage Models Variants for Model-Based Testing
Hamza Samih, Hélène Le Guen, Ralf Bogusch, Mathieu Acher, Benoit Baudry |
ICTSS | 5 |
| 2014 | INCREMENT: A Mixed MDE-IR Approach for Regulatory Requirements Modeling and Analysis
Nicolas Sannier, Benoit Baudry |
REFSQ | 2 |
| 2014 | Customization and 3D printing: a challenging playground for software product linesabstract3D printing is gaining more and more momentum to build customized product in a wide variety of fields. We conduct an exploratory study of Thingiverse, the most popular Website for sharing user-created 3D design files, in order to establish a possible connection with software product line (SPL) engineering. We report on the socio-technical aspects and current practices for modeling variability, implementing variability, configuring and deriving products, and reusing artefacts. We provide hints that SPL-alike techniques are practically used in 3D printing and thus relevant. Finally, we discuss why the customization in the 3D printing field represents a challenging playground for SPL engineering. Mathieu Acher, Benoit Baudry, Olivier Barais, Jean-Marc Jézéquel |
SPLC | 2 |
| 2014 | Moving toward product line engineering in a nuclear industry consortiumabstractNuclear power plants are some of the most sophisticated and complex energy systems ever designed. These systems perform safety critical functions and must conform to national safety institutions and international regulations. In many cases, regulatory documents provide very high level and ambiguous requirements that leave a large margin for interpretation. As the French nuclear industry is now seeking to spread its activities outside France, it is but necessary to master the ins and the outs of the variability between countries safety culture and regulations. This sets both an industrial and a scientific challenge to introduce and propose a product line engineering approach to an unaware industry whose safety culture is made of interpretations, specificities, and exceptions. Sana Ben Nasr, Nicolas Sannier, Mathieu Acher, Benoit Baudry |
SPLC | 4 |
| 2014 | Slicing-Based Techniques for Visualizing Large MetamodelsabstractIn model-driven engineering, a model describes an aspect of a system. A model conforms to a metamodel that defines the concepts and relationships of a given domain. Metamodels are thus corner-stones of various meta-modeling activities that require a good understanding of the metamodels or parts of them. Current metamodel editing tools are based on standard visualization and navigation features, such as physical zooms. However, as soon as metamodels become larger, navigating through large metamodels becomes a tedious task that hinders their understanding. In this work, we promote the use of model slicing techniques to build visualization techniques dedicated to metamodels. We propose an approach based on model slicing, inspired from program slicing, to build interactive visualization techniques dedicated to metamodels. These techniques permit users to focus on metamodel elements of interest, which aims at improving the understand ability. This approach is implemented in a metamodel visualizer, called Explen. Arnaud Blouin, Naouel Moha, Benoit Baudry, Houari Sahraoui |
VISSOFT | 3 |
| 2014 | Model-based testing of global properties on large-scale distributed systems
Gerson Sunyé, Eduardo C. de Almeida, Yves Le Traon, Benoit Baudry, Jean-Marc Jézéquel |
Inf. Softw. Technol. | 4 |
| 2013 | From comparison matrix to Variability Model: The Wikipedia case studyabstractProduct comparison matrices (PCMs) provide a convenient way to document the discriminant features of a family of related products and now abound on the internet. Despite their apparent simplicity, the information present in existing PCMs can be very heterogeneous, partial, ambiguous, hard to exploit by users who desire to choose an appropriate product. Variability Models (VMs) can be employed to formulate in a more precise way the semantics of PCMs and enable automated reasoning such as assisted configuration. Yet, the gap between PCMs and VMs should be precisely understood and automated techniques should support the transition between the two. In this paper, we propose variability patterns that describe PCMs content and conduct an empirical analysis of 300+ PCMs mined from Wikipedia. Our findings are a first step toward better engineering techniques for maintaining and configuring PCMs. Nicolas Sannier, Mathieu Acher, Benoit Baudry |
ASE | 3 |
| 2013 | Automatically Searching for Metamodel Well-Formedness Rules in Examples and Counter-Examples
Martin Faunes, Juan José Cadavid, Benoit Baudry, Houari Sahraoui, Benoît Combemale |
MoDELS | 3 |
| 2013 | Empirical evidence of large-scale diversity in API usage of object-oriented softwareabstractIn this paper, we study how object-oriented classes are used across thousands of software packages. We concentrate on “usage diversity”, defined as the different statically observable combinations of methods called on the same object. We present empirical evidence that there is a significant usage diversity for many classes. For instance, we observe in our dataset that Java's String is used in 2460 manners. We discuss the reasons of this observed diversity and the consequences on software engineering knowledge and research. Diego Mendez 0002, Benoit Baudry, Martin Monperrus |
SCAM | 2 |
| 2013 | Reifying Concurrency for Executable Metamodeling
Benoît Combemale, Julien Deantoni, Matias Vara Larsen, Frédéric Mallet, Olivier Barais, Benoit Baudry, Robert B. France |
SLE | 6 |
| 2013 | Generating counterexamples of model-based software product lines: an exploratory studyabstractModel-based Software Product Line (MSPL) engineering aims at deriving customized models corresponding to individual products of a family. MSPL approaches usually promote the joint use of a variability model, a base model expressed in a specific formalism, and a realization layer that maps variation points to model elements. The design space of an MSPL is extremely complex to manage for the engineer, since the number of variants may be exponential and the derived product models have to be conformant to numerous well-formedness and business rules. In this paper, the objective is to provide a way to generate MSPLs, called counterexamples, that can produce invalid product models despite a valid configuration in the variability model. We provide a systematic and automated process, based on the Common Variability Language (CVL), to randomly search the space of MSPLs for a specific formalism. We validate the effectiveness of this process for three formalisms at different scales (up to 247 metaclasses and 684 rules). We also explore and discuss how counterexamples could guide practitioners when customizing derivation engines, when implementing checking rules that prevent early incorrect CVL models, or simply when specifying an MSPL. João Bosco Ferreira Filho, Olivier Barais, Mathieu Acher, Benoit Baudry, Jérôme Le Noir |
SPLC | 4 |
| 2013 | Soa Antipatterns: an Approach for their Specification and DetectionabstractLike any other large and complex software systems, Service-Based Systems (SBSs) must evolve to fit new user requirements and execution contexts. The changes resulting from the evolution of SBSs may degrade their design and quality of service (QoS) and may often cause the appearance of common poor solutions in their architecture, called antipatterns, in opposition to design patterns, which are good solutions to recurring problems. Antipatterns resulting from these changes may hinder the future maintenance and evolution of SBSs. The detection of antipatterns is thus crucial to assess the design and QoS of SBSs and facilitate their maintenance and evolution. However, methods and techniques for the detection of antipatterns in SBSs are still in their infancy despite their importance. In this paper, we introduce a novel and innovative approach supported by a framework for specifying and detecting antipatterns in SBSs. Using our approach, we specify 10 well-known and common antipatterns, including Multi Service and Tiny Service, and automatically generate their detection algorithms. We apply and validate the detection algorithms in terms of precision and recall two systems developed independently, (1) Home-Automation, an SBS with 13 services, and (2) FraSCAti, an open-source implementation of the Service Component Architecture (SCA) standard with more than 100 services. This validation demonstrates that our approach enables the specification and detection of Service Oriented Architecture (SOA) antipatterns with an average precision of 90% and recall of 97.5%. Francis Palma, Mathieu Nayrolles, Naouel Moha, Yann-Gaël Guéhéneuc, Benoit Baudry, Jean-Marc Jézéquel |
Int. J. Cooperative Inf. Syst. | 5 |
| 2013 | Usage and testability of AOP: An empirical study of AspectJ
Freddy Muñoz, Benoit Baudry, Romain Delamare, Yves Le Traon |
Inf. Softw. Technol. | 2 |
| 2013 | Automating the maintenance of nonfunctional system properties using demonstration-based model transformationabstractABSTRACT Domain‐Specific Modeling Languages (DSMLs) are playing an increasingly significant role in software development. By raising the level of abstraction using notations that are representative of a specific domain, DSMLs allow the core essence of a problem to be separated from irrelevant accidental complexities, which are typically found at the implementation level in source code. In addition to modeling the functional aspects of a system, a number of nonfunctional properties (e.g., quality of service constraints and timing requirements) also need to be integrated into models in order to reach a complete specification of a system. This is particularly true for domains that have distributed real time and embedded needs. Given a base model with functional components, maintaining the nonfunctional properties that crosscut the base model has become an essential modeling task when using DSMLs. The task of maintaining nonfunctional properties in DSMLs is traditionally supported by manual model editing or by using model transformation languages. However, these approaches are challenging to use for those unfamiliar with the specific details of a modeling transformation language and the underlying metamodel of the domain, which presents a7 steep learning curve for many users. This paper presents a demonstration‐based approach to automate the maintenance of nonfunctional properties in DSMLs. Instead of writing model transformation rules explicitly, users demonstrate how to apply the nonfunctional properties by directly editing the concrete model instances and simulating a single case of the maintenance process. By recording a user's operations, an inference engine analyzes the user's intention and generates generic model transformation patterns automatically, which can be refined by users and then reused to automate the same evolution and maintenance task in other models. Using this approach, users are able to automate the maintenance tasks without learning a complex model transformation language. In addition, because the demonstration is performed on model instances, users are isolated from the underlying abstract metamodel definitions. Our demonstration‐based approach has been applied to several scenarios, such as auto scaling and model layout. The specific contribution in this paper is the application of the demonstration‐based approach to capture crosscutting concerns representative of aspects at the modeling level. Several examples are presented across multiple modeling languages to demonstrate the benefits of our approach. Copyright © 2013 John Wiley & Sons, Ltd. Yu Sun 0002, Jeffrey G. Gray, Romain Delamare, Benoit Baudry, Jules White |
J. Softw. Evol. Process. | 4 |
| 2013 | Automated measurement of models of requirements
Martin Monperrus, Benoit Baudry, Joël Champeau, Brigitte Hoeltzener, Jean-Marc Jézéquel |
Softw. Qual. J. | 2 |
| 2012 | A categorical model of model merging and weavingabstractModel driven engineering advocates the separation of concerns during the design time of a system, which leads to the creation of several different models, using several different syntaxes. However, to reason on the overall system, we need to compose these models. Unfortunately, composition of models is done in an ad hoc way, preventing comparison, capitalisation and reuse of the composition operators. In order to improve comprehension and allow comparison of merging and weaving operators, we use category theory to propose a unified framework to formally define merging and weaving of models. We successfully use this framework to compare them, both through the way they are transformed in the formalism, and through several properties, such as completeness or non-redundancy. Finally, we validate this framework by checking that it correctly identifies three tools as performing merging or weaving of models. Jonathan Y. Marchand, Benoît Combemale, Benoit Baudry |
MiSE | 3 |
| 2012 | Specification and Detection of SOA Antipatterns
Naouel Moha, Francis Palma, Mathieu Nayrolles, Benjamin Joyen Conseil, Yann-Gaël Guéhéneuc, Benoit Baudry, Jean-Marc Jézéquel |
ICSOC | 6 |
| 2012 | Searching the Boundaries of a Modeling Space to Test MetamodelsabstractModel-driven software development relies on metamodels to formally capture modeling spaces. Metamodels specify concepts and relationships between them in order to represent either a specific business domain model or the input and output domains for operations on models (e.g., model refinement). In all cases, a metamodel is a finite description of a possibly infinite set of models, i.e. the set of all models which structure conforms to the description specified in the metamodel. However, there is currently no systematic method to test that a metamodel captures all the correct models of the domain and no more. In this paper, we focus on the automatic selection of a set of models in the modeling space captured by a metamodel. The selected set should both cover as many representative situations as possible and be kept small as possible for further manual analysis. We use simulated annealing to select a set of models that satisfies those two objectives and report on results using two metamodels from two different domains. Juan José Cadavid, Benoit Baudry, Houari Sahraoui |
ICST | 2 |
| 2012 | Minimum Pairwise Coverage Using Constraint Programming TechniquesabstractThis paper presented the global constraint pairwise that can be used to enforce the presence of a given pair within a set of test cases or configurations. It also introduced several optimizations for implementing a method that computes the minimum set of test cases that covers pairwise. In addition, the method, seen as a constraint optimization problem, provides a way to compromise between time and efficiency by allowing anytime interruption or time-contract execution. Our approach has been implemented and evaluated on several instances of a test configurations generation problems [5] where input variables are boolean only. We envision to address other instances of these problem where the variables take their values in larger finite domains. Arnaud Gotlieb, Aymeric Hervieu, Benoit Baudry |
ICST | 3 |
| 2012 | A Vision for Behavioural Model-Driven Validation of Software Product Lines
Xavier Devroey, Maxime Cordy, Gilles Perrouin, Eun-Young Kang 0001, Pierre-Yves Schobbens, Patrick Heymans, Axel Legay, Benoit Baudry |
ISoLA (1) | 8 |
| 2012 | Formally Defining and Iterating Infinite Models
Benoît Combemale, Xavier Thirioux, Benoit Baudry |
MoDELS | 3 |
| 2012 | Managing Execution Environment Variability during Software Testing: An Industrial Experience
Aymeric Hervieu, Benoit Baudry, Arnaud Gotlieb |
ICTSS | 2 |
| 2012 | Bridging the Chasm between Executable Metamodeling and Models of Computation
Benoît Combemale, Cécile Hardebolle, Christophe Jacquet, Frédéric Boulanger, Benoit Baudry |
SLE | 5 |
| 2012 | An approach for semantic enrichment of software product linesabstractSoftware Product Lines (SPLs) have evolved and gained attention as one of the most promising approaches for software reuse. Feature models are the main technique to represent domain variability in SPLs. However, there are other domain aspects, besides variability, which cannot be expressed in a feature model. Also, these diagrams were not designed to facilitate information retrieval, interoperability and inference. In contrast, ontologies seem to be the best solution to meet these requirements. Therefore, this work presents an approach for semantic enrichment of SPLs using ontologies. Our proposal provides methods to add domain information besides variability description, and a top-ontology that specifies generic concepts and relations in an SPL, working as a guide model for information addition. The proposed approach reuses the existing SPL feature model, adding semantic descriptions in a less intrusive way than modifying the feature model notation. João Bosco Ferreira Filho, Olivier Barais, Benoit Baudry, Windson Viana, Rossana M. de Castro Andrade |
SPLC (2) | 3 |
| 2012 | Modeling modeling modeling
Pierre-Alain Muller, Frédéric Fondement, Benoit Baudry, Benoît Combemale |
Softw. Syst. Model. | 3 |
| 2012 | Reusable model transformations
Sagar Sen, Naouel Moha, Vincent Mahé, Olivier Barais, Benoit Baudry, Jean-Marc Jézéquel |
Softw. Syst. Model. | 5 |
| 2012 | Pairwise testing for software product lines: comparison of two approaches
Gilles Perrouin, Sebastian Oster, Sagar Sen, Jacques Klein, Benoit Baudry, Yves Le Traon |
Softw. Qual. J. | 5 |
| 2011 | Estimating footprints of model operationsabstractWhen performed on a model, a set of operations (e.g., queries or model transformations) rarely uses all the information present in the model. Unintended underuse of a model can indicate various problems: the model may contain more detail than necessary or the operations may be immature or erroneous. Analyzing the footprints of the operations - i.e., the part of a model actually used by an operation - is a simple technique to diagnose and analyze such problems. However, precisely calculating the footprint of an operation is expensive, because it requires analyzing the operation's execution trace. Cédric Jeanneret, Martin Glinz, Benoit Baudry |
ICSE | 3 |
| 2011 | Tailored Shielding and Bypass Testing of Web ApplicationsabstractUser input validation is a technique to counter attacks on web applications. In typical client-server architectures, this validation is performed on the client side. This is inefficient because hackers bypass these checks and directly send malicious data to the server. User input validation thus has to be duplicated from the client-side (HTML pages) to the server-side (PHP or JSP etc.). We present a black-box approach for shielding and testing web application against bypass attacks. We automatically analyze HTML pages in order to extract all the constraints on user inputs in addition to the JavaScript validation code. Then, we leverage these constraints for an automated synthesis of a shield, a reverse-proxy tool that protects the server side. The originality and main contribution of this paper is to offer a solution specifically tailored to the web application, through a preliminary learning/analysis step. An experimental study on several open-source web-applications evaluates the effectiveness of the protection tool and the different flaws detected by the testing too and the impact of the shield on performance. Tejeddine Mouelhi, Yves Le Traon, Erwan Abgrall, Benoit Baudry, Sylvain Gombault |
ICST | 4 |
| 2011 | PACOGEN: Automatic Generation of Pairwise Test Configurations from Feature ModelsabstractFeature models are commonly used to specify variability in software product lines. Several tools support feature models for variability management at different steps in the development process. However, tool support for test configuration generation is currently limited. This test generation task consists in systematically selecting a set of configurations that represent a relevant sample of the variability space and that can be used to test the product line. In this paper we propose \pw tool to analyze feature models and automatically generate a set of configurations that cover all pair wise interactions between features. \pw tool relies on constraint programming to generate configurations that satisfy all constraints imposed by the feature model and to minimize the set of the tests configurations. This work also proposes an extensive experiment, based on the state-of-the art SPLOT feature models repository, showing that \pw tool scales over variability spaces with millions of configurations and covers pair wise with less configurations than other available tools. Aymeric Hervieu, Benoit Baudry, Arnaud Gotlieb |
ISSRE | 2 |
| 2011 | Modeling Model Slicers
Arnaud Blouin, Benoît Combemale, Benoit Baudry, Olivier Beaudoux |
MoDELS | 3 |
| 2011 | Guest Editorial for Special Section on Mutation Testing
Benoit Baudry, Jeremy S. Bradbury, Gordon Fraser 0001 |
Inf. Softw. Technol. | 1 |
| 2011 | Model-driven generative development of measurement software
Martin Monperrus, Jean-Marc Jézéquel, Benoit Baudry, Joël Champeau, Brigitte Hoeltzener |
Softw. Syst. Model. | 3 |
| 2011 | An approach for testing pointcut descriptors in AspectJabstractAbstract Aspect‐oriented programming (AOP) promises better software quality through enhanced modularity. Crosscutting concerns are encapsulated in separate units called aspects and are introduced at specific points in the base program at compile time or runtime. However, aspect‐oriented mechanisms also introduce new risks for reliability that must be tackled by specific testing techniques in order to fully benefit from the use of AOP. This paper focuses on the pointcut descriptor (PCD) that declares the set of points in the base program's execution where the crosscutting concern must be woven. A fault in the PCD can have a ripple effect and result in many different faults. New behavior may be added in unexpected places, or places where new behavior should be added may be missed. When implementing aspect‐oriented programs with AspectJ, JUnit is most commonly used to test the program. However, JUnit does not offer any mechanism to look for faults specifically located in the PCD. As a consequence, these faults can be detected only through complex test scenarios and side effects that are difficult to trigger and observe. This paper proposes to monitor the execution of advices in an aspect‐oriented program and use this information to build test cases that target faults in PCDs. The AdviceTracer tool has been developed to automatically monitor and store all information related to advice executions. It also offers a set of operations that can be used to check the presence or absence of advices at specific points in the execution. These operations improve the definition of an oracle for PCD test cases. An empirical study is performed to compare JUnit and AdviceTracer for testing PCDs in terms of the complexity of test cases and their ability to detect faults. The study is performed on a Healthwatcher system that has 93 classes and 19 PCDs. It reveals that test cases that use AdviceTracer to test PCDs are easier to write (shorter test cases and written in less time than with JUnit) and detect more faults. Copyright © 2011 John Wiley & Sons, Ltd. Romain Delamare, Benoit Baudry, Sudipto Ghosh 0001, Yves Le Traon |
Softw. Test. Verification Reliab. | 2 |
| 2010 | Automated and Scalable T-wise Test Case Generation Strategies for Software Product LinesabstractSoftware Product Lines (SPL) are difficult to validate due to combinatorics induced by variability across their features. This leads to combinatorial explosion of the number of derivable products. Exhaustive testing in such a large space of products is infeasible. One possible option is to test SPLs by generating test cases that cover all possible T feature interactions (T-wise). T-wise dramatically reduces the number of test products while ensuring reasonable SPL coverage. However, automatic generation of test cases satisfying T-wise using SAT solvers raises two issues. The encoding of SPL models and T-wise criteria into a set of formulas acceptable by the solver and their satisfaction which fails when processed “all-at-once'”. We propose a scalable toolset using Alloy to automatically generate test cases satisfying T-wise from SPL models. We define strategies to split T-wise combinations into solvable subsets. We design and compute metrics to evaluate strategies on Aspect OPTIMA, a concrete transactional SPL. Gilles Perrouin, Sagar Sen, Jacques Klein, Benoit Baudry, Yves Le Traon |
ICST | 4 |
| 2010 | Variability Modeling and QoS Analysis of Web Services OrchestrationsabstractThe ever-growing choice in diverse services is making service orchestration variability an essential aspect of a composite web service. Influence of this variation on the Quality of Service (QoS) of a composite service is critical and the focus of our work. In this paper, we present a methodology to first model orchestration variability using a feature diagram (FD). The FD specifies a product line of orchestrations represented as configurations of invoked/rejected atomic services. Second, due to the potentially large set of configurations we employ combinatorial testing techniques to automatically generate configurations covering all valid pair wise interactions between services. Third, we analyze QoS variation for each configuration using probabilistic models of QoS. Using a crisis management system case study we experimentally show that pair wise generation covers all QoS outliers and eliminates analysis of > 75% of all possible configurations. The QoS analysis of the pair wise configurations reveals unsafe/ineffective configurations, helps determine realistic Service Level Agreements (SLAs), and provides valuable feedback to help remodel an orchestration. Ajay Kattepur, Sagar Sen, Benoit Baudry, Albert Benveniste, Claude Jard |
ICWS | 3 |
| 2010 | Vidock: A Tool for Impact Analysis of Aspect Weaving on Test Cases
Romain Delamare, Freddy Muñoz, Benoit Baudry, Yves Le Traon |
ICTSS | 3 |
| 2009 | Inquiring the usage of aspect-oriented programming: An empirical studyabstractBack in 2001, the MIT announced aspect-oriented programming as a key technology in the next 10 years. Nowadays, 8 years later, AOP is not widely adopted. Several reasons can explain this distrust in front of AOP, and one of them is the lack of robust tools for analysis, testing and maintenance. In order to develop dedicated solutions for assisting the development with AOP, and increase its adoption, we need to understand how it is actually used. In this paper we analyze 38 aspect-oriented open source projects with respect to the impact of aspects on the projects, and to coverage of the language features. This reveals that AOP is currently used in a cautious way. This work is a first step to built support and development tools dedicated to actual practices for AOP, based on empirical usage profiles. Freddy Muñoz, Benoit Baudry, Romain Delamare, Yves Le Traon |
ICSM | 2 |
| 2009 | A Test-Driven Approach to Developing Pointcut Descriptors in AspectJabstractAspect-oriented programming (AOP) languages introduce new constructs that can lead to new types of faults, which must be targeted by testing techniques. In particular, AOP languages such as AspectJ use a pointcut descriptor (PCD) that provides a convenient way to declaratively specify a set of joinpoints in the program where the aspect should be woven. However, a major difficulty when testing that the PCD matches the intended set of joinpoints is the lack of precise specification for this set other than the PCD itself. In this paper, we propose a test-driven approach for the development and validation of the PCD. We developed a tool, AdviceTracer, which enriches the JUnit API with new types of assertions that can be used to specify the expected joinpoints. In order to validate our approach, we also developed a mutation tool that systematically injects faults into PCDs. Using these two tools, we perform experiments to validate that our approach can be applied for specifying expected joinpoints and for detecting faults in the PCD. Romain Delamare, Benoit Baudry, Sudipto Ghosh 0001, Yves Le Traon |
ICST | 2 |
| 2009 | Transforming and Selecting Functional Test Cases for Security Policy TestingabstractIn this paper, we consider typical applications in which the business logic is separated from the access control logic, implemented in an independent component, called the Policy Decision Point (PDP). The execution of functions in the business logic should thus include calls to the PDP, which grants or denies the access to the protected resources/functionalities of the system, depending on the way the PDP has been configured. The task of testing the correctness of the implementation of the security policy is tedious and costly. In this paper, we propose a new approach to reuse and automatically transform existing functional test cases for specifically testing the security mechanisms. The method includes a three-step technique based on mutation applied to security policies (RBAC, XACML, OrBAC) and AOP for transforming automatically functional test cases into security policy test cases. The method is applied to Java programs and provides tools for performing the steps from the dynamic analyses of impacted test cases to their transformation. Three empirical case studies provide fruitful results and a first proof of concepts for this approach, e.g. by comparing its efficiency to an error-prone manual adaptation task. Tejeddine Mouelhi, Yves Le Traon, Benoit Baudry |
ICST | 3 |
| 2009 | Evaluating Context Descriptions and Property Definition Patterns for Software Formal Validation
Philippe Dhaussy, Pierre Yves Pillain, Stephen Creff, Amine Raji, Yves Le Traon, Benoit Baudry |
MoDELS | 6 |
| 2009 | Modeling Modeling
Pierre-Alain Muller, Frédéric Fondement, Benoit Baudry |
MoDELS | 3 |
| 2009 | Meta-model Pruning
Sagar Sen, Naouel Moha, Benoit Baudry, Jean-Marc Jézéquel |
MoDELS | 3 |
| 2009 | Composing Models for Detecting Inconsistencies: A Requirements Engineering Perspective
Gilles Perrouin, Erwan Brottier, Benoit Baudry, Yves Le Traon |
REFSQ | 3 |
| 2009 | Qualifying input test data for model transformations
Franck Fleurey, Benoit Baudry, Pierre-Alain Muller, Yves Le Traon |
Softw. Syst. Model. | 2 |
| 2008 | Improving maintenance in AOP through an interaction specification frameworkabstractThe invasiveness of aspects is beneficial to modularize crosscutting concerns that require the modification of the data or control flow. However, it introduces subtle errors that are hard to locate and fix in case of evolution. In this paper we illustrate this issue by evolving a program implemented using aspects. Interaction issues, between aspects and the program, emerge from this evolution. We locate them through manual inspection and test execution. This tedious process motivates the need for an abstract specification of intended interactions. To tackle this issue, we propose a framework for specifying the types of invasiveness pattern that are allowed of forbidden in the program. We have also implemented a tool that automatically checks whether the specification is satisfied by the aspects. Freddy Muñoz, Benoit Baudry, Olivier Barais |
ICSM | 2 |
| 2008 | On Combining Multi-formalism Knowledge to Select Models for Model Transformation TestingabstractTesting remains a major challenge for model transformation development. Test models that are used as test data for model transformations, are constrained by various sources of knowledge that is expressed in different formalisms. Thus, in order to automatically generate test models it is necessary to interpret these different sources of knowledge and combine them into a consistent set of information that can be used for model synthesis. In this paper, we identify sources of testing knowledge and present our tool Cartier that uses Alloy as the first-order relational logic language to represent combined knowledge in the form of constraints. The constraints are solved leading to a selection of qualified test models from the input domain of a model transformation. We illustrate our approach using the Unified Modeling Language class diagram to relational database management systems transformation as a running example. Sagar Sen, Benoit Baudry, Jean-Marie Mottu |
ICST | 2 |
| 2008 | Test-Driven Assessment of Access Control in Legacy ApplicationsabstractIf access control policy decision points are not neatly separated from the business logic of a system, the evolution of a security policy likely leads to the necessity of changing the system's code base. This is often the case with legacy systems. We present a test- driven methodology to assess the flexibility of a system, a property that describes the degree of coupling between the access control logic and the business logic of a system. A low flexibility indicates that a modification of the policy will lead to substantial changes of the code. In this paper, we analyze the notion of flexibility which is related to the presence of hidden and implicit security mechanisms in the business logic. We detail how testing can be used for detecting such mechanisms and how it may drive the incremental evolution of a security policy. We use several case studies to illustrate and validate the methodology. Yves Le Traon, Tejeddine Mouelhi, Alexander Pretschner, Benoit Baudry |
ICST | 4 |
| 2008 | A Model-Based Framework for Security Policy Specification, Deployment and Testing
Tejeddine Mouelhi, Franck Fleurey, Benoit Baudry, Yves Le Traon |
MoDELS | 3 |
| 2007 | Model-Driven Engineering for Requirements AnalysisabstractRequirements engineering (RE) encompasses a set of activities for eliciting, modelling, agreeing, communicating and validating requirements that precisely define the problem domain for a software system. Several tools and methods exist to perform each of these activities, but they mainly remain separate, making it difficult to capture the global consistency of large requirement documents. In this paper we introduce model-driven engineering (MDE) as a possible technical solution to integrate these activities in a common framework. First, we dicuss how RE can leverage the two main techniques for MDE: metamodelling and model transformation. Then, we introduce a metamodel for requirements and present how we have implemented this metamodel to make it executable and usable through a constrained natural language for requirements definition. Benoit Baudry, Clémentine Nebut, Yves Le Traon |
EDOC | 1 |
| 2007 | Producing a Global Requirement Model from Multiple Requirement Specificationsabstractcollection of partial specifications produced by different stakeholders. Obtaining a global specification is a fundamental step of a requirement analysis process. Merging requirement specifications is indeed a way to reveal inconsistencies between them. We propose in this paper a model-driven mechanism for that purpose. It takes as inputs a set of texts or models which conform to input requirement languages and produces a global requirements model. This mechanism is integrated in a platform called R2A which stands for "requirements to analysis". The R2A core element is its core requirement metamodel which has been defined for capturing the global requirements model. We illustrate our approach with requirement specifications expressed in a constrained natural language. This platform and its mechanism have been completely implemented with MDE (Model Driven Engineering) technologies. As such, it is a good example of how MDE technologies can contribute to requirements engineering as a technical solution. Erwan Brottier, Benoit Baudry, Yves Le Traon, David Touzet, Bertrand Nicolas |
EDOC | 2 |
| 2007 | Providing Support for Model Composition in MetamodelsabstractIn aspect-oriented modeling (AOM), a design is described using a set of design views. It is sometimes necessary to compose the views to obtain an integrated view that can be analyzed by tools. Analysis can uncover conflicts and interactions that give rise to undesirable emergent behavior. Design models tend to have complex structures and thus manual model composition can be arduous and error- prone. Tools that automate significant parts of model composition are needed if AOM is to gain industrial acceptance. One way of providing automated support for composing models written in a particular language is to define model composition behavior in the metamodel defining the language. In this paper we show how this can be done by extending the UML metamodel with behavior describing symmetric, signature-based composition of UML model elements. We also describe an implementation of the metamodel that supports systematic composition of UML class models. Robert B. France, Franck Fleurey, Y. Raghu Reddy, Benoit Baudry, Sudipto Ghosh 0001 |
EDOC | 4 |
| 2007 | Testing Security Policies: Going Beyond Functional TestingabstractWhile important efforts are dedicated to system functional testing, very few works study how to test specifically security mechanisms, implementing a security policy. This paper introduces security policy testing as a specific target for testing. We propose two strategies for producing security policy test cases, depending if they are built in complement of existing functional test cases or independently from them. Indeed, any security policy is strongly connected to system functionality: testing functions includes exercising many security mechanisms. However, testing functionality does not intend at putting to the test security aspects. We thus propose test selection criteria to produce tests from a security policy. To quantify the effectiveness of a set of test cases to detect security policy flaws, we adapt mutation analysis and define security policy mutation operators. A library case study, a 3-tiers architecture, is used to obtain experimental trends. Results confirm that security must become a specific target of testing to reach a satisfying level of confidence in security mechanisms. Yves Le Traon, Tejeddine Mouelhi, Benoit Baudry |
ISSRE | 3 |
| 2007 | Model-Driven Engineering for Software Migration in a Large Industrial Context
Franck Fleurey, Erwan Breton, Benoit Baudry, Alain Nicolas, Jean-Marc Jézéquel |
MoDELS | 3 |
| 2006 | Improving test suites for efficient fault localizationabstractThe need for testing-for-diagnosis strategies has been identified for a long time, but the explicit link from testing to diagnosis (fault localization) is rare. Analyzing the type of information needed for efficient fault localization, we identify the attribute (called Dynamic Basic Block) that restricts the accuracy of a diagnosis algorithm. Based on this attribute, a test-for-diagnosis criterion is proposed and validated through rigorous case studies: it shows that a test suite can be improved to reach a high level of diagnosis accuracy. So, the dilemma between a reduced testing effort (with as few test cases as possible) and the diagnosis accuracy (that needs as much test cases as possible to get more information) is partly solved by selecting test cases that are dedicated to diagnosis. Benoit Baudry, Franck Fleurey, Yves Le Traon |
ICSE | 1 |
| 2006 | Metamodel-based Test Generation for Model Transformations: an Algorithm and a ToolabstractIn a model-driven development context (MDE), model transformations allow memorizing and reusing design know-how, and thus automate parts of the design and refinement steps of a software development process. A model transformation program is a specific program, in the sense it manipulates models as main parameters. Each model must be an instance of a "metamodel", a metamodel being the specification of a set of models. Programming a model transformation is a difficult and error-prone task, since the manipulated data are clearly complex. In this paper, we focus on generating input test data (called test models) for model transformations. We present an algorithm to automatically build test models from a metamodel Erwan Brottier, Franck Fleurey, Jim Steel, Benoit Baudry, Yves Le Traon |
ISSRE | 4 |
| 2006 | Reusable MDA Components: A Testing-for-Trust Approach
Jean-Marie Mottu, Benoit Baudry, Yves Le Traon |
MoDELS | 2 |
| 2006 | Design by Contract to Improve Software VigilanceabstractDesign by Contract is a lightweight technique for embedding elements of formal specification (such as invariants, pre and postconditions) into an object-oriented design. When contracts are made executable, they can play the role of embedded, online oracles. Executable contracts allow components to be responsive to erroneous states and, thus, may help in detecting and locating faults. In this paper, we define Vigilance as the degree to which a program is able to detect an erroneous state at runtime. Diagnosability represents the effort needed to locate a fault once it has been detected. In order to estimate the benefit of using Design by Contract, we formalize both notions of Vigilance and Diagnosability as software quality measures. The main steps of measure elaboration are given, from informal definitions of the factors to be measured to the mathematical model of the measures. As is the standard in this domain, the parameters are then fixed through actual measures, based on a mutation analysis in our case. Several measures are presented that reveal and estimate the contribution of contracts to the overall quality of a system in terms of vigilance and diagnosability. Yves Le Traon, Benoit Baudry, Jean-Marc Jézéquel |
IEEE Trans. Software Eng. | 2 |
| 2005 | Measuring design testability of a UML class diagram
Benoit Baudry, Yves Le Traon |
Inf. Softw. Technol. | 1 |
| 2005 | From genetic to bacteriological algorithms for mutation-based testingabstractThe level of confidence in a software component is often linked to the quality of its test cases. This quality can in turn be evaluated with mutation analysis: faults are injected into the software component (making mutants of it) to check the proportion of mutants detected (‘killed’) by the test cases. But while the generation of a set of basic test cases is easy, improving its quality may require prohibitive effort. This paper focuses on the issue of automating the test optimization. The application of genetic algorithms would appear to be an interesting way of tackling it. The optimization problem is modelled as follows: a test case can be considered as a predator while a mutant program is analogous to a prey. The aim of the selection process is to generate test cases able to kill as many mutants as possible, starting from an initial set of predators, which is the test cases set provided by the programmer. To overcome disappointing experimentation results, on .Net components and unit Eiffel classes, a slight variation on this idea is studied, no longer at the ‘animal’ level (lions killing zebras, say) but at the bacteriological level. The bacteriological level indeed better reflects the test case optimization issue: it mainly differs from the genetic one by the introduction of a memorization function and the suppression of the crossover operator. The purpose of this paper is to explain how the genetic algorithms have been adapted to fit with the issue of test optimization. The resulting algorithm differs so much from genetic algorithms that it has been given another name: bacteriological algorithm. Copyright © 2005 John Wiley & Sons, Ltd. Benoit Baudry, Franck Fleurey, Jean-Marc Jézéquel, Yves Le Traon |
Softw. Test. Verification Reliab. | 1 |
| 2004 | A UML-Based Concept for High Concurrency: The Real-Time ObjectabstractReal-time (RT) applications are designed to control systems that are inherently parallel. For easing development, abstractions are mandatory to model this concurrency. To achieve this goal, the real-time object paradigm is proposed. It is used to demonstrate how to separate functional and concurrency concerns ensuring also high-level abstraction for parallelism modeling Sébastien Gérard, Chokri Mraidha, François Terrier, Benoit Baudry |
ISORC | 4 |
| 2004 | From Testing to Diagnosis: An Automated Approach
Franck Fleurey, Yves Le Traon, Benoit Baudry |
ASE | 3 |
| 2003 | From diagnosis to diagnosability: axiomatization, measurement and application
Yves Le Traon, Farid Ouabdesselam, Chantal Robach, Benoit Baudry |
J. Syst. Softw. | 4 |
| 2002 | Genes and Bacteria for Automatic Test Cases Optimization in the .NET EnvironmentabstractThe level of confidence in a software component is often linked to the quality of its test cases. This quality can in turn be evaluated with mutation analysis: faulty components (mutants) are systematically generated to check the proportion of mutants detected ("killed") by the test cases. But while the generation of basic test cases set is easy, improving its quality may require prohibitive effort. We focus on the issue of automating the test optimization. We looked at genetic algorithms to solve this problem and modeled it as follows: a test case can be considered as a predator while a mutant program is analogous to a prey. The aim of the selection process is to generate test cases able to kill as many mutants as possible. To overcome disappointing experimentation results on the studied .NET system, we propose a slight variation on this idea, no longer at the "animal" level (lions killing zebras) but at the bacteriological level. The bacteriological level indeed better reflects the test case optimization issue: it introduces a memorization function and suppresses the crossover operator. We describe this model and show how it behaves on the case study. Benoit Baudry, Franck Fleurey, Jean-Marc Jézéquel, Yves Le Traon |
ISSRE | 1 |
| 2002 | Automatic Test Cases Optimization Using a Bacteriological Adaptation Model: Application to .NET ComponentabstractIn this paper, we present several complementary computational intelligence techniques that we explored in the field of .Net component testing. Mutation testing serves as the common backbone for applying classical and new artificial intelligence (AI) algorithms. With mutation tools, we know how to estimate the revealing power of test cases. With AI, we aim at automatically improving test case efficiency. We therefore looked first at genetic algorithms (GA) to solve the problem of test. The aim of the selection process is to generate test cases able to kill as many mutants as possible. We then propose a new AI algorithm that fits better to the test optimization problem, called bacteriological algorithm (BA): BAs behave better that GAs for this problem. However, between GAs and BAs, a family of intermediate algorithms exists: we explore the whole spectrum of these intermediate algorithms to determine whether an algorithm exists that would be more efficient than BAs.: the approaches are compared on a .Net system. Benoit Baudry, Franck Fleurey, Jean-Marc Jézéquel, Yves Le Traon |
ASE | 1 |
| 2001 | Towards a 'Safe' Use of Design Patterns to Improve OO Software TestabilityabstractDesign-for-testability is a very important issue in software engineering. It becomes crucial in the case of OO designs where control flows are generally not hierarchical, but are diffuse and distributed over the whole architecture. We introduce the concept of a "testing conflict" when potentially concurrent client/supplier relationships between the same classes along different paths exist in a system. Such conflicts may be hard to test, especially when dynamic binding and polymorphism are involved. We describe the conflicts using topological class configuration diagrams. An overall architecture is represented as a combination of the initial design and several patterns. We focus on the design patterns as coherent subsets in the architecture, and we explain how their use can provide a way for limiting the complexity of testing for conflicts, and of confining their effects to the classes involved in the pattern. Benoit Baudry, Yves Le Traon, Gerson Sunyé, Jean-Marc Jézéquel |
ISSRE | 1 |
| 2000 | Building Trust into OO Components Using a Genetic AnalogyabstractDespite the growing interest for component based systems, few works tackle the question of the trust we can bring into a component. The paper presents a method and a tool for building trustable OO components. It is particularly adapted to a design-by-contract approach, where the specification is systematically derived into executable assertions (invariant properties, pre/postconditions of methods). A component is seen as an organic set composed of a specification, a given implementation and its embedded test cases. We propose an adaptation of mutation analysis to the OO paradigm that checks the consistency between specification/implementation and tests. Faulty programs, called "mutants", are generated by systematic fault injection in the implementation. The quality of tests is related to the mutation score, i.e. the proportion of faulty programs it detects. The main contribution is to show how a similar idea can be used in the same context to address the problem of effective test optimization. To map the genetic analogy to the test optimization problem, we consider mutant programs to be detected as the initial preys population and test cases as the predators population. The test selection consists of mutating the "predator" test cases and crossing them over in order to improve their ability to kill the prey population. The feasibility of component validation using such a "Darwinian" model and its usefulness for test optimization are studied. Benoit Baudry, Vu Le Hanh, Jean-Marc Jézéquel, Yves Le Traon |
ISSRE | 1 |