Jens Dietrich 0001

dblp:82/2970 · DBLP profile ↗
← Back
58ranked-venue papers
18as first author
23since 2021 · last 2025
0000-0001-9019-6550ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 46 · 14 first-author · 19 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Levels of Binary Equivalence for the Comparison of Binaries from Alternative Builds
abstract
In response to challenges in software supply chain security, several organisations have created infrastructures to independently build commodity open source projects and release the resulting binaries for Java/Maven and other software ecosystems. Build platform variability can strengthen security as it facilitates the detection of compromised build environments. Furthermore, by improving the security posture of the build platform and collecting provenance information during the build, the resulting artifacts can be used with greater trust. Such offerings are now available from Google, Oracle and RedHat. The availability of multiple binaries built from the same sources creates new challenges and opportunities, and raises questions such as: “Does build A confirm the integrity of build B?” or “Can build A reveal a compromised build B?”. To answer such questions requires a notion of equivalence between binaries. We demonstrate that the obvious approach based on bitwise equality has significant shortcomings in practice, and propose an alternative approach based on levels of equivalence, inspired by clone detection types. We demonstrate the value of these new levels through several experiments. For this purpose, we construct a dataset consisting of Java binaries (jar files) built from the same sources independently by different providers, resulting in 14,156 pairs of binaries in total. We then compare the compiled class files in those jar files and find that for$\mathbf{3, 7 5 0}$pairs of jars ($\mathbf{2 6. 4 9 \%}$) there is at least one such file that is different, also forcing the jar files and their cryptographic hashes to be different. However, based on the new equivalence levels, we can still establish that many of them are practically equivalent; the number of pairs of jars with non-equivalent classes drops to 13.65 % in some cases. We evaluate several candidate equivalence relations on a semi-synthetic dataset that provides oracles consisting of pairs of binaries that either should be, or must not be equivalent. This technique has been applied to evaluate artifacts built from source and used within Oracle's Graal Development Kit for Micronaut (GDK) product.
Jens Dietrich 0001, Tim White, Behnaz Hassanshahi, Paddy Krishnan
ICSME1
2025 Syntest-ACR: Automated Crash Reproduction for Javascript
abstract
Automated Crash Reproduction (ACR) is an area of software testing research that aims to reproduce software crashes to improve developers' ability to debug programs. There has been little progress in applying ACR techniques to JavaScript, as the highly dynamic nature of JavaScript poses challenges for program analysis and synthesis. We present SynTest-ACR, the first tool for ACR in JavaScript, applying artificial intelligence techniques to evolve suitable reproduction cases. We have evaluated SynTest-ACR against the CrashJS dataset consisting of 453 crashes. As a baseline, we ported the state-of-the-art searchguiding fitness function from EvoCrash for Java, finding that it performs much worse when applied to JavaScript programs, and through comprehensively designing and evaluating alternative fitness functions more suitable for JS ACR we obtain an 18.9% increase in reproduction rate over this baseline for Syntest-ACR.
Philip Oliver, Jens Dietrich 0001, Craig Anslow, Michael Homer
ICSME2
2025 JDala - A Simple Capability System for Java
abstract
Dala is a novel capability-based programming model that ensures data-race freedom while also supporting efficient inter-thread communication. While Dala has been designed to inform the design of future programming languages, the question arises whether existing languages can be retrofitted with Dala capabilities. We report such a design called JDala. In JDala, Dala capabilities are added to Java using annotations and interpreted using bytecode instrumentation. With some examples we demonstrate that by adding three simple annotations to the language, we can avoid concurrency bugs like deadlocks and unexpected program behaviour resulting from shallow immutability of Java standard library APIs. JDala demo: https://youtu.be/QddK1q35h-U
Quinten Smit, Jens Dietrich 0001, Michael Homer, Andrew Fawcet, James Noble 0001
ICSME2
2025 Towards Cross-Build Differential Testing
abstract
Recent concerns about software supply chain security have led to the emergence of different binaries built from the same source code. This will sometimes result in binaries that are not identical and therefore have different cryptographic hashes. The question arises whether those binaries are still equivalent, i.e., whether they have the same behaviour. We explore whether differential testing can be used to provide evidence for non-equivalence. We study this for 3,541 pairs of binaries built for the same Maven artifact version, distributed on Maven Central, Google Assured Open Source Software and/or Oracle Build-From-Source. We use EVOSUITE to generate tests for the baseline binary from Maven Central, run these tests against this baseline binary and any available alternately built binaries, and compare the results for consistency. We argue that any differences may indicate variations in program behaviour and could, therefore, be used to detect compromised binaries or failures at runtime. Although our preliminary experiments did not reveal any compromised builds, our approach successfully identified three build configuration errors that caused changes in runtime behaviour. These findings underscore the potential of our method to uncover subtle build differences, highlighting opportunities for improvement.
Jens Dietrich 0001, Tim White, Valerio Terragni, Behnaz Hassanshahi
ICST1
2025 DALEQ - Explainable Equivalence for Java Bytecode
abstract
The security of software builds has attracted increased attention in recent years in response to incidents like solarwinds and xz. Now, several companies including Oracle and Google rebuild open source projects in a secure environment and publish the resulting binaries through dedicated repositories. This practice enables direct comparison between these rebuilt binaries and the original ones produced by developers and published in repositories such as Maven Central. These binaries are often not bitwise identical; however, in most cases, the differences can be attributed to variations in the build environment, and the binaries can still be considered equivalent. Establishing such equivalence, however, is a labor-intensive and error-prone process.While there are some tools that can be used for this purpose, they all fall short of providing provenance, i.e. readable explanation of why two binaries are equivalent, or not. To address this issue, we present daleq, a tool that disassembles Java byte code into a relational database, and can normalise this database by applying datalog rules. Those databases can then be used to infer equivalence between two classes. Notably, equivalence statements are accompanied with datalog proofs recording the normalisation process. We demonstrate the impact of daleq in an industrial context through a large-scale evaluation involving 2,714 pairs of jars, comprising 265,690 class pairs. In this evaluation, daleq is compared to two existing bytecode transformation tools. Our findings reveal a significant reduction in the manual effort required to assess non-bitwise equivalent artifacts, which would otherwise demand intensive human inspection. Furthermore, the results show that daleq outperforms existing tools by identifying more artifacts rebuilt from the same code as equivalent, even when no behavioral differences are present.
Jens Dietrich 0001, Behnaz Hassanshahi
ASE1
2025 Popularity and Innovation in Maven Central
abstract
Maven Central is a large popular repository of Java components that has evolved over the last 20 years. The distribution of dependencies indicates that the repository is dominated by a relatively small number of components other components depend on. The question is whether those elites are static, or change over time, and how this relates to innovation in the Maven ecosystem. We study those questions using several metrics. We find that elites are dynamic, and that the rate of innovation is slowing as the repository ages but remains healthy.
Nkiru Ede, Jens Dietrich 0001, Ulrich Zülicke
MSR2
2025 An extended study of syntactic breaking changes in the wild
abstract
Abstract Libraries assist in accelerating the development of software applications by providing reusable functionalities. Libraries and applications that declare these libraries as dependencies become their clients. However, as libraries evolve, maintaining the dependencies in client projects can be challenging if the new version contains breaking changes. Yet, limited research focuses on analyzing the impact of breaking changes on client projects when updating dependencies in the wild. Hence, we conduct an empirical analysis using Java projects built using Maven to investigate the impact of breaking changes introduced between two library versions. Our dataset included 18,415 Maven artifacts, declaring 142,355 direct dependencies, out of which 71.60% were not up-to-date. We automatically updated these dependencies and discovered that 11.58% of the dependency updates resulted in breaking changes that affected the client, and almost half of them were introduced during a non-major update. We analyzed the changes in the libraries that contributed towards these breaking changes, and our results indicate that changes in transitive dependencies were a significant factor in introducing breaking changes. We further investigated if it was common for clients to use functionalities of transitive dependencies directly without declaring them. This showed that over half of the clients use transitive functionality. Therefore, we analyzed actions suggested to resolve these breaking changes introduced by transitive dependencies under the discussions on open-source platforms, and the frequently suggested action was to exclude the transitive dependency from the project configuration.
Dhanushka Jayasuriya, Samuel Ou, Saakshi Hegde, Valerio Terragni, Jens Dietrich 0001, Kelly Blincoe
Empir. Softw. Eng.5
2024 CrashJS: A NodeJS Benchmark for Automated Crash Reproduction
abstract
Software bugs often lead to software crashes, which cost US companies upwards of $2.08 trillion annually. Automated Crash Reproduction (ACR) aims to generate unit tests that successfully reproduce a crash. The goal of ACR is to aid developers with debugging, providing them with another tool to locate where a bug is in a program. The main approach ACR currently takes is to replicate a stack trace from an error thrown within a program. Currently, ACR has been developed for C, Java, and Python, but there are no tools targeting JavaScript programs. To aid the development of JavaScript ACR tools, we propose CrashJS: a benchmark dataset of 453 Node.js crashes from several sources. CrashJS includes a mix of real-world and synthesised tests, multiple projects, and different levels of complexity for both crashes and target programs.
Philip Oliver, Jens Dietrich 0001, Craig Anslow, Michael Homer
MSR2
2024 Keep Me Updated: An Empirical Study on Embedded JavaScript Engines in Android Apps
abstract
Although JavaScript (JS) has been widely used in mobile development, little is known about the security implications of utilizing JS engines shipped as native app libraries. In this paper, we conduct an empirical study by designing a JS-Inspector pipeline to identify the embedded JS engines in Android apps and assess their security. We investigate over 65,000 Android apps released between Jan 2018 and July 2023. The results show that many popular apps use embedded JS engines, and their engines remain outdated for extended periods. Moreover, approximately 85% of apps have not received updates since their initial release. As such, over 70% of the identified embedded engines are vulnerable to known exploits. We further present case studies of popular apps catering to millions of users. By exploiting their unpatched JS engines through various strategies, such as man-in-the-middle attacks, intent abuse, and malicious mini-apps, we can easily seize control of the targeted apps and execute arbitrary code. This work highlights critical security concerns associated with embedded JS engines. It emphasizes the urgency for timely updates and enhanced security measures during app development.
Elliott Wen, Jiaxiang Zhou, Xiapu Luo, Giovanni Russello, Jens Dietrich 0001
MSR5
2024 Enhancing Security through Modularization: A Counterfactual Analysis of Vulnerability Propagation and Detection Precision
abstract
In today's software development landscape, the use of third-party libraries is near-ubiquitous; leveraging third-party libraries can significantly accelerate development, allowing teams to implement complex functionalities without reinventing the wheel. However, one significant cost of reusing code is security vulnerabilities. Vulnerabilities in third-party libraries have allowed attackers to breach databases, conduct identity theft, steal sensitive user data, and launch mass phishing campaigns. Notorious examples of vulnerabilities in libraries from the past few years include log4shell, solarwinds, event-stream, lodash, and equifax. Existing software composition analysis (SCA) tools track the propagation of vulnerabilities from libraries through dependencies to downstream clients and alert those clients. Due to their design, many existing tools are highly imprecise―they create alerts for clients even when the flagged vulnerabilities are not exploitable. Library developers occasionally release new versions of their software with refactorings that improve modularity. In this work, we explore the impacts of modularity improvements on vulnerability detection. In addition to generally improving the nonfunctional properties of the code, refactoring also has several security-related beneficial side effects: (1) it improves the precision of existing (fast and stable) SCAs; and (2) it protects from vulnerabilities that are exploitable when the vulnerable code is present and not even reachable, as in gadget chain attacks. Our primary contribution is thus to quantify, using a novel simulation-based counterfactual vulnerability analysis, two main ways that improved modularity can boost security. We propose a modularization method using a DAG partitioning algorithm, and statically measure properties of systems that we (synthetically) modularize. In our experiments, we find that modularization can improve precision of Software Composition Analysis (SCA) tools to 71%, up from 35%. Furthermore, migrating to modularized libraries results in 78% of clients no longer being vulnerable to attacks referencing inactive dependencies. We further verify that the results of our modularization reflect the structures that are already implicit in the projects (but for which no modularity boundaries are enforced).
Mohammad Mahdi Abdollahpour, Jens Dietrich 0001, Patrick Lam 0001
SCAM2
2023 On the Effect of Instrumentation on Test Flakiness
abstract
Test flakiness is a problem that affects testing and processes that rely on it. Several factors cause or influence the flakiness of test outcomes. Test execution order, randomness and concurrency are some of the more common and well-studied causes. Some studies mention code instrumentation as a factor that causes or affects test flakiness. However, evidence for this issue is scarce. In this study, we attempt to systematically collect evidence for the effects of instrumentation on test flakiness. We experiment with common types of instrumentation for Java programs—namely, application performance monitoring, coverage and profiling instrumentation. We then study the effects of instrumentation on a set of nine programs obtained from an existing dataset used to study test flakiness, consisting of popular GitHub projects written in Java. We observe cases where realworld instrumentation causes flakiness in a program. However, this effect is rare. We also discuss a related issue—how instrumentation may interfere with flakiness detection and prevention.
Shawn Rasheed, Jens Dietrich 0001, Amjed Tahir
AST2
2023 On Leveraging Tests to Infer Nullable Annotations
Jens Dietrich 0001, David J. Pearce 0001, Mahin Chandramohan
ECOOP1
2023 Efficient Sink-Reachability Analysis via Graph Reduction (Extended Abstract)
abstract
We study a variation of the elementary graph reachability problem, called the sink-reachability problem, which can be found in many applications such as static program analysis, social network analysis, large scale web graph analysis, XML document link path analysis, and the study of gene regulation relationships. To scale sink-reachablity analysis to large graphs, we develop a highly scalable sink-reachability preserving graph reduction strategy for input sink graphs, by using a composition framework. That is, individual sink-reachability preserving condensation operators, each running in linear time, are pipelined together to produce graph reduction algorithms that result in close to maximum reduction, while keeping the computation efficient. Experiments on large real-world sink graphs demonstrate that our compositional approach achieves a reduction rate of up to 99.74% for vertices and a rate of up to 99.46% for edges.
Jens Dietrich 0001, Lijun Chang, Lyndon M. Henry, Catherine McCartin, Bernhard Scholz
ICDE1
2023 Understanding Breaking Changes in the Wild
abstract
Modern software applications rely heavily on the usage of libraries, which provide reusable functionality, to accelerate the development process. As libraries evolve and release new versions, the software systems that depend on those libraries (the clients) should update their dependencies to use these new versions as the new release could, for example, include critical fixes for security vulnerabilities. However, updating is not always a smooth process, as it can result in software failures in the clients if the new version includes breaking changes. Yet, there is little research on how these breaking changes impact the client projects in the wild. To identify if changes between two library versions cause breaking changes at the client end, we perform an empirical study on Java projects built using Maven. For the analysis, we used 18,415 Maven artifacts, which declared 142,355 direct dependencies, of which 71.60% were not up-to-date. We updated these dependencies and found that 11.58% of the dependency updates contain breaking changes that impact the client. We further analyzed these changes in the library which impact the client projects and examine if libraries have adhered to the semantic versioning scheme when introducing breaking changes in their releases. Our results show that changes in transitive dependencies were a major factor in introducing breaking changes during dependency updates and almost half of the detected client impacting breaking changes violate the semantic versioning scheme by introducing breaking changes in non-Major updates.
Dhanushka Jayasuriya, Valerio Terragni, Jens Dietrich 0001, Samuel Ou, Kelly Blincoe
ISSTA3
2023 WasmSlim: Optimizing WebAssembly Binary Distribution via Automatic Module Splitting
abstract
Many web applications have adopted WebAssembly thanks to its near-native performance and high portability. However, as a WebAssembly application grows complex, it starts suffering from bloated file size and extensive startup time. This leads to poor usability, especially on low-end devices. Existing works attempt to address this issue by refactoring the codebase into several smaller shared libraries and dynamically linking them on demand. However, this approach requires source code access and extensive human effort, which are not always feasible. In this work, we present WasmSlim, a novel web server middleware that optimizes WebAssembly binary distribution without needing source code and user intervention. Our system exploits an observation that many functions are rarely used in the application's life cycle. Therefore, our system first serves users with an instrumented binary to collect execution profiles. It then performs a binary-level transformation to generate a slim main module and secondary modules. The main module contains frequently executed functions and patchable method stubs to load secondary modules on demand. Based on the module loading events, our system can also constantly refine the module partitioning scheme to suppress the module loading latency. Our preliminary experiments show that our system on average reduces the binary size by 69% and improves the startup speed by 71%.
Elliott Wen, Jens Dietrich 0001
SANER2
2023 Test flakiness' causes, detection, impact and responses: A multivocal review
abstract
Flaky tests (tests with non-deterministic outcomes) pose a major challenge for software testing. They are known to cause significant issues, such as reducing the effectiveness and efficiency of testing and delaying software releases. In recent years, there has been an increased interest in flaky tests, with research focusing on different aspects of flakiness, such as identifying causes, detection methods and mitigation strategies. Test flakiness has also become a key discussion point for practitioners (in blog posts, technical magazines, etc.) as the impact of flaky tests is felt across the industry. This paper presents a multivocal review that investigates how flaky tests, as a topic, have been addressed in both research and practice. Out of 560 articles we reviewed, we identified and analysed a total of 200 articles that are focused on flaky tests (composed of 109 academic and 91 grey literature articles/posts) and structured the body of relevant research and knowledge using four different dimensions: causes, detection, impact and responses. For each of those dimensions, we provide categorization and classify existing research, discussions, methods and tools With this, we provide a comprehensive and current snapshot of existing thinking on test flakiness, covering both academic views and industrial practices, and identify limitations and opportunities for future research.
Amjed Tahir, Shawn Rasheed, Jens Dietrich 0001, Negar Hashemi, Lu Zhang 0023
J. Syst. Softw.3
2022 A study of single statement bugs involving dynamic language features
abstract
Dynamic language features are widely available in programming languages to implement functionality that can adapt to multiple usage contexts, enabling reuse. Functionality such as data binding, object-relational mapping and user interface builders can be heavily dependent on these features. However, their use has risks and downsides as they affect the soundness of static analyses and techniques that rely on such analyses (such as bug detection and automated program repair). They can also make software more error-prone due to potential difficulties in understanding reflective code, loss of compile-time safety and incorrect API usage. In this paper, we set out to quantify some of the effects of using dynamic language features in Java programs - that is, the error-proneness of using those features with respect to a particular type of bug known as single statement bugs. By mining 2,024 GitHub projects, we found 139 single statement bug instances (falling under 10 different bug patterns), with the highest number of bugs belonging to three specific patterns: Wrong Function Name, Same Function More Args and Change Identifier Used. These results can help practitioners to quantify the risk of using dynamic techniques over alternatives (such as code generation). We hope this classification raises attention on choosing dynamic APIs that are likely to be error-prone, and provides developers a better understanding when designing bug detection tools for such feature.
Li Sui, Shawn Rasheed, Amjed Tahir, Jens Dietrich 0001
ICPC4
2022 Flaky Test Sanitisation via On-the-Fly Assumption Inference for Tests with Network Dependencies
abstract
Flaky tests cause significant problems as they can interrupt automated build processes that rely on all tests succeeding and undermine the trustworthiness of tests. Numerous causes of test flakiness have been identified, and program analyses exist to detect such tests. Typically, these methods produce advice to developers on how to refactor tests in order to make test outcomes deterministic. We argue that one source of flakiness is the lack of assumptions that precisely describe under which circumstances a test is meaningful. We devise a sanitisation technique that can isolate flaky tests quickly by inferring such assumptions on-the-fly, allowing automated builds to proceed as flaky tests are ignored. We demonstrate this approach for Java and Groovy programs by implementing it as extensions for three popular testing frameworks (JUnit4, JUnit5 and Spock) that can transparently inject the inferred assumptions. If JUnit5 is used, those extensions can be deployed without refactoring project source code. We demonstrate and evaluate the utility of our approach using a set of six popular real-world programs, addressing known test flakiness issues in these programs caused by dependencies of tests on network availability. We find that our method effectively sanitises failures induced by network connectivity problems with high precision and recall.
Jens Dietrich 0001, Shawn Rasheed, Amjed Tahir
SCAM1
2022 SecretHunter: A Large-scale Secret Scanner for Public Git Repositories
abstract
Collaborative software development platforms like GitHub have gained tremendous popularity. Unfortunately, many users have reportedly leaked authentication secrets (e.g., textual passwords and API keys) in public Git repositories and caused security incidents and finical loss. Recently, several tools were built to investigate the secret leakage in GitHub. However, these tools could only discover and scan a limited portion of files in GitHub due to platform API restrictions and band-width limitations. In this paper, we present SecretHunter, a real-time large-scale comprehensive secret scanner for GitHub. SecretHunter resolves the file discovery and retrieval difficulty via two major improvements to the Git cloning process. Firstly, our system will retrieve file metadata from repositories before cloning file contents. The early metadata access can help identify newly committed files and enable many bandwidth optimizations such as filename filtering and object deduplication. Secondly, SecretHunter adopts a reinforcement learning model to analyze file contents being downloaded and infer whether the file is sensitive. If not, the download process can be aborted to conserve bandwidth. We conduct a one-month empirical study to evaluate SecretHunter. Our results show that SecretHunter discovers 57% more leaked secrets than state-of-the-art tools. SecretHunter also reduces 85% bandwidth consumption in the object retrieval process and can be used in low-bandwidth settings (e.g., 4G connections).
Elliott Wen, Jia Wang 0009, Jens Dietrich 0001
TrustCom3
2022 VizAPI: Visualizing Interactions between Java Libraries and Clients
abstract
Software projects make use of libraries extensively. Libraries make available intended API surfaces—sets of exposed library interfaces that library developers expect clients to use. However, in practice, clients only use small fractions of intended API surfaces of libraries. We have implemented the VizAPI tool, which shows a visualization that includes both static and dynamic interactions between clients, the libraries they use, and those libraries’ transitive dependencies (all written in Java). We then present some usage scenarios of VizAPI, targetted at library upgrades. One application, by client developers, is to answer a query about upstream code: will their code be affected by breaking changes in library APIs? Or, library developers can use VizAPI to find out about downstream code: which APIs in their source code are commonly used by clients?
Sruthi Venkatanarayanan, Jens Dietrich 0001, Craig Anslow, Patrick Lam 0001
VISSOFT2
2022 Efficient Sink-Reachability Analysis via Graph Reduction
abstract
The reachability problem on directed graphs, asking whether two vertices are connected via a directed path, is an elementary problem that has been well-studied. In this paper, we study a variation of the elementary reachability problem, called thesink-reachabilityproblem, which can be found in many applications such as static program analysis, social network analysis, large scale web graph analysis, XML document link path analysis, and the study of gene regulation relationships. To scale sink-reachablity analysis to large graphs, we develop a highly scalablesink-reachability preservinggraph reduction strategy for input sink graphs, by using acompositionframework. That is, individual sink-reachability preserving condensation operators, each running in linear time, are pipelined together to produce graph reduction algorithms that result in close to maximum reduction, while keeping the computation efficient. Experiments on large real-world sink graphs demonstrate the efficiency and effectiveness of our compositional approach to sink-reachability preserving graph reduction with a reduction rate of up to 99.74 percent for vertices and a rate of up to 99.46 percent for edges.
Jens Dietrich 0001, Lijun Chang, Lyndon M. Henry, Catherine McCartin, Bernhard Scholz
IEEE Trans. Knowl. Data Eng.1
2021 Caught in the Web: DoS Vulnerabilities in Parsers for Structured Data
Shawn Rasheed, Jens Dietrich 0001, Amjed Tahir
ESORICS (1)2
2021 A Partial Reproduction of A Guided Genetic Algorithm for Automated Crash Reproduction
abstract
This paper is a partial reproduction of work by Soltani et al. which presented EvoCrash, a tool for replicating software failures in Java by reproducing stack traces. EvoCrash uses a guided genetic algorithm to generate JUnit test cases capable of reproducing failures more reliably than existing coverage-based solutions. In this paper, we present the findings of our reproduction of the initial study exploring the effectiveness of EvoCrash and comparison to three existing solutions: STAR, JCHARMING, and MuCrash. We further explored the capabilities of EvoCrash on different programs to check for selection bias. We found that we can reproduce the crashes covered by EvoCrash in the original study while reproducing two additional crashes not reported as reproduced. We also find that EvoCrash was unsuccessful in reproducing several crashes from the JCHARMING paper, which were excluded from the original study. Both EvoCrash and JCHARMING could reproduce 73% of the crashes from the JCHARMING paper. We found that there was potentially some selection bias in the dataset for EvoCrash. We also found that some crashes had been reported as non-reproducible even when EvoCrash could reproduce them. We suggest this may be due to EvoCrash becoming stuck in a local optimum.
Philip Oliver, Michael Homer, Jens Dietrich 0001, Craig Anslow
ICSME3
2020 Technical Lag of Dependencies in Major Package Managers
abstract
Background: Third party libraries used by a project (dependencies) can easily become outdated over time, a phenomenon called technical lag. Keeping dependencies up to date induces a significant overhead in terms of the resources (e.g, developer time), but necessary to maintain software quality. Aims: This study provides a large scale analysis of technical lag across the major package managers currently in use. Method: We conducted a mixed-methods study using open-source project data obtained from 14 package managers using the libraries.io dataset. Results: The majority of fixed version declarations, along with a significant number of flexible declarations, are outdated. Fixed declarations are not regularly updated, except in major updates, so they quickly lag. Despite the prevalence of breaking changes in updates, downgrading declarations to earlier versions are rare. Conclusions: Technical lag is prevalent but preventable across package managers - semantic versioning based declaration ranges would remove the majority of lag. Further tooling uptake is also recommended to minimise technical lag.
Jacob Stringer, Amjed Tahir, Kelly Blincoe, Jens Dietrich 0001
APSEC4
2020 On the recall of static call graph construction in practice
abstract
Static analyses have problems modelling dynamic language features soundly while retaining acceptable precision. The problem is well-understood in theory, but there is little evidence on how this impacts the analysis of real-world programs. We have studied this issue for call graph construction on a set of 31 real-world Java programs using an oracle of actual program behaviour recorded from executions of built-in and synthesised test cases with high coverage, have measured the recall that is being achieved by various static analysis algorithms and configurations, and investigated which language features lead to static analysis false negatives.
Li Sui, Jens Dietrich 0001, Amjed Tahir, Georgios Fourtounis 0001
ICSE2
2020 A Hybrid Analysis to Detect Java Serialisation Vulnerabilities
abstract
Serialisation related security vulnerabilities have recently been reported for numerous Java applications. Since serialisation presents both soundness and precision challenges for static analysis, it can be difficult for analyses to precisely pinpoint serialisation vulnerabilities in a Java library. In this paper, we propose a hybrid approach that extends a static analysis with fuzzing to detect serialisation vulnerabilities. The novelty of our approach is in its use of a heap abstraction to direct fuzzing for vulnerabilities in Java libraries. This guides fuzzing to produce results quickly and effectively, and it validates static analysis reports automatically. Our approach shows potential as it can detect known serialisation vulnerabilities in the Apache Commons Collections library.
Shawn Rasheed, Jens Dietrich 0001
ASE2
2020 A large scale study on how developers discuss code smells and anti-pattern in Stack Exchange sites
Amjed Tahir, Jens Dietrich 0001, Steve Counsell, Sherlock A. Licorish, Aiko Fallas Yamashita
Inf. Softw. Technol.2
2019 Generating Mock Skeletons for Lightweight Web-Service Testing
abstract
Modern application development allows applications to be composed using lightweight HTTP services. Testing such an application requires the availability of services that the application makes requests to. However, access to dependent services during testing may be restrained. Simulating the behaviour of such services is, therefore, useful to address their absence and move on application testing. This paper examines the appropriateness of Symbolic Machine Learning algorithms to automatically synthesise HTTP services' mock skeletons from network traffic recordings. These skeletons can then be customised to create mocks that can generate service responses suitable for testing. The mock skeletons have human-readable logic for key aspects of service responses, such as headers and status codes, and are highly accurate.
Thilini Bhagya, Jens Dietrich 0001, Hans W. Guesgen
APSEC2
2019 Dependency versioning in the wild
abstract
Many modern software systems are built on top of existing packages (modules, components, libraries). The increasing number and complexity of dependencies has given rise to automated dependency management where package managers resolve symbolic dependencies against a central repository. When declaring dependencies, developers face various choices, such as whether or not to declare a fixed version or a range of versions. The former results in runtime behaviour that is easier to predict, whilst the latter enables flexibility in resolution that can, for example, prevent different versions of the same package being included and facilitates the automated deployment of bug fixes. We study the choices developers make across 17 different package managers, investigating over 70 million dependencies. This is complemented by a survey of 170 developers. We find that many package managers support - and the respective community adapts - flexible versioning practices. This does not always work: developers struggle to find the sweet spot between the predictability of fixed version dependencies, and the agility of flexible ones, and depending on their experience, adjust practices. We see some uptake of semantic versioning in some package managers, supported by tools. However, there is no evidence that projects switch to semantic versioning on a large scale. The results of this study can guide further research into better practices for automated dependency management, and aid the adaptation of semantic versioning.
Jens Dietrich 0001, David J. Pearce 0001, Jacob Stringer, Amjed Tahir, Kelly Blincoe
MSR1
2019 Man vs machine: a study into language identification of stack overflow code snippets
abstract
Software engineers produce large amounts of publicly accessible data that enables researchers to mine knowledge, fostering a better understanding of the field. Knowledge extraction often relies on meta data. This meta data can either be harvested from user-provided tags, or inferred by algorithms from the respective data. The question arises to which extent either type of meta data can be trusted and relied upon. We study this problem in the context of language identification of code snippets posted on Stack Overflow. We analyse the consistency between user-provided tags and the classification obtained with GitHub linguist, an industry-strength automated language recognition tool. We find that the results obtained by both approaches are often not consistent. This indicates that both have to be used with great care. Our results also suggest that developers may not follow the evolutionary path of programming languages beyond one step when seeking or providing answers to software engineering challenges encountered.
Jens Dietrich 0001, Markus Luczak-Rösch, Elroy Dalefield
MSR1
2019 CorpusVis - Visualizing Software Metrics at Scale
abstract
We do not know fully understand how software violates metrics based principles, particularly in large systems. Systems are restricted by structural and static deficiencies that we can aim to reduce by providing developers with effective visualizations of their code. We developed CorpusVis a widget-based application to explore software metrics of Java software systems from the Qualitas Corpus. Through an evaluation of the visualization techniques we identified what visualizations were effective and which ones did not scale well for large software systems. Our application helps to reduce the structural and static deficiencies in developers code which enables developers to spend less time maintaining legacy systems and learn to develop more effective code for future systems.
Jack Slater, Craig Anslow, Jens Dietrich 0001, Leonel Merino
VISSOFT3
2018 On the Soundness of Call Graph Construction in the Presence of Dynamic Language Features - A Benchmark and Tool Evaluation
Li Sui, Jens Dietrich 0001, Michael Emery, Shawn Rasheed, Amjed Tahir
APLAS2
2018 Can you tell me if it smells?: A study on how developers discuss code smells and anti-patterns in Stack Overflow
abstract
This paper investigates how developers discuss code smells and anti-patterns over Stack Overflow to understand better their perceptions and understanding of these two concepts. Understanding developers' perceptions of these issues are important in order to inform and align future research efforts and direct tools vendors in the area of code smells and anti-patterns. In addition, such insights could lead the creation of solutions to code smells and anti-patterns that are better fit to the realities developers face in practice. We applied both quantitative and qualitative techniques to analyse discussions containing terms associated with code smells and anti-patterns. Our findings show that developers widely use Stack Overflow to ask for general assessments of code smells or anti-patterns, instead of asking for particular refactoring solutions. An interesting finding is that developers very often ask their peers 'to smell their code' (i.e., ask whether their own code 'smells' or not), and thus, utilize Stack Overflow as an informal, crowd-based code smell/anti-pattern detector. We conjecture that the crowd-based detection approach considers contextual factors, and thus, tends to be more trusted by developers over automated detection tools. We also found that developers often discuss the downsides of implementing specific design patterns, and 'flag' them as potential anti-patterns to be avoided. Conversely, we found discussions on why some anti-patterns previously considered harmful should not be flagged as anti-patterns. Our results suggest that there is a need for: 1) more context-based evaluations of code smells and anti-patterns, and 2) better guidelines for making trade-offs when applying design patterns or eliminating smells/anti-patterns in industry.
Amjed Tahir, Aiko Fallas Yamashita, Sherlock A. Licorish, Jens Dietrich 0001, Steve Counsell
EASE4
2018 GHTraffic: A Dataset for Reproducible Research in Service-Oriented Computing
abstract
We present GHTraffic, a dataset of significant size comprising HTTP transactions extracted from GitHub data and augmented with synthetic transaction data. The dataset facilitates reproducible research on many aspects of service-oriented computing. This paper discusses use cases for such a dataset and extracts a set of requirements from these use cases. We then discuss the design of GHTraffic, and the methods and tool used to construct it. We conclude our contribution with some selective metrics that characterise GHTraffic.
Thilini Bhagya, Jens Dietrich 0001, Hans W. Guesgen, Steve Versteeg
ICWS2
2018 Visualizing Design Erosion: How Big Balls of Mud are Made
abstract
Software systems are not static, they have to undergo frequent changes to stay fit for purpose, and in the process of doing so, their complexity increases. It has been observed that this process often leads to the erosion of the systems design and architecture and with it, the decline of many desirable quality attributes, such as maintainability. This process can be captured in terms of antipatterns - atomic violations of widely accepted design principles. We present a visualisation that exposes the design of evolving Java programs, highlighting instances of selected antipatterns including their emergence and cancerous growth. This visualisation assists software engineers and architects in assessing, tracing and therefore combating design erosion. We evaluated the effectiveness of the visualisation in four case studies with ten participants.
David Baum, Jens Dietrich 0001, Craig Anslow, Richard Müller 0002
VISSOFT2
2017 On the Use of Mined Stack Traces to Improve the Soundness of Statically Constructed Call Graphs
abstract
Static program analysis is a cornerstone of modern software engineering - it is used to detect bugs and security vulnerabilities early before software is deployed. While there is a large body of research into the scalability and the precision of static analysis, the (un) soundness of static analysis is a critical issue that has not attracted the same level of attention by the research community. In this paper we investigate the question whether information harvested from stack traces obtained from the GitHub issue tracker and Stack Overflow Q&A forums can be used in order to complement statically built call graphs. For this purpose, we extract reflective call graph edges from parsed stack traces, and check whether these edges are correctly computed by Doop, a widely used tool for static analysis with built-in support for reflection analysis. We do find edges that Doop misses when analysing real-world programs, even when reflection analysis is enabled. This suggests that mining techniques are a useful tool to test and improve the soundness of static analysis.
Li Sui, Jens Dietrich 0001, Amjed Tahir
APSEC2
2017 Evil Pickles: DoS Attacks Based on Object-Graph Engineering
abstract
In recent years, multiple vulnerabilities exploiting the serialisation APIs of various programming languages, including Java, have been discovered. These vulnerabilities can be used to devise in- jection attacks, exploiting the presence of dynamic programming language features like reflection or dynamic proxies. In this paper, we investigate a new type of serialisation-related vulnerabilit- ies for Java that exploit the topology of object graphs constructed from classes of the standard library in a way that deserialisation leads to resource exhaustion, facilitating denial of service attacks. We analyse three such vulnerabilities that can be exploited to exhaust stack memory, heap memory and CPU time. We discuss the language and library design features that enable these vulnerabilities, and investigate whether these vulnerabilities can be ported to C#, Java- Script and Ruby. We present two case studies that demonstrate how the vulnerabilities can be used in attacks on two widely used servers, Jenkins deployed on Tomcat and JBoss. Finally, we propose a mitigation strategy based on contract injection.
Jens Dietrich 0001, Kamil Jezek, Shawn Rasheed, Amjed Tahir, Alex Potanin
ECOOP1
2017 Contracts in the Wild: A Study of Java Programs
abstract
The use of formal contracts has long been advocated as an approach to develop programs that are provably correct. However, the reality is that adoption of contracts has been slow in practice. Despite this, the adoption of lightweight contracts — typically utilising runtime checking — has progressed. In the case of Java, built-in features of the language (e.g. assertions and exceptions) can be used for this. Furthermore, a number of libraries which facilitate contract checking have arisen. In this paper, we catalogue 25 techniques and tools for lightweight contract checking in Java, and present the results of an empirical study looking at a dataset extracted from the 200 most popular projects found on Maven Central, constituting roughly 351,034 KLOC. We examine (1) the extent to which contracts are used and (2) what kind of contracts are used. We then investigate how contracts are used to safeguard code, and study problems in the context of two types of substitutability that can be guarded by contracts: (3) unsafe evolution of APIs that may break client programs and (4) violations of Liskovs Substitution Principle (LSP) when methods are overridden. We find that: (1) a wide range of techniques and constructs are used to represent contracts, and often the same program uses different techniques at the same time; (2) overall, contracts are used less than expected, with significant differences between programs; (3) projects that use contracts continue to do so, and expand the use of contracts as they grow and evolve; and, (4) there are cases where the use of contracts points to unsafe subtyping (violations of Liskov's Substitution Principle) and unsafe evolution.
Jens Dietrich 0001, David J. Pearce 0001, Kamil Jezek, Premek Brada
ECOOP1
2017 Parallel Symmetric Class Expression Learning
abstract
In machine learning, one often encounters data sets where a general pattern is violated by a relatively small number of exceptions (for example, a rule that says that all birds can fly is violated by examples such as penguins). This complicates the concept learning process and may lead to the rejection of some simple and expressive rules that cover many cases. In this paper we present an approach to this problem in description logic learning by computing partial descriptions (which are not necessarily entirely complete) of both positive and negative examples and combining them. Our Symmetric Parallel Class Expression Learning approach enables the generation of general rules with exception patterns included. We demonstrate that this algorithm provides significantly better results (in terms of metrics such as accuracy, search space covered, and learning time) than standard approaches on some typical data sets. Further, the approach has the added benefit that it can be parallelised relatively simply, leading to much faster exploration of the search tree on modern computers.
An Cong Tran, Jens Dietrich 0001, Hans W. Guesgen, Stephen R. Marsland
J. Mach. Learn. Res.2
2016 A Note on the Soundness of Difference Propagation
Jens Dietrich 0001, Nicholas Hollingum, Bernhard Scholz
FTfJP@ECOOP1
2016 Magic with Dynamo -- Flexible Cross-Component Linking for Java with Invokedynamic
abstract
Modern software systems are not built from scratch. They use functionality provided by libraries. These libraries evolve and often upgrades are deployed without the systems being recompiled. In Java, this process is particularly error-prone due to the mismatch between source and binary compatibility, and the lack of API stability in many popular libraries. We propose a novel approach to mitigate this problem based on the use of invokedynamic instructions for cross-component method invocations. The dispatch mechanism of invokedynamic is used to provide on-the-fly signature adaptation. We show how this idea can be used to construct a Java compiler that produces more resilient bytecode. We present the dynamo compiler, a proof-of-concept implemented as a javac post compiler. We evaluate our approach using several benchmark examples and two case studies showing how the dynamo compiler can prevent certain types of linkage and stack overflow errors that have been observed in real-world systems.
Kamil Jezek, Jens Dietrich 0001
ECOOP2
2016 A Web-Based Environment for Introductory Programming based on a Bi-Directional Layered Notional Machine
abstract
No abstract available.
Li Sui, Jens Dietrich 0001, Eva Heinrich, Manfred Meyer
ITiCSE2
2016 Antipattern and Code Smell False Positives: Preliminary Conceptualization and Classification
abstract
Anti-patterns and code smells are archetypes used for describing software design shortcomings that can negatively affect software quality, in particular maintainability. Tools, metrics and methodologies have been developed to identify these archetypes, based on the assumption that they can point at problematic code. However, recent empirical studies have shown that some of these archetypes are ubiquitous in real world programs, and many of them are found not to be as detrimental to quality as previously conjectured. We are therefore interested in revisiting common anti-patterns and code smells, and building a catalogue of cases that constitute candidates for "false positives". We propose a preliminary classification of such false positives with the aim of facilitating a better understanding of the effects of anti-patterns and code smells in practice. We hope that the development and further refinement of such a classification can support researchers and tool vendors in their endeavour to develop more pragmatic, context-relevant detection and analysis tools for anti-patterns and code smells.
Francesca Arcelli Fontana, Jens Dietrich 0001, Bartosz Walter, Aiko Fallas Yamashita, Marco Zanoni
SANER2
2016 What Java developers know about compatibility, and why this matters
Jens Dietrich 0001, Kamil Jezek, Premek Brada
Empir. Softw. Eng.1
2015 Giga-scale exhaustive points-to analysis for Java in under a minute
abstract
Computing a precise points-to analysis for very large Java programs remains challenging despite the large body of research on points-to analysis. Any approach must solve an underlying dynamic graph reachability problem, for which the best algorithms have near-cubic worst-case runtime complexity, and, hence, previous work does not scale to programs with millions of lines of code. In this work, we present a novel approach for solving the field-sensitive points-to problem for Java with the means of (1) a transitive-closure data-structure, and (2) a pre-computed set of potentially matching load/store pairs to accelerate the fix-point calculation. Experimentation on Java benchmarks validates the superior performance of our approach over the standard context-free language reachability implementations. Our approach computes a points-to index for the OpenJDK with over 1.5 billion tuples in under a minute.
Jens Dietrich 0001, Nicholas Hollingum, Bernhard Scholz
OOPSLA1
2015 Circular dependencies and change-proneness: An empirical study
abstract
Advice that circular dependencies between programming artefacts should be avoided goes back to the earliest work on software design, and is well-established and rarely questioned. However, empirical studies have shown that real-world (Java) programs are riddled with circular dependencies between artefacts on different levels of abstraction and aggregation. It has been suggested that additional heuristics could be used to distinguish between bad and harmless cycles, for instances by relating them to the hierarchical structure of the packages within a program, or to violations of additional design principles. In this study, we try to explore this question further by analysing the relationship between different kinds of circular dependencies between Java classes, and their change frequency. We find that (1) the presence of cycles can have a significant impact on the change proneness of the classes near these cycles and (2) neither subtype knowledge nor the location of the cycle within the package containment tree are suitable criteria to distinguish between critical and harmless cycles.
Tosin Daniel Oyetoyan, Jean-Rémy Falleri, Jens Dietrich 0001, Kamil Jezek
SANER3
2015 How Java APIs break - An empirical study
Kamil Jezek, Jens Dietrich 0001, Premek Brada
Inf. Softw. Technol.2
2013 Improving Predictive Specificity of Description Logic Learners by Fortification
abstract
The predictive accuracy of a learning algorithm can be split into specificity and sensitivity, amongst other decompositions. Sensitivity, also known as completeness, is the ratio of true positives to the total number of positive examples, while specificity is the ratio of true negative to the total negative examples. In top-down learning methods of inductive logic programming, there is generally a bias towards sensitivity, since the learning starts from the most general rule (everything is positive) and specialises by excluding some of the negative examples. While this is often useful, it is not always the best choice: for example, in novelty detection, where the negative examples are rare and often varied, they may well be ignored by the learning. In this paper we introduce a method that attempts to remove the bias towards sensitivity by fortifying the model by computing and then including in the model some descriptions of the negative data even if they are considered redundant by the normal learning algorithm. We demonstrate the method on a set of standard datasets for description logic learning and show that the predictive accuracy increases.
An Cong Tran, Jens Dietrich 0001, Hans W. Guesgen, Stephen R. Marsland
ACML2
2013 On the Automation of Dependency-Breaking Refactorings in Java
abstract
The architecture and design of object-oriented systems is often described in terms of dependencies between program artefacts like classes, packages and libraries. In particular, there are many heuristics that mandate that certain dependencies should be avoided, and there are increasingly popular frameworks that focus on actively managed dependencies between artefacts, such as the Spring framework and the OSGi dynamic module system. However, empirical studies have shown that many real-world Java programs are riddled with dependency related problems. The question arises whether it is possible to automatically refurbish Java programs by removing unwanted dependencies without affecting the functionality of the program. We present an algorithm and a proof-of-concept implementation that does this. Our approach uses several refactorings: move class, type generalisation, service locator, and in lining and is validated on the Qualitas Corpus set of Java programs.
Syed Muhammad Ali Shah, Jens Dietrich 0001, Catherine McCartin
ICSM2
2010 A Formal Framework to Optimise Component Dependency Resolution
abstract
Dependency resolution (DR) uses a component's explicitly declared requirements to calculate systems where all dependencies are satisfied. There can be many configurations to choose from when resolving dependencies. DR should aim to identify and return an optimal component configuration. This becomes a significant challenge when diverse and sometimes conflicting criteria such as user preferences, contextual constraints and functional requirements must be considered. In this paper we present a framework in which to represent and compose such criteria. This is achieved by defining criteria as ranking systems over complete lattices and composing them in different ways. We present a depth first branch and bound algorithm for this framework, and an example problem that demonstrates the frameworks application. The presented framework will enable the formal definition and composition of criteria to optimise dependency resolution.
Graham Jenson, Jens Dietrich 0001, Hans W. Guesgen
APSEC2
2010 The Qualitas Corpus: A Curated Collection of Java Code for Empirical Studies
abstract
In order to increase our ability to use measurement to support software development practise we need to do more analysis of code. However, empirical studies of code are expensive and their results are difficult to compare. We describe the Qualitas Corpus, a large curated collection of open source Java systems. The corpus reduces the cost of performing large empirical studies of code and supports comparison of measurements of the same artifacts. We discuss its design, organisation, and issues associated with its development.
Ewan D. Tempero, Craig Anslow, Jens Dietrich 0001, Ted Han, Markus Lumpe, Hayden Melton, James Noble 0001
APSEC3
2010 Use Cases for Abnormal Behaviour Detection in Smart Homes
An Cong Tran, Stephen R. Marsland, Jens Dietrich 0001, Hans W. Guesgen, Paul Lyons
ICOST3
2009 Layered Government and E-Citizenship: Objectives and Technical Challenges in the EU
abstract
A typical citizen engages with governmental agencies at several hierarchical levels - local, regional and national. Federal and international communities add even more layers atop of them. We denote this as the concept of layered government. A citizen within such a community expects to find similar procedures and HCI interfaces to accomplish tasks, e.g., voting registration or business registration. Such commonly designed single points of contact are emerging as real challenges and are often mandated by legal regulations. We use the examples of the German E-Government Initiative and the EU Services Directive to define technical problem domains and propose concepts that address them.
Vladimir Stantchev, Marten Schönherr, Jens Dietrich 0001
ICIW3
2008 Requirements for Rich Internet Application Design Methodologies
Jevon M. Wright, Jens Dietrich 0001
WISE2
2008 Using social networking and semantic web technology in software engineering - Use cases, patterns, and a case study
Jens Dietrich 0001, Nathan Jones, Jevon M. Wright
J. Syst. Softw.1
2007 A Formal Contract Language for Plugin-based Software Engineering
abstract
Plugin-based application design has become increasingly popular in recent years, and has contributed to the success of a range of very different applications including Mozilla Firefox and the Eclipse development environment. Using plugins is a promising approach to build complex systems that have to be reconfigured at runtime, and several plugin based general purpose runtime environments are currently under development. Plugin-based design is based on the idea that plugins provide additional functionality extending the capabilities of a core product. While this is often understood as providing services by implementing abstract classes or interfaces defined in the core product, modern plugin-based systems like Eclipse use a much wider definition of service. We propose to consider these services as typed resources and introduce a contract language that can be used to define contracts between plugins providing and consuming services. This language is based on the Semantic Web Rule Language (SWRL) that has a well-defined syntax and semantics. These contracts can then be used in order to validate complex, plugin-based applications.
Jens Dietrich 0001, John G. Hosking, Jonathan Giles
ICECCS1
2007 Towards a web of patterns
Jens Dietrich 0001, Chris Elgar
J. Web Semant.1
2004 A Rule-Based System for eCommerce Applications
Jens Dietrich 0001
KES1