VLDB 2026 Research / reviewers in the wild / expert
Ezekiel O. Soremekun
dblp:200/2864 · also Ezekiel Olamide Soremekun
· DBLP profile ↗
17ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-0039-8106ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 4 first-author · 10 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding End-User Perception of Transfer Risks in Smart Contracts
Yustynn Panicker, Ezekiel O. Soremekun, Sudipta Chattopadhyay 0001, Sumei Sun |
CHI | 2 |
| 2025 | Automatic Data Repair without Format SpecificationsabstractIn data processing, datasets are expected to adhere to specific formats. However, inconsistencies due to human error, data corruption, or partial transmission can render these datasets nonconforming, hindering automated processing. This necessitates manual data repair, a time-consuming and errorprone task, especially when formal specifications are unavailable. To address this challenge, we introduce $\epsilon$REPAIR, a novel format-free approach to automating data repair. $\epsilon$REPAIR leverages parser feedback to detect and correct data inconsistencies, making it a versatile solution for data cleansing. In evaluation, $\epsilon$REPAIR achieves $2.6 \times$ higher-quality repairs than its closest competitor, DDMax, in terms of the number of edits required to restore corrupted data, while reducing data loss by $2.8 \times$ compared to DDMax, with only $1.4 \times$ runtime overhead. This work presents a practical, robust, and flexible formatfree data repair alternative to DDMax. Its applications extend to domains such as data science, software development, and other human-centric systems, where handling diverse and inconsistent datasets is critical. Zijian Luo, Lukas Kirschner, Ezekiel O. Soremekun, Rahul Gopinath |
ISSRE | 3 |
| 2025 | Directed Grammar-Based Test GenerationabstractTo effectively test complex software, it is important to generategoal-specific inputs, i.e., inputs that achieve a specific testing goal. For instance, developers may intend to target one or more testing goal(s) during testing – generate complex inputs or trigger new or error-prone behaviors.Problem:However, most state-of-the-art test generators are not designed totarget specific goals. Notably, grammar-based test generators, which (randomly) producesyntactically valid inputsvia an input specification (i.e., grammar) have a low probability of achieving an arbitrary testing goal.Aim: This work addresses this challenge by proposing an automated test generation approach (calledFDLOOP) which iteratively learns relevant input properties from existing inputs to drive the generation of goal-specific inputs.Method: The main idea of our approach is to leveragetest feedbackto generategoal-specific inputsvia a combination ofevolutionary testing and grammar learning.FDLOOPautomatically learns a mapping between input structures and a specific testing goal, such mappings allow to generate inputs that target the goal-at-hand. Given a testing goal,FDLOOPiteratively selects, evolves and learn the input distribution of goal-specific test inputs via test feedback and a probabilistic grammar. We concretizeFDLOOPfor four testing goals, namely unique code coverage, input-to-code complexity, program failures (exceptions) and long execution time. We evaluateFDLOOPusing three (3) well-known input formats (JSON, CSS and JavaScript) and 20 open-source software.Results: In most (86%) settings,FDLOOPoutperforms all five tested baselines namely the baseline grammar-based test generators (random, probabilistic and inverse-probabilistic methods), EvoGFuzz and DynaMOSA.FDLOOPis (up to) twice (2X) as effective as the best baseline (EvoGFuzz) in inducing erroneous behaviors. In addition, we show that the main components ofFDLOOP(i.e., input mutator, grammar mutator and test feedbacks) contribute positively to its effectiveness. We also observed thatFDLOOPis effective across varying parameter settings – the number of initial seed inputs, the number of generated inputs, the number of input generations and varying random seed values.Implications: Finally, our evaluation demonstrates thatFDLOOPeffectively achieves single testing goals (revealing erroneous behaviors, generating complex inputs, or inducing long execution time) and scales to multiple testing goals. Lukas Kirschner, Ezekiel O. Soremekun |
IEEE Trans. Software Eng. | 2 |
| 2024 | Distribution-aware fairness test generation
Sai Sathiesh Rajan, Ezekiel O. Soremekun, Yves Le Traon, Sudipta Chattopadhyay 0001 |
J. Syst. Softw. | 2 |
| 2023 | Evaluating the Impact of Experimental Assumptions in Automated Fault LocalizationabstractMuch research on automated program debugging often assumes that bug fix location(s) indicate the faults' root causes and that root causes of faults lie within single code elements (statements). It is also often assumed that the number of statements a developer would need to inspect before finding the first faulty statement reflects debugging effort. Although intuitive, these three assumptions are typically used (55% of experiments in surveyed publications make at least one of these three assumptions) without any consideration of their effects on the debugger's effectiveness and potential impact on developers in practice. To deal with this issue, we perform controlled experimentation, split testing in particular, using 352 bugs from 46 open-source C programs, 19 Automated Fault Localization (AFL) techniques (18 statistical debugging formulas and dynamic slicing), two (2) state-of-the-art automated program repair (APR) techniques (GenProg and Angelix) and 76 professional developers. Our results show that these assumptions conceal the difficulty of debugging. They make AFL techniques appear to be (up to 38%) more effective, and make APR tools appear to be (2X) less effective. We also find that most developers (83%) consider these assumptions to be unsuitable for debuggers and, perhaps worse, that they may inhibit development productivity. The majority (66%) of developers prefer debugging diagnoses without these assumptions twice as much as with the assumptions. Our findings motivate the need to assess debuggers conservatively, i.e., without these assumptions. Ezekiel O. Soremekun, Lukas Kirschner, Marcel Böhme, Mike Papadakis |
ICSE | 1 |
| 2023 | Towards Backdoor Attacks and Defense in Robust Machine Learning ModelsabstractThe introduction of robust optimisation has pushed the state-of-the-art in defending against adversarial attacks . Notably, the state-of-the-art projected gradient descent (PGD) -based training method has been shown to be universally and reliably effective in defending against adversarial inputs. This robustness approach uses PGD as a reliable and universal “first-order adversary”. However, the behaviour of such optimisation has not been studied in the light of a fundamentally different class of attacks called backdoors. In this paper, we study how to inject and defend against backdoor attacks for robust models trained using PGD-based robust optimisation. We demonstrate that these models are susceptible to backdoor attacks. Subsequently, we observe that backdoors are reflected in the feature representation of such models. Then, this observation is leveraged to detect such backdoor-infected models via a detection technique called AEGIS. Specifically, given a robust Deep Neural Network (DNN) that is trained using PGD-based first-order adversarial training approach, AEGIS uses feature clustering to effectively detect whether such DNNs are backdoor-infected or clean. In our evaluation of several visible and hidden backdoor triggers on major classification tasks using CIFAR-10, MNIST and FMNIST datasets, AEGIS effectively detects PGD-trained robust DNNs infected with backdoors. AEGIS detects such backdoor-infected models with 91.6% accuracy (11 out of 12 tested models), without any false positives . Furthermore, AEGIS detects the targeted class in the backdoor-infected model with a reasonably low (11.1%) false positive rate. Our investigation reveals that salient features of adversarially robust DNNs could be promising to break the stealthy nature of backdoor attacks. Ezekiel O. Soremekun, Sakshi Udeshi, Sudipta Chattopadhyay 0001 |
Comput. Secur. | 1 |
| 2023 | Mutation Testing in Evolving Systems: Studying the Relevance of Mutants to Code EvolutionabstractContext:When software evolves, opportunities for introducing faults appear. Therefore, it is important to test the evolved program behaviors during each evolution cycle. However, while software evolves, its complexity is also evolving, introducing challenges to the testing process. To deal with this issue, testing techniques should be adapted to target the effect of the program changes instead of the entire program functionality. To this end,commit-aware mutation testing, a powerful testing technique, has been proposed. Unfortunately, commit-aware mutation testing is challenging due to the complex program semantics involved. Hence, it is pertinent to understand the characteristics, predictability, and potential of the technique. Objective:We conduct an exploratory study to investigate the properties ofcommit-relevant mutants, i.e., the test elements of commit-aware mutation testing, by proposing a general definition and an experimental approach to identify them. We thus aim at investigating the prevalence, location, and comparative advantages of commit-aware mutation testing over time (i.e., the program evolution). We also investigate the predictive power of several commit-related features in identifying and selecting commit-relevant mutants to understand the essential properties for its best-effort application case. Method:Our commit-relevant definition relies on the notion of observational slicing, approximated by higher-order mutation. Specifically, our approach utilizes the impact of mutants, effects of one mutant on another in capturing and analyzing the implicit interactions between the changed and unchanged code parts. The study analyses millions of mutants (over 10 million), 288 commits, five (5) different open-source software projects involving over 68,213 CPU days of computation and sets a ground truth where we perform our analysis. Results:Our analysis shows that commit-relevant mutants arelocated mainly outside of program commit change(81%), suggesting a limitation in previous work. We also note that effective selection of commit-relevant mutants has the potential of reducing the number of mutants by up to 93%. In addition, we demonstrate that commit relevant mutation testing is significantly more effective and efficient than state-of-the-art baselines, i.e., random mutant selection and analysis of only mutants within the program change. In our analysis of the predictive power of mutants and commit-related features (e.g., number of mutants within a change, mutant type, and commit size) in predicting commit-relevant mutants, we found that mostproxy features do not reliably predict commit-relevant mutants. Conclusion:This empirical study highlights the properties of commit-relevant mutants and demonstrates the importance of identifying and selecting commit-relevant mutants when testing evolving software systems. Milos Ojdanic, Ezekiel O. Soremekun, Renzo Degiovanni, Mike Papadakis, Yves Le Traon |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | GraphCode2Vec: Generic Code Embedding via Lexical and Program Dependence Analysesabstractpeer reviewed Wei Ma 0014, Ezekiel O. Soremekun, Jie Zhang 0050, Mike Papadakis, Maxime Cordy, Xiaofei Xie, Yves Le Traon |
MSR | 3 |
| 2022 | IntJect: Vulnerability Intent Bug SeedingabstractStudying and exposing software vulnerabilities is important to ensure software security, safety, and reliability. Software engineers often inject vulnerabilities into their programs to test the reliability of their test suites, vulnerability detectors, and security measures. However, state-of-the-art vulnerability injection methods only capture code syntax/patterns, they do not learn the intent of the vulnerability and are limited to the syntax of the original dataset. To address this challenge, we propose the first intent-based vulnerability injection method that learns both the program syntax and vulnerability intent. Our approach applies a combination of NLP methods and semantic-preserving program mutations (at the bytecode level) to inject code vulnerabilities. Given a dataset of known vulnerabilities (containing benign and vulnerable code pairs), our approach proceeds by employing semantic-preserving program mutations to transform the existing dataset to semantically similar code. Then, it learns the intent of the vulnerability via neural machine translation (Seq2Seq) models. The key insight is to employ Seq2Seq to learn the intent (context) of the vulnerable code in a manner that is agnostic of the specific program instance. We evaluate the performance of our approach using 1275 vulnerabilities belonging to five (5) CWEs from the Juliet test suite. We examine the effectiveness of our approach in producing compilable and vulnerable code. Our results show that IntJECT is effective, almost all (99%) of the code produced by our approach is vulnerable and compilable. We also demonstrate that the vulnerable programs generated by IntJECT are semantically similar to the withheld original vulnerable code. Finally, we show that our mutation-based data transformation approach outperforms its alternatives, namely data obfuscation and using the original data. Benjamin Petit, Ahmed Khanfir, Ezekiel O. Soremekun, Gilles Perrouin, Mike Papadakis |
QRS | 3 |
| 2022 | Inputs From HellabstractGrammarscan serve asproducersfor structured test inputs that are syntactically correct by construction. A probabilistic grammar assigns probabilities to individual productions, thus controlling the distribution of input elements. Using the grammars as input parsers, we show how tolearn input distributions from input samples,allowing to create inputs that aresimilarto the sample; byinvertingthe probabilities, we can create inputs that aredissimilarto the sample. This allows for threetest generation strategies: 1) “Common inputs”–by learning from common inputs, we can create inputs that aresimilarto the sample; this is useful for regression testing. 2) “Uncommon inputs”–learning from common inputs and inverting probabilities yields inputs that arestrongly dissimilarto the sample; this is useful for completing a test suite with “inputs from hell” that test uncommon features, yet are syntactically valid. 3) “Failure-inducing inputs”–learning from inputs that caused failures in the past gives us inputs that share similar features and thus also have ahigh chance of triggering bugs; this is useful for testing the completeness of fixes. Our evaluation on three common input formats (JSON, JavaScript, CSS) shows the effectiveness of these approaches. Results show that “common inputs” reproduced 96 percent of the methods induced by the samples. In contrast, for almost all subjects (95 percent), the “uncommon inputs” covered significantly different methods from the samples. Learning from failure-inducing samples reproduced all exceptions (100 percent) triggered by the failure-inducing samples and discovered new exceptions not found in any of the samples learned from. Ezekiel O. Soremekun, Esteban Pavese, Nikolas Havrikov, Lars Grunske, Andreas Zeller |
IEEE Trans. Software Eng. | 1 |
| 2022 | Astraea: Grammar-Based Fairness TestingabstractSoftware often produces biased outputs. In particular, machine learning (ML) based software is known to produce erroneous predictions when processing discriminatory inputs. Such unfair program behavior can be caused by societal bias. In the last few years, Amazon, Microsoft and Google have provided software services that produce unfair outputs, mostly due to societal bias (e.g. gender or race). In such events, developers are saddled with the task of conducting fairness testing. Fairness testing is challenging; developers are tasked with generating discriminatory inputs that reveal and explain biases. We propose a grammar-based fairness testing approach (called ASTRAEA) which leverages context-free grammars to generate discriminatory inputs that reveal fairness violations in software systems. Using probabilistic grammars, ASTRAEA also provides fault diagnosis by isolating the cause of observed software bias. ASTRAEAs diagnoses facilitate the improvement of ML fairness. ASTRAEA was evaluated on 18 software systems that provide three major natural language processing (NLP) services. In our evaluation, ASTRAEA generated fairness violations at a rate of about 18%. ASTRAEA generated over 573K discriminatory test cases and found over 102K fairness violations. Furthermore, ASTRAEA improves software fairness by about 76% via model-retraining, on average. Ezekiel O. Soremekun, Sakshi Udeshi, Sudipta Chattopadhyay 0001 |
IEEE Trans. Software Eng. | 1 |
| 2021 | Locating faults with program slicing: an empirical analysis
Ezekiel O. Soremekun, Lukas Kirschner, Marcel Böhme, Andreas Zeller |
Empir. Softw. Eng. | 1 |
| 2020 | Debugging inputsabstractWhen a program fails to process an input, it need not be the program code that is at fault. It can also be that the input data is faulty, for instance as result of data corruption. To get the data processed, one then has to debug the input data---that is, (1) identify which parts of the input data prevent processing, and (2) recover as much of the (valuable) input data as possible. In this paper, we present a general-purpose algorithm called ddmax that addresses these problems automatically. Through experiments, ddmax maximizes the subset of the input that can still be processed by the program, thus recovering and repairing as much data as possible; the difference between the original failing input and the "maximized" passing input includes all input fragments that could not be processed. To the best of our knowledge, ddmax is the first approach that fixes faults in the input data without requiring program analysis. In our evaluation, ddmax repaired about 69% of input files and recovered about 78% of data within one minute per input. Lukas Kirschner, Ezekiel O. Soremekun, Andreas Zeller |
ICSE | 2 |
| 2020 | Abstracting failure-inducing inputsabstractA program fails. Under which circumstances does the failure occur? Starting with a single failure-inducing input ("The input ((4)) fails") and an input grammar, the DDSET algorithm uses systematic tests to automatically generalize the input to an abstract failure-inducing input that contains both (concrete) terminal symbols and (abstract) nonterminal symbols from the grammar—for instance, "(( ))", which represents any expression in double parentheses. Such an abstract failure-inducing input can be used (1) as a debugging diagnostic, characterizing the circumstances under which a failure occurs ("The error occurs whenever an expression is enclosed in double parentheses"); (2) as a producer of additional failure-inducing tests to help design and validate fixes and repair candidates ("The inputs ((1)), ((3 * 4)), and many more also fail"). In its evaluation on real-world bugs in JavaScript, Clojure, Lua, and UNIX command line utilities, DDSET’s abstract failure-inducing inputs provided to-the-point diagnostics, and precise producers for further failure inducing inputs. Rahul Gopinath, Alexander Kampmann, Nikolas Havrikov, Ezekiel O. Soremekun, Andreas Zeller |
ISSTA | 4 |
| 2020 | When does my program do this? learning circumstances of software behaviorabstractA program fails. Under which circumstances does the failure occur? Our Alhazenapproach starts with a run that exhibits a particular behavior and automatically determines input features associated with the behavior in question: (1) We use a grammar to parse the input into individual elements. (2) We use a decision tree learner to observe and learn which input elements are associated with the behavior in question. (3) We use the grammar to generate additional inputs to further strengthen or refute hypotheses as learned associations. (4) By repeating steps 2 and 3, we obtain a theory that explains and predicts the given behavior. In our evaluation using inputs for find, grep, NetHack, and a JavaScript transpiler, the theories produced by Alhazen predict and produce failures with high accuracy and allow developers to focus on a small set of input features: “grep fails whenever the --fixed-strings option is used in conjunction with an empty search string.” Alexander Kampmann, Nikolas Havrikov, Ezekiel O. Soremekun, Andreas Zeller |
ESEC/SIGSOFT FSE | 3 |
| 2017 | Detecting information flow by mutating input dataabstractAnalyzing information flow is central in assessing the security of applications. However, static and dynamic analyses of information flow are easily challenged by non-available or obscure code. We present a lightweight mutation-based analysis that systematically mutates dynamic values returned by sensitive sources to assess whether the mutation changes the values passed to sensitive sinks. If so, we found a flow between source and sink. In contrast to existing techniques, mutation-based flow analysis does not attempt to identify the specific path of the flow and is thus resilient to obfuscation. In its evaluation, our MUTAFLOW prototype for Android programs showed that mutation-based flow analysis is a lightweight yet effective complement to existing tools. Compared to the popular FlowDroid static analysis tool, MutaFlow requires less than 10% of source code lines but has similar accuracy; on 20 tested real-world apps, it is able to detect 75 flows that FlowDroid misses. Björn Mathis, Vitalii Avdiienko, Ezekiel O. Soremekun, Marcel Böhme, Andreas Zeller |
ASE | 3 |
| 2017 | Where is the bug and how is it fixed? an experiment with practitionersabstractResearch has produced many approaches to automatically locate, explain, and repair software bugs. But do these approaches relate to the way practitioners actually locate, understand, and fix bugs? To help answer this question, we have collected a dataset named DBGBENCH --- the correct fault locations, bug diagnoses, and software patches of 27 real errors in open-source C projects that were consolidated from hundreds of debugging sessions of professional software engineers. Moreover, we shed light on the entire debugging process, from constructing a hypothesis to submitting a patch, and how debugging time, difficulty, and strategies vary across practitioners and types of errors. Most notably, DBGBENCH can serve as reality check for novel automated debugging and repair techniques. Marcel Böhme, Ezekiel O. Soremekun, Sudipta Chattopadhyay 0001, Emamurho Ugherughe, Andreas Zeller |
ESEC/SIGSOFT FSE | 2 |