VLDB 2026 Research / reviewers in the wild / expert
Shanto Rahman
dblp:183/4623
· DBLP profile ↗
10ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ranking Relevant Tests for Order-Dependent Flaky TestsabstractOne major challenge of regression testing are flaky tests, i.e., tests that may pass in one run but fail in another run for the same version of code. One prominent category of flaky tests is order-dependent (OD) flaky tests, which can pass or fail depending on the order in which the tests are run. To help developers debug and fix OD tests, prior work attempts to automatically find OD-relevant tests, which are tests that determine whether an OD test passes or fails, depending on whether the OD-relevant tests run before or after the OD test. Prior work found OD-relevant tests by running different tests before the OD test, without considering each test's likelihood of being OD-relevant tests. We propose RankF to rank tests in order of likelihood of being OD-relevant tests, finding the first OD-relevant test for a given OD test more quickly. We propose two ranking approaches, each requiring different information. Our first approach,$\boldsymbol{RankF}_L$, relies on training a large-language model to analyze test code. Our second approach,$\boldsymbol{RankF}_O$, relies on analyzing prior test-order execution information. We evaluate our approaches on 155 OD tests across 24 open-source projects. We compare RankF against baselines from prior work, where we find that RankF finds the first OD-relevant test for an OD test faster than the best baseline; depending on the type of OD-relevant test, RankF takes 9.4 to 14.1 seconds on median, compared to the baseline's 34.2 to 118.5 seconds on median. Shanto Rahman, Bala Naren Chanumolu, Suzzana Rafi, August Shi, Wing Lam |
ICSE | 1 |
| 2025 | Understanding and Improving Flaky Test ClassificationabstractRegression testing is an essential part of software development, but it suffers from the presence of flaky tests - tests that pass and fail non-deterministically when run on the same code. These unpredictable failures waste developers’ time and often hide real bugs. Prior work showed that fine-tuned large language models (LLMs) can classify flaky tests into different categories with very high accuracy. However, we find that prior approaches over-estimated the accuracy of the models due to incorrect experimental design and unrealistic datasets - making the flaky test classification problem seem simpler than it is. In this paper, we first show how prior flaky test classifiers over-estimate the prediction accuracy due to 1) flawed experiment design and 2) mis-representation of the real distribution of flaky (and non-flaky) tests in their datasets. After we fix the experimental design and construct a more realistic dataset (which we name FlakeBench), the prior state-of-the-art model shows a steep drop in F1-score, from 81.82% down to 56.62%. Motivated by these observations, we develop a new training strategy to fine-tune a flaky test classifier, FlakyLens, that improves the classification F1-score to 65.79% (9.17pp higher than the state-of-the-art). We also compare FlakyLens against recent pre-trained LLMs, such as CodeLlama and DeepSeekCoder, on the same classification task. Our results show that FlakyLens consistently outperforms these models, highlighting that general-purpose LLMs still fall short on this specialized task. Using our improved flaky test classifier, we identify the important tokens in the test code that influence the models in making correct or incorrect predictions. By leveraging attribution scores computed per code token in each test, we investigate the tokens that have higher impact on the flaky test classifier’s decision-making per flaky test category. To assess the influence of these important tokens, we introduce adversarial perturbation using these important tokens into the tests and observe whether the model’s predictions change. Our findings show that, when introducing perturbations using the most important tokens, the classification accuracy can change by as much as -18.37pp. These results highlight that these models still struggle to generalize beyond their training data and rely on identifying category-specific tokens (instead of understanding their semantic context), calling for further research into more robust training methodologies. Shanto Rahman, Saikat Dutta 0001, August Shi |
Proc. ACM Program. Lang. | 1 |
| 2025 | UTFix: Change Aware Unit Test Repairing using LLMabstractSoftware updates, including bug repair and feature additions, are frequent in modern applications but they often leave test suites outdated, resulting in undetected bugs and increased chances of system failures. A recent study by Meta revealed that 14%-22% of software failures stem from outdated tests that fail to reflect changes in the codebase. This highlights the need to keep tests in sync with code changes to ensure software reliability. In this paper, we present UTFix , a novel approach for repairing unit tests when their corresponding focal methods undergo changes. UTFix addresses two critical issues: assertion failure and reduced code coverage caused by changes in the focal method. Our approach leverages language models to repair unit tests by providing contextual information such as static code slices, dynamic code slices, and failure messages. We evaluate UTFix on our generated synthetic benchmark (Syn-Bench), and real-world benchmark. In our experiment, UTFix successfully repaired 89.2% of assertion failures and achieved 100% code coverage for 96 tests out of 369 unit tests. On the real-world benchmarks, UTFix repaired 60% of assertion failures while achieving 100% code coverage for 19 out of 30 unit tests. To the best of our knowledge, this is the first comprehensive study focused on unit test in evolving Python projects. Our contributions include the development of UTFix , the creation of Syn-Bench and real-world benchmarks, and the demonstration of the effectiveness of LLM-based methods in addressing unit test failures due to software evolution. Shanto Rahman, Sachit Kuhar, Berk Çirisci, Pranav Garg 0001, Shiqi Wang 0002, Xiaofei Ma 0001, Anoop Deoras, Baishakhi Ray |
Proc. ACM Program. Lang. | 1 |
| 2024 | FlakeSync: Automatically Repairing Async Flaky TestsabstractRegression testing is an important part of the development process but suffers from the presence of flaky tests. Flaky tests nondeterministically pass or fail when run on the same code, misleading developers about the correctness of their changes. A common type of flaky tests are async flaky tests that flakily fail due to timing-related issues such as asynchronous waits that do not return in time or different thread interleavings during execution. Developers commonly try to repair async flaky tests by inserting or increasing some wait time, but such repairs are unreliable. Shanto Rahman, August Shi |
ICSE | 1 |
| 2024 | Quantizing Large-Language Models for Predicting Flaky TestsabstractA major challenge in regression testing practice is the presence of flaky tests, which non-deterministically pass or fail when run on the same code. Previous research identified multiple categories of flaky tests. Prior research has also de-veloped techniques for automatically detecting which tests are flaky or categorizing flaky tests, but these techniques generally involve repeatedly rerunning tests in various ways, making them costly to use. Although several recent approaches have utilized large-language models (LLMs) to predict which tests are flaky or predict flaky-test categories without needing to rerun tests, they are costly to use due to relying on a large neural network to perform feature extraction and prediction. We propose FlakyQ to improve the effectiveness of LLM-based flaky-test prediction by quantizing LLM's weights. The quantized LLM can extract features from test code more efficiently. To make up for loss in prediction performance due to quantization, we further train a traditional ML classifier (e.g., a random forest) to learn from the quantized LLM-extracted features and do the same prediction. The final model has similar prediction performance while running faster than the non-quantized LLM. Our evaluation finds that FlakyQ classifiers consistently improves prediction time over the non-quantized LLM classifier, saving 25.4% in prediction time over all tests, along with a 48.4 % reduction in memory usage. Furthermore, prediction performance is equal or better than the non-quantized LLM classifier. Shanto Rahman, Abdelrahman Baz, Sasa Misailovic, August Shi |
ICST | 1 |
| 2024 | Automatically Reproducing Timing-Dependent Flaky-Test FailuresabstractWhen developers run tests after making code changes, they may encounter test failures from flaky tests, which are tests that can non-deterministically pass or fail on the same version of code. Prior work has found “timing dependence” to be a top cause of this non-determinism, i.e., tests may pass or fail depending on the timing of asynchronous callbacks or different thread interleavings that can occur when thread executions run faster or slower relative to others. Similar to how one debugs and fixes normal test failures, developers need to be able to reliably reproduce flaky-test failures. However, many of these failures can be extremely unlikely to occur (e.g., failing only once out of 10,000 runs in prior work), making it costly for developers to reproduce the failures. We present FlakeRake, an automated approach for reproducing timing-dependent (TD) flaky-test failures by inserting well-placed sleep calls, which temporarily pauses one thread or task and allows another to overtake it. When applied to an existing dataset of known flaky-test failures, FlakeRake is able to reproduce the exact same failure at least once for 136 failures, whereas simply rerunning each test 10,000 times reproduces only 115 failures or rerunning the entire test suites 10,000 times reproduces only 127 failures. For each failure that can be reproduced, we find that FlakeRake can reliably reproduce (>50% of the time) 107 failures, while rerunning just the flaky test or the entire test suite could not reliably reproduce any failure. We also find that if a developer needs to reproduce a failure six or more times, using FlakeRake (including the one-time cost to search for sleep calls) takes less time to reproduce that many failures than continually rerunning just the flaky test. Lastly, we inspect the sleep locations that FlakeRake outputs and provide insights for how one should cope with TD flaky tests. Shanto Rahman, Aaron Massey, Wing Lam, August Shi, Jonathan Bell 0001 |
ICST | 1 |
| 2023 | Optimizing Continuous Development by Detecting and Preventing Unnecessary Content GenerationabstractContinuous development (CD) helps developers quickly release and update their software. To enact CD, developers customize their CD builds to perform several tasks, including compiling, testing, static analysis checks, etc. However, as developers add more tasks to their builds, the builds take longer to run, therefore slowing down the entire CD process. Furthermore, developers may unknowingly include tasks into their builds whose results are not used (e.g., generating coverage files that are never read or uploaded anywhere), therefore wasting build runtime doing unnecessary tasks. We propose OptCD, a technique to dynamically detect unnecessary work within CD builds. Our intuition is that unnecessary work can be identified by the generation of files that are not used by any other task within the build. OptCD runs alongside a CD build, tracking the generated files during the build and which files are read/written. Files that are written to but are never read from are unnecessary content from a build. Based on the names of the unnecessary files, OptCD then maps the files to the specific build tasks responsible for generating or writing to those files. Finally, OptCD leverages ChatGPT to suggest changing the build configuration to disable generating these unnecessary files. Our evaluation of OptCD on 22 open-source projects finds that 95.6% of projects generate at least one unused directory, a directory whose contents are all unnecessarily generated. OptCD identifies the correct task that generates 92.0% of the unused directories. Further, OptCD can produce a patch for the CD configuration file to prevent generating 72.0% of the unused directories. Using the patches, we reduce the runtime by 7.0% on average for the projects we studied. We submitted 26 pull requests for the unused directories that we could disable. Developers have accepted 12 of them, with five rejected, and nine still pending. Talank Baral, Shanto Rahman, Bala Naren Chanumolu, Basak Balci, Tuna Tuncer, August Shi, Wing Lam |
ASE | 2 |
| 2017 | A Statement Level Bug Localization Technique using Statement Dependency Graph
Shanto Rahman, Mostafijur Rahman, Kazi Sakib |
ENASE | 1 |
| 2016 | An Appropriate Method Ranking Approach for Localizing Bugs using Minimized Search SpaceabstractIn automatic software bug localization, source code analysis is usually used to localize the buggy code without manual intervention. However, due to considering irrelevant source code, localization accuracy may get biased. In this paper, a Method level Bug localization using Minimized search space (MBuM) is proposed for improving the accuracy, which considers only the liable source code for generating a bug. The relevant search space for a bug is extracted using the execution trace of the source code. By processing these relevant source code and the bug report, code and bug corpora are generated. Afterwards, MBuM ranks the source code methods based on the textual similarity between the bug and code corpora. To do so, modified Vector Space Model (mVSM) is used which incorporates the size of a method with Vector Space Model. Rigorous experimental analysis using different case studies are conducted on two large scale open source projects namely Eclipse and Mozilla. Experiments show that MBuM outperforms existing bug localization techniques. Shanto Rahman, Kazi Sakib |
ENASE | 1 |
| 2016 | Bilateral histogram equalization with pre-processing for contrast enhancementabstractVarious methods have been proposed for enhancing the images. Some of those perform well in some specific application areas but most of the techniques suffer from artifacts due to over enhancement. To overcome this problem, we have introduced a new image enhancement technique namely Bilateral Histogram Equalization with Pre-processing (BHEP) which uses Harmonic mean to divide the histogram of the image. We have performed both qualitative and quantitative measurements for experiments and the results show that BHEP creates less artifacts in several standard images than the existing state-of-the-art image enhancement techniques. Feroz Mahmud Amil, Mostafijur Rahman, Shanto Rahman, Emon Kumar Dey, Mohammad Shoyaib |
SNPD | 3 |