VLDB 2026 Research / reviewers in the wild / expert
Shibbir Ahmed
dblp:241/9346
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-1183-883XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Precise LLM-based Semantic Slicing for Deadlock Detection in Concurrent Java ProgramsabstractFormal verification of concurrent programs is essential in mission-critical domains, where deadlocks and concurrency bugs can cause catastrophic failures. However, traditional model-checking approaches face severe scalability issues and often rely on static slicing techniques that fail to capture semantic dependencies, resulting in large, non-executable slices. This paper introduces a novel method leveraging large language models (LLMs) pre-trained on Java bytecode, specifically byteBERT, to generate precise, executable program slices. This approach enables efficient deadlock detection in concurrent Java programs by combining LLM-driven semantic slicing with rigorous model checking via Java Pathfinder. Through a case study on a code repository, we demonstrate an 87% reduction in executable slice size. The results indicate that bytecode-trained LLMs substantially enhance both the scalability and precision of automated software test workflows for concurrent systems. Taythir G. Martin, Rodion Podorozhny, Shibbir Ahmed |
ICPC | 3 |
| 2025 | Can Large Language Models Challenge CNNs in Medical Image Analysis?abstractThis study presents a multimodal AI framework designed for precisely classifying medical diagnostic images. Utilizing publicly available datasets, the proposed system compares the strengths of convolutional neural networks (CNNs) and different large language models (LLMs). This in-depth comparative analysis highlights key differences in diagnostic performance, execution efficiency, and environmental impacts. Model evaluation was based on accuracy, F1-score, average execution time, average energy consumption, and estimated CO2emission. The findings indicate that although CNN-based models can outperform various multimodal techniques that incorporate both images and contextual information, applying additional filtering on top of LLMs can lead to substantial performance gains. These findings highlight the transformative potential of multimodal AI systems to enhance the reliability, efficiency, and scalability of medical diagnostics in clinical settings. Shibbir Ahmed, Shahnewaz Karim Sakib, Anindya Bijoy Das |
ICIP | 1 |
| 2024 | Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentabstractDeep learning models are trained with certain assumptions about the data during the development stage and then used for prediction in the deployment stage. It is important to reason about the trustworthiness of the model's predictions with unseen data during deployment. Existing methods for specifying and verifying traditional software are insufficient for this task, as they cannot handle the complexity of DNN model architecture and expected outcomes. In this work, we propose a novel technique that uses rules derived from neural network computations to infer data preconditions for a DNN model to determine the trustworthiness of its predictions. Our approach, DeepInfer involves introducing a novel abstraction for a trained DNN model that enables weakest precondition reasoning using Dijkstra's Predicate Transformer Semantics. By deriving rules over the inductive type of neural network abstract representation, we can overcome the matrix dimensionality issues that arise from the backward non-linear computation from the output layer to the input layer. We utilize the weakest precondition computation using rules of each kind of activation function to compute layer-wise precondition from the given postcondition on the final output of a deep neural network. We extensively evaluated DeepInfer on 29 real-world DNN models using four different datasets collected from five different sources and demonstrated the utility, effectiveness, and performance improvement over closely related work. DeepInfer efficiently detects correct and incorrect predictions of high-accuracy models with high recall (0.98) and high F-1 score (0.84) and has significantly improved over the prior technique, SelfChecker. The average runtime overhead of DeepInfer is low, 0.22 sec for all the unseen datasets. We also compared runtime overhead using the same hardware settings and found that DeepInfer is 3.27 times faster than SelfChecker, the state-of-the-art in this area. Shibbir Ahmed, Hongyang Gao, Hridesh Rajan |
ICSE | 1 |
| 2023 | Design by Contract for Deep Learning APIsabstractDeep Learning (DL) techniques are increasingly being incorporated in critical software systems today. DL software is buggy too. Recent work in SE has characterized these bugs, studied fix patterns, and proposed detection and localization strategies. In this work, we introduce a preventative measure. We propose design by contract for DL libraries, DL Contract for short, to document the properties of DL libraries and provide developers with a mechanism to identify bugs during development. While DL Contract builds on the traditional design by contract techniques, we need to address unique challenges. In particular, we need to document properties of the training process that are not visible at the functional interface of the DL libraries. To solve these problems, we have introduced mechanisms that allow developers to specify properties of the model architecture, data, and training process. We have designed and implemented DL Contract for Python-based DL libraries and used it to document the properties of Keras, a well-known DL library. We evaluate DL Contract in terms of effectiveness, runtime overhead, and usability. To evaluate the utility of DL Contract, we have developed 15 sample contracts specifically for training problems and structural bugs. We have adopted four well-vetted benchmarks from prior works on DL bug detection and repair. For the effectiveness, DL Contract correctly detects 259 bugs in 272 real-world buggy programs, from well-vetted benchmarks provided in prior work on DL bug detection and repair. We found that the DL Contract overhead is fairly minimal for the used benchmarks. Lastly, to evaluate the usability, we conducted a survey of twenty participants who have used DL Contract to find and fix bugs. The results reveal that DL Contract can be very helpful to DL application developers when debugging their code. Shibbir Ahmed, Sayem Mohammad Imtiaz, Syeda Khairunnesa Samantha, Breno Dantas Cruz, Hridesh Rajan |
ESEC/SIGSOFT FSE | 1 |
| 2023 | What kinds of contracts do ML APIs need?
Syeda Khairunnesa Samantha, Shibbir Ahmed, Sayem Mohammad Imtiaz, Hridesh Rajan, Gary T. Leavens |
Empir. Softw. Eng. | 2 |