VLDB 2026 Research / reviewers in the wild / expert
Md. Adnan Arefeen
dblp:247/5721
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0001-6486-8181ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | iRAG: Advancing RAG for Videos with an Incremental ApproachabstractRetrieval-augmented generation (RAG) systems combine the strengths of language generation and information retrieval to power many real-world applications like chatbots. Use of RAG for understanding of videos is appealing but there are two critical limitations. One-time, upfront conversion of all content in large corpus of videos into text descriptions entails high processing times. Also, not all information in the rich video data is typically captured in the text descriptions. Since user queries are not known apriori, developing a system for video to text conversion and interactive querying of video data is challenging. Md. Adnan Arefeen, Biplob Debnath, Md. Yusuf Sarwar Uddin, Srimat T. Chakradhar |
CIKM | 1 |
| 2023 | FactionFormer: Context-Driven Collaborative Vision Transformer Models for Edge IntelligenceabstractEdge Intelligence has received attention in the recent times for its potential towards improving responsiveness, reducing the cost of data transmission, enhancing security and privacy, and enabling autonomous decisions by edge devices. However, edge devices lack the power and compute resources necessary to execute most Al models. In this paper, we present FactionFormer, a novel method to deploy resource-intensive deep-learning models, such as vision transformers (ViT), on resource-constrained edge devices. Our method is based on a key observation: edge devices are often deployed in settings where they encounter only a subset of the classes that the resource-intensive Al model is trained to classify, and this subset changes across deployments. Therefore, we automatically identify this subset as a faction, devise on-the fly a bespoke resource-efficient ViT called a modelette for the faction, and set up an efficient processing pipeline consisting of a modelette on the device, a wireless network such as 5G, and the resource-intensive ViT model on an edge server, all of which work collaboratively to do the inference. For several ViT models pre-trained on benchmark datasets, FactionFormer’s modelettes are up to 4× smaller than the corresponding baseline models in terms of the number of parameters, and they can infer up to 2.5× faster than the baseline setup where every input is processed by the resource-intensive ViT on the edge server. Our work is the first of its kind to propose a device-edge collaborative inference framework where bespoke deep learning models for the device are automatically devised on-the-fly for most frequently encountered subset of classes. Sumaiya Tabassum Nimi, Md. Adnan Arefeen, Md. Yusuf Sarwar Uddin, Biplob Debnath, Srimat T. Chakradhar |
SMARTCOMP | 2 |
| 2022 | FrameHopper: Selective Processing of Video Frames in Detection-driven Real-Time Video AnalyticsabstractDetection-driven real-time video analytics require continuous detection of objects contained in the video frames using deep learning models like YOLOV3, EfficientDet, etc. However, running these detectors on each and every frame in resource-constrained edge devices is computationally intensive. By taking the temporal correlation between consecutive video frames into account, we note that detection outputs tend to be overlapping in successive frames. Elimination of “similar” consecutive frames (the same set of objects with slightly offset bounding boxes) will lead to a negligible drop in performance while offering significant performance benefits by reducing overall computation and communication costs. The key technical questions are, therefore, (a) how to identify which frames to be processed by the object detector, and (b) how many successive frames can be skipped (called skip-length) once a frame is selected to be processed. The overall goal of the process is to keep the error due to skipping frames as small as possible. We introduce a novel error vs processing rate optimization problem with respect to the object detection task that balances between the error rate and the fraction of frames actually passed and processed. Subsequently, we propose an off-line Reinforcement Learning (RL)-based algorithm to determine these skip-lengths as a state-action policy of the RL agent from a recorded video and then deploy the agent online for live video streams. To this end, we develop FrameHopper, an edge-cloud collaborative video analytics framework, that runs a lightweight trained RL agent on the camera and passes filtered frames to the cloud/edge server where the object detection model runs for a set of applications. We have tested our approach on a number of live videos captured from real-life scenarios and show that FrameHopper processes only a handful of frames but produces detection results closer to the "oracle" solution and outperforms recent state-of-the-art solutions in most cases. Md. Adnan Arefeen, Sumaiya Tabassum Nimi, Md. Yusuf Sarwar Uddin |
DCOSS | 1 |
| 2022 | Chimera: Context-Aware Splittable Deep Multitasking Models for Edge IntelligenceabstractDesign of multitasking deep learning models has mostly focused on improving the accuracy of the constituent tasks, but the challenges of efficiently deploying such models in a device-edge collaborative setup (that is common in 5G deployments) has not been investigated. Towards this end, in this paper, we propose an approach called Chimera1for training (done Offline) and deployment (done Online) of multitasking deep learning models that are splittable across the device and edge. In the offline phase, we train our multi-tasking setup such that features from a pre-trained model for one of the tasks (called the Primary task) are extracted and task-specific sub-models are trained to generate the other (Secondary) tasks' outputs through a knowledge distillation like training strategy to mimic the outputs of pre-trained models for the tasks. The task-specific sub-models are designed to be significantly lightweight than the original pre-trained models for the Secondary tasks. Once the sub-models are trained, during deployment, for given deployment context, characterized by the configurations, we search for the optimal (in terms of both model performance and cost) deployment strategy for the generated multitasking model, through finding one or multiple suitable layer(s) for splitting the model, so that inference workloads are distributed between the device and the edge server and the inference is done in a collaborative manner. Extensive experiments on benchmark computer vision tasks demonstrate that Chimera generates splittable multitasking models that are at least ~ 3 x parameter efficient than the existing such models, and the end-to-end device-edge collaborative inference becomes ~ 1.35 x faster with our choice of context-aware splitting decisions. Sumaiya Tabassum Nimi, Md. Adnan Arefeen, Md. Yusuf Sarwar Uddin, Biplob Debnath, Srimat T. Chakradhar |
SMARTCOMP | 2 |
| 2022 | Neural Network-Based Undersampling TechniquesabstractMachine learning models have gained popularity nowadays for their potential to solve real-life issues when trained on pertinent data. In many cases, the real-life data are class imbalanced and hence the corresponding machine learning models trained on the data tend to perform poorly on metrics like precision, recall, AUC, F1, and G-mean score. Since class imbalance issue poses serious challenges to the performance of trained models, a multitude of research works have addressed this issue. Two common data-based sampling techniques have mostly been proposed-undersampling the data of the majority class and oversampling the data of the minority class. In this article, we focus on the former approach. We propose two novel algorithms that employ neural network-based approaches to remove majority samples that are found to reside in the vicinity of the minority samples, thereby undersampling the former to remove (or alleviate) the imbalance issue. We delineate the proposed algorithms and then test the proposed algorithms on some publicly available imbalanced datasets. We then compare the performance of our proposed algorithms to other popular undersampling algorithms. Finally, we conclude that our proposed algorithms outperform most of the existing undersampling approaches on most performance metrics. Md. Adnan Arefeen, Sumaiya Tabassum Nimi, Mohammad Sohel Rahman |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | TransJury: Towards Explainable Transfer Learning through Selection of Layers from Deep Neural NetworksabstractTraining a neural network model from scratch is a computationally intensive operation. To alleviate this issue, researchers often employ "transfer learning" that transfers knowledge from a source data distribution to a target data distribution, instead of training the whole model from scratch. Typically, the last few layers of a pretrained convolutional neural network (CNN) are chosen for many transfer learning tasks where the outputs of those selected layers are combined to construct a feature space based on which a task-specific classification n etwork i s t rained o r fi ne-tuned. Th is arbitrary way of selecting layers, however, often fails to achieve the desired accuracy for the target task. What we need is an intelligent way of selecting layers from a pretrained model for a given task so that the additional overhead of successive training remains low. To this end, we propose a novel method, called TransJury, to find t he m ost s ignificant la yers from a pretrained mo del for transfer learning along with preserving the knowledge for the source domain. Through extensive experimentation on several target domain datasets, we show the supremacy of our approach in terms of lower training overhead and improved accuracy. By deploying MobileNet-v2, a lightweight CNN model pretrained on the ImageNet dataset on an edge device, we also discuss the future direction of this research. Md. Adnan Arefeen, Sumaiya Tabassum Nimi, Md. Yusuf Sarwar Uddin, Yugyung Lee |
IEEE BigData | 1 |
| 2021 | A Lightweight Relu-Based Feature Fusion For Aerial Scene ClassificationabstractIn this paper, we propose a transfer-learning based model construction technique for the aerial scene classification problem. The core of our technique is a layer selection strategy, named ReLU-Based Feature Fusion (RBFF), that extracts feature maps from a pretrained CNN-based single-object image classification model, namely MobileNetV2, and constructs a model for the aerial scene classification task. RBFF stacks features extracted from the batch normalization layer of a few selected blocks of MobileNetV2, where the candidate blocks are selected based on the characteristics of the ReLU activation layers present in those blocks. The feature vector is then compressed into a low-dimensional feature space using dimension reduction algorithms on which we train a low-cost SVM classifier for the classification of the aerial images. We validate our choice of selected features based on the significance of the extracted features with respect to our classification pipeline. RBFF remarkably does not involve any training of the base CNN model except for a few parameters for the classifier, which makes the technique very cost-effective for practical deployments. The constructed model despite being lightweight outperforms several recently proposed models in terms of accuracy for a number of aerial scene datasets. Md. Adnan Arefeen, Sumaiya Tabassum Nimi, Md. Yusuf Sarwar Uddin, Zhu Li 0001 |
ICIP | 1 |
| 2021 | Towards resource-efficient detection-driven processing of multi-stream videosabstractDetection-driven video analytics is resource hungry as it depends on running object detectors on video frames. Running an object detection engine (i.e., deep learning models such as YOLO and EfficientDet) for each frame makes video analytics pipelines difficult to achieve real-time processing. In this paper, we leverage selective processing of frames and batching of frames to reduce the overall cost of running detection models on live videos. We discuss several factors that hinder the real-time processing of detection-driven video analytics. We propose a system with configurable knobs and show how to achieve the stability of the system using a Lyapunov-based control strategy. In our setup, heterogeneous edge devices (e.g. mobile phones, cameras) stream videos to a low-resource edge server where frames are selectively processed in batches and the detection results are sent to the cloud or to the edge device for further application-aware processing. Preliminary results on controlling different knobs, such as frame skipping, frame size, and batch size show interesting insights to achieve real-time processing of multi-stream video streams with low overhead and low overall information loss. Md. Adnan Arefeen, Md. Yusuf Sarwar Uddin |
MobiCom | 1 |
| 2021 | EARLIN: Early Out-of-Distribution Detection for Resource-Efficient Collaborative Inference
Sumaiya Tabassum Nimi, Md. Adnan Arefeen, Md. Yusuf Sarwar Uddin, Yugyung Lee |
ECML/PKDD (1) | 2 |