Tanvir Mahmud

dblp:255/3041 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0003-0529-2826ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
abstract
Video-to- video synthesis poses significant challenges in maintaining character consistency, smooth temporal tran-sitions, and preserving visual quality during fast motion. While recent fully cross-frame self-attention mechanisms have improved character consistency across multiple frames, they come with high computational costs and often include re-dundant operations, especially for videos with higher frame rates. To address these inefficiencies, we propose an adaptive motion-guided cross-frame attention mechanism that selectively reduces redundant computations. This enables a greater number of cross-frame attentions over more frames within the same computational budget, thereby enhancing both video quality and temporal coherence. Our method leverages optical flow to focus on moving regions while sparsely attending to stationary areas, allowing for the joint editing of more frames without increasing computational demands. Traditional frame interpolation techniques struggle with motion blur and flickering in intermediate frames, which compromises visual fidelity. To mitigate this, we intro-duce KV-caching for jointly edited frames, reusing keys and values across intermediate frames to preserve visual quality and maintain temporal consistency throughout the video. With our adaptive cross-frame self-attention approach, we achieve a threefold increase in the number of keyframes processed compared to existing methods, all within the same computational budget as fully cross-frame attention base-lines. This results in significant improvements in prediction accuracy and temporal consistency, outperforming state-of-the-art approaches. Code is made publicly available at https://github.com/tanvir-utexaslAdaVEltree/main.
Tanvir Mahmud, Mustafa Munir, Radu Marculescu, Diana Marculescu
WACV1
2024 T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
abstract
Visual sound source localization poses a significant chal-lenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggle to accu-rately distinguish the semantic regions of each sounding object, particularly in multi-source mixtures. These methods often rely on audio-visual correspondence as guidance, which can lead to substantial performance drops in com-plex multi-source localization scenarios. The lack of access to individual source sounds in multi-source mixtures during training exacerbates the difficulty of learning effective audio-visual correspondence for localization. To ad-dress this limitation, in this paper, we propose incorpo-rating the text modality as an intermediate feature guide using tri-modal joint embedding models (e.g., Audio Clip) to disentangle the semantic audio-visual source correspon-dence in multi-source mixtures. Our framework, dubbed T-VSL, begins by predicting the class of sounding enti-ties in mixtures. Subsequently, the textual representation of each sounding source is employed as guidance to dis-entangle fine-grained audio-visual source correspondence from multi-source mixtures, leveraging the tri-modal Audio-CLIP embedding. This approach enables our framework to handle a flexible number of sources and exhibits promising zero-shot transferability to unseen classes during test time. Extensive experiments conducted on the MUSIC, VG-GSound, and VGGSound-Instruments datasets demonstrate significant performance improvements over state-of-the-art methods. Code is released at https://github.com/enyac-group/T-VSL/tree/main.
Tanvir Mahmud, Yapeng Tian, Diana Marculescu
CVPR1
2024 PaPr: Training-Free One-Step Patch Pruning with Lightweight ConvNets for Faster Inference
Tanvir Mahmud, Burhaneddin Yaman, Chun-Hao Liu, Diana Marculescu
ECCV (23)1
2024 OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
abstract
Audio separation in real-world scenarios, where mixtures contain a variable number of sources, presents significant challenges due to limitations of existing models, such as over-separation, under-separation, and dependence on predefined training sources. We propose OpenSep, a novel framework that leverages large language models (LLMs) for automated audio separation, eliminating the need for manual intervention and overcoming source limitations. OpenSep uses textual inversion to generate captions from audio mixtures with off-the-shelf audio captioning models, effectively parsing the sound sources present. It then employs few-shot LLM prompting to extract detailed audio properties of each parsed source, facilitating separation in unseen mixtures. Additionally, we introduce a multi-level extension of the mix-and-separate training framework to enhance modality alignment by separating single source sounds and mixtures simultaneously. Extensive experiments demonstrate OpenSep’s superiority in precisely separating new, unseen, and variable sources in challenging mixtures, outperforming SOTA baseline methods. Code is released at https://github.com/tanvir-utexas/OpenSep.git.
Tanvir Mahmud, Diana Marculescu
EMNLP1
2024 Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
abstract
Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervision signal for single source separation cases during training. However, in the case of language-conditional audio separation, we do have access to corresponding text descriptions for each audio mixture in our training data, which can be seen as (rough) representations of the audio samples in the language modality. That raises the curious question of how to generate supervision signal for single-source audio extraction by leveraging the fact that single-source sounding language entities can be easily extracted from the text description. To this end, in this paper, we propose a generic bi-modal separation framework which can enhance the existing unsupervised frameworks to separate single-source signals in a target modality (i.e., audio) using the easily separable corresponding signals in the conditioning modality (i.e., language), without having access to single-source samples in the target modality during training. We empirically show that this is well within reach if we have access to a pretrained joint embedding model between the two modalities (i.e., CLAP). Furthermore, we propose to incorporate our framework into two fundamental scenarios to enhance separation performance. First, we show that our proposed methodology significantly improves the performance of purely unsupervised baselines by reducing the distribution shift between training and test samples. In particular, we show that our framework can achieve 71% boost in terms of Signal-to-Distortion Ratio (SDR) over the baseline, reaching 97.5% of the supervised learning performance. Second, we show that we can further improve the performance of the supervised learning itself by 17% if we augment it by our proposed weakly-supervised framework. Our framework achieves this by making large corpora of unsupervised data available to the supervised learning model as well as utilizing a natural, robust regularization mechanism through weak supervision from the language modality, and hence enabling a powerful semi-supervised framework for audio separation. Code is released at https://github.com/microsoft/BiModalAudioSeparation.
Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu
ICLR1
2024 SSVOD: Semi-Supervised Video Object Detection with Sparse Annotations
abstract
Despite significant progress in semi-supervised learning for image object detection, several key issues are yet to be addressed for video object detection: (1) Achieving good performance for supervised video object detection greatly depends on the availability of annotated frames. (2) Despite having large inter-frame correlations in a video, collecting annotations for a large number of frames per video is expensive, time-consuming, and often redundant. (3) Existing semi-supervised techniques on static images can hardly exploit the temporal motion dynamics inherently present in videos. In this paper, we introduce SSVOD, an end-to-end semi-supervised video object detection framework that exploits motion dynamics of videos to utilize large-scale unlabeled frames with sparse annotations. To selectively assemble robust pseudo-labels across groups of frames, we introduce flow-warped predictions from nearby frames for temporal-consistency estimation. In particular, we introduce cross-IoU and cross-divergence based selection methods over a set of estimated predictions to include robust pseudo-labels for bounding boxes and class labels, respectively. To strike a balance between confirmation bias and uncertainty noise in pseudo-labels, we propose confidence threshold based combination of hard and soft pseudo-labels. Our method achieves significant performance improvements over existing methods on ImageNet-VID, Epic-KITCHENS, and YouTube-VIS datasets. Codes are available at https://github.com/enyacgroup/SSVOD.git.
Tanvir Mahmud, Chun-Hao Liu, Burhaneddin Yaman, Diana Marculescu
WACV1
2024 MD-CardioNet: A Multi-Dimensional Deep Neural Network for Cardiovascular Disease Diagnosis From Electrocardiogram
abstract
Automated classification of cardiovascular diseases from electrocardiogram (ECG) signals using deep learning has gained significant interest due to its wide range of applications. However, existing deep learning approaches often overlook inter-channel shared information or lose time-sequence dependent information when considering 1D and 2D ECG representations, respectively. Moreover, besides considering spatial dimension, it is necessary to understand the context of the signals from a global feature space. We propose MD-CardioNet, an efficient deep learning architecture that captures temporal, spatial, and volumetric features from multi-lead ECG signals using multidimensional (1D, 2D, and 3D) convolutions to address these challenges. Sequential feature extractors capture time-dependent information, while a 2D convolution is applied to form an image representation from the multi-channel ECG signal, extracting inter-channel features. Additionally, a volumetric feature extraction network is designed to incorporate intra-channel, inter-channel, and inter-filter global space information. To reduce computational complexity, we introduce a practical knowledge distillation framework that reduces the number of trainable parameters by up to eight times ( from 4,304,910 parameters to 94,842 parameters) while maintaining satisfactory performance compatible with the other existing approaches. The proposed architecture is evaluated on a large publicly available dataset containing ECG signals from over 10,000 patients, achieving an accuracy of 97.3% in classifying six heartbeat rhythms. Our results surpass the performance of some state-of-the-art approaches. This paper presents a novel deep-learning approach for ECG classification that addresses the limitations of existing methods. The experimental results highlight the robustness and accuracy of MD-CardioNet in cardiovascular disease classification, offering valuable insights for future research in this field.
Md Toki Tahmid, Muhammad Ehsanul Kader, Tanvir Mahmud, Shaikh Anowarul Fattah
IEEE J. Biomed. Health Informatics3
2023 CLIP4VideoCap: Rethinking Clip for Video Captioning with Multiscale Temporal Fusion and Commonsense Knowledge
abstract
In this paper, we propose CLIP4VideoCap for video captioning based on large-scale pre-trained CLIP image and text encoders together with multi-scale temporal reasoning and commonsense knowledge. In addition to the CLIP-image encoder operating on successive video frames, we introduce a knowledge distillation-based learning scheme that aims to exploit the CLIP-text encoder to generate rich textual knowledge from the image features. For improved temporal reasoning over the video, we propose a multi-scale temporal fusion scheme that accumulates temporal features from different temporal windows. In addition, we integrate various commonsense aspects in the caption generation which greatly enhances the caption quality by extracting the commonsense features from the video in the intermediate phase. Combining these strategies, we achieve state-of-the-art performance on the benchmark MSR-VTT dataset confirming that our framework significantly outperforms existing approaches.
Tanvir Mahmud, Yaling Qing, Diana Marculescu
ICASSP1
2023 AVE-CLIP: AudioCLIP-based Multi-window Temporal Transformer for Audio Visual Event Localization
abstract
An audio-visual event (AVE) is denoted by the correspondence of the visual and auditory signals in a video segment. Precise localization of the AVEs is very challenging since it demands effective multi-modal feature correspondence to ground the short and long range temporal interactions. Existing approaches struggle in capturing the different scales of multi-modal interaction due to ineffective multi-modal training strategies. To overcome this limitation, we introduce AVE-CLIP, a novel framework that integrates the AudioCLIP pre-trained on large-scale audio-visual data with a multi-window temporal transformer to effectively operate on different temporal scales of video frames. Our contributions are three-fold: (1) We introduce a multi-stage training framework to incorporate AudioCLIP pre-trained with audio-image pairs into the AVE localization task on video frames through contrastive fine-tuning, effective mean video feature extraction, and multi-scale training phases. (2) We propose a multi-domain attention mechanism that operates on both temporal and feature domains over varying timescales to fuse the local and global feature variations. (3) We introduce a temporal refining scheme with event-guided attention followed by a simple-yet-effective post processing step to handle significant variations of the background over diverse events. Our method achieves state-of-the-art performance on the publicly available AVE dataset with 5.9% mean accuracy improvement which proves its superiority over existing approaches.
Tanvir Mahmud, Diana Marculescu
WACV1
2021 CovTANet: A Hybrid Tri-Level Attention-Based Network for Lesion Segmentation, Diagnosis, and Severity Prediction of COVID-19 Chest CT Scans
abstract
Rapid and precise diagnosis of COVID-19 is one of the major challenges faced by the global community to control the spread of this overgrowing pandemic. In this article, a hybrid neural network is proposed, named CovTANet, to provide an end-to-end clinical diagnostic tool for early diagnosis, lesion segmentation, and severity prediction of COVID-19 utilizing chest computer tomography (CT) scans. A multiphase optimization strategy is introduced for solving the challenges of complicated diagnosis at a very early stage of infection, where an efficient lesion segmentation network is optimized initially, which is later integrated into a joint optimization framework for the diagnosis and severity prediction tasks providing feature enhancement of the infected regions. Moreover, for overcoming the challenges with diffused, blurred, and varying shaped edges of COVID lesions with novel and diverse characteristics, a novel segmentation network is introduced, namely tri-level attention-based segmentation network. This network has significantly reduced semantic gaps in subsequent encoding-decoding stages, with immense parallelization of multiscale features for faster convergence providing considerable performance improvement over traditional networks. Furthermore, a novel tri-level attention mechanism has been introduced, which is repeatedly utilized over the network, combining channel, spatial, and pixel attention schemes for faster and efficient generalization of contextual information embedded in the feature map through feature recalibration and enhancement operations. Outstanding performances have been achieved in all three tasks through extensive experimentation on a large publicly available dataset containing 1110 chest CT-volumes, which signifies the effectiveness of the proposed scheme at the current stage of the pandemic.
Tanvir Mahmud, Md. Jahin Alam, Sakib Chowdhury, Shams Nafisa Ali, Md Maisoon Rahman, Shaikh Anowarul Fattah, Mohammad Saquib
IEEE Trans. Ind. Informatics1
2020 ResCovNet: A Deep Learning-Based Architecture For COVID-19 Detection From Chest CT Scan Images
abstract
Automatic disease detection using machine learning-based techniques from X-ray and computed tomography (CT) can play a major role in the frontline to assist medical professionals during the current outbreak of COVID-19. Fast diagnosis of the disease is the key to reduce the uncontrollable spread of this life-threatening disease, where machine learning-based applications can contribute greatly by predicting the situation of patients so that professionals can decide accordingly. The major drawbacks of detecting COVID-19 are its similarities with different types of pneumonia, and the absence of properly labeled data. Considering the ResNet152V2 as a backbone network, an efficient architecture, namely ResCovNet is proposed to detect COVID-19 accurately from chest CT scan images by separating it from three types of pneumonia and normal cases. Otsu's thresholding is applied in the pre-processing step to strengthen the features for the classification network. With the use of proposed architecture, a very satisfactory classification accuracy of 88.1% is achieved to separate COVID-19 from all other four classes. Evaluating the performance of this study by 3-fold cross-validation, and comparison with related works prove that this adroit algorithm provides an effective way to be implemented as a diagnostic tool in the COVID-19 screening.
Ankan Ghosh Dastider, Mohseu Rashid Subah, Farhan Sadik, Tanvir Mahmud, Shaikh Anowarul Fattah
TENCON4
2020 Transfer Learning Based Method for COVID-19 Detection From Chest X-ray Images
abstract
Radiology examination of chest radiography or chest X-ray (CXR), is currently performed manually by radiologists. With the onset of the COVID-19 pandemic, there is now a need to automate this process which is currently one of the key methods of primary detection of the SARS-Cov-2 virus. This will lead to shorter diagnosis time and less human error. In this study, we try to perform three-class image classification on a dataset of chest X-rays of confirmed COVID-19 patients(408 images), confirmed pneumonia patients(4273 images), and chest X-rays of healthy people(1590 images). In total the dataset consists of 6271 people. We aim to use a Convolutional Neural Network(CNN) and transfer learning to perform this image classification task. Our model is based on a pre-trained InceptionV3 network with weights trained on the ImageNet dataset. We fine-tune the layers of the Inception network to train it to our specific task. We try fine-tuning the network to different extents by freezing a different number of layers and then comparing accuracy for each variation of the network. To evaluate the performance of our network we use several metrics which include Classification accuracy, Precision, Sensitivity, and Specificity. Our proposed method achieves an accuracy of 96.33% on a 3-class classification task (Normal, COVID-19, Pneumonia) and an accuracy of 99.39% on a 2-class (COVID and Non-COVID) classification task.
Nayeeb Rashid, Md Adnan Faisal Hossain, Mumtahina Islam Sukanya, Tanvir Mahmud, Shaikh Anowarul Fattah
TENCON5
2020 A Multi-Model Based Ensembling Approach to Detect COVID-19 from Chest X-Ray Images
abstract
Since the onset of COVID-19, radiographic image analysis coupled with artificial intelligence (AI) has become popular due to insufficient RT-PCR test kits. In this paper, an automated AI-assisted COVID-19 diagnosis scheme is proposed utilizing the ensembling approach of multiple convolutional neural networks (CNNs). Two different strategies have been carried out for ensembling: A feature level fusionbased ensembling method and a decision level ensembling method. Several traditional CNN architectures are tested and finally in the ensembling operation, MobileNet, InceptionV3, DenseNet201, DenseNet121 and Xception are used. To handle the computational complexity of multiple networks, transfer learning strategy is incorporated through ImageNet pre-trained weight initialization. For feature-level ensembling scheme, global averages of the convolutional feature maps generated from multiple networks are aggregated and undergo through fully connected layers for combined optimization. Additionally, for decision level ensembling scheme, final prediction generated from multiple networks are converged into a single prediction by utilizing the maximum voting criterion. Both strategies perform better than any individual network. Outstanding performances have been achieved through extensive experimentation on a public database with 96% accuracy on 3-class (COVID-19/normal/pneumonia) diagnosis and 89.21% on 4-class (COVID-19/normal/viral pneumonia/bacterial pneumonia) diagnosis.
Oishy Saha, Jarin Tasnim, Md. Tanvir Raihan, Tanvir Mahmud, Istak Ahmmed, Shaikh Anowarul Fattah
TENCON4