Mrigank Rochan

dblp:12/10720 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-9513-6573ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 20 · 8 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Test-Time Adaptation for Video Highlight Detection Using Meta-Auxiliary Learning and Cross-Modality Hallucinations
abstract
Existing video highlight detection methods, although advanced, struggle to generalize well to all test videos. These methods typically employ a generic highlight detection model for each test video, which is suboptimal as it fails to account for the unique characteristics and variations of individual test videos. Such fixed models do not adapt to the diverse content, styles, or audio and visual qualities present in new, unseen test videos, leading to reduced highlight detection performance. In this paper, we propose Highlight-TTA, a test-time adaptation framework for video highlight detection that addresses this limitation by dynamically adapting the model during testing to better align with the specific characteristics of each test video, thereby improving generalization and highlight detection performance. Highlight-TTA is jointly optimized with an auxiliary task, cross-modality hallucinations, alongside the primary highlight detection task. We utilize a meta-auxiliary training scheme to enable effective adaptation through the auxiliary task while enhancing the primary task. During testing, we adapt the trained model using the auxiliary task on the test video to further enhance its highlight detection performance. Extensive experiments with three state-of-the-art highlight detection models and three benchmark datasets show that the introduction of Highlight-TTA to these models improves their performance, yielding superior results.
Sujoy Paul, Mrigank Rochan
WACV3
2026 A metamorphic testing perspective on knowledge distillation for language models of code: Does the student deeply mimic the teacher?
abstract
Transformer-based language models of code have achieved state-of-the-art performance across a wide range of software analytics tasks, but their practical deployment remains limited due to high computational costs, slow inference speeds, and significant environmental impact. To address these challenges, recent research has increasingly explored knowledge distillation as a method for compressing a large language model of code (the teacher) into a smaller model (the student) while maintaining performance. However, the degree to which a student model deeply mimics the predictive behavior and internal representations of its teacher remains largely unexplored, as current accuracy-based evaluation provides only a surface-level view of model quality and often fails to capture more profound discrepancies in behavioral fidelity between the teacher and student models. To address this gap, we empirically show that the student model often fails to deeply mimic the teacher model, resulting in up to 285% greater performance drop under adversarial attacks, which is not captured by traditional accuracy-based evaluation. In addition, to capture discrepancies in behavioral fidelity, we propose MetaCompress , a metamorphic testing framework that systematically evaluates behavioral fidelity by comparing the outputs of teacher and student models under a set of behavior-preserving metamorphic relations. We evaluate MetaCompress on two widely studied tasks, clone detection and vulnerability prediction, using compressed versions of popular language models of code, CodeBERT and GraphCodeBERT , obtained via three different knowledge distillation techniques: Compressor , AVATAR , and MORPH . The results show that MetaCompress identifies up to 62% behavioral discrepancies in student models, underscoring the need for behavioral fidelity evaluation within the knowledge distillation pipeline. Furthermore, the ablation study indicates that MetaCompress is robust and effectively detects behavioral fidelity divergence, even when both the teacher and student models are assessed using transformed inputs. These results position MetaCompress as a practical framework for evaluating the behavioral fidelity of compressed language models of code derived through knowledge distillation.
Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy
J. Syst. Softw.2
2025 Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?
abstract
Large Language Model (LLM)-generated data is increasingly used in software analytics, but it is unclear how this data compares to human-written data, particularly when models are exposed to adversarial scenarios. Adversarial attacks can compromise the reliability and security of software systems, so understanding how LLM-generated data performs under these conditions, compared to human-written data, which serves as the benchmark for model performance, can provide valuable insights into whether LLM-generated data offers similar robustness and effectiveness. To address this gap, we systematically evaluate and compare the quality of human-written and LLM-generated data for fine-tuning robust pre-trained models (PTMs) in the context of adversarial attacks. We evaluate the robustness of six widely used PTMs, fine-tuned on human-written and LLM-generated data, before and after adversarial attacks. This evaluation employs nine state-of-the-art (SOTA) adversarial attack techniques across three popular software analytics tasks: clone detection, code summarization, and sentiment analysis in code review discussions. Additionally, we analyze the quality of the generated adversarial examples using eleven similarity metrics. Our findings reveal that while PTMs fine-tuned on LLM-generated data perform competitively with those fine-tuned on human-written data, they exhibit less robustness against adversarial attacks in software analytics tasks. Our study underscores the need for further exploration into enhancing the quality of LLM-generated training data to develop models that are both high-performing and capable of withstanding adversarial attacks in software analytics.
Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy
EASE2
2025 Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence
abstract
With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user experiences by allowing quick access to relevant content across diverse domains. Existing methods typically rely either on expensive manually labeled frame-level annotations, or on a large external dataset of videos for weak supervision through category information. To overcome this, we focus on unsupervised video highlight detection, eliminating the need for manual annotations. We propose a novel unsupervised approach which capitalizes on the premise that significant moments tend to recur across multiple videos of the similar category in both audio and visual modalities. Surprisingly, audio remains under-explored, especially in unsupervised algorithms, despite its potential to detect key moments. Through a clustering technique, we identify pseudo-categories of videos and compute audio pseudo-highlight scores for each video by measuring the similarities of audio features among audio clips of all the videos within each pseudo-category. Similarly, we also compute visual pseudo-highlight scores for each video using visual features. Then, we combine audio and visual pseudo-highlights to create the audio-visual pseudo ground-truth highlight of each video for training an audio-visual highlight detection network. Extensive experiments and ablation studies on three benchmarks showcase the superior performance of our method over prior work.
Sujoy Paul, Mrigank Rochan
WACV3
2025 Investigating adversarial attacks in software analytics via machine learning explainability
Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy
Softw. Qual. J.2
2024 Effects of Range-based LiDAR Point Cloud Density Manipulation on 3D Object Detection
abstract
In recent years, much progress has been made in LiDAR-based 3D object detection mainly due to advances in detector architecture designs and availability of large-scale LiDAR datasets. Existing 3D object detectors tend to perform well on the point cloud regions closer to the LiDAR sensor as opposed to on regions that are farther away. In this paper, we investigate this problem from the data perspective instead of detector architecture design. We observe that there is a learning bias in detection models towards the dense objects near the sensor and show that the detection performance can be improved by simply manipulating the input point cloud density at different distance ranges without modifying the detector architecture and without data augmentation. We propose a model-free point cloud density adjustment pre-processing mechanism that uses iterative MCMC optimization to estimate near-optimal parameters for altering the point density at different distance ranges. We conduct experiments using four state-of-the-art LiDAR 3D object detectors on two public LiDAR datasets, namely Waymo and ONCE. Our results demonstrate that our range-based point cloud density manipulation technique can improve the performance of the existing detectors, which in turn could potentially inspire future detector designs.
Eduardo R. Corral-Soto, Alaap Grandhi, Yannis Y. He, Mrigank Rochan
IV4
2024 Gradual Batch Alternation for Effective Domain Adaptation in LiDAR-Based 3D Object Detection
abstract
We address the challenge of domain adaptation in LiDAR-based 3D object detection by introducing a simple yet effective training strategy known as Gradual Batch Alternation. This method enables adaptation from a well-labeled source domain to an insufficiently labeled target domain. Initially, training commences with alternating batches of samples from both the source and target domains. As the training progresses, we gradually reduce the number of samples from the source domain. Consequently, the model undergoes a gradual transition towards the target domain, resulting in improved adaptation. Domain adaptation experiments for 3D object detection on four benchmark autonomous driving datasets, namely ONCE, PandaSet, Waymo, and nuScenes, demonstrate significant performance gains over prior works and strong baselines.
Mrigank Rochan, Xingxin Chen, Alaap Grandhi, Eduardo R. Corral-Soto
IV1
2023 Domain Adaptation in LiDAR Semantic Segmentation via Hybrid Learning with Alternating Skip Connections
abstract
In this paper we address the challenging problem of domain adaptation in LiDAR semantic segmentation. We consider the setting where we have a fully-labeled data set from source domain and a target domain with a few labeled and many unlabeled examples. We propose a domain adaptation framework that mitigates the issue of domain shift and produces appealing performance on the target domain. To this end, we develop a GAN-based image-to-image translation engine that has generators with alternating connections, and couple it with a LiDAR semantic segmentation network. Our framework is hybrid in nature in the sense that our model learning is composed of self-supervision, semi-supervision and unsupervised learning. Extensive experiments on benchmark LiDAR semantic segmentation data sets demonstrate that our method achieves superior performance in comparison to state-of-the-art baselines and prior arts.
Eduardo R. Corral-Soto, Mrigank Rochan, Yannis Y. He, Xingxin Chen, Shubhra Aich
IV2
2022 Contrastive Learning for Unsupervised Video Highlight Detection
abstract
Video highlight detection can greatly simplify video browsing, potentially paving the way for a wide range of ap-plications. Existing efforts are mostly fully-supervised, requiring humans to manually identify and label the interesting moments (called highlights) in a video. Recent weakly supervised methods forgo the use of highlight annotations, but typically require extensive efforts in collecting external data such as web-crawled videos for model learning. This observation has inspired us to consider unsupervised highlight detection where neither frame-level nor video-level annotations are available in training. We propose a simple contrastive learning framework for unsupervised highlight detection. Our framework encodes a video into a vector representation by learning to pick video clips that help to distinguish it from other videos via a contrastive objective using dropout noise. This inherently allows our framework to identify video clips corresponding to highlight of the video. Extensive empirical evaluations on three highlight detection benchmarks demonstrate the superior performance of our approach.
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang 0003
CVPR2
2022 Unsupervised Domain Adaptation in LiDAR Semantic Segmentation with Self-Supervision and Gated Adapters
abstract
In this paper, we focus on a less explored, but more realistic and complex problem of domain adaptation in LiDAR semantic segmentation. There is a significant drop in performance of an existing segmentation model when training (source domain) and testing (target domain) data originate from different LiDAR sensors. To overcome this shortcoming, we propose an unsupervised domain adaptation framework that leverages unlabeled target domain data for self-supervision, coupled with an unpaired mask transfer strategy to mitigate the impact of domain shifts. Furthermore, we introduce the gated adapter module with a small number of parameters into the network to account for target domain-specific information. Experiments adapting from both real-to-real and synthetic-to-real LiDAR semantic segmentation benchmarks demonstrate the significant improvement over prior arts.
Mrigank Rochan, Shubhra Aich, Eduardo R. Corral-Soto, Amir Nabatchian
ICRA1
2022 Referring Segmentation in Images and Videos With Cross-Modal Self-Attention Network
abstract
We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods.
Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Xiaoqin Zhang 0002, Yang Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 AdaCrowd: Unlabeled Scene Adaptation for Crowd Counting
abstract
We address the problem of image-based crowd counting. In particular, we propose a new problem calledunlabeled scene-adaptive crowd counting. Given a new target scene, we would like to have a crowd counting model specifically adapted to this particular scene based on the target data that capture some information about the new scene. In this paper, we propose to use one or more unlabeled images from the target scene to perform the adaptation. In comparison with the existing problem setups (e.g. fully supervised), our proposed problem setup is closer to the real-world applications of crowd counting systems. We introduce a novelAdaCrowdframework to solve this problem. Our framework consists of a crowd counting network and a guiding network. The guiding network predicts some parameters in the crowd counting network based on the unlabeled images from a particular scene. This allows our model to adapt to different target scenes. The experimental results on several challenging benchmark datasets demonstrate the effectiveness of our proposed approach compared with other alternative methods. Code is available athttps://github.com/maheshkkumar/adacrowd
Mahesh Kumar Krishna Reddy, Mrigank Rochan, Yiwei Lu 0001, Yang Wang 0003
IEEE Trans. Multim.2
2021 Joint Visual and Audio Learning for Video Highlight Detection
abstract
In video highlight detection, the goal is to identify the interesting moments within an unedited video. Although the audio component of the video provides important cues for highlight detection, the majority of existing efforts focus almost exclusively on the visual component. In this paper, we argue that both audio and visual components of a video should be modeled jointly to retrieve its best moments. To this end, we propose an audio-visual network for video highlight detection. At the core of our approach lies a bimodal attention mechanism, which captures the interaction between the audio and visual components of a video, and produces fused representations to facilitate highlight detection. Furthermore, we introduce a noise sentinel technique to adaptively discount a noisy visual or audio modality. Empirical evaluations on two benchmark datasets demonstrate the superior performance of our approach over the state-of-the-art methods.
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang 0003
ICCV2
2020 Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Yang Wang 0003
BMVC1
2020 Adaptive Video Highlight Detection by Learning from User History
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye, Yang Wang 0003
ECCV (21)1
2020 Few-Shot Scene Adaptive Crowd Counting Using Meta-Learning
abstract
We consider the problem of few-shot scene adaptive crowd counting. Given a target camera scene, our goal is to adapt a model to this specific scene with only a few labeled images of that scene. The solution to this problem has potential applications in numerous real-world scenarios, where we ideally like to deploy a crowd counting model specially adapted to a target camera. We accomplish this challenge by taking inspiration from the recently introduced learning-to-learn paradigm in the context of few-shot regime. In training, our method learns the model parameters in a way that facilitates the fast adaptation to the target scene. At test time, given a target scene with a small number of labeled data, our method quickly adapts to that scene with a few gradient updates to the learned parameters. Our extensive experimental results show that the proposed approach outperforms other alternatives in few-shot scene adaptive crowd counting.
Mahesh Kumar Krishna Reddy, Mohammad Asiful Hossain, Mrigank Rochan, Yang Wang 0003
WACV3
2019 Video-Based Person Re-Identification using Refined Attention Networks
abstract
We consider the problem of video-based person reidentification. The goal is to identify a person from videos captured under different cameras. In this paper, we propose an efficient attention based model for person re-identifying from videos. Our method generates an attention score for each frame based on frame-level features. The attention scores of all frames in a video are used to produce a weighted feature vector for the input video. This video-level feature vector is refined iteratively for re-identifying persons from videos. Unlike most existing deep learning methods that use global or spatial representation, our approach focuses on attention scores. Extensive experiments on three benchmark datasets demonstrate that our method achieves the state-of-the-art performance.
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
AVSS2
2019 Non-Local Attentive Temporal Network for Video-Based Person Re-Identification
abstract
Given a video containing a person, the goal of person re-identification is to identify the same person from videos captured under different cameras. A common approach for tackling this problem is to first extract image features for all frames in the video. These frame-level features are then combined (e.g. via temporal pooling) to form a video-level feature vector. The video-level features of two input videos are then compared by calculating the distance between them. More recently, attention-based learning mechanism has been proposed for this problem. In particular, recurrent neural networks have been used to generate the attention scores of frames in a video. However, the limitation of RNN-based approach is that it is difficult for RNNs to capture long-range dependencies in videos. Inspired by the success of non-local neural networks, we propose a novel non-local temporal attention model in this paper. Our model can effectively capture long-range and global dependencies among the frames of the videos. Extensive experiments on three different benchmark datasets (i.e. iLIDS-VID, PRID-2011 and SDU-VID) show that our proposed method outperforms other state-of-the-art approaches.
Shivansh Rao, Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
AVSS4
2019 Video Summarization by Learning From Unpaired Data
abstract
We consider the problem of video summarization. Given an input raw video, the goal is to select a small subset of key frames from the input video to create a shorter summary video that best describes the content of the original video. Most of the current state-of-the-art video summarization approaches use supervised learning and require labeled training data. Each training instance consists of a raw input video and its ground truth summary video curated by human annotators. However, it is very expensive and difficult to create such labeled training examples. To address this limitation, we propose a novel formulation to learn video summarization from unpaired data. We present an approach that learns to generate optimal video summaries using a set of raw videos (V) and a set of summary videos (S), where there exists no correspondence between V and S. We argue that this type of data is much easier to collect. Our model aims to learn a mapping function F : V -> S such that the distribution of resultant summary videos from F(V) is similar to the distribution of S with the help of an adversarial objective. In addition, we enforce a diversity constraint on F(V) to ensure that the generated video summaries are visually diverse. Experimental results on two benchmark datasets indicate that our proposed approach significantly outperforms other alternative methods.
Mrigank Rochan, Yang Wang 0003
CVPR1
2019 Cross-Modal Self-Attention Network for Referring Image Segmentation
abstract
We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this area treat the language expression and the input image separately in their representations. They do not sufficiently capture long-range correlations between these two modalities. In this paper, we propose a cross-modal self-attention (CMSA) module that effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the input image. In addition, we propose a gated multi-level fusion module to selectively integrate self-attentive cross-modal features corresponding to different levels in the image. This module controls the information flow of features at different levels. We validate the proposed approach on four evaluation datasets. Our proposed approach consistently outperforms existing state-of-the-art methods.
Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Yang Wang 0003
CVPR2
2019 Convolutional Temporal Attention Model for Video-Based Person Re-Identification
abstract
The goal of video-based person re-identification is to match two input videos, so that the distance of the two videos is small if two videos contain the same person. A common approach for person re-identification is to first extract image features for all frames in the video, then aggregate all the features to form a video-level feature. The video-level features of two videos can then be used to calculate the distance of the two videos. In this paper, we propose a temporal attention approach for aggregating frame-level features into a video-level feature vector for re-identification. Our method is motivated by the fact that not all frames in a video are equally informative. We propose a fully convolutional temporal attention model for generating the attention scores. Fully convolutional network (FCN) has been widely used in semantic segmentation for generating 2D output maps. In this paper, we formulate video based person reidentification as a sequence labeling problem like semantic segmentation. We establish a connection between them and modify FCN to generate attention scores to represent the importance of each frame. Extensive experiments on three different benchmark datasets (i.e. iLIDS-VID, PRID-2011 and SDU-VID) show that our proposed method outperforms other state-of-the-art approaches.
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
ICME2
2018 Future Semantic Segmentation with Convolutional LSTM
Seyed Shahabeddin Nabavi, Mrigank Rochan, Yang Wang 0003
BMVC2
2018 Video Summarization Using Fully Convolutional Sequence Networks
Mrigank Rochan, Linwei Ye, Yang Wang 0003
ECCV (12)1
2017 Adapting Object Detectors from Images to Weakly Labeled Videos
Omit Chanda, Eu Wern Teh, Mrigank Rochan, Yang Wang 0003
BMVC3
2017 Salient Object Detection using a Context-Aware Refinement Network
Md. Amirul Islam, Mahmoud Kalash, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
BMVC3
2017 Person Re-Identification by Localizing Discriminative Regions
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
BMVC2
2017 Gated Feedback Refinement Network for Dense Image Labeling
abstract
Effective integration of local and global contextual information is crucial for dense labeling problems. Most existing methods based on an encoder-decoder architecture simply concatenate features from earlier layers to obtain higher-frequency details in the refinement stages. However, there are limits to the quality of refinement possible if ambiguous information is passed forward. In this paper we propose Gated Feedback Refinement Network (G-FRNet), an end-to-end deep learning framework for dense labeling tasks that addresses this limitation of existing methods. Initially, G-FRNet makes a coarse prediction and then it progressively refines the details by efficiently integrating local and global contextual information during the refinement stages. We introduce gate units that control the information passed forward in order to filter out ambiguity. Experiments on three challenging dense labeling datasets (CamVid, PASCAL VOC 2012, and Horse-Cow Parsing) show the effectiveness of our method. Our proposed approach achieves state-of-the-art results on the CamVid and Horse-Cow Parsing datasets, and produces competitive results on the PASCAL VOC 2012 dataset.
Md. Amirul Islam, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
CVPR2
2016 Attention Networks for Weakly Supervised Object Localization
Eu Wern Teh, Mrigank Rochan, Yang Wang 0003
BMVC2
2016 Weakly supervised object localization and segmentation in videos
Mrigank Rochan, Shafin Rahman, Neil D. B. Bruce, Yang Wang 0003
Image Vis. Comput.1
2015 Weakly supervised localization of novel objects using appearance transfer
abstract
We consider the problem of localizing unseen objects in weakly labeled image collections. Given a set of images annotated at the image level, our goal is to localize the object in each image. The novelty of our proposed work is that, in addition to building object appearance model from the weakly labeled data, we also make use of existing detectors of some other object classes (which we call “familiar objects”). We propose a method for transferring the appearance models of the familiar objects to the unseen object. Our experimental results on both image and video datasets demonstrate the effectiveness of our approach.
Mrigank Rochan, Yang Wang 0003
CVPR1
2014 Examining visual saliency prediction in naturalistic scenes
abstract
Given the significant number of potential applications, visual saliency has increasingly become an area of interest in image and vision research. Many different strategies for predicting visual saliency have been proposed, that differ in their composition or rationale, and with a significant focus on improving performance across standard benchmarks. Recent benchmarks considering a large number of algorithms have further provided an understanding of the behavior of different algorithms. Performance evaluation has primarily focused on indoor and outdoor images of urban environments, many of which are composed, and contain salient objects. In this work, we test the performance of a number of the better performing algorithms on data derived from naturalistic scenes. In addition, given the strong connection to human vision, we test a putative model for early visual processing in primates tied to spectral energy and normalization. Results demonstrate significant differences between common datasets, and natural images. Performance analysis of the second-order contrast model also provides additional insight concerning the role of spectral energy in determining saliency. Finally we include analysis that demonstrates statistical properties of images that tend to imply common gaze patterns across observers.
Shafin Rahman, Mrigank Rochan, Yang Wang 0003, Neil D. B. Bruce
ICIP2