Majid Rabbani

dblp:25/6015 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
10since 2021 · last 2025
0009-0008-6289-0500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
abstract
Prevalent text-to-video retrieval systems mainly adopt embedding models for feature extraction and compute cosine similarities for ranking.However, this design presents two limitations.Low-quality text-video data pairs could compromise the retrieval, yet are hard to identify and examine.Cosine similarity alone provides no explanation for the ranking results, limiting the interpretability.We ask that can we interpret the ranking results, so as to assess the retrieval models and examine the text-video data?This work proposes X-CoT, an explainable retrieval framework upon LLM CoT reasoning in place of the embedding model-based similarity ranking.We first expand the existing benchmarks with additional video annotations to support semantic understanding and reduce data bias.We also devise a retrieval CoT consisting of pairwise comparison steps, yielding detailed reasoning and complete ranking.X-CoT empirically improves the retrieval performance and produces detailed rationales.It also facilitates the model behavior and data quality analysis.Code and data are available at: github.com
Prasanna Reddy Pulakurthi, Jiamian Wang, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Zhiqiang Tao
EMNLP3
2025 Structured Policy Optimization: Enhance Large Vision-Language Model via Self-Referenced Dialogue
Can Qin, Yihao Feng, Zeyuan Chen 0001, Ran Xu 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
ICCV7
2025 Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation
abstract
This work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (SPM), and a novel reweighting strategy are introduced to enhance performance. SPM shuffles and blends image patches to generate diverse and challenging augmentations, while the reweighting strategy prioritizes reliable pseudo-labels to mitigate label noise. These techniques are particularly effective on smaller datasets like PACS, where overfitting and pseudo-label noise pose greater risks. State-of-the-art results are achieved on three major benchmarks: PACS, VisDA-C, and DomainNet-126. Notably, on PACS, improvements of 7.3% (79.4% to 86.7%) and 7.2% are observed in single-target and multi-target settings, respectively, while gains of 2.8% and 0.7% are attained on DomainNet-126 and VisDA-C. This combination of advanced augmentation and robust pseudo-label reweighting establishes a new benchmark for SFDA. The code is available at: https://github.com/PrasannaPulakurthi/SPM.
Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard, Sohail A. Dianat, Celso de Melo, Raghuveer M. Rao
ICIP2
2025 Re-Imagining Multimodal Instruction Tuning: A Representation View
abstract
Multimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior.
Yiyang Liu 0003, James Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001
ICLR5
2025 Latent Chain-of-Thought for Visual Reasoning
abstract
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such as SFT, PPO, and GRPO may not generalize well across unseen reasoning tasks and heavily rely on a biased reward model. To address this challenge, we reformulate reasoning in LVLMs as posterior inference and propose a scalable training algorithm based on amortized variational inference. By leveraging diversity-seeking reinforcement learning algorithms, we introduce a novel sparse reward function for token-level learning signals that encourage diverse, high-likelihood latent CoT, overcoming deterministic sampling limitations and avoiding reward hacking. Additionally, we implement a Bayesian inference-scaling strategy that replaces costly Best-of-N and Beam Search with a marginal likelihood to efficiently rank optimal rationales and answers. We empirically demonstrate that the proposed method enhances the state-of-the-art LVLMs on four reasoning benchmarks, in terms of effectiveness, generalization, and interpretability.
Hang Hua, Jiebo Luo 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS6
2024 Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
abstract
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and re-silient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the de-termination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% ~ 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five bench-mark datasets, including MSRVTT, LSMDC, DiDeMo, VA-TEX, and Charades. Code and models are available here.
Jiamian Wang, Pichao Wang, Dongfang Liu, Sohail A. Dianat, Raghuveer M. Rao, Majid Rabbani, Zhiqiang Tao
CVPR7
2024 AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu
ECCV (65)4
2024 Enhancing GAN Performance Through Neural Architecture Search and Tensor Decomposition
abstract
Generative Adversarial Networks (GANs) have emerged as a powerful tool for generating high-fidelity content. This paper presents a new training procedure that leverages Neural Architecture Search (NAS) to discover the optimal architecture for image generation while employing the Maximum Mean Discrepancy (MMD) repulsive loss for adversarial training. Moreover, the generator network is compressed using tensor decomposition to reduce its computational footprint and inference time while preserving its generative performance. Experimental results show improvements of 34% and 28% in the FID score on the CIFAR-10 and STL-10 datasets, respectively, with corresponding footprint reductions of 14× and 31× compared to the best FID score method reported in the literature. The implementation code is available at: https://github.com/PrasannaPulakurthi/MMD-AdversarialNAS.
Prasanna Reddy Pulakurthi, Mahsa Mozaffari, Sohail A. Dianat, Majid Rabbani, Jamison Heard, Raghuveer M. Rao
ICASSP4
2024 Image Translation as Diffusion Visual Programmers
abstract
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model within the GPT architecture, orchestrating a coherent sequence of visual programs ($i.e.$, computer vision models) for various pro-symbolic steps, which span RoI identification, style transfer, and position manipulation, facilitating transparent and controllable image translation processes. Extensive experiments demonstrate DVP’s remarkable performance, surpassing concurrent arts. This success can be attributed to several key features of DVP: First, DVP achieves condition-flexible translation via instance normalization, enabling the model to eliminate sensitivity caused by the manual guidance and optimally focus on textual descriptions for high-quality content generation. Second, the frame work enhances in-context reasoning by deciphering intricate high-dimensional concepts in feature spaces into more accessible low-dimensional symbols ($e.g.$, [Prompt], [RoI object]), allowing for localized, context-free editing while maintaining overall coherence. Last but not least, DVP improves systemic controllability and explainability by offering explicit symbolic representations at each programming stage, empowering users to intuitively interpret and modify results. Our research marks a substantial step towards harmonizing artificial image translation processes with cognitive intelligence, promising broader applications.
Cheng Han 0001, James Liang, Qifan Wang 0001, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Ying Nian Wu, Dongfang Liu
ICLR4
2024 Diffusion-Inspired Truncated Sampler for Text-Video Retrieval
abstract
Prevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval
Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao
NeurIPS6
2013 Learning to Produce 3D Media From a Captured 2D Video
abstract
Due to the advances in display technologies and commercial success of 3D motion pictures in recent years, there is renewed interest in enabling consumers to create 3D content. While new 3D content can be created using more advanced capture devices (i.e., stereo cameras), most people still own 2D capture devices. Further, enormously large collections of captured media exist only in 2D. We present a system for producing pseudo-stereo images from captured 2D videos. Our system employs a two-phase procedure where the first phase detects “good” pseudo-stereo images frames from a 2D video, which was captured a priori without any constraints on camera motion or content. We use a trained classifier to detect pairs of video frames that are suitable for constructing pseudo-stereo images. In particular, for a given frame at time t, we determine if exists such that It+t̅and Itcan form an acceptable pseudo-stereo image. Moreover, even if t̂ is determined, generating a good pseudo-stereo image from 2D captured video frames can be nontrivial since in many videos, professional or amateur, both foreground and background objects may undergo complex motion. Independent foreground motions from different scene objects define different epipolar geometries that cause the conventional method of generating pseudo-stereo images to fail. To address this problem, the second phase of the proposed system further recomposes the frame pairs to ensure consistent 3D perception for objects for such cases. In this phase, final left and right pseudo-stereo images are created by recompositing different regions of the initial frame pairs to ensure a consistent camera geometry. We verify the performance of our method for producing pseudo-stereo media from captured 2D videos in a psychovisual evaluation using both professional movie clips and amateur home videos.
Jiebo Luo 0001, Andrew C. Gallagher, Majid Rabbani
IEEE Trans. Multim.4
2011 Learning to produce 3D media from a captured 2D video
abstract
Due to the advances in display technologies and the commercial success of 3D motion pictures in recent years, there is renewed interest in enabling consumers to create 3D content. While new 3D content can be created using more advanced capture devices (i.e., stereo cameras), most people still use 2D capture devices. Furthermore, enormously large collections of captured media exist only in 2D. We present a system for producing stereo images from captured 2D videos. Our system detects "good" stereo frames from a 2D video, which was captured a priori without any constraints on camera motion or content. We use a trained classifier to detect pairs of video frames that are suitable for constructing stereo images. In particular, for a given frame It at time t, we determine if t̂ exists such that It+t̂ and It can form an acceptable stereo image. We verify the performance of our method for producing stereo media from captured 2D videos in a psychovisual evaluation using both professional movie clips and amateur home videos. To the best of our knowledge, detecting good stereo pairs from a captured 2D video has been adequately addressed in the literature.
Jiebo Luo 0001, Andrew C. Gallagher, Majid Rabbani
ACM Multimedia4
2002 An overview of the JPEG 2000 still image compression standard
Majid Rabbani, Rajan L. Joshi
Signal Process. Image Commun.1
2000 The continuing evolution of digital cameras and digital photography systems
abstract
Electronic photography became popular when it focused on getting images into PCs, rather than onto TV screens. As desktop computers became image enabled, the sales of digital cameras to both professionals and consumers have grown rapidly. In the past, the primary emphasis was on increasing pixel count. Now that cameras costing less than US $1,000 provide more than 2 million pixels, users are learning that high pixel count is a necessary, but not sufficient condition for photographic quality images. As a result, the focus is shifting to designing optics, sensors, and digital processing to produce cameras having higher ISO speed, lower noise, wider dynamic range, improved tone and color reproduction, and fewer artifacts. Digital image processing can provide product differentiation by both enhancing image quality and providing new features. As pixel count increases, image compression becomes more important. While most digital photography systems today use the current DCT-based JPEG compression, the new wavelet-based JPEG2000 compression standard offers many valuable features for future systems. To support a range of digital photography systems, JPEG2000 will support a range of color interchange spaces and a flexible metadata mechanism.
Kenneth A. Parulski, Majid Rabbani
ISCAS2
1999 Biased reconstruction for JPEG decoding
abstract
Assuming a Laplacian distribution, there exists a well known method for optimally biasing the reconstruction levels for the quantized ac discrete cosine transform (DCT) coefficients in the JPEG decoder. This, however, requires an estimate of the Laplacian distribution parameter. We derive a new, maximum likelihood estimate of the Laplacian parameter using only the quantized coefficients available at the decoder. We quantify the benefits of biased reconstruction through extensive simulations and demonstrate that such improvements are very close to the best possible resulting from centroid reconstruction.
Jeff Price 0001, Majid Rabbani
IEEE Signal Process. Lett.2