VLDB 2026 Research / reviewers in the wild / expert
Siddharth Roheda
dblp:226/5579
· DBLP profile ↗
9ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0002-6195-8517ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CART: Compositional AutoRegressive Transformer for Image GenerationabstractWe propose a novel Auto-Regressive (AR) image generation approach that models images as hierarchical compositions of interpretable visual layers. While AR models have achieved transformative success in language modeling, replicating this success in vision remains challenging due to inherent spatial dependencies in images. Addressing the unique challenges of vision tasks, our method (CART) adds image details iteratively via semantically meaningful decompositions. We demonstrate the flexibility and generality of CART by applying it across three distinct decomposition strategies: (i) Base-Detail Decomposition (Mumford-Shah smoothness), (ii) Intrinsic Decomposition (albedo/shading), and (iii) Specularity Decomposition (diffuse/specular). This “next-detail" strategy outperforms traditional “next-token" and “next-scale" approaches, improving controllability, semantic interpretability, and resolution scalability. Experiments show CART generates visually compelling results while enabling structured image manipulation, opening new directions for controllable generative modeling via physically or perceptually motivated image factorization. Siddharth Roheda, Rohit Chowdhury, Aniruddha Bala, Rohan Jaiswal |
AAAI | 1 |
| 2025 | DCT-Shield: A Robust Frequency Domain Defense Against Malicious Image EditingabstractAdvancements in diffusion models have enabled effortless image editing via text prompts, raising concerns about image security. Attackers with access to user images can exploit these tools for malicious edits. Recent defenses attempt to protect images by adding a limited noise in the pixel space to disrupt the functioning of diffusion-based editing models. However, the adversarial noise added by previous methods is easily noticeable to the human eye. Moreover, most of these methods are not robust to purification techniques like JPEG compression under a feasible pixel budget. We propose a novel optimization approach that introduces adversarial perturbations directly in the frequency domain by modifying the Discrete Cosine Transform (DCT) coefficients of the input image. By leveraging the JPEG pipeline, our method generates adversarial images that effectively prevent malicious image editing. Extensive experiments across a variety of tasks and datasets demonstrate that our approach introduces fewer visual artifacts while maintaining similar levels of edit protection and robustness to noise purification techniques. Aniruddha Bala, Rohit Chowdhury, Rohan Jaiswal, Siddharth Roheda |
ICCV | 4 |
| 2024 | MR-VNet: Media Restoration using Volterra NetworksabstractThis research paper presents a novel class of restoration network architecture based on the Volterra series formulation. By incorporating non-linearity into the system response function through higher order convolutions instead of traditional activation functions, we introduce a general framework for image/video restoration. Through extensive experimentation, we demonstrate that our proposed architecture achieves state-of-the-art (SOTA) performance in the field of Image/video Restoration. Moreover, we establish that the recently introduced Non-Linear Activation Free Network (NAF-NET) can be considered a special case within the broader class of Volterra Neural Networks. These findings highlight the potential of Volterra Neural Networks as a versatile and powerful tool for addressing complex restoration tasks in computer vision. Siddharth Roheda, Amit Satish Unde, Loay Rashid |
CVPR | 1 |
| 2024 | Volterra Neural Networks (VNNs)abstractThe importance of inference in Machine Learning (ML) has led to an explosive number of different proposals, particularly in Deep Learning. In an attempt to reduce the complexity of Convolutional Neural Networks, we propose a Volterra filter-inspired Network architecture. This architecture introduces controlled non-linearities in the form of interactions between the delayed input samples of data. We propose a cascaded implementation of Volterra Filtering so as to significantly reduce the number of parameters required to carry out the same classification task as that of a conventional Neural Network. We demonstrate an efficient parallel implementation of this Volterra Neural Network (VNN), along with its remarkable performance while retaining a relatively simpler and potentially more tractable structure. Furthermore, we show a rather sophisticated adaptation of this network to nonlinearly fuse the RGB (spatial) information and the Optical Flow (temporal) information of a video sequence for action recognition. The proposed approach is evaluated on UCF-101 and HMDB-51 datasets for action recognition, and is shown to outperform state of the art CNN approaches. Siddharth Roheda, Hamid Krim |
J. Mach. Learn. Res. | 1 |
| 2023 | Fast Optimal Transport for Latent Domain AdaptationabstractIn this paper, we address the problem of unsupervised Domain Adaptation. The need for such an adaptation arises when the distribution of the target data differs from that which is used to develop the model and the ground truth information of the target data is unknown. We propose an algorithm that uses optimal transport theory with a verifiably efficient and implementable solution to learn the best latent feature representation. This is achieved by minimizing the cost of transporting the samples from the target domain to the distribution of the source domain. Siddharth Roheda, Ashkan Panahi, Hamid Krim |
ICIP | 1 |
| 2021 | Event driven sensor fusion
Siddharth Roheda, Hamid Krim, Zhi-Quan Luo, Tianfu Wu 0001 |
Signal Process. | 1 |
| 2020 | Conquering the CNN Over-Parameterization Dilemma: A Volterra Filtering Approach for Action Recognition
Siddharth Roheda, Hamid Krim |
AAAI | 1 |
| 2020 | Commuting Conditional GANS for Multi-Modal FusionabstractThis paper presents a data driven approach to multi-modal fusion where a hidden latent sub-space between the different modalities is learned. The hidden space is estimated via a bank of Conditional GANs which also commute with each other, leading to an output that lies in a common subspace. Experimental results show improved detection performance compared to existing fusion techniques in ideal as well as noisy sensor condition. Siddharth Roheda, Hamid Krim, Benjamin S. Riggan |
ICASSP | 1 |
| 2018 | Cross-Modality Distillation: A Case for Conditional Generative Adversarial NetworksabstractIn this paper, we propose to use a Conditional Generative Adversarial Network (CGAN) for distilling (i.e. transferring) knowledge from sensor data and enhancing low-resolution target detection. In unconstrained surveillance settings, sensor measurements are often noisy, degraded, corrupted, and even missing/absent, thereby presenting a significant problem for multi-modal fusion. We therefore specifically tackle the problem of a missing modality in our attempt to propose an algorithm based on CGANs to generate representative information from the missing modalities when given some other available modalities. Despite modality gaps, we show that one can distill knowledge from one set of modalities to another. Moreover, we demonstrate that it achieves better performance than traditional approaches and recent teacher-student models. Siddharth Roheda, Benjamin S. Riggan, Hamid Krim, Liyi Dai |
ICASSP | 1 |