Nancy Mehta

dblp:305/3859 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-1249-8577ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Complexity Experts are Task-Discriminative Learners for Any Image Restoration
abstract
Recent advancements in all-in-one image restoration models have revolutionized the ability to address diverse degradations through a unified framework. However, parameters tied to specific tasks often remain inactive for other tasks, making mixture-of-experts (MoE) architectures a natural extension. Despite this, MoEs often show inconsistent behavior, with some experts unexpectedly generalizing across tasks while others struggle within their intended scope. This hinders leveraging MoEs’ computational benefits by bypassing irrelevant experts during inference. We attribute this undesired behavior to the uniform and rigid architecture of traditional MoEs. To address this, we introduce “complexity experts” – flexible expert blocks with varying computational complexity and receptive fields. A key challenge is assigning tasks to each expert, as degradation complexity is unknown in advance. Thus, we execute tasks with a simple bias toward lower complexity. To our surprise, this preference effectively drives task-specific allocation, assigning tasks to experts with the appropriate complexity. Extensive experiments validate our approach, demonstrating the ability to bypass irrelevant experts during inference while maintaining superior performance. The proposed MoCE-IR model outperforms state-of-the-art methods, affirming its efficiency and practical applicability. The source code and models are publicly available at eduardzamfir.github.io/MoCE-IR/
Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan, Danda Pani Paudel, Yulun Zhang 0001, Radu Timofte
CVPR3
2025 Color Matching Using Hypernetwork-Based Kolmogorov-Arnold Networks
Artem V. Nikonorov, Georgy Perevozchikov, Andrei Korepanov, Nancy Mehta, Mahmoud Afifi, Egor Ershov, Radu Timofte
ICCV4
2025 LeMoRe: Learn More Details for Lightweight Semantic Segmentation
abstract
Lightweight semantic segmentation is essential for many downstream vision tasks. Unfortunately, existing methods often struggle to balance efficiency and performance due to the complexity of feature modeling. Many of these existing approaches are constrained by rigid architectures and implicit representation learning, often characterized by parameter-heavy designs and a reliance on computationally intensive Vision Transformer-based frameworks. In this work, we introduce an efficient paradigm by synergizing explicit and implicit modeling to balance computational efficiency with representational fidelity. Our method combines well-defined Cartesian directions with explicitly modeled views and implicitly inferred intermediate representations, efficiently capturing global dependencies through a nested attention mechanism. Extensive experiments on challenging datasets, including ADE20K, CityScapes, Pascal Context, and COCO-Stuff, demonstrate that LeMoRe strikes an effective balance between performance and efficiency. https://github.com/miannaeem-lab/LeMoRe
Mian Muhammad Naeem Abid, Nancy Mehta, Zongwei Wu, Radu Timofte
ICIP2
2025 USWformer: Efficient Sparse Wavelet Transformer for Underwater Image Enhancement
abstract
Transformer-based methods have shown great promise in underwater image enhancement (UIE) tasks due to their capability to model long-range dependencies, which are vital for reconstructing clear images. While numerous effective attention mechanisms have been devised to handle the computational requirements of transformers, they frequently incorporate redundant information and noisy interactions from irrelevant regions. Additionally, the current methods focusing solely on the raw pixel space constrains the exploration of the underwater image frequency dynamics, thus hindering the models from fully leveraging their potential for producing high-quality images. To address these challenges, we propose USWformer, an efficient UIE Sparse Wavelet Transformer Network (1.19 M parameters) to eliminate the redundant features in both the spatial and frequency domains. The USWformer consists of two fundamental components: a Sparse Wavelet Self-Attention (SWSA) block and a Multi-scale Wavelet Feed-Forward Network (MWFN). The SWSA block selectively preserves essential attention scores from the keys corresponding to each query, adjusting the feature details. MWFN further diminishes the feature redundancy in the aggregated features thereby improving the enhancement of the underwater images. We assess the efficacy of our approach across benchmark datasets comprising synthetic and real-world under-water images, showcasing its superiority via thorough ablation studies and comparative analyses.
Nancy Mehta, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
WACV2
2024 Rawformer: Unpaired Raw-to-Raw Translation for Learnable Camera ISPs
Georgy Perevozchikov, Nancy Mehta, Mahmoud Afifi, Radu Timofte
ECCV (36)2
2024 See More Details: Efficient Image Super-Resolution by Experts Mining
abstract
Reconstructing high-resolution (HR) images from low-resolution (LR) inputs poses a significant challenge in image super-resolution (SR). While recent approaches have demonstrated the efficacy of intricate operations customized for various objectives, the straightforward stacking of these disparate operations can result in a substantial computational burden, hampering their practical utility. In response, we introduce SeemoRe, an efficient SR model employing expert mining. Our approach strategically incorporates experts at different levels, adopting a collaborative methodology. At the macro scale, our experts address rank-wise and spatial-wise informative features, providing a holistic understanding. Subsequently, the model delves into the subtleties of rank choice by leveraging a mixture of low-rank experts. By tapping into experts specialized in distinct key factors crucial for accurate SR, our model excels in uncovering intricate intra-feature details. This collaborative approach is reminiscent of the concept of “see more", allowing our model to achieve an optimal performance with minimal computational costs in efficient settings.
Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yulun Zhang 0001, Radu Timofte
ICML3
2024 Probing Attention-Driven Normalizing Flow Network for Low-Light Image Enhancement
Nancy Mehta, K. N. Prakash, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
ICPR (32)2
2024 Spectroformer: Multi-Domain Query Cascaded Transformer Network For Underwater Image Enhancement
abstract
Underwater images often suffer from color distortion, haze, and limited visibility due to light refraction and absorption in water. These challenges significantly impact autonomous underwater vehicle applications, necessitating efficient image enhancement techniques. To address these challenges, we propose a Multi-Domain Query Cascaded Transformer Network for underwater image enhancement. Our approach includes a novel Multi-Domain Query Cascaded Attention mechanism that integrates localized transmission features and global illumination features. To improve feature propagation from the encoder to the decoder, we propose a Spatio-Spectro Fusion-Based Attention Block. Additionally, we introduce a Hybrid Fourier-Spatial Up-sampling Block, which uniquely combines Fourier and spatial upsampling techniques to enhance feature resolution effectively. We evaluate our method on benchmark synthetic and real-world underwater image datasets, demonstrating its superiority through extensive ablation studies and comparative analysis. The testing code is available at: https://github.com/Mdraqibkhan/Spectroformer.
Md Raqib Khan, Nancy Mehta, Shruti S. Phutke, Santosh Kumar Vipparthi, Sukumar Nandi, M. Subrahmanyam 0001
WACV3
2023 Gated Multi-Resolution Transfer Network for Burst Restoration and Enhancement
abstract
Burst image processing is becoming increasingly popular in recent years. However, it is a challenging task since individual burst images undergo multiple degradations and often have mutual misalignments resulting in ghosting and zipper artifacts. Existing burst restoration methods usually do not consider the mutual correlation and non-local contextual information among burst frames, which tends to limit these approaches in challenging cases. Another key challenge lies in the robust up-sampling of burst frames. The existing up-sampling methods cannot effectively utilize the advantages of single-stage and progressive up-sampling strategies with conventional and/or recent up-samplers at the same time. To address these challenges, we propose a novel Gated Multi-Resolution Transfer Network (GMTNet) to reconstruct a spatially precise high-quality image from a burst of low-quality raw images. GMT-Net consists of three modules optimized for burst processing tasks: Multi-scale Burst Feature Alignment (MBFA) for feature denoising and alignment, Transposed-Attention Feature Merging (TAFM) for multi-frame feature aggregation, and Resolution Transfer Feature Up-sampler (RTFU) to up-scale merged features and construct a high-quality output image. Detailed experimental analysis on five datasets validate our approach and sets a state-of-the-art for burst super-resolution, burst denoising, and low-light burst enhancement. Our codes and models are available at https://github.com/nanmehta/GMTNet.
Nancy Mehta, Akshay Dudhane, M. Subrahmanyam 0001, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan
CVPR1
2021 MSAR-Net: Multi-scale attention based light-weight image super-resolution
Nancy Mehta, M. Subrahmanyam 0001
Pattern Recognit. Lett.1