Sayak Nag

dblp:205/2508 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 32% Trustworthy machine learning · 24% Segmentation and scene understanding · 20%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
scene graph generation
1.522025
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation · CVPR 2025
Unbiased Scene Graph Generation in Videos · CVPR 2023
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Computer vision › 3D vision
3d scene reconstruction
0.912025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.912025
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation · CVPR 2025
Machine learning › Generative modeling › diffusion model › diffusion model inference
diffusion-based refinement
0.912025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Computer vision › 3D vision
novel view synthesis
0.912025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.912025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation · CVPR 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
Reconstruction Guided Meta-Learning for Few Shot Open Set Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Trustworthy machine learning › open-world recognition › open-set recognition
few-shot open-set recognition
0.712023
Reconstruction Guided Meta-Learning for Few Shot Open Set Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › Segmentation and scene understanding › scene graph generation
unbiased scene graph generation
0.712023
Unbiased Scene Graph Generation in Videos · CVPR 2023
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
video scene graph generation
0.712023
Unbiased Scene Graph Generation in Videos · CVPR 2023
Data mining › predictive modeling › classification
imbalanced classification
0.412020
Boosting with Lexicographic Programming: Addressing Class Imbalance without Cost Tuning · IEEE Trans. Knowl. Data Eng. 2020
Machine learning › Generative modeling
diffusion model
0.312025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Machine learning › Generative modeling › diffusion model › video diffusion model
latent video diffusion model
0.312025
Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes · ICCV 2025
Machine learning › Trustworthy machine learning › open-world recognition
open-set recognition
0.212023
Reconstruction Guided Meta-Learning for Few Shot Open Set Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

semantic uncertainty quantification · 0.9post-processing · 0.9per-pixel entropy · 0.9multimodal large language model · 0.9latent video diffusion model · 0.9fourier-style transfer · 0.9transformer · 0.7memory-guided training · 0.7gaussian mixture model · 0.7exemplar reconstruction · 0.7lexicographic linear programming · 0.4dual formulation · 0.4
YearPublicationVenuePosition
2025 Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantification in SGG for its practical viability. In this paper, we introduce a novel Conformal Prediction based framework, adaptive to any existing SGG method, for quantifying their predictive uncertainty by constructing well-calibrated prediction sets over their generated scene graphs. These scene graph prediction sets are designed to achieve statistically rigorous coverage guarantees under exchangeability assumptions. Additionally, to ensure the prediction sets contain the most practically interpretable scene graphs, we propose an effective MLLM-based post-processing strategy for selecting the most visually and semantically plausible scene graphs within each set. We show that our proposed approach can produce diverse possible scene graphs from an image, assess the reliability of SGG methods, and improve overall SGG performance.
Sayak Nag, Udita Ghosh, Calvin-Khang Ta, Sarosij Bose, Amit K. Roy-Chowdhury
CVPR1
2025 Uncertainty-Aware Diffusion-Guided Refinement of 3D Scenes
abstract
Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D reconstruction methods render incoherent and blurry views. This problem is exacerbated when the unseen regions are far away from the input camera. In this work, we address these inherent limitations in existing single image-to-3D scene feedforward networks. To alleviate the poor performance due to insufficient information beyond the input image's view, we leverage a strong generative prior in the form of a pre-trained latent video diffusion model, for iterative refinement of a coarse scene represented by optimizable Gaussian parameters. To ensure that the style and texture of the generated images align with that of the input image, we incorporate on-the-fly Fourier-style transfer between the generated images and the input image. Additionally, we design a semantic uncertainty quantification module that calculates the per-pixel entropy and yields uncertainty maps used to guide the refinement process from the most confident pixels while discarding the remaining highly uncertain ones. We conduct extensive experiments on real-world scene datasets, including in-domain RealEstate-10K and out-of-domain KITTI-v2, showing that our approach can provide more realistic and high-fidelity novel view synthesis results compared to existing state-of-the-art methods.
Sarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang, Konstantinos Karydis, Amit K. Roy-Chowdhury
ICCV3
2025 ODES: Online Domain Adaptation with Expert Guidance for Medical Image Segmentation
Md Shazid Islam, Sayak Nag, Arindam Dutta, Sk Miraj Ahmed, Fahim Faisal Niloy, Shreyangshu Bera, Amit K. Roy-Chowdhury
MICCAI (4)2
2023 Unbiased Scene Graph Generation in Videos
abstract
The task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG. Existing methods for dynamic SGG have primarily focused on capturing spatio-temporal context using complex architectures without addressing the challenges mentioned above, especially the long-tailed distribution of relationships. This often leads to the generation of biased scene graphs. To address these challenges, we introduce a new framework called TEMPURA: TEmporal consistency and Memory Prototype guided UnceR-tainty Attenuation for unbiased dynamic SGG. TEMPURA employs object-level temporal consistencies via transformer-based sequence modeling, learns to synthesize unbiased relationship representations using memory-guided training, and attenuates the predictive uncertainty of visual relations using a Gaussian Mixture Model (GMM). Extensive experiments demonstrate that our method achieves significant (up to 10% in some cases) performance gain over existing methods highlighting its superiority in generating more unbiased scene graphs. Code: https://github.com/sayaknag/unbiasedSGG.git
Sayak Nag, Kyle Min 0001, Subarna Tripathi, Amit K. Roy-Chowdhury
CVPR1
2023 Semantics Guided Contrastive Learning of Transformers for Zero-shot Temporal Activity Detection
abstract
Zero-shot temporal activity detection (ZSTAD) is the problem of simultaneous temporal localization and classification of activity segments that are previously unseen during training. This is achieved by transferring the knowledge learned from semantically-related seen activities. This ability to reason about unseen concepts without supervision makes ZSTAD very promising for applications where the acquisition of annotated training videos is difficult. In this paper, we design a transformer-based framework titled TranZAD, which streamlines the detection of unseen activities by casting ZSTAD as a direct set-prediction problem, removing the need for hand-crafted designs and manual post-processing. We show how a semantic information-guided contrastive learning strategy can effectively train TranZAD for the zero-shot setting, enabling the efficient transfer of knowledge from the seen to the unseen activities. To reduce confusion between unseen activities and unrelated background information in videos, we introduce a more efficient method of computing the background class embedding by dynamically adapting it as part of the end-to-end learning. Additionally, unlike existing work on ZSTAD, we do not assume the knowledge of which classes are unseen during training and use the visual and semantic information of only the seen classes for the knowledge transfer. This makes TranZAD more viable for practical scenarios, which we evaluate by conducting extensive experiments on Thumos’14 and Charades.
Sayak Nag, Orpaz Goldstein, Amit K. Roy-Chowdhury
WACV1
2023 Reconstruction Guided Meta-Learning for Few Shot Open Set Recognition
abstract
In many applications, we are constrained to learn classifiers from very limited data (few-shot classification). The task becomes even more challenging if it is also required to identify samples from unknown categories (open-set classification). Learning a good abstraction for a class with very few samples is extremely difficult, especially under open-set settings. As a result, open-set recognition has received limited attention in the few-shot setting. However, it is a critical task in many applications like environmental monitoring, where the number of labeled examples for each class is limited. Existing few-shot open-set recognition (FSOSR) methods rely on thresholding schemes, with some considering uniform probability for open-class samples. However, this approach is often inaccurate, especially for fine-grained categorization, and makes them highly sensitive to the choice of a threshold. To address these concerns, we propose Reconstructing Exemplar-based Few-shot Open-set ClaSsifier (ReFOCS). By using a novel exemplar reconstruction-based meta-learning strategy ReFOCS streamlines FSOSR eliminating the need for a carefully tuned threshold by learning to be self-aware of the openness of a sample. The exemplars, act as class representatives and can be either provided in the training dataset or estimated in the feature domain. By testing on a wide variety of datasets, we show ReFOCS to outperform multiple state-of-the-art methods.
Sayak Nag, Dripta S. Raychaudhuri, Sujoy Paul, Amit K. Roy-Chowdhury
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Boosting with Lexicographic Programming: Addressing Class Imbalance without Cost Tuning
abstract
A large amount of research effort has been dedicated to adapting boosting for imbalanced classification. However, boosting methods are yet to be satisfactorily immune to class imbalance, especially for multi-class problems. This is because most of the existing solutions for handling class imbalance rely on expensive cost set tuning for determining the proper level of compensation. We show that the assignment of weights to the component classifiers of a boosted ensemble can be thought of as a game ofTug of Warbetween the classes in the margin space. We then demonstrate how this insight can be used to attain a good compromise between the rare and abundant classes without having to resort to cost set tuning, which has long been the norm for imbalanced classification. The solution is based on a lexicographic linear programming framework which requires two stages. Initially, class-specific component weight combinations are found so as to minimize a hinge loss individually for each of the classes. Subsequently, the final component weights are assigned so that the maximum deviation from the class-specific minimum loss values (obtained in the previous stage) is minimized. Hence, the proposal is not only restricted to two-class situations, but is also readily applicable to multi-class problems. Additionally, we also derive the dual formulation corresponding to the proposed framework. Experiments conducted on artificial and real-world imbalanced datasets as well as on challenging applications such as hyperspectral image classification and ImageNet classification establish the efficacy of the proposal.
Shounak Datta, Sayak Nag, Swagatam Das
IEEE Trans. Knowl. Data Eng.2