VLDB 2026 Research / reviewers in the wild / expert
Pranjay Shyam
dblp:274/1191
· DBLP profile ↗
11ranked-venue papers
10as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIME-VAD: Text-Informed Magnitude Enhancement Feature Learning for Vehicle Accident Detection and AnticipationabstractVehicular accidents pose a substantial risk to drivers, underscoring the persistent and vital need for heightening safety measures. Early accident anticipation mechanisms are imperative for proactive measures, while detection accuracy is pivotal for prompt response and effective post-accident mitigation. Accurate and early anticipation of accidents for automated driving assistance systems in vehicles or CCTV in cities remains a complex task due to the intricate spatial-temporal interactions within traffic videos. This study presents text-informed magnitude enhancement in contrastive multiple-instance feature learning for vehicle accident detection and anticipation (TIME-VAD). Text is a better representative of concepts when compared to images in video, thus multi-modal learning is suitable. Also, the traditional assumption about feature magnitude of accidents and normal frames in magnitude based multiple-instance learning using weak supervision may not hold. This has led to the development of a novel weak-supervised learning strategy involving magnitude enhancement from textual concepts. For a better frame-level perception of accident risks in videos, dynamic temporal attentions are refined using the proposed dilated temporal conv-attention (DTCA) block. In-depth component-level analysis is performed to showcase the model’s efficacy while elucidating its operational mechanisms. Evaluation is conducted on three benchmark datasets, considering both earliness and accuracy-related metrics. Extensive experiments demonstrate that our TIME-VAD model outperforms the existing models. Compared to the previous top-performing supervised model achieving 84.7% accuracy, TIME-VAD achieves a 94.44% accuracy (measured by ROC-AUC) on the largest DoTA dataset. Notably, our model also excels in measuring how early it detects accidents compared to previous methods. The code will be released onhttps://github.com/sumitmishra209/TIME-VAD Sumit Mishra, Medhavi Mishra, Pranjay Shyam, Dong-Soo Har |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | PAIR : Perception Aided Image Restoration for Natural Driving ConditionsabstractWe present a two-stage mechanism for generic image restoration in natural driving conditions, where multiple non-linear degradations simultaneously impact perception for humans and driving assistance systems. Our approach overcomes the limitations of utilizing a single neural network that incurs excessive computational overhead and yields sub-optimal recovery. The proposed first stage comprises computationally inexpensive image processing operations applied at a patch level using a lightweight convolutional neural network (CNN) that determines their intensity of operation. This patch size is guided by the receptive field of the CNN, allowing for dynamic restoration of non-linear and non-homogeneous degradation profiles. The second stage leverages a lightweight end-to-end neural network functioning as an inpainting network. It identifies inadequately restored regions and leverages global semantic and structural information to fill the affected areas. This approach enhances the restoration process by considering the entire image and addresses the remainder of localized deficiencies. In addition, we integrate dense perception tasks such as semantic and depth estimation during the optimization cycle to ensure restored images that are perceptually pleasing and conducive for downstream perception tasks. Since datasets covering diverse degradation scenarios for high- and low-level perception tasks are lacking, we utilize a synthetic data augmentation technique to generate non-homogeneous non-linear degradation profiles. Experiments on images captured in adverse weather conditions demonstrate the efficacy of our approach, yielding higher perceptual quality in restored images and improved performance in downstream perception tasks under adverse driving conditions. Importantly, our method offers computational efficiency compared to end-to-end image restoration algorithms, making it suitable for real-time applications. Pranjay Shyam, Hyunjin Yoo |
WACV | 1 |
| 2024 | Lightweight Thermal Super-Resolution and Object Detection for Robust Perception in Adverse Weather ConditionsabstractIn this work, we examine the potential application of thermal cameras in improving perception capabilities in adverse weather conditions like snow, night-time driving, and haze, focusing on retaining the performance of Advanced Driver Assistance Systems (ADAS), thus enhancing its functionality and safety characteristics. While thermal sensors offer the advantage of robust information capture in adverse weather conditions, their integration is plagued with issues surrounding poor feature capture in normal conditions, low imaging resolution, and high sensor costs. We address the former by formulating the problem definition as information switching wherein thermal images are selected when visible images are degraded. Furthermore, we consider a single object detector for RGB and thermal images to ensure low latency. We propose utilizing a learnable projection function that translates the thermal image into RGB color space, thus providing minimal modifications to the underlying object detector. We address the issues of low imaging resolution and cost by proposing a novel procedure that combines super-resolution and object detection, enabling the utilization of low-resolution and low-cost uncooled thermal imaging sensors. To ensure the complete pipeline meets the actual deployment requirements of real-time inference on resource-constrained devices, we introduce a lightweight super-resolution algorithm, implementing optimizations within the network structure followed by global pruning. In addition, to improve the feature representations extracted by lightweight encoders, we propose a bidirectional feature pyramid network to enhance the feature representation. We demonstrate the efficacy of the proposed mechanism through extensive simulated evaluations on automotive datasets such as FLIR, KAIST, DENSE, and Freiburg Thermal. Pranjay Shyam, Hyunjin Yoo |
WACV | 1 |
| 2023 | Overcoming Degradation Imbalance for Consistent Image Dehazing
Pranjay Shyam, Hyunjin Yoo |
BMVC | 1 |
| 2022 | GIQE: Generic Image Quality Enhancement via Nth Order Iterative DegradationabstractVisual degradations caused by motion blur, raindrop, rain, snow, illumination, and fog deteriorate image quality and, subsequently, the performance of perception algorithms deployed in outdoor conditions. While degradation-specific image restoration techniques have been extensively studied, such algorithms are domain sensitive and fail in real scenarios where multiple degradations exist simultaneously. This makes a case for blind image restoration and reconstruction algorithms as practically relevant. However, the absence of a dataset diverse enough to encapsulate all variations hinders development for such an algorithm. In this paper, we utilize a synthetic degradation model that recursively applies sets of random degradations to generate naturalistic degradation images of varying complexity, which are used as input. Furthermore, as the degradation intensity can vary across an image, the spatially invariant convolutional filter cannot be applied for all degradations. Hence to enable spatial variance during image restoration and reconstruction, we design a transformer-based architecture to benefit from the long-range dependencies. In addition, to reduce the computational cost of transformers, we propose a multi-branch structure coupled with modifications such as a complimentary feature selection mechanism and the replacement of a feed-forward network with lightweight multiscale convolutions. Finally, to improve restoration and reconstruction, we integrate an auxiliary decoder branch to predict the degradation mask to ensure the underlying network can localize the degradation information. From empirical analysis on 10 datasets covering rain drop removal, deraining, dehazing, image enhancement, and deblurring, we demonstrate the efficacy of the proposed approach while obtaining SoTA performance. Pranjay Shyam, Kyung-Soo Kim 0001, Kuk-Jin Yoon |
CVPR | 1 |
| 2022 | Multi-Source Domain Alignment for Domain Invariant Segmentation in Unknown TargetsabstractSemantic segmentation provides scene understanding capability by performing pixel-wise classification of objects within an image. However, the sensitivity of such algorithms towards domain changes requires fine-tuning using an annotated dataset for each novel domain, which is expensive to construct and inefficient. We highlight that irrespective of the training dataset, structural properties of scenes remain the same hence domain sensitivity arises from training methodology. Thus, in this paper, we propose a domain alignment approach wherein multiple synthetic source domains are used to train an underlying segmentation network such that it performs consistently in unknown real target domains. Towards this end, we propose a pixel-wise supervised contrastive learning framework that enforces constraints in latent space resulting in features belonging to the same class being clustered closely and away from different classes. This approach allows for better capturing of global and local semantics while providing domain invariant properties. Our approach can be easily incorporated into prior semantic segmentation approaches without the significant computational overhead. We empirically demonstrate the efficacy of the proposed approach on GTAV → Cityscapes, GTAV+Synthia → Cityscapes, and GTAV+Synthia+Synscapes → Cityscapes scenarios and report state-of-the-art (SoTA) performance without requiring access to images from the target domain. Pranjay Shyam, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
IROS | 1 |
| 2022 | Infra Sim-to-Real: An efficient baseline and dataset for Infrastructure based Online Object Detection and Tracking using Domain AdaptationabstractIncreasing usage of traffic cameras provides an opportunity to utilize them for smart city applications. However, the efficacy of such systems is determined by their ability to detect and track objects of interest from diverse viewpoints accurately. This is challenging due to the diverse viewpoints, elevations, and distinct properties of camera sensors. Thus, to ensure robust performance, the training dataset should cover many variations, including viewpoints, illumination changes, and diverse weather conditions. However, constructing such a dataset is expensive in terms of data collection and annotation. This paper proposes an unsupervised domain adaptation approach wherein a synthetic dataset is generated using a simulator and subsequently used to ensure performance consistency of multi-object-tracking (MOT) algorithms across a diverse range of manually annotated natural scenes. Towards this end, we emphasize achieving domain invariant object detection by combining image stylization and class-balancing augmentation. Furthermore, we extend the robust detection algorithm to track detected objects across a large time scale using feature embeddings generated by the detector. Based on qualitative and quantitative results, we demonstrate the viability of such a system that is invariant to illumination, weather, viewpoint, and scene changes while providing a baseline for future research. Codebase and datasets would be made available at https://github.com/pranjay-dev/IS2R. Pranjay Shyam, Sumit Mishra, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
IV | 1 |
| 2021 | Towards Domain Invariant Single Image DehazingabstractPresence of haze in images obscures underlying information, which is undesirable in applications requiring accurate environment information. To recover such an image, a dehazing algorithm should localize and recover affected regions while ensuring consistency between recovered and its neighboring regions. However owing to fixed receptive field of convolutional kernels and non uniform haze distribution, assuring consistency between regions is difficult. In this paper, we utilize an encoder-decoder based network architecture to perform the task of dehazing and integrate an spatially aware channel attention mechanism to enhance features of interest beyond the receptive field of traditional conventional kernels. To ensure performance consistency across diverse range of haze densities, we utilize greedy localized data augmentation mechanism. Synthetic datasets are typically used to ensure a large amount of paired training samples, however the methodology to generate such samples introduces a gap between them and real images while accounting for only uniform haze distribution and overlooking more realistic scenario of non-uniform haze distribution resulting in inferior dehazing performance when evaluated on real datasets. Despite this, the abundance of paired samples within synthetic datasets cannot be ignored. Thus to ensure performance consistency across diverse datasets, we train the proposed network within an adversarial prior-guided framework that relies on a generated image along with its low and high frequency components to determine if properties of dehazed images matches those of ground truth. We preform extensive experiments to validate the dehazing and domain invariance performance of proposed framework across diverse domains and report state-of-the-art (SoTA) results. The source code with pretrained models will be available at https://github.com/PS06/DIDH. Pranjay Shyam, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
AAAI | 1 |
| 2021 | Lightweight HDR Camera ISP for Robust Perception in Dynamic Illumination Conditions via Fourier Adversarial Networks
Pranjay Shyam, Sandeep Singh Sengar, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
BMVC | 1 |
| 2021 | Adversarially-trained Hierarchical Feature Extractor for Vehicle Re-identificationabstractVehicle Re-identification (Re-ID) aims to retrieve all instances of query vehicle images present in an image pool. However viewpoint, illumination, and occlusion variations along with subtle differences between two unique images pose a significant challenge towards achieving an effective system. In this paper, we emphasize upon enhancing the performance of visual feature based ReID system by improving feature embedding quality and propose (1) an attention-guided hierarchical feature extractor (HFE) that leverages the structure of a backbone CNN to extract coarse and fine-grained features and (2) to train the proposed network within a hard negative adversarial framework that generates samples exhibiting extreme variations, encouraging the network to extract important distinguishing features across varying scales. To demonstrate the effectiveness of the proposed framework we use VERI-Wild, VRIC and Veri-776 datasets that exhibit extreme intra-class and minute inter-class differences and achieve state-of-the-art (SoTA) performance. Codes related to this paper are publicly available at https://github.com/PS06/VReID. Pranjay Shyam, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
ICRA | 1 |
| 2020 | Dynamic Anchor Selection for Improving Object LocalizationabstractAnchor boxes act as potential object localization candidates allow single-stage detectors to achieve real-time performance, at the cost of localization accuracy when compared to state-of-the-art two-stage detectors. Therefore, correct selection of the scale and aspect ratio associated with an anchor box is crucial for detector performance. In this work, we propose a novel architecture called DANet for improving the localization performance of single-stage object detectors, while maintaining real-time inference. The proposed network achieves this by predicting (1) the combination of aspect ratio and scale per feature map based on object density and (2) localization confidence per anchor box. We evaluate the proposed network using the benchmark dataset. On the MS COCO dataset, DANet achieves 30.9% AP at 51.8 fps using ResNet-18 and 45.3% AP at 7.4 fps using ResNeXt-101. The code and models will be available at https://github.com/PS06/AnchorNet. Pranjay Shyam, Kuk-Jin Yoon, Kyung-Soo Kim 0001 |
ICRA | 1 |