VLDB 2026 Research / reviewers in the wild / expert
Paras Maharjan
dblp:246/5951
· DBLP profile ↗
9ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0007-3950-706XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Amplitude, Phase, and Gradient Recovery From Compressed SAR ImagesabstractCompressing Synthetic Aperture Radar (SAR) images presents unique challenges due to the high dynamic range and inherent acquisition noise in the amplitude signal, as well as the noise-sensitive and limited information content in the phase signal. Traditional compression methods, such as JPEG and JPEG2000, although widely used, often fail to preserve SAR image quality due to their susceptibility to compression artifacts. The continuous capture of high-resolution raw SAR images over extended periods on drones and Unmanned Aerial Vehicles (UAVs), combined with constraints on computational resources, bandwidth, and onboard storage, further complicates the problem. An effective and efficient compression pipeline is essential for either onboard storage or real-time transmission to ground stations. In this work, we propose a hybrid solution for complex-valued SAR image compression by utilizing the Versatile Video Coding (VVC) framework as a backbone compression engine and employing a deep learning-based method that operates jointly in the pixel and transform domains for deblocking and reconstructing SAR amplitude and phase images. Specifically, we design task-specific compression artifact removal networks calledAmpResandAngResfor amplitude and phase reconstruction, respectively. Additionally, we introduce theGradResnetwork to learn gradients for SAR Scale-Invariant Feature Transform (SAR-SIFT), resulting in robust orientation and magnitude estimations that improve downstream tasks such as keypoints detection and matching in noisy and compressed scenarios. Experimental results demonstrate that our approach achieves a 10% Bjøntegaard Delta (BD)-Rate savings over VVC for amplitude recovery, along with notable improvement in phase reconstruction, and delivers an average of 34% improvement in SAR-SIFT repeatability. Paras Maharjan, Zhu Li 0001, Chris McGuiness |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Distributed Polarimetric SAR Compression with Joint Deblocking Using Side InformationabstractQuad-pol Synthetic Aperture Radar (SAR) images consist of four polarization schemes(HH, HV, VH, and VV), each having different responses from different types of terrain, foliage, and buildings. These polarimetric SAR images show correlations between cross-polarizations (HV and VH) and co-polarizations (HH and VV). Real-time transmission or storage of these large raw SAR images from UAVs in hostile electromagnetic environments becomes impractical due to file size, bandwidth constraints, and limited bit rate. To overcome this challenge, we propose a Distributed Polarimetric SAR Compression and Deblocking network (DPCD) that can exploit polarimetric correlations to improve compression efficiency. Our approach enables distributed and independent encoding of quad-pol SAR images while utilizing a learning-based method for joint frequency and spatial domain deblocking of the primary polarization, integrating feature-level information from Side Information (SI). Experimental results on the NGA SAR dataset show that DPCD outperforms state-of-the-art single-polarization compression techniques, particularly at lower bitrates, delivering significant bitrate savings without compromising image quality. Paras Maharjan, Sayush Maharjan, Zhu Li 0001, Neil Rogers, George York |
ISCAS | 1 |
| 2024 | E2SIFT: Neuromorphic SIFT via Direct Feature Pyramid Recovery from EventsabstractIn recent years, event cameras have achieved significant attention due to their advantages over conventional cameras. Event cameras have high dynamic range, no motion blur, and high temporal resolution. Contrary to traditional cameras which generate intensity frames, event cameras output a stream of asynchronous events based on brightness change. There is extensive ongoing research on performing computer vision tasks like object detection, classification, etc via the event camera. However, due to the unconventional output format of the event camera, it is difficult to perform computer vision tasks directly on the event stream. Mostly, works reconstruct the intensity image from the event stream and then perform such tasks. An important and crucial task is feature detection and description. Scale-invariant feature transform (SIFT) is a widely-used scale-invariant keypoint detector and descriptor that is invariant to transformations like scale, rotation, noise, and illumination. In this work, given an event voxel, we directly generate the LoG pyramid for SIFT keypoint detection. We fit a 3rd-degree polynomial and calculate the polynomial roots to compute the scale-space extrema response for SIFT keypoint detection. Since the extrema computation is performed after LoG thresholding, the solution is computationally less expensive. Experimental results validate the effectiveness of our system. Chris Henry, Paras Maharjan, Zhu Li 0001, George York |
ICIP | 2 |
| 2024 | End-to-End Compression of Complex-Valued SAR ImagesabstractCompression of Synthetic Aperture Radar (SAR) images presents significant challenges due to inherent noise in the data acquisition process. SAR images consist of two components: “Amplitude” and “Phase.” The amplitude represents the magnitude of the radar signal, similar to the intensity of an optical image but with noise. Traditional SAR compression methods often apply existing optical image compression tools like JPEG, JPEG2000, etc. directly to the amplitude component. However, compressing the phase component is more problematic due to low information content, high noise level, and sensitivity to the quantization errors introduced by compression algorithms. A more effective approach involves transforming the amplitude and phase into representations that are easier to compress. Complex-valued SAR images achieve this by expressing the amplitude and phase as In-phase (I) and Quadrature-phase (Q) components. In this context, we propose an end-to-end deep learning pipeline specifically designed to compress complex-valued SAR images, ensuring that both amplitude and phase information are efficiently preserved and compressed. In addition, we applied frequency domain processing using subsampled 4×4 2D-DCT on complex-valued SAR images. This technique reduces the spatial resolution without information loss, effectively reducing the inference time on the CPU by almost twofold while maintaining acceptable performance metrics. Paras Maharjan, Corey Marrs, Zhu Li 0001 |
MMSP | 1 |
| 2023 | Fast LoG SIFT Keypoint DetectorabstractScale-invariant feature transform (SIFT) is a classical computer vision technique for scale-invariant keypoint detection and feature extraction. SIFT exhibits invariance to various transformations such as scale, rotation, noise, and illumination, making it applicable in a wide range of applications like object recognition, image matching and stitching, environment mapping, navigation, robotics, camera calibration, and more. A key contribution of SIFT is its utilization of the Difference-of-Gaussian (DoG) feature pyramid, which approximates the scale-space response of the Laplacian-of-Gaussian (LoG) filter. The DoG feature pyramid is computed by taking the separable Gaussian filtering and stacking the difference of Gaussian blurred images. In this paper, we propose a novel approach called “Fast LoG” filtering, which offers direct computation of the LoG filter to model the scale-space response solution. The “Fast LoG” filter is achieved by decomposing the LoG filter into two separable filters via SVD, and the scale-space response is computed by a direct polynomial fitting and differentiation, which is analytically more accurate. The polynomial fitting and differentiation only happen after the LoG peak strength thresholding, therefore the overall complexity is low compared with the DoG-based SIFT. The experimental results show that the keypoint generated by the Fast LoG method matches the SIFT keypoints, and per-pixel filtering complexity is lower. Paras Maharjan, Lyle Vanfossan, Zhu Li 0001, Jialie Shen 0001 |
MMSP | 1 |
| 2023 | Complex-valued SAR Image Compression: A Novel Approach for Amplitude and Phase RecoveryabstractSynthetic Aperture Radar (SAR) is an active and coherent imaging system that utilizes radio waves to illuminate the Earth’s surface, generating complex-valued images of the ground. It is widely employed in environmental monitoring and surveillance due to its ability to function effectively in diverse weather and lighting conditions, including day and night. SAR achieves this through the emission of coherent radio frequency signals and the capture of slant range and azimuth information. These details are stored as a complex signal in the form of Inphase (I) and Quadrature (Q) components, known as complex-valued SAR. The complex values (I/Q) can be transformed into amplitude and phase information. SAR images have a substantial size and a wide dynamic range, which pose challenges for storage and transmission. Thus, the adoption of compression techniques becomes essential to manage the large size of SAR images. In contrast to existing methods that primarily compress amplitude information, our study presents a novel pipeline to compress complex-valued SAR images using a prediction-recovery framework that utilizes a conventional image compression algorithm, High-Efficiency Video Codec (HEVC), to encode the original image to a bitstream. Subsequently, the framework decodes a prediction of the I/Q channels and then recovers the phase and amplitude via SARRecoveryNet, a deep-learning-based network, to effectively remove compression artifacts and recover information that suits different applications. By directly compressing the I/Q values as a predictor, and recovery via a deep learning module, our method preserves both the amplitude and phase information of the SAR image as part of the loss function in the training. Experimental results demonstrate that our proposed approach significantly enhances complex-valued SAR image compression in terms of rate-distortion (RD) performance and visual perception. The code will be made available on GitHub after the paper is accepted. Paras Maharjan, Zhu Li 0001 |
VCIP | 1 |
| 2022 | DCT-Based Residual Network for NIR Image ColorizationabstractColorization of Near-Infrared (NIR) image is a challenging problem due to the lack of low-level clues available in the luminance channel of visible images. Recently, deep learning has witnessed remarkable progress in the NIR image colorization approaches. However, we have observed that most research focuses on designing deeper and wider architectures to improve the quality of RGB images at the expense of computational burden and speed. Few studies have adopted lightweight but effective modules to improve the efficiency of NIR image colorization without affecting its performance. In this paper, we propose the Discrete Cosine Transform (DCT)-based Residual Network (DCT-RCAN) for NIR image colorization. Specifically, the output of our network is four coefficients generated by the Two-Dimensional (2D) 4×4 DCT of RGB images, which reduces the training difficulty of our network by explicitly separating low-frequency and high-frequency details into four subgroups. We adopt the Residual in Residual (RIR) module as a basic module in our network, which can reduce the complexity of the model. Thus, our method can focus on more crucial underlying patterns in channel dimension in a lightweight manner. Extensive experiments validate that our DCT-RCAN is computationally efficient and demonstrate competitive results against state-of-the-art NIR image colorization methods. Hongcheng Jiang, Paras Maharjan, Zhu Li 0001, George York |
ICIP | 2 |
| 2021 | DCTResNet: Transform Domain Image Deblocking for Motion Blur ImagesabstractPixel recovery with deep learning has shown to be very effective for a variety of low-level vision tasks like image super-resolution, denoising, and deblurring. Most existing works operate in the spatial domain, and there are few works that exploit the transform domain for image restoration tasks. In this paper, we present a transform domain approach for image deblocking using a deep neural network called DCTResNet. Our application is compressed video motion deblur, where the input video frame has blocking artifacts that make the deblurring task very challenging. Specifically, we use a block-wise Discrete Cosine Transform (DCT) to decompose the image into its low and high-frequency sub-band images and exploit the strong sub-band specific features for more effective deblocking solutions. Since JPEG also uses DCT for image compression, using DCT sub-band images for image deblocking helps to learn the JPEG compression prior to effectively correct the blocking artifacts. Our experimental results show that both PSNR and SSIM for DCTResNet perform more favorably than other state-of-the-art (SOTA) methods, while significantly faster in inference time. Paras Maharjan, Ning Xu 0001, Yuyan Song, Zhu Li 0001 |
VCIP | 1 |
| 2019 | Improving Extreme Low-Light Image Denoising via Residual LearningabstractTaking a satisfactory picture in a low-light environment remains a challenging problem. Low-light imaging mainly suffers from noise due to the low signal-to-noise ratio. Many methods have been proposed for the task of image denoising, but they fail to work under extremely low-light conditions. Recently, deep learning based approaches have been presented that have higher objective quality than traditional methods, but they usually have high computational cost which makes them impractical to use in real-time applications or where the processing power is limited. In this paper, we propose a new residual learning based deep neural network for end-to-end extreme low-light image denoising that can not only significantly reduce the computational cost but also improve the quality over existing methods in both objective and subjective metrics. Specifically, in one setting we achieved 29x speedup with higher PSNR. Subjectively, our method provides better color reproduction and preserves more detailed texture information compared to state-of-the-art methods. Paras Maharjan, Li Li 0040, Zhu Li 0001, Ning Xu 0001, Chongyang Ma, Yue Li 0015 |
ICME | 1 |