VLDB 2026 Research / reviewers in the wild / expert
Jinglei Shi
dblp:243/6582
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-2926-0415ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion GuidanceabstractClean images are crucial for visual tasks such as small object detection, especially at high resolutions. However, real-world images are often degraded by adverse weather, and weather restoration methods may sacrifice high-frequency details critical for analyzing small objects. A natural solution is to apply super-resolution (SR) after weather removal to recover both clarity and fine structures. However, simply cascading restoration and SR struggle to bridge their inherent conflict: removal aims to remove high-frequency weather-induced noise, while SR aims to hallucinate high-frequency textures from existing details, leading to inconsistent restoration contents. In this paper, we take deraining as a case study and propose DHGM, a Diffusion-based High-frequency Guided Model for generating clean and high-resolution images. DHGM integrates pre-trained diffusion priors with high-pass filters to simultaneously remove rain artifacts and enhance structural details. Extensive experiments demonstrate that DHGM achieves superior performance over existing methods, with lower costs. Jinglei Shi, Heng Guo 0003, Zhanyu Ma |
AAAI | 2 |
| 2026 | Robust 2.5D Feature Matching in Light Fields via a Learnable Parameterized Depth-Degraded ProjectionabstractDue to the loss of 3D information, accurate and robust 2D image feature matching remains challenging for many computer vision applications. This paper introduces a 2.5D feature that uses the disparity value from the light field Fourier disparity layer (FDL) as a rough proxy of scene depth. Without explicit depth estimation, a parameterized depth-degraded projection is proposed to construct the geometric transformation of paired features between two light fields. Then, we propose a parameterized learning solution to calculate the depth-degraded projection. This solution estimates a global constant fundamental matrix, a variable disparity-guided translation vector, and a depth compensation term using a very simple network. Although the 0.5D relative disparity provided by the FDL does not represent precise depth, it can also significantly reduce the depth ambiguity in feature matching. Therefore, the proposed solution achieves accurate feature-matching results by minimizing the sum of reprojection errors across all matching candidates. On the public light field feature-matching dataset, the proposed solution outperforms existing 2D image feature-matching solutions and light field feature-matching algorithms in terms of matching accuracy and robustness. The code is available online. Haiyan Jin, Zhaolin Xiao, Jinglei Shi, Xiaoran Jiang |
IEEE Trans. Image Process. | 4 |
| 2025 | A Polarization-Aided Transformer for Image Deblurring via Motion Vector DecompositionabstractEffectively leveraging motion information is crucial for the image deblurring task. Existing methods typically build deep-learning models to restore a clean image by estimating blur patterns over the entire movement. This suggests that the blur caused by rotational motion components is processed together with the translational one. Exploring the movement without separation leads to limited performance for complex motion deblurring, especially rotational motion. In this paper, we propose Motion Decomposition Transformer (MDT), a transformer-based architecture augmented with polarized modules for deblurring via motion vector decomposition. MDT consists of a Motion Decomposition Module (MDM) for extracting hybrid rotation and translation features and a Radial Stripe Attention Solver (RSAS) for sharp image reconstruction with enhanced rotational information. Specifically, the MDM uses a deformable Cartesian convolutional branch to capture translational motion, complemented by a polar-system branch to capture rotational motion. The RSAS employs radial stripe windows and angular relative positional encoding in the polar system to enhance rotational information. This design preserves translational details while keeping computational costs lower than dual-coordinate design. Experimental results on 6 image deblurring datasets show that MDT outperforms state-of-the-art methods, particularly in handling blur caused by complex motions with significant rotational components. The code and pre-trained models are available at https://github.com/Calvin11311/MDT. Duosheng Chen, Shihao Zhou 0003, Jinshan Pan, Jinglei Shi, Lishen Qu, Jufeng Yang |
CVPR | 4 |
| 2025 | No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image RecognitionabstractOver the last decade, many notable methods have emerged to tackle the computational resource challenge of the high resolution image recognition (HRIR). They typically focus on identifying and aggregating a few salient regions for classification, discarding sub-salient areas for low training consumption. Nevertheless, many HRIR tasks necessitate the exploration of wider regions to model objects and contexts, which limits their performance in such scenarios. To address this issue, we present a DBPS strategy to enable training with more patches at low consumption. Specifically, in addition to a fundamental buffer that stores the embeddings of most salient patches, DBPS further employs an auxiliary buffer to recycle those sub-salient ones. To reduce the computational cost associated with gradients of sub-salient patches, these patches are primarily used in the forward pass to provide sufficient information for classification. Meanwhile, only the gradients of the salient patches are back-propagated to update the entire network. Moreover, we design a Multiple Instance Learning (MIL) architecture that leverages aggregated information from salient patches to filter out uninformative background within sub-salient patches for better accuracy. Besides, we introduce the random patch drop to accelerate training process and uncover informative regions. Experiment results demonstrate the superiority of our method in terms of both accuracy and training consumption against other advanced methods. The code is available in the https://github.com/Qinrong-Nku/DBPS. Rong Qin 0001, Xin Liu 0091, Jinglei Shi, Jufeng Yang |
CVPR | 5 |
| 2025 | Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty EstimationabstractOver the last decade, significant efforts have been dedicated to designing efficient models for the challenge of ultra-high resolution (UHR) semantic segmentation. These models mainly follow the dual-stream architecture and generally fall into three subcategories according to the improvement objectives, i.e., dual-stream ensemble, selective zoom, and complementary learning. However, most of them overly concentrate on crafting complex pipelines to pursue one of the above objectives separately, limiting the model performance in both accuracy and inference consumption. In this paper, we suggest simultaneously achieving these objectives by estimating resolution-biased uncertainties in low resolution stream. Here, the resolution-biased uncertainty refers to the degree of prediction unreliability primarily caused by resolution loss from down-sampling operations. Specifically, we propose a dual-stream UHR segmentation framework, where an estimator is used to assess resolution-biased uncertainties through the entropy map and high-frequency feature residual. The framework also includes a selector, an ensembler, and a complementer to boost the model with obtained estimations. They share the uncertainty estimations as the weights to choose difficult regions as the inputs for UHR stream, perform weighted fusion between distinct streams, and enhance the learning for important pixels, respectively. Experiment results demonstrate that our method achieves a satisfactory balance between accuracy and inference consumption against other state-of-the-art (SOTA) methods. The code is available in the https://github.com/Qinrong-NKU/RUE. Rong Qin 0001, Jinglei Shi, Jufeng Yang |
CVPR | 3 |
| 2025 | Devil is in the Uniformity: Exploring Diverse Learners Within Transformer for Image RestorationabstractTransformer-based approaches have gained significant attention in image restoration, where the core component, i.e, Multi-Head Attention (MHA), plays a crucial role in capturing diverse features and recovering high-quality results. In MHA, heads perform attention calculation independently from uniform split subspaces, and a redundancy issue is triggered to hinder the model from achieving satisfactory outputs. In this paper, we propose to improve MHA by exploring diverse learners and introducing various interactions between heads, which results in a Hierarchical multI-head atteNtion driven Transformer model, termed HINT, for image restoration. HINT contains two modules, i.e., the Hierarchical Multi-Head Attention (HMHA) and the Query-Key Cache Updating (QKCU) module, to address the redundancy problem that is rooted in vanilla MHA. Specifically, HMHA extracts diverse contextual features by employing heads to learn from subspaces of varying sizes and containing different information. Moreover, QKCU, comprising intra- and inter-layer schemes, further reduces the redundancy problem by facilitating enhanced interactions between attention heads within and across layers. Extensive experiments are conducted on 12 benchmarks across 5 image restoration tasks, including low-light enhancement, dehazing, desnowing, denoising, and deraining, to demonstrate the superiority of HINT. The source code is available in the supplementary materials. Shihao Zhou 0003, Dayu Li, Jinshan Pan, Juncheng Zhou, Jinglei Shi, Jufeng Yang |
ICCV | 5 |
| 2025 | FlareX: A Physics-Informed Dataset for Lens Flare Removal via 2D Synthesis and 3D RenderingabstractLens flare occurs when shooting towards strong light sources, significantly degrading the visual quality of images. Due to the difficulty in capturing flare-corrupted and flare-free image pairs in the real world, existing datasets are typically synthesized in 2D by overlaying artificial flare templates onto background images. However, the lack of flare diversity in templates and the neglect of physical principles in the synthesis process hinder models trained on these datasets from generalizing well to real-world scenarios. To address these challenges, we propose a new physics-informed method for flare data generation, which consists of three stages: parameterized template creation, the laws of illumination-aware 2D synthesis, and physical engine-based 3D rendering, which finally gives us a mixed flare dataset that incorporates both 2D and 3D perspectives, namely FlareX. This dataset offers 9,500 2D templates derived from 95 flare patterns and 3,000 flare image pairs rendered from 60 3D scenes. Furthermore, we design a masking approach to obtain real-world flare-free images from their corrupted counterparts to measure the performance of the model on real-world images. Extensive experiments demonstrate the effectiveness of our method and dataset. Lishen Qu, Jinshan Pan, Shihao Zhou 0003, Jinglei Shi, Duosheng Chen, Jufeng Yang |
NeurIPS | 5 |
| 2025 | ICTNet: Image Complexity-Aware Two-Branch Network With Enhanced Decoding for Real-Time Segmentation
Jinglei Shi, Teodor Boyadzhiev, Jufeng Yang |
IEEE Trans. Multim. | 2 |
| 2024 | Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationabstractTransformer-based approaches have achieved promising performance in image restoration tasks, given their ability to model long-range dependencies, which is crucial for recovering clear images. Though diverse efficient attention mechanism designs have addressed the intensive computations associated with using transformers, they often involve redundant information and noisy interactions from irrelevant regions by considering all available tokens. In this work, we propose an Adaptive Sparse Transformer (AST) to mitigate the noisy interactions of irrelevant areas and remove feature redundancy in both spatial and channel domains. AST comprises two core designs, i.e., an Adaptive Sparse Self-Attention (ASSA) block and a Feature Refinement Feed-forward Network (FRFN). Specifically, ASSA is adaptively computed using a two-branch paradigm, where the sparse branch is introduced to filter out the negative impacts of low query-key matching scores for aggregating features, while the dense one ensures sufficient information flow through the network for learning discriminative representations. Meanwhile, FRFN employs an enhance-and-ease scheme to eliminate feature redundancy in channels, enhancing the restoration of clear latent images. Experimental results on commonly used benchmarks have demonstrated the versatility and competitive performance of our method in several tasks, including rain streak removal, real haze removal, and raindrop removal. The code and pre-trained models are available at https://github.com/joshyZhou/AST. Shihao Zhou 0003, Duosheng Chen, Jinshan Pan, Jinglei Shi, Jufeng Yang |
CVPR | 4 |
| 2024 | Seeing the Unseen: A Frequency Prompt Guided Transformer for Image Restoration
Shihao Zhou 0003, Jinshan Pan, Jinglei Shi, Duosheng Chen, Lishen Qu, Jufeng Yang |
ECCV (16) | 3 |
| 2024 | ICFRNet: Image Complexity Prior Guided Feature Refinement for Real-time Semantic SegmentationabstractIn this paper, we leverage image complexity as a prior for refining segmentation features to achieve accurate real-time semantic segmentation. The design philosophy is based on the observation that different pixel regions within an image exhibit varying levels of complexity, with higher complexities posing a greater challenge for accurate segmentation. We thus introduce image complexity as prior guidance and propose the Image Complexity prior-guided Feature Refinement Network (ICFRNet). This network aggregates both complexity and segmentation features to produce an attention map for refining segmentation features within an Image Complexity Guided Attention (ICGA) module. We optimize the network in terms of both segmentation and image complexity prediction tasks with a combined loss function. Experimental results on the Cityscapes and CamViD datasets have shown that our ICFRNet achieves higher accuracy with a competitive efficiency for real-time segmentation. Teodor Boyadzhiev, Jinglei Shi, Jufeng Yang |
ICME | 3 |
| 2024 | To Err Like Human: Affective Bias-Inspired Measures for Visual Emotion Recognition EvaluationabstractAccuracy is a commonly adopted performance metric in various classification tasks, which measures the proportion of correctly classified samples among all samples. It assumes equal importance for all classes, hence equal severity for misclassifications. However, in the task of emotional classification, due to the psychological similarities between emotions, misclassifying a certain emotion into one class may be more severe than another, e.g., misclassifying 'excitement' as 'anger' apparently is more severe than as 'awe'. Albeit high meaningful for many applications, metrics capable of measuring these cases of misclassifications in visual emotion recognition tasks have yet to be explored. In this paper, based on Mikel's emotion wheel from psychology, we propose a novel approach for evaluating the performance in visual emotion recognition, which takes into account the distance on the emotion wheel between different emotions to mimic the psychological nuances of emotions. Experimental results in semi-supervised learning on emotion recognition and user study have shown that our proposed metrics is more effective than the accuracy to assess the performance and conforms to the cognitive laws of human emotions. The code is available at https://github.com/ZhaoChenxi-nku/ECC. Chenxi Zhao 0002, Jinglei Shi, Liqiang Nie, Jufeng Yang |
NeurIPS | 2 |
| 2024 | Learning Kernel-Modulated Neural Representation for Efficient Light Field CompressionabstractLight fields capture 3D scene information by recording light rays emitted from a scene at various orientations. They offer a more immersive perception, compared with classic 2D images, but at the cost of huge data volumes. In this paper, we design a compact neural network representation for the light field compression task. In the same vein as the deep image prior, the neural network takes randomly initialized noise as input and is trained in a supervised manner in order to best reconstruct the target light field Sub-Aperture Images (SAIs). The network is composed of two types of complementary kernels: descriptive kernels (descriptors) that store scene description information learned during training, and modulatory kernels (modulators) that control the rendering of different SAIs from the queried perspectives. To further enhance compactness of the network meanwhile retain high quality of the decoded light field, we propose modulator allocation and apply kernel tensor decomposition techniques, followed by non-uniform quantization and lossless entropy coding. Extensive experiments demonstrate that our method outperforms other state-of-the-art (SOTA) methods by a significant margin in the light field compression task. Moreover, after adapting descriptors, the modulators learned from one light field can be transferred to new light fields for rendering dense views, showing the potential of the solution for view synthesis. Jinglei Shi, Christine Guillemot |
IEEE Trans. Image Process. | 1 |
| 2023 | JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance FieldsabstractThis paper presents JAWS, an optimization-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an implicit-neural-representation (INR) in a way to compute a clip that shares the same cinematic features as the reference clip. We propose a general formulation of a camera optimization problem in an INR that computes extrinsic and intrinsic camera parameters as well as timing. By leveraging the differentiability of neural representations, we can back-propagate our designed cinematic losses measured on proxy estimators through a NeRF network to the proposed cinematic parameters directly. We also introduce specific enhancements such as guidance maps to improve the overall quality and efficiency. Results display the capacity of our system to replicate well known camera sequences from movies, adapting the framing, camera parameters and timing of the generated video clip to maximize the similarity with the reference clip. Robin Courant, Jinglei Shi, Éric Marchand, Marc Christie |
CVPR | 3 |
| 2023 | Light Field Compression Via Compact Neural Scene RepresentationabstractIn this paper, we propose a novel light field compression method based on a low rank-constrained neural scene representation. While most existing methods directly compress the light field views, our method first learns a Multi-Layer Perceptron (MLP)-based Neural Radiance Field (NeRF) from the input views. To be able to efficiently compress the NeRF scene representation, the weights of the MLP are optimized under a low-rank constraint using the Alternating Direction Method of Multipliers (ADMM) optimization method. The weights of NeRF are then decomposed into Tensor Train (TT) components which allow us to distill original NeRF network into a slimmer one. The slim NeRF is then refined using a quantization-aware training procedure. Experimental results show that this low rank-constrained NeRF-based light field compression method can achieve better rate-distortion than reference methods, while keeping the free-viewpoint reconstruction capability. Jinglei Shi, Christine Guillemot |
ICASSP | 1 |
| 2022 | Axial refocusing precision model with light fields
Zhaolin Xiao, Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
Signal Process. Image Commun. | 2 |
| 2022 | An Untrained Neural Network Prior for Light Field CompressionabstractDeep generative models have proven to be effective priors for solving a variety of image processing problems. However, the learning of realistic image priors, based on a large number of parameters, requires a large amount of training data. It has been shown recently, with the so-called deep image prior (DIP), that randomly initialized neural networks can act as good image priors without learning. In this paper, we propose a deep generative model for light fields, which is compact and which does not require any training data other than the light field itself. To show the potential of the proposed generative model, we develop a complete light field compression scheme with quantization-aware learning and entropy coding of the quantized weights. Experimental results show that the proposed method yields very competitive results compared with state-of-the-art light field compression methods, both in terms of PSNR and MS-SSIM metrics. Xiaoran Jiang, Jinglei Shi, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2021 | A learning-based view extrapolation method for axial super-resolution
Zhaolin Xiao, Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
Neurocomputing | 2 |
| 2020 | Learning Fused Pixel and Feature-Based View Reconstructions for Light FieldsabstractIn this paper, we present a learning-based framework for light field view synthesis from a subset of input views. Building upon a light-weight optical flow estimation network to obtain depth maps, our method employs two reconstruction modules in pixel and feature domains respectively. For the pixel-wise reconstruction, occlusions are explicitly handled by a disparity-dependent interpolation filter, whereas inpainting on disoccluded areas is learned by convolutional layers. Due to disparity inconsistencies, the pixel-based reconstruction may lead to blurriness in highly textured areas as well as on object contours. On the contrary, the feature-based reconstruction well performs on high frequencies, making the reconstruction in the two domains complementary. End-to-end learning is finally performed including a fusion module merging pixel and feature-based reconstructions. Experimental results show that our method achieves state-of-the-art performance on both synthetic and real-world datasets, moreover, it is even able to extend light fields' baseline by extrapolating high quality views without additional training. Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
CVPR | 1 |
| 2019 | A Learning Based Depth Estimation Framework for 4D Densely and Sparsely Sampled Light FieldsabstractThis paper proposes a learning based solution to disparity (depth) estimation for either densely or sparsely sampled light fields. Disparity between stereo pairs among a sparse subset of anchor views is first estimated by a fine-tuned FlowNet 2.0 network adapted to disparity prediction task. These coarse estimates are fused by exploiting the photo-consistency warping error, and refined by a Multi-view Stereo Refinement Network (MSRNet). The propagation of disparity from anchor viewpoints towards other viewpoints is performed by an occlusion-aware soft 3D reconstruction method. The experiments show that, both for dense and sparse light fields, our algorithm outperforms significantly the state-of-the-art algorithms, especially for subpixel accuracy. Xiaoran Jiang, Jinglei Shi, Christine Guillemot |
ICASSP | 2 |
| 2019 | A Framework for Learning Depth From a Flexible Subset of Dense and Sparse Light Field ViewsabstractIn this paper, we propose a learning-based depth estimation framework suitable for both densely and sparsely sampled light fields. The proposed framework consists of three processing steps: initial depth estimation, fusion with occlusion handling, and refinement. The estimation can be performed from a flexible subset of input views. The fusion of initial disparity estimates, relying on two warping error measures, allows us to have an accurate estimation in occluded regions and along the contours. In contrast with methods relying on the computation of cost volumes, the proposed approach does not need any prior information on the disparity range. Experimental results show that the proposed method outperforms state-of-the-art light fields depth estimation methods, including prior methods based on deep neural architectures. Jinglei Shi, Xiaoran Jiang, Christine Guillemot |
IEEE Trans. Image Process. | 1 |