VLDB 2026 Research / reviewers in the wild / expert
Xuelin Shen
dblp:215/7115
· DBLP profile ↗
21ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0002-0877-7143ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Malice Hides in Equivalence: Attacking Text-Image Alignment Assessment with Subtle Text Variations
Kang Xiao, Xuelin Shen, Baoliang Chen, Wenhan Yang, Meng Wang 0001 |
ISCAS | 2 |
| 2026 | Towards Generalized Image Coding for Machine Through Meta Adversarial Adaptation
Xuelin Shen, Kangsheng Yin, Wenhan Yang |
Int. J. Comput. Vis. | 1 |
| 2025 | Fast Omni-Directional Image Super-Resolution: Adapting the Implicit Image Function with Pixel and Semantic-Wise Spherical Geometric PriorsabstractIn the context of Omni-Directional Image (ODI) Super-Resolution (SR), the unique challenge arises from the non-uniform oversampling characteristics caused by EquiRectangular Projection (ERP). Considerable efforts in designing complex spherical convolutions or polyhedron reprojection offer significant performance improvements but at the expense of cumbersome processing procedures and slower inference speeds. Under these circumstances, this paper proposes a new ODI-SR model characterized by its capacity to perform Fast and Arbitrary-scale ODI-SR processes, denoted as FAOR. The key innovation lies in adapting the implicit image function from the planar image domain to the ERP image domain by incorporating spherical geometric priors at both the latent representation and image reconstruction stages, in a low-overhead manner. Specifically, at the latent representation stage, we adopt a pair of pixel-wise and semantic-wise sphere-to-planar distortion maps to perform affine transformations on the latent representation, thereby incorporating it with spherical properties. Moreover, during the image reconstruction stage, we introduce a geodesic-based resampling strategy, aligning the implicit image function with spherical geometrics without introducing additional parameters. As a result, the proposed FAOR outperforms the state-of-the-art ODI-SR models with a much faster inference speed. Extensive experimental results and ablation studies have demonstrated the effectiveness of our design. Xuelin Shen, Silin Zheng, Kang Xiao, Wenhan Yang |
AAAI | 1 |
| 2025 | Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP SupervisionabstractThe image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre-training models and incorporates adaptive multi-objective optimization tailored to support both human visual perception and machine vision simultaneously with a single bitstream, denoted as Unified and Generalized Image Coding for Machine (UG-ICM). Specifically, to get rid of the reliance between compression models with downstream task supervision, we introduce Contrastive Language-Image Pre-training (CLIP) models into the training constraint for improved generalization. Global-to-instance-wise CLIP supervision is applied to help obtain hierarchical semantics that make models more generalizable for the tasks relying on the information of different granularity. Furthermore, for supporting both human and machine visions with only a unifying bitstream, we incorporate a conditional decoding strategy that takes as conditions human or machine preferences, enabling the bitstream to be decoded into different versions for corresponding preferences. As such, our proposed UG-ICM is fully trained in a self-supervised manner, i.e., without awareness of any specific downstream models and tasks. The extensive experiments have shown that the proposed UG-ICM is capable of achieving remarkable improvements in various unseen machine analytics tasks, while simultaneously providing perceptually satisfying images. Kangsheng Yin, Xuelin Shen, Yu-Lin He, Wenhan Yang, Shiqi Wang 0001 |
AAAI | 3 |
| 2025 | Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained ModelsabstractThe improved semantic understanding of vision-language pretrained (VLP) models has made it increasingly difficult to protect publicly posted images from being exploited by search engines and other similar tools. In this context, this paper seeks to protect users' privacy by implementing defenses at the image compression stage to prevent exploitation. Specifically, we propose a flexible coding method, termed Privacy-Shielded Image Compression (PSIC), that can produce bitstreams with multiple decoding options. By default, the bitstream is decoded to preserve satisfactory perceptual quality while preventing interpretation by VLP models. Our method also retains the original image compression functionality. With a customizable input condition, the proposed scheme can reconstruct the image that preserves its full semantic information. A Conditional Latent Trigger Generation (CLTG) module is proposed to produce bias information based on customizable conditions to guide the decoding process into different reconstructed versions, and an Uncertainty-Aware Encryption-Oriented (UAEO) optimization function is designed to leverage the soft labels inferred from the target VLP model's uncertainty on the training data. This paper further incorporates an adaptive multi-objective optimization strategy to obtain improved encrypting performance and perceptual quality simultaneously within a unified training process. The proposed scheme is plug-and-play and can be seamlessly integrated into most existing Learned Image Compression (LIC) models. Extensive experiments across multiple downstream tasks have demonstrated the effectiveness of our design. Xuelin Shen, Jiayin Xu, Kangsheng Yin, Wenhan Yang |
ICML | 1 |
| 2025 | End-to-End Low-Light Enhancement for Object Detection with Learned Metadata from RAWsabstractAlthough RAW images offer advantages over sRGB by avoiding ISP-induced distortion and preserving more information in low-light conditions, their widespread use is limited due to high storage costs, transmission burdens, and the need for significant architectural changes for downstream tasks. To address the issues, this paper explores a new raw-based machine vision paradigm, termed Compact RAW Metadata-guided Image Refinement (CRM-IR). In particular, we propose a Machine Vision-oriented Image Refinement (MV-IR) module that refines sRGB images to better suit machine vision preferences, guided by learned raw metadata. Such a design allows the CRM-IR to focus on extracting the most essential metadata from raw images to support downstream machine vision tasks, while remaining plug-and-play and fully compatible with existing imaging pipelines, without any changes to model architectures or ISP modules. We implement our CRM-IR scheme on various object detection networks, and extensive experiments under low-light conditions demonstrate that it can significantly improve performance with an additional bitrate cost of less than $10^{-3}$ bits per pixel. Xuelin Shen, Haifeng Jiao, Yu-Lin He, Wenhan Yang |
NeurIPS | 1 |
| 2025 | Prompt-Guided Alignment with Information Bottleneck Makes Image Compression Also a RestorerabstractLearned Image Compression (LIC) models face critical challenges in real-world scenarios due to various environmental degradations, such as fog and rain. Due to the distribution mismatch between degraded inputs and clean training data, well-trained LIC models suffer from reduced compression efficiency, while retraining dedicated models for diverse degradation types is costly and impractical. Our method addresses the above issue by leveraging prompt learning under the information bottleneck principle, enabling compact extraction of shared components between degraded and clean images for improved latent alignment and compression efficiency. In detail, we propose an Information Bottleneck-constrained Latent Representation Unifying (IB-LRU) scheme, in which a Probabilistic Prompt Generator (PPG) is deployed to simultaneously capture the distribution of different degradations. Such a design dynamically guides the latent-representation process at the encoder through a gated modulation process. Moreover, to promote the degradation distribution capture process, the probabilistic prompt learning is guided by the Information Bottleneck (IB) principle. That is,IB constrains the information encoded in the prompt to focus solely on degradation characteristics while avoiding the inclusion of redundant image contextual information. We apply our IB-LRU method to a variety of state-of-the-art LIC backbones, and extensive experiments under various degradation scenarios demonstrate the effectiveness of our design. Our code will be publicly available. Xuelin Shen, Jiayin Xu, Wenhan Yang |
NeurIPS | 1 |
| 2025 | Monotonic and Invertible Network: A General Framework for Learning IQA Model from Mixed Datasets
Baoliang Chen, Kang Xiao, Xuelin Shen, Shiqi Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Breaking Boundaries: Unifying Imaging and Compression for HDR Image CompressionabstractHigh Dynamic Range (HDR) images present unique challenges for Learned Image Compression (LIC) due to their complex domain distribution compared to Low Dynamic Range (LDR) images. In coding practice, HDR-oriented LIC typically adopts preprocessing steps (e.g., perceptual quantization and tone mapping operation) to align the distributions between LDR and HDR images, which inevitably comes at the expense of perceptual quality. To address this challenge, we rethink the HDR imaging process which involves fusing multiple exposure LDR images to create an HDR image and propose a novel HDR image compression paradigm, Unifying Imaging and Compression (HDR-UIC). The key innovation lies in establishing a seamless pipeline from image capture to delivery and enabling end-to-end training and optimization. Specifically, a Mixture-ATtention (MAT)-based compression backbone merges LDR features while simultaneously generating a compact representation. Meanwhile, the Reference-guided Misalignment-aware feature Enhancement (RME) module mitigates ghosting artifacts caused by misalignment in the LDR branches, maintaining fidelity without introducing additional information. Furthermore, we introduce an Appearance Redundancy Removal (ARR) module to optimize coding resource allocation among LDR features, thereby enhancing the final HDR compression performance. Extensive experimental results demonstrate the efficacy of our approach, showing significant improvements over existing state-of-the-art HDR compression schemes. Our code is available at: https://github.com/plf1999/HDR-UIC. Xuelin Shen, Linfeng Pan, Zhangkai Ni, Yu-Lin He, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2025 | Transferring From Distortion to Perception-Oriented Optimization: Just-Noticeable-Distortion-Based Domain AdaptationabstractTheperception-distortion- tradeoffreveals the limitation of current low-level deep learning paradigms,i.e., minimizing reconstruction distortion does not guarantee improved perceptual quality. Acknowledging the lack of a reliableperception-oriented optimization function, we are motivated to explore a flexible approach for enhancing perceptual quality by steering thetradeoffto prioritizeperception. To this end, we reconsider theperception-distortionfunction by incorporating the Just-Noticeable-Distortion (JND) mechanism. We mathematically demonstrate that in the common image restoration process, altering the optimization target from natural images to distorted images—where the distortion intensity is constrained by the JND threshold and the distortion type aligns with that arising from the restorer itself—effectively obtained improvedperceptionindices without any changes to the restorer or optimization function. Accordingly, to facilitate various low-level learning models, we are motivated to construct the first large-scale CNN-oriented JND image dataset. Our dataset comprises 500 natural images and 4,500 degraded versions generated by a series of autoencoders, as well as the actual JND judgment results collected through rigorous subjective testing from twenty volunteers. Finally, a learning-based JND inference model is established on the proposed dataset and employed in the proposed JND-based adaptation scheme, where the inferred JND images serve as pseudo-ground truth for the training or fine-tuning processes of low-level vision models. Extensive experiments on image super-resolution and end-to-end image compression across multiple models have shown encouraging improvements in perceptual quality, demonstrating the effectiveness of the proposed scheme. Our dataset is available at:https://github.com/ohq17/CNN-Oriented-JND-Dataset. Xuelin Shen, Haoqiao Ou, Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2024 | Image Coding for Analytics via Adversarially Augmented AdaptationabstractImage Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. This paper proposes a novel adversarially augmented adaptation route that achieves a better trade-off between the utility of the human and machine perspectives by making slight changes to the image manifold. In detail, a targeted adversarial attack is employed to generate subtle image perturbations that are nearly imperceptible to humans but significantly improve machine analytic performance. These perturbed images would be subsequently employed as ground truth to guide training/fine-tuning of an end-to-end image compression network. Note that, our method is a plug-and-play framework that does not rely on any change in existing architecture or loss functions. Extensive experimental results demonstrate the superiority of the proposed scheme over conventional ICM frameworks and the effectiveness of our design. Xuelin Shen, Kangsheng Yin, Xu Wang 0006, Yu-Lin He, Shiqi Wang 0001, Wenhan Yang |
ICASSP | 1 |
| 2024 | Image Coding For Machine Via Analytics-Driven Appearance Redundancy ReductionabstractAmong various technical approaches in machine vision coding, Image Coding for Machine (ICM) stands out for its capability to simultaneously fulfill both human perception and machine vision needs. However, it is often criticized for its lack of efficiency regarding rate-analytics performance. In this paper, we propose an Appearance Redundancy Reduction (ARR) module, designed to function as a plug-in for existing ICM frameworks, aiming to further enhance the coding efficiency regarding rate analytics without any changes to the ICM itself. To be specific, our work pays additional attention to the intrinsic correlation between the low-level image structure and high-level vision analytics, and subsequently proposes a novel colour quantization mechanism to squeeze out the analytics-free redundant appearance information. Moreover, a differentiable soften quantization operation is derived to enable end-to-end training within the ICM framework. Extensive experimental results have shown that integrating the proposed ARR module yields substantial improvements regarding rate-analytic performance, even surpassing the performance of the feature coding paradigm, while maintaining the generalizability across different tasks and acceptable perceptual representation. Xuelin Shen, Haoqiao Ou, Wenhan Yang |
ICIP | 1 |
| 2024 | Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment EnhancementabstractFull-reference image quality assessment (FR-IQA) models generally operate by measuring the visual differences between a degraded image and its reference. However, existing FR-IQA models including both the classical ones (e.g., PSNR and SSIM) and deep-learning based measures (e.g., LPIPS and DISTS) still exhibit limitations in capturing the full perception characteristics of the human visual system (HVS). In this paper, instead of designing a new FR-IQA measure, we aim to explore a generalized human visual attention estimation strategy to mimic the process of human quality rating and enhance existing IQA models. In particular, we model human attention generation by measuring the statistical dependency between the degraded image and the reference image. The dependency is captured in a training-free manner by our proposed sliced maximal information coefficient and exhibits surprising generalization in different IQA measures. Experimental results verify the performance of existing IQA models can be consistently improved when our attention module is incorporated. The source code is available at https://github.com/KANGX99/SMIC. Kang Xiao, Xu Wang 0006, Yu-Lin He, Baoliang Chen, Xuelin Shen |
ICME | 5 |
| 2024 | Multi-exposure embeddings for graph learning: Towards high dynamic range image saliency predictionabstractAbstract Identifying saliency in high dynamic range (HDR) images is a fundamentally important issue in HDR imaging, and plays critical roles towards comprehensive scene understanding. Most of existing studies leverage hand‐crafted features for HDR image saliency prediction, lacking the capabilities of fully exploiting the characteristics of HDR image (i.e. wider luminance range and richer colour gamut). Here, systematical studies are carried out on HDR image saliency prediction by proposing a new framework to single out the contributions from multi‐exposure images. Specifically, inspired by the mechanism of HDR imaging, the method first utilizes graph neural networks to model the relations among multi‐exposure images and the tone‐mapped image obtained from an HDR image, enabling more discriminative saliency‐related feature representations. Subsequently, the saliency features driven by global semantic knowledge are aggregated from the tone‐mapped image through enhancing global context‐aware semantic information. Finally, a fusion module is designed to integrate saliency‐oriented feature representations originated from multi‐exposure images and the tone‐mapped image, producing the saliency maps of HDR images. Moreover, a new challenging HDR eye fixation database (HDR‐EYEFix) is created, expecting to further contribute the research on HDR image saliency prediction. Experiment results show that the method obtains superior performance compared to the state‐of‐the‐art methods. Jun Xing, Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
IET Image Process. | 3 |
| 2024 | Towards 360$^{\circ }$ image compression for machines via modulating pixel significance
Silin Zheng, Xuelin Shen, Qiudan Zhang, Zhuo Chen 0006, Wenhan Yang, Xu Wang 0006 |
Multim. Tools Appl. | 2 |
| 2023 | Hybrid Prior-Based Diminished Reality for Indoor Panoramic Images
Jiashu Liu, Qiudan Zhang, Xuelin Shen, Wenhui Wu 0001, Xu Wang 0006 |
CGI (3) | 3 |
| 2023 | Salient Object Detection on 360° Omnidirectional Image with Bi-Branch Hybrid Projection NetworkabstractWith the advent of panoramic cameras, modeling saliency in 360° omnidirectional images becomes very urgent and challenging. However, severe distortions limit the prediction accuracy of 360° saliency model. In this paper, we devise a bi-branch hybrid projection network (HPNet), which exploits characteristics of equirectangular projection (ERP) and cubic map projection (CMP) formats to predict salient objects in 360° omnidirectional images. Specifically, an ERP image and a CMP image are first fed into a bi-branch network to aggregate the comprehensive features of the omnidirectional image. Subsequently, to explore the coherence among ERP and CMP images, we design a hybrid projection feature fusion module to efficiently combine CMP and ERP features extracted from different layers. Ultimately, a progressive prediction module is developed to refine the features and locate salient objects incrementally, and then produce the final saliency map for the 360° omnidirectional image. Experimental results illustrate that our model is superior to the existing advanced methods in two publicly available datasets. Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
MMSP | 3 |
| 2022 | Semi-automatic Data Annotation System for Multi-Target Multi-Camera Vehicle TrackingabstractMulti-target multi-camera tracking (MTMCT) plays an important role in intelligent video analysis, surveillance video retrieval, and other application scenarios. Nowadays, the deep-learning-based MTMCT has been the mainstream and has achieved fascinating improvements regarding tracking accuracy and efficiency. However, according to our investigation, the lacking of datasets focusing on real-world application scenarios limits the further improvements for current learning-based MTMCT models. Specifically, the learning-based MTMCT models training by common datasets usually cannot achieve satisfactory results in real-world application scenarios. Motivated by this, this paper presents a semi-automatic data annotation system to facilitate the real-world MTMCT dataset establishment. The proposed system first employs a deep-learning-based single-camera trajectory generation method to automatically extract trajectories from surveillance videos. Subsequently, the system provides a recommendation list in the following manual cross-camera trajectory matching process. The recommendation list is generated based on side information, including camera location, timestamp relation, and background scene. In the experimental stage, extensive results further demonstrate the efficiency of the proposed system. Haohong Liao, Silin Zheng, Xuelin Shen, Mark Junjie Li, Xu Wang 0006 |
DSAA | 3 |
| 2021 | Just Noticeable Distortion Profile Inference: A Patch-Level Structural Visibility Learning ApproachabstractIn this paper, we propose an effective approach to infer the just noticeable distortion (JND) profile based on patch-level structural visibility learning. Instead of pixel-level JND profile estimation, the image patch, which is regarded as the basic processing unit to better correlate with the human perception, can be further decomposed into three conceptually independent components for visibility estimation. In particular, to incorporate the structural degradation into the patch-level JND model, a deep learning-based structural degradation estimation model is trained to approximate the masking of structural visibility. In order to facilitate the learning process, a JND dataset is further established, including 202 pristine images and 7878 distorted images generated by advanced compression algorithms based on the upcoming Versatile Video Coding (VVC) standard. Extensive experimental results further show the superiority of the proposed approach over the state-of-the-art. Our dataset is available at: https://github.com/ShenXuelin-CityU/PWJNDInfer. Xuelin Shen, Zhangkai Ni, Wenhan Yang, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2020 | Just Noticeable Distortion Based Perceptually Lossless Intra CodingabstractPerceptual video coding plays a very important role in video codec optimization aiming at removing the perceptual redundancies in video content. In this paper, a just noticeable distortion (JND) guided perceptually lossless coding framework is proposed for Versatile Video Coding (VVC) intra coding. Within this framework, a pattern-based pixel wise JND model is employed to guide the distortion distribution, and subsequently the most appropriate quantization parameter is chosen for each Coding Tree Unit (CTU). The content adaptive Laplacian distribution based D-Q model based on a two pass coding framework is established to derive the most proper QP that satisfies the perceptually lossless coding criteria. The whole framework is integrated into the H.266/VVC intra coding framework. Experimental results demonstrate that the proposed scheme can achieve high accuracy prediction and efficient perceptually lossless intra coding, leading to around 10% bitrate savings comparing with the frame level QP derivation scheme. Xuelin Shen, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong, Guopu Zhu |
ICASSP | 1 |
| 2017 | Feature based inter prediction optimization for non-translational video coding in cloudabstractVisual features of images and video frames have become pervasive and maturely developed in extensive research fields such as computer vision and visual search. In more and more cases, the visual feature becomes necessary information which needs to be transmitted and stored at server side in cloud. Among visual features, the local feature descriptors extracted by SIFT can represent both translational and non-translational motion, such as orientation and zooming. On the other hand, only translational motion can be represented by the Motion Vector (MV) in current MV based block video coding standard. Inspired by these properties, a method that utilizes the available feature to optimize inter prediction video coding is proposed in this paper. In this method, the localization, orientation and scale parameters of matching features extracted by SIFT are delivered to inter prediction to provide non-translational motion estimation (ME) and optimized merge mode. Experimental results have shown that the proposed method can efficiently improve the coding performance according to the accurate feature-matching. Xuelin Shen, Jun Wang 0015, Peilin Chen 0001, Fan Liang 0001 |
VCIP | 1 |