Xiaoyi He

dblp:220/3273 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0006-4440-8034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 84% Image and video coding · 16%
Artificial intelligence
2 papers
Representation and self-supervised learning · 70% Deep learning architectures and training · 30%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration › multi-task image restoration
all-in-one image restoration
1.922026
Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Degradation-Aware Residual-Conditioned Optimal Transport for Unified Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Image and video processing
image restoration
1.922026
Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Degradation-Aware Residual-Conditioned Optimal Transport for Unified Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Image and video processing › image restoration
artifact removal
0.412020
Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC · IEEE Trans. Multim. 2020
Image and video coding › video compression › video codec
HEVC
0.412020
Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC · IEEE Trans. Multim. 2020
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.312026
Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Image and video processing › image restoration › unsupervised image restoration
unpaired image restoration
0.312025
Degradation-Aware Residual-Conditioned Optimal Transport for Unified Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training
convolutional neural network
0.112020
Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC · IEEE Trans. Multim. 2020

Methods — techniques the papers use, named apart from their topics

wasserstein barycenter · 2.0residual subspace disentanglement · 2.0contrastive learning · 2.0residual-conditioned transport map · 0.9optimal transport · 0.9iterative training · 0.9convolutional neural network · 0.9
YearPublicationVenuePosition
2026 Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image Restoration
abstract
Despite substantial advances in all-in-one image restoration for addressing diverse degradations within a unified model, existing methods remain vulnerable to out-of-distribution degradations, thereby limiting their generalization in real-world scenarios. To tackle the challenge, this work is motivated by the intuition that multisource degraded feature distributions are induced by different degradation-specific shifts from an underlying degradation-agnostic distribution, and recovering such a shared distribution is thus crucial for achieving generalization across degradations. With this insight, we propose BaryIR, a representation learning framework that aligns multisource degraded features in the Wasserstein barycenter (WB) space, which models a degradation-agnostic distribution by minimizing the average of Wasserstein distances to multisource degraded distributions. We further introduce residual subspaces, whose embeddings are mutually contrasted while remaining orthogonal to the WB embeddings. Consequently, BaryIR explicitly decouples two orthogonal spaces: a WB space that encodes the degradation-agnostic invariant contents shared across degradations, and residual subspaces that adaptively preserve the degradation-specific knowledge. This disentanglement mitigates overfitting to in-distribution degradations and enables adaptive restoration grounded on the degradation-agnostic shared invariance. Extensive experiments demonstrate that BaryIR performs competitively against state-of-the-art all-in-one methods. Notably, BaryIR generalizes well to unseen degradations (e.g., types and levels) and shows remarkable robustness in learning generalized features, even when trained on limited degradation types and evaluated on real-world data with mixed degradations.
Xiaole Tang, Xiaoyi He, Xiang Gu 0005, Jian Sun 0009
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Degradation-Aware Residual-Conditioned Optimal Transport for Unified Image Restoration
abstract
Unified, or more formally, all-in-one image restoration has emerged as a practical and promising low-level vision task for real-world applications. In this context, the key issue lies in how to deal with different types of degraded images simultaneously. Existing methods fit joint regression models over multi-domain degraded-clean image pairs of different degradations. However, due to the severe ill-posedness of inverting heterogeneous degradations, they often struggle with thoroughly perceiving the degradation semantics and rely on paired data for supervised training, yielding suboptimal restoration maps with structurally compromised results and lacking practicality for real-world or unpaired data. To break the barriers, we present a Degradation-Aware Residual-Conditioned Optimal Transport (DA-RCOT) approach that models (all-in-one) image restoration as an optimal transport (OT) problem for unpaired and paired settings, introducing the transport residual as a degradation-specific cue for both the transport cost and the transport map. Specifically, we formalize image restoration with a residual-guided OT objective by exploiting the degradation-specific patterns of the Fourier residual in the transport cost. More crucially, we design the transport map for restoration as a two-pass DA-RCOT map, in which the transport residual is computed in the first pass and then encoded as multi-scale residual embeddings to condition the second-pass restoration. This conditioning process injects intrinsic degradation knowledge (e.g., degradation type and level) and structural information from the multi-scale residual embeddings into the OT map, which thereby can dynamically adjust its behaviors for all-in-one restoration. Extensive experiments across five degradations demonstrate the favorable performance of DA-RCOT as compared to state-of-the-art methods, in terms of distortion measures, perceptual quality, and image structure preservation. Notably, DA-RCOT delivers superior adaptability to real-world scenarios even with mixed degradations and shows distinctive robustness to both degradation levels and the number of degradations.
Xiaole Tang, Xiang Gu 0005, Xiaoyi He, Jian Sun 0009
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Exploring target-related information with reliable global pixel relationships for robust RGB-T tracking
Tianlu Zhang, Xiaoyi He, Yongjiang Luo, Qiang Zhang 0020, Jungong Han
Pattern Recognit.2
2024 AMNet: Learning to Align Multi-Modality for RGB-T Tracking
abstract
RGB-T tracking has attracted increasing attention recently due to the all-weather and all-day working capability. However, most current RGB-T trackers usually assume that RGB data and thermal infrared (TIR) data are well spatially aligned, which is difficult to be achieved in practice. Such spatial misalignment between RGB data and TIR data may lead to the ineffective cross-modal information propagation during multi-modal feature fusion, thus reducing the tracking performance. In addition, due to the discrepancy in imaging characteristics of RGB images and TIR images, there also exist great differences between the information captured by the two modality data. The differences in characteristics of RGB and TIR modalities in different local areas will cause a single fusion strategy to be unable to fully explore the complementary information within multi-modal data. For that, we propose an RGB-T tracker, referred to as AMNet, to specifically solve such two problems with two dedicated modules, i.e., a Mutual-interacted Spatial Alignment (MSA) module and an Information Matching Fusion (IMF) module. The former spatially aligns the two modality data through three essential parts, including interactions of multi-modal features, prediction of cross-modal offset map, and enhancement of the aligned features. While the latter first discriminates different types of local regions by employing several intra-modal attention modules and then uses a divide-and-conquer fusion strategy to exploit such discriminative information within RGB and TIR features of different cases for tracking. We validate the effectiveness of our AMNet with extensive experiments on three RGB-T benchmarks, which achieves new state-of-the-art performance.
Tianlu Zhang, Xiaoyi He, Qiang Jiao, Qiang Zhang 0020, Jungong Han
IEEE Trans. Circuits Syst. Video Technol.2
2020 Adaptive lossless compression of skeleton sequences
Weiyao Lin, Tushar Shankar Shinde, Wenrui Dai, Mingzhou Liu 0001, Xiaoyi He, Anil Kumar Tiwari, Hongkai Xiong
Signal Process. Image Commun.5
2020 Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVC
abstract
This article addresses neural network based post-processing for the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC). We first propose a partition-aware convolution neural network (CNN) that utilizes the partition information produced by the encoder to assist in the post-processing. In contrast to existing CNN-based approaches, which only take the decoded frame as input, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the artifacts introduced by HEVC are efficiently reduced. We further introduce an adaptive-switching neural network (ASN) that consists of multiple independent CNNs to adaptively handle the variations in content and distortion within compressed-video frames, providing further reduction in visual artifacts. Additionally, an iterative training procedure is proposed to train these independent CNNs attentively on different local patch-wise classes. Experiments on benchmark sequences demonstrate the effectiveness of our partition-aware and adaptive-switching neural networks.
Weiyao Lin, Xiaoyi He, Xintong Han, Dong Liu 0002, John See, Junni Zou, Hongkai Xiong, Feng Wu 0001
IEEE Trans. Multim.2
2018 Enhancing HEVC Compressed Videos with a Partition-Masked Convolutional Neural Network
abstract
In this paper, we propose a partition-masked Convolution Neural Network (CNN) to achieve compressed-video enhancement for the state-of-the-art coding standard, High Efficiency Video Coding (HECV). More precisely, our method utilizes the partition information produced by the encoder to guide the quality enhancement process. In contrast to existing CNN-based approaches, which only take the decoded frame as the input to the CNN, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the degradation introduced by HEVC is reduced more efficiently. Experimental results show that our approach leads to over 9.76% BD-rate saving on benchmark sequences, which achieves the state-of-the-art performance.
Xiaoyi He, Qiang Hu 0003, Xiaoyun Zhang 0001, Weiyao Lin, Xintong Han
ICIP1