Kepeng Xu

dblp:247/4181 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-0650-2442ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation Learning
abstract
High-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic Range (SDR) content to HDR. Existing methods primarily rely on fixed tone mapping operators, which struggle to handle the diverse appearances and degradations commonly present in real-world SDR content. To address this limitation, we propose a generalized SDR-to-HDR framework that enhances robustness by learning attribute-disentangled representations. Central to our approach is Realistic Attribute-Disentangled Representation Learning (RealRep), which explicitly disentangles luminance and chrominance components to capture intrinsic content variations across different SDR distributions. Furthermore, we design a Luma-/Chroma-aware negative exemplar generation strategy that constructs degradation-sensitive contrastive pairs, effectively modeling tone discrepancies across SDR styles. Building on these attribute-level priors, we introduce the Degradation-Domain Aware Controlled Mapping Network (DDACMNet), a lightweight, two-stage framework that performs adaptive hierarchical mapping guided by a control-aware normalization mechanism. DDACMNet dynamically modulates the mapping process via degradation-conditioned features, enabling robust adaptation across diverse degradation domains. Extensive experiments demonstrate that RealRep consistently outperforms state-of-the-art methods in both generalization and perceptually faithful HDR color gamut reconstruction.
Li Xu 0008, Kepeng Xu, Lin Zhang 0040, Gang He 0002, Yu-Wing Tai
AAAI3
2025 Beyond Feature Mapping GAP: Integrating Real HDRTV Priors for Superior SDRTV-to-HDRTV Conversion
abstract
The rise of HDR-WCG display devices has highlighted the need to convert SDRTV to HDRTV, as most video sources are still in SDR. Existing methods primarily focus on designing neural networks to learn a single-style mapping from SDRTV to HDRTV. However, the limited information in SDRTV and the diversity of styles in real-world conversions render this process an ill-posed problem, thereby constraining the performance and generalization of these methods. Inspired by generative approaches, we propose a novel method for SDRTV to HDRTV conversion guided by real HDRTV priors. Despite the limited information in SDRTV, introducing real HDRTV as reference priors significantly constrains the solution space of the originally high-dimensional ill-posed problem. This shift transforms the task from solving an unreferenced prediction problem to making a referenced selection, thereby markedly enhancing the accuracy and reliability of the conversion process. Specifically, our approach comprises two stages: the first stage employs a Vector Quantized Generative Adversarial Network to capture HDRTV priors, while the second stage matches these priors to the input SDRTV content to recover realistic HDRTV outputs. We evaluate our method on public datasets, demonstrating its effectiveness with significant improvements in both objective and subjective metrics across real and synthetic datasets.
Gang He 0002, Kepeng Xu, Li Xu 0008, Wenxin Yu 0001, Xianyun Wu
IJCAI2
2025 FCKT: Fine-Grained Cross-Task Knowledge Transfer with Semantic Contrastive Learning for Targeted Sentiment Analysis
abstract
In this paper, we address the task of targeted sentiment analysis , which involves two sub-tasks, i.e., identifying specific aspects from reviews and determining their corresponding senti-ments. Aspect extraction forms the foundation for sentiment prediction, highlighting the critical dependency between these two tasks for effective cross-task knowledge transfer. While most existing studies adopt a multi-task learning paradigm to align task-specific features in the latent space, they predominantly rely on coarse-grained knowledge transfer. Such approaches lack fine-grained control over aspect-sentiment relationships, often assuming uniform sentiment polarity within related aspects. This oversimplification neglects contextual cues that differentiate sentiments, leading to negative transfer. To overcome these limitations, we propose FCKT, a fine-grained cross-task knowledge transfer framework tailored for TSA. By explicitly incorporating aspect-level information into sentiment prediction, our framework achieves fine-grained knowledge transfer, effectively mitigating negative transfer and enhancing task performance. Extensive experiments on three real-world datasets, including comparisons with various baselines and large language models (LLMs), demonstrate the effectiveness of FCKT. The source code is available on https://github.com/cwei01/FCKT.
Wei Chen 0061, Zhao Zhang 0011, Kepeng Xu, Fuzhen Zhuang
IJCAI4
2025 Unleashing the Potential of Transformer Flow for Photorealistic Face Restoration
abstract
Face restoration is a challenging task due to the need to remove artifacts and restore details. Traditional methods usually use generative model prior to achieve face restoration, but the restored results are still insufficient in terms of realism and details. In this paper, we introduce OmniFace, a novel face restoration framework that leverages Transformer-based diffusion flow. By exploiting the scaling property of Transformer, OmniFace achieves high-resolution restoration with exceptional realism and detail. The framework integrates three key components: (1) a Transformer-driven vector estimation network, (2) a representation aligned ControlNet, and (3) an adaptive training strategy for face restoration. The inherent scaling law of Transformer architectures enables the restoration of high-quality faces at high resolution. The controlnet combined with pre-trained diffusion representation can be easily trained. The adaptive training strategy provides a vector field that is more suitable for face restoration. Comprehensive experiments demonstrate that OmniFace outperforms existing techniques in terms of restoration quality across multiple benchmark datasets, especially in restoring photographic-level texture details in high-resolution scenes.
Kepeng Xu, Li Xu 0008, Gang He 0002, Wei Chen 0062, Xianyun Wu, Wenxin Yu 0001
IJCAI1
2025 NeRI: Implicit Neural Representation for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) remains challenging due to the weak spatial features of targets and their susceptibility to background clutter. Recent studies have improved detection performance through the embedding of additional spatial representations. However, these feature prompting methods rely on discretely sampled feature spaces, which weaken high-frequency information and consequently limit their representational efficiency. To overcome this, we propose NeRI, a network that leverages the potential of implicit neural representations (INRs) through a continuous formulation to learn mappings from spatial coordinates to the high-frequency structural representations of targets. Specifically, these mappings are realized through INR Blocks (INRBs) integrated into different encoder layers, providing continuous spatial guidance from multi-scale inputs and enabling more accurate localization and distinction. In addition, to better model the distinction between foreground and background, we construct a hybrid U-shaped block (HUB) that combines a U-shaped Transformer block (UTB) and multi-scale convolution block (MCB). The UTB component effectively increases network depth and facilitates long-range dependency modeling across different scales, while the MCB employs convolutions with varying receptive fields to capture fine-grained local information, thereby enabling the two components to fully exploit their complementary strengths. Finally, we propose a simple yet effective spatial–semantic fusion (SSF) module that reweights and integrates spatial information from diverse layers to enhance the expressive power of the features. The proposed NeRI offers a robust solution for the accurate separation of targets from backgrounds. Experimental validation, conducted on three public datasets (i.e., NUDT-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrates the superior performance of NeRI compared to other methods. Open-source implementations will be available at https://github.com/Shangwei-Deng/NeRI.
Shangwei Deng, Qianwen Ma, Shangqi Deng, Ziqian Chen, Ruoqi Lian, Bincheng Li, Kepeng Xu, Xiaobo Li 0004, Haofeng Hu
IEEE Trans. Geosci. Remote. Sens.8
2025 IOVarNet: Inner-Outer Variation Synergy Network for Infrared Small Target Detection
abstract
Sparsity and weak characteristics of targets pose significant challenges in infrared small target detection (IRSTD). For convolutional neural network-based methods, the increase in semantic information during propagation is often accompanied by the degradation of spatial features, which hampers the performance of IRSTD. In this paper, we proposed the Inner-outer Variation Synergy Network (IOVarNet) for IRSTD, which explicitly enhances the spatial response of targets by reinforcing their structural representations across different layers of the network. Specifically, IOVarNet leverages the delicate Total Variation-inspired Module, which takes the form of a partial differential equation, and incorporates it through the Inner-outer Variational Synergy architecture to supplement the target’s structural information at both the inner and outer layers of the encoder and decoder. Besides, the Dual Space Attention mechanism was introduced to enhance the semantic distinction between the target and background, while fusing spatial features from different layers. Experimental validation was conducted on three public datasets (i.e., NUDT-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrating the performance superiority of IOVarNet over other methods. Open-source implementations will be available at https://github.com/Shangwei-Deng/IOVarNet.
Shangwei Deng, Qianwen Ma, Bincheng Li, Liaoran Jin, Kepeng Xu, Shangqi Deng, Xiaobo Li 0004, Haofeng Hu
IEEE Trans. Geosci. Remote. Sens.5
2024 Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
Kepeng Xu, Li Xu 0008, Gang He 0002, Wenxin Yu 0001, Yunsong Li 0001
IJCAI1
2024 An End-to-End Real-World Camera Imaging Pipeline
abstract
pipeline still faces challenges including the lack of joint optimization in system components, computational redundancies, and optical distortions such as lens shading.In light of this, we propose an end-to-end camera imaging pipeline (RealCamNet) to enhance realworld camera imaging performance.Our methodology diverges from conventional, fragmented multi-stage image signal processing towards end-to-end architecture.This architecture facilitates joint optimization across the full pipeline and the restoration of coordinate-biased distortions.RealCamNet is designed for highquality conversion from RAW to RGB and compact image compression.Specifically, we deeply analyze coordinate-dependent optical distortions, e.g., vignetting and dark shading, and design a novel 2804
Kepeng Xu, Zijia Ma, Li Xu 0008, Gang He 0002, Yunsong Li 0001, Wenxin Yu 0001, Taichu Han, Cheng Yang 0016
ACM Multimedia1
2022 SDRTV-to-HDRTV via Hierarchical Dynamic Context Feature Mapping
abstract
In this work, we address the task of SDR videos to HDR videos(SDRTV-to-HDRTV conversion). Previous approaches use global feature modulation for SDRTV-to-HDRTV conversion. Feature modulation scales and shifts the features in the original feature space, which has limited mapping capability. In addition, the global image mapping cannot restore detail in HDR frames due to the luminance differences in different regions of SDR frames. To resolve the appeal, we propose a two-stage solution. The first stage is a hierarchical Dynamic Context feature mapping (HDCFM) model. HDCFM learns the SDR frame to HDR frame mapping function via hierarchical feature modulation (HME and HM ) module and a dynamic context feature transformation (DYCT) module. The HME estimates the feature modulation vector, HM is capable of hierarchical feature modulation, consisting of global feature modulation in series with local feature modulation, and is capable of adaptive mapping of local image features. The DYCT module constructs a feature transformation module in conjunction with the context, which is capable of adaptively generating a feature transformation matrix for feature mapping. Compared with simple feature scaling and shifting, the DYCT module can map features into a new feature space and thus has a more excellent feature mapping capability. In the second stage, we introduce a patch discriminator-based context generation model PDCG to obtain subjective quality enhancement of over-exposed regions. The proposed method can achieve state-of-the-art objective and subjective quality results. Specifically, HDCFM achieves a PSNR gain of 0.81 dB at about 100K parameters. The number of parameters is 1/14th of the previous state-of-the-art methods. The test code will be released on https://github.com/cooperlike/HDCFM.
Gang He 0002, Kepeng Xu, Li Xu 0008, Chang Wu 0001, Ming Sun 0008, Yu-Wing Tai
ACM Multimedia2
2020 Interactive Separation Network For Image Inpainting
abstract
Image inpainting, also known as image completion, is the process of filling in the missing region of an incomplete image to make the repaired image visually plausible. Strided convolutional layer learns high-level representations while reducing the computational complexity, but fails to preserve existing detail from the original images (eg, texture, sharp transients), therefore it degrades the generative model in image inpainting task. To reduce the erosion of high-resolution components of images meanwhile maintaining the semantic representation, this paper designs a brand-new network called Interactive Separation Network that progressively decomposites the features into two streams and fuses them. Besides, the rationality of network design and the efficiency of proposed network is demonstrated in the ablation study. To the best of our knowledge, the experimental results of proposed method are superior to state-of-the-art inpainting approaches.
Siyuan Li 0004, Xin Cheng 0004, Kepeng Xu, Wenxin Yu 0001, Gang He 0002, Jinjia Zhou
ICIP5
2020 LPI-Net: Lightweight Inpainting Network with Pyramidal Hierarchy
Siyuan Li 0004, Kepeng Xu, Wenxin Yu 0001, Ning Jiang 0002
ICONIP (4)3
2019 Target-Based Attention Model for Aspect-Level Sentiment Analysis
Wei Chen 0062, Wenxin Yu 0001, Yunye Zhang, Kepeng Xu, Fengwei Zhang, Yibo Fan, Gang He 0002
ICONIP (3)5
2019 Inpainting with Sketch Reconstruction and Comprehensive Feature Selection
Siyuan Li 0004, Zhijing Li 0003, Kepeng Xu, Matthieu Claisse, Wenxin Yu 0001, Gang He 0001, Gang He 0002, Yibo Fan
ICONIP (5)4
2019 IRSNET: An Inception-Resnet Feature Reconstruction Model for Building Segmentation
Kepeng Xu, Li Nie, Wenxin Yu 0001, Yunye Zhang, Wei Chen 0062, Siyuan Li 0004, Shangwei Deng, Yibo Fan, Hui Zhang 0051, Valentin Bouillon
ICONIP (5)1