VLDB 2026 Research / reviewers in the wild / expert
Yuan Chen 0012
dblp:55/6958-12
· DBLP profile ↗
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-5344-958XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BITMNet: A Degradation-Aware Mixture-of-Experts Framework for Blind Inverse Tone Mapping
Wenyou Zhang, Yang Zhao 0002, Fangxing Zhang, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Joint Resolution and Rendering Artifacts Removal for Cloud Gaming ImageabstractWith the rapid development of the cloud gaming industry, low-quality rendering and rescaling strategies are commonly employed to mitigate the high costs of cloud-based computation and bandwidth. As a result, client-side images often contain artifacts such as mixed aliasing and resolution distortions, which cannot be effectively handled by current super-resolution models. In response, this paper proposes a cloud gaming image enhancement (CGIE) model to tackle both rendering and resolution degradations. Initially, this paper builds a large dataset by rendering and rescaling paired data with different qualities from collected 3D game scenes. Subsequently, a lightweight dual-branch enhancement network is designed, which consists of a high-frequency branch primarily focused on detail enhancement and a sampling-space branch aimed at enlarging the receptive field and perceiving multi-scale aliasing artifacts. Experimental results demonstrate the superior anti-aliasing and image enhancement performance of the proposed method across various real-world cloud games and even mobile games. The dataset and codes are available at https://github.com/YCheno/CGIE/. Yang Zhao 0002, Yuan Chen 0012, Lin Li 0053, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Unlocking Cross-Domain Synergies for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation semantic segmentation (UDASS) aims to perform dense prediction on the unlabeled target domain by training the model on a labeled source domain. In this field, self-training approaches have demonstrated strong competitiveness and advantages. However, existing methods often rely on additional training data (such as reference datasets or depth maps) to rectify the unreliable pseudo-labels, ignoring the cross-domain interaction between the target and source domains. To address this issue, in this paper, we propose a novel method for unsupervised domain adaptation semantic segmentation, termed Unlocking Cross-Domain Synergies (UCDS). Specifically, in the UCDS network, we design a new Dynamic Self-Correction (DSC) module that effectively transfers source domain knowledge and generates high-confidence pseudo-labels without additional training resources. Unlike the existing methods, DSC proposes a Dynamic Noisy Label Detection method for the target domain. To correct the noisy pseudo-labels, we design a Dual Bank mechanism that explores the reliable and unreliable predictions of the source domain, and conducts cross-domain synergy through Weighted Reassignment Self-Correction and Negative Correction Prevention strategies. To enhance the discriminative ability of features and amplify the dissimilarity of different categories, we propose Discrepancy-based Contrastive Learning (DCL). The DCL selects positive and negative samples in the source and target domains based on the semantic discrepancies among different categories, effectively avoiding the numerous false negative samples found in existing methods. Extensive experimental results on three commonly used datasets demonstrate the superiority of the proposed UCDS in comparison with the state-of-the-art methods. The project and code are available at https://github.com/wqh011128/UCDS. Qihang Wu, Bo Jiang 0002, Yuan Chen 0012, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Multi-Layer Gaussian Splatting for Single-Image Feed-Forward Spatial Scene ReconstructionabstractRecently, 3D Gaussian Splatting (3DGS) has achieved remarkable results in 3D reconstruction and view synthesis tasks. However, single-view feed-forward 3DGS still faces significant challenges. Current state-of-the-art (SOTA) single-view 3DGS methods typically employ a small number of layers (1-2 layers) with Gaussian Splatting (GS) representations at the same resolution as the input image to address the irregularity of GS data. However, such shallow and uniform GS primitive distributions is difficult to represent occluded regions and important spatial details. Inspired by multi-plane images, this paper proposes a Multi-Layer Gaussian Splatting (MLGS) representation, which consists of shallow base GS layers for visible content and multiple occlusion GS layers dedicated to reconstructing occluded regions. The proposed MLGS representation explicitly decouples the learning processes of visible and occluded content while enhancing occlusion prediction through the following components. First, spatial stratification of GS is achieved by estimating the depth distribution range of GS primitives across different layers, forcing GS to learn spatial content reconstruction at different depths. Second, a mask-guided mechanism is proposed to effectively isolate occlusion regions and guide inpainting using spatially context-aware features. Finally, a gated convolution block is designed to dynamically modulate feature fusion to enhance reconstruction fidelity. With separate loss supervision for base and occlusion layers, MLGS enables geometrically plausible scene completion. Experiments on RealEstate10K, KITTI, and NYUv2 datasets demonstrate that the proposed method achieves SOTA performance for single-image spatial scene reconstruction. Shanding Diao, Yang Zhao 0002, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001, Ronggang Wang |
ACM Multimedia | 3 |
| 2025 | Local Texture Pattern Estimation for Image Detail Super-ResolutionabstractIn the image super-resolution (SR) field, recovering missing high-frequency textures has always been an important goal. However, deep SR networks based on pixel-level constraints tend to focus on stable edge details and cannot effectively restore random high-frequency textures. It was not until the emergence of the generative adversarial network (GAN) that GAN-based SR models achieved realistic texture restoration and quickly became the mainstream method for texture SR. However, GAN-based SR models still have some drawbacks, such as relying on a large number of parameters and generating fake textures that are inconsistent with ground truth. Inspired by traditional texture analysis research, this paper proposes a novel SR network based on local texture pattern estimation (LTPE), which can restore fine high-frequency texture details without GAN. A differentiable local texture operator is first designed to extract local texture structures, and a texture enhancement branch is used to predict the high-resolution local texture distribution based on the LTPE. Then, the predicted high-resolution texture structure map can be used as a reference for the texture fusion SR branch to obtain high-quality texture reconstruction. Finally, $L_{1}$L1 loss and Gram loss are simultaneously used to optimize the network. Experimental results demonstrate that the proposed method can effectively recover high-frequency texture without using GAN structures. In addition, the restored high-frequency details are constrained by local texture distribution, thereby reducing significant errors in texture generation. Yang Zhao 0002, Yuan Chen 0012, Nannan Li 0001, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Multimodal Remote Sensing Image Registration via Modality Perception and Self-Supervised Position EstimationabstractMulti-modal remote sensing images registration ensures that images from different sensors or modalities are spatial and informational consistent for effective comparison and analysis. However, due to the non-linear modality gaps that exist between images, making it difficult to focus only on the spatial position differences of the images and ignore the modality gaps. In this paper, to address this issue, we propose a new framework for Multi-Modal remote sensing image Registration, named MMRNet. The proposed framework comprises the following main aspects. First, a novel self-supervised Positional Misalignment Estimator (PME) is designed for multi-modal image registration. PME is able to efficiently overcome the modality gaps and learn the positional differences between multi-modal images more reliably, optimizing the registration loss by minimizing the positional differences directly. Then, a new paradigm of modality translation, termed Modality Perception Module (MPM), is introduced to effectively learn modality gaps and perform modality translation in the case of positional misalignment. Finally, we further design the modality perception guidance loss to supervise the modality translation task, which can encourage the fidelity of the generated pseudo-modality images. Our registration network integrates both rigid registration model and non-rigid registration model. Experimental results demonstrate that the proposed registration framework can obtain obviously superior performance in both rigid and non-rigid image registration tasks on optical-SAR data, optical-map data and optical-infrared data. The code and relevant dataset will be made publicly available at https://github.com/Ahuer-Lei/MMRNet. Yun Xiao 0003, Bo Jiang 0002, Yuan Chen 0012, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane ImagesabstractWith the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of occlusion exposure, which occurs when the occluded contents become visible in other views. Recently, the single-view multiplane images (MPI) representation has shown promising performance for planar video stereoscopy. However, the MPI still lacks real details that are occluded in the current frame, resulting in blurry artifacts in occlusion exposure regions. In fact, planar videos can leverage complementary information from adjacent frames to predict a more complete scene representation for the current frame. Therefore, this paper extends the MPI from still frames to the temporal domain, introducing the temporal MPI (TMPI). By extracting complementary information from adjacent frames based on optical flow guidance, obscured regions in the current frame can be effectively repaired. Additionally, a new module called masked optical flow warping (MOFW) is introduced to improve the propagation of pixels along optical flow trajectories. Experimental results demonstrate that the proposed method can generate high-quality stereoscopic or light-field videos from a single view and reproduce better occluded details than other state-of-the-art (SOTA) methods. https://github.com/Dio3ding/TMPI Shanding Diao, Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
AAAI | 2 |
| 2024 | MGDR: Multi-modal Graph Disentangled Representation for Brain Disease Prediction
Bo Jiang 0002, Xixi Wan, Yuan Chen 0012, Zhengzheng Tu, Yumiao Zhao, Jin Tang 0001 |
MICCAI (2) | 4 |
| 2024 | Blind Video Bit-Depth ExpansionabstractWith the rapid development of high-bit-depth display devices, bit-depth expansion (BDE) algorithms that extend low-bit-depth images to high-bit-depth images have received increasing attention. Due to the sensitivity of bit-depth distortions to tiny numerical changes in the least significant bits, the nuanced degradation differences in the training process may lead to varying degradation data distributions, causing the trained models to overfit specific types of degradations. This paper focuses on the problem of blind video BDE, proposing a degradation prediction and embedding framework, and designing a video BDE network based on a recurrent structure and dual-frame alignment fusion. Experimental results demonstrate that the proposed model can outperform some state-of-the-art (SOTA) models in terms of banding artifact removal and color correction, avoiding overfitting to specific degradations and obtaining better generalization ability across multiple datasets. https://github.com/duanpanjun/BVBDE Panjun Duan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
ACM Multimedia | 3 |
| 2024 | Regional Traditional Painting Generation Based on Controllable Disentanglement ModelabstractAutomatic generation of painting images is an interesting and difficult task, especially for regional traditional paintings with unique cultural styles while lacking large-scale training sets. In this paper, a hierarchical painting generation method is proposed, which can disentangle the generation of content and style. By mimicking the human painting process, the proposed method introduces multiple content blocks first and gradually generates image contents. In each block, a spatial self-modulation module is proposed to inject local details while preserving the global layout. After the preliminary generation of contents, a series of style blocks are presented to gradually adjust the artistic style. In the style block, an edge-oriented style-modulation module is proposed, which focuses on the lines and edges. In addition, edge adversarial training is used to further improve the quality of generated lines. To train and evaluate the proposed method, we construct datasets for five types of Chinese folk paintings. Experimental results demonstrate that the proposed method can generate high-quality and diverse painting images. More importantly, it can disentangle content and style sufficiently, so that the generation of specific contents or styles can be controlled freely. The datasets and source codes is available at https://github.com/Ritsu-mio/HPGN. Yang Zhao 0002, Huaen Li, Zhao Zhang 0001, Yuan Chen 0012, Qing Liu 0022 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | ADRNet: Affine and Deformable Registration Networks for Multimodal Remote Sensing ImagesabstractMulti-modal remote sensing images registration ensures the consistency of the spatial positions for different images. It can provide the accurate geographic information and supports the fusion of multi-source data for geospatial analyses and applications. Rigid registration method shows high performance in dealing with large-scale deformation, but it is difficult to achieve high-precision image registration. In contrast, non-rigid registration method is suitable for processing local differences, but cannot effectively deal with large-scale deformation differences. Therefore, the combination of rigid and non-rigid registration methods becomes a necessary strategy to address such issues. In this paper, we propose a novel ADRNet method for multi-modal remote sensing images registration. The proposed ADRNet method contains three main modules: affine registration module, deformable registration module, and spatial transformer module that integrates the affine and deformable transformation parameters to obtain the final aligned images. Meanwhile, we design a new feature enhancement module and an attention module with dilated convolutions which have different dilation rates, which are used to alleviate the limitations imposed by receptive fields in the convolution operation. Moreover, we propose a specific symmetric loss function to optimize the whole network from the perspective of inverse consistency. To assess the efficiency and performance of the network, we extend the experimental data, ranging from cross-modal images in a conventional viewpoint to cross-modal images in a remote sensing viewpoint. The experimental results show that our method exhibits excellent performance for the images with different viewpoints and deformation scales. The relevant code will be released at: https://github.com/Ahuer-Lei/ADRNet. Yun Xiao 0003, Yuan Chen 0012, Bo Jiang 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Cross-view Resolution and Frame Rate Joint Enhancement for Binocular VideoabstractWith the popular of stereo video and free-viewpoint video, binocular and multi-view video enhancement has attracted increasing attention. Current binocular video enhancement methods mainly focus on stereo super-resolution. In this paper, we tend to discuss a new binocular video resolution and frame-rate enhancement scenario to fully utilize the cross-view complementary information. Specifically, one view is captured with high resolution (HR) and low frame-rate (LFR), while the other viewpoint records low resolution (LR) and high frame-rate (HFR) video. Then, a binocular video joint enhancement network, which adopts dual-branch structure with cross-view guidance, is proposed to jointly reconstruct HR and HFR stereo videos. The proposed framework can reduce the capture, storage, compression, and transmission cost of normal HR and HFR stereo videos. Compared with single-view super-resolution and video frame interpolation techniques, the proposed method can recover more realistic HR details and intermediate motion by using cross-view reference. Experimental results on stereo video datasets demonstrate the effectiveness of the proposed joint resolution and frame-rate enhancement framework. Panda Pan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
ACM Multimedia | 3 |
| 2023 | Fast Blind Decontouring NetworkabstractContouring artifacts usually appear in large and smooth flat areas, which are caused by many widely used processes such as bit-depth expansion, compression, image sharpening and contrast enhancement. Unfortunately, recent decontouring methods were mainly designed for specific and non-blind degradations, which significantly reduces the generalization ability of these methods when applied to complex and various real-world false contours. Therefore, this paper explores the blind decontouring problem by proposing a blind decontouring network (BDCN). Instead of directly training a decontouring network with mixed degradations, the proposed model consists of two independent modules, i.e., a flat region detection module (FDM) and a decontouring module (DCM). The FDM is designed to extract flat region masks robust to various false contours, which can preserve texture details from global smoothing. Then, the task of DCM becomes simply smoothing different contouring artifacts. Both the FDM and DCM are designed with a lightweight architecture and reparameterization strategy. Experimental results on both synthetic and real-world contouring artifacts demonstrate the effectiveness and generalization of the proposed method. Yang Zhao 0002, Wei Jia 0001, Yuan Chen 0012, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Learning Deep Blind Quality Assessment for Cartoon ImagesabstractAlthough the cartoon industry has developed rapidly in recent years, few studies pay special attention to cartoon image quality assessment (IQA). Unfortunately, applying blind natural IQA algorithms directly to cartoons often leads to inconsistent results with subjective visual perception. Hence, this brief proposes a blind cartoon IQA method based on convolutional neural networks (CNNs). Note that training a robust CNN depends on manually labeled training sets. However, for a large number of cartoon images, it is very time-consuming and costly to manually generate enough mean opinion scores (MOSs). Therefore, this brief first proposes a full reference (FR) cartoon IQA metric based on cartoon-texture decomposition and then uses the estimated FR index to guide the no-reference IQA network. Moreover, in order to improve the robustness of the proposed network, a large-scale dataset is established in the training stage, and a stochastic degradation strategy is presented, which randomly implements different degradations with random parameters. Experimental results on both synthetic and real-world cartoon image datasets demonstrate the effectiveness and robustness of the proposed method. Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Cartoon Image Processing: A Survey
Yang Zhao 0002, Diya Ren, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
Int. J. Comput. Vis. | 3 |
| 2022 | Multiframe Joint Enhancement for Early Interlaced VideosabstractEarly interlaced videos usually contain multiple and interlacing and complex compression artifacts, which significantly reduce the visual quality. Although the high-definition reconstruction technology for early videos has made great progress in recent years, related research on deinterlacing is still lacking. Traditional methods mainly focus on simple interlacing mechanism, and cannot deal with the complex artifacts in real-world early videos. Recent interlaced video reconstruction deep deinterlacing models only focus on single frame, while neglecting important temporal information. Therefore, this paper proposes a multiframe deinterlacing network joint enhancement network for early interlaced videos that consists of three modules, i.e., spatial vertical interpolation module, temporal alignment and fusion module, and final refinement module. The proposed method can effectively remove the complex artifacts in early videos by using temporal redundancy of multi-fields. Experimental results demonstrate that the proposed method can recover high quality results for both synthetic dataset and real-world early interlaced videos. At the same time, the method also won the first place in the MSU Deinterlacer Benchmark. The code is available at: https://github.com/anymyb/MFDIN. Yang Zhao 0002, Yanbo Ma, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
IEEE Trans. Image Process. | 3 |
| 2021 | Lighter but Efficient Bit-Depth Expansion NetworkabstractWith the development of display technology, bit-depth expansion (BDE) has emerged as a basic process to display low-bit-depth image and video resources on high-bit-depth monitors. Most current BDE methods are based on traditional algorithms, and the few existing methods based on deep neural networks still suffer from loss of pixel-level details or from high computational cost. This paper proposes a lightweight but efficient BDE network that can effectively improve the capacity of shallow network by introducing a residual-block-in-residual-block structure. Furthermore, the proposed network adopts residual network architecture and dilated convolution to balance the preservation of pixel-level information and the expansion of the receptive field. Hence, the proposed method can also totally remove significant artifacts from very low-bit-depth images. Experimental results demonstrate that the proposed method can achieve performance comparable to or even better than that of some state-of-the-art methods while having much lighter architecture and fewer parameters. Yang Zhao 0002, Ronggang Wang, Yuan Chen 0012, Wei Jia 0001, Xiaoping Liu 0003, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Estimated Exposure Guided Reconstruction Model for Low-Light Image Enhancement
Xiaona Liu, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
PRCV (1) | 3 |
| 2020 | Adversarial-learning-based image-to-image transformation: A survey
Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003 |
Neurocomputing | 1 |
| 2020 | Blind Quality Assessment for Cartoon ImagesabstractCurrent blind image quality assessment (BIQA) algorithms are mainly designed for natural images. Unfortunately, cartoon and cartoon-like images are quite different from natural images. Hence, recent BIQA methods are not very robust to cartoon images. In this paper, we propose a specific BIQA algorithm designed for cartoon images, which consists of the following terms. First, a cartoon image is divided into edge areas and nonedge areas via a Tchebichef moment (TM)-based process. Second, a multiorder sharpness statistic term is used to measure the quality of the edges, and a sharpness statistic prior model of high-quality (HQ) cartoon images is built. Finally, a local encoding statistic term is adopted to describe the textural complexity in the nonedge areas, and a texture statistic prior model is also established. The experimental results on the cartoon image datasets demonstrate that the proposed method can accurately evaluate the visual quality of cartoon images and is more suitable for cartoon scenarios than some traditional BIQA algorithms. Yuan Chen 0012, Yang Zhao 0002, Shujie Li 0002, Wangmeng Zuo, Wei Jia 0001, Xiaoping Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |