VLDB 2026 Research / reviewers in the wild / expert
Xianhang Cheng
dblp:266/4256
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0006-9495-7330ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Image and video processing · 50% Visual content generation and editing · 25% Geometric modeling and processing · 25% | |
| Artificial intelligence
2 papers |
3D vision · 31% Segmentation and scene understanding · 30% Learning paradigms · 30% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › video frame interpolation › interpolation
kernel interpolation |
1.0 | 2 | 2022 | Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Video Frame Interpolation via Deformable Separable Convolution · AAAI 2020 |
Image and video processing
video frame interpolation |
1.0 | 2 | 2022 | Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Video Frame Interpolation via Deformable Separable Convolution · AAAI 2020 |
Geometric modeling and processing
shape correspondence |
1.0 | 1 | 2026 | Learned Universal Interoperable Virtual Try-ON · ACM Trans. Graph. 2026 |
Visual content generation and editing
virtual try-on |
1.0 | 1 | 2026 | Learned Universal Interoperable Virtual Try-ON · ACM Trans. Graph. 2026 |
Machine learning › Learning paradigms › class imbalance
long-tailed learning |
0.6 | 1 | 2022 | Predicate Correlation Learning for Scene Graph Generation · IEEE Trans. Image Process. 2022 |
Computer vision › Segmentation and scene understanding
scene graph generation |
0.6 | 1 | 2022 | Predicate Correlation Learning for Scene Graph Generation · IEEE Trans. Image Process. 2022 |
Computer vision › 3D vision
3d human reconstruction |
0.3 | 1 | 2026 | Learned Universal Interoperable Virtual Try-ON · ACM Trans. Graph. 2026 |
Computer vision › 3D vision › 3d human reconstruction
SMPL |
0.3 | 1 | 2026 | Learned Universal Interoperable Virtual Try-ON · ACM Trans. Graph. 2026 |
Machine learning › Trustworthy machine learning › interpretability
visual explanation |
0.2 | 1 | 2022 | Predicate Correlation Learning for Scene Graph Generation · IEEE Trans. Image Process. 2022 |
Methods — techniques the papers use, named apart from their topics
simulator-driven fitting · 2.0foundation model features · 2.0diffusion model · 2.0deformable separable convolution · 1.0predicate correlation matrix · 0.6predicate correlation learning · 0.6loss function design · 0.6coord-conv · 0.6adaptive kernel estimation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learned Universal Interoperable Virtual Try-ONabstractTo enable large-scale reuse of real-world 3D assets-where garments and characters rarely share skeletons, templates, or dense correspondences-we present a fully automated virtual try-on system that dresses complex, multi-layer garments onto diverse, arbitrarily posed humanoids. Our key idea is to use SMPL as an intermediate proxy and decompose clothing-to-body transfer into two correspondence tasks with distinct challenges: (1) clothing-to-SMPL (partial-to-complete alignment) and (2) body-to-SMPL (large pose/shape variation and stylization). We address clothing-to-SMPL using a geometry-driven correspondence model, and introduce a diffusion-based body-to-SMPL correspondence approach that leverages multi-view consistent appearance features together with a pretrained 2D foundation model. Using these correspondences, we register SMPL/SMPL+D (Displacement) to the garment and target body and then perform simulator-driven fitting by transferring the garment along a smooth SMPL→SMPL+D transition, producing physically plausible draping on the target. Our system handles complex garment topology (including non-manifold meshes) and generalizes to a wide range of humanoid characters (e.g., humans, robots, cartoons, and creatures) while remaining computationally practical. Upon draping, our system also supports fast customization of clothing size. We show that our system can produce high-quality 3D clothing fittings without any human labor, even when 2D clothing sewing patterns are not available. Our project page is: https://cao-cong0.github.io/LUIVITON-Learned-Universal-Interoperable-VIrtual-Try-ON/. Xianhang Cheng, Yujian Zheng, Zhenhui Lin, Meriem Chkir, Hao Li 0015 |
ACM Trans. Graph. | 2 |
| 2024 | oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning CompilationabstractWith the rapid development of deep learning models and hardware support for dense computing, the deep learning (DL) workload characteristics changed significantly from a few hot spots on compute-intensive operations to a broad range of operations scattered across the models. Accelerating a few compute-intensive operations using the expert-tuned implementation of primitives doesn't fully exploit the performance potential of AI hardware. Various efforts have been made to compile a full deep neural network (DNN) graph. One of the biggest challenges is to achieve high-performance tensor compilation by generating expert-level performance code for the dense compute-intensive operations and applying compilation optimization at the scope of DNN computation graph across multiple compute-intensive operations. We present oneDNN Graph Compiler, a tensor compiler that employs a hybrid approach of using techniques from both compiler optimization and expert-tuned kernels for high-performance code generation of the deep neural network graph. oneDNN Graph Compiler addresses unique optimization challenges in the deep learning domain, such as low-precision computation, aggressive fusion of graph operations, optimization for static tensor shapes and memory layout, constant weight optimization, and memory buffer reuse. Experimental results demonstrate significant performance gains over existing tensor compiler and primitives library for performance-critical DNN computation graphs and end-to-end models on Intel® Xeon® Scalable Processors. Zhennan Qin, Yijie Mei, Jingze Cui, Yunfei Song, Ciyong Chen, Longsheng Du, Xianhang Cheng, Baihui Jin, Jason Ye, Eric Lin, Dan Lavery |
CGO | 9 |
| 2022 | Multiple Video Frame Interpolation via Enhanced Deformable Separable ConvolutionabstractGenerating non-existing frames from a consecutive video sequence has been an interesting and challenging problem in the video processing field. Typical kernel-based interpolation methods predict pixels with a single convolution process that convolves source frames with spatially adaptive local kernels, which circumvents the time-consuming, explicit motion estimation in the form of optical flow. However, when scene motion is larger than the pre-defined kernel size, these methods are prone to yield less plausible results. In addition, they cannot directly generate a frame at an arbitrary temporal position because the learned kernels are tied to the midpoint in time between the input frames. In this paper, we try to solve these problems and propose a novel non-flow kernel-based approach that we refer to as enhanced deformable separable convolution (EDSC) to estimate not only adaptive kernels, but also offsets, masks and biases to make the network obtain information from non-local neighborhood. During the learning process, different intermediate time step can be involved as a control variable by means of an extension of coord-conv trick, allowing the estimated components to vary with different input temporal information. This makes our method capable to produce multiple in-between frames. Furthermore, we investigate the relationships between our method and other typical kernel- and flow-based methods. Experimental results show that our method performs favorably against the state-of-the-art methods across a broad range of datasets. Code will be publicly available on URL: https://github.com/Xianhang/EDSC-pytorch. Xianhang Cheng, Zhenzhong Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Predicate Correlation Learning for Scene Graph GenerationabstractFor a typical Scene Graph Generation (SGG) method in image understanding, there usually exists a large gap in the performance of the predicates’ head classes and tail classes. This phenomenon is mainly caused by the semantic overlap between different predicates as well as the long-tailed data distribution. In this paper, a Predicate Correlation Learning (PCL) method for SGG is proposed to address the above problems by taking the correlation between predicates into consideration. To measure the semantic overlap between highly correlated predicate classes, a Predicate Correlation Matrix (PCM) is defined to quantify the relationship between predicate pairs, which is dynamically updated to remove the matrix’s long-tailed bias. In addition, PCM is integrated into a predicate correlation loss function (LPC) to reduce discouraging gradients of unannotated classes. The proposed method is evaluated on several benchmarks, where the performance of the tail classes is significantly improved when built on existing methods. Leitian Tao, Li Mi, Nannan Li 0004, Xianhang Cheng, Yaosi Hu, Zhenzhong Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Video Frame Interpolation via Deformable Separable ConvolutionabstractLearning to synthesize non-existing frames from the original consecutive video frames is a challenging task. Recent kernel-based interpolation methods predict pixels with a single convolution process to replace the dependency of optical flow. However, when scene motion is larger than the pre-defined kernel size, these methods yield poor results even though they take thousands of neighboring pixels into account. To solve this problem in this paper, we propose to use deformable separable convolution (DSepConv) to adaptively estimate kernels, offsets and masks to allow the network to obtain information with much fewer but more relevant pixels. In addition, we show that the kernel-based methods and conventional flow-based methods are specific instances of the proposed DSepConv. Experimental results demonstrate that our method significantly outperforms the other kernel-based interpolation methods and shows strong performance on par or even better than the state-of-the-art algorithms both qualitatively and quantitatively. Xianhang Cheng |
AAAI | 1 |
| 2020 | A Multi-Scale Position Feature Transform Network for Video Frame InterpolationabstractGiven a video sequence, video frame interpolation aims to synthesize an in-between frame of two consecutive frames. In this paper, we propose a multi-scale position feature transform (MS-PFT) network for video frame interpolation where two parallel prediction networks and one optimization network are designed to predict the features of target frame and generate the final interpolation result, respectively. To increase the fidelity of the synthesised frames, we propose to apply a position feature transform (PFT) layer in the residual blocks of the prediction networks to estimate scaling factors which help evaluate different degrees of the importance of deep features around a target pixel. A PFT layer utilizes optical flow to extract and generate position features and then adjusts the learning process of our model. We further extend our model into a multi-scale structure in which each scale of the network shares the same parameters to maximise the efficiency of our network with model size unchanged. The experiments show that our method can handle the challenging scenarios like occlusion and large motion effectively and the proposed method outperforms those state-of-the-art approaches on different datasets. Xianhang Cheng, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |