EDBT 2026 Demo / reviewers in the wild / expert
Leyuan Liu 0001
dblp:76/8615-1
· DBLP profile ↗
24ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0002-8050-8677ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR-GS: Spatio-Temporal Geometry Alignment and Generative Refinement for Sparse-View 4D Gaussian SplattingabstractRecent 4D Gaussian Splatting (4DGS) methods for reconstructing dynamic scenes have achieved unified spatiotemporal modeling under densely captured multi-view inputs. Nevertheless, reconstruction from sparse-view video sequences remains highly challenging due to unreliable Gaussian initialization and insufficient photometric supervision. We present STAR-GS, a novel framework for sparse-view 4D reconstruction that integrates Spatio-Temporal Geometry Alignment and Generative View Refinement. To address unreliable Gaussian initialization, we develop the Geometry Alignment that combines feed-forward geometry prediction with cross-temporal joint camera optimization, resolving inter-frame similarity ambiguities and establishing a unified global Gaussian initialization. To compensate for insufficient supervision, we further incorporate the Generative View Refinement based on a reference-conditioned single-step diffusion model, which synthesizes high-fidelity novel views to provide dense and temporally consistent photometric guidance for 4DGS optimization. Extensive experiments demonstrate that STAR-GS significantly improves reconstruction quality under sparse-view settings. Ablation studies further validate the effectiveness and complementary contributions of the proposed components. Yunqi Gao, Zhanfeng Liao, Dongbo Zhou, Leyuan Liu 0001 |
ICMR | 4 |
| 2026 | GD-Head: Reconstructing High-quality 3D Avatar Head from a Single Image Using Geometry-guided Diffusion ModelsabstractExisting single-image 3D head reconstruction methods often fall short of the demanding requirements for high-quality digital avatars in immersive virtual reality applications. To bridge this gap, we propose GD-Head—a novel diffusion-based framework for high-fidelity and detail-preserving 3D head texture reconstruction from a single image. GD-Head operates in three coherent stages: first, it employs an advanced geometric model to apply weak perspective projection to the input image, through which realistic facial texture details are directly extracted and occluded regions are identified and masked in gray for subsequent inpainting. Second, the 3D head model is parameterized into 2D UV space via cylindrical UV mapping. Third, a geometry-guided diffusion model is utilized to realistically inpaint the missing texture regions. By reframing 3D texture reconstruction as a 2D image inpainting task, GD-Head consistently produces complete and seamless texture maps that accurately reflect the subject’s facial expressions and surface details, essential for believable avatar representation in VR environments. Extensive experiments on six benchmark datasets and an additional in-the-wild collection demonstrate that GD-Head achieves state-of-the-art performance in reconstruction fidelity and robustness. To support reproducibility and further research, both the code and pre-trained models will be made publicly available. Leyuan Liu 0001, Yufei Qian, Jingying Chen 0001 |
ICMR | 1 |
| 2025 | Biomarker Discovery for ASD via HMM-Based EEG Microstate AnalysisabstractDiscovering biomarkers for Autism Spectrum Disorder (ASD) is essential for elucidating its etiology, enabling early diagnosis, and refining treatment strategies. Electroencephalogram (EEG) microstates reflect the brain's overall dynamic changes, aiding in exploring differences in brain function patterns between ASD and Typically Developing (TD) groups. To this end, this study proposes an adaptive EEG microstate analysis approach based on Hidden Markov Models (HMMs) for the discovery of ASD biomarkers. Specifically, the proposed method, within the HMM framework, adaptively extracts millisecondscale transient brain microstate patterns that recur over time and models microstates using a multivariate Gaussian distribution rather than static topological structures. Resting-state EEG from 178 children aged 3 to 6 are used to validate the proposed approach. The analysis of the four microstates (#1, #2, #3, and #4) reveals significant differences between ASD and TD groups. Temporally, the ASD group shows difficulty in microstate transitions, primarily between microstates #2 and #3. In the frequency and spatial domains, TD individuals exhibit stronger brain region activation and interaction in microstates #2 and #4, whereas the ASD group shows reduced activity. Notably, during microstate #3, the ASD group demonstrates higher spectral power and channel coherence. Additionally, Ttests on intergroup feature differences and the results of the ASD discrimination task (accuracy: 88.89%) further confirm the potential of microstate features in assessing ASD tendencies. Overall, this study holds the potential to reveal novel insights into the neural mechanisms underlying ASD and identify valuable biomarkers for clinical assessment and diagnosis. Dan Chen 0001, Meiqi Zhou, Tengfei Gao, Jingying Chen 0001, Naiqian Mao, Leyuan Liu 0001 |
BIBM | 6 |
| 2025 | ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single ImageabstractWith 3D data rapidly emerging as an important form of multimedia information, 3D human mesh recovery technology has also advanced accordingly. However, current methods mainly focus on handling humans wearing tight clothing and perform poorly when estimating body shapes and poses under diverse clothing, especially loose garments. To this end, we make two key insights: (1) tailoring clothing to fit the human body can mitigate the adverse impact of clothing on 3D human mesh recovery, and (2) utilizing human visual information from large foundational models can enhance the generalization ability of the estimation. Based on these insights, we propose ClothHMR, to accurately recover 3D meshes of humans in diverse clothing. ClothHMR primarily consists of two modules: clothing tailoring (CT) and FHVM-based mesh recovering (MR). The CT module employs body semantic estimation and body edge prediction to tailor the clothing, ensuring it fits the body silhouette. The MR module optimizes the initial parameters of the 3D human mesh by continuously aligning the intermediate representations of the 3D mesh with those inferred from the foundational human visual model (FHVM). ClothHMR can accurately recover 3D meshes of humans wearing diverse clothing, precisely estimating their body shapes and poses. Experimental results demonstrate that ClothHMR significantly outperforms existing state-of-the-art methods across benchmark datasets and in-the-wild images. Additionally, a web application for online fashion and shopping powered by ClothHMR is developed, illustrating that ClothHMR can effectively serve real-world usage scenarios. The code and model for ClothHMR are available at: https://github.com/starVisionTeam/ClothHMR. Yunqi Gao, Leyuan Liu 0001, Yuhan Li 0009, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
ICMR | 2 |
| 2025 | HumanPrinter: Reconstructing 3D Human from a Single Image Like a 3D PrinterabstractHigh-fidelity 3D human reconstruction is essential for numerous applications. Existing reconstruction methods still suffer from several limitations. Implicit-function-based methods often produce artifacts, particularly when handling complex poses and loose-fitting clothing. Existing deformation-based methods require the entire body mesh to be input to the network for deformation, resulting in reconstructed results that are not ideal in detail. We propose HumanPrinter, a novel method for reconstructing high-fidelity 3D clothed human models from a single RGB image. Drawing inspiration from 3D printing, HumanPrinter reconstructs the human mesh layer by layer. HumanPrinter slices the coarse mesh resulting from deformation based on the estimated SMPL-X mesh into multiple vertically stacked polygons. The network then regresses vertex offsets from visual cues extracted from the input image to deform these polygons. Finally, these deformed polygons are stitched together and further refined to achieve a complete and detailed 3D human mesh. Polygon-based deformations associate the deformations of each vertex with its adjacent vertex so that the HumanPrinter produces fewer artifacts. By reducing the number of input polygons and increasing the number of deformable vertices, the layered reconstruction method can make the network more focused on local details. Through experiments on three datasets and visual results on in-the-wild data, we demonstrate that HumanPrinter performs competitive reconstruction quality compared to current state-of-the-art methods. Leyuan Liu 0001, Shen Chen 0005, Jingying Chen 0001 |
ACM Multimedia | 1 |
| 2025 | Beyond boundaries: Hierarchical-contrast unsupervised temporal action localization with high-coupling feature learning
Yuanyuan Liu 0004, Leyuan Liu 0001, Wujie Zhou, Chang Tang |
Pattern Recognit. | 5 |
| 2024 | VS: Reconstructing Clothed 3D Human from Single Image via Vertex ShiftabstractVarious applications require high-fidelity and artifact free 3D human reconstructions. However, current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper, we propose a two-stage deformation method named Vertex Shift (VS) for reconstructing clothed 3D humans from single images. Specifically, VS first stretches the estimated SMPL-X mesh into a coarse 3D human model using shift fields inferred from normal maps, then refines the coarse 3D human model into a detailed 3D human model via a graph convolutional network embedded with implicit-function-learned features. This “stretch-refine” strategy addresses large deformations required for reconstructing loose clothing and delicate deformations for recovering intricate and detailed surfaces, achieving high-fidelity reconstructions that faithfully convey the pose, clothing, and surface details from the input images. The graph convolutional network's ability to exploit neighborhood vertices coupled with the advantages inherited from the deformation methods ensure VS rarely produces artifacts like distortions and non-human shapes and never produces artifacts like holes, broken parts, and dismembered limbs. As a result, VS can reconstruct highfidelity and artifact-less clothed 3D humans from single images, even under scenarios of challenging poses and loose clothing. Experimental results on three benchmarks and two in-the-wild datasets demonstrate that VS significantly outperforms current state-of-the-art methods. The code and models of VS are available for research purposes at https://github.com/starVisionTeam/VS. Leyuan Liu 0001, Yuhan Li 0009, Yunqi Gao, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
CVPR | 1 |
| 2024 | An Avatar-Based Intervention System for Children with Autism Spectrum Disorder
Leyuan Liu 0001, Yuanjian You, Zhichen He, Jingying Chen 0001 |
PRCV (9) | 1 |
| 2024 | SeIF: Semantic-Constrained Deep Implicit Function for Single-Image 3D Head ReconstructionabstractVarious applications require realistic, artifact-free, and animatable 3D avatars. However, traditional 3D morphable models (3DMMs) produce animatable 3D heads but fail to capture accurate geometries and details, while existing deep implicit functions have been shown to achieve realistic reconstructions but suffer from artifacts and struggle to yield 3D heads that are easy to animate. To reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images, we leverage semantics to bridge the best properties of 3DMMs and deep implicit functions and propose SeIF—a semantic-constrained deep implicit function. First, SeIF derives fine-grained semantics from a standard 3DMM (e.g., FLAME) and samples a semantic code for each query point in the query space to provide a soft constraint to the deep implicit function. The reconstruction results show that this semantic constraint does not weaken the powerful representation ability of the deep implicit function while significantly suppressing artifacts. Second, SeIF predicts a more accurate semantic code for each query point and utilizes the semantic codes to uniformize the structure of reconstructed 3D head meshes with the standard 3DMM. Since our reconstructed 3D head meshes have the same structure as the 3DMM, 3DMM-based animation approaches can be easily transferred to animate our reconstructed 3D heads. As a result, SeIF can reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images of individuals with diverse ages, genders, races, and facial expressions. Quantitative and qualitative experimental results on seven datasets show that SeIF outperforms existing state-of-the-art methods by a large margin. The code and models of SeIF are available for research purposes athttps://github.com/starVisionTeam/SeIF. Leyuan Liu 0001, Jianchi Sun, Changxin Gao, Jingying Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Single-image clothed 3D human reconstruction guided by a well-aligned parametric body model
Leyuan Liu 0001, Yunqi Gao, Jianchi Sun, Jingying Chen 0001 |
Multim. Syst. | 1 |
| 2021 | HEI-Human: A Hybrid Explicit and Implicit Method for Single-View 3D Clothed Human Reconstruction
Leyuan Liu 0001, Jianchi Sun, Yunqi Gao, Jingying Chen 0001 |
PRCV (2) | 1 |
| 2020 | PlugNet: Degradation Aware Scene Text Recognition Supervised by a Pluggable Super-Resolution Unit
Yongqiang Mou, Jingying Chen 0001, Leyuan Liu 0001, Yaohong Huang |
ECCV (15) | 5 |
| 2019 | Progressive Pose Normalization Generative Adversarial Network for Frontal Face Synthesis and Face Recognition Under Large PoseabstractThis paper proposes a Progressive Pose-Normalization Generative Adversarial Network (PPN-GAN) for frontal face synthesis and face recognition. The key idea is to normalize a profile face progressively: starting from inferring an intermediate face that has a small view difference to the profile face, and then increasing the view difference step by step, until the frontal view of the profile face is recovered. In addition to the progressive strategy, an additional identity discriminator and identity-aware losses in both the image and feature spaces are also incorporated into the GAN for identity preserving. Experimental results show that our method not only produces compelling perceptual results but also outperforms the state-of-the-art methods on face recognition under large-pose. Leyuan Liu 0001, Jingying Chen 0001 |
ICIP | 1 |
| 2018 | Light YOLO for High-Speed Gesture RecognitionabstractThis paper proposes an efficient model named Light YOLO for hand gesture recognition on the embedded platforms. Light YOLO improves accuracy, speed, and model size, in three aspects. To deal with the small scale gestures in practical applications, we strengthen the YOLOv2 with a spatial refinement module to obtain fine-grained features. To accelerate the refined network, we propose a selective-dropout channel pruning approach to prune the redundancy convolution kernels in the network. Moreover, we introduce a dataset for hand gesture recognition in complex scenes. The experimental results on this dataset show that the proposed Light YOLO significantly improve the YOLOv2 network, i.e., accuracy from 96.80% to 98.06%, speed form 40PFS to 125FPS, and size form 250M to 4MB. Zihan Ni, Nong Sang, Changxin Gao, Leyuan Liu 0001 |
ICIP | 5 |
| 2018 | Semi-supervised Learning of Deep Difference Features for Facial Expression Recognition
Ruyi Xu, Jingying Chen 0001, Leyuan Liu 0001 |
PRCV (3) | 4 |
| 2018 | Deep peak-neutral difference feature for facial expression recognition
Jingying Chen 0001, Ruyi Xu, Leyuan Liu 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Book Page Identification Using Convolutional Neural Networks Trained by Task-Unrelated Dataset
Leyuan Liu 0001, Huabing Zhou, Jingying Chen 0001 |
ICIG (1) | 1 |
| 2017 | A low-cost real-time face tracking system for ITSs and SDASsabstractSummary It is important to track people's face efficiently and accurately in many Intelligent Transportation Systems (ITSs) and Safety Driving Assistant Systems (SDASs). This paper presents a high‐performance and low‐cost real‐time face tracking system, which runs on general onboard computer with very low CPU consumption. The proposed face tracking system is composed of four modules: the motion detector, face detector, face tracker, and face validator. The motion detector extracts motion areas by using a spatial‐temporal bi‐differential method with a very low computational cost. The face detector integrates motions into a cascade face detection framework to reject most of non‐face scanning‐windows to ensure efficient face localization. The face tracker fuses motion feature with color feature to alleviate the drifting problem during tracking. The face validator builds face appearance models online and identifies each specific tracked face to avoid confusion. Experimental results on three challenging video sequences show that the proposed face tracking system outperforms the state‐of‐the‐art face tracker and consumes only 5–13% CPU resources of a low‐spec onboard computer while processing in real time. Copyright © 2016 John Wiley & Sons, Ltd. Leyuan Liu 0001, Jingying Chen 0001, Changxin Gao, Nong Sang |
Softw. Pract. Exp. | 1 |
| 2016 | Temporally aligned pooling representation for video-based person re-identificationabstractThis paper proposes an effective Temporally Aligned Pooling Representation (TAPR) for video-based person re-identification. To extract the motion information from a sequence, we propose to track the superpixels of the lowest portions of human. To perform temporal alignment of videos, we propose to select the “best” walking cycle from the noisy motion information according to the intrinsic periodicity property of walking persons, that is fitted sinusoid in our implementation. To describe the video data in the selected walking cycle, we first divide the cycle into several segments according to the sinusoid, and then describe each segment by temporally aligned pooling. Extensive experimental results on the public datasets demonstrate the effectiveness of the proposed method compared with the state-of-the-art approaches. Changxin Gao, Jin Wang 0019, Leyuan Liu 0001, Jin-Gang Yu, Nong Sang |
ICIP | 3 |
| 2016 | Multi-person Visual Focus of Attention from Head Pose on a Natural Classroom
Yuanyuan Liu 0004, Leyuan Liu 0001, Jingying Chen 0001, Chunyan Su, Kun Zhang 0031 |
ICPRAM | 2 |
| 2016 | Robust head pose estimation using Dirichlet-tree distribution enhanced random forests
Yuanyuan Liu 0004, Jingying Chen 0001, Zhiming Su, Zhenzhen Luo, Nan Luo, Leyuan Liu 0001, Kun Zhang 0031 |
Neurocomputing | 6 |
| 2014 | Dirichlet-tree Distribution Enhanced Random Forests for Head Pose EstimationabstractHead pose estimation is important in human-machine interfaces. However, illumination variation, occlusion and low image resolution make the estimation task difficult. Hence, a Dirichlet-tree distribution enhanced Random Forests approach (D-RF) is proposed in this paper to estimate head pose efficiently and robustly under various conditions. First, PCA based sub-features space from Gabor features and histogram distributions of the facial patches are extracted to eliminate the influence of occlusion and noise. Then, the D-RF is proposed to estimate the head pose in a coarse-to-fine way. In order to improve the discrimination capability of the approach, an adaptive Gaussian mixture model is introduced in the tree distribution. The proposed method has been evaluated with different data sets spanning from -90° to 90° in vertical and horizontal directions under various conditions. The experimental results demonstrate the approachâs robustness and efficiency. Yuanyuan Liu 0004, Jingying Chen 0001, Leyuan Liu 0001, Yujiao Gong, Nan Luo |
ICPRAM | 3 |
| 2011 | Metrics for Objective Evaluation of Background Subtraction AlgorithmsabstractAlthough a large number of background subtraction (BS) algorithms have been proposed, relevant objective metrics for evaluating these algorithms are still lacking. In this paper, empirical discrepancy metrics, which quantify the spatial accuracy and temporal stability of estimated masks by taking into account the potential inaccuracy of reference masks, the location of the pixel errors relative to the border of reference masks as well as the type of errors, are presented for evaluating the performance of BS algorithms. To validate the proposed metrics, they are applied to tune the optimal parameters of LBP-based background subtraction algorithm, and the experimental results confirm the efficiency of them. Leyuan Liu 0001, Nong Sang |
ICIG | 1 |
| 2010 | Saliency Based on Multi-scale Ratio of DissimilarityabstractRecently, many vision applications tend to utilize saliency maps derived from input images to guide them to focus on processing salient regions in images. In this paper, we propose a simple and effective method to quantify the saliency for each pixel in images. Specially, we define the saliency for a pixel in a ratio form, where the numerator measures the number of dissimilar pixels in its center-surround and the denominator measures the total number of pixels in its center-surround. The final saliency is obtained by combining these ratios of dissimilarity over multiple scales. For images, the saliency map generated by our method not only has a high quality in resolution also looks more reasonable. Finally, we apply our saliency map to extract the salient regions in images, and compare the performance with some state-of-the-art methods over an established ground-truth which contains 1000 images. Rui Huang 0001, Nong Sang, Leyuan Liu 0001, Qiling Tang |
ICPR | 3 |