EDBT 2026 Demo / reviewers in the wild / expert
Yuru Pei
dblp:93/6606
· DBLP profile ↗
46ranked-venue papers
17as first author
17since 2021 · last 2025
0000-0001-8520-3509ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 15 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCCS: Deep Neural Spectral Clustering for Self-Supervised Subcellular Structure SegmentationabstractSubcellular structure segmentation is a fundamental task in biological imaging. Existing self-supervised representation learning combined with classical k-means clustering achieved unsupervised image segmentation, but it was constrained by time-consuming test-time pixel-wise feature extraction and clustering synchronization. This study introduces SCCS, a lightweight graph neural network-based spectral clustering framework for end-to-end subcellular structure segmentation upon superpixel graphs, greatly relieving the computational complexity in test-time numerical spectral clustering and inter-graph label inconsistency. Specifically, SCCS exploits the self-supervised masked autoencoder for representation learning and the construction of superpixel graphs (spG). Unlike per-graph scalar affinity-based spectral clustering, the proposed SCCS parameterizes the mapping from learned deep spG representations to coordinates in the spectral embedding space and the clustering assignments. The SCCS is optimized under unsupervised eigendecomposition and incremental clustering criteria, which synchronize the intra- and inter-graph spectral clustering. The proposed approach is evaluated on a publicly available volumetric electron microscopy dataset. Experiments demonstrate the effectiveness and performance gains of the proposed SCCS over the state-of-the-art in discovering a variety of subcellular structures. Jimao Jiang, Diya Sun, Tianbing Wang, Yuru Pei |
AAAI | 4 |
| 2025 | Foundation Model-Guided Pseudo-Labeling for Computed Tomography-Based Segmentation of Masticatory MusclesabstractSegmenting masticatory muscles from CT scans remains a challenging task due to intertwined tissues and image artifacts. While large vision foundation models (FMs) have achieved notable success in medical image segmentation, they typically depend on extensive and high-quality annotated datasets for fine-tuning or adapter training. In this paper, we introduce a novel FM-based pseudo-labeling framework for segmenting masticatory muscles from CT scans without the need for manual pixel-level annotations. Our method employs an adaptive pseudolabeling strategy that integrates FM-generated labels with selftraining. By leveraging the robustness of self-training to noisy data, the framework improves the reliability of pseudo-labels and helps differentiate closely intertwined muscle structures. To further enhance segmentation performance, we incorporate a structure-aware enhancement module that combines contrastive learning and the consistency regularization of semantic similarity relationships. This reinforces patch-level feature embeddings and improves semantic correlation consistency in label inference. Our approach supports simultaneous segmentation of multi-category muscles and achieves reliable contour delineation from adjacent tissues. Experimental results demonstrate that our approach outperforms compared state-of-the-art methods in masticatory muscle segmentation. The source code is available at http-s://anonymous.4open.science/r/pl_fuse_bba-895C/. Tianlin Li, Yuru Pei |
BIBM | 2 |
| 2025 | Neural Deformation Prior and Spatial Correlation Regularization for One-Shot Craniofacial Landmark LocalizationabstractPseudo-labeling plays a critical role in semisupervised learning (SSL). Existing one-shot landmark localization approaches relied on unsupervised registration-based label transfer or similarity maps for pseudo-labeling, which has been combined with consistency regularization for detector training. However, unsupervised registration can misinterpret semantic correspondences, leading to noisy and unreliable pseudo-labels. Such pseudo labels mislead SSL and result in incorrect consensus in landmark localization. In this paper, we introduce a novel neural deformation prior-guided pseudo-labeling for one-shot craniofacial landmark localization. Specifically, we introduce a self-supervised volumetric registration model for attribute transfer and pseudo labeling, guided by a prior deformation field and sparse matched landmarks with confidence scoring. Moreover, we present learnable spatial correlation regularization to mitigate landmark perturbations during co-teaching of landmark detectors. By leveraging neural deformation priors and spatial correlation regularization, the proposed approach improves pseudo-labeling and enhances landmark-wise interdependencies, even in regions affected by image artifacts. Extensive experiments on clinical and public CT image datasets demonstrate our method achieves craniofacial landmark localization with performance gains over state-of-theart landmark detection methods. The source code is available at https://anonymous.4open.science/r/NPSCT-D3C1/. Kaichen Nie, Tianmin Xu, Yuru Pei |
BIBM | 3 |
| 2025 | Foundation Model-Based Deformable Registration of Multi-Modal Remote Sensing ImagesabstractDeformable image registration and correspondence play a critical role in multi-modal remote sensing (MRS) image analysis. However, challenges such as domain gaps and heterogeneous noise in MRS images often hinder performance. Existing methods typically address these issues through carefully designed strategies for image style separation and task-specific feature extraction. In this paper, we introduce an unsupervised domain-adaptive registration framework for establishing semantic correspondences in MRS images. Our approach leverages frozen pre-trained foundation models, DINOv2 and Stable Diffusion, to extract domain-adaptable features, effectively addressing sensor-specific radiometric and geometric distortions. To favor the semantically rich and domain-agnostic feature channels, we propose a learnable Channel Attention Adapter (CAA) that captures channel-wise dependencies. Additionally, we introduce a neural displacement field (DF) decoder, which eliminates the need for computationally expensive nearest neighbor searching or iterative optimization. We conduct extensive experiments to demonstrate the effectiveness of the proposed CAA-based feature enhancement and the neural DF decoder in deformable registration of MRS images. Our method outperforms state-of-the-art approaches on publicly available MRS image datasets. The code is available at https://github.com/orgxicv/CA-DF. Linsi Wu, Kaichen Nie, Xuefei Lv, Yuru Pei |
ICIP | 6 |
| 2025 | DiffStain: Conditioned Diffusion-Based Semantic Virtual Staining with Mask Guidance
Yikai Han, Jimao Jiang, Yuru Pei |
MICCAI (13) | 3 |
| 2025 | CS2C: Collaborative Spatial and Spectral Neural Clustering for Organelle Segmentation from Volumetric Electron Microscopy
Jimao Jiang, Yuru Pei |
MICCAI (3) | 2 |
| 2025 | Toward Semantically-Consistent Deformable 2D-3D Registration for 3D Craniofacial Structure Estimation From a Single-View Lateral Cephalometric RadiographabstractThe deep neural networks combined with the statistical shape model have enabled efficient deformable 2D-3D registration and recovery of 3D anatomical structures from a single radiograph. However, the recovered volumetric image tends to lack the volumetric fidelity of fine-grained anatomical structures and explicit consideration of cross-dimensional semantic correspondence. In this paper, we introduce a simple but effective solution for semantically-consistent deformable 2D-3D registration and detailed volumetric image recovery by inferring a voxel-wise registration field between the cone-beam computed tomography and a single lateral cephalometric radiograph (LC). The key idea is to refine the initial statistical model-based registration field with craniofacial structural details and semantic consistency from the LC. Specifically, our framework employs a self-supervised scheme to learn a voxel-level refiner of registration fields to provide fine-grained craniofacial structural details and volumetric fidelity. We also present a weakly supervised semantic consistency measure for semantic correspondence, relieving the requirements of volumetric image collections and annotations. Experiments showcase that our method achieves deformable 2D-3D registration with performance gains over state-of-the-art registration and radiograph-based volumetric reconstruction methods. The source code is available at https://github.com/Jyk-122/SC-DREG. Yikun Jiang, Yuru Pei, Tianmin Xu, Xiaoru Yuan, Hongbin Zha |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Generalizable Structure-Aware INF: Biplanar-View CT Reconstruction via Disentangled Implicit Neural Field
Bei Huang, Yuru Pei |
ACCV (1) | 2 |
| 2024 | Stochastic Anomaly Simulation for Maxilla Completion from Cone-Beam Computed Tomography
Yixiao Guo, Yuru Pei, Zhi-bo Zhou, Tianmin Xu, Hongbin Zha |
MICCAI (8) | 2 |
| 2023 | Bi-Graph Reasoning for Masticatory Muscle Segmentation From Cone-Beam Computed TomographyabstractAutomated segmentation of masticatory muscles is a challenging task considering ambiguous soft tissue attachments and image artifacts of low-radiation cone-beam computed tomography (CBCT) images. In this paper, we propose a bi-graph reasoning model (BGR) for the simultaneous detection and segmentation of multi-category masticatory muscles from CBCTs. The BGR exploits the local and long-range interdependencies of regions of interest and category-specific prior knowledge of masticatory muscles by reasoning on the category graph and the region graph. The category graph of the learnable muscle prior knowledge handles high-level dependencies of muscle categories, enhancing the feature representation with noise-agnostic category knowledge. The region graph models both local and global dependencies of the candidate muscle regions of interest. The proposed BGR accommodates the high-level dependencies and enhances the region features in the presence of entangled soft tissue and image artifacts. We evaluated the proposed approach by segmenting masticatory muscles on clinically acquired CBCTs. Extensive experimental results show that the BGR effectively segments masticatory muscles with state-of-the-art accuracy. Yicheng Zhong, Yuru Pei, Kaichen Nie, Yungeng Zhang, Tianmin Xu, Hongbin Zha |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Deep Supervoxel Mapping Learning for Dense Correspondence of Cone-Beam Computed Tomography
Kaichen Nie, Yuru Pei, Diya Sun, Tianmin Xu |
PRCV (2) | 2 |
| 2022 | Unsupervised random forest for affinity estimationabstractThis paper presents an unsupervised clustering random-forest-based metric for affinity estimation in large and high-dimensional data. The criterion used for node splitting during forest construction can handle rank-deficiency when measuring cluster compactness. The binary forest-based metric is extended to continuous metrics by exploiting both the common traversal path and the smallest shared parent node. The proposed forest-based metric efficiently estimates affinity by passing down data pairs in the forest using a limited number of decision trees. A pseudo-leaf-splitting (PLS) algorithm is introduced to account for spatial relationships, which regularizes affinity measures and overcomes inconsistent leaf assign-ments. The random-forest-based metric with PLS facilitates the establishment of consistent and point-wise correspondences. The proposed method has been applied to automatic phrase recognition using color and depth videos and point-wise correspondence. Extensive experiments demonstrate the effectiveness of the proposed method in affinity estimation in a comparison with the state-of-the-art. Yunai Yi, Diya Sun, Peixin Li, Tae-Kyun Kim 0001, Tianmin Xu, Yuru Pei |
Comput. Vis. Media | 6 |
| 2022 | Dense correspondence of deformable volumetric images via deep spectral embedding and descriptor learning
Diya Sun, Yuru Pei, Yungeng Zhang, Tianmin Xu, Tianbing Wang, Hongbin Zha |
Medical Image Anal. | 2 |
| 2022 | Deep Volumetric Descriptor Learning for Dense Correspondence of Cone-Beam Computed Tomography via Spectral Maps
Diya Sun, Yungeng Zhang, Yuru Pei, Peixin Li, Kaichen Nie, Tianmin Xu, Tianbing Wang, Hongbin Zha |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Spectral Embedding Approximation and Descriptor Learning for Craniofacial Volumetric Image Correspondence
Diya Sun, Yungeng Zhang, Yuru Pei, Tianmin Xu, Hongbin Zha |
MICCAI (4) | 3 |
| 2021 | Learning Dual Transformer Network for Diffeomorphic Registration
Yungeng Zhang, Yuru Pei, Hongbin Zha |
MICCAI (4) | 2 |
| 2021 | Robust 3D face reconstruction from single noisy depth image through semantic consistencyabstractAbstract This paper addresses the 3D face reconstruction and semantic annotation from a single‐view noisy depth image. A deep neural network‐based coarse‐to‐fine framework is presented to take advantage of 3D morphable model (3DMM) regression and per‐vertex geometry refinement. The low‐dimensional subspace coefficients of the 3DMM initialize the global facial geometry, being prone to be over‐smooth because of the low‐pass characteristics of the shape subspace. The proposed geometry refinement subnetwork predicts per‐vertex displacements to enrich local details, which is learned from unlabelled noisy depth images based on the registration‐like loss. In order to guarantee the semantic correspondence between the resultant 3D face and the depth image, a semantic consistency constraint is introduced to adapt an annotation model learned from the synthetic data to real noisy depth images. The resultant depth annotations are required to be consistent with the label propagation from the coarse and refined parametric 3D faces. The proposed coarse‐to‐fine reconstruction scheme and the semantic consistency constraint are evaluated on the depth‐based 3D face reconstruction and semantic annotation. The series of experiments demonstrate that the proposed approach achieves the performance improvements over compared methods regarding 3D face reconstruction and depth image annotation. Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Hongbin Zha |
IET Comput. Vis. | 2 |
| 2020 | Fully Convolutional Network for Consistent Voxel-Wise CorrespondenceabstractIn this paper, we propose a fully convolutional network-based dense map from voxels to invertible pair of displacement vector fields regarding a template grid for the consistent voxel-wise correspondence. We parameterize the volumetric mapping using a convolutional network and train it in an unsupervised way by leveraging the spatial transformer to minimize the gap between the warped volumetric image and the template grid. Instead of learning the unidirectional map, we learn the nonlinear mapping functions for both forward and backward transformations. We introduce the combinational inverse constraints for the volumetric one-to-one maps, where the pairwise and triple constraints are utilized to learn the cycle-consistent correspondence maps between volumes. Experiments on both synthetic and clinically captured volumetric cone-beam CT (CBCT) images show that the proposed framework is effective and competitive against state-of-the-art deformable registration techniques. Yungeng Zhang, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha |
AAAI | 2 |
| 2020 | An Unsupervised Approach for 3D Face Reconstruction from a Single Depth Image
Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha |
CGI | 2 |
| 2020 | Face Denoising and 3D Reconstruction from A Single Depth ImageabstractThe reconstruction of 3D face shapes and expressions from a single depth image obtained by a consumer depth camera is a challenging issue considering device-specific noise, the data missing, and the lack of textual constraints. In order to relieve the computationally-intensive nonlinear optimization of traditional template-fitting-based methods, we aim to build an end-to-end regression framework between a depth image and a 3D face encoded by the identity, the expression, and the pose parameters. Concerning the lack of paired depth images and 3D faces, we utilize the unsupervised CycleGAN-based network to adapt the regression model learned from the synthetic data to the real-captured noisy depth images. Instead of separate depth image denoising and 3D face inference, we present a task-specific coupled loss for end-to-end 3D face estimation. We propose a three-tier constraint for the shape consistency in the joint embedding, the depth image, and the surface space to avoid shape distortions in the unsupervised domain adaptation network. We report promising qualitative results for the task of the face denoising and 3D face reconstruction from a single depth image. Yicheng Zhong, Yuru Pei, Peixin Li, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha |
FG | 2 |
| 2020 | Automatic Tooth Segmentation and Dense Correspondence of 3D Dental Model
Diya Sun, Yuru Pei, Peixin Li, Guangying Song, Yuke Guo, Hongbin Zha, Tianmin Xu |
MICCAI (4) | 2 |
| 2020 | Automatic Tooth Segmentation and 3D Reconstruction from Panoramic and Lateral Radiographs
Mochen Yu, Yuke Guo, Diya Sun, Yuru Pei, Tianmin Xu |
PRCV (1) | 4 |
| 2018 | Dense Correspondence of Cone-Beam Computed Tomography Images Using Oblique Clustering Forest
Diya Sun, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha |
BMVC | 2 |
| 2018 | Consistent Correspondence of Cone-Beam CT Images Using Volume Functional Maps
Yungeng Zhang, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha |
MICCAI (1) | 2 |
| 2018 | Incremental Feature Forest for Real-Time SLAM on Mobile Devices
Yuke Guo, Yuru Pei |
PRCV (1) | 2 |
| 2018 | Spatially Consistent Supervoxel Correspondences of Cone-Beam Computed Tomography ImagesabstractEstablishing dense correspondences of cone-beam computed tomography (CBCT) images is a crucial step for the attribute transfer and morphological variation assessment in clinical orthodontics. In this paper, a novel method, unsupervised spatially consistent clustering forest, is proposed to tackle the challenges for automatic supervoxel-wise correspondences of CBCT images. A complexity analysis of the proposed method with respect to the clustering hypotheses is provided with a data-dependent learning guarantee. The learning bound considers both the sequential tree traversals determined by questions stored in branch nodes and the clustering compactness of leaf nodes. A novel tree-pruning algorithm, guided by the learning bound, is also proposed to remove locally inconsistent leaf nodes. The resulting forest yields spatially consistent affinity estimations, thanks to the pruning penalizing trees with inconsistent leaf assignments and the combinational contextual feature channels used to learn the forest. A forest-based metric is utilized to derive the pairwise affinities and dense correspondences of CBCT images. The proposed method has been applied to the label propagation of clinically captured CBCT images. In the experiments, the method outperforms variants of both supervised and unsupervised forest-based methods and state-of-the-art label-propagation methods, achieving the mean dice similarity coefficients of 0.92, 0.89, 0.94, and 0.93 for the mandible, the maxilla, the zygoma arch, and the teeth data, respectively. Yuru Pei, Yunai Yi, Gengyu Ma, Tae-Kyun Kim 0001, Yuke Guo, Tianmin Xu, Hongbin Zha |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Mixed Metric Random Forest for Dense Correspondence of Cone-Beam Computed Tomography Images
Yuru Pei, Yunai Yi, Gengyu Ma, Yuke Guo, Tianmin Xu, Hongbin Zha |
MICCAI (1) | 1 |
| 2016 | Volumetric reconstruction of craniofacial structures from 2D lateral cephalograms by regression forestabstractThe 3D reconstruction is an essential step to measure the craniofacial morphological changes from the historical growth database with only 2D cephalograms. In this paper, we propose a novel regression-forest-based method to estimate the volumetric intensity images from a lateral cephalogram. The regression forest can produce a prediction of the volumetric craniofacial structure as a mixture of Gaussian by weighted aggregating the distributions from trees. The dense anatomical structure can be reconstructed with no time-consuming digitally-reconstructed-radiographs (DRR) in the online testing process. The experiments demonstrate the proposed method can reconstruct volumetric intensity images from the lateral cephalograms effectively. Yuru Pei, Fanfan Dai, Tianmin Xu, Hongbin Zha, Gengyu Ma |
ICIP | 1 |
| 2016 | Anatomical structure similarity estimation by random forestabstractThe morphological similarity of anatomical structures is essential to the study of the species evolution. In this paper, we investigate the unsupervised shape similarity analysis by a random-forest-based metric. The dense continuous deformation fields are employed as the shape descriptors. The forest is built when given the unlabeled deformation fields, where the leaves can be seen as an optimal clustering of the data set. The salient region is defined based on the dominant feature channels determined by the forests. The pairwise shape distance is computed efficiently with just binary comparisons stored in tree branches. We have applied our method to several skeletal data sets, including the skulls, teeth, radii, and metatarsals. Our experiments demonstrate the proposed method can handle the taxonomic classification effectively. Yuru Pei, Lei Kou, Hongbin Zha |
ICIP | 1 |
| 2016 | Fast 3D hand estimation for mobile interactionsabstractThe ubiquitous hand gesture plays an important role in the natural human machine interaction (HMI). Recently, the consumer color and depth cameras have been used to estimate hand shapes and postures for the mid-air HMI. Under the observation that 3D hand contours possess much information of hand postures, we estimate 3D hand contours from infrared images with a limited computation complexity for the HMI on mobile devices. A variant of the dynamic programming (vDP) algorithm is proposed to handle complex self-occlusions in 3D hand estimations, where a set of heuristic rules are introduced to avoid finger missing. Furthermore, the constraints are used to reduce the searching space in contour alignments. Given 3D hand contours, a set of hand gestures, including touching, swiping, and pinching, can be applied to mid-air interactions. The proposed method is much faster than the traditional depth estimation of the whole hand, and can achieve up to 500 Hz on PC, and 100 Hz on mobile devices. Yuru Pei, Gengyu Ma |
ICPR | 1 |
| 2015 | Multi-modal Brain Image Registration Based on Subset Definition and Manifold-to-Manifold DistanceabstractImage registration is an important procedure in multi-modal brain image processing. The main challenge is the variations of intensity distributions in different image modalities. The efficient SSD based method cannot handle this kind of variations. And other approaches based on modality independent descriptors and metrics are usually time-consuming. In this article, we propose a novel similarity metric based on manifold-to-manifold distance imposed on the subset of original images. We define a subset for a compact representation of the original image. Manifold learning technique is employed to reveal the intrinsic structure of the sampled data. Instead of comparing the images in the original feature space, we use the manifold-to-manifold distance to measure the difference. By minimizing the distance between the manifolds, we iteratively obtain the optimal registration of the original image pair. Experiment results show that our approach is effective to deal with the multi-modal image registration on the BrainWeb dataset. Yuru Pei, Hongbin Zha |
ICIG (2) | 2 |
| 2014 | Enhanced Random Forest with Image/Patch-Level Learning for Image UnderstandingabstractImage understanding is an important research domain in the computer vision due to its wide real-world applications. For an image understanding framework that uses the Bag-of-Words model representation, the visual codebook is an essential part. Random forest (RF) as a tree-structure discriminative codebook has been a popular choice. However, the performance of the RF can be degraded if the local patch labels are poorly assigned. In this paper, we tackle this problem by a novel way to update the RF codebook learning for a more discriminative codebook with the introduction of the soft class labels, estimated from the pLSA model based on a feedback scheme. The feedback scheme is performed on both the image and patch levels respectively, which is in contrast to the state-of-the-art RF codebook learning that focused on either image or patch level only. Experiments on 15-Scene and C-Pascal datasets had shown the effectiveness of the proposed method in image understanding task. Wai Lam Hoo, Tae-Kyun Kim 0001, Yuru Pei, Chee Seng Chan |
ICPR | 3 |
| 2013 | Anatomical Structure Sketcher for Cephalograms by Bimodal Deep LearningabstractLateral cephalogram X-ray (LCX) images are essential to provide patientspecific morphological information of anatomical structures. The automatic annotation of anatomical structures in cephalograms has been performed in the biomedical engineering for nearly twenty years. Most systems only handle a portion of salient craniofacial landmark set [1, 2, 3]. Although model-based methods can produce a full set of markers [5, 7], the pattern fitting can fail to converge in blurry images. It is challenging to annotate LCX images with high fidelity. In this work, we propose a novel cephalogram sketcher system as shown in Fig. 1 for the automatic anatomical-structure annotation, especially for the blemished images due to structure overlappings and devicespecific distortions during projection. Firstly, we introduce an hierarchical extension of a pictorial model to detect anatomical structures. Secondly, the bimodal deep Boltzmann machine (DBM) is employed to sketch the structure contours. Specifically, the contour sketcher takes advantages of the path in the DBM to extract the contour definitions from the patch textures by alternating Gibbs sampling. Given a cephalogram I, the structure definition S, and the parameters Θ = (Θq,Θr) with respect to the intraand inter-layer correlations, the posterior probability distribution according to the Bayes rule is defined as P(S|I,Θ) ∝ P(I|S,Θ)P(S|Θ), where P(S|Θ) is a shape prior distribution. P(I|S,Θ) is the image likelihood given the hierarchical architecture and the model parameters. The likelihood can be factorized as a product of likelihoods of local structures. Yuru Pei, Hongbin Zha, Tianmin Xu |
BMVC | 1 |
| 2013 | Unsupervised Random Forest Manifold Alignment for LipreadingabstractLip reading from visual channels remains a challenging topic considering the various speaking characteristics. In this paper, we address an efficient lip reading approach by investigating the unsupervised random forest manifold alignment (RFMA). The density random forest is employed to estimate affinity of patch trajectories in speaking facial videos. We propose novel criteria for node splitting to avoid the rank-deficiency in learning density forests. By virtue of the hierarchical structure of random forests, the trajectory affinities are measured efficiently, which are used to find embeddings of the speaking video clips by a graph-based algorithm. Lip reading is formulated as matching between manifolds of query and reference video clips. We employ the manifold alignment technique for matching, where the L∞-norm-based manifold-to-manifold distance is proposed to find the matching pairs. We apply this random forest manifold alignment technique to various video data sets captured by consumer cameras. The experiments demonstrate that lip reading can be performed effectively, and outperform state-of-the-arts. Yuru Pei, Tae-Kyun Kim 0001, Hongbin Zha |
ICCV | 1 |
| 2012 | Random-sampling-based spatial-temporal feature for consumer video concept classificationabstractConcept classification for consumer videos is a challenging task considering the co-occurrence of a variety objects and arbitrary motions in video segments. In this paper, we present a novel video concept classification framework with random-sampling-based spatialtemporal features. Short-term random-sampled point tracks are obtained within video segments. The spatial-temporal features are extracted from these tracks. Concept codebooks are constructed using Multiple Instance Learning upon the spatial-temporal features. The SVM classifiers are trained over codebook-based histograms for an online concept detection. We performed experiments on a video database taken from YouTube. The experimental results demonstrate that the consumer videos can be efficiently assigned concept labels by our approach. Anjun Wei, Yuru Pei, Hongbin Zha |
ICIP | 2 |
| 2012 | Unsupervised Image Matching Based on Manifold AlignmentabstractThis paper challenges the issue of automatic matching between two image sets with similar intrinsic structures and different appearances, especially when there is no prior correspondence. An unsupervised manifold alignment framework is proposed to establish correspondence between data sets by a mapping function in the mutual embedding space. We introduce a local similarity metric based on parameterized distance curves to represent the connection of one point with the rest of the manifold. A small set of valid feature pairs can be found without manual interactions by matching the distance curve of one manifold with the curve cluster of the other manifold. To avoid potential confusions in image matching, we propose an extended affine transformation to solve the nonrigid alignment in the embedding space. The comparatively tight alignments and the structure preservation can be obtained simultaneously. The point pairs with the minimum distance after alignment are viewed as the matchings. We apply manifold alignment to image set matching problems. The correspondence between image sets of different poses, illuminations, and identities can be established effectively by our approach. Yuru Pei, Fengchun Huang, Fuhao Shi, Hongbin Zha |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Visyllable-specific facial transition motion embedding and extractionabstractThe visual facial appearances are important to the speaking perception. The effective and reasonable extraction of the transition motions between the keyframes is desirable to the facial speech animation. In this paper, we present the visyllable-specific transition motion embedding with the temporal extension of the Laplacian eigenmaps (TLE). By imposing the temporal constraints, the TLE-based embedding preserves the possible transitions between the keyshapes inside the visyllable sequence. Given the keyframe pair, the in-between transition motions can be extracted in the latent space by the shortest path searching algorithm. Our experiments demonstrate an effective engine for embedding and extracting the transition motions specific to the visyllables. Yuru Pei, Hongbin Zha |
ICIP | 1 |
| 2009 | Interactive modeling of 3D facial expressions with hierarchical Gaussian process latent variable modelsabstractThe natural expressions play an important role in the daily communication. The efficient and intuitive facial expression editing based on the limited constraints is desirable in the facial animation. In this paper, we present an interactive 3D facial expression editing system with the hierarchical Gaussian process latent variable model (HGPLVM). The hierarchical model incorporates the joint work of the local facial features to produce the natural expressions. To deal with the holistic expression modeling from the local constraints, the inverse mapping between the low-level feature nodes and the high-level facial region nodes is established by the RBF regression model in the latent space. A propagation algorithm is introduced to predict the holistic facial configurations. The experiments demonstrate the 3D facial expressions satisfying the user constraints can be produced efficiently. Fuhao Shi, Yuru Pei, Hongbin Zha |
ICIP | 2 |
| 2009 | 3D facial expression editing based on the dynamic graph modelabstractTo model a detailed 3D expressive face based on the limited user constraints is a challenge work. In this paper, we present the facial expression editing technique based on a dynamic graph model. The probabilistic relations between facial expressions and the complex combination of local facial features, as well as the temporal behaviors of facial expressions are represented by the hierarchical dynamic Bayesian network. Given limited user-constraints on the sparse feature mesh, the system can infer the basis expression probabilities, which are used to locate the corresponding expressive mesh in the shape space spanned by the basis models. The experiments demonstrate the 3D dense facial meshes corresponding to the user-constraints can be synthesized effectively. Yuru Pei, Hongbin Zha |
ICME | 1 |
| 2008 | Creating a face model from an unknown skull based on the tissue mapabstractThe craniofacial reconstruction is employed as an initialization of the identification in forensics. In this paper we present a tissue map based craniofacial reconstruction technique. The reconstruction is formulated as the superimposition of the selected tissue map onto the novel skull. The key problem is the accurate map registration, which is implemented as a warping guided by 2D feature patterns. Given a novel skull, the feature patterns are extracted automatically under an energy minimization framework. The feature configuration on the warped tissue map is expected to resemble that on the novel skull. The target facial model is reconstructed by a smooth interpolation of the point cloud, which results from a simple addition of range images. The presented experiments demonstrate the facial outlook can be reconstructed from the tissue map feasibly and efficiently. Yuru Pei, Hongbin Zha, Zhongbiao Yuan |
ICIP | 1 |
| 2008 | Facial feature estimation from the local structural diversity of skullsabstractIn forensics, the craniofacial reconstruction is employed as an initialization of the identification from skulls. It is a challenging work to develop such a system due to the ambiguity in the relationship between the shape of the skull and the face. In this paper, we present a facial feature estimation method based on the local structural diversity of skulls. A mapping system between the skull structural measurements and the facial feature shapes is established via a RBF regression model. The PCA subspaces are established for the local facial features and the skull structures. Moreover, we investigate the attribute vector of the facial feature polyhedron and the distance graph of the skull structure as the shape descriptors. The experiments demonstrate the feature outlooks can be estimated feasibly and efficiently. Yuru Pei, Hongbin Zha, Zhongbiao Yuan |
ICPR | 1 |
| 2008 | The Craniofacial Reconstruction from the Local Structural Diversity of SkullsabstractAbstract The craniofacial reconstruction is employed as an initialization of the identification from skulls in forensics. In this paper, we present a two‐level craniofacial reconstruction framework based on the local structural diversity of the skulls. On the low level, the holistic reconstruction is formulated as the superimposition of the selected tissue map on the novel skull. The crux is the accurate map registration, which is implemented as a warping guided by the 2D feature curve patterns. The curve pattern extraction under an energy minimization framework is proposed for the automatic feature labeling on the skull depth map. The feature configuration on the warped tissue map is expected to resemble that on the novel skull. In order to make the reconstructed faces personalized, on the high level, the local facial features are estimated from the skull measurements via a RBF model. The RBF model is learnt from a dataset of the skull and the face feature pairs extracted from the head volume data. The experiments demonstrate the facial outlooks can be reconstructed feasibly and efficiently. Yuru Pei, Hongbin Zha, Zhongbiao Yuan |
Comput. Graph. Forum | 1 |
| 2007 | Stylized synthesis of facial speech motionsabstractAbstract Stylized synthesis of facial speech motions is central to facial animation. Most synthesis algorithms put emphasis on the reasonable concatenation of captured motion segments. The dynamic modeling of speech units, e.g. visemes and visyllables (the visual appearance of a syllable), has not drawn much attention. In this paper, we address the fundamental issues regarding the stylized dynamic modeling of visyllables. The decomposable generalized model is learnt for the stylized motion synthesis. The visyllable modeling includes two parts: (1) A dynamic model for each kind of visyllable that is learnt based on a Gaussian Process Dynamical Model; (2) A multilinear model based unified mapping between the high dimensional observation space and low dimensional latent space. The dynamic visyllable model embeds the high dimensional motion data, and constructs the dynamic mapping in the latent space simultaneously. To generalize the visyllable model from several instances, the mapping coefficient matrices are assembled to a tensor, which is decomposed into independent modes, e.g. identity and uttering styles. Therefore, with the linear combination of components in each mode, the novel stylized motions can be synthesized. Copyright © 2007 John Wiley & Sons, Ltd. Yuru Pei, Hongbin Zha |
Comput. Animat. Virtual Worlds | 1 |
| 2007 | Transferring of Speech Movements from Video to 3D Face SpaceabstractWe present a novel method for transferring speech animation recorded in low quality videos to high resolution 3D face models. The basic idea is to synthesize the animated faces by an interpolation based on a small set of 3D key face shapes which span a 3D face space. The 3D key shapes are extracted by an unsupervised learning process in 2D video space to form a set of 2D visemes which are then mapped to the 3D face space. The learning process consists of two main phases: 1) Isomap-based nonlinear dimensionality reduction to embed the video speech movements into a low-dimensional manifold and 2) K-means clustering in the low-dimensional space to extract 2D key viseme frames. Our main contribution is that we use the Isomap-based learning method to extract intrinsic geometry of the speech video space and thus to make it possible to define the 3D key viseme shapes. To do so, we need only to capture a limited number of 3D key face models by using a general 3D scanner. Moreover, we also develop a skull movement recovery method based on simple anatomical structures to enhance 3D realism in local mouth movements. Experimental results show that our method can achieve realistic 3D animation effects with a small number of 3D key face models. Yuru Pei, Hongbin Zha |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2006 | Vision Based Speech Animation Transferring with Underlying Anatomical Structure
Yuru Pei, Hongbin Zha |
ACCV (1) | 1 |
| 2004 | Tissue map based craniofacial reconstruction and facial deformation using RBF networkabstractIn this paper we present a novel craniofacial reconstruction method employing statistical tissue thickness information. The tissue thickness data gotten from CT images are represented as 2D tissue maps. The input (target) skull model is parameterized onto a 2D planar map and the landmarks are utilized to train a RBFN (radial basis function network), which realizes warping of planar maps between the target tissue and the generic tissue. The generic tissue is aligned onto the target skull by applying the trained network onto it, and thus the target facial map can be obtained by a simple addition of the warped generic maps. Finally, we interactively deform the model based on a RBFN to make the facial meshes more personalized, and map the texture from orthogonal photos onto the reconstructed model to improve rendering effects. Experiment results show that the proposed approach is helpful in improving the recognition ability in forensic applications. Yuru Pei, Hongbin Zha, Zhongbiao Yuan |
ICIG | 1 |