VLDB 2026 Research / reviewers in the wild / expert
Ming Zeng 0008
dblp:52/2761-8
· DBLP profile ↗
42ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-5056-0706ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 6 first-author · 22 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BoostPoint: Boosting Point Cloud Backbones with Image Pre-Training for 3D UnderstandingabstractNowadays, pre-training models on largescale datasets and fine-tuning models on task-specific datasets have become common paradigms, achieving impressive success in natural language processing and 2D vision. Nonetheless, the potential of this paradigm has not been fully explored in 3D vision due to the scale of the datasets. To overcome this, we propose BoostPoint, a novel pipeline that uses large-scale rendered images as 3D point cloud model inputs for pretraining and uses general 3D tasks for fine-tuning. In BoostPoint, we propose a novel learning-free image-topoint (I2P) module to transform raw pixels into required inputs. Specifically, we view pixels as unorganized points, including essential raw features (e.g., color) and positional information (e.g., coordinates). Employing simple linear iterative clustering (SLIC), the I2P module effectively groups these unorganized points into superpixels, facilitating point cloud backbone pretraining. Furthermore, we employ a modality-agnostic debiasing mechanism during pre-training to prevent negative transfer in downstream tasks. Extensive finetuning experiments show that BoostPoint provides significant improvements to 3D point cloud backbones for 3D point cloud classification and part segmentation. Honggu Zhou, Yakai Zhang, Haohan Li, Xiaoling Gu, Ming Zeng 0008, Zizhao Wu |
Comput. Vis. Media | 5 |
| 2025 | RHanDS: Refining Malformed Hands for Generated Images with Decoupled Structure and Style GuidanceabstractAlthough diffusion models can generate high-quality human images, their applications are limited by the instability in generating hands with correct structures. In this paper, we introduce RHanDS, a conditional diffusion-based framework designed to refine malformed hands by utilizing decoupled structure and style guidance. The hand mesh reconstructed from the malformed hand offers structure guidance for correcting the structure of the hand, while the malformed hand itself provides style guidance for preserving the style of the hand. To alleviate the mutual interference between style and structure guidance, we introduce a two-stage training strategy and build a series of multi-style hand datasets. In the first stage, we use paired hand images for training to ensure stylistic consistency in hand refining. In the second stage, various hand images generated based on human meshes are used for training, enabling the model to gain control over the hand structure. Experimental results demonstrate that RHanDS can effectively refine hand structure while preserving consistency in hand style. Ming Zeng 0008, Xubin Li, Tiezheng Ge, Bo Zheng 0007 |
AAAI | 4 |
| 2025 | Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual SignalsabstractThis paper addresses the gap in predicting turn-taking and backchannel actions in humanmachine conversations using multi-modal signals (linguistic, acoustic, and visual).To overcome the limitation of existing datasets, we propose an automatic data collection pipeline that allows us to collect and annotate over 210 hours of human conversation videos.From this, we construct a Multi-Modal Face-to-Face (MM-F2F) human conversation dataset, including over 1.5M words and corresponding turntaking and backchannel annotations from approximately 20M frames.Additionally, we present an end-to-end framework that predicts the probability of turn-taking and backchannel actions from multi-modal signals.The proposed model emphasizes the interrelation between modalities and supports any combination of text, audio, and video inputs, making it adaptable to a variety of realistic scenarios.Our experiments show that our approach achieves state-of-the-art performance on turn-taking and backchannel prediction tasks, achieving a 10% increase in F1-score on turn-taking and a 33% increase on backchannel prediction.Our dataset and code are publicly available online to ease of subsequent research.The code and dataset are available at https://github.com/Linyx1125/MM-F2F. Yinglin Zheng, Ming Zeng 0008, Wangzheng Shi |
ACL (1) | 3 |
| 2025 | EmoHuman: Fine-Grained Emotion-Controlled Talking Head Generation via Audio-Text Multimodal DetanglingabstractAudio-driven talking head generation has made significant strides in creating realistic and lip-synchronized portraits. However, most existing approaches overlook facial expressions, with only a few attempting to model facial emotions explicitly, often leading to unnatural results. To address this gap, we introduce EmoHuman, an audio-to-video synthesis method that generates emotionally nuanced talking head videos without relying on intermediate 3D representations or facial landmarks. EmoHuman decouples content, emotion, and emotional intensity from the multimodal information of the audio and the corresponding textual content through an Audio Emotion Decoupling Module. The content features are used to drive a powerful video diffusion model, generating synchronized lip movements, while emotion and emotional intensity govern the simulation of facial expressions, resulting in more realistic video outputs. Extensive experiments demonstrate that EmoHuman outperforms state-of-the-art methods in image and video quality, expression correlation, and lip-synchronization accuracy. Qifeng Dai, Huidong Feng, Wendi Cui, Xinqi Cai, Yinglin Zheng, Ming Zeng 0008 |
ICMR | 6 |
| 2025 | Consistent Human Animation with Pseudo Multi-View Anchoring and Cross-Granularity Integration
Jintai Wang, Yinglin Zheng, Qifeng Dai, Ming Zeng 0008 |
ICMR | 5 |
| 2025 | HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided AnimationabstractHair transfer is increasingly valuable across domains such as social media, gaming, advertising, and entertainment. While significant progress has been made in single-image hair transfer, video-based hair transfer remains challenging due to the need for temporal consistency, spatial fidelity, and dynamic adaptability. In this work, we propose HairShifter, a novel ''Anchor Frame + Animation'' framework that unifies high-quality image hair transfer with smooth and coherent video animation. At its core, HairShifter integrates a Image Hair Transfer (IHT) module for precise per-frame transformation and a Multi-Scale Gated SPADE Decoder to ensure seamless spatial blending and temporal coherence. Our method maintains hairstyle fidelity across frames while preserving non-hair regions. Extensive experiments demonstrate that HairShifter achieves state-of-the-art performance in video hairstyle transfer, combining superior visual quality, temporal consistency, and scalability. The code will be publicly available. We believe this work will open new avenues for video-based hairstyle transfer and establish a robust baseline in this field. Wangzheng Shi, Yinglin Zheng, Jianmin Bao, Ming Zeng 0008, Dong Chen 0003 |
ACM Multimedia | 5 |
| 2025 | PointHuman: Learning high-fidelity and generalizable human neural radiance fields using guidance of fine-grained semantics-enriched geometry
Jintai Wang, Huidong Feng, Qifeng Dai, Yinglin Zheng, Ming Zeng 0008 |
Comput. Graph. | 6 |
| 2025 | SPAC-Net: Rethinking Point Cloud Completion With Structural PriorabstractPoint cloud completion aims to infer a complete shape from its partial observation. Many approaches utilize a pure encoder-decoder paradigm in which complete shape can be directly predicted by shape priors learned from partial scans, however, these methods suffer from the loss of details inevitably due to the feature abstraction issues. In this paper, we propose a novel framework, termed SPAC-Net, that aims to rethink the completion task under the guidance of a new structural prior, we call it interface. Specifically, our method first investigates Marginal Detector (MAD) module to localize the interface, defined as the intersection between the known observation and the missing parts. Based on the interface, our method predicts the coarse shape by learning the displacement from the points in interface move to their corresponding position in missing parts. Furthermore, we devise an additional Structure Supplement (SSP) module before the upsampling stage to enhance the structural details of the coarse shape, enabling the upsampling module to focus more on the upsampling task. Extensive experiments have been conducted on several challenging benchmarks, and the results demonstrate that our method outperforms existing state-of-the-art approaches. Zizhao Wu, Cheng Zhang 0041, Genfu Yang, Ming Zeng 0008, Yunhai Wang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Interpretable procedural material graph generation via diffusion models from reference images
Xiaoyu Lv, Zizhao Wu, Jiamin Xu, Xiaoling Gu, Ming Zeng 0008, Weiwei Xu 0003 |
Vis. Comput. | 5 |
| 2024 | Multi-Modal Gait Recognition with Unidirectional Cross-modal AlignmentabstractGait recognition represents a pivotal challenge in visual signal comprehension, encompassing multi-modal spatial-temporal information, and exhibiting considerable complexity. The prevailing gait recognition methods are categorized into appearance-based and model-based, which commonly use silhouette and skeleton as input respectively. Certain recent studies have endeavored to integrate both of these modalities, achieving a certain degree of success in doing so. However, the aforementioned methods either rely solely on a modal of information or lack a profound consideration of the intricate interplay among multiple modal. Consequently, the potential inherent in gait’s spatial-temporal information remains untapped. To address this issue, this study strives to enhance the efficiency of data utilization through the integration of multi-modal information within a multi-modal fusion network. Furthermore, we propose a scheme aimed at enhancing the consistency between distinct modal features. Experiments on the widely used gait dataset CASIA-B have shown that our model has significantly improved under complex gait conditions, with overall performance reaching the most advanced level. Hengda Li, Yinglin Zheng, Qifeng Dai, Jintai Wang, Ming Zeng 0008 |
ICME | 6 |
| 2024 | High-fidelity instructional fashion image editingabstractInstructional image editing has received a significant surge of attention recently. In this work, we are interested in the challenging problem of instructional image editing within the particular fashion realm, a domain with significant potential demand in both commercial and personal contexts. This specific domain presents heightened challenges owing to the stringent quality requirements. It necessitates not only the creation of vivid details in alignment with instructions, but also the preservation of precise attributes unrelated to the text guidance. Naive extensions of existing image editing methods produce noticeable artifacts. In order to achieve high-fidelity fashion editing, we propose a novel framework, leveraging the generative prior of a pre-trained human generator and performing edit in the latent space. In addition, we introduce a novel CLIP-based loss to better align the generated target with the instruction. Extensive experiments demonstrate that our approach outperforms prior works including GAN-based editing as well as diffusion-based editing by a large margin, showing impressive visual quality. Yinglin Zheng, Ting Zhang 0002, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008 |
Graph. Model. | 5 |
| 2024 | Contrastive disentanglement for self-supervised motion style transfer
Zizhao Wu, Siyuan Mao, Cheng Zhang 0041, Yigang Wang, Ming Zeng 0008 |
Multim. Tools Appl. | 5 |
| 2023 | EMEF: Ensemble Multi-Exposure Image FusionabstractAlthough remarkable progress has been made in recent years, current multi-exposure image fusion (MEF) research is still bounded by the lack of real ground truth, objective evaluation function, and robust fusion strategy. In this paper, we study the MEF problem from a new perspective. We don’t utilize any synthesized ground truth, design any loss function, or develop any fusion strategy. Our proposed method EMEF takes advantage of the wisdom of multiple imperfect MEF contributors including both conventional and deep learning-based methods. Specifically, EMEF consists of two main stages: pre-train an imitator network and tune the imitator in the runtime. In the first stage, we make a unified network imitate different MEF targets in a style modulation way. In the second stage, we tune the imitator network by optimizing the style code, in order to find an optimal fusion result for each input pair. In the experiment, we construct EMEF from four state-of-the-art MEF methods and then make comparisons with the individuals and several other competitive methods on the latest released MEF benchmark dataset. The promising experimental results demonstrate that our ensemble framework can “get the best of all worlds”. The code is available at https://github.com/medalwill/EMEF. Renshuai Liu, Haitao Cao 0006, Yinglin Zheng, Ming Zeng 0008 |
AAAI | 5 |
| 2023 | MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingabstractThis paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a masked image. Such incorporation enjoys two vital benefits. First, masked self-distillation targets local patch representation learning, which is complementary to vision-language contrastive focusing on text-related representation. Second, masked self-distillation is also consistent with vision-language contrastive from the perspective of training objective as both utilize the visual encoder for feature aligning, and thus is able to learn local semantics getting indirect supervision from the language. We provide specially designed experiments with a comprehensive analysis to validate the two benefits. Symmetrically, we also introduce the local semantic supervision into the text branch, which further improves the pretraining performance. With extensive experiments, we show that MaskCLIP, when applied to various challenging downstream tasks, achieves superior results in linear probing, finetuning, and zeroshot performance with the guidance of the language encoder. Code will be release at https://github.com/LightDXY/MaskCLIP. Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 0002, Dongdong Chen 0001, Hao Yang 0036, Ming Zeng 0008, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu |
CVPR | 7 |
| 2023 | Vertex position estimation with spatial-temporal transformer for 3D human reconstructionabstractReconstructing 3D human pose and body shape from monocular images or videos is a fundamental task for comprehending human dynamics. Frame-based methods can be broadly categorized into two fashions: those regressing parametric model parameters (e.g., SMPL) and those exploring alternative representations (e.g., volumetric shapes, 3D coordinates). Non-parametric representations have demonstrated superior performance due to their enhanced flexibility. However, when applied to video data, these non-parametric frame-based methods tend to generate inconsistent and unsmooth results. To this end, we present a novel approach that directly regresses the 3D coordinates of the mesh vertices and body joints with a spatial–temporal Transformer. In our method, we introduce a SpatioTemporal Learning Block (STLB) with Spatial Learning Module (SLM) and Temporal Learning Module (TLM), which leverages spatial and temporal information to model interactions at a finer granularity, specifically at the body token level. Our method outperforms previous state-of-the-art approaches on Human3.6M and 3DPW benchmark datasets. Xiangjun Zhang, Yinglin Zheng, Wenjin Deng, Qifeng Dai, Wangzheng Shi, Ming Zeng 0008 |
Graph. Model. | 7 |
| 2023 | MusicFace: Music-driven expressive singing face synthesisabstractIt remains an interesting and challenging problem to synthesize a vivid and realistic singing face driven by music. In this paper, we present a method for this task with natural motions for the lips, facial expression, head pose, and eyes. Due to the coupling of mixed information for the human voice and backing music in common music audio signals, we design a decouple-and-fuse strategy to tackle the challenge. We first decompose the input music audio into a human voice stream and a backing music stream. Due to the implicit and complicated correlation between the two-stream input signals and the dynamics of the facial expressions, head motions, and eye states, we model their relationship with an attention scheme, where the effects of the two streams are fused seamlessly. Furthermore, to improve the expressivenes of the generated results, we decompose head movement generation in terms of speed and direction, and decompose eye state generation into short-term blinking and long-term eye closing, modeling them separately. We have also built a novel dataset, SingingFace, to support training and evaluation of models for this task, including future work on this topic. Extensive experiments and a user study show that our proposed method is capable of synthesizing vivid singing faces, qualitatively and quantitatively better than the prior state-of-the-art. Wenjin Deng, Hengda Li, Jintai Wang, Yinglin Zheng, Yiwei Ding, Xiaohu Guo, Ming Zeng 0008 |
Comput. Vis. Media | 8 |
| 2023 | 3D Talking Face With Personalized Pose DynamicsabstractRecently, we have witnessed a boom in applications for 3D talking face generation. However, most existing 3D face generation methods can only generate 3D faces with a static head pose, which is inconsistent with how humans perceive faces. Only a few articles focus on head pose generation, but even these ignore the attribute of personality. In this article, we propose a unified audio-driven approach to endow 3D talking faces with personalized pose dynamics. To achieve this goal, we establish an original person-specific dataset, providing corresponding head poses and face shapes for each video. Our framework is composed of two separate modules: PoseGAN and PGFace. Given an input audio, PoseGAN first produces a head pose sequence for the 3D head, and then, PGFace utilizes the audio and pose information to generate natural face models. With the combination of these two parts, a 3D talking head with dynamic head movement can be constructed. Experimental evidence indicates that our method can generate person-specific head pose sequences that are in sync with the input audio and that best match with the human experience of talking heads. Saifeng Ni, Zhipeng Fan 0001, Ming Zeng 0008, Madhukar Budagavi, Xiaohu Guo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | General Facial Representation Learning in a Visual-Linguistic MannerabstractHow to learn a universal facial representation that boosts all face analysis tasks? This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learning. On one hand, the framework involves a contrastive loss to learn high-level semantic meaning from image-text pairs. On the other hand, we propose exploring low-level information simultaneously to further enhance the face representation by adding a masked image modeling. We perform pre-training on LAION-FACE, a dataset containing a large amount of face image-text pairs, and evaluate the representation capability on multiple downstream tasks. We show that FaRL achieves better transfer performance compared with previous pre-trained models. We also verify its superiority in the low-data regime. More importantly, our model surpasses the state-of-the-art methods on face analysis tasks including face parsing and face alignment. Yinglin Zheng, Hao Yang 0036, Ting Zhang 0002, Jianmin Bao, Dongdong Chen 0001, Yangyu Huang, Lu Yuan 0001, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
CVPR | 9 |
| 2022 | I²R-Net: Intra- and Inter-Human Relation Network for Multi-Person Pose EstimationabstractIn this paper, we present the Intra- and Inter-Human Relation Networks I²R-Net for Multi-Person Pose Estimation. It involves two basic modules. First, the Intra-Human Relation Module operates on a single person and aims to capture Intra-Human dependencies. Second, the Inter-Human Relation Module considers the relation between multiple instances and focuses on capturing Inter-Human interactions. The Inter-Human Relation Module can be designed very lightweight by reducing the resolution of feature map, yet learn useful relation information to significantly boost the performance of the Intra-Human Relation Module. Even without bells and whistles, our method can compete or outperform current competition winners. We conduct extensive experiments on COCO, CrowdPose, and OCHuman datasets. The results demonstrate that the proposed model surpasses all the state-of-the-art methods. Concretely, the proposed method achieves 77.4% AP on CrowPose dataset and 67.8% AP on OCHuman dataset respectively, outperforming existing methods by a large margin. Additionally, the ablation study and visualization analysis also prove the effectiveness of our model. Yiwei Ding, Wenjin Deng, Yinglin Zheng, Meihong Wang, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008 |
IJCAI | 9 |
| 2022 | Controllable Facial Caricaturization With Localized Deformation and Personalized Semantic AttentionsabstractThe facial caricature shows the distinct characteristics of a person via exaggerations of both shape and appearance. This paper presents a novel framework that automatically generates vivid facial caricatures by encoding personalized semantic information. To this end, we first design a part-based scheme for geometry warping, which composes local semantic deformation into a global warping field, equipped with sufficient warping freedom of different facial components. Second, under the scheme of Part-based Warping, we design a photo-to-caricature translation network called PbWarpGAN, and adopt several novel losses to capture the personalized characteristics of each input face and preserve its identity better. Third, based on PbWarpGAN, we develop a user-friendly interface by introducing an attention scheme on each facial component, allowing ordinary users to adjust the automatically generated caricature by PbWarpGAN according to their preference conveniently. Experimental results show that our PbWarpGAN is more effective in capturing personalized characteristics than counterparts, and provides an efficient tool for caricature designing application. Ming Zeng 0008, Yinglin Zheng, Jinpeng Lin, Jing Liao 0001, Zizhao Wu, Wenjin Deng |
IEEE Trans. Multim. | 1 |
| 2021 | FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningabstractIn this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and eye blinks that are in-sync with the input audio signal. We note that the synthetic face attributes include not only explicit ones such as lip motions that have high correlations with speech, but also implicit ones such as head poses and eye blinks that have only weak correlation with the input audio. To model such complicated relationships among different face attributes with input audio, we propose a FACe Implicit Attribute Learning Generative Adversarial Network (FACIAL-GAN), which integrates the phonetics-aware, context-aware, and identity-aware information to synthesize the 3D face animation with realistic motions of lips, head poses, and eye blinks. Then, our Rendering-to-Video network takes the rendered face images and the attention map of eye blinks as input to generate the photorealistic output video frames. Experimental results and user studies show our method can generate realistic talking face videos with not only synchronized lip motions, but also natural head movements and eye blinks, with better qualities than the results of state-of-the-art methods. Yifan Zhao 0002, Ming Zeng 0008, Saifeng Ni, Madhukar Budagavi, Xiaohu Guo |
ICCV | 4 |
| 2021 | Exploring Temporal Coherence for More General Video Face Forgery DetectionabstractAlthough current face manipulation techniques achieve impressive performance regarding quality and controllability, they are struggling to generate temporal coherent face videos. In this work, we explore to take full advantage of the temporal coherence for video face forgery detection. To achieve this, we propose a novel end-to-end framework, which consists of two major stages. The first stage is a fully temporal convolution network (FTCN). The key insight of FTCN is to reduce the spatial convolution kernel size to 1, while maintaining the temporal convolution kernel size un-changed. We surprisingly find this special design can benefit the model for extracting the temporal features as well as improve the generalization capability. The second stage is a Temporal Transformer network, which aims to explore the long-term temporal coherence. The proposed frame-work is general and flexible, which can be directly trained from scratch without any pre-training models or external datasets. Extensive experiments show that our framework outperforms existing methods and remains effective when applied to detect new sorts of face forgery videos. Yinglin Zheng, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
ICCV | 4 |
| 2021 | Real-Time Masked Face Revealing for Video ConferenceabstractVideo conferencing is an essential way for contactless conversation, which conveys abundant multimedia signals. Especially under COVID-19, the video conference has been becoming a common way for daily communications. However, for the sake of plague prevention, it usually happens that the people attending the video conference are wearing a mouth mask, leading to inconvenient communication due to incomplete facial information. To tackle this problem, we develop a novel system that reveals the masked faces in real-time, making each participant feel like the others are mask-free. Moreover, we map the audio to 3DMM expression to guide the generation of various mouth shapes utilizing multi-modal information. Extensive experiments validate the revealing effectiveness and better user experience of the system. Furthermore, by applying lightweight networks design, the proposed system can run in real-time. Jinpeng Lin, Yinglin Zheng, Wenjin Deng, Ming Zeng 0008 |
ICME | 5 |
| 2020 | VH3D-LSFM: Video-Based Human 3D Pose Estimation with Long-Term and Short-Term Pose Fusion Mechanism
Wenjin Deng, Yinglin Zheng, Zizhao Wu, Ming Zeng 0008 |
PRCV (1) | 6 |
| 2020 | Joint learning for face alignment and face transfer with depth image
Xiaoli Wang 0002, Yinglin Zheng, Ming Zeng 0008, Wei Lu 0015 |
Multim. Tools Appl. | 3 |
| 2019 | Face Parsing With RoI Tanh-WarpingabstractFace parsing computes pixel-wise label maps for different semantic components (e.g., hair, mouth, eyes) from face images. Existing face parsing literature have illustrated significant advantages by focusing on individual regions of interest (RoIs) for faces and facial components. However,the traditional crop-and-resize focusing mechanism ignores all contextual area outside the RoIs, and thus is not suitable when the component area is unpredictable, e.g. hair. Inspired by the physiological vision system of human, we propose a novel RoI Tanh-warping operator that combines the central vision and the peripheral vision together. It addresses the dilemma between a limited sized RoI for focusing and an unpredictable area of surrounding context for peripheral information. To this end, we propose a novel hybrid convolutional neural network for face parsing. It uses hierarchical local based method for inner facial components and global methods for outer facial components. The whole framework is simple and principled, and can be trained end-to-end. To facilitate future research of face parsing, we also manually relabel the training data of the HELEN dataset and will make it public. Experiments on both HELEN and LFW-PL benchmarks demonstrate that our method surpasses state-of-the-art methods. Jinpeng Lin, Hao Yang 0036, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001, Lu Yuan 0001 |
CVPR | 4 |
| 2019 | Efficient L0 resampling of point sets
Ming Zeng 0008, Jinpeng Lin, Zizhao Wu, Xinguo Liu |
Comput. Aided Geom. Des. | 2 |
| 2018 | Geometric Primitives Based RGB-D SLAM for Low-texture EnvironmentabstractVisual Simultaneous Localization and Mapping (SLAM) is very important in various applications such as AR, Robotics, etc. A challenging problem in SLAM is the inferior tracking performance in the low-texture environment due to their low-level feature based tactic. This paper presents a novel SLAM system which leverages feature-wise alignment and layout information to handle this challenge. The key idea is to utilize the points, lines and planes with both feature-wise constraints and layout consistency constraints to obtain robust motion estimation. Firstly, we extract the points, lines and planes from color images and depth images. Then, we find some matches between these geometric primitives. Lastly, we propose a unified solution to represent the local primitives information and the layout context among the extracted primitives, thus to stabilize the camera motion estimation. Experiments on TUM datasets demonstrate that our method outperforms other RGB-D SLAM systems in the low-texture environment, and provides comparable results in the richly-textured environment. Penglei Ji, Ming Zeng 0008, Xinguo Liu |
CASA | 2 |
| 2018 | Accurate geometry modeling of vasculatures using implicit fitting with 2D radial basis functions
Qingqi Hong, Qingde Li, Beizhan Wang, Kunhong Liu 0001, Fan Lin, Juncong Lin, Zhihong Zhang 0001, Ming Zeng 0008 |
Comput. Aided Geom. Des. | 9 |
| 2018 | Joint analysis of shapes and images via deep domain adaptation
Zizhao Wu, Yunhui Zhang, Ming Zeng 0008, Fei-wei Qin, Yigang Wang |
Comput. Graph. | 3 |
| 2016 | Video segmentation with L0 gradient minimization
Yuanli Feng, Ming Zeng 0008, Xinguo Liu |
Comput. Graph. | 3 |
| 2016 | Spatially constrained level-set tracking and segmentation of non-rigid objects
Ming Zeng 0008, Xinguo Liu |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Shape completion for depth image via repeated objects registration
Ming Zeng 0008, Liujuan Cao, Kunhui Lin, Huailin Dong, Cheng Wang 0003 |
Neurocomputing | 1 |
| 2015 | Estimation of human body shape and cloth field in front of a kinect
Ming Zeng 0008, Liujuan Cao, Huailin Dong, Kunhui Lin, Meihong Wang, Jing Tong |
Neurocomputing | 1 |
| 2014 | Feature-preserving filtering with L0 gradient minimization
Ming Zeng 0008, Xinguo Liu |
Comput. Graph. | 2 |
| 2014 | SCAPE-based human performance reconstruction
Jiaxiang Zheng, Ming Zeng 0008, Xinguo Liu |
Comput. Graph. | 2 |
| 2013 | High Quality Binocular Facial Performance Capture from Partially Blurred Image SequenceabstractExisting methods on passive facial performance capture assume that the input images are well captured. They merely consider how to deal with motion blurred images in the input sequence, which is very common in the image capture process. This paper presents a collection of novel algorithms and a thereby resulting system to reconstruct high quality facial dynamic geometry even from a partially blurred image sequence. In our method, we adopt binocular cameras to capture a stereo sequence. With this sequence, we first estimate depth map for each frame using a state-of-the-art stereo matching method. Then, based on the estimated depth map sequence, we track facial motion by leveraging constraints of both optical flow and geometry. In this step, a blur detection and regularization algorithm are devised to adaptively keep both shape and details. Finally, we synthesize temporal mesoscopic geometry on the blurred region from clear image texture of neighboring frames. We conduct extensive experiments on several sequences containing facial performance with blurred region, and the results demonstrate the effectiveness and robustness of our algorithms. Ming Zeng 0008, Bojun Liang, Xinguo Liu |
CAD/Graphics | 2 |
| 2013 | Dynamic Human Surface Reconstruction Using a Single KinectabstractThis paper presents a system for robust dynamic human surface reconstruction using a single Kinect. The single Kinect provides a self-occluded and noisy RGBD data. Thus it is challenging to track the whole human surface robustly. To overcome both incompleteness and data noise, we adopt a template to confine the shape in the un-seen part, and propose a two-stage tracking pipeline. The first stage tracks the articulated motion of the human, which improves robustness of tracking by introducing more constraints between the surface points. The second stage tracks movements of non-articulated motion. For long sequences, we stabilize the human surface in the un-seen part by directly warping the surface from the first frame to the current frame according to sequentially tracked correspondences, preventing surface from collapsing caused by error accumulation. We demonstrate our method by several real captured RGBD data, containing complex human motion. The reconstruction results show the effectiveness and robustness of our method. Ming Zeng 0008, Jiaxiang Zheng, Xinguo Liu |
CAD/Graphics | 1 |
| 2013 | Templateless Quasi-rigid Shape Modeling with Implicit Loop-ClosureabstractThis paper presents a method for quasi-rigid objects modeling from a sequence of depth scans captured at different time instances. As quasi-rigid objects, such as human bodies, usually have shape motions during the capture procedure, it is difficult to reconstruct their geometries. We represent the shape motion by a deformation graph, and propose a model-to-part method to gradually integrate sampled points of depth scans into the deformation graph. Under an as-rigid-as-possible assumption, the model-to-part method can adjust the deformation graph non-rigidly, so as to avoid error accumulation in alignment, which also implicitly achieves loop-closure. To handle the drift and topological error for the deformation graph, two algorithms are introduced. First, we use a two-stage registration to largely keep the rigid motion part. Second, in the step of graph integration, we topology-adaptively integrate new parts and dynamically control the regularization effect of the deformation graph. We demonstrate the effectiveness and robustness of our method by several depth sequences of quasi-rigid objects, and an application in human shape modeling. Ming Zeng 0008, Jiaxiang Zheng, Xinguo Liu |
CVPR | 1 |
| 2013 | Octree-based fusion for realtime 3D reconstruction
Ming Zeng 0008, Fukai Zhao, Jiaxiang Zheng, Xinguo Liu |
Graph. Model. | 1 |
| 2012 | A Memory-Efficient KinectFusion Using Octree
Ming Zeng 0008, Fukai Zhao, Jiaxiang Zheng, Xinguo Liu |
CVM | 1 |
| 2012 | Video-driven state-aware facial animationabstractABSTRACT It is important in computer animation to synthesize expressive facial animation for avatars from videos. Some traditional methods track a set of semantic feature points on the face to drive the avatar. However, these methods usually suffer from inaccurate detection and sparseness of the feature points and fail to obtain high‐level understanding of facial expressions, leading to less expressive and even wrong expressions on the avatar. In this paper, we propose a state‐aware synthesis framework. Instead of simply fitting 3D face to the 2D feature points, we use expression states obtained by a set of low‐cost classifiers (based on local binary pattern and support vector machine) on the face texture to guide the face fitting procedure. Our experimental results show that the proposed hybrid framework enjoys the advantages of the original methods based on feature point and the awareness of the expression states of the classifiers and thus vivifies and enriches the face expressions of the avatar. Copyright © 2012 John Wiley & Sons, Ltd. Ming Zeng 0008, Xinguo Liu, Hujun Bao |
Comput. Animat. Virtual Worlds | 1 |