VLDB 2026 Research / reviewers in the wild / expert
Peng Zhang 0080
dblp:21/1048-80
· DBLP profile ↗
14ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MaTe3D: Mask-Guided Text-based 3D-Aware Portrait Editing
Kangneng Zhou, Daiheng Gao, Xuan Wang 0009, Jie Zhang 0090, Peng Zhang 0080, Xusen Sun, Longhao Zhang, Shiqi Yang 0002, Bang Zhang, Liefeng Bo, Yaxing Wang, Ming-Ming Cheng |
Int. J. Comput. Vis. | 5 |
| 2025 | VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorabstractAudio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, facial motion, head pose generation, and video quality. However, no model has yet led or tied on all these metrics due to the one-to-many mapping between audio and motion. In this paper, we propose VividTalk, a two-stage generic framework that supports generating high-visual quality talking head videos with all the above properties. Specifically, in the first stage, we map the audio to mesh by learning two motions, including non-rigid facial motion and rigid head motion. For facial motion, both blendshape and vertex are adopted as the intermediate representation to maximize the representation ability of the model. For head motion, a novel learnable head pose codebook with a two-phase training mechanism is proposed. In the second stage, we proposed a dual branch motion-vae and a generator to transform the meshes into dense motion and synthesize high-quality video frame-by-frame. Extensive experiments show that the proposed VividTalk can generate high-visual quality talking head videos with lip-sync and realistic enhanced by a large margin, and outperforms previous state-of-the-art works in objective and subjective comparisons. The code will be publicly released upon publication. Xusen Sun, Longhao Zhang, Hao Zhu 0004, Peng Zhang 0080, Bang Zhang, Xinya Ji, Kangneng Zhou, Daiheng Gao, Liefeng Bo, Xun Cao |
3DV | 4 |
| 2025 | Exploring Timeline Control for Facial Motion GenerationabstractThis paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arranged in temporal intervals, allowing precise control over the timing of each action. To model the timeline control capability, We first annotate the time intervals of facial actions in natural facial motion sequences at a frame-level granularity. This process is facilitated by Toeplitz Inverse Covariance-based Clustering to minimize human labor. Based on the annotations, we propose a diffusion-based generation model capable of generating facial motions that are natural and accurately aligned with input timelines. Our method supports text-guided motion generation by using ChatGPT to convert text into timelines. Experimental results show that our method can annotate facial action intervals with satisfactory accuracy, and produces natural facial motions accurately aligned with timelines. Yifeng Ma 0001, Jinwei Qi, Chaonan Ji, Peng Zhang 0080, Bang Zhang, Zhidong Deng, Liefeng Bo |
CVPR | 4 |
| 2025 | Animate Anyone 2: High-Fidelity Character Image Animation with Environment AffordanceabstractRecent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable associations between characters and their environments. To address this limitation, we introduce Animate Anyone 2, aiming to animate characters with environment affordance. Beyond extracting motion signals from source video, we additionally capture environmental representations as conditional inputs. The environment is formulated as the region with the exclusion of characters and our model generates characters to populate these regions while maintaining coherence with the environmental context. We propose a shape-agnostic mask strategy that more effectively characterizes the relationship between character and environment. Furthermore, to enhance the fidelity of object interactions, we leverage an object guider to extract features of interacting objects and employ spatial blending for feature injection. We also introduce a pose modulation strategy that enables the model to handle more diverse motion patterns. Experimental results demonstrate the superior performance of the proposed method. Guangyuan Wang, Dechao Meng, Lian Zhuo, Peng Zhang 0080, Bang Zhang, Liefeng Bo |
ICCV | 7 |
| 2025 | OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingabstractAlthough significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity's speaking and facial movement styles, including speech characteristics, head motion, and facial dynamics. Our framework adopts a dual-branch diffusion transformer (DiT) architecture, with one branch dedicated to audio generation and the other to video synthesis.
At the shallow layers, cross-modal fusion modules are introduced to integrate information between the two modalities. In deeper layers, each modality is processed independently, with the generated audio decoded by a vocoder and the video rendered using a GAN-based high-quality visual renderer. Leveraging DiT’s in-context learning capability through a masked-infilling strategy, our model can simultaneously capture both audio and visual styles without requiring explicit style extraction modules. Thanks to the efficiency of the DiT backbone and the optimized visual renderer, OmniTalker achieves real-time inference at 25 FPS.
To the best of our knowledge, OmniTalker is the first one-shot framework capable of jointly modeling speech and facial styles in real time. Extensive experiments demonstrate its superiority over existing methods in terms of generation quality, particularly in preserving style consistency and ensuring precise audio-video synchronization, all while maintaining efficient inference. Zhongjian Wang, Peng Zhang 0080, Jinwei Qi, Sheng Xu 0007, Bang Zhang |
NeurIPS | 2 |
| 2024 | Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement LearningabstractTrustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers from hallucinations, one primary cause of which is the conflict between contextual and parametric knowledge. We deem that retrieval-augmented language models have the inherent capabilities of supplying response according to both contextual and parametric knowledge. Inspired by aligning language models with human preference, we take the first step towards aligning retrieval-augmented language models to a status where it responds relying merely on the external evidence and disregards the interference of parametric knowledge. Specifically, we propose a reinforcement learning based algorithm Trustworthy-Alignment, theoretically and experimentally demonstrating large language models' capability of reaching a trustworthy status without explicit supervision on how to respond. Our work highlights the potential of large language models on exploring its intrinsic abilities by its own and expands the application scenarios of alignment from fulfilling human preference to creating trustworthy agents. Zongmeng Zhang, Jinhua Zhu 0001, Wengang Zhou 0001, Xiang Qi, Peng Zhang 0080, Houqiang Li |
ICML | 6 |
| 2022 | Text/Speech-Driven Full-Body AnimationabstractDue to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and corresponding speech, our system synthesizes face and body animations simultaneously, which are then skinned and rendered to obtain a video stream output. We adopt a learning-based approach for synthesizing facial animation and a graph-based approach to animate the body, which generates high-quality avatar animation efficiently and robustly. Our results demonstrate the generated avatar animations are realistic, diverse and highly text/speech-correlated. Wenlin Zhuang, Jinwei Qi, Peng Zhang 0080, Bang Zhang |
IJCAI | 3 |
| 2022 | DART: Articulated Hand Model with Diverse Accessories and Rich TexturesabstractHand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits its power to synthesize photorealistic hand data. In this paper, we extend MANO with Diverse Accessories and Rich Textures, namely DART. DART is composed of 50 daily 3D accessories which varies in appearance and shape, and 325 hand-crafted 2D texture maps covers different kinds of blemishes or make-ups. Unity GUI is also provided to generate synthetic hand data with user-defined settings, e.g., pose, camera, background, lighting, textures, and accessories. Finally, we release DARTset, which contains large-scale (800K), high-fidelity synthetic hand images, paired with perfect-aligned 3D labels. Experiments demonstrate its superiority in diversity. As a complement to existing hand datasets, DARTset boosts the generalization in both hand pose estimation and mesh recovery tasks. Raw ingredients (textures, accessories), Unity GUI, source code and DARTset are publicly available at dart2022.github.io. Daiheng Gao, Yuliang Xiu, Kailin Li 0001, Lixin Yang 0001, Feng Wang 0072, Peng Zhang 0080, Bang Zhang, Cewu Lu |
NeurIPS | 6 |
| 2021 | Learning Position and Target Consistency for Memory-Based Video Object SegmentationabstractThis paper studies the problem of semi-supervised video object segmentation(VOS). Multiple works have shown that memory-based approaches can be effective for video object segmentation. They are mostly based on pixel-level matching, both spatially and temporally. The main shortcoming of memory-based approaches is that they do not take into account the sequential order among frames and do not exploit object-level knowledge from the target. To address this limitation, we propose to Learn position and target Consistency framework for Memory-based video object segmentation, termed as LCM. It applies the memory mechanism to retrieve pixels globally, and meanwhile learns position consistency for more reliable segmentation. The learned location response promotes a better discrimination between target and distractors. Besides, LCM introduces an object-level relationship from the target to maintain target consistency, making LCM more robust to error drifting. Experiments show that our LCM achieves state-of-the-art performance on both DAVIS and Youtube-VOS benchmark. And we rank the 1st in the DAVIS 2020 challenge semi-supervised VOS task. Peng Zhang 0080, Bang Zhang, Rong Jin 0001 |
CVPR | 2 |
| 2021 | A Virtual Character Generation and Animation System for E-Commerce Live StreamingabstractVirtual character has been widely adopted in many areas, such as virtual assistant, virtual customer service, robotics and etc. In this paper, we focus on its application in e-commerce live streaming. Particularly, we propose a virtual character generation and animation system that supports e-commerce live streaming with virtual characters as anchors. The system offers a virtual character face generation tool based on a weakly supervised 3D face reconstruction method. The method takes a single photo as input and generates a 3D face model with both similarity and aesthetics considered. It does not require 3D face annotation data due to the assist of differentiable neural rendering technique which seamlessly integrates rendering into a deep learning based 3D face reconstruction framework. Moreover, the system provides two animation approaches which support two different ways of live stream respectively. The first approach is based on real-time motion capture. An actor's performance is captured in real-time via a monocular camera, and then utilized for animating a virtual anchor. The second approach is text driven animation, in which the human-like animation is automatically generated based on a text script. The relationship between text script and animation is learned based on the training data which can be accumulated via the motion capture based animation. To our best knowledge, the presented work is the first sophisticated virtual character generation and animation system that is designed for e-commerce live streaming and actually deployed on an online shopping platform with millions of daily audiences. Bang Zhang, Peng Zhang 0080, Jinwei Qi, Daiheng Gao, Haiming Zhao, Xiaoduan Feng, Qi Wang 0148, Lian Zhuo |
ACM Multimedia | 3 |
| 2015 | SOM: Semantic obviousness metric for image quality assessmentabstractImage quality assessment (IQA) tries to estimate human perception based image visual quality in an objective manner. Existing approaches target this problem with or without reference images. For no-reference image quality assessment, there is no given reference image or any knowledge of the distortion type of the image. Previous approaches measure the image quality from signal level rather than semantic analysis. They typically depend on various features to represent local characteristic of an image. In this paper we propose a new no-reference (NR) image quality assessment (IQA) framework based on semantic obviousness. We discover that semantic-level factors affect human perception of image quality. With such observation, we explore semantic obviousness as a metric to perceive objects of an image. We propose to extract two types of features, one to measure the semantic obviousness of the image and the other to discover local characteristic. Then the two kinds of features are combined for image quality estimation. The principles proposed in our approach can also be incorporated with many existing IQA algorithms to boost their performance. We evaluate our approach on the LIVE dataset. Our approach is demonstrated to be superior to the existing NR-IQA algorithms and comparable to the state-of-the-art full-reference IQA (FR-IQA) methods. Cross-dataset experiments show the generalization ability of our approach. Peng Zhang 0080, Wengang Zhou 0001, Lei Wu 0017, Houqiang Li |
CVPR | 1 |
| 2011 | A novel tracking-by-encoding scheme based on linear programming matchingabstractIn this paper, we present a new object tracking scheme. Different from conventional tracking process, the proposed method is combined with video encoding and can be easily embedded into video encoding system. The only information the tracker need is Summed Absolute Difference (SAD) generated by motion estimation during encoding. These SADs are used as soft votes in a fast Hough-Transform (HT) algorithm in parameter space of translation. This can be much faster than conventional tracking algorithm such as particle filter and more accurate than recent compress-domain approaches. Taking the advantage of HT, this approach is also robust to unstable camera, background moving and partial occlusion. The parameter space can be also extended to affine transform space with those soft votes to track scale change and partially rotation moving. We solve the latter multi-parameter problem by Linear Programming (LP) algorithm. The tracking result can also direct encoding behavior: detecting Region of Interest (ROI), adjusting QP, and etc. Peng Zhang 0080, Houqiang Li |
ISCAS | 2 |
| 2011 | Peak Tree: A New Tool for Multiscale Hierarchical Representation and Peak Detection of Mass Spectrometry DataabstractPeak detection is one of the most important steps in mass spectrometry (MS) analysis. However, the detection result is greatly affected by severe spectrum variations. Unfortunately, most current peak detection methods are neither flexible enough to revise false detection results nor robust enough to resist spectrum variations. To improve flexibility, we introduce peak tree to represent the peak information in MS spectra. Each tree node is a peak judgment on a range of scales, and each tree decomposition, as a set of nodes, is a candidate peak detection result. To improve robustness, we combine peak detection and common peak alignment into a closed-loop framework, which finds the optimal decomposition via both peak intensity and common peak information. The common peak information is derived and loopily refined from the density clustering of the latest peak detection result. Finally, we present an improved ant colony optimization biomarker selection method to build a whole MS analysis system. Experiment shows that our peak detection method can better resist spectrum variations and provide higher sensitivity and lower false detection rates than conventional methods. The benefits from our peak-tree-based system for MS disease analysis are also proved on real SELDI data. Peng Zhang 0080, Houqiang Li, Stephen T. C. Wong, Xiaobo Zhou 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | MAP spatial pyramid mean shift for object trackingabstractMean Shift is popular in object tracking due to its simplicity and efficiency. It finds local maximum of the similarity
measure between the target model and target candidate, and works well in many situations. However, it suffers from two
aspects. First, Mean Shift tracker ignores background knowledge. As a result, it may fail when the background color is
similar to that of the target or the initial target region contains too much background. Second, Mean Shift tracker omits
the geometric structure with a global color histogram as the target model. Therefore, it may not work in the case of
partial occlusion. To solve the first problem, we introduce background color histogram into a MAP formulation. To
address the second problem, we divide the target into hierarchical blocks. These blocks are described with a histogram
each but tracked as a whole. The two threads lead to a new algorithm, named MAP spatial pyramid (MAP-SP) Mean
Shift. The efficiency of MAP-SP Mean Shift is demonstrated via comparative experiments on both standard and our own
video sequences Xiaobo Han, Peng Zhang 0080, Houqiang Li |
VCIP | 2 |