EDBT 2026 Demo / reviewers in the wild / expert
Yinglin Zheng
dblp:264/8457
· DBLP profile ↗
20ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0003-4671-6111ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual SignalsabstractThis paper addresses the gap in predicting turn-taking and backchannel actions in humanmachine conversations using multi-modal signals (linguistic, acoustic, and visual).To overcome the limitation of existing datasets, we propose an automatic data collection pipeline that allows us to collect and annotate over 210 hours of human conversation videos.From this, we construct a Multi-Modal Face-to-Face (MM-F2F) human conversation dataset, including over 1.5M words and corresponding turntaking and backchannel annotations from approximately 20M frames.Additionally, we present an end-to-end framework that predicts the probability of turn-taking and backchannel actions from multi-modal signals.The proposed model emphasizes the interrelation between modalities and supports any combination of text, audio, and video inputs, making it adaptable to a variety of realistic scenarios.Our experiments show that our approach achieves state-of-the-art performance on turn-taking and backchannel prediction tasks, achieving a 10% increase in F1-score on turn-taking and a 33% increase on backchannel prediction.Our dataset and code are publicly available online to ease of subsequent research.The code and dataset are available at https://github.com/Linyx1125/MM-F2F. Yinglin Zheng, Ming Zeng 0008, Wangzheng Shi |
ACL (1) | 2 |
| 2025 | EmoHuman: Fine-Grained Emotion-Controlled Talking Head Generation via Audio-Text Multimodal DetanglingabstractAudio-driven talking head generation has made significant strides in creating realistic and lip-synchronized portraits. However, most existing approaches overlook facial expressions, with only a few attempting to model facial emotions explicitly, often leading to unnatural results. To address this gap, we introduce EmoHuman, an audio-to-video synthesis method that generates emotionally nuanced talking head videos without relying on intermediate 3D representations or facial landmarks. EmoHuman decouples content, emotion, and emotional intensity from the multimodal information of the audio and the corresponding textual content through an Audio Emotion Decoupling Module. The content features are used to drive a powerful video diffusion model, generating synchronized lip movements, while emotion and emotional intensity govern the simulation of facial expressions, resulting in more realistic video outputs. Extensive experiments demonstrate that EmoHuman outperforms state-of-the-art methods in image and video quality, expression correlation, and lip-synchronization accuracy. Qifeng Dai, Huidong Feng, Wendi Cui, Xinqi Cai, Yinglin Zheng, Ming Zeng 0008 |
ICMR | 5 |
| 2025 | Consistent Human Animation with Pseudo Multi-View Anchoring and Cross-Granularity Integration
Jintai Wang, Yinglin Zheng, Qifeng Dai, Ming Zeng 0008 |
ICMR | 2 |
| 2025 | HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided AnimationabstractHair transfer is increasingly valuable across domains such as social media, gaming, advertising, and entertainment. While significant progress has been made in single-image hair transfer, video-based hair transfer remains challenging due to the need for temporal consistency, spatial fidelity, and dynamic adaptability. In this work, we propose HairShifter, a novel ''Anchor Frame + Animation'' framework that unifies high-quality image hair transfer with smooth and coherent video animation. At its core, HairShifter integrates a Image Hair Transfer (IHT) module for precise per-frame transformation and a Multi-Scale Gated SPADE Decoder to ensure seamless spatial blending and temporal coherence. Our method maintains hairstyle fidelity across frames while preserving non-hair regions. Extensive experiments demonstrate that HairShifter achieves state-of-the-art performance in video hairstyle transfer, combining superior visual quality, temporal consistency, and scalability. The code will be publicly available. We believe this work will open new avenues for video-based hairstyle transfer and establish a robust baseline in this field. Wangzheng Shi, Yinglin Zheng, Jianmin Bao, Ming Zeng 0008, Dong Chen 0003 |
ACM Multimedia | 2 |
| 2025 | PointHuman: Learning high-fidelity and generalizable human neural radiance fields using guidance of fine-grained semantics-enriched geometry
Jintai Wang, Huidong Feng, Qifeng Dai, Yinglin Zheng, Ming Zeng 0008 |
Comput. Graph. | 5 |
| 2024 | Multi-Modal Gait Recognition with Unidirectional Cross-modal AlignmentabstractGait recognition represents a pivotal challenge in visual signal comprehension, encompassing multi-modal spatial-temporal information, and exhibiting considerable complexity. The prevailing gait recognition methods are categorized into appearance-based and model-based, which commonly use silhouette and skeleton as input respectively. Certain recent studies have endeavored to integrate both of these modalities, achieving a certain degree of success in doing so. However, the aforementioned methods either rely solely on a modal of information or lack a profound consideration of the intricate interplay among multiple modal. Consequently, the potential inherent in gait’s spatial-temporal information remains untapped. To address this issue, this study strives to enhance the efficiency of data utilization through the integration of multi-modal information within a multi-modal fusion network. Furthermore, we propose a scheme aimed at enhancing the consistency between distinct modal features. Experiments on the widely used gait dataset CASIA-B have shown that our model has significantly improved under complex gait conditions, with overall performance reaching the most advanced level. Hengda Li, Yinglin Zheng, Qifeng Dai, Jintai Wang, Ming Zeng 0008 |
ICME | 2 |
| 2024 | High-fidelity instructional fashion image editingabstractInstructional image editing has received a significant surge of attention recently. In this work, we are interested in the challenging problem of instructional image editing within the particular fashion realm, a domain with significant potential demand in both commercial and personal contexts. This specific domain presents heightened challenges owing to the stringent quality requirements. It necessitates not only the creation of vivid details in alignment with instructions, but also the preservation of precise attributes unrelated to the text guidance. Naive extensions of existing image editing methods produce noticeable artifacts. In order to achieve high-fidelity fashion editing, we propose a novel framework, leveraging the generative prior of a pre-trained human generator and performing edit in the latent space. In addition, we introduce a novel CLIP-based loss to better align the generated target with the instruction. Extensive experiments demonstrate that our approach outperforms prior works including GAN-based editing as well as diffusion-based editing by a large margin, showing impressive visual quality. Yinglin Zheng, Ting Zhang 0002, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008 |
Graph. Model. | 1 |
| 2024 | FaceRefiner: High-Fidelity Facial Texture Refinement With Differentiable Rendering-Based Style TransferabstractRecent facial texture generation methods prefer to use deep networks to synthesize image content and then fill in the UV map, thus generating a compelling full texture from a single image. Nevertheless, the synthesized texture UV map usually comes from a space constructed by the training data or the 2D face generator, which limits the methods' generalization ability for in-the-wild input images. Consequently, their facial details, structures and identity may not be consistent with the input. In this paper, we address this issue by proposing a style transfer-based facial texture refinement method named FaceRefiner. FaceRefiner treats the 3D sampled texture asstyleand the output of a texture generation method ascontent. The photo-realistic style is then expected to be transferred from the style image to the content image. Different from current style transfer methods that only transfer high and middle level information to the result, our style transfer method integrates differentiable rendering to also transfer low level (or pixel level) information in the visible face regions. The main benefit of suchmulti-levelinformation transfer is that, the details, structures and semantics in the input can thus be well preserved. The extensive experiments on Multi-PIE, CelebA and FFHQ datasets demonstrate that our refinement method can improve the texture quality and the face identity preserving ability, compared with state-of-the-arts. Baoping Cheng, Yao Cheng 0005, Haocheng Zhang, Renshuai Liu, Yinglin Zheng, Jing Liao 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | EMEF: Ensemble Multi-Exposure Image FusionabstractAlthough remarkable progress has been made in recent years, current multi-exposure image fusion (MEF) research is still bounded by the lack of real ground truth, objective evaluation function, and robust fusion strategy. In this paper, we study the MEF problem from a new perspective. We don’t utilize any synthesized ground truth, design any loss function, or develop any fusion strategy. Our proposed method EMEF takes advantage of the wisdom of multiple imperfect MEF contributors including both conventional and deep learning-based methods. Specifically, EMEF consists of two main stages: pre-train an imitator network and tune the imitator in the runtime. In the first stage, we make a unified network imitate different MEF targets in a style modulation way. In the second stage, we tune the imitator network by optimizing the style code, in order to find an optimal fusion result for each input pair. In the experiment, we construct EMEF from four state-of-the-art MEF methods and then make comparisons with the individuals and several other competitive methods on the latest released MEF benchmark dataset. The promising experimental results demonstrate that our ensemble framework can “get the best of all worlds”. The code is available at https://github.com/medalwill/EMEF. Renshuai Liu, Haitao Cao 0006, Yinglin Zheng, Ming Zeng 0008 |
AAAI | 4 |
| 2023 | MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image PretrainingabstractThis paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a masked image. Such incorporation enjoys two vital benefits. First, masked self-distillation targets local patch representation learning, which is complementary to vision-language contrastive focusing on text-related representation. Second, masked self-distillation is also consistent with vision-language contrastive from the perspective of training objective as both utilize the visual encoder for feature aligning, and thus is able to learn local semantics getting indirect supervision from the language. We provide specially designed experiments with a comprehensive analysis to validate the two benefits. Symmetrically, we also introduce the local semantic supervision into the text branch, which further improves the pretraining performance. With extensive experiments, we show that MaskCLIP, when applied to various challenging downstream tasks, achieves superior results in linear probing, finetuning, and zeroshot performance with the guidance of the language encoder. Code will be release at https://github.com/LightDXY/MaskCLIP. Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang 0002, Dongdong Chen 0001, Hao Yang 0036, Ming Zeng 0008, Weiming Zhang 0001, Lu Yuan 0001, Dong Chen 0003, Fang Wen 0001, Nenghai Yu |
CVPR | 3 |
| 2023 | Vertex position estimation with spatial-temporal transformer for 3D human reconstructionabstractReconstructing 3D human pose and body shape from monocular images or videos is a fundamental task for comprehending human dynamics. Frame-based methods can be broadly categorized into two fashions: those regressing parametric model parameters (e.g., SMPL) and those exploring alternative representations (e.g., volumetric shapes, 3D coordinates). Non-parametric representations have demonstrated superior performance due to their enhanced flexibility. However, when applied to video data, these non-parametric frame-based methods tend to generate inconsistent and unsmooth results. To this end, we present a novel approach that directly regresses the 3D coordinates of the mesh vertices and body joints with a spatial–temporal Transformer. In our method, we introduce a SpatioTemporal Learning Block (STLB) with Spatial Learning Module (SLM) and Temporal Learning Module (TLM), which leverages spatial and temporal information to model interactions at a finer granularity, specifically at the body token level. Our method outperforms previous state-of-the-art approaches on Human3.6M and 3DPW benchmark datasets. Xiangjun Zhang, Yinglin Zheng, Wenjin Deng, Qifeng Dai, Wangzheng Shi, Ming Zeng 0008 |
Graph. Model. | 2 |
| 2023 | MusicFace: Music-driven expressive singing face synthesisabstractIt remains an interesting and challenging problem to synthesize a vivid and realistic singing face driven by music. In this paper, we present a method for this task with natural motions for the lips, facial expression, head pose, and eyes. Due to the coupling of mixed information for the human voice and backing music in common music audio signals, we design a decouple-and-fuse strategy to tackle the challenge. We first decompose the input music audio into a human voice stream and a backing music stream. Due to the implicit and complicated correlation between the two-stream input signals and the dynamics of the facial expressions, head motions, and eye states, we model their relationship with an attention scheme, where the effects of the two streams are fused seamlessly. Furthermore, to improve the expressivenes of the generated results, we decompose head movement generation in terms of speed and direction, and decompose eye state generation into short-term blinking and long-term eye closing, modeling them separately. We have also built a novel dataset, SingingFace, to support training and evaluation of models for this task, including future work on this topic. Extensive experiments and a user study show that our proposed method is capable of synthesizing vivid singing faces, qualitatively and quantitatively better than the prior state-of-the-art. Wenjin Deng, Hengda Li, Jintai Wang, Yinglin Zheng, Yiwei Ding, Xiaohu Guo, Ming Zeng 0008 |
Comput. Vis. Media | 5 |
| 2022 | General Facial Representation Learning in a Visual-Linguistic MannerabstractHow to learn a universal facial representation that boosts all face analysis tasks? This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learning. On one hand, the framework involves a contrastive loss to learn high-level semantic meaning from image-text pairs. On the other hand, we propose exploring low-level information simultaneously to further enhance the face representation by adding a masked image modeling. We perform pre-training on LAION-FACE, a dataset containing a large amount of face image-text pairs, and evaluate the representation capability on multiple downstream tasks. We show that FaRL achieves better transfer performance compared with previous pre-trained models. We also verify its superiority in the low-data regime. More importantly, our model surpasses the state-of-the-art methods on face analysis tasks including face parsing and face alignment. Yinglin Zheng, Hao Yang 0036, Ting Zhang 0002, Jianmin Bao, Dongdong Chen 0001, Yangyu Huang, Lu Yuan 0001, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
CVPR | 1 |
| 2022 | I²R-Net: Intra- and Inter-Human Relation Network for Multi-Person Pose EstimationabstractIn this paper, we present the Intra- and Inter-Human Relation Networks I²R-Net for Multi-Person Pose Estimation. It involves two basic modules. First, the Intra-Human Relation Module operates on a single person and aims to capture Intra-Human dependencies. Second, the Inter-Human Relation Module considers the relation between multiple instances and focuses on capturing Inter-Human interactions. The Inter-Human Relation Module can be designed very lightweight by reducing the resolution of feature map, yet learn useful relation information to significantly boost the performance of the Intra-Human Relation Module. Even without bells and whistles, our method can compete or outperform current competition winners. We conduct extensive experiments on COCO, CrowdPose, and OCHuman datasets. The results demonstrate that the proposed model surpasses all the state-of-the-art methods. Concretely, the proposed method achieves 77.4% AP on CrowPose dataset and 67.8% AP on OCHuman dataset respectively, outperforming existing methods by a large margin. Additionally, the ablation study and visualization analysis also prove the effectiveness of our model. Yiwei Ding, Wenjin Deng, Yinglin Zheng, Meihong Wang, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008 |
IJCAI | 3 |
| 2022 | Improving Person Re-identification with Semantically Aligned Appearance TransformerabstractPerson re-identification (re-id) has achieved significant improvement under the convolution neural network (CNN)-based methods. But it suffers from information loss on details caused by convolution and downsampling operators. Recently, there has been a growing interest in conducting image classification tasks by applying transformer-based methods to overcome these limitations. However, it is still limited by the body misalignment problem caused by pose/viewpoint variations, imperfect person detection, occlusion, etc. To address these problems, we leverage the estimation of the dense semantics of a person image to construct a set of densely human appearance semantically aligned images (DHASA-images), where the same spatial positions have the same semantics across different images. In this paper, we take both original images and DHASA-images into consideration to obtain the discriminative representation of pedestrians. We propose a novel approach named appearance semantically aligned guiding network (ASAG-Net) for person reid based on the transformer. The network is composed of two subnetworks: one for feature extraction from two kinds of images and another for feature fusion and obtain the final feature final feature generation. To the best of our knowledge, we are the first to make use of densely human appearance semantically aligned to strengthening the transformer-based person re-id method. We demonstrate the proposed method through extensive experiments and achieve superior results on three large-scale person re-id datasets, Market1501, DukeMTMC-reID, and MSMT17. Yinglin Zheng, Zhaodong Tan, Wenjin Deng |
IJCNN | 2 |
| 2022 | Controllable Facial Caricaturization With Localized Deformation and Personalized Semantic AttentionsabstractThe facial caricature shows the distinct characteristics of a person via exaggerations of both shape and appearance. This paper presents a novel framework that automatically generates vivid facial caricatures by encoding personalized semantic information. To this end, we first design a part-based scheme for geometry warping, which composes local semantic deformation into a global warping field, equipped with sufficient warping freedom of different facial components. Second, under the scheme of Part-based Warping, we design a photo-to-caricature translation network called PbWarpGAN, and adopt several novel losses to capture the personalized characteristics of each input face and preserve its identity better. Third, based on PbWarpGAN, we develop a user-friendly interface by introducing an attention scheme on each facial component, allowing ordinary users to adjust the automatically generated caricature by PbWarpGAN according to their preference conveniently. Experimental results show that our PbWarpGAN is more effective in capturing personalized characteristics than counterparts, and provides an efficient tool for caricature designing application. Ming Zeng 0008, Yinglin Zheng, Jinpeng Lin, Jing Liao 0001, Zizhao Wu, Wenjin Deng |
IEEE Trans. Multim. | 2 |
| 2021 | Exploring Temporal Coherence for More General Video Face Forgery DetectionabstractAlthough current face manipulation techniques achieve impressive performance regarding quality and controllability, they are struggling to generate temporal coherent face videos. In this work, we explore to take full advantage of the temporal coherence for video face forgery detection. To achieve this, we propose a novel end-to-end framework, which consists of two major stages. The first stage is a fully temporal convolution network (FTCN). The key insight of FTCN is to reduce the spatial convolution kernel size to 1, while maintaining the temporal convolution kernel size un-changed. We surprisingly find this special design can benefit the model for extracting the temporal features as well as improve the generalization capability. The second stage is a Temporal Transformer network, which aims to explore the long-term temporal coherence. The proposed frame-work is general and flexible, which can be directly trained from scratch without any pre-training models or external datasets. Extensive experiments show that our framework outperforms existing methods and remains effective when applied to detect new sorts of face forgery videos. Yinglin Zheng, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008, Fang Wen 0001 |
ICCV | 1 |
| 2021 | Real-Time Masked Face Revealing for Video ConferenceabstractVideo conferencing is an essential way for contactless conversation, which conveys abundant multimedia signals. Especially under COVID-19, the video conference has been becoming a common way for daily communications. However, for the sake of plague prevention, it usually happens that the people attending the video conference are wearing a mouth mask, leading to inconvenient communication due to incomplete facial information. To tackle this problem, we develop a novel system that reveals the masked faces in real-time, making each participant feel like the others are mask-free. Moreover, we map the audio to 3DMM expression to guide the generation of various mouth shapes utilizing multi-modal information. Extensive experiments validate the revealing effectiveness and better user experience of the system. Furthermore, by applying lightweight networks design, the proposed system can run in real-time. Jinpeng Lin, Yinglin Zheng, Wenjin Deng, Ming Zeng 0008 |
ICME | 3 |
| 2020 | VH3D-LSFM: Video-Based Human 3D Pose Estimation with Long-Term and Short-Term Pose Fusion Mechanism
Wenjin Deng, Yinglin Zheng, Zizhao Wu, Ming Zeng 0008 |
PRCV (1) | 2 |
| 2020 | Joint learning for face alignment and face transfer with depth image
Xiaoli Wang 0002, Yinglin Zheng, Ming Zeng 0008, Wei Lu 0015 |
Multim. Tools Appl. | 2 |