Wenjin Deng

dblp:276/3639 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2023
0000-0002-9213-6022ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video
abstract
For few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos. Previous works fail to generate high-fidelity dubbing results. To address the above problem, this paper proposes a Deformation Inpainting Network (DINet) for high-resolution face visually dubbing. Different from previous works relying on multiple up-sample layers to directly generate pixels from latent embeddings, DINet performs spatial deformation on feature maps of reference images to better preserve high-frequency textural details. Specifically, DINet consists of one deformation part and one inpainting part. In the first part, five reference facial images adaptively perform spatial deformation to create deformed feature maps encoding mouth shapes at each frame, in order to align with input driving audio and also the head poses of input source images. In the second part, to produce face visually dubbing, a feature decoder is responsible for adaptively incorporating mouth movements from the deformed feature maps and other attributes (i.e., head pose and upper facial expression) from the source feature maps together. Finally, DINet achieves face visually dubbing with rich textural details. We conduct qualitative and quantitative comparisons to validate our DINet on high-resolution videos. The experimental results show that our method outperforms state-of-the-art works.
Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, Yu Ding 0001
AAAI3
2023 ROI-demand Traffic Prediction: A Pre-train, Query and Fine-tune Framework
abstract
Traffic prediction has drawn increasing attention due to its essential role in smart city applications. To achieve precise predictions, a large number of approaches have been proposed to model spatial dependencies and temporal dynamics. Despite their superior performance, most existing studies focus datasets that are usually in large geographic scales, e.g., citywide, while ignoring the results on specific regions. However, in many scenarios, for example, route planning on time-dependent road networks, only small regions are of interest. We name the task of answering forecasting requests from any query region of interest (ROI) as ROI-demand traffic prediction (RTP). In this paper, we make a primary observation that existing methods fail to jointly achieve effectiveness and efficiency for RTP. To address this issue, a novel model-agnostic framework based on pre-Training, Querying and fine-Tuning, named TQT, is proposed, which first customizes input data given an ROI, and then makes fast adaptation from pre-trained traffic prediction backbone models by fine-tuning. We evaluate TQT on two real-world traffic datasets, performing both flow and speed prediction tasks. Extensive experiment results demonstrate the effectiveness and efficiency of the proposed method.
Yue Cui 0001, Shuhao Li 0001, Wenjin Deng, Zhaokun Zhang, Jing Zhao 0040, Kai Zheng 0001, Xiaofang Zhou 0001
ICDE3
2023 Vertex position estimation with spatial-temporal transformer for 3D human reconstruction
abstract
Reconstructing 3D human pose and body shape from monocular images or videos is a fundamental task for comprehending human dynamics. Frame-based methods can be broadly categorized into two fashions: those regressing parametric model parameters (e.g., SMPL) and those exploring alternative representations (e.g., volumetric shapes, 3D coordinates). Non-parametric representations have demonstrated superior performance due to their enhanced flexibility. However, when applied to video data, these non-parametric frame-based methods tend to generate inconsistent and unsmooth results. To this end, we present a novel approach that directly regresses the 3D coordinates of the mesh vertices and body joints with a spatial–temporal Transformer. In our method, we introduce a SpatioTemporal Learning Block (STLB) with Spatial Learning Module (SLM) and Temporal Learning Module (TLM), which leverages spatial and temporal information to model interactions at a finer granularity, specifically at the body token level. Our method outperforms previous state-of-the-art approaches on Human3.6M and 3DPW benchmark datasets.
Xiangjun Zhang, Yinglin Zheng, Wenjin Deng, Qifeng Dai, Wangzheng Shi, Ming Zeng 0008
Graph. Model.3
2023 MusicFace: Music-driven expressive singing face synthesis
abstract
It remains an interesting and challenging problem to synthesize a vivid and realistic singing face driven by music. In this paper, we present a method for this task with natural motions for the lips, facial expression, head pose, and eyes. Due to the coupling of mixed information for the human voice and backing music in common music audio signals, we design a decouple-and-fuse strategy to tackle the challenge. We first decompose the input music audio into a human voice stream and a backing music stream. Due to the implicit and complicated correlation between the two-stream input signals and the dynamics of the facial expressions, head motions, and eye states, we model their relationship with an attention scheme, where the effects of the two streams are fused seamlessly. Furthermore, to improve the expressivenes of the generated results, we decompose head movement generation in terms of speed and direction, and decompose eye state generation into short-term blinking and long-term eye closing, modeling them separately. We have also built a novel dataset, SingingFace, to support training and evaluation of models for this task, including future work on this topic. Extensive experiments and a user study show that our proposed method is capable of synthesizing vivid singing faces, qualitatively and quantitatively better than the prior state-of-the-art.
Wenjin Deng, Hengda Li, Jintai Wang, Yinglin Zheng, Yiwei Ding, Xiaohu Guo, Ming Zeng 0008
Comput. Vis. Media2
2022 I²R-Net: Intra- and Inter-Human Relation Network for Multi-Person Pose Estimation
abstract
In this paper, we present the Intra- and Inter-Human Relation Networks I²R-Net for Multi-Person Pose Estimation. It involves two basic modules. First, the Intra-Human Relation Module operates on a single person and aims to capture Intra-Human dependencies. Second, the Inter-Human Relation Module considers the relation between multiple instances and focuses on capturing Inter-Human interactions. The Inter-Human Relation Module can be designed very lightweight by reducing the resolution of feature map, yet learn useful relation information to significantly boost the performance of the Intra-Human Relation Module. Even without bells and whistles, our method can compete or outperform current competition winners. We conduct extensive experiments on COCO, CrowdPose, and OCHuman datasets. The results demonstrate that the proposed model surpasses all the state-of-the-art methods. Concretely, the proposed method achieves 77.4% AP on CrowPose dataset and 67.8% AP on OCHuman dataset respectively, outperforming existing methods by a large margin. Additionally, the ablation study and visualization analysis also prove the effectiveness of our model.
Yiwei Ding, Wenjin Deng, Yinglin Zheng, Meihong Wang, Jianmin Bao, Dong Chen 0003, Ming Zeng 0008
IJCAI2
2022 Improving Person Re-identification with Semantically Aligned Appearance Transformer
abstract
Person re-identification (re-id) has achieved significant improvement under the convolution neural network (CNN)-based methods. But it suffers from information loss on details caused by convolution and downsampling operators. Recently, there has been a growing interest in conducting image classification tasks by applying transformer-based methods to overcome these limitations. However, it is still limited by the body misalignment problem caused by pose/viewpoint variations, imperfect person detection, occlusion, etc. To address these problems, we leverage the estimation of the dense semantics of a person image to construct a set of densely human appearance semantically aligned images (DHASA-images), where the same spatial positions have the same semantics across different images. In this paper, we take both original images and DHASA-images into consideration to obtain the discriminative representation of pedestrians. We propose a novel approach named appearance semantically aligned guiding network (ASAG-Net) for person reid based on the transformer. The network is composed of two subnetworks: one for feature extraction from two kinds of images and another for feature fusion and obtain the final feature final feature generation. To the best of our knowledge, we are the first to make use of densely human appearance semantically aligned to strengthening the transformer-based person re-id method. We demonstrate the proposed method through extensive experiments and achieve superior results on three large-scale person re-id datasets, Market1501, DukeMTMC-reID, and MSMT17.
Yinglin Zheng, Zhaodong Tan, Wenjin Deng
IJCNN4
2022 Controllable Facial Caricaturization With Localized Deformation and Personalized Semantic Attentions
abstract
The facial caricature shows the distinct characteristics of a person via exaggerations of both shape and appearance. This paper presents a novel framework that automatically generates vivid facial caricatures by encoding personalized semantic information. To this end, we first design a part-based scheme for geometry warping, which composes local semantic deformation into a global warping field, equipped with sufficient warping freedom of different facial components. Second, under the scheme of Part-based Warping, we design a photo-to-caricature translation network called PbWarpGAN, and adopt several novel losses to capture the personalized characteristics of each input face and preserve its identity better. Third, based on PbWarpGAN, we develop a user-friendly interface by introducing an attention scheme on each facial component, allowing ordinary users to adjust the automatically generated caricature by PbWarpGAN according to their preference conveniently. Experimental results show that our PbWarpGAN is more effective in capturing personalized characteristics than counterparts, and provides an efficient tool for caricature designing application.
Ming Zeng 0008, Yinglin Zheng, Jinpeng Lin, Jing Liao 0001, Zizhao Wu, Wenjin Deng
IEEE Trans. Multim.7
2021 Real-Time Masked Face Revealing for Video Conference
abstract
Video conferencing is an essential way for contactless conversation, which conveys abundant multimedia signals. Especially under COVID-19, the video conference has been becoming a common way for daily communications. However, for the sake of plague prevention, it usually happens that the people attending the video conference are wearing a mouth mask, leading to inconvenient communication due to incomplete facial information. To tackle this problem, we develop a novel system that reveals the masked faces in real-time, making each participant feel like the others are mask-free. Moreover, we map the audio to 3DMM expression to guide the generation of various mouth shapes utilizing multi-modal information. Extensive experiments validate the revealing effectiveness and better user experience of the system. Furthermore, by applying lightweight networks design, the proposed system can run in real-time.
Jinpeng Lin, Yinglin Zheng, Wenjin Deng, Ming Zeng 0008
ICME4
2020 VH3D-LSFM: Video-Based Human 3D Pose Estimation with Long-Term and Short-Term Pose Fusion Mechanism
Wenjin Deng, Yinglin Zheng, Zizhao Wu, Ming Zeng 0008
PRCV (1)1