EDBT 2026 Demo / reviewers in the wild / expert
Haozhe Wu
dblp:254/1647
· DBLP profile ↗
17ranked-venue papers
5as first author
13since 2021 · last 2024
0000-0002-3036-6930ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DanceCamera3D: 3D Camera Movement Synthesis with Music and DanceabstractChoreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have show-cased the feasibility of dance synthesis. However, camera movement synthesis with music and dance remains an un-solved challenging problem due to the scarcity of paired data. Thus, we present DCM, a new multi-modal 3D dataset, which for the first time combines camera movement with dance motion and music audio. This dataset encom-passes 108 dance sequences (3.2 hours) of paired dance-camera-music data from the anime community, covering 4 music genres. With this dataset, we uncover that dance camera movement is multifaceted and human-centric, and possesses multiple influencing factors, making dance camera synthesis a more challenging task compared to camera or dance synthesis alone. To overcome these difficulties, we propose DanceCamera3D, a transformer-based diffusion model that incorporates a novel body attention loss and a condition separation strategy. For evaluation, we devise new metrics measuring camera movement quality, diversity, and dancer fidelity. Utilizing these metrics, we conduct extensive experiments on our DCM dataset, providing both quantitative and qualitative evidence showcasing the effectiveness of our DanceCamera3D model. Code and video demos are available at https://github.com/Carmenw1203/DanceCamera3D-Official. Zixuan Wang 0026, Jia Jia 0001, Shikun Sun, Haozhe Wu, Jiaqing Zhou, Jiebo Luo 0001 |
CVPR | 4 |
| 2023 | An Urban Electric Load Forecasting Model Using Discrepancy Compensation and Short-term Sampling Contrastive Loss in LSTMabstractUrban electric load forecasting is an important content of urban electric system planning and dispatching. However, the problem of data imbalance in urban electric load forecasting leads to poor performance of single-model based methods. Existing multi-models based methods not only increase the cost of constructing models but also separate common time series features among different electric load profiles samples. In this paper, we propose an urban electric load forecasting model (DCSC-LSTM), which introduces the discrepancy compensation module and the short-term sampling contrastive loss to the Long Short-Term Memory. The DCSC-LSTM model uses the discrepancy compensation module to learn the discrepancies between samples with different electric load profiles, and the short-term sampling contrastive loss to regularize the training of the model. A series of experiments were conducted to validate the design of the DCSC-LSTM model. The experiments results show that the DCSC-LSTM model achieves competitive performance in the task of forecasting electric load. Runhuan Chen, Hua Dai 0003, Geng Yang 0002, Haozhe Wu |
CSCWD | 5 |
| 2023 | HSFV-based Action Recognition Using Recurrent Neural NetworksabstractHuman action recognition is one of the basic problems in the field of computer vision, which has a wide range of applications in video surveillance, sports analysis, medical care. Since the skeleton data can not be easily affected by background and views and has small computational cost, skeleton-based action recognition has attracted a lot of researchers’ interest. In recent years, some researchers have proposed to use the joints connection methods to mine the intrinsic correlation between skeleton joints as the feature representation of motion skeleton, and good results are achieved. However, there is still the problem of confusion when distinguishing actions with similar motion segments. To solve this problem, a human skeleton feature vector model (HSFV) is proposed in this paper. By constructing the feature extraction reference frame, the model uses OKS indicator to calculate and generate the feature vector describing the human posture. The feature vectors are input into the recurrent neural network, the experimental results on public and self-built datasets show that the human skeleton feature vector model based action recognition method proposed in this paper can distinguish actions with similar motion segments, and has the advantages of simple training and small calculation cost. It has broad application prospects in the fields of process detection, motion evaluation and so on. Haozhe Wu, Hua Dai 0003, Guineng Zheng, Xiaofei Ji |
CSCWD | 1 |
| 2023 | A Human Pose Similarity Calculation Method Based on Partition Weighted OKS ModelabstractImage-based calculation of human pose similarity is one of the computer vision research fields. Most existing research uses the human skeleton joint to calculate the human pose similarity, but usually does not consider the influence of inaccurate recognition of skeleton joint on similarity calculation caused by the complex environment (such as the occlusion of body parts, etc.). We propose a human pose similarity calculation method based on partition weighted OKS model. Due to the influence of external factors such as occlusion, the skeleton joint extracted by the human pose estimation algorithm is inaccurate, which leads to the decrease of the accuracy of the human pose similarity calculation. We propose the partition rule of human skeleton joints and the dynamic strategy adjustment of partition weight. The partition weighted OKS model and a human pose similarity calculation method based on the partition weighted OKS model are given. The experimental results on datasets show that the proposed method for human pose similarity calculation is superior to the traditional one. Hua Dai 0003, Haozhe Wu, Geng Yang 0002, Guineng Zheng |
CSCWD | 3 |
| 2023 | Shuffled Autoregression for Motion InterpolationabstractThis work aims to provide a deep-learning solution for the motion interpolation task. Previous studies solve it with geometric weight functions. Some other works propose neural networks for different problem settings with consecutive pose sequences as input. However, motion interpolation is a more complex problem that takes isolated poses (e.g., only one start pose and one end pose) as input. When applied to motion interpolation, these deep learning methods have limited performance since they do not leverage the flexible dependencies between interpolation frames as the original geometric formulas do. To realize this interpolation characteristic, we propose a novel framework, referred to as Shuffled AutoRegression, which expands the autoregression to generate in arbitrary (shuffled) order and models any inter-frame dependencies as a directed acyclic graph. We further propose an approach to constructing a particular kind of dependency graph, with three stages assembled into an end-to-end spatial-temporal motion Transformer. Experimental results on one of the current largest datasets show that our model generates vivid and coherent motions from only one start frame to one end frame and outperforms competing methods by a large margin. The proposed model is also extensible to multiple keyframes’ motion interpolation tasks and other areas’ interpolation. Shuo Huang 0005, Jia Jia 0001, Zongxin Yang, Wei Wang 0010, Haozhe Wu, Yi Yang 0001, Junliang Xing |
ICASSP | 5 |
| 2023 | MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style TransferabstractSentiment plays an essential role in people’s perception of images. To incorporate the sentiment information into the image style transfer task for better sentiment-aware performance, we introduce a new task named sentiment-aware image style transfer. To solve this problem, we first introduce a novel Multi-Sentiment Semantics Space (MSS-Space) to capture the non-deterministic and complicated nature of sentiment semantics. With the MSS-Space, we establish tight associations between the visual attributes of images and the multi-sentiment semantics by minimizing their distance in MSS-Space and then propose the Multi-Sentiment Style Transfer Net (MSNet). Experiments demonstrate that, compared with three competing models, our proposed MSNet generates more explicit images and better preserves the integrity of salient objects, local details, and multi-sentiment. In particular, our model outperforms the state-of-the-art by +28.72% in terms of the top-3 accuracy on average. Shikun Sun, Jia Jia 0001, Haozhe Wu, Zijie Ye, Junliang Xing |
ICASSP | 3 |
| 2023 | Salient Co-Speech Gesture Synthesizing with Discrete Motion RepresentationabstractSynthesizing co-speech gestures is challenging because the mapping from speech to gesticulation is inherently non-deterministic. When giving talks, people conduct not only gentle and rhythmic motions but also abrupt and salient gesticulations. Most previous research efforts, however, ignore this nature of co-speech gestures and synthesize deterministic results, producing over-smoothed movements with limited expressiveness. To address this issue, we propose a new co-speech gesture generation approach that produces high-quality salient gesticulations. Specifically, we build a discrete motion representation (DMR) space to bridge the speech-gesture mapping and the gesture generation stages. The incorporation of DMR enables random sampling in motion space and avoids the over-smooth problem in speech-gesture mapping. Based on DMR, we devise a novel multi-modal co-speech gesture synthesis model with temporal attention (MCGT). MCGT explicitly models DMR’s categorical distribution conditioned on the speech context, which captures complex context patterns and produces more salient gesticulations in sync with the context. In addition, we construct a new benchmark for evaluating salient motion quality in co-speech gestures, containing a large-scale co-speech gesture dataset with salient gesticulations. We also introduce a new metric, referred to as salient motion similarity, to evaluate the salient motion quality. Experiments demonstrate superior results from our approach over several competing baselines. Zijie Ye, Jia Jia 0001, Haozhe Wu, Shuo Huang 0005, Shikun Sun, Junliang Xing |
ICASSP | 3 |
| 2023 | Prosody Modeling with 3D Visual Information for Expressive Video Dubbing
Shansong Liu, Xu Li 0015, Haozhe Wu, Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
INTERSPEECH | 4 |
| 2023 | Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD SpaceabstractCreating realistic 3D facial animation is crucial for various applications in the movie production and gaming industry, especially with the burgeoning demand in the metaverse. However, prevalent methods such as blendshape-based approaches and facial rigging techniques are time-consuming, labor-intensive, and lack standardized configurations, making facial animation production challenging and costly. In this paper, we propose a novel self-supervised framework, Versatile Face Animator, which combines facial motion capture with motion retargeting in an end-to-end manner, eliminating the need for blendshapes or rigs. Our method has the following two main characteristics: 1) we propose an RGBD animation module to learn facial motion from raw RGBD videos by hierarchical motion dictionaries and animate RGBD images rendered from 3D facial mesh coarse-to-fine, enabling facial animation on arbitrary 3D characters regardless of their topology, textures, blendshapes, and rigs; and 2) we introduce a mesh retarget module to utilize RGBD animation to create 3D facial animation by manipulating facial mesh with controller transformations, which are estimated from dense optical flow fields and blended together with geodesic-distance-based weights. Comprehensive experiments demonstrate the effectiveness of our proposed framework in generating impressive 3D facial animation results, highlighting its potential as a promising solution for the cost-effective and efficient production of facial animation in the metaverse. Haoyu Wang 0009, Haozhe Wu, Junliang Xing, Jia Jia 0001 |
ACM Multimedia | 2 |
| 2023 | Speech-Driven 3D Face Animation with Composite and Regional Facial MovementsabstractSpeech-driven 3D face animation poses significant challenges due to the intricacy and variability inherent in human facial movements. This paper emphasizes the importance of considering both the composite and regional natures of facial movements in speech-driven 3D face animation. The composite nature pertains to how speech-independent factors globally modulate speech-driven facial movements along the temporal dimension. Meanwhile, the regional nature alludes to the notion that facial movements are not globally correlated but are actuated by local musculature along the spatial dimension. It is thus indispensable to incorporate both natures for engendering vivid animation. To address the composite nature, we introduce an adaptive modulation module that employs arbitrary facial movements to dynamically adjust speech-driven facial movements across frames on a global scale. To accommodate the regional nature, our approach ensures that each constituent of the facial features for every frame focuses on the local spatial movements of 3D faces. Moreover, we present a non-autoregressive backbone for translating audio to 3D facial movements, which maintains high-frequency nuances of facial movements and facilitates efficient inference. Comprehensive experiments and user studies demonstrate that our method surpasses contemporary state-of-the-art approaches both qualitatively and quantitatively. Haozhe Wu, Songtao Zhou, Jia Jia 0001, Junliang Xing |
ACM Multimedia | 1 |
| 2022 | GroupDancer: Music to Multi-People Dance Synthesis with Style CollaborationabstractDifferent people dance in different styles. So when multiple people dance together, the phenomenon of style collaboration occurs: people need to seek common points while reserving differences in various dancing periods. Thus, we introduce a novel Music-driven Group Dance Synthesis task. Compared with single-people dance synthesis explored by most previous works, modeling the style collaboration phenomenon and choreographing for multiple people are more complicated and challenging. Moreover, the lack of sufficient records for conducting multi-people choreography in prior datasets further aggravates this problem. To address these issues, we construct a rich-annotated 3D Multi-Dancer Choreography dataset (MDC) and newly devise a metric SCEU for style collaboration evaluation. To our best knowledge, MDC is the first 3D dance dataset that collects both individual and collaborated music-dance pairs. Based on MDC, we present a novel framework, GroupDancer, consisting of three stages: Dancer Collaboration, Motion Choreography and Motion Transition. The Dancer Collaboration stage determines when and which dancers should collaborate their dancing styles from music. Afterward, the Motion Choreography stage produces a motion sequence for each dancer. Finally, the Motion Transition stage fills the gaps between the motions to achieve fluent and natural group dance. To make GroupDancer trainable from end to end and able to synthesize group dance with style collaboration, we propose mixed training and selective updating strategies. Comprehensive evaluations on the MDC dataset demonstrate that the proposed GroupDancer model can synthesize quite satisfactory group dance synthesis results with style collaboration. Zixuan Wang 0026, Jia Jia 0001, Haozhe Wu, Junliang Xing, Jinghe Cai, Guowen Chen |
ACM Multimedia | 3 |
| 2021 | PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective FeedbackabstractSecond language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency. Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002 |
CHI | 9 |
| 2021 | Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisabstractPeople talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the "solemn" style is more standardized and seldomly exhibits exaggerated motions. Due to such huge differences between different styles, it is necessary to incorporate the talking style into audio-driven talking face synthesis framework. In this paper, we propose to inject style into the talking face synthesis framework through imitating arbitrary talking style of the particular reference video. Specifically, we systematically investigate talking styles with our collected Ted-HD dataset and construct style codes as several statistics of 3D morphable model (3DMM) parameters. Afterwards, we devise a latent-style-fusion (LSF) model to synthesize stylized talking faces by imitating talking styles from the style codes. We emphasize the following novel characteristics of our framework: (1) It doesn't require any annotation of the style, the talking style is learned in an unsupervised manner from talking videos in the wild. (2) It can imitate arbitrary styles from arbitrary videos, and the style codes can also be interpolated to generate new styles. Extensive experiments demonstrate that the proposed framework has the ability to synthesize more natural and expressive talking styles compared with baseline methods. Haozhe Wu, Jia Jia 0001, Haoyu Wang 0009, Yishun Dou, Qingshan Deng |
ACM Multimedia | 1 |
| 2020 | Mining Unfollow Behavior in Large-Scale Online Social Networks via Spatial-Temporal InteractionabstractOnline Social Networks (OSNs) evolve through two pervasive behaviors: follow and unfollow, which respectively signify relationship creation and relationship dissolution. Researches on social network evolution mainly focus on the follow behavior, while the unfollow behavior has largely been ignored. Mining unfollow behavior is challenging because user's decision on unfollow is not only affected by the simple combination of user's attributes like informativeness and reciprocity, but also affected by the complex interaction among them. Meanwhile, prior datasets seldom contain sufficient records for inferring such complex interaction. To address these issues, we first construct a large-scale real-world Weibo1 dataset, which records detailed post content and relationship dynamics of 1.8 million Chinese users. Next, we define user's attributes as two categories: spatial attributes (e.g., social role of user) and temporal attributes (e.g., post content of user). Leveraging the constructed dataset, we systematically study how the interaction effects between user's spatial and temporal attributes contribute to the unfollow behavior. Afterwards, we propose a novel unified model with heterogeneous information (UMHI) for unfollow prediction. Specifically, our UMHI model: 1) captures user's spatial attributes through social network structure; 2) infers user's temporal attributes through user-posted content and unfollow history; and 3) models the interaction between spatial and temporal attributes by the nonlinear MLP layers. Comprehensive evaluations on the constructed dataset demonstrate that the proposed UMHI model outperforms baseline methods by 16.44 on average in terms of precision. In addition, factor analyses verify that both spatial attributes and temporal attributes are essential for mining unfollow behavior. Haozhe Wu, Jia Jia 0001, Yaohua Bu, Xiangnan He 0001, Tat-Seng Chua |
AAAI | 1 |
| 2020 | Rethinking the Distribution Gap of Person Re-identification with Camera-Based Batch Normalization
Zijie Zhuang, Longhui Wei, Lingxi Xie, Hengheng Zhang, Haozhe Wu, Haizhou Ai, Qi Tian 0001 |
ECCV (12) | 6 |
| 2020 | Cross-VAE: Towards Disentangling Expression from Identity For Human FacesabstractFacial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discriminative features that are invariant of identities for expression recognition. However, such disentanglement normally requires annotations of both expression and identity on one large dataset, which is often unavailable. Our solution is to extend conditional VAE to a crossed version named Cross-VAE, which is able to use partially labeled data to disentangle expression from identity. We emphasis the following novel characteristics of our Cross-VAE: (1) It is based on an independent assumption that the two latent representations' distributions are orthogonal. This ensures both encoded representations to be disentangled and expressive. (2) It utilizes a symmetric training procedure where the output of each encoder is fed as the condition of the other. Thus two partially labeled sets can be jointly used. Extensive experiments show that our proposed method is capable of encoding expressive and disentangled features for facial expression. Compared with the baseline methods, our model shows an improvement of 3.56% on average in terms of accuracy. Haozhe Wu, Jia Jia 0001, Lingxi Xie, Guo-Jun Qi, Yuanchun Shi, Qi Tian 0001 |
ICASSP | 1 |
| 2020 | ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action UnitabstractDance and music are two highly correlated artistic forms. Synthesizing dance motions has attracted much attention recently. Most previous works conduct music-to-dance synthesis via directly music to human skeleton keypoints mapping. Meanwhile, human choreographers design dance motions from music in a two-stage manner: they firstly devise multiple choreographic dance units (CAUs), each with a series of dance motions, and then arrange the CAU sequence according to the rhythm, melody and emotion of the music. Inspired by these, we systematically study such two-stage choreography approach and construct a dataset to incorporate such choreography knowledge. Based on the constructed dataset, we design a two-stage music-to-dance synthesis framework ChoreoNet to imitate human choreography procedure. Our framework firstly devises a CAU prediction model to learn the mapping relationship between music and CAU sequences. Afterwards, we devise a spatial-temporal inpainting model to convert the CAU sequence into continuous dance motions. Experimental results demonstrate that the proposed ChoreoNet outperforms baseline methods (0.622 in terms of CAU BLEU score and 1.59 in terms of user study score). Zijie Ye, Haozhe Wu, Jia Jia 0001, Yaohua Bu, Wei Chen 0071 |
ACM Multimedia | 2 |