Hanqin Wang

dblp:152/9443 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask Learning
abstract
We introduce T2V2 (**T**ext to **V**oice and **V**oice to **T**ext), a unified non-autoregressive model capable of performing both automatic speech recognition (ASR) and text-to-speech (TTS) synthesis within the same framework. T2V2 uses a shared Conformer backbone with rotary positional embeddings to efficiently handle these core tasks, with ASR trained using Connectionist Temporal Classification (CTC) loss and TTS using masked language modeling (MLM) loss. The model operates on discrete tokens, where speech tokens are generated by clustering features from a self-supervised learning model. To further enhance performance, we introduce auxiliary tasks: CTC error correction to refine raw ASR outputs using contextual information from speech embeddings, and unconditional speech MLM, enabling classifier free guidance to improve TTS. Our method is self-contained, leveraging intermediate CTC outputs to align text and speech using Monotonic Alignment Search, without relying on external aligners. We perform extensive experimental evaluation to verify the efficacy of the T2V2 framework, achieving state-of-the-art performance on TTS task and competitive performance in discrete ASR.
Nabarun Goswami, Hanqin Wang, Tatsuya Harada
ICLR2
2025 ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model
abstract
Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow generation speed limits their application potential. In this paper, we introduce a novel autoregressive model that achieves real-time generation of highly synchronized lip movements and realistic head poses and eye blinks by learning a mapping from speech to a multi-scale motion codebook. Furthermore, our model can adapt to unseen speaking styles, enabling the creation of 3D talking avatars with unique personal styles beyond the identities seen during training. Extensive evaluations and user studies demonstrate that our method outperforms existing approaches in lip synchronization accuracy and perceived quality. Demos and codes are available at https://xg-chu.site/project_artalk/.
Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, Tatsuya Harada
SIGGRAPH Asia4
2025 Visual signatures for music mood and timbre
Hanqin Wang, Alexei Sourin
Vis. Comput.1
2025 Sound signatures for images and geometric shapes
Hanqin Wang, Alexei Sourin
Vis. Comput.1
2024 An Objective Metric Towards Music Mood Visualization
abstract
Music visualization is presented in different formats. The mainstream music visualization remains focused on providing animated images for visual pleasure or for enhancing the perception of sheet music with colors. Static visualization of music, on the other hand, has the potential to provide quick insights into the music’s attributes, particularly its mood. However, despite our previous work exploring several approaches to music mood visualization, a major problem still exists in how to objectively evaluate the mood of music pieces and the generated images. Thus, this paper proposes IMEMNextan improvement to the existing framework IMEMNet [1], which serves as a novel and objective metric for music mood visualization. Not only does it allow for a user study-free evaluation of music mood visualization, but the new network also enables for creation of datasets of music-image pairs with the same mood. The new IMEMNext network achieved a significant performance improvement compared to its original implementation.
Hanqin Wang, Alexei Sourin
CW1
2023 Deep Learning-based Visualization of Music Mood
abstract
Music has been visualized in different forms. Majority of the existing methods of music visualization utilized only a select few parameters of the music, of which pitch and frequency were the most visualized. Also, current musk visualizers usually produce animated visual backgrounds for the music being played. Visualization of music as static images was only addressed by some people who claimed to have a neurological condition called synesthesia or chromesthesia. In this paper, we use artificial intelligence to simulate such music visualization. Specifically, we consider how the mood of music can be visualized in static images using deep learning techniques. We consider two approaches. First, we use the deep learning network to generate abstract paintings based on the music sentiments obtained from the Spotify library. In the second approach, two different deep learning networks are used to both classify the musk and to generate the respective landscape images representing its mood. The results are analyzed and compared as well as subjected to user tests.
Hanqin Wang, Alexei Sourin
CW1
2022 Feasibility Study on Interactive Geometry Sonification
abstract
We performed a feasibility study on new ways of using sound in interactive computer graphics to improve visual interaction as well as to replace or augment haptic interaction. We considered using sound during surface interaction tasks performed with a desktop haptic device and an optical hand tracking device and evaluated the efficiency of such interactive geometry sonification in terms of its precision and speed. We also considered scenarios of using sound for easing navigation while moving along a path or surface.
Hanqin Wang, Alexei Sourin
CW1
2014 Modeling for user interaction by influence transfer effect in online social networks
abstract
User interaction is one of the most important features of online social networks, and is the basis of research of user behavior analysis, information spreading model, etc. However, existing approaches focus on the interactions between adjacent nodes, which do not fully take the interactions and relationship between local region users into consideration as well as the details of interaction process. In this paper, we find that there exists influence transfer effect in the process of user interactions, and present a regional user interaction model to analyze and understand interactions between users in a local region by influence transfer effect. Based on real data from Sina Weibo, we validate the effectiveness of our model by the experiments of user type classification, influential user identification and zombie user identification in online social networks. The experimental results show that our model present better performance than the PageRank based method and machine learning method.
Qindong Sun, Hanqin Wang, Liansheng Sui
LCN4