Hanting Wang

dblp:329/0753 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
abstract
In recent years, large language models have achieved significant success in generative tasks (e.g., speech cloning and audio generation) related to speech, audio, music, and other signal domains. A crucial element of these models is the discrete acoustic codecs, which serve as an intermediate representation replacing the mel-spectrogram. However, there exist several gaps between discrete codecs and downstream speech language models. Specifically, 1) Due to the reconstruction paradigm of the Codec model and the structure of residual vector quantization, the initial channel of the codebooks contains excessive information, making it challenging to directly generate acoustic tokens from weakly supervised signals such as text in downstream tasks. 2) Achieving good reconstruction performance requires the utilization of numerous codebooks, which increases the burden on downstream speech language models. Consequently, leveraging the characteristics of speech language models, we propose Language-Codec. In the Language-Codec, we introduce a Masked Channel Residual Vector Quantization (MCRVQ) mechanism along with improved fourier transform structures, refined discriminator design to address the aforementioned gaps. We compare our method with competing audio compression algorithms and observe significant outperformance across extensive evaluations. Furthermore, we also validate the efficiency of the Language-Codec on downstream speech language models. The source code and pretrained models will be open-sourced after the paper is accepted. Codes are available at https://github.com/jishengpeng/Languagecodec.
Shengpeng Ji, Minghui Fang 0002, Jialong Zuo, Ziyue Jiang 0004, Dingdong Wang, Hanting Wang, Hai Huang 0013, Zhou Zhao 0001
ACL (1)6
2025 Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
abstract
We introduce a novel, training-free approach for enhancing alignment in Transformer-based Text-Guided Diffusion Models (TGDMs). Existing TGDMs often struggle to generate semantically aligned images, particularly when dealing with complex text prompts or multi-concept attribute binding challenges. Previous U-Net-based methods primarily optimized the latent space, but their direct application to Transformer-based architectures has shown limited effectiveness. Our method addresses these challenges by directly optimizing cross-attention maps during the generation process. Specifically, we introduce Self-Coherence Guidance, a method that dynamically refines attention maps using masks derived from previous denoising steps, ensuring precise alignment without additional training. To validate our approach, we constructed more challenging benchmarks for evaluating coarse-grained attribute binding, fine-grained attribute binding, and style binding. Experimental results demonstrate the superior performance of our method, significantly surpassing other state-of-the-art methods across all evaluated tasks. Our code is available at https://scg-diffusion.github.io/scg-diffusion.
Shulei Wang, Hai Huang 0013, Hanting Wang, Sihang Cai, WenKang Han, Tao Jin 0004, Jingyuan Chen 0003, Jieming Zhu, Zhou Zhao 0001
CVPR4
2025 Open-Set Cross Modal Generalization via Multimodal Unified Representation
abstract
This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions, addressing the limitations of prior closed-set cross-modal evaluations. OSCMG requires not only cross-modal knowledge transfer but also robust generalization to unseen classes within new modalities, a scenario frequently encountered in real-world applications. Existing multimodal unified representation work lacks consideration for open-set environments. To tackle this, we propose MICU, comprising two key components: Fine-Coarse Masked multimodal InfoNCE (FCMI) and Cross modal Unified Jigsaw Puzzles (CUJP). FCMI enhances multimodal alignment by applying contrastive learning at both holistic semantic and temporal levels, incorporating masking to enhance generalization. CUJP enhances feature diversity and model uncertainty by integrating modality-agnostic feature selection with self-supervised learning, thereby strengthening the model's ability to handle unknown categories in open-set tasks. Extensive experiments on CMG and the newly proposed OSCMG validate the effectiveness of our approach. The code is available at https://github.com/haihuangcode/CMG.
Hai Huang 0013, Yan Xia 0006, Shulei Wang, Hanting Wang, Minghui Fang 0002, Shengpeng Ji, Sashuai Zhou, Tao Jin 0004, Zhou Zhao 0001
ICCV4
2025 Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
abstract
Domain Generalization (DG) aims to enhance model robustness in unseen or distributionally shifted target domains through training exclusively on source domains. Although existing DG techniques, such as data manipulation, learning strategies, and representation learning, have shown significant progress, they predominantly address single-modal data. With the emergence of numerous multi-modal datasets and increasing demand for multi-modal tasks, a key challenge in Multi-modal Domain Generalization (MMDG) has emerged: enabling models trained on multi-modal sources to generalize to unseen target distributions within the same modality set. Due to the inherent differences between modalities, directly transferring methods from single-modal DG to MMDG typically yields sub-optimal results. These methods often exhibit randomness during generalization due to the invisibility of target domains and fail to consider inter-modal consistency. Applying these methods independently to each modality in the MMDG setting before combining them can lead to divergent generalization directions across different modalities, resulting in degraded generalization capabilities. To address these challenges, we propose a novel approach that leverages Unified Representations to map different paired modalities together, effectively adapting DG methods to MMDG by enabling synchronized multi-modal improvements within the unified space. Additionally, we introduce a supervised disentanglement framework that separates modal-general and modal-specific information, further enhancing the alignment of unified representations. Extensive experiments on benchmark datasets, including EPIC-Kitchens and Human-Animal-Cartoon, demonstrate the effectiveness and superiority of our method in enhancing multi-modal domain generalization.
Hai Huang 0013, Yan Xia 0006, Sashuai Zhou, Hanting Wang, Shulei Wang, Zhou Zhao 0001
ICCV4
2025 IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
abstract
Bridge models in image restoration construct a diffusion process from degraded to clear images. However, existing methods typically require training a bridge model from scratch for each specific type of degradation, resulting in high computational costs and limited performance. This work aims to efficiently leverage pretrained generative priors within existing image restoration bridges to eliminate this requirement. The main challenge is that standard generative models are typically designed for a diffusion process that starts from pure noise, while restoration tasks begin with a low-quality image, resulting in a mismatch in the state distributions between the two processes. To address this challenge, we propose a transition equation that bridges two diffusion processes with the same endpoint distribution. Based on this, we introduce the IRBridge framework, which enables the direct utilization of generative models within image restoration bridges, offering a more flexible and adaptable approach to image restoration. Extensive experiments on six image restoration tasks demonstrate that IRBridge efficiently integrates generative priors, resulting in improved robustness and generalization performance. Code will be available at GitHub.
Hanting Wang, Tao Jin 0004, Shulei Wang, Hai Huang 0013, Shengpeng Ji, Zhou Zhao 0001
ICML1
2025 TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
abstract
Image restoration under adverse weather conditions has been extensively explored, leading to numerous high-performance methods. In particular, recent advances in All-in-One approaches have shown impressive results by training on multi-task image restoration datasets. However, most of these methods rely on dedicated network modules or parameters for each specific degradation type, resulting in a significant parameter overhead. Moreover, the relatedness across different restoration tasks is often overlooked. In light of these issues, we propose a parameter-efficient All-in-One image restoration framework that leverages task-aware enhanced prompts to tackle various adverse weather degradations. Specifically, we adopt a two-stage training paradigm consisting of a pretraining phase and a prompt-tuning phase to mitigate parameter conflicts across tasks. We first employ supervised learning to acquire general restoration knowledge, and then adapt the model to handle specific degradation via trainable soft prompts. Crucially, we enhance these task-specific prompts in a task-aware manner. We apply low-rank decomposition to these prompts to capture both task-general and task-specific characteristics, and impose contrastive constraints to better align them with the actual inter-task relatedness. These enhanced prompts not only improve the parameter efficiency of the restoration model but also enable more accurate task modeling, as evidenced by t-SNE analysis. Experimental results on different restoration tasks demonstrate that the proposed method achieves superior performance with only 2.75M parameters.
Hanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang 0013, Qifei Zhang 0001, Tao Jin 0004
ACM Multimedia1
2025 A QoS Prediction Framework via Utility Maximization and Region-Aware Matrix Factorization
abstract
With the surge of Web services, users are more concerned about Quality-of-Service (QoS) information when choosing Web services with similar functionalities. Today, effectively and accurately predicting QoS values is a tough challenge. Typically, traditional methods only use the QoS values provided by users to predict the missing QoS values, ignoring the arbitrariness of some users in providing observed QoS values and failing to consider the existence of anomalous QoS values with contingencies caused by some unstable Web services. Taking into account the above, this article proposes HyLoReF-us, a new framework for QoS prediction. HyLoReF-us uses the user reputation to measure the trustworthiness of users and the service reputation to measure the stability of web services. First, considering the utility generated by the invocation between users and Web services, HyLoReF-us employs a Logit model to calculate the user reputation and service reputation. Second, after combining the location information of users and services, as well as their reputations, HyLoReF-us obtains QoS predictions through an improved Matrix Factorization (MF) model. Finally, a series of experiments were conducted on the standard WS-DREAM dataset. Experimental results show that HyLoReF-us outperforms current state-of-the-art or baseline methods at Matrix Densities (MD) from 5% to 30%.
Yugen Du, Guoxing Tang, Yingwei Luo, Hanting Wang
IEEE Trans. Serv. Comput.6
2024 MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
abstract
Zero-shot text-to-speech (TTS) has gained significant attention due to its powerful voice cloning capabilities, requiring only a few seconds of unseen speaker voice prompts.However, all previous work has been developed for cloud-based systems.Taking autoregressive models as an example, although these approaches achieve high-fidelity voice cloning, they fall short in terms of inference speed, model size, and robustness.Therefore, we propose MobileSpeech, which is a fast, lightweight, and robust zero-shot text-tospeech system based on mobile devices for the first time.Specifically: 1) leveraging discrete codec, we design a parallel speech mask decoder module called SMD, which incorporates hierarchical information from the speech codec and weight mechanisms across different codec layers during the generation process.Moreover, to bridge the gap between text and speech, we introduce a high-level probabilistic mask that simulates the progression of information flow from less to more during speech generation.2) For speaker prompts, we extract fine-grained prompt duration from the prompt speech and incorporate text, prompt speech by cross attention in SMD.We demonstrate the effectiveness of MobileSpeech on multilingual datasets at different levels, achieving state-ofthe-art results in terms of generating speed and speech quality.MobileSpeech achieves RTF of 0.09 on a single A100 GPU and we have successfully deployed MobileSpeech on mobile devices.Audio samples are available at https://mobilespeech.github.io/ .
Shengpeng Ji, Ziyue Jiang 0001, Hanting Wang, Jialong Zuo, Zhou Zhao 0001
ACL (1)3
2024 HyLoReF: A Reputation Based QoS Prediction Framework using Hybrid Location Information
abstract
With the proliferation of Web services, users pay more attention to Quality of Service (QoS) information when choosing Web services with similar functionalities. Predicting QoS values effectively and accurately is a difficult challenge. In the real world, some users are strongly subjective in submitting QoS observations and some Web services suffer from instability caused by bugs. Therefore, this paper uses user reputation to measure the reliability of users and service reputation to measure the stability of Web services. We propose HyLoReF, a QoS prediction framework based on reputation and hybrid location information. HyLoReF uses Logit model to compute user reputation and service reputation, and combines them into Matrix Factorization (MF) model to improve the QoS prediction accuracy. Experimental results show that HyLoReF outperforms baseline methods and state-of-the-art models on the Web services standard dataset WS-DREAM [1] when the Matrix Density (MD) is in the interval of 5% to 30%.
Yugen Du, Hanting Wang, Yingwei Luo, Benchi Ma, Guoxing Tang
ICWS4
2024 Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
abstract
In the burgeoning field of Audio-Visual Speech Recognition (AVSR), extant research has predominantly concentrated on the training paradigms tailored for high-quality resources. However, owing to the challenges inherent in real-world data collection, audio-visual data are frequently affected by modality-distortion, which encompasses audio-visual asynchrony, video noise and audio noise. The recognition accuracy of existing AVSR method is significantly compromised when multiple modality-distortion coexist in low-resource data. In light of the above challenges, we propose PCD: cluster-Prompt with Contrastive Decomposition, a robust framework for modality-distortion speech recognition, specifically devised to transpose the pre-trained knowledge from high-resource domain to the targeted domain by leveraging contrast-augmented prompts. In contrast to previous studies, we take into consideration the possibility of various types of distortion in both the audio and visual modalities. Concretely, we design bespoke prompts to delineate each modality-distortion, guiding the model to achieve speech recognition applicable to various distortion scenarios with quite few learnable parameters. To materialize the prompt mechanism, we employ multiple cluster-based strategies that better suits the pre-trained audio-visual model. Additionally, we design a contrastive decomposition mechanism to restrict the explicit relationships among various modality conditions, given their shared task knowledge and disparate modality priors. Extensive results on LRS2 dataset demonstrate that PCD achieves state-of-the-art performance for audio-visual speech recognition under the constraints of distorted resources. Code is available at https://github.com/ballooncatt/PCD.
Xize Cheng, Xiaoda Yang, Hanting Wang, Zhou Zhao 0001, Tao Jin 0004
ACM Multimedia4
2022 Collaborative Web Service Quality Prediction via Network Biased Matrix Factorization
abstract
Facing a large number of candidate Web service with the same function, user wishes to get the most appropriate one.Quality-of-Service (QoS) which represents non-functional attributes of Web services, has become a major concern for choosing service.But it is time-consuming and resource-consuming to assess all the QoS values by invoking candidate services one by one.Thus, QoS prediction is considered an effective method to obtain QoS information.Although most of QoS prediction methods claim be able to capture the interaction between users and services, few of them take account non-interaction factors, especially the factors arising from the network environment.In this paper, the non-interaction factors from the network environment are referred as network bias, and a network biased matrix factorization (NBMF) method is proposed for QoS prediction.The method packages network bias into a linear regression model and puts the user-service interaction into a matrix factorization model, which is more sophisticated in adapting diversified circumstance, particularly in complex network environment.In addition, extensive experiments are conduct on real-world QoS dataset, and the result prove that the NBMF method achieves better performance than other state-ofthe-art methods.
Wenhao Zhong, Yugen Du, Chuang Shan, Hanting Wang
SEKE4