Haonan Cheng

dblp:198/3969 · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
abstract
The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458× fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets.
Yuankun Xie, Ruibo Fu, Songjun Cao, Haonan Cheng, Long Ye
AAAI7
2026 Knowledge-enhanced Chinese multimodal hate speech detection
Qingbao Huang, Pijian Li, Xingmao Zhang, Shizhen Chen, Haonan Cheng, Zhiyue Liu
Expert Syst. Appl.6
2026 OpenST: Toward open-set source tracing for neural codec deepfake audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Songjun Cao, Chenxing Li, Haonan Cheng, Long Ye
Neurocomputing9
2026 Anchor-Based Multimodal Verification: A Dynamic Query Framework for Fake News Forensics in Short Videos
abstract
The proliferation of maliciously altered short videos on social media platforms poses a significant threat to information security ecosystems, eroding public trust in digital media. Despite recent advancements in detecting fake video news, significant challenges remain in the forensic analysis of short videos, leading to issues of bias. First, as technology rapidly advances, fake videos are becoming increasingly semantically convincing, undermining the effectiveness of current classification methods. Second, the heterogeneous nature of video modalities (visual, textual, audio) creates critical challenges for models to learn discriminative feature representations. To address these challenges, we propose a dynamic query framework for fake news forensics in short videos, termed the Semantic Guided Adaptive Network (SGAN). Our approach is motivated by the need to utilize superficial alignment to identify suspicious manipulations through anchor-based verification and to leverage the adaptive capability of learnable queries to learn the heterogeneous boundary in each modality. Specifically, SGAN comprises a verification module and a flexible query learning module. The verification module employs text as the anchor to verify detailed context, mining fine-grained information while emphasizing key features, thereby providing candidate manipulations for downstream modules. The query learning module leverages learnable queries to map heterogeneous forensic features and integrates them through multi-level fusion for decision-making. Extensive experiments conducted on two widely used datasets demonstrate the effectiveness and generalization of the proposed method.
Pijian Li, Qingbao Huang, Feng Shuang 0002, Yi Cai 0001, Haonan Cheng, Qing Li 0001
IEEE Trans. Inf. Forensics Secur.5
2026 RD-VTA: Rule-Data Guided Video-to-Audio Generation for Fine-Grained Footstep Sound
abstract
It is challenging to implement visually guided fine-grained footstep sounds based on a limited number of samples in complex scenes. This is due to the interference of redundant information in the complex background of the visual scene for audiovisual mapping. As well as the complex coupling of sound features makes audiovisual fine-grained linear mapping difficult. To address the mentioned problems, we propose an automated video-to-audio generation method (RD-VTA) for footstep sound that incorporates data-driven and rule-based modelling approaches. First, we design a data-driven masked footstep sound generation network (DM-AFSG) to acquire audiovisual temporal concordance. The network is capable of separating visual sound objects, reducing background redundant interference, and generating initial target sounds that capture temporal cues. Secondly, a rule-based fine-grained footstep sound adjustment method (RT-AFSG) is designed based on visual guides such as material, motion type and displacement distance. The proposed RT-AFSG effectively achieves diverse sounds with a limited number of sound samples through sound texture analysis and modification. Moreover, it constructs the mapping relationship between different visual cues and footstep sounds, and realizes the fine variation of footstep sounds. To adequately validate the effectiveness of the method in terms of audiovisual temporal consistency and content granularity, we perform objective synchronization metrics and subjective human evaluation on the footsteps audiovisual dataset VAFoot. The experimental results show that the method obtains an average of 5% improvement in sound synchronization performance and significantly outperforms several existing methods in terms of sound content granularity. We encourage readers to watch and listen to the footstep sound results on our demo website:https://quinntt.github.io/RD-VTA/.
Qiutang Qi, Haonan Cheng, Hengyan Huang, Long Ye, Shaobin Li
IEEE Trans. Multim.2
2026 VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) have achieved significant advancements in multimodal understanding, reasoning, and interaction. However, they still suffer from hallucination, where the generated text often deviates from the factual content of the input image. To mitigate this issue, prior studies have primarily employed direct preference optimization (DPO) for human preference alignment. However, these approaches treat all textual words equally, neglecting the varying significance of individual words in grounding text generation to image content. This limitation hinders fine-grained semantic alignment and consequently constrains their effectiveness in hallucination suppression. To address this limitation, we propose a vision-guided lexical DPO method, called VGL-DPO. Specifically, we quantify the significance of words in positive preference data based on their relevance to the visual input and dynamically assign different weights to different words during training. This facilitates more precise optimization by emphasizing critical words that contribute to factual grounding. Additionally, we leverage the importance differences between high-significance words in positive and negative preference data to adaptively adjust the weight of the negative preference loss. This dynamic reweighting mechanism further refines the model’s ability to suppress hallucinated content while reinforcing factual accuracy. Extensive experiments across various models demonstrate that our method outperforms existing state-of-the-art methods in reducing hallucination and enhancing factual accuracy.
Siyuan Li 0001, Feng Wang 0063, Simeng Qin, Ranjie Duan, Haonan Cheng, Long Ye
ACM Trans. Multim. Comput. Commun. Appl.5
2026 Implement Referring Expression Comprehension by Extending Auto-focus Lens to Locked Vision Model
abstract
Referring Expression Comprehension (REC) aims to achieve fine-grained cross-modal content alignment. The traditional two-stage approaches, by decomposing REC into localization (region proposal) and comprehension (expression-based ranking), lead to the isolation of continuous image information and heavily rely on the quality of the proposals. In this article, we propose a point-based two-stage framework for REC to quickly achieve localization by inserting a language-modulated auto-focus module into the locked vision model. Specifically, we redefine REC as two processes: point-based cross-modal comprehension and point-based instance localization. For the comprehension stage, we reconstruct the raw annotations into soft masks at the feature point level as a metric of cross-modal correlation. With this indirect metric, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions. Remarkably, soft masks are shape-independent, which means our method is extremely general. By switching different vision models, different types of predictions (e.g., localization and segmentation) can be obtained. Experiments on multiple benchmarks demonstrate the feasibility and potential of our point-based paradigm. Our code will be public at https://github.com/VILAN-Lab/PBREC-AF .
Shiyi Zheng, Peizhi Zhao, Qingbao Huang, Yi Cai 0001, Haonan Cheng, Qi Wu 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Look Around Before Locating: Considering Content and Structure Information for Visual Grounding
abstract
As a long-term challenge and fundamental requirement in vision and language tasks, visual grounding aims to localize a target referred by a natural language query. The regional annotations form a superficial correlation between the subject of expression and some common visual entities, which hinder models from comprehending the linguistic content and structure. However, current one-stage methods struggle to uniformly model the visual and linguistic structure due to the structural gap between continuous image patches and discrete text tokens. In this paper, we propose a semi-structured reasoning framework for visual grounding to gradually comprehend the linguistic content and structure. Specifically, we devise a cross-modal content alignment module to effectively align unlabeled contextual information into a stable semantic space corrected by token-level prior knowledge obtained with CLIP. A multi-branch modulated localization module is also established to obtain modulation grounding by linguistic structure. Through a soft split mechanism, our method can destructure the expression into a fixed semi-structure (i.e., subject and context) while ensuring the completeness of linguistic content. Our method is thus capable of building a semi-structured reasoning system to effectively comprehend the linguistic content and structure by content alignment and structure modulated grounding. Experimental results on five widely-used datasets validate the performance improvements of our proposed method.
Shiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He, Haonan Cheng, Yi Cai 0001, Qingbao Huang
AAAI5
2025 Pop-Diffuseq: Controllable Symbolic Music Multi-Instrument Infilling and Accompaniment Generation with Long-Axis Attention
abstract
Controllability is a major challenge in music infilling and accompaniment tasks. Solutions based on transformer decoders have been widely adopted, while data-driven approaches with full self-attention result in high costs and unsatisfied outcomes for fine-grained control. Existing diffusion methods rely on trained classifiers, unconditional frameworks, or solo track, etc. To address these issues, we explore novel methods to enhance the controllability and quality of music model while reducing computational complexity. Firstly, we improve the classifier-free diffusion for multi-instrumental pop music. Secondly, we design a long-axis attention algorithm that combines long with axial attention to acquire the feature correlations of multi-dimensional attributes. Additionally, we contribute a pop band dataset with melody, style and mood labels handcrafted by musicians. After experiments on the benchmark dataset, our method demonstrates high-quality controllable results and outperforms existing state-of-the-art models. The GPU memory of our model is 26.9% lower than Diffuseq under the same hyperparameters.
Haonan Cheng, Long Ye, Qin Zhang 0009
ICME2
2025 FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-Attributes
Haonan Cheng, Hengyan Huang, Long Ye
ACM Multimedia1
2025 Exploring news intent and its application: A theory-driven approach
Zhengjia Wang 0001, Danding Wang, Qiang Sheng 0001, Juan Cao 0001, Haonan Cheng
Inf. Process. Manag.6
2025 Generalization enhancement strategy based on ensemble learning for open domain image manipulation detection
Haonan Cheng, L. Niu, L. Ye
J. Vis. Commun. Image Represent.1
2025 Visual primitives as words: Alignment and interaction for compositional zero-shot learning
Feng Shuang 0002, Jiahuan Li, Qingbao Huang, Wenye Zhao, Dongsheng Xu 0001, Haonan Cheng
Pattern Recognit.7
2025 DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based image captioning
Dongsheng Xu 0001, Qingbao Huang, Xingmao Zhang, Haonan Cheng, Yi Cai 0001
Pattern Recognit.4
2025 Noise-Informed Diffusion-Generated Image Detection With Anomaly Attention
abstract
With the rapid development of image generation technologies, especially the advancement of Diffusion Models, the quality of synthesized images has significantly improved, raising concerns among researchers about information security. To mitigate the malicious abuse of diffusion models, diffusion-generated image detection has proven to be an effective countermeasure. However, a key challenge for forgery detection is generalising to diffusion models not seen during training. In this paper, we address this problem by focusing on image noise. We observe that images from different diffusion models share similar noise patterns, distinct from genuine images. Building upon this insight, we introduce a novel Noise-Aware Self-Attention (NASA) module that focuses on noise regions to capture anomalous patterns. To implement a SOTA detection model, we incorporate NASA into Swin Transformer, forming an novel detection architecture NASA-Swin. Additionally, we employ a cross-modality fusion embedding to combine RGB and noise images, along with a channel mask strategy to enhance feature learning from both modalities. Extensive experiments demonstrate the effectiveness of our approach in enhancing detection capabilities for diffusion-generated images. When encountering unseen generation methods, our approach achieves the state-of-the-art performance.
Weinan Guan, Wei Wang 0025, Bo Peng 0002, Ziwen He, Jing Dong 0003, Haonan Cheng
IEEE Trans. Inf. Forensics Secur.6
2024 DNIT: Enhancing Day-Night Image-to-Image Translation through Fine-Grained Feature Handling (Student Abstract)
abstract
Existing image-to-image translation methods perform less satisfactorily in the "day-night" domain due to insufficient scene feature study. To address this problem, we propose DNIT, which performs fine-grained handling of features by a nighttime image preprocessing (NIP) module and an edge fusion detection (EFD) module. The NIP module enhances brightness while minimizing noise, facilitating extracting content and style features. Meanwhile, the EFD module utilizes two types of edge images as additional constraints to optimize the generator. Experimental results show that we can generate more realistic and higher-quality images compared to other methods, proving the effectiveness of our DNIT.
Haonan Cheng, Long Ye
AAAI2
2024 Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio Generation
abstract
Cross-modal binaural audio generation is an important task and has broad applications such as game sound development and auditory assistance for the visually impaired. However, existing datasets lack binaural samples with abundant visual venues. As a consequence, state-of-the-art cross-modal binaural audio generation methods have weak generalization. To support research on building robust binaural audio generation, we construct BinauralMusic dataset consisting of 5,462 performance video clips with binaural audio from 9 musical instrument categories. The performance venues involve indoor closed places such as shopping mall, hotel, bedroom, as well as outdoor open areas such as field, garden and seashore. Experiments show that the performance of the cross-modal binaural audio generation model can be significantly improved by 10.62% by using the BinauralMusic dataset as training material. Moreover, different from previous datasets, the BinauralMusic dataset can also support other audio-visual cross-modal learning tasks, including visually guided sound source localization and separation.
Haonan Cheng, Long Ye
ICASSP3
2024 An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio Detection
abstract
Partially spoofed audio detection is a challenging task, lying in the need to accurately locate the authenticity of audio at the frame level. To address this issue, we propose a fine-grained partially spoofed audio detection method, namely Temporal Deepfake Location (TDL), which can effectively capture information of both features and locations. Specifically, our approach involves two novel parts: embedding similarity module and temporal convolution operation. To enhance the identification between the real and fake features, the embedding similarity module is designed to generate an embedding space that can separate the real frames from fake frames. To effectively concentrate on the position information, temporal convolution operation is proposed to calculate the frame-specific similarities among neighboring frames, and dynamically select informative neighbors to convolution. Extensive experiments show that our method outperform baseline models in ASVspoof2019 Partial Spoof dataset and demonstrate superior performance even in the cross-dataset scenario.
Yuankun Xie, Haonan Cheng, Long Ye
ICASSP2
2024 FSD: An Initial Chinese Dataset for Fake Song Detection
abstract
Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio DeepFake Detection (ADD), the field of song deepfake detection lacks specialized datasets or methods for song authenticity verification. In this paper, we initially construct a Chinese Fake Song Detection (FSD) dataset to investigate the field of song deepfake detection. The fake songs in the FSD dataset are generated by five state-of-the-art singing voice synthesis and singing voice conversion methods. Our initial experiments on FSD revealed the ineffectiveness of existing speech-trained ADD models for the task of song deepfake detection. Thus, we employ the FSD dataset for the training of ADD models. We subsequently evaluate these models under two scenarios: one with the original songs and another with separated vocal tracks. Experiment results show that song-trained ADD models exhibit a 38.58% reduction in average equal error rate compared to speech-trained ADD models on the FSD test set.
Yuankun Xie, Xiaolin Lu, Zhenghao Jiang, Haonan Cheng, Long Ye
ICASSP6
2024 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Haonan Cheng, Long Ye, Jianhua Tao 0001
INTERSPEECH6
2024 Artifact feature purification for cross-domain detection of AI-generated images
Zheling Meng, Bo Peng 0002, Jing Dong 0003, Tieniu Tan, Haonan Cheng
Comput. Vis. Image Underst.5
2024 DiffuseRoll: multi-track multi-attribute music generation based on diffusion model
Haonan Cheng, Long Ye
Multim. Syst.3
2024 MusicECAN: An Automatic Denoising Network for Music Recordings With Efficient Channel Attention
abstract
In this work, we address the long-standing problem of automatic recorded music denoising. In previous audio denoising research, the primary focus has been on speech, and music denoising works only considered noise types in indoor conversation scenarios or old gramophone recordings, neglecting the amateur music recording scenario. To this end, we first propose MusicECAN, an automatic music denoising method designed to filter out additional noise components in recorded music. The novel architecture comprises two key components, namely, a feature learning module and a noise filtering module, which can efficiently but effectively model, refine and denoise the noisy input. Specifically, in order to capture sufficient noisy music information, an ECA-U-SAM based feature learning module is designed by incorporating an efficient channel attention (ECA) mechanism into the traditional U-Net model with a supervised attention module (SAM). To train our MusicECAN, we collect M&N, a dataset containing various clean music and noise recordings. Through the combination of different clean-noise recording pairs, we can effectively simulate possible music performance environments with various background noise. Extensive quantitative and qualitative comparisons demonstrate that our MusicECAN outperforms the state-of-the-art audio denoising methods.
Haonan Cheng, Zhicheng Lian, Long Ye, Qin Zhang 0009
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Domain Generalization via Aggregation and Separation for Audio Deepfake Detection
abstract
In this paper, we propose an Aggregation and Separation Domain Generalization (ASDG) method for Audio DeepFake Detection (ADD). Fake speech generated from different methods exhibits varied amplitude and frequency distributions rather than genuine speech. In addition, the spoofing attacks in training sets may not keep pace with the evolving diversity of real-world deepfake distributions. In light of this, we attempt to learn an ideal feature space that can aggregate real speech and separate fake speech to achieve better generalizability in the detection of unseen target domains. Specifically, we first propose a feature generator based on Lightweight Convolutional Neural Networks (LCNN), which is employed for generating a feature space and categorizing the feature into real and fake. Meanwhile, single-side domain adversarial learning is leveraged to make only the real speech from different domains indistinguishable, which enables the distribution of real speech to be aggregated in the feature space. Furthermore, a triplet loss is adopted to separate the distribution of fake speech while aggregating the distribution of real speech. Finally, in order to test the generalizability of the model, we train it with three different English datasets and evaluate in harsh conditions: cross-language and noisy datasets. The extensive experiments show that ASDG outperforms the baseline models in cross-domain tasks and decreases Equal Error Rate (EER) by up to 39.24% when compared to that of RawNet2. It is proved that the proposed Aggregation and Separation Domain Generalization method can be an effective strategy to improve the model generalizability.
Yuankun Xie, Haonan Cheng, Long Ye
IEEE Trans. Inf. Forensics Secur.2
2023 Learning A Self-Supervised Domain-Invariant Feature Representation for Generalized Audio Deepfake Detection
Yuankun Xie, Haonan Cheng, Long Ye
INTERSPEECH2
2023 RD-FGFS: A Rule-Data Hybrid Framework for Fine-Grained Footstep Sound Synthesis from Visual Guidance
abstract
Existing methods are difficult to synthesize fine-grained footsteps based on video frames only. This is due to the complicated nonlinear mapping relationships between motion states, spatial locations and different footstep sounds. Aiming to address this issue, we propose a Rule-Data guided Fine-Grained Footstep Sound (RD-FGFS) synthesis method. To the best of our knowledge, our work takes the first step in integrating data-driven and rule modeling approaches for visually aligned footstep sound synthesis. Firstly, we design a learning-based footstep sound generation network (FSGN) architecture driven by pose and flow features. The FSGN is proposed for generating an initial target sound which captures timing cues. Secondly, a rule-based fine-grained footstep sound adjustment (FGFSA) method is designed based on the visual guidance, namely ground material, movement type, and displacement distance. The proposed FGFSA effectively constructs a mapping relationship between different visual cues and footstep sounds, enabling fine-grained variations of footstep sounds. Experimental results show that our method improves the visual and sound synchronization results of footsteps and achieves impressive performance in footstep sound fine-grained control.
Qiutang Qi, Haonan Cheng, Yang Wang 0191, Long Ye, Shaobin Li
ACM Multimedia2
2023 Lightweight Scene-aware Rain Sound Simulation for Interactive Virtual Environments
abstract
We present a lightweight and efficient rain sound synthesis method for interactive virtual environments. Existing rain sound simulation methods require massive superposition of scene-specific precomputed rain sounds, which is excessive memory consumption for virtual reality systems (e.g. video games) with limited audio memory budgets. Facing this issue, we reduce the audio memory budgets by introducing a lightweight rain sound synthesis method which is only based on eight physically-inspired basic rain sounds. First, in order to generate sufficiently various rain sounds with limited sound data, we propose an exponential moving average based frequency domain additive (FDA) synthesis method to extend and modify the pre-computed basic rain sounds. Each rain sound is generated in the frequency domain before conversion back to the time domain, allowing us to extend the rain sound which is free of temporal distortions and discontinuities. Next, we introduce an efficient binaural rendering method to simulate the 3D perception that coheres with the visual scene based on a set of Near-Field Transfer Functions (NFTF). Various results demonstrate that the proposed method drastically decreases the memory cost (77 times compressed) and overcomes the limitations of existing methods in terms of interaction.
Haonan Cheng, Shiguang Liu, Jiawan Zhang
VR1
2023 PQG-A2SA: Performance Quantification Guided Audio-to-Score Alignment for Orchestral Music
abstract
Audio-to-score alignment is a multi-modal task that aims at generating an accurate mapping between symbolic and signal-level representations of musical signals, which is important for music performance analysis and retrieval. Among numerous music genres, orchestral music is a category of music with complex performance characteristics such as multi-instrument, non-percussive instrument and music expressiveness. However, previous methods do not take sufficient account of the performance characteristics of orchestral music, leading to limitations in alignment accuracy on orchestral music of these methods. To solve this problem, we present a performance quantification guided audio-to-score alignment (PQG-A2SA) method with high alignment accuracy for orchestral music at note-level. Specially, the PQG-A2SA contains two parts, namely an Inter Onset Interval (IOI) guided conditionally-constrained Dynamic Time Wrapping (DTW) and an articulation guided onset and offset detection. Different from the previous work, the IOI-guided conditionally-constrained DTW is designed to achieve a preliminary mapping between symbolic and chord-level representations of musical signals. In the second module, the onset and offset detection model under different musical articulations are established, thus refining the alignment results. We provide extensive experimental validation and analysis of our method. Our PQG-A2SA method can improve 9.0% in onset align rate and 17.5% in offset align rate at most compared with the state-of-the-art methods.
Zhicheng Lian, Haonan Cheng, Jiawan Zhang
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Global-Local Similarity Function for Automatic Playlist Generation
abstract
This paper proposes the Global-Local Similarity Function (GLSF) to exploit the multi-scale cues in track sequences for automatic playlist generation (APG). Unlike previous neighborhood-based methods only looking on local similarities for a given playlist, GLTS is constructed by first modeling the fine-grained audio features of each track, then capturing the long-term relations among consecutive tracks. Specifically, the fine-grained audio features are captured beat-by-beat to represent the rhythmic variation of music. The long- term relations are modeled by a designed track distance constraint (TD-constraint) to alleviate the incoherences and un- smooth transition in track sequences. The fine-grained audio features and TD-constraint are aggregated as the final GLSF by a simple distance function. Objective and subjective evaluations show that GLSF-based APG achieves better smooth transition and ensure the long-term content consistency among the tracks. Furthermore, GLSF yields a better understanding of the sequential relationship between tracks and propose a promising way to improve APG algorithms.1
Haonan Cheng, Ruyu Zhang, Long Ye
ICME1
2022 Towards an End-to-End Visual-to-Raw-Audio Generation With GAN
abstract
Automatically synthesizing sounds for different visual contents poses a challenge and there is a strong need to facilitate the direct creation of realistic sounds. Different from previous works, in this paper, we propose a novel deep learning based approach, which formulates sound simulation as a regression problem. This allows us to circumvent the complexity of the acoustic theory by a novel, general-purpose neural sound synthesis (V2RA) network. Moreover, the end-to-end architecture of V2RA ensures full training without any extra inputs, which thereby greatly improves the scalability and reusability over previous works. In contrast to conventional visual-to-audio generation methods, the V2RA problem is established and solved by generative adversarial networks (GANs). Furthermore, our network architecture can directly predict synchronized raw audio signals (unlike most existing approaches that handle the audio through spectrograms) and generate sound in real time. To evaluate the performance of the neural network generator, we specifically introduce two quantitative scores. Various experiments demonstrate that our V2RA network can produce compelling sound results, which thus provides a viable solution for applications such as sound design and dubbing.
Shiguang Liu, Haonan Cheng
IEEE Trans. Circuits Syst. Video Technol.3
2019 Haptic Force Guided Sound Synthesis in Multisensory Virtual Reality (VR) Simulation for Rigid-Fluid Interaction
abstract
This paper tackles a challenging problem for interactive rigid-fluid interaction sound synthesis. One core issue of the rigid-fluid interaction in multisensory VR system is how to balance the algorithm efficiency, result authenticity and result synchronization. Since the sampling rate of audio is far greater than visual and haptic modalities, sound synthesis for a multisensory VR system is more difficult than visual simulation and haptic rendering, which still remains an open challenge until now. Therefore, this paper focuses on developing an efficient sound synthesis method tailored for a multisensory system. To improve the result authenticity while ensuring real time performance and result synchronization, we propose a novel haptic force guided granular sound synthesis method tailored for sounding in multisensory VR systems. To the best of our knowledge, this is the first step that exploits haptic force feedback from the tactile channel for guiding sound synthesis in a multisensory VR system. Specifically, we propose a modified spectral granular sound synthesis method, which can ensure real time simulation and improve the result authenticity as well. Then, to balance the algorithm efficiency and result synchronization, we design a multi-force (MF) granulation algorithm which avoids repeated analysis of fluid particle motion and thereby improves the synchronization performance. Various results show that the proposed sound synthesis method effectively overcomes the limitations of existing methods in terms of audio modality, which has great potential to provide powerful technological support for building a more immersive multisensory VR system.
Haonan Cheng, Shiguang Liu
VR1
2019 Liquid-solid interaction sound synthesis
Haonan Cheng, Shiguang Liu
Graph. Model.1
2019 Physically-based statistical simulation of rain sound
abstract
A typical rainfall scenario contains tens of thousands of dynamic sound sources. A characteristic of the large-scale scene is the strong randomness in raindrop distribution, which makes it notoriously expensive to synthesize such sounds with purely physical methods. Moreover, the raindrops hitting different surfaces (liquid or various solids) can emit distinct sounds, for which prior methods with unified impact sound models are ill-suited. In this paper, we present a physically-based statistical simulation method to synthesize realistic rain sound, which respects surface materials. We first model the raindrop sound with two mechanisms, namely the initial impact and the subsequent pulsation of entrained bubbles. Then we generate material sound textures (MSTs) based on a specially designed signal decomposition and reconstruction model. This allows us to distinguish liquid surface with bubble sound and different solid surfaces with MSTs. Furthermore, we build a basic rain sound (BR-sound) bank with the proposed raindrop sound clustering method based on a statistical model, and design a sound source activator for simulating spatial propagation in an efficient manner. This novel method drastically decreases the computational cost while producing convincing sound results. Various experiments demonstrate the effectiveness of our sound simulation model.
Shiguang Liu, Haonan Cheng, Yiying Tong
ACM Trans. Graph.2
2017 Efficient sound synthesis for natural scenes
abstract
This paper presents a novel framework to generate the sound of outdoor natural scenes, such as waterfall, ocean, etc. Our method firstly simulates liquid with a grid-based method. Then combined with the movement of liquid, we generate seed-particles which represent bubbles, foams or splashes. Next, we assign each seed-particles a radius with a new radius distribution model. By calculating the bubbles' pressure wave we generate the sound. Experiments demonstrated that our novel framework can efficiently synthesize the sounds for natural scenes.
Kai Wang 0015, Haonan Cheng, Shiguang Liu
VR2