Yinfeng Yu

dblp:237/3612 · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0003-3089-4140ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
Kexue Wang, Yinfeng Yu
ICIC (13)2
2026 EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence
abstract
Emotional talking head video generation aims to synthesize expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic information. Although high-level semantics enhance emotional expressiveness, using semantic guidance alone may compromise lip-sync precision. Furthermore, mainstream generation methods struggle to capture temporal dependencies across frames, resulting in degraded temporal coherence. Therefore, we propose an Emotion-Aware Diffusion model-based Network, called EAD-Net. We introduce SyncNet supervision and Temporal REPresentation Alignment (TREPA) to strengthen audio-visual alignment under multi-modal guidance. We further introduce a Spatio-Temporal Directional Attention (STDA) module to enhance fine-grained feature interactions via directional context aggregation, thereby improving audio-visual alignment. A Temporal Frame graph Reasoning Module (TFRM) is designed to explicitly model inter-frame dependencies, ensuring temporally coherent motion transitions without abrupt artifacts. To enhance emotional expressiveness, a large language model is employed to extract textual descriptions from real videos, serving as high-level semantic guidance. Experiments on the HDTF and MEAD datasets demonstrate that our method outperforms existing methods in terms of lip-sync accuracy, temporal consistency and emotional accuracy.
Yinfeng Yu, Shengjie Shen
ICMR2
2026 Beyond textual knowledge: Leveraging multimodal knowledge bases for enhancing vision-and-language navigation
Yinfeng Yu
Inf. Process. Manag.2
2025 LGFormer: A Local-Global Dynamic Attention Window Transformer for Speech Emotion Recognition
abstract
Speech emotion recognition is important in intelligent human-computer interaction, but modeling to handle long-range dependencies and local emotional cues remains challenging. This paper proposes a Local-Global Dynamic Attention Window-based Transformer model (LGFormer). The Local module dynamically divides the window based on temporal significance, capturing fine-grained localized information and synthesizing features using the Global module. The model introduces a novel attention mechanism to optimize computational efficiency, making it particularly suitable for scenarios with limited computational resources. We evaluate the method on the IEMOCAP and MELD datasets, achieving 2.8% and 1.1% improvements in weighted accuracy and F1 score, respectively. Comparison experiments with multiple benchmark algorithms validate the effectiveness of the model.
Yinfeng Yu, Wendong Zheng
CSCWD2
2025 Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
abstract
Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effectively leveraging multimodal cues to guide navigation. While prior works have explored basic fusion of visual and audio data, they often overlook deeper perceptual context. To address this, we propose the Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation (DMTF-AVN). Our approach uses a multi-target architecture coupled with a refined Transformer mechanism to filter and selectively fuse cross-modal information. Extensive experiments on the Replica and Matterport3D datasets demonstrate that DMTF-AVN achieves state-of-the-art performance, outperforming existing methods in success rate (SR), path efficiency (SPL), and scene adaptation (SNA). Furthermore, the model exhibits strong scalability and generalizability, paving the way for advanced multimodal fusion strategies in robotic navigation. The code and videos are available at https://github.com/zzzmmm-svg/DMTF.
Yinfeng Yu, Meiling Zhu
ECAI1
2025 Audio-Driven Talking Head Generation with Emotion Based on FLAME Geometry Model
Yinfeng Yu, Shengjie Shen
ICANN (4)3
2025 Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) integrates diverse modalities—text, audio, and video—to comprehensively analyze and understand individuals’ emotional states. However, the real-world prevalence of incomplete data poses significant challenges to MSA, mainly due to the randomness of modality missing. Moreover, the heterogeneity issue in multimodal data has yet to be effectively addressed. To tackle these challenges, we introduce the Modality-Invariant Bidirectional Temporal Representation Distillation Network (MITR-DNet) for Missing Multimodal Sentiment Analysis. MITR-DNet employs a distillation approach, wherein a complete modality teacher model guides a missing modality student model, ensuring robustness in the presence of modality missing. Simultaneously, we developed the Modality-Invariant Bidirectional Temporal Representation Learning Module (MIB-TRL) to mitigate heterogeneity.
Yinfeng Yu, Xinxin Jiao
ICASSP3
2025 Landmark-Guided Knowledge for Vision-and-Language Navigation
Meiling Zhu, Yinfeng Yu
ICIC (6)3
2025 A Speech Enhancement Method Based on Training Lifetime Knowledge Distillation
Yinfeng Yu
ICIC (17)3
2025 Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
abstract
Speech self-supervised learning (SSL) has made great progress in various speech processing tasks, but there is still room for improvement in speech enhancement (SE). This paper presents BSP-MPNet, a dual-path framework that combines self-supervised features with magnitude-phase information for SE. The approach starts by applying the perceptual contrast stretching (PCS) algorithm to enhance the magnitude-phase spectrum. A magnitude-phase 2D coarse (MP-2DC) encoder then extracts coarse features from the enhanced spectrum. Next, a feature-separating self-supervised learning (FS-SSL) model generates self-supervised embeddings for the magnitude and phase components separately. These embeddings are fused to create cross-domain feature representations. Finally, two parallel RNN-enhanced multi-attention (REMA) mask decoders refine the features, apply them to the mask, and reconstruct the speech signal. We evaluate BSP-MPNet on the VoiceBank+DEMAND and WHAMR! datasets. Experimental results show that BSP-MPNet outperforms existing methods under various noise conditions, providing new directions for self-supervised speech enhancement research.
Alimjan Mattursun, Yinfeng Yu, Chunyang Ma
ICME3
2025 Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
ICONIP (5)2
2025 PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control
Tianheng Zhu, Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
ICONIP (4)2
2025 AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
abstract
This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the Fast-Speech 2 architecture while addressing the challenge of local context modeling, which is crucial for capturing intricate speech features such as pauses, stress, and intonation. By embedding a phrase structure parser into the model and introducing a local convolution module, AMNet enhances the model’s sensitivity to local information. Additionally, AMNet decouples tonal characteristics from phonemes, providing explicit guidance for tone modeling, which improves tone accuracy and pronunciation. Experimental results demonstrate that AMNet outperforms baseline models in subjective and objective evaluations. The proposed model achieves superior Mean Opinion Scores (MOS), lower Mel Cepstral Distortion (MCD), and improved fundamental frequency fitting F 0(R2), confirming its ability to generate high-quality, natural, and expressive Mandarin speech.
Yubing Cao, Yinfeng Yu
IJCNN2
2025 Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
abstract
Multimodal emotion recognition (MER) seeks to integrate various modalities to predict emotional states accurately. However, most current research focuses solely on the fusion of audio and text features, overlooking the valuable information in emotion labels. This oversight could potentially hinder the performance of existing methods, as emotion labels harbor rich, insightful information that could significantly aid MER. We introduce a novel model called Label Signal-Guided Multimodal Emotion Recognition (LSGMER) to overcome this limitation. This model aims to fully harness the power of emotion label information to boost the classification accuracy and stability of MER. Specifically, LSGMER employs a Label Signal Enhancement module that optimizes the representation of modality features by interacting with audio and text features through label embeddings, enabling it to capture the nuances of emotions precisely. Furthermore, we propose a Joint Objective Optimization(JOO) approach to enhance classification accuracy by introducing the Attribution-Prediction Consistency Constraint (APC), which strengthens the alignment between fused features and emotion categories. Extensive experiments conducted on the IEMOCAP and MELD datasets have demonstrated the effectiveness of our proposed LSGMER model.
Xuechun Shao, Yinfeng Yu
IJCNN2
2025 DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion
abstract
Current Audio-Visual Source Separation methods primarily adopt two design strategies. The first strategy involves fusing audio and visual features at the bottleneck layer of the encoder, followed by processing the fused features through the decoder. However, when there is a significant disparity between the two modalities, this approach may lead to the loss of critical information. The second strategy avoids direct fusion and instead relies on the decoder to handle the interaction between audio and visual features. Nonetheless, if the encoder fails to integrate information across modalities adequately, the decoder may be unable to effectively capture the complex relationships between them. To address these issues, this paper proposes a dynamic fusion method based on a gating mechanism that dynamically adjusts the modality fusion degree. This approach mitigates the limitations of solely relying on the decoder and facilitates efficient collaboration between audio and visual features. Additionally, an audio attention module is introduced to enhance the expressive capacity of audio features, thereby further improving model performance. Experimental results demonstrate that our method achieves significant performance improvements on two benchmark datasets, validating its effectiveness and advantages in Audio-Visual Source Separation tasks.
Yinfeng Yu
ICMR1
2025 DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a challenging task where an agent must understand language instructions and navigate unfamiliar environments using visual cues. The agent must accurately locate the target based on visual information from the environment and complete tasks through interaction with the surroundings. Despite significant advancements in this field, two major limitations persist: (1) Many existing methods input complete language instructions directly into multi-layer Transformer networks without fully exploiting the detailed information within the instructions, thereby limiting the agent's language understanding capabilities during task execution; (2) Current approaches often overlook the modeling of object relationships across different modalities, failing to effectively utilize latent clues between objects, which affects the accuracy and robustness of navigation decisions. We propose a Dual Object Perception-Enhancement Network (DOPE) to address these issues to improve navigation performance. First, we design a Text Semantic Extraction (TSE) to extract relatively essential phrases from the text and input them into the Text Object Perception-Augmentation (TOPA) to fully leverage details such as objects and actions within the instructions. Second, we introduce an Image Object Perception-Augmentation (IOPA), which performs additional modeling of object information across different modalities, enabling the model to more effectively utilize latent clues between objects in images and text, enhancing decision-making accuracy. Extensive experiments on the R2R and REVERIE datasets validate the efficacy of the proposed approach.
Yinfeng Yu
ICMR1
2025 Phoneme-Controlled LLM with Self-Supervised Speech Prompts for Mispronunciation Detection
abstract
Pronunciation Error Detection and Diagnosis (MDD) is a key technology in Computer-Assisted Pronunciation Training (CAPT) and Computer-Assisted Language Learning (CALL). Recently large language models (LLMs) have shown strong performance in multimodal tasks. This paper proposes a new MDD framework called S-TATLLM which combines the advantages of an incremental self-supervised model (based on a local-global feature extraction structure CGSL using multi-head self-attention and convolution) and large language models to build an end-to-end multimodal pronunciation error detection system. By introducing phoneme-level control information and a text-audio-text embedding approach the system guides the language model to focus on easily confused pronunciation errors thus improving detection performance. S-TATLLM achieves a recall of 99.44% an F1 score of 0.8281 and a diagnosis accuracy (DER) of 92.03% which are better than the wav2vec2-CTC method with 60.84% 0.6164 and 70.74% and the AEL method with 73.71% 0.7749 and 81.60%.
Zhengping Song, Zaokere Kadeer, Mulati Kahaer, Xudong Pang, Yinfeng Yu, Aishan Wumaier
MMAsia5
2025 ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
abstract
Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into consistency models, enabling efficient one-step generation. However, these approaches introduce additional training costs and rely heavily on the performance of pre-trained teacher models. In this paper, we propose ECTSpeech, a simple and effective one-step speech synthesis framework that, for the first time, incorporates the Easy Consistency Tuning (ECT) strategy into speech synthesis. By progressively tightening consistency constraints on a pre-trained diffusion model, ECTSpeech achieves high-quality one-step generation while significantly reducing training complexity. In addition, we design a multi-scale gate module (MSGate) to enhance the denoiser’s ability to fuse features at different scales. Experimental results on the LJSpeech dataset demonstrate that ECTSpeech achieves audio quality comparable to state-of-the-art methods under single-step sampling, while substantially reducing the model’s training cost and complexity.
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
MMAsia2
2025 DP-GaussTalk: Dual-Path Audio-Driven Feature Fusion for 3D Gaussian-Based Talking Head Synthesis
abstract
Audio-driven 3D talking head generation has emerged as a significant research area in artificial intelligence, with broad applications in virtual assistants, digital media production, and online education. However, existing approaches often suffer from limitations in audio-visual synchronization accuracy, fine-grained facial expression reconstruction, and real-time rendering efficiency. To address these challenges, we propose DP-GaussTalk, a novel framework for generating 3D Gaussian-based talking heads driven by speech input. The proposed method utilizes a WavLM-based audio feature extractor to obtain multi-scale acoustic representations. A Temporal Audio Compressor (TACo) is introduced to refine audio features, enhancing the control precision over 3D facial deformations. Furthermore, a dual-path cross-attention mechanism is designed to align audio and visual modalities effectively, thereby improving lip-sync accuracy and facial expression fidelity. A HyperRestore module is also integrated to enhance the visual quality of the synthesized outputs. Experimental evaluations on several benchmark datasets demonstrate that DP-GaussTalk outperforms state-of-the-art methods in terms of visual realism, synchronization accuracy, and inference speed, offering a practical and scalable solution for real-time 3D talking head generation.
Yinfeng Yu
SMC3
2025 FGHFN: High-Resolution Fusion Network with Frequency-Domain Guidance for Remote Sensing Semantic Segmentation
abstract
When performing semantic segmentation on high-resolution remote sensing images, existing methods face a trade-off between capturing spatial details and modeling global context efficiently. To address this, we propose FGHFN, a High-Resolution Fusion Network with Frequency-Domain Guidance. The encoder extracts multi-scale local features, while the HRFusion module dynamically merges adjacent-resolution features via channel gating, preserving critical edges and textures. At the decoder, the Frequency-domain Global Filter (FFGF) models global context with convolutional cost, and the Strip Fusion Block (SFB) aligns cross-layer receptive fields through strip convolution. Experiments on LoveDA, Vaihingen, and Potsdam show mIoU scores of 55.65%, 84.65%, and 87.61%, demonstrating FGHFN’s efficacy.
Yinfeng Yu
SMC2
2025 Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks
abstract
Audio-visual navigation represents a significant area of research in which intelligent agents utilize egocentric visual and auditory perceptions to identify audio targets. Conventional navigation methodologies typically adopt a staged modular design, which involves first executing feature fusion, then utilizing Gated Recurrent Unit (GRU) modules for sequence modeling, and finally making decisions through reinforcement learning. While this modular approach has demonstrated effectiveness, it may also lead to redundant information processing and inconsistencies in information transmission between the various modules during the feature fusion and GRU sequence modeling phases. This paper presents IRCAM-AVN (Iterative Residual Cross-Attention Mechanism for Audiovisual Navigation), an end-to-end framework that integrates multimodal information fusion and sequence modeling within a unified IRCAM module, thereby replacing the traditional separate components for fusion and GRU. This innovative mechanism employs a multi-level residual design that concatenates initial multimodal sequences with processed information sequences. This methodological shift progressively optimizes the feature extraction process while reducing model bias and enhancing the model’s stability and generalization capabilities. Empirical results indicate that intelligent agents employing the iterative residual cross-attention mechanism exhibit superior navigation performance.
Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
SMC2
2025 EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
abstract
This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3–5 minutes of training video to synthesize high-quality facial animations. The framework comprises two key stages: static Gaussian initialization and audio-driven deformation. In the first stage, a multi-resolution hash triplane and a Kolmogorov-Arnold Network (KAN) are used to extract spatial features and construct a compact 3D Gaussian representation. In the second stage, we propose an Efficient Spatial-Audio Attention (ESAA) module to fuse audio and spatial cues, while KAN predicts the corresponding Gaussian deformations. Extensive experiments demonstrate that EGSTalker achieves rendering quality and lip-sync accuracy comparable to state-of-the-art methods, while significantly outperforming them in inference speed. These results highlight EGSTalker’s potential for real-time multimedia applications.
Tianheng Zhu, Yinfeng Yu, Fuchun Sun 0001, Wendong Zheng
SMC2
2024 PCQ: Emotion Recognition in Speech via Progressive Channel Querying
Yinfeng Yu, Xinxin Jiao
ICIC (3)3
2024 MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
abstract
Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using Multi-Spatial Fusion and Hierarchical Cooperative Attention on spectrograms and raw audio. We employ the Multi-Spatial Fusion module (MF) to efficiently identify emotion-related spectrogram regions and integrate Hubert features for higher-level acoustic information. Our approach also includes a Hierarchical Cooperative Attention module (HCA) to merge features from various auditory levels. We evaluate our method on the IEMOCAP dataset and achieve 2.6% and 1.87% improvements on the weighted accuracy and unweighted accuracy, respectively. Extensive experiments demonstrate the effectiveness of the proposed method.
Xinxin Jiao, Yinfeng Yu
ICME3
2024 ECMISM: Speech Recognition via Enhancing Conformer Models with Innovative Scoring Matrices
Yinfeng Yu, Miaomiao Xu
ICPR (28)3
2024 Collaborative Transformer Decoder Method for Uyghur Speech Recognition in-Vehicle Environment
Yinfeng Yu, Miaomiao Xu, Alimjan Mattursun
ICPR (33)3
2024 VNet: A GAN-Based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
abstract
Since the introduction of Generative Adversarial Networks (GANs) in speech synthesis, remarkable achievements have been attained. In a thorough exploration of vocoders, it has been discovered that audio waveforms can be generated at speeds exceeding real-time while maintaining high fidelity, achieved through the utilization of GAN-based models. Typically, the inputs to the vocoder consist of band-limited spectral information, which inevitably sacrifices high-frequency details. To address this, we adopt the full-band Mel spectrogram information as input, aiming to provide the vocoder with the most comprehensive information possible. However, previous studies have revealed that the use of full-band spectral information as input can result in the issue of over-smoothing, compromising the naturalness of the synthesized speech. To tackle this challenge, we propose VNet, a GAN-based neural vocoder network that incorporates full-band spectral information and introduces a Multi-Tier Discriminator (MTD) comprising multiple sub-discriminators to generate high-resolution signals. Additionally, we introduce an asymptotically constrained method that modifies the adversarial loss of the generator and discriminator, enhancing the stability of the training process. Through rigorous experiments, we demonstrate that the VNet model is capable of generating high-fidelity speech and significantly improving the performance of the vocoder.
Yubing Cao, Yinfeng Yu
SMC4
2024 BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network Based on Self-Supervised Embedding
abstract
Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportunities for improvement. In this study, we introduce a novel cross-domain feature fusion and multi-attention speech enhancement network, termed BSS-CFFMA, which leverages self-supervised embeddings. BSS-CFFMA comprises a multi-scale cross-domain feature fusion (MSCFF) block and a residual hybrid multi-attention (RHMA) block. The MSCFF block effectively integrates cross-domain features, facilitating the extraction of rich acoustic information. The RHMA block, serving as the primary enhancement module, utilizes three distinct attention modules to capture diverse attention representations and estimate high-quality speech signals. We evaluate the performance of the BSS-CFFMA model through comparative and ablation studies on the VoiceBank-DEMAND dataset, achieving SOTA results. Furthermore, we select three types of data from the WHAMR! dataset, a collection specifically designed for speech enhancement tasks, to assess the capabilities of BSS-CFFMA in tasks such as denoising only, dereverberation only, and simultaneous denoising and dereverberation. This study marks the first attempt to explore the effectiveness of self-supervised embedding-based speech enhancement methods in complex tasks encompassing derever-beration and simultaneous denoising and dereverberation. The demo implementation of BSS-CFFMA is available online2,
Alimjan Mattursun, Yinfeng Yu
SMC3
2024 Heterogeneous Space Fusion and Dual-Dimension Attention: A New Paradigm for Speech Enhancement
abstract
Self-supervised learning has demonstrated impressive performance in speech tasks, yet there remains ample opportunity for advancement in the realm of speech enhancement research. In addressing speech tasks, confining the attention mechanism solely to the temporal dimension poses limitations in effectively focusing on critical speech features. Taking into account the aforementioned issues, our study introduces a novel speech enhancement framework, HFSDA, which skillfully integrates heterogeneous spatial features and incorporates a dual-dimension attention mechanism to significantly enhance speech clarity and quality in noisy environments. By leveraging self-supervised learning embeddings in tandem with Short-Time Fourier Transform (STFT) spectrogram features, our model excels at capturing both high-level semantic information and detailed spectral data, enabling a more thorough analysis and refinement of speech signals. Furthermore, we employ the innovative Omni-dimensional Dynamic Convolution (ODConv) technology within the spectrogram input branch, enabling enhanced extraction and integration of crucial information across multiple dimensions. Additionally, we refine the Conformer model by enhancing its feature extraction capabilities not only in the temporal dimension but also across the spectral domain. Extensive experiments on the VCTK-DEMAND dataset show that HFSDA is comparable to existing state-of-the-art models, confirming the validity of our approach.
Yinfeng Yu
SMC3
2023 SRTNET: Time Domain Speech Enhancement via Stochastic Refinement
abstract
Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose SRTNet, a novel method for speech enhancement via Stochastic Refinement in complete Time b domain. Specifically, we design a joint network consisting of a deterministic module and a stochastic module, which makes up the "enhance-and-refine" paradigm. We theoretically demonstrate the feasibility of our method and experimentally prove that our method achieves faster training, faster sampling and higher quality. Our code is available at https://github.com/zhibinQiu/SRTNet.git
Zhibin Qiu, Mengfan Fu, Yinfeng Yu, Fuchun Sun 0001, Hao Huang 0009
ICASSP3
2023 Measuring Acoustics with Collaborative Multiple Agents
abstract
As humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly used to characterize environment acoustics as a function of the scene geometry, materials, and source/receiver locations. Traditionally, RIRs are measured by setting up a loudspeaker and microphone in the environment for all source/receiver locations, which is time-consuming and inefficient. We propose to let two robots measure the environment's acoustics by actively moving and emitting/receiving sweep signals. We also devise a collaborative multi-agent policy where these two robots are trained to explore the environment's acoustics while being rewarded for wide exploration and accurate prediction. We show that the robots learn to collaborate and move to explore environment acoustics while minimizing the prediction error. To the best of our knowledge, we present the very first problem formulation and solution to the task of collaborative environment acoustics measurements with multiple agents.
Yinfeng Yu, Changan Chen, Le-le Cao, Fangkai Yang, Fuchun Sun 0001
IJCAI1
2023 Echo-Enhanced Embodied Visual Navigation
abstract
Visual navigation involves a movable robotic agent striving to reach a point goal (target location) using vision sensory input. While navigation with ideal visibility has seen plenty of success, it becomes challenging in suboptimal visual conditions like poor illumination, where traditional approaches suffer from severe performance degradation. We propose E3VN (echo-enhanced embodied visual navigation) to effectively perceive the surroundings even under poor visibility to mitigate this problem. This is made possible by adopting an echoer that actively perceives the environment via auditory signals. E3VN models the robot agent as playing a cooperative Markov game with that echoer. The action policies of robot and echoer are jointly optimized to maximize the reward in a two-stream actor-critic architecture. During optimization, the reward is also adaptively decomposed into the robot and echoer parts. Our experiments and ablation studies show that E3VN is consistently effective and robust in point goal navigation tasks, especially under nonideal visibility.
Yinfeng Yu, Le-le Cao, Fuchun Sun 0001, Chao Yang 0026, Huicheng Lai, Wenbing Huang 0001
Neural Comput.1
2022 Pay Self-Attention to Audio-Visual Navigation
Yinfeng Yu, Le-le Cao, Fuchun Sun 0001
BMVC1
2022 Sound Adversarial Audio-Visual Navigation
Yinfeng Yu, Wenbing Huang 0001, Fuchun Sun 0001, Changan Chen, Yikai Wang 0001
ICLR1