Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongkang Yin

dblp:364/0146 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
1.012026
WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability · AAAI 2026
Natural language and speech › Speech recognition and synthesis
speaker diarization
1.012026
WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability · AAAI 2026

Methods — techniques the papers use, named apart from their topics

token-level diarization · 1.0speaker similarity matrix supervision · 1.0speaker adapters · 1.0
YearPublicationVenuePosition
2026 WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability
abstract
Speaker diarization is a fundamental task in speech processing aims to determine 'who speaks when'. When combined with ASR, it enables speaker-labeled transcription with broad practical value. Most existing methods rely on frame-level classification, but the high cost of annotating mixed-speaker audio limits the availability of large-scale, accurately labeled datasets. As a result, even state-of-the-art models struggle with imprecise speaker boundary detection and semantic segmentation errors, which degrade timestamp accuracy and downstream ASR performance. To address these challenges, we propose WhisperDiari, a unified framework for speaker diarization and ASR. We first construct LibriDiari, a dataset derived from LibriSpeech, containing 2–4 speaker mixed audio annotated with transcripts and speaker labels. WhisperDiari builds on the Whisper model, incorporating speaker adapters and Speaker Similarity Matrix Supervision to enhance speaker representation. In addition, a dedicated speaker decoder fuses speaker embeddings with contextual semantics from Whisper's decoder, enabling token-level diarization. This design effectively resolves segmentation ambiguity, aligns diarization with semantic units, and jointly models 'who speaks what and when', producing accurate, timestamped transcripts. We train the model on LibriDiari and evaluate it on both LibriDiari and the real-world AMI corpus. Experimental results demonstrate that WhisperDiari consistently outperforms state-of-the-art open-source baselines.
Yongkang Yin, Yuexian Zou
AAAI1
2025 Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection
abstract
Audio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the acoustic features fused with the facial features from the same frame in the manner of vector concatenating or simple projecting. Such solution faces the challenges when there are more than one faces in the frame or overlapping speeches occur since there are lack of information alignment between active speech and the face of taking person. Based on this observation, in this study, we adopt a new solution to establish the relationships between the audio and face information using a heterogeneous graph explicitly. Specifically, we propose AFs-Net, which is able to capture both the relationships between the audio and each candidate’s face, and also the interactions between the faces of the candidates themselves within the same intra-frame. As the result, the graph with attention is trained to learn the importance (attention coefficient) between adjacent nodes. Additionally, we impose consistency constraints that bring speech features closer to speaker characteristics, while aligning non-speech features with non-speaker characteristics, further enhancing the audio-faces alignment. Our frame-level modeling approach supports both streaming applications and real-time operation. Experiments show that our method achieves state-of-the-art (SOTA) performance across multiple datasets.
Yongkang Yin, Xusheng Yang, Yuexian Zou
ICASSP1
2025 FoleyMaster: High-Quality Video-to-Audio Synthesis via MLLM-Augmented Prompt Tuning and Joint Semantic-Temporal Adaptation
Yuehan Jin, Xianwei Zhuang, Yongkang Yin, Yuexian Zou
INTERSPEECH6
2024 AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li 0015, Ying Shan, Yuexian Zou
INTERSPEECH1