EDBT 2026 Demo / reviewers in the wild / expert
Yongkang Yin
dblp:364/0146
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
1.0 | 1 | 2026 | WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability · AAAI 2026 |
Natural language and speech › Speech recognition and synthesis
speaker diarization |
1.0 | 1 | 2026 | WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text Adaptability · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
token-level diarization · 1.0speaker similarity matrix supervision · 1.0speaker adapters · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityabstractSpeaker diarization is a fundamental task in speech processing aims to determine 'who speaks when'. When combined with ASR, it enables speaker-labeled transcription with broad practical value. Most existing methods rely on frame-level classification, but the high cost of annotating mixed-speaker audio limits the availability of large-scale, accurately labeled datasets. As a result, even state-of-the-art models struggle with imprecise speaker boundary detection and semantic segmentation errors, which degrade timestamp accuracy and downstream ASR performance. To address these challenges, we propose WhisperDiari, a unified framework for speaker diarization and ASR. We first construct LibriDiari, a dataset derived from LibriSpeech, containing 2–4 speaker mixed audio annotated with transcripts and speaker labels. WhisperDiari builds on the Whisper model, incorporating speaker adapters and Speaker Similarity Matrix Supervision to enhance speaker representation. In addition, a dedicated speaker decoder fuses speaker embeddings with contextual semantics from Whisper's decoder, enabling token-level diarization. This design effectively resolves segmentation ambiguity, aligns diarization with semantic units, and jointly models 'who speaks what and when', producing accurate, timestamped transcripts. We train the model on LibriDiari and evaluate it on both LibriDiari and the real-world AMI corpus. Experimental results demonstrate that WhisperDiari consistently outperforms state-of-the-art open-source baselines. Yongkang Yin, Yuexian Zou |
AAAI | 1 |
| 2025 | Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker DetectionabstractAudio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the acoustic features fused with the facial features from the same frame in the manner of vector concatenating or simple projecting. Such solution faces the challenges when there are more than one faces in the frame or overlapping speeches occur since there are lack of information alignment between active speech and the face of taking person. Based on this observation, in this study, we adopt a new solution to establish the relationships between the audio and face information using a heterogeneous graph explicitly. Specifically, we propose AFs-Net, which is able to capture both the relationships between the audio and each candidate’s face, and also the interactions between the faces of the candidates themselves within the same intra-frame. As the result, the graph with attention is trained to learn the importance (attention coefficient) between adjacent nodes. Additionally, we impose consistency constraints that bring speech features closer to speaker characteristics, while aligning non-speech features with non-speaker characteristics, further enhancing the audio-faces alignment. Our frame-level modeling approach supports both streaming applications and real-time operation. Experiments show that our method achieves state-of-the-art (SOTA) performance across multiple datasets. Yongkang Yin, Xusheng Yang, Yuexian Zou |
ICASSP | 1 |
| 2025 | FoleyMaster: High-Quality Video-to-Audio Synthesis via MLLM-Augmented Prompt Tuning and Joint Semantic-Temporal Adaptation
Yuehan Jin, Xianwei Zhuang, Yongkang Yin, Yuexian Zou |
INTERSPEECH | 6 |
| 2024 | AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li 0015, Ying Shan, Yuexian Zou |
INTERSPEECH | 1 |