EDBT 2026 Demo / reviewers in the wild / expert
Han Yin
dblp:147/6049
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsabstractThe Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers. Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li |
AAAI | 1 |
| 2026 | AdaSched: A Performance-Driven Cluster Scheduler for Deep Learning Workloads using Deep Reinforcement Learning
Han Yin, Jialun Li, Xuan Mo, Weigang Wu |
CCGrid | 1 |
| 2026 | IPSM-GCN: Intelligent Perception Structure-Motion Collaborative Graph Convolutional Networks for Skeleton-Based Action Recognition
Chengming Xie, Han Yin |
ICIC (18) | 3 |
| 2025 | MIMO Backscatter Communications Based on Cyclic Delay DiversityabstractBackscatter communication has emerged as a promising technology for future Internet of Things (IoT) due to its ultra-low power consumption. However, its performance is limited by the double fading effect in the backscatter channel. In this paper, we incorporate multiple-input multiple-output (MIMO) technique into backscatter communications, where the sinusoidal radio frequency (RF) source, backscatter device (BD), and receiver are all equipped with multiple antennas. Furthermore, cyclic delay diversity (CDD) is employed across antennas at the BD to achieve array gain as well as transmit diversity gain. To cope with the frequency-selectivity introduced by CDD, two frequency-domain signal detection schemes are utilized based on the maximum ratio combining (MRC) and minimum mean square error (MMSE) criteria, and the closed-form expressions for the bit error rates (BERs) of backscatter communications are derived. The transmit beamformer at the RF source is optimized to minimize the BER of the backscatter communication. With properly designed beamforming, the overall backscatter link can be effectively strengthened, thereby enabling more reliable backscatter communications. Simulation results show that deploying multiple antennas at both the RF source and the receiver significantly improves the BER performance, while the CDD-based transmission scheme can help to achieve additional array gain as well as transmit diversity gain for backscatter communications. Han Yin, Qianqian Zhang 0001, Hao Chen 0070, Ying-Chang Liang |
GLOBECOM | 1 |
| 2025 | Exploring Text-Queried Sound Event Detection with Audio Source SeparationabstractIn sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. Han Yin, Jisheng Bai, Yang Xiao 0019, Hui Wang 0030, Yafeng Chen, Rohan Kumar Das, Chong Deng |
ICASSP | 1 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 5 |
| 2025 | EnvSDD: Benchmarking Environmental Sound Deepfake Detection
Han Yin, Yang Xiao 0019, Rohan Kumar Das, Jisheng Bai, Haohe Liu, Wenwu Wang 0001, Mark D. Plumbley |
INTERSPEECH | 1 |
| 2025 | Diversified generation of commonsense reasoning questions
Jianxing Yu, Shiqi Wang 0016, Han Yin, Wei Liu 0061, Yanghui Rao, Qinliang Su |
Expert Syst. Appl. | 3 |
| 2025 | Disaggregated State Management in Apache Flink 2.0abstractWe present Apache Flink 2.0, an evolution of the popular stream processing system's architecture that decouples computation from state management. Flink 2.0 relies on a remote distributed file system (DFS) for primary state storage and uses local disks as a secondary cache, with state updates streamed continuously and directly to the DFS. To address the latency implications of remote storage, Flink 2.0 incorporates an asynchronous runtime execution model. Furthermore, Flink 2.0 introduces ForSt, a novel state store featuring a unified file system that enables faster and lightweight checkpointing, recovery, and reconfiguration with minimal intrusion to the existing Flink runtime architecture. Using a comprehensive set of Nexmark benchmarks and a large-scale stateful production workload, we evaluate Flink 2.0's large-state processing, checkpointing, and recovery mechanisms. Our results show significant performance improvements and reduced resource utilization compared to the baseline Flink 1.20 implementation. Specifically, we observe up to 94% reduction in checkpoint duration, up to 49× faster recovery after failures or a rescaling operation, and up to 50% cost savings. Zhaoqian Lan, Yanfei Lei, Han Yin, Kaitian Hu, Paris Carbone, Vasiliki Kalavri |
Proc. VLDB Endow. | 5 |
| 2025 | Multi-granularity acoustic information fusion for sound event detection
Han Yin, Jisheng Bai, Mou Wang, Susanto Rahardja, Dongyuan Shi, Woon-Seng Gan |
Signal Process. | 1 |
| 2025 | Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
Yuanjian Chen, Yang Xiao 0019, Han Yin, Yadong Guan, Xubo Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | DRLPG: Reinforced Opponent-Aware Order Pricing for Hub Mobility ServicesabstractA modern service model known as the “hub-oriented” model has emerged with the development of mobility services. This model allows users to request vehicles from multiple companies (agents) simultaneously through a unified entry (a ‘hub’). In contrast to conventional services, the “hub-oriented” model emphasizes pricing competition. To address this scenario, an agent should consider its competitors when developing its pricing strategy. In this paper, we introduce DRLPG, a mixed opponent-aware pricing method, which consists of two main components: the two-stage guarantor and the end-to-end deep reinforcement learning (DRL) module, as well as interaction mechanisms. In the guarantor, we design a prediction-decision framework. Specifically, we propose a new objective function for the spatiotemporal neural network in the prediction stage and utilize a traditional reinforcement learning method in the decision stage, respectively. In the end-to-end DRL framework, we explore the adoption of conventional DRL in the “hub-oriented” scenario. Finally, a meta-decider and an experience-sharing mechanism are proposed to combine both methods and leverage their advantages. We conduct extensive experiments on real data, and DRLPG achieves an average improvement of 99.9% and 61.1% in the peak and low peak periods, respectively. Our results demonstrate the effectiveness of our approach compared to the baseline. Zuohan Wu, Chen Zhang 0013, Han Yin, Libin Zheng 0001, Huaijie Zhu, Wei Liu 0061 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation InferenceabstractThis paper focuses on detecting clickbait posts on the Web.These posts often use eye-catching disinformation in mixed modalities to mislead users to click for profit.That affects the user experience and thus would be blocked by content provider.To escape detection, malicious creators use tricks to add some irrelevant nonbait content into bait posts, dressing them up as legal to fool the detector.This content often has biased relations with non-bait labels, yet traditional detectors tend to make predictions based on simple co-occurrence rather than grasping inherent factors that lead to malicious behavior.This spurious bias would easily cause misjudgments.To address this problem, we propose a new debiased method based on causal inference.We first employ a set of features in multiple modalities to characterize the posts.Considering these features are often mixed up with unknown biases, we then disentangle three kinds of latent factors from them, including the invariant factor that indicates intrinsic bait intention; the causal factor which reflects deceptive patterns in a certain scenario, and non-causal noise.By eliminating the noise that causes bias, we can use invariant and causal factors to build a robust model with good generalization ability.Experiments on three popular datasets show the effectiveness of our approach. Jianxing Yu, Shiqi Wang 0016, Han Yin, Zhenlong Sun, Ruobing Xie, Bo Zhang 0056, Yanghui Rao |
EMNLP | 3 |
| 2024 | Audiolog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive LearningabstractPrevious studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a large language models (LLMs)-powered audio logging system with hybrid token-semantic contrastive learning. Specifically, we propose to fine-tune the pre-trained hierarchical token-semantic audio Transformer by incorporating contrastive learning between hybrid acoustic representations. We then leverage LLMs to generate audio logs that summarize textual descriptions of the acoustic environment. Finally, we evaluate the AudioLog system on two datasets with both scene and event annotations. Experiments show that the proposed system achieves exceptional performance in acoustic scene classification and sound event detection, surpassing existing methods in the field. Further analysis of the prompts to LLMs demonstrates that AudioLog can effectively summarize long audio sequences1. To the best of our knowledge, this approach is the first attempt to leverage LLMs for summarizing long audio sequences. Jisheng Bai, Han Yin, Mou Wang, Dongyuan Shi, Woon-Seng Gan, Susanto Rahardja |
ICME | 2 |
| 2023 | 3D Audio Signal Processing Systems for Speech Enhancement and Sound Localization and DetectionabstractThe L3DAS23 of ICASSP Signal Processing Grand Challenge encourages research on 3D audio signal processing, such as 3D speech enhancement (SE) and 3D sound localization and detection (SELD). In this paper, we propose a two-stage system based on DPRNN and UNet for the SE task and a Conformer-based system for the SELD task. The proposed SE and SELD systems are evaluated on the L3DAS23 blind test sets. Results show that the proposed methods achieve state-of-the-art performance for 3D SE and SELD. Jisheng Bai, Siwei Huang, Han Yin, Yafei Jia, Mou Wang |
ICASSP | 3 |
| 2022 | Meces: Latency-efficient Rescaling via Prioritized State Migration for Stateful Distributed Stream Processing Systems
Rong Gu 0001, Han Yin, Weichang Zhong, Chunfeng Yuan, Yihua Huang 0001 |
USENIX ATC | 2 |
| 2021 | Towards Efficient Large-Scale Interprocedural Program Static Analysis on Distributed Data-Parallel ComputationabstractStatic program analysis has been widely applied along the whole process of the program development for bug detection, code optimization, testing, etc. Although researchers have made significant work in static program analysis, it is still challenging to perform sophisticated interprocedural analysis on large-scale modern software. The underlying reason is that interprocedural analysis for large-scale modern software is highly computation- and memory-intensive, leading to poor efficiency and scalability. In this article, we introduce an efficient distributed and scalable solution for sophisticated static analysis. Specifically, we propose a data-parallel algorithm and a join-process-filter computation model for the CFL-reachability-based interprocedural analysis. Based on that, an efficient distributed static analysis engine called BigSpa is developed, which is composed of an offline batch static program analysis system and an online incremental static program analysis system. The BigSpa system has high generality and can support all kinds of static analysis tasks that can be expressed as CFL reachability problems. The performance of BigSpa is evaluated on real-world large-scale software datasets. Our experiments show that the offline batch system can exceed an order of magnitude compared with the most advanced analysis tools available on performance, and for incremental analysis with small batch updates on the same data sets, the online analysis system can achieve near real-time response, which is very fast and flexible. Rong Gu 0001, Zhiqiang Zuo 0002, Han Yin, Zhaokang Wang, Linzhang Wang, Xuandong Li, Yihua Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |