VLDB 2026 Research / reviewers in the wild / expert
Yulong Wan
dblp:148/9752
· DBLP profile ↗
11ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language ModelsabstractEmotional support is a core capability in human-AI interaction, with applications including psychological counseling, role play, and companionship. However, existing evaluations of large language models (LLMs) often rely on short, static dialogues and fail to capture the dynamic and long-term nature of emotional support. To overcome this limitation, we shift from snapshot-based evaluation to trajectory-based assessment, adopting a user-centered perspective that evaluates models based on their ability to improve and stabilize user emotional states over time. Our framework constructs a large-scale benchmark consisting of 328 emotional contexts and 1,152 disturbance events, simulating realistic emotional shifts under evolving dialogue scenarios. To encourage psychologically grounded responses, we constrain model outputs using validated emotion regulation strategies such as situation selection and cognitive reappraisal. User emotional trajectories are modeled as a first-order Markov process, and we apply causally-adjusted emotion estimation to obtain unbiased emotional state tracking. Based on this framework, we introduce three trajectory-level metrics: Baseline Emotional Level (BEL), Emotional Trajectory Volatility (ETV), and Emotional Centroid Position (ECP). These metrics collectively capture user emotional dynamics over time and support comprehensive evaluation of long-term emotional support performance of LLMs. Extensive evaluations across a diverse set of LLMs reveal significant disparities in emotional support capabilities and provide actionable insights for model development. Zhouxing Tan, Ruochong Xiong, Yulong Wan, Jinlong Ma, Hanlin Xue, Qichun Deng, Haifeng Jing, Zhengtong Zhang, Depei Liu, Shiyuan Luo |
AAAI | 3 |
| 2026 | Efficient Transcoder Adaptation for Fine-Tuned Models: Revealing Medical Reasoning Mechanisms in Large Language ModelsabstractLarge language models (LLMs) suffer from a lack of decision-making transparency, limiting their deployment in high-stakes domains such as healthcare. We propose a mechanistic interpretability framework that introduces two novel paradigms: Medical Fine-Tuning with Frozen Attention Layers (FTFA) and Posterior Adaptation Transcoders (PAT). FTFA freezes attention layers while fine-tuning only feed-forward network (FFN) parameters, enabling PAT to efficiently adapt pre-trained transcoders on the same data. This approach achieves over 1000× efficiency improvement compared to training transcoders from scratch. We theoretically justify this methodology and demonstrate its cost-effectiveness for cross-domain transfer. Transcoders are sparse autoencoders that replace MLP layers to provide interpretable feature representations. By substituting MLP layers of both base Gemma2-2b and its medical fine-tuned variant with per-layer transcoders, we enable feature-level attribution analysis. Through systematic pruning and node merging of resulting attribution graphs, we construct human-interpretable decision pathways. Our analysis reveals that LLMs employ two parallel mechanisms for medical diagnosis: pattern matching and multi-hop reasoning, with fine-tuned models demonstrating enhanced correct reasoning patterns. This work provides a practical framework for training transcoders on fine-tuned models at minimal cost, enabling broader application of mechanistic interpretability across domains and potentially guiding model training through transcoder-based analysis. Zhouxing Tan, Hanlin Xue, Yulong Wan, Ruochong Xiong |
AAAI | 3 |
| 2026 | Verifiable LLM-Generated Text Detection via Projected Semantic-Structural DistributionsabstractRuochong Xiong, Qien Li, Wangwang Lian, Yulong Wan, Hanlin Xue, Zhouxing Tan, Han Yang, Fengyu Lu, Junfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruochong Xiong, Qien Li, Wangwang Lian, Yulong Wan, Hanlin Xue, Zhouxing Tan, Fengyu Lu |
ACL (1) | 4 |
| 2025 | EGLC: Enhancing Global Localization Capability for medical image segmentation
Yulong Wan, Dongming Zhou 0001 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Multi-channel multi-speaker transformer for speech recognitionabstractWith the development of teleconferencing and in-vehicle voice assistants, far-field multi-speaker speech recognition has become a hot research topic. Recently, a multi-channel transformer (MCT) has been proposed, which demonstrates the ability of the transformer to model far-field acoustic environments. However, MCT cannot encode high-dimensional acoustic features for each speaker from mixed input audio because of the interference between speakers. Based on these, we propose the multi-channel multi-speaker transformer (M2Former) for far-field multi-speaker ASR in this paper. Experiments on the SMS-WSJ benchmark show that the M2Former outperforms the neural beamformer, MCT, dual-path RNN with transform-average-concatenate and multi-channel deep clustering based end-to-end systems by 9.2%, 14.3%, 24.9%, and 52.2% respectively, in terms of relative word error rate reduction. Hongbin Suo, Yulong Wan |
INTERSPEECH | 4 |
| 2023 | Task-Agnostic Structured Pruning of Speech Representation Models
Haoyu Wang 0014, Siyuan Wang 0002, Weiqiang Zhang 0001, Hongbin Suo, Yulong Wan |
INTERSPEECH | 5 |
| 2023 | Robust Audio Anti-spoofing Countermeasure with Joint Training of Front-end and Back-end Models
Xingming Wang, Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 4 |
| 2023 | SEF-Net: Speaker Embedding Free Target Speaker Extraction Network
Bang Zeng, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 3 |
| 2023 | Outlier-aware Inlier Modeling and Multi-scale Scoring for Anomalous Sound Detection via Multitask LearningabstractThis paper proposes an approach for anomalous sound detection that incorporates outlier exposure and inlier modeling within a unified framework by multitask learning. While outlier exposure-based methods can extract features efficiently, it is not robust. Inlier modeling is good at generating robust features, but the features are not very effective. Recently, serial approaches are proposed to combine these two methods, but it still requires a separate training step for normal data modeling. To overcome these limitations, we use multitask learning to train a conformer-based encoder for outlier-aware inlier modeling. Moreover, our approach provides multi-scale scores for detecting anomalies. Experimental results on the MIMII and DCASE 2020 task 2 datasets show that our approach outperforms state-of-the-art single-model systems and achieves comparable results with top-ranked multi-system ensembles. Yucong Zhang, Hongbin Suo, Yulong Wan, Ming Li 0026 |
INTERSPEECH | 3 |
| 2015 | Phonotactic language recognition using dynamic pronunciation and language branch discriminative information
Xianliang Wang, Yulong Wan, Lin Yang 0016, Ruohua Zhou, Yonghong Yan 0002 |
Speech Commun. | 2 |
| 2014 | Language recognition system using language branch discriminative informationabstractThis paper presents our study of using language branch discriminative information effectively for language recognition. Language branch variability (LBV) method based on factor analysis techniques is proposed. In LBV method, language branch variability factor is obtained by concatenating low-dimensional factors in the language branch variability spaces. Language models are trained within language branches and between languages. Experiments on NIST 2011 Language Recognition Evaluation (LRE) 30s, 10s and 03s tasks show the proposed LBV method provides stable improvement compared to the state-of-art total variability (TV) approach. In 30-second task, it gains relative improvement by 14.6% in equal error rate (EER) and 12.9% in minimum decision cost value (minDCF), and in new metrics of NIST 2011 LRE, it leads to relative improvement of 7.2%-17.7%. Xianliang Wang, Yulong Wan, Lin Yang 0016, Ruohua Zhou, Yonghong Yan 0002 |
ICASSP | 2 |