VLDB 2026 Research / reviewers in the wild / expert
Yudong Yang
dblp:38/5425
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology optimization of piezoelectric structures based on multi-modal input neural network
Jianhua Xiang, Yudong Yang, Xiwen Cai, Yongfeng Zheng |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | CCSEMO: A Chinese Counseling Speech Emotion Dataset annotated via a unified standardized annotation framework
Zhixing Guo, Yudong Yang, Hangbin Yu, Rongfeng Su |
Pattern Recognit. | 3 |
| 2025 | Audio-centric Video Understanding Benchmark without Text ShortcutabstractYudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Yixuan Li, Peihan Li, Yifan Jiang, Wei Li, Zejun Ma, Chao Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Peihan Li, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
EMNLP | 1 |
| 2025 | Enabling Auditory Large Language Models for Automatic Speech Quality EvaluationabstractSpeech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN. Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang, Jimin Zhuang, Xianzhao Chen, Xiaohai Tian, Guangzhi Sun, Lu Lu 0015, Chao Zhang 0031 |
ICASSP | 3 |
| 2025 | Improving LLM Video Understanding with 16 Frames Per SecondabstractHuman vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual information loss. In this paper, we introduce F-16, the first multimodal LLM designed for high-frame-rate video understanding. By increasing the frame rate to 16 FPS and compressing visual tokens within each 1-second clip, F-16 efficiently captures dynamic visual features while preserving key semantic information.
Experimental results demonstrate that higher frame rates considerably enhance video understanding across multiple benchmarks, providing a new approach to improving video LLMs beyond scaling model size or training data. F-16 achieves state-of-the-art performance among 7-billion-parameter video LLMs on both general and fine-grained video understanding benchmarks, such as Video-MME and TemporalBench. Furthermore, F-16 excels in complex spatiotemporal tasks, including high-speed sports analysis (*e.g.*, basketball, football, gymnastics, and diving), outperforming SOTA proprietary visual models like GPT-4o and Gemini-1.5-pro.
Additionally, we introduce a novel decoding method for F-16 that enables highly efficient low-frame-rate inference without requiring model retraining. We will release the source code, model checkpoints, and data at [https://github.com/bytedance/F-16](https://github.com/bytedance/F-16). Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 4 |
| 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language ModelabstractWhile recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding. This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks. To enhance its reasoning abilities, we develop a reasoning-intensive dataset featuring challenging audio-visual questions with step-by-step solutions. We also propose process direct preference optimization (pDPO), which leverages contrastive step selection to achieve efficient step-level reward modelling tailored for multimodal inputs. Additionally, we introduce RivaBench, the first reasoning-intensive video understanding benchmark, featuring over 4,000 high-quality, expert-curated question-answer pairs across scenarios such as standup comedy, academic presentations, and synthetic video detection. video-SALMONN-o1 achieves 3-8% accuracy improvements over the LLaVA-OneVision baseline across different video reasoning benchmarks. Besides, pDPO achieves 6-8% improvements compared to the supervised fine-tuning model on RivaBench. Enhanced reasoning enables video-SALMONN-o1 zero-shot synthetic video detection capabilities. Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 2 |
| 2025 | Addressing Task Conflicts in Stuttering Detection via MMoE-Based Multi-Task Learning
Xingfeng Li 0009, Yudong Yang |
INTERSPEECH | 3 |
| 2025 | Emotion-Guided Graph Attention Networks for Speech-Based Depression Detection under Emotion-Inducting Tasks
Yuqiu Zhou, Yongjie Zhou, Yudong Yang, Shuzhi Zhao, Rongfeng Su |
INTERSPEECH | 3 |
| 2025 | Visual-tactile fusion learning for material recognition based on channel switching and dual cross-attention
Wei Sun 0028, Qiaokang Liang, Yudong Yang |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | KGMV-net: Knowledge-guided multi-view network for audio-visual dysarthria severity assessment
Yudong Yang, Guorong Xu, Xiaoxia Du, Rongfeng Su |
Knowl. Based Syst. | 2 |
| 2024 | An Audio-Textual Diffusion Model for Converting Speech Signals into Ultrasound Tongue Imaging DataabstractAcoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the general patterns of tongue motions, and thus the quality of generated UTI data is limited. To address this issue, this paper proposes an audio-textual diffusion model for the UTI data generation task. In this model, the inherent acoustic characteristics of individuals related to the tongue motion details are encoded by using wav2vec 2.0, while the ASR transcriptions related to the universality of tongue motions are encoded by using BERT. UTI data are then generated by using a diffusion module. Experimental results showed that the proposed diffusion model could generate high-quality UTI data with clear tongue contour that is crucial for the linguistic analysis and clinical assessment. The codes and examples can be found on the website1. Yudong Yang, Rongfeng Su |
ICASSP | 1 |
| 2024 | Optical Flow Guided Tongue Trajectory Generation for Diffusion-based Acoustic to Articulatory Inversion
Yudong Yang, Rongfeng Su, Rukiye Ruzi, Manwa L. Ng, Shaofeng Zhao |
INTERSPEECH | 1 |
| 2019 | Virtual Wires: Rethinking WiFi networksabstractWiFi is the dominant means for home Internet access, yet is frequently a performance bottleneck. Without reliable, satisfactory performance at the last hop, end-to-end quality of service (QoS) efforts will fail. Three major reasons for WiFi bottlenecking performance are its: 1) inherent wireless channel characteristics, 2) approach to access control of the shared broadcast channel, and 3) impact on transport layer protocols, such as TCP, that operate end-to-end, and over-react to the loss or delay caused by the single WiFi link. In this paper, we leverage the philosophy of centralization in modern networking and present our cross layer design to address the problem. Specifically, we introduce centralized control at the point of entry/egress into the WiFi network. Based on network conditions measured from buffer sizes, airtime and throughput, flows are scheduled to the optimal utility. Unlike most existing WiFi QoS approaches, our design only relies on transparent modifications, requiring no changes to the network (including link layer) protocols, applications, or user intervention. Through extensive experimental investigation, we show that our design significantly enhances the reliability and predictability of WiFi performance, providing a “virtual wire”-like link to the targeted application. Yudong Yang, Yuming Jiang 0001, Vishal Misra, Dan Rubenstein |
LANMAN | 1 |
| 2017 | Improving Throughput and Fairness of Convergecast in Vehicular NetworksabstractDelivering data from source vehicles to infrastructures, or convergecast, is a fundamental operation in vehicular networks. However, the network capacity of vehicular network is always limited because of scarce inter-vehicle contacts. Thus, throughput maximization of convergecast in vehicular networks is of great importance. The unique characteristics of vehicular networks, however, present great challenges including frequent connection unavailability and opportunistic contacts. We propose an approach called ConvergeCode for improving the convergecast throughput in vehicular networks, which employs random linear coding for packet delivery. A vehicle randomly combines all received coded data and forwards it to any contacted vehicles. Through extensive empirical study based on the two large datasets of real GPS traces, we make the key observation that significant throughput gain can be achieved by using network coding but a serious fairness issue arises. In this paper, we study the problem of maximizing the throughput of convergecast in vehicular networks at the same time enhancing the fairness among different source nodes. We first formulate the problem of allocating inter-vehicle contacts as a lexicographical max-min multi-source flow problem, and then develop an efficient approximation algorithm with ε-approximation guarantee. Simulations based on real vehicular GPS traces have been performed and results show that the throughput is improved by 74-110 percent while the lexicographical max-min fairness is achieved. Ruobing Jiang, Yanmin Zhu 0006, Yudong Yang |
IEEE Trans. Mob. Comput. | 3 |
| 2013 | Price differentiation and control in the Kelly mechanism
Yudong Yang, Richard T. B. Ma, John C. S. Lui |
Perform. Evaluation | 1 |
| 2001 | HTML Page Analysis Based on Visual CuesabstractWe present an approach to automatically analyzing semantic structure of HTML pages based on detecting visual similarities of content objects on Web pages. The approach is developed based on the observation that in most Web pages, layout styles of subtitles or records of the same content category are consistent and there are apparent separation boundaries between different categories. Thus these subtitles should have similar appearances if they are rendered in visual browsers and different categories can be separated clearly In our approach we first measure visual similarities of HTML content objects. Then we apply a pattern detection algorithm to detect frequent patterns of visual similarity and use a number of heuristics to choose the most possible patterns. By grouping items according to these patterns, we finally build a hierarchical representation (tree) of an HTML document with "visual consistency" inferred semantics. Preliminary experimental results show promising performances of the method with real Web pages. Yudong Yang, HongJiang Zhang |
ICDAR | 1 |