EDBT 2026 Demo / reviewers in the wild / expert
Jimin Zhuang
dblp:388/2298
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 41% Vision and language · 41% Language models and text generation · 14% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model › multimodal large language model
audio-visual large language model |
0.9 | 1 | 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model · ICML 2025 |
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
0.9 | 1 | 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model · ICML 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model · ICML 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2025 | Improving LLM Video Understanding with 16 Frames Per Second · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
multimodal learning · 1.7visual token compression · 0.9supervised fine-tuning · 0.9speculative decoding · 0.9contrastive step selection · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio-centric Video Understanding Benchmark without Text ShortcutabstractYudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Yixuan Li, Peihan Li, Yifan Jiang, Wei Li, Zejun Ma, Chao Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Peihan Li, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
EMNLP | 2 |
| 2025 | Enabling Auditory Large Language Models for Automatic Speech Quality EvaluationabstractSpeech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN. Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang, Jimin Zhuang, Xianzhao Chen, Xiaohai Tian, Guangzhi Sun, Lu Lu 0015, Chao Zhang 0031 |
ICASSP | 6 |
| 2025 | Improving LLM Video Understanding with 16 Frames Per SecondabstractHuman vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual information loss. In this paper, we introduce F-16, the first multimodal LLM designed for high-frame-rate video understanding. By increasing the frame rate to 16 FPS and compressing visual tokens within each 1-second clip, F-16 efficiently captures dynamic visual features while preserving key semantic information.
Experimental results demonstrate that higher frame rates considerably enhance video understanding across multiple benchmarks, providing a new approach to improving video LLMs beyond scaling model size or training data. F-16 achieves state-of-the-art performance among 7-billion-parameter video LLMs on both general and fine-grained video understanding benchmarks, such as Video-MME and TemporalBench. Furthermore, F-16 excels in complex spatiotemporal tasks, including high-speed sports analysis (*e.g.*, basketball, football, gymnastics, and diving), outperforming SOTA proprietary visual models like GPT-4o and Gemini-1.5-pro.
Additionally, we introduce a novel decoding method for F-16 that enables highly efficient low-frame-rate inference without requiring model retraining. We will release the source code, model checkpoints, and data at [https://github.com/bytedance/F-16](https://github.com/bytedance/F-16). Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 3 |
| 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language ModelabstractWhile recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding. This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks. To enhance its reasoning abilities, we develop a reasoning-intensive dataset featuring challenging audio-visual questions with step-by-step solutions. We also propose process direct preference optimization (pDPO), which leverages contrastive step selection to achieve efficient step-level reward modelling tailored for multimodal inputs. Additionally, we introduce RivaBench, the first reasoning-intensive video understanding benchmark, featuring over 4,000 high-quality, expert-curated question-answer pairs across scenarios such as standup comedy, academic presentations, and synthetic video detection. video-SALMONN-o1 achieves 3-8% accuracy improvements over the LLaVA-OneVision baseline across different video reasoning benchmarks. Besides, pDPO achieves 6-8% improvements compared to the supervised fine-tuning model on RivaBench. Enhanced reasoning enables video-SALMONN-o1 zero-shot synthetic video detection capabilities. Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 3 |