VLDB 2026 Research / reviewers in the wild / expert
Yinghao Ma
dblp:248/7435
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 35% Generative modeling · 17% Speech recognition and synthesis · 17% | |
| Computer graphics and multimedia
4 papers |
Audio and music processing · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 44% Electronic design automation · 44% Integrated circuit design · 13% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
audio-language model |
0.9 | 1 | 2025 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix · NeurIPS 2025 |
Machine learning › Generative modeling › music generation
symbolic music generation |
0.9 | 1 | 2025 | MuPT: A Generative Symbolic Music Pretrained Transformer · ICLR 2025 |
Audio and music processing › music technology › computer music
symbolic music modeling |
0.9 | 1 | 2025 | MuPT: A Generative Symbolic Music Pretrained Transformer · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
continual pre-training |
0.8 | 1 | 2024 | D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model training |
0.8 | 1 | 2024 | D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
scaling laws |
0.8 | 1 | 2024 | D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised audio representation learning |
0.8 | 1 | 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training · ICLR 2024 |
Audio and music processing › music information retrieval
music understanding |
0.8 | 1 | 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training · ICLR 2024 |
Electronic design automation › logic synthesis › logic representation
majority-inverter graph |
0.8 | 1 | 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays · Sci. China Inf. Sci. 2024 |
Emerging computing paradigms
memristive computing |
0.8 | 1 | 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays · Sci. China Inf. Sci. 2024 |
Audio and music processing
music information retrieval |
0.7 | 1 | 2023 | MARBLE: Music Audio Representation Benchmark for Universal Evaluation · NeurIPS 2023 |
Medical and health informatics › mental health
mental health assessment |
0.3 | 1 | 2025 | Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression Detection · ACM Multimedia 2025 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.2 | 1 | 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays · Sci. China Inf. Sci. 2024 |
Methods — techniques the papers use, named apart from their topics
scaling laws · 1.7large language model · 1.7chain-of-thought · 1.7residual vector quantization · 1.5masked language modelling · 1.5constant-q transform · 1.5segment-level fusion · 0.9mean class variance · 0.9adaptive time-window selection · 0.9scaling law fitting · 0.8memristor logic synthesis · 0.8domain-specific pre-training · 0.8representation learning · 0.7pre-trained music language models · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from the high-level semantic features of multimodal data (video, audio, and text) generated by pretrained Large Language Models (LMMs), to enhance their representations, and the latter reliably integrates multiple predictions to further improve the robustness of emotion and intent understanding. Our RLF method achieved first place on Track 2 (Mandarin) of MEIJU, with performance scores for emotion, intent, and joint recognition reaching 0.7285, 0.7456, and 0.7370. Cheng Lu 0005, Yuyun Liu, Yinghao Ma, Jiahao Luo, Yuan Zong, Wenming Zheng |
ICASSP | 5 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 3 |
| 2025 | Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression DetectionabstractMajor Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675. Yuyun Liu, Kaifei Zhang, Yinghao Ma, Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
ACM Multimedia | 3 |
| 2025 | OmniBench: Towards The Future of Universal Omni-Language ModelsabstractRecent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/. Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002 |
NeurIPS | 3 |
| 2025 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their MixabstractWe introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area. Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001 |
NeurIPS | 2 |
| 2024 | Mertech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model with Multi-Task FinetuningabstractInstrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled data and inherent class imbalance issues. In this paper, we propose to apply a self-supervised learning model pre-trained on large-scale unlabeled music data and finetune it on IPT detection tasks. This approach addresses data scarcity and class imbalance challenges. Recognizing the significance of pitch in capturing the nuances of IPTs and the importance of onset in locating IPT events, we investigate multi-task finetuning with pitch and onset detection as auxiliary tasks. Additionally, we apply a post-processing approach for event-level prediction, where an IPT activation initiates an event only if the onset output confirms an onset in that frame. Our method outperforms prior approaches in both frame-level and event-level metrics across multiple IPT benchmark datasets. Further experiments demonstrate the efficacy of multi-task finetuning on each IPT class.1 Dichucheng Li, Yinghao Ma, Weixing Wei, Qiuqiang Kong, Yulun Wu 0002, Mingjin Che, Emmanouil Benetos, Wei Li 0012 |
ICASSP | 2 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 4 |
| 2024 | D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsabstractContinual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1\% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law. Haoran Que, Ge Zhang 0009, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan 0003, Jie Fu 0001, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007 |
NeurIPS | 6 |
| 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays
Zhouchao Gan, Yinghao Ma, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 4 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 2 |
| 2021 | FlowDiviner: Spatio-Temporal Network Traffic Prediction Method Based on Graph Neural NetworkabstractNo abstract available. Wenting Wei, Yinghao Ma |
APNet | 3 |