EDBT 2026 Demo / reviewers in the wild / expert
Wenhao Huang 0001
dblp:51/11-1
· DBLP profile ↗
38ranked-venue papers
6as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CriticLean: Critic-Guided Reinforcement Learning for Mathematical FormalizationabstractZhongyuan Peng, Yifan Yao, Kaijing Ma, Shuyue Guo, Yizhe Li, Yichi Zhang, Chenchen Zhang, Yifan Zhang, Zhouliang Yu, Luming Li, Minghao Liu, Yihang Xia, Jiawei Shen, Yuchen Wu, Yixin Cao, Zhaoxiang Zhang, Wenhao Huang, Jiaheng Liu, Ge Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhongyuan Peng, Kaijing Ma, Shuyue Guo, Yichi Zhang 0010, Zhouliang Yu, Luming Li, Minghao Liu 0003, Yihang Xia, Yixin Cao 0002, Zhaoxiang Zhang 0001, Wenhao Huang 0001, Ge Zhang 0009 |
ACL (1) | 17 |
| 2025 | Can MLLMs Understand the Deep Implication Behind Chinese Images?abstractAs the capabilities of Multimodal Large Language Models (MLLMs) improve, the need for higher-order evaluation of them is increasing. However, there is a lack of work evaluating MLLM for higher-order perception and understanding of Chinese visual content. To address this, we introduce the CII-Bench, which aims to assess MLLMs’ such capabilities for Chinese images. To ensure the authenticity of the Chinese context, images in CII-Bench are sourced from the Chinese Internet and manually reviewed, with corresponding answers also manually crafted. Additionally, CII-Bench incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, which can deeply reflect the model’s understanding of Chinese traditional culture. Through experiments on multiple MLLMs using CII-Bench, significant findings emerged. There is a large gap between MLLMs and humans in performance. The highest MLLM accuracy is 64.4%, while the human average is 78.2% and the peak is 81.0%. MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. Moreover, most models have higher accuracy when image emotion hints are added to the prompts. We believe CII-Bench will help MLLMs better understand Chinese semantics and specific images, and move forward the development of expert artificial general intelligence (AGI). Our project is publicly available at https://cii-bench.github.io. Chenhao Zhang 0005, Yuelin Bai, Xeron Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Xingwei Qu, Qixuan Zhao, Yiming Liang, Feiteng Fang, Min Yang 0007, Wenhao Huang 0001, Chenghua Lin 0002, Ge Zhang 0009, Shiwen Ni |
ACL (1) | 18 |
| 2025 | PopAlign: Diversifying Contrasting Patterns for a More Comprehensive AlignmentabstractZekun Moore Wang, Shenzhi Wang, King Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zekun Moore Wang, Shenzhi Wang, King Zhu, Ke Xu 0001, Jie Fu 0001, Wangchunshu Zhou, Wenhao Huang 0001 |
ACL (1) | 8 |
| 2025 | Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data CurationabstractWe introduce Presto, a novel video diffusion model designed to generate 15-second videos with long-range coherence and rich content. Extending video generation methods to maintain scenario diversity over long durations presents significant challenges. To address this, we propose a Segmented Cross-Attention (SCA) strategy, which splits hidden states into segments along the temporal dimension, allowing each segment to cross-attend to a corresponding sub-caption. SCA requires no additional parameters, enabling seamless incorporation into current DiT-based architectures. To facilitate high-quality long video generation, we build the LongTake-HD dataset, consisting of 261k content-rich videos with scenario coherence, annotated with an overall video caption and five progressive sub-captions. Experiments show that our Presto achieves 78.5% on the VBench Semantic Score and 100% on the Dynamic Degree, outperforming existing state-of-the-art video generation methods. This demonstrates that our proposed Presto significantly enhances content richness, maintains long-range coherence, and captures intricate textual details. More details are displayed on our project page: presto-video.github.io. Qiuyue Wang, Wenhao Huang 0001, Huan Yang 0005 |
CVPR | 5 |
| 2025 | MIO: A Foundation Model on Multimodal TokensabstractZekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001 |
EMNLP | 17 |
| 2025 | SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsabstractThe increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded in factual information (e.g. common and domain-specific knowledge). In this work, we introduce SimpleVQA, the first comprehensive multi-modal benchmark to evaluate the factuality ability of MLLMs to answer natural language short questions. SimpleVQA is characterized by six key features: it covers multiple tasks and multiple scenarios, ensures high quality and challenging queries, maintains static and timeless reference answers, and is straightforward to evaluate. Our approach involves categorizing visual question-answering items into 9 different tasks around objective events or common knowledge and situating these within 9 topics. Rigorous quality control processes are implemented to guarantee high-quality, concise, and clear answers, facilitating evaluation with minimal variance via an LLM-as-a-judge scoring system. Using SimpleVQA, we perform a comprehensive assessment of leading 18 MLLMs and 8 text-only LLMs, delving into their image comprehension and text generation abilities by identifying and analyzing error cases. Xianfu Cheng, Wei Zhang 0384, Jian Yang 0030, Xiangyuan Guan, Xianjie Wu, Xiang Li 0117, Ge Zhang 0009, Yuying Mai, Yutao Zeng, Zhoufutu Wen, Baorui Wang, Weixiao Zhou, Yunhong Lu, Hangyuan Ji, Tongliang Li, Wenhao Huang 0001, Zhoujun Li 0001 |
ICCV | 19 |
| 2025 | Steering Protein Family Design through Profile Bayesian FlowabstractProtein family design emerges as a promising alternative by combining the advantages of de novo protein design and mutation-based directed evolution.In this paper, we propose ProfileBFN, the Profile Bayesian Flow Networks, for specifically generative modeling of protein families. ProfileBFN extends the discrete Bayesian Flow Network from an MSA profile perspective, which can be trained on single protein sequences by regarding it as a degenerate profile, thereby achieving efficient protein family design by avoiding large-scale MSA data construction and training. Empirical results show that ProfileBFN has a profound understanding of proteins. When generating diverse and novel family proteins, it can accurately capture the structural characteristics of the family. The enzyme produced by this method is more likely than the previous approach to have the corresponding function, offering better odds of generating diverse proteins with the desired functionality. Jingjing Gong, Siyu Long, Yuxuan Song 0002, Wenhao Huang 0001, Ziyao Cao, Hao Zhou 0012, Wei-Ying Ma |
ICLR | 6 |
| 2025 | KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning TasksabstractIn this paper, we introduce Knowledge-Orthogonal Reasoning (KOR), a concept aimed at minimizing reliance on domain-specific knowledge, enabling more accurate evaluation of models' reasoning abilities in out-of-distribution settings.
Based on this concept, we propose the Knowledge-Orthogonal Reasoning Benchmark (KOR-Bench), encompassing five task categories: Operation, Logic, Cipher, Puzzle, and Counterfactual.
KOR-Bench emphasizes models' effectiveness in applying new rule descriptions to solve novel rule-driven questions. O1-Preview and O1-Mini achieve accuracies of 72.88\% and 70.16\%, surpassing Claude-3.5-Sonnet and GPT-4o (58.96\% and 58.00\%), highlighting the effectiveness of KOR-Bench.
We perform detailed analyses, identifying bottlenecks in the Cipher task with Stepwise Prompting, where two rounds of Self-Correction yield optimal results.
We evaluate performance across three integrated tasks, explore the impact of Tricks on the Puzzle task, and visualize rule-focused attention. Additionally, we conduct an ablation study on dataset size, benchmark correlations, and zero-shot and three-shot "only questions" experiments.
KOR-Bench aims to enhance reasoning evaluation and support further research in this area. Kaijing Ma, Xeron Du, Yunran Wang, Haoran Zhang 0013, Zhoufutu Wen, Xingwei Qu, Jian Yang 0037, Minghao Liu 0003, Xiang Yue, Wenhao Huang 0001, Ge Zhang 0009 |
ICLR | 11 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 26 |
| 2025 | FlexWorld: Progressively Expanding 3D Scenes for Flexible-View ExplorationabstractGenerating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework that progressively constructs a persistent 3D Gaussian splatting representation by synthesizing and integrating new 3D content. To handle novel view synthesis under large camera variations, we leverage an advanced pre-trained video model fine-tuned on accurate depth-estimated training pairs. By combining geometry-aware scene integration and optimization, FlexWorld refines the scene representation, producing visually consistent 3D scenes with flexible viewpoints. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Additionally, FlexWorld supports extrapolating from existing 3D scenes, further extending its applicability. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes that enable 360° rotations and zooming exploration. Our code is available at https://github.com/ML-GSAI/FlexWorld. Luxi Chen, Min Zhao 0013, Ge Zhang 0009, Wenhao Huang 0001, Hao Sun 0002, Ji-Rong Wen, Chongxuan Li |
NeurIPS | 6 |
| 2025 | OmniBench: Towards The Future of Universal Omni-Language ModelsabstractRecent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/. Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002 |
NeurIPS | 22 |
| 2025 | MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsabstractThe advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research. Yuanxing Zhang, Noah Wang, Ge Zhang 0009, Jian Yang 0037, Yanghai Wang, Xintao Wang 0002, Houyi Li, Wei Ji 0011, Pengfei Wan 0001, Wenhao Huang 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 14 |
| 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning EvaluationabstractRecent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments. Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009 |
NeurIPS | 28 |
| 2025 | DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent DiffusionabstractStory visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to create a coherent sequence of subject-consistent frames based solely on a story. To this end, we propose DreamStory, an automatic open-domain story visualization framework by leveraging the LLMs and a novel multi-subject consistent diffusion model. DreamStory consists of (1) an LLM acting as a story director and (2) an innovative Multi-Subject consistent Diffusion model (MSD) for generating consistent multi-subject across the images. First, DreamStory employs the LLM to generate descriptive prompts for subjects and scenes aligned with the story, annotating each scene's subjects for subsequent subject-consistent generation. Second, DreamStory utilizes these detailed subject descriptions to create portraits of the subjects, with these portraits and their corresponding textual information serving as multimodal anchors (guidance). Finally, the MSD uses these multimodal anchors to generate story scenes with consistent multi-subject. Specifically, the MSD includes Masked Mutual Self-Attention (MMSA) and Masked Mutual Cross-Attention (MMCA) modules. MMSA module ensures detailed appearance consistency with reference images, while MMCA captures key attributes of subjects from their reference text to ensure semantic consistency. Both modules employ masking mechanisms to restrict each scene's subjects to referencing the multimodal information of the corresponding subject, effectively preventing blending between multiple subjects. To validate our approach and promote progress in story visualization, we established a benchmark, DS-500, which can assess the overall performance of the story visualization framework, subject-identification accuracy, and the consistency of the generation model. Extensive experiments validate the effectiveness of DreamStory in both subjective and objective evaluations. Huiguo He, Huan Yang 0005, Zixi Tuo, Qiuyue Wang, Wenhao Huang 0001, Hongyang Chao, Jian Yin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM AgentsabstractQisen Yang, Zekun Wang, Honghui Chen, Shenzhi Wang, Yifan Pu, Xin Gao, Wenhao Huang, Shiji Song, Gao Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Qisen Yang, Honghui Chen, Shenzhi Wang, Yifan Pu, Wenhao Huang 0001, Shiji Song, Gao Huang 0001 |
ACL (1) | 7 |
| 2024 | CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor GenerationabstractMetaphor is a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication. This paper introduces a large-scale high quality annotated Chinese Metaphor Corpus, which comprises around 28K sentences drawn from a diverse range of Chinese literary sources, such as poems, prose, song lyrics, etc. To ensure the accuracy and consistency of our annotations, we introduce a comprehensive set of guidelines. These guidelines address the facets of metaphor annotation, including identifying tenors, vehicles, and grounds to handling the complexities of similes, personifications, juxtapositions, and hyperboles. Breaking tradition, our approach to metaphor generation emphasizes tenors and their distinct features rather than the conventional combination of tenors and vehicles. By integrating “ground” as a CoT (Chain of Thoughts) input, we are able to generate metaphors that resonate more with real-world intuition. We test generative models such as Belle, Baichuan, and Chinese-alpaca-33B using our annotated corpus. These models are able to generate creative and fluent metaphor sentences more frequently induced by selected samples from our dataset, demonstrating the value of our corpus for Chinese metaphor research. Yujie Shao, Xinrong Yao, Xingwei Qu, Chenghua Lin 0002, Shi Wang 0002, Wenhao Huang 0001, Ge Zhang 0009, Jie Fu 0001 |
LREC/COLING | 6 |
| 2024 | MORE-3S: Multimodal-based Offline Reinforcement Learning with Shared Semantic SpacesabstractDrawing upon the intuition that aligning different modalities to the same semantic embedding space would allow models to understand states and actions more easily, we propose a new perspective to the offline reinforcement learning (RL) challenge. More concretely, we transform it into a supervised learning task by integrating multimodal and pre-trained language models. Our approach incorporates state information derived from images and action-related data obtained from text, thereby bolstering RL training performance and promoting long-term strategic thinking. We emphasize the contextual understanding of language and demonstrate how decision-making in RL can benefit from aligning states’ and actions’ representation with languages’ representation. Our method significantly outperforms current baselines as evidenced by evaluations conducted on Atari and OpenAI Gym environments. This contributes to advancing offline RL performance and efficiency while providing a novel perspective on offline RL. Tianyu Zheng, Ge Zhang 0009, Xingwei Qu, Ming Kuang, Wenhao Huang 0001, Zhaofeng He 0001 |
LREC/COLING | 5 |
| 2024 | MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIabstractWe introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
CVPR | 19 |
| 2024 | MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageabstractMachine Translation (MT) has developed rapidly since the release of Large Language Models and current MT evaluation is performed through comparison with reference human translations or by predicting quality scores from human-labeled data.However, these mainstream evaluation methods mainly focus on fluency and factual reliability, whilst paying little attention to figurative quality.In this paper, we investigate the figurative quality of MT and propose a set of human evaluation metrics focused on the translation of figurative language.We additionally present a multilingual parallel metaphor corpus generated by postediting.Our evaluation protocol is designed to estimate four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality.In doing so, we observe that translations of figurative expressions display different traits from literal ones. Ge Zhang 0009, Tyler Loakman, Wenhao Huang 0001, Chenghua Lin 0002 |
EMNLP | 5 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 17 |
| 2024 | MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningabstractWe introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuning dataset. MathInstruct is compiled from 13 math datasets with intermediate rationales, six of which have rationales newly curated by us. It presents a unique hybrid of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and also ensures extensive coverage of diverse fields in math. The hybrid of CoT and PoT not only unleashes the potential of tool use but also allows different thought processes for different math problems. As a result, the MAmmoTH series substantially outperform existing open-source models on nine mathematical reasoning datasets across all scales with an average accuracy gain between 16% and 32%. Remarkably, our MAmmoTH-7B model reaches 33% on MATH (a competition-level dataset), which exceeds the best open-source 7B model (WizardMath) by 23%, and the MAmmoTH-34B model achieves 44% accuracy on MATH, even surpassing GPT-4’s CoT result. Our work underscores the importance of diverse problem coverage and the use of hybrid rationales in developing superior math generalist models. Xiang Yue, Xingwei Qu, Ge Zhang 0009, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
ICLR | 5 |
| 2024 | II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language ModelsabstractThe rapid advancements in the development of multimodal large language models (MLLMs) have consistently led to new breakthroughs on various benchmarks. In response, numerous challenging and comprehensive benchmarks have been proposed to more accurately assess the capabilities of MLLMs. However, there is a dearth of exploration of the higher-order perceptual capabilities of MLLMs. To fill this gap, we propose the Image Implication understanding Benchmark, II-Bench, which aims to evaluate the model's higher-order perception of images. Through extensive experiments on II-Bench across multiple MLLMs, we have made significant findings. Initially, a substantial gap is observed between the performance of MLLMs and humans on II-Bench. The pinnacle accuracy of MLLMs attains 74.8%, whereas human accuracy averages 90%, peaking at an impressive 98%. Subsequently, MLLMs perform worse on abstract and complex images, suggesting limitations in their ability to understand high-level semantics and capture image details. Finally, it is observed that most models exhibit enhanced accuracy when image sentiment polarity hints are incorporated into the prompts. This observation underscores a notable deficiency in their inherent understanding of image sentiment. We believe that II-Bench will inspire the community to develop the next generation of MLLMs, advancing the journey towards expert artificial general intelligence (AGI). II-Bench is publicly available at https://huggingface.co/datasets/m-a-p/II-Bench. Feiteng Fang, Xeron Du, Chenhao Zhang 0005, Noah Wang, Yuelin Bai, Qixuan Zhao, Liyang Fan, Chengguang Gan, Hongquan Lin, Jiaming Li 0004, Yuansheng Ni, Haihong Wu, Yaswanth Narsupalli, Zhigang Zheng, Chengming Li 0004, Xiping Hu, Ruifeng Xu 0001, Xiaojun Chen 0006, Min Yang 0007, Ruibo Liu, Wenhao Huang 0001, Ge Zhang 0009, Shiwen Ni |
NeurIPS | 24 |
| 2022 | PEMP: Leveraging Physics Properties to Enhance Molecular Property PredictionabstractMolecular property prediction is essential for drug discovery. In recent years, deep learning methods have been introduced to this area and achieved state-of-the-art performances. However, most of existing methods ignore the intrinsic relations between molecular properties which can be utilized to improve the performances of corresponding prediction tasks. In this paper, we propose a new approach, namely Physics properties Enhanced Molecular Property prediction (PEMP), to utilize relations between molecular properties revealed by previous physics theory and physical chemistry studies. Specifically, we enhance the training of the chemical and physiological property predictors with related physics property prediction tasks. We design two different methods for PEMP, respectively based on multi-task learning and transfer learning. Both methods include a model-agnostic molecule representation module and a property prediction module. In our implementation, we adopt both the state-of-the-art molecule embedding models under the supervised learning paradigm and the pretraining paradigm as the molecule representation module of PEMP, respectively. Experimental results on public benchmark MoleculeNet show that the proposed methods have the ability to outperform corresponding state-of-the-art models. Yuancheng Sun, Weizhi Ma, Wenhao Huang 0001, Kang Liu 0001, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
CIKM | 4 |
| 2020 | Real-time Transportation Prediction Correction using Reconstruction Error in Deep LearningabstractIn online complex systems such as transportation system, an important work is real-time traffic prediction. Due to the data shift, data model inconsistency, and sudden change of traffic patterns (like transportation accident), the prediction result derived from an offline-built model would be unreliable. Retraining the model is usually not time affordable for online prediction, especially when the prediction model is very complex and costs a lot of training time (for example, deep neural networks). A real-time prediction correction strategy would be of great value under this situation. Traditionally, the prediction correction usually relies on the prediction error in several previous time intervals. They assume that the error pattern is similar in the current time interval, so that it is time-delayed to some extent. In this article, we propose the prediction correction strategy using the reconstruction error in the deep neural network. The reconstruction error can reflect the model’s ability on feature representation and then determine the fitness of an input data to the model. We first build the relationship between reconstruction error and prediction error. From the perspective of the prediction interval, we demonstrate that the reconstruction error is in positive relation with the prediction interval. Thus the prediction result is more reliable when the reconstruction error is smaller. Then we propose two mechanisms of real-time prediction correction using the reconstruction error. The data driven prediction correction approach selects several training instances with similar reconstruction errors to the current instance and using their average prediction error in correcting the prediction result. The model-driven approach builds several component deep neural networks in training. The component training set for each network is selected according to the reconstruction error of training instances. For a predicting instance, it first computes the reconstruction error of the sample in each component network and then averages the results by the reconstruction error and prediction interval. The model-driven approach is actually a reconstruction error-based deep neural network ensemble approach. Finally, a series of experiments demonstrated that reconstruction error based prediction correction approaches are effective in several prediction problems in transportation including traffic flow prediction on road, traffic flow prediction in entrance and exit station and travel time prediction. Besides the high overall accuracy, our approach can also provide many observations of using the reconstruction error in transportation prediction. Shuai Liu 0018, Guojie Song, Wenhao Huang 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2019 | Learning Personalized End-to-End Goal-Oriented DialogabstractMost existing works on dialog systems only consider conversation content while neglecting the personality of the user the bot is interacting with, which begets several unsolved issues. In this paper, we present a personalized end-to-end model in an attempt to leverage personalization in goal-oriented dialogs. We first introduce a PROFILE MODEL which encodes user profiles into distributed embeddings and refers to conversation history from other similar users. Then a PREFERENCE MODEL captures user preferences over knowledge base entities to handle the ambiguity in user requests. The two models are combined into the PERSONALIZED MEMN2N. Experiments show that the proposed model achieves qualitative performance improvements over state-of-the-art methods. As for human evaluation, it also outperforms other approaches in terms of task completion rate and user satisfaction. Liangchen Luo, Wenhao Huang 0001, Qi Zeng 0001, Zaiqing Nie, Xu Sun 0001 |
AAAI | 2 |
| 2019 | Text Assisted Insight Ranking Using Context-Aware Memory Network
Qi Zeng 0001, Liangchen Luo, Wenhao Huang 0001 |
AAAI | 3 |
| 2015 | Learning Common Metrics for Homogenous Tasks in Traffic Flow PredictionabstractNearest neighbor based nonparametric regression is a classic data-driven method for traffic flow prediction in intelligent transportation systems (ITS). Performances of those models depend heavily on the similarity or distance metric used to search nearest neighborhood. Metric learning algorithms have been developed to learn the distance metrics from data in recent years. In real-world transportation application, multiple forecasting tasks are set since there are lots of road sections and detector points in the traffic network. Previous works tend to learn only one global metric to be used for all the tasks or learn multiple local metrics for each task which may lead to under-fitting or over-fitting problem. To balance these two kinds of methods and improve the generalization of learned metrics, we propose a common metric learning algorithm under the intuition that homogenous tasks tend to have similar local metrics. Then the learned common metrics are used in common metric KNN (CM-KNN) for traffic flow prediction. Experimental results show that our algorithm to learn common metrics are reasonable and CM-KNN method for traffic flow prediction outperforms other competing methods. Haikun Hong, Xiabing Zhou, Wenhao Huang 0001, Xingxing Xing, Kaigui Bian, Kunqing Xie |
ICMLA | 3 |
| 2015 | Improving deep neural network ensembles using reconstruction errorabstractEnsemble learning of neural network is a learning paradigm where ensembles of several neural networks show improved generalization capabilities that outperform those of single networks. For deep learning of multi-layer neural networks, ensemble learning is still applicable. In addition, characteristics of deep neural networks can provide potential opportunities to improve the performance of traditional neural network ensembles. In this paper, we propose an ensemble criterion of deep neural networks that is based on the reconstruction error and present two strategies to solve the most important issues in ensemble learning of neural networks: component dataset sampling and output averaging. Component training datasets are selected according to the reconstruction error instead of random bootstrap sampling or re-weighting. Moreover, for each testing instance, we can compute the reconstruction error yielded by the sub-model simultaneously with the output. The reconstruction error is used as the weights in output averaging. From the perspectives of prediction interval and confidence interval, we demonstrated that smaller reconstruction error could ensure smaller prediction interval. We also incorporate the famous structure ensemble approach “Dropout” into the proposed approach to achieve the best performance. We conduct experiments on classification and regression datasets to validate the effectiveness of our approach. Wenhao Huang 0001, Haikun Hong, Kaigui Bian, Xiabing Zhou, Guojie Song, Kunqing Xie |
IJCNN | 1 |
| 2015 | Probabilistic dynamic causal model for temporal dataabstractLearning temporal causal structures between time series is one of key tools for analyzing time series data. Most previous works focuse on learning with static temporal causal relationships. However, in many real world applications, such as climate environment and transportation system, the causal structures vary dramatically over time. In this paper, we propose a probabilistic dynamic causal (PDC) model based on Lasso-Granger to uncover the dynamic temporal dependencies. Specifically, the PDC model infers different state varying of temporal data and causal structures of each state in one unified model. We devise the expectation-maximization (EM) algorithm to infer the model parameters. Furthermore, to address the smoothness of state varying in adjacent time, we extend the PDC model with a regularization term encouraging states to be similar in adjacent time. Though it may slightly decrease the precision on training data, it improves the generalization capability of the model. We conduct experiments on synthetic dataset as well as two real-world datasets of climate and traffic to evaluate the effectiveness of the PDC model. Experimental results show that the proposed model is effective in discovering the dynamic causal factors of Particulate Matter 2.5 (PM2.5) and traffic spatial causalities. Xiabing Zhou, Wenhao Huang 0001, Weisong Hu, Sizhen Du, Guojie Song, Kunqing Xie |
IJCNN | 2 |
| 2015 | Mining Dependencies Considering Time Lag in Spatio-Temporal Traffic Data
Xiabing Zhou, Haikun Hong, Xingxing Xing, Wenhao Huang 0001, Kaigui Bian, Kunqing Xie |
WAIM | 4 |
| 2014 | Deep process neural network for temporal deep learningabstractProcess neural network is widely used in modeling temporal process inputs in neural networks. Traditional process neural network is usually limited in structure of single hidden layer due to the unfavorable training strategies of neural network with multiple hidden layers and complex temporal weights in process neural network. Deep learning has emerged as an effective pre-training method for neural network with multiple hidden layers. Though deep learning is usually limited in static inputs, it provided us a good solution for training neural network with multiple hidden layers. In this paper, we extended process neural network to deep process neural network. Two basic structures of deep process neural network are discussed. One is the accumulation first deep process neural network and the other is accumulation last deep process neural network. We could build any architecture of deep process neural network based on those two structures. Temporal process inputs are represented as sequences in this work for the purpose of unsupervised feature learning with less prior knowledge. Based on this, we proposed learning algorithms for two basic structures inspired by the numerical learning approach for process neural network and the auto-encoder in deep learning. Finally, extensive experiments demonstrated that deep process neural network is effective in tasks with temporal process inputs. Accuracy of deep process neural network is higher than traditional process neural network while time complexity is near in the task of traffic flow prediction in highway system. Wenhao Huang 0001, Haikun Hong, Guojie Song, Kunqing Xie |
IJCNN | 1 |
| 2014 | Dynamic boosting in deep learning using reconstruction errorabstractDeep learning has attracted a lot of attention in research and industry in recent years. Behind the success of deep learning, there is much space for improvement. It is difficult to identify if a testing sample can be represented by the deep network effectively before we examining the final result. In this paper, we proposed a dynamic boosting strategy according to reconstruction error in deep networks. We use reconstruction error to determine whether the result is reliable or not. From the perspective of prediction interval, we demonstrated that with the increase of reconstruction error, the prediction interval would become bigger. Therefore, the classification result is not reliable when the reconstruction error exceeds the predetermined threshold. Since we can record the reconstruction error as well as the classification error for all training samples in training set. We can learn an extra boosting model besides the deep network in training set to improve the performance of the model. An important factor in learning the boosting model is to determine an appropriate threshold for selecting training samples. In testing, we first examine whether the reconstruction error of a testing sample exceeds the threshold to determine if we should use the boosting model. If the boosting model is used, the final result is the average of the output of the deep network and the boosting model. We conducted experiments on two widely used classification datasets and an air quality dataset. From the experiments, we see that our boosting strategy is effective in improving the performance of classification. We tested several boosting models in this paper. They can all reduce the test error to some extent under appropriate parameter settings. Wenhao Huang 0001, Weisong Hu, Haikun Hong, Guojie Song, Kunqing Xie |
IJCNN | 1 |
| 2014 | Metric-Based Multi-Task Grouping Neural Network for Traffic Flow Forecasting
Haikun Hong, Wenhao Huang 0001, Guojie Song, Kunqing Xie |
ISNN | 2 |
| 2014 | A Spatial-temporal Topic Segmentation Model for Human Mobile Behavior
Xingxing Xing, Weisong Hu, Wenhao Huang 0001, Guojie Song, Kunqing Xie |
WAIM | 4 |
| 2014 | Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask LearningabstractTraffic flow prediction is a fundamental problem in transportation modeling and management. Many existing approaches fail to provide favorable results due to being: 1) shallow in architecture; 2) hand engineered in features; and 3) separate in learning. In this paper we propose a deep architecture that consists of two parts, i.e., a deep belief network (DBN) at the bottom and a multitask regression layer at the top. A DBN is employed here for unsupervised feature learning. It can learn effective features for traffic flow prediction in an unsupervised fashion, which has been examined and found to be effective for many areas such as image and audio classification. To the best of our knowledge, this is the first paper that applies the deep learning approach to transportation research. To incorporate multitask learning (MTL) in our deep architecture, a multitask regression layer is used above the DBN for supervised prediction. We further investigate homogeneous MTL and heterogeneous MTL for traffic flow prediction. To take full advantage of weight sharing in our deep architecture, we propose a grouping method based on the weights in the top layer to make MTL more effective. Experiments on transportation data sets show good performance of our deep architecture. Abundant experiments show that our approach achieved close to 5% improvements over the state of the art. It is also presented that MTL can improve the generalization performance of shared tasks. These positive results demonstrate that deep learning and MTL are promising in transportation research. Wenhao Huang 0001, Guojie Song, Haikun Hong, Kunqing Xie |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Deep Architecture for Traffic Flow Prediction
Wenhao Huang 0001, Haikun Hong, Weisong Hu, Guojie Song, Kunqing Xie |
ADMA (2) | 1 |
| 2011 | Discrete Trajectory Prediction on Mobile Data
Wenhao Huang 0001, Guojie Song, Kunqing Xie |
APWeb | 2 |
| 2010 | Anchor Points Seeking of Large Urban Crowd Based on the Mobile Billing Data
Wenhao Huang 0001, Zhengbin Dong, Guojie Song, Kunqing Xie |
ADMA (1) | 1 |