EDBT 2026 Demo / reviewers in the wild / expert
Xiang Yue
dblp:176/6718
· DBLP profile ↗
52ranked-venue papers
11as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 9 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional VideosabstractKairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li, Yuanhan Zhang, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li 0080, Yuanhan Zhang, Ziwei Liu 0002 |
ACL (1) | 5 |
| 2026 | Temporal Sampling for Forgotten Reasoning in LLMsabstractYuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian, Luyao Niu, Bill Yuchen Lin, Xiang Yue, Radha Poovendran. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian, Luyao Niu, Bill Y. Lin, Xiang Yue, Radha Poovendran |
ACL (1) | 7 |
| 2025 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleabstractJiawei Guo, Tianyu Zheng, Yizhi Li, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianyu Zheng, Yuelin Bai, Bo Li 0080, Yubo Wang 0019, King Zhu, Graham Neubig, Wenhu Chen, Xiang Yue |
ACL (1) | 10 |
| 2025 | Evaluating Language Models as Synthetic Data GeneratorsabstractSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan 0002, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig |
ACL (1) | 3 |
| 2025 | MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkabstractXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 0019, Kai Zhang 0033, Shengbang Tong, Yuxuan Sun 0002, Botao Yu, Ge Zhang 0009, Huan Sun 0001, Yu Su 0001, Wenhu Chen, Graham Neubig |
ACL (1) | 1 |
| 2025 | Evaluating Vision-Language Models as Evaluators in Path PlanningabstractDespite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they still contribute to the planning framework as a helpful plan evaluator? In this work, we generalize this question to consider LLMs augmented with visual understanding, i.e., Vision-Language Models (VLMs). We introduce PathEval, a novel benchmark evaluating VLMs as plan evaluators in complex path-planning scenarios. Succeeding in the benchmark requires a VLM to be able to abstract traits of optimal paths from the scenario description, demonstrate precise low-level perception on each path, and integrate this information to decide the better path. Our analysis of state-of-the-art VLMs reveals that these models face significant challenges on the benchmark. We observe that the VLMs can precisely abstract given scenarios to identify the desired traits and exhibit mixed performance in integrating the provided information. Yet, their vision component presents a critical bottleneck, with models struggling to perceive low-level details about a path. Our experimental results show that this issue cannot be trivially addressed via end-to-end fine-tuning; rather, task-specific discriminative adaptation of these vision encoders is needed for these VLMs to become effective path evaluators.12 Mohamed Aghzal, Xiang Yue, Erion Plaku, Ziyu Yao 0002 |
CVPR | 2 |
| 2025 | VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearchabstractVision-Language Models have made significant progress on many perception-focused tasks.However, their progress on reasoningfocused tasks remains limited due to the lack of high-quality and diverse training data.In this work, we aim to address the scarcity of reasoning-focused multimodal datasets.We propose VisualWebInstruct, a novel approach that leverages search engines to create a diverse and high-quality dataset spanning multiple disciplines, including mathematics, physics, finance, and chemistry, etc. Starting with a meticulously selected set of 30,000 seed images, we employ Google Image Search to identify websites containing similar images.We collect and process HTML data from over 700K unique URLs.Through a pipeline of content extraction, filtering, and synthesis, we construct a dataset of approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs and the remaining comprising text-based QA pairs.Models finetuned on VisualWebInstruct demonstrate significant performance improvements: (1) finetuning on Llava-OV results in 10-20 absolute points improvement across benchmarks, and (2) fine-tuning from MAmmoTH-VL yields a 5 absolute points gain across benchmarks.Our best model, MAmmoTH-VL2, achieves the best known performance with SFT without RL within the 10B parameter class on MMMU-Pro (40.7),MathVerse (42.6), and DynaMath (55.7).These results highlight the effectiveness of our dataset in enhancing the reasoning capabilities of vision-language models for complex multimodal tasks. Yiming Jia, Xiang Yue, Bo Li 0080, Ping Nie, Wenhu Chen |
EMNLP | 3 |
| 2025 | MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World TasksabstractWe present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users.
Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation.
In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MM-Bench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks.
Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions. Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang 0068, Yubo Wang 0019, Yuansheng Ni, Ziyan Jiang, Wang Zhu 0001, Bohan Lyu 0001, Dongfu Jiang, Hexiang Hu, Xiang Yue, Wenhu Chen |
ICLR | 15 |
| 2025 | Harnessing Webpage UIs for Text-Rich Visual UnderstandingabstractText-rich visual understanding—the ability to interpret both textual content and visual elements within a scene—is crucial for multimodal large language models (MLLMs) to effectively interact with structured environments. We propose leveraging webpage UIs as a naturally structured and diverse data source to enhance MLLMs’ capabilities in this area. Existing approaches, such as rule-based extraction, multimodal model captioning, and rigid HTML parsing, are hindered by issues like noise, hallucinations, and limited generalization. To overcome these challenges, we introduce MultiUI, a dataset of 7.3 million samples spanning various UI types and tasks, structured using enhanced accessibility trees and task taxonomies. By scaling multimodal instructions from web UIs through LLMs, our dataset enhances generalization beyond web domains, significantly improving performance in document understanding, GUI comprehension, grounding, and advanced agent tasks. This demonstrates the potential of structured web data to elevate MLLMs’ proficiency in processing text-rich visual environments and generalizing across domains. Junpeng Liu 0001, Tianyue Ou, Yifan Song 0002, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, Xiang Yue |
ICLR | 9 |
| 2025 | KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning TasksabstractIn this paper, we introduce Knowledge-Orthogonal Reasoning (KOR), a concept aimed at minimizing reliance on domain-specific knowledge, enabling more accurate evaluation of models' reasoning abilities in out-of-distribution settings.
Based on this concept, we propose the Knowledge-Orthogonal Reasoning Benchmark (KOR-Bench), encompassing five task categories: Operation, Logic, Cipher, Puzzle, and Counterfactual.
KOR-Bench emphasizes models' effectiveness in applying new rule descriptions to solve novel rule-driven questions. O1-Preview and O1-Mini achieve accuracies of 72.88\% and 70.16\%, surpassing Claude-3.5-Sonnet and GPT-4o (58.96\% and 58.00\%), highlighting the effectiveness of KOR-Bench.
We perform detailed analyses, identifying bottlenecks in the Cipher task with Stepwise Prompting, where two rounds of Self-Correction yield optimal results.
We evaluate performance across three integrated tasks, explore the impact of Tricks on the Puzzle task, and visualize rule-focused attention. Additionally, we conduct an ablation study on dataset size, benchmark correlations, and zero-shot and three-shot "only questions" experiments.
KOR-Bench aims to enhance reasoning evaluation and support further research in this area. Kaijing Ma, Xeron Du, Yunran Wang, Haoran Zhang 0013, Zhoufutu Wen, Xingwei Qu, Jian Yang 0037, Minghao Liu 0003, Xiang Yue, Wenhao Huang 0001, Ge Zhang 0009 |
ICLR | 10 |
| 2025 | MixEval-X: Any-to-any Evaluations from Real-world Data MixtureabstractPerceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying protocols and maturity levels; and (2) significant query, grading, and generalization biases. To address these, we introduce MixEval-X, the first any-to-any, real-world benchmark designed to optimize and standardize evaluations across diverse input and output modalities. We propose multi-modal benchmark mixture and adaptation-rectification pipelines to reconstruct real-world task distributions, ensuring evaluations generalize effectively to real-world use cases. Extensive meta-evaluations show our approach effectively aligns benchmark samples with real-world task distributions. Meanwhile, MixEval-X's model rankings correlate strongly with that of crowd-sourced real-world evaluations (up to 0.98) while being much more efficient. We provide comprehensive leaderboards to rerank existing models and organizations and offer insights to enhance understanding of multi-modal evaluations and inform future research. Jinjie Ni, Deepanway Ghosal, Bo Li 0080, Junhao Zhang 0001, Xiang Yue, Fuzhao Xue, Yuntian Deng, Zian Zheng 0001, Kaichen Zhang, Mahir Shah, Kabir Jain, Yang You 0001, Michael Shieh |
ICLR | 6 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 23 |
| 2025 | Pangea: A Fully Open Multilingual Multimodal LLM for 39 LanguagesabstractDespite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented.
This paper introduces PANGEA, a multilingual multimodal LLM trained on PANGEAINS, a diverse 6M instruction dataset spanning 39 languages. PANGEAINS features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage.
To rigorously assess models' capabilities, we introduce PANGEABENCH, a holistic evaluation suite encompassing 14 datasets covering 47 languages.
Results show that PANGEA significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum. Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim, Jean de Dieu Nyandwi, Simran Khanuja, Anjali Kantharuban, Lintang Sutawika, Sathyanarayanan Ramamoorthy, Graham Neubig |
ICLR | 1 |
| 2025 | Overtrained Language Models Are Harder to Fine-TuneabstractLarge language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work, we challenge this assumption and show that extended pre-training can make models harder to fine-tune, leading to degraded final performance. We term this phenomenon \textbf{catastrophic overtraining}. For example, the instruction-tuned OLMo-1B model pre-trained on 3T tokens leads to over 2\% worse performance on multiple standard LLM benchmarks than its 2.3T token counterpart. Through controlled experiments and theoretical analysis, we show that catastrophic overtraining arises from a systematic increase in the broad sensitivity of pre-trained parameters to modifications, including but not limited to fine-tuning. Our findings call for a critical reassessment of pre-training design that considers the downstream adaptability of the model. Jacob Mitchell Springer, Sachin Goyal, Kaiyue Wen, Tanishq Kumar, Xiang Yue, Sadhika Malladi, Graham Neubig, Aditi Raghunathan |
ICML | 5 |
| 2025 | Underestimated Privacy Risks for Minority Populations in Large Language Model UnlearningabstractLarge Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods. Although certified unlearning offers strong privacy guarantees, its restrictive assumptions make it unsuitable for LLMs, giving rise to various heuristic approaches typically assessed through empirical evaluations. These standard evaluations randomly select data for removal, apply unlearning techniques, and use membership inference attacks (MIAs) to compare unlearned models against models retrained without the removed data. However, to ensure robust privacy protections for every data point, it is essential to account for scenarios in which certain data subsets face elevated risks. Prior research suggests that outliers, particularly including data tied to minority groups, often exhibit higher memorization propensity which indicates they may be more difficult to unlearn. Building on these insights, we introduce a complementary, minority-aware evaluation framework to highlight blind spots in existing frameworks. We substantiate our findings with carefully designed experiments, using canaries with personally identifiable information (PII) to represent these minority subsets and demonstrate that they suffer at least 20\% higher privacy leakage across various unlearning methods, MIAs, datasets, and LLM scales. Our proposed minority-aware evaluation framework marks an essential step toward more equitable and comprehensive assessments of LLM unlearning efficacy. Rongzhe Wei, Mufei Li, Mohsen Ghassemi, Eleonora Kreacic, Xiang Yue, Bo Li 0026, Vamsi K. Potluru, Pan Li 0005, Eli Chien |
ICML | 6 |
| 2025 | Demystifying Long Chain-of-Thought ReasoningabstractScaling inference compute has become a key driver of advanced reasoning in large language models (LLMs). A proven approach for scaling inference compute is to generate long chains-of-thought (CoTs), enabling models to engage in structured reasoning strategies such as backtracking and error correction. Reinforcement learning (RL) has emerged as a crucial method for developing these capabilities, yet the conditions under which long CoTs emerge remain unclear, and RL training requires careful design choices. In this study, we systematically investigate the underlying mechanics of long CoT reasoning—examining the factors that enable models to generate extended reasoning trajectories. Through extensive supervised fine-tuning (SFT) and RL experiments, we identify three key findings: 1) while SFT is not strictly necessary, it significantly simplifies training and improves efficiency; 2) reasoning capabilities tend to emerge with increased training compute but are not guaranteed, making reward shaping essential for stabilizing CoT length growth; and 3) scaling verifiable reward signals is critical for RL, and we find that leveraging noisy, web-extracted solutions with filtering mechanisms shows promising potential, particularly in out-of-distribution (OOD) reasoning tasks such as STEM problem-solving. These insights provide practical guidance for optimizing training strategies to enhance long CoT reasoning in LLMs. Shiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig, Xiang Yue |
ICML | 5 |
| 2025 | OpusLM: A Family of Open Unified Speech Language Models
Jinchuan Tian, Yifan Peng 0003, Jiatong Shi, Siddhant Arora, Shikhar Bharadwaj, Takashi Maekaku, Yusuke Shinohara, Keita Goto, Xiang Yue, Chao-Han Huck Yang, Shinji Watanabe 0001 |
INTERSPEECH | 10 |
| 2025 | JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware EvaluationabstractShota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa |
NAACL (Long Papers) | 6 |
| 2025 | Deep learning for recognition and detection of plant diseases and pests
Xiang Yue, Kai Qi, Xinyi Na, Fuhao Yang |
Neural Comput. Appl. | 1 |
| 2024 | VIEScore: Towards Explainable Metrics for Conditional Image Synthesis EvaluationabstractIn the rapidly advancing field of conditional image generation research, challenges such as limited explainability lie in effectively evaluating the performance and capabilities of various models.This paper introduces VIESCORE, a Visual Instruction-guided Explainable metric for evaluating any conditional image generation tasks.VIESCORE leverages general knowledge from Multimodal Large Language Models (MLLMs) as the backbone and does not require training or fine-tuning.We evaluate VIESCORE on seven prominent tasks in conditional image tasks and found: (1) VIESCORE (GPT4-o) achieves a high Spearman correlation of 0.4 with human evaluations, while the human-to-human correlation is 0.45.(2) VI-ESCORE (with open-source MLLM) is significantly weaker than GPT-4o and GPT-4v in evaluating synthetic images.(3) VIESCORE achieves a correlation on par with human ratings in the generation tasks but struggles in editing tasks.With these results, we believe VIESCORE shows its great potential to replace human judges in evaluating image synthesis tasks. Max Ku, Dongfu Jiang, Cong Wei 0001, Xiang Yue, Wenhu Chen |
ACL (1) | 4 |
| 2024 | Trial and Error: Exploration-Based Trajectory Optimization of LLM AgentsabstractLarge Language Models (LLMs) have become integral components in various autonomous agent systems.In this study, we present an exploration-based trajectory optimization approach, referred to as ETO.This learning method is designed to enhance the performance of open LLM agents.Contrary to previous studies that exclusively train on successful expert trajectories, our method allows agents to learn from their exploration failures.This leads to improved performance through an iterative optimization framework.During the exploration phase, the agent interacts with the environment while completing given tasks, gathering failure trajectories to create contrastive trajectory pairs.In the subsequent training phase, the agent utilizes these trajectory preference pairs to update its policy using contrastive learning methods like DPO (Rafailov et al., 2023).This iterative cycle of exploration and training fosters continued improvement in the agents.Our experiments on three complex tasks demonstrate that ETO consistently surpasses baseline performance by a large margin.Furthermore, an examination of task-solving efficiency and potential in scenarios lacking expert trajectory underscores the effectiveness of our approach.1 Yifan Song 0002, Da Yin, Xiang Yue, Sujian Li, Bill Y. Lin |
ACL (1) | 3 |
| 2024 | Machine Unlearning of Pre-trained Large Language ModelsabstractThis study investigates the concept of the 'right to be forgotten' within the context of large language models (LLMs).We explore machine unlearning as a pivotal solution, with a focus on pre-trained models-a notably under-researched area.Our research delineates a comprehensive framework for machine unlearning in pretrained LLMs, encompassing a critical analysis of seven diverse unlearning methods.Through rigorous evaluation using curated datasets from arXiv, books, and GitHub, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 10 5 times more computationally efficient than retraining.Our results show that integrating gradient ascent with gradient descent on in-distribution data improves hyperparameter robustness.We also provide detailed guidelines for efficient hyperparameter tuning in the unlearning process.Our findings advance the discourse on ethical AI practices, offering substantive insights into the mechanics of machine unlearning for pretrained LLMs and underscoring the potential for responsible AI development.1 Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang 0001, Zezhou Cheng, Xiang Yue |
ACL (1) | 7 |
| 2024 | MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIabstractWe introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence. Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
CVPR | 1 |
| 2024 | MAmmoTH: Building Math Generalist Models through Hybrid Instruction TuningabstractWe introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, our meticulously curated instruction tuning dataset. MathInstruct is compiled from 13 math datasets with intermediate rationales, six of which have rationales newly curated by us. It presents a unique hybrid of chain-of-thought (CoT) and program-of-thought (PoT) rationales, and also ensures extensive coverage of diverse fields in math. The hybrid of CoT and PoT not only unleashes the potential of tool use but also allows different thought processes for different math problems. As a result, the MAmmoTH series substantially outperform existing open-source models on nine mathematical reasoning datasets across all scales with an average accuracy gain between 16% and 32%. Remarkably, our MAmmoTH-7B model reaches 33% on MATH (a competition-level dataset), which exceeds the best open-source 7B model (WizardMath) by 23%, and the MAmmoTH-34B model achieves 44% accuracy on MATH, even surpassing GPT-4’s CoT result. Our work underscores the importance of diverse problem coverage and the use of hybrid rationales in developing superior math generalist models. Xiang Yue, Xingwei Qu, Ge Zhang 0009, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen |
ICLR | 1 |
| 2024 | Data Engineering for Scaling Language Models to 128K ContextabstractWe study continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in particular *the ability to utilize information at arbitrary input locations*, is a capability that is mostly already acquired through large-scale pretraining, and that this capability can be readily extended to contexts substantially longer than seen during training (e.g., 4K to 128K) through lightweight continual pretraining on appropriate data mixture. We investigate the *quantity* and *quality* of the data for continual pretraining: (1) for quantity, we show that 500 million to 5 billion tokens are enough to enable the model to retrieve information anywhere within the 128K context; (2) for quality, our results equally emphasize *domain balance* and *length upsampling*. Concretely, naïvely upsampling longer data on certain domains like books, a common practice of existing work, gives suboptimal performance; a balanced domain mixture is equally important. We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K. Our recipe outperforms strong open-source long-context models and closes the gap to frontier models like GPT-4 128K. Rameswar Panda, Xinyao Niu, Xiang Yue, Hannaneh Hajishirzi, Hao Peng 0018 |
ICML | 4 |
| 2024 | TableLlama: Towards Open Large Generalist Models for TablesabstractTianshu Zhang, Xiang Yue, Yifei Li, Huan Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tianshu Zhang 0001, Xiang Yue, Yifei Li 0005, Huan Sun 0001 |
NAACL-HLT | 2 |
| 2024 | MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesabstractEvaluating large language models (LLMs) is challenging. Traditional ground-truth- based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User- facing evaluation, such as Chatbot Arena, provides reliable signals but is costly and slow. In this work, we propose MixEval, a new paradigm for establishing efficient, gold-standard LLM evaluation by strategically mixing off-the-shelf bench- marks. It bridges (1) comprehensive and well-distributed real-world user queries and (2) efficient and fairly-graded ground-truth-based benchmarks, by matching queries mined from the web with similar queries from existing benchmarks. Based on MixEval, we further build MixEval-Hard, which offers more room for model improvement. Our benchmarks’ advantages lie in (1) a 0.96 model ranking correlation with Chatbot Arena arising from the highly impartial query distribution and grading mechanism, (2) fast, cheap, and reproducible execution (6% of the time and cost of MMLU), and (3) dynamic evaluation enabled by the rapid and stable data update pipeline. We provide extensive meta-evaluation and analysis for our and existing LLM benchmarks to deepen the community’s understanding of LLM evaluation and guide future research directions. Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, Yang You 0001 |
NeurIPS | 3 |
| 2024 | MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding BenchmarkabstractIn the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates part of the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16\% to 33\% compared to MMLU, but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5\% in MMLU to just 2\% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is more discriminative benchmark to better track progress in the field. Yubo Wang 0019, Xueguang Ma, Ge Zhang 0009, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang 0068, Alex Zhuang, Rongqi Fan, Xiang Yue, Wenhu Chen |
NeurIPS | 16 |
| 2024 | Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of GeneralizationabstractWe study whether transformers can learn to *implicitly* reason over parametric knowledge, a skill that even the most capable language models struggle with. Focusing on two representative reasoning types, composition and comparison, we consistently find that transformers *can* learn implicit reasoning, but only through *grokking*, i.e., extended training far beyond overfitting. The levels of generalization also vary across reasoning types: when faced with out-of-distribution examples, transformers fail to systematically generalize for composition but succeed for comparison. We delve into the model's internals throughout training, conducting analytical experiments that reveal: 1) the mechanism behind grokking, such as the formation of the generalizing circuit and its relation to the relative efficiency of generalizing and memorizing circuits, and 2) the connection between systematicity and the configuration of the generalizing circuit. Our findings guide data and training setup to better induce implicit reasoning and suggest potential improvements to the transformer architecture, such as encouraging cross-layer knowledge sharing. Furthermore, we demonstrate that for a challenging reasoning task with a large search space, GPT-4-Turbo and Gemini-1.5-Pro based on non-parametric memory fail badly regardless of prompting styles or retrieval augmentation, while a fully grokked transformer can achieve near-perfect accuracy, showcasing the power of parametric memory for complex reasoning. Boshi Wang, Xiang Yue, Yu Su 0001, Huan Sun 0001 |
NeurIPS | 2 |
| 2024 | MAmmoTH2: Scaling Instructions from the WebabstractInstruction tuning improves the reasoning abilities of large language models (LLMs), with data quality and scalability being the crucial factors. Most instruction tuning data come from human crowd-sourcing or GPT-4 distillation. We propose a paradigm to efficiently harvest 10 million naturally existing instruction data from the pre-training web corpus to enhance LLM reasoning. Our approach involves (1) recalling relevant documents, (2) extracting instruction-response pairs, and (3) refining the extracted pairs using open-source LLMs. Fine-tuning base LLMs on this dataset, we build MAmmoTH2 models, which significantly boost performance on reasoning benchmarks. Notably, MAmmoTH2-7B’s (Mistral) performance increases from 11% to 36.7% on MATH and from 36% to 68.4% on GSM8K without training on any in-domain data. Further training MAmmoTH2 on public instruction tuning datasets yields MAmmoTH2-Plus, achieving state-of-the-art performance on several reasoning and chatbot benchmarks. Our work demonstrates how to harvest large-scale, high-quality instruction data without costly human annotation or GPT-4 distillation, providing a new paradigm for building better instruction tuning data. Xiang Yue, Tianyu Zheng, Ge Zhang 0009, Wenhu Chen |
NeurIPS | 1 |
| 2023 | Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeabstractXiang Yue, Huseyin Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, Robert Sim. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiang Yue, Huseyin A. Inan, Julia McAnallen, Hoda Shajari, Huan Sun 0001, David Levitan, Robert Sim |
ACL (1) | 1 |
| 2023 | DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassabstractDifferentially private stochastic gradient descent (DP-SGD) adds noise to gradients in back-propagation, safeguarding training data from privacy leakage, particularly membership inference. It fails to cover (inference-time) threats like embedding inversion and sensitive attribute inference. It is also costly in storage and computation when used to fine-tune large pre-trained language models (LMs). Minxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 0001, Huan Sun 0001 |
CCS | 2 |
| 2023 | Roll Up Your Sleeves: Working with a Collaborative and Engaging Task-Oriented Dialogue SystemabstractLingbo Mo, Shijie Chen, Ziru Chen, Xiang Deng, Ashley Lewis, Sunit Singh, Samuel Stevens, Chang-You Tai, Zhen Wang, Xiang Yue, Tianshu Zhang, Yu Su, Huan Sun. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Lingbo Mo, Ziru Chen, Xiang Deng 0001, Ashley Lewis, Sunit Singh, Samuel Stevens 0001, Chang-You Tai, Zhen Wang 0041, Xiang Yue, Tianshu Zhang 0001, Yu Su 0001, Huan Sun 0001 |
SIGDIAL | 10 |
| 2023 | Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyabstractDifferentially private (DP) learning, notably DP stochastic gradient descent (DP-SGD), has limited applicability in fine-tuning gigantic pre-trained language models (LMs) for natural language processing tasks. The culprit is the perturbation of gradients (as gigantic as entire models), leading to significant efficiency and accuracy drops. Minxin Du, Xiang Yue, Sherman S. M. Chow, Huan Sun 0001 |
WWW | 2 |
| 2022 | Synthetic Question Value Estimation for Domain Adaptation of Question AnsweringabstractSynthesizing QA pairs with a question generator (QG) on the target domain has become a popular approach for domain adaptation of question answering (QA) models.Since synthetic questions are often noisy in practice, existing work adapts scores from a pretrained QA (or QG) model as criteria to select highquality questions.However, these scores do not directly serve the ultimate goal of improving QA performance on the target domain.In this paper, we introduce a novel idea of training a question value estimator (QVE) that directly estimates the usefulness of synthetic questions for improving the target-domain QA performance.By conducting comprehensive experiments, we show that the synthetic questions selected by QVE can help achieve better target-domain QA performance, in comparison with existing techniques.We additionally show that by using such questions and only around 15% of the human annotations on the target domain, we can achieve comparable performance to the fully-supervised baselines.1 Xiang Yue, Ziyu Yao 0002, Huan Sun 0001 |
ACL (1) | 1 |
| 2021 | CliniQG4QA: Generating Diverse Questions for Domain Adaptation of Clinical Question AnsweringabstractClinical question answering (QA) aims to automatically answer questions from medical professionals based on clinical texts. Studies show that neural QA models trained on one corpus may not generalize well to new clinical texts from a different institute or a different patient group, where largescale QA pairs are not readily available for model retraining. To address this challenge, we propose a simple yet effective framework, CliniQG4QA, which leverages question generation (QG) to synthesize QA pairs on new clinical contexts and boosts QA models without requiring manual annotations. In order to generate diverse types of questions that are essential for training QA models, we further introduce a seq2seq-based question phrase prediction (QPP) module that can be used together with most existing QG models to diversify the generation. Our comprehensive experiment results show that the QA corpus generated by our framework can improve QA models on the new contexts (up to 8% absolute gain in terms of Exact Match), and that the QPP module plays a crucial role in achieving the gain.11Our dataset and code are available at: https://github.com/sunlabosu/CliniQG4QA/. Xiang Yue, Xinliang Frederick Zhang, Ziyu Yao 0002, Simon M. Lin, Huan Sun 0001 |
BIBM | 1 |
| 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ RetrievalabstractWe present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: . Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001 |
EMNLP (1) | 3 |
| 2021 | Tensor decomposition with relational constraints for predicting multiple types of microRNA-disease associationsabstractMicroRNAs (miRNAs) play crucial roles in multifarious biological processes associated with human diseases. Identifying potential miRNA-disease associations contributes to understanding the molecular mechanisms of miRNA-related diseases. Most of the existing computational methods mainly focus on predicting whether a miRNA-disease association exists or not. However, the roles of miRNAs in diseases are prominently diverged, for instance, Genetic variants of miRNA (mir-15) may affect the expression level of miRNAs leading to B cell chronic lymphocytic leukemia, while circulating miRNAs (including mir-1246, mir-1307-3p, etc.) have potentials to detecting breast cancer in the early stage. In this paper, we aim to predict multi-type miRNA-disease associations instead of taking them as binary. To this end, we innovatively represent miRNA-disease-type triples as a tensor and introduce tensor decomposition methods to solve the prediction task. Experimental results on two widely-adopted miRNA-disease datasets: HMDD v2.0 and HMDD v3.2 show that tensor decomposition methods improve a recent baseline in a large scale (up to $38\%$ in Top-1F1). We then propose a novel method, Tensor Decomposition with Relational Constraints (TDRC), which incorporates biological features as relational constraints to further the existing tensor decomposition methods. Compared with two existing tensor decomposition methods, TDRC can produce better performance while being more efficient. Feng Huang 0004, Xiang Yue, Zhankun Xiong, Zhouxin Yu, Shichao Liu 0002, Wen Zhang 0008 |
Briefings Bioinform. | 2 |
| 2020 | Clinical Reading Comprehension: A Thorough Analysis of the emrQA DatasetabstractMachine reading comprehension has made great progress in recent years owing to largescale annotated datasets.In the clinical domain, however, creating such datasets is quite difficult due to the domain expertise required for annotation.Recently, Pampari et al. (2018) tackled this issue by using expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering (QA) based on clinical notes.In this paper, we provide an indepth analysis of this dataset and the clinical reading comprehension (CliniRC) task.From our qualitative analysis, we find that (i) emrQA answers are often incomplete, and (ii) emrQA questions are often answerable without using domain knowledge.From our quantitative experiments, surprising results include that (iii) using a small sampled subset (5%-20%), we can obtain roughly equal performance compared to the model trained on the entire dataset, (iv) this performance is close to human expert's performance, and (v) BERT models do not beat the best performing base model.Following our analysis of the emrQA, we further explore two desired aspects of CliniRC systems: the ability to utilize clinical domain knowledge and to generalize to unseen questions and contexts.We argue that both should be considered when creating future datasets.1 Xiang Yue, Bernal Jimenez Gutierrez, Huan Sun 0001 |
ACL | 1 |
| 2020 | Clinical Phrase Mining with Language ModelsabstractA vast amount of vital clinical data is available within unstructured texts such as discharge summaries and procedure notes in Electronic Medical Records (EMRs). Automatically transforming such unstructured data into structured units is crucial for effective data analysis in the field of clinical informatics. Recognizing phrases that reveal important medical information in a concise and thorough manner is a fundamental step in this process. Existing systems that are built for opendomain texts are designed to detect mostly non-medical phrases, while tools designed specifically for extracting concepts from clinical texts are not scalable to large corpora and often leave out essential context surrounding those detected clinical concepts. We address these issues by proposing a framework, CliniPhrase, which adapts domain-specific deep neural network based language models (such as ClinicalBERT) to effectively and efficiently extract high-quality phrases from clinical documents with a limited amount of training data. Experimental results on the MIMIC-III dataset show that our method can outperform the current state-of-the-art techniques by up to 18% in terms of F1measure while being very efficient (up to 48 times faster). Kaushik Mani, Xiang Yue, Bernal Jimenez Gutierrez, Yungui Huang, Simon M. Lin, Huan Sun 0001 |
BIBM | 2 |
| 2020 | Towards Making the Most of Context in Neural Machine TranslationabstractDocument-level machine translation manages to outperform sentence level models by a small margin, but have failed to be widely adopted. We argue that previous research did not make a clear use of the global context, and propose a new document-level NMT framework that deliberately models the local context of each sentence with the awareness of the global context of the document in both source and target languages. We specifically design the model to be able to deal with documents containing any number of sentences, including single sentences. This unified approach allows our model to be trained elegantly on standard datasets without needing to train on sentence and document level data separately. Experimental results demonstrate that our model outperforms Transformer baselines and previous document-level NMT models with substantial margins of up to 2.1 BLEU on state-of-the-art baselines. We also provide analyses which show the benefit of context far beyond the neighboring two or three sentences, which previous studies have typically incorporated. Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen 0001, Alexandra Birch |
IJCAI | 2 |
| 2020 | Graph embedding on biomedical networks: methods, applications and evaluationsabstractMOTIVATION: Graph embedding learning that aims to automatically learn low-dimensional node representations, has drawn increasing attention in recent years. To date, most recent graph embedding methods are evaluated on social and information networks and are not comprehensively studied on biomedical networks under systematic experiments and analyses. On the other hand, for a variety of biomedical network analysis tasks, traditional techniques such as matrix factorization (which can be seen as a type of graph embedding methods) have shown promising results, and hence there is a need to systematically evaluate the more recent graph embedding methods (e.g. random walk-based and neural network-based) in terms of their usability and potential to further the state-of-the-art. RESULTS: We select 11 representative graph embedding methods and conduct a systematic comparison on 3 important biomedical link prediction tasks: drug-disease association (DDA) prediction, drug-drug interaction (DDI) prediction, protein-protein interaction (PPI) prediction; and 2 node classification tasks: medical term semantic type classification, protein function prediction. Our experimental results demonstrate that the recent graph embedding methods achieve promising results and deserve more attention in the future biomedical graph analysis. Compared with three state-of-the-art methods for DDAs, DDIs and protein function predictions, the recent graph embedding methods achieve competitive performance without using any biological features and the learned embeddings can be treated as complementary representations for the biological features. By summarizing the experimental results, we provide general guidelines for properly selecting graph embedding methods and setting their hyper-parameters for different biomedical tasks. AVAILABILITY AND IMPLEMENTATION: As part of our contributions in the paper, we develop an easy-to-use Python package with detailed instructions, BioNEV, available at: https://github.com/xiangyue9607/BioNEV, including all source code and datasets, to facilitate studying various graph embedding methods on biomedical tasks. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiang Yue, Zhen Wang 0041, Jingong Huang, Srinivasan Parthasarathy 0001, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Wen Zhang 0008, Ping Zhang 0016, Huan Sun 0001 |
Bioinform. | 1 |
| 2019 | LncRNA-miRNA interaction prediction from the heterogeneous network through graph embedding ensemble learningabstractLncRNA-miRNA interactions play crucial roles in gene regulatory networks and can reveal functions of lncRNAs and miRNAs. Although several methods have been proposed to infer interactions on the lncRNA-miRNA interaction network, few attentions have been paid to fully exploiting the structure of lncRNA-miRNA interaction network. In this paper, we propose a Graph Embedding Ensemble Learning method (abbreviated as “GEEL”) to predict lncRNA-miRNA interactions. First, we collect lncRNA sequences and miRNA sequences to calculate lncRNA-lncRNA sequence similarity and miRNA-miRNA sequence similarity, and then we combine them with the known lncRNA-miRNA interactions to construct a heterogeneous network, which takes lncRNAs and miRNAs as nodes. We adopt graph embedding methods to learn representations of lncRNAs and miRNAs from the heterogeneous network, and then merge the representations of lncRNAs and miRNAs to represent the lncRNA-miRNA pairs. Random forest classifiers are built based on the merged representations to predict lncRNA-miRNA interactions. We consider five different graph embedding methods, and evaluate the corresponding models. Further, we use individual graph embedding method-based model as base predictors and build a high-level ensemble model. The experimental results show that GEEL achieves AUPR score of 0.7004 and AUC score of 0.9537, and outperforms base predictors and other state-of-the-art methods. In conclusion, GEEL is an effective tool for lncRNA-miRNA interaction prediction. Shuang Zhou 0009, Xiang Yue, Xinran Xu, Shichao Liu 0002, Wen Zhang 0008, Yanqing Niu |
BIBM | 2 |
| 2019 | SurfCon: Synonym Discovery on Privacy-Aware Clinical DataabstractUnstructured clinical texts contain rich health-related information. To better utilize the knowledge buried in clinical texts, discovering synonyms for a medical query term has become an important task. Recent automatic synonym discovery methods leveraging raw text information have been developed. However, to preserve patient privacy and security, it is usually quite difficult to get access to large-scale raw clinical texts. In this paper, we study a new setting named synonym discovery on privacy-aware clinical data (i.e., medical terms extracted from the clinical texts and their aggregated co-occurrence counts, without raw clinical texts). To solve the problem, we propose a new framework SurfCon that leverages two important types of information in the privacy-aware clinical data, i.e., the surface form information, and the global context information for synonym discovery. In particular, the surface form module enables us to detect synonyms that look similar while the global context module plays a complementary role to discover synonyms that are semantically similar but in different surface forms, and both allow us to deal with the OOV query issue (i.e., when the query is not found in the given data). We conduct extensive experiments and case studies on publicly available privacy-aware clinical data, and show that SurfCon can outperform strong baseline methods by large margins under various settings. Zhen Wang 0041, Xiang Yue, Soheil Moosavinasab, Yungui Huang, Simon M. Lin, Huan Sun 0001 |
KDD | 2 |
| 2019 | Prediction of drug-disease associations based on ensemble meta paths and singular value decompositionabstractBACKGROUND: In the field of drug repositioning, it is assumed that similar drugs may treat similar diseases, therefore many existing computational methods need to compute the similarities of drugs and diseases. However, the calculation of similarity depends on the adopted measure and the available features, which may lead that the similarity scores vary dramatically from one to another, and it will not work when facing the incomplete data. Besides, supervised learning based methods usually need both positive and negative samples to train the prediction models, whereas in drug-disease pairs data there are only some verified interactions (positive samples) and a lot of unlabeled pairs. To train the models, many methods simply treat the unlabeled samples as negative ones, which may introduce artificial noises. Herein, we propose a method to predict drug-disease associations without the need of similarity information, and select more likely negative samples. RESULTS: In the proposed EMP-SVD (Ensemble Meta Paths and Singular Value Decomposition), we introduce five meta paths corresponding to different kinds of interaction data, and for each meta path we generate a commuting matrix. Every matrix is factorized into two low rank matrices by SVD which are used for the latent features of drugs and diseases respectively. The features are combined to represent drug-disease pairs. We build a base classifier via Random Forest for each meta path and five base classifiers are combined as the final ensemble classifier. In order to train out a more reliable prediction model, we select more likely negative ones from unlabeled samples under the assumption that non-associated drug and disease pair have no common interacted proteins. The experiments have shown that the proposed EMP-SVD method outperforms several state-of-the-art approaches. Case studies by literature investigation have found that the proposed EMP-SVD can mine out many drug-disease associations, which implies the practicality of EMP-SVD. CONCLUSIONS: The proposed EMP-SVD can integrate the interaction data among drugs, proteins and diseases, and predict the drug-disease associations without the need of similarity information. At the same time, the strategy of selecting more reliable negative samples will benefit the prediction. Guangsheng Wu, Juan Liu 0007, Xiang Yue |
BMC Bioinform. | 3 |
| 2018 | Prediction of Drug-Disease Associations and Their Effects by Signed Network-Based Nonnegative Matrix Factorization
Wen Zhang 0008, Feng Huang 0004, Xiang Yue, Xiaoting Lu, Weitai Yang, Zhishuai Li |
BIBM | 3 |
| 2018 | Sequence-based bacterial small RNAs prediction using ensemble learning strategiesabstractBACKGROUND: Bacterial small non-coding RNAs (sRNAs) have emerged as important elements in diverse physiological processes, including growth, development, cell proliferation, differentiation, metabolic reactions and carbon metabolism, and attract great attention. Accurate prediction of sRNAs is important and challenging, and helps to explore functions and mechanism of sRNAs. RESULTS: In this paper, we utilize a variety of sRNA sequence-derived features to develop ensemble learning methods for the sRNA prediction. First, we compile a balanced dataset and four imbalanced datasets. Then, we investigate various sRNA sequence-derived features, such as spectrum profile, mismatch profile, reverse compliment k-mer and pseudo nucleotide composition. Finally, we consider two ensemble learning strategies to integrate all features for building ensemble learning models for the sRNA prediction. One is the weighted average ensemble method (WAEM), which uses the linear weighted sum of outputs from the individual feature-based predictors to predict sRNAs. The other is the neural network ensemble method (NNEM), which trains a deep neural network by combining diverse features. In the computational experiments, we evaluate our methods on these five datasets by using 5-fold cross validation. WAEM and NNEM can produce better results than existing state-of-the-art sRNA prediction methods. CONCLUSIONS: WAEM and NNEM have great potential for the sRNA prediction, and are helpful for understanding the biological mechanism of bacteria. Guifeng Tang, Jingwen Shi, Wenjian Wu, Xiang Yue, Wen Zhang 0008 |
BMC Bioinform. | 4 |
| 2018 | Predicting drug-disease associations by using similarity constrained matrix factorizationabstractBACKGROUND: Drug-disease associations provide important information for the drug discovery. Wet experiments that identify drug-disease associations are time-consuming and expensive. However, many drug-disease associations are still unobserved or unknown. The development of computational methods for predicting unobserved drug-disease associations is an important and urgent task. RESULTS: In this paper, we proposed a similarity constrained matrix factorization method for the drug-disease association prediction (SCMFDD), which makes use of known drug-disease associations, drug features and disease semantic information. SCMFDD projects the drug-disease association relationship into two low-rank spaces, which uncover latent features for drugs and diseases, and then introduces drug feature-based similarities and disease semantic similarity as constraints for drugs and diseases in low-rank spaces. Different from the classic matrix factorization technique, SCMFDD takes the biological context of the problem into account. In computational experiments, the proposed method can produce high-accuracy performances on benchmark datasets, and outperform existing state-of-the-art prediction methods when evaluated by five-fold cross validation and independent testing. CONCLUSION: We developed a user-friendly web server by using known associations collected from the CTD database, available at http://www.bioinfotech.cn/SCMFDD/ . The case studies show that the server can find out novel associations, which are not included in the CTD database. Wen Zhang 0008, Xiang Yue, Weiran Lin, Wenjian Wu, Ruoqi Liu, Feng Huang 0004 |
BMC Bioinform. | 2 |
| 2018 | Manifold regularized matrix factorization for drug-drug interaction prediction
Wen Zhang 0008, Yanlin Chen 0002, Dingfang Li, Xiang Yue |
J. Biomed. Informatics | 4 |
| 2018 | SFPEL-LPI: Sequence-based feature projection ensemble learning for predicting LncRNA-protein interactionsabstractLncRNA-protein interactions play important roles in post-transcriptional gene regulation, poly-adenylation, splicing and translation. Identification of lncRNA-protein interactions helps to understand lncRNA-related activities. Existing computational methods utilize multiple lncRNA features or multiple protein features to predict lncRNA-protein interactions, but features are not available for all lncRNAs or proteins; most of existing methods are not capable of predicting interacting proteins (or lncRNAs) for new lncRNAs (or proteins), which don't have known interactions. In this paper, we propose the sequence-based feature projection ensemble learning method, "SFPEL-LPI", to predict lncRNA-protein interactions. First, SFPEL-LPI extracts lncRNA sequence-based features and protein sequence-based features. Second, SFPEL-LPI calculates multiple lncRNA-lncRNA similarities and protein-protein similarities by using lncRNA sequences, protein sequences and known lncRNA-protein interactions. Then, SFPEL-LPI combines multiple similarities and multiple features with a feature projection ensemble learning frame. In computational experiments, SFPEL-LPI accurately predicts lncRNA-protein associations and outperforms other state-of-the-art methods. More importantly, SFPEL-LPI can be applied to new lncRNAs (or proteins). The case studies demonstrate that our method can find out novel lncRNA-protein interactions, which are confirmed by literature. Finally, we construct a user-friendly web server, available at http://www.bioinfotech.cn/SFPEL-LPI/. Wen Zhang 0008, Xiang Yue, Guifeng Tang, Wenjian Wu, Feng Huang 0004 |
PLoS Comput. Biol. | 2 |
| 2017 | Predicting small RNAs in bacteria via sequence learning ensemble methodabstractBacterial small non-coding RNAs (sRNAs) play important roles in various physiological processes, and predicting sRNAs is an important task. In this paper, we develop a computational method for the sRNA prediction by using sRNA sequence-derived features. We investigate a variety of sRNA sequence-derived features, and evaluate the usefulness of features for the sRNA prediction. Then, we develop the sequence learning ensemble method, which uses the linear weighted sum of outputs from the individual feature-based predictors to predict sRNAs, and the genetic algorithm is adopted to optimize the parameters in the ensemble system. In the computational experiments, we compile a balanced dataset and four imbalanced datasets, and evaluate our method on these datasets by using 5-fold cross validation. The sequence learning ensemble method can achieve AUC scores greater than 0.9, and outperforms existing state-of-the-art sRNA prediction methods. In conclusion, the proposed method has a great potential for sRNA prediction. The source codes, datasets and supplementary are available in http://www.bioinfotech.cn/BIBM2017/SLEM. Wen Zhang 0008, Jingwen Shi, Guifeng Tang, Wenjian Wu, Xiang Yue, Dingfang Li |
BIBM | 5 |
| 2017 | Predicting drug-disease associations based on the known association bipartite networkabstractRecent studies show that drug-disease associations provide important information for drug discovery and drug repositioning. Wet experimental identification of drug-disease associations is time-consuming and labor-intensive. Therefore, the development of computational methods that predict drug-disease associations is an urgent task. In this paper, we propose a novel computational method named NTSIM, which only uses known drug-disease associations to predict unobserved associations. First of all, known drug-disease associations are represented as a drug-disease bipartite network, and a novel similarity measure named linear neighborhood similarity (LNS) is proposed to calculate drug-drug similarity and disease-disease similarity based on the bipartite network. Then, we predict unobserved drug-disease associations in the similarity-based graph by using label propagation process. In the computational experiments, this proposed method achieves high-accuracy performances, and outperforms representative state-of-the-art methods: PREDICT, TL-HGBI and LRSSL. Our studies reveal that known drug-disease associations can provide enough information to build the high-accuracy prediction models; linear neighbor similarity (LNS) can lead to better performances than other similarity measures such as Jaccard similarity, Gauss similarity and cosine similarity; the bipartite network-derived features outperform the drug biological features and disease semantic features. Wen Zhang 0008, Xiang Yue, Yanlin Chen 0002, Weiran Lin, Bolin Li, Xiaohong Li 0003 |
BIBM | 2 |