Siru Ouyang

dblp:282/4345 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
0009-0001-1331-424XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cell-o1 : training LLMs to solve single-cell reasoning puzzles with reinforcement learning
abstract
Abstract Motivation Large language models (LLMs) have demonstrated strong general reasoning abilities, but applying them to domain-specific tasks such as analysing single-cell RNA sequencing data remains a challenge. A central task in this domain is cell type annotation, which is critical for understanding cellular heterogeneity. Although recent foundation models attempt to automate this process, they typically annotate cells independently, without considering batch-level context or providing explanatory reasoning. To address this limitation, we introduce the CellPuzzles benchmark, which reformulates cell type annotation as a batch-level reasoning task. CellPuzzles spans diverse tissues, diseases, and donor conditions, and requires reasoning across the batch-level cellular context to ensure label uniqueness. Results We find that off-the-shelf LLMs struggle on this task, with the best baseline (OpenAI o1) achieving only 19.0% batch-level accuracy. To fill this gap, we propose Cell-o1, a 7B LLM trained via supervised fine-tuning on distilled reasoning traces, followed by reinforcement learning with batch-level rewards. Cell-o1 achieves state-of-the-art performance, outperforming OpenAI o1 by over 73% and generalizing well across contexts. Further analysis of training dynamics and reasoning behaviors provides insights into batch-level annotation performance and emergent expert-like reasoning. Availability and Implementation Code and data are available at https://github.com/ncbi-nlp/cell-o1.
Yin Fang, Qiao Jin 0001, Guangzhi Xiong, Bowen Jin, Xianrui Zhong, Siru Ouyang, Yifan Yang 0006, Aidong Zhang 0001, Jiawei Han 0001, Zhiyong Lu
Bioinform.6
2026 Discourse-Aware Language Representation
abstract
Recent Transformer-based language representation techniques have commonly adopted a straightforward approach to modeling textual context as a linear sequence of successive tokens. However, this sequential modeling strategy falls short in actively exploring intermediate structures present in natural languages and does not account for the rich interactive relationships between sentences. To overcome these limitations, we propose a discourse-aware framework that bridges the gap between sequential contextualization and the interactive nature of conversational reading comprehension. Concretely, we first divide the context into elementary discourse units (EDUs), ensuring that each unit contains precisely one condition. Then, we systematically explore three instantiations for modeling discourse features: sequential EDU encoding, discourse-aware masking, and discourse graph network. These techniques allow us to capture the nuanced interactions within the discourse. To assess the efficacy of our methodologies, we perform experiments on three conversational reading comprehension tasks: multi-turn response selection, conversational question answering, and conversational machine reading. Experimental results demonstrate the superiority of our proposed approach. Moreover, analysis reveals that the discourse-aware approach enables the model to effectively capture intricate relationships within the context and fosters reasoning interpretability. Additionally, our method exhibits efficacy across various backbone PLMs and diverse domains.
Zhuosheng Zhang 0001, Siru Ouyang, Hai Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Synergizing Unsupervised Episode Detection with LLMs for Large-Scale News Events
abstract
State-of-the-art automatic event detection struggles with interpretability and adaptability to evolving large-scale key events-unlike episodic structures, which excel in these areas.Often overlooked, episodes represent cohesive clusters of core entities performing actions at a specific time and location; a partially ordered sequence of episodes can represent a key event.This paper introduces a novel task, episode detection, which identifies episodes within a news corpus of key event articles.Detecting episodes poses unique challenges, as they lack explicit temporal or locational markers and cannot be merged using semantic similarity alone.While large language models (LLMs) can aid with these reasoning difficulties, they suffer with long contexts typical of news corpora.To address these challenges, we introduce EpiMine, an unsupervised framework that identifies a key event's candidate episodes by leveraging natural episodic partitions in articles, estimated through shifts in discriminative term combinations.These candidate episodes are more cohesive and representative of true episodes, synergizing with LLMs to better interpret and refine them into final episodes.We apply EpiMine to our three diverse, real-world event datasets annotated at the episode level, where it achieves a 59.2% average gain across all metrics compared to baselines.
Priyanka Kargupta, Yunyi Zhang 0001, Yizhu Jiao, Siru Ouyang, Jiawei Han 0001
ACL (1)4
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR1
2025 ChemAgent: Self-updating Memories in Large Language Models Improves Chemical Reasoning
abstract
Chemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large language models (LLMs) encounter difficulties handling domain-specific formulas, executing reasoning steps accurately, and integrating code ef- effectively when tackling chemical reasoning tasks. To address these challenges, we present ChemAgent, a novel framework designed to improve the performance of LLMs through a dynamic, self-updating library. This library is developed by decomposing chemical tasks into sub-tasks and compiling these sub-tasks into a structured collection that can be referenced for future queries. Then, when presented with a new problem, ChemAgent retrieves and refines pertinent information from the library, which we call memory, facilitating effective task decomposition and the generation of solutions. Our method designs three types of memory and a library-enhanced reasoning component, enabling LLMs to improve over time through experience. Experimental results on four chemical reasoning datasets from SciBench demonstrate that ChemAgent achieves performance gains of up to 46% (GPT-4), significantly outperforming existing methods. Our findings suggest substantial potential for future applications, including tasks such as drug discovery and materials science. Our code can be found at https://github.com/gersteinlab/ChemAgent.
Xiangru Tang, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang 0001, Yilun Zhao 0001, Arman Cohan, Mark Gerstein
ICLR6
2025 Retrieval And Structuring Augmented Generation with Large Language Models
abstract
Large Language Models (LLMs) have revolutionized natural language processing with their remarkable capabilities in text generation and reasoning. However, these models face critical challenges when deployed in real-world applications, including hallucination generation, outdated knowledge, and limited domain expertise. Retrieval And Structuring (RAS) Augmented Generation addresses these limitations by integrating dynamic information retrieval with structured knowledge representations. This survey (1) examines retrieval mechanisms including sparse, dense, and hybrid approaches for accessing external knowledge; (2) explore text structuring techniques such as taxonomy construction, hierarchical classification, and information extraction that transform unstructured text into organized representations; and (3) investigate how these structured representations integrate with LLMs through prompt-based methods, reasoning frameworks, and knowledge embedding techniques. It also identifies technical challenges in retrieval efficiency, structure quality, and knowledge integration, while highlighting research opportunities in multimodal retrieval, cross-lingual structures, and interactive systems. This comprehensive overview provides researchers and practitioners with insights into RAS methods, applications, and future directions.
Pengcheng Jiang, Siru Ouyang, Yizhu Jiao, Ming Zhong 0005, Runchu Tian, Jiawei Han 0001
KDD (2)2
2025 FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models
abstract
Large language models (LLMs) have gained significant attention in chemistry. However, most existing datasets center on molecular-level property prediction and overlook the role of fine-grained functional group (FG) information. Incorporating FG-level data can provide valuable prior knowledge that links molecular structures with textual descriptions, which can be used to build more interpretable, structure-aware LLMs for reasoning on molecule-related tasks. Moreover, LLMs can learn from such fine-grained information to uncover hidden relationships between specific functional groups and molecular properties, thereby advancing molecular design and drug discovery. Here, we introduce FGBench, a dataset comprising 625K molecular property reasoning problems with functional group information. Functional groups are precisely annotated and localized within the molecule, which ensures the dataset's interoperability thereby facilitating further multimodal applications. FGBench includes both regression and classification tasks on 245 different functional groups across three categories for molecular property reasoning: (1) single functional group impacts, (2) multiple functional group interactions, and (3) direct molecular comparisons. In the benchmark of state-of-the-art LLMs on 7K curated data, the results indicate that current LLMs struggle with FG-level property reasoning, highlighting the need to enhance reasoning capabilities in LLMs for chemistry tasks. We anticipate that the methodology employed in FGBench to construct datasets with functional group-level information will serve as a foundational framework for generating new question–answer pairs, enabling LLMs to better understand fine-grained molecular structure–property relationships. The dataset and evaluation code are available at this \href{https://github.com/xuanliugit/FGBench}{link}.
Xuan Liu 0009, Siru Ouyang, Xianrui Zhong, Jiawei Han 0001, Huimin Zhao 0007
NeurIPS2
2025 RAST: Reasoning Activation in LLMs via Small-model Transfer
abstract
Reinforcement learning (RL) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs), as evidenced by recent successes such as OpenAI's o1 and Deepseek-R1. However, applying RL at scale remains intimidatingly resource-intensive, requiring multiple model copies and extensive GPU workloads. On the other hand, while being powerful, recent studies suggest that RL does not fundamentally endow models with new knowledge; rather, it primarily reshapes the model's output distribution to activate reasoning capabilities latent in the base model. Building on this insight, we hypothesize that the changes in output probabilities induced by RL are largely model-size invariant, opening the door to a more efficient paradigm: training a small model with RL and transferring its induced probability shifts to larger base models. To verify our hypothesis, we conduct a token-level analysis of decoding trajectories and find high alignment in RL-induced output distributions across model scales, validating our hypothesis. Motivated by this, we propose RAST, a simple yet effective method that transfers reasoning behaviors by injecting RL-induced probability adjustments from a small RL-trained model into larger models. Experiments across multiple mathematical reasoning benchmarks show that RAST substantially and consistently enhances the reasoning capabilities of base models while requiring significantly lower GPU memory than direct RL training, sometimes even yielding better performance than the RL-trained counterparts. Our findings offer new insights into the nature of RL-driven reasoning and practical strategies for scaling its benefits without incurring its full computational cost. The project page of RAST is available at https://ozyyshr.github.io/RAST/.
Siru Ouyang, Zilin Xiao, Minhao Jiang, Yu Meng 0001, Jiawei Han 0001
NeurIPS1
2025 Multimodal Search in Chemical Documents and Reactions
abstract
We present a multimodal search tool for retrieval of chemical reactions, molecular structures, and associated text from scientific literature.Queries may combine molecular diagrams, textual descriptions, and reaction data, allowing users to connect different chemical information representations.Indexing includes chemical diagram extraction and parsing, extraction of reaction data from text in tabular form, and cross-modal linking of diagrams with their mentions in text.We describe the system's architecture and retrieval features, along with expert assessments of the system.Our demo highlights the workflow and search components.Online demo: https://www.cs.rit.edu/
Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong 0005, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han 0001
SIGIR7
2024 Fact-Driven Logical Reasoning for Machine Reading Comprehension
abstract
Recent years have witnessed an increasing interest in training machines with reasoning ability, which deeply relies on accurately and clearly presented clue forms. The clues are usually modeled as entity-aware knowledge in existing studies. However, those entity-aware clues are primarily focused on commonsense, making them insufficient for tasks that require knowledge of temporary facts or events, particularly in logical reasoning for reading comprehension. To address this challenge, we are motivated to cover both commonsense and temporary knowledge clues hierarchically. Specifically, we propose a general formalism of knowledge units by extracting backbone constituents of the sentence, such as the subject-verb-object formed ``facts''. We then construct a supergraph on top of the fact units, allowing for the benefit of sentence-level (relations among fact groups) and entity-level interactions (concepts or actions inside a fact). Experimental results on logical reasoning benchmarks and dialogue modeling datasets show that our approach improves the baselines substantially, and it is general across backbone models. Code is available at https://github.com/ozyyshr/FocalReasoner.
Siru Ouyang, Zhuosheng Zhang 0001, Hai Zhao 0001
AAAI1
2024 ActionIE: Action Extraction from Scientific Literature with Programming Languages
abstract
Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, Jiawei Han. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong 0005, Tingfeng Luo, Qirong Ho, Hao Peng 0009, Heng Ji 0001, Jiawei Han 0001
ACL (1)3
2024 Structured Chemistry Reasoning with Large Language Models
abstract
Large Language Models (LLMs) excel in diverse areas, yet struggle with complex scientific reasoning, especially in the field of chemistry. Different from the simple chemistry tasks (e.g., molecule classification) addressed in previous studies, complex chemistry problems require not only vast knowledge and precise calculation, but also compositional reasoning about rich dynamic interactions of different concepts (e.g., temperature changes). Our study shows that even advanced LLMs, like GPT-4, can fail easily in different ways. Interestingly, the errors often stem not from a lack of domain knowledge within the LLMs, but rather from the absence of an effective reasoning *structure* that guides the LLMs to elicit the right knowledge, incorporate the knowledge in step-by-step reasoning, and iteratively refine results for further improved quality. On this basis, we introduce StructChem, a simple yet effective prompting strategy that offers the desired guidance and substantially boosts the LLMs' chemical reasoning capability. Testing across four chemistry areas---quantum chemistry, mechanics, physical chemistry, and kinetics---StructChem substantially enhances GPT-4's performance, with up to 30% peak improvement. Our analysis also underscores the unique difficulties of precise grounded reasoning in science with LLMs, highlighting a need for more research in this area.
Siru Ouyang, Zhuosheng Zhang 0001, Xuan Liu 0009, Yejin Choi 0001, Jiawei Han 0001, Lianhui Qin
ICML1
2024 Ontology Enrichment for Effective Fine-grained Entity Typing
abstract
Fine-grained entity typing (FET) is the task of identifying specific entity types at a fine-grained level for entity mentions based on their contextual information. Conventional methods for FET require extensive human annotation, which is time-consuming and costly given the massive scale of data. Recent studies have been developing weakly supervised or zero-shot approaches. We study the setting of zero-shot FET where only an ontology is provided. However, most existing ontology structures lack rich supporting information and even contain ambiguous relations, making them ineffective in guiding FET. Recently developed language models, though promising in various few-shot and zero-shot NLP tasks, may face challenges in zero-shot FET due to their lack of interaction with task-specific ontology. In this study, we propose øurs, where we (1) enrich each node in the ontology structure with two categories of extra information:instance information for training sample augmentation andtopic information to relate types with contexts, and (2) develop a coarse-to-fine typing algorithm that exploits the enriched information by training an entailment model with contrasting topics and instance-based augmented training samples. Our experiments show that øurs achieves high-quality fine-grained entity typing without human annotation, outperforming existing zero-shot methods by a large margin and rivaling supervised methods. øurs also enjoys strong transferability to unseen and finer-grained types. We will open source this work upon acceptance.
Siru Ouyang, Jiaxin Huang 0001, Pranav Pillai, Yunyi Zhang 0001, Yu Zhang 0044, Jiawei Han 0001
KDD1
2024 Automated Mining of Structured Knowledge from Text in the Era of Large Language Models
abstract
Massive amount of unstructured text data are generated daily, ranging from news articles to scientific papers. How to mine structured knowledge from the text data remains a crucial research question. Recently, large language models (LLMs) have shed light on the text mining field with their superior text understanding and instruction-following ability. There are typically two ways of utilizing LLMs: fine-tune the LLMs with human-annotated training data, which is labor intensive and hard to scale; prompt the LLMs in a zero-shot or few-shot way, which cannot take advantage of the useful information in the massive text data. Therefore, it remains a challenge on automated mining of structured knowledge from massive text data in the era of large language models.
Yunyi Zhang 0001, Ming Zhong 0005, Siru Ouyang, Yizhu Jiao, Sizhe Zhou, Linyi Ding, Jiawei Han 0001
KDD3
2023 Compositional Data Augmentation for Abstractive Conversation Summarization
abstract
Recent abstractive conversation summarization systems generally rely on large-scale datasets with annotated summaries.However, collecting and annotating these conversations can be a time-consuming and labor-intensive task.To address this issue, in this work, we present a sub-structure level compositional data augmentation method, COMPO, for generating diverse and high-quality pairs of conversations and summaries.Specifically, COMPO first extracts conversation structures like topic splits and action triples as basic units.Then we organize these semantically meaningful conversation snippets compositionally to create new training instances.Additionally, we explore noise-tolerant settings in both self-training and joint-training paradigms to make the most of these augmented samples.Our experiments on benchmark datasets, SAMSum and Di-alogSum, show that COMPO substantially outperforms prior baseline methods by achieving a nearly 10% increase of ROUGE scores with limited data.We have publically released our code at https://github.com/ ozyyshr/Compo.
Siru Ouyang, Jiaao Chen, Jiawei Han 0001, Diyi Yang
ACL (1)1
2023 Instruct and Extract: Instruction Tuning for On-Demand Information Extraction
abstract
Large language models with instructionfollowing capabilities open the door to a wider group of users.However, when it comes to information extraction -a classic task in natural language processing -most task-specific systems cannot align well with long-tail ad hoc extraction use cases for non-expert users.To address this, we propose a novel paradigm, termed On-Demand Information Extraction, to fulfill the personalized demands of real-world users.Our task aims to follow the instructions to extract the desired content from the associated text and present it in a structured tabular format.The table headers can either be userspecified or inferred contextually by the model.To facilitate research in this emerging area, we present a benchmark named INSTRUCTIE, inclusive of both automatically generated training data, as well as the human-annotated test set.Building on INSTRUCTIE, we further develop an On-Demand Information Extractor, ODIE.Comprehensive evaluations on our benchmark reveal that ODIE substantially outperforms the existing open-source models of similar size.Our code and dataset are released on https://github.com/yzjiao/On-Demand-IE.
Yizhu Jiao, Ming Zhong 0005, Ruining Zhao, Siru Ouyang, Heng Ji 0001, Jiawei Han 0001
EMNLP5
2023 The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions
abstract
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Siru Ouyang, Shuohang Wang, Ming Zhong 0005, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001
EMNLP1
2022 Two-Hop Relay Deployment Based on User Trajectory in Wireless Networks
abstract
Abstract The traditional relay deployment problem typically assumes that the locations of users are known and stationary, which is not realistic in practice. The prevalence of mobile devices has made it possible to collect user trajectory and account for user movement while deploying relays. Under this background, a novel problem trajectory-based relay deployment (TBRD) is put forward. This problem considers communication-related metrics and is aimed at maximizing user connection time as users roam through the target area under relay resource constraints, which is more reasonable than the goal of expanding the relay coverage. To figure out the TBRD, we first propose the concept demand nodes (DNs), which are virtual weighted nodes representing the locations where users frequently pass or stay for a long period. Next, we design a Demand Node Generation algorithm that can transform the continuous historical user trajectory into a number of discrete DNs. By generating DNs, we convert the TBRD problem into a demand node coverage (DNC) problem, which is proved to be NP-complete. Followed by that, we introduce an approximation algorithm, named Submodular Iterative Deployment Algorithm, which solves the DNC problem with the approximation factor $1-\frac{1}{\sqrt{e\cdot (1-1/k)}}$, where $e$ is the mathematical constant, and $k$ is the relay number constraint. Finally, five real trajectory datasets are used to evaluate our proposed algorithm, and the simulation results demonstrate that our algorithm can obtain high coverage for users in motion, which can lead to better user experience. In addition, we also analyze the impact of different parameters on the coverage performance, and under this circumstance, we may safely come to the conclusion that our work is at the leading edge to utilize user trajectories for relay deployment in wireless networks.
Siru Ouyang, Xiaofeng Gao 0001, Guihai Chen
Comput. J.2
2021 Smoothing Dialogue States for Open Conversational Machine Reading
abstract
Conversational machine reading (CMR) requires machines to communicate with humans through multi-turn interactions between two salient dialogue states of decision making and question generation processes.In open CMR settings, as the more realistic scenario, the retrieved background knowledge would be noisy, which results in severe challenges in the information transmission.Existing studies commonly train independent or pipeline systems for the two subtasks.However, those methods are trivial by using hard-label decisions to activate question generation, which eventually hinders the model performance.In this work, we propose an effective gating strategy by smoothing the two dialogue states in only one decoder and bridge decision making and question generation to provide a richer dialogue state reference.Experiments on the OR-ShARC dataset show the effectiveness of our method, which achieves new state-of-the-art results.
Zhuosheng Zhang 0001, Siru Ouyang, Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita
EMNLP (1)2