Siyuan Wang 0025

dblp:12/9626-25 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0001-7357-2913ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
abstract
Zheng Jia, Shengbin Yue, Wei Chen, Siyuan Wang, Yidong Liu, Zejun Li, Yun Song, Zhongyu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zheng Jia, Shengbin Yue, Wei Chen 0088, Siyuan Wang 0025, Yidong Liu, Yun Song, Zhongyu Wei
ACL (1)4
2025 Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
abstract
Recent advancements in Large Language Models (LLMs) have led to significant breakthroughs in various natural language processing tasks. However, generating factually consistent responses in knowledge-intensive scenarios remains a challenge due to issues such as hallucination, difficulty in acquiring long-tailed knowledge, and limited memory expansion. This paper introduces SMART, a novel multi-agent framework that leverages external knowledge to enhance the interpretability and factual consistency of LLM-generated responses. SMART comprises four specialized agents, each performing a specific sub-trajectory action to navigate complex knowledge-intensive tasks. We propose a multi-agent co-training paradigm, Long-Short Trajectory Learning, which ensures synergistic collaboration among agents while maintaining fine-grained execution by each agent. Extensive experiments on five knowledge-intensive tasks demonstrate SMART's superior performance compared to widely adopted knowledge internalization and knowledge enhancement methods. Our framework can extend beyond knowledge-intensive tasks to more complex scenarios.
Shengbin Yue, Siyuan Wang 0025, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei
AAAI2
2025 HAF-RM: A Hybrid Alignment Framework for Reward Model Training
abstract
Shujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang, Shengbin Yue, Zengfeng Huang, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shujun Liu, Yuhang Lai, Siyuan Wang 0025, Shengbin Yue, Zengfeng Huang, Xuanjing Huang 0001, Zhongyu Wei
ACL (1)4
2025 Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
abstract
Siyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li, Zhihao Fan, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Siyuan Wang 0025, Dianyi Wang, Chengxing Zhou, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei
ACL (1)1
2025 AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
abstract
Artificial intelligence has significantly revolutionized healthcare, particularly through large language models (LLMs) that demonstrate superior performance in static medical question answering benchmarks. However, evaluating the potential of LLMs for real-world clinical applications remains challenging due to the intricate nature of doctor-patient interactions. To address this, we introduce AI Hospital, a multi-agent framework emulating dynamic medical interactions between Doctor as player and NPCs including Patient and Examiner. This setup allows for more practical assessments of LLMs in simulated clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and multiple evaluation strategies to quantify the performance of LLM-driven Doctor agents on symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance medical interaction capabilities through iterative discussions. Despite improvements, current LLMs (including GPT-4) still exhibit significant performance gaps in multi-turn interactive scenarios compared to non-interactive scenarios. Our findings highlight the need for further research to bridge these gaps and improve LLMs’ clinical decision-making capabilities. Our data, code, and experimental results are all open-sourced at https://github.com/LibertFan/AI_Hospital.
Zhihao Fan, Jialong Tang, Wei Chen 0088, Siyuan Wang 0025, Zhongyu Wei, Fei Huang 0002
COLING5
2025 Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
abstract
This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs). We utilize a multi-agent system to reframe new evolving instances with high confidence that extend existing benchmarks. Towards a more scalable, robust and fine-grained evaluation, we implement six reframing operations to construct evolving instances testing LLMs against diverse queries, shortcut biases and probing their problem-solving sub-abilities. With this framework, we extend datasets across general and specific tasks, through various iterations. Experimental results show a performance decline in most LLMs against their original results under scalable and robust evaluations, offering a more accurate reflection of model capabilities alongside our fine-grained evaluation. Besides, our framework widens performance discrepancies both between different models and within the same model across various tasks, facilitating more informed model selection for specific tasks. We hope this framework contributes the research community for continuously evolving benchmarks alongside LLM development.
Siyuan Wang 0025, Zhuohan Long, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei
COLING1
2024 Multi-Objective Forward Reasoning and Multi-Reward Backward Refinement for Product Review Summarization
abstract
Product review summarization aims to generate a concise summary based on product reviews to facilitate purchasing decisions. This intricate task gives rise to three challenges in existing work: factual accuracy, aspect comprehensiveness, and content relevance. In this paper, we first propose an FB-Thinker framework to improve the summarization ability of LLMs with multi-objective forward reasoning and multi-reward backward refinement. To enable LLM with these dual capabilities, we present two Chinese product review summarization datasets, Product-CSum and Product-CSum-Cross, for both instruction-tuning and cross-domain evaluation. Specifically, these datasets are collected via GPT-assisted manual annotations from an online forum and public datasets. We further design an evaluation mechanism Product-Eval, integrating both automatic and human evaluation across multiple dimensions for product summarization. Experimental results show the competitiveness and generalizability of our proposed framework in the product review summarization tasks.
Siyuan Wang 0025, Ruofei Lai, Xinyu Zhang 0018, Xuanjing Huang 0001, Zhongyu Wei
LREC/COLING2
2024 LawLLM: Intelligent Legal System with Legal Reasoning and Verifiable Retrieval
Shengbin Yue, Shujun Liu, Chenchen Shen, Siyuan Wang 0025, Yun Song, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei
DASFAA (5)5
2024 Unifying Structure Reasoning and Language Pre-Training for Complex Reasoning Tasks
abstract
Recent pre-trained language models (PLMs) equipped with foundation reasoning skills have shown remarkable performance on downstream complex tasks. However, the significant structure reasoning skill has been rarely studied, which involves modeling implicit structure information within the text and performing explicit logical reasoning over them to deduce the conclusion. This paper proposes a unified learning framework that combines explicit structure reasoning and language pre-training to endow PLMs with the structure reasoning skill. It first identifies several elementary structures within contexts to construct structured queries and performs step-by-step reasoning along the queries to identify the answer entity. The fusion of textual semantics and structure reasoning is achieved by using contextual representations learned by PLMs to initialize the representation space of structures, and performing stepwise reasoning on this semantic representation space. Experimental results on four datasets demonstrate that the proposed model achieves significant improvements in complex reasoning tasks involving diverse structures, and shows transferability to downstream tasks with limited training data and effectiveness for complex reasoning of KGs modality.
Siyuan Wang 0025, Zhongyu Wei, Jiarong Xu, Taishan Li, Zhihao Fan
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Visual Explanation for Open-Domain Question Answering With BERT
abstract
Open-domain question answering (OpenQA) is an essential but challenging task in natural language processing that aims to answer questions in natural language formats on the basis of large-scale unstructured passages. Recent research has taken the performance of benchmark datasets to new heights, especially when these datasets are combined with techniques for machine reading comprehension based on Transformer models. However, as identified through our ongoing collaboration with domain experts and our review of literature, three key challenges limit their further improvement: (i) complex data with multiple long texts, (ii) complex model architecture with multiple modules, and (iii) semantically complex decision process. In this paper, we present VEQA, a visual analytics system that helps experts understand the decision reasons of OpenQA and provides insights into model improvement. The system summarizes the data flow within and between modules in the OpenQA model as the decision process takes place at the summary, instance and candidate levels. Specifically, it guides users through a summary visualization of dataset and module response to explore individual instances with a ranking visualization that incorporates context. Furthermore, VEQA supports fine-grained exploration of the decision flow within a single module through a comparative tree visualization. We demonstrate the effectiveness of VEQA in promoting interpretability and providing insights into model enhancement through a case study and expert evaluation.
Zekai Shao 0001, Shuran Sun, Yuheng Zhao, Siyuan Wang 0025, Zhongyu Wei, Tao Gui, Cagatay Turkay, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.4
2023 Query Structure Modeling for Inductive Logical Reasoning Over Knowledge Graphs
abstract
Siyuan Wang, Zhongyu Wei, Meng Han, Zhihao Fan, Haijun Shan, Qi Zhang, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Haijun Shan, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)1
2022 A Structure-Aware Argument Encoder for Literature Discourse Analysis
abstract
Existing research for argument representation learning mainly treats tokens in the sentence equally and ignores the implied structure information of argumentative context. In this paper, we propose to separate tokens into two groups, namely framing tokens and topic ones, to capture structural information of arguments. In addition, we consider high-level structure by incorporating paragraph-level position information. A novel structure-aware argument encoder is proposed for literature discourse analysis. Experimental results on both a self-constructed corpus and a public corpus show the effectiveness of our model. Resources are available at https://github.com/lemuria-wchen/SAE.
Yinzi Li, Wei Chen 0088, Zhongyu Wei, Yujun Huang, Chujun Wang, Siyuan Wang 0025, Qi Zhang 0001, Xuanjing Huang 0001, Libo Wu
COLING6
2022 Locate Then Ask: Interpretable Stepwise Reasoning for Multi-hop Question Answering
abstract
Multi-hop reasoning requires aggregating multiple documents to answer a complex question. Existing methods usually decompose the multi-hop question into simpler single-hop questions to solve the problem for illustrating the explainable reasoning process. However, they ignore grounding on the supporting facts of each reasoning step, which tends to generate inaccurate decompositions. In this paper, we propose an interpretable stepwise reasoning framework to incorporate both single-hop supporting sentence identification and single-hop question generation at each intermediate step, and utilize the inference of the current hop for the next until reasoning out the final result. We employ a unified reader model for both intermediate hop reasoning and final hop inference and adopt joint optimization for more accurate and robust multi-hop reasoning. We conduct experiments on two benchmark datasets HotpotQA and 2WikiMultiHopQA. The results show that our method can effectively boost performance and also yields a better interpretable reasoning process without decomposition supervision.
Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Qi Zhang 0001, Xuanjing Huang 0001
COLING1
2022 Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text Retrieval
abstract
Existing research for image text retrieval mainly relies on sentence-level supervision to distinguish matched and mismatched sentences for a query image. However, semantic mismatch between an image and sentences usually happens in finer grain, i.e., phrase level. In this paper, we explore to introduce additional phrase-level supervision for the better identification of mismatched units in the text. In practice, multi-grained semantic labels are automatically constructed for a query image in both sentence-level and phrase-level. We construct text scene graphs for the matched sentences and extract entities and triples as the phrase-level labels. In order to integrate both supervision of sentence-level and phrase-level, we propose Semantic Structure Aware Multimodal Transformer (SSAMT) for multi-modal representation learning. Inside the SSAMT, we utilize different kinds of attention mechanisms to enforce interactions of multi-grained semantic units in both sides of vision and language. For the training, we propose multi-scale matching from both global and local perspectives, and penalize mismatched phrases. Experimental results on MS-COCO and Flickr30K show the effectiveness of our approach compared to some state-of-the-art models.
Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001, Jianqing Fan
ICMR4
2022 From LSAT: The Progress and Challenges of Complex Reasoning
abstract
Complex reasoning aims to draw a correct inference based on complex rules. As a hallmark of human intelligence, it involves a degree of explicit reading comprehension, interpretation of logical knowledge and complex rule application. In this paper, we take a step forward in complex reasoning by systematically studying the three challenging and domain-general tasks of the Law School Admission Test (LSAT), including analytical reasoning, logical reasoning and reading comprehension. We propose a hybrid reasoning system to integrate these three tasks and achieve impressive overall performance on the LSAT tests. The experimental results demonstrate that our system endows itself a certain complex reasoning ability, especially the fundamental reading comprehension and challenging logical reasoning capacities. Further analysis also shows the effectiveness of combining the pre-trained models with the task-specific reasoning module, and integrating symbolic knowledge into discrete interpretable reasoning steps in complex reasoning. We further shed a light on the potential future directions, like unsupervised symbolic knowledge extraction, model interpretability, few-shot learning and comprehensive benchmark for complex reasoning.
Siyuan Wang 0025, Zhongkun Liu, Wanjun Zhong, Ming Zhou 0001, Zhongyu Wei, Zhumin Chen, Nan Duan 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Fine-Grained Element Identification in Complaint Text of Internet Fraud
abstract
Existing system dealing with online complaint provides a final decision without explanations. We propose to analyse the complaint text of internet fraud in a fine-grained manner. Considering the complaint text includes multiple clauses with various functions, we propose to identify the role of each clause and classify them into different types of fraud element. We construct a large labeled dataset originated from a real finance service platform. We build an element identification model on top of BERT and propose additional two modules to utilize the context of complaint text for better element label classification, namely, global context encoder and label refiner. Experimental results show the effectiveness of our model.
Siyuan Wang 0025, Jingchao Fu, Lei Chen 0082, Zhongyu Wei, Heng Ye, Liaosa Xu, Weiqiang Wang 0002, Xuanjing Huang 0001
CIKM2
2021 TCIC: Theme Concepts Learning Cross Language and Vision for Image Captioning
abstract
Existing research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concepts extended Image Captioning (TCIC) framework that incorporates theme concepts to represent high-level cross-modality semantics. In practice, we model theme concepts as memory vectors and propose Transformer with Theme Nodes (TTN) to incorporate those vectors for image captioning. Considering that theme concepts can be learned from both images and captions, we propose two settings for their representations learning based on TTN. On the vision side, TTN is configured to take both scene graph based features and theme concepts as input for visual representation learning. On the language side, TTN is configured to take both captions and theme concepts as input for text representation re-construction. Both settings aim to generate target captions with the same transformer-based decoder. During the training, we further align representations of theme concepts learned from images and corresponding captions to enforce the cross-modality learning. Experimental results on MS COCO show the effectiveness of our approach compared to some state-of-the-art models.
Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001
IJCAI3
2021 Mask Attention Networks: Rethinking and Strengthen Transformer
abstract
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001
NAACL-HLT5
2020 An Enhanced Knowledge Injection Model for Commonsense Generation
abstract
Commonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics.
Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang
COLING4
2020 PathQG: Neural Question Generation from Facts
abstract
Existing research for question generation encodes the input text as a sequence of tokens without explicitly modeling fact information.These models tend to generate irrelevant and uninformative questions.In this paper, we explore to incorporate facts in the text for question generation in a comprehensive way.We present a novel task of question generation given a query path in the knowledge graph constructed from the input text.We divide the task into two steps, namely, query representation learning and query-based question generation.We formulate query representation learning as a sequence labeling problem for identifying the involved facts to form a query and employ an RNN-based generator for question generation.We first train the two modules jointly in an end-to-end fashion, and further enforce the interaction between these two modules in a variational framework.We construct the experimental datasets on top of SQuAD and results show that our model outperforms other state-of-the-art approaches, and the performance margin is larger when target questions are complex.Human evaluation also proves that our model is able to generate relevant and informative questions. 1
Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Zengfeng Huang, Weijian Sun, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP (1)1
2019 A Multi-Agent Communication Framework for Question-Worthy Phrase Extraction and Question Generation
abstract
Question generation aims to produce questions automatically given a piece of text as input. Existing research follows a sequence-to-sequence fashion that constructs a single question based on the input. Considering each question usually focuses on a specific fragment of the input, especially in the scenario of reading comprehension, it is reasonable to identify the corresponding focus before constructing the question. In this paper, we propose to identify question-worthy phrases first and generate questions with the assistance of these phrases. We introduce a multi-agent communication framework, taking phrase extraction and question generation as two agents, and learn these two tasks simultaneously via message passing mechanism. The results of experiments show the effectiveness of our framework: we can extract question-worthy phrases, which are able to improve the performance of question generation. Besides, our system is able to extract more than one question worthy phrases and generate multiple questions accordingly.
Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Yang Liu 0004, Xuanjing Huang 0001
AAAI1
2019 Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning
abstract
Existing research for visual captioning usually employs a CNN-RNN architecture that combines a CNN for image encoding with a RNN for caption generation, where the vocabulary is constructed from the entire training dataset as the decoding space.Such approaches typically suffer from the problem of generating N-grams which occur frequently in the training set but are irrelevant to the given image.To tackle this problem, we propose to construct an image-grounded vocabulary that leverages image semantics for more effective caption generation.More concretely, a two-step approach is proposed to construct the vocabulary by incorporating both visual information and relationships among words.Two strategies are then explored to utilize the constructed vocabulary for caption generation.One constrains the generator to select words from the image-grounded vocabulary only and the other integrates the vocabulary information into the RNN cell during the caption generation process.Experimental results on two public datasets show the effectiveness of our framework compared to state-of-the-art models.Our code is available on Github 1 .
Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Xuanjing Huang 0001
ACL (1)3
2018 A Reinforcement Learning Framework for Natural Question Generation using Bi-discriminators
abstract
Visual Question Generation (VQG) aims to ask natural questions about an image automatically. Existing research focus on training model to fit the annotated data set that makes it indifferent from other language generation tasks. We argue that natural questions need to have two specific attributes from the perspectives of content and linguistic respectively, namely, natural and human-written. Inspired by the setting of discriminator in adversarial learning, we propose two discriminators, one for each attribute, to enhance the training. We then use the reinforcement learning framework to incorporate scores from the two discriminators as the reward to guide the training of the question generator. Experimental results on a benchmark VQG dataset show the effectiveness and robustness of our model compared to some state-of-the-art models in terms of both automatic and human evaluation metrics.
Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Yang Liu 0004, Xuanjing Huang 0001
COLING3