Fanyu Wang

dblp:329/1821 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2027
0000-0002-9937-8534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Consistency-driven evidence reconstruction for retrieval-augmented generation
Huihui Shao, Shuaiyu Zhang, Fanyu Wang, Zhenping Xie
Expert Syst. Appl.3
2026 LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues
abstract
Fanyu Wang, Xiaoxi Kang, Paul Burgess, Aashish Srivastava, Chetan Arora, Adnan Trakic, Lay-Ki Soon, Md Khalid Hossain, Lizhen Qu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fanyu Wang, Xiaoxi Kang, Paul Burgess, Aashish Srivastava, Chetan Arora 0002, Adnan Trakic, Lay-Ki Soon, Md Khalid Hossain, Lizhen Qu
ACL (1)1
2026 Integrated Sensing and Communication with Two-Sided Sensing
Fanyu Wang, Yao Liu 0007, Lawrence Ong, Aylin Yener
ISIT1
2026 S2CR: A self-supervised self-consistency reasoning framework coupled to retrieval-augmented generation
abstract
Existing approaches to the self-consistency exploration in Large Language Models (LLMs) primarily rely on post-hoc selection, overlooking the inherent logical structures essential for reasoning. To address this gap, we pioneer a new perspective by introducing S 2 CR , a self-supervised reasoning framework that presents the LLM consistency from internal modeling while coupling retrieval-augmented generation (RAG) for knowledge integration and process supervision. This framework operates in four stages: Information Retrieval patches the parametric knowledge of LLMs and provides factual grounding for evaluation; Response Generation generates multiple candidate responses for input to explore diverse reasoning paths; Consistency Evaluation quantifies logical consistency by aligning extracted triples from both the generated responses and retrieval information; and Duality Synergy Optimization (DSOP) further bolsters the consistency performance through two complementary modules, Introspection-Driven Self-correction Guidance (IDSG) and Fine-Grained Consensus Alignment (FGCA). Experiments conducted on three public datasets POPQA, Biography, and ALCE-ASQA demonstrate that S 2 CR achieves objective quantification of internal self-consistency and significantly improves performance ranging from 3.19% to 23.49% over Baseline ⋄ across diverse foundational LLMs, e.g. , GPT-3.5-turbo, GPT-4o, and open-source models LLaMA3-8B and LLaMA3-70B.
Huihui Shao, Fanyu Wang, Shuaiyu Zhang, Zhenping Xie
Inf. Process. Manag.2
2026 DGNet: Dynamic Graph Networks for Multivariate Time Series Prediction
Yongqiang Cheng 0009, Shaohao Tan, Fanyu Wang, Jicen Tan
Mach. Learn.6
2025 Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs
abstract
Acceptance criteria (ACs) play a critical role in software development by clearly defining the conditions under which a software feature satisfies stakeholder expectations. However, manually creating accurate, comprehensive, and unambiguous acceptance criteria is challenging, particularly in user interface-intensive applications, due to the reliance on domain-specific knowledge and visual context that is not always captured by textual requirements alone. To address these challenges, we propose RAGcceptance_M2RE, a novel approach that leverages Retrieval-Augmented Generation (RAG) to generate acceptance criteria from multi-modal requirements data, including both textual documentation and visual UI information. We systematically evaluated our approach in an industrial case study involving an education-focused software system used by approximately 100,000 users. The results indicate that integrating multi-modal information significantly enhances the relevance, correctness, and comprehensibility of the generated ACs. Moreover, practitioner evaluations confirm that our approach effectively reduces manual effort, captures nuanced stakeholder intent, and provides valuable criteria that domain experts may overlook, demonstrating practical utility and significant potential for industry adoption. This research underscores the potential of multi-modal RAG techniques in streamlining software validation processes and improving development efficiency. We also make our implementation and a dataset available.
Fanyu Wang, Chetan Arora 0002, Yonghui Liu 0001, Kaicheng Huang, Chakkrit Tantithamthavorn, Aldeida Aleti, Dishan Sambathkumar, David Lo 0001
ASE1
2025 From Domain Documents to Requirements: Retrieval-Augmented Generation in the Space Industry
abstract
Requirements engineering (RE) in the space industry is inherently complex, demanding high precision, alignment with rigorous standards, and adaptability to mission-specific constraints. Smaller space organisations and new entrants often struggle to derive actionable requirements from extensive, unstructured documents such as mission briefs, interface specifications, and regulatory standards. In this innovation opportunity paper, we explore the potential of Retrieval-Augmented Generation (RAG) models to support and (semi-)automate requirements generation in the space domain. We present a modular, AI-driven approach that preprocesses raw space mission documents, classifies them into semantically meaningful categories, retrieves contextually relevant content from domain standards, and synthesises draft requirements using large language models (LLMs). We apply the approach to a real-world mission document from the space domain to demonstrate feasibility and assess early outcomes in collaboration with our industry partner, Starbound Space Solutions. Our preliminary results indicate that the approach can reduce manual effort, improve coverage of relevant requirements, and support lightweight compliance alignment. We outline a roadmap toward broader integration of AI in RE workflows, intending to lower barriers for smaller organisations to participate in large-scale, safety-critical missions.
Chetan Arora 0002, Fanyu Wang, Chakkrit Tantithamthavorn, Aldeida Aleti, Shaun Kenyon
RE2
2025 S2AF: An action framework to self-check the Understanding Self-Consistency of Large Language Models
Huihui Shao, Fanyu Wang, Zhenping Xie
Neural Networks2
2024 Optimizing LLMs for Code Generation: Which Hyperparameter Settings Yield the Best Results?
abstract
Large Language Models (LLMs), such as GPT models, are increasingly used in software engineering for various tasks, such as code generation, requirements management, and debugging. While automating these tasks has garnered significant attention, a systematic study on the impact of varying hyperparameters on code generation outcomes remains unexplored. This study aims to assess LLMs' code generation performance by exhaustively exploring the impact of various hyperparameters. Hyperparameters for LLMs are adjustable settings that affect the model's behaviour and performance. Specifically, we investigated how changes to the hyperparameters-temperature, top probability (top_p), frequency penalty, and presence penalty-affect code generation outcomes. We systematically adjusted all hyperparameters together, exploring every possible combination by making small increments to each hyperparameter at a time. This exhaustive approach was applied to 13 Python code generation tasks, yielding one of four outcomes for each hyperparameter combination: no output from the LLM, non-executable code, code that fails unit tests, or correct and functional code. We analysed these outcomes for a total of 14,742 generated Python code segments, focusing on correctness, to determine how the hyperparameters influence the LLM to arrive at each outcome. Using correlation coefficient and regression tree analyses, we ascertained which hyperparameters influence which aspect of the LLM. Our results indicate that optimal performance is achieved with a temperature below 0.5, top probability below 0.75, frequency penalty above -1 and below 1.5, and presence penalty above -1. We make our dataset and results available to facilitate replication.
Chetan Arora 0002, Ahnaf Ibn Sayeed, Sherlock A. Licorish, Fanyu Wang, Christoph Treude
APSEC4
2022 An Adversarial Multi-task Learning Method for Chinese Text Correction with Semantic Detection
Fanyu Wang, Zhenping Xie
ICANN (2)1