Gexiang Fang

dblp:345/8099 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0008-0967-1333ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning
abstract
We revisit retrieval-augmented generation (RAG) by embedding retrieval control directly into generation.Instead of treating retrieval as an external intervention, we express retrieval decisions within token-level decoding, enabling end-to-end coordination without additional controllers or classifiers.Under the paradigm of Retrieval as Generation, we propose GRIP (Generation-guided Retrieval with Information Planning), a unified framework in which the model regulates retrieval behavior through control-token emission.Central to GRIP is Self-Triggered Information Planning, which allows the model to decide when to retrieve, how to reformulate queries, and when to terminate, all within a single autoregressive trajectory.This design tightly couples retrieval and reasoning and supports dynamic multi-step inference with on-the-fly evidence integration.To supervise these behaviors, we construct a structured training set covering answerable, partially answerable, and multi-hop queries, each aligned with specific token patterns.Experiments on five QA benchmarks show that GRIP surpasses strong RAG baselines and is competitive with GPT-4o while using substantially fewer parameters.
Bo Li 0099, Gexiang Fang, Shikun Zhang, Wei Ye 0004
ACL (1)3
2026 An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
Chengli Xing, Zhengran Zeng, Gexiang Fang, Rui Xie 0003, Wei Ye 0004, Shikun Zhang
ICPC3
2023 Leveraging Conditional Statement to Generate Acceptance Tests Automatically via Traceable Sequence Generation
abstract
In software development, testing is critical in guaranteeing software quality, with test case design at the core of the testing phase. However, generating effective test cases requires deep expertise and significant time and effort. Therefore, much prior research has turned to methods of Natural Language Processing, utilizing generative deep learning methods to automate test case generation. These earlier studies, however, have largely ignored the crucial role of conditional statements within software requirements - a factor we believe is indispensable for generating high-quality test cases. To bridge this gap, we introduce a pioneering approach for automatically deriving antecedents and consequents from requirements, termed as Traceable Sequence Generation (TSG). The TSG generates conditional statements first and then generates corresponding test cases by constructing a Cause-Effect-Graph. To verify the effectiveness of TSG, we constructed a requirement-to-test-case dataset, called Code Test Case Eval (CTCE). The dataset also includes annotated conditional statements for each segment of the requirement text, so we can utilize them to improve test case generation easily. Our experimental results indicate that TSG notably surpasses traditional and NLP-based methods, excelling in conditional statement extraction and generating high-coverage test cases.
Gexiang Fang, Dongdong Du
QRS2