Fengyu Lu

dblp:321/6725 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Verifiable LLM-Generated Text Detection via Projected Semantic-Structural Distributions
abstract
Ruochong Xiong, Qien Li, Wangwang Lian, Yulong Wan, Hanlin Xue, Zhouxing Tan, Han Yang, Fengyu Lu, Junfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruochong Xiong, Qien Li, Wangwang Lian, Yulong Wan, Hanlin Xue, Zhouxing Tan, Fengyu Lu
ACL (1)8
2025 MedKRG: A Knowledge and Retrieval-Guided Framework for Electronic Health Record Synthesis
abstract
Diagnosis based on electronic health records (EHRs) often struggles with data scarcity and privacy concerns, particularly when dealing with rare diseases. Previous approaches using artificial intelligence tools, such as large language models (LLMs), to generate synthetic EHRs often produce factually hallucinatory content. To address these issues, we propose MedKRG, a knowledge and retrieval-guided data synthesis approach that augments limited EHR samples. Specifically, MedKRG creates EHRs with targeted diseases by first retrieving EHRs similar to a given seed, spanning both common and specific conditions, and then substituting their key clinical attributes under the supervision of a knowledge graph. The core modules of our framework are a Transformer encoder (TE) trained on diverse medical sources and a powerful LLM. Specifically, KeGES uses an LLM to reconstruct a seed EHR from similar EHRs retrieved by TE, revisits the reconstructed content using a medical graph, and finally produces novel EHRs. This loop - retrieval, reconstruction, and revisiting repeats until the synthesized EHRs achieve comprehensive coverage of the knowledge related to the targeted disease. We perform comprehensive evaluations on Jarvis datasets to assess the effectiveness of synthetic data and the security of patient privacy. Experimental results demonstrate improved diagnostic accuracy across generic and medical LLMs, without the leakage of patient information.
Fengyu Lu, Jiaxin Duan
BIBM1
2025 PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction
Jiaxin Duan, Fengyu Lu
CogSci2
2025 Efficient Multi-dimensional Optimization in Abstractive Summarization via Mixture-of-Learnable Prompts Tuning
Fengyu Lu, Jiaxin Duan
CogSci1
2025 Learning Markup Language Model for Composite Relationships Extraction
abstract
Generative approaches to relational triplet extraction (RTE) developed on the sequence model have effectively handled monotonous relations. However, the nuanced complexities of composite relationships among entities significantly challenge their performance. In this paper, we adapt pre-trained language models (PLMs) to extract complex relations in RTE, for which we propose a novel learning framework, named MarkET. Specifically, MarkET uses markup language with a hierarchical syntax structure to concisely encode diverse relationships, which allows a PLM to seamlessly integrate the extraction of relational triplets with its advanced language-generating capabilities. To further ensure the validity of model outputs, MarkET trains the PLM through a two-stage curriculum, consisting of the bootstrap tasks of entity and relation recognition and the primary task RTE, each with incrementally increased difficulty along the training steps. We conduct extensive experiments on four public datasets, and our approach is approved superior to previous SOTA across diverse settings and metrics.
Fengyu Lu, Jiaxin Duan
ICASSP1
2024 Alleviating Exposure Bias in Abstractive Summarization via Sequentially Generating and Revising
abstract
Abstractive summarization commonly suffers from exposure bias caused by supervised teacher-force learning, that a model predicts the next token conditioned on the accurate pre-context during training while on its preceding outputs at inference. Existing solutions bridge this gap through un- or semi-supervised holistic learning yet still leave the risk of error accumulation while generating a summary. In this paper, we attribute this problem to the limitation of unidirectional autoregressive text generation and introduce post-processing steps to alleviate it. Specifically, we reformat abstractive summarization to sequential generation and revision (SeGRe), i.e., a model in the revision phase re-inputs the generated summary and refines it by contrasting it with the source document. This provides the model additional opportunities to assess the flawed summary from a global view and thereby modify inappropriate expressions. Moreover, we train the SeGRe model with a regularized minimum-risk policy to ensure effective generation and revision. A lot of comparative experiments are implemented on two well-known datasets, exhibiting the new or matched state-of-the-art performance of SeGRe.
Jiaxin Duan, Fengyu Lu
LREC/COLING2
2024 Prophecy Distillation for Boosting Abstractive Summarization
abstract
Abstractive summarization models learned with maximum likelihood estimation (MLE) have long been guilty of generating unfaithful facts alongside ambiguous focus. Improved paradigm under the guidance of reference-identified words, i.e., guided summarization, has exhibited remarkable advantages in overcoming this problem. However, it suffers limited real applications since the prophetic guidance is practically agnostic at inference. In this paper, we introduce a novel teacher-student framework, which learns a regular summarization model to mimic the behavior of being guided by prophecy for boosting abstractive summaries. Specifically, by training in probability spaces to follow and distinguish a guided teacher model, a student model learns the key to generating teacher-like quality summaries without any guidance. We refer to this process as prophecy distillation, and it breaks the limitations of both standard and guided summarization. Through extensive experiments, we show that our method achieves new or matched state-of-the-art on four well-known datasets, including ROUGE scores, faithfulness, and saliency awareness. Human evaluations are also carried out to evidence these merits. Furthermore, we conduct empirical studies to analyze how the hyperparameters setting and the guidance choice affect TPG performance.
Jiaxin Duan, Fengyu Lu
LREC/COLING2
2024 Alleviating Hallucinations Via Supportive Window Indexing in Abstractive Summarization
abstract
Abstractive summarization models learned with maximum likelihood estimation (MLE) have been proven to produce hallucinatory content, which heavily limits their real-world applicability. Preceding studies attribute this problem to the semantic insensitivity of MLE, and they compensate for it with additional unsupervised learning objectives that maximize the metrics of document-summary inferring, however, resulting in unstable and expensive model training. In this paper, we propose a novel supportive windows indexing grounded summarization (SWIGS) paradigm, where an input document is split into several windows, and a summarization model orderly generates the indices of supportive windows before each summary sentence. Because the supportive windows locate the source information closely related to the summary sentence to be generated, pointing out their indices at first helps to ground the evidence-based summary generation, thus alleviating groundless hallucinations. We create only supervised objectives to learn the SWIGS model and conduct extensive experiments on two well-known datasets to validate its effectiveness. Results vouch for the superiority of SWIGS as it outperforms previous methods regarding multiple metrics.
Jiaxin Duan, Fengyu Lu
ICASSP2
2024 Iterative Autoregressive Generation for Abstractive Summarization
abstract
Abstractive summarization suffers from exposure bias caused by the teacher-forced maximum likelihood estimation (MLE) learning, that an autoregressive language model predicts the next token distribution conditioned on the exact pre-context during training while on its own predictions at inference. Preceding resolutions for this problem straightforwardly augment the pure token-level MLE with summary-level objectives. Although this method effectively exposes a model to its prediction errors during summary-level learning, such errors accumulate in the unidirectional autoregressive generation and further limit the learning efficiency. To address this problem, we imitate the human behavior of revising a manuscript multiple times after writing it and introduce a novel iterative autoregressive summarization (IARSum) paradigm, which iteratively rewrites a generated summary to approximate an errorless version. Concretely, IARSum performs iterative revisions after summarization, where the output of the previous revision is taken as the input for the next, and a minimum-risk training strategy is used to ensure that the original summary is effectively polished in every revision round. We conduct extensive experiments on two widely used datasets and show the new or matched state-of-the-art performance of IARSum.
Jiaxin Duan, Fengyu Lu
ICASSP2
2024 ConFit: Contrastive Fine-Tuning of Text-to-Text Transformer for Relation Classification
Jiaxin Duan, Fengyu Lu
NLPCC (2)2
2024 Adaptive semi-supervised learning from stronger augmentation transformations of discrete text information
Xuemiao Zhang, Zhouxing Tan, Fengyu Lu, Rui Yan 0001
Knowl. Inf. Syst.3
2023 MVP: Optimizing Multi-view Prompts for Medical Dialogue Summarization
abstract
Medical dialogue summarization (MDS) is commonly known as summarizing patients’ electronic health records (EHRs) from doctor-patient dialogues, and the automating of this task can significantly liberate doctors from trivial recordings. Recently, state-of-the-art abstractive summarization systems in open domains are typically adapted from pre-trained language models (PLMs), providing a promising paradigm for MDS to follow. However, such large models with millions of parameters tend to overfit the limited number of training samples available from the medical community. This makes the generated EHRs always contain hallucinatory facts that never appeared in the input dialogue. To address these problems, we propose MVP, a prompt learning method with multi-view prompts to adapt large PLMs for medical dialogue summarization. It learns from the input dialogue: a single input-specific prompt (InP) captures the global dialogue features to ensure faithful results, and several distinct item-specific prompts (ItPs) extract local dialogue information for prompting the generation of each EHR item. We freeze all PLM parameters and only tune the two types of prompts under a multi-level contrastive learning framework, thus effectively avoiding the overfitting problem. Experimental results on two public datasets demonstrate the superior performances of our method.
Jiaxin Duan, Fengyu Lu
BIBM2
2023 Towards Efficient Medical Dialogue Summarization with Compacting-Abstractive Model
abstract
Medical dialogue summarization (MDS) refers to automatically generating electronic health records (EHRs) from doctor-patient dialogues to relieve doctors from the recording burden. The long-term nature of medical dialogue makes it more efficient to handle MDS with a two-stage abstractive summarization model, which first extracts salient content from the source text and then generates an abstract summary based on that. However, this commonly used extractive-abstractive paradigm struggles to identify absolute salient statements in the first stage while discards all information mistakenly considered unimportant, heavily limiting the second-stage summarization performance. In this paper, we introduce a novel two-stage model for MDS with a compact-then-abstract workflow to ensure data efficiency and information integrity. After predicting EHR-related utterances in a medical dialogue, our model adaptively compresses the remaining into a soft context and generates an EHR according to both the prediction and compression results. This way, we loosen the traditional first-stage extraction to a hybrid of extraction and compression, which makes the input context compact and avoids suffering from losing information for producing an EHR. Extensive experiments on two public datasets show that the proposed model significantly outperforms the state-of-the-art counterparts w.r.t. multiple metrics.
Jiaxin Duan, Fengyu Lu
BIBM2
2023 A Factual Aware Two-Stage Model for Medical Dialogue Summarization
abstract
Medical dialogue summarization (MDS) is commonly known as generating electronic health records (EHR) from doctor-patient dialogues to relieve doctors from trivial recordings. Because of their excellent performance on summarization tasks, it is advisable to employ pre-trained language models (PLM) for MDS. However, most of these models are not designed to handle such lengthy dialogues and struggle with domain-specific characteristics. To address this problem, we propose a two-stage summarization model that first constructs compact contexts by selecting salient utterances and then generates EHRs with delexicalization and lexicalization. A REINFORCE algorithm with a multiple reward strategy is employed to connect the two modules to increase faithfulness and adaptively control the length of the extracted context. We implement our model using publicly available PLMs without changing their nature, and we pre-train each module separately before alternately fine-tuning them with the reinforcement objective. Extensive experiments on two public datasets show that our proposed model significantly outperforms state-of-the-art comparison models w.r.t. ROUGE score and terminology matching rate.
Fengyu Lu, Jiaxin Duan
BIBM1
2022 A Gaussian Mixture Model for Dialogue Generation with Dynamic Parameter Sharing Strategy
abstract
Existing dialog models are trained with data in an encoder-decoder framework with the same parameters, ignoring the multinomial distribution nature in the dataset. In fact, model improvement and development commonly requires fine-grained modeling on individual data subsets. However, collecting a labeled fine-grained dialogue dataset often requires expert-level domain knowledge and therefore is difficult to scale in the real world. As we focus on better modeling multinomial data for dialog generation, we study an approach that combines the unsupervised clustering and generative model together with a GMM (Gaussian Mixture Model) based encoder-decoder framework. Specifically, our model samples from the prior and recognition distributions over the latent variables by a Gaussian mixture network and the latent layer with the capability to form multiple clusters. We also introduce knowledge distillation to guide and improve the clustering results. Finally, we use a dynamic parameter sharing strategy conditioned on different labels to train different decoders. Experimental results on a widely used dialogue dataset verify the effectiveness of the proposed method.
Qingqing Zhu, Zhouxing Tan, Jiaxin Duan, Fengyu Lu
ICASSP5