VLDB 2026 Research / reviewers in the wild / expert
Sufeng Duan
dblp:230/3655
· DBLP profile ↗
13ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-1210-8954ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic ForgettingabstractLarge Vision-Language Models (LVLMs) demonstrate impressive general-purpose capabilities but often suffer from catastrophic forgetting when incorporating specialized knowledge. To address this plasticity-stability dilemma, we introduce Structured Dialogue Fine-Tuning (SDFT), a data-centric approach that injects domain-specific concepts while preserving foundational abilities. Distinct from parameter-constrained continual learning methods, SDFT leverages a three-phase dialogue structure: Foundation Preservation reinforces pre-trained visual-linguistic alignment through captioning tasks; Contrastive Disambiguation uses carefully designed counterfactual examples to establish precise semantic boundaries; and Knowledge Specialization embeds specialized information via chain-of-thought reasoning. Evaluations across personalized entity recognition, abstract concept understanding, and biomedical domains show that SDFT significantly outperforms state-of-the-art baselines, including EWC-LoRA and O-LoRA. The results confirm SDFT’s effectiveness in balancing specialized knowledge acquisition and general capability retention. Yijie Hong, Xiaofei Yin, Xinzhong Wang, Huijia Zhu, Sufeng Duan |
ICMR | 6 |
| 2026 | OpenImplicit: Benchmarking Implicit Reasoning in MLLMs via Open-Ended Evaluation
Jidong Li, Xiaofei Yin, Shuheng Zhou 0001, Haodong Zhao, Sufeng Duan, Gongshen Liu, Huijia Zhu |
ICMR | 7 |
| 2025 | ALIS: Aligned LLM Instruction Security Strategy for Unsafe Input PromptabstractIn large language models, existing instruction tuning methods may fail to balance the performance with robustness against attacks from user input like prompt injection and jailbreaking. Inspired by computer hardware and operating systems, we propose an instruction tuning paradigm named Aligned LLM Instruction Security Strategy (ALIS) to enhance model performance by decomposing user inputs into irreducible atomic instructions and organizing them into instruction streams which will guide the response generation of model. ALIS is a hierarchical structure, in which user inputs and system prompts are treated as user and kernel mode instructions respectively. Based on ALIS, the model can maintain security constraints by ignoring or rejecting the input instructions when user mode instructions attempt to conflict with kernel mode instructions. To build ALIS, we also develop an automatic instruction generation method for training ALIS, and give one instruction decomposition task and respective datasets. Notably, the ALIS framework with a small model to generate instruction streams still improve the resilience of LLM to attacks substantially without any lose on general capabilities. Xinhao Song, Sufeng Duan, Gongshen Liu |
COLING | 2 |
| 2025 | McEval: Massively Multilingual Code EvaluationabstractCode large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval. Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001 |
ICLR | 17 |
| 2025 | Improving semi-autoregressive machine translation with the guidance of syntactic dependency parsing structure
Sufeng Duan, Gongshen Liu |
Neurocomputing | 2 |
| 2024 | Improving Non-autoregressive Machine Translation with Error Exposure and Consistency Regularization
Sufeng Duan, Gongshen Liu |
NLPCC (3) | 2 |
| 2024 | MO-Transformer: Extract High-Level Relationship Between Words for Neural Machine TranslationabstractIn this paper, we propose an explanation of representation for self-attention network (SAN) based neural sequence encoders, which regards the information captured by the model and the encoding of the model as graph structure and the generation of these graph structures respectively. The proposed explanation applies to existing works on SAN-based models and can explain the relationship among the ability to capture the structural or linguistic information, depth of model, and length of sentence, and can also be extended to other models such as recurrent neural network based models. We also propose a revisited multigraph called Multi-order-Graph (MoG) based on our explanation to model the graph structures in the SAN-based model as subgraphs in MoG and convert the encoding of the SAN-based model to the generation of MoG. Based on our explanation, we further introduce an MO-Transformer by enhancing the ability to capture multiple subgraphs of different orders and focusing on subgraphs of high orders. Experimental results on multiple neural machine translation tasks show that the MO-Transformer can yield effective performance improvement. Sufeng Duan, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | SDPSAT: Syntactic Dependency Parsing Structure-Guided Semi-Autoregressive Machine Translation
Yuran Zhao, Jianming Guo, Sufeng Duan, Gongshen Liu |
ICONIP (8) | 4 |
| 2023 | Syntax-Aware Data Augmentation for Neural Machine TranslationabstractData augmentation is an effective method for the performance enhancement of neural machine translation (NMT) by generating additional bilingual data. In this paper, we propose a novel data augmentation strategy for neural machine translation. Unlike existing data augmentation methods that simply modify words with the same probability across different sentences, we introduce a sentence-specific probability approach for word selection based on the syntactic roles of words in the sentence. Our motivation is to consider a linguistics-motivated method to obtain more ingenious language generation rather than relying on computation-motivated approaches only. We argue that high-quality aligned bilingual data is crucial for NMT, and only computation-motivated data augmentation is insufficient to provide good enough extra enhancement data. Our approach leverages dependency parse trees of input sentences to determine the selection probability of each word in the sentence using three different functions to calculate probabilities for words with different depths. Besides, our method also revises the probability for words considering the sentence length. We evaluate our methods on multiple translation tasks. The experimental results demonstrate that our proposed data augmentation method does effectively boost existing sentence-independent methods for significant improvement of performance on translation tasks. Furthermore, an ablation study shows that our method does select fewer essential words and preserves the syntactic structure. Sufeng Duan, Hai Zhao 0001, Dongdong Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Multi-Grained Evidence Inference for Multi-Choice Reading ComprehensionabstractMulti-choice Machine Reading Comprehension (MRC) is a major and challenging task for machines to answer questions according to provided options. Answers in multi-choice MRC cannot be directly extracted in the given passages, and essentially require machines capable of reasoning from accurate extracted evidence. However, the critical evidence may be as simple as just one word or phrase, while it is hidden in the given redundant, noisy passage with multiple linguistic hierarchies from phrase, fragment, sentence until the entire passage. We thus propose a novel general-purpose model enhancement which integrates multi-grained evidence comprehensively, namedMulti-grainedevidence inferencer (Mugen), to make up for the inability.Mugenextracts three different granularities of evidence: coarse-, middle- and fine-grained evidence, and integrates evidence with the original passages, achieving significant and consistent performance improvement on four multi-choice MRC benchmarks. Hai Zhao 0001, Sufeng Duan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | SG-Net: Syntax Guided Transformer for Language RepresentationabstractUnderstanding human language is one of the key themes of artificial intelligence. For language representation, the capacity of effectively modeling the linguistic knowledge from the detail-riddled and lengthy texts and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to all words without explicit constraint, which results in inaccurate concentration on some dispensable words. In this work, we propose using syntax to guide the text modeling by incorporating explicit syntactic constraints into attention mechanisms for better linguistically motivated word representations. In detail, for self-attention network (SAN) sponsored Transformer-based encoder, we introduce syntactic dependency of interest (SDOI) design into the SAN to form an SDOI-SAN with syntax-guided self-attention. Syntax-guided network (SG-Net) is then composed of this extra SDOI-SAN and the SAN from the original Transformer encoder through a dual contextual architecture for better linguistics inspired representation. The proposed SG-Net is applied to typical Transformer encoders. Extensive experiments on popular benchmark tasks, including machine reading comprehension, natural language inference, and neural machine translation show the effectiveness of the proposed SG-Net design. Zhuosheng Zhang 0001, Yuwei Wu 0003, Junru Zhou, Sufeng Duan, Hai Zhao 0001, Rui Wang 0015 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | SG-Net: Syntax-Guided Machine Reading ComprehensionabstractFor machine reading comprehension, the capacity of effectively modeling the linguistic knowledge from the detail-riddled and lengthy passages and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to all words without explicit constraint, which results in inaccurate concentration on some dispensable words. In this work, we propose using syntax to guide the text modeling by incorporating explicit syntactic constraints into attention mechanism for better linguistically motivated word representations. In detail, for self-attention network (SAN) sponsored Transformer-based encoder, we introduce syntactic dependency of interest (SDOI) design into the SAN to form an SDOI-SAN with syntax-guided self-attention. Syntax-guided network (SG-Net) is then composed of this extra SDOI-SAN and the SAN from the original Transformer encoder through a dual contextual architecture for better linguistics inspired representation. To verify its effectiveness, the proposed SG-Net is applied to typical pre-trained language model BERT which is right based on a Transformer encoder. Extensive experiments on popular benchmarks including SQuAD 2.0 and RACE show that the proposed SG-Net design helps achieve substantial performance improvement over strong baselines. Zhuosheng Zhang 0001, Yuwei Wu 0003, Junru Zhou, Sufeng Duan, Hai Zhao 0001, Rui Wang 0015 |
AAAI | 4 |
| 2020 | Attention Is All You Need for Chinese Word SegmentationabstractTaking greedy decoding algorithm as it should be, this work focuses on further strengthening the model itself for Chinese word segmentation (CWS), which results in an even more fast and more accurate CWS model.Our model consists of an attention only stacked encoder and a light enough decoder for the greedy segmentation plus two highway connections for smoother training, in which the encoder is composed of a newly proposed Transformer variant, Gaussian-masked Directional (GD) Transformer, and a biaffine attention scorer.With the effective encoder design, our model only needs to take unigram features for scoring.Our model is evaluated on SIGHAN Bakeoff benchmark datasets.The experimental results show that with the highest segmentation speed, the proposed model achieves new state-of-the-art or comparable performance against strong baselines in terms of strict closed test setting. Sufeng Duan, Hai Zhao 0001 |
EMNLP (1) | 1 |