VLDB 2026 Research / reviewers in the wild / expert
Jingfeng Yang 0001
dblp:50/371-1
· DBLP profile ↗
19ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0001-9713-1792ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Large Language Models Function Calling and Interpretability via Guided-Structured TemplatesabstractHy Dang, Tianyi Liu, Zhuofeng Wu, Jingfeng Yang, Haoming Jiang, Tao Yang, Pei Chen, Zhengyang Wang, Helen Wang, Huasheng Li, Bing Yin, Meng Jiang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hy Dang, Zhuofeng Wu 0005, Jingfeng Yang 0001, Haoming Jiang, Helen Wang, Huasheng Li, Meng Jiang 0001 |
EMNLP | 4 |
| 2025 | LongLeader: A Comprehensive Leaderboard for Large Language Models in Long-context ScenariosabstractPei Chen, Hongye Jin, Cheng-Che Lee, Rulin Shao, Jingfeng Yang, Mingyu Zhao, Zhaoyu Zhang, Qin Lu, Kaiwen Men, Ning Xie, Huasheng Li, Bing Yin, Han Li, Lingyun Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hongye Jin, Cheng-Che Lee, Rulin Shao, Jingfeng Yang 0001, Kaiwen Men, Huasheng Li |
NAACL (Long Papers) | 5 |
| 2025 | Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-TrainingabstractYuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao, Qing Ping, Tianyi Liu, Binxuan Huang, Zheng Li, Zhengyang Wang, Pei Chen, Ruijie Wang, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Bing Yin, Chao Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yuchen Zhuang, Jingfeng Yang 0001, Haoming Jiang, Xin Liu 0039, Kewei Cheng, Sanket Lokegaonkar, Yifan Gao 0001, Qing Ping, Binxuan Huang, Zheng Li 0018, Ruijie Wang 0004, Rongzhi Zhang, Nasser Zalmout, Priyanka Nigam, Chao Zhang 0014 |
NAACL (Long Papers) | 2 |
| 2024 | Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive BenchmarkabstractFenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag, Bing Yin, Yining Hua, Xuan Zhou, Omid Rohanian, Anshul Thakur, Lei Clifton, David A. Clifton. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zheng Li 0018, Hongjian Zhou, Qingyu Yin, Jingfeng Yang 0001, Xianfeng Tang, Chen Luo 0003, Ming Zeng 0001, Haoming Jiang, Yifan Gao 0001, Priyanka Nigam, Sreyashi Nag, Yining Hua, Omid Rohanian, Anshul Thakur, Lei A. Clifton, David A. Clifton |
EMNLP | 5 |
| 2024 | LLM Maybe LongLM: SelfExtend LLM Context Window Without TuningabstractIt is well known that LLMs cannot generalize well to long contexts whose lengths are larger than the training sequence length. This poses challenges when employing LLMs for processing long input sequences during inference. In this work, we argue that LLMs themselves have inherent capabilities to handles s long contexts without fine-tuning. To achieve this goal, we propose SelfExtend to extend the context window of LLMs by constructing bi-level attention information: the grouped attention and the neighbor attention. The grouped attention captures the dependencies among tokens that are far apart, while neighbor attention captures dependencies among adjacent tokens within a specified range. The two-level attentions are computed based on the original model’s self-attention mechanism during inference. With minor code modification, our SelfExtend can effortlessly extend existing LLMs’ context window without any fine-tuning. We conduct comprehensive experiments on multiple benchmarks and the results show that our SelfExtend can effectively extend existing LLMs’ context window length. Hongye Jin, Jingfeng Yang 0001, Zhimeng Jiang, Zirui Liu 0001, Chia-Yuan Chang 0002, Huiyuan Chen, Xia Ben Hu |
ICML | 3 |
| 2024 | MEMORYLLM: Towards Self-Updatable Large Language ModelsabstractExisting Large Language Models (LLMs) usually remain static after deployment, which might make it hard to inject new knowledge into the model. We aim to build models containing a considerable portion of self-updatable parameters, enabling the model to integrate new knowledge effectively and efficiently. To this end, we introduce MEMORYLLM, a model that comprises a transformer and a fixed-size memory pool within the latent space of the transformer. MEMORYLLM can self-update with text knowledge and memorize the knowledge injected earlier. Our evaluations demonstrate the ability of MEMORYLLM to effectively incorporate new knowledge, as evidenced by its performance on model editing benchmarks. Meanwhile, the model exhibits long-term information retention capacity, which is validated through our custom-designed evaluations and long-context benchmarks. MEMORYLLM also shows operational integrity without any sign of performance degradation even after nearly a million memory updates. Our code and model are open-sourced at https://github.com/wangyu-ustc/MemoryLLM. Yu Wang 0170, Yifan Gao 0001, Xiusi Chen, Haoming Jiang, Jingfeng Yang 0001, Qingyu Yin, Zheng Li 0018, Jingbo Shang, Julian J. McAuley |
ICML | 6 |
| 2024 | Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language ModelsabstractOnline shopping is a complex multi-task, few-shot learning problem with a wide and evolving range of entities, relations, and tasks. However, existing models and benchmarks are commonly tailored to specific tasks, falling short of capturing the full complexity of online shopping. Large Language Models (LLMs), with their multi-task and few-shot learning abilities, have the potential to profoundly transform online shopping by alleviating task-specific engineering efforts and by providing users with interactive conversations. Despite the potential, LLMs face unique challenges in online shopping, such as domain-specific concepts, implicit knowledge, and heterogeneous user behaviors. Motivated by the potential and challenges, we propose Shopping MMLU, a diverse multi-task online shopping benchmark derived from real-world Amazon data. Shopping MMLU consists of 57 tasks covering 4 major shopping skills: concept understanding, knowledge reasoning, user behavior alignment, and multi-linguality, and can thus comprehensively evaluate the abilities of LLMs as general shop assistants. With Shoppping MMLU, we benchmark over 20 existing LLMs and uncover valuable insights about practices and prospects of building versatile LLM-based shop assistants. Shopping MMLU can be publicly accessed at https://github.com/KL4805/ShoppingMMLU. In addition, with Shopping MMLU, we are hosting a competition in KDD Cup 2024 with over 500 participating teams. The winning solutions and the associated workshop can be accessed at our website https://amazon-kddcup24.github.io/. Yilun Jin, Zheng Li 0018, Tianyu Cao 0001, Yifan Gao 0001, Pratik Jayarao, Xin Liu 0039, Ritesh Sarkhel, Xianfeng Tang, Wenju Xu, Jingfeng Yang 0001, Qingyu Yin, Priyanka Nigam, Yi Xu 0011, Kai Chen 0005, Qiang Yang 0001, Meng Jiang 0001 |
NeurIPS | 14 |
| 2024 | Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and BeyondabstractThis article presents a comprehensive and practical guide for practitioners and end-users working with Large Language Models (LLMs) in their downstream Natural Language Processing (NLP) tasks. We provide discussions and insights into the usage of LLMs from the perspectives of models, data, and downstream tasks. First, we offer an introduction and brief summary of current language models. Then, we discuss the influence of pre-training data, training data, and test data. Most importantly, we provide a detailed discussion about the use and non-use cases of large language models for various natural language processing tasks, such as knowledge-intensive tasks, traditional natural language understanding tasks, generation tasks, emergent abilities, and considerations for specific tasks. We present various use cases and non-use cases to illustrate the practical applications and limitations of LLMs in real-world scenarios. We also try to understand the importance of data and the specific challenges associated with each NLP task. Furthermore, we explore the impact of spurious biases on LLMs and delve into other essential considerations, such as efficiency, cost, and latency, to ensure a comprehensive understanding of deploying LLMs in practice. This comprehensive guide aims to provide researchers and practitioners with valuable insights and best practices for working with LLMs, thereby enabling the successful implementation of these models in a wide range of NLP tasks. A curated list of practical guide resources of LLMs, regularly updated, can be found at https://github.com/Mooler0410/LLMsPracticalGuide . An LLMs evolutionary tree, editable yet regularly updated, can be found at llmtree.ai . Jingfeng Yang 0001, Hongye Jin, Ruixiang Tang, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Xia Ben Hu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Multi-VALUE: A Framework for Cross-Dialectal English NLPabstractDialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/. Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang |
ACL (1) | 3 |
| 2023 | Choice Fusion As Knowledge For Zero-Shot Dialogue State TrackingabstractWith the demanding need for deploying dialogue systems in new domains with less cost, zero-shot dialogue state tracking (DST), which tracks user’s requirements in task-oriented dialogues without training on desired domains, draws attention increasingly. Although prior works have leveraged question-answering (QA) data to reduce the need for in-domain training in DST, they fail to explicitly model knowledge transfer and fusion for tracking dialogue states. To address this issue, we propose CoFunDST, which is trained on domain-agnostic QA datasets and directly uses candidate choices of slot-values as knowledge for zero-shot dialogue-state generation, based on a T5 pre-trained language model. Specifically, CoFunDST selects highly-relevant choices to the reference context and fuses them to initialize the decoder to constrain the model outputs. Our experimental results show that our proposed model achieves outperformed joint goal accuracy compared to existing zero-shot DST approaches in most domains on the MultiWOZ 2.1. Extensive analyses demonstrate the effectiveness of our proposed approach for improving zero-shot DST learning from QA. Ruolin Su, Jingfeng Yang 0001, Ting-Wei Wu, Biing-Hwang Juang |
ICASSP | 2 |
| 2023 | On the Vulnerabilities of Text-to-SQL ModelsabstractAlthough it has been demonstrated that Natural Language Processing (NLP) algorithms are vulnerable to deliberate attacks, the question of whether such weaknesses can lead to software security threats is under-explored. To bridge this gap, we conducted vulnerability tests on Text-to-SQL systems that are commonly used to create natural language interfaces to databases. We showed that the Text-to-SQL modules within six commercial applications can be manipulated to produce malicious code, potentially leading to data breaches and Denial of Service attacks.1This is the first demonstration that NLP models can be exploited as attack vectors in the wild. In addition, experiments using four open-source language models verified that straightforward backdoor attacks on Text-to-SQL systems achieve a 100% success rate without affecting their performance. The aim of this work is to draw the community’s attention to potential software security issues associated with NLP algorithms and encourage exploration of methods to mitigate against them. Xutan Peng, Jingfeng Yang 0001, Mark Stevenson 0001 |
ISSRE | 3 |
| 2023 | Enhancing User Intent Capture in Session-Based Recommendation with Attribute PatternsabstractThe goal of session-based recommendation in E-commerce is to predict the next item that an anonymous user will purchase based on the browsing and purchase history. However, constructing global or local transition graphs to supplement session data can lead to noisy correlations and user intent vanishing. In this work, we propose the Frequent Attribute Pattern Augmented Transformer (FAPAT) that characterizes user intents by building attribute transition graphs and matching attribute patterns. Specifically, the frequent and compact attribute patterns are served as memory to augment session representations, followed by a gate and a transformer block to fuse the whole session information. Through extensive experiments on two public benchmarks and 100 million industrial data in three domains, we demonstrate that FAPAT consistently outperforms state-of-the-art methods by an average of 4.5% across various evaluation metrics (Hits, NDCG, MRR). Besides evaluating the next-item prediction, we estimate the models' capabilities to capture user intents via predicting items' attributes and period-item recommendations. Xin Liu 0039, Zheng Li 0018, Yifan Gao 0001, Jingfeng Yang 0001, Tianyu Cao 0001, Yangqiu Song |
NeurIPS | 4 |
| 2023 | Mutually-paced Knowledge Distillation for Cross-lingual Temporal Knowledge Graph ReasoningabstractThis paper investigates cross-lingual temporal knowledge graph reasoning problem, which aims to facilitate reasoning on Temporal Knowledge Graphs (TKGs) in low-resource languages by transfering knowledge from TKGs in high-resource ones. The cross-lingual distillation ability across TKGs becomes increasingly crucial, in light of the unsatisfying performance of existing reasoning methods on those severely incomplete TKGs, especially in low-resource languages. However, it poses tremendous challenges in two aspects. First, the cross-lingual alignments, which serve as bridges for knowledge transfer, are usually too scarce to transfer sufficient knowledge between two TKGs. Second, temporal knowledge discrepancy of the aligned entities, especially when alignments are unreliable, can mislead the knowledge distillation process. We correspondingly propose a mutually-paced knowledge distillation model MP-KD, where a teacher network trained on a source TKG can guide the training of a student network on target TKGs with an alignment module. Concretely, to deal with the scarcity issue, MP-KD generates pseudo alignments between TKGs based on the temporal information extracted by our representation module. To maximize the efficacy of knowledge transfer and control the noise caused by the temporal knowledge discrepancy, we enhance MP-KD with a temporal cross-lingual attention mechanism to dynamically estimate the alignment strength. The two procedures are mutually paced along with model training. Extensive experiments on twelve cross-lingual TKG transfer tasks in the EventKG benchmark demonstrate the effectiveness of the proposed MP-KD method. Ruijie Wang 0004, Zheng Li 0018, Jingfeng Yang 0001, Tianyu Cao 0001, Chao Zhang 0014, Tarek F. Abdelzaher |
WWW | 3 |
| 2022 | TableFormer: Robust Transformer Modeling for Table-Text EncodingabstractUnderstanding tables is an important aspect of natural language understanding.Existing models for table understanding require linearization of the table structure, where row or column order is encoded as an unwanted bias.Such spurious biases make the model vulnerable to row and column order perturbations.Additionally, prior work has not thoroughly modeled the table structures or table-text alignments, hindering the table-text understanding ability.In this work, we propose a robust and structurally aware table-text encoding architecture TABLEFORMER, where tabular structural biases are incorporated completely through learnable attention biases.TABLEFORMER is (1) strictly invariant to row and column orders, and, (2) could understand tables better due to its tabular inductive biases.Our evaluations showed that TABLEFORMER outperforms strong baselines in all settings on SQA, WTQ and TABFACT table reasoning datasets, and achieves state-of-the-art performance on SQA, especially when facing answer-invariant row and column order perturbations (6% improvement over the best baseline), because previous SOTA models' performance drops by 4% -6% when facing such perturbations while TABLEFORMER is not affected.1 Jingfeng Yang 0001, Aditya Gupta 0001, Shyam Upadhyay, Luheng He, Rahul Goel, Shachi Paul |
ACL (1) | 1 |
| 2022 | SUBS: Subtree Substitution for Compositional Semantic ParsingabstractAlthough sequence-to-sequence models often achieve good performance in semantic parsing for i.i.d.data, their performance is still inferior in compositional generalization.Several data augmentation methods have been proposed to alleviate this problem.However, prior work only leveraged superficial grammar or rules for data augmentation, which resulted in limited improvement.We propose to use subtree substitution for compositional data augmentation, where we consider subtrees with similar semantic functions as exchangeable.Our experiments showed that such augmented data led to significantly better performance on SCAN and GEOQUERY, and reached new SOTA on compositional split of GEOQUERY. Jingfeng Yang 0001, Le Zhang 0012, Diyi Yang |
NAACL-HLT | 1 |
| 2021 | Frustratingly Simple but Surprisingly Strong: Using Language-Independent Features for Zero-shot Cross-lingual Semantic ParsingabstractThe availability of corpora has led to significant advances in training semantic parsers in English.Unfortunately, for languages other than English, annotated data is limited and so is the performance of the developed parsers.Recently, pretrained multilingual models have been proven useful for zero-shot cross-lingual transfer in many NLP tasks.What else does it require to apply a parser trained in English to other languages for zero-shot cross-lingual semantic parsing?Will simple language-independent features help?To this end, we experiment with six Discourse Representation Structure (DRS) semantic parsers in English, and generalize them to Italian, German and Dutch, where there are only a small number of manually annotated parses available.Extensive experiments show that despite its simplicity, adding Universal Dependency (UD) relations and Universal POS tags (UPOS) as model-agnostic features achieves surprisingly strong improvement on all parsers.We have publicly released our code at https://github.com/GT-SALT/ Jingfeng Yang 0001, Federico Fancellu, Bonnie L. Webber, Diyi Yang |
EMNLP (1) | 1 |
| 2020 | Planning and Generating Natural and Diverse Disfluent Texts as Augmentation for Disfluency DetectionabstractExisting approaches to disfluency detection heavily depend on human-annotated data.Numbers of data augmentation methods have been proposed to alleviate the dependence on labeled data.However, current augmentation approaches such as random insertion or repetition fail to resemble training corpus well and usually resulted in unnatural and limited types of disfluencies.In this work, we propose a simple Planner-Generator based disfluency generation model to generate natural and diverse disfluent texts as augmented data, where the Planner decides on where to insert disfluent segments and the Generator follows the prediction to generate corresponding disfluent segments.We further utilize this augmented data for pretraining and leverage it for the task of disfluency detection.Experiments demonstrated that our two-stage disfluency generation model outperforms existing baselines; those disfluent sentences generated significantly aided the task of disfluency detection and led to state-of-the-art performance on Switchboard corpus.We have publicly released our code at https://github.com/GT-SALT/ Disfluency-Generation-and-Detection. Jingfeng Yang 0001, Diyi Yang, Zhaoran Ma |
EMNLP (1) | 1 |
| 2018 | Toward Fast and Accurate Neural Discourse SegmentationabstractDiscourse segmentation, which segments texts into Elementary Discourse Units, is a fundamental step in discourse analysis.Previous discourse segmenters rely on complicated hand-crafted features and are not practical in actual use.In this paper, we propose an endto-end neural segmenter based on BiLSTM-CRF framework.To improve its accuracy, we address the problem of data insufficiency by transferring a word representation model that is trained on a large corpus.We also propose a restricted self-attention mechanism in order to capture useful information within a neighborhood.Experiments on the RST-DT corpus show that our model is significantly faster than previous methods, while achieving new stateof-the-art performance.1 Yizhong Wang, Sujian Li, Jingfeng Yang 0001 |
EMNLP | 3 |
| 2017 | Tag-Enhanced Tree-Structured Neural Networks for Implicit Discourse Relation ClassificationabstractIdentifying implicit discourse relations between text spans is a challenging task because it requires understanding the meaning of the text. To tackle this task, recent studies have tried several deep learning methods but few of them exploited the syntactic information. In this work, we explore the idea of incorporating syntactic parse tree into neural networks. Specifically, we employ the Tree-LSTM model and Tree-GRU model, which is based on the tree structure, to encode the arguments in a relation. And we further leverage the constituent tags to control the semantic composition process in these tree-structured neural networks. Experimental results show that our method achieves state-of-the-art performance on PDTB corpus. Yizhong Wang, Sujian Li, Jingfeng Yang 0001, Xu Sun 0001, Houfeng Wang |
IJCNLP(1) | 3 |