Jiangtao Feng

dblp:183/0908 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement Learning
abstract
YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lihao Wang, Lei Bai, Lijun Wu, Wei-Ying Ma, Hao Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lei Bai 0001, Lijun Wu 0003, Wei-Ying Ma, Hao Zhou 0012
ACL (1)3
2025 Retro-R1: LLM-based Agentic Retrosynthesis
abstract
Retrosynthetic planning is a fundamental task in chemical discovery. Due to the vast combinatorial search space, identifying viable synthetic routes remains a significant challenge--even for expert chemists. Recent advances in Large Language Models (LLMs), particularly equipped with reinforcement learning, have demonstrated strong human-like reasoning and planning abilities, especially in mathematics and code problem solving. This raises a natural question: Can the reasoning capabilities of LLMs be harnessed to develop an AI chemist capable of learning effective policies for multi-step retrosynthesis? In this study, we introduce Retro-R1, a novel LLM-based retrosynthesis agent trained via reinforcement learning to design molecular synthesis pathways. Unlike prior approaches, which typically rely on single-turn, question-answering formats, Retro-R1 interacts dynamically with plug-in single-step retrosynthesis tools and learns from environmental feedback. Experimental results show that Retro-R1 achieves a 55.79\% pass@1 success rate, surpassing the previous state of the art by 8.95\%. Notably, Retro-R1 demonstrates strong generalization to out-of-domain test cases, where existing methods tend to fail despite their high in-domain performance. Our work marks a significant step toward equipping LLMs with advanced, chemist-like reasoning abilities, highlighting the promise of reinforcement learning for enabling data-efficient, generalizable, and sophisticated scientific problem-solving in LLM-based agents.
Jiangtao Feng, Hongli Yu, Yuxuan Song 0002, Shufei Zhang, Lei Bai 0001, Wei-Ying Ma, Hao Zhou 0012
NeurIPS2
2024 BjTT: A Large-Scale Multimodal Dataset for Traffic Prediction
abstract
Traffic prediction plays a significant role in Intelligent Transportation Systems (ITS). Although many datasets have been introduced to support the study of traffic prediction, most of them only provide time-series traffic data. However, urban transportation systems are always susceptible to various factors, including unusual weather and traffic accidents. Therefore, relying solely on historical data for traffic prediction greatly limits the accuracy of the prediction. In this paper, we introduce Beijing Text-Traffic (BjTT), a large-scale multimodal dataset for traffic prediction. BjTT comprises over 32,000 time-series traffic records, capturing velocity and congestion levels on more than 1,200 roads within the 5th ring area of Beijing. Meanwhile, each piece of traffic data is coupled with a text describing the traffic system (including time, location, and events). We detail the data collection and processing procedures and present a statistical analysis of the BjTT dataset. Furthermore, we conduct comprehensive experiments on the dataset with state-of-the-art traffic prediction methods and text-guided generative models, which reveal the unique characteristics of the BjTT. The dataset is available athttps://github.com/ChyaZhang/BjTT.
Yong Zhang 0029, Qitan Shao, Jiangtao Feng, Bo Li 0128, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.4
2023 DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu 0003, Lingpeng Kong
ICLR3
2023 Compositional Exemplars for In-context Learning
abstract
Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we systematically formulate in-context example selection as a subset selection problem, and optimize it in an end-to-end fashion. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through carefully-designed contrastive learning to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, phraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation and semantic parsing. Extensive experiments demonstrate the effectiveness, transferability, compositionality of CEIL, shedding new lights on in-context leaning. Our code is released at https://github.com/HKUNLP/icl-ceil.
Jiacheng Ye, Zhiyong Wu 0003, Jiangtao Feng, Tao Yu 0009, Lingpeng Kong
ICML3
2023 CAB: Comprehensive Attention Benchmarking on Long Sequence Modeling
abstract
Transformer has achieved remarkable success in language, image, and speech processing. Recently, various efficient attention architectures have been proposed to improve transformer’s efficiency while largely preserving its efficacy, especially in modeling long sequences. A widely-used benchmark to test these efficient methods’ capability on long-range modeling is Long Range Arena (LRA). However, LRA only focuses on the standard bidirectional (or noncausal) self attention, and completely ignores cross attentions and unidirectional (or causal) attentions, which are equally important to downstream applications. In this paper, we propose Comprehensive Attention Benchmark (CAB) under a fine-grained attention taxonomy with four distinguishable attention patterns, namely, noncausal self, causal self, noncausal cross, and causal cross attentions. CAB collects seven real-world tasks from different research areas to evaluate efficient attentions under the four attention patterns. Among these tasks, CAB validates efficient attentions in eight backbone networks to show their generalization across neural architectures. We conduct exhaustive experiments to benchmark the performances of nine widely-used efficient attention architectures designed with different philosophies on CAB. Extensive experimental results also shed light on the fundamental problems of efficient attentions, such as efficiency length against vanilla attention, performance consistency across attention patterns, the benefit of attention mechanisms, and interpolation/extrapolation on long-context language modeling.
Shuyang Jiang, Jiangtao Feng, Lingpeng Kong
ICML3
2023 Attentive Multi-Layer Perceptron for Non-autoregressive Generation
Shuyang Jiang, Jiangtao Feng, Lingpeng Kong
ECML/PKDD (2)3
2022 ZeroGen: Efficient Zero-shot Learning via Dataset Generation
abstract
There is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs).In this paper, we study a flexible and efficient zero-short learning method, ZEROGEN.Given a zero-shot task, we first generate a dataset from scratch using PLMs in an unsupervised manner.Then, we train a tiny task model (e.g., LSTM) under the supervision of the synthesized dataset.This approach allows highly efficient inference as the final task model only has orders of magnitude fewer parameters comparing to PLMs (e.g., GPT2-XL).Apart from being annotation-free and efficient, we argue that ZEROGEN can also provide useful insights from the perspective of datafree model-agnostic knowledge distillation, and unreferenced text generation evaluation.Experiments and analysis on different NLP tasks, namely, text classification, question answering, and natural language inference, show the effectiveness of ZEROGEN.
Jiacheng Ye, Jiahui Gao 0002, Qintong Li, Hang Xu 0004, Jiangtao Feng, Zhiyong Wu 0003, Tao Yu 0009, Lingpeng Kong
EMNLP5
2022 CoNT: Contrastive Neural Text Generation
abstract
Recently, contrastive learning attracts increasing interests in neural text generation as a new solution to alleviate the exposure bias problem. It introduces a sequence-level training signal which is crucial to generation tasks that always rely on auto-regressive decoding. However, previous methods using contrastive learning in neural text generation usually lead to inferior performance. In this paper, we analyse the underlying reasons and propose a new Contrastive Neural Text generation framework, CoNT. CoNT addresses bottlenecks that prevent contrastive learning from being widely adopted in generation tasks from three aspects -- the construction of contrastive examples, the choice of the contrastive loss, and the strategy in decoding. We validate CoNT on five generation tasks with ten benchmarks, including machine translation, summarization, code comment generation, data-to-text generation and commonsense generation. Experimental results show that CoNT clearly outperforms its baseline on all the ten benchmarks with a convincing margin. Especially, CoNT surpasses previous the most competitive contrastive learning method for text generation, by 1.50 BLEU on machine translation and 1.77 ROUGE-1 on summarization, respectively. It achieves new state-of-the-art on summarization, code comment generation (without external data) and data-to-text generation.
Chenxin An, Jiangtao Feng, Kai Lv 0001, Lingpeng Kong, Xipeng Qiu, Xuanjing Huang 0001
NeurIPS2
2021 Learning Logic Rules for Document-Level Relation Extraction
abstract
Document-level relation extraction aims to identify relations between entities in a whole document.Prior efforts to capture long-range dependencies have relied heavily on implicitly powerful representations learned through (graph) neural networks, which makes the model less transparent.To tackle this challenge, in this paper, we propose LogiRE, a novel probabilistic model for document-level relation extraction by learning logic rules.Lo-giRE treats logic rules as latent variables and consists of two modules: a rule generator and a relation extractor.The rule generator is to generate logic rules potentially contributing to final predictions, and the relation extractor outputs final predictions based on the generated logic rules.Those two modules can be efficiently optimized with the expectationmaximization (EM) algorithm.By introducing logic rules into neural networks, LogiRE can explicitly capture long-range dependencies as well as enjoy better interpretation.Empirical results show that LogiRE significantly outperforms several strong baselines in terms of relation performance (∼1.8 F1 score) and logical consistency (over 3.3 logic score).Our code is available at https://github.com/rudongyu/LogiRE.
Dongyu Ru, Changzhi Sun, Jiangtao Feng, Hao Zhou 0012, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005
EMNLP (1)3
2020 Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information
abstract
We investigate the following question for machine translation (MT): can we develop a single universal MT model to serve as the common seed and obtain derivative and improved models on arbitrary language pairs?We propose mRASP, an approach to pre-train a universal multilingual neural machine translation model.Our key idea in mRASP is its novel technique of random aligned substitution, which brings words and phrases with similar meanings across multiple languages closer in the representation space.We pre-train a mRASP model on 32 language pairs jointly with only public datasets.The model is then fine-tuned on downstream language pairs to obtain specialized MT models.We carry out extensive experiments on 42 translation directions across a diverse settings, including low, medium, rich resource, and as well as transferring to exotic language pairs.Experimental results demonstrate that mRASP achieves significant performance improvement compared to directly training on those target pairs.It is the first time to verify that multiple lowresource language pairs can be utilized to improve rich resource MT.Surprisingly, mRASP is even able to improve the translation quality on exotic languages that never occur in the pretraining corpus.Code, data, and pre-trained models are available at https://github. com/linzehui/mRASP.
Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou 0012, Lei Li 0005
EMNLP (1)5
2018 Geometric Relationship between Word and Context Representations
abstract
Pre-trained distributed word representations have been proven to be useful in various natural language processing (NLP) tasks. However, the geometric basis of word representations and their relations to the representations of word's contexts has not been carefully studied yet. In this study, we first investigate such geometric relationship under a general framework, which is abstracted from some typical word representation learning approaches, and find out that only the directions of word representations are well associated to their context vector representations while the magnitudes are not. In order to make better use of the information contained in the magnitudes of word representations, we propose a hierarchical Gaussian model combined with maximum a posteriori estimation to learn word representations, and extend it to represent polysemous words. Our word representations have been evaluated on multiple NLP tasks, and the experimental results show that the proposed model achieved promising results, comparing to several popular word representations.
Jiangtao Feng, Xiaoqing Zheng
AAAI1
2018 RNN-Based Sequence-Preserved Attention for Dependency Parsing
abstract
Recurrent neural networks (RNN) combined with attention mechanism has proved to be useful for various NLP tasks including machine translation, sequence labeling and syntactic parsing. The attention mechanism is usually applied by estimating the weights (or importance) of inputs and taking the weighted sum of inputs as derived features. Although such features have demonstrated their effectiveness, they may fail to capture the sequence information due to the simple weighted sum being used to produce them. The order of the words does matter to the meaning or the structure of the sentences, especially for syntactic parsing, which aims to recover the structure from a sequence of words. In this study, we propose an RNN-based attention to capture the relevant and sequence-preserved features from a sentence, and use the derived features to perform the dependency parsing. We evaluated the graph-based and transition-based parsing models enhanced with the RNN-based sequence-preserved attention on the both English PTB and Chinese CTB datasets. The experimental results show that the enhanced systems were improved with significant increase in parsing accuracy.
Yi Zhou 0018, Junying Zhou, Lu Liu 0009, Jiangtao Feng, Haoyuan Peng, Xiaoqing Zheng
AAAI4
2017 Learning Context-Specific Word/Character Embeddings
abstract
Unsupervised word representations have demonstrated improvements in predictive generalization on various NLP tasks. Most of the existing models are in fact good at capturing the relatedness among words rather than their ''genuine'' similarity because the context representations are often represented by a sum (or an average) of the neighbor's embeddings, which simplifies the computation but ignores an important fact that the meaning of a word is determined by its context, reflecting not only the surrounding words but also the rules used to combine them (i.e. compositionality). On the other hand, much effort has been devoted to learning a single-prototype representation per word, which is problematic because many words are polysemous, and a single-prototype model is incapable of capturing phenomena of homonymy and polysemy. We present a neural network architecture to jointly learn word embeddings and context representations from large data sets. The explicitly produced context representations are further used to learn context-specific and multi-prototype word embeddings. Our embeddings were evaluated on several NLP tasks, and the experimental results demonstrated the proposed model outperformed other competitors and is applicable to intrinsically "character-based" languages.
Xiaoqing Zheng, Jiangtao Feng, Haoyuan Peng
AAAI2
2016 Context-Specific and Multi-Prototype Character Representations
Xiaoqing Zheng, Jiangtao Feng, Mengxiao Lin
IJCAI2