Jianing Wang 0002

dblp:85/1466-2 · also Jia-ning Wang 0002 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0001-6006-053XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
abstract
Offline evaluation of LLMs is crucial in understanding their capacities, though current methods remain underexplored in existing research. In this work, we focus on the offline evaluation of the chain-of-thought capabilities and show how to optimize LLMs based on the proposed evaluation method. To enable offline feedback with rich knowledge and reasoning paths, we use knowledge graphs (KGs) (e.g., Wikidata5M) to provide feedback on the generated chain of thoughts. Due to the heterogeneity between LLM reasoning and KG structures, direct interaction and feedback from knowledge graphs on LLM behavior are challenging, as they require accurate entity linking and grounding of LLM-generated chains of thought in the KG. To address the above challenge, we propose an offline chain-of-thought evaluation framework, OCEAN, which models chain-of-thought reasoning in LLMs as a Markov Decision Process (MDP), and evaluate the policy’s alignment with KG preference modeling. To overcome the reasoning heterogeneity and grounding problems, we leverage on-policy KG exploration and reinforcement learning to model a KG policy that generates token-level likelihood distributions for LLM-generated chain-of-thought reasoning paths, simulating KG reasoning preference. Then we incorporate the knowledge-graph feedback on the validity and alignment of the generated reasoning paths into inverse propensity scores and propose KG-IPS estimator. Theoretically, we prove the unbiasedness of the proposed KG-IPS estimator and provide a lower bound on its variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance chain-of-thought alignment. Our empirical study shows that OCEAN can be efficiently optimized for generating chain-of-thought reasoning paths with higher estimated values without affecting LLMs’ general abilities in downstream tasks or their internal knowledge.
Junda Wu, Xintong Li 0001, Ruoyu Wang 0038, Yu Xia 0007, Yuxin Xiong, Jianing Wang 0002, Tong Yu 0001, Xiang Chen 0010, Branislav Kveton, Lina Yao 0001, Jingbo Shang, Julian J. McAuley
ICLR6
2024 Boosting Language Models Reasoning with Chain-of-Knowledge Prompting
abstract
Recently, Chain-of-Thought (CoT) prompting has delivered success on complex reasoning tasks, which aims at designing a simple prompt like "Let's think step by step" or multiple incontext exemplars with well-designed rationales to elicit Large Language Models (LLMs) to generate intermediate reasoning steps.However, the generated rationales often come with hallucinations, making unfactual and unfaithful reasoning chains.To mitigate this brittleness, we propose a novel Chain-of-Knowledge (CoK) prompting, where we aim at eliciting LLMs to generate explicit pieces of knowledge evidence in the form of structure triple.This is inspired by our human behaviors, i.e., we can draw a mind map or knowledge map as the reasoning evidence in the brain before answering a complex question.Benefiting from CoK, we additionally introduce a F 2 -Verification method to estimate the reliability of the reasoning chains in terms of factuality and faithfulness.For the unreliable response, the wrong evidence can be indicated to prompt the LLM to rethink.Extensive experiments demonstrate that our method can further improve the performance of commonsense, factual, symbolic, and arithmetic reasoning tasks 1 .
Jianing Wang 0002, Qiushi Sun, Xiang Li 0067, Ming Gao 0001
ACL (1)1
2024 TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills
abstract
Code pre-trained models (CodePTMs) have recently demonstrated a solid capacity to process various code intelligence tasks, e.g., code clone detection, code translation, and code summarization. The current mainstream method that deploys these models to downstream tasks is to fine-tune them on individual tasks, which is generally costly and needs sufficient data for large models. To tackle the issue, in this paper, we present TransCoder, a unified Transferable fine-tuning strategy for Code representation learning. Inspired by human inherent skills of knowledge generalization, TransCoder drives the model to learn better code-related knowledge like human programmers. Specifically, we employ a tunable prefix encoder to first capture cross-task and cross-language transferable knowledge, subsequently applying the acquired knowledge for optimized downstream adaptation. Besides, our approach confers benefits for tasks with minor training sample sizes and languages with smaller corpora, underscoring versatility and efficacy. Extensive experiments conducted on representative datasets clearly demonstrate that our method can lead to superior performance on various code-related tasks and encourage mutual reinforcement, especially in low-resource scenarios. Our codes are available at https://github.com/QiushiSun/TransCoder.
Qiushi Sun, Nuo Chen 0002, Jianing Wang 0002, Ming Gao 0001, Xiang Li 0067
LREC/COLING3
2024 Rethinking the Role of Structural Information: How It Enhances Code Representation Learning?
abstract
Code pre-trained models (CodePTMs) have recently exhibited remarkable accomplishments in the realm of software engineering. However, there are still limited advancements in understanding the inner mechanism of these models, as well as their sensitivity to samples of varying quality. Codes have a more rigid and structured syntax compared to natural languages; hence, leveraging and understanding structural information becomes essential for analyzing, interpreting, and utilizing CodePTMs. While previous studies have verified models’ ability to acquire knowledge from code structure through techniques such as attention analysis and probing tasks, the specific roles it plays in downstream tasks have yet to be explored. In this work, we propose a set of novel and practical methods for probing and exploiting the structural information within the code. In particular, dataflow perturbation experiments are first employed to explore the sensitivity of models with varying levels of structural information when confronted with input changes. Based on our findings, structure-aware exemplars selection strategies are proposed for both code generation and understanding, aiming to recover the model performance at minimal cost under perturbed conditions. Moreover, efficient fine-tuning can be achieved by utilizing exemplars instead of full fine-tuning.
Qiushi Sun, Nuo Chen 0002, Jianing Wang 0002, Xiaoli Li 0001
IJCNN3
2024 CoRAL: Collaborative Retrieval-Augmented Large Language Models Improve Long-tail Recommendation
abstract
The long-tail recommendation is a challenging task for traditional recommender systems, due to data sparsity and data imbalance issues. The recent development of large language models (LLMs) has shown their abilities in complex reasoning, which can help to deduce users' preferences based on very few previous interactions. However, since most LLM-based systems rely on items' semantic meaning as the sole evidence for reasoning, the collaborative information of user-item interactions is neglected, which can cause the LLM's reasoning to be misaligned with task-specific collaborative information of the dataset. To further align LLMs' reasoning to task-specific user-item interaction knowledge, we introduce collaborative retrieval-augmented LLMs, CoRAL, which directly incorporate collaborative evidence into the prompts. Based on the retrieved user-item interactions, the LLM can analyze shared and distinct preferences among users, and summarize the patterns indicating which types of users would be attracted by certain items. The retrieved collaborative evidence prompts the LLM to align its reasoning with the user-item interaction patterns in the dataset. However, since the capacity of the input prompt is limited, finding the minimally-sufficient collaborative information for recommendation tasks can be challenging. We propose to find the optimal interaction set through a sequential decision-making process and develop a retrieval policy learned through a reinforcement learning (RL) framework, CoRAL. Our experimental results show that CoRAL can significantly improve LLMs' reasoning abilities on specific recommendation tasks. Our analysis also reveals that CoRAL can more efficiently explore collaborative information through reinforcement learning.
Junda Wu, Cheng-Chun Chang, Tong Yu 0001, Zhankui He, Jianing Wang 0002, Yupeng Hou, Julian J. McAuley
KDD5
2023 Uncertainty-Aware Self-Training for Low-Resource Neural Sequence Labeling
abstract
Neural sequence labeling (NSL) aims at assigning labels for input language tokens, which covers a broad range of applications, such as named entity recognition (NER) and slot filling, etc. However, the satisfying results achieved by traditional supervised-based approaches heavily depend on the large amounts of human annotation data, which may not be feasible in real-world scenarios due to data privacy and computation efficiency issues. This paper presents SeqUST, a novel uncertain-aware self-training framework for NSL to address the labeled data scarcity issue and to effectively utilize unlabeled data. Specifically, we incorporate Monte Carlo (MC) dropout in Bayesian neural network (BNN) to perform uncertainty estimation at the token level and then select reliable language tokens from unlabeled data based on the model confidence and certainty. A well-designed masked sequence labeling task with a noise-robust loss supports robust training, which aims to suppress the problem of noisy pseudo labels. In addition, we develop a Gaussian-based consistency regularization technique to further improve the model robustness on Gaussian-distributed perturbed representations. This effectively alleviates the over-fitting dilemma originating from pseudo-labeled augmented data. Extensive experiments over six benchmarks demonstrate that our SeqUST framework effectively improves the performance of self-training, and consistently outperforms strong baselines by a large margin in low-resource scenarios.
Jianing Wang 0002, Chengyu Wang 0001, Jun Huang 0007, Ming Gao 0001, Aoying Zhou
AAAI1
2023 HugNLP: A Unified and Comprehensive Library for Natural Language Processing
abstract
In this paper, we introduce HugNLP, a unified and comprehensive library for natural language processing (NLP) with the prevalent backend of Hugging Face Transformers, which is designed for NLP researchers to easily utilize off-the-shelf algorithms and develop novel methods with user-defined models and tasks in real-world scenarios. HugNLP consists of a hierarchical structure including models, processors and applications that unifies the learning process of pre-trained language models (PLMs) on different NLP tasks. Additionally, we present some featured NLP applications to show the effectiveness of HugNLP, such as knowledge-enhanced PLMs, universal information extraction, low-resource mining, and code understanding and generation, etc. The source code will be released on GitHub (https://github.com/HugAILab/HugNLP).
Jianing Wang 0002, Nuo Chen 0002, Qiushi Sun, Wenkang Huang, Chengyu Wang 0001, Ming Gao 0001
CIKM1
2023 Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question Generation
abstract
The task of Question Generation over Knowledge Bases (KBQG) aims to convert a logical form into a natural language question.For the sake of expensive cost of large-scale question annotation, the methods of KBQG under low-resource scenarios urgently need to be developed.However, current methods heavily rely on annotated data for fine-tuning, which is not well-suited for few-shot question generation.The emergence of Large Language Models (LLMs) has shown their impressive generalization ability in few-shot tasks.Inspired by Chain-of-Thought (CoT) prompting, which is an in-context learning strategy for reasoning, we formulate KBQG task as a reasoning problem, where the generation of a complete question is split into a series of sub-question generation.Our proposed prompting method KQG-CoT first selects supportive logical forms from the unlabeled data pool taking account of the characteristics of the logical form.Then, we construct a task-specific prompt to guide LLMs to generate complicated questions based on selective logic forms.To further ensure prompt quality, we extend KQG-CoT into KQG-CoT+ via sorting the logical forms by their complexity.We conduct extensive experiments over three public KBQG datasets.The results demonstrate that our prompting method consistently outperforms other prompting baselines on the evaluated datasets.Remarkably, our KQG-CoT+ method could surpass existing fewshot SoTA results of the PathQuestions dataset by 18.25, 10.72, and 10.18 absolute points on BLEU-4, METEOR, and ROUGE-L, respectively.
Yuanyuan Liang, Jianing Wang 0002, Hanlun Zhu, Weining Qian, Yunshi Lan
EMNLP2
2023 ParaSum: Contrastive Paraphrasing for Low-Resource Extractive Text Summarization
Moming Tang, Chengyu Wang 0001, Jianing Wang 0002, Cen Chen 0001, Ming Gao 0001, Weining Qian
KSEM (3)3
2023 UKT: A Unified Knowledgeable Tuning Framework for Chinese Information Extraction
Jiyong Zhou, Chengyu Wang 0001, Jianing Wang 0002, Yukang Xie, Jun Huang 0007, Ying Gao 0004
NLPCC (2)4
2022 KECP: Knowledge Enhanced Contrastive Prompting for Few-shot Extractive Question Answering
abstract
Extractive Question Answering (EQA) is one of the most essential tasks in Machine Reading Comprehension (MRC), which can be solved by fine-tuning the span selecting heads of Pre-trained Language Models (PLMs). However, most existing approaches for MRC may perform poorly in the few-shot learning scenario. To solve this issue, we propose a novel framework named Knowledge Enhanced Contrastive Prompt-tuning (KECP). Instead of adding pointer heads to PLMs, we introduce a seminal paradigm for EQA that transforms the task into a non-autoregressive Masked Language Modeling (MLM) generation problem. Simultaneously, rich semantics from the external knowledge base (KB) and the passage context support enhancing the query’s representations. In addition, to boost the performance of PLMs, we jointly train the model by the MLM and contrastive learning objectives. Experiments on multiple benchmarks demonstrate that our method consistently outperforms state-of-the-art approaches in few-shot settings by a large margin.
Jianing Wang 0002, Chengyu Wang 0001, Minghui Qiu, Qiuhui Shi, Jun Huang 0007, Ming Gao 0001
EMNLP1
2022 SpanProto: A Two-stage Span-based Prototypical Network for Few-shot Named Entity Recognition
abstract
Few-shot Named Entity Recognition (NER) aims to identify named entities with very little annotated data.Previous methods solve this problem based on token-wise classification, which ignores the information of entity boundaries, and inevitably the performance is affected by the massive non-entity tokens.To this end, we propose a seminal span-based prototypical network (SpanProto) that tackles few-shot NER via a two-stage approach, including span extraction and mention classification.In the span extraction stage, we transform the sequential tags into a global boundary matrix, enabling the model to focus on the explicit boundary information.For mention classification, we leverage prototypical learning to capture the semantic representations for each labeled span and make the model better adapt to novel-class entities.To further improve the model performance, we split out the false positives generated by the span extractor but not labeled in the current episode set, and then present a margin-based loss to separate them from each prototype region.Experiments over multiple benchmarks demonstrate that our model outperforms strong baselines by a large margin. 1
Jianing Wang 0002, Chengyu Wang 0001, Chuanqi Tan, Minghui Qiu, Songfang Huang, Jun Huang 0007, Ming Gao 0001
EMNLP1
2022 Knowledge Prompting in Pre-trained Language Model for Natural Language Understanding
abstract
Knowledge-enhanced Pre-trained Language Model (PLM) has recently received significant attention, which aims to incorporate factual knowledge into PLMs.However, most existing methods modify the internal structures of fixed types of PLMs by stacking complicated modules, and introduce redundant and irrelevant factual knowledge from knowledge bases (KBs).In this paper, to address these problems, we introduce a seminal knowledge prompting paradigm and further propose a knowledge-prompting-based PLM framework KP-PLM.This framework can be flexibly combined with existing mainstream PLMs.Specifically, we first construct a knowledge sub-graph from KBs for each context.Then we design multiple continuous prompts rules and transform the knowledge sub-graph into natural language prompts.To further leverage the factual knowledge from these prompts, we propose two novel knowledge-aware self-supervised tasks including prompt relevance inspection and masked prompt modeling.Extensive experiments on multiple natural language understanding (NLU) tasks show the superiority of KP-PLM over other state-of-the-art methods in both full-resource and low-resource settings 1 .
Jianing Wang 0002, Wenkang Huang, Minghui Qiu, Qiuhui Shi, Xiang Li 0067, Ming Gao 0001
EMNLP1
2021 TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text Classification
abstract
Recent studies have shown that prompts improve the performance of large pre-trained language models for few-shot text classification.Yet, it is unclear how the prompting knowledge can be transferred across similar NLP tasks for the purpose of mutual reinforcement.Based on continuous prompt embeddings, we propose TransPrompt, a transferable prompting framework for few-shot learning across similar tasks.In TransPrompt, we employ a multitask meta-knowledge acquisition procedure to train a meta-learner that captures cross-task transferable knowledge.Two de-biasing techniques are further designed to make it more task-agnostic and unbiased towards any tasks.After that, the meta-learner can be adapted to target tasks with high accuracy.Extensive experiments show that TransPrompt outperforms single-task and cross-task strong baselines over multiple NLP tasks and datasets.We further show that the meta-learner can effectively improve the performance on previously unseen tasks.TransPrompt also outperforms strong fine-tuning baselines when learning with full training sets.
Chengyu Wang 0001, Jianing Wang 0002, Minghui Qiu, Jun Huang 0007, Ming Gao 0001
EMNLP (1)2