VLDB 2026 Research / reviewers in the wild / expert
Fangkai Jiao
dblp:264/9981
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-0670-6990ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 48% Question answering and dialogue systems · 14% Knowledge representation and reasoning · 6% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 67% Recommender systems · 26% Knowledge graphs · 7% |
Topics — the 24 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
preference optimization |
1.6 | 2 | 2025 | Preference Optimization for Reasoning with Pseudo Feedback · ICLR 2025 Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing · EMNLP 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Preference Optimization for Reasoning with Pseudo Feedback · ICLR 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks · ACL (1) 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.8 | 1 | 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing · EMNLP 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › reasoning about action and change
plan inference |
0.8 | 1 | 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing · EMNLP 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model |
0.8 | 1 | 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing · EMNLP 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.8 | 1 | 2024 | UNK-VQA: A Dataset and a Probe Into the Abstention Ability of Multi-Modal Large Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | UNK-VQA: A Dataset and a Probe Into the Abstention Ability of Multi-Modal Large Models · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking |
0.7 | 1 | 2023 | Enhanced Multi-Domain Dialogue State Tracker With Second-Order Slot Interactions · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.7 | 1 | 2023 | Enhanced Multi-Domain Dialogue State Tracker With Second-Order Slot Interactions · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Machine learning › Graph learning
heterogeneous graph learning |
0.6 | 1 | 2022 | Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph Learning · SIGIR 2022 |
Recommender systems › fashion recommendation
fashion compatibility modeling |
0.6 | 1 | 2022 | Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph Learning · SIGIR 2022 |
Information retrieval › interactive information retrieval › conversational information seeking › conversational search
conversational image retrieval |
0.5 | 1 | 2021 | Conversational Image Search · IEEE Trans. Image Process. 2021 |
Information retrieval
image retrieval |
0.5 | 1 | 2021 | Conversational Image Search · IEEE Trans. Image Process. 2021 |
Information retrieval
retrieval models |
0.5 | 1 | 2021 | Conversational Image Search · IEEE Trans. Image Process. 2021 |
Natural language and speech › Information extraction and text analysis
evidence extraction |
0.4 | 1 | 2020 | A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.4 | 1 | 2020 | A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.4 | 1 | 2020 | A Self-Training Method for Machine Reading Comprehension with Soft Evidence Extraction · ACL 2020 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.3 | 1 | 2025 | Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks · ACL (1) 2025 |
Natural language and speech › Language models and text generation
code generation |
0.3 | 1 | 2025 | Preference Optimization for Reasoning with Pseudo Feedback · ICLR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.3 | 1 | 2025 | Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2022 | Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph Learning · SIGIR 2022 |
Methods — techniques the papers use, named apart from their topics
direct preference optimization · 1.6test-case-based feedback · 0.9self-consistency · 0.9output matching · 0.9monte carlo tree search · 0.9fine-tuning · 0.9critic model · 0.9alignment · 0.9trajectory collection · 0.8few-shot evaluation · 0.8multi-modal user embedding · 0.6metapath-guided heterogeneous graph learning · 0.6contrastive regularization · 0.6multimodal hierarchical graph neural network · 0.5memory network · 0.5gated neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging TasksabstractLarge language models excel at problemsolving but often struggle with complex reasoning and factual accuracy.While chainof-thought and retrieval-augmented generation help break down problems and retrieve knowledge, they still falter on challenging tasks like competitive programming due to frequent reasoning errors and irrelevant retrieval.To address this, we introduce Critic-guided planning with Retrieval-augmentation, CR-Planner, a novel framework that leverages fine-tuned critic models to guide both reasoning and retrieval processes through planning.CR-Planner iteratively selects and executes sub-goals, guided by critic models.A sub-goal critic identifies promising sub-goals from reasoning, query generation, and retrieval, while an execution critic evaluates outputs of sub-goal executions.We employ Monte Carlo Tree Search to collect data for critic training, allowing systematic exploration of action sequences and effective navigation toward the final answer.We evaluate CR-Planner on challenging domain-knowledgeintensive and reasoning-heavy tasks, including competitive programming, theorem-driven math reasoning, and complex domain retrieval problems.It significantly outperforms baselines, demonstrating effectiveness in both reasoning and retrieval.Our code is available at https://github.com/xingxuanli/CR-Planner. Xingxuan Li, Weiwen Xu, Fangkai Jiao, Shafiq R. Joty, Lidong Bing |
ACL (1) | 4 |
| 2025 | Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningabstractChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen, Yuchen Hu, Bosheng Ding, Ruirui Chen, Shafiq Joty. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen 0075, Bosheng Ding, Ruirui Chen 0002, Shafiq R. Joty |
ACL (1) | 3 |
| 2025 | Preference Optimization for Reasoning with Pseudo FeedbackabstractPreference optimization techniques, such as Direct Preference Optimization (DPO), are frequently employed to enhance the reasoning capabilities of large language models (LLMs) in domains like mathematical reasoning and coding, typically following supervised fine-tuning. These methods rely on high-quality labels for reasoning tasks to generate preference pairs; however, the availability of reasoning datasets with human-verified labels is limited.
In this study, we introduce a novel approach to generate pseudo feedback for reasoning tasks by framing the labeling of solutions to reason problems as an evaluation against associated \emph{test cases}.
We explore two forms of pseudo feedback based on test cases: one generated by frontier LLMs and the other by extending self-consistency to multi-test-case.
We conduct experiments on both mathematical reasoning and coding tasks using pseudo feedback for preference optimization, and observe improvements across both tasks. Specifically, using Mathstral-7B as our base model, we improve MATH results from 58.3 to 68.6, surpassing both NuminaMath-72B and GPT-4-Turbo-1106-preview. In GSM8K and College Math, our scores increase from 85.6 to 90.3 and from 34.3 to 42.3, respectively. Building on Deepseek-coder-7B-v1.5, we achieve a score of 24.3 on LiveCodeBench (from 21.1), surpassing Claude-3-Haiku. Fangkai Jiao, Geyang Guo, Nancy F. Chen, Shafiq R. Joty, Furu Wei |
ICLR | 1 |
| 2025 | The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and DefenseabstractThe vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise.
However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evaluations, often with minimal effort.
This dual high performance in both attack and defense gives rise to a fundamental and perplexing paradox.
To gain a deep understanding of this issue and thus further help strengthen the trustworthiness of VLLMs, this paper makes three key contributions:
i) One tentative explanation for VLLMs being prone to jailbreak attacks--inclusion of vision inputs, as well as its in-depth analysis.
ii) The recognition of a largely ignored problem in existing VLLM defense mechanisms--over-prudence.
The problem causes these defense methods to exhibit unintended abstention, even in the presence of benign inputs, thereby undermining their reliability in faithfully defending against attacks.
iii) A simple safety-aware method--LLM-Pipeline.
Our method repurposes the more advanced guardrails of LLMs on the fly, serving as an effective alternative detector prior to VLLM response.
Last but not least, we find that the two representative evaluation methods for jailbreak often exhibit chance agreement.
This limitation makes it potentially misleading when evaluating attack strategies or defense mechanisms.
We believe the findings from this paper offer useful insights to rethink the foundational development of VLLM safety with respect to benchmark datasets, defense strategies, and evaluation methods. Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli |
NeurIPS | 2 |
| 2024 | Learning Planning-based Reasoning by Trajectories Collection and Process Reward SynthesizingabstractLarge Language Models (LLMs) have demonstrated significant potential in handling complex reasoning tasks through step-by-step rationale generation.However, recent studies have raised concerns regarding the hallucination and flaws in their reasoning process.Substantial efforts are being made to improve the reliability and faithfulness of the generated rationales.Some approaches model reasoning as planning, while others focus on annotating for process supervision.Nevertheless, the planning-based search process often results in high latency due to the frequent assessment of intermediate reasoning states and the extensive exploration space.Additionally, supervising the reasoning process with human annotation is costly and challenging to scale for LLM training.To address these issues, in this paper, we propose a framework to learn planning-based reasoning through Direct Preference Optimization (DPO) on collected trajectories, which are ranked according to our synthesized process rewards.Our results on challenging logical reasoning benchmarks demonstrate the effectiveness of our learning framework, showing that our 7B model can surpass the strong counterparts like GPT-3.5-Turbo. Fangkai Jiao, Chengwei Qin, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty |
EMNLP | 1 |
| 2024 | Exploring Self-supervised Logic-enhanced Training for Large Language ModelsabstractFangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy Chen, Shafiq Joty. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty |
NAACL-HLT | 1 |
| 2024 | SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural ReasoningabstractBin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, AiTi Aw, Nancy Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bin Wang 0040, Zhengyuan Liu, Fangkai Jiao, AiTi Aw, Nancy F. Chen |
NAACL-HLT | 4 |
| 2024 | UNK-VQA: A Dataset and a Probe Into the Abstention Ability of Multi-Modal Large ModelsabstractTeaching Visual Question Answering (VQA) models to refrain from answering unanswerable questions is necessary for building a trustworthy AI system. Existing studies, though have explored various aspects of VQA but somewhat ignored this particular attribute. This paper aims to bridge the research gap by contributing a comprehensive dataset, called UNK-VQA. The dataset is specifically designed to address the challenge of questions that models do not know. To this end, we first augment the existing data via deliberate perturbations on either the image or question. In specific, we carefully ensure that the question-image semantics remain close to the original unperturbed distribution. By this means, the identification of unanswerable questions becomes challenging, setting our dataset apart from others that involve mere image replacement. We then extensively evaluate the zero- and few-shot performance of several emerging multi-modal large models and discover their significant limitations when applied to our dataset. Additionally, we also propose a straightforward method to tackle these unanswerable questions. This dataset, we believe, will serve as a valuable benchmark for enhancing the abstention capability of VQA models, thereby leading to increased trustworthiness of AI systems. We have made the dataset available to facilitate further exploration in this area. Fangkai Jiao, Zhiqi Shen 0002, Liqiang Nie, Mohan Kankanhalli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Enhanced Multi-Domain Dialogue State Tracker With Second-Order Slot InteractionsabstractDialogue state tracking (DST) is often used to track the system's understanding of the user goal in task-oriented dialogue systems. Existing DST methods mainly fall into two categories according to their adopted model structure: non-hierarchical and hierarchical models. The former takes the whole dialogue history as inputs during each conversation round, while the latter leverages both an utterance encoder and a dialogue encoder to efficiently model the long-term dialogue dependency. However, few of them exploit the second-order slot interaction, which refers to the pair-wise semantic relationships between different slots. As a result, these methods fall short in the context understanding throughout conversations, leading to sub-optimal performance. Towards this end, in this paper, we present a novel hierarchy-based DST framework equipped with a well-designed value copy mechanism. In particular, to model the second-order slot interaction, we firstly encode the utterance via a state reuse module to yield slot-sensitive context representation. We then selectively and effectively copy the filled values from other slots to attain more accurate state tracking. In order to evaluate the effectiveness of the proposed method, we perform extensive experiments on the widely adopted benchmark dataset MultiWOZ2.1. Our experimental results demonstrate the superiority in context understanding, as well as the strong generalization capability under a zero-shot setting compared with several DST baselines. Fangkai Jiao, Minlie Huang, Liqiang Nie |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph LearningabstractFashion Compatibility Modeling (FCM) is a new yet challenging task, which aims to automatically access the matching degree among a set of complementary items. Most of existing methods evaluate the fashion compatibility from the common perspective, but overlook the user's personal preference. Inspired by this, a few pioneers study the Personalized Fashion Compatibility Modeling (PFCM). Despite their significance, these PFCM methods mainly concentrate on the user and item entities, as well as their interactions, but ignore the attribute entities, which contain rich semantics. To address this problem, we propose to fully explore the related entities and their relations involved in PFCM to boost the PFCM performance. This is, however, non-trivial due to the heterogeneous contents of different entities, embeddings for new users, and various high-order relations. Towards these ends, we present a novel metapath-guided personalized fashion compatibility modeling, dubbed as MG-PFCM. In particular, we creatively build a heterogeneous graph to unify the three types of entities (i.e., users, items, and attributes) and their relations (i.e., user-item interactions, item-item matching relations, and item-attribute association relations). Thereafter, we design a multi-modal content-oriented user embedding module to learn user representations by inheriting the contents of their interacted items. Meanwhile, we define the user-oriented and item-oriented metapaths, and perform the metapath-guided heterogeneous graph learning to enhance the user and item embeddings. In addition, we introduce the contrastive regularization to improve the model performance. We conduct extensive experiments on the real-world benchmark dataset, which verifies the superiority of our proposed scheme over several cutting-edge baselines. As a byproduct, we have released our source codes to benefit other researchers. Weili Guan, Fangkai Jiao, Xuemeng Song, Haokun Wen, Chung-Hsing Yeh, Xiaojun Chang |
SIGIR | 2 |
| 2021 | Conversational Image SearchabstractConversational image search, a revolutionary search mode, is able to interactively induce the user response to clarify their intents step by step. Several efforts have been dedicated to the conversation part, namely automatically asking the right question at the right time for user preference elicitation, while few studies focus on the image search part given the well-prepared conversational query. In this paper, we work towards conversational image search, which is much difficult compared to the traditional image search task, due to the following challenges: 1) understanding complex user intents from a multimodal conversational query; 2) utilizing multiform knowledge associated images from a memory network; and 3) enhancing the image representation with distilled knowledge. To address these problems, in this paper, we present a novel contextuaL imAge seaRch sCHeme (LARCH for short), consisting of three components. In the first component, we design a multimodal hierarchical graph-based neural network, which learns the conversational query embedding for better user intent understanding. As to the second one, we devise a multi-form knowledge embedding memory network to unify heterogeneous knowledge structures into a homogeneous base that greatly facilitates relevant knowledge retrieval. In the third component, we learn the knowledge-enhanced image representation via a novel gated neural network, which selects the useful knowledge from retrieved relevant one. Extensive experiments have shown that our LARCH yields significant performance over an extended benchmark dataset. As a side contribution, we have released the data, codes, and parameter settings to facilitate other researchers in the conversational image search community. Liqiang Nie, Fangkai Jiao, Wenjie Wang 0007, Yinglong Wang 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | A Self-Training Method for Machine Reading Comprehension with Soft Evidence ExtractionabstractNeural models have achieved great success on machine reading comprehension (MRC), many of which typically consist of two components: an evidence extractor and an answer predictor.The former seeks the most relevant information from a reference text, while the latter is to locate or generate answers from the extracted evidence.Despite the importance of evidence labels for training the evidence extractor, they are not cheaply accessible, particularly in many non-extractive MRC tasks such as YES/NO question answering and multi-choice MRC.To address this problem, we present a Self-Training method (STM), which supervises the evidence extractor with auto-generated evidence labels in an iterative process.At each iteration, a base MRC model is trained with golden answers and noisy evidence labels.The trained model will predict pseudo evidence labels as extra supervision in the next iteration.We evaluate STM on seven datasets over three MRC tasks.Experimental results demonstrate the improvement on existing MRC models, and we also analyze how and why such a self-training method works in MRC. Yilin Niu, Fangkai Jiao, Mantong Zhou, Jingfang Xu, Minlie Huang |
ACL | 2 |