VLDB 2026 Research / reviewers in the wild / expert
Jiangming Liu
dblp:154/8222
· DBLP profile ↗
23ranked-venue papers
13as first author
12since 2021 · last 2026
0000-0002-2178-2385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 13 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human PreferenceabstractHumans increasingly query Large Language Models (LLMs) to accomplish personal tasks according to their individual preferences. However, these preferences are often unconsciously veiled during conversation. To address this, LLMs have to elicit human preferences through multi-turn dialogue, where tasks are accomplished via iterative clarifying questions and final response generated by LLMs as effective questioners. Existing approaches based on self-taught reasoning have two limitations: 1) they struggle to avoid generating irrelevant questions and 2) the final responses to tasks are misled by the conversations. To overcome these limitations, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization. TO-GATE comprises two key components: a clarification resolver, which generates optimal questioning trajectories to produce effective elicitation questions, and a summarizer, which ensures task-aligned final responses. Experimental results show that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation benchmarks. Yulin Dou, Jiangming Liu |
AAAI | 2 |
| 2026 | Structural entropy guided relation extraction on adaptive graph structure
Jiangming Liu, Kun Yue, Liang Duan, Jiande Wu |
Pattern Recognit. | 2 |
| 2025 | QuanSIRA: The Quantitative Investment Risk Modeling in Stock Markets with Pre-trained Language ModelsabstractStock market analysis is important for investors to make financial decisions. Stock price prediction is widely investigated with the help of superiority of pre-trained language models. Recent works have developed several datasets for stock price predictions. However, investment risk, considered an essential factor for investors, is rarely discussed due to the serious challenge of quantification over the risk. In this work, we propose to quantify investment risk and introduce the dataset QuanSIRA that can be used for investment risk modeling. The experimental results show that the model built on pre-trained language models obtained F1 scores of 69.75 and 65.01 in the in-stock benchmark and the cross-stock benchmark of investment risk prediction task. Zheyang Luo, Jiangming Liu |
IJCNN | 2 |
| 2025 | Adaptive Federated Distillation for Multi-Domain Non-IID Textual DataabstractThe widespread success of pre-trained language models has established a new training paradigm, where a global PLM is fine-tuned using task-specific data from local clients. The local data are highly different from each other and can not capture the global distribution of the whole data in real world. To address the challenges of non-IID data in real environments, privacy-preserving federated distillation has been proposed and highly investigated. However, previous experimental non-IID scenarios are primarily identified with the label (output) diversity, without considering the diversity of language domains (input) that is crucial in natural language processing. In this paper, we introduce a comprehensive set of multi-domain non-IID scenarios and propose a unified benchmarking framework that includes diverse data. The benchmark can be used to evaluate the federated learning framework in a real environment. To this end, we propose an Adaptive Federated Distillation (AdaFD) framework designed to address multi-domain non-IID challenges in both homogeneous and heterogeneous settings. Experimental results demonstrate that our models capture the diversity of local clients and achieve better performance compared to the existing works. The code for this paper is available at: https://github.com/jiahaoxiao1228/AdaFD. Jiahao Xiao, Jiangming Liu |
IJCNN | 2 |
| 2024 | Model-Agnostic Cross-Lingual Training for Discourse Representation Structure ParsingabstractDiscourse Representation Structure (DRS) is an innovative semantic representation designed to capture the meaning of texts with arbitrary lengths across languages. The semantic representation parsing is essential for achieving natural language understanding through logical forms. Nevertheless, the performance of DRS parsing models remains constrained when trained exclusively on monolingual data. To tackle this issue, we introduce a cross-lingual training strategy. The proposed method is model-agnostic yet highly effective. It leverages cross-lingual training data and fully exploits the alignments between languages encoded in pre-trained language models. The experiments conducted on the standard benchmarks demonstrate that models trained using the cross-lingual training method exhibit significant improvements in DRS clause and graph parsing in English, German, Italian and Dutch. Comparing our final models to previous works, we achieve state-of-the-art results in the standard benchmarks. Furthermore, the detailed analysis provides deep insights into the performance of the parsers, offering inspiration for future research in DRS parsing. Jiangming Liu |
LREC/COLING | 1 |
| 2024 | Soft Well-Formed Semantic Parsing with Score-Based SelectionabstractSemantic parsing is the task of translating natural language into a structured, formal semantic representation that can be interpreted by machines. These semantic representations are organized with complex structures. While various models have been developed for semantic parsing, there has been limited focus on generating semantic representations with well-formed structures. In this study, we introduce a score-based method to select well-formed outputs from candidates generated by beam search algorithms. Our experiments focus on parsing texts into discourse representation structures, which are innovative semantic representations designed to capture the meaning of texts with arbitrary lengths across languages. Our experimental results demonstrate that models utilizing the proposed method can reduce the number of ill-formed outputs and improve F1 scores in English. Furthermore, our final model achieves significant improvements in German, Italian and Dutch zero-shot DRS parsing by effectively preventing ill-formed outputs. Jiangming Liu |
LREC/COLING | 1 |
| 2024 | Uncovering what, why and How: A Comprehensive Benchmark for Causation Understanding of Video AnomalyabstractVideo anomaly understanding (VAU) aims to automat-ically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial manufacturing. While existing VAU benchmarks primarily concentrate on anomaly detection and localization, our focus is on more practicality, prompting us to raise the following crucial questions: “what anomaly occurred?”,”why did it happen?”, and “how severe is this abnormal event?”. In pursuit of these answers, we present a comprehensive benchmark for Causation Understanding of Video Anomaly (CUVA). Specifically, each instance of the proposed benchmark involves three sets of human annotations to indicate the”what”, “why” and “how” of an anomaly, including 1) anomaly type, start and end times, and event descriptions, 2) natural language explanations for the cause of an anomaly, and 3) free text reflecting the effect of the abnormality. In addition, we also introduce MMEval, a novel evaluation metric designed to better align with human preferences for CUVA, facilitating the measurement of existing LLMs in comprehending the underlying cause and corresponding effect of video anoma-lies. Finally, we propose a novel prompt-based method that can serve as a baseline approach for the challenging CUVA. We conduct extensive experiments to show the superiority of our evaluation metric and the prompt-based approach. Our code and dataset are available at https://github.com/fesvhtr/CUVA. Binzhu Xie, Guoshun Nan, Junrui Xu, Hangyu Liu 0001, Sicong Leng, Jiangming Liu, Hehe Fan, Dajiu Huang, Linli Chen, Xuhuan Li, Jianhang Chen, Qimei Cui, Xiaofeng Tao 0001 |
CVPR | 9 |
| 2024 | Controllable Multi-Behavior Recommendation for In-Game Skins with Large Sequential ModelabstractOnline games often house virtual shops where players can acquire character skins. Our task is centered on tailoring skin recommendations across diverse scenarios by analyzing historical interactions such as clicks, usage, and purchases. Traditional multi-behavior recommendation models employed for this task are limited. They either only predict skins based on a single type of behavior or merely recommend skins for target behavior type/task. These models lack the ability to control predictions of skins that are associated with different scenarios and behaviors. To overcome these limitations, we utilize the pretraining capabilities of Large Sequential Models (LSMs) coupled with a novel stimulus prompt mechanism and build a controllable multi-behavior recommendation (CMBR) model. In our approach, the pretraining ability is used to encapsulate users' multi-behavioral sequences into the representation of users' general interests. Subsequently, our designed stimulus prompt mechanism stimulates the model to extract scenario-related interests, thus generating potential skin purchases (or clicks and other interactions) for users. To the best of our knowledge, this is the first work to provide controlled multi-behavior recommendations, and also the first to apply the pretraining capabilities of LSMs in game domain. Through offline experiments and online A/B tests, we validate our method significantly outperforms baseline models, exhibiting about a tenfold improvement on various metrics during the offline test. Yanjie Gou, Yuanzhou Yao, Zhao Zhang 0011, Yiqing Wu, Fuzhen Zhuang, Jiangming Liu, Yongjun Xu 0001 |
KDD | 7 |
| 2023 | FedID: Federated Interactive Distillation for Large-Scale Pretraining Language ModelsabstractThe growing concerns and regulations surrounding the protection of user data privacy have necessitated decentralized training paradigms.To this end, federated learning (FL) is widely studied in user-related natural language processing (NLP).However, it suffers from several critical limitations including extensive communication overhead, inability to handle heterogeneity, and vulnerability to white-box inference attacks.Federated distillation (FD) is proposed to alleviate these limitations, but its performance is faded by confirmation bias.To tackle this issue, we propose Federated Interactive Distillation (FedID), which utilizes a small amount of labeled data retained by the server to further rectify the local models during knowledge transfer.Additionally, based on the GLUE benchmark, we develop a benchmarking framework across multiple tasks with diverse data distributions to contribute to the research of FD in NLP community.Experiments show that our proposed Fe-dID framework achieves the best results in homogeneous and heterogeneous federated scenarios.The code for this paper is available at: https://github.com/maxinge8698/FedID. Xinge Ma, Jiangming Liu, Jin Wang 0008, Xuejie Zhang 0002 |
EMNLP | 2 |
| 2023 | Semantic Candidate Retrieval for Few-Shot Entity Linking
Jianyong Chen, Jiangming Liu, Jin Wang 0008, Xuejie Zhang 0002 |
NLPCC (3) | 2 |
| 2021 | Text Generation from Discourse Representation StructuresabstractWe propose neural models to generate text from formal meaning representations based on Discourse Representation Structures (DRSs).DRSs are document-level representations which encode rich semantic detail pertaining to rhetorical relations, presupposition, and co-reference within and across sentences.We formalize the task of neural DRS-to-text generation and provide modeling solutions for the problems of condition ordering and variable naming which render generation from DRSs non-trivial.Our generator relies on a novel sibling treeLSTM model which is able to accurately represent DRS structures and is more generally suited to trees with wide branches.We achieve competitive performance (59.48 BLEU) on the GMB benchmark against several strong baselines. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
NAACL-HLT | 1 |
| 2021 | Universal Discourse Representation Structure ParsingabstractAbstract We consider the task of crosslingual semantic parsing in the style of Discourse Representation Theory (DRT) where knowledge from annotated corpora in a resource-rich language is transferred via bitext to guide learning in other languages. We introduce Universal Discourse Representation Theory (UDRT), a variant of DRT that explicitly anchors semantic representations to tokens in the linguistic input. We develop a semantic parsing framework based on the Transformer architecture and utilize it to obtain semantic resources in multiple languages following two learning schemes. The Many-to-One approach translates non-English text to English, and then runs a relatively accurate English parser on the translated text, while the One-to-Many approach translates gold standard English to non-English text and trains multiple parsers (one per language) on the translations. Experimental results on the Parallel Meaning Bank show that our proposal outperforms strong baselines by a wide margin and can be used to construct (silver-standard) meaning banks for 99 languages. Jiangming Liu, Shay B. Cohen, Mirella Lapata, Johan Bos |
Comput. Linguistics | 1 |
| 2020 | DRTS Parsing with Structure-Aware Encoding and DecodingabstractDiscourse representation tree structure (DRTS) parsing is a novel semantic parsing task which has been concerned most recently.State-of-the-art performance can be achieved by a neural sequence-to-sequence model, treating the tree construction as an incremental sequence generation problem.Structural information such as input syntax and the intermediate skeleton of the partial output has been ignored in the model, which could be potentially useful for the DRTS parsing.In this work, we propose a structural-aware model at both the encoder and decoder phase to integrate the structural information, where graph attention network (GAT) is exploited for effectively modeling.Experimental results on a benchmark dataset show that our proposed model is effective and can obtain the best performance in the literature. Qiankun Fu, Yue Zhang 0004, Jiangming Liu, Meishan Zhang |
ACL | 3 |
| 2020 | Dscorer: A Fast Evaluation Metric for Discourse Representation Structure ParsingabstractDiscourse representation structures (DRSs) are scoped semantic representations for texts of arbitrary length.Evaluation of the accuracy of predicted DRSs plays a key role in developing semantic parsers and improving their performance.DRSs are typically visualized as nested boxes, in a way that is not straightforward to process automatically.COUNTER, an evaluation algorithm for DRSs, transforms them to clauses and measures clause overlap by searching for variable mappings between two DRSs.Unfortunately, COUNTER is computationally costly (with respect to memory and CPU time) and does not scale with longer texts.We introduce DSCORER, an efficient new metric which converts box-style DRSs to graphs and then measures the overlap of n-grams in the graphs.Experiments show that DSCORER computes accuracy scores that correlate with scores from COUNTER at a fraction of the time. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL | 1 |
| 2020 | Multi-Step Inference for Reasoning Over ParagraphsabstractComplex reasoning over text requires understanding and chaining together free-form predicates and logical connectives.Prior work has largely tried to do this either symbolically or with black-box transformers.We present a middle ground between these two extremes: a compositional model reminiscent of neural module networks that can perform chained logical reasoning.This model first finds relevant sentences in the context and then chains them together using neural modules.Our model gives significant performance improvements (up to 29% relative error reduction when combined with a reranker) on ROPES, a recentlyintroduced complex reasoning dataset. Jiangming Liu, Matt Gardner 0001, Shay B. Cohen, Mirella Lapata |
EMNLP (1) | 1 |
| 2019 | Discourse Representation Parsing for Sentences and DocumentsabstractWe introduce a novel semantic parsing task based on Discourse Representation Theory (DRT; Kamp and Reyle 1993).Our model operates over Discourse Representation Tree Structures which we formally define for sentences and documents.We present a general framework for parsing discourse structures of arbitrary length and granularity.We achieve this with a neural model equipped with a supervised hierarchical attention mechanism and a linguistically-motivated copy strategy.Experimental results on sentence-and documentlevel benchmarks show that our model outperforms competitive baselines by a wide margin. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL (1) | 1 |
| 2018 | Discourse Representation Structure ParsingabstractWe introduce an open-domain neural semantic parser which generates formal meaning representations in the style of Discourse Representation Theory (DRT; Kamp and Reyle 1993).We propose a method which transforms Discourse Representation Structures (DRSs) to trees and develop a structure-aware model which decomposes the decoding process into three stages: basic DRS structure prediction, condition prediction (i.e., predicates and relations), and referent prediction (i.e., variables).Experimental results on the Groningen Meaning Bank (GMB) show that our model outperforms competitive baselines by a wide margin. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL (1) | 1 |
| 2018 | Learning Domain Representation for Multi-Domain Sentiment ClassificationabstractQi Liu, Yue Zhang, Jiangming Liu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Qi Liu 0049, Yue Zhang 0004, Jiangming Liu |
NAACL-HLT | 3 |
| 2017 | Shift-Reduce Constituent Parsing with Neural Lookahead FeaturesabstractTransition-based models can be fast and accurate for constituent parsing. Compared with chart-based models, they leverage richer features by extracting history information from a parser stack, which consists of a sequence of non-local constituents. On the other hand, during incremental parsing, constituent information on the right hand side of the current word is not utilized, which is a relative weakness of shift-reduce parsing. To address this limitation, we leverage a fast neural model to extract lookahead features. In particular, we build a bidirectional LSTM model, which leverages full sentence information to predict the hierarchy of constituents that each word starts and ends. The results are then passed to a strong transition-based constituent parser as lookahead features. The resulting parser gives 1.3% absolute improvement in WSJ and 2.3% in CTB compared to the baseline, giving the highest reported accuracies for fully-supervised parsing. Jiangming Liu |
Trans. Assoc. Comput. Linguistics | 1 |
| 2017 | In-Order Transition-based Constituent ParsingabstractBoth bottom-up and top-down strategies have been used for neural transition-based constituent parsing. The parsing strategies differ in terms of the order in which they recognize productions in the derivation tree, where bottom-up strategies and top-down strategies take post-order and pre-order traversal over trees, respectively. Bottom-up parsers benefit from rich features from readily built partial parses, but lack lookahead guidance in the parsing process; top-down parsers benefit from non-local guidance for local decisions, but rely on a strong encoder over the input to predict a constituent hierarchy before its construction. To mitigate both issues, we propose a novel parsing system based on in-order traversal over syntactic trees, designing a set of transition actions to find a compromise between bottom-up constituent information and top-down lookahead information. Based on stack-LSTM, our psycholinguistically motivated constituent parsing system achieves 91.8 F1 on the WSJ benchmark. Furthermore, the system achieves 93.6 F1 with supervised reranking and 94.2 F1 with semi-supervised reranking, which are the best results on the WSJ benchmark. Jiangming Liu |
Trans. Assoc. Comput. Linguistics | 1 |
| 2015 | An Empirical Comparison Between N-gram and Syntactic Language Models for Word OrderingabstractSyntactic language models and N-gram language models have both been used in word ordering.In this paper, we give an empirical comparison between N-gram and syntactic language models on word order task.Our results show that the quality of automatically-parsed training data has a relatively small impact on syntactic models.Both of syntactic and N-gram models can benefit from large-scale raw text.Compared with N-gram models, syntactic models give overall better performance, but they require much more training time.In addition, the two models lead to different error distributions in word ordering.A combination of the two models integrates the advantages of each model, achieving the best result in a standard benchmark. Jiangming Liu, Yue Zhang 0004 |
EMNLP | 1 |
| 2014 | Case Frame Constraints for Hierarchical Phrase-Based Translation: Japanese-Chinese as an Example
Jiangming Liu, Jin An Xu |
NLPCC | 1 |
| 2013 | An Approach of Hybrid Hierarchical Structure for Word Similarity Computing by HowNet
Jiangming Liu, Jin An Xu |
IJCNLP | 1 |