VLDB 2026 Research / reviewers in the wild / expert
Hanmeng Liu
dblp:269/4615
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-1320-9973ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 46% Question answering and dialogue systems · 19% Knowledge representation and reasoning · 19% | |
| Theoretical computer science
1 paper |
Logic in computer science · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
1.1 | 2 | 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023 LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning · IJCAI 2020 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
1.1 | 2 | 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023 LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning · IJCAI 2020 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
1.0 | 1 | 2026 | ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect category sentiment analysis |
0.5 | 1 | 2021 | Solving Aspect Category Sentiment Analysis as a Text Generation Task · EMNLP (1) 2021 |
Computer vision › Image recognition and object detection › object detection
contextual reasoning |
0.5 | 1 | 2021 | Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts · AAAI 2021 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.5 | 1 | 2021 | Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts · AAAI 2021 |
Logic in computer science
first-order logic |
0.3 | 1 | 2026 | ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Methods — techniques the papers use, named apart from their topics
neuro-symbolic framework · 2.0knowledge distillation · 2.0chain-of-thought · 2.0pre-trained language model · 1.2natural language inference · 0.7seq2seq language model · 0.5benchmark construction · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReEfBench: Quantifying the Reasoning Efficiency of LLMsabstractTest-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscure whether performance gains stem from genuine reasoning or mere verbosity.To address this, (1) we propose a novel neuro-symbolic framework for the nonintrusive, comprehensive process-centric evaluation of reasoning grounded in First-Order Logic.(2) Through this lens, we identify four distinct behavioral prototypes and diagnose the failure modes.(3) We examine the impact of inference mode, training strategy, and model scale.Our analysis reveals that extended token generation is not a prerequisite for deep reasoning.Furthermore, we reveal critical constraints: mixing long and short CoT data in training risks premature saturation and collapse, while distillation into smaller models captures behavioral length but fails to replicate logical efficacy due to intrinsic capacity limits. Zhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu, Yue Zhang 0004 |
ACL (1) | 4 |
| 2026 | Constraint-aware spatio-temporal graph completion for mobile sensing networks
Xiulai Li, Hanmeng Liu, Xiaozhang Liu |
Knowl. Based Syst. | 4 |
| 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language UnderstandingabstractNLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research. Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Natural Language Inference in Context - Investigating Contextual Reasoning over Long TextsabstractNatural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long texts. Consisting of 8,325 expert-designed "context-hypothesis" pairs with gold labels, ConTRoL is a passage-level NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries. Hanmeng Liu, Leyang Cui, Jian Liu 0030, Yue Zhang 0004 |
AAAI | 1 |
| 2021 | Solving Aspect Category Sentiment Analysis as a Text Generation TaskabstractAspect category sentiment analysis has attracted increasing research attention.The dominant methods make use of pre-trained language models by learning effective aspect category-specific representations, and adding specific output layers to its pre-trained representation.We consider a more direct way of making use of pre-trained language models, by casting the ACSA tasks into natural language generation tasks, using natural language sentences to represent the output.Our method allows more direct use of pre-trained knowledge in seq2seq language models by directly following the task setting during pre-training.Experiments on several benchmarks show that our method gives the best reported results, having large advantages in few-shot and zero-shot settings. Jian Liu 0030, Zhiyang Teng, Leyang Cui, Hanmeng Liu, Yue Zhang 0004 |
EMNLP (1) | 4 |
| 2020 | LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical ReasoningabstractMachine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging machine reading datasets have been proposed. Though various challenges such as evidence integration and commonsense knowledge have been integrated, one of the fundamental capabilities in human reading, namely logical reasoning, is not fully investigated. We build a comprehensive dataset, named LogiQA, which is sourced from expert-written questions for testing human Logical reasoning. It consists of 8,678 QA instances, covering multiple types of deductive reasoning. Results show that state-of-the-art neural models perform by far worse than human ceiling. Our dataset can also serve as a benchmark for reinvestigating logical AI under the deep learning NLP setting. The dataset is freely available at https://github.com/lgw863/LogiQA-dataset. Jian Liu 0030, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang 0001, Yue Zhang 0004 |
IJCAI | 3 |