Hanmeng Liu

dblp:269/4615 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-1320-9973ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 46% Question answering and dialogue systems · 19% Knowledge representation and reasoning · 19%
Theoretical computer science
1 paper
Logic in computer science · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning
1.122023
LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning · IJCAI 2020
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
1.122023
LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning · IJCAI 2020
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
1.012026
ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › sentiment analysis › aspect-based sentiment analysis
aspect category sentiment analysis
0.512021
Solving Aspect Category Sentiment Analysis as a Text Generation Task · EMNLP (1) 2021
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.512021
Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts · AAAI 2021
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference
0.512021
Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts · AAAI 2021
Logic in computer science
first-order logic
0.312026
ReEfBench: Quantifying the Reasoning Efficiency of LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation
pre-trained language model
0.212023
LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding · IEEE ACM Trans. Audio Speech Lang. Process. 2023

Methods — techniques the papers use, named apart from their topics

neuro-symbolic framework · 2.0knowledge distillation · 2.0chain-of-thought · 2.0pre-trained language model · 1.2natural language inference · 0.7seq2seq language model · 0.5benchmark construction · 0.5
YearPublicationVenuePosition
2026 ReEfBench: Quantifying the Reasoning Efficiency of LLMs
abstract
Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscure whether performance gains stem from genuine reasoning or mere verbosity.To address this, (1) we propose a novel neuro-symbolic framework for the nonintrusive, comprehensive process-centric evaluation of reasoning grounded in First-Order Logic.(2) Through this lens, we identify four distinct behavioral prototypes and diagnose the failure modes.(3) We examine the impact of inference mode, training strategy, and model scale.Our analysis reveals that extended token generation is not a prerequisite for deep reasoning.Furthermore, we reveal critical constraints: mixing long and short CoT data in training risks premature saturation and collapse, while distillation into smaller models captures behavioral length but fails to replicate logical efficacy due to intrinsic capacity limits.
Zhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu, Yue Zhang 0004
ACL (1)4
2026 Constraint-aware spatio-temporal graph completion for mobile sensing networks
Xiulai Li, Hanmeng Liu, Xiaozhang Liu
Knowl. Based Syst.4
2023 LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding
abstract
NLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research.
Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Natural Language Inference in Context - Investigating Contextual Reasoning over Long Texts
abstract
Natural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long texts. Consisting of 8,325 expert-designed "context-hypothesis" pairs with gold labels, ConTRoL is a passage-level NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries.
Hanmeng Liu, Leyang Cui, Jian Liu 0030, Yue Zhang 0004
AAAI1
2021 Solving Aspect Category Sentiment Analysis as a Text Generation Task
abstract
Aspect category sentiment analysis has attracted increasing research attention.The dominant methods make use of pre-trained language models by learning effective aspect category-specific representations, and adding specific output layers to its pre-trained representation.We consider a more direct way of making use of pre-trained language models, by casting the ACSA tasks into natural language generation tasks, using natural language sentences to represent the output.Our method allows more direct use of pre-trained knowledge in seq2seq language models by directly following the task setting during pre-training.Experiments on several benchmarks show that our method gives the best reported results, having large advantages in few-shot and zero-shot settings.
Jian Liu 0030, Zhiyang Teng, Leyang Cui, Hanmeng Liu, Yue Zhang 0004
EMNLP (1)4
2020 LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
abstract
Machine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging machine reading datasets have been proposed. Though various challenges such as evidence integration and commonsense knowledge have been integrated, one of the fundamental capabilities in human reading, namely logical reasoning, is not fully investigated. We build a comprehensive dataset, named LogiQA, which is sourced from expert-written questions for testing human Logical reasoning. It consists of 8,678 QA instances, covering multiple types of deductive reasoning. Results show that state-of-the-art neural models perform by far worse than human ceiling. Our dataset can also serve as a benchmark for reinvestigating logical AI under the deep learning NLP setting. The dataset is freely available at https://github.com/lgw863/LogiQA-dataset.
Jian Liu 0030, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang 0001, Yue Zhang 0004
IJCAI3