Zhaohe Liao

dblp:263/9800 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0004-0040-8353ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 37% Video understanding and tracking · 20% Trustworthy machine learning · 16%
Databases, data mining, and information retrieval
1 paper
Spatial and temporal data management · 67% Data mining · 33%
Computer graphics and multimedia
2 papers
Multimedia systems and quality of experience · 57% Image and video processing · 43%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video question answering
2.632026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering · ICML 2025
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering · CVPR 2024
Machine learning › Trustworthy machine learning
interpretability
2.132026
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention · ACL (1) 2026
COIN-Matting: Confounder Intervention for Image Matting · ECCV (19) 2024
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
1.822026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering · CVPR 2024
Natural language and speech › Language models and text generation › large language model reasoning › inference-time reasoning
latent reasoning
1.012026
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment · WWW 2026
Natural language and speech › Question answering and dialogue systems
multimodal question answering
1.012026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Multimedia systems and quality of experience
video quality assessment
1.012026
Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment · WWW 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
0.912025
Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering · ICML 2025
Natural language and speech › Question answering and dialogue systems
answer aggregation
0.812024
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering · CVPR 2024
Spatial and temporal data management › spatial crowdsourcing
ridesharing
0.812024
Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing · ICDE 2024
Spatial and temporal data management
spatial crowdsourcing
0.812024
Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing · ICDE 2024
Data mining › crowdsourcing
worker selection
0.812024
Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing · ICDE 2024
Image and video processing
image matting
0.812024
COIN-Matting: Confounder Intervention for Image Matting · ECCV (19) 2024
Computer vision › Vision and language
video-language understanding
0.212024
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering · CVPR 2024
Graph algorithms and graph theory › graph algorithms
route planning
0.212024
Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing · ICDE 2024

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 3.0causal inference · 1.5additional travel distance prediction · 1.5intervention · 1.0interpretability analysis · 1.0graph-driven reasoning · 1.0retrieval-augmented generation · 0.9few-shot prompting · 0.9video alignment · 0.8large language model question decomposition · 0.8contrastive learning · 0.8confounder intervention · 0.8
YearPublicationVenuePosition
2026 Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
abstract
Shuochen Chang, Tong Bai, Xiaofeng Zhang, Qianli Ma, Qingyang Liu, Zhaohe Liao, Yibo Miao, Li Niu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuochen Chang, Tong Bai, Xiaofeng Zhang 0006, Qianli Ma 0008, Qingyang Liu 0008, Zhaohe Liao, Yibo Miao, Li Niu 0002
ACL (1)6
2026 Bridging Visual Dynamics and Narrative Reasoning: Multimodal Large Language Models for Short Drama Quality Assessment
Qingyang Liu 0008, Jiangtong Li, Zelin Peng, Shaobo Wang 0001, Zhaohe Liao, Shuochen Chang, Bingjie Gao, Mu Liu, Jidong Jiang, Li Niu 0002
WWW5
2026 Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering
abstract
Video Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts.
Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering
abstract
Video Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition.
Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001
ICML1
2024 Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
abstract
Despite the recent progress made in Video Question-Answering (VideoQA), these methods typically function as black-boxes, making it difficult to understand their reasoning processes and perform consistent compositional reasoning. To address these challenges, we propose a model-agnostic Video Alignment and Answer Aggregation (VA3) framework, which is capable of enhancing both compositional consistency and accuracy of existing VidQA methods by integrating video aligner and answer aggregator modules. The video aligner hierarchically selects the relevant video clips based on the question, while the answer ag-gregator deduces the answer to the question based on its sub-questions, with compositional consistency ensured by the information flow along question decomposition graph and the contrastive learning strategy. We evaluate our framework on three settings of the AGQA-Decomp dataset with three baseline methods, and propose new metrics to measure the compositional consistency of VidQA methods more comprehensively. Moreover, we propose a large language model (LLM) based automatic question decomposition pipeline to apply our framework to any VidQA dataset. We extend MSVD and NExT-QA datasets with it to evaluate our VA3framework on broader scenarios. Extensive experiments show that our framework improves both compositional consistency and accuracy of existing methods, leading to more interpretable real-world VidQA models.
Zhaohe Liao, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001
CVPR1
2024 COIN-Matting: Confounder Intervention for Image Matting
Zhaohe Liao, Jiangtong Li, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Li Niu 0002, Liqing Zhang 0001
ECCV (19)1
2024 Cross Online Ride-Sharing for Multiple-Platform Cooperations in Spatial Crowdsourcing
abstract
The last few years have seen the wide applications of ride-sharing, a transportation service that allows users to share their travel routes. A typical problem for ride-sharing is to find an optimal route for each worker to serve the dynamically arriving requests with different objectives. Previous studies focus on the route planning on a single platform. However, a single platform may have an uneven distribution of supply and demand, which causes the platform to lose requests from lack of available workers. Luckily, some ride-sharing platforms provide the same service, which enables their collaborations. The inter-platform collaborations on ride-sharing can ease the worker shortages and greatly improve the service quality, but have not been studied yet. In this paper, we propose a Cross Online Ride-sharing (CORS) problem, which allows a platform to borrow the available workers from other platforms to serve its own requests. We first design two algorithms to select the optimal available worker from other platforms, ROWS and DOWS. ROWS randomly picks an available worker, while DOWS selects the optimal worker with the minimum additional travel distance calculated based on his/er predicted destination direction. Then, we design an efficient CORS framework that embeds the proposed optimal worker selection algorithms for the CORS problem. Extensive experiments on real and synthetic datasets demonstrate the effectiveness and efficiency of our algorithms.
Yurong Cheng, Zhaohe Liao, Xiaosong Huang, Yi Yang 0032, Xiangmin Zhou, Ye Yuan 0001, Guoren Wang
ICDE2