Zhaolu Kang

dblp:371/2900 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0009-0000-1163-1615ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models
abstract
While recent multimodal models have shown progress in vision-language tasks, small-scale variants still struggle with the fine-grained temporal reasoning required for video understanding. We introduce ReasonAct, a method that enhances video reasoning in smaller models through a three-stage training process: first building a foundation with text-only reasoning, then fine-tuning on video, and finally refining with temporal-aware reinforcement learning. We build upon Temporal Group Relative Policy Optimization (T-GRPO) by incorporating temporal consistency modeling into policy optimization. We also propose a biomechanically-motivated sub-action decomposition mechanism that provides graduated rewards for constituent action phases. Through experiments on HMDB51, UCF-101, and Kinetics-400, our 3B-parameter model achieves 67.2%, 94.1%, and 78.9% accuracy respectively, demonstrating improvements of 17.9, 15.8, and 12.3 points over baselines. Ablation studies validate that our progressive training enables smaller models to achieve competitive video reasoning performance while maintaining computational efficiency.
Jiaxin Liu 0002, Zhaolu Kang
AAAI2
2026 Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
abstract
Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 10.8% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.
Shujian Gao, Songtao Jiang, Haoxiang Xia, Zhaolu Kang, Yemin Wang, Zuozhu Liu
AAAI7
2026 NeuReasoner: Towards Explainable, Controllable, and Unified Reasoning via Mixture-of-Neurons
abstract
Large Reasoning Models (LRMs) have recently achieved remarkable success in complex reasoning tasks.However, closer scrutiny reveals persistent failure modes compromising performance and cost: I) Intra-step level, marked by calculation or derivation errors; II) Inter-step level, involving oscillation and stagnation; and III) Instance level, causing maladaptive over-thinking.Existing endeavors target isolated levels without unification, while their black-box nature and reliance on RL hinder explainability and controllability.To bridge these gaps, we conduct an in-depth white-box analysis, identifying key neurons (Mixture of Neurons, MoN) and their fluctuation patterns associated with distinct failures.Building upon these insights, we propose NeuReasoner, an explainable, controllable, and unified reasoning framework driven by MoN.Technically, NeuReasoner integrates lightweight MLPs for failure detection with a special token-triggered self-correction mechanism learned via SFT.During inference, special tokens are inserted upon failure detection to actuate controllable remedial behaviors.Extensive evaluations across six benchmarks, six backbone models (8B∼70B) against nine competitive baselines, demonstrate that NeuReasoner achieves performance gains of up to 27.0% while reducing token consumption by 19.6% ∼ 63.3%.... Correct Answer...
Haonan Dong, Kehan Jiang, Haoran Ye, Zhaolu Kang, Guojie Song
ACL (1)5
2026 FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
abstract
Yujie Feng, Hao Wang, Jian Li, Xu Chu, Zhaolu Kang, Yiran Liu, Yasha Wang, Philip S. Yu, Xiao-Ming Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaolu Kang, Yasha Wang, Philip S. Yu, Xiao-Ming Wu 0003
ACL (1)5
2026 LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models
abstract
Jian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Xi Yang, Changjinli, Yonghua Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Changjin Li, Yonghua Lin
ACL (1)3
2026 FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models
abstract
Recent Large Reasoning Models (LRMs) likeDeepSeek-R1 have demonstrated remarkable success in complex reasoning tasks, exhibiting human-like patterns in exploring multiple alternative solutions.Upon closer inspection, however, we uncover a surprising phenomenon: The First is The Best, where alternative solutions are not merely suboptimal but potentially detrimental.This observation challenges widely accepted test-time scaling laws, leading us to hypothesize that errors within the reasoning path scale concurrently with test time.Through comprehensive empirical analysis, we characterize errors as a forest-structured Forest of Errors (FoE) and conclude that FoE makes the First the Best, which is underpinned by rigorous theoretical analysis.Leveraging these insights, we propose RED, a self-guided efficient reasoning framework comprising two components: I) Refining First, which suppresses FoE growth in the first solution; and II) Discarding Subs, which prunes subsequent FoE via dualconsistency.Extensive experiments across five benchmarks and six backbone models demonstrate that RED outperforms eight competitive baselines, achieving performance gains of up to 19.0% while reducing token consumption by 37.7% ∼ 70.4%.Moreover, comparative experiments on FoE metrics shed light on how RED achieves effectiveness.... ... ...
Kehan Jiang, Haonan Dong, Zhaolu Kang, Zhengzhou Zhu, Guojie Song
ACL (1)3
2026 MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
abstract
Pengxiang Zhao, Guangyi Liu, Yaozhen Liang, Weiqing He, Zhengxi Lu, WenHao Wang, Yuehao Huang, Yuxiang Chai, Zhaolu Kang, Yaxuan Guo, Hao Wang, Kexin Zhang, Liang Liu, Yong Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengxiang Zhao, Yaozhen Liang, Weiqing He, Zhengxi Lu, WenHao Wang, Yuehao Huang, Yuxiang Chai, Zhaolu Kang, Yaxuan Guo
ACL (1)9
2026 Symmetry-Aware Causal Inference for Robust Neural PDE Solvers
abstract
While neural networks are increasingly prevalent in solving partial differential equations (PDEs), the high cost of generating precise training data through numerical simulations poses a significant challenge to model efficiency and accuracy. To address this, we propose LiPS, a data augmentation framework based on Lie point symmetries. By leveraging the inherent mathematical symmetries of PDEs, LiPS generates physically consistent training samples that enrich the data distribution without violating underlying physical laws. We further introduce an innovative architecture that integrates contrastive learning with causal reasoning. During self-supervised pre-training, the model generates causal attention maps to identify regions critical for PDE evolution, which the causal reasoning engine then uses to refine latent space representations. Extensive experiments on multiple benchmarks, including Navier-Stokes and Spherical Shallow Water equations, validate the superiority of LiPS over state-of-the-art baselines. Notably, our approach achieves up to a 38% reduction in Mean Squared Error (MSE) in out-of-distribution generalization scenarios, demonstrating exceptional robustness to unseen physical environments.
Yuanming Xie, Yanzhuo Xiang, Haochen You, Nuoya Liu, Fangzhou Liu 0002, Zhaolu Kang
ICMR7
2026 What Should I Cite? A RAG Benchmark for Academic Citation Prediction
abstract
With the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation prediction aims to automatically suggest appropriate references, helping scholars navigate the expanding scientific literature. Here we present CiteRAG, the first comprehensive retrieval-augmented generation (RAG)-integrated benchmark for evaluating large language models on academic citation prediction, featuring a multi-level retrieval strategy, specialized retrievers, and generators. Our benchmark makes four core contributions: (1) We establish two instances of the citation prediction task with different granularity. Task 1 focuses on coarse-grained list-specific citation prediction, while Task 2 targets fine-grained position-specific citation prediction. To enhance these two tasks, we build a dataset containing 7,267 instances for Task 1 and 8,541 instances for Task 2, enabling comprehensive evaluation of both retrieval and generation. (2) We construct a three-level large-scale corpus with 554k papers spanning many major subfields, using an incremental pipeline. (3) We propose a multi-level hybrid RAG approach to citation prediction, fine-tuning embedding models with contrastive learning to capture complex citation relationships, paired with specialized generation models. (4) We conduct extensive experiments across state-of-the-art language models, including closed-source APIs, open-source models, and our fine-tuned generators, demonstrating the effectiveness of our framework. Our open-source toolkit enables reproducible evaluation and focuses on academic literature, providing the first comprehensive evaluation framework for citation prediction and serving as a methodological template for other scientific domains. Our source code and data are released at https://github.com/LQgdwind/CiteRAG.
Leqi Zheng, Jiajun Zhang 0012, Canzhi Chen, Chaokun Wang, Hongwei Li 0032, Yuying Li 0006, Yaoxin Mao, Shannan Yan, Zixin Song, Zhiyuan Feng, Zhaolu Kang, Zirong Chen, Hang Zhang 0032, Qiang Liu 0006, Liang Wang 0001, Ziyang Liu 0004
WWW11
2025 MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
abstract
Xiaolong Wang, Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou, Yunghwei Lai, Ziyue Wang, Yawen Wang, Kaiyu Huang, Yile Wang, Peng Li, Yang Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zhaolu Kang, Wangyuxuan Zhai, Xinyue Lou, Yunghwei Lai
EMNLP2
2025 JurisCTC: Enhancing Legal Judgment Prediction via Cross-Domain Transfer and Contrastive Learning
abstract
In recent years, Unsupervised Domain Adaptation (UDA) has gained significant attention in the field of Natural Language Processing (NLP) owing to its ability to enhance model generalization across diverse domains. However, its application for knowledge transfer between distinct legal domains remains largely unexplored. To address the challenges posed by lengthy and complex legal texts and the limited availability of large-scale annotated datasets, we propose JurisCTC, a novel model designed to improve the accuracy of Legal Judgment Prediction (LJP) tasks. Unlike existing approaches, JurisCTC facilitates effective knowledge transfer across various legal domains and employs contrastive learning to distinguish samples from different domains. Specifically, for the LJP task, we enable knowledge transfer between civil and criminal law domains. Compared to other models and specific large language models (LLMs), JurisCTC demonstrates notable advancements, achieving peak accuracies of 76.59% and 78.83%, respectively.1
Zhaolu Kang, Hongtian Cai, Xiangyang Ji, Jinzhe Li, Nanfei Gu
IJCNN1
2024 CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models
abstract
Fuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li, Xiaolong Wang, Siyu Wang, Ziyue Wang, Xiaoyue Mi, Peng Li, Ning Ma, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fuwen Luo, Chi Chen 0005, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li 0009, Xiaolong Wang 0014, Ziyue Wang 0002, Xiaoyue Mi, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)4