Zehua Liu

dblp:l/ZehuaLiu · DBLP profile ↗
← Back
30ranked-venue papers
13as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepOR: A Deep Reasoning Foundation Model for Optimization Modeling
abstract
Optimization modeling plays a critical role in supporting optimal decision-making across various domains. Previous works have demonstrated that large language models (LLMs) tailored for optimization modeling have significantly automated and simplified this process. However, these models typically employ a straightforward input-output paradigm and struggle with challenging instances. In contrast, recent advances in general-purpose reasoning LLMs (RLLMs), such as DeepSeek-R1, have shown impressive capabilities in complex domains like mathematics and coding. In this paper, we introduce DeepOR, the first RLLM specifically designed for optimization modeling. Instead of directly outputting solutions, DeepOR explicitly performs multiple intermediate reasoning steps. To adapt a base LLM into an RLLM, we begin by synthesizing long chain-of-thought (CoT) data guided by a flowchart, which is automatically generated using a self-exploration algorithm. Once the training data are prepared, we employ supervised fine-tuning on the base LLM to endow it with reasoning capabilities tailored for optimization modeling. To fully leverage the model's reasoning potential, we further apply reinforcement learning with reward-shaping derived from solver feedback. Experimental results on benchmarks confirm that DeepOR consistently and significantly outperforms existing state-of-the-art approaches.
Ziyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan, Jingyan Zhu, Jingrong Xie 0001, Lilin Xu, Han Wu 0004, Wing Yin Yu, Zehua Liu, Xiaojin Fu, Gang Chen 0001, Dongxiang Zhang
AAAI10
2026 Elevating descriptive excellence: Object-centric dense video captioning
Zehua Liu, Huicheng Zheng, Yun Lan
Neurocomputing1
2026 Integrating frequency-aware mamba with diffusion for 4D volumetric image synthesis
Yangyang Shi, Beiji Zou 0001, Xiaonian Deng, Yucong Zhang, Zehua Liu, Xiaoyan Kui, Weixin Si
Pattern Recognit.5
2025 BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving
abstract
Teng Wang, Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu, Xiongwei Han, Wei Shi, Ruifeng She, Fangzhou Zhu, Tao Zhong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu 0004, Xiongwei Han, Ruifeng She, Fangzhou Zhu, Tao Zhong 0004
ACL (1)4
2025 Fast Intra-Operative Angiographic Parametric Imaging for Surgical Outcome Prediction of Interventional Cerebral Aneurysms Treatment
abstract
Endovascular intervention of cerebral aneurysms remains challenging due to the lack of intra-operative indicators that reflect hemodynamic alterations following stent implantation. This limitation makes it difficult to predict surgical outcomes and increases dependence on surgeon experience. In this study, we propose a fast and practical approach for converting intraoperative digital subtraction angiography (DSA) images into angiographic parametric imaging (API), enabling timely visualization of cerebral blood flow during surgery. We develop a stylusbased regions of interest (ROI) analysis interface. The proposed method has been validated against multiple reference standards, including CT perfusion (CTP), computational fluid dynamics (CFD), and 4D flow MRI, demonstrating strong consistency in hemodynamic quantification. Notably, unlike traditional methods that often take several hours, the proposed method has an average processing time of only 4.47 minutes, highlighting its suitability for intra-operative use. Furthermore, we designed a qualitative assessment strategy based on pre- and intra-operative ROI comparison to predict complication risk. The proposed method achieved an 87.5 % accuracy in predicting post-operative complications across 40 clinical cases, demonstrating particularly high sensitivity in detecting ischemia risk. These results demonstrate that the proposed method provides a reliable, efficient, and interpretable solution for intra-operative hemodynamic assessment, with strong potential for integration into clinical workflows. Code and data are available at: https://github.com/LZH970328/API.git.
Zehua Liu, Jianping Lv, Weixin Si
BIBM1
2025 TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular Registration
abstract
Transarterial chemoembolization (TACE) is a preferred treatment option for hepatocellular carcinoma and other liver malignancies, yet it remains a highly challenging procedure due to complex intra-operative vascular navigation and anatomical variability. Accurate and robust 2D-3D vessel registration is essential to guide microcatheter and instruments during TACE, enabling precise localization of vascular structures and optimal therapeutic targeting. To tackle this issue, we develop a coarse-to-fine registration strategy. First, we introduce a global alignment module, structure-aware perspective n-point (SA-PnP), to establish correspondence between 2D and 3D vessel structures. Second, we propose TempDiffReg, a temporal diffusion model that performs vessel deformation iteratively by leveraging temporal context to capture complex anatomical variations and local structural changes. We collected data from 23 patients and constructed 626 paired multi-frame samples for comprehensive evaluation. Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art (SOTA) methods in both accuracy and anatomical plausibility. Specifically, our method achieves a mean squared error (MSE) of 0.63 mm and a mean absolute error (MAE) of 0.51 mm in registration accuracy, representing$66.7\%$lower MSE and$17.7\%$lower MAE compared to the most competitive existing approaches. It has the potential to assist less-experienced clinicians in safely and efficiently performing complex TACE procedures, ultimately enhancing both surgical outcomes and patient care. Code and data are available at: https://github.com/LZH970328/TempDiffReg.git
Zehua Liu, Shihao Zou, Jincai Huang 0003, Weixin Si
BIBM1
2025 Decoupling Training-Free Guided Diffusion by ADMM
abstract
In this paper, we consider the conditional generation problem by guiding off-the-shelf unconditional diffusion models with differentiable loss functions in a plug-and-play fashion. While previous research has primarily focused on balancing the unconditional diffusion model and the guided loss through a tuned weight hyperparameter, we propose a novel framework that distinctly decouples these two components. Specifically, we introduce two variables x and z, to represent the generated samples governed by the unconditional generation model and the guidance function, respectively. This decoupling reformulates conditional generation into two manageable subproblems, unified by the constraint x = z. Leveraging this setup, we develop a new algorithm based on the Alternating Direction Method of Multipliers (ADMM) to adaptively balance these components. Additionally, we establish the equivalence between the diffusion reverse step and the proximal operator of ADMM and provide a detailed convergence analysis of our algorithm under certain mild assumptions. Our experiments demonstrate that our proposed method ADMMDiff consistently generates high-quality samples while ensuring strong adherence to the conditioning criteria. It outperforms existing methods across a range of conditional generation tasks, including image generation with various guidance and controllable motion synthesis.
Youyuan Zhang, Zehua Liu, Zenan Li, James J. Clark, Xujie Si
CVPR2
2025 CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
Zehua Liu, Xiaolou Li, Chen Chen 0075, Lantian Li, Dong Wang 0013
INTERSPEECH1
2025 A Blockchain-Based Verifiable Data Quality Assessment Scheme
abstract
This paper presents a blockchain-based scheme to address the transparency and verifiability issues in data standard verification caused by external service providers or server-dominated evaluation processes. The scheme includes several components: task publishing and strategy deployment algorithms, low-quality data user detection and reliable verification algorithms, adaptive privacy-preserving intersection verification algorithms, and task-related data homogeneity and content diversity evaluation verification algorithms. This scheme ensures secure data storage and access, while also enabling public verification of evaluation results. Empirical outcomes indicate that the method conforms to practical requirements in terms of productivity and expense.
Keqi Xiong, Zehua Liu, Jiayong Wei, Huimin Gong
IWCMC2
2025 Activation-Guided Consensus Merging for Large Language Models
abstract
Recent research has increasingly focused on reconciling the reasoning capabilities of System 2 with the efficiency of System 1. While existing training-based and prompt-based approaches face significant challenges in terms of efficiency and stability, model merging emerges as a promising strategy to integrate the diverse capabilities of different Large Language Models (LLMs) into a unified model. However, conventional model merging methods often assume uniform importance across layers, overlooking the functional heterogeneity inherent in neural components. To address this limitation, we propose \textbf{A}ctivation-Guided \textbf{C}onsensus \textbf{M}erging (\textbf{ACM}), a plug-and-play merging framework that determines layer-specific merging coefficients based on mutual information between activations of pre-trained and fine-tuned models. ACM effectively preserves task-specific capabilities without requiring gradient computations or additional training. Extensive experiments on Long-to-Short (L2S) and general merging tasks demonstrate that ACM consistently outperforms all baseline methods. For instance, in the case of Qwen-7B models, TIES-Merging equipped with ACM achieves a \textbf{55.3\%} reduction in response length while simultaneously improving reasoning accuracy by \textbf{1.3} points. We submit the code with the paper for reproducibility, and it will be publicly available.
Shuqi Liu 0001, Zehua Liu, Qintong Li, Xiongwei Han, Zhijiang Guo, Han Wu 0004, Linqi Song
NeurIPS3
2025 Cost-effective Tangible Rehearsal Interface for Microsurgical Clipping of Intracranial Aneurysm
abstract
Microsurgical clipping (MC) is widely used for the treatment of intracranial aneurysms (IA). However, it is a high risk for neurosurgeons to perform this operation due to the complex intracranial anatomy, limited microscopic view, and confined operational space. To tackle the above issues, meticulous preoperative rehearsal is in urgent need to improve the neurosurgeons’ perception of patient-specific anatomy while designing the optimal surgical plan, which can greatly enhance the patients’ safety. Most existing MC simulators only support limited interaction methods, such as haptic devices, leading to substantial differences between the training experience and actual surgical situations. To this end, we present a mixed reality (MR) interface for IA MC rehearsal, which can provide neurosurgeons with a more immersive experience. Firstly, considering the labor-intensive labeling cost of reconstructing a 3D full-brain vascular model from CTA images, the simulator employs a cost-effective specific-to-general vessel stitching technique to generate 3D lesion-specific IA geometric models, which replaces normal vascular segments in a standard brain with a personalized-specific operating region by a scaling-constrained iterative closest point (ICP) algorithm. Besides, we design a marker-based tracking method allowing accessible natural human-computer interaction using real surgical instruments which can fuse virtual anatomy and real operation environments, enhancing the users’ spatial perception and tangible stimuli. Additionally, to ensure simulation stability while providing visually plausible clipping operations, we develop a collision distance constrained position-based dynamics (PBD) method with low-resolution sampling particles to simulate the deformation of aneurysm vessels. Quantitative experiments demonstrate the accuracy of our surgical instruments tracking and vessel deformation simulation, which can also fulfill the real-time performance of MC rehearsal. User study indicates that virtual rehearsals significantly improve spatial awareness and dexterity in handling aneurysms, and have the great potential to be applied in practical applications.
Wei Cao 0008, Zehua Liu, Jianping Lv, Weixin Si
VR4
2024 Full Bayesian Significance Testing for Neural Networks
abstract
Significance testing aims to determine whether a proposition about the population distribution is the truth or not given observations. However, traditional significance testing often needs to derive the distribution of the testing statistic, failing to deal with complex nonlinear relationships. In this paper, we propose to conduct Full Bayesian Significance Testing for neural networks, called nFBST, to overcome the limitation in relationship characterization of traditional approaches. A Bayesian neural network is utilized to fit the nonlinear and multi-dimensional relationships with small errors and avoid hard theoretical derivation by computing the evidence value. Besides, nFBST can test not only global significance but also local and instance-wise significance, which previous testing methods don't focus on. Moreover, nFBST is a general framework that can be extended based on the measures selected, such as Grad-nFBST, LRP-nFBST, DeepLIFT-nFBST, LIME-nFBST. A range of experiments on both simulated and real data are conducted to show the advantages of our method.
Zehua Liu, Zimeng Li 0002, Jingyuan Wang 0001
AAAI1
2024 Full Bayesian Significance Testing for Neural Networks in Traffic Forecasting
Zehua Liu, Jingyuan Wang 0001, Zimeng Li 0002
IJCAI1
2024 CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
Chen Chen 0075, Zehua Liu, Xiaolou Li, Lantian Li, Dong Wang 0013
INTERSPEECH2
2024 Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, Chen Chen 0075, Lantian Li, Li Guo 0004, Dong Wang 0013
INTERSPEECH2
2023 A Method for Identifying the Timeliness of Manufacturing Data Based on Weighted Timeliness Graph
Zehua Liu, Xuefeng Ding 0002, Yuming Jiang 0004, Dasha Hu
ADMA (1)1
2023 Learning with Logical Constraints but without Shortcut Satisfaction
Zenan Li, Zehua Liu, Yuan Yao 0001, Jingwei Xu 0001, Taolue Chen 0001, Xiaoxing Ma, Jian Lu 0001
ICLR2
2023 HDRNet: High-Dimensional Regression Network for Point Cloud Registration
abstract
Abstract Abstract‐3D point cloud registration is a crucial topic in the reverse engineering, computer vision and robotics fields. The core of this problem is to estimate a transformation matrix for aligning the source point cloud with a target point cloud. Several learning‐based methods have achieved a high performance. However, they are challenged with both partial overlap point clouds and multiscale point clouds, since they use the singular value decomposition (SVD) to find the rotation matrix without fully considering the scale information. Furthermore, previous networks cannot effectively handle the point clouds having large initial rotation angles, which is a common practical case. To address these problems, this paper presents a learning‐based point cloud registration network, namely HDRNet, which consists of four stages: local feature extraction, correspondence matrix estimation, feature embedding and fusion and parametric regression. HDRNet is robust to noise and large rotation angles, and can effectively handle the partial overlap and multi‐scale point clouds registration. The proposed model is trained on the ModelNet40 dataset, and compared with ICP, SICP, FGR and recent learning‐based methods (PCRNet, IDAM, RGMNet and GMCNet) under several settings, including its performance on moving to invisible objects, with higher success rates. To verify the effectiveness and generality of our model, we also further tested our model on the Stanford 3D scanning repository.
Zehua Liu
Comput. Graph. Forum3
2022 A Worker Selection Scheme for Vehicle Crowdsourcing Blockchain
abstract
With the improvement of the computing power and storage capacity of vehicular equipment, as well as the growing privacy concerns over sharing sensitive raw data, federated learning can be a promising solution for realizing distributed vehicular crowdsourcing services. The global learning model can be used as crowdsourcing tasks and assigned to vehicles that utilize their local training models. In such a distributed scenario, it is essential to ensure the completion of high-quality training tasks and the security of sensitive data. We propose a vehicular crowdsourcing blockchain to achieve secure reputation management of vehicles in a distributed manner, providing a safe and trusted solution for vehicle crowdsourcing services. We design a worker selection scheme, which combines the reputation of workers with the amount of data as an indicator of worker reliability. We further propose an effective incentive machine to enable more workers with high reputations to participate in crowdsourcing tasks and to contribute higher quality local data. Simulation results show that the worker selection scheme improves the matching rate by 10% and significantly improves the total utility. The proposed scheme is deployed on the IBM Hyperledger Fabric platform to observe its real-world running time and overall performance.
Xinran Ma, Shulin Sun, Zehua Liu
ICSS3
2022 VFMVAC: View-filtering-based multi-view aggregating convolution for 3D shape recognition and retrieval
Zehua Liu
Pattern Recognit.1
2021 From Coarse to Fine: Hierarchical Multi-scale Temporal Information Modeling via Sub-group Convolution for Video Action Recognition
abstract
In the video action recognition task, it is essential to model the temporal information. Since different actions have different durations, capturing multi-scale temporal features is very crucial. In this paper, we propose a multi-scale modeling (MSM) module to exploit temporal information for action recognition, which is composed of a multi-scale temporal convolution (MTC) block and a multi-scale hierarchical convolution (MHC) block. MTC uses convolutions of multiple temporal depths to capture features of different temporal scales to enhance the connection between frames. MHC employs group convolution to obtain more fine-grained multi-scale features in the channel dimension. In MHC, the convolutions are formulated as a hierarchical structure to expand the receptive fields as well as help implement a series of sub-group convolutions, which can help to realize the modeling of distinctive and long-term information. The two components of MSM are complementary in temporal modeling. Finally, we evaluated our method on several action recognition benchmarks including Kinetics, UCF10l, and HMDB51, and obtained competitive results, which verified the effectiveness of the proposed method in temporal modeling,
Fengwen Cheng, Huicheng Zheng, Zehua Liu
IJCNN3
2021 Dense Video Captioning with Hierarchical Attention-Based Encoder-Decoder Networks
abstract
Dense video captioning is a challenging task with the goal of localizing and describing all events in an untrimmed video, taking into account both visual and text information. Although existing methods have made some achievements, most of them suffer from missing details and inferior captioning. Recent progress has been made in using object features to supplement more detailed information. However, due to the considerable number of objects in the video, the representation of learning objects is often noisy, which may interfere with the generation of correct captions. We also notice that realworld video-text data involve different granularity levels, such as objects/words and events/sentences. Therefore, we propose the hierarchical video-text attention-based encoder-decoder networks for dense video captioning. The proposed method successfully considers the hierarchy in the video and text and exploits the most relevant visual and text features when generating caption. Specially, we design a hierarchical attention encoder for learning complex visual information: an object attention module focusing on the most relevant objects and an event attention module modeling the long-range temporal context. A corresponding decoder has been built for translating multi-level features into the linguistic description, i.e., a word attention module to exploit the most correlated text features and a sentence attention module to leverage high-level semantic information. The proposed hierarchical attention mechanism achieves state-of-the-art performance on the ActivityNet Captions dataset.
Mingjing Yu, Huicheng Zheng, Zehua Liu
IJCNN3
2005 On organizing and accessing geospatial and georeferenced Web resources using the G-Portal system
abstract
In order to organise and manage geospatial and georeferenced information on the Web making them convenient for searching and browsing, a digital portal known as G-Portal has been designed and implemented. Compared to other digital libraries, G-Portal is unique for several of its features. It maintains metadata resources in XML with flexible resource schemas. Logical groupings of metadata resources as projects and layers are possible to allow the entire metadata collection to be partitioned differently for users with different information needs. These metadata resources can be displayed in both the classification-based and map-based interfaces provided by G-Portal. G-Portal further incorporates both a query module and an annotation module for users to search metadata and to create additional knowledge for sharing respectively. G-Portal also includes a resource classification module that categorizes resources into one or more hierarchical category trees based on user-defined classification schemas. This paper gives an overview of the G-Portal design and implementation. The portal features will be illustrated using a collection of high school geography examination-related resources.
Ee-Peng Lim, Zehua Liu, Dion Hoe-Lian Goh, Yin Leng Theng, Wee Keong Ng
Inf. Process. Manag.2
2005 Applying scenario-based design and claims analysis to the design of a digital library of geography examination resources
Yin Leng Theng, Dion Hoe-Lian Goh, Ee-Peng Lim, Zehua Liu, Natalie Lee-San Pang, Patricia Bao-Bao Wong
Inf. Process. Manag.4
2004 Unloading Unwanted Information: From Physical Websites to Personalized Web Views
Zehua Liu, Wee Keong Ng, Ee-Peng Lim
APWeb1
2004 An Automated Algorithm for Extracting Website Skeleton
Zehua Liu, Wee Keong Ng, Ee-Peng Lim
DASFAA1
2004 Regional Differentiation of Chinese Tourism Web sites
Jie Zhang 0022, Shufei Lu, Minghua Wen, Zehua Liu, Yun-xia Feng
ENTER4
2004 Towards building logical views of websites
Zehua Liu, Wee Keong Ng, Ee-Peng Lim, Feifei Li 0001
Data Knowl. Eng.1
2004 A Java-based digital library portal for geography education
abstract
G-Portal is a Java-based digital library system for managing the metadata of geography related resources on the Web. In addition to providing a flexible repository subsystem to accommodate metadata of different formats using XML and XML Schemas, G-Portal organizes metadata into projects and layers, and supports an integrated and synchronized classification and map-based interfaces over the stored metadata. G-Portal also includes a classification subsystem that creates category structures and classifies metadata resources into categories based on user-specified classification schemas. Furthermore, G-Portal users can annotate resources and make their annotations available to others. In this paper, we describe the design and implementation of G-Portal and elaborate how Java is used to implement its features. G-Portal has been designed to be modular and some of the modules can be used as stand-alone tools. In this paper, we use UML notation to describe the detailed design of G-Portal and highlight some of the design decisions.
Zehua Liu, Ee-Peng Lim, Dion Hoe-Lian Goh, Yin Leng Theng, Wee Keong Ng
Sci. Comput. Program.1
2002 Wiccap Data Model: Mapping Physical Websites to Logical Views
Zehua Liu, Feifei Li 0001, Wee Keong Ng
ER1