Caihua Liu

dblp:137/8563 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MIRNet: Integrating Constrained Graph-Based Reasoning with Pre-training for Diagnostic Medical Imaging
abstract
Automated interpretation of medical images demands robust modeling of complex visual-semantic relationships while addressing annotation scarcity, label imbalance, and clinical plausibility constraints. We introduce MIRNet (Medical Image Reasoner Network), a novel framework that integrates self-supervised pre-training with constrained graph-based reasoning. Tongue image diagnosis is a particularly challenging domain that requires fine-grained visual and semantic understanding. Our approach leverages self-supervised masked autoencoder (MAE) to learn transferable visual representations from unlabeled data; employs graph attention networks (GAT) to model label correlations through expert-defined structured graphs; enforces clinical priors via constraint-aware optimization using KL divergence and regularization losses; and mitigates imbalance using asymmetric loss (ASL) and boosting ensembles. To address annotation scarcity, we also introduce TongueAtlas-4K, a comprehensive expert-curated benchmark comprising 4,000 images annotated with 22 diagnostic labels–representing the largest public dataset in tongue analysis. Validation shows our method achieves state-of-the-art performance. While optimized for tongue diagnosis, the framework readily generalizes to broader diagnostic medical imaging tasks.
Shufeng Kong, Nuan Cui, Yihan Meng, Yuanyuan Wei 0011, Feifan Chen, Yingheng Wang, Zhuo Cai 0004, Yuzheng Li, Zibin Zheng, Caihua Liu
AAAI14
2026 scDBic: a novel deep learning-based biclustering algorithm for analyzing scRNA-seq data
abstract
MOTIVATION: Clustering single-cell RNA sequencing (scRNA-seq) data plays a vital role in the study of cellular heterogeneity. Many algorithms have been developed to cluster scRNA-seq data. However, traditional clustering algorithms often fail to capture local consistency, whereas biclustering algorithms suffer from issues such as cell loss, poor adaptability to high-dimensional data, and iterative selection challenges. RESULTS: In this paper, we introduce scDBic, a novel deep learning-based biclustering algorithm specialized for scRNA-seq data. It comprises three main steps: cell clustering with a deep autoencoder, gene clustering, and identification of key gene clusters using the reverse strategy. The key idea is that the deep autoencoder captures the main information of gene expression and the reverse strategy identifies the key genes of cell groups. Therefore, cell clustering performance can be improved. The results demonstrate that our algorithm not only discovers cell groups in scRNA-seq data but also identifies the key genes of the cell groups. Furthermore, the clustering performance of our algorithm is better than that of traditional clustering and biclustering algorithms. This novel technique can be directly applied to discover cell groups and identify key genes in cell groups. AVAILABILITY AND IMPLEMENTATION: The source code and test data are freely available at GitHub (https://github.com/Xiaoqi-Tang/scDBic) and archived on Zenodo (DOI: 10.5281/zenodo.18676401).
Xiaoqi Tang, Caihua Liu, Chaowang Lan
Bioinform.2
2026 ST-AlignTrack: Spatio-temporal alignment transformer network for multi-camera multi object tracking
Caihua Liu, Xu Qu
Pattern Recognit. Lett.1
2025 Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
abstract
Existing video captioning methods merely provide shallow or simplistic representations of object behaviors, resulting in superficial and ambiguous descriptions. However, object behavior is dynamic and complex. To comprehensively capture the essence of object behavior, we propose a dynamic action semantic-aware graph transformer. Firstly, a multi-scale temporal modeling module is designed to flexibly learn long and short-term latent action features. It not only acquires latent action features across time scales, but also considers local latent action details, enhancing the coherence and sensitiveness of latent action representations. Secondly, a visual-action semantic aware module is proposed to adaptively capture semantic representations related to object behavior, enhancing the richness and accurateness of action representations. By harnessing the collaborative efforts of these two modules, we can acquire rich behavior representations to generate human-like natural descriptions. Finally, this rich behavior representations and object representations are used to construct a temporal objects-action graph, which is fed into the graph transformer to model the complex temporal dependencies between objects and actions. To avoid adding complexity in the inference phase, the behavioral knowledge of objects is distilled into a simple network through knowledge distillation. The experimental results on MSVD and MSR-VTT datasets demonstrate that the proposed method achieves significant performance improvements across multiple metrics.
Caihua Liu, Wenjing Xue, Xia Feng
ICASSP1
2024 ILP-FORMER: Solving Integer Linear Programming with Sequence to Multi-Label Learning
abstract
Integer Linear Programming (ILP) is an essential class of combinatorial optimization problems (COPs). Its inherent NP-hardness has fostered considerable efforts towards the development of heuristic strategies. An emerging approach involves leveraging data-driven methods to automatically learn these heuristics. For example, using deep (reinforcement) learning to recurrently reoptimize an initial solution with Large Neighborhood Search (LNS) has demonstrated exceptional performance across numerous applications. A pivotal challenge within LNS lies in identifying an optimal subset of variables for reoptimization at each stage. Existing methods typically learn a policy to select a subset, either by maintaining a fixed cardinality or by decomposing the subset into independent binary decisions for each variable. However, such strategies overlook the modeling of LNS’s sequential processes and fail to explore the correlations inherent in variable selection. To overcome these shortcomings, we introduce ILP-FORMER, an innovative model that reimagines policy learning as a sequence-to-multi-label classification (MLC) problem. Our approach uniquely integrates a causal transformer encoder to capture the sequential nature of LNS. Additionally, we employ an MLC decoder with contrastive learning to exploit the correlations in variable selection. Our extensive experiments confirm that ILP-FORMER delivers state-of-the-art anytime performance on several ILP benchmarks. Furthermore, ILP-FORMER exhibits impressive generalization capabilities when dealing with larger problem instances.
Shufeng Kong, Caihua Liu, Carla P. Gomes
UAI2
2023 ScaleMix: Intra- And Inter-Layer Multiscale Feature Combination for Change Detection
abstract
Change detection (CD) aims at finding change objects from bi-temporal images, which has wide applications in different vision tasks. Previous CD methods focus more on fusing inter-layer multiscale features while ignoring the intra-layer multiscale characteristics, which hurts the integrity of change objects with different sizes. In this paper, we propose to mix intra- and inter-layer multiscale features to generate more complete change regions. To realize intra-layer multi-scale, we propose inception difference module (IDM), which employs convolutional filters with different sizes, absolute differences, and residual connections to capture intra-layer multiscale characteristics. To capture inter-layer multiscale, we propose a residual network refinement module (RNR) to fuse the features from the highest layer to the lowest layer and generate finely detailed change predictions. Our method can capture complete changes of different sizes by considering the multiscale characteristics of intra- and inter-layer simultaneously. Experiments on two benchmark datasets reveal that our method outperforms six state-of-the-art change detectors.
Qingyi Zhao, Ruofei Wang, Caihua Liu, Sihua Gao
ICASSP4
2023 Exploring Leximin Principle for Fair Core-Selecting Combinatorial Auctions: Payment Rule Design and Implementation
abstract
Core-selecting combinatorial auctions (CAs) restrict the auction result in the core such that no coalitions could improve their utilities by engaging in collusion. The minimum-revenue-core (MRC) rule is a widely used core-selecting payment rule to maximize the total utilities of all bidders. However, the MRC rule can suffer from severe unfairness since it ignores individuals' utilities. To address this limitation, we propose to explore the leximin principle to achieve fairness in core-selecting CAs since the leximin principle prefers to maximize the utility of the worst-off; the resulting bidder-leximin-optimal (BLO) payment rule is then theoretically analyzed and an effective algorithm is further provided to compute the BLO outcome. Moreover, we conduct extensive experiments to show that our algorithm returns fairer utility distributions and is faster than existing algorithms of core-selecting payment rules.
Hao Cheng 0014, Shufeng Kong, Yanchen Deng, Caihua Liu, Bo An 0001, Chong-Jun Wang
IJCAI4
2023 Disentangled Attribute Features Vision Transformer for Pedestrian Attribute Recognition
Caihua Liu, Jiaxian Guo, Sichu Chen, Xia Feng
PRCV (6)1
2023 Salient Feature Enhanced Multi-object Tracking with Soft-Sparse Attention in Transformer
Caihua Liu, Xu Qu, Xiaoyi Ma, Sichu Chen
PRCV (12)1
2023 FGPTQ-ViT: Fine-Grained Post-training Quantization for Vision Transformers
Caihua Liu, Hongyang Shi, Xinyu He 0002
PRCV (9)1
2023 A latent topic-aware network for dense video captioning
abstract
Abstract Multiple events in a long untrimmed video possess the characteristics of similarity and continuity. These characteristics can be considered as a kind of topic semantic information, which probably behaves as same sports, similar scenes, same objects etc. Inspired by this, a novel latent topic‐aware network (LTNet) is proposed in this article. The LTNet explores potential themes within videos and generates more continuous captions. Firstly, a global visual topic finder is employed to detect the similarity among events and obtain latent topic‐level features. Secondly, a latent topic‐oriented relation learner is designed to further enhance the topic‐level representations by capturing the relationship between each event and the video themes. Benefiting from the finder and the learner, the caption generator is capable of predicting more accurate and coherent descriptions. The effectiveness of our proposed method is demonstrated on ActivityNet Captions and YouCook2 datasets, where LTNet shows a relative performance of over 3.03% and 0.50% in CIDEr score respectively.
Tao Xu 0015, Yuanyuan Cui, Xinyu He 0002, Caihua Liu
IET Comput. Vis.4
2023 A Review of the State of the Art of Data Quality in Healthcare
abstract
Effective implementation of strategic data-driven health analysis initiatives is heavily dependent on the quality of the electronic medical records that serve as the foundation from which to improve clinical decisions and, in turn, the quality of care. Although there is a large body of research on the quality of healthcare data, a systematical understanding of the methods used to address the issues of data quality is missing. This study analyzes research articles in health information systems/healthcare informatics on data quality to derive a set of dimensions for understanding data quality. Issues related to each dimension are identified and methods used to address them summarized. The issues and methods can inform healthcare professionals of how to improve data practices.
Caihua Liu, Amir Talaei-Khoei, Veda C. Storey, Guo Chao Peng
J. Glob. Inf. Manag.1
2023 Valuing Your Patient's Opinion: Online Patient Reviews and Power Distance
abstract
Online reviews have a continuing impact across all industries. Even industries with highly skilled workers are affected by online reviews, despite large gaps in experience and skill between reviewers and reviewees. The authors conducted a study amongst physicians in Nevada and China to measure the perception of online patient reviews from the perspective of healthcare providers to explore whether this skill gap affected the perception of online reviews. The authors distributed and collected survey responses from over 200 physicians and used structural equation modeling techniques to evaluate the relationships. These findings show that physician perception of online patient reviews is partially mediated by power distance, direct effects exist between the relationships identified in our model, and that cross-cultural effects are present between physician responses across Nevada and China. This study expands the existing work in the field of review evaluations by operationalizing social-psychological distance into the construct of power distance within the context of healthcare.
Alan T. Yang, Caihua Liu, Amir Talaei-Khoei, Guo Chao Peng
J. Glob. Inf. Manag.2
2022 Hierarchical Multimodal Attention Network Based on Semantically Textual Guidance for Video Captioning
Caihua Liu, Xiaoyi Ma, Xinyu He 0002, Tao Xu 0015
ICONIP (3)1
2022 Deep Attentive Belief Propagation: Integrating Reasoning and Learning for Solving Constraint Optimization Problems
abstract
Belief Propagation (BP) is an important message-passing algorithm for various reasoning tasks over graphical models, including solving the Constraint Optimization Problems (COPs). It has been shown that BP can achieve state-of-the-art performance on various benchmarks by mixing old and new messages before sending the new one, i.e., damping. However, existing methods on tuning a static damping factor for BP not only is laborious but also harms their performance. Moreover, existing BP algorithms treat each variable node's neighbors equally when composing a new message, which also limits their exploration ability. To address these issues, we seamlessly integrate BP, Gated Recurrent Units (GRUs), and Graph Attention Networks (GATs) within the massage-passing framework to reason about dynamic weights and damping factors for composing new BP messages. Our model, Deep Attentive Belief Propagation (DABP), takes the factor graph and the BP messages in each iteration as the input and infers the optimal weights and damping factors through GRUs and GATs, followed by a multi-head attention layer. Furthermore, unlike existing neural-based BP variants, we propose a novel self-supervised learning algorithm for DABP with a smoothed solution cost, which does not require expensive training labels and also avoids the common out-of-distribution issue through efficient online learning. Extensive experiments show that our model significantly outperforms state-of-the-art baselines.
Yanchen Deng, Shufeng Kong, Caihua Liu, Bo An 0001
NeurIPS3
2021 A Fully Dynamic Context Guided Reasoning and Reconsidering Network for Video Captioning
Xia Feng, Xinyu He 0002, Caihua Liu
PRICAI (1)4
2021 Text-Image Retrieval With Salient Features
abstract
In recent years, deep learning has achieved remarkable results in the text-image retrieval task. However, only global image features are considered, and the vital local information is ignored. This results in a failure to match the text well. Considering that object-level image features can help the matching between text and image, this article proposes a text-image retrieval method that fuses salient image feature representation. Fusion of salient features at the object level can improve the understanding of image semantics and thus improve the performance of text-image retrieval. The experimental results show that the method proposed in the paper is comparable to the latest methods, and the recall rate of some retrieval results is better than the current work.
Xia Feng, Zhiyi Hu, Caihua Liu, Andrew W. H. Ip
J. Database Manag.3
2021 An Investigation to the Industry 4.0 Readiness of Manufacturing Enterprises: The Ongoing Problems of Information Systems Strategic Misalignment
abstract
The visions of what constitutes Industry 4.0 is an industry based on gains in efficiency and productivity enhancements supported by integrated, smart information systems. This has caused information systems strategic misalignment that present a severe barrier to national and organizational aspirations. This paper studies the readiness of manufacturing companies for Industry 4.0 by using a case study of Chinese multinational enterprise in the aluminum production sector. The research design follows a rigorous grounded theory approach, which consisted of 41 semi-structured interviews in 7 different company branches. Based on this case study, the paper proposes an IS strategic misalignment model that identifies three levels of misalignment that need to be resolved before the vision of the smart industry can be realized. Six main categories of causes and five main categories of consequences of IS strategic misalignment are presented. This study contributes to the IS alignment literature and provides important implications for the achievement of Industry 4.0 in practice.
Guo Chao Peng, Caihua Liu
J. Glob. Inf. Manag.4
2020 Context Visual Information-based Deliberation Network for Video Captioning
abstract
Video captioning automatically and accurately generates a textual description for a video. The typical methods following the encoder-decoder architecture directly utilize hidden states to predict words. Nevertheless, these methods do not amend the inaccurate hidden states before feeding those states into word prediction. This leads to a cascade of errors in generating word by word. In this paper, the context visual information-based deliberation network is proposed, abbreviated as CVI-DeINet. Its key idea is to introduce a deliberator into the encoder-decoder framework. The encoder-decoder first generates a raw hidden state sequence. Unlike the existing methods, the raw hidden state is no longer directly used for word prediction but is fed into the deliberator to generate the refined hidden state. The words are then predicted according to the refined hidden states and the contextual visual features. The results on two datasets show that the proposed method significantly outperforms the state-of-the-art methods.
Xueyong Li, Caihua Liu
ICPR3
2019 Crowd Counting via Conditional Generative Adversarial Networks
Yinong Duan, Caihua Liu
PRCV (2)4
2019 Dropout non-negative matrix factorization
Zhicheng He 0001, Jie Liu 0007, Caihua Liu, Airu Yin, Yalou Huang
Knowl. Inf. Syst.3
2017 Multi-granularity sequence labeling model for acronym expansion identification
Jie Liu 0007, Caihua Liu, Yalou Huang
Inf. Sci.2
2016 Convolutional neural random fields for action recognition
Caihua Liu, Jie Liu 0007, Zhicheng He 0001, Qinghua Hu, Yalou Huang
Pattern Recognit.1