Xiaohao Cai

dblp:16/10261 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-0924-2834ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiFi-Mamba: Dual-Stream ?-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
abstract
Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision tasks offer promising long-range modeling capabilities with linear-time complexity, their direct application to MRI reconstruction inherits two key limitations: (1) insensitivity to high-frequency anatomical details; and (2) reliance on redundant multi-directional scanning. To address these limitations, we introduce High-Fidelity Mamba (HiFi-Mamba), a novel dual-stream Mamba-based architecture comprising stacked ?-Laplacian (WL) and HiFi-Mamba blocks. Specifically, the WL block performs fidelity-preserving spectral decoupling, producing complementary low- and high-frequency streams. This separation enables the HiFi-Mamba block to focus on low-frequency structures, enhancing global feature modeling. Concurrently, the HiFi-Mamba block selectively integrates high-frequency features through adaptive state-space modulation, preserving comprehensive spectral details. To eliminate the scanning redundancy, the HiFi-Mamba block adopts a streamlined unidirectional traversal strategy that preserves long-range modeling capability with improved computational efficiency. Extensive experiments on standard MRI reconstruction benchmarks demonstrate that HiFi-Mamba consistently outperforms state-of-the-art CNN-based, Transformer-based, and other Mamba-based models in reconstruction accuracy while maintaining a compact and efficient model design.
Pengcheng Fang, Yuxia Chen, Yingxuan Ren, Fangfang Tang, Xiaohao Cai, Shanshan Shan, Feng Liu 0005
AAAI7
2026 MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation
abstract
Recent advances in transformer-based text-to-motion generation have significantly improved motion quality. However, achieving both real-time performance and long-horizon scalability remains an open challenge. In this paper, we present MOGO (Motion Generation with One-pass), a novel autoregressive framework for efficient and scalable 3D human motion generation. MOGO consists of two key components. First, we introduce MoSA-VQ, a motion scale-adaptive residual vector quantization module that hierarchically discretizes motion sequences through learnable scaling parameters, enabling dynamic allocation of representation capacity and producing compact yet expressive multi-level representations. Second, we design the RQHC-Transformer, a residual quantized hierarchical causal transformer that decodes motion tokens in a single forward pass. Each transformer block aligns with one quantization level, allowing hierarchical abstraction and temporally coherent generation with strong semantic flow. Compared to diffusion- and LLM-based approaches, MOGO achieves lower inference latency while preserving high motion fidelity. Moreover, its hierarchical latent design enables seamless and controllable infinite-length motion generation, with stable transitions and the ability to adaptively incorporate updated control signals at arbitrary points in time. To further enhance generalization and interpretability, we introduce Textual Condition Alignment (TCA), which leverages large language models with Chain-of-Thought reasoning to bridge the gap between real-world prompts and training data. TCA not only improves zero-shot performance on unseen datasets but also enriches motion comprehension for in-distribution prompts through explicit intent decomposition. Extensive experiments on HumanML3D, KIT-ML, and the unseen CMP dataset demonstrate that MOGO outperforms prior methods in generation quality, inference efficiency, and temporal scalability.
Tengjiao Sun, Pengcheng Fang, Xiaohao Cai, Hansung Kim 0001
AAAI4
2025 EmoPerso: Enhancing Personality Detection with Self-Supervised Emotion-Aware Modelling
abstract
Personality detection from text is commonly performed by analysing users' social media posts. However, existing methods heavily rely on large-scale annotated datasets, making it challenging to obtain high-quality personality labels. Moreover, most studies treat emotion and personality as independent variables, overlooking their interactions. In this paper, we propose a novel self-supervised framework, EmoPerso, which improves personality detection through emotion-aware modelling. EmoPerso first leverages generative mechanisms for synthetic data augmentation and rich representation learning. It then extracts pseudo-labeled emotion features and jointly optimizes them with personality prediction via multi-task learning. A cross-attention module is employed to capture fine-grained interactions between personality traits and the inferred emotional representations. To further refine relational reasoning, EmoPerso adopts a self-taught strategy to enhance the model's reasoning capabilities iteratively. Extensive experiments on two benchmark datasets demonstrate that EmoPerso surpasses state-of-the-art models. The source code is available at https://github.com/slz0925/EmoPerso.
Lingzhi Shen, Xiaohao Cai, Muhammad Imran Razzak, Guanming Chen, Shoaib Jameel
CIKM2
2025 LL4G: Self-Supervised Dynamic Optimization for Graph-Based Personality Detection
abstract
Graph-based personality detection constructs graph structures from textual data, particularly social media posts. Current methods often struggle with sparse or noisy data and rely on static graphs, limiting their ability to capture dynamic changes between nodes and relationships. This paper introduces LL4G, a self-supervised framework leveraging large language models (LLMs) to optimize graph neural networks (GNNs). LLMs extract rich semantic features to generate node representations and to infer explicit and implicit relationships. The graph structure adaptively adds nodes and edges based on input data, continuously optimizing itself. The GNN then uses these optimized representations for joint training on node reconstruction, edge prediction, and contrastive learning tasks. This integration of semantic and structural information generates robust personality profiles. Experimental results on Kaggle and Pandora datasets show LL4G outperforms state-of-the-art models.
Lingzhi Shen, Xiaohao Cai, Guanming Chen, Muhammad Imran Razzak, Shoaib Jameel
ICME3
2025 Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
abstract
Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar.
Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001
ICRA6
2025 Less but Better: Parameter-Efficient Fine-Tuning of Large Language Models for Personality Detection
abstract
Personality detection automatically identifies an individual’s personality from various data sources, such as social media texts. However, as the parameter scale of language models continues to grow, the computational cost becomes increasingly difficult to manage. Fine-tuning also grows more complex, making it harder to justify the effort and reliably predict outcomes. We introduce a novel parameter-efficient fine-tuning framework, PersLLM, to address these challenges. In PersLLM, a large language model (LLM) extracts high-dimensional representations from raw data and stores them in a dynamic memory layer. PersLLM then updates the downstream layers with a replaceable output network, enabling flexible adaptation to various personality detection scenarios. By storing the features in the memory layer, we eliminate the need for repeated complex computations by the LLM. Meanwhile, the lightweight output network serves as a proxy for evaluating the overall effectiveness of the framework, improving the predictability of results. Experimental results on key benchmark datasets like Kaggle and Pandora show that PersLLM significantly reduces computational cost while maintaining competitive performance and strong adaptability.
Lingzhi Shen, Xiaohao Cai, Guanming Chen, Muhammad Imran Razzak, Shoaib Jameel
IJCNN3
2025 tCURLoRA: Tensor CUR Decomposition Based Low-Rank Parameter Adaptation and Its Application in Medical Image Segmentation
Guanghua He, Wangang Cheng, Hancan Zhu, Xiaohao Cai, Gaohang Yu
MICCAI (16)4
2025 CALM: Culturally Self-Aware Language Models
abstract
Cultural awareness in language models is the capacity to understand and adapt to diverse cultural contexts. However, most existing approaches treat culture as static background knowledge, overlooking its dynamic and evolving nature. This limitation reduces their reliability in downstream tasks that demand genuine cultural sensitivity. In this work, we introduce CALM, a novel framework designed to endow language models with cultural self-awareness. CALM disentangles task semantics from explicit cultural concepts and latent cultural signals, shaping them into structured cultural clusters through contrastive learning. These clusters are then aligned via cross-attention to establish fine-grained interactions among related cultural features and are adaptively integrated through a Mixture-of-Experts mechanism along culture-specific dimensions. The resulting unified representation is fused with the model's original knowledge to construct a culturally grounded internal identity state, which is further enhanced through self-prompted reflective learning, enabling continual adaptation and self-correction. Extensive experiments conducted on multiple cross-cultural benchmark datasets demonstrate that CALM consistently outperforms state-of-the-art methods.
Lingzhi Shen, Xiaohao Cai, Muhammad Imran Razzak, Guanming Chen, Shoaib Jameel
NeurIPS2
2025 GAMED: Knowledge Adaptive Multi-Experts Decoupling for Multimodal Fake News Detection
abstract
Multimodal fake news detection often involves modelling heterogeneous data sources, such as vision and language. Existing detection methods typically rely on fusion effectiveness and cross-modal consistency to model the content, complicating understanding how each modality affects prediction accuracy. Additionally, these methods are primarily based on static feature modelling, making it difficult to adapt to the dynamic changes and relationships between different data modalities. This paper develops a significantly novel approach, GAMED, for multimodal modelling, which focuses on generating distinctive and discriminative features through modal decoupling to enhance cross-modal synergies, thereby optimizing overall performance in the detection process. GAMED leverages multiple parallel expert networks to refine features and pre-embed semantic knowledge to improve the experts' ability in information selection and viewpoint sharing. Subsequently, the feature distribution of each modality is adaptively adjusted based on the respective experts' opinions. GAMED also introduces a novel classification technique to dynamically manage contributions from different modalities, while improving the explainability of decisions. Experimental results on the Fakeddit and Yang datasets demonstrate that GAMED performs better than recently developed state-of-the-art models. The source code can be accessed at https://github.com/slz0925/GAMED.
Lingzhi Shen, Xiaohao Cai, Muhammad Imran Razzak, Guanming Chen, Shoaib Jameel
WSDM3
2024 Detect Closer Surfaces That Can be Seen: New Modeling and Evaluation in Cross-Domain 3D Object Detection
abstract
The performance of domain adaptation technologies has not yet reached an ideal level in the current 3D object detection field for autonomous driving, which is mainly due to significant differences in the size of vehicles, as well as the environments they operate in when applied across domains. These factors together hinder the effective transfer and application of knowledge learned from specific datasets. Since the existing evaluation metrics are initially designed for evaluation on a single domain by calculating the 2D or 3D overlap between the prediction and ground-truth bounding boxes, they often suffer from the overfitting problem caused by the size differences among datasets. This raises a fundamental question related to the evaluation of the 3D object detection models’ cross-domain performance: Do we really need models to maintain excellent performance in their original 3D bounding boxes after being applied across domains? From a practical application perspective, one of our main focuses is actually on preventing collisions between vehicles and other obstacles, especially in cross-domain scenarios where correctly predicting the size of vehicles is much more difficult. In other words, as long as a model can accurately identify the closest surfaces to the ego vehicle, it is sufficient to effectively avoid obstacles. In this paper, we propose two metrics to measure 3D object detection models’ ability of detecting the closer surfaces to the sensor on the ego vehicle, which can be used to evaluate their cross-domain performance more comprehensively and reasonably. Furthermore, we propose a refinement head, named EdgeHead, to guide models to focus more on the learnable closer surfaces, which can greatly improve the cross-domain performance of existing models not only under our new metrics, but even also under the original BEV/3D metrics. Our code is available at https://github.com/Galaxy-ZRX/EdgeHead.
Ruixiao Zhang 0001, Yihong Wu 0004, Juheon Lee, Xiaohao Cai, Adam Prügel-Bennett
ECAI4
2024 Non-negative subspace feature representation for few-shot learning in medical imaging
Keqiang Fan, Xiaohao Cai, Mahesan Niranjan
Image Vis. Comput.2
2024 3D orientation field transform
abstract
Abstract Vascular structure enhancement is very useful in image processing and computer vision. The enhancement of the presence of the structures like tubular networks in given images can improve image-dependent diagnostics and can also facilitate tasks like segmentation. The two-dimensional (2D) orientation field transform has been proved to be effective at enhancing 2D contours and curves in images by means of top-down processing. It, however, has no counterpart in 3D images due to the extremely complicated orientation in 3D against 2D. Given the rising demand and interest in handling 3D images, we experiment with modularising the concept and generalise the algorithm to 3D curves. In this work, we propose a 3D orientation field transform. It is a vascular structure enhancement algorithm that can cleanly enhance images having very low signal-to-noise ratio, and push the limits of 3D image quality that can be enhanced computationally. This work also utilises the benefits of modularity and offers several combinative options that each yield moderately better enhancement results in different scenarios. In principle, the proposed 3D orientation field transform can naturally tackle any number of dimensions. As a special case, it is also ideal for 2D images, owning a simpler methodology compared to the previous 2D orientation field transform. The concise structure of the proposed 3D orientation field transform also allows it to be mixed with other enhancement algorithms, and as a preliminary filter to other tasks like segmentation and detection. The effectiveness of the proposed method is demonstrated with synthetic 3D images and real-world transmission electron microscopy tomograms ranging from 2D curve enhancement to, the more important and interesting, 3D ones. Extensive experiments and comparisons with existing related methods also demonstrate the excellent performance of the proposed 3D orientation field transform.
Wai-Tsun Yeung, Xiaohao Cai, Zizhen Liang, Byung-Ho Kang
Pattern Anal. Appl.2
2023 A Bilevel Formalism for the Peer-Reviewing Problem
abstract
Due to the large number of submissions that more and more conferences experience, finding an automatized way to well distribute the submitted papers among reviewers has become necessary. We model the peer-reviewing matching problem as a bilevel programming (BP) formulation. Our model consists of a lower-level problem describing the reviewers’ perspective and an upper-level problem describing the editors’. Every reviewer is interested in minimizing their overall effort, while the editors are interested in finding an allocation that maximizes the quality of the reviews and follows the reviewers’ preferences the most. To the best of our knowledge, the proposed model is the first one that formulates the peer-reviewing matching problem by considering two objective functions, one to describe the reviewers’ viewpoint and the other to describe the editors’ viewpoint. We demonstrate that both the upper-level and lower-level problems are feasible and that our BP model admits a solution under mild assumptions. After studying the properties of the solutions, we propose a heuristic to solve our model and compare its performance with the relevant state-of-the-art methods. Extensive numerical results show that our approach can find fairer solutions with competitive quality and less effort from the reviewers.(Our code website: https://github.com/Galaxy-ZRX/Bilevel-Review.)
Gennaro Auricchio, Ruixiao Zhang 0001, Jie Zhang 0008, Xiaohao Cai
ECAI4
2023 TransNet: A Transfer Learning-Based Network for Human Action Recognition
abstract
Human action recognition (HAR) is a high-level and significant research area in computer vision due to its ubiquitous applications. The main limitations of the current HAR models are their complex structures and lengthy training time. In this paper, we propose a simple yet versatile and effective end-to-end deep learning architecture, coined as TransNet, for HAR. TransNet decomposes the complex 3D-CNNs into 2D- and 1D-CNNs, where the 2D- and 1D-CNN components extract spatial features and temporal patterns in videos, respectively. Benefiting from its concise architecture, TransNet is ideally compatible with any pretrained state-of-the-art 2D-CNN models in other fields, being transferred to serve the HAR task. In other words, it naturally leverages the power and success of transfer learning for HAR, bringing huge advantages in terms of efficiency and effectiveness. Extensive experimental results and the comparison with the state-of-the-art models demonstrate the superior performance of the proposed TransNet in HAR in terms of flexibility, model complexity, training speed and classification accuracy.
Khaled Alomar, Xiaohao Cai
ICMLA2
2023 IIHT: Medical Report Generation with Image-to-Indicator Hierarchical Transformer
Keqiang Fan, Xiaohao Cai, Mahesan Niranjan
ICONIP (6)2
2020 Wavelet-based segmentation on the sphere
Xiaohao Cai, Christopher G. R. Wallis, Jennifer Y. H. Chan, Jason D. McEwen
Pattern Recognit.1
2020 3D Segmentation of Trees Through a Flexible Multiclass Graph Cut Algorithm
abstract
Developing a robust algorithm for automatic individual tree crown (ITC) detection from airborne laser scanning (ALS) data sets is important for tracking the responses of trees to anthropogenic change. Such approaches allow the size, growth, and mortality of individual trees to be measured, enabling forest carbon stocks and dynamics to be tracked and understood. Many algorithms exist for structurally simple forests, including coniferous forests and plantations. Finding a robust solution for structurally complex, species-rich tropical forests remains a challenge; existing segmentation algorithms often perform less well than simple area-based approaches when estimating plot-level biomass. Here, we describe a multiclass graph cut (MCGC) approach to tree crown delineation. This uses local 3D geometry and density information, alongside knowledge of crown allometries, to segment ITCs from airborne light detection and ranging point clouds. Our approach robustly identifies trees in the top and intermediate layers of the canopy, but cannot recognize small trees. From these 3D crowns, we are able to measure individual tree biomass. Comparing these estimates with those from permanent inventory plots, our algorithm can produce robust estimates of hectare-scale carbon density, demonstrating the power of ITC approaches in monitoring forests. The flexibility of our method to add additional dimensions of information, such as spectral reflectance, make this approach an obvious avenue for future development and extension to other sources of 3D data, such as structure from motion data sets.
Jonathan Williams 0003, Carola-Bibiane Schönlieb, Tom Swinfield, Juheon Lee, Xiaohao Cai, Lan Qie, David Coomes
IEEE Trans. Geosci. Remote. Sens.5
2015 Mapping individual trees from airborne multi-sensor imagery
abstract
Individual tree species mapping is important to understand forest dynamics and species distribution patterns. Airborne LiDAR with hyperspectral imaging has been extensively used to extract biophysical traits of vegetation and detect species. However, its application for individual tree mapping is limited due to technical problems. To address the problems, this paper presents effective and efficient algorithms in terms of tackling co-alingment of LiDAR and hyperspectral datasets, classifying individual trees, thus detecting tree species and leaf chemistry from the tree mapping.
Juheon Lee, Xiaohao Cai, Carola-Bibiane Schönlieb, David Coomes
IGARSS2
2015 Variational image segmentation model coupled with image restoration achievements
Xiaohao Cai
Pattern Recognit.1
2015 Nonparametric Image Registration of Airborne LiDAR, Hyperspectral and Photographic Imagery of Wooded Landscapes
abstract
There is much current interest in using multisensor airborne remote sensing to monitor the structure and biodiversity of woodlands. This paper addresses the application of nonparametric (NP) image-registration techniques to precisely align images obtained from multisensor imaging, which is critical for the successful identification of individual trees using object recognition approaches. NP image registration, in particular, the technique of optimizing an objective function, containing similarity and regularization terms, provides a flexible approach for image registration. Here, we develop a NP registration approach, in which a normalized gradient field is used to quantify similarity, and curvature is used for regularization (NGF-Curv method). Using a survey of woodlands in southern Spain as an example, we show that NGF-Curv can be successful at fusing data sets when there is little prior knowledge about how the data sets are interrelated (i.e., in the absence of ground control points). The validity of NGF-Curv in airborne remote sensing is demonstrated by a series of experiments. We show that NGF-Curv is capable of aligning images precisely, making it a valuable component of algorithms designed to identify objects, such as trees, within multisensor data sets.
Juheon Lee, Xiaohao Cai, Carola-Bibiane Schönlieb, David Coomes
IEEE Trans. Geosci. Remote. Sens.2
2013 Vessel Segmentation in Medical Imaging Using a Tight-Frame-Based Algorithm
abstract
Tight-frame, a generalization of orthogonal wavelets, has been used successfully in various problems in image processing, including inpainting, impulse noise removal, and superresolution image restoration. Segmentation is the process of identifying object outlines within images. There are quite a few efficient algorithms for segmentation such as model-based approaches, pattern recognition techniques, tracking-based approaches, and artificial intelligence--based approaches. In this paper, we propose applying the tight-frame approach to automatically identify tube-like structures in medical imaging, with the primary application of segmenting blood vessels in magnetic resonance angiography images. Our method iteratively refines a region that encloses the potential boundary of the vessels. At each iteration, we apply the tight-frame algorithm to denoise and smooth the potential boundary and sharpen the region. The cost per iteration is proportional to the number of pixels in the image. We prove that the iteration converges in a finite number of steps to a binary image whereby the segmentation of the vessels can be done straightforwardly. Numerical experiments on synthetic and real two-dimensional (2D) and three-dimensional (3D) images demonstrate that our method is more accurate when compared with some representative segmentation methods, and it usually converges within a few iterations.
Xiaohao Cai, Raymond Chan 0001, Serena Morigi, Fiorella Sgallari
SIAM J. Imaging Sci.1
2013 A Two-Stage Image Segmentation Method Using a Convex Variant of the Mumford-Shah Model and Thresholding
abstract
The Mumford--Shah model is one of the most important image segmentation models and has been studied extensively in the last twenty years. In this paper, we propose a two-stage segmentation method based on the Mumford--Shah model. The first stage of our method is to find a smooth solution $g$ to a convex variant of the Mumford--Shah model. Once $g$ is obtained, then in the second stage the segmentation is done by thresholding $g$ into different phases. The thresholds can be given by the users or can be obtained automatically using any clustering methods. Because of the convexity of the model, $g$ can be solved efficiently by techniques like the split-Bregman algorithm or the Chambolle--Pock method. We prove that our method is convergent and that the solution $g$ is always unique. In our method, there is no need to specify the number of segments $K$ ($K\geq2$) before finding $g$. We can obtain any $K$-phase segmentations by choosing $(K-1)$ thresholds after $g$ is found in the first stage, and in the second stage there is no need to recompute $g$ if the thresholds are changed to reveal different segmentation features in the image. Experimental results show that our two-stage method performs better than many standard two-phase or multiphase segmentation methods for very general images, including antimass, tubular, MRI, noisy, and blurry images.
Xiaohao Cai, Raymond Chan 0001, Tieyong Zeng
SIAM J. Imaging Sci.1