Hao Xiong 0001

dblp:46/5036-1 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-6842-1667ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CondFoodGen: A Conditional Two-Stream Network for Controllable Food Image Generation
abstract
Food image generation is an important research direction in food computing, aiming to produce highly realistic images that accurately capture the visual characteristics of various dishes while adhering to specified input conditions. Existing methods that rely solely on textual descriptions struggle to handle the large intra-class variability of food, often resulting in limited diversity and accuracy. Although some approaches incorporate additional conditions, they generally lack optimizations for food-specific challenges, leading to inconsistencies in texture, shape, and color fidelity. To address these limitations, we propose CondFoodGen, a diffusion-based two-stream network for controllable food image generation. The architecture consists of a control stream and a generation stream, where the control stream provides conditional guidance to regulate the generation process. To optimize bidirectional interactions between the two streams, we introduce the Bidirectional Adaptive Gating (BAG) mechanism, which not only guides synthesis but also adaptively refines control representations through feedback from the generation stream. In addition, we propose the Wavelet-Guided Hierarchical Attention (WGHA) module, which combines wavelet-based multi-frequency analysis with hierarchical attention to enhance fine-grained texture fidelity and structural realism. A progressive multi-stage training strategy further stabilizes optimization and enables seamless integration of conditional guidance with bidirectional interaction. Extensive experiments on three food image datasets demonstrate that CondFoodGen consistently generates high-quality and diverse images. Compared with the best existing food image generation methods, our approach achieves an average improvement of about 11.0% across three evaluation metrics and compared to the leading conditional generation approaches, the average improvement reaches 16.2%. The source code, trained models, and supplementary materials are publicly available at https://github.com/housujuan123/CondFoodGen.
Mengyao Zhao, Hao Xiong 0001, Weiqing Min, Sujuan Hou, Mengmeng Zhang 0008, Shuqiang Jiang
IEEE Trans. Image Process.2
2025 BraTS-UMamba: Adaptive Mamba UNet with Dual-Band Frequency Based Feature Enhancement for Brain Tumor Segmentation
Haoran Yao, Hao Xiong 0001, Hualei Shen, Shlomo Berkovsky
MICCAI (16)2
2025 DSDGF-Nutri: A Decoupled Self-Distillation Network with Gating Fusion For Food Nutritional Assessment
abstract
Accurate assessment of food nutrition is essential for promoting healthy eating habits. While recent deep learning approaches have enhanced vision-based nutritional estimation through RGB-D multi-modal fusion, they often overlook fine-grained surface components (e.g., oil and sugar) that significantly influence nutritional values. Some recent approaches have improved accuracy by incorporating ingredient data, but their reliance on such input during inference limits practical applicability, as ingredient details are often unavailable in real-world settings. To address this limitation, we propose DSDGF-Nutri, a novel Decoupled Self-Distillation network with Gating Fusion for food Nutri tional assessment. Our method leverages ingredient knowledge during training but relies solely on RGB-D inputs at inference. Specifically, DSDGF-Nutri introduces: (1) a self-distillation mechanism with gating fusion that transfers ingredient-aware features to the RGB-D network, enabling robust prediction without test-time ingredient input, and (2) a multi-task decoupling architecture with task-specific decoders to minimize cross-task interference. Extensive evaluations on two benchmark datasets demonstrate DSDGF-Nutri outperforms existing methods, achieving state-of-the-art results. This work establishes a new paradigm of multimodal fusion in nutritional assessment by unifying scientific measurements with scalable computer vision applications.
Sujuan Hou, Zhihui Feng, Hao Xiong 0001, Weiqing Min, Peng Li 0081, Shuqiang Jiang
ACM Multimedia3
2025 Distributed Radar Imaging with Parallel Cross-Attention for Continuous Human Motion Recognition
abstract
Radar imaging provides non-contact, privacy-preserving, and environmentally robust monitoring for continuous human motion recognition (HMR) by leveraging diverse information embedded in various radar signal domains. However, current research has not effectively integrated multi-radar and multi-domain imaging to fully exploit the benefits of distributed radar systems. To bridge this gap, we propose a multi-radar, multi-domain parallel cross-attention model with four key components: intra-domain cross-radar weight sharing encoders specific to each domain for consistent feature extraction and parameter reduction, domain-level parallel cross-attention (DLPCAN) modules to fuse domain-specific features and enhance feature representation robustness in each radar, a source-level attention fusion (SLAF) module to highlight significant features from multiple radar inputs, and two bi-directional gated recurrent unit (BiGRU) modules to capture temporal information. The model is trained using connectionist temporal classification (CTC) loss for effective sequence prediction. By integrating data from multiple radar nodes and domains, our approach significantly improves continuous HMR performance compared to single radar systems and single domain data. Comparative evaluations demonstrate that our model outperforms state-of-the-art radar imaging-based HMR solutions.
Jianqiao Zhang 0003, Yijie Gao, Hao Xiong 0001, Jiquan Ma, Qiangguo Jin, ChangYang Li, Peng Cheng 0002, Hui Cui 0002
VTC2025-Spring3
2025 Uncertainty-guided attention learning for malaria parasite detection in thick blood smears
abstract
Malaria may seriously threaten an individual's health and wellbeing, and early screening is pivotal for timely treatment and recovery. In malaria screening, thick blood smears are exploited to count the parasites and assess the severity of the disease. Parasites are tiny objects that can be found in high resolution blood smear images, which renders them difficult for detection. Other than using object detection based methods, prior works also applied image classification techniques to this problem. They first extracted image patches from blood smears as parasite candidates and then utilized convolutional neural networks to classify these patches as parasites or non-parasites. However, these approaches overlook the fact that the blood smear images may contain noises, errors, and background artifacts, which introduces uncertainty and makes the model predictions less stable. In this work, we propose an uncertainty-guided attention learning based network for malaria parasite detection from thick blood smears, which incorporates pixel attention mechanism to identify more fine-grained and pixel-wise informative features, to improve the classification capability of our model. We further put uncertainty estimation on channels of the feature map to guide pixel attention learning, such that the features from channels with higher uncertainty are considered unreliable and are thus restrictively exploited by pixel attention learning. To estimate channel-wise uncertainty, we introduce the Bayesian channel attention, which reformulates the traditional channel attention under the Bayesian framework. As a result, it denotes channel uncertainties with estimated variances that guide the pixel attention learning. We compared to several state-of-the-art baselines on two public datasets using parasite-level and patient-level evaluations. The proposed method demonstrates superior performance with respect to most metrics on two datasets, especially achieving highest average precision (AP) scores in both parasite and patient-level scenarios.
Hao Xiong 0001, Zhiyong Wang 0001, Roneel V. Sharan, Shlomo Berkovsky
Neural Networks1
2024 Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation
abstract
Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation. In this work, we propose a novel hybrid network that involves three main branches: the Multi-Layer Perception (MLP) branch, the Convolutional Neural Network (CNN) branch, and a Fusion branch. The MLP and CNN branches aim to learn global and local features, respectively. To fuse these, the fusion branch introduces a novel hierarchical fusion that performs multi-layered fusions that generate high-level representations to enhance segmentation. Our evaluation with two datasets shows strong performance of the proposed method compared to state-of-the-art baselines.
Chengyu Yuan, Hao Xiong 0001, Guoqing Shangguan, Hualei Shen, Kun Qian 0003, Bin Hu 0001, Björn W. Schuller, Yoshiharu Yamamoto, Shlomo Berkovsky
ICASSP2
2024 AuxSegCount: Auxiliary Seg-Attention Based Network for Wheat Ears Counting in Field Conditions
abstract
Accurate wheat ears counting is crucial to the wheat yield estimation. Existing counting methods explore various architectures using density maps for training, and a few works also incorporate an auxiliary network. However, the small wheat ear with background noises makes its detection hard, and the auxiliary network methods tend to ignore interactions with its main network. To mitigate these issues, we propose a novel framework AuxSegCount which includes a segmentation auxiliary network and a main network for wheat ears counting. Unlike density map, the segmentation mask provides more local contexts of wheat ears. We therefore utilise it and introduce the segmentation attention module (SAM) that aims to capture local features around wheat ears. To promote the interactions, we further present the BiAttention Fusion Module (BiFM) that exploits both global information and local contexts into the main network. The experimental results on two datasets show the superiority of our method.
Hao Xiong 0001, Hecang Zang, Hualei Shen
ICME2
2024 A Multi-information Dual-Layer Cross-Attention Model for Esophageal Fistula Prognosis
Jianqiao Zhang 0003, Hao Xiong 0001, Qiangguo Jin, Tian Feng 0001, Jiquan Ma, Ping Xuan, Peng Cheng 0002, Zhiyu Ning, ChangYang Li, Hui Cui 0002
MICCAI (5)2
2024 Fast Online Adaptation of Visual SLAM via Variational Information Transfer and Preservation
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Shlomo Berkovsky, Zhiyong Wang 0001
MMAsia2
2023 Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning
abstract
Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance degradation on the unseen data of target domain. Hence, we propose an online adaptation approach to continuously adapt a pre-trained visual SLAM model to changing environments in a self-supervised manner. To preserve pre-learned knowledge against catastrophic forgetting, we perform updating on a novel adapter proposed rather than fine-tuning the whole model for adaptation. The adapter includes a cross-domain feature translation module that translates pre-learned features into translated features suitable for adaptation. Ideally, the translated new features should not only contain pre-learned knowledge but also substantially distinct from pre-learned features since these two features represent different domains. We thus introduce cycle-consistent contrastive learning to maximize the dissimilarity between these two features by enlarging the distance between them in the feature space. Besides, our contrastive learning method exploiting cycle-consistency contraint enables the translated features to be transferred back to the pre-learned ones, which helps the translated features better preserve pre-learned knowledge. Comprehensive experiments on both synthetic and real-world datasets demonstrate superior adaptation performance of our proposed method over several state-of-the-art baselines.
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Zhihui Wang 0001, Zhiyong Wang 0001
ICRA2
2022 Weak label based Bayesian U-Net for optic disc segmentation in fundus images
Hao Xiong 0001, Sidong Liu, Roneel V. Sharan, Enrico W. Coiera, Shlomo Berkovsky
Artif. Intell. Medicine1
2022 Family informatics
abstract
While families have a central role in shaping individual choices and behaviors, healthcare largely focuses on treating individuals or supporting self-care. However, a family is also a health unit. We argue that family informatics is a necessary evolution in scope of health informatics. To deal with the needs of individuals, we must ensure technologies account for the role of their families and may require new classes of digital service. Social networks can help conceptualize the structure, composition, and behavior of families. A family network can be seen as a multiagent system with distributed cognition. Digital tools can address family needs in (1) sensing and monitoring; (2) communicating and sharing; (3) deciding and acting; and (4) treating and preventing illness. Family informatics is inherently multidisciplinary and has the potential to address unresolved chronic health challenges such as obesity, mental health, and substance abuse, support acute health challenges, and to improve the capacity of individuals to manage their own health needs.
Enrico W. Coiera, Kathleen Yin, Roneel V. Sharan, Saba Akbar, Satya Vedantam, Hao Xiong 0001, Jenny Waldie, Annie Y. S. Lau
J. Am. Medical Informatics Assoc.6
2022 Identifying daily activities of patient work for type 2 diabetes and co-morbidities: a deep learning and wearable camera approach
abstract
OBJECTIVE: People are increasingly encouraged to self-manage their chronic conditions; however, many struggle to practise it effectively. Most studies that investigate patient work (ie, tasks involved in self-management and contexts influencing such tasks) rely on self-reports, which are subject to recall and other biases. Few studies use wearable cameras and deep learning to capture and classify patient work activities automatically. MATERIALS AND METHODS: We propose a deep learning approach to classify activities of patient work collected from wearable cameras, thereby studying self-management routines more effectively. Twenty-six people with type 2 diabetes and comorbidities wore a wearable camera for a day, generating more than 400 h of video across 12 daily activities. To classify these video images, a weighted ensemble network that combines Linear Discriminant Analysis, Deep Convolutional Neural Networks, and Object Detection algorithms is developed. Performance of our model is assessed using Top-1 and Top-5 metrics, compared against manual classification conducted by 2 independent researchers. RESULTS: Across 12 daily activities, our model achieved on average the best Top-1 and Top-5 scores of 81.9 and 86.8, respectively. Our model also outperformed other non-ensemble techniques in terms of Top-1 and Top-5 scores for most activity classes, demonstrating the superiority of leveraging weighted ensemble techniques. CONCLUSIONS: Deep learning can be used to automatically classify daily activities of patient work collected from wearable cameras with high levels of accuracy. Using wearable cameras and a deep learning approach can offer an alternative approach to investigate patient work, one not subjected to biases commonly associated with self-report methods.
Hao Xiong 0001, Hoai Nam Phan, Kathleen Yin, Shlomo Berkovsky, Joshua Jung, Annie Y. S. Lau
J. Am. Medical Informatics Assoc.1
2021 A diversified shared latent variable model for efficient image characteristics extraction and modelling
Hao Xiong 0001, Yuan Yan Tang, Fionn Murtagh, Leszek Rutkowski, Shlomo Berkovsky
Neurocomputing1
2021 Prediction of anxiety disorders using a feature ensemble based bayesian neural network
Hao Xiong 0001, Shlomo Berkovsky, Mia Romano, Roneel V. Sharan, Sidong Liu, Enrico W. Coiera, Lauren F. McLellan
J. Biomed. Informatics1
2017 A Diversified Generative Latent Variable Model for WiFi-SLAM
abstract
WiFi-SLAM aims to map WiFi signals within an unknown environment while simultaneously determining the location of a mobile device. This localization method has been extensively used in indoor, space, undersea, and underground environments. For the sake of accuracy, most methods label the signal readings against ground truth locations. However, this is impractical in large environments, where it is hard to collect and maintain the data. Some methods use latent variable models to generate latent-space locations of signal strength data, an advantage being that no prior labeling of signal strength readings and their physical locations is required. However, the generated latent variables cannot cover all wireless signal locations and WiFi-SLAM performance is significantly degraded. Here we propose the diversified generative latent variable model (DGLVM) to overcome these limitations. By building a positive-definite kernel function, a diversity-encouraging prior is introduced to render the generated latent variables non-overlapping, thus capturing more wireless signal measurements characteristics. The defined objective function is then solved by variational inference. Our experiments illustrate that the method performs WiFi localization more accurately than other label-free methods.
Hao Xiong 0001, Dacheng Tao
AAAI1
2016 Diversified Dynamical Gaussian Process Latent Variable Model for Video Repair
abstract
Videos can be conserved on different media. However, storing on media such as films and hard disks can suffer from unexpected data loss, for instance from physical damage. Repair of missing or damaged pixels is essential for video maintenance and preservation. Most methods seek to fill in missing holes by synthesizing similar textures from local or global frames. However, this can introduce incorrect contexts, especially when the missing hole or number of damaged frames is large. Furthermore, simple texture synthesis can introduce artifacts in undamaged and recovered areas. To address aforementioned problems, we propose the diversified dynamical Gaussian process latent variable model (D2GPLVM) for considering the variety in existing videos and thus introducing a diversity encouraging prior to inducing points. The aim is to ensure that the trained inducing points, which are a smaller set of all observed undamaged frames, are more diverse and resistant for context-aware and artifacts-free based video repair. The defined objective function in our proposed model is initially not analytically tractable and must be solved by variational inference. Finally, experimental testing illustrates the robustness and effectiveness of our method for damaged video repair.
Hao Xiong 0001, Tongliang Liu, Dacheng Tao
AAAI1
2016 Robust foreground object segmentation from handheld camera videos with occlusion map
Hao Xiong 0001, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.1
2016 Dual Diversified Dynamical Gaussian Process Latent Variable Model for Video Repairing
abstract
In this paper, we propose a dual diversified dynamical Gaussian process latent variable model ( [Formula: see text]GPLVM) to tackle the video repairing issue. For preservation purposes, videos have to be conserved on media. However, storing on media, such as films and hard disks, can suffer from unexpected data loss, for instance, physical damage. So repairing of missing or damaged pixels is essential for better video maintenance. Most methods seek to fill in missing holes by synthesizing similar textures from local patches (the neighboring pixels), consecutive frames, or the whole video. However, these can introduce incorrect contexts, especially when the missing hole or number of damaged frames is large. Furthermore, simple texture synthesis can introduce artifacts in undamaged and recovered areas. To address aforementioned problems, we introduce two diversity encouraging priors to both of inducing points and latent variables for considering the variety in existing videos. In [Formula: see text]GPLVM, the inducing points constitute a smaller subset of observed data, while latent variables are a low-dimensional representation of observed data. Since they have a strong correlation with the observed data, it is essential that both of them can capture distinct aspects of and fully represent the observed data. The dual diversity encouraging priors ensure that the trained inducing points and latent variables are more diverse and resistant for context-aware and artifacts-free-based video repairing. The defined objective function in our proposed model is initially not analytically tractable and must be solved by variational inference. Finally, experimental testing results illustrate the robustness and effectiveness of our method for damaged video repairing.
Hao Xiong 0001, Tongliang Liu, Dacheng Tao, Heng Tao Shen
IEEE Trans. Image Process.1