Xiyao Ma

dblp:254/1076 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-5411-4527ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Disentangling for Transfer: Boosting Limited Modalities via Information-Theoretic Regularization and Cross-Modal Reconstruction
abstract
Missing critical modalities in medical imaging poses significant challenges for AI-driven diagnostic systems, particularly in scenarios where limited modalities must suffice for downstream tasks. Existing approaches often fail to fully leverage privileged features available only at training or address the information gap between privileged and limited modalities, resulting in suboptimal performance. To address this, we propose a unified, dual-stage Disentanglement-AligNmenT framEwork (DANTE), which uses InformationTheoretic Regularization and Cross-Modal Reconstruction to decompose full-modality information into alignable and privileged-exclusive components. In the first stage, a self-supervised pre-training strategy based on cross-modal reconstruction acts as a proxy task to implicitly incentivize disentangled representations. In the second stage, we present an information-theoretic regularization to explicitly maximize the transfer of privileged knowledge through two novel modules: (1) a Mutual Alignment Module that employs multilevel bidirectional alignment between limited-modality features and alignable features, enhancing cross-modal representation consistency; (2) a Privileged Compaction Module that restricts the privileged-exclusive information flow, promoting the integration of task-relevant content into alignable representations. Experimental results on three challenging medical datasets demonstrate that DANTE achieves state-of-the-art performance, demonstrating its effectiveness in leveraging privileged guidance under modality scarcity, and exhibits broad applicability across diverse medical imaging scenarios.
Zhiyun Zhang, Yan-Jie Zhou, Yujian Hu, Xiyao Ma, Zhouhang Yuan, Hongkun Zhang, Minfeng Xu
AAAI4
2026 Toward Precise Guidance: A Novel Cross-Dimensional Mapping Framework for 3-D Cerebrovascular Surgical Navigation
Haining Zhao 0002, Shiqi Liu 0004, Ji-Chang Luo, Xiao-Hu Zhou, Zeng-Guang Hou, Li-Qun Jiao, Xiyao Ma, Lin-Sen Zhang, Xiaoliang Xie
IEEE Trans Autom. Sci. Eng.8
2025 Advancing Efficiency and Accuracy: A Dynamic Anatomy-Aware 3D Vessel Segmentation Framework
abstract
Fast and accurate segmentation of three-dimensional vasculature significantly enhances the precision and safety of endovascular surgeries. However, current 3D segmentation methods suffer from low accuracy due to a lack of global information and severe time consumption, making them impractical for clinical scenarios. In this paper, we propose a novel Dynamic Anatomy-aware Inference (DAI) framework, which leverages the anatomical prior of vascular structures to facilitate segmentation performance. In this framework, a Target Window Search (TWS) method is proposed to sample potential patches containing target vasculature, which leads to less computational burden and more consistent input data distribution for the embedded segmentation models. Then, a Spatial Fuse Module (SFM) is designed to encode the features of sampled patches based on their topological relationships, so that the long-range dependencies of target vasculature are effectively captured. Furthermore, comprehensive experiments are conducted to validate the effectiveness of proposed methods. Compared with the prevalent sliding window inference framework, a variety of models embedded in DAI achieve significant improvements in terms of both efficiency and accuracy: 13-147 × reductions in inference time (↓), 2-9% increases in Dice (↑), 4-16% increases in mIoU (↑), 57-83% decreases in HD (↓), 21-56% increases in clDice (↑).
Haining Zhao 0002, Shiqi Liu 0004, Ji-Chang Luo, Xiaohu Zhou, Jiaxing Wang 0001, Zeng-Guang Hou, Li-Qun Jiao, Xiyao Ma, Xiaoliang Xie
IEEE Trans Autom. Sci. Eng.9
2024 MEND: Meta Demonstration Distillation for Efficient and Effective In-Context Learning
abstract
Large Language models (LLMs) have demonstrated impressive in-context learning (ICL) capabilities, where a LLM makes predictions for a given test input together with a few input-output pairs (demonstrations). Nevertheless, the inclusion of demonstrations poses a challenge, leading to a quadratic increase in the computational overhead of the self-attention mechanism. Existing solutions attempt to condense lengthy demonstrations into compact vectors. However, they often require task-specific retraining or compromise LLM's in-context learning performance. To mitigate these challenges, we present Meta Demonstration Distillation (MEND), where a language model learns to distill any lengthy demonstrations into vectors without retraining for a new downstream task. We exploit the knowledge distillation to enhance alignment between MEND and MEND, achieving both efficiency and effectiveness concurrently. MEND is endowed with the meta-knowledge of distilling demonstrations through a two-stage training process, which includes meta-distillation pretraining and fine-tuning. Comprehensive evaluations across seven diverse ICL settings using decoder-only (GPT-2) and encoder-decoder (T5) attest to MEND's prowess. It not only matches but often outperforms the Vanilla ICL as well as other state-of-the-art distillation models, while significantly reducing the computational demands. This innovation promises enhanced scalability and efficiency for the practical deployment of large language models.
Yichuan Li 0001, Xiyao Ma, Sixing Lu, Kyumin Lee, Chenlei Guo
ICLR2
2023 Towards Flexible and Universal: A Novel Endpoint-based Framework for Vessel Structural Information Extraction
abstract
In computer-assisted intravascular interventional surgery, extracting detailed information of target vessels from X-ray angiographic images can be meaningful in improving safety and effectiveness. However, large amounts of effort have been dedicated to segmenting the whole blood vessels from the background while ignoring the internal structure, which is limited in clinical application. In this paper, we propose a flexible and universal endpoint-based framework for vessel structural information extraction. The framework first localizes all the endpoints of target vessel segments through a Coarse-to-Fine Keypoint Detection Network (CFKD-Net), in which the designed Multi-branch Feature Aggregation (MFA) module captures both in-patch and cross-patch information to help recognize the points of interest based on global structure. A novel MaskMSELoss is also proposed to disambiguate those irrelevant responses. Then a designed VEssel Segmentation and Analysis (VESA) algorithm will generate the segmentation mask and morphological analysis for each vessel segment simply based on the endpoints. It can also be flexibly applied to analyze variant blood vessels which are not pre-defined before. Extensive experiments on two different coronary artery datasets consistently demonstrate that this framework can achieve state-of-the-art detection performance and successfully extract and analyze target vessel segments. Since the framework shows excellent performance on the coronary arteries with severe deformation and strong noise, it is highly promising for analyzing other vascular images.
Xiyao Ma, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Xinkai Qu, Wenzheng Han, Ming Wang 0001, Lin-Sen Zhang
ACM Multimedia1
2023 Communication-Efficient and Attack-Resistant Federated Edge Learning With Dataset Distillation
abstract
Federated Edge Learning considers a large amount of distributed edge nodes collectively train a global gradient-based model for edge computing in the Artificial Internet of Things, which significantly promotes the development of cloud computing. However, current federated learning algorithms take tens of communication rounds transmitting unwieldy model weights under ideal circumstances and hundreds when data is poorly distributed. This drawback directly results in expensive communication overhead for edge devices. Inspired by recent work on dataset distillation and distributed one-shot learning, we propose Distilled One-Shot Federated Learning (DOSFL) to significantly reduce the communication cost while achieving comparable performance. In just one round, each client distills their private dataset, sends the synthetic data to the server, and collectively trains a global model. The distilled data look like noise and are only useful to the specific model weights,i.e.,become useless after the model updates. With this weight-less and gradient-less design, the total communication cost of DOSFL is up to three orders of magnitude less than FedAvg while preserving up to 99% performance of centralized training on both vision and language tasks with different models including CNN, LSTM, Transformer,etc. We demonstrate that an eavesdropping attacker cannot properly train a good model using the leaked distilled data, without knowing the initial model weights. DOSFL serves as an inexpensive method to quickly converge on a performant pre-trained model with less than 0.1% communication cost of traditional methods.
Xiyao Ma, Dapeng Oliver Wu, Xiaolin Li 0001
IEEE Trans. Cloud Comput.2
2023 A Novel Spatial Position Prediction Navigation System Makes Surgery More Accurate
abstract
During intravascular interventional surgery, the 3D surgical navigation system can provide doctors with 3D spatial information of the vascular lumen, reducing the impact of missing dimension caused by digital subtraction angiography (DSA) guidance and further improving the success rate of surgeries. Nevertheless, this task often comes with the challenge of complex registration problems due to vessel deformation caused by respiratory motion and high requirements for the surgical environment because of the dependence on external electromagnetic sensors. This article proposes a novel 3D spatial predictive positioning navigation (SPPN) technique to predict the real-time tip position of surgical instruments. In the first stage, we propose a trajectory prediction algorithm integrated with instrumental morphological constraints to generate the initial trajectory. Then, a novel hybrid physical model is designed to estimate the trajectory's energy and mechanics. In the second stage, a point cloud clustering algorithm applies multi-information fusion to generate the maximum probability endpoint cloud. Then, an energy-weighted probability density function is introduced using statistical analysis to achieve the prediction of the 3D spatial location of instrument endpoints. Extensive experiments are conducted on 3D-printed human artery and vein models based on a high-precision electromagnetic tracking system. Experimental results demonstrate the outstanding performance of our method, reaching 98.2% of the achievement ratio and less than 3 mm of the average positioning accuracy. This work is the first 3D surgical navigation algorithm that entirely relies on vascular interventional robot sensors, effectively improving the accuracy of interventional surgery and making it more accessible for primary surgeons.
Lin-Sen Zhang, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Chao-Nan Wang, Xinkai Qu, Wenzheng Han, Xiyao Ma
IEEE Trans. Medical Imaging9
2022 Contrastive Knowledge Graph Attention Network for Request-Based Recipe Recommendation
abstract
To improve daily customer experience, kitchen assistant becomes one of the enabled service in intelligent voice assistants, presenting personalized and relevant recipes to satisfy customer requests. Current solutions for recipe recommendation suffers from two limitations: First, user-recipe interactions are modeled in a uniform manner, which neglects the diversity of user preferences on recipe adoptions, inherently hurting the model performance. Second, users may interact with recipe randomly, resulting in inevitable data noise issue. In this work, we alleviate the foregoing issues by proposing contrastive knowledge graph attention network for recipe recommendation, where a knowledge graph attention-based recommender helps learn fine-grained user and recipe embeddings by modeling diversified user preferences from user behaviors. Moreover, a contrastive learning module that integrates unsupervised and supervised contrastive learning is proposed to improve model robustness. The experimental results on two real-world datasets show that the proposed approach outperforms the state-of-the-art baseline methods.
Xiyao Ma, Zheng Gao 0002, Mohamed Abdelhady
ICASSP1
2022 Towards Automated Segmentation of Human Abdominal Aorta and Its Branches Using a Hybrid Feature Extraction Module with LSTM
Bo Zhang 0104, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Xiyao Ma, Lin-Sen Zhang
ICONIP (7)7
2022 HCL: Hybrid Contrastive Learning for Graph-based Recommendation
abstract
Graph-based collaborative filtering for recommendation has attracted great attention recently, due to its effectiveness of capturing high-order proximity among users and items. To further improve its model robustness and alleviate label-sparsity issue, contrastive learning has been introduced to polish user and item representation by contrasting different views of user/item nodes, learning necessary and robust representation for recommendation. However, we argue that prior contrastive learning approaches only explore its unsupervised intrinsic nature as a plug-in without leveraging available user-item interactions, failing to exploit the huge potential of contrastive learning. In this paper, to alleviate the above issues, we propose Hybrid Contrastive Learning for graph-based recommendation that integrates unsupervised and supervised contrastive learning. Specifically, to improve model robustness, we first present bipartite graph augmentation operations from the perspectives of node attributes and topology to generate incomplete and noisy graph views. Then, we propose a hybrid contrastive learning module that conducts unsupervised and supervised contrastive learning together. Last, we present an approach to perform hybrid contrastive learning permutationally among multiple views. Extensive experiments show that our proposed model not only outperforms state-of-the-art baselines significantly on two public datasets and one internal dataset, but also demonstrates superiority regarding to model robustness over other strong baselines.
Xiyao Ma, Zheng Gao 0002, Mohamed Abdelhady
IJCNN1
2022 Contrastive Co-training for Diversified Recommendation
abstract
Beyond accuracy, diversity has become a crucial factor to evaluate a recommendation system as higher diversity helps mitigate echo chamber issue and improve user satisfaction. Recently, great success has been made to improve diversity, but the approaches often sacrifice much lower accuracy. Herein this work, we propose contrastive co-training for diversified recommendation that improves diversity greatly and achieves comparable or even better accuracy. Specifically, we keep two user-item graph views for recommendation and contrastive learning, respectively. Pseudo edges are predicted from current graph view to augment the other graph view by mining the novel items that users might be highly interested in. However, merely leveraging co-training hurts the accuracy since the pseudo labels are sometimes noisy. Therefore, we propose diversified contrastive learning that not only is robust to the noisy pseudo edges but also improves the diversity further by alleviating the popularity and category biases by re-balancing item-level popularity and category-level advantage. The extensive experiments on three public datasets show the superiority of our proposed model in terms of accuracy and diversity compared with strong baselines.
Xiyao Ma, Zheng Gao 0002, Mohamed Abdelhady
IJCNN1
2022 DSP-Net: Deeply-Supervised Pseudo-Siamese Network for Dynamic Angiographic Image Matching
Xiyao Ma, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Yan-Jie Zhou, Lin-Sen Zhang, Chao-Nan Wang
MICCAI (8)1
2022 A Novel Fusion Network for Morphological Analysis of Common Iliac Artery
Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Zeng-Guang Hou, Yan-Jie Zhou, Xiyao Ma
MICCAI (8)7
2022 Beyond Class-Level Privacy Leakage: Breaking Record-Level Privacy in Federated Learning
abstract
Federated learning (FL) enables multiple clients to collaboratively build a global learning model without sharing their own raw data for privacy protection. Unfortunately, recent research still found privacy leakage in FL, especially on image classification tasks, such as the reconstruction of class representatives. Nevertheless, such analysis on image classification tasks is not applicable to uncover the privacy threats against natural language processing (NLP) tasks, whose records composed of sequential texts cannot be grouped as class representatives. The finer (record-level) granularity in NLP tasks not only makes it more challenging to extract individual text records, but also exposes more serious threats. This article presents the first attempt to explore the record-level privacy leakage against NLP tasks in FL. We propose a framework to investigate the exposure of the records of interest in federated aggregations by leveraging the perplexity of language modeling. Through monitoring the exposure patterns, we propose two correlation attacks to identify the corresponding clients when extracting their specific records. Extensive experimental results demonstrate the effectiveness of the proposed attacks. We have also examined several countermeasures and shown that they are ineffective to mitigate such attacks, and hence further research is expected.
Xiaoyong Yuan, Xiyao Ma, Lan Zhang 0005, Yuguang Fang, Dapeng Oliver Wu
IEEE Internet Things J.2
2020 Improving Question Generation with Sentence-Level Semantic Matching and Answer Position Inferring
abstract
Taking an answer and its context as input, sequence-to-sequence models have made considerable progress on question generation. However, we observe that these approaches often generate wrong question words or keywords and copy answer-irrelevant words from the input. We believe that lacking global question semantics and exploiting answer position-awareness not well are the key root causes. In this paper, we propose a neural question generation model with two general modules: sentence-level semantic matching and answer position inferring. Further, we enhance the initial state of the decoder by leveraging the answer-aware gated fusion mechanism. Experimental results demonstrate that our model outperforms the state-of-the-art (SOTA) models on SQuAD and MARCO datasets. Owing to its generality, our work also improves the existing models significantly.
Xiyao Ma, Qile Zhu, Xiaolin Li 0001
AAAI1
2020 A Batch Normalized Inference Network Keeps the KL Vanishing Away
abstract
Variational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks.However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse".Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint.We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive.Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters.Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently.We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE).Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE.
Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li 0001, Dapeng Oliver Wu
ACL4
2019 Adaptive Leader-Follower Formation Control and Obstacle Avoidance via Deep Reinforcement Learning
abstract
We propose a deep reinforcement learning (DRL) methodology for the tracking, obstacle avoidance, and formation control of nonholonomic robots. By separating vision-based control into a perception module and a controller module, we can train a DRL agent without sophisticated physics or 3D modeling. In addition, the modular framework averts daunting retrains of an image-to-action end-to-end neural network, and provides flexibility in transferring the controller to different robots. First, we train a convolutional neural network (CNN) to accurately localize in an indoor setting with dynamic foreground/background. Then, we design a new DRL algorithm named Momentum Policy Gradient (MPG) for continuous control tasks and prove its convergence. We also show that MPG is robust at tracking varying leader movements and can naturally be extended to problems of formation control. Leveraging reward shaping, features such as collision and obstacle avoidance can be easily integrated into a DRL controller.
George Pu, Xiyao Ma, Runhan Sun, Hsi-Yuan Chen, Xiaolin Li 0001
IROS4