Kai Zhang 0039

dblp:55/957-39 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-6322-6096ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
abstract
Low-light image enhancement remains a persistent challenge in computer vision, where state-of-the-art models are often hampered by hardware constraints and computational inefficiency, particularly at high resolutions. While foundational architectures like transformers and diffusion models have advanced the field, their computational complexity limits their deployment on edge devices. We introduce ExpoMamba, a novel architecture that integrates a frequency-aware state-space model within a modified U-Net. ExpoMamba is designed to address mixed-exposure challenges by decoupling the modeling of amplitude (intensity) and phase (structure) in the frequency domain. This allows for targeted enhancement, making it highly effective for real-time applications, including downstream tasks like object detection and segmentation. Our experiments on six benchmark datasets show that ExpoMamba is up to 2-3x faster than competing models and achieves a 6.8\% PSNR improvement, establishing a new state-of-the-art in efficient, high-quality low-light enhancement. Source code: https://www.github.com/eashanadhikarla/ExpoMamba.
Eashan Adhikarla, Kai Zhang 0039, John W. Nicholson 0001, Brian D. Davison 0001
WACV2
2026 Semi-supervised Fetal Brain Parcellation via Hierarchical Learning Framework
Kai Zhang 0039, Fangmei Zhu, Zhongxiang Ding, Geng Chen 0001, Dinggang Shen
Medical Image Anal.2
2026 Super-Resolution Reconstruction of Fetal Brain MRI With Multi-View Interpolation Weight Learning
abstract
Super-resolution reconstruction (SRR) of isotropic fetal brain MR images is critical for prenatal examinations but is hindered by fetal motion and misalignment of thick-slice scans. To address these challenges comprehensively, we introduce an innovative deep learning model, namely 3D-WISE, a 3D Weighted Interpolation for Super-resolution Estimation of fetal brain MRI. The model generates high-quality isotropic fetal brain MR images by learning the interpolation weights to correct misalignments between slices and volumes. These misalignments are estimated by extracting deep features from multiple motion-corrupted stacks. Specifically, 3D-WISE incorporates two key components: (1) a weight learning module for multi-view interpolation and (2) a feature extraction module guided by multi-type attention mechanisms. The weight learning module first maps motion-corrupted thick-slice stacks into latent feature spaces. The resulting features are then fed to an implicit decoding block to estimate interpolation weights of the surrounding points for a given coordinate. We further enhance our approach by incorporating convolutional block attention and atlas-induced cross-attention mechanisms. Extensive experiments on two benchmark datasets show that our 3D-WISE achieves remarkably improved performance compared to the widely adopted registration-reconstruction framework. We also extend the experiments on anatomical structure reconstruction and achieve promising results, highlighting the significant potential of our 3D-WISE for fetal brain MR images SRR in clinical settings.
DengQiang Jia, Kai Zhang 0039, Lingnan Kong, Fangmei Zhu, Zhongxiang Ding, Geng Chen 0001, Dinggang Shen
IEEE J. Biomed. Health Informatics3
2025 Brain-Heart-Gut Guided Multi-constraint Knowledge Distillation for Early Alzheimer's Disease Diagnosis
Shilun Zhao, Shuwei Bai, Kai Zhang 0039, Yin Xu 0001, Ya Zhang 0002, Kaicong Sun, Dinggang Shen
MICCAI (15)5
2025 Scaling up biomedical vision-language models: Fine-tuning, instruction tuning, and multi-modal learning
Cheng Peng 0009, Kai Zhang 0039, Mengxian Lyu, Lichao Sun 0001, Yonghui Wu 0001
J. Biomed. Informatics2
2025 Secure Embedding Aggregation for Cross-Silo Federated Representation Learning
abstract
Representation learning plays a pivotal role in modern applications by enabling high-quality embeddings that support various downstream tasks such as recommendation, clustering, and personalized services. In federated representation learning (FRL), a central server collaborates withNclients, each holding private data, to jointly learn representations of entities (e.g., users in a social network). However, existing embedding aggregation protocols often fall short in either ensuring privacy protections or fully leveraging aggregation opportunities, leaving sensitive data exposed or vulnerable to collusion. To address these challenges, we propose SecEA, a secure embedding aggregation protocol that fully exploits all potential aggregation opportunities across all entities among clients while providing provable privacy guarantees. SecEA defends both local entities and their embeddings—ensuring computational security against a curious server and statistical privacy against up toTN/2 colluding clients. Comprehensive experiments on various representation learning tasks in cross-silo scenarios demonstrate that SecEA incurs a negligible performance loss (within 5%) compared to protocols with weaker or no privacy guarantees, and its additional computational latency significantly diminishes when training deeper models on larger datasets. A parallel mechanism is also included, which helps further improve the efficiency linearly. These results underscore that SecEA not only provides full privacy protections for both entity and embedding, but also preserves the utility of the learned representations.
Jiaxiang Tang, Jinbao Zhu, Kai Zhang 0039, Lichao Sun 0001, Changyu Dong
IEEE Trans. Inf. Forensics Secur.4
2025 Development-Driven Diffusion Model for Longitudinal Prediction of Fetal Brain MRI With Unpaired Data
abstract
Longitudinal magnetic resonance imaging (MRI) is essential for studying the early development of the brain, as it allows to observe and analyze how the brain changes over time. Unfortunately, existing cohort research suffers from lacking sufficient MRI data for studying the development of fetal brains. Apart from research data, a viable alternative is the use of large-scale clinical fetal brain MRI data, which is currently the primary source for longitudinal studies. Although clinical data has several benefits, it is impeded by the inherent drawback of incomplete data. In the context of clinical practice, nearly all subjects undergo only one MRI scan throughout their entire pregnancy, resulting in a lack of longitudinal data for any fetus. To address this issue and obtain longitudinal clinical fetal brain MRI data, we propose to generate MR images for two adjacent gestational weeks (GWs) within one subject, thereby bridging the information gap between three consecutive GWs. This fetal MRI prediction task suffers from two significant challenges, including 1) heterogeneous generation and 2) the lack of paired training data at adjacent GWs. To tackle these two challenges, we propose a new approach, called the Development-driven Diffusion Model (DDM). Specifically, our approach first involves training a conditional diffusion model using population development information spanning all GWs. This allows the model to generate images at various GWs. Next, during the inference stage, we incorporate individual development information of a specific subject using a specially designed perception feature guidance module. The DDM enables the generated 3D MR images to encompass both the general characteristics representative of the targeted GWs, as well as the distinct feature specific to each individual. To assess the efficacy of our approach, extensive experiments were carried out on a large-scale clinical dataset obtained from three different medical centers. The experimental results unequivocally establish the effectiveness of DDM for generating longitudinal MR images of fetal brains.
Kai Zhang 0039, Geng Chen 0001, Fangmei Zhu, Zhongxiang Ding, Dinggang Shen
IEEE Trans. Medical Imaging1
2024 Unified-EGformer: Exposure Guided Lightweight Transformer for Mixed-Exposure Image Enhancement
Eashan Adhikarla, Kai Zhang 0039, Rosaura G. VidalMata, Manjushree B. Aithal, Nikhil Ambha Madhusudhana, John W. Nicholson 0001, Lichao Sun 0001, Brian D. Davison 0001
ICPR (29)2
2024 FedSecurity: A Benchmark for Attacks and Defenses in Federated Learning and Federated LLMs
abstract
This paper introduces FedSecurity, an end-to-end benchmark that serves as a supplementary component of the FedML library for simulating adversarial attacks and corresponding defense mechanisms in Federated Learning (FL). FedSecurity eliminates the need for implementing the fundamental FL procedures, e.g., FL training and data loading, from scratch, thus enables users to focus on developing their own attack and defense strategies. It contains two key components, including FedAttacker that conducts a variety of attacks during FL training, and FedDefender that implements defensive mechanisms to counteract these attacks. FedSecurity has the following features: i) It offers extensive customization options to accommodate a broad range of machine learning models (e.g., Logistic Regression, ResNet, and GAN) and FL optimizers (e.g., FedAVG, FedOPT, and FedNOVA); ii) it enables exploring the effectiveness of attacks and defenses across different datasets and models; and iii) it supports flexible configuration and customization through a configuration file and some APIs. We further demonstrate FedSecurity's utility and adaptability through federated training of Large Language Models (LLMs) to showcase its potential on a wide range of complex applications.
Baturalp Buyukates, Zijian Hu 0001, Weizhao Jin, Lichao Sun 0001, Chulin Xie, Yuhang Yao 0003, Kai Zhang 0039, Qifan Zhang 0002, Carlee Joe-Wong, Amir Salman Avestimehr, Chaoyang He 0001
KDD11
2023 Developing Large Pre-trained Model for Breast Tumor Segmentation from Ultrasound Images
Meiyu Li, Kaicong Sun, Yuning Gu, Kai Zhang 0039, Yiqun Sun, Zhenhui Li, Dinggang Shen
MICCAI (7)4
2023 Adversarial Attack and Defense on Graph Data: A Survey
abstract
Deep neural networks (DNNs) have been widely applied to various applications, including image classification, text generation, audio recognition, and graph data analysis. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. Though there are several works about adversarial attack and defense strategies on domains such as images and natural language processing, it is still difficult to directly transfer the learned knowledge to graph data due to its representation structure. Given the importance of graph analysis, an increasing number of studies over the past few years have attempted to analyze the robustness of machine learning models on graph data. Nevertheless, existing research considering adversarial behaviors on graph data often focuses on specific types of attacks with certain assumptions. In addition, each work proposes its own mathematical formulation, which makes the comparison among different methods difficult. Therefore, this review is intended to provide an overall landscape of more than 100 papers on adversarial attack and defense strategies for graph data, and establish a unified formulation encompassing most graph adversarial learning models. Moreover, we also compare different graph attacks and defenses along with their contributions and limitations, as well as summarize the evaluation metrics, datasets and future trends. We hope this survey can help fill the gap in the literature and facilitate further development of this promising new field We also have created an online resource to keep track of relevant research on the basis of this survey athttps://github.com/safe-graph/graph-adversarial-learning-literature.
Lichao Sun 0001, Yingtong Dou, Carl Yang 0001, Kai Zhang 0039, Ji Wang 0002, Philip S. Yu, Lifang He 0001, Bo Li 0026
IEEE Trans. Knowl. Data Eng.4
2022 Spatial-Temporal Graph Data Mining for IoT-Enabled Air Mobility Prediction
abstract
Big data analytics and mining have the potential to enable real-time decision making and control in a range of Internet of Things (IoT) application domains, such as the Internet of Vehicles, the Internet of Wings, and the Airport of Things. The prediction toward air mobility, which is essential to the studies of air traffic management, has been a challenging task due to the complex spatial and temporal dependencies in air traffic data with highly nonlinear and variational patterns. Existing works for air traffic prediction only focus on either modeling static traffic patterns of individual flight or temporal correlation, with no or limited addressing of the spatial impact, namely, the propagation of traffic perturbation among airports. In this article, we propose to leverage the concept of graph and model the airports as nodes with time-series features and conduct data mining on graph-structured data. To be specific, first, airline on-time performance (AOTP) data is preprocessed to generate a temporal graph data set, which includes three features: 1) the number; 2) average delay; and 3) average taxiing time of departure and arrival flights. Then, a spatial–temporal graph neural networks model is implemented to forecast the mobility level at each airport over time, where a combination of graph convolution and time-dimensional convolution is used to capture the spatial and temporal correlation simultaneously. Experiments on the data set demonstrate the advantage of the model on spatial–temporal air mobility prediction, together with the impact of different priors on adjacency matrices and the effectiveness of the temporal attention mechanism. Finally, we analyze the prediction performance and discuss the capability of our model. The prediction framework proposed in this work has the potential to be generalized to other spatial–temporal tasks in IoT.
Yushan Jiang, Shuteng Niu, Kai Zhang 0039, Chengtao Xu, Dahai Liu, Houbing Song
IEEE Internet Things J.3
2022 Enabling the interpretability of pretrained venue representations using semantic categories
Ning An 0004, Meng Chen 0003, Li Lian, Kai Zhang 0039, Xiaohui Yu 0001, Yilong Yin
Knowl. Based Syst.5
2021 Learning to Detect: A Data-driven Approach for Network Intrusion Detection
abstract
With massive data being generated daily and the ever-increasing interconnectivity of the world’s Internet infrastructures, a machine learning based intrusion detection system (IDS) has become a vital component to protect our economic and national security. In this paper, we perform a comprehensive study on NSL-KDD, a network traffic dataset, by visualizing patterns and employing different learning-based models to detect cyber attacks. Unlike previous shallow learning and deep learning models that use the single learning model approach for intrusion detection, we adopt a hierarchy strategy, in which the intrusion and normal behavior are classified firstly, and then the specific types of attacks are classified. We demonstrate the advantage of the unsupervised representation learning model in binary intrusion detection tasks. Besides, we alleviate the data imbalance problem with SVM-SMOTE oversampling technique in 4-class classification and further demonstrate the effectiveness and the drawback of the oversampling mechanism with a deep neural network as a base model.
Zachary Tauscher, Yushan Jiang, Kai Zhang 0039, Jian Wang 0061, Houbing Song
IPCCC3
2021 Federated Variational Learning for Anomaly Detection in Multivariate Time Series
abstract
Anomaly detection has been a challenging task given high-dimensional multivariate time series data generated by networked sensors and actuators in Cyber-Physical Systems (CPS). Besides the highly nonlinear, complex, and dynamic nature of such time series, the lack of labeled data impedes data exploitation in a supervised manner and thus prevents an accurate detection of abnormal phenomenons. On the other hand, the collected data at the edge of the network is often privacy sensitive and large in quantity, which may hinder the centralized training at the main server. To tackle these issues, we propose an unsupervised time series anomaly detection framework in a federated fashion to continuously monitor the behaviors of interconnected devices within a network and alert for abnormal incidents so that countermeasures can be taken before undesired consequences occur. To be specific, we leave the training data distributed at the edge to learn a shared Variational Autoencoder (VAE) based on Convolutional Gated Recurrent Unit (ConvGRU) model, which jointly captures feature and temporal dependencies in the multivariate time series data for representation learning and downstream anomaly detection tasks. Experiments on three real-world networked sensor datasets illustrate the advantage of our approach over other state-of-the-art models. We also conduct extensive experiments to demonstrate the effectiveness of our detection framework under non-federated and federated settings in terms of overall performance and detection latency.
Kai Zhang 0039, Yushan Jiang, Lee Seversky, Chengtao Xu, Dahai Liu, Houbing Song
IPCCC1
2020 UAV Swarm Communication Aware Formation Control via Deep Q Network
abstract
We propose a DQN based reinforcement learning method in swarm communication constraint based formation control with target searching function. A decentralized communication performance indicator is applied in evaluating the UAV's formation control in simulating a more realistic wireless communication environment between each UAV. A target searching model based on communication aware formation control is presented. The simulation results show that the trained model with state observation space and one thrust action space could be applied in the larger swarm system's group formation and target point tracking.
Chengtao Xu, Kai Zhang 0039, Houbing Song
IPCCC2
2020 Spatio-Temporal Data Mining for Aviation Delay Prediction
abstract
To accommodate the unprecedented increase of commercial airlines over the next ten years, the Next Generation Air Transportation System (NextGen) has been implemented in the USA that records large-scale Air Traffic Management (ATM) data to make air travel safer, more efficient, and more economical. A key role of collaborative decision making for air traffic scheduling and airspace resource management is the accurate prediction of flight delay. There has been a lot of attempts to apply data-driven methods such as machine learning to forecast flight delay situation using air traffic data of departures and arrivals. However, most of them omit en-route spatial information of airlines and temporal correlation between serial flights which results in inaccuracy prediction. In this paper, we present a novel aviation delay prediction system based on stacked Long Short-Term Memory (LSTM) networks for commercial flights. The system learns from historical trajectories from automatic dependent surveillance-broadcast (ADS-B) messages and uses the correlative geolocations to collect indispensable features such as climatic elements, air traffic, airspace, and human factors data along posterior routes. These features are integrated and then are fed into our proposed regression model. The latent spatio-temporal patterns of data are abstracted and learned in the LSTM architecture. Compared with previous schemes, our approach is demonstrated to be more robust and accurate for large hub airports.
Kai Zhang 0039, Yushan Jiang, Dahai Liu, Houbing Song
IPCCC1