Benteng Ma

dblp:199/1966 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Operationalizing Uncertainty via Structural Coupling for Reliable 3D Dental Point Cloud Segmentation
Benteng Ma
ICIC (1)1
2026 TSAR: A two-stage approach to motion artifact reduction in OCTA images
Benteng Ma, Xiaomeng Li 0001, Dongping Shao, Chubin Ou, Lin An, Kwang-Ting Cheng
Pattern Recognit.1
2026 Topology-Preserving retinal vascular segmentation via sparse persistent homology and MoE convolution
Benteng Ma, Xiaomeng Li 0001, Bin Pu, Kwang-Ting Cheng
Pattern Recognit.1
2026 Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event Prediction
abstract
Deep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD.
Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu
IEEE Trans. Cybern.7
2026 Multi-Scale Collaborative Distillation Graph Neural Networks for Session-Based Recommendation
abstract
Session-based recommendation (SBR) in service computing is pivotal in predicting a user's next action based on their current anonymous session. While Graph Neural Network (GNN)-based methods have shown promise in capturing intricate item transformation relationships within sessions, they often fall short in accurately modeling user preferences. This is primarily due to the common practice of solely considering the last item in the session as the user's current interest, neglecting potentially valuable information embedded in other session items which is essential for capturing user global preferences. Moreover, existing models typically optimize performance solely through cross-entropy loss between predicted items and ground truth labels, while overlooking latent valuable knowledge embedded in intermediate features and item-item relationships that lends support to the model in accurately capturing and modeling user preferences. To address these shortcomings, we propose Multi-Scale Collaborative Distillation (MSCD) for SBR. Our approach introduces a current interest adaptive selection module, which dynamically selects appropriate item embeddings as session-local embeddings by evaluating the importance of each item within the session. This allows for a more accurate capture of the user's current true preferences. Additionally, we propose collaborative knowledge distillation, where multiple models are trained concurrently, enabling the transfer of three types of knowledge including response-based, feature-based, and relationship-based knowledge between models, thereby enriching the model's understanding of user preferences. Experimental evaluations conducted on three popular SBR datasets demonstrate that our MSCD model outperforms recent state-of-the-art methods in terms of recommendation accuracy. Our codes are available at:https://github.com/lonely-ice/MSCD.
Jianping Gou, Youhui Cheng, Benteng Ma, Lan Du 0002, Xin Luo 0001, Zhang Yi 0001
IEEE Trans. Serv. Comput.3
2025 Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure Detection
abstract
In medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods.
Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001
AAAI5
2025 Personalized Federated Side-Tuning for Medical Image Classification
Jiayi Chen 0006, Benteng Ma, Yongsheng Pan, Bin Pu, Hengfei Cui, Yong Xia 0001
MICCAI (14)2
2025 Contrastive Neuron Pruning for Backdoor Defense
abstract
Recent studies have revealed that deep neural networks (DNNs) are susceptible to backdoor attacks, in which attackers insert a pre-defined backdoor into a DNN model by poisoning a few training samples. A small subset of neurons in DNN is responsible for activating this backdoor and pruning these backdoor-associated neurons has been shown to mitigate the impact of such attacks. Current neuron pruning techniques often face challenges in accurately identifying these critical neurons, and they typically depend on the availability of labeled clean data, which is not always feasible. To address these challenges, we propose a novel defense strategy called Contrastive Neuron Pruning (CNP). This approach is based on the observation that poisoned samples tend to cluster together and are distinguishable from benign samples in the feature space of a backdoored model. Given a backdoored model, we initially apply a reversed trigger to benign samples, generating multiple positive (benign-benign) and negative (benign-poisoned) feature pairs from the backdoored model. We then employ contrastive learning on these pairs to improve the separation between benign and poisoned features. Subsequently, we identify and prune neurons in the Batch Normalization layers that show significant response differences to the generated pairs. By removing these backdoor-associated neurons, CNP effectively defends against backdoor attacks while requiring the pruning of only about 1% of the total neurons. Comprehensive experiments conducted on various benchmarks validate the efficacy of CNP, demonstrating its robustness and effectiveness in mitigating backdoor attacks compared to existing methods.
Benteng Ma, Dongnan Liu, Yanning Zhang 0001, Tom Weidong Cai, Yong Xia 0001
IEEE Trans. Image Process.2
2025 Active Learning Based on Temporal Difference of Gradient Flow in Thoracic Disease Diagnosis
abstract
Given the significant advancements in thoracic disease diagnosis due to deep learning, there is a reliance on the availability of numerous annotated samples, which, however, can hardly be guaranteed due to the resource-intensive nature of medical image annotation. Active learning has been introduced to mitigate annotation costs by selecting a subset of uncertain samples for annotation and training. Existing active learning methods encounter two primary challenges: 1) overlooking the impact of samples on the dynamics of model training during data selection, and 2) suffering from high costs of data evaluation and selection. To tackle both issues, we propose a novel metric called Temporal Difference of Gradient Flow (TDGF) for data selection in active learning. Each round of active learning involves three steps: model training, data selection, and data annotation. First, we train a target model, a proxy model, and a historical proxy model on the labeled set. Second, the TDGF scores of unlabeled samples are evaluated based on the surrogate gradient flow, i.e., the TDGF w.r.t the final fully-connected layer between the proxy and historical proxy models, and top-K samples with the highest TDGF scores are selected. Third, the selected samples are annotated, and the labeled pool and unlabeled pool are updated. Comparative experiments have been conducted on two public chest radiograph datasets, i.e., ChestX-ray14 and CheXpert. Our results suggest that the proposed TDGF metric is prone to selecting hard and uncertain samples, and the use of proxy models and surrogate gradient flow substantially reduces the complexity of TDGF calculation. More importantly, the results also indicate that our TDGF-based method outperforms classical and state-of-the-art active learning methods in thoracic disease diagnosis.
Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Jingfeng Zhang, Yong Xia 0001
IEEE J. Biomed. Health Informatics2
2024 Think Twice Before Selection: Federated Evidential Active Learning for Medical Image Analysis with Domain Shifts
abstract
Federated learning facilitates the collaborative learning of a global model across multiple distributed medical in-stitutions without centralizing data. Nevertheless, the ex-pensive cost of annotation on local clients remains an ob-stacle to effectively utilizing local data. To mitigate this issue, federated active learning methods suggest leveraging local and global model predictions to select a relatively small amount of informative local data for annotation. However, existing methods mainly focus on all local data sampled from the same domain, making them un-reliable in realistic medical scenarios with domain shifts among different clients. In this paper, we make the first at-tempt to assess the informativeness of local data derived from diverse domains and propose a novel methodology termed Federated Evidential Active Learning (FEAL) to calibrate the data evaluation under domain shift. Specif-ically, we introduce a Dirichlet prior distribution in both local and global models to treat the prediction as a distribution over the probability simplex and capture both aleatoric and epistemic uncertainties by using the Dirichlet-based evidential model. Then we employ the epistemic uncer-tainty to calibrate the aleatoric uncertainty. Afterward, we design a diversity relaxation strategy to reduce data re-dundancy and maintain data diversity. Extensive experi-ments and analysis on five real multi-center medical image datasets demonstrate the superiority of FEAL over the state-of-the-art active learning methods in federated sce-narios with domain shifts. The code will be available at https://github.com/JiayiChen815/FEAL.
Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Yong Xia 0001
CVPR2
2024 FedEvi: Improving Federated Medical Image Segmentation via Evidential Weight Aggregation
Jiayi Chen 0006, Benteng Ma, Hengfei Cui, Yong Xia 0001
MICCAI (10)2
2024 Enhancing Federated Learning Performance Fairness via Collaboration Graph-Based Reinforcement Learning
Yuexuan Xia, Benteng Ma, Qi Dou 0001, Yong Xia 0001
MICCAI (10)2
2024 VNAS: Variational Neural Architecture Search
Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao
Int. J. Comput. Vis.1
2024 Momentum recursive DARTS
Benteng Ma, Yanning Zhang 0001, Yong Xia 0001
Pattern Recognit.1
2024 Exploratory Training for Universal Lesion Detection: Enhancing Lesion Mining Quality Through Temporal Verification
abstract
Universal lesion detection (ULD) has great value in clinical practice as it can detect various lesions across multiple organs. Deep learning-based detectors have great potential but require high-quality annotated training data. In practice, due to cost, expertise requirements, and the diverse nature of lesions, incomplete annotations are encountered. Directly training ULD detectors under this condition can yield suboptimal results. Leading pseudo-label methods rely on a dynamic lesion-mining mechanism operating at the mini-batch level to address this issue. However, the quality of mined lesions is inconsistent across different iterations, potentially limiting performance enhancement. Inspired by the observation that deep models learn concepts with increasing complexity, we propose an exploratory-training-based ULD (ET-ULD) method to assess the reliability of mined lesions over time. Our approach uses a teacher-student detection model where the teacher mines suspicious lesions, which are then combined with incomplete annotations to train the student. On top of that, we design a bounding-box bank to record the mining timestamps. Each image is trained in several rounds, allowing us to get a sequence of timestamps for the mined lesions. If a mined lesion consistently appears, it is likely to be a true lesion, otherwise, it may just be a noise. This serves as a crucial criterion for selecting reliable mined lesions for retraining. Experimental results show that ET-ULD surpass existing state-of-the-art methods on two distinct lesion image datasets. Notably, on the DeepLesion dataset, ET-ULD achieved a 5.4% improvement in Average Precision (AP) over the previous methods, demonstrating its superior performance.
Geng Chen 0001, Benteng Ma, ChangYang Li, Jingfeng Zhang, Yong Xia 0001
IEEE J. Biomed. Health Informatics3
2023 Federated adaptive reweighting for medical image classification
Benteng Ma, Geng Chen 0001, ChangYang Li, Yong Xia 0001
Pattern Recognit.1
2023 Inter-layer transition in neural architecture search
Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao
Pattern Recognit.1
2022 FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis
abstract
In recent years, the security of AI systems has drawn increasing research attention, especially in the medical imaging realm. To develop a secure medical image analysis (MIA) system, it is a must to study possible backdoor attacks (BAs), which can embed hidden malicious behaviors into the system. However, designing a unified BA method that can be applied to various MIA systems is challenging due to the diversity of imaging modalities (e.g., X-Ray, CT, and MRI) and analysis tasks (e.g., classification, detection, and segmentation). Most existing BA methods are designed to attack natural image classification models, which apply spatial triggers to training images and inevitably corrupt the semantics of poisoned pixels, leading to the failures of attacking dense prediction models. To address this issue, we propose a novel Frequency-Injection based Backdoor Attack method (FIBA) that is capable of delivering attacks in various MIA tasks. Specifically, FIBA leverages a trigger function in the frequency domain that can inject the low-frequency information of a trigger image into the poisoned image by linearly combining the spectral amplitude of both images. Since it preserves the semantics of the poisoned image pixels, FIBA can perform attacks on both classification and dense prediction models. Experiments on three benchmarks in MIA (i.e., ISIC-2019 [4] for skin lesion classification, KiTS-19 [17] for kidney tumor segmentation, and EAD-2019 [1] for endoscopic artifact detection), validate the effectiveness of FIBA and its superiority over stateof-the-art methods in attacking MIA models and bypassing backdoor defense. Source code will be available at code.
Benteng Ma, Jing Zhang 0037, Shanshan Zhao 0001, Yong Xia 0001, Dacheng Tao
CVPR2
2022 GPU-Oriented Designs of Constant False Alarm Rate Detectors for Fast Target Detection in Radar Images
abstract
Constant false alarm rate (CFAR) detector is a class of widely used methods for target detection in radar images. Classical CFAR detectors perform target detection on a pixel-by-pixel basis using certain sliding windows for estimating clutter statistics, which run fast for small images. However, as the image size gets large, the time cost of these detectors will increase significantly since the time complexity with respect toN×N-pixel image isO(N2). In practice, radar images, such as those in synthetic aperture radar (SAR), usually have very large numbers of pixels (which can be on the order of 10000 × 10000), making the classical CFAR detectors very time-consuming when applied to these images. In this paper, we present graphics processing unit (GPU)-oriented Designs for speeding up CFAR detectors, including smallest/greatest-of CFAR and order-statistic CFAR. The proposed designs implement CFAR detectors via tensor operations, including tensor convolution, shift, and boolean operation, which can be fast operated by GPU. Experiment results show that the proposed GPU-oriented CFAR detectors running on a high-performance Nvidia RTX 3090 GPU can be thousands of times faster than the classical CFAR detectors, and realize real-time target detection in large-size radar images. Examples using SAR and range-Doppler images are provided as illustrative applications of the proposed GPU CFAR detectors to target detection in radar images.
Huizhang Yang, Tao Zhang 0027, Yaomin He, Yihua Dan, Junjun Yin 0001, Benteng Ma, Jian Yang 0011
IEEE Trans. Geosci. Remote. Sens.6
2020 Auto Learning Attention
abstract
Attention modules have been demonstrated effective in strengthening the representation ability of a neural network via reweighting spatial or channel features or stacking both operations sequentially. However, designing the structures of different attention operations requires a bulk of computation and extensive expertise. In this paper, we devise an Auto Learning Attention (AutoLA) method, which is the first attempt on automatic attention design. Specifically, we define a novel attention module named high order group attention (HOGA) as a directed acyclic graph (DAG) where each group represents a node, and each edge represents an operation of heterogeneous attentions. A typical HOGA architecture can be searched automatically via the differential AutoLA method within 1 GPU day using the ResNet-20 backbone on CIFAR10. Further, the searched attention module can generalize to various backbones as a plug-and-play component and outperforms popular manually designed channel and spatial attentions for many vision tasks, including image classification on CIFAR100 and ImageNet, object detection and human keypoint detection on COCO dataset. The code will be released.
Benteng Ma, Jing Zhang 0037, Yong Xia 0001, Dacheng Tao
NeurIPS1
2020 Autonomous deep learning: A genetic DCNN designer for image classification
Benteng Ma, Yong Xia 0001, Yanning Zhang 0001
Neurocomputing1