Siteng Ma

dblp:321/6893 · DBLP profile ↗
← Back
16ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero-shot learning
Siteng Ma, Haocheng Tang, Jisheng Chu, Wenrui Li 0001
Neurocomputing1
2025 ActiViz: Understanding Sample Selection in Active Learning through Boundary Visualization
abstract
The performance of Active Learning (AL) methods varies widely, influenced by the query strategy, model, and dataset, with the reasons for variation in performance still unclear and insufficiently studied. However, commonly used metrics like accuracy, precision, and recall provide only limited analytical perspectives. No research has effectively uncovered or explained the reasons behind these performance variations, leaving a gap in understanding of the factors that influence the success or failure of AL methods. To address this issue, we propose a novel method and tool leveraging Voronoi Diagrams to visualize AL processes by illustrating interactions between classification decision boundary changes and queried samples across AL iterations. We perform experiments on synthetic and real-world datasets to validate the effectiveness of our method and analyze various AL query strategies. By visualizing the AL process, we illustrate how different query strategies progressively select samples and influence performance in each iteration. This reveals the potential benefits of adapting query strategies at different learning stages to improve active learning efficiency.
Honghui Du, Dairui Liu, Siteng Ma, Brian Mac Namee, Ruihai Dong
CIKM4
2025 Is Complete Labeling Necessary? Understanding Active Learning in Longitudinal Medical Imaging
abstract
Detecting changes in longitudinal medical imaging using deep learning requires a substantial amount of accurately labeled data. However, labeling these images is notably more costly and time-consuming than labeling other image types, as it requires labeling across various time points, where new lesions can be minor, and subtle changes are easily missed. Deep Active Learning (DAL) has shown promise in minimizing labeling costs by selectively querying the most informative samples, but existing studies have primarily focused on static tasks like classification and segmentation. Consequently, the conventional DAL approach cannot be directly applied to change detection tasks, which involve identifying subtle differences across multiple images. In this study, we propose a novel DAL framework, named Longitudinal Medical Imaging Active Learning (LMI-AL), tailored specifically for longitudinal medical imaging. By pairing and differencing all 2D slices from baseline and follow-up 3D images, LMI-AL iteratively selects the most informative pairs for labeling using DAL, training a deep learning model with minimal manual annotation. Experimental results demonstrate that, with less than 8% of the data labeled, LMI-AL can achieve performance comparable to models trained on fully labeled datasets. We also provide a detailed analysis of the method’s performance, as guidance for future research. The code is publicly available at https://github.com/HelenMa9998/LongitudinalAL.
Siteng Ma, Honghui Du, Prateek Mathur, Brendan S. Kelly, Ronan P. Killeen, Aonghus Lawlor, Ruihai Dong
IJCNN1
2025 Spatial Aggregation for Semi-supervised Active Learning in 3D Medical Image Segmentation
Siteng Ma, Honghui Du, Dairui Liu, Kathleen M. Curran, Aonghus Lawlor, Ruihai Dong
MICCAI (8)1
2024 Masked Angle-Aware Autoencoder for Remote Sensing Images
Zhihao Li 0005, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
ECCV (8)3
2024 Breaking the Barrier: Selective Uncertainty-Based Active Learning for Medical Image Segmentation
abstract
Active learning (AL) has found wide applications in medical image segmentation, aiming to alleviate the annotation workload and enhance performance. Conventional uncertainty-based AL methods, such as entropy and Bayesian, often rely on an aggregate of all pixel-level metrics. However, in imbalanced settings, these methods tend to neglect the significance of target regions, eg., lesions, and tumors. Moreover, uncertainty-based selection introduces redundancy. These factors lead to unsatisfactory performance, and in many cases, even underperform random sampling. To solve this problem, we introduce a novel approach called the Selective Uncertainty-based AL, avoiding the conventional practice of summing up the metrics of all pixels. Through a filtering process, our strategy prioritizes pixels within target areas and those near decision boundaries. This resolves the aforementioned disregard for target areas and redundancy. Our method showed substantial improvements across five different uncertainty-based methods and two distinct datasets, utilizing fewer labeled data to reach the supervised baseline and consistently achieving the highest overall performance. Our code is available at https://github.com/HelenMa9998/Selective_Uncertainty_AL.
Siteng Ma, Haochang Wu, Aonghus Lawlor, Ruihai Dong
ICASSP1
2024 Adaptive Curriculum Query Strategy for Active Learning in Medical Image Classification
Siteng Ma, Honghui Du, Kathleen M. Curran, Aonghus Lawlor, Ruihai Dong
MICCAI (11)1
2024 Weakly supervised object localization via knowledge distillation based on foreground-background contrast
Siteng Ma, Biao Hou, Zhihao Li 0005, Zitong Wu, Xianpeng Guo, Chen Yang 0027, Licheng Jiao
Neurocomputing1
2023 DO-FAM: Disentangled Non-Linear Latent Navigation For Facial Attribute Manipulation
abstract
Facial attribute manipulation (FAM) aims to edit the semantic attributes of facial images according to the user’s requirements. Unfortunately, the majority of existing FAM methods struggle in meeting at least one of the two requirements: high reconstruction quality and high irrelevance preservation. To alleviate these two limitations, we propose a novel Disentangled nOn-linear latent navigation framework for FAM, termed DO-FAM. To promote the reconstruction quality, we leverage hypernetworks to fine-tune a pre-trained StyleGAN2 generator. To decouple entangled attributes, we propose a novel Disentangled nOn-Linear Latent transformation module, named DOLL, which consists of three components: (1) a decomposer to factorize input latent codes into two parts: attribute-related and attribute-unrelated; (2) a non-linear Latent Transformation Network (LTNet) to navigate the attribute-related latent codes to the target one with respect to the designed attribute(s); and (3) a latent classifier tasked with predicting latent codes’ attributes to guide the latent code navigation. Extensive experimental results on a widely-used benchmark facial editing dataset, CelebA-HQ, demonstrate the superiority of our method over state-of-the-art methods.
Yifan Yuan 0001, Siteng Ma, Hongming Shan, Junping Zhang
ICASSP2
2023 Adaptive Nonlinear Latent Transformation for Conditional Face Editing
abstract
Recent works for face editing usually manipulate the latent space of StyleGAN via the linear semantic directions. However, they usually suffer from the entanglement of facial attributes, need to tune the optimal editing strength, and are limited to binary attributes with strong supervision signals. This paper proposes a novel adaptive nonlinear latent transformation for disentangled and conditional face editing, termed AdaTrans. Specifically, our AdaTrans divides the manipulation process into several finer steps; i.e., the direction and size at each step are conditioned on both the facial attributes and the latent codes. In this way, AdaTrans describes an adaptive nonlinear transformation trajectory to manipulate the faces into target attributes while keeping other attributes unchanged. Then, AdaTrans leverages a predefined density model to constrain the learned trajectory in the distribution of latent codes by maximizing the likelihood of transformed latent code. Moreover, we also propose a disentangled learning strategy under a mutual information framework to eliminate the entanglement among attributes, which can further relax the need for labeled data. Consequently, AdaTrans enables a controllable face editing with the advantages of disentanglement, flexibility with non-binary attributes, and high fidelity. Extensive experimental results on various facial attributes demonstrate the qualitative and quantitative effectiveness of the proposed AdaTrans over existing state-of-the-art methods, especially in the most challenging scenarios with a large age gap and few labeled examples. The source code is available at https://github.com/Hzzone/AdaTrans.
Zhizhong Huang, Siteng Ma, Junping Zhang, Hongming Shan
ICCV2
2023 Adaptive Adversarial Samples Based Active Learning for Medical Image Classification
Siteng Ma, Aonghus Lawlor, Ruihai Dong
ICPRAM1
2023 Contrastive Learning Based on Multiscale Hard Features for Remote-Sensing Image Scene Classification
abstract
The overwhelming majority of models for remote sensing image (RSI) scene classification generally require the weights pre-trained on natural images for initialization before formal training. However, differences in imaging mechanisms lead to huge discrepancies between natural images and RSIs, and the strong visual representation learned from massive natural images limits the performance of models when inferencing RSIs. To address this issue, the well-established self-supervised contrastive learning paradigm in the natural image field is introduced to the RSI field. We propose a contrastive learning method based on multi-scale hard features, MHCL, which aims to use finite RSIs to learn sufficient visual representations in an unsupervised contrastive manner, thus provide a powerful upstream pre-trained model for fine-tuning downstream scene classification task. Multi-level features extracted by intermediate layers of each encoder’s backbone are first gathered, and then a hard features transformation method is proposed to create hard positive features and diverse queues that save hard negatives, thereby enriching the finite scene information in small-scale RSIs. Furthermore, we redesign the multi-scale hard features joint contrastive loss to boost the model to explore sufficient invariant representations by additionally pulling hard positive pairs closer and pushing hard negative pairs farther away in the embedding space. Extensive experiments demonstrate that the upstream pre-training model generated by MHCL achieves competitive transferred performance on three popular scene classification datasets, outperforming the traditional model pre-trained on ImageNet and models pre-trained by other state-of-the-art contrastive learning methods. Our code will be released at: https://github.com/benesakitam/MHCL.
Zhihao Li 0005, Biao Hou, Xianpeng Guo, Siteng Ma, Yanyu Cui, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2023 Unsupervised Prototype-Wise Contrastive Learning for Domain Adaptive Semantic Segmentation in Remote Sensing Image
abstract
Labeling data in the field of remote sensing is time-consuming and labor-intensive, making domain adaptation between different domains an urgently needed solution. To address the domain gap between diverse datasets in the remote sensing domain, numerous methods tailored for domain adaptation in high-resolution remote sensing imagery have emerged. Some of the existing methods focus on reducing the domain gap at either the feature level or the pixel level, often overlooking their underlying connection. To tackle this issue, we introduce a prototype-wise contrastive feature alignment paradigm (PCFA) aimed at bridging the representations between the feature and pixel levels. By dynamically updating, we acquire prototype information encompassed by different mini-batches and employ an optimal transport mechanism to reasonably apply the prototype feature distribution in guiding the learning of target domain features. We conduct extensive domain adaptation semantic segmentation (DASS) experiments on the ISPRS Vaihingen and Potsdam datasets, achieving an improvement about 4%~5% in mIoU (mean Intersection over Union) compared to previous methods using the DeepLabV2 framework.
Siteng Ma, Biao Hou, Xianpeng Guo, Zitong Wu, Zhihao Li 0005, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2023 Automatic Aug-Aware Contrastive Proposal Encoding for Few-Shot Object Detection of Remote Sensing Images
abstract
In the annotation of remote sensing images (RSIs), the effectiveness of common object detection methods trained on only a few samples decreases instantly, which has prompted increasing research on the few-shot problem in remote sensing. RSIs often exhibit suboptimal performance in few-shot scenarios due to the intricate nature of scene information interference and the high degree of cosine similarity, both of which present significant challenges to their effectiveness. In this paper, a two-stage detection framework based on fine-tuning is selected to deal with the common problems in the few-shot task of remote sensing domain. Considering the excessive scale variation of instances in remote sensing datasets, we introduce an automatically learned aug-aware search module to provide an intelligent data augmentation solution for Faster R-CNN using different optimal augmentation policies searched by the network to fit the current dataset. We introduce a contrastive RoI branch to better classify novel class proposal features that are easily confused by the base class. We named our work AACE and conducted extensive experiments on two common object detection datasets in remote sensing, NWPU VHR 10 and DIOR, on which AACE achieved about 2.30% and 2.61% improvement, respectively, in the number of shots listed in the paper, compared to other algorithms.
Siteng Ma, Biao Hou, Zitong Wu, Zhihao Li 0005, Xianpeng Guo, Bo Ren 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 VR-FAM: Variance-Reduced Encoder with Nonlinear Transformation for Facial Attribute Manipulation
abstract
Facial attribute manipulation (FAM) aims to infer desired facial images by modifying specific attributes while keeping others unchanged. Existing works suffer from the entanglement of facial attributes, leading to unexpected artifacts and the loss of facial identity information after editing. To alleviate these issues, we propose a novel FAM framework based on StyleGAN, termed VR-FAM, which can meet the requirements of FAM—editing ability, distortion, and fidelity. First, we propose a variance-reduced encoder to make the latent space close to the one of StyleGAN. Second, we present a nonlinear latent transformation network, which can convert the source latent code to target latent code in line with the nonlinear latent space of StyleGAN. Experimentally, we evaluate the proposed FAM framework on the benchmark FFHQ dataset and demonstrate the improvement gain over the recently published models in terms of edit accuracy and fidelity.
Siteng Ma, Junping Zhang
ICASSP2
2022 Multicrop Fusion Strategy Based on Prototype Assignment for Remote Sensing Image Scene Classification
abstract
The gap between self-supervised visual representation learning and supervised learning is gradually closing. Self-supervised learning does not rely on a large amount of labeled data and reduces the loss of human labeled information. Compared with natural images, remote sensing images require rich samples and human annotation by experts. Moreover, many algorithms have poor interpretability and unconvincing results. Therefore, this paper proposes a self-supervised method based on prototype assignment by designing a pretext task so that the network maps features to prototypes in the process of learning, swaps the code corresponding to the obtained features, combines them with another data-enhancing feature, and then optimizes the network. The prototype is introduced to explain the clustering idea embodied in the whole process. Considering the existence of the scene information-rich characteristic of remote sensing images, we introduce multiple views with different resolutions to capture more detailed information on the images. Finally, if the data enhancement method is not powerful enough, the network can easily fall into an overfitting state, which prevents the network from learning subtle differences and detailed information. To address this shortcoming, we propose a fusion strategy to flatten the decision boundary of the framework so that the model can also learn the soft similarity between sample pairs. We name the whole framework MFPC. In extensive experiments conducted on three common remote sensing image datasets (i.e., UCMerced, AID, and NWPU45), MFPC achieves a maximum improvement of 4.3% over some existing self-supervised algorithms, indicating that it can achieve good results.
Siteng Ma, Biao Hou, Xianpeng Guo, Zhihao Li 0005, Zitong Wu, Shuang Wang 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1