Sudong Cai

dblp:259/4988 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-5446-5618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 InfoARD: Enhancing Adversarial Robustness Distillation With Attack-Strength Adaptation and Mutual-Information Maximization
abstract
Adversarial distillation (AD) aims to mitigate deep neural networks' inherent vulnerability to adversarial attacks, thereby providing robust protection for compact models through teacher-student interactions. Despite advancements, existing AD studies still suffer from insufficient robustness due to the limitations of fixed attack strength and attention region shifts. To address these challenges, we propose a strength-adaptive Info-maximizing Adversarial Robustness Distillation paradigm, namely "InfoARD", which strategically incorporates the Attack-Strength Adaptation (ASA) and Mutual-Information Maximization (MIM) to enhance adversarial robustness against adversarial attacks and perturbations. Unlike previous adversarial training (AT) methods that utilize fixed attack strength, the ASA mechanism is designed to capture smoother and generalized classification boundaries by dynamically tailoring the attack strength based on the characteristics of individual instances. Benefiting from mutual information constraints, our MIM strategy ensures the student model effectively learns from various levels of feature representations and attention patterns, thereby deepening the student model's understanding of the teacher model's decision-making processes. Furthermore, a comprehensive multi-granularity distillation is conducted to capture knowledge across multiple dimensions, enabling a more effective transfer of knowledge from the teacher model to the student model. Note that our InfoARD can be seamlessly integrated into existing AD frameworks, further boosting the adversarial robustness of deep learning models. Extensive experiments on various challenging datasets consistently demonstrate the effectiveness and robustness of our InfoARD, surpassing previous state-of-the-art methods.
Ruihan Liu, Jieyi Cai, Yishu Liu 0001, Sudong Cai, Bingzhi Chen, Yulan Guo, Mohammed Bennamoun
IEEE Trans. Image Process.4
2026 Toward Bidirectional Adaptability for Few-Shot Class-Incremental Learning With Forward-Backward Knowledge Transfer
abstract
The development of Deep Neural Networks (DNNs) has enabled AI-driven models to excel in recognizing a limited set of classes within static environments. As AI systems progress, few-shot class-incremental learning (FSCIL) aims to expand their understanding of novel classes from minimal samples while retaining knowledge of previously encountered ones. However, most existing FSCIL models face significant challenges, includinginadequate adaptabilityandcatastrophic forgetting, which hinder their ability to maintain robust forward and backward learning capabilities. To address these issues, this paper proposes a novel Forward-Backward Knowledge Transfer (FBKT) paradigm, which strategically integrates forward distribution adaptation (FDA) and backward semantic alignment (BSA) mechanisms to achieve bidirectional adaptability in knowledge transfer. The FDA mechanism enhances forward adaptability by expanding and reserving the embedding space for new classes using semantic-irrelevant masked images as virtual negative classes, thereby mitigating data overfitting. It also employs self-supervised representation learning to utilize semantic-relevant local embeddings as additional positive samples, fostering class separation and generalization. Meanwhile, the BSA mechanism ensures the semantic consistency of previously learned classes across sessions during class-incremental learning, promoting smoother backward adaptability and reducing model degradation. Extensive experiments conducted on multiple benchmark datasets consistently highlight the superior performance and effectiveness of our FBKT compared to state-of-the-art methods.
Bingzhi Chen, Sudong Cai, Xiaozhao Fang, Mohammed Bennamoun, Shengli Xie 0001
IEEE Trans. Multim.3
2025 Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative Debiasing
abstract
Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias.
Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen
NeurIPS4
2024 AdaShift: Learning Discriminative Self-Gated Neural Feature Activation With an Adaptive Shift Factor
abstract
Nonlinearities are decisive in neural representation learning. Traditional Activation (Act) functions im-pose fixed inductive biases on neural networks with ori-ented biological intuitions. Recent methods leverage self-gated curves to compensate for the rigid traditional Act paradigms in fitting flexibility. However, substantial improvements are still impeded by the norm-induced mis-matched feature recalibrations (see Section 1), i.e., the actual importance of a feature can be inconsistent with its explicit intensity such that violates the basic intention of a direct self-gated feature re-weighting. To address this problem, we propose to learn discriminative neural feature Act with a novel prototype, namely, AdaShift, which en-hances typical self-gated Act by incorporating an adaptive shift factor into the re-weighting function of Act. AdaShift casts dynamic translations on the inputs of are-weighting function by exploiting comprehensive feature-filter context cues of different ranges in a simple yet effective manner. We obtain the new intuitions of AdaShift by rethinking the feature-filter relationships from a common Softmax-based classification and by generalizing the new observations to a common learning layer that encodes features with updatable filters. Our practical AdaShifts, built upon the new Act prototype, demonstrate significant improvements to the popular/SOTA Act functions on different vision benchmarks. By simply replacing ReLU with AdaShifts, ResNets can match advanced Transformer counterparts (e.g., ResNet-50 vs. Swin-T) with lower cost andfewer parameters.
Sudong Cai
CVPR1
2024 RGB road scene material segmentation
abstract
We introduce RGB road scene material segmentation, i.e. , per-pixel segmentation of materials in real-world driving views with pure RGB images , as a novel computer vision task by building a benchmark dataset and by deriving a new method. Our dataset, KITTI-Materials, is based on the well-established KITTI dataset and consists of 1000 frames covering 24 different road scenes of urban/suburban landscapes, carefully annotated with one of 20 material categories for every pixel. It is the first dataset for RGB material segmentation in real driving scenes. Through careful analysis of KITTI-Materials, we identify the extraction and fusion of texture and image context as the key to accurate modeling of road scene material appearance. For this, we introduce R oad scene M aterial S egmentation Net work ( RMSNet ) as a baseline method for this challenging task. RMSNet encodes multi-scale hierarchical features with efficient Transformer layers. We construct the decoder of RMSNet based on a novel efficient self-attention model, which we refer to as SAMixer which adaptively fuses texture and context cues across multiple feature levels. Extensive experiments on KITTI-Materials validate the effectiveness of our RMSNet. We believe our work lays a solid foundation for further studies on RGB road scene material segmentation.
Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino
Image Vis. Comput.1
2023 IIEU: Rethinking Neural Feature Activation from Decision-Making
abstract
Nonlinear Activation (Act) models which help fit the underlying mappings are critical for neural representation learning. Neuronal behaviors inspire basic Act functions, e.g., Softplus and ReLU. We instead seek improved explainable Act models by re-interpreting neural feature Act from a new philosophical perspective of Multi-Criteria Decision-Making (MCDM). By treating activation models as selective feature re-calibrators that suppress/emphasize features according to their importance scores measured by feature-filter similarities, we propose a set of specific properties of effective Act models with new intuitions. This helps us identify the unexcavated yet critical problem of mismatched feature scoring led by the differentiated norms of the features and filters. We present the Instantaneous Importance Estimation Units (IIEUs), a novel class of interpretable Act models that address the problem by re-calibrating the feature with the Instantaneous Importance (II) score (which we refer to as) estimated with the adaptive norm-decoupled feature-filter similarities, capable of modeling the cross-layer and -channel cues at a low cost. The extensive experiments on various vision benchmarks demonstrate the significant improvements of our IIEUs over the SOTA Act models and validate our interpretation of feature Act. By replacing the popular/SOTA Act models with IIEUs, the small ResNet-26s outperform/match the large ResNet-101s on ImageNet with far fewer parameters and computations.
Sudong Cai
ICCV1
2022 RGB Road Scene Material Segmentation
Sudong Cai, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino
ACCV (2)1
2019 Ground-to-Aerial Image Geo-Localization With a Hard Exemplar Reweighting Triplet Loss
abstract
The task of ground-to-aerial image geo-localization can be achieved by matching a ground view query image to a reference database of aerial/satellite images. It is highly challenging due to the dramatic viewpoint changes and unknown orientations. In this paper, we propose a novel in-batch reweighting triplet loss to emphasize the positive effect of hard exemplars during end-to-end training. We also integrate an attention mechanism into our model using feature-level contextual information. To analyze the difficulty level of each triplet, we first enforce a modified logistic regression to triplets with a distance rectifying factor. Then, the reference negative distances for corresponding anchors are set, and the relative weights of triplets are computed by comparing their difficulty to the corresponding references. To reduce the influence of extreme hard data and less useful simple exemplars, the final weights are pruned using upper and lower bound constraints. Experiments on two benchmark datasets show that the proposed approach significantly outperforms the state-of-the-art methods.
Sudong Cai, Yulan Guo, Salman Khan 0001, Jiwei Hu, GongJian Wen
ICCV1