EDBT 2026 Demo / reviewers in the wild / expert
Xinlong Ding
dblp:359/7522
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0005-7130-6917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Scale Spatial Channel Joint Representation for General Multi-Modality Image Fusion With Self-SupervisionabstractThe rapid advancement of multi-modality image fusion technology enables researchers to simultaneously acquire information from different modalities within a single fused image. In existing methods, some general approaches can implement both infrared and visible image fusion (IVIF) and medical image fusion (MIF) in the same framework. Nevertheless, these methods often ignore the learning of specific features in different modalities, resulting in unsatisfactory performance in fused results. To overcome this issue, we propose a multi-scale joint framework with self-supervision for general multi-modality image fusion, abbreviated as SCSFusion. It enables more targeted and robust implementation of IVIF and MIF. Specifically, in the fusion network, a joint attention module is employed to parallelly capture self-attention features in spatial and channel domains, which can keep fused results accurate in visual representation. Meanwhile, we utilize source images of different modalities to generate visual-focused maps as pseudo labels for self-supervised training of the fusion results. It effectively preserves the salient details in each fused image from being disrupted by other extracted information. Moreover, a medical dataset with segmentation labels, termed M2DF, is reorganized for fusion and down-stream tasks in MIF. With the help of M2DF, a pre-trained segmentation model can be cascaded with the fusion network, aiming to obtain high-level semantic features from inputs and enhance the data generalization in our general framework. We have conducted extensive experiments and analyses on SCSFusion in M$\rm ^{3}$FD, FMB, and M2DF datasets, respectively. The results indicate that the fused images generated by SCSFusion can not only achieve visually appealing results and superior performance metrics in MIF and IVIF, but also exhibit satisfactory performance in down-stream tasks. Jiawei Li 0016, Jiansheng Chen 0001, Jinyuan Liu 0001, Xinlong Ding, Huimin Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image FusionabstractInfrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate interference on the effectiveness of fusion results. To investigate the robustness of fusion models, in this paper, we propose a novel adversarial attack resilient network, called A2RNet. Specifically, we develop an adversarial paradigm with an anti-attack loss function to implement adversarial attacks and training. It is constructed based on the intrinsic nature of IVIF and provide a robust foundation for future research advancements. We adopt a Unet as the pipeline with a transformer-based defensive refinement module (DRM) under this paradigm, which guarantees fused image quality in a robust coarse-to-fine manner. Compared to previous works, our method mitigates the adverse effects of adversarial perturbations, consistently maintaining high-fidelity fusion results. Furthermore, the performance of downstream tasks can also be well maintained under adversarial attacks. Jiawei Li 0016, Jiansheng Chen 0002, Xinlong Ding, Jinyuan Liu 0001, Bochao Zou, Huimin Ma 0001 |
AAAI | 4 |
| 2025 | Kaleidoscopic Background Attack: Disrupting Pose Estimation With Multi-Fold Radial Symmetry TexturesabstractCamera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric scenarios with sparse inputs, the accuracy of pose estimation can be significantly influenced by background textures that occupy major portions of the images across different viewpoints. In light of this, we introduce the Kaleidoscopic Background Attack (KBA), which uses identical segments to form discs with multi-fold radial symmetry. These discs maintain high similarity across different viewpoints, enabling effective attacks on pose estimation models even with natural texture segments. Additionally, a projected orientation consistency loss is proposed to optimize the kaleidoscopic segments, leading to significant enhancement in the attack effectiveness. Experimental results show that optimized adversarial kaleidoscopic backgrounds can effectively attack various camera pose estimation models. Xinlong Ding, Jiawei Li 0016, Bochao Zou, Huimin Ma 0001 |
ICCV | 1 |
| 2025 | DADet: Safeguarding Image Conditional Diffusion Models Against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection
Xinlong Ding, Jiawei Li 0016, Yudong Zhang 0008, Rongquan Wang, Huimin Ma 0001, Jiansheng Chen 0001 |
ICCV | 2 |
| 2024 | Transferable Adversarial Attacks for Object Detection Using Object-Aware Significant Feature DistortionabstractTransferable black-box adversarial attacks against classifiers by disturbing the intermediate-layer features have been extensively studied in recent years. However, these methods have not yet achieved satisfactory performances when directly applied to object detectors. This is largely because the features of detectors are fundamentally different from that of the classifiers. In this study, we propose a simple but effective method to improve the transferability of adversarial examples for object detectors by leveraging the properties of spatial consistency and limited equivariance of object detectors’ features. Specifically, we combine a novel loss function and deliberately designed data augmentation to distort the backbone features of object detectors by suppressing significant features corresponding to objects and amplifying the surrounding vicinal features corresponding to object boundaries. As such the target object and background area on the generated adversarial samples are more likely to be confused by other detectors. Extensive experimental results show that our proposed method achieves state-of-the-art black-box transferability for untargeted attacks on various models, including one/two-stage, CNN/Transformer-based, and anchor-free/anchor-based detectors. Xinlong Ding, Yining Qin, Huimin Ma 0001 |
AAAI | 1 |
| 2024 | Step Vulnerability Guided Mean Fluctuation Adversarial Attack against Conditional Diffusion ModelsabstractThe high-quality generation results of conditional diffusion models have brought about concerns regarding privacy and copyright issues. As a possible technique for preventing the abuse of diffusion models, the adversarial attack against diffusion models has attracted academic attention recently. In this work, utilizing the phenomenon that diffusion models are highly sensitive to the mean value of the input noise, we propose the Mean Fluctuation Attack (MFA) to introduce mean fluctuations by shifting the mean values of the estimated noises during the reverse process. In addition, we reveal that the vulnerability of different reverse steps against adversarial attacks actually varies significantly. By modeling the step vulnerability and using it as guidance to sample the target steps for generating adversarial examples, the effectiveness of adversarial attacks can be substantially enhanced. Extensive experiments show that our algorithm can steadily cause the mean shift of the predicted noises so as to disrupt the entire reverse generation process and degrade the generation results significantly. We also demonstrate that the step vulnerability is intrinsic to the reverse process by verifying its effectiveness in an attack method other than MFA. Code and Supplementary is available at https://github.com/yuhongwei22/MFA Jiansheng Chen 0001, Xinlong Ding, Yudong Zhang 0008, Ting Tang, Huimin Ma 0001 |
AAAI | 3 |
| 2024 | Enhancing Adversarial Transferability in Object Detection with Bidirectional Feature DistortionabstractPrevious works have shown that perturbing internal-layer features can significantly enhance the transferability of black-box attacks in classifiers. However, these methods have not achieved satisfactory performance when applied to detectors due to the inherent differences in features between detectors and classifiers. In this paper, we introduce a concise and practical untargeted adversarial attack in a label-free manner, which leverages only the feature extracted from the backbone model. By implicitly suppressing the critical feature elements for detection while enhancing the candidate object-relevant elements corresponding to possible detection boxes, we conduct a Bidirectional Feature Distortion Attack (BFDA). Experimental results show that BFDA achieves state-of-the-art black-box transferability on various detector architectures. Xinlong Ding, Huimin Ma 0001 |
ICASSP | 1 |
| 2024 | Invisible Pedestrians: Synthesizing Adversarial Clothing Textures To Evade Industrial Camera-Based 3D DetectionabstractRecent studies have explored attacking perception models of autonomous driving systems and creating adversarial textures to evade 2D vision or infrared pedestrian detectors. However, purely camera-based 3D detectors remain to be explored, especially in complex real-world traffic scenarios and road conditions where pedestrians are prevalent. In this paper, we propose a pipeline that accurately renders 3D human models onto 2D scenes while maintaining physical realism and spatial coordinate constraints. By leveraging this pipeline, we design a soft-weighted confidence loss to optimize adversarial textures on clothing, effectively suppressing predicted boxes near the human model. A comprehensive evaluation of adversarial textures is conducted on a test set comprising 100 scenes with complex environments. Our attack can consistently evade detection by the industrial camera-based 3D detector SMOKE across multiple viewpoints and scenes. Xinlong Ding, Jintai Du, Huimin Ma 0001 |
ICME | 1 |
| 2024 | Analyzing Behavior and Intention in Multi-Agent Systems Using Graph Neural NetworksabstractMulti-agent behavior and intention analysis has been applied in many aspects of our daily lives such as driving maneuver anticipation and assistive driving perception. However, the research in this field suffers from the lack of publicly available datasets, which is mainly caused by the low availability and high complexity of the multi-agent behavior and intention data. In this paper, we propose MBI, a dataset that contains five categories of agents in eight scenarios based on an open-source simulated platform called HarFang3D DogFight SandBox. In MBI, five types of behaviors and five types of intentions are collected. Additionally, we benchmark different Graph Neural Networks (GNNs) on the proposed dataset to test their performance on the tasks of analyzing different behaviors and intentions in multi-agent systems. We also propose a new method called D2RGAT and find it can achieve the best results on MBI. Jintai Du, Xinlong Ding, Jiehui Wu, Huimin Ma 0001 |
ICME | 4 |
| 2023 | Defending Against Universal Patch Attacks by Restricting Token Attention in Vision TransformersabstractPrevious works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and the attention of the network can largely influence the classification result. Adversarial patches usually produce large feature norm for the corresponding shallow token vectors which can attract the attention anomalously. Inspired by this, we propose a restriction operation on the attention matrix, which effectively reduces the influence of the patch region. Experiments on ImageNet validate that our proposal can effectively improve ViT’s robustness towards white-box universal patch attacks while maintaining satisfactory classification accuracy for clean samples. Huimin Ma 0001, Xinlong Ding |
ICASSP | 5 |