Jin Xiao 0001

dblp:62/4290-1 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Quadruplex-depth based multi-view stereo network with wave-shaped depth cells and Epipolar Transformer
Boyang Song, Jin Xiao 0001, Xiaoguang Hu, Baochang Zhang 0001
Eng. Appl. Artif. Intell.2
2025 Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models
abstract
Diffusion models (DMs) are advanced generative models, yet recent research reveals their vulnerability to backdoor attacks, which establish hidden associations between input patterns and targeted model behavior, potentially causing malicious outputs during inference. These attacks pose significant risks, including model owner reputation damage and harmful content generation. However, existing defense methods often fail against backdoor attacks on diffusion models. To address this gap, we propose Diff-Cleanse, a two-stage defense framework. The first stage introduces a novel trigger inversion method for backdoor detection, and the second stage applies a structural pruning-based method for backdoor removal. Experiments on 373 models poisoned by three state-of-the-art attacks show that Diff-Cleanse achieves > 97% detection accuracy, completely remove backdoors and maintains the models’ benign performance. Code is available at https://github.com/shymuel/diff-cleanse.
Jin Xiao 0001, Xiaoguang Hu, Tianyou Chen
ICME2
2025 When Epipolar Transformers Meets Implicit Neural Super-Resolution in Multi-View Stereo
abstract
Learning-based Multi-View Stereo (MVS) methods heavily rely on feature extraction and cost volume construction to bridge 2D semantics and 3D spatial associations. However, many recent studies have increasingly shifted focus away from these two critical steps, often adopting a multi-stage cascaded framework that can propagate errors from earlier stages. To this end, we introduce TTINS-MVSNet, a novel progressive refinement framework that integrates a one-stage MVS module enhanced by two types of Epipolar Transformer (ET) and subsequent depth optimization modules. Specifically, we incorporate an intra-view ET into context-aware feature extraction and an inter-view ET into visibility-aware cost aggregation. A significant innovation in our approach is the introduction of implicit neural super-resolution module, designed to recover finer details. Experimental results on benchmark datasets demonstrate that our method outperforms all previous approaches with similar structures, highlighting its effectiveness and generalization capability. The code will be available after publication.
Boyang Song, Jin Xiao 0001, Xiaoguang Hu, Guofeng Zhang 0002
ICME2
2025 A three-stage model for camouflaged object detection
Tianyou Chen, Hui Ruan, Jin Xiao 0001, Xiaoguang Hu
Neurocomputing4
2025 An edge-aware high-resolution framework for camouflaged object detection
Jingyuan Ma, Tianyou Chen, Jin Xiao 0001, Xiaoguang Hu, Yingxun Wang
Image Vis. Comput.3
2025 Enhancing point cloud analysis via neighbor aggregation correction based on cross-stage structure correlation
Jin Xiao 0001, Xiaoguang Hu, Boyang Song, Tianyou Chen, Baochang Zhang 0001
Vis. Comput.2
2025 3D Reconstruction based on multi-view stereo in the deep learning era: a survey and comparison of methods
Boyang Song, Jin Xiao 0001, Xiaoguang Hu, Baochang Zhang 0001
Vis. Comput.3
2023 Adaptive fusion network for RGB-D salient object detection
Tianyou Chen, Jin Xiao 0001, Xiaoguang Hu, Guofeng Zhang 0002
Neurocomputing2
2023 Boundary-guided context-aware network for camouflaged object detection
Jin Xiao 0001, Tianyou Chen, Xiaoguang Hu, Guofeng Zhang 0002
Neural Comput. Appl.1
2022 Accurate Instance Segmentation Via Collaborative Learning
abstract
We propose an instance segmentation model, named CoMask, that effectively alleviates the scale variation issue and addresses the precise localization. Specifically, we develop a multi-scale feature extraction module (MSFEM) to exploit multi-scale spatial cues. Besides, the channel attention mechanism is also adopted to further enhance the discriminating ability. Equipped with MSFEMs, multi-scale and multi-level features can be extracted to better characterize objects of various sizes and provide affluent high-level semantic information. For precise localization, we propose a collaborative learning framework to compute coarse masks and regresses position-sensitive dense offsets. The foreground confidence of each position is then assigned as the weight of the corresponding bounding box to calculate a weighted average. Thus, we can mitigate interference of background regions. After obtaining the final regressed bounding boxes, finer foreground masks can be calculated. We conduct experiments on MS COCO dataset. Experimental results validate that CoMask is competitive compared with state-of-the-art models.
Tianyou Chen, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002
ICASSP3
2022 A survey of visual navigation: From geometry to embodied AI
Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002
Eng. Appl. Artif. Intell.3
2022 Implicit neural refinement based multi-view stereo network with adaptive correlation
Boyang Song, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002, Tianyou Chen
Image Vis. Comput.3
2022 Boundary-guided network for camouflaged object detection
Tianyou Chen, Jin Xiao 0001, Xiaoguang Hu, Guofeng Zhang 0002
Knowl. Based Syst.2
2022 CFIDNet: cascaded feature interaction decoder for RGB-D salient object detection
Tianyou Chen, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002
Neural Comput. Appl.3
2022 Spatiotemporal context-aware network for video salient object detection
Tianyou Chen, Jin Xiao 0001, Xiaoguang Hu, Guofeng Zhang 0002
Neural Comput. Appl.2
2021 Canet: Context-Aware Loss for Descriptor Learning
abstract
Research on designing local feature descriptors has gradually shifted to deep learning. Different from other computer vision tasks, the biggest challenge for local descriptor learning lies with the formulation of loss functions. Existing methods solve the problem by leveraging Siamese loss or triplet loss and improve the performance of the learned descriptors by a significant margin. However, the widely used Siamese loss and triplet loss cannot fully utilize the context information. In this paper, we propose a novel loss function to introduce more context information to facilitate training. After incorporating the proposed loss function into training, our learned descriptor demonstrates state-of-the-art performance in patch verification, image matching and patch retrieval benchmarks. The pretrained model will be publicly available at https://github.com/clelouch/CANet.
Tianyou Chen, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002, Hui Ruan
ICASSP3
2021 BPFINet: Boundary-aware progressive feature integration network for salient object detection
Tianyou Chen, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002
Neurocomputing3
2021 BINet: Bidirectional interactive network for salient object detection
Tianyou Chen, Xiaoguang Hu, Jin Xiao 0001, Guofeng Zhang 0002
Neurocomputing3
2017 The invariant features-based target tracking across multiple cameras
Jin Xiao 0001, Xiaoguang Hu
Multim. Tools Appl.1
2012 Quaternion switching filter for impulse noise reduction in color image
Xiaoguang Hu, Jin Xiao 0001
Signal Process.3