Ziyun Cai

dblp:179/6081 · DBLP profile ↗
← Back
38ranked-venue papers
22as first author
30since 2021 · last 2026
0000-0001-6822-915XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 12 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Unsupervised domain adaptation without source data for visual classification via adaptive confidence-driven mechanism
Ziyun Cai, Jie Song 0014, Yawen Huang, Changhui Hu 0001
Expert Syst. Appl.1
2026 Irrelevant feature filtering module for deep multi-view generative clustering
Xiaoyuan Jing, Yong-Fang Yao, Wei Liu 0200, Fei Wu 0004, Changhui Hu 0001, Ziyun Cai
Inf. Sci.7
2026 Compressive self-attention transformer for low-light enhancement and zero-element pixels restoration
Changhui Hu 0001, Donghang Jing, Kerui Hu, Tiesheng Chen, Ziyun Cai, Fei Wu 0004, Xiaoyuan Jing
Pattern Recognit.6
2025 Enhancing Federated Domain Adaptation via Multi-Granular Fine-Grained Alignment
abstract
Traditional unsupervised multi-source domain adaptation usually assumes that all source domain data can be utilized during training. Unfortunately, due to practical concerns such as privacy, data storage, and computational costs, data from different source domains are often isolated from each other. To address this issue, we propose a federated domain adaptation framework based on fine-grained alignment. This method achieves domain adaptation at the model level through iterative training of source and target domains, thereby avoiding the direct use of source domain data. Specifically, our approach employs specialized techniques at various stages—model construction, pseudo-label generation, and model training—to handle fine-grained features that are often overlooked. This enables the model to effectively remove irrelevant information and learn more discriminative features, thus narrowing the distribution gap between domains. Extensive experimental results demonstrate the effectiveness of our proposed method across multiple datasets.
Ziyun Cai, Shangshang Song, Jie Song 0014, Yawen Huang, Changhui Hu 0001, Xiaoyuan Jing
ICASSP1
2025 Source-Free Domain Adaptation via Transformer-based Object-centric Perception
abstract
In this paper, we investigate the Source-Free Domain Adaptation (SFDA), where a well-trained model adapts to an unlabeled target domain without access to source data. Previous SFDA methods mainly relied on convolutional neural networks, which struggle with domain shifts due to their local focus. To address this, we propose the Object-centric Perception Source-Free Transformer (OP-SFT), which leverages the self-attention mechanism of Transformers to focus on relevant target regions, improving adaptability to domain shifts. We also introduce self-supervised knowledge distillation to enhance semantic perception and a confidence-based k-means clustering method for more accurate pseudo-label generation. Extensive experiments demonstrate that our OP-SFT achieves significant adaptation performance across four widely-used domain adaptation benchmark datasets compared to other state-of-the-art baselines. The code is available at https://github.com/Weilong-Gao/OP-SFT.
Ziyun Cai, Weilong Gao, Yawen Huang, Jie Song 0014, Changhui Hu 0001, Tengfei Zhang 0001
ICME1
2025 Make Multi-source Task Greater Again: Adaptive Causal Diffusion Strategy
abstract
Multi-source Domain Adaptation (MSDA) aims to adapt models trained on multiple labeled source domains to an unlabeled target domain. Recent MSDA methods based on Generative Adversarial Networks (GANs) implicitly capture the image distribution, which can lead to limited sample fidelity and result in misalignment of pixel-level information between the sources and the target domain. Moreover, when samples from different sources interact during training, significant misalignment across various source domains can occur. In this study, we introduce a novel MSDA framework called Adaptive Causal Diffusion Networks (ACDN) to address these challenges. ACDN integrates a diffusive domain adaptation model for effective, high-fidelity adaptation between the source and target domains, incorporating Granger-causal inference to ensure that the assigned weights for each source domain are closely related to their respective contributions to the decision-making process. Experimental results show that ACDN outperforms existing methods significantly across real-world domain adaptation benchmarks.
Ziyun Cai, Yawen Huang, Jie Song 0014, Changhui Hu 0001, Tengfei Zhang 0001
ICME1
2025 Multi-Scale Tubularity-Aware U-Net
abstract
U-Net architectures have made great progress in dealing with semantic segmentation tasks. However, existing frameworks have not yet possessed the ability of capturing sufficient local and contextual dependencies of tubular structures. The reasons are two-fold. First, traditional square convolutions are inherently limited to model irregular pixel changes due to their fixed geometric structures. Second, there exist semantic gaps among the multi-scale tubularity features and between stages of their encoding and the decoding. To mitigate these issues, we propose multi-scale tubularity-aware U-Net, by coupling a novel tubularity deformable convolution (TdConv) embedding and a dual attention Transformer (DaTrans) alternative to skip connection. On the one hand, TdConv embedding iteratively learns the deformation offsets of convolution itself in both directions along the tubular structure. On the other hand, DaTrans connection endows skip connections with attention mechanism from both multi-scale local pixel and cross-scale global semantic perspectives. Hinging on the local irregularity perception and the global semantic association, our method enables to analyze tubular structures appeared in complex contexts and at different scales. Extensive experiments show that our approach outperforms state-of-the-art techniques, including different U-Net variants, for various datasets on several tasks including road extraction and vessel segmentation.
Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang
ICME3
2025 Adaptive margin for unsupervised domain adaptation without source data
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Changhui Hu 0001, Xiaoyuan Jing
Comput. Vis. Image Underst.1
2025 Multi-Source Domain Adaptation by Causal-Guided Adaptive Multimodal Diffusion Networks
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Yefeng Zheng 0001, Dong Yue 0001
Int. J. Comput. Vis.1
2025 Sample-pair learning network for extremely imbalanced classification
abstract
In data classification, class-balanced data is ideal, but real datasets are often imbalanced, necessitating rebalancing through methods like resampling. In recent years, some new generative model-based resampling methods have been proposed. However, when facing extreme class imbalance, where the minority class is strongly underrepresented and on its own does not contain enough information to conduct the generative process. Some deep learning methods have been proposed to solve extremely imbalanced classification problems, but some of them are only used for specific datasets. Therefore, we proposed a novel deep learning method that combines a generative strategy with multi-task joint learning, termed sample-pair learning network (SPLN), for extremely imbalanced classification. The network consists of data preprocessing and multi-task joint learning modules. During data preprocessing, the training set is expanded by constructing positive and negative sample-pairs, then rebalanced using a strategy combining attention and resampling, termed undersampling based on attention power values (APVUS). The multi-task joint learning module employs a Siamese convolutional subnetwork to measure the similarity between sample-pairs and a multi-layer perceptron to recognize the category of single samples. The module can reduce the risk of overfitting caused by excessive noise in the training set. Finally, we designed a voting model based on the Siamese convolutional subnetwork to infer the categories of test samples. Experimental results demonstrate that our approach outperforms state-of-the-art generative model-based methods and is effective and general for extremely imbalanced classification.
Linjun Chen, Xiaoyuan Jing, Runhang Chen, Fei Wu 0004, Yongchang Ding, Changhui Hu 0001, Ziyun Cai
Neurocomputing7
2025 Tracking Mamba for Road Extraction From Satellite Imagery
abstract
Automated road extraction from satellite imagery for dynamic map updating has become a crucial research focus in remote sensing, where existing state-of-the-art Transformer-based methods exhibit two critical limitations: (1) inadequate precision in capturing tubular road patterns and (2) suboptimal computational efficiency on standard GPUs. To address these challenges, we propose TrMamba, a new Tracking-based Mamba architecture that combines the original Mamba’s efficiency with enhanced tubular road pattern recognition through two key innovations: a tubular road pattern tracking mechanism for continuous road feature extraction and a tracking selective scanning module for directional context modeling via adaptive attention. By integrating a novel tubular tracking mechanism into the Mamba’s selective scanning process, TrMamba fundamentally improves the original paradigm and achieves superior road topology encoding, as demonstrated by extensive experiments showing state-of-the-art performance across multiple remote sensing benchmarks in both accuracy and computational efficiency. The source code is available at: https://github.com/Apheliosa/TrMamba.
Jie Song 0014, Ziyun Cai, Liang Xiao 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 UPT-Flow: Multi-scale transformer-guided normalizing flow for low-light image enhancement
Lintao Xu, Changhui Hu 0001, Xiaoyuan Jing, Ziyun Cai, Xiaobo Lu
Pattern Recognit.5
2025 DUSA-UNet: Dual Sparse Attentive U-Net for Multiscale Road Network Extraction
abstract
The challenges of road network segmentation demand an algorithm capable of adapting to the sparse and irregular shapes, as well as the diverse context, which often leads traditional encoding-decoding methods and simple Transformer embeddings to failure. We introduce a computationally efficient and powerful framework for elegant road-aware segmentation. Our method, called DUSA-UNet, effectively encodes fine-grained local road connectivity and holistic global topological semantics while decoding multiscale road network information. DUSA-UNet offers a novel alternative to the U-Net architecture by integrating connectivity attention, which can exploit intra-road interactions across multi-level sampling features with reduced computational complexity. This local interaction serves as valuable prior information for learning global interactions between road networks and the background through another integrality attention mechanism. The two forms of sparse attention are arranged alternatively and complementarily, and trained jointly, resulting in performance improvements without significant increases in computational complexity. Extensive experiments on various datasets with different resolutions, including Massachusetts, DeepGlobe, SpaceNet, and Large-Scale remote sensing images, demonstrate that DUSA-UNet outperforms state-of-the-art techniques. Our approach represents a significant advancement in the field of road network extraction, providing a computationally feasible solution that achieves high-quality segmentation results.
Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang, Yefeng Zheng 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 JTE-CFlow for Low-Light Enhancement and Zero-Element Pixels Restoration With Application to Night Traffic Monitoring Images
abstract
We observe that the low-light RGB images, as well as night traffic monitoring (NTM) images, contain lots of color pixels with zeros caused by the low-light, which means that the low-light images suffer both information weakness and information loss of zero-element pixels. In this paper, we propose a novel flow-based generative method JTE-CFlow for low-light image enhancement, which consists of a joint-attention transformer based conditional encoder (JTE) and a map-wise cross affine coupling flow (CFlow). Specifically, JTE executes short-range and long-range operations by RRDBs (i.e., residual-in-residual dense blocks) and JATs (i.e., joint-attention transformer blocks) in series connection. JAT achieves weak information amplification and information loss restoration of zero-element pixels by the integration of self-attention and specific-attention with sharing the same value vectors, where the query and key vectors of specific-attention are from the zero-map feature of the low-light image. On the other hand, CFlow develops a map-wise cross affine coupling (MCAC) layer to perform cross learning for the flow feature, and a multiplication coupling network (MCN) to learn the transformation parameters of MCAC. JTE-CFlow learns to map the subtraction of outputs of CFlow and JTE (i.e., the residual code) into a standard normal distribution, and the inverse network of CFlow takes the latent feature of the low-light image as its input to infer the enhanced image. Experiments show that JTE-CFlow outperforms most SOTA methods on 7 mainstream low-light datasets with the same architecture, and can be applied to enhance NTM images. The source code and pre-trained models are available athttps://github.com/NJUPT-IPR-HuYin/JTE-CFlow.
Changhui Hu 0001, Lintao Xu, Yanyong Guo, Ziyun Cai, Xiaoyuan Jing, Pan Liu 0013
IEEE Trans. Intell. Transp. Syst.5
2024 Local weight coupled network: multi-modal unequal semi-supervised domain adaptation
Ziyun Cai, Jie Song 0014, Tengfei Zhang 0001, Changhui Hu 0001, Xiaoyuan Jing
Multim. Tools Appl.1
2024 Swin transformer and ResNet based deep networks for low-light image enhancement
Lintao Xu, Changhui Hu 0001, Fei Wu 0004, Ziyun Cai
Multim. Tools Appl.5
2024 Attention Cycle-consistent universal network for More Universal Domain Adaptation
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Xiaoyuan Jing, Yefeng Zheng 0001, Ling Shao 0001
Pattern Recognit.1
2023 Learnable Snake R-CNN for Instance-Level Biomedical Image Segmentation
abstract
Precisely knowing each instance’s position and extents is a critical first step in many biological applications. State-of-the-art techniques rely either on deep learning models designed to predict segmentation masks on each Region of Interest (RoI) or on classic active contour methods. The former struggles to precisely delineating boundaries and tends to output masks at low resolutions when the cells/nuclei are very irregular while the latter often needs good initialization and manual setting of parameters, thus limiting their usefulness. To bridge this gap, we introduce Snake R-CNN, a new level of the learnable active contour model that predict boundary on each RoI in a sequent way. To do so, for each RoI, we reformulate the contour deformation task in terms of a hidden state evolution problem and update the evolution process using energy minimization. We learn snake parameterizations per instance in an end-to-end manner, and demonstrate its effectiveness for contour inferences of various cell/nucleus types where consistently higher performances were obtained for comparison against state-of-the-arts.
Jie Song 0014, Ziyun Cai, Yurong Song, Guoping Jiang, Zhichao Lian, Liang Xiao 0001
ICIP2
2023 Flow Learning Based Dual Networks for Low-Light Image Enhancement
Changhui Hu 0001, Weilin Yi, Ziyun Cai, Mingliang Zhai, Wankou Yang
Neural Process. Lett.4
2023 Domain embedding transfer for unequal RGB-D image recognition
Ziyun Cai, Xiaoyuan Jing, Ling Shao 0001
Pattern Recognit.1
2023 Single-/Multi-Source Domain Adaptation via domain separation: A simple but effective method
Ziyun Cai, Tengfei Zhang 0001, Changhui Hu 0001, Xiaoyuan Jing
Pattern Recognit. Lett.1
2022 Dual Re-Weighting Network for Multi-Source Domain Adaptation
abstract
In this paper, we propose a novel framework called Du-al Re-weighting Multi-source Network (DRMN) to address the task of Multi-source Domain Adaptation (MSDA). Two challenges exist in MSDA: i) the domain discrepancies a-mong the multiple source domains, and ii) the domain mis-match between target and source domains. We propose du-al re- weighting mechanisms including source distribution re-weighting and sample selected re-weighting. Source distribution re-weighting mechanism can match the estimated source label distribution and the unknown target label distribution to adapt the classifier. Sample selected re-weighting mechanism can select highly confident target data as pseudo-labeled sam-ples to integrate the information from different sources, and further improve the classification performance. We find that DRMN can show competitive performance with respect to the state-of-the-art on different real-world datasets.
Ziyun Cai, Tengfei Zhang 0001, Xiaoyuan Jing
ICME1
2022 Unequal adaptive visual recognition by learning from multi-modal data
Ziyun Cai, Tengfei Zhang 0001, Xiaoyuan Jing, Ling Shao 0001
Inf. Sci.1
2022 Dual contrastive universal adaptation network for multi-source visual recognition
Ziyun Cai, Tengfei Zhang 0001, Fumin Ma, Xiaoyuan Jing
Knowl. Based Syst.1
2022 A survey of deep domain adaptation based on label set classification
Ziyun Cai, Tengfei Zhang 0001, Baoyun Wang
Multim. Tools Appl.2
2022 Co-embedding: a semi-supervised multi-view representation learning approach
Xiaodong Jia 0005, Xiaoyuan Jing, Xiaoke Zhu, Ziyun Cai, Changhui Hu 0001
Neural Comput. Appl.4
2022 Visual-Depth Matching Network: Deep RGB-D Domain Adaptation With Unequal Categories
abstract
Existing domain adaptation (DA) methods generally assume that different domains have identical label space, and the training data are only sampled from a single domain. This unrealistic assumption is quite restricted for real-world applications, since it neglects the more practical scenario, where the source domain can contain the categories that are not shared by the target domain, and the training data can be collected from multiple modalities. In this article, we address a more difficult but practical problem, which recognizes RGB images through training on RGB-D data under the label space inequality scenario. There are three challenges in this task: 1) source and target domains are affected by the domain mismatch issue, which results in that the trained models perform imperfectly on the test data; 2) depth images are absent in the target domain (e.g., target images are captured by smartphones), when the source domain contains both the RGB and depth data. It makes the ordinary visual recognition approaches hardly applied to this task; and 3) in the real world, the source and target domains always have different numbers of categories, which would result in a negative transfer bottleneck being more prominent. Toward tackling the above challenges, we formulate a deep model, called visual-depth matching network (VDMN), where two new modules and a matching component can be trained in an end-to-end fashion jointly to identify the common and outlier categories effectively. The significance of VDMN is that it can take advantage of depth information and handle the domain distribution mismatch under label inequality simultaneously. The experimental results reveal that VDMN exceeds the state-of-the-art performance on various DA datasets, especially under the label inequality scenario.
Ziyun Cai, Xiaoyuan Jing, Ling Shao 0001
IEEE Trans. Cybern.1
2021 Grand Unified Domain Adaptation
Ziyun Cai, Tengfei Zhang 0001, Xiaoyuan Jing, Ling Shao 0001
BMVC1
2021 Dual Contrastive Universal Adaptation Network
abstract
We study Universal Domain Adaptation (UniDA) problem, which is recently proposed. Different from existing domain adaptation (DA) methods, e.g., Closed set, Open set and Partial DA, UniDA does not need any prior knowledge about the overlap across the source and target label sets. We have two challenges in UniDA problem: i) Domain shift. ii) Category shift. Towards tackling above challenges, we formulate a universal adaptation network called Dual Contrastive Network (DCN), where a contrastive module and a transferability rule are included. The experimental results reveal that DCN can work stably on different UniDA settings and exceeds the state-of-the-art performance across five benchmarks against existing DA methods.
Ziyun Cai, Jie Song 0014, Tengfei Zhang 0001, Xiaoyuan Jing, Ling Shao 0001
ICME1
2021 Semi-Supervised Multi-View Deep Discriminant Representation Learning
abstract
Learning an expressive representation from multi-view data is a key step in various real-world applications. In this paper, we propose a semi-supervised multi-view deep discriminant representation learning (SMDDRL) approach. Unlike existing joint or alignment multi-view representation learning methods that cannot simultaneously utilize the consensus and complementary properties of multi-view data to learn inter-view shared and intra-view specific representations, SMDDRL comprehensively exploits the consensus and complementary properties as well as learns both shared and specific representations by employing the shared and specific representation learning network. Unlike existing shared and specific multi-view representation learning methods that ignore the redundancy problem in representation learning, SMDDRL incorporates the orthogonality and adversarial similarity constraints to reduce the redundancy of learned representations. Moreover, to exploit the information contained in unlabeled data, we design a semi-supervised learning framework by combining deep metric learning and density clustering. Experimental results on three typical multi-view learning tasks, i.e., webpage classification, image classification, and document classification demonstrate the effectiveness of the proposed approach.
Xiaodong Jia 0005, Xiaoyuan Jing, Xiaoke Zhu, Songcan Chen, Bo Du 0001, Ziyun Cai, Zhenyu He 0001, Dong Yue 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Scale-fusion framework for improving video-based person re-identification performance
Li Cheng 0006, Xiaoyuan Jing, Xiaoke Zhu, Fei Ma 0004, Changhui Hu 0001, Ziyun Cai, Fumin Qi
Neural Comput. Appl.6
2019 Classification complexity assessment for hyper-parameter optimization
Ziyun Cai, Yang Long 0001, Ling Shao 0001
Pattern Recognit. Lett.1
2018 Adaptive Visual-Depth Fusion Transfer
Ziyun Cai, Yang Long 0001, Xiaoyuan Jing, Ling Shao 0001
ACCV (4)1
2018 Adaptive RGB Image Recognition by Visual-Depth Embedding
abstract
Recognizing RGB images from RGB-D data is a promising application, which significantly reduces the cost while can still retain high recognition rates. However, existing methods still suffer from the domain shifting problem due to conventional surveillance cameras and depth sensors are using different mechanisms. In this paper, we aim to simultaneously solve the above two challenges: 1) how to take advantage of the additional depth information in the source domain? 2) how to reduce the data distribution mismatch between the source and target domains? We propose a novel method called adaptive Visual- Depth Embedding (aVDE) which learns the compact shared latent space between two representations of labeled RGB and depth modalities in the source domain first. Then the shared latent space can help the transfer of the depth information to the unlabeled target dataset. At last, aVDE models two separate learning strategies for domain adaptation (feature matching and instance reweighting) in a unified optimization problem, which matches features and reweights instances jointly across the shared latent space and the projected target domain for an adaptive classifier. We test our method on five pairs of datasets for object recognition and scene classification, the results of which demonstrates the effectiveness of our proposed method.
Ziyun Cai, Yang Long 0001, Ling Shao 0001
IEEE Trans. Image Process.1
2017 RGB-D data fusion in complex space
abstract
Most of the RGB-D fusion methods extract features from RGB data and depth data separately and then simply concatenate them or encode these two kinds of features. Such frameworks cannot explore the correlation between the RGB pixels and their corresponding depth pixels. Motivated by the physical concept that range data correspond to the phase change and color information corresponds to the intensity, we first project raw RGB-D data into a complex space and then jointly extract features from the fused RGB-D images. Consequently, the correlated and individual parts of the RGB-D information in the new feature space are well combined. Experimental results of SIFT and fused images trained CNNs on two RGB-D datasets show that our proposed RGB-D fusion method can achieve competing performance against the classical fusion methods.
Ziyun Cai, Ling Shao 0001
ICIP1
2017 Performance evaluation of deep feature learning for RGB-D image/video classification
Ling Shao 0001, Ziyun Cai, Li Liu 0004, Ke Lu 0002
Inf. Sci.2
2017 RGB-D datasets using microsoft kinect or similar sensors: a survey
abstract
RGB-D data has turned out to be a very useful representation of an indoor scene for solving fundamental computer vision problems. It takes the advantages of the color image that provides appearance information of an object and also the depth image that is immune to the variations in color, illumination, rotation angle and scale. With the invention of the low-cost Microsoft Kinect sensor, which was initially used for gaming and later became a popular device for computer vision, high quality RGB-D data can be acquired easily. In recent years, more and more RGB-D image/video datasets dedicated to various applications have become available, which are of great importance to benchmark the state-of-the-art. In this paper, we systematically survey popular RGB-D datasets for different applications including object recognition, scene classification, hand gesture recognition, 3D-simultaneous localization and mapping, and pose estimation. We provide the insights into the characteristics of each important dataset, and compare the popularity and the difficulty of those datasets. Overall, the main goal of this survey is to give a comprehensive description about the available RGB-D datasets and thus to guide researchers in the selection of suitable datasets for evaluating their algorithms.
Ziyun Cai, Jungong Han, Li Liu 0004, Ling Shao 0001
Multim. Tools Appl.1
2015 Latent Structure Preserving Hashing
Ziyun Cai, Li Liu 0004, Mengyang Yu, Ling Shao 0001
BMVC1