EDBT 2026 Demo / reviewers in the wild / expert
Jinjing Zhu
dblp:251/2954
· DBLP profile ↗
17ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Application of machine learning to mechanical properties of composite materials: A ten-year review (2015-2025)
Jinjing Zhu, Ruixiang Bai, Yaoxing Xu, Jiantong Wang, Longchao He, Zhenkun Lei |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Exploring the Vulnerabilities of Federated Learning: A Deep Dive Into Gradient Inversion AttacksabstractFederated Learning (FL) has emerged as a promising privacy-preserving collaborative model training paradigm without sharing raw data. However, recent studies have revealed that private information can still be leaked through shared gradient information and attacked by Gradient Inversion Attacks (GIA). While many GIA methods have been proposed, a detailed analysis, evaluation, and summary of these methods are still lacking. Although various survey papers summarize existing privacy attacks in FL, few studies have conducted extensive experiments to unveil the effectiveness of GIA and their associated limiting factors in this context. To fill this gap, we first undertake a systematic review of GIA and categorize existing methods into three types, i.e., optimization-based GIA (OP-GIA), generation-based GIA (GEN-GIA), and analytics-based GIA (ANA-GIA). Then, we comprehensively analyze and evaluate the three types of GIA in FL, providing insights into the factors that influence their performance, practicality, and potential threats. Our findings indicate that OP-GIA is the most practical attack setting despite its unsatisfactory performance, while GEN-GIA has many dependencies and ANA-GIA is easily detectable, making them both impractical. Finally, we offer a three-stage defense pipeline to users when designing FL frameworks and protocols for better privacy protection and share some future research directions from the perspectives of attackers and defenders that we believe should be pursued. We hope that our study can help researchers design more robust FL frameworks to defend against these attacks. Pengxin Guo 0001, Runxi Wang, Shuang Zeng, Jinjing Zhu, Haoning Jiang, Yuyin Zhou, Hui Xiong 0001, Liangqiong Qu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | CLIP the divergence: Language-guided unsupervised domain adaptation
Jinjing Zhu |
Pattern Recognit. | 1 |
| 2025 | PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial AugmentationabstractRecently, Depth Anything Models (DAMs) [47], [48] - a type of depth foundation models – have demonstrated impressive zero-shot capabilities across diverse perspective images. Despite its success, it remains an open question regarding DAMs’ performance on panorama images that enjoy a large field-of-view (180° × 360°) but suffer from spherical distortions. To address this gap, we conduct an empirical analysis to evaluate the performance of DAMs on panoramic images and identify their limitations. For this, we undertake comprehensive experiments to assess the performance of DAMs from three key factors: panoramic representations, 360° camera positions for capturing scenarios, and spherical spatial transformations. This way, we reveal some key findings, e.g., DAMs are sensitive to spatial transformations. We then propose a semi-supervised learning (SSL) framework to learn a panoramic DAM, dubbed PanDA. Under the umbrella of SSL, PanDA first learns a teacher model by fine-tuning DAM through joint training on synthetic indoor and outdoor panoramic datasets. Then, a student model is trained using large-scale unlabeled data, leveraging pseudo-labels generated by the teacher model. To enhance PanDA’s generalization capability, Möbius transformation-based spatial augmentation (MTSA) is proposed to impose consistency regularization between the predicted depth maps from the original and spatially transformed ones. This subtly improves the student model’s robustness to various spatial transformations, even under severe distortions. Extensive experiments demonstrate that PanDA exhibits remarkable zero-shot capability across diverse scenes, and outperforms the data-specific panoramic depth estimation methods on two popular real-world benchmarks. Project page: https://caozidong.github.io/PanDA_Depth/. Zidong Cao, Jinjing Zhu, Weiming Zhang 0001, Hao Ai, Haotian Bai, Hengshuang Zhao, Lin Wang 0025 |
CVPR | 2 |
| 2025 | Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationabstractWith the superior sensitivity of event cameras to high-speed motion and extreme lighting conditions, event-based monocular depth estimation has gained popularity to predict structural information about surrounding scenes in challenging environments. However, the scarcity of labeled event data constrains prior supervised learning methods. Unleashing the promising potential of the existing RGB-based depth foundation model, DAM, we propose Depth Any Event stream (EventDAM) to achieve high-performance event based monocular depth estimation in an annotation-free manner. EventDAM effectively combines paired dense RGB images with sparse event data by incorporating three key cross-modality components: Sparsity-aware Feature Mixture (SFM), Sparsity-aware Feature Distillation (SFD), and Sparsity-invariant Consistency Module (SCM). With the proposed sparsity metric, SFM mixes features from RGB images and event data to generate auxiliary depth predictions, while SFD facilitates adaptive feature distillation. Furthermore, SCM ensures output consistency across varying sparsity levels in event data, thereby endowing EventDAM with zero shot capabilities across diverse scenes. Extensive experiments across a variety of benchmark datasets, compared to approaches using diverse input modalities, robustly substantiate the generalization and zero-shot capabilities of EventDAM. Jinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu, James T. Kwok, Hui Xiong 0001 |
ICCV | 1 |
| 2025 | An Event-tailored State-Space Based Model for Pedestrian DetectionabstractEvent cameras, as emerging bio-inspired sensors, endow us with a unique scene perception capability with sub-millisecond latency in challenging environments, such as high-dynamic range and motion blur, as to which a plausible yet efficient exploration on spatiotemporal characteristics of the sparse, asynchronous event data remains an open problem. Event-based pedestrian detection, considered as a promising alternate for road safety in autonomous driving, is chosen as the testbed in this paper for pursuing a specific event-tailored spatiotemporal model. Note that, heterogeneous architectures are generally used in literature, such as building on a CNN/Transformer-style model for capturing the spatial features and a RNN/LSTM model for mining the temporal coherence, respectively. However, existing methods still face significant limitations, particularly as deployed in multi-rate dynamic environments, characterized by pronounced sparsity patterns in slow-motion or other scenarios. As such, a homogeneous neural network for robust pedestrian detection is proposed, with Event-tailored Recurrent Spatiotemporal State-Space Module (ERS3M) as the core innovation, for a joint meticulous modeling of spatiotemporal sparsity and dynamics over event data. On one hand, inspired by Vision Mamba, ERS3M is equipped with adaptive spatiotemporal state propagation as well as multi-directional compensatory scanning, enabling elaborate detection even as to observation intervals with extremely limited events triggered. On the other hand, ERS3 M is augmented with an additional block termed Temporal-Entropy Synergy, offering a collaborative spatiotemporal event purification mechanism, so as to enhance the probability credibility of event streams in visual semantics considering their complicated dynamics. Finally, ERS3M ends with an aliasing-alleviated S5 block to transit information between consecutive time steps, facilitating the temporal consistent pedestrian detection. Evaluations on the PEDRo dataset demonstrate that, the proposed detection method with ERS3 M as backbone has achieved a comparable or even superior performance to state-of-the-art approaches in terms of both accuracy and efficiency. Liuyi Li, Jian Wang 0145, Jinjing Zhu, Wenze Shao |
ACM Multimedia | 4 |
| 2025 | ST$^2$360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth Estimationabstract360-degree monocular depth estimation plays a crucial role in scene understanding owing to its 180-degree by 360-degree field-of-view (FoV). To mitigate the distortions brought by equirectangular projection, existing methods typically divide 360-degree images into distortion-less perspective patches. However, since these patches are processed independently, depth inconsistencies are often introduced due to scale drift among patches. Recently, video depth estimation (VDE) models have leveraged temporal consistency for stable depth predictions across frames. Inspired by this, we propose to represent a 360-degree image as a sequence of perspective frames, mimicking the viewpoint adjustments users make when exploring a 360-degree scenario in virtual reality. Thus, the spatial consistency among perspective depth patches can be enhanced by exploiting the temporal consistency inherent in VDE models. To this end, we introduce a training-free pipeline for 360-degree monocular depth estimation, called ST²360D. Specifically, ST²360D transforms a 360-degree image into perspective video frames, predicts video depth maps using VDE models, and seamlessly merges these predictions into a complete 360-degree depth map. To generate sequenced perspective frames that align with VDE models, we propose two tailored strategies. First, a spherical-uniform sampling (SUS) strategy is proposed to facilitate uniform sampling of perspective views across the sphere, avoiding oversampling in polar regions typically with limited structural details. Second, a latitude-guided scanning (LGS) strategy is introduced to organize the frames into a coherent sequence, starting from the equator, prioritizing low-latitude slices, and progressively moving toward higher latitudes. Extensive experiments demonstrate that ST²360D achieves strong zero-shot capability on several datasets, supporting resolutions up to 4K. Zidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang, Yuanhuiyi Lyu, Hui Xiong 0001 |
NeurIPS | 2 |
| 2025 | Source-Free Cross-Modal Knowledge Transfer by Unleashing the Potential of Task-Irrelevant DataabstractSource-free cross-modal knowledge transfer is a crucial yet challenging task, which aims to transfer knowledge from one source modality (e.g., RGB) to the target modality (e.g., depth or infrared) with no access to the task-relevant (TR) source data due to memory and privacy concerns. A recent attempt leverages the paired task-irrelevant (TI) data and directly matches the features from them to eliminate the modality gap. However, it ignores a pivotal clue that the paired TI data could be utilized to effectively estimate the source data distribution and better facilitate knowledge transfer to the target modality. To this end, we propose a novel yet concise framework to unlock the potential of paired TI data for enhancing source-free cross-modal knowledge transfer. Our work is buttressed by two key technical components. Firstly, to better estimate the source data distribution, we introduce a Task-irrelevant data-Guided Modality Bridging (TGMB) module. It translates the target modality data into the source-like images based on paired TI data and the guidance of the available source model to alleviate two key gaps: 1) inter-modality gap between the paired TI data; 2) intra-modality gap between TI and TR target data. We then propose a Task-irrelevant data-Guided Knowledge Transfer (TGKT) module that transfers knowledge from the source model to the target model by leveraging the paired TI data. Notably, due to the unavailability of labels for the TR target data and its less reliable prediction from the source model, our TGKT model incorporates a self-supervised pseudo-labeling approach to enable the target model to learn from its predictions. Extensive experiments show that our method achieves state-of-the-art performance on three datasets (RGB-to-depth and RGB-to-infrared). Jinjing Zhu, Yucheng Chen 0002, Lin Wang 0025 |
IEEE Trans. Image Process. | 1 |
| 2024 | Open-Set Semi-Supervised Learning by Distribution AlignmentabstractSemi-Supervised Learning (SSL) has been shown to be effective in the closed-set case where the label spaces in labeled and unlabeled data are the same. However, in open-set SSL, its performance is seriously degraded since unlabeled data contains some classes not seen in the labeled data, leading to the distribution mismatch between labeled and unlabeled data. To solve this problem, we propose a Distribution Aligned Openset SSL (DAOSSL) method, which aims to explicitly reduce the empirical distribution mismatch between the labeled and unlabeled data. Specifically, we first introduce a progressive separation mechanism that utilizes a coarse-to-fine pipeline to weigh the unlabeled data. Based on this weighting strategy, we then propose a weighted distribution alignment approach to minimize the distribution discrepancy between the labeled and unlabeled data. These two strategies can be easily integrated into existing deep SSL approaches for open-set SSL tasks. The effectiveness of the proposed DAOSSL method is demonstrated through empirical studies, which show that the method is able to successfully reduce the distribution mismatch between labeled and unlabeled data, resulting in performance improvement in open-set SSL tasks. Qiao Xiao, Jinjing Zhu, Boqian Wu |
IJCNN | 2 |
| 2024 | Towards Dynamic and Small Objects Refinement for Unsupervised Domain Adaptative Nighttime Semantic SegmentationabstractNighttime semantic segmentation plays a crucial role in practical applications, such as autonomous driving, where it frequently encounters difficulties caused by inadequate illumination conditions and the absence of well-annotated datasets. Moreover, semantic segmentation models trained on daytime datasets often face difficulties in generalizing effectively to nighttime conditions. Unsupervised domain adaptation (UDA) has shown the potential to address the challenges and achieved remarkable results for nighttime semantic segmentation. However, existing methods still face limitations in 1) their reliance on style transfer or relighting models, which struggle to generalize to complex nighttime environments, and 2) their ignorance of dynamic and small objects like vehicles and poles, which are difficult to be directly learned from other domains. This paper proposes a novel UDA method that refines both label and feature levels for dynamic and small objects for nighttime semantic segmentation. First, we propose a dynamic and small object refinement module to complement the knowledge of dynamic and small objects from the source domain to target the nighttime domain. These dynamic and small objects are normally context-inconsistent in under-exposed conditions. Then, we design a feature prototype alignment module to reduce the domain gap by deploying contrastive learning between features and prototypes of the same class from different domains, while re-weighting the categories of dynamic and small objects. Extensive experiments on three benchmark datasets demonstrate that our method outperforms prior arts by a large margin for nighttime segmentation. Project page: https://rorisis.github.io/DSRNSS/. Jingyi Pan 0001, Sihang Li 0003, Yucheng Chen 0002, Jinjing Zhu, Lin Wang 0025 |
IROS | 4 |
| 2024 | A Versatile Framework for Unsupervised Domain Adaptation Based on Instance WeightingabstractDespite the progress made in domain adaptation, solving Unsupervised Domain Adaptation (UDA) problems with a general method under complex conditions caused by label shifts between domains remains a challenging task. In this work, we comprehensively investigate four distinct UDA settings including closed set domain adaptation, partial domain adaptation, open set domain adaptation, and universal domain adaptation, where shared common classes between source and target domains coexist alongside domain-specific private classes. The prominent challenges inherent in diverse UDA settings center around the discrimination of common/private classes and the precise measurement of domain discrepancy. To surmount these challenges effectively, we propose a novel yet effective method called Learning Instance Weighting for Unsupervised Domain Adaptation (LIWUDA), which caters to various UDA settings. Specifically, the proposed LIWUDA method constructs a weight network to assign weights to each instance based on its probability of belonging to common classes, and designs Weighted Optimal Transport (WOT) for domain alignment by leveraging instance weights. Additionally, the proposed LIWUDA method devises a Separate and Align (SA) loss to separate instances with low similarities and align instances with high similarities. To guide the learning of the weight network, Intra-domain Optimal Transport (IOT) is proposed to enforce the weights of instances in common classes to follow a uniform distribution. Through the integration of those three components, the proposed LIWUDA method demonstrates its capability to address all four UDA settings in a unified manner. Experimental evaluations conducted on four benchmark datasets substantiate the effectiveness of the proposed LIWUDA method. The code is available at https://github.com/JinjingZhu/LIWUDA. Jinjing Zhu, Feiyang Ye 0001, Qiao Xiao, Pengxin Guo 0001, Yu Zhang 0006, Qiang Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | SEPT: Towards Scalable and Efficient Visual Pre-trainingabstractRecently, the self-supervised pre-training paradigm has shown great potential in leveraging large-scale unlabeled data to improve downstream task performance. However, increasing the scale of unlabeled pre-training data in real-world scenarios requires prohibitive computational costs and faces the challenge of uncurated samples. To address these issues, we build a task-specific self-supervised pre-training framework from a data selection perspective based on a simple hypothesis that pre-training on the unlabeled samples with similar distribution to the target task can bring substantial performance gains. Buttressed by the hypothesis, we propose the first yet novel framework for Scalable and Efficient visual Pre-Training (SEPT) by introducing a retrieval pipeline for data selection. SEPT first leverage a self-supervised pre-trained model to extract the features of the entire unlabeled dataset for retrieval pipeline initialization. Then, for a specific target task, SEPT retrievals the most similar samples from the unlabeled dataset based on feature similarity for each target instance for pre-training. Finally, SEPT pre-trains the target model with the selected unlabeled samples in a self-supervised manner for target data finetuning. By decoupling the scale of pre-training and available upstream data for a target task, SEPT achieves high scalability of the upstream dataset and high efficiency of pre-training, resulting in high model architecture flexibility. Results on various downstream tasks demonstrate that SEPT can achieve competitive or even better performance compared with ImageNet pre-training while reducing the size of training samples by one magnitude without resorting to any extra annotations. Huabin Zheng, Huaping Zhong, Jinjing Zhu, Conghui He, Lin Wang 0025 |
AAAI | 4 |
| 2023 | Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic SegmentationabstractThe ability of scene understanding has sparked active research for panoramic image semantic segmentation. However, the performance is hampered by distortion of the equirectangular projection (ERP) and a lack of pixel-wise annotations. For this reason, some works treat the ERP and pinhole images equally and transfer knowledge from the pinhole to ERP images via unsupervised domain adaptation (UDA). However, they fail to handle the domain gaps caused by: 1) the inherent differences between camera sensors and captured scenes; 2) the distinct image formats (e.g., ERP and pinhole images). In this paper, we propose a novel yet flexible dual-path UDA framework, DPPASS, taking ERP and tangent projection (TP) images as inputs. To reduce the domain gaps, we propose cross-projection and intra-projection training. The cross-projection training includes tangent-wise feature contrastive training and prediction consistency training. That is, the former formulates the features with the same projection locations as positive examples and vice versa, for the models' awareness of distortion, while the latter ensures the consistency of cross-model predictions between the ERP and TP. Moreover, adversarial intra-projection training is proposed to reduce the inherent gap, between the features of the pinhole images and those of the ERP and TP images, respectively. Importantly, the TP path can be freely removed after training, leading to no additional inference cost. Extensive experiments on two benchmarks show that our DPPASS achieves + 1.06% mIoU increment than the state-of-the-art approaches. https://vlis2022.github.io/cvpr23/DPPASS Xu Zheng 0002, Jinjing Zhu, Yexin Liu, Zidong Cao, Chong Fu 0001, Lin Wang 0025 |
CVPR | 2 |
| 2023 | Patch-Mix Transformer for Unsupervised Domain Adaptation: A Game PerspectiveabstractEndeavors have been recently made to leverage the vision transformer (ViT) for the challenging unsupervised domain adaptation (UDA) task. They typically adopt the cross-attention in ViT for direct domain alignment. However, as the performance of cross-attention highly relies on the quality of pseudo labels for targeted samples, it becomes less effective when the domain gap becomes large. We solve this problem from a game theory's perspective with the proposed model dubbed as PMTrans, which bridges source and target domains with an intermediate domain. Specifically, we propose a novel ViT-based module called PatchMix that effectively builds up the intermediate domain, i.e., probability distribution, by learning to sample patches from both domains based on the game-theoretical models. This way, it learns to mix the patches from the source and target domains to maximize the cross entropy (CE), while exploiting two semi-supervised mixup losses in the feature and label spaces to minimize it. As such, we interpret the process of UDA as a min-max CE game with three players, including the feature extractor, classifier, and PatchMix, to find the Nash Equilibria. Moreover, we leverage attention maps from ViT to re-weight the label of each patch by its importance, making it possible to obtain more domain-discriminative feature representations. We conduct extensive experiments on four benchmark datasets, and the results show that PMTrans significantly surpasses the ViT-based and CNN-based SoTA methods by +3.6% on Office-Home, +1.4% on Office-31, and +17.7% on DomainNet, respectively. https://vlis2022.github.io/cvpr23/PMTrans Jinjing Zhu, Haotian Bai, Lin Wang 0025 |
CVPR | 1 |
| 2023 | A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic SegmentationabstractIn this paper, we strive to answer the question ‘how to collaboratively learn convolutional neural network (CNN)-based and vision transformer (ViT)-based models by selecting and exchanging the reliable knowledge between them for semantic segmentation?’ Accordingly, we propose an online knowledge distillation (KD) framework that can simultaneously learn compact yet effective CNN-based and ViT-based models with two key technical breakthroughs to take full advantage of CNNs and ViT while compensating their limitations. Firstly, we propose heterogeneous feature distillation (HFD) to improve students’ consistency in low-layer feature space by mimicking heterogeneous features between CNNs and ViT. Secondly, to facilitate the two students to learn reliable knowledge from each other, we propose bidirectional selective distillation (BSD) that can dynamically transfer selective knowledge. This is achieved by 1) region-wise BSD determining the directions of knowledge transferred between the corresponding regions in the feature space and 2) pixel-wise BSD discerning which of the prediction knowledge to be transferred in the logit space. Extensive experiments on three benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art online distillation methods by a large margin, and shows its efficacy in learning collaboratively between ViT-based and CNN-based models. Jinjing Zhu, Yunhao Luo 0001, Xu Zheng 0002, Hao Wang 0005, Lin Wang 0025 |
ICCV | 1 |
| 2022 | Selective Partial Domain Adaptation
Pengxin Guo 0001, Jinjing Zhu, Yu Zhang 0006 |
BMVC | 2 |
| 2019 | Atrial Fibrillation Detection using Different Duration ECG Signals with SE-ResNetabstractAtrial fibrillation (AF) is the most common cardiac arrhythmia that has significant effects on associated morbidity and mortality. Therefore, early detection of AF by electrocardiogram (ECG) is substantial and necessary for effective treatments of AF. Here, we creatively propose a signal processing method for ECG signals and develop a 94-layer deep neural network to classify 4 rhythm signals using 8528 different duration single-lead ECGs from The PhysioNet Challenge 2017. And the model is trained and validated on different databases, respectively. We utilize F1 score to evaluate the performance of our method and obtain the mean F1 score of 85%. Finally, the experiment results demonstrate that our algorithm achieves excellent performance for AF and other arrhythmia detection and is also superior to existing algorithms. Jinjing Zhu |
MMSP | 1 |