EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Cao
dblp:253/2218
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Notes about the linear complexity of quaternary cyclotomic sequences of order four
Vladimir Edemskiy, Zeyu Cao |
Inf. Process. Lett. | 2 |
| 2025 | Cross-View-Prediction: Exploring Contrastive Feature for Hyperspectral Image ClassificationabstractThis paper presents a self-supervised feature learning framework for hyperspectral image (HSI) classification. Our approach constructs two distinct views of the raw HSI data through a cross-representation learning strategy and learns semantically consistent representations via a contrastive learning objective. Specifically, we introduce four cross-channel prediction-based augmentation techniques that exploit the highdimensional characteristics of hyperspectral data to generate diverse views. The proposed contrastive network learns discriminative and robust features by maximizing mutual information and minimizing conditional entropy across different views. Despite its simplicity, our Cross-View Prediction framework achieves state-of-the-art performance in unsupervised HSI classification when combined with a linear SVM classifier. The code is available at https://github.com/new-captain/CrossView-Prediction. Anyu Zhang, Zeyu Cao |
EUC | 2 |
| 2025 | DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single ImageabstractReconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, complex deformations, and the ambiguity of the single-view setting. The previous state-of-the-art, Decaf, employs a global fitting optimization guided by contact and deformation estimation networks trained on studio-collected data with 3D annotations. However, Decaf suffers from a time-consuming optimization process and limited generalization capability due to its reliance on 3D annotations of hand-face interaction data. To address these issues, we present DICE, the first end-to-end method for Deformation-aware hand-face Interaction reCovEry from a single image. DICE estimates the poses of hands and faces, contacts, and deformations simultaneously using a Transformer-based architecture. It features disentangling the regression of local deformation fields and global mesh vertex locations into two network branches, enhancing deformation and contact estimation for precise and robust hand-face mesh recovery. To improve generalizability, we propose a weakly-supervised training approach that augments the training set using in-the-wild images without 3D ground-truth annotations, employing the depths of 2D keypoints estimated by off-the-shelf models and adversarial priors of poses for supervision. Our experiments demonstrate that DICE achieves state-of-the-art performance on a standard benchmark and in-the- wild data in terms of accuracy and physical plausibility. Additionally, our method operates at an interactive rate (20 fps) on an Nvidia 4090 GPU, whereas Decaf requires more than 15 seconds for a single image. The code will be available at: https://github.com/Qingxuan-Wu/DICE. Qingxuan Wu, Zhiyang Dou, Sirui Xu 0002, Soshi Shimada, Chen Wang 0049, Zhengming Yu, Yuan Liu 0025, Cheng Lin 0001, Zeyu Cao, Taku Komura, Vladislav Golyanik, Christian Theobalt, Wenping Wang 0001, Lingjie Liu |
ICLR | 9 |
| 2025 | Diffusion-Decided Views in Contrastive Learning for Hyperspectral Image ClassificationabstractSelf-supervised learning has made significant strides in hyperspectral image classification. For contrastive learning, traditional data augmentation methods will lead to biases with real-world spectra, limited variability, loss of information. To address those limitations, this study introduces diffusion models to generate realistic synthetic samples as views input to contrastive learning. This diffusion-based augmentation enables our contrastive learning network to learn more robust spectral-spatial representations. Experimental results on two publicly available datasets indicate that this approach outperforms numerous existing methods. And the ablation study demonstrates the effectiveness of employing diffusion models. Xiaorun Li, Shuhan Chen, Zeyu Cao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Combining Contrastive Learning and Diffusion Model for Hyperspectral Image ClassificationabstractIn recent years, self-supervised learning has made significant strides in hyperspectral image classification [1]. However, different approaches come with distinct strengths and limitations. Contrastive learning excels at extracting key information from large volumes of redundant data, but its training objective can inadvertently increase intra-class feature distance. To address this limitation, we leverage diffusion models for their proven ability to refine and aggregate features by modeling complex data distributions. Specifically, diffusion models’ inherent denoising and generative process are theoretically well-suited to enhance intra-class compactness by learning to reconstruct clean, representative features from perturbed inputs. We propose the new method - ContrastDM. This approach generates synthetic features, improving and enriching feature representation, and partially addressing the issue of sample sparsity. Classification experiments on three publicly available datasets demonstrate that ContrastDM significantly outperforms state-of-the-art methods. Xiaorun Li, Shuhan Chen, Zeyu Cao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Enhancing Hyperspectral Image Classification in Small Sample Settings Using Conditional Denoising Diffusion Implicit ModelabstractEnhancing classification accuracy in hyperspectral image classification under small sample scenarios remains a critical challenge. This study introduces a novel framework based on Conditional Denoising Diffusion Implicit Models (DDIM), featuring an adaptive synthetic sample generation strategy and label noise training. Our approach dynamically allocates synthetic samples for each class based on classification difficulty and generation reliability, while label noise training enhances the model's generalization capabilities. As a flexible data augmentation method, our framework can be seamlessly integrated with various feature extraction techniques, demonstrating significant potential for further performance improvements. Validated on the Salinas and Pavia University datasets, our method highlights the effectiveness of diffusion models as a robust tool for addressing data scarcity, paving the way for broader applications in hyperspectral image classification and remote sensing. Zeyu Cao, Shuhan Chen |
IEEE Signal Process. Lett. | 1 |
| 2024 | EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang 0003, Wenjia Wang 0009, Yuan Liu 0025, Taku Komura, Wenping Wang 0001, Lingjie Liu |
ECCV (2) | 3 |
| 2024 | Contrastdm: Combining Contrastive Learning and Diffusion Model for Hyperspectral Image ClassificationabstractIn recent years, self-supervised learning has made rapid progress in the field of hyperspectral image classification(HSIC). However, different methods have their own advantages and disadvantages. Contrastive learning can extract key information from a large amount of redundant data, but its training method can lead to significant intra-class variance. Diffusion models can enhance features and aggregate similar ones, but they have a high computational cost due to the sampling process. We combine contrastive learning and diffusion models by training the diffusion model with low-dimensional features extracted from contrastive learning. This generates new artificial features to improve feature representation. We name this new method ContrastDM. The results of classification experiments on two publicly available datasets demonstrate that ContrastDM significantly outperforms the compared state-of-the-art (SOTA) methods. Zeyu Cao |
IGARSS | 3 |
| 2024 | Residual Spatio-Temporal Attention Based Prototypical Network for Rare Arrhythmia Classification
Zeyu Cao, Fengyi Guo, Ying An |
ISBRA (3) | 1 |
| 2024 | AutoPSV: Automated Process-Supervised VerifierabstractIn this work, we propose a novel method named \textbf{Auto}mated \textbf{P}rocess-\textbf{S}upervised \textbf{V}erifier (\textbf{\textsc{AutoPSV}}) to enhance the reasoning capabilities of large language models (LLMs) by automatically annotating the reasoning steps.
\textsc{AutoPSV} begins by training a verification model on the correctness of final answers, enabling it to generate automatic process annotations.
This verification model assigns a confidence score to each reasoning step, indicating the probability of arriving at the correct final answer from that point onward.
We detect relative changes in the verification's confidence scores across reasoning steps to automatically annotate the reasoning process, enabling error detection even in scenarios where ground truth answers are unavailable.
This alleviates the need for numerous manual annotations or the high computational costs associated with model-induced annotation approaches.
We experimentally validate that the step-level confidence changes learned by the verification model trained on the final answer correctness can effectively identify errors in the reasoning steps.
We demonstrate that the verification model, when trained on process annotations generated by \textsc{AutoPSV}, exhibits improved performance in selecting correct answers from multiple LLM-generated outputs.
Notably, we achieve substantial improvements across five datasets in mathematics and commonsense reasoning. The source code of \textsc{AutoPSV} is available at \url{https://github.com/rookie-joe/AutoPSV}. Jianqiao Lu, Zhiyang Dou, Hongru Wang 0003, Zeyu Cao, Jianbo Dai, Yunlong Feng, Zhijiang Guo |
NeurIPS | 4 |
| 2023 | TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerabstractIn this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caused by redundant tokens. We propose token reduction strategies based on two important aspects, i.e., the 3D geometry structure and 2D image feature, where we hierarchically recover the mesh geometry with priors from body structure and conduct token clustering to pass fewer but more discriminative image feature tokens to the Transformer. Our method massively reduces the number of tokens involved in high-complexity interactions in the Transformer. This leads to a significantly reduced computational cost while still achieving competitive or even higher accuracy in shape recovery. Extensive experiments across a wide range of benchmarks validate the superior effectiveness of the proposed method. We further demonstrate the generalizability of our method on hand mesh recovery. Visit our project page at https://frank-zy-dou.github.io/projects/Tore/index.html. Zhiyang Dou, Qingxuan Wu, Cheng Lin 0001, Zeyu Cao, Qiangqiang Wu, Weilin Wan 0001, Taku Komura, Wenping Wang 0001 |
ICCV | 4 |
| 2023 | Privacy Matters: Vertical Federated Linear Contextual Bandits for Privacy Protected RecommendationabstractRecent awareness of privacy protection and compliance requirement resulted in a controversial view of recommendation system due to personal data usage. Therefore, privacy-protected recommendation emerges as a novel research direction. In this paper, we first formulate this problem as a vertical federated learning problem, i.e., features are vertically distributed over different departments. We study a contextual bandit learning problem for recommendation in the vertical federated setting. To this end, we carefully design a customized encryption scheme named orthogonal matrix-based mask mechanism (O3M). O3M mechanism, a tailored component for contextual bandits by carefully exploiting their shared structure, can ensure privacy protection while avoiding expensive conventional cryptographic techniques. We further apply the mechanism to two commonly-used bandit algorithms, LinUCB and LinTS, and instantiate two practical protocols for online recommendation. The proposed protocols can perfectly recover the service quality of centralized bandit algorithms while achieving a satisfactory runtime efficiency, which is theoretically proved and analysed in this paper. By conducting extensive experiments on both synthetic and real-world datasets, we show the superiority of the proposed method in terms of privacy protection and recommendation performance. Zeyu Cao, Bingzhe Wu, Shu Zhang 0013, Hangyu Li 0002, Ouyang Wen, Yu Rong 0001, Peilin Zhao |
KDD | 1 |
| 2022 | ALPN: Active-Learning-Based Prototypical Network for Few-Shot Hyperspectral Imagery ClassificationabstractWith the development of deep learning, the benchmark of hyperspectral imagery classification is constantly improving, but there are still significant challenges for hyperspectral imagery classification of few-shot scenes. This letter proposes an active-learning-based prototypical network (ALPN), which uses the prototypical network to extract representative features from a few samples. Moreover, it combines semisupervised clustering and active learning methods to select and request labels from valuable examples actively. In this way, the feature extraction ability of the network is gradually optimized. The experimental results validated that the classification accuracy and robustness of ALPN significant exceeded the comparison baselines. Furthermore, because it can be regarded as a sample selection method, ALPN can be easily combined with other models to obtain better classification results. Xiaorun Li, Zeyu Cao, Liaoying Zhao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | ROBYOL: Random-Occlusion-Based BYOL for Hyperspectral Image ClassificationabstractWith the development of deep learning, hyperspectral image classification (HSIC) has improved rapidly in recent years. Unsupervised feature learning algorithms play an important role in extracting features from hyperspectral images (HSIs). This letter proposed a random-occlusion-based Bootstrap-Your-Own-Latent network (ROBYOL), combining a new augmentation method and a superior contrastive learning algorithm for feature extraction. The proposed method consists of a self-supervised learning part for feature extracting and a classifier part as the downstream task. It can be proved by the experimental results that the feature extraction ability of the network is effective in this way. Furthermore, the influence of different occlusion strategies is also studied, including changing occlusion area and occlusion value, and we proposed translucent occlusion. Our results with two well-known HSIs reveal that proper occlusion strategies can improve hyperspectral classification results effectively. Xiaorun Li, Zeyu Cao, Liaoying Zhao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | ContrastNet: Unsupervised feature learning by autoencoder and prototypical contrastive learning for hyperspectral imagery classification
Zeyu Cao, Xiaorun Li, Yueming Feng, Shuhan Chen, Chaoqun Xia, Liaoying Zhao |
Neurocomputing | 1 |
| 2019 | Object Detection in VHR Image Using Transfer Learning with Deformable ConvolutionabstractIn the field of deep learning, finetuning the pretrained networks to get a good classifier is a common way of transfer learning. Unlike the traditional way, we insert deformable convolutional layers into the pretrained networks, and finetune the new networks. As a result, we find it performs as well as the normal one in classification, and when we construct a plane detection pipeline based on the two classifiers respectively, the one with deformable convolution shows a better result than the other. Zeyu Cao, Xiaorun Li, Liaoying Zhao |
IGARSS | 1 |