EDBT 2026 Demo / reviewers in the wild / expert
Jingyi Cao
dblp:249/3431
· DBLP profile ↗
19ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extended r-adaptive isogeometric analysis for weak-discontinuous problemsabstract• An extended r -adaptive IGA framework is proposed for weak-discontinuous problems. • A Gaussian monitor redistributes geometric resolution using level-set interfaces. • Control-point relocation enhances interface resolution without mesh refinement. • Up to 65.7% error reduction is achieved with negligible computational overhead. Jingyi Cao, Ye Ji 0001, Matthias Möller, Chungang Zhu |
Comput. Aided Des. | 1 |
| 2026 | Joint regenerator placement, routing, and wavelength assignment in survivable translucent WDM networks
Jingyi Cao, Weichang Zheng, Mingcong Yang, Ruiling Wu, Yongbing Zhang 0001 |
Comput. Networks | 2 |
| 2025 | Rethinking Contrastive Learning in Graph Anomaly Detection: A Clean-View PerspectiveabstractGraph anomaly detection aims to identify unusual patterns in graph-based data, with wide applications in fields such as web security and financial fraud detection. Existing methods typically rely on contrastive learning, assuming that a lower similarity between a node and its local subgraph indicates abnormality. However, these approaches overlook a crucial limitation: the presence of interfering edges invalidates this assumption, since it introduces disruptive noise that compromises the contrastive learning process. Consequently, this limitation impairs the ability to effectively learn meaningful representations of normal patterns, leading to suboptimal detection performance. To address this issue, we propose a Clean-View Enhanced Graph Anomaly Detection framework (CVGAD), which includes a multi-scale anomaly awareness module to identify key sources of interference in the contrastive learning process. Moreover, to mitigate bias from the one-step edge removal process, we introduce a novel progressive purification module. This module incrementally refines the graph by iteratively identifying and removing interfering edges, thereby enhancing model performance. Extensive experiments on five benchmark datasets validate the effectiveness of our approach. Di Jin 0001, Jingyi Cao, Xiaobao Wang, Bingdao Feng, Dongxiao He, Longbiao Wang, Jianwu Dang 0001 |
IJCAI | 2 |
| 2025 | Exploiting Self-Refining Normal Graph Structures for Robust Defense against Unsupervised Adversarial AttacksabstractDefending against adversarial attacks on graphs has become increasingly important. Graph refinement to enhance the quality and robustness of representation learning is a critical area that requires thorough investigation. We observe that representations learned from attacked graphs are often ineffective for refinement due to perturbations that cause the endpoints of perturbed edges to become more similar, complicating the defender's ability to distinguish them. To address this challenge, we propose a robust unsupervised graph learning framework that utilizes cleaner graphs to learn effective representations. Specifically, we introduce an anomaly detection model based on contrastive learning to obtain a rough graph excluding a large number of perturbed structures. Subsequently, we then propose the Graph Pollution Degree (GPD), a mutual information-based measure that leverages the encoder's representation capability on the rough graph to assess the trustworthiness of the predicted graph and refine the learned representations. Extensive experiments on four benchmark datasets demonstrate that our method outperforms nine state-of-the-art defense models, effectively defending against adversarial attacks and enhancing node classification performance. Bingdao Feng, Di Jin 0001, Xiaobao Wang, Dongxiao He, Jingyi Cao, Zhen Wang 0004 |
IJCAI | 5 |
| 2024 | MV-MOE: A Visual Mixture-of-Experts Model for Optical-SAR Image MatchingabstractOptical and Synthetic Aperture Radar (SAR) matching produces spatial and semantic correspondences of the input images, playing a pivotal role in the registration process. However, due to the difference in radiation characteristics and geometric properties, even the same target may manifest distinctive morphological and feature expressions in the cross-modal images. Consistent feature extraction remains a challenge for optical-SAR image matching. Therefore, based on the salient image patterns (keypoint, line, and block), a Visual Mixture-of-Experts method for optical-SAR image Matching (MV-MOE) is proposed. It facilitates adaptive graphical representation of multi-modal images across various scenes through the multi-task learning framework. With the aid of the attention mechanism, the task-related basic features are reconstructed into matching-related features, yielding similarity along with spatial offset vectors. Additionally, we employ a multi-level feature extraction backbone based on the visual retentive block, enhancing local feature perception with the preservation capability of the recurrent network structure. Experiments demonstrate the advantages of the proposed method on multi-modal image matching and its contribution to the subsequent registration. Jingyi Cao, Yanan You, Jun Liu 0014 |
IGARSS | 1 |
| 2024 | Cross-Domain Activity Recognition Based on Stacked Transfer NetworkabstractHuman activity recognition (HAR) based on wearable sensors is a hot topic in health detection and motion management. Nonetheless, conventional identification methods necessitate substantial labeled datasets, and the acquisition of high-quality labeled data crucial for human activity recognition is both time-consuming and costly. To tackle this problem, transfer learning is used to annotating unlabeled or a few labeled target domains using labeled source domains. Meanwhile, individuals exhibit significant differences in amplitude, angles, and other aspects when performing the same action. Due to the limited expressive capacity of time series, effectively capturing various types of differences simultaneously poses a challenge, impacting the effectiveness of transfer. Therefore, in this paper, we propose a stacked transfer network (STN) for 2D modeling and feature extraction of sensor data in steps. It adaptively decomposes complex temporal variations into multiple intra- and inter-periodic variations using the Fast Fourier Transform (FFT), allowing the model to focus more on the trends of the activty and less on the effects of individual discrepancy. Our comprehensive cross-domain activity recognition (CDAR) experiments on three large public activity recognition datasets (i.e., OPPORTUNITY, PAMAP2, and UCI-DSADS) show that the STN achieves high accuracy in activity recognition ttransfer. Junhuai Li, Jingyi Cao, Yuxing Zhi, Huaijun Wang, Ting Cao 0002, Rong Fei |
IJCNN | 2 |
| 2024 | Adaptive filtering under multi-peak noise
Qizhen Wang, Bangyuan Li, Jingyi Cao |
Signal Process. | 4 |
| 2024 | TSK: A Trustworthy Semantic Keypoint Detector for Remote Sensing ImagesabstractKeypoint detection aims to automatically locate the most significant and informative points in remote sensing images (RSIs), which directly affects the accuracy of matching and registration. In contrast to the handcrafted keypoint detectors that heavily rely on the morphological gradient of corner, line, and ridge, the learning-based detectors emphasize obtaining reliable keypoints from deep features. However, the limited accuracy of semantics undermines the reliability of keypoints, especially in challenging scenarios characterized by repeated textures and boundaries. Therefore, a novel trustworthy semantic keypoint (TSK) detector is proposed for RSIs. It utilizes a lightweight multiscale feature extraction and fusion network, along with a saliency keypoint localization mechanism, to facilitate keypoint detection. Notably, the TSK detector employed explicit semantics, which is refined with multiple learning strategies about repeatability and representability across the multigranularity reasoning spaces, namely, pixel window, neighbor window, and existence entity. Finally, several metrics about repeatability, matching, and registration are used to evaluate the performance of the TSK detector and other competitive methods. Four RSI datasets, including MICGE, HRSCD, OSCD, and SZTAKI, are used to verify performances. TSK detector achieves competitive performance against existing methods. Jingyi Cao, Yanan You, Jun Liu 0014 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Divide and Conquer: a Two-Step Method for High Quality Face De-identification with Model ExplainabilityabstractFace de-identification involves concealing the true identity of a face while retaining other facial characteristics. Current target-generic methods typically disentangle identity features in the latent space, using adversarial training to balance privacy and utility. However, this pattern often leads to a trade-off between privacy and utility, and the latent space remains difficult to explain. To address these issues, we propose IDeudemon, which employs a "divide and conquer" strategy to protect identity and preserve utility step by step while maintaining good explainability. In Step I, we obfuscate the 3D disentangled ID code calculated by a parametric NeRF model to protect identity. In Step II, we incorporate visual similarity assistance and train a GAN with adjusted losses to preserve image utility. Thanks to the powerful 3D prior and delicate generative designs, our approach could protect the identity naturally, produce high quality details and is robust to different poses and expressions. Extensive experiments demonstrate that the proposed IDeudemon outperforms previous state-of-the-art methods. Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001 |
ICCV | 3 |
| 2023 | CloudLoc-NeRF: Point-cloud Assisted Volume Location for Neural Radiance FieldsabstractRealistic rendering results are available to be generated by volume-based neural rendering methods like NeRF. However, the existing schemes take the color vector as the unique supervision information, which leads to ambiguous prediction results of the volume density in the same spatial position through different rendering rays. This problem is common in large-scale scene reconstruction based on remote sensing images taken by UAVs and other equipment. To this end, we contrive to integrate spatial point cloud and multi-view image information, making the sparse point cloud the calibration for key features and regions. Therefore, the CloudLoc-NeRF is proposed. For volume information extraction, the multi-resolution hash coding and voxel are adopted to estimate the district for ray marching and extract volume features efficiently. For the point cloud, annular sampling and plane coding are used to combine image features of the training views and the point cloud. The regions with high feature response in multiple modal data should correspond to the regions with high volume density. In addition, an optimization method based on point cloud density is proposed. The weight parameter of volume density confidence is constructed to symbolize the correlation between density distribution and point cloud density. We verified the performance of our method on NVSF and the wide-area scene reconstruction dataset. Experiments showed that CloudLoc-NeRF accurately expresses the details of the rendered scene and produces better view synthesis results. Jingyi Cao, Yanan You, Songzhi Gao, Jun Liu 0014 |
IGARSS | 1 |
| 2023 | Achieving Privacy-Preserving Multi-View Consistency with Advanced 3D-Aware Face De-identificationabstractThe widespread application of face recognition technology has exacerbated privacy threats. Face de-identification is an effective means of protecting visual privacy by concealing identity information. While deep learning-based methods have greatly improved de-identification results, most existing algorithms rely on 2D generative models that struggle to produce identity-consistent results for multiple views. In this paper, we focus on identity disentanglement within the latest 3D-aware face generation model, and propose an advanced face de-identification framework that can be applied to various scenarios. Our proposed framework disentangles identity from other facial features, modifies only the former and generates the de-identified face using a 3D generator. This approach results in high-quality, identity-consistent de-identification that preserves other facial features. We demonstrate our approach on StyleNeRF, one of the most widely-used style-based neural radiation field models. Through extensive experiments, we demonstrate the effectiveness of our approach in achieving face de-identification both for a single image and group images with the same identity. Our work is a significant step forward in the field of face de-identification, opening up new possibilities for practical applications. Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001 |
MMAsia | 1 |
| 2022 | Make Object Connect: A Pose Estimation Network for UAV Images of the Outdoor SceneabstractAs the basics of 3D vision, pose estimation with 2D images is of significance in 3D reconstruction, UAV positioning, and other fields. However, the related works focus on the natural images and pay less attention to the wide-coverage UAV remote sensing (RS) images. In fact, the relationship between objects in UAV images can benefit pose estimation. Therefore, aiming at the outdoor scene captured by the UAV monocular camera, a novel pose estimation network that emphasizes the association between objects is proposed. The multi-scale visual features extracted by the convolutional neural network (CNN) are manipulated by the object-agnostic segmentation model to indicate the existing space of all possible objects in the whole scene. The features of all possible objects are embedded into vectors, and then processed with a graph convolution network (GCN) for relationship analysis. Based on the known sparse point cloud and the optimized features of 2D images, the camera pose is regressed iteratively by 3D visual geometry. To verify the feasibility of the network, experiments are conducted on the Extended CMU Seasons and the simulation UAV dataset. Results prove that our network emphasizes more features on the small objects and obtains superior pose estimation results. Jingyi Cao, Yanan You, Le Xia, Jun Liu 0014 |
IGARSS | 1 |
| 2022 | IdentityMask: Deep Motion Flow Guided Reversible Face Video De-IdentificationabstractUnprecedented video collection and sharing have exacerbated privacy concerns and led to increasing interest in privacy-preserving tools. A satisfactory video de-identification tool should be able to remove sensitive identity information from face videos while maintaining useful information for other identity-agnostic tasks. Meanwhile, it is necessary to allow the authority to inspect real identity when abnormal events are detected. Existing methods only focus on the study of de-identification, and lack the desired recovery ability when granting permissions. Furthermore, they all process the videos frame by frame, which hardly benefit from motion and inter-frame information. In this paper, we propose a modular architecture for reversible face video de-identification, called IdentityMask, which leverages deep motion flow to avoid per-frame evaluation. Our framework consists of two processes: the de-identification process provides a protective mask for identity information, while the recovery process can remove the protective mask if and only if the right key is provided. To this end, a Protection Module and a Recovery Module are built as two major functional modules, both based on an identity disentanglement network and guided by a crucial Motion Flow Module. An Affine Transformation Module provides simple but reliable assistance. Extensive experiments on a diverse natural video dataset (gender, ethnicity, age, etc.) demonstrate the effectiveness of the proposed framework for reversible face video de-identification. Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Personalized and Invertible Face De-identification by Disentangled Identity Information ManipulationabstractThe popularization of intelligent devices including smartphones and surveillance cameras results in more serious privacy issues. De-identification is regarded as an effective tool for visual privacy protection with the process of concealing or replacing identity information. Most of the existing de-identification methods suffer from some limitations since they mainly focus on the protection process and are usually non-reversible. In this paper, we propose a personalized and invertible de-identification method based on the deep generative model, where the main idea is introducing a user-specific password and an adjustable parameter to control the direction and degree of identity variation. Extensive experiments demonstrate the effectiveness and generalization of our proposed framework for both face de-identification and recovery. Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001 |
ICCV | 1 |
| 2021 | MoatNet: Registration for Multi-Temporal Optical Remote Sensing Images Using Deep Convolutional FeaturesabstractImage registration is an important technique that has been widely used in many areas. It is an indispensable premise for remote sensing image tasks like change detection and image fusion. In this paper, we propose a deep learning framework to generate descriptors for key points and then combine the descriptors constructed with FAST key points for accurate image registration. During the process of training, we adopt a novel loss function named Moat Loss (ML) to train our model, which is accordingly called MoatNet. Experiments show that our method is more robust than traditional algorithms like SIFT and is more accurate than the end-to-end deep learning methods in much more complex cases. Yanan You, Jingyi Cao |
IGARSS | 3 |
| 2021 | Fusion Detection of Closed Water in Medium-Low Resolution Remote Sensing ImageryabstractAiming at the closed water detection in remote sensing imagery at medium-low resolution, this paper proposes a novel method for closed water detection based on fusion detection which conducts detection via informative fused images blended by Synthetic Aperture Radar (SAR) and optical images. Firstly, it utilizes SAR and optical image pairs containing the same closed water object to generate aligned image pairs according to latitude and longitude information. Next, generative adversarial network (GAN) is adopted to fuse two categories of images. At last, a target detection network driven by optical image samples is used to detect the closed water on the fused image. The experiment result on Sentinel-1&2 shows that the proposed method can effectively make up for the shortage of SAR image in closed water detection and improve the detection performance. Yuanyong Ning, Yanan You, Jingyi Cao, Fang Liu 0026 |
IGARSS | 3 |
| 2021 | Deep Motion Flow Aided Face Video De-identificationabstractAdvances in cameras and web technology have made it easy to capture and share large amounts of face videos over to an unknown audience with uncontrollable purposes. These raise increasing concerns about unwanted identity-relevant computer vision devices invading the characters's privacy. Previous de-identification methods rely on designing novel neural networks and processing face videos frame by frame, which ignore the data feature in redundancy and continuity. Besides, these techniques are incapable of well-balancing privacy and utility, and per-frame evaluation is easy to cause flicker. In this paper, we present deep motion flow, which can create remarkable de-identified face videos with a good privacy-utility tradeoff. It calculates the relative dense motion flow between every two adjacent original frames and runs the high quality image anonymization only on the first frame. The de-identified video will be obtained based on the anonymous first frame via the relative dense motion flow. Extensive experiments demonstrate the effectiveness of our proposed de-identification method. Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Jingyi Cao, Li Song 0001 |
VCIP | 4 |
| 2020 | Change Detection Network of Nearshore Ships for Multi-Temporal Optical Remote Sensing ImagesabstractShip change detection is of great significance for maritime safety supervision, wharf vessel management, and vessel life cycle analysis. Nowadays, the ship change information is obtained through the difference between the multiple independent object detection results of multi-temporal images. However, it neglects the temporal correlation of the concerned features, impeding further improvement of the detection accuracy of change detection. Therefore, based on the sequential network, namely convolution LSTM, we built an end-to-end ship change detection (SCD) R-CNN network. The network extracts abstract semantic information reflecting the changed features of ships, and the time correlation between features of the different phases is established. Specifically, the changed features are utilized to guide the judgement of ship change. It is verified from the RS images that the proposed network avoids the misjudgment caused by the errors of object detection in the conventional method. In addition, a higher efficiency is revealed, maintaining the accuracy of the change detection of ship targets. Jingyi Cao, Yanan You, Yuanyong Ning |
IGARSS | 1 |
| 2020 | A Hybrid Model for Natural Face De-Identiation with Adjustable PrivacyabstractAs more and more personal photos are shared and tagged in social media, security and privacy protection are becoming an unprecedentedly focus of attention. Avoiding privacy risks such as unintended verification, becomes increasingly challenging. To enable people to enjoy uploading photos without having to consider these privacy concerns, it is crucial to study techniques that allow individuals to limit the identity information leaked in visual data. In this paper, we propose a novel hybrid model consists of two stages to generate visually pleasing de-identified face images according to a single input. Meanwhile, we successfully preserve visual similarity with the original face to retain data usability. Our approach combines latest advances in GAN-based face generation with well-designed adjustable randomness. In our experiments we show visually pleasing de-identified output of our method while preserving a high similarity to the original image content. Moreover, our method adapts well to the verificator of unknown structure, which further improves the practical value in our real life. Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Yunhui Zhu, Jingyi Cao, Li Song 0001 |
VCIP | 5 |