EDBT 2026 Demo / reviewers in the wild / expert
Yingying Zhu 0001
dblp:40/5552-1
· DBLP profile ↗
50ranked-venue papers
17as first author
29since 2021 · last 2026
0000-0002-3475-6186ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 23 · 11 first-author · 12 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MRGeo: Robust Cross-View Geo-Localization of Corrupted Images via Spatial and Channel Feature EnhancementabstractCross-view geo-localization (CVGL) aims to accurately localize street-view images through retrieval of corresponding geo-tagged satellite images. While prior works have achieved nearly perfect performance on certain standard datasets, their robustness in real-world corrupted environments remains under-explored. This oversight causes severe performance degradation or failure when images are affected by corruption such as blur or weather, significantly limiting practical deployment. To address this critical gap, we introduce MRGeo, the first systematic method designed for robust CVGL under corruption. MRGeo employs a hierarchical defense strategy that enhances the intrinsic quality of features and then enforces a robust geometric prior. Its core is the Spatial-Channel Enhancement Block, which contains: (1) a Spatial Adaptive Representation Module that models global and local features in parallel and uses a dynamic gating mechanism to arbitrate their fusion based on feature reliability; and (2) a Channel Calibration Module that performs compensatory adjustments by modeling multi-granularity channel dependencies to counteract information loss. To prevent spatial misalignment under severe corruption, a Region-level Geometric Alignment Module imposes a geometric structure on the final descriptors, ensuring coarse-grained consistency. Comprehensive experiments on both robustness benchmark and standard datasets demonstrate that MRGeo not only achieves an average R@1 improvement of 2.92% across three comprehensive robustness benchmarks (CVUSA-C-ALL, CVACT_val-C-ALL, and CVACT_test-C-ALL) but also establishes superior performance in cross-area evaluation, thereby demonstrating its robustness and generalization capability. Songsong Ouyang, Yingying Zhu 0001 |
AAAI | 4 |
| 2026 | Learning A Bank of Transferable Prompts for Vision-Language Models
Zhongwei Huang, Chong Wang 0001, Endai Huang, Ran Zhou 0002, Haitao Gan, Yingying Zhu 0001, Xiaoyu Shen 0001 |
ICMR | 7 |
| 2025 | BGHR: Bridging the Gap Between HBox-Supervised and RBox-Supervised Oriented Object Detection via Adaptive Fine-Grained Sample MiningabstractOriented object detection is crucial for complex scenes such as aerial images and industrial inspection, providing precise delineation by minimizing background interference. Recently, the weakly-supervised detector paradigm H2RBox has demonstrated promise in learning rotated bounding box (RBox) from the more readily available horizontal bounding box (HBox), alleviating the scarcity and high cost of RBox annotations. However, these H2RBox-based methods have primarily focused on the gap in orientation information between HBox- and RBox-supervised approaches, overlooking the gap in training sample selection. In response, we propose the Adaptive Fine-grained Sample Mining (AFSM) strategy, which improves the selection of fine-grained training samples in HBox-supervised methods. AFSM assigns the best-matching prediction RBox to each ground truth (GT) HBox and selects positive samples based on these paired boxes. Furthermore, to effectively filter the best-matching prediction RBox for AFSM, we introduce the Prediction Rbox Assignment (PRA) scheme, employing Kullback-Leibler Divergence (KLD) as a localization quality metric. Additionally, we introduce an improved self-supervised branch loss (Lss) to address the symmetry of weakly-supervised branch prediction boxes. Incorporating these core components (AFSM, PRA, and Lss), we develop an end-to-end network architecture (BGHR) to further bridge the gap between HBox- and RBox-supervised oriented object detection. Extensive experiments on DOTA-v1.0 and DIOR-R demonstrate that BGHR achieves state-of-the-art performance compared to HBox-supervised methods without additional overhead. Even when benchmarked against fully supervised FCOS, our method still exhibits a slight performance advantage. Chenlin Fu, Yingying Zhu 0001 |
AAAI | 2 |
| 2025 | Adaptive Optimization Strategy for Semi-supervised Arbitrary-oriented Object DetectionabstractSemi-supervised arbitrary-oriented object detection (Semi-AOOD) has received increasing attention for its capability to improve detection performance by leveraging unlabeled data. Most existing works are based on the pseudo-label framework, which focuses on enhancing the consistency of the teacher-student network for better performance. However, they neglect the unique geometric characteristic of arbitrary-oriented objects during the feature extraction and overlook the potentially interfering information in pseudo-labels that may influence the optimization direction. In this paper, we propose a novel semi-supervised arbitrary-oriented object detection model called AODet (Adaptive Optimization Detector) to tackle these problems. First, we introduce an object adaptive strategy (OAS) to highlight the important regions while adaptively extracting the geometric features of arbitrary-oriented objects. Second, we develop an adaptive pseudo-label refinement strategy (APRS) to reduce the classification scores of low-confident categories in pseudo-labels. Extensive experiments on the DOTA dataset demonstrate the effectiveness of our AODet. Jiecong Chen, Chenlin Fu, Yingying Zhu 0001 |
ICME | 3 |
| 2025 | Rethinking Cross-view Object Geo-Localization: Towards Many-to-Many Real-world LocalizationabstractCross-view Object Geo-localization (CVOGL) determines the geographic location of objects in the ground-view image by matching them with corresponding objects in geo-tagged satellite imagery. Current research is limited to single-object localization, which significantly constrains CVOGL’s practical applications. We advance the CVOGL setting to more realistic scenarios by introducing a novel multi-object localization setting. This more realistic setting bridges the gap between current research and practical applications, which has never been explored before. Furthermore, we propose MTMGeo, a new model that delivers superior performance in Cross-view Multi-object Geo-localization (CVMOGL) while maintaining competitive accuracy in Cross-view Single-object Geo-localization (CVSOGL). At the core of MTMGeo is the Neighborhood Attention Cross-view Fusion Module (NA-CFM), which enhances object differentiation by leveraging neighborhood information to dynamically weight query objects. Through extensive evaluation on two public datasets, our method demonstrates state-of-the-art performance, achieving consistent and significant improvements across various CVOGL tasks. Qingwang Zhang, Yingying Zhu 0001 |
ICME | 3 |
| 2025 | PLGeo: A Patch-level Framework to Overcome Orientation Discrepancies in Cross-view Geo-localizationabstractCross-view geo-localization(CVGL) aims to determine the location of a ground-view image by referencing geo-tagged satellite-view images. Existing methods assume known ground-view image orientation-an unrealistic constraint in real-world scenarios where cameras have arbitrary orientations and limited fields of view (FOV).Unknown orientation cross-view geo-localization (UOCVGL) better reflects real-world applications but introduces severe feature alignment challenges due to misalignments in both viewpoints and orientations, significantly degrading localization accuracy. To address this challenge, we propose PLGeo, a novel method for UOCVGL, which includes two key components: (1) a Patch-wise Similarity Enhancement Component, which computes patch-level similarities between corresponding patches and refines alignments using learned attention weights, improving accuracy and mitigating issues caused by varying orientations and reduced FOV in ground-view images; and (2) an Attention-guided Patch Matching Component, which refines intra-domain feature matching within the same view by emphasizing stronger correspondences and suppressing weaker ones. We comprehensively evaluate PLGeo on several benchmark datasets under different settings, including unknown orientation, limited FOVs, robust datasets, the UAV dataset, north-aligned setting, and few-shot scenarios. Experimental results demonstrate that PLGeo consistently outperforms state-of-the-art methods, exhibiting remarkable robustness and generalization ability even in challenging real-world conditions. Code is available at https://github.com/1203ll/PLGeo. Yiru Li, Yingying Zhu 0001 |
ACM Multimedia | 2 |
| 2025 | CVGL: Causal Learning and Geometric TopologyabstractCross-view geo-localization (CVGL) aims to estimate the geographic location of a street image by matching it with a corresponding aerial image. This is critical for autonomous navigation and mapping in complex real-world scenarios. However, the task remains challenging due to significant viewpoint differences and the influence of confounding factors. To tackle these issues, we propose the Causal Learning and Geometric Topology (CLGT) framework, which integrates two key components: a Causal Feature Extractor (CFE) that mitigates the influence of confounding factors by leveraging causal intervention to encourage the model to focus on stable, task-relevant semantics; and a Geometric Topology Fusion (GT Fusion) module that injects Bird’s Eye View (BEV) road topology into street features to alleviate cross-view inconsistencies caused by extreme perspective changes. Additionally, we introduce a Data-Adaptive Pooling (DA Pooling) module to enhance the representation of semantically rich regions. Extensive experiments on CVUSA, CVACT, and their robustness-enhanced variants (CVUSA-C-ALL and CVACT-C-ALL) demonstrate that CLGT achieves state-of-the-art performance, particularly under challenging real-world corruptions. Songsong Ouyang, Yingying Zhu 0001 |
NeurIPS | 2 |
| 2025 | NSGHG: Neural surface guided generalizable human Gaussian splatting for sparse view synthesis
Yingying Zhu 0001 |
Neurocomputing | 3 |
| 2024 | Aligning Geometric Spatial Layout in Cross-View Geo-Localization via Feature RecombinationabstractCross-view geo-localization holds significant potential for various applications, but drastic differences in viewpoints and visual appearances between cross-view images make this task extremely challenging. Recent works have made notable progress in cross-view geo-localization. However, existing methods either ignore the correspondence between geometric spatial layout in cross-view images or require high costs or strict constraints to achieve such alignment. In response to these challenges, we propose a Feature Recombination Module (FRM) that explicitly establishes the geometric spatial layout correspondences between two views. Unlike existing methods, FRM aligns geometric spatial layout by directly recombining features, avoiding image preprocessing, and introducing no additional computational and parameter costs. This effectively reduces ambiguities caused by geometric misalignments between ground-level and aerial-level images. Furthermore, it is not sensitive to frameworks and applies to both CNN-based and Transformer-based architectures. Additionally, as part of the training procedure, we also introduce a novel weighted (B+1)-tuple loss (WBL) as optimization objective. Compared to the widely used weighted soft margin ranking loss, this innovative loss enhances convergence speed and final performance. Based on the two core components (FRM and WBL), we develop an end-to-end network architecture (FRGeo) to address these limitations from a different perspective. Extensive experiments show that our proposed FRGeo not only achieves state-of-the-art performance on cross-view geo-localization benchmarks, including CVUSA, CVACT, and VIGOR, but also is significantly superior or competitive in terms of computational complexity and trainable parameters. Our project homepage is at https://zqwlearning.github.io/FRGeo. Qingwang Zhang, Yingying Zhu 0001 |
AAAI | 2 |
| 2024 | CLIP-based Cross-Level Semantic Interaction and Recombination Network for Composed Image RetrievalabstractComposed image retrieval seeks to retrieve a target image that fulfills both modalities based on the user’s query, which comprises a modified text and a reference image. Although existing studies introduce novel multi-modal feature fusion techniques at the global or local level, they ignore exploring the cross-level semantic correspondence and recombining the original and target features in the reference image and the modified text, and the problems of modal inconsistency. To address these issues, we propose a CLIP-based Cross-Level Semantic Interaction and Recombination Network (SeIR). Specifically, we first resort the image and text encoders of the CLIP pre-trained model to narrow the modal gap between image and text, and extract both global and local features. We also introduce a cross-modal attention mechanism to screen out original and target features by exploring the semantic correlation between the cross-level image and text features. Subsequently, to alleviate the modal difference between the generated composed query representation and the target image, we utilize an affine transformation technique to recombine the original and target features from the image and text. Extensive experiments on the FashionIQ and CIRR benchmark datasets demonstrate the competitive performance of the proposed SeIR compared to the state-of-the-art methods. Yingying Zhu 0001 |
ECAI | 3 |
| 2024 | Benchmarking the Robustness of Cross-View Geo-Localization Models
Qingwang Zhang, Yingying Zhu 0001 |
ECCV (87) | 2 |
| 2024 | Contextual Interaction Enhancement Network for Smoke DetectionabstractSmoke detection can warn fires and prevent them. Visual-based smoke detection has gained more attention. However, there is a lack of publicly available smoke datasets in this area. To fill this gap, in this paper, we build a smoke dataset based on various realistic environments that contains 24776 smoke RGB images. Furthermore, a Contextual Interaction Enhancement Network for smoke detection is presented. Re-parameterized large kernel convolutions are utilized in the backbone to increase the receptive field of smoke feature extraction. And a Contextual Interaction Enhancement Module (CIEM)-neck is proposed to get aggregated features for better feature assignment to the decoupled head. Experimental results show that our method has excellent performance, achieving state-of-the-art results on the proposed Smoke 24776 dataset at 83.8% precision, 77.3% [email protected] and 49.8% mAP. Additionally, even if it is specialized for smoke detection, our method has still competitive performance and generalization on famous MS-COCO with 66% [email protected] and 47.0% mAP. Dataset is available at https://github.com/linjiefengFutureMediaSZU/Smoke24776. Jiefeng Lin, Chenlin Fu, Yingying Zhu 0001 |
ICME | 4 |
| 2024 | C2F-CCPE: Coarse-to-Fine Cross-View Camera Pose EstimationabstractImage-based cross-view localization often yields imprecise camera pose estimations due to the limited sampling density in the satellite image database. Current cross-view camera pose estimation approaches often overlook the importance of multi-scale features. In this paper, we introduce a novel coarse-to-fine cross-view camera pose estimation method (C2F-CCPE) that leverages multi-scale feature fusion and a localization and orientation feature fusion module (LOFFM) to enhance performance in localization and directional prediction. The proposed method outperforms state-of-the-art approaches on two benchmark datasets, addressing limitations in existing methods that focus on single-scale feature representations. And C2F-CCPE captures global and local information simultaneously, improving robustness against occlusions and enhancing precision in complex scenes. LOFFM further aggregates directional and positional information, enabling the network to deeply comprehend location features. Quantitative and qualitative experiments on two datasets not only demonstrate that our proposed method outperforms state-of-the-art models but also verify the effectiveness of our proposed method. Yingying Zhu 0001 |
ICME | 3 |
| 2024 | Libra-SOD: Balanced label assignment for small object detection
Yingying Zhu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | KLDet: Detecting Tiny Objects in Remote Sensing Images via Kullback-Leibler DivergenceabstractRemote sensing images (RSIs) frequently contain quite a few tiny objects with a finite number of pixels to study. The limited spatial information poses a challenge for extracting discriminative features for representing the characteristics of tiny objects. Existing solutions mainly focus on aggregating contextual information at different levels, while rarely touching the step that is crucial for model training, i.e., label assignment. Tiny instances occupy fairly small regions of images and have limited overlaps to priors (anchors or dots), which is a dilemma for traditional label assignment strategies. Despite being simple and effective, the mainstream Intersection over Union (IoU)-based label assignment strategy struggles to accurately measure the localization of tiny bounding boxes. In contrast, the Kullback–Leibler divergence (KLD) localization metric accurately reflects minor offsets of tiny bounding boxes. More importantly, KLD is able to measure non-overlapping bounding boxes, providing an advantage in mining more potential positive samples of tiny objects. In this article, from a cost-efficient point of view, we detect tiny objects through KLD in the form of single-stage framework. Specifically, we model the parameterized bounding box as a 2-D Gaussian distribution (Bbox2Gaussian) in order to use KLD as a localization metric. Then, we propose an adaptive online training sample mining (Ali-TSM) strategy based on inter-distribution similarity, which selects high-quality positive samples by considering localization and classification rather than just centroid distance or IoU. Finally, task-level attention (TlA) is introduced to guide the model in freely selecting the appropriate features for the classification or regression task. We conducted extensive experiments on four popular public datasets. Compared to the baseline, KLDet improves performance on Tiny Object Detection in Aerial Images (AI-TOD) and object DetectIon in Optical Remote sensing image (DIOR) by 4.1 AP and 6.7 mAP. On VisDrone and Small Object Detection dAtasets (SODA-D), KLDet exhibits superior performance than baseline. The code is available athttps://github.com/TinyOD/mmdet-kldet. Yingying Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | RaFPN: Relation-Aware Feature Pyramid Network for Dense Image PredictionabstractIntuitively, relations among objects assist a model in performing inference under constrained environments. However, the top-down information flow in the Feature pyramid network (FPN) dilutes the relation features contained in the non-adjacent layers. Such a defect reduces the accuracy of detectors, especially for small or obscured objects. To adequately exploit the relations among object instances, we propose the relation-aware feature pyramid network (RaFPN), a simple but effective balanced multi-scale feature module for dense image prediction. RaFPN models the relations among objects by computing the similarity between pixels located on cross-scale features. The result is then delivered to FPN to guide the detector in completing accurate inference. Specifically, we first generate a pair of cross-scale aggregated features based on the channel importance of the output features from FPN. After that, the relation among the cross-scale objects is extracted by a bi-directional interaction mechanism. Finally, relation features are injected directly into each layer of the feature pyramid to avoid dilution. In this way, the relation among instances can adequately guide the detector for dense prediction. Our RaFPN pushes the performance bound of Faster RCNN by 2.0 AP (average precision), outperforming the recent state-of-the-art FPN-based improvements. Notably, for dense prediction tasks such as instance, semantic, and panoptic segmentation, our method brings consistent boosts to them as well. Yingying Zhu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Cross-modal Semantically Augmented Network for Image-text MatchingabstractImage-text matching plays an important role in solving the problem of cross-modal information processing. Since there are nonnegligible semantic differences between heterogenous pairwise data, a crucial challenge is how to learn a unified representation. Existing methods mainly rely on the alignment between regional image features and corresponding entity words. However, the regional features in the image are often more concerned with the foreground entity information, and the attribute information of the entities and the relational information are ignored. How to effectively integrate entity-attribute alignment and relationship alignment has not been fully studied. Therefore, we propose a Cross-Modal Semantically Augmented Network for Image-Text Matching (CMSAN), which combines the relationships between entities in the image with the semantics of relational words in the text. CMSAN (1) proposes an adaptive word-type prediction model that classifies the words into four types, i.e., entity word, attribute word, relation word, and unnecessary word. It can align different image features at multiple levels. CMSAN (2) designs a sophisticated relationship alignment module and an entity-attribute alignment module that maximizes the exploitation of the semantic information, which enables the model to have more discriminative power and further improves the matching accuracy. Yiru Li, Ying Li 0016, Yingying Zhu 0001, Gang Wang 0029, Jun Yue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Visual Place Recognition Datasets for Indoor SpacesabstractVisual place recognition (VPR) is widely cast as a challenging image retrieval problem. Recently, many studies in this area have achieved superior results. However, most existing works consider outdoor spaces rather than indoor spaces. To fill this gap, this paper for the first time constructs a benchmark based on realistic environments for indoor VPR, involving three new indoor datasets which contains 25, 233 RGB images in total. These datasets cover typical indoor environments with over 250 places and provide a wide range of challenging cases. Moreover, this paper introduces a patch relation module based on spatial coordinate position of image patches and a global average pooling pyramid to get discriminative and robust features. Extensive experiments are conducted on our datasets to validate the effectiveness of the proposed method. The results indicate that indoor VPR in realistic setting is still challenging, fostering new research in this direction. Our dataset and code will be released at https://github.com/Dauntless-Wind/indoor-visual-place-recognition. Zemian Guo, Yingying Zhu 0001 |
ICME | 2 |
| 2023 | Visual-Linguistic Alignment and Composition for Image Retrieval with Text FeedbackabstractIn this paper, we focus on the task of image retrieval with text feedback, which maintains two key challenges. One is the misalignment problem between different modalities, and the other is to selectively alter the corresponding attributes on the reference image according to the textual words. To this end, we propose a novel visual-linguistic alignment and composition network (ACNet) consisting of two key components: the modality alignment module (MAM) and the relation composition module (RCM). Specifically, the MAM performs alignment between the features from different modalities by applying image-text contrastive loss. The RCM correlates the image regions with their corresponding words and then adaptively modifies the specific regions of the reference image conditioned on textual semantics. Quantitative and Qualitative experiments on three datasets not only demonstrate that our ACNet outperforms state-of-the-art models, but also verify the effectiveness of our method. Dafeng Li, Yingying Zhu 0001 |
ICME | 2 |
| 2023 | Large-Scale Image Retrieval with Deep Attentive Global FeaturesabstractHow to obtain discriminative features has proved to be a core problem for image retrieval. Many recent works use convolutional neural networks to extract features. However, clutter and occlusion will interfere with the distinguishability of features when using convolutional neural network (CNN) for feature extraction. To address this problem, we intend to obtain high-response activations in the feature map based on the attention mechanism. We propose two attention modules, a spatial attention module and a channel attention module. For the spatial attention module, we first capture the global information and model the relation between channels as a region evaluator, which evaluates and assigns new weights to local features. For the channel attention module, we use a vector with trainable parameters to weight the importance of each feature map. The two attention modules are cascaded to adjust the weight distribution for the feature map, which makes the extracted features more discriminative. Furthermore, we present a scale and mask scheme to scale the major components and filter out the meaningless local features. This scheme can reduce the disadvantages of the various scales of the major components in images by applying multiple scale filters, and filter out the redundant features with the MAX-Mask. Exhaustive experiments demonstrate that the two attention modules are complementary to improve performance, and our network with the three modules outperforms the state-of-the-art methods on four well-known image retrieval datasets. Yingying Zhu 0001, Yinghao Wang, Zemian Guo |
Int. J. Neural Syst. | 1 |
| 2023 | Learning relation-based features for fine-grained image retrieval
Yingying Zhu 0001, Zhanyuan Yang, Xiufan Lu |
Pattern Recognit. | 1 |
| 2023 | Cross-View Image Synthesis From a Single Image With Progressive Parallel GANabstractCross-view image synthesis aims to synthesize a ground-view image covering the same geographic region for a given single aerial-view image (or vice versa). Existing approaches typically tackle this challenging task by relaxing the single-image constraint and using a ground-truth semantic map as additional input to aid synthesis. However, this is nearly infeasible in practice. In this paper, we investigate how to generate a detail-enriched and structurally accurate ground-level image from only a single aerial-level input image, in which there are no other prior knowledge except for the input image. Towards this goal, we propose a novel Progressive Parallel Generative Adversarial Network (PPGAN) that starts from generating low-resolution outputs and progressively produces ground images at higher resolutions as the network propagates forward. In this manner, our PPGAN decomposes the task into several manageable sub-tasks, which helps to generate detail-enriched and structurally accurate ground images. During progressive generation, the PPGAN employs a parallel generation paradigm that enables the generator to produce multi-resolution images in parallel, thereby avoiding excessive time cost on training. Furthermore, for effective information propagation across multi-resolution images, a feature fusion module (FFM) is devised to mitigate the domain gap between cross-level image features, which enables a balance of detail and structural information synthesis. Additionally, the proposed Channel-Space Attention Selection Module (CSASM) learns the mapping relationship between cross-view images in a larger scale space to enhance the quality of the output image. Quantitative and qualitative experiments demonstrate that, our method requires only one input image without the aid of additional inputs, but is capable of synthesizing detail-enriched and structurally accurate ground images and outperforms the existing state-of-the-art methods on two famous benchmarks. Yingying Zhu 0001, Shihai Chen, Xiufan Lu, Jianyong Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Few-Shot Classification with Contrastive Learning
Zhanyuan Yang, Yingying Zhu 0001 |
ECCV (20) | 3 |
| 2022 | DMPCANet: A Low Dimensional Aggregation Network for Visual Place RecognitionabstractVisual place recognition (VPR) aims to estimate the geographical location of a query image by finding its nearest reference images from a large geo-tagged database. Most of the existing methods adopt convolutional neural networks to extract feature maps from images. Nevertheless, such feature maps are high-dimensional tensors, and it is a challenge to effectively aggregate them into a compact vector representation for efficient retrieval. To tackle this challenge, we develop an end-to-end convolutional neural network architecture named DMPCANet. The network adopts the regional pooling module to generate feature tensors of the same size from images of different sizes. The core component of our network, the Differentiable Multilinear Principal Component Analysis (DMPCA) module, directly acts on tensor data and utilizes convolution operations to generate projection matrices for dimensionality reduction, thereby reducing the dimensionality to one sixteenth. This module can preserve crucial information while reducing data dimensions. Experiments on two widely used place recognition datasets demonstrate that our proposed DMPCANet can generate low-dimensional discriminative global descriptors and achieve the state-of-the-art results. Yinghao Wang, Yingying Zhu 0001 |
ICMR | 4 |
| 2022 | It's Okay to Be Wrong: Cross-View Geo-Localization With Step-Adaptive Iterative RefinementabstractCross-view image geo-localization is a challenging task of estimating the geospatial location of a street view image by matching it with a database of geotagged aerial/satellite images, and vice versa. Compared to existing CNN-based approaches that attempt to generate discriminative representations in a single step for this task, in this paper, we instead advocate endowing the network with the capability of progressive self-correcting. Towards this target, we propose a novel step-adaptive iterative refinement network (SIRNet), which decomposes the complex learning process into several refinement steps while adapting the refinement steps specifically for each input. Specifically, the SIRNet takes the output of the backbone as a rough network prediction and iteratively refines it via an iterative refinement module (IRM). The IRM cascades several refinement blocks sharing the same structure for progressive self-correcting. For each refinement block, the goal is to improve the output of the previous refinement block under the guidance of height-wise context. In this way, the IRM is capable of improving the rough network prediction step by step, and the refined features are increasingly focused on more discriminative scene regions as they are iteratively refined. In addition, considering different characteristics of input images, we devise an adaptive step estimation (ASE) mechanism, which enables our SIRNet to adapt the number of refinement steps to each input automatically. Concretely, the ASE is performed by comparing features at adjacent refinement steps, estimating whether the next step brings improvements, and finally making a halting decision at each refinement step. With the ASE, our SIRNet becomes a dynamic architecture that considers different characteristics of the inputs when performing the iterative refinement. Extensive experiments demonstrate that our SIRNet performs favorably against state-of-the-art methods on the CVUSA and the CVACT datasets. Furthermore, quantitative and qualitative experimental results demonstrate our approach’s wide applicability, impressive generalization ability, and robustness. Xiufan Lu, Yingying Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Geographic Semantic Network for Cross-View Image Geo-LocalizationabstractThe task of cross-view image geo-localization aims to determine the geo-location (Global Positioning System (GPS) coordinates) of a query ground-view image by matching the image with GPS-tagged aerial (or satellite) images in the reference dataset. Due to the dramatic domain gap between the ground and aerial images, the problem is challenging. The existing approaches mainly adopt convolutional neural networks (CNNs) to learn discriminative features. However, these CNN-based methods mainly leverage appearance and semantic information but fail to jointly model the appearance, positional, and orientation properties of scene objects, which belong to the spatial hierarchy. Since spatial hierarchy information is crucial for cross-view feature correspondence, in this article, we propose an end-to-end network architecture, dubbed GeoNet. GeoNet consists of a ResNetX module and a GeoCaps module. On the one hand, the ResNetX module is developed to learn powerful intermediate feature maps and allows the stable propagation of gradients in deep CNNs. On the other hand, the GeoCaps module utilizes the capsule network to encapsulate the intermediate feature maps into several capsules, whose length and orientation represent the existence probability and spatial hierarchy information of scene objects, respectively. Moreover, by using a dynamic routing-by-agreement mechanism, the GeoCaps module is capable of modeling parts-to-whole relationships between scene objects, which is viewpoint invariant and capable of bridging the cross-view domain gap. In addition to GeoNet, we introduce a simple yet effective metric learning method, based on which two weighted soft margin loss functions with online batch hard sample mining are devised. These functions not only speed up convergence but also improve the generalization ability of the network. Extensive experiments on three well-known datasets demonstrate that our GeoNet not only achieves state-of-the-art results for the ground-to-aerial and aerial-to-ground geo-localization tasks but also outperforms competing approaches for the few-shot geo-localization task. Yingying Zhu 0001, Xiufan Lu, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Associative Segmentation for Instances and Semantics by Perceiving Neighborhood in Point CloudsabstractIn the process of human visual perception, when facing complex scenes, people often rely on the neighborhood information of an object to aid in understanding scene, which is applied to robot visual perception as well. In this paper, we propose a neighborhood-aware module (NAM) which captures rich instance-level contextual dependencies to reduce the search space of possible categories for segmentation tasks. To address the instance and semantic segmentation associatively, we further design a Associative segmentation module (ASM) to make two tasks promote each other and get a win-win solution. Experimental results on two well-known dataset (S3DIS and ShapeNet) show that our NAM is capable of caturing contextual dependency and ASM boosts the performance by enabling semantic segmentation and instance segmentation to take advantage of each other. Our method largely outperforms the state-of-the-art methods in 3D instance segmentation, as well as achieving a significant improvement in 3D semantic segmentation. Yingying Zhu 0001 |
ICME | 1 |
| 2021 | Fine-Grained Image Retrieval Via Multiple Part-Level Feature EnsembleabstractFine-Grained Image Retrieval (FGIR) has become a crucial research area of fine-grained analysis. Despite the extensive progress in FGIR, the main problem remains. Existing methods rely mainly on object-level representations, which are interrupted by background clutters. Instead of object-level features, in this paper, we propose to learn part-level representations. We first present a novel Attention-Activation-based Part Detector (AAPD) without any part-level annotations to extract part-level features and to remove background noises. AAPD not only localizes the discriminative part by attention mechanism automatically but also selects the part with high activation values in an unsupervised way. Then we propose a novel unified Multiple Part-level Feature Ensemble (MPFE) framework to assemble serval part-level features extracted by AAPDs. Finally, the proposed MPFE is evaluated on two widely-used benchmark datasets, including CUB-200-2011 and Cars-196. The MPFE Framework is capable of localizing discriminative parts and achieves state-of-the-art performances on two datasets. Yingying Zhu 0001, Xiufan Lu |
ICME | 2 |
| 2021 | Cross-view Geo-localization with Layer-to-Layer TransformerabstractIn this work, we address the problem of cross-view geo-localization, which estimates the geospatial location of a street view image by matching it with a database of geo-tagged aerial images. The cross-view matching task is extremely challenging due to drastic appearance and geometry differences across views. Unlike existing methods that predominantly fall back on CNN, here we devise a novel layer-to-layer Transformer (L2LTR) that utilizes the properties of self-attention in Transformer to model global dependencies, thus significantly decreasing visual ambiguities in cross-view geo-localization. We also exploit the positional encoding of the Transformer to help the L2LTR understand and correspond geometric configurations between ground and aerial images. Compared to state-of-the-art methods that impose strong assumptions on geometry knowledge, the L2LTR flexibly learns the positional embeddings through the training objective. It hence becomes more practical in many real-world scenarios. Although Transformer is well suited to our task, its vanilla self-attention mechanism independently interacts within image patches in each layer, which overlooks correlations between layers. Instead, this paper proposes a simple yet effective self-cross attention mechanism to improve the quality of learned representations. Self-cross attention models global dependencies between adjacent layers and creates short paths for effective information flow. As a result, the proposed self-cross attention leads to more stable training, improves the generalization ability, and prevents the learned intermediate features from being overly similar. Extensive experiments demonstrate that our L2LTR performs favorably against state-of-the-art methods on standard, fine-grained, and cross-dataset cross-view geo-localization tasks. The code is available online. Xiufan Lu, Yingying Zhu 0001 |
NeurIPS | 3 |
| 2020 | Regional Relation Modeling for Visual Place RecognitionabstractIn the process of visual perception, humans perceive not only the appearance of objects existing in a place but also their relationships (e.g. spatial layout). However, the dominant works on visual place recognition are always based on the assumption that two images depict the same place if they contain enough similar objects, while the relation information is neglected. In this paper, we propose a regional relation module which models the regional relationships and converts the convolutional feature maps to the relational feature maps. We further design a cascaded pooling method to get discriminative relation descriptors by preventing the influence of confusing relations and preserving as much useful information as possible. Extensive experiments on two place recognition benchmarks demonstrate that training with the proposed regional relation module improves the appearance descriptors and the relation descriptors are complementary to appearance descriptors. When these two kinds of descriptors are concatenated together, the resulting combined descriptors outperform the state-of-the-art methods. Yingying Zhu 0001, Zhou Zhao 0001 |
SIGIR | 1 |
| 2020 | A Filter Model Based on Hidden Generalized Mixture Transition Distribution Model for Intrusion Detection System in Vehicle Ad Hoc NetworksabstractVehicle ad hoc networks (VANETs) are considered to be the next big thing that will remarkably change our lives, since this kind of technology is able to make our lives and roads safer. Due to the very fast move and high dynamic in VANETs, it is important to quickly ascertain the reliability of information. Although intrusion detection system (IDS) has been proposed as a reliable approach to protect VANETs against attacks, its overhead is serious, which spends too much time on detection, especially when the number of vehicles increases. Thus, in this paper, we propose a novel filter model based on a hidden generalized mixture transition distribution model (HgMTD) in VANETs, called FM-HgMTD, which can quickly filter the messages from neighboring vehicles so as to reduce the overhead and detection time. It adopts a well-known multi-objective optimization (NSGA-II) algorithm combined with an expectation-maximization (EM) algorithm to forecast the future states of neighboring vehicles and then to filter out malicious messages, by monitoring the change of the state pattern of each neighboring vehicle. In addition, a timeliness method is used to maintain the accuracy of the forecast. The experiments show that IDS with the proposed FM-HgMTD has better performance than other available IDSs in terms of detection rate, detection time, and overhead. Junwei Liang 0004, Qiuzhen Lin, Jianyong Chen, Yingying Zhu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | GEOCAPSNET: Ground to Aerial View Image Geo-Localization using Capsule NetworkabstractThe task of cross-view image geo-localization aims to determine the geo-location (GPS coordinates) of a query ground-view image by matching it with the GPS-tagged aerial (satellite) images in a reference dataset. Due to the dramatic changes of viewpoint, matching the cross-view images is challenging. In this paper, we propose the GeoCapsNet based on the capsule network for ground-to-aerial image geo-localization. The network first extracts features from both ground and aerial images via standard convolution layers and the capsule layers further encode the features to model the spatial feature hierarchies and enhance the representation power. Moreover, we introduce a simple and effective weighted soft-margin triplet loss with online batch hard sample mining, which can greatly improve the image retrieval accuracy. Experimental results show that our GeoCapsNet significantly outperforms the state-of-the-art approaches on two benchmark datasets. Chen Chen 0001, Yingying Zhu 0001, Jianmin Jiang |
ICME | 3 |
| 2019 | Intelligent Image Retrieval Based on Multi-swarm of Particle Swarm Optimization and Relevance Feedback
Yingying Zhu 0001, Yishan Chen 0004, Wenlong Han, Zhenkun Wen |
ICONIP (2) | 1 |
| 2019 | Learning Discriminative Features for Image RetrievalabstractDiscriminative local features obtained from activations of convolutional neural networks have proven to be essential for image retrieval. To improve retrieval performance, many recent works aim to obtain more powerful and discriminative features. In this work, we propose a new attention layer to assess the importance of local features and assign higher weights to those more discriminative. Furthermore, we present a scale and mask module to filter out the meaningless local features and scale the major components. This module not only reduces the impact of the various scales of the major components in images by scaling them on the feature maps, but also filters out the redundant and confusing features with the MAX-Mask. Finally, the features are aggregated into the image representation. Experimental evaluations demonstrate that the proposed method outperforms the state-of-the-art methods on standard image retrieval datasets. Yinghao Wang, Chen Chen 0001, Yingying Zhu 0001 |
ICMR | 4 |
| 2019 | Hybrid feature-based analysis of video's affective content using protagonist detection
Yingying Zhu 0001, Min Tong, Zhengbo Jiang, Shenghua Zhong, Qi Tian 0001 |
Expert Syst. Appl. | 1 |
| 2018 | A Multi-indicator Feature Selection for CNN-Driven Stock Index Prediction
Yingying Zhu 0001 |
ICONIP (5) | 2 |
| 2018 | An Adaptive Box-Normalization Stock Index Trading Strategy Based on Reinforcement Learning
Yingying Zhu 0001, Jianmin Jiang |
ICONIP (3) | 1 |
| 2018 | Attention-based Pyramid Aggregation Network for Visual Place RecognitionabstractVisual place recognition is challenging in the urban environment and is usually viewed as a large scale image retrieval task. The intrinsic challenges in place recognition exist that the confusing objects such as cars and trees frequently occur in the complex urban scene, and buildings with repetitive structures may cause over-counting and the burstiness problem degrading the image representations. To address these problems, we present an Attention-based Pyramid Aggregation Network (APANet), which is trained in an end-to-end manner for place recognition. One main component of APANet, the spatial pyramid pooling, can effectively encode the multi-size buildings containing geo-information. The other one, the attention block, is adopted as a region evaluator for suppressing the confusing regional features while highlighting the discriminative ones. When testing, we further propose a simple yet effective PCA power whitening strategy, which significantly improves the widely used PCA whitening by reasonably limiting the impact of over-counting. Experimental evaluations demonstrate that the proposed APANet outperforms the state-of-the-art methods on two place recognition benchmarks, and generalizes well on standard image retrieval datasets. Yingying Zhu 0001, Lingxi Xie, Liang Zheng 0001 |
ACM Multimedia | 1 |
| 2018 | Haze removal method for natural restoration of images with sky
Yingying Zhu 0001, Gaoyang Tang, Xiaoyan Zhang 0002, Jianmin Jiang, Qi Tian 0001 |
Neurocomputing | 1 |
| 2017 | Study of subjective and objective quality assessment for screen content imagesabstractIn this paper, we present the results of a recent large-scale subjective study of image quality on a collection of screen contents distorted by a variety of application-relevant processes. With the development of multi-device interactive multimedia applications, metrics to predict the visual quality of screen content images (SCIs) as perceived by subjects are becoming fundamentally important. For developing the objective image quality assessment (IQA) method, there is a need for large-scale public database with diversity of distorted types and scene contents, and available subjective scores of distorted SCIs. The resulting Immersive Media Laboratory screen content image quality database (IML-SCIQD) contains 1250 distorted SCIs from 25 reference SCIs with 10 distortion types. Each image was rated by 35 human observers, and the different mean opinion scores (DMOS) were obtained after data processing. The performance comparison of 17 state-of-the-arts, publicly available IQA algorithms are evaluated on the new database. The database will be available online in our project website. Xu Wang 0006, Yingying Zhu 0001, Yun Zhang 0002, Jianmin Jiang, Sam Kwong |
ICIP | 3 |
| 2017 | Adaptive Dehaze Method for Aerial Image Processing
Rong-Qin Xu, Shenghua Zhong, Gaoyang Tang, Jiaxin Wu 0001, Yingying Zhu 0001 |
PSIVT | 5 |
| 2017 | Interpretation of users' feedback via swarmed particles for content-based image retrieval
Yingying Zhu 0001, Jianmin Jiang, Wenlong Han, Qi Tian 0001 |
Inf. Sci. | 1 |
| 2017 | An improved NSGA-III algorithm for feature selection used in intrusion detection
Yingying Zhu 0001, Junwei Liang 0004, Jianyong Chen, Zhong Ming 0001 |
Knowl. Based Syst. | 1 |
| 2016 | Visual Orientation Inhomogeneity Based Convolutional Neural NetworksabstractThe details of oriented visual stimuli are better resolved when they are horizontal or vertical rather than oblique. This "oblique effect" has been researched and confirmed in numerous research studies, including behavioral studies and neurophysiological and neuroimaging findings. Although the "oblique effect" has influence in many fields, little research integrated it into computational models. In this paper, we try to explore this inhomogeneity of visual orientation based on Convolutional neural networks (CNNs) in image recognition. We validate that visual orientation inhomogeneity CNNs can achieve comparable performance with higher computational efficiency on various datasets. We can also get the conclusion that, compared with the cardinal information, oblique information is indeed less useful in natural color image recognition. Through the exploration of the proposed model on image recognition, we gain more understanding of the inhomogeneity of visual orientation. It also illuminates a wide range of opportunities for integrating the inhomogeneity of visual orientation with other computational models. Shenghua Zhong, Jiaxin Wu 0001, Yingying Zhu 0001, Peiqi Liu, Jianmin Jiang, Yan Liu 0004 |
ICTAI | 3 |
| 2016 | Large-scale video copy retrieval with temporal-concentration SIFT
Yingying Zhu 0001, Qi Tian 0001 |
Neurocomputing | 1 |
| 2015 | A Temporal-Compress and Shorter SIFT Research on Web VideosabstractThe large-scale video data on the web contain a lot of semantics, which are an important part of semantic web. Video descriptors can usually represent somewhat the semantics. Thus, they play a very important role in web multimedia content analysis, such as Scale-invariant feature transform (SIFT) feature. In this paper, we proposed a new video descriptor, called a temporal-compress and shorter SIFT(TC-S-SIFT) which can efficiently and effectively represent the semantics of web videos. By omitting the least discriminability orientation in three stages of standard SIFT on every representative frame, the dimensions of the shorter SIFT are reduced from 128-dimension to 96-dimension to save space storage. Then, the SIFT can be compressed by tracing SIFT features on video temporal domain, which highly compress the quantity of local features to reduce visual redundancy, and keep basically the robustness and discrimination. Experimental results show our method can yield comparable accuracy and compact storage size. Yingying Zhu 0001, Chuanhua Jiang, Zhijiao Xiao, Shenghua Zhong |
KSEM | 1 |
| 2015 | A solution of dynamic VMs placement problem for energy consumption optimization based on evolutionary game theory
Zhijiao Xiao, Jianmin Jiang, Yingying Zhu 0001, Zhong Ming 0001, Shenghua Zhong, Shubin Cai |
J. Syst. Softw. | 3 |
| 2008 | Video scene classification and segmentation based on Support Vector MachineabstractVideo scene classification and segmentation are fundamental steps for multimedia retrieval, indexing and browsing. In this paper, a robust scene classification and segmentation approach based on support vector machine (SVM) is presented, which extracts both audio and visual features and analyzes their inter-relations to identify and classify video scenes. Our system works on content from a diverse range of genres by allowing sets of features to be combined and compared automatically without the use of thresholds. With the temporal behaviors of different scene classes, SVM classifier can effectively classify presegmented video clips into one of the predefined scene classes. After identifying scene classes, the scene change boundary can be easily detected The experimental results show that the proposed system not only improves precision and recall, but also performs better than the other classification systems using the decision tree (DT), K nearest neighbor (K-NN) and neural network (NN). Yingying Zhu 0001, Zhong Ming 0001, Jun Zhang 0003 |
IJCNN | 1 |
| 2007 | Evaluation of Dependable Embedded System
Xiaofeng Liang, Yingying Zhu 0001 |
ICIC (3) | 3 |
| 2003 | Video Browsing and Retrieval Based on Multimodal IntegrationabstractThe rapid growth of multimedia data requires more effective content-based video browsing and retrieval. We present a system developed for video browsing and retrieval based on multimedia integration. First, a basic structure of the system is defined. Second, a robust scene segmentation method is presented, which analyzes audio and visual information and accounts for their interrelations and coincidence to semantically identify video scenes. We then extract text from key frames with video OCR technique and extract text transcriptions by speech recognition to classify video scenes and form the full-text indices. Finally, natural language understanding technique is used to automatically classify video scenes on the basis of the texts obtained from close caption, video OCR process and speech recognition. In this way, we have developed the content-based video database system which integrates multimodality to browse and retrieve video data. The experimental results show that multimodal integration is effective for video scene segmentation. Our system built on the idea of multimodal integration makes content-based browsing and retrieval of video data, key-frame-based video abstract and search by keywords practical. Yingying Zhu 0001, Dongru Zhou |
Web Intelligence | 1 |