EDBT 2026 Demo / reviewers in the wild / expert
Yingmei Wei
dblp:70/1797
· DBLP profile ↗
27ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0003-4568-551XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciceVPR: Stable cross-image correlation enhanced model for visual place recognition
Shanshan Wan, Yingmei Wei, Lai Kang, Tianrui Shen, Haixuan Wang, Yee-Hong Yang |
Neurocomputing | 2 |
| 2026 | Facade parsing via joint structural priors and phased deep learning
Yuning Huang, Weize Quan, Dong-Ming Yan 0001, Jie Jiang 0017, Yingmei Wei |
Neurocomputing | 6 |
| 2026 | BGC-Net: Bilateral Graph Convolutional Network for Weakly Supervised Semantic Segmentation of Large-Scale Point CloudsabstractWeakly-supervised point cloud semantic segmentation (WS-PCS) has attracted increasing attention due to the challenge of sparse annotations. A central problem is how to effectively extract informative features from the annotated points, enabling reliable supervision. Although many existing works extend 2D graph convolution to 3D point cloud data, 2D convolution inherently assumes feature localization, which is an assumption that does not hold in point clouds, and lacks consistent semantic offsets. To address this, we propose a novel Bilateral Graph Convolutional (BGC) method, which refines graph edges into two categories: regular edges and offset edges, providing improved guidance for WS-PCS. Firstly, we create the Local Bilateral Relations (LBR) module to learn the relational features of edges in local point cloud graphs, encompassing both regular and offset edges. To the best of our knowledge, we are the first to utilize offset edges to capture irregular semantic offsets in point cloud data. Secondly, we propose the Adaptive Pooling (AP) module, which adaptively pools edge information learned from LBR, enhancing the feature characterization ability by incorporating salient and pervasive features. Finally, we design BGC as BGC-Net and evaluate its performance against recent networks on four datasets, achieving state-of-the-art results. Lixin Zhan, Yukun Du, Jie Jiang 0017, Yingmei Wei, Tianjian Zhou, Ziyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud CompletionabstractPoint cloud completion aims to reconstruct the complete 3D shape from incomplete point clouds, and it is crucial for tasks such as 3D object detection and segmentation. Despite the continuous advances in point cloud analysis techniques, feature extraction methods are still confronted with apparent limitations. The sparse sampling of point clouds, used as inputs in most methods, often results in a certain loss of global structure information. Meanwhile, traditional local feature extraction methods usually struggle to capture the intricate geometric details. To overcome these drawbacks, we introduce PointCFormer, a transformer framework optimized for robust global retention and precise local detail capture in point cloud completion. This framework embraces several key advantages. First, we propose a relation-based local feature extraction method to perceive local delicate geometry characteristics. This approach establishes a fine-grained relationship metric between the target point and its k-nearest neighbors, quantifying each neighboring point's contribution to the target point's local features. Secondly, we introduce a progressive feature extractor that integrates our local feature perception method with self-attention. Starting with a denser sampling of points as input, it iteratively queries long-distance global dependencies and local neighborhood relationships. This extractor maintains enhanced global structure and refined local details, without generating substantial computational overhead. Additionally, we develop a correction module after generating point proxies in the latent space to reintroduce denser information from the input points, enhancing the representation capability of the point proxies. PointCFormer demonstrates state-of-the-art performance on several widely used benchmarks. Weize Quan, Dong-Ming Yan 0001, Jie Jiang 0017, Yingmei Wei |
AAAI | 5 |
| 2025 | Exploiting Event Temporal Dynamics and Sparsity Characteristics for RGB-Event Fusion Semantic SegmentationabstractFrame-based semantic segmentation faces information loss due to the limited dynamic range of conventional cameras. Event cameras, with their high dynamic range and temporal resolution, offer a promising solution. Unlike RGB images, event cameras produce sparse, asynchronous event streams, prompting a reconsideration of the event generation mechanism and a comprehensive exploration of their characteristics. Previous event representation methods have been limited to fixed time windows, neglecting the rich temporal dynamics inherent in event streams. Furthermore, the inherent noise and calibration deficiencies in event data present significant challenges for RGB-event fusion. We propose an Event-driven Fusion Network (EFNet) to improve semantic segmentation by leveraging event camera characteristics. To address the first challenge, we introduce a Dual-Temporal Event Integration (DT-EI) module, which leverages temporal dynamics to generate and integrate dual-temporal event representations, capturing objects with varying motion speeds. For the second challenge, we propose an Event-Count Guided Recalibration and Fusion (EGRF) module, which utilizes event sparsity to guide feature recalibration through Event-Count Attention Maps, combined with a bidirectional cross-attention mechanism for adaptive fusion of event and image features. Experimental results demonstrate that EFNet outperforms state-of-the-art methods in event-based semantic segmentation. Yingmei Wei, Yanming Guo, Jiangming Chen |
ICMR | 2 |
| 2025 | FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
Yingmei Wei |
PRCV (5) | 4 |
| 2025 | DJIST: Decoupled joint image and sequence training framework for sequential visual place recognition
Shanshan Wan, Lai Kang, Yingmei Wei, Tianrui Shen, Haixuan Wang |
Neurocomputing | 3 |
| 2025 | EntroCap: Zero-shot image captioning with entropy-based retrieval
Yuxiang Xie, Shiwei Zou, Yingmei Wei, Xidao Luan |
Neurocomputing | 4 |
| 2025 | Few-shot event-based action recognitionabstractDespite the evident superiority of event cameras in practical vision applications (e.g., action recognition), owing to their distinctive sensing mechanism, existing event-based action recognition methods rely heavily on large-scale training data. However, the expensive cost of camera deployment and the requirement of data privacy protection make it challenging to collect substantial data in real-world scenarios. To address this limitation, we explore a novel yet practical task, Few-Shot Event-Based Action Recognition (FSEAR), which aims at leveraging a minimal number of intractable event action data for model training and accurately classifying unlabeled data into a specific category. Accordingly, we design a new framework for FSEAR, including a Noise-Aware Event Encoder (NAE) and a Distilled Prototypical Distance Fusion (DPDF). The former efficiently filters noise within the spatiotemporal domain while retaining vital information related to action timing. The latter conducts multi-scale measurements across geometric, directional, and distributional dimensions. These two modules benefit mutually and thus effectively exploit the potential characteristics of event data. Extensive experiments on four distinct event action recognition datasets have demonstrated the significant advantages of our model over other few-shot learning methods. Our code and models will be publicly released. Zanxi Ruan, Nan Pu, Jiangming Chen, Songqun Gao, Yanming Guo, Qiuyu Kong, Yuxiang Xie, Yingmei Wei |
Neural Networks | 8 |
| 2025 | Enhancing spatial perception and contextual understanding for 3D dense captioning
Yuxiang Xie, Shiwei Zou, Yingmei Wei, Xidao Luan |
Neural Networks | 4 |
| 2025 | Sorted Texture-Aware Glance and Gaze Network for Hyperspectral Image Classification With Low Training SamplesabstractHyperspectral images (HSI) provide a wealth of information surpassing human visual capabilities, enabling precise identification of remote sensing targets. However, it faces significant challenges, including insufficient long-range dependency modeling, difficulties in data collection, and the tendency of models to get trapped in local optima during training. To overcome these obstacles, we present the sorted texture-aware glance and gaze network (ST-GGNet) tailored for HSI classification. First, we propose the glance and gaze attention (GGA) mechanism, which employs feature interaction-based long-term modeling to minimize information loss across spectral bands and focus on critical land cover features within HSI. Subsequently, the sorted texture-aware module (STM) is introduced to deeply mine and efficiently utilizes detailed texture and spectral information, thereby enhancing accuracy even with limited training data. Additionally, we propose the budding growth optimization algorithm (BGO), which integrates a budding growth mechanism to help the model discover better solutions, boosting optimization and classification performance. Experimental evaluations conducted on four public HSI datasets—Pavia University, Salinas, Houston, and WHU-Longkou—demonstrate the superior performance of ST-GGNet compared to nine state-of-the-art (SOTA) classification methods. Specifically, under limited training samples, ST-GGNet achieves overall accuracies (OA) of 99.42%, 96.88%, 96.86%, and 97.74%; average accuracies (AA) of 98.90%, 98.01%, 97.07%, and 92.48%; and Kappa coefficients of 99.24%, 96.53%, 96.59%, and 97.03% respectively. The findings reveal that ST-GGNet not only maintains strong robustness and generalization but also effectively suppresses noise and excels at distinguishing spatially similar adjacent land covers, especially in low-samples scenarios, consistently outperforming existing SOTA methods. We have released our code and models at https://github.com/Pluviophile-sy/ST-GGNet. Taiyong Li, Jialei Zhan, Jialang Liu, Xuan Xiong, Weiwei Cai 0001, Exian Liu, Yingmei Wei, Yaowen Hu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | ARBiBench: Benchmarking and Analyzing Adversarial Robustness of Binarized Convolutional Neural NetworksabstractBinarized convolutional neural networks (BCNNs), which restrict the weights and activations of the model to +1 or −1, provide notable reductions in memory requirements and enhanced model inference speed during deployment. Current research on BCNNs primarily revolves around addressing the performance degradation resulting from binarization. However, the investigation of the effects of extreme discretization on the robustness of BCNNs has been largely overlooked, despite its critical relevance to real-world applications. To this end, we propose ARBiBench, a comprehensive benchmark for evaluating the adversarial robustness of BCNNs in the image classification task. The key contributions of ARBiBench include: 1) systematically evaluating the robustness of seven influential BCNN methods across various architectures; 2) rigorous validation of diverse adversarial attack methods; and 3) novel empirical findings showing that BCNNs exhibit weaker robustness than full-precision networks on small datasets but surprisingly stronger robustness on large-scale datasets. Leveraging Information Bottleneck theory, we further demonstrate how data scale and model capacity collectively determine BCNNs’ adversarial robustness. These findings not only challenge conventional assumptions about BCNN security, but also provide new insights for developing robust yet efficient neural network architectures. Li Liu 0002, Bowen Peng, Zhen Liu 0004, Longguang Wang, Yingmei Wei |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Refining Pseudo Labeling via Multi-Granularity Confidence Alignment for Unsupervised Cross Domain Object DetectionabstractMost state-of-the-art object detection methods suffer from poor generalization due to the domain shift between training and testing datasets. To resolve this challenge, unsupervised cross domain object detection is proposed to learn an object detector for an unlabeled target domain by transferring knowledge from an annotated source domain. Promising results have been achieved via Mean Teacher, however, pseudo labeling which is the bottleneck of mutual learning remains to be further explored. In this study, we find that confidence misalignment of the predictions, including category-level overconfidence, instance-level task confidence inconsistency, and image-level confidence misfocusing, leading to the injection of noisy pseudo labels in the training process, will bring suboptimal performance. Considering the above issue, we present a novel general framework termed Multi-Granularity Confidence Alignment Mean Teacher (MGCAMT) for unsupervised cross domain object detection, which alleviates confidence misalignment across category-, instance-, and image-levels simultaneously to refine pseudo labeling for better teacher-student learning. Specifically, to align confidence with accuracy at category level, we propose Classification Confidence Alignment (CCA) to model category uncertainty based on Evidential Deep Learning (EDL) and filter out the category incorrect labels via an uncertainty-aware selection strategy. Furthermore, we design Task Confidence Alignment (TCA) to mitigate the instance-level misalignment between classification and localization by enabling each classification feature to adaptively identify the optimal feature for regression. Finally, we develop imagery Focusing Confidence Alignment (FCA) adopting another way of pseudo label learning, i.e., we use the original outputs from the Mean Teacher network for supervised learning without label assignment to achieve a balanced perception of the image's spatial layout. When these three procedures are integrated into a single framework, they mutually benefit to improve the final performance from a cooperative learning perspective. Extensive experiments across multiple scenarios demonstrate that our method outperforms large foundational models, and surpasses other state-of-the-art approaches by a large margin. Jiangming Chen, Li Liu 0002, Wanxia Deng, Zhen Liu 0004, Yu Liu 0012, Yingmei Wei, Yongxiang Liu |
IEEE Trans. Image Process. | 6 |
| 2024 | Emergence of collective adaptive response based on visual variation
Jingtao Qi, Yingmei Wei, Huaxi Zhang 0002, Yandong Xiao |
Inf. Sci. | 3 |
| 2024 | MCCG: A ConvNeXt-Based Multiple-Classifier Method for Cross-View Geo-LocalizationabstractThe key to crossview geolocalization is to match images of the same target from different viewpoints, e.g., images from drones and satellites. It is a challenging problem due to the changing appearance of objects from variable viewpoints. Most existing methods focus mainly on extracting global features or on segmenting feature maps, causing the loss of information contained in the images. To address the above issues, we propose a new ConvNeXt-based method called MCCG, which stands for Multiple Classifier for Cross-view Geolocalization. The proposed method captures rich discriminative information by cross-dimension interaction and acquires multiple feature representations, realizing a comprehensive feature representation. Additionally, the robustness of the model is improved crediting the multiple feature representations exploiting more contextual information despite position shifting or scale variations. Extensive experiments on the widely used public benchmarks University-1652 and SUES-200 demonstrate that the proposed method achieves state-of-the-art performance in both drone-view target localization and drone navigation applications by over 3% compared to existing methods. Our code and model are available athttps://github.com/mode-str/crossview. Tianrui Shen, Yingmei Wei, Lai Kang, Shanshan Wan, Yee-Hong Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Variational Information Bottleneck for Cross Domain Object DetectionabstractCross domain object detection leverages a labeled source domain to learn an object detector which performs well in a novel unlabeled target domain. Most existing works mainly align the distribution utilizing the entire image knowledge ignoring the obstacles of task-uncorrelated information to alleviate the domain discrepancy. To tackle this issue, we propose a novel module called Variational Instance Disentanglement (VID) based on information theory which aims to decouple the information of task-correlated while filtering out the task-uncorrelated factors at the instance level. Notably, the proposed VID can be used as a plug-and-play module without bringing extra network parameter cost. We equip it with adversarial network and self-training network forming Variational Instance Disentanglement Adversarial Network (VIDAN) and Variational Instance Disentanglement Self-training Network (VIDSN), respectively. Extensive experiments on multiple widely-used scenarios show that the proposed method improves the performance of the popular frameworks and outperforms state-of-the-art methods. Jiangming Chen, Wanxia Deng, Tianpeng Liu, Yingmei Wei, Li Liu 0002 |
ICME | 5 |
| 2023 | Emergence of Adaptation of Collective Behavior Based on Visual PerceptionabstractUnmanned swarms are widespread used in the IoT. The ability of unmanned swarms to achieve adaptive collective behavior in complicated mission scenarios is a prerequisite for meeting mission objectives. However, classical collective behavior models often use the velocity and position of neighbors as inputs to be constructed from a phenomenological perspective. This complicates the construction of unmanned swarms and is incompatible with biological perception. Therefore, this article proposes an observation-orientation-decision-action (OODA) framework for the construction of adaptive collective behavior based on visual perception, inspired by biological collective behavior formations. The model contains no explicit alignment, and no information exchange occurs between individuals. Instead, individuals make decisions based purely on the sight distance corresponding to different relative orientations. Based on adaptability evaluation metrics defined at the collective level, experiments, including coordinated collective motion, single disturbed individual, single external disturbance, narrow passage, and multiple external disturbances scenarios show that the group can respond adaptively to different scenarios with a stable crystal structure while avoiding collisions. In addition to particle simulations, different scenarios provide validation using the Webots robot simulator. As a result, this approach compensates for the inadequacies of existing models and provides technical support for the application of unmanned swarms in various IoT scenarios. Jingtao Qi, Yingmei Wei, Huaxi Zhang 0002, Yandong Xiao |
IEEE Internet Things J. | 3 |
| 2023 | Dual adaptive learning multi-task multi-view for graph network representation learning
Beibei Han, Yingmei Wei, Qingyong Wang, Shanshan Wan |
Neural Networks | 2 |
| 2022 | A Novel Video Copy Sub-sequence Detection and Location MethodabstractWith the rapid growth of the internet and multimedia technology, there is an exponential growth of copy video, which causes some certain impact on video retrieval and copyright protection. Therefore, it becomes increasingly important to find copy videos and locate subsequent clips in a large-scale video database. In this paper, an efficient method is proposed to solve the current problem of video copy detection when the test video contains both copy video clips and non-copy video clips. This paper presents a method of judging the video copy sub-sequence based on the distance between the test video keyframe and the reference video keyframe. First, AlexNet is used to extract the features of the keyframe. Second, it judges whether each test video keyframe is a copy frame according to the distance. Then, the location of the copy sub-sequence is determined by finding the continuous copy frame, and then the video location of the copy sub-sequence is determined. Experimental results show that the proposed method can achieve 86.29% in recall and 95.71% in precision. Yuxiang Xie, Xidao Luan, Yancheng Zhao, Yingmei Wei |
IEEE Big Data | 6 |
| 2022 | The emergence of collective obstacle avoidance based on a visual perception mechanism
Jingtao Qi, Yandong Xiao, Yingmei Wei, Wansen Wu |
Inf. Sci. | 4 |
| 2017 | Two-view underwater 3D reconstruction for cameras with unknown poses under flat refractive interfaces
Lai Kang, Lingda Wu, Yingmei Wei, Songyang Lao, Yee-Hong Yang |
Pattern Recognit. | 3 |
| 2016 | Interactive Visual Analysis on Large Attributed NetworksabstractIncreasing scale leaves a challenging problem for visualizing large attributed networks. Hierarchical aggregation is a promising solution. Existing methods mainly focus on the topological structure but ignore vertex properties; moreover, the inherent hierarchy restricts network navigation process. This paper proposes an user-specified visualization method with a content-based clustering algorithm to explore large attributed networks. The content-based algorithm is able to locate major structures and cluster network based on structural and attribute similarities. Then a novel visualization system is introduced that allows navigation of large networks at any level-of-detail. The user-specified interaction strategy enables user to manipulate cluster metrics and built hierarchy based on interest. Case study demonstrates that the proposed method is effective to extract global knowledge about the network as well as locate critical nodes and major structures. Xiaolei Du, Yingmei Wei, Lingda Wu |
CW | 2 |
| 2015 | Fusing Sorted Random Projections for Robust Texture and Material ClassificationabstractThis paper presents a conceptually simple, and robust, yet highly effective, approach to both texture classification and material categorization. The proposed system is composed of three components: 1) local, highly discriminative, and robust features based on sorted random projections (RPs), built on the universal and information-preserving properties of RPs; 2) an effective bag-of-words global model; and 3) a novel approach for combining multiple features in a support vector machine classifier. The proposed approach encompasses the simplicity, broad applicability, and efficiency of the three methods. We have tested the proposed approach on eight popular texture databases, including Flickr Materials Database, a highly challenging materials database. We compare our method with 13 recent state-of-the-art methods, and the experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for Columbia-Utrecht, 97.16% for Brodatz, 99.30% for University of Maryland Database, and 99.29% for Kungliga Tekniska högskolan-textures under varying illumination, pose, and scale. Moreover, the proposed approach significantly outperforms the current state-of-the-art approach in materials categorization, with an improvement to classification accuracy of 67%. Li Liu 0002, Paul W. Fieguth, Dewen Hu, Yingmei Wei, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | A highly accurate dense approach for homography estimation using modified differential evolution
Lai Kang, Lingda Wu, Yingmei Wei, Hanchen Song |
Eng. Appl. Artif. Intell. | 3 |
| 2013 | BRINT: A binary rotation invariant and noise tolerant texture descriptorabstractLocal Binary Pattern (LBP) and its variants are effective and popular descriptors for texture classification. Most LBP like descriptors have disadvantages including sensitiveness to noise and inability to capture long distance texture information. In this paper we propose a simple, efficient, yet robust multi-resolution descriptor to texture classification - Binary Rotation Invariant and Noise Tolerant (BRINT). The proposed descriptor is very fast to build, very compact while remaining robust to illumination variations, rotation changes and noise. We develop a novel and simple strategy - averaging before binarization - to compute a local binary descriptor based on the conventional LBP approach. Points are sampled in a circular neighborhood, but keeping the number of bins in a single-scale LBP histogram constant and small by averaging over several contiguous pixels in the circle. There is no need for pre-training, no texton dictionary, and no tuning of parameters to deal with different datasets. Experiments on the Outex test suite demonstrate that the proposed approach is very robust to noise and significantly outperforms the state-of-the-art in terms of classifying noise corrupted textures. Li Liu 0002, Paul W. Fieguth, Yingmei Wei |
ICIP | 5 |
| 2013 | RBRIEF: a robust descriptor based on random binary comparisonsabstractThe authors propose a robust descriptor based on BRIEF, called RBRIEF. Unlike the original BRIEF, the proposed descriptor is also robust to scale and in‐plane rotation transformations. Furthermore, the authors use first derivative as sample function to do binary comparisons which has proven to be better compared against the function of intensity used in BRIEF. In the feature matching stage, the authors use Hamming distance to evaluate the descriptor similarity. As a result, the performance of the proposed descriptor outperforms SURF, BRIEF and ORB using standard benchmarks. In particular, the experiments demonstrate the proposed descriptor's superior performance in the presence of image blur, JPEG compression and light changes. Furthermore, the descriptor exhibits robust performance using only relatively few bits compared to other descriptors. Lingda Wu, Hanchen Song, Yingmei Wei |
IET Comput. Vis. | 4 |
| 2013 | Applications of structure from motion: a surveyabstractStructure from motion (SfM) has been an active research area in computer vision for decades and numerous practical applications are benefiting from this research. While no previous work has tried to summarize the applications appearing in the literature, this paper deals with a comprehensive overview of recent applications of SfM by classifying them into 10 categories, namely augmented reality, autonomous navigation/guidance, motion capture, hand-eye calibration, image/video processing, image-based 3D modeling, remote sensing, image organization/browsing, segmentation and recognition, and military applications. The goal is to provide insights for researchers to position their work more appropriately in the context of existing techniques, and to perceive both new applications and relevant research problems. Yingmei Wei, Lai Kang, Lingda Wu |
J. Zhejiang Univ. Sci. C | 1 |