VLDB 2026 Research / reviewers in the wild / expert
Kaige Li
dblp:242/6142
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-1716-4381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian UpdatingabstractImage geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainable abilities, they often struggle with confirmation bias, i.e., committing to early, potentially incorrect guesses driven by visual clues with varied geographic likelihoods. In this paper, we propose GeoBayes, a novel training-free framework that formulates geolocalization as a Maximum a Posteriori (MAP) estimation task over multiple geographic hypotheses and performs probabilistic thought via sequential Bayesian reasoning. GeoBayes treats each visual object and its associated geographic clues as probabilistic evidence, integrating them iteratively through a Hypothesize–Verify–Update loop. At each step, it evaluates how new evidence supports existing hypotheses and updates their posterior probabilities, gradually converging on the most probable location. This allows GeoBayes to explicitly quantify and fuse the varied geographic probabilities implied by various visual elements, reducing the risk of overcommitting to misleading clues. Furthermore, considering the natural hierarchy of geographic labels (e.g., country, city), GeoBayes introduces a state memory mechanism that stores hypotheses, inference context, and evidence scores across levels. This design enables the framework to propagate prior knowledge across levels of the geographic hierarchy and incorporate geographic structural constraints into the Bayesian update process, achieving a coarse-to-fine geo-localization. Experiments on IM2GPS3k and YFCC4K show that GeoBayes improves MLLM-based geo-localization accuracy without extra training. This demonstrates the effectiveness of probabilistic reasoning for robust and interpretable geo-localization. Kaige Li, Junhao Fang, Qichuan Geng, Zhong Zhou |
AAAI | 3 |
| 2026 | Out-of-Distribution Semantic Segmentation With Disentangled and Calibrated RepresentationabstractOut-of-distribution (OoD) semantic segmentation aims to recognize pixels of classes undefined in the training dataset. Existing methods mostly focus on training the model to fit real OoD data samples to identify OoD pixels, which requires extra data collection and annotation efforts. By contrast, synthesizing OoD data with training data provides a more resource-efficient alternative. However, synthetic data generated from controlled settings lacks diversity, causing the model to suffer from overfitting. To this end, we propose a disentangled representation learning (DRL) method to guide the model to disentangle semantic-related and semantic-unrelated features from synthetic OoD data. DRL encourages the model to utilize the former to identify semantic categories, rather than overfitting to such semantic-unrelated features as synthetic artificiality. Specifically, DRL first incorporates two disentanglers to extract the semantic-related and -unrelated features and then applies a shuffle and reconstruction mechanism to regularize the disentangled features. Furthermore, to facilitate disentangling, we propose a pixel-wise feature similarity calibration (PSC) module, which utilizes more accurate ID-OoD similarity to calibrate inaccurate ID-OoD similarity learned exclusively from ID data. Thus, PSC delivers accurate and stable pixel-wise features for effective disentangling. Extensive experiments illustrate that the proposed method exhibits strong generalization ability. It attains 74.04% AuPRC and 20.82% FPR on Road Anomaly, 69.85% AuPRC and 5.78% FPR on Fishyscapes LostAndFound Validation Set, using SegFormer with the MiT-B5 backbone. Source code is available at https://github.com/WanMotion/DisentangledOoDSeg. Maoxian Wan, Kaige Li, Qichuan Geng, Binyi Su, Zhong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Quantification of Large Language Model DistillationabstractSunbowen Lee, Junting Zhou, Chang Ao, Kaige Li, Xeron Du, Sirui He, Haihong Wu, Tianci Liu, Jiaheng Liu, Hamid Alinejad-Rokny, Min Yang, Yitao Liang, Zhoufutu Wen, Shiwen Ni. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sunbowen Lee, Junting Zhou, Chang Ao, Kaige Li, Xeron Du, Sirui He 0001, Haihong Wu, Tianci Liu 0011, Hamid Alinejad-Rokny, Min Yang 0007, Yitao Liang, Zhoufutu Wen, Shiwen Ni |
ACL (1) | 4 |
| 2025 | Incremental Few-Shot Semantic Segmentation Via Multi-Level Switchable Visual Prompts
Maoxian Wan, Kaige Li, Qichuan Geng, Zhong Zhou |
ICCV | 2 |
| 2025 | Boosting Human Pose Estimation via Heatmap Refinement
Ling Jiang 0001, Zhuocheng Liu, Kaige Li, Wei Wu 0008 |
MMM (1) | 3 |
| 2025 | MAP: Masked Adversarial Perturbation for Boosting Black-Box Attack TransferabilityabstractThe transferability of adversarial examples is vital for black-box attacks, as it enables the adversary to deceive the target model without knowing its internals. Despite numerous methods focusing on transferability, they still struggle with transferring across models with distinct architectural components (e.g., CNNs and ViTs). In this work, we argue that the limited adversarial perturbation diversity leads to overfitting of the surrogate model, which acts as a key factor in reducing transferability. To this end, we propose a Masked Adversarial Perturbation (MAP) method to boost adversarial transferability across various architectures from a novel perspective of diversifying perturbation. Specifically, MAP randomly masks perturbation patches during iterations and compels the remaining ones to retain the attack effect, which diversifies perturbations to mitigate their overfitting to the surrogate model. Naturally, MAP spreads perturbation over local patches to alleviate their co-adaptation and prevent perturbations from overly relying on specific patterns. Consequently, it can deceive convolution operation and self-attention mechanism indiscriminately by attacking their basic input units, i.e., a single patch, showing superior transferability over previous methods. Extensive experiments illustrate that MAP consistently and significantly boosts diverse black-box attacks to achieve state-of-the-art performance. Kaige Li, Maoxian Wan, Qichuan Geng, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 1 |
| 2025 | LangLoc: Language-Driven Localization via Formatted Spatial Description GenerationabstractExisting localization methods commonly employ vision to perceive scene and achieve localization in GNSS-denied areas, yet they often struggle in environments with complex lighting conditions, dynamic objects or privacy-preserving areas. Humans possess the ability to describe various scenes using natural language, effectively inferring their location by leveraging the rich semantic information in these descriptions. Harnessing language presents a potential solution for robust localization. Thus, this study introduces a new task, Language-driven Localization, and proposes a novel localization framework, LangLoc, which determines the user's position and orientation through textual descriptions. Given the diversity of natural language descriptions, we first design a Spatial Description Generator (SDG), foundational to LangLoc, which extracts and combines the position and attribute information of objects within a scene to generate uniformly formatted textual descriptions. SDG eliminates the ambiguity of language, detailing the spatial layout and object relations of the scene, providing a reliable basis for localization. With generated descriptions, LangLoc effortlessly achieves language-only localization using text encoder and pose regressor. Furthermore, LangLoc can add one image to text input, achieving mutual optimization and feature adaptive fusion across modalities through two modality-specific encoders, cross-modal fusion, and multimodal joint learning strategies. This enhances the framework's capability to handle complex scenes, achieving more accurate localization. Extensive experiments on the Oxford RobotCar, 4-Seasons, and Virtual Gallery datasets demonstrate LangLoc's effectiveness in both language-only and visual-language localization across various outdoor and indoor scenarios. Notably, LangLoc achieves noticeable performance gains when using both text and image inputs in challenging conditions such as overexposure, low lighting, and occlusions, showcasing its superior robustness. Changhao Chen, Kaige Li, Yuan Xiong, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 3 |
| 2024 | Exploring Scale-Aware Features for Real-Time Semantic Segmentation of Street ScenesabstractReal-time semantic segmentation of street scenes is an essential and challenging task for autonomous driving systems, which needs to achieve both high accuracy and efficiency. Moreover, numerous objects and stuff at different scales in street scenes further increase the difficulty of this task. To address this challenge, we develop a lightweight and high-accuracy network termed Scale-Aware Network (SANet), which aims to selectively aggregate multi-scale features while maintaining high efficiency. In SANet, we first design a Selective Context Encoding (SCE) module, which considers the intrinsic differences of various pixels to selectively encode private contexts for each pixel, thus learning more desirable contextual features while reducing redundancy. With the context embedding in hand, we then design a Selective Feature Fusion (SFF) module to recursively fuses them with multiple features at different levels or scales to generate scale-aware features, where each feature map contains scale-specific information. Extensive experiments on challenging street scene datasets, i.e., Cityscapes and CamVid, illustrate that our SANet achieves a leading trade-off between segmentation accuracy and speed. Concretely, our method yields$78.1\%$mIoU at$109.0$FPS on the Cityscapes test set and$77.2\%$mIoU at$250.4$FPS on the CamVid test set. Code will be available at https://github.com/kaigelee/SANet. Kaige Li, Qichuan Geng, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Multimodality Adaptive Transformer and Mutual Learning for Unsupervised Domain Adaptation Vehicle Re-IdentificationabstractUnsupervised Domain Adaptation Vehicle Re-Identification (UDA vehicle re-ID) aims to enable the model trained in the source domain dataset to adapt to the target domain data and obtain accurate re-identification results, which has received widespread attention due to its practicality in the field of intelligent transportation systems. Most current UDA vehicle re-ID research ignores the mining and utilization of attribute information. Meanwhile, the Convolutional Neural Networks-based (CNN-based) network will cause the loss of fine-grained information, reducing the expression and generalization ability of vehicle features. To alleviate such issues, we are motivated by the Transformer, which can exploit distinguishable attribute information and fuse multimodal features effectively. Therefore, this paper proposes a Multimodality Adaptive Transformer Network (MATNet) to intensify the ability to learn vehicle fine-grained features related to attributes. Moreover, the noise contained in pseudo-labels assigned by cluster algorithms interferes with the performance of the UDA vehicle re-ID method. We also design the Dual Mutual Dynamic Update Pseudo-Label generation strategy (DMDU) to improve the accuracy of pseudo-labels and alleviate error accumulation. The strategy is based on mutual learning, which can effectively utilize the congruous and particular knowledge of the two models to generate pseudo-labels. Extensive experiments on two large-scale public datasets, including VeRi-776 and VehicleID, illustrate that our method outperforms the state-of-the-art methods. Xin Zhang 0116, Yunan Ling, Kaige Li, Zhong Zhou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Context and Spatial Feature Calibration for Real-Time Semantic SegmentationabstractContext modeling or multi-level feature fusion methods have been proved to be effective in improving semantic segmentation performance. However, they are not specialized to deal with the problems of pixel-context mismatch and spatial feature misalignment, and the high computational complexity hinders their widespread application in real-time scenarios. In this work, we propose a lightweight Context and Spatial Feature Calibration Network (CSFCN) to address the above issues with pooling-based and sampling-based attention mechanisms. CSFCN contains two core modules: Context Feature Calibration (CFC) module and Spatial Feature Calibration (SFC) module. CFC adopts a cascaded pyramid pooling module to efficiently capture nested contexts, and then aggregates private contexts for each pixel based on pixel-context similarity to realize context feature calibration. SFC splits features into multiple groups of sub-features along the channel dimension and propagates sub-features therein by the learnable sampling to achieve spatial feature calibration. Extensive experiments on the Cityscapes and CamVid datasets illustrate that our method achieves a state-of-the-art trade-off between speed and accuracy. Concretely, our method achieves 78.7% mIoU with 70.0 FPS and 77.8% mIoU with 179.2 FPS on the Cityscapes and CamVid test sets, respectively. The code is available at https://nave.vr3i.com/ and https://github.com/kaigelee/CSFCN. Kaige Li, Qichuan Geng, Maoxian Wan, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 1 |