Jianzhe Gao

dblp:363/7431 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
abstract
Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental reasoning and local scene comprehension, existing UAV agents typically adopt mono-granularity frameworks that struggle to balance these two aspects. To address this limitation, this work proposes a History-Enhanced Two-Stage Transformer (HETT) framework, which integrates the two aspects through a coarse-to-fine navigation pipeline. Specifically, HETT first predicts coarse-grained target positions by fusing spatial landmarks and historical context, then refines actions via fine-grained visual analysis. In addition, a historical grid map is designed to dynamically aggregate visual features into a structured spatial memory, enhancing comprehensive scene awareness. Additionally, the CityNav dataset annotations are manually refined to enhance data quality. Experiments on the refined CityNav dataset show that HETT delivers significant performance gains, while extensive ablation studies further verify the effectiveness of each component.
Xichen Ding, Jianzhe Gao, Wenguan Wang
AAAI2
2026 FADMB: Fully attention-based dual memory bank network for weakly supervised video anomaly detection
Zhiming Luo, Shuheng Huang, Jianzhe Gao, Shaozi Li
Pattern Recognit.4
2025 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
Jianzhe Gao, Wenguan Wang
ICCV1
2024 Mask Matching Network for Self-supervised Few-shot Medical Image Segmentation
abstract
Existing few-shot segmentation methods have achieved remarkable progress in medical image segmentation. However, many existing methods yield incomplete and discontinuous boundary predictions. In contrast, the Segment Anything Model (SAM) consistently produces clear, continuous, and comprehensive segmentation boundaries. Building on this observation, we propose a new two-step network called Mask Matching Network (MMNet) to introduce extra knowledge learned by SAM in natural images for few-shot medical image segmentation. Firstly, Q-Net has been utilized to locate some Regions of Interest (RoI) as prompts for SAM, allowing for the automatic generation of masks without relying on manual prompts. Secondly, we propose a novel Mask Matching Module (MMM), which considers both feature similarity and volume similarity as guidance to collaboratively mine the final segmentation from proposal masks. MMNet achieves state-of-the-art performance with remarkable improvements on two widely used datasets, abdominal MR (ABD) and cardiac MR (CMR), under two different settings.
Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li
ICME3
2024 FDAPNet: Feature Denoising and Aggregation Prediction Network for Polyp Segmentation
abstract
Automatic polyp segmentation can assist doctors in screening colonoscopy images, which is crucial for preventing colorectal cancer. Although existing methods have achieved promising results, they are susceptible to interference from surrounding tissues and lack adaptability to different sizes of polyps. To address these issues, we propose a feature denoising and aggregation prediction network (FDAPNet) for polyp segmentation. Specifically, we first design a feature denoising module to eliminate the background noise in shallow features, allowing the network to focus more on the target area. Then, a channel attention module is proposed to automatically emphasize relevant features while suppressing irrelevant ones, which can enhance the network’s ability to capture important information. Moreover, we propose a multi-scale feature aggregation prediction module to better deal with polyps of various sizes. We conduct extensive experiments on five public datasets. The quantitative and qualitative results demonstrate that our FDAPNet outperforms other state-of-the-art polyp segmentation methods.
Jianzhe Gao, Lijun Bao
IJCNN2
2024 A Collaborative Framework Using Multimodal Data and Adaptive Noise for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection in video aims to identify unusual behaviors that are crucial for public safety. Recently, there has been an increase in reconstruction or prediction-based methods that integrate diverse modal features to enhance anomaly detection. However, they use methods that independently or directly fusion multimodal features without fully considering the collaborative potential between multimodal features, which are susceptible to interference from semantic differences, thereby impacting detection performance. In contrast, we design a collaborative framework using multimodal data and adaptive noise for behavior anomaly detection. Our framework detects anomalies by analyzing the contrastive differences between two modalities alongside single-frame reconstruction errors. Specifically, we first learn the correlation between RGB and skeletal modalities for normal behavior through contrastive learning and use inter-modal contrast difference to detect motion anomalies. Additionally, we propose a single-frame reconstruction network that adaptively adds noise based on the importance of foreground features to detect appearance anomalies. Anomalies often occur in the motion foreground, and increasing noise in this area can make it more difficult to reconstruct anomalies. Extensive experiments validate the state-of-the-art performance of our method on three public datasets.
Jianzhe Gao, Kejia Zhang 0003, Yifan He 0002, Zhiming Luo, Shaozi Li
IJCNN2
2024 QueryNet: A Unified Framework for Accurate Polyp Segmentation and Detection
Jiaxing Chai, Zhiming Luo, Jianzhe Gao, Licun Dai, Yingxin Lai, Shaozi Li
MICCAI (8)3
2024 A Multilevel Guidance-Exploration Network and Behavior-Scene Matching Method for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection aims to identify unusual human actions, playing a crucial role in intelligent surveillance and other areas. The current mainstream methods still adopt reconstruction or future frame prediction techniques. However, reconstructing or predicting low-level pixel features easily enables the network to achieve overly strong generalization ability, allowing anomalies to be reconstructed or predicted as effectively as normal data. Different from their methods, inspired by the Student-Teacher Network, we propose a novel framework called the Multilevel Guidance-Exploration Network (MGENet), which detects anomalies through the difference in high-level representation between the Guidance and Exploration network. Specifically, we first utilize the Normalizing Flow that takes skeletal keypoints as input to guide an RGB encoder, which takes unmasked RGB frames as input, to explore latent motion features. Then, the RGB encoder guides the mask encoder, which takes masked RGB frames as input, to explore the latent appearance feature. Additionally, we design a Behavior-Scene Matching Module to detect scene-related behavioral anomalies. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on ShanghaiTech and UBnormal datasets, with AUC of 86.9% and 74.3%, respectively. The code is available at https://github.com/molu-ggg/GENet.
Zhiming Luo, Jianzhe Gao, Yingxin Lai, Yifan He 0002, Shaozi Li
ACM Multimedia3
2024 CPNet: Cross Prototype Network for Few-Shot Medical Image Segmentation
Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li
PRCV (15)2
2023 TPNet: Enhancing Weakly Supervised Polyp Frame Detection with Temporal Encoder and Prototype-Based Memory Bank
Jianzhe Gao, Zhiming Luo, Shaozi Li
PRCV (12)1