Zheng Ge

dblp:231/1007 · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 4 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 14 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
abstract
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Xiangyu Zhang, Heung-Yeung Shum. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Zhewei Huang, Hebin Zhou, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Zheng Ge, Xiangyu Zhang 0005, Harry Shum
ACL (1)17
2026 PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
abstract
Xiangfeng Wang, Hangyu Guo, Yanlin Lai, Mitt Huang, Liang Zhao, Chengyuan Yao, Yinmin Zhang, Qi Han, Xiaoxiaoren, Chun Yuan, Tong Xu, Zheng Ge, Xiangyu Zhang, Daxin Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiangfeng Wang 0005, Hangyu Guo, Yanlin Lai, Mitt Huang, Chengyuan Yao, Yinmin Zhang, Xiaoxiao Ren, Chun Yuan 0003, Tong Xu 0001, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang
ACL (1)12
2026 Active perception: Gaze-guided thinking for chart understanding
abstract
Answering questions about charts presents a unique challenge for Vision-Language Models (VLMs). Unlike natural images, charts are structured artifacts governed by explicit visual grammar that demands pixel-level accuracy in visual perception. While recent VLMs demonstrate impressive reasoning abilities on chart tasks, a critical gap remains: their reasoning operates abstractly, disconnected from precise visual grounding. We introduce Active Perception, a framework that enables Gaze-Guided Thinking, a reasoning pattern that explicitly anchors abstract inference to concrete visual locations through coordinate-based operations (Locate, Trace, Extract, Compare). To instill this capability, we propose Skill Cultivation, a two-stage training strategy: Stage I injects coordinate-aware primitives via Supervised Fine-Tuning on ChartQAGaze-14K, our synthesized dataset of 14K coordinate-annotated reasoning chains; Stage II internalizes these skills into adaptive strategies via Reinforcement Learning with outcome-based rewards. Building upon Qwen2.5-VL-7B, Active Perception achieves state-of-the-art performance on ChartQA, improving overall accuracy from 78.96% to 82.44%, with particularly notable gains on the challenging Human split (75.76% to 81.28%). Qualitative analysis reveals emergent systematic chart-reading behaviors that mirror human visual strategies, demonstrating the effectiveness of spatially grounded reasoning for structured visual understanding. • We identify a critical gap in current VLMs for chart understanding: reasoning operates abstractly without precise visual grounding, limiting accurate data extraction from structured visualizations. • We propose Gaze-Guided Thinking , a reasoning pattern that anchors abstract inference to concrete visual locations through Coordinate Primitives ( Locate , Trace , Extract , Compare ), mimicking human chart scanning behavior. • We introduce Skill Cultivation , a two-stage training strategy combining SFT on ChartQAGaze-14K (14K coordinate-annotated reasoning chains) with outcome-based RL to inject and internalize spatially-grounded reasoning. • Active Perception achieves 82.44% overall accuracy on ChartQA (a 3.48% absolute improvement), with particularly strong performance on the Human split (81.28%, +5.52%), demonstrating that explicit visual grounding substantially enhances structured visual understanding. • Qualitative analysis reveals emergent human-like chart reading behaviors, where models systematically leverage coordinates for precise value extraction and spatial reasoning.
Xin Huang 0027, Hongbing Li, Zejia Weng, Jia Wang 0025, Yeqing Shen, Haolong Yan, Kaijun Tan, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Osamu Yoshie
Neurocomputing11
2026 Reliability in Statistically Dependent Networks: Bounds, Linear Programming, and Scalability
abstract
Network reliability and robustness are important to understand and design resilient networks. While in the state-of-the-art models, independent failures at the links are assumed, we present an analytical method for determining reliability bounds with arbitrary dependency structures. In addition to the analytical bounds, an optimization approach based on linear programming (LP) is introduced. Since the direct application of LP to large networks encounters computational limitations, a graph decomposition method is developed to decompose large networks into smaller, manageable subgraphs. These subgraphs can then be analyzed using either the analytical bounds or LP. To better understand the impact of links on the reliability, we investigate specific conditions under which adding so-called bridge connections can improve reliability. The numerical results show that both the LP method and the analytical approach provide precise upper and lower reliability bounds after decomposition. For the bridge network, the calculated best-case bound is about 50% higher than the i.i.d. case, while the worst-case bound is about 24% lower. The proposed framework is also suitable for analyzing large-scale networks with complex dependencies. To this end, an algorithm is developed that automatically decomposes large graphs and incorporates the proposed analytical and LP methods to compute the overall probability bounds for arbitrary dependencies.
Zheng Ge, Eduard A. Jorswieck
IEEE Trans. Commun.1
2025 Taming Teacher Forcing for Masked Autoregressive Video Generation
abstract
We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation.
Yuang Peng, Kun Yan 0004, Runpei Dong, Duomin Wang, Zheng Ge, Nan Duan 0001, Xiangyu Zhang 0005
CVPR7
2025 Dependency and Link Diversity Placement for Reliable Wireless Diamond Networks
abstract
Dependency of link failures in wireless networks impact the end-to-end reliability. Placing independent links can improve the guaranteed reliability significantly. However, it is costly and resource in-efficient. Therefore, we study an idealized wireless diamond network with dependent link failures and compute worst-case reliability for different link diversity placement strategies. A complete characterization of the reliability as a function of the marginal probabilities is provided. Numerical results show the interesting switching behaviour of optimal placement and achieved reliability as a function of the marginal probabilities.
Zheng Ge, Eduard A. Jorswieck
ICC1
2025 DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
abstract
Personalized image generation holds great promise in assisting humans in everyday work and life due to its impressive function in creatively generating personalized content. However, current evaluations either are automated but misalign with humans or require human evaluations that are time-consuming and expensive. In this work, we present DreamBench++, a human-aligned benchmark that advanced multimodal GPT models automate. Specifically, we systematically design the prompts to let GPT be both human-aligned and self-aligned, empowered with task reinforcement. Further, we construct a comprehensive dataset comprising diverse images and prompts. By benchmarking 7 modern generative models, we demonstrate that \dreambench results in significantly more human-aligned evaluation, helping boost the community with innovative findings.
Yuang Peng, Haomiao Tang, Zekun Qi, Runpei Dong, Chunrui Han, Zheng Ge, Xiangyu Zhang 0005, Shutao Xia
ICLR8
2025 Reconstructive Visual Instruction Tuning
abstract
This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that exclusively supervise text outputs, ROSS prompts LMMs to supervise visual outputs via reconstructing input images. By doing so, it capitalizes on the inherent richness and detail present within input images themselves, which are often lost in pure text supervision. However, producing meaningful feedback from natural images is challenging due to the heavy spatial redundancy of visual signals. To address this issue, ROSS employs a denoising objective to reconstruct latent representations of input images, avoiding directly regressing exact raw RGB values. This intrinsic activation design inherently encourages LMMs to maintain image detail, thereby enhancing their fine-grained comprehension capabilities and reducing hallucinations. Empirically, ROSS consistently brings significant improvements across different visual encoders and language models. In comparison with extrinsic assistance state-of-the-art alternatives that aggregate multiple visual experts, ROSS delivers competitive performance with a single SigLIP visual encoder, demonstrating the efficacy of our vision-centric supervision tailored for visual outputs. The code will be made publicly available upon acceptance.
Anlin Zheng, Tiancai Wang, Zheng Ge, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001
ICLR5
2025 Unhackable Temporal Reward for Scalable Video MLLMs
abstract
In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the “anti-scaling law”, where more data and larger models lead to worse performance. This study unmasks the culprit: “temporal hacking”, a phenomenon where models shortcut by fixating on select frames, missing the full video narrative. In this work, we systematically establish a comprehensive theory of temporal hacking, defining it from a reinforcement learning perspective, introducing the Temporal Perplexity (TPL) score to assess this misalignment, and proposing the Unhackable Temporal Rewarding (UTR) framework to mitigate the temporal hacking. Both theoretically and empirically, TPL proves to be a reliable indicator of temporal modeling quality, correlating strongly with frame activation patterns. Extensive experiments reveal that UTR not only counters temporal hacking but significantly elevates video comprehension capabilities. This work not only advances video-AI systems but also illuminates the critical importance of aligning proxy rewards with true objectives in MLLM development.
En Yu, Kangheng Lin, Yana Wei, Zining Zhu 0004, Jianjian Sun, Zheng Ge, Xiangyu Zhang 0005, Jingyu Wang 0001, Wenbing Tao
ICLR8
2025 Perception in Reflection
abstract
We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve perfect perception initially. Specifically, we propose Reflective Perception (RePer), a dual-model reflection mechanism that systematically alternates between policy and critic models, enables iterative refinement of visual perception. This framework is powered by Reflective Perceptual Learning (RPL), which reinforces intrinsic reflective capabilities through a methodically constructed visual reflection dataset and reflective unlikelihood training Comprehensive experimental evaluation demonstrates RePer's quantifiable improvements in image understanding, captioning precision, and hallucination reduction. Notably, RePer achieves strong alignment between model attention patterns and human visual focus, while RPL optimizes fine-grained and free-form preference alignment. These advancements establish perception in reflection as a robust paradigm for future multimodal agents, particularly in tasks requiring complex reasoning and multi-step manipulation. Project Page: [https://weiyana.github.io/Perception-in-Reflection](https://weiyana.github.io/Perception-in-Reflection)
Yana Wei, Kangheng Lin, En Yu, Yuang Peng, Runpei Dong, Jianjian Sun, Zheng Ge, Xiangyu Zhang 0005, Vishal M. Patel
ICML9
2025 Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
abstract
The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning. We introduce a two-stage paradigm built on Qwen2.5-VL-7B: a massive linguistic cold-start fine-tuning, followed by multimodal reinforcement learning (RL) spanning nearly 1,000 steps—surpassing all previous open-source efforts in scale. This pioneering work reveals three fundamental insights: 1) Behavior transfer emerges surprisingly early in cold start due to linguistic mental imagery. 2) Cold start broadly memorizes visual behaviors, while RL critically discerns and scales up effective patterns. 3) Transfer strategically favors high-utility behaviors such as visual reflection. Our resulting model, Open-Vision-Reasoner (OVR), achieves state-of-the-art performance on a suite of reasoning benchmarks, including 95.3% on MATH500, 51.8% on MathVision and 54.6% on MathVerse. We release our model, data, and training dynamics to catalyze the development of more capable, behavior-aligned multimodal reasoners.
Yana Wei, Jianjian Sun, Kangheng Lin, Jisheng Yin, Jingcheng Hu, Yinmin Zhang, En Yu, Zejia Weng, Jia Wang 0025, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Vishal M. Patel
NeurIPS13
2025 GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
abstract
With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary, making it difficult to obtain the comprehensive environment information needed for agent training and evaluation. This limitation hinders systematic investigation and benchmarking of agent navigation capabilities. To address this limitation, we introduce GUI Exploration Lab, a simulation environment engine for GUI agent navigation research that enables flexible definition and composition of screens, icons, and navigation graphs, while providing full access to environment information for comprehensive agent training and evaluation. Through extensive experiments, we find that supervised fine-tuning enables effective memorization of fundamental knowledge, serving as a crucial foundation for subsequent training. Building on this, single-turn reinforcement learning further enhances generalization to unseen scenarios. Finally, multi-turn reinforcement learning encourages the development of exploration strategies through interactive trial and error, leading to further improvements in screen navigation performance. We validate our methods on both static and interactive benchmarks, demonstrating that our findings generalize effectively to real-world scenarios. These findings demonstrate the advantages of reinforcement learning approaches in GUI navigation and offer practical guidance for building more capable and generalizable GUI agents.
Haolong Yan, Yeqing Shen, Xin Huang 0027, Jia Wang 0025, Kaijun Tan, Zhixuan Liang, Zheng Ge, Osamu Yoshie, Xiangyu Zhang 0005, Daxin Jiang
NeurIPS8
2025 Perception-R1: Pioneering Perception Policy with Reinforcement Learning
abstract
Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual perplexity is a major factor in determining the effectiveness of RL. We also observe that reward design plays a crucial role in further approaching the upper limit of model perception. To leverage these findings, we propose Perception-R1, a scalable RL framework using GRPO during MLLM post-training. With a standard Qwen2-VL-2B-Instruct, Perception-R1 achieves +4.2% on RefCOCO+, +17.9% on PixMo-Count, +4.2% on PageOCR, and notably, 31.9% AP on COCO2017 val for the first time, establishing a strong baseline for perception policy learning.
En Yu, Kangheng Lin, Jisheng Yin, Yana Wei, Yuang Peng, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Jingyu Wang 0001, Wenbing Tao
NeurIPS10
2025 DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
abstract
Multimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple modalities introduces heterogeneity in both the model and training data, creating unique systems challenges.
Yinmin Zhong, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu 0001, Daxin Jiang, Xin Jin 0008
SIGCOMM6
2025 Reliability of Load Balancing and Packet Duplication in Dependent Diamond Networks
abstract
In recent years, Multi-Connectivity (MC) has emerged as a promising approach to enhancing the reliability and resilience of communication networks. The effectiveness of MC depends on the statistical dependencies between the links and heavily on the selection of an appropriate MC strategy. We introduce Packet Duplication (PD) and Load Balancing (LB) in this paper. Each of these strategies offers distinct advantages, but a comprehensive comparison under varying network conditions and dependencies is missing. To address this gap, we investigate an idealized diamond network element, applying two strategies. Analytical results are provided, highlighting the mathematical relationship between effective rate and reliability, considering worst-, best-, and independent and identically distributed (i.i.d.)-dependencies between links. Numerical simulations demonstrate the dichotomy of the optimal strategy under different network scenarios. Interestingly, LB outperforms PD for the worst-case. Additionally, we extend our analysis by deriving reliability bounds as functions of the effective rate for arbitrary network structures in both PD and LB.
Zheng Ge, Shashank Jhansale, Lara Jüschke, Lars C. Wolf, Eduard A. Jorswieck
VTC2025-Fall1
2024 Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
Zhi Cai, Guodong Wang 0006, Zheng Ge, Xiangyu Zhang 0005, Di Huang 0001
BMVC5
2024 ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Chunrui Han, Zheng Ge, Li Yi 0001, Kaisheng Ma
ECCV (43)6
2024 Vary: Scaling up the Vision Vocabulary for Large Vision-Language Model
Lingyu Kong, Jinyue Chen, Zheng Ge, Jianjian Sun, Chunrui Han, Xiangyu Zhang 0005
ECCV (4)5
2024 Merlin: Empowering Multimodal LLMs with Foresight Minds
En Yu, Yana Wei, Dongming Wu 0005, Lingyu Kong, Tiancai Wang, Zheng Ge, Xiangyu Zhang 0005, Wenbing Tao
ECCV (4)9
2024 DreamLLM: Synergistic Multimodal Comprehension and Creation
abstract
This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative modeling of both language and image posteriors by direct sampling in the raw multimodal space. This approach circumvents the limitations and information loss inherent to external feature extractors like CLIP, and a more thorough multimodal understanding is obtained. Second, DreamLLM fosters the generation of raw, interleaved documents, modeling both text and image contents, along with unstructured layouts. This allows DreamLLM to learn all conditional, marginal, and joint multimodal distributions effectively. As a result, DreamLLM is the first MLLM capable of generating free-form interleaved content. Comprehensive experiments highlight DreamLLM's superior performance as a zero-shot multimodal generalist, reaping from the enhanced learning synergy. Project page: https://dreamllm.github.io.
Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jianjian Sun, Xiangwen Kong, Xiangyu Zhang 0005, Kaisheng Ma, Li Yi 0001
ICLR5
2024 ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
En Yu, Zheng Ge, Jianjian Sun, Yuang Peng, Runpei Dong, Chunrui Han, Xiangyu Zhang 0005
IJCAI3
2024 OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
Jinyue Chen, Lingyu Kong, Zheng Ge, Jianjian Sun, Chunrui Han, Xiangyu Zhang 0005
ACM Multimedia5
2024 Self-Supervised Visual Preference Alignment
abstract
This paper makes the first attempt towards unsupervised preference alignment in Vision-Language Models (VLMs). We generate chosen and rejected responses with regard to the original and augmented image pairs, and conduct preference alignment with direct preference optimization. It is based on a core idea: properly designed augmentation to the image input will induce VLM to generate false but hard negative responses, which helps the model to learn from and produce more robust and powerful answers. The whole pipeline no longer hinges on supervision from GPT-4 or human involvement during alignment, and is highly efficient with few lines of code. With only 8k randomly sampled unsupervised data, it achieves 90% relative score to GPT-4 on complex reasoning in LLaVA-Bench, and improves LLaVA-7B/13B by 6.7%/5.6% score on complex multi-modal benchmark MM-Vet. Visualizations shows its improved ability to align with user-intentions. A series of ablations are firmly conducted to reveal the latent mechanism of the approach, which also indicates its potential towards further scaling.
Zheng Ge, Xiangyu Zhang 0005
ACM Multimedia3
2023 BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal Stereo
abstract
Restricted by the ability of depth perception, all Multi-view 3D object detection methods fall into the bottleneck of depth accuracy. By constructing temporal stereo, depth estimation is quite reliable in indoor scenarios. However, there are two difficulties in directly integrating temporal stereo into outdoor multi-view 3D object detectors: 1) The construction of temporal stereos for all views results in high computing costs. 2) Unable to adapt to challenging outdoor scenarios. In this study, we propose an effective method for creating temporal stereo by dynamically determining the center and range of the temporal stereo. The most confident center is found using the EM algorithm. Numerous experiments on nuScenes have shown the BEVStereo's ability to deal with complex outdoor scenarios that other stereo-based methods are unable to handle. For the first time, a stereo-based approach shows superiority in scenarios like a static ego vehicle and moving objects. BEVStereo achieves the new state-of-the-art in the camera-only track of nuScenes dataset while maintaining memory efficiency. Codes have been released.
Han Bao 0008, Zheng Ge, Jianjian Sun
AAAI3
2023 BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object Detection
abstract
In this research, we propose a new 3D object detector with a trustworthy depth estimation, dubbed BEVDepth, for camera-based Bird's-Eye-View~(BEV) 3D object detection. Our work is based on a key observation -- depth estimation in recent approaches is surprisingly inadequate given the fact that depth is essential to camera 3D detection. Our BEVDepth resolves this by leveraging explicit depth supervision. A camera-awareness depth estimation module is also introduced to facilitate the depth predicting capability. Besides, we design a novel Depth Refinement Module to counter the side effects carried by imprecise feature unprojection. Aided by customized Efficient Voxel Pooling and multi-frame mechanism, BEVDepth achieves the new state-of-the-art 60.9% NDS on the challenging nuScenes test set while maintaining high efficiency. For the first time, the NDS score of a camera model reaches 60%. Codes have been released.
Zheng Ge, Guanyi Yu, Zengran Wang, Yukang Shi, Jianjian Sun
AAAI2
2023 Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection Generalization
abstract
In this paper, we analyse the generalization ability of binary classifiers for the task of deepfake detection. We find that the stumbling block to their generalization is caused by the unexpected learned identity representation on images. Termed as the Implicit Identity Leakage, this phenomenon has been qualitatively and quantitatively verified among various DNNs. Furthermore, based on such understanding, we propose a simple yet effective method named the ID-unaware Deepfake Detection Model to reduce the influence of this phenomenon. Extensive experimental results demonstrate that our method outperforms the state-of-the-art in both in-dataset and cross-dataset evaluation. The code is available at https://github.com/megvii-research/CADDM.
Shichao Dong 0001, Jin Wang 0039, Renhe Ji, Jiajun Liang, Haoqiang Fan, Zheng Ge
CVPR6
2023 MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D Perception
abstract
This paper proposes an efficient multi-camera to Bird’s-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from poor efficiency or rely on device-specific operators, hindering the broad application of BEV models. In contrast, our method generates BEV features efficiently with only convolutions and matrix multiplications (MatMul). Specifically, we propose describing the BEV feature as the MatMul of image feature and a sparse Feature Transporting Matrix (FTM). A Prime Extraction module is then introduced to compress the dimension of image features and reduce FTM’s sparsity. Moreover, we propose the Ring & Ray Decomposition to replace the FTM with two matrices and reformulate our pipeline to reduce calculation further. Compared to existing methods, MatrixVT enjoys a faster speed and less memory footprint while remaining deploy-friendly. Extensive experiments on nuScenes and Waymo benchmarks demonstrate that our method is highly efficient but obtains results on par with the SOTA method in object detection and map segmentation tasks.
Zheng Ge, Xiangyu Zhang 0005
ICCV2
2023 Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?
Runpei Dong, Zekun Qi, Linfeng Zhang 0001, Jianjian Sun, Zheng Ge, Li Yi 0001, Kaisheng Ma
ICLR6
2023 Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative Pretraining
abstract
Mainstream 3D representation learning approaches are built upon contrastive or generative modeling pretext tasks, where great improvements in performance on various downstream tasks have been achieved. However, we find these two paradigms have different characteristics: (i) contrastive models are data-hungry that suffer from a representation over-fitting issue; (ii) generative models have a data filling issue that shows inferior data scaling capacity compared to contrastive models. This motivates us to learn 3D representations by sharing the merits of both paradigms, which is non-trivial due to the pattern difference between the two paradigms. In this paper, we propose contrast with reconstruct (ReCon) that unifies these two paradigms. ReCon is trained to learn from both generative modeling teachers and cross-modal contrastive teachers through ensemble distillation, where the generative student is used to guide the contrastive student. An encoder-decoder style ReCon-block is proposed that transfers knowledge through cross attention with stop-gradient, which avoids pretraining over-fitting and pattern difference issues. ReCon achieves a new state-of-the-art in 3D representation learning, e.g., 91.26% accuracy on ScanObjectNN. Codes have been released at https://github.com/qizekun/ReCon.
Zekun Qi, Runpei Dong, Guofan Fan, Zheng Ge, Xiangyu Zhang 0005, Kaisheng Ma, Li Yi 0001
ICML4
2022 Dense Teacher: Dense Pseudo-Labels for Semi-supervised Object Detection
Zheng Ge, Weixin Mao, Jian Sun 0001
ECCV (9)2
2021 OTA: Optimal Transport Assignment for Object Detection
abstract
Recent advances in label assignment in object detection mainly seek to independently define positive/negative training samples for each ground-truth (gt) object. In this paper, we innovatively revisit the label assignment from a global perspective and propose to formulate the assigning procedure as an Optimal Transport (OT) problem – a well-studied topic in Optimization Theory. Concretely, we define the unit transportation cost between each demander (anchor) and supplier (gt) pair as the weighted summation of their classification and regression losses. After formulation, finding the best assignment solution is converted to solve the optimal transport plan at minimal transportation costs, which can be solved via Sinkhorn-Knopp Iteration. On COCO, a single FCOS-ResNet-50 detector equipped with Optimal Transport Assignment (OTA) can reach 40.7% mAP under 1× scheduler, outperforming all other existing assigning methods. Extensive experiments conducted on COCO and CrowdHuman further validate the effectiveness of our proposed OTA, especially its superiority in crowd scenarios. The code is available at https://github.com/Megvii-BaseDetection/OTA.
Zheng Ge, Osamu Yoshie, Jian Sun 0001
CVPR1
2021 Delving deep into the imbalance of positive proposals in two-stage object detection
abstract
Imbalance issue is a major yet unsolved bottleneck for the current object detection models. In this work, we observe two crucial yet never discussed imbalance issues. The first imbalance lies in the large number of low-quality RPN proposals, which makes the R-CNN module (i.e., post-classification layers) become highly biased towards the negative proposals in the early training stage. The second imbalance stems from the unbalanced ground-truth numbers across different testing images, resulting in the imbalance of the number of potentially existing positive proposals in testing phase. To tackle these two imbalance issues, we incorporates two innovations into Faster R-CNN: 1) an R-CNN Gradient Annealing (RGA) strategy to enhance the impact of positive proposals in the early training stage. 2) a set of Parallel R-CNN Modules (PRM) with different positive/negative sampling ratios during training on one same backbone. Our RGA and PRM can totally bring 2.0% improvements on AP on COCO minival. Experiments on CrowdHuman further validates the effectiveness of our innovations across various kinds of object detection tasks.
Zheng Ge, Zequn Jie, Chengzheng Li, Osamu Yoshie
Neurocomputing1
2021 LLA: Loss-aware label assignment for dense pedestrian detection
abstract
Label assignment has been widely studied in general object detection because of its great impact on detectors’ performance. In the field of dense pedestrian detection, human bodies are often heavily entangled, making label assignment more important. However, none of the existing label assignment method focuses on crowd scenarios. Motivated by this, we propose Loss-aware Label Assignment (LLA) to boost the performance of pedestrian detectors in crowd scenarios. Concretely, LLA first calculates classification (cls) and regression (reg) losses between each anchor and ground-truth (GT) pair. A joint loss is then defined as the weighted summation of cls and reg losses as the assigning indicator. Finally, anchors with top K minimum joint losses for a certain GT box are assigned as its positive anchors. Anchors that are not assigned to any GT box are considered negative. LLA is simple but effective. Experiments on CrowdHuman and CityPersons show that such a simple label assigning strategy can boost MR by 9.53% and 5.47% on two famous one-stage detectors – RetinaNet and FCOS, becoming the first one-stage detector that surpasses Faster R-CNN in crowd scenarios.
Zheng Ge, Osamu Yoshie
Neurocomputing1
2020 NMS by Representative Region: Towards Crowded Pedestrian Detection by Proposal Pairing
abstract
Although significant progress has been made in pedestrian detection recently, pedestrian detection in crowded scenes is still challenging. The heavy occlusion between pedestrians imposes great challenges to the standard Non-Maximum Suppression (NMS). A relative low threshold of intersection over union (IoU) leads to missing highly overlapped pedestrians, while a higher one brings in plenty of false positives. To avoid such a dilemma, this paper proposes a novel Representative Region NMS (R2NMS) approach leveraging the less occluded visible parts, effectively removing the redundant boxes without bringing in many false positives. To acquire the visible parts, a novel Paired-Box Model (PBM) is proposed to simultaneously predict the full and visible boxes of a pedestrian. The full and visible boxes constitute a pair serving as the sample unit of the model, thus guaranteeing a strong correspondence between the two boxes throughout the detection pipeline. Moreover, convenient feature integration of the two boxes is allowed for the better performance on both full and visible pedestrian detection tasks. Experiments on the challenging CrowdHuman and CityPersons benchmarks sufficiently validate the effectiveness of the proposed approach on pedestrian detection in the crowded situation.
Zheng Ge, Zequn Jie, Osamu Yoshie
CVPR2
2020 PS-RCNN: Detecting Secondary Human Instances in a Crowd via Primary Object Suppression
abstract
Detecting human bodies in highly crowded scenes is a challenging problem. Two main reasons result in such a problem: 1). weak visual cues of heavily occluded instances can hardly provide sufficient information for accurate detection; 2). heavily occluded instances are easier to be suppressed by Non-Maximum-Suppression (NMS). To address these two issues, we introduce a variant of two-stage detectors called PS-RCNN. PS-RCNN first detects slightly/none occluded objects by an R-CNN [1] module (referred as P-RCNN), and then suppress the detected instances by human-shaped masks so that the features of heavily occluded instances can stand out. After that, PS-RCNN utilizes another R-CNN module specialized in heavily occluded human detection (referred as S-RCNN) to detect the rest missed objects by P-RCNN. Final results are the ensemble of the outputs from these two RCNNs. Moreover, we introduce a High Resolution RoI Align (HRRA) module to retain as much of fine-grained features of visible parts of the heavily occluded humans as possible. Our PS-RCNN significantly improves recall and AP by 4.49% and 2.92% respectively on CrowdHuman [2], compared to the baseline. Similar improvements on Widerperson [3] are also achieved by the PS-RCNN.
Zheng Ge, Zequn Jie, Osamu Yoshie
ICME1
2020 DualBox: Generating BBox Pair with Strong Correspondence via Occlusion Pattern Clustering and Proposal Refinement
abstract
Despite the rapid development of pedestrian detection, the problem of dense pedestrian detection is still unsolved, especially the upper limit of Recall caused by Non-Maximum-Suppression (NMS). Out of this reason, R2NMS [1] is proposed to simultaneously detect full and visible body bounding boxes, by replacing the full body BBoxes with less occluded visible body BBoxes in the NMS algorithm, achieving a higher recall. However, the P-RPN and P-RCNN modules proposed in R2NMS for simultaneous high quality full and visible body prediction require non-trivial positive/negative assigning strategies for anchor BBoxes. To simplify the prerequisites and improve the utility of R2NMS, we incorporate clustering analysis into the learning of visible body proposals from full body proposals. Furthermore, to reduce the computation complexity caused by the large number of potential visible body proposals, we introduce a novel occlusion pattern prediction branch on top of the R-CNN module (i.e. F-RCNN) to select the best matched visible proposals for each full body proposals and then feed them into another R-CNN module (i.e. V-RCNN). Incorporated with R2NMS, our DualBox model can achieve competitive performance while only requires few hyper-parameters. We validate the effectiveness of the proposed approach on the CrowdHuman [2] and CityPersons [3] datasets. Experimental results show that our approach achieves promising performance for detecting both non-occluded and occluded pedestrians, especially heavily occluded ones.
Zheng Ge, Chuyu Hu, Baiqiao Qiu, Osamu Yoshie
ICPR1
2018 Fast Portrait Matting Using Spatial Detail-Preserving Network
Shaofan Cai, Biao Leng, Guanglu Song, Zheng Ge
ICONIP (6)4