Fudong Ge

dblp:160/4136 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-5759-4159ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 55% 3D vision · 14% Representation and self-supervised learning · 14%
Computer graphics and multimedia
1 paper
Rendering · 67% Geometric modeling and processing · 33%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
detection transformer
1.012026
Integrating Diverse Assignment Strategies into DETRs · AAAI 2026
Computer vision › Image recognition and object detection › object detection › detector training
label assignment
1.012026
Integrating Diverse Assignment Strategies into DETRs · AAAI 2026
Computer vision › Image recognition and object detection
object detection
1.012026
Integrating Diverse Assignment Strategies into DETRs · AAAI 2026
Geometric modeling and processing › 3d reconstruction
3d scene reconstruction
1.012026
HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes · AAAI 2026
Rendering › neural rendering
dynamic gaussian splatting
1.012026
HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes · AAAI 2026
Rendering
gaussian splatting
1.012026
HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes · AAAI 2026
Machine learning › Representation and self-supervised learning
vector quantization
0.812024
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization · NeurIPS 2024
Machine learning › Generative modeling › variational autoencoder
vector-quantized variational autoencoder
0.812024
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization · NeurIPS 2024
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
0.212024
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

one-to-many assignment · 1.0low-rank adaptation · 1.0hybrid supervision · 1.0hierarchical decomposition · 1.0anchor-based gaussian compression · 1.0vector quantization · 0.8token decoder · 0.8VQ-VAE · 0.8
YearPublicationVenuePosition
2026 HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes
abstract
This paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene modeling, face a critical scalability bottleneck. Inspired by recent advances in the 3DGS-based compression beyond autonomous driving, we address this challenge by leveraging the compression capability of anchor-driven methods. However, this is non-trivial as our exploratory experiments reveal that the direct application of this paradigm to dynamic, large-scale urban scenes results in performance degradation. We attribute this phenomenon to the hierarchical anchor design that severely loses dynamic information. To this end, we propose Hierarchical Dynamic Gaussian Splatting (HDGS), a novel framework designed to adapt the anchor-based Gaussian paradigm to 4D urban environments. We first establish a local support network to reinforce inter-anchor consistency, mitigating geometric and appearance fractures caused by supervision attenuation in deep hierarchies. Then, we handle heterogeneous object motion via coarse-to-fine decomposition, where high-level anchors model coarse dynamics and low-level anchors refine them with residual deformations. Third, we introduce a hybrid supervision scheme that fuses global geometric constraints and local pixel-level cues to alleviate geometrically inconsistent reconstruction under sparse LiDAR. Extensive experiments show that HDGS reduces storage by 69.0% while maintaining or even improving rendering fidelity compared to state-of-the-art methods.
Fudong Ge, Hanshi Wang, Weiming Hu 0004
AAAI1
2026 Integrating Diverse Assignment Strategies into DETRs
abstract
Label assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich supervisory signals, they often introduce complex, architecture-specific modifications and typically focus on a single auxiliary strategy, lacking a unified and scalable design. In this paper, we first systematically investigate the effects of ``one-to-many'' supervision and reveal a surprising insight that performance gains are driven not by the sheer quantity of supervision, but by the diversity of the assignment strategies employed. This finding suggests that a more elegant, parameter-efficient approach is attainable. Building on this insight, we propose LoRA-DETR, a flexible and lightweight framework that seamlessly integrates diverse assignment strategies into any DETR-style detector. Our method augments the primary network with multiple Low-Rank Adaptation (LoRA) branches during training, each instantiating a different one-to-many assignment rule. These branches act as auxiliary modules that inject rich, varied supervisory gradients into the main model and are discarded during inference, thus incurring no additional computational cost. This design promotes robust joint optimization while maintaining the architectural simplicity of the original detector. Extensive experiments on different baselines validate the effectiveness of our approach. Our work presents a new paradigm for enhancing detectors, demonstrating that diverse ``one-to-many'' supervision can be integrated to achieve state-of-the-art results without compromising model elegance.
Hanshi Wang, Fudong Ge, Guan Luo, Weiming Hu 0004
AAAI4
2026 Exponential Stabilization of Coupled Semilinear Reaction-Diffusion PDEs Using Mobile Collocated Actuator-Sensor Pairs
Fudong Ge, YangQuan Chen, Zhiqiang Zuo 0001
IEEE Trans Autom. Sci. Eng.1
2024 BEV2PR: BEV-Enhanced Visual Place Recognition with Structural Cues
abstract
In this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) For the methods relying on LiDAR sensors, the integration of LiDAR in robotic systems has led to increased expenses, while the alignment of data between different sensors is also a major challenge. 2) Other image-/camera-based methods, involving integrating RGB images and their derived variants (e.g., pseudo depth images, pseudo 3D point clouds), exhibit several limitations, such as the failure to effectively exploit the explicit spatial relationships between different objects. To tackle the above issues, we design a new BEV-enhanced VPR framework, namely BEV2PR, generating a composite descriptor with both visual cues and spatial awareness based on a single camera. The key points lie in: 1) We use BEV features as an explicit source of structural knowledge in constructing global features. 2) The lower layers of the pretrained backbone from BEV generation are shared for visual and structural streams in VPR, facilitating the learning of fine-grained local features in the visual stream. 3) The complementary visual and structural features can jointly enhance VPR performance. Our BEV2PR framework enables consistent performance improvements over several popular aggregation modules for RGB global features. The experiments on our collected VPR-NuScenes dataset demonstrate an absolute gain of 2.47% on Recall@1 for the strong Conv-AP baseline to achieve the best performance in our setting, and notably, a 18.06% gain on the hard set. The code and dataset will be available at https://github.com/FudongGe/BEV2PR.
Fudong Ge, Shuhan Shen, Weiming Hu 0004
IROS1
2024 VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
abstract
Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{generating} the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. \emph{The question is how to align the PV features with the generative models to facilitate the map estimation}. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on \url{https://github.com/Z1zyw/VQ-Map}.
Fudong Ge, Guan Luo, Bing Li 0001, Zhaoxiang Zhang 0001, Haibin Ling, Weiming Hu 0004
NeurIPS3
2024 Observer-Based Boundary Stabilization of Coupled Semilinear Reaction-Diffusion Neural Networks With Spatially Varying Coefficients via Event-Triggered Controller
abstract
The aim of this article is to propose an observer-based event-triggered Robin boundary control strategy for the exponential stabilization of the coupled semilinear reaction-diffusion neural networks with spatially varying coefficients. Toward this aim, we design an observer to estimate the value of system states by using some of these system values as the available measurement. An observer-based event-triggered boundary stabilizer is then presented to exponentially stabilize the considered systems with the Zeno behavior being excluded. Throughout this article, the main used method is backstepping, which yields an explicit expression of the control formulae. Moreover, we see that the proposed event-triggered boundary control scheme can ensure the desired level of control performance with fewer control law updates. A numerical example is finally given to illustrate the effectiveness of our proposed method.
Fudong Ge, YangQuan Chen
IEEE Trans. Neural Networks Learn. Syst.1
2019 Event-triggered boundary feedback control for networked reaction-subdiffusion processes with input uncertainties
Fudong Ge, YangQuan Chen
Inf. Sci.1