Xingtong Ge

dblp:303/9641 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting
abstract
Implicit neural representations (INRs) have achieved remarkable success in image representation and compression, but they require substantial training time and memory. Meanwhile, recent 2D Gaussian Splatting (GS) methods (\textit{e.g.}, GaussianImage) offer promising alternatives through efficient primitive-based rendering. However, these methods require excessive Gaussian primitives to maintain high visual fidelity. To exploit the potential of GS-based approaches, we present GaussianImage++, which utilizes limited Gaussian primitives to achieve impressive representation and compression performance. Firstly, we introduce a distortion-driven densification mechanism. It progressively allocates Gaussian primitives according to signal intensity. Secondly, we employ context-aware Gaussian filters for each primitive, which assist in the densification to optimize Gaussian primitives based on varying image content. Thirdly, we integrate attribute-separated learnable scalar quantizers and quantization-aware training, enabling efficient compression of primitive attributes. Experimental results demonstrate the effectiveness of our method. In particular, GaussianImage++ outperforms GaussianImage and INRs-based COIN in representation and compression performance while maintaining real-time decoding and low memory usage.
Xingtong Ge, Tongda Xu, Dailan He, Jun Zhang 0004, Yan Wang 0105
AAAI3
2025 CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
abstract
Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, which leads to suboptimal rate-distortion results. In this paper, we propose a stereo image compression framework, named CAMSIC. CAMSIC independently transforms each image to latent representation and employs a powerful decoder-free Transformer entropy model to capture both spatial and disparity dependencies, by introducing a novel content-aware masked image modeling (MIM) technique. Our content-aware MIM facilitates efficient bidirectional interaction between prior information and estimated tokens, which naturally obviates the need for an extra Transformer decoder. Experiments show that our stereo image codec achieves state-of-the-art rate-distortion performance on two stereo image datasets Cityscapes and InStereo2K with fast encoding and decoding speed.
Shenyuan Gao, Zhening Liu 0001, Jiawei Shao, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Jun Zhang 0004
AAAI5
2025 MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes
abstract
4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, notably the requirement of millions of 4D Gaussians, each with extensive associated attributes, leading to substantial memory and storage cost. This paper introduces a memory-efficient framework for 4DGS. We streamline the color attribute by decomposing it into a per-Gaussian direct color component with only 3 parameters and a shared lightweight alternating current color predictor. This approach eliminates the need for spherical harmonics coefficients, which typically involve up to 144 parameters in classic 4DGS, thereby creating a memory-efficient 4D Gaussian representation. Furthermore, we introduce an entropy-constrained Gaussian deformation technique that uses a deformation field to expand the action range of each Gaussian and integrates an opacity-based entropy loss to limit the number of Gaussians, thus forcing our model to use as few Gaussians as possible to fit a dynamic scene well. With simple half-precision storage and zip compression, our framework achieves a storage reduction by approximately 190$\times$ and 125$\times$ on the Technicolor and Neural 3D Video datasets, respectively, compared to the original 4DGS. Meanwhile, it maintains comparable rendering speeds and scene representation quality, setting a new standard in the field. Code is available at https://github.com/Xinjie-Q/MEGA.
Zhening Liu 0001, Yifan Zhang 0004, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang 0105, Zehong Lin, Shuicheng Yan, Jun Zhang 0004
ICCV4
2025 Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior
abstract
Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512$\times$512 ImageNet images, revealing that: 1) DPS’s conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS’s conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in https://github.com/tongdaxu/Rethinking-Diffusion-Posterior-Sampling-From-Conditional-Score-Estimator-to-Maximizing-a-Posterior.
Tongda Xu, Xiyan Cai, Xingtong Ge, Dailan He, Ya-Qin Zhang, Yan Wang 0105
ICLR4
2025 Spatio-temporal graph neural networks for missing data completion in traffic prediction
abstract
Missing traffic data completion is a key part of the construction of a smart city. However, due to cost constraints and other reasons, many locations do not have sensors to record traffic data. Most research methods do not systematically consider filling in missing traffic data. This study explores a new spatio-temporal feature extraction layer that includes spatio-temporal feature fusion, graph learning on an adaptive adjacency matrix, and a gated recurrent unit with a mask for missing traffic data completion. This idea is based on a hypothesis: missing data can be inferred from the spatio-temporal features of other nearby recorded sensor nodes. Therefore, we propose an end-to-end traffic model dealing with missing data - missing traffic data completion graph neural networks (MTC-GNN). Experiments demonstrate that the proposed model can learn spatio-temporal patterns and fill in speed from traffic data with various missing ratios and outperform existing models.
Jiahui Chen 0010, Yi Yang 0019, Xingtong Ge
Int. J. Geogr. Inf. Sci.5
2024 Task-Aware Encoder Control for Deep Video Compression
abstract
Prior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DvC decoder. Empirical evidence demonstrates that our method is applica-ble across multiple tasks with various existing pre-trained Dv'Cs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bi-trate for different tasks, with only one pre-trained decoder.
Xingtong Ge, Jixiang Luo, Tongda Xu, Guo Lu, Dailan He, Yan Wang 0105, Jun Zhang 0004, Hongwei Qin
CVPR1
2024 Boosting Neural Representations for Videos with a Conditional Decoder
abstract
Implicit neural representations (INRs) have emerged as a promising approach for video storage and processing, showing remarkable versatility across various video tasks. However, existing methods often fail to fully leverage their representation capabilities, primarily due to inadequate alignment of intermediate features during target frame decoding. This paper introduces a universal boosting framework for current implicit video representation approaches. Specifically, we utilize a conditional decoder with a temporal-aware affine transform module, which uses the frame index as a prior condition to effectively align intermediate features with target frames. Besides, we introduce a sinusoidal NeRV-like block to generate diverse intermediate features and achieve a more balanced parameter distribution, thereby enhancing the model's capacity. With a high-frequency information-preserving reconstruction loss, our approach successfully boosts multiple baseline INRs in the reconstruction quality and convergence speed for video regression, and exhibits superior inpainting and interpolation results. Further, we integrate a consistent entropy minimization technique and develop video codecs based on these boosted INRs. Experiments on the UVG dataset confirm that our enhanced codecs significantly outperform baseline INRs and offer competitive rate-distortion performance compared to traditional and learning-based codecs. Code is available at htt ps://github.com/Xin j ieQ/Boosting-NeRV.
Dailan He, Xingtong Ge, Tongda Xu, Yan Wang 0105, Hongwei Qin, Jun Zhang 0004
CVPR4
2024 GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Xingtong Ge, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin, Guo Lu, Jun Zhang 0004
ECCV (9)2
2024 Preprocessing Enhanced Image Compression for Machine Vision
abstract
Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.
Guo Lu, Xingtong Ge, Tianxiong Zhong, Qiang Hu 0003, Jing Geng 0002
IEEE Trans. Circuits Syst. Video Technol.2
2021 Construction of Spatiotemporal Knowledge Graph for Emergency Decision Making
abstract
Disaster emergency decision-making often involves a large amount of spatiotemporal information, and current emergency spatiotemporal analysis methods are often difficult to associate incidents, features and related knowledge at the same time. At present, the research on the temporal and spatial semantics of the knowledge graph has not been combined with the emergency field. This paper proposes to use spatiotemporal knowledge graph as an information management framework, construct conceptual and instance levels for emergency decision-making, define knowledge graph data model at the conceptual level; preprocess multi-source data at the instance level and map it into the content of the knowledge graph, and further Use SWRL to effectively extend OWL semantics and realize rule representation. Finally, this article uses several SPARQL query cases to illustrate how to obtain the required knowledge from the knowledge graph, so as to realize the intelligent service of emergency decision-making.
Jiahui Chen 0010, Xingtong Ge, Weichao Li 0004
IGARSS2