VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Shan
dblp:127/8709
· DBLP profile ↗
11ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenesabstract3D Gaussian Splatting (3DGS) has shown promising results for Novel View Synthesis. However, while it is quite effective when based on high-quality images, its performance declines as image quality degrades, due to lack of resolution, motion blur, noise, compression artifacts, or other factors common in real-world data collection. While some solutions have been proposed for specific types of degradation, general techniques are still missing. To address the problem, we propose a robust HQGS that significantly enhances the 3DGS under various degradation scenarios. We first analyze that 3DGS lacks sufficient attention in some detailed regions in low-quality scenes, leading to the absence of Gaussian primitives in those areas and resulting in loss of detail in the rendered images. To address this issue, we focus on leveraging edge structural information to provide additional guidance for 3DGS, enhancing its robustness. First, we introduce an edge-semantic fusion guidance module that combines rich texture information from high-frequency edge-aware maps with semantic information from images. The fused features serve as prior guidance to capture detailed distribution across different regions, bringing more attention to areas with detailed edge information and allowing for a higher concentration of Gaussian primitives to be assigned to such areas. Additionally, we present a structural cosine similarity loss to complement pixel-level constraints, further improving the quality of the rendered images. Extensive experiments demonstrate that our method offers better robustness and achieves the best results across various degraded scenes. Source code and trained models are publicly available at: \url{https://github.com/linxin0/HQGS}. Shi Luo, Xiaojun Shan, Chao Ren 0002, Lu Qi 0001, Ming-Hsuan Yang 0001, Nuno Vasconcelos |
ICLR | 3 |
| 2025 | EgoPrivacy: What Your First-Person Camera Says About You?abstractWhile the rapid proliferation of wearable cameras has raised significant concerns about egocentric video privacy, prior work has largely overlooked the unique privacy threats posed to the camera wearer. This work investigates the core question: How much privacy information about the camera wearer can be inferred from their first-person view videos? We introduce EgoPrivacy, the first large-scale benchmark for the comprehensive evaluation of privacy risks in egocentric vision. EgoPrivacy covers three types of privacy (demographic, individual, and situational), defining seven tasks that aim to recover private information ranging from fine-grained (e.g., wearer's identity) to coarse-grained (e.g., age group). To further emphasize the privacy threats inherent to egocentric vision, we propose Retrieval-Augmented Attack, a novel attack strategy that leverages ego-to-exo retrieval from an external pool of exocentric videos to boost the effectiveness of demographic privacy attacks. An extensive comparison of the different attacks possible under all threat models is presented, showing that private information of the wearer is highly susceptible to leakage. For instance, our findings indicate that foundation models can effectively compromise wearer privacy even in zero-shot settings by recovering attributes such as identity, scene, gender, and race with 70–80% accuracy. Our code and data are available at https://github.com/williamium3000/ego-privacy. Yijiang Li, Genpei Zhang, Yi Li 0051, Xiaojun Shan, Dashan Gao 0001, Jiancheng Lyu, Ning Bi, Nuno Vasconcelos |
ICML | 5 |
| 2025 | EARTH: Epidemiology-Aware Neural ODE with Continuous Disease Transmission GraphabstractEffective epidemic forecasting is critical for public health strategies and efficient medical resource allocation, especially in the face of rapidly spreading infectious diseases. However, existing deep-learning methods often overlook the dynamic nature of epidemics and fail to account for the specific mechanisms of disease transmission. In response to these challenges, we introduce an innovative end-to-end framework called Epidemiology-Aware Neural ODE with Continuous Disease Transmission Graph (EARTH) in this paper. To learn continuous and regional disease transmission patterns, we first propose EANO, which seamlessly integrates the neural ODE approach with the epidemic mechanism, considering the complex spatial spread process during epidemic evolution. Additionally, we introduce GLTG to model global infection trends and leverage these signals to guide local transmission dynamically. To accommodate both the global coherence of epidemic trends and the local nuances of epidemic transmission patterns, we build a cross-attention approach to fuse the most meaningful information for forecasting. Through the smooth synergy of both components, EARTH offers a more robust and flexible approach to understanding and predicting the spread of infectious diseases. Extensive experiments show EARTH superior performance in forecasting real-world epidemics compared to state-of-the-art methods. The code is available at https://github.com/GuanchengWan/EARTH. Guancheng Wan, Zewen Liu 0005, Xiaojun Shan, Max S. Y. Lau, B. Aditya Prakash, Wei Jin 0009 |
ICML | 3 |
| 2025 | OverLayBench: A Benchmark for Layout-to-Image Generation with Dense OverlapsabstractDespite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary challenges: (1) large overlapping regions and (2) overlapping instances with minimal semantic distinction. Through both qualitative examples and quantitative analysis, we demonstrate how these factors degrade generation quality. To systematically assess this issue, we introduce OverLayScore, a novel metric that quantifies the complexity of overlapping bounding boxes. Our analysis reveals that existing benchmarks are biased toward simpler cases with low OverLayScore values, limiting their effectiveness in evaluating models under more challenging conditions. To reduce this gap, we present OverLayBench, a new benchmark featuring balanced OverLayScore distributions and high-quality annotations. As an initial step toward improved performance on complex overlaps, we also propose CreatiLayout-AM, a model trained on a curated amodal mask dataset. Together, our contributions establish a foundation for more robust layout-to-image generation under realistic and challenging scenarios. Haiyang Xu 0002, Xiang Zhang 0015, Ethan Armand, Divyansh Srivastava, Xiaojun Shan, Jianwen Xie, Zhuowen Tu |
NeurIPS | 7 |
| 2025 | HoliGS: Holistic Gaussian Splatting for Embodied View SynthesisabstractWe propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dynamic NeRF pipelines, which struggle with training overhead in minute-long captures, our method leverages invertible Gaussian Splatting deformation networks to reconstruct large-scale, dynamic environments accurately. Specifically, we decompose each scene into a static background plus time-varying objects, each represented by learned Gaussian primitives undergoing global rigid transformations, skeleton-driven articulation, and subtle non-rigid deformations via an invertible neural flow. This hierarchical warping strategy enables robust free-viewpoint novel-view rendering from various embodied camera trajectories by attaching Gaussians to a complete canonical foreground shape (e.g., egocentric or third-person follow), which may involve substantial viewpoint changes and interactions between multiple actors. Our experiments demonstrate that HoliGS achieves superior reconstruction quality on challenging datasets while significantly reducing both training and rendering time compared to state-of-the-art monocular deformable NeRFs. These results highlight a practical and scalable solution for EVS in real-world scenarios. The source code will be released. Botao Ye, Xiaojun Shan, Weijie Lyu, Lu Qi 0001, Kelvin C. K. Chan, Yinxiao Li, Ming-Hsuan Yang 0001 |
NeurIPS | 4 |
| 2024 | DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving ScenesabstractWe present DrivingGaussian, an efficient and effective framework for surrounding dynamic autonomous driving scenes. For complex scenes with moving objects, we first sequentially and progressively model the static background of the entire scene with incremental static 3D Gaussians. We then leverage a composite dynamic Gaussian graph to handle multiple moving objects, individually reconstructing each object and restoring their accurate positions and occlusion relationships within the scene. We further use a LiDAR prior for Gaussian Splatting to reconstruct scenes with greater details and maintain panoramic consistency. DrivingGaussian outperforms existing methods in dynamic driving scene reconstruction and enables photorealistic surround-view synthesis with high-fidelity and multi-camera consistency. Our project page is at: https:/github.com/VDIGPKU/DrivingGaussian. Xiaojun Shan, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2024 | AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data BiasesabstractTandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the high-throughput analysis of protein composition in biological tissues. Despite the development of several deep learning methods for predicting amino acid sequences (peptides) responsible for generating the observed mass spectra, training data biases hinder further advancements of \emph{de novo} peptide sequencing. Firstly, prior methods struggle to identify amino acids with Post-Translational Modifications (PTMs) due to their lower frequency in training data compared to canonical amino acids, further resulting in unsatisfactory peptide sequencing performance. Secondly, various noise and missing peaks in mass spectra reduce the reliability of training data (Peptide-Spectrum Matches, PSMs). To address these challenges, we propose AdaNovo, a novel and domain knowledge-inspired framework that calculates Conditional Mutual Information (CMI) between the mass spectra and amino acids or peptides, using CMI for robust training against above biases. Extensive experiments indicate that AdaNovo outperforms previous competitors on the widely-used 9-species benchmark, meanwhile yielding 3.6\% - 9.4\% improvements in PTMs identification. The supplements contain the code. Jun Xia 0001, Shaorong Chen, Xiaojun Shan, Wenjie Du 0003, Zhangyang Gao, Cheng Tan 0012, Bozhen Hu, Jiangbin Zheng 0002, Stan Z. Li |
NeurIPS | 4 |
| 2023 | SAMPLING: Scene-adaptive Hierarchical Multiplane Images Representation for Novel View Synthesis from a Single ImageabstractRecent novel view synthesis methods obtain promising results for relatively small scenes, e.g., indoor environments and scenes with a few objects, but tend to fail for unbounded outdoor scenes with a single image as input. In this paper, we introduce SAMPLING, a Scene-adaptive Hierarchical Multiplane Images Representation for Novel View Synthesis from a Single Image based on improved multiplane images (MPI). Observing that depth distribution varies significantly for unbounded outdoor scenes, we employ an adaptive-bins strategy for MPI to arrange planes in accordance with each scene image. To represent intricate geometry and multi-scale details, we further introduce a hierarchical refinement branch, which results in high-quality synthesized novel views. Our method demonstrates considerable performance gains in synthesizing large-scale unbounded outdoor scenes using a single image on the KITTI dataset and generalizes well to the unseen Tanks and Temples dataset. The code and models will be made available at https://pkuvdig.github.io/SAMPLING/. Xiaojun Shan, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2023 | Enhancing Knowledge Transfer for Task Incremental Learning with Data-free SubnetworkabstractAs there exist competitive subnetworks within a dense network in concert with Lottery Ticket Hypothesis, we introduce a novel neuron-wise task incremental learning method, namely Data-free Subnetworks (DSN), which attempts to enhance the elastic knowledge transfer across the tasks that sequentially arrive. Specifically, DSN primarily seeks to transfer knowledge to the new coming task from the learned tasks by selecting the affiliated weights of a small set of neurons to be activated, including the reused neurons from prior tasks via neuron-wise masks. And it also transfers possibly valuable knowledge to the earlier tasks via data-free replay. Especially, DSN inherently relieves the catastrophic forgetting and the unavailability of past data or possible privacy concerns. The comprehensive experiments conducted on four benchmark datasets demonstrate the effectiveness of the proposed DSN in the context of task-incremental learning by comparing it to several state-of-the-art baselines. In particular, DSN enables the knowledge transfer to the earlier tasks, which is often overlooked by prior efforts. Qiang Gao 0003, Xiaojun Shan, Fan Zhou 0002 |
NeurIPS | 2 |
| 2014 | Dehazing algorithm's high performance and parallel computing for GF-1 Satellite ImagesabstractRecent years, China has launched a series of remote sensing satellites, such as HJ-1, FY3, GF-1, etc. Making the extreme rapid growth of high resolution remote sensing images. In this case, preliminary data processing is becoming more and more important, dehazing is an important step of radiometric processing in prfeliminary data processing. However, the existing dehazing methods are quite slow. In this paper, we are going to discuss a faster dehazing algorithm based on large-scale median filtering, by using the multi-resolution remote sensing image data of GF-1 satellite, we can highly optimize the algorithm itself..Parallel computing techniques are used to improve the processing speed, and the proposed algorithm is implemented in standard C++. On the computer with Intel core i7-3770 and CPU 3.4 GHz, our algorithm can decrease the dehazing time from one hour to less than five minutes for a 4-band, 10000×10000 pixel size, 16 bits unsigned integer data type image. Xiaojun Shan, Zheng Zhang 0021 |
DSAA | 2 |
| 2013 | Modis data-based spatial consistency correction of low-resolution multi-source remote sensing imageryabstractThe spatial consistency of low-resolution multi-source remote sensing data is of great importance for their combination in global change research. Currently, many methods are developed for the precise geometric correction of single kind of low-resolution data. However, the spatial consistency correction method is still need to be developed when many different kinds of low-resolution sensors' data are taken into consideration altogether, which is aimed to make their spectral data become consistent in geo-location. MODIS surface reflectance products, as they are of high accuracy of geo-location and data quality among low-resolution data, the spectral data of the multi-source low-resolution sensors are corrected to be consistency with it, which contain the level 1B data of NOAA/AVHRR, FY-3/VIRR, FY-3/MERSI, FY-2/VISSR. The proposed method in this paper can conducted spatial consistency correction on multi-source low-resolution remote sensing data precisely, efficiently, and automatically, which is based on contour point coarse matching and contour precise matching. Yongquan Zhao, Xiaojun Shan |
IGARSS | 2 |