VLDB 2026 Research / reviewers in the wild / expert
Xiao Pan 0001
dblp:01/6672-1
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 37% Vision and language · 24% Representation and self-supervised learning · 21% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 72% Rendering · 28% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | InsightEdit: Towards Better Instruction Following for Image Editing · CVPR 2025 |
Visual content generation and editing
image editing |
0.9 | 1 | 2025 | InsightEdit: Towards Better Instruction Following for Image Editing · CVPR 2025 |
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing |
0.9 | 1 | 2025 | InsightEdit: Towards Better Instruction Following for Image Editing · CVPR 2025 |
Computer vision › 3D vision
human rendering |
0.7 | 1 | 2023 | TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering · ICCV 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
0.6 | 1 | 2022 | In-N-Out Generative Learning for Dense Unsupervised Video Segmentation · ACM Multimedia 2022 |
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation |
0.6 | 1 | 2022 | In-N-Out Generative Learning for Dense Unsupervised Video Segmentation · ACM Multimedia 2022 |
Computer vision › Video understanding and tracking
video object segmentation |
0.6 | 1 | 2022 | In-N-Out Generative Learning for Dense Unsupervised Video Segmentation · ACM Multimedia 2022 |
Computer vision › Video understanding and tracking › multi-camera video analysis
multiview video understanding |
0.2 | 1 | 2023 | TransHuman: A Transformer-based Human Representation for Generalizable Neural Human Rendering · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual correspondence learning |
0.2 | 1 | 2022 | In-N-Out Generative Learning for Dense Unsupervised Video Segmentation · ACM Multimedia 2022 |
Methods — techniques the papers use, named apart from their topics
two-stream bridging · 1.7multimodal large language model · 1.7transformer · 1.3deformable radiance fields · 1.3SMPL · 1.3vision transformer · 0.6generative learning · 0.6contrastive learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InsightEdit: Towards Better Instruction Following for Image EditingabstractIn this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, Instruct-Diffusion, and SmartEdit have explored end-to-end editing. However, two limitations still remain: First, existing datasets suffer from low resolution, poor background consistency, and overly simplistic instructions. Second, current approaches mainly condition on the text while the rich image information is underexplored, therefore inferior in complex instruction following and maintaining background consistency. Targeting these issues, we first curated the AdvancedEdit dataset using a novel data construction pipeline, formulating a large-scale dataset with high visual quality, complex instructions, and good background consistency. Then, to further inject the rich image information, we introduce a two-stream bridging mechanism utilizing both the textual and visual features reasoned by the powerful Multimodal Large Language Models (MLLM) to guide the image editing process more precisely. Extensive results demonstrate that our approach, InsightEdit, achieves state-of-the-art performance, excelling in complex instruction following and maintaining high background consistency with the original image. The project page is https://poppyxu.github.io/InsightEditweb/. Yingjing Xu, Jiazhi Wang, Xiao Pan 0001 |
CVPR | 4 |
| 2025 | FRPGS: Fast, Robust, and Photorealistic Monocular Dynamic Scene Reconstruction With Deformable 3D GaussiansabstractDynamic reconstruction technology presents significant promise for applications in visual and interactive fields. Current techniques utilizing 3D Gaussian Splatting show favorable results and fast reconstruction speed. However, as scene expanding, using individual Gaussian structure (i) leads to instability in large-scale dynamic reconstruction, marked by abrupt deformation, and (ii) the heuristic densification of individuals suffers significant redundancy. Tackling these issues, we propose a jointed Gaussian representation method named FRPGS, which learns the global information and the deformation using center Gaussians and generates the neural Gaussians around them for local detail. Specifically, FRPGS employs center Gaussians initialized from point clouds, which are learned with a deformation field for representing global relationships and dynamic motion over time. Then, for each center Gaussian, attribute networks generate neural Gaussians that move under the linked center Gaussian driving, thereby ensuring structural integrity during movement within this joint-based representation. Finally, to reduce Gaussian redundancy, a densification strategy is developed based on the average cumulative gradient of the associated neural Gaussians, imposing strict limits on the growing of center Gaussians without compromising accuracy. Additionally, we established a large-scale dynamic indoor dataset at the MuLong Laboratory of ZTE Corporation. Evaluations demonstrate that FRPGS significantly outperforms state-of-the-art methods in both training efficiency and reconstruction quality, achieving over a 50% (up to 74%) improvement in efficiency on an RTX 4090. FRPGS also supports the 4K resolution reconstruction of 60 frames simultaneously. Xiao Pan 0001, Daquan Feng, Wenzhe Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GD-NeRF: Generative Detail Compensation for One-shot Generalizable Neural Radiance FieldsabstractIn this article, we focus on the one-shot novel view synthesis task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous One-shot Generalizable Neural Radiance Field (OG-NeRF) methods solve this task in a finetuning-free manner, yet suffer from the blurry issue due to the encoder-only architecture that highly relies on the limited reference image. On the other hand, recent diffusion-based image-to-3D methods show vivid plausible results via distilling pre-trained 2D diffusion models, yet require tedious per-scene optimization. Targeting these issues, we propose GD-NeRF, a generative detail compensation framework that is both capable of producing vivid plausible details and is finetuning-free. Following a coarse-to-fine strategy, it is mainly composed of a One-stage Parallel Pipeline (OPP) and a Diffusion-based 3D-consistent Enhancer (Diff3DE). At the coarse stage, OPP first efficiently integrates the GAN model into the existing OG-NeRF pipeline for injecting primary in-distribution details. Then, at the fine stage, Diff3DE further leverages the pre-trained diffusion models to complement rich out-distribution details while maintaining decent 3D consistency. Extensive experiments on both the synthetic and real-world datasets show that GD-NeRF noticeably improves the vivid details while eliminating the need for per-scene finetuning. Xiao Pan 0001, Zongxin Yang, Shuai Bai, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Dynamic gradient reactivation for backward compatible person re-identification
Xiao Pan 0001, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Pattern Recognit. | 1 |
| 2023 | TransHuman: A Transformer-based Human Representation for Generalizable Neural Human RenderingabstractIn this paper, we focus on the task of generalizable neural human rendering which trains conditional Neural Radiance Fields (NeRF) from multi-view videos of different characters. To handle the dynamic human motion, previous methods have primarily used a SparseConvNet (SPC)-based human representation to process the painted SMPL. However, such SPC-based representation i) optimizes under the volatile observation space which leads to the pose-misalignment between training and inference stages, and ii) lacks the global relationships among human parts that is critical for handling the incomplete painted SMPL. Tackling these issues, we present a brand-new framework named TransHuman, which learns the painted SMPL under the canonical space and captures the global relationships between human parts with transformers. Specifically, TransHuman is mainly composed of Transformer-based Human Encoding (TransHE), Deformable Partial Radiance Fields (DPaRF), and Fine-grained Detail Integration (FDI). TransHE first processes the painted SMPL under the canonical space via transformers for capturing the global relationships between human parts. Then, DPaRF binds each output token with a deformable radiance field for encoding the query point under the observation space. Finally, the FDI is employed to further integrate fine-grained information from reference images. Extensive experiments on ZJU-MoCap and H36M show that our TransHuman achieves a significantly new state-of-the-art performance with high efficiency. Project page: https://pansanity666.github.io/TransHuman/ Xiao Pan 0001, Zongxin Yang, Chang Zhou 0005, Yi Yang 0001 |
ICCV | 1 |
| 2023 | Dynamic graph transformer for 3D object detection
Siyuan Ren 0001, Xiao Pan 0001, Binling Nie |
Knowl. Based Syst. | 2 |
| 2022 | In-N-Out Generative Learning for Dense Unsupervised Video SegmentationabstractIn this paper, we focus on unsupervised learning for Video Object Segmentation (VOS) which learns visual correspondence (i.e., the similarity between pixel-level features) from unlabeled videos. Previous methods are mainly based on the contrastive learning paradigm, which optimize either in image level or pixel level. Image-level optimization (e.g., the spatially pooled feature of ResNet) learns robust high-level semantics but is sub-optimal since the pixel-level features are optimized implicitly. By contrast, pixel-level optimization is more explicit, however, it is sensitive to the visual quality of training data and is not robust to object deformation. To complementarily perform these two levels of optimization in a unified framework, we propose the In-aNd-Out (INO) generative learning from a purely generative perspective with the help of naturally designed class tokens and patch tokens in Vision Transformer (ViT). Specifically, for image-level optimization, we force the out-view imagination from local to global views on class tokens, which helps capture high-level semantics, and we name it as out-generative learning. As to pixel-level optimization, we perform in-view masked image modeling on patch tokens, which recovers the corrupted parts of an image via inferring its fine-grained structure, and we term it as in-generative learning. To discover the temporal information better, we additionally force the inter-frame consistency from both feature and affinity matrix levels. Extensive experiments on DAVIS-2017 val and YouTube-VOS 2018 val show that our INO outperforms previous state-of-the-art methods by significant margins. Xiao Pan 0001, Peike Li, Zongxin Yang, Huiling Zhou, Chang Zhou 0005, Hongxia Yang, Jingren Zhou 0001, Yi Yang 0001 |
ACM Multimedia | 1 |
| 2022 | SFGN: Representing the sequence with one super frame for video person re-identification
Xiao Pan 0001, Hao Luo 0004, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Knowl. Based Syst. | 1 |