VLDB 2026 Research / reviewers in the wild / expert
Tianshuo Xu
dblp:304/1328
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
3D vision · 41% Generative modeling · 28% Efficient and distributed learning · 15% | |
| Computer graphics and multimedia
3 papers |
Image and video processing · 61% Rendering · 34% Visual content generation and editing · 5% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.1 | 2 | 2025 | Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion · CVPR 2025 From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model · ICRA 2024 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.9 | 1 | 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving · NeurIPS 2025 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving · NeurIPS 2025 |
Computer vision › 3D vision
3d scene reconstruction |
0.9 | 1 | 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving · NeurIPS 2025 |
Computer vision › 3D vision › 3d scene understanding › semantic scene completion
4d occupancy forecasting |
0.9 | 1 | 2025 | Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models · ICRA 2025 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.9 | 1 | 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving · NeurIPS 2025 |
Computer vision › 3D vision › 3d scene understanding
semantic scene completion |
0.9 | 1 | 2025 | Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models · ICRA 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
sharpness-aware minimization |
0.9 | 1 | 2025 | Systematic Investigation of Sparse Perturbed Sharpness-Aware Minimization Optimizer · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.9 | 1 | 2025 | Systematic Investigation of Sparse Perturbed Sharpness-Aware Minimization Optimizer · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Image and video processing › image restoration
image inpainting |
0.9 | 1 | 2025 | Orpaint: a zero-shot inpainting model for oracle bone inscription rubbings with visual mamba block · Sci. China Inf. Sci. 2025 |
Rendering
inverse rendering |
0.9 | 1 | 2025 | Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion · CVPR 2025 |
Rendering
neural rendering |
0.9 | 1 | 2025 | Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion · CVPR 2025 |
Machine learning › Generative modeling › image generation
conditional image generation |
0.8 | 1 | 2024 | From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model · ICRA 2024 |
Machine learning › Generative modeling › image generation › conditional image synthesis
layout-to-image generation |
0.8 | 1 | 2024 | From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model · ICRA 2024 |
Machine learning › Generative modeling
scene generation |
0.8 | 1 | 2024 | From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model · ICRA 2024 |
Image and video processing
image restoration |
0.8 | 1 | 2024 | Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video processing › super-resolution
image super-resolution |
0.8 | 1 | 2024 | Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video processing › super-resolution › image super-resolution
perceptual super-resolution |
0.8 | 1 | 2024 | Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning |
0.5 | 1 | 2021 | CDP: Towards Optimal Filter Pruning via Class-wise Discriminative Power · ACM Multimedia 2021 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | CDP: Towards Optimal Filter Pruning via Class-wise Discriminative Power · ACM Multimedia 2021 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.2 | 1 | 2024 | From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model · ICRA 2024 |
Machine learning › Deep learning architectures and training
loss function design |
0.2 | 1 | 2024 | Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2021 | CDP: Towards Optimal Filter Pruning via Class-wise Discriminative Power · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.7cycle consistency · 1.7visual mamba · 0.9variational autoencoder · 0.9stochastic gradient descent · 0.9prune and dilate block · 0.9motion separation · 0.9large language model · 0.9fisher information · 0.9dynamic-static decoupling · 0.9dynamic sparse training · 0.9entropy-based quantification · 0.8contrastive learning · 0.8GAN-based super-resolution · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream DiffusionabstractRendering and inverse rendering are pivotal tasks in both computer vision and graphics. The rendering equation is the core of the two tasks, as an ideal conditional distribution transfer function from intrinsic properties to RGB images. Despite achieving promising results of existing rendering methods, they merely approximate the ideal estimation for a specific scene and come with a high computational cost. Additionally, the inverse conditional distribution transfer is intractable due to the inherent ambiguity. To address these challenges, we propose a data-driven method that jointly models rendering and inverse rendering as two conditional generation tasks within a single diffusion framework. Inspired by UniDiffuser, we utilize two distinct time schedules to model both tasks, and with a tailored dual streaming module, we achieve cross-conditioning of two pre-trained diffusion models. This unified approach, named Uni-Renderer, allows the two processes to facilitate each other through a cycle-consistent constrain, mitigating ambiguity by enforcing consistency between intrinsic properties and rendered images. Combined with a meticulously prepared dataset, our method effectively decomposition of intrinsic properties and demonstrating a strong capability to recognize changes during rendering. Zhifei Chen, Tianshuo Xu, Wenhang Ge, Leyi Wu, Dongyu Yan, Luozhou Wang, Shunsi Zhang, Ying-Cong Chen |
CVPR | 2 |
| 2025 | Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language ModelsabstractLarge Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode occupancy as input for the LLM and address the category imbalances associated with occupancy, we propose Motion Separation Variational Autoencoder (MS-VAE). This innovative approach utilizes prior knowledge to distinguish dynamic objects from static scenes before inputting them into a tailored Variational Autoencoder (VAE). This separation enhances the model's capacity to concentrate on dynamic trajectories while effectively reconstructing static scenes. The efficacy of Occ-LLM has been validated across key tasks, including 4D occupancy forecasting, self-ego planning, and occupancybased scene question answering. Comprehensive evaluations demonstrate that Occ-LLM significantly surpasses existing state-of-the-art methodologies, achieving gains of about 6% in Intersection over Union (IoU) and 4% in mean Intersection over Union (mIoU) for the task of 4D occupancy forecasting. These findings highlight the transformative potential of Occ-LLM in reshaping current paradigms within robotic and autonomous driving. Tianshuo Xu, Hao Lu 0009, Xu Yan 0005, Yingjie Cai, Ying-Cong Chen |
ICRA | 1 |
| 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingabstractLarge reconstruction model has remarkable progress, which can directly predict 3D or 4D representations for unseen scenes and objects. However, current work has not systematically explored the potential of large reconstruction models in the field of autonomous driving. To achieve this, we introduce the Large 4D Gaussian Reconstruction Model (DrivingRecon). With an elaborate and simple framework design, it not only ensures efficient and high-quality reconstruction, but also provides potential for downstream tasks. There are two core contributions: firstly, the Prune and Dilate Block (PD-Block) is proposed to prune redundant and overlapping Gaussian points and dilate Gaussian points for complex objects. Then, dynamic and static decoupling is tailored to better learn the temporary-consistent geometry across different time. Experimental results demonstrate that DrivingRecon significantly improves scene reconstruction quality compared to existing methods. Furthermore, we explore applications of DrivingRecon in model pre-training, vehicle type adaptation, and scene editing. Our code will be available. Hao Lu 0009, Tianshuo Xu, Wenzhao Zheng, Dalong Du, Masayoshi Tomizuka, Kurt Keutzer, Ying-Cong Chen |
NeurIPS | 2 |
| 2025 | Orpaint: a zero-shot inpainting model for oracle bone inscription rubbings with visual mamba block
Zijie Meng, Yuan-Ze Zeng, Xiang Chang, Tianshuo Xu, Fei Chao 0001, Xixin Cao, Changjing Shang, Qiang Shen 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Systematic Investigation of Sparse Perturbed Sharpness-Aware Minimization OptimizerabstractDeep neural networks often suffer from poor generalization due to complex and non-convex loss landscapes. Sharpness-Aware Minimization (SAM) is a popular solution that smooths the loss landscape by minimizing the maximized change of training loss when adding a perturbation to the weight. However, indiscriminate perturbation of SAM on all parameters is suboptimal and results in excessive computation, double the overhead of common optimizers like Stochastic Gradient Descent (SGD). In this paper, we propose Sparse SAM (SSAM), an efficient and effective training scheme that achieves sparse perturbation by a binary mask. To obtain the sparse mask, we provide two solutions based on Fisher information and dynamic sparse training, respectively. We investigate the impact of different masks, including unstructured, structured, and $N$N:$M$M structured patterns, as well as explicit and implicit forms of implementing sparse perturbation. We theoretically prove that SSAM can converge at the same rate as SAM, i.e., $O(\log T/\sqrt{T})$O(logT/T) . Sparse SAM has the potential to accelerate training and smooth the loss landscape effectively. Extensive experimental results on CIFAR and ImageNet-1K confirm that our method is superior to SAM in terms of efficiency, and the performance is preserved or even improved with a perturbation of merely 50% sparsity. Peng Mi, Li Shen 0008, Tianhe Ren, Yiyi Zhou, Tianshuo Xu, Xiaoshuai Sun, Tongliang Liu, Rongrong Ji, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | NADM: Noise-Aware Diffusion Model for Landscape Painting Video GenerationabstractLandscape painting is a gem of cultural and artistic heritage that showcases the splendor of nature through the deep observations and imaginations of its painters. Limited by traditional techniques, these artworks were confined to static imagery in ancient times, leaving the dynamism of landscapes and the subtleties of artistic sentiment to the viewer's imagination. Recently, emerging text-to-video (T2V) diffusion methods have shown significant promise in video generation, providing hope for the creation of dynamic landscape paintings. However, current T2V methods focus on generating natural videos, emphasizing the capture of details and the authenticity of physical laws. In contrast, landscape painting videos emphasize the overall dynamic aesthetic. Besides, challenges, such as the lack of specific datasets, the intricacy of artistic styles, and the creation of extensive, high-quality videos pose difficulties for these models in generating landscape painting videos. In this article, we propose landscape painting videos-high definition (LPV-HD), a novel T2V dataset for landscape painting videos, and noise-aware diffusion model (NADM), a T2V model that utilizes Stable Diffusion. Specifically, we present a motion module featuring a dual attention mechanism to capture the dynamic transformations of landscape imageries, alongside a noise adapter to leverage unsupervised contrastive learning in the latent space to ensure the overall beauty of the landscape painting video. Following the generation of keyframes, we employ optical flow for frame interpolation to enhance video smoothness. Our method not only retains the essence of the landscape painting imageries but also achieves dynamic transitions, significantly advancing the field of artistic video generation. Source code and dataset are available at https://github.com/llzlh21/NADM. Ding-Ming Liu, Shao-Wei Li, Ruo-Yan Zhou, Lili Liang, Yongguan Hong, Yuan-Ze Zeng, Xiang Chang, Lijiang Li, Tianshuo Xu, Fei Chao 0001, Changjing Shang, Qiang Shen 0001 |
IEEE Trans. Cybern. | 9 |
| 2024 | From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion ModelabstractWe explore Bird’s-Eye View (BEV) generation, converting a BEV map into its corresponding multi-view street images. Valued for its unified spatial representation aiding multi-sensor fusion, BEV is pivotal for various autonomous driving applications. Creating accurate street-view images from BEV maps is essential for portraying complex traffic scenarios and enhancing driving algorithms. Concurrently, diffusion-based conditional image generation models have demonstrated remarkable outcomes, adept at producing diverse, high-quality, and condition-aligned results. Nonetheless, the training of these models demands substantial data and computational resources. Hence, exploring methods to fine-tune these advanced models, like Stable Diffusion, for specific conditional generation tasks emerges as a promising avenue. In this paper, we introduce a practical framework for generating images from a BEV layout. Our approach comprises two main components: the Neural View Transformation and the Street Image Generation. The Neural View Transformation phase converts the BEV map into aligned multi-view semantic segmentation maps by learning the shape correspondence between the BEV and perspective views. Subsequently, the Street Image Generation phase utilizes these segmentations as a condition to guide a fine-tuned latent diffusion model. This finetuning process ensures both view and style consistency. Our model leverages the generative capacity of large pretrained diffusion models within traffic contexts, effectively yielding diverse and condition-coherent street view images. Tianshuo Xu, Fulong Ma, Ying-Cong Chen |
ICRA | 2 |
| 2024 | CPE COIN++: Towards Optimized Implicit Neural Representation Compression Via Chebyshev Positional Encoding
Haocheng Chu, Shaohui Dai, Wenqi Ding, Tianshuo Xu, Pingyang Dai, Shengchuan Zhang, Yan Zhang 0109, Xiang Chang, Chih-Min Lin, Fei Chao 0001, Changjiang Shang, Qiang Shen 0001 |
PRCV (9) | 5 |
| 2024 | Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive OptimizationabstractPSNR-oriented models are a critical class of super-resolution models with applications across various fields. However, these models tend to generate over-smoothed images, a problem that has been analyzed previously from the perspectives of models or loss functions, but without taking into account the impact of data properties. In this paper, we present a novel phenomenon that we term the center-oriented optimization (COO) problem, where a model's output converges towards the center point of similar high-resolution images, rather than towards the ground truth. We demonstrate that the strength of this problem is related to the uncertainty of data, which we quantify using entropy. We prove that as the entropy of high-resolution images increases, their center point will move further away from the clean image distribution, and the model will generate over-smoothed images. Implicitly optimizing the COO problem, perceptual-driven approaches such as perceptual loss, model structure optimization, or GAN-based methods can be viewed. We propose an explicit solution to the COO problem, called Detail Enhanced Contrastive Loss (DECLoss). DECLoss utilizes the clustering property of contrastive learning to directly reduce the variance of the potential high-resolution distribution and thereby decrease the entropy. We evaluate DECLoss on multiple super-resolution benchmarks and demonstrate that it improves the perceptual quality of PSNR-oriented models. Moreover, when applied to GAN-based methods, such as RaGAN, DECLoss helps to achieve state-of-the-art performance, such as 0.093 LPIPS with 24.51 PSNR on 4× downsampled Urban100, validating the effectiveness and generalization of our approach. Tianshuo Xu, Lijiang Li, Peng Mi, Xiawu Zheng, Fei Chao 0001, Rongrong Ji, Yonghong Tian 0001, Qiang Shen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | CDP: Towards Optimal Filter Pruning via Class-wise Discriminative PowerabstractNeural network pruning has shown promising performance in reducing computational complexity and facilitate the deployment of deep neural networks on resource-limited edge devices. Most existing pruning methods focus on the indicators of the filter's weight, gradient, or feature map and regard the weak or similar filters as network redundancy. In contrast, the representation of discriminative power is also a fundamental attribute that analog neural networks to have extraordinary performance in various tasks. However, such representation is neglected in existing works. Alternatively, we propose a novel filter pruning strategy via class-wise discriminative power (CDP). Unlike the previous methods, CDP treats the filters that always yield large or small activation values as redundant and reserves the filters that show different magnitudes in activations as they yield high discriminative power. We further propose to obtain such discriminative power by employing the widely-used Term Frequency-Inverse Document Frequency (TF-IDF) on feature representations across classes. Specifically, the output of a filter is considered as a word, and the whole feature map is considered as a document. Then, TF-IDF is used to generate the relevant score between words and all documents. If a filter has low TF-IDF scores is less discriminate and can be pruned. Thus, the filters with high TF-IDF scores are reserved. To our best knowledge, this is the first work that prunes neural networks through class-wise discriminative power and measures such power by introducing TF-IDF in feature representation among different classes. Without any iterative process, CDP achieves better compression trade-offs comparing to the state-of-the-art compression algorithms. For instance, in VGG-16, we achieve a 68.05%-FLOPs reduction, with a 94.86% Top-1 accuracy on CIFAR-10. Specifically, we compress a 90.12%-FLOPs reduction VGG-16, even retains 93.30% Top-1 accuracy on CIFAR-10. The code is available at https://github.com/Tianshuo-Xu/CDP-Towards-Optimal-Filter-Pruning-via-Class-wise-Discriminative-Power.git Tianshuo Xu, Yuhang Wu 0004, Xiawu Zheng, Teng Xi, Errui Ding, Fei Chao 0001, Rongrong Ji |
ACM Multimedia | 1 |