EDBT 2026 Demo / reviewers in the wild / expert
Xiongfei Su
dblp:216/2648
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-7733-3159ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
7 papers |
Image and video processing · 77% Computational photography and imaging · 23% | |
| Artificial intelligence
4 papers |
Generative modeling · 60% Efficient and distributed learning · 29% Video understanding and tracking · 11% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.0 | 3 | 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator · NeurIPS 2025 DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution · NeurIPS 2025 Binarized Diffusion Model for Image Super-Resolution · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › few-step generation
one-step diffusion |
1.7 | 2 | 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator · NeurIPS 2025 DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution · NeurIPS 2025 |
Image and video processing › super-resolution › image super-resolution › generative image super-resolution
diffusion-based super-resolution |
1.7 | 2 | 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator · NeurIPS 2025 DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution · NeurIPS 2025 |
Image and video processing › super-resolution
image super-resolution |
1.6 | 2 | 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator · NeurIPS 2025 Binarized Diffusion Model for Image Super-Resolution · NeurIPS 2024 |
Computational photography and imaging
non-line-of-sight imaging |
1.6 | 2 | 2025 | Dual-branch Graph Feature Learning for NLOS Imaging · AAAI 2025 Plug-and-Play Algorithms for Dynamic Non-line-of-sight Imaging · ACM Trans. Graph. 2024 |
Image and video processing › image restoration
image dehazing |
0.9 | 1 | 2025 | Prior-guided Hierarchical Harmonization Network for Efficient Image Dehazing · AAAI 2025 |
Image and video processing
image restoration |
0.9 | 1 | 2025 | Prior-guided Hierarchical Harmonization Network for Efficient Image Dehazing · AAAI 2025 |
Computational photography and imaging › non-line-of-sight imaging
NLOS reconstruction |
0.9 | 1 | 2025 | Dual-branch Graph Feature Learning for NLOS Imaging · AAAI 2025 |
Image and video processing › super-resolution
video super-resolution |
0.9 | 1 | 2025 | DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Binarized Diffusion Model for Image Super-Resolution · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
network binarization |
0.8 | 1 | 2024 | Binarized Diffusion Model for Image Super-Resolution · NeurIPS 2024 |
Image and video processing
image reconstruction |
0.8 | 1 | 2024 | Plug-and-Play Algorithms for Dynamic Non-line-of-sight Imaging · ACM Trans. Graph. 2024 |
Image and video processing › image restoration › inverse problem › inverse problem regularization
plug-and-play priors |
0.8 | 1 | 2024 | Plug-and-Play Algorithms for Dynamic Non-line-of-sight Imaging · ACM Trans. Graph. 2024 |
Computer vision › Video understanding and tracking
video reconstruction |
0.7 | 1 | 2023 | Adaptive Deep PnP Algorithm for Video Snapshot Compressive Imaging · Int. J. Comput. Vis. 2023 |
Image and video processing › compressive sensing
compressive imaging |
0.7 | 1 | 2023 | Adaptive Deep PnP Algorithm for Video Snapshot Compressive Imaging · Int. J. Comput. Vis. 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator · NeurIPS 2025 |
Image and video processing › image restoration
denoising |
0.2 | 1 | 2024 | Plug-and-Play Algorithms for Dynamic Non-line-of-sight Imaging · ACM Trans. Graph. 2024 |
Methods — techniques the papers use, named apart from their topics
perceptual loss · 1.7latent-pixel training · 1.7knowledge distillation · 1.7fine-tuning · 1.7diffusion discriminator · 1.7vision transformer · 0.9histogram equalization · 0.9graph neural network · 0.9dark channel prior · 0.9bright channel prior · 0.9u-net · 0.8diffusion model · 0.8binarization · 0.8deep plug-and-play · 0.7adaptive algorithm · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prior-guided Hierarchical Harmonization Network for Efficient Image DehazingabstractImage dehazing is a crucial task that involves the enhancement of degraded images to recover their sharpness and textures. While vision Transformers have exhibited impressive results in diverse dehazing tasks, their quadratic complexity and lack of dehazing priors pose significant drawbacks for real-world applications. In this paper, guided by triple priors, Bright Channel Prior (BCP), Dark Channel Prior (DCP), and Histogram Equalization (HE), we propose a Prior-guided Hierarchical Harmonization Network (PGHHNet) for image dehazing. PGHNet is built upon the UNet-like architecture with an efficient encoder and decoder, consisting of two module types: (1) Prior aggregation module that injects BCP/DCP and selects diverse contexts with gating attention. (2) Feature harmonization modules that subtract low-frequency components from spatial and channel aspects and learn more informative feature distributions to equalize the feature maps. Inspired by observing the sparsity of BCP/DCP and the histogram equalization, we harmonize the deep features using a histogram equation-guided module and further leverage BCP/DCP to guide spatial attention through a sandwich module as the bottleneck. Comprehensive experiments demonstrate that our model efficiently attains the highest level of performance among existing methods across four different datasets for image dehazing tasks. Xiongfei Su, Yuning Cui 0001, Yulun Zhang 0001, Zheng Chen 0014, Zongliang Wu, Zedong Wang, Yuanlong Zhang, Xin Yuan 0002 |
AAAI | 1 |
| 2025 | Dual-branch Graph Feature Learning for NLOS ImagingabstractThe domain of non-line-of-sight (NLOS) imaging is advancing rapidly, offering the capability to reveal occluded scenes that are not directly visible. However, contemporary NLOS systems face several significant challenges: (1) The computational and storage requirements are profound due to the inherent three-dimensional grid data structure, which restricts practical application. (2) The simultaneous reconstruction of albedo and depth information requires a delicate balance using hyperparameters in the loss function, rendering the concurrent reconstruction of texture and depth information difficult. This paper introduces the innovative methodology, DG-NLOS, which integrates an albedo-focused reconstruction branch dedicated to albedo information recovery and a depth-focused reconstruction branch that extracts geometrical structure, to overcome these obstacles. The dual-branch framework segregates content delivery to the respective reconstructions, thereby enhancing the quality of the retrieved data. To our knowledge, we are the first to employ the GNN as a fundamental component to transform dense NLOS grid data into sparse structural features for efficient reconstruction. Comprehensive experiments demonstrate that our method attains the highest level of performance among existing methods across synthetic and real data. Xiongfei Su, Lina Liu 0010, Zheng Chen 0014, Yulun Zhang 0001, Juntian Ye, Feihu Xu, Xin Yuan 0002 |
AAAI | 1 |
| 2025 | Unfolding Framework with Complex-Valued Deformable Attention for High-Quality Computer-Generated Hologram GenerationabstractComputer-generated holography (CGH) has gained wide attention with deep learning-based algorithms. However, due to its nonlinear and ill-posed nature, challenges remain in achieving accurate and stable reconstruction. Specifically, (i) the widely used end-to-end networks treat the reconstruction model as a black box, ignoring underlying physical relationships, which reduces interpretability and flexibility. (ii) CNN-based CGH algorithms have limited receptive fields, hindering their ability to capture long-range dependencies and global context. (iii) Angular spectrum method (ASM)-based models are constrained to finite near-fields. In this paper, we propose a Deep Unfolding Network (DUN) that decomposes gradient descent into two modules: an adaptive bandwidth-preserving model (ABPM) and a phase-domain complex-valued denoiser (PCD), providing more flexibility. ABPM allows for wider working distances compared to ASM-based methods. At the same time, PCD leverages its complex-valued deformable self-attention module to capture global features and enhance performance, achieving a PSNR over 35 dB. Experiments on simulated and real data show state-of-the-art results. Code is available at https://github.com/HannahZhang1926/Complex-Valued-Deformable-Transformer-for-CGH. Haomiao Zhang, Zhangyuan Li, Yanling Piao, Xiaodong Wang 0026, Xiongfei Su, Xin Yuan 0002 |
ICME | 7 |
| 2025 | DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-ResolutionabstractDiffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training overhead on video data and stringent fidelity demands. To tackle the above issues, we propose DOVE, an efficient one-step diffusion model for real-world VSR. DOVE is obtained by fine-tuning a pretrained video diffusion model (*i.e.*, CogVideoX). To effectively train DOVE, we introduce the latent-pixel training strategy. The strategy employs a two-stage scheme to gradually adapt the model to the video super-resolution task. Meanwhile, we design a video processing pipeline to construct a high-quality dataset tailored for VSR, termed HQ-VSR. Fine-tuning on this dataset further enhances the restoration capability of DOVE. Extensive experiments show that DOVE exhibits comparable or superior performance to multi-step diffusion-based VSR methods. It also offers outstanding inference efficiency, achieving up to a **28$\times$** speed-up over existing methods such as MGLD-VSR. Code is available at: https://github.com/zhengchen1999/DOVE. Zheng Chen 0014, Zichen Zou, Xiongfei Su, Xin Yuan 0002, Yulun Zhang 0001 |
NeurIPS | 4 |
| 2025 | Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion DiscriminatorabstractDiffusion models have demonstrated excellent performance for real-world image super-resolution (Real-ISR), albeit at high computational costs. Most existing methods are trying to derive one-step diffusion models from multi-step counterparts through knowledge distillation (KD) or variational score distillation (VSD). However, these methods are limited by the capabilities of the teacher model, especially if the teacher model itself is not sufficiently strong. To tackle these issues, we propose a new One-Step \textbf{D}iffusion model with a larger-scale \textbf{D}iffusion \textbf{D}iscriminator for SR, called D$^3$SR. Our discriminator is able to distill noisy features from any time step of diffusion models in the latent space. In this way, our diffusion discriminator breaks through the potential limitations imposed by the presence of a teacher model. Additionally, we improve the perceptual loss with edge-aware DISTS (EA-DISTS) to enhance the model's ability to generate fine details. Our experiments demonstrate that, compared with previous diffusion-based methods requiring dozens or even hundreds of steps, our D$^3$SR attains comparable or even superior results in both quantitative metrics and qualitative evaluations. Moreover, compared with other methods, D$^3$SR achieves at least $3\times$ faster inference speed and reduces parameters by at least 30\%. Jianze Li, Jiezhang Cao, Zichen Zou, Xiongfei Su, Xin Yuan 0002, Yulun Zhang 0001, Xiaokang Yang 0001 |
NeurIPS | 4 |
| 2025 | Frequency-Prompted Image Restoration to Enhance Perception in Intelligent Transportation SystemsabstractHigh perceptual image quality is crucial for intelligent transportation systems (ITS), including autonomous vehicles, digital twins, and surveillance infrastructure. However, images captured in adverse weather conditions or dynamic environments often suffer from various visibility degradations. To address this issue, image restoration aims to recover missing details and remove distortions from degraded observations, thereby enhancing the usability of visual data in intelligent transportation applications. Inspired by the success of prompt learning in natural language processing, recent studies have explored prompt-based approaches for various image restoration tasks. However, most of these methods operate in the spatial domain. Given the importance of frequency learning in image restoration, particularly in reducing the spectral discrepancy between degraded and sharp image pairs, this study investigates the use of frequency prompts through a plug-and-play mechanism consisting of a prompt generation module and a prompt integration module. Specifically, the prompt generation module encodes frequency information by aggregating pre-defined learnable parameters, guided by the implicitly decomposed spectra of the input features. The learned prompts are then integrated into the feature spectra via dual-dimensional attention, dynamically guiding the reconstruction process and enabling more effective frequency-aware learning. To validate the effectiveness of the proposed plug-in module, we integrate it into both CNN-based and Transformer-based backbones. Extensive experiments demonstrate that the CNN-based variant achieves state-of-the-art performance on 15 datasets across five representative image restoration tasks. Furthermore, it generalizes well to composite degradation scenarios. The Transformer-based model performs competitively with state-of-the-art methods under two all-in-one image restoration settings. Finally, the effectiveness of our models in enhancing perception for ITS is empirically verified. Yuning Cui 0001, Xiongfei Su, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Binarized Diffusion Model for Image Super-ResolutionabstractAdvanced diffusion models (DMs) perform impressively in image super-resolution (SR), but the high memory and computational costs hinder their deployment. Binarization, an ultra-compression algorithm, offers the potential for effectively accelerating DMs. Nonetheless, due to the model structure and the multi-step iterative attribute of DMs, existing binarization methods result in significant performance degradation. In this paper, we introduce a novel binarized diffusion model, BI-DiffSR, for image SR. First, for the model structure, we design a UNet architecture optimized for binarization. We propose the consistent-pixel-downsample (CP-Down) and consistent-pixel-upsample (CP-Up) to maintain dimension consistent and facilitate the full-precision information transfer. Meanwhile, we design the channel-shuffle-fusion (CS-Fusion) to enhance feature fusion in skip connection. Second, for the activation difference across timestep, we design the timestep-aware redistribution (TaR) and activation function (TaA). The TaR and TaA dynamically adjust the distribution of activations based on different timesteps, improving the flexibility and representation alability of the binarized module. Comprehensive experiments demonstrate that our BI-DiffSR outperforms existing binarization methods. Code is released at: https://github.com/zhengchen1999/BI-DiffSR. Zheng Chen 0014, Haotong Qin, Xiongfei Su, Xin Yuan 0002, Linghe Kong, Yulun Zhang 0001 |
NeurIPS | 4 |
| 2024 | Plug-and-Play Algorithms for Dynamic Non-line-of-sight ImagingabstractNon-line-of-sight (NLOS) imaging has the ability to recover 3D images of scenes outside the direct line of sight, which is of growing interest for diverse applications. Despite the remarkable progress, NLOS imaging of dynamic objects is still challenging. It requires a large amount of multibounce photons for the reconstruction of single-frame data. To overcome this obstacle, we develop a computational framework for dynamic time-of-flight NLOS imaging based on plug-and-play (PnP) algorithms. By combining imaging forward model with the deep denoising network from the computer vision community, we show a 4 frames-per-second (fps) 3D NLOS video recovery (128 × 128 × 512) in post-processing. Our method leverages the temporal similarity among adjacent frames and incorporates sparse priors and frequency filtering. This enables higher-quality reconstructions for complex scenes. Extensive experiments are conducted to verify the superior performance of our proposed algorithm both through simulations and real data. Juntian Ye, Yu Hong 0004, Xiongfei Su, Xin Yuan 0002, Feihu Xu |
ACM Trans. Graph. | 3 |
| 2023 | Multi-scale Iterative Model-guided Unfolding Network for NLOS ReconstructionabstractAbstract Non‐line‐of‐sight (NLOS) imaging can reconstruct hidden objects by analyzing diffuse reflection of relay surfaces, and is potentially used in autonomous driving, medical imaging and national defense. Despite the challenges of low signal‐to‐noise ratio (SNR) and ill‐conditioned problem, NLOS imaging has developed rapidly in recent years. While deep neural networks have achieved impressive success in NLOS imaging, most of them lack flexibility when dealing with multiple spatial‐temporal resolution and multi‐scene images in practical applications. To bridge the gap between learning methods and physical priors, we present a novel end‐to‐end Multi‐scale Iterative Model‐guided Unfolding (MIMU), with superior performance and strong flexibility. Furthermore, we overcome the lack of real training data with a general architecture that can be trained in simulation. Unlike existing encoder‐decoder architectures and generative adversarial networks, the proposed method allows for only one trained model adaptive for various dimensions, such as various sampling time resolution, various spatial resolution and multiple channels for colorful scenes. Simulation and real‐data experiments verify that the proposed method achieves better reconstruction results both in quality and quantity than existing methods. Xiongfei Su, Yu Hong 0004, Juntian Ye, Feihu Xu, Xin Yuan 0002 |
Comput. Graph. Forum | 1 |
| 2023 | Adaptive Deep PnP Algorithm for Video Snapshot Compressive Imaging
Zongliang Wu, Chengshuai Yang, Xiongfei Su, Xin Yuan 0002 |
Int. J. Comput. Vis. | 3 |