EDBT 2026 Demo / reviewers in the wild / expert
Xiaotong Luo
dblp:169/3960
· DBLP profile ↗
21ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpikingIR: A Novel Converted Spiking Neural Network for Efficient Image RestorationabstractImage restoration has made great progress with the rise of deep learning, but its energy consumption limits its real-world applications. Spiking Neural Networks (SNNs) are seen as energy-efficient alternatives to Artificial Neural Networks (ANNs). Applying SNNs to image restoration (IR) remains challenging, primarily due to the limited information capacity of spike-based signals. This limitation leads to quantization errors and information loss, while IR tasks are highly sensitive to output precision and error. Thus, the restoration performance suffers significantly. To address this challenge, we propose SpikingIR, an ANN-to-SNN conversion framework for IR that reduces information loss and quantization error. SpikingIR mainly consists of two components: Convolutional Pixel Mapping (CPM) and Membrane Potential Reuse Neuron (MPRN), which are designed to alleviate quantization errors and information loss in the output and intermediate layers, respectively. Specifically, CPM maps discrete outputs into a continuous space, better aligning with pixel-level details. From the perspective of information entropy, we show that outputs of CPM contain more information than the original outputs. MPRN introduces a post-processing step with relaxed firing conditions to extract residual membrane potential, reducing information waste. Furthermore, we fine-tune the converted model to jointly optimize both accuracy and energy efficiency. Experimental results demonstrate that SpikingIR achieves performance comparable to ANN counterparts across various IR benchmarks while reducing energy consumption by up to 50%. Yang Ouyang, Xiaotong Luo, Yanyun Qu |
AAAI | 3 |
| 2026 | Diffusion Once and Done: Degradation-Aware LoRA for All-in-One Image RestorationabstractDiffusion models have revealed powerful potential in all-in-one image restoration (AiOIR), which is talented in generating abundant texture details. The existing AiOIR methods either retrain a diffusion model or fine-tune the pretrained diffusion model with extra conditional guidance. However, they often suffer from high inference costs and limited adaptability to diverse degradation types. In this paper, we propose an efficient AiOIR method, Diffusion Once and Done (DOD), which aims to achieve superior restoration performance with only one-step sampling of Stable Diffusion (SD) models. Specifically, multi-degradation feature modulation is first introduced to capture different degradation prompts with a pretrained diffusion model. Then, parameter-efficient conditional low-rank adaptation integrates the prompts to enable the fine-tuning of the SD model for adapting to different degradation types. Besides, a high-fidelity detail enhancement module is integrated into the decoder of SD to improve structural and textural details. Experiments demonstrate that our method outperforms existing diffusion-based restoration approaches in both visual quality and inference efficiency. Ni Tang, Xiaotong Luo, Liangtai Zhou, Dongxiao Zhang, Yanyun Qu |
AAAI | 2 |
| 2026 | RAIPP: Anchoring visual drift via foundation model priors for unsupervised incremental anomaly detection
Jiangshan Zhao, Deyu Zeng, Xiaotong Luo, Zongze Wu 0001 |
Pattern Recognit. | 4 |
| 2025 | S3SR: Towards Efficient Image Super-Resolution with Selective State Space ModelabstractThough Transformer-based image super-resolution (SR) has made remarkable progress, the burdensome computation complexity hinders its applications in memory-limited devices. Existing efficient Transformer-based image SR methods mainly focus on designing efficient local window self-attention mechanisms to improve computational efficiency. However, the limited receptive field of local windows often fails to capture global contextual information effectively. Recently, the Selective State Space Model, e.g., Mamba, has shown powerful potential for long-range dependencies modeling with linear complexity. In this work, we propose a selective state space model for efficient image SR, dubbed S3SR. Specifically, we design the Local-then-Global Fusion Block as the core component, which employs different convolution and a 2D cross scan mechanism to take advantage of local patch texture and global relevance. Extensive experiments have demonstrated the superiority of our S3SR, which even outperforms the efficient Transformer-based SR methods, using less computational cost but with a larger global receptive field. Xiaotong Luo, Zekun Ai, Yanyun Qu |
ICME | 2 |
| 2025 | Farewell to CycleGAN: Single GAN with decoupled constraint for unpaired image dehazing
Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
Neurocomputing | 1 |
| 2024 | SkipDiff: Adaptive Skip Diffusion Model for High-Fidelity Perceptual Image Super-resolutionabstractIt is well-known that image quality assessment usually meets with the problem of perception-distortion (p-d) tradeoff. The existing deep image super-resolution (SR) methods either focus on high fidelity with pixel-level objectives or high perception with generative models. The emergence of diffusion model paves a fresh way for image restoration, which has the potential to offer a brand-new solution for p-d trade-off. We experimentally observed that the perceptual quality and distortion change in an opposite direction with the increase of sampling steps. In light of this property, we propose an adaptive skip diffusion model (SkipDiff), which aims to achieve high-fidelity perceptual image SR with fewer sampling steps. Specifically, it decouples the sampling procedure into coarse skip approximation and fine skip refinement stages. A coarse-grained skip diffusion is first performed as a high-fidelity prior to obtaining a latent approximation of the full diffusion. Then, a fine-grained skip diffusion is followed to further refine the latent sample for promoting perception, where the fine time steps are adaptively learned by deep reinforcement learning. Meanwhile, this approach also enables faster sampling of diffusion model through skipping the intermediate denoising process to shorten the effective steps of the computation. Extensive experimental results show that our SkipDiff achieves superior perceptual quality with plausible reconstruction accuracy and a faster sampling speed. Xiaotong Luo, Yuan Xie 0006, Yanyun Qu, Yun Fu 0001 |
AAAI | 1 |
| 2024 | AdaFormer: Efficient Transformer with Adaptive Token Sparsification for Image Super-resolutionabstractEfficient transformer-based models have made remarkable progress in image super-resolution (SR). Most of these works mainly design elaborate structures to accelerate the inference of the transformer, where all feature tokens are propagated equally. However, they ignore the underlying characteristic of image content, i.e., various image regions have distinct restoration difficulties, especially for large images (2K-8K), failing to achieve adaptive inference. In this work, we propose an adaptive token sparsification transformer (AdaFormer) to speed up the model inference for image SR. Specifically, a texture-relevant sparse attention block with parallel global and local branches is introduced, aiming to integrate informative tokens from the global view instead of only in fixed local windows. Then, an early-exit strategy is designed to progressively halt tokens according to the token importance. To estimate the plausibility of each token, we adopt a lightweight confidence estimator, which is constrained by an uncertainty-guided loss to obtain a binary halting mask about the tokens. Experiments on large images have illustrated that our proposal reduces nearly 90% latency against SwinIR on Test8K, while maintaining a comparable performance. Xiaotong Luo, Zekun Ai, Qiuyuan Liang, Ding Liu 0001, Yuan Xie 0006, Yanyun Qu, Yun Fu 0001 |
AAAI | 1 |
| 2024 | Data-Free Learning for Lightweight Multi-Weather Image RestorationabstractImage restoration has made a remarkable performance with the large-scale training data and increasing model capacity. However, the burdensome model complexity hinders the mode deployment on resource-constrained devices. Besides, the training data may be unavailable due to some constraints, which undoubtedly affects the efficient model learning. In this paper, we propose an effective data-free model compression framework for lightweight multi-weather image restoration, which consists of data generation and model distillation stages. Specifically, a data generator is first utilized to synthesize degradation-aware samples from a latent distribution. Then, the on-the-shelf teacher model provides a pseudo-label to supervise the training of the student model. To ensure the diversity of the training data, adversarial learning is adopted to maximize the dependency between teacher and student models. Moreover, we adopt a contrastive regularization constraint to further improve model representation. Experimental results show that our proposal achieves comparable performance with the student model trained with the original data and some unsupervised methods for image dehazing and deraining tasks. Hongzhan Huang, Xiaotong Luo, Yanyun Qu |
ISCAS | 3 |
| 2024 | SkipVSR: Adaptive Patch Routing for Video Super-Resolution with Inter-Frame MaskabstractDeep neural networks have revealed enormous potential in video super-resolution (VSR), yet the expensive computational expense limits their deployment on resource-limited devices and actual scenarios, especially for restoring multiple frames simultaneously. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To accelerate the inference of VSR models, we propose a scalable method based on adaptive patch routing to achieve practical speedup. Specifically, we design a confidence estimator to predict the aggregation performance of each block for adjacent patch information. It learns to dynamically perform block skipping, i.e., choose which basic blocks of the VSR network to execute during inference so as to reduce total computation to the maximum extent without degrading reconstruction accuracy dramatically. However, we observe that skipping error would be amplified as the hidden states propagate along with recurrent networks. To alleviate the issue, we design temporal feature alignment to guarantee the performance. This proposal essentially proposes an adaptive routing scheme for each patch. Extensive experiments demonstrate that our method can not only accelerate inference but also provide strong quantitative and qualitative results. Built upon the BasicVSR model, our method achieves a speedup of 20% on average, going as high as 50% for some images, while even maintaining competitive performance on REDS4. Zekun Ai, Xiaotong Luo, Yanyun Qu, Yuan Xie 0006 |
ACM Multimedia | 2 |
| 2024 | UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models Priorabstract3D semantic segmentation using an adapting model trained from a source domain with or without accessing unlabeled target-domain data is the fundamental task in computer vision, containing domain adaptation and domain generalization.
The essence of simultaneously solving cross-domain tasks is to enhance the generalizability of the encoder.
In light of this, we propose a groundbreaking universal method with the help of off-the-shelf Visual Foundation Models (VFMs) to boost the adaptability and generalizability of cross-domain 3D semantic segmentation, dubbed $\textbf{UniDSeg}$.
Our method explores the VFMs prior and how to harness them, aiming to inherit the recognition ability of VFMs.
Specifically, this method introduces layer-wise learnable blocks to the VFMs, which hinges on alternately learning two representations during training: (i) Learning visual prompt. The 3D-to-2D transitional prior and task-shared knowledge is captured from the prompt space, and then (ii) Learning deep query. Spatial Tunability is constructed to the representation of distinct instances driven by prompts in the query space.
Integrating these representations into a cross-modal learning framework, UniDSeg efficiently mitigates the domain gap between 2D and 3D modalities, achieving unified cross-domain 3D semantic segmentation.
Extensive experiments demonstrate the effectiveness of our method across widely recognized tasks and datasets, all achieving superior performance over state-of-the-art methods. Remarkably, UniDSeg achieves 57.5\%/54.4\% mIoU on ``A2D2/sKITTI'' for domain adaptive/generalized tasks. Code is available at https://github.com/Barcaaaa/UniDSeg. Mingwei Xing, Yachao Zhang 0001, Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
NeurIPS | 4 |
| 2024 | Global semantic enhancement network for video captioning
Xuemei Luo, Xiaotong Luo, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003 |
Pattern Recognit. | 2 |
| 2023 | Memory-Friendly Scalable Super-Resolution via Rewinding Lottery Ticket HypothesisabstractScalable deep Super-Resolution (SR) models are increasingly in demand, whose memory can be customized and tuned to the computational recourse of the platform. The existing dynamic scalable SR methods are not memory-friendly enough because multi-scale models have to be saved with a fixed size for each model. Inspired by the success of Lottery Tickets Hypothesis (LTH) on image classification, we explore the existence of unstructured scalable SR deep models, that is, we find gradual shrinkage subnetworks of extreme sparsity named winning tickets. In this paper, we propose a Memory-friendly Scalable SR framework (MSSR). The advantage is that only a single scalable model covers multiple SR models with different sizes, instead of reloading SR models of different sizes. Concretely, MSSR consists of the forward and backward stages, the former for model compression and the latter for model expansion. In the forward stage, we take advantage of LTH with rewinding weights to progressively shrink the SR model and the pruning-out masks that form nested sets. Moreover, stochastic self-distillation (SSD) is conducted to boost the performance of sub-networks. By stochastically selecting multiple depths, the current model inputs the selected features into the corresponding parts in the larger model and improves the performance of the current model based on the feedback results of the larger model. In the backward stage, the smaller SR model could be expanded by recovering and fine-tuning the pruned parameters according to the pruning-out masks obtained in the forward. Extensive experiments show the effectiveness of MMSR. The smallest-scale sub-network could achieve the sparsity of 94% and outperforms the compared lightweight SR methods. Xiaotong Luo, Ming Hong, Yanyun Qu, Yuan Xie 0006, Zongze Wu 0001 |
CVPR | 2 |
| 2023 | Joint Feature Aggregation for Stereo Image Super-resolutionabstractStereo image super-resolution (Stereo SR) has been a newly rising and challenging problem with the popular application of dual cameras, which can be used to promote the SR performance by adding auxiliary information from another viewpoint. Most of the existing excellent works have concentrated on leveraging the intrinsic feature correlation of two view images via exploring the non-local attention mechanism. However, they only perform interaction once for feature registration and fusion accompanied by the complex view transition constraint, which cannot fully take advantage of the information in the stereo image pairs. In this paper, we propose a joint feature aggregation network for Stereo SR to calibrate single-view features and integrate cross-view knowledge effectively. Specifically, we introduce a self-calibrated feature extractor to excavate multi-scale and multi-direction features within the single-view image. What’s more, we design an adaptive fusion module with the cross-view attention mechanism, which is utilized to mine and fuse the long-range dependencies between the stereo image pairs so as to get rid of the inflexible cycle constraints. Extensive experimental results demonstrate that our proposal successfully achieves superior performance against the state-of-the-art methods on four datasets. Zekun Ai, Xiaotong Luo, Yanyun Qu |
ICME | 2 |
| 2023 | Learning Re-sampling Methods with Parameter Attribution for Image Super-resolutionabstractSingle image super-resolution (SISR) has made a significant breakthrough benefiting from the prevalent rise of deep neural networks and large-scale training samples. The mainstream deep SR models primarily focus on network architecture design as well as optimization schemes, while few pay attention to the training data. In fact, most of the existing SR methods train the model on uniformly sampled patch pairs from the whole image. However, the uneven image content makes the training data present an unbalanced distribution, i.e., the easily reconstructed region (smooth) occupies the majority of the data, while the hard reconstructed region (edge or texture) has rarely few samples. Based on this phenomenon, we consider rethinking the current paradigm of merely using uniform data sampling way for training SR models. In this paper, we propose a simple yet effective Bi-Sampling Parameter Attribution (BSPA) method for accurate image SR. Specifically, the bi-sampling consists of uniform sampling and inverse sampling, which is introduced to reconcile the unbalanced inherent data bias. The former aims to keep the intrinsic data distribution, and the latter is designed to enhance the feature extraction ability of the model on the hard samples. Moreover, integrated gradient is introduced to attribute the contribution of each parameter in the alternate models trained by both sampling data so as to filter the trivial parameters for further dynamic refinement. By progressively decoupling the allocation of parameters, the SR model can learn a more compact representation. Extensive experiments on publicly available datasets demonstrate that our proposal can effectively boost the performance of baseline methods from the data re-sampling view. Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
NeurIPS | 1 |
| 2023 | Lattice Network for Lightweight Image RestorationabstractDeep learning has made unprecedented progress in image restoration (IR), where residual block (RB) is popularly used and has a significant effect on promising performance. However, the massive stacked RBs bring about burdensome memory and computation cost. To tackle this issue, we aim to design an economical structure for adaptively connecting pair-wise RBs, thereby enhancing the model representation. Inspired by the topological structure of lattice filter in signal processing theory, we elaborately propose the lattice block (LB), where couple butterfly-style topological structures are utilized to bridge pair-wise RBs. Specifically, each candidate structure of LB relies on the combination coefficients learned through adaptive channel reweighting. As a basic mapping block, LB can be plugged into various IR models, such as image super-resolution, image denoising, image deraining, etc. It can avail the construction of lightweight IR models accompanying half parameter amount reduced, while keeping the considerable reconstruction accuracy compared with RBs. Moreover, a novel contrastive loss is exploited as a regularization constraint, which can further enhance the model representation without increasing the inference expenses. Experiments on several IR tasks illustrate that our method can achieve more favorable performance than other state-of-the-art models with lower storage and computation. Xiaotong Luo, Yanyun Qu, Yuan Xie 0006, Yulun Zhang 0001, Cuihua Li, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Self-Mimic Mutual-Distillation for Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification is a newly rising and challenging problem, as there is a significant gap between the visible and infrared images. Though recent methods rapidly narrow the gap, the intra-modality variance is often ignored before inter-modality alignment. In this paper, we study this problem in the knowledge distillation perspective and design a self-mimic mutual-distillation method to reduce the discrepancy of each person from intra-modality feature alignment to cross-modality feature alignment. For intra-modality feature alignment, the self-mimic mechanism is implemented to simultaneously learn globally viewed, stable, and distinguish prototypes for each ID and minimize the intra-modality discrepancy. For inter-modality feature alignment, the mutual distillation is conducted to minimize the cross-modality distribution discrepancy of each person. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that the proposed method achieves the best performance, outperforming state-of-the-art methods by a large margin without adding extra network parameters to the baseline. Especially, on the SYSU-MM01 dataset, our method achieves 64.8% Rank-1 and 60.2% mAP with significant gains over the latest related method. Demao Zhang, Ming Hong, Zheng Wang 0007, Zhizhong Zhang 0001, Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
ICME | 6 |
| 2022 | Adjustable Memory-efficient Image Super-resolution via Individual Kernel SparsityabstractThough single image super-resolution (SR) has witnessed incredible progress, the increasing model complexity impairs its applications in memory-limited devices. To solve this problem, prior arts have aimed to reduce the number of model parameters and sparsity has been exploited, which usually enforces the group sparsity constraint on the filter level and thus is not arbitrarily adjustable for satisfying the customized memory requirements. In this paper, we propose an individual kernel sparsity (IKS) method for memory-efficient and sparsity-adjustable image SR to aid deep network deployment in memory-limited devices. IKS performs model sparsity in the weight level that implicitly allocates the user-defined target sparsity to each individual kernel. To induce the kernel sparsity, a soft thresholding operation is used as a gating constraint for filtering the trivial weights. To achieve adjustable sparsity, a dynamic threshold learning algorithm is proposed, in which the threshold is updated by associated training with the network weight and is adaptively decayed with the guidance of the desired sparsity. This work essentially provides a dynamic parameter reassignment scheme with a given resource budget for an off-the-shelf SR model. Extensive experimental results demonstrate that IKS imparts considerable sparsity with negligible effect on SR quality. The code is available at: https://github.com/RaccoonDML/IKS. Xiaotong Luo, Mingliang Dai, Yulun Zhang 0001, Yuan Xie 0006, Ding Liu 0001, Yanyun Qu, Yun Fu 0001, Junping Zhang |
ACM Multimedia | 1 |
| 2021 | Boosting Lightweight Single Image Super-resolution via Joint-distillationabstractThe rising of deep learning has facilitated the development of single image super-resolution (SISR). However, the growing burdensome model complexity and memory occupation severely hinder its practical deployments on resource-limited devices. In this paper, we propose a novel joint-distillation (JDSR) framework to boost the representation of various off-the-shelf lightweight SR models. The framework includes two stages: the superior LR generation and the joint-distillation learning. The superior LR is obtained from the HR image itself. With less than $300$K parameters, the peer network using superior LR as input can achieve comparable SR performance with large models, e.g., RCAN, with 15M parameters, which enables it as the input of peer network to save the training expense. The joint-distillation learning consists of internal self-distillation and external mutual learning. The internal self-distillation aims to achieve model self-boosting by transferring the knowledge from the deeper SR output to the shallower one. Specifically, each intermediate SR output is supervised by the HR image and the soft label from subsequent deeper outputs. To shrink the capacity gap between shallow and deep layers, a soft label generator is designed in a progressive backward fusion way with meta-learning for adaptive weight fine-tuning. The external mutual learning focuses on obtaining interaction information from a peer network in the process. Moreover, a curriculum learning strategy and a performance gap threshold are introduced for balancing the convergence rate of the original SR model and its peer network. Comprehensive experiments on benchmark datasets demonstrate that our proposal improves the performance of recent lightweight SR models by a large margin, with the same model architecture and inference expense. Xiaotong Luo, Qiuyuan Liang, Ding Liu 0001, Yanyun Qu |
ACM Multimedia | 1 |
| 2020 | LatticeNet: Towards Lightweight Image Super-Resolution with Lattice Block
Xiaotong Luo, Yuan Xie 0006, Yulun Zhang 0001, Yanyun Qu, Cuihua Li, Yun Fu 0001 |
ECCV (22) | 1 |
| 2019 | Joint-attention Discriminator for Accurate Super-resolution via Adversarial TrainingabstractTremendous progress has been witnessed on single image super-resolution (SR), where existing deep SR models achieve impressive performance in objective criteria, e.g., PSNR and SSIM. However, most of the SR methods are limited in visual perception, for example, they look too smooth. Generative adversarial network (GAN) favors SR visual effects over most of the deep SR models but is poor in objective criteria. In order to trade off the objective and subjective SR performance, we design a joint-attention discriminator with which GAN improves the SR performance in PSNR and SSIM, as well as maintaining the visual effect compared with non-attention GAN based SR models. The joint-attention discriminator contains dense channel-wise attention and cross-layer attention blocks. The former is applied in the shallow layers of the discriminator for channel-wise weighting combination of feature maps. The latter is employed to select feature maps in some middle and deep layers for effective discrimination. Extensive experiments are conducted on six benchmark datasets and the experimental results show that our proposed discriminator combining with different generators can achieve more realistic visual performances. Yuan Xie 0006, Xiaotong Luo, Yanyun Qu, Cuihua Li |
ACM Multimedia | 3 |
| 2015 | IBS: an illustrator for the presentation and visualization of biological sequencesabstractUNLABELLED: Biological sequence diagrams are fundamental for visualizing various functional elements in protein or nucleotide sequences that enable a summarization and presentation of existing information as well as means of intuitive new discoveries. Here, we present a software package called illustrator of biological sequences (IBS) that can be used for representing the organization of either protein or nucleotide sequences in a convenient, efficient and precise manner. Multiple options are provided in IBS, and biological sequences can be manipulated, recolored or rescaled in a user-defined mode. Also, the final representational artwork can be directly exported into a publication-quality figure. AVAILABILITY AND IMPLEMENTATION: The standalone package of IBS was implemented in JAVA, while the online service was implemented in HTML5 and JavaScript. Both the standalone package and online service are freely available at http://ibs.biocuckoo.org. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wenzhong Liu, Yubin Xie, Jiyong Ma, Xiaotong Luo, Zhixiang Zuo, Urs Lahrmann, Qi Zhao 0009, Yueyuan Zheng, Yong Zhao 0013, Yu Xue 0001, Jian Ren 0002 |
Bioinform. | 4 |