VLDB 2026 Research / reviewers in the wild / expert
Qiaosi Yi
dblp:249/8335
· DBLP profile ↗
15ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-9720-4322ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA ApproachabstractDiffusion prior-based methods have shown impressive results in real-world image super-resolution (SR). However, most existing methods entangle pixel-level and semantic-level SR objectives in the training process, struggling to balance pixel-wise fidelity and perceptual quality. Meanwhile, users have varying preferences on SR results, thus it is demanded to develop an adjustable SR model that can be tailored to different fidelity-perception preferences during inference without re-training. We present Pixel-level and Semantic-level Adjustable SR (PiSA-SR), which learns two LoRA modules upon the pre-trained stable-diffusion (SD) model to achieve improved and adjustable SR results. We first formulate the SD-based SR problem as learning the residual between the low-quality input and the high-quality output, then show that the learning objective can be decoupled into two distinct LoRA weight spaces: one is characterized by the ℓ2-loss for pixel-level regression, and another is characterized by the LPIPS and classifier score distillation losses to extract semantic information from pre-trained classification and SD models. In its default setting, PiSA-SR can be performed in a single diffusion step, achieving leading real-world SR results in both quality and efficiency. By introducing two adjustable guidance scales on the two LoRA modules to control the strengths of pixel-wise fidelity and semantic-level details during inference, PiSA-SR can offer flexible SR results according to user preference without re-training. The source code of our method can be found at https://github.com/csslc/PiSA-SR. Lingchen Sun, Rongyuan Wu, Zhiyuan Ma 0002, Shuaizheng Liu, Qiaosi Yi, Lei Zhang 0006 |
CVPR | 5 |
| 2025 | Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu 0009, Rongyuan Wu, Lingchen Sun, Yuhui Wu 0001, Lei Zhang 0006 |
ICCV | 1 |
| 2025 | DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-ResolutionabstractBenefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Although this randomness is sometimes seen as a limitation, it also introduces a wider perceptual quality range, which can be exploited to improve Real-ISR performance. To this end, we introduce Direct Perceptual Preference Optimization for Real-ISR (DP²O-SR), a framework that aligns generative models with perceptual preferences without requiring costly human annotations. We construct a hybrid reward signal by combining full-reference and no-reference image quality assessment (IQA) models trained on large-scale human preference datasets. This reward encourages both structural fidelity and natural appearance. To better utilize perceptual diversity, we move beyond the standard best-vs-worst selection and construct multiple preference pairs from outputs of the same model. Our analysis reveals that the optimal selection ratio depends on model capacity: smaller models benefit from broader coverage, while larger models respond better to stronger contrast in supervision. Furthermore, we propose hierarchical preference optimization, which adaptively weights training pairs based on intra-group reward gaps and inter-group diversity, enabling more efficient and stable learning. Extensive experiments across both diffusion- and flow-based T2I backbones demonstrate that DP²O-SR significantly improves perceptual quality and generalizes well to real-world benchmarks. Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Tianhe Wu, Qiaosi Yi, Shuai Li 0014, Lei Zhang 0006 |
NeurIPS | 6 |
| 2025 | DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow EditingabstractLeveraging the powerful generation capability of large-scale pretrained text-to-image models, training-free methods have demonstrated impressive image editing results. Conventional diffusion-based methods, as well as recent rectified flow (RF)-based methods, typically reverse synthesis trajectories by gradually adding noise to clean images, during which the noisy latent at the current timestep is used to approximate that at the next timesteps, introducing accumulated drift and degrading reconstruction accuracy. Considering the fact that in RF the noisy latent is estimated through direct interpolation between Gaussian noises and clean images at each timestep, we propose Direct Noise Alignment (DNA), which directly refines the desired Gaussian noise in the noise domain, significantly reducing the error accumulation in previous methods. Specifically, DNA estimates the velocity field of the interpolated noised latent at each timestep and adjusts the Gaussian noise by computing the difference between the predicted and expected velocity field. We validate the effectiveness of DNA and reveal its relationship with existing RF-based inversion methods. Additionally, we introduce a Mobile Velocity Guidance (MVG) to control the target prompt-guided generation process, balancing image background preservation and target object editability. DNA and MVG collectively constitute our proposed method, namely DNAEdit. Finally, we introduce DNA-Bench, a long-prompt benchmark, to evaluate the performance of advanced image editing models. Experimental results demonstrate that our DNAEdit achieves superior performance to state-of-the-art text-guided editing methods. Our code, model, and benchmark will be made publicly available. Minghan Li 0001, Shuai Li 0014, Yuhui Wu 0001, Qiaosi Yi, Lei Zhang 0006 |
NeurIPS | 5 |
| 2025 | Textual prompt guided image restoration
Qiuhai Yan, Aiwen Jiang, Long Peng 0003, Qiaosi Yi |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | High-Frequency Modulated Transformer for Multi-Contrast MRI Super-ResolutionabstractAccelerating the MRI acquisition process is always a key issue in modern medical practice, and great efforts have been devoted to fast MR imaging. Among them, multi-contrast MR imaging is a promising and effective solution that utilizes and combines information from different contrasts. However, existing methods may ignore the importance of the high-frequency priors among different contrasts. Moreover, they may lack an efficient method to fully utilize the information from the reference contrast. In this paper, we propose a lightweight and accurate High-frequency Modulated Transformer (HFMT) for multi-contrast MRI super-resolution. The key ideas of HFMT are high-frequency prior enhancement and its fusion with global features. Specifically, we employ an enhancement module to enhance and amplify the high-frequency priors in the reference and target modalities. In addition, we utilize the Rectangle Window Transformer Block (RWTB) to capture global information in the target contrast. Meanwhile, we propose a novel cross-attention mechanism to fuse the high-frequency enhanced features with the global features sequentially, which assists the network in recovering clear texture details from the low-resolution inputs. Extensive experiments show that our proposed method can reconstruct high-quality images with fewer parameters and faster inference time. Juncheng Li 0003, Hanhui Yang, Qiaosi Yi, Minhua Lu, Jun Shi 0004, Tieyong Zeng |
IEEE Trans. Medical Imaging | 3 |
| 2024 | HFF-Net: A High-Frequency Fidelity Model for Accelerated Parallel MRI ReconstructionabstractMagnetic Resonance Imaging (MRI) plays a crucial role in diagnosing and treating various diseases. However, the long acquisition time of MRI scans often leads to patient discomfort and motion artifacts. Consequently, accelerating MRI speed is essential. Researchers have combined Deep Learning with Compressed Sensing and Parallel Imaging to advance MRI. However, many existing methods fail to effectively recover the fine details and structures in Magnetic Resonance images. To address these challenges, we propose a novel model for accelerated parallel MRI reconstruction. Our model incorporates a high-frequency fidelity method into the reconstruction process, explicitly emphasizing the recovery of high-frequency information. Additionally, we consider the joint priori distribution between the reconstructed images from each coil. Using the variable splitting approach, the proposed model is unrolled as an end-to-end network termed HFF-Net. Experimental results demonstrate that our method outperforms state-of-the-art techniques, yielding high-quality MR images with enhanced detail and fine structure recovery. Zhenggang Yang, Faming Fang, Qiaosi Yi, Guixu Zhang, Fang Li 0004 |
ICME | 3 |
| 2024 | HFGN: High-Frequency residual Feature Guided Network for fast MRI reconstruction
Faming Fang, Le Hu, Qiaosi Yi, Tieyong Zeng, Guixu Zhang |
Pattern Recognit. | 4 |
| 2023 | Frequency Learning via Multi-Scale Fourier Transformer for MRI ReconstructionabstractSince Magnetic Resonance Imaging (MRI) requires a long acquisition time, various methods were proposed to reduce the time, but they ignored the frequency information and non-local similarity, so that they failed to reconstruct images with a clear structure. In this article, we propose Frequency Learning via Multi-scale Fourier Transformer for MRI Reconstruction (FMTNet), which focuses on repairing the low-frequency and high-frequency information. Specifically, FMTNet is composed of a high-frequency learning branch (HFLB) and a low-frequency learning branch (LFLB). Meanwhile, we propose a Multi-scale Fourier Transformer (MFT) as the basic module to learn the non-local information. Unlike normal Transformers, MFT adopts Fourier convolution to replace self-attention to efficiently learn global information. Moreover, we further introduce a multi-scale learning and cross-scale linear fusion strategy in MFT to interact information between features of different scales and strengthen the representation of features. Compared with normal Transformers, the proposed MFT occupies fewer computing resources. Based on MFT, we design a Residual Multi-scale Fourier Transformer module as the main component of HFLB and LFLB. We conduct several experiments under different acceleration rates and different sampling patterns on different datasets, and the experiment results show that our method is superior to the previous state-of-the-art method. Qiaosi Yi, Faming Fang, Guixu Zhang, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Efficient and Accurate Multi-Scale Topological Network for Single Image DehazingabstractSingle image dehazing is a challenging ill-posed problem that has drawn significant attention in the last few years. Recently, convolutional neural networks have achieved great success in image dehazing. However, it is still difficult for these increasingly complex models to recover accurate details from the hazy image. In this paper, we pay attention to the feature extraction and utilization of the input image itself. To achieve this, we propose a Multi-scale Topological Network (MSTN) to fully explore the features at different scales. Meanwhile, we design a Multi-scale Feature Fusion Module (MFFM) and an Adaptive Feature Selection Module (AFSM) to achieve the selection and fusion of features at different scales, so as to achieve progressive image dehazing. This topological network provides a large number of search paths that enable the network to extract abundant image features as well as strong fault tolerance and robustness. In addition, ASFM and MFFM can adaptively select important features and ignore interference information when fusing different scale representations. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods. Qiaosi Yi, Juncheng Li 0003, Faming Fang, Aiwen Jiang, Guixu Zhang |
IEEE Trans. Multim. | 1 |
| 2021 | Structure-Preserving Deraining with Residue Channel Prior GuidanceabstractSingle image deraining is important for many high-level computer vision tasks since the rain streaks can severely degrade the visibility of images, thereby affecting the recognition and analysis of the image. Recently, many CNN-based methods have been proposed for rain removal. Although these methods can remove part of the rain streaks, it is difficult for them to adapt to real-world scenarios and restore high-quality rain-free images with clear and accurate structures. To solve this problem, we propose a Structure-Preserving Deraining Network (SPDNet) with RCP guidance. SPDNet directly generates high-quality rain-free images with clear and accurate structures under the guidance of RCP but does not rely on any rain-generating assumptions. Specifically, we found that the RCP of images contains more accurate structural information than rainy images. Therefore, we introduced it to our deraining network to protect structure information of the rain-free image. Meanwhile, a Wavelet-based Multi-Level Module (WMLM) is proposed as the backbone for learning the background information of rainy images and an Interactive Fusion Module (IFM) is designed to make full use of RCP information. In addition, an iterative guidance strategy is proposed to gradually improve the accuracy of RCP, refining the result in a progressive path. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Code: https://github.com/Joyies/SPDNet Qiaosi Yi, Juncheng Li 0003, Qinyan Dai, Faming Fang, Guixu Zhang, Tieyong Zeng |
ICCV | 1 |
| 2021 | Feedback Network for Mutually Boosted Stereo Image Super-Resolution and Disparity EstimationabstractUnder stereo settings, the problem of image super-resolution (SR) and disparity estimation are interrelated that the result of each problem could help to solve the other. The effective exploitation of correspondence between different views facilitates the SR performance, while the high-resolution (HR) features with richer details benefit the correspondence estimation. According to this motivation, we propose a Stereo Super-Resolution and Disparity Estimation Feedback Network (SSRDE-FNet), which simultaneously handles the stereo image super-resolution and disparity estimation in a unified framework and interact them with each other to further improve their performance. Specifically, the SSRDE-FNet is composed of two dual recursive sub-networks for left and right views. Besides the cross-view information exploitation in the low-resolution (LR) space, HR representations produced by the SR process are utilized to perform HR disparity estimation with higher accuracy, through which the HR features can be aggregated to generate a finer SR result. Afterward, the proposed HR Disparity Information Feedback (HRDIF) mechanism delivers information carried by HR disparity back to previous layers to further refine the SR image reconstruction. Extensive experiments demonstrate the effectiveness and advancement of SSRDE-FNet. Qinyan Dai, Juncheng Li 0003, Qiaosi Yi, Faming Fang, Guixu Zhang |
ACM Multimedia | 3 |
| 2021 | MSNet: A novel end-to-end single image dehazing network with multiple inter-scale dense skip-connectionsabstractAbstract Dehazing is a challenging ill‐posed image restoration task. Various prior‐based and learning‐based methods have been proposed. Among them, end‐to‐end deep models achieve great success on performance improvement. However, most of them are concentrated on feature learning within the same block scale in isolation, and cannot perform associated analysis well on feature characteristics of different scales. Inter‐scale information reuse which is especially beneficial to image restoration is often neglected. Therefore, in this paper, a novel end‐to‐end network with multiple inter‐scale dense skip‐connections for image dehazing is proposed. Sufficient complementary information combination is considered through dense inter‐scale skip‐connections among encoder and decoder block layers. Besides avoiding gradient vanishing, a kind of bottleneck residual block is proposed to control the importance of local gradients at different scales over global learning process. Extensive comparisons and ablation studies on public dehazing datasets and real‐world images have been conducted. The experiment results demonstrate that the proposed novel elements can ensure more stable training process and superior testing performance with great improvements on PSNR and SSIM. Authors' haze‐removal results consistently comply satisfactorily with real situations, having much higher definition and contrast without colour distortion than those from the state‐of‐the‐art methods compared in this paper. Qiaosi Yi, Aiwen Jiang, Xiaolin Deng, Changhong Liu |
IET Image Process. | 1 |
| 2020 | Cumulative Rain Density Sensing Network for Single Image DerainabstractThis paper focuses on single image derain, which aims to restore clear image from single rain image. Through full consideration of different frequency information preservation and the complicated interactions between rain-streaks and background, a novel end-to-end cumulative rain-density sensing network (CRDNet) is proposed for adaptive rain-streaks removal. An effective W-Net with powerful learning ability is proposed as a key component to recover rain-invariant low-frequency signals. A cumulative rain-density classifier with a novel cost-sensitive label encoding strategy is proposed as an auxiliary network to improve discriminative power of extracted high-frequency rain-streaks through multi-task training. The proposed CRDNet has been compared with state-of-the-art methods on two public datasets. The quantitative and visual experimental results demonstrate that it can achieve excellent performance with great improvement. Related source code and models are available on github https://github.com/peylnog/CRDNet. Long Peng 0003, Aiwen Jiang, Qiaosi Yi, Mingwen Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Static Crowd Scene Analysis via Deep Network with Multi-branch Dilated Convolution BlocksabstractIn this paper, we have proposed a static crowd scene analysis network via multi-branch dilated convolution block, called MDBNet. It focuses on a joint task of estimating crowd count and high-quality density map from static single image. The proposed MDBNet follows one-stage object detection framework, and consists of two parts: pre-trained convolutional layers as the front end for high-level feature extraction and cascaded multi-branch dilated convolution block as the back end for context information aggregation on different ranges. Pixel-wise objectness probabilities are predicted and regressed to generate density map. The proposed MDBNet is an easy training model with strong learning ability. We have tested it on two public datasets (ShanghaiTech dataset and the UFC_CC_50 dataset). On almost all evaluation criterions, the proposed method has achieved superior performance. Especially on structure quality criterions, including our newly introduced spatial adjusted mutual information measurement, the MDBNet reports a new state-of-the-art performance. The source code will be distributed depending on publication of our work. Aiwen Jiang, Qiaosi Yi, Xiaolin Deng, Jianyi Wan, Mingwen Wang 0001 |
IJCNN | 3 |