Qingsen Yan

dblp:206/9166 · DBLP profile ↗
← Back
79ranked-venue papers
29as first author
65since 2021 · last 2026
0000-0003-1010-3540ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 17 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 16 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
abstract
The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-world deployment. In the paper, we devise a ''filter-correlate-compress'' framework to accelerate the MLLM by systematically optimizing multimodal context length during prefilling. The framework first implements FiCoCo-V, a training-free method operating within the vision encoder. It employs a redundancy-based token discard mechanism that uses a novel integrated metric to accurately filter out redundant visual tokens. To mitigate information loss, the framework introduces a correlation-based information recycling mechanism that allows preserved tokens to selectively recycle information from correlated discarded tokens with a self-preserving compression, thereby preventing the dilution of their own core content. The framework's FiCoCo-L variant further leverages task-aware textual priors to perform token reduction directly within the LLM decoder. Extensive experiments demonstrate that the FiCoCo series effectively accelerates a range of MLLMs, achieves up to 14.7× FLOPs reduction with 93.6% performance retention. Our methods consistently outperform state-of-the-art training-free approaches, showcasing effectiveness and generalizability across model architectures, sizes, and tasks without requiring retraining.
Xuyang Liu 0002, Pengxiang Ding, Honggang Chen, Qingsen Yan, Siteng Huang
AAAI8
2026 Alternating exposure control network for real-world environments
Chenyuan Zhao, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Eng. Appl. Artif. Intell.3
2026 History-aware adaptive teacher for cross-domain object detection
Yaoqi Hu, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Expert Syst. Appl.5
2026 AGNet: Attention guided network for single HDR reconstruction
Zhou Gong, Weiyu Zhou, Qingsen Yan
Neurocomputing4
2026 FSCFNet: Lightweight neural networks via multi-dimensional importance-aware optimization
Mengyang Nie, Jinqiu Sun, Hongsong Guoyang, Axi Niu, Yaoqi Hu, Qingsen Yan, Yu Zhu 0004
Neurocomputing6
2026 Generative morphodynamic forecasting enables robust zero-shot volumetric medical segmentation
Duwei Dai, Caixia Dong, Guowei Dai 0001, Bowen Qin, Qingsen Yan
Medical Image Anal.8
2026 AdaPrompt-IR: Adaptive learning to perceive degradation semantic and prompting for all-in-one image restoration
Wei Sun 0036, Qianzhou Wang, Qingsen Yan, Yanning Zhang 0001
Pattern Recognit.5
2026 DT-RSRGAN: An one-off domain translation generative model for real image super-resolution
Shaolin Su, Yu Zhu 0004, Lingmei Zhang, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Pattern Recognit.5
2026 HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
Kangbiao Shi, Yixu Feng, Tao Hu 0013, Peng Wu 0015, Guansong Pang, Qingsen Yan
IEEE Trans. Circuits Syst. Video Technol.7
2026 High Dynamic Range Imaging via Spatial-Frequency Interaction
abstract
publicly available.High Dynamic Range (HDR) imaging aims to reconstruct scenes with a wide range of luminance by fusing multi-exposure Low Dynamic Range (LDR) images. In dynamic scenes with pronounced foreground motion or camera jitter, especially under challenging conditions including extremely low or high luminance, widespread saturation, and substantial motion, existing approaches often encounter ghosting artifacts, spatial misalignment, and degradation of fine structural details. Traditional techniques based on handcrafted priors struggle to generalize to complex motion patterns, while most deep learning-based methods operate exclusively in the spatial domain, limiting their ability to capture global contextual cues and restore high-frequency structures that are better represented in the frequency domain. To address these challenges, we introduce a Dual-Domain Parallel Fusion Network with Prompt Refinement (DDPF-PR), which jointly leverages spatial and frequency-domain features for enhanced HDR reconstruction. Specifically, the proposed framework consists of a Bi-Domain Interaction Module(BDIM), which integrates spatial features for local detail and frequency features for global structure to suppress ghosting artifacts caused by motion. In addition, a Prompt Refinement Module(PRM) is designed to recover fine details in degraded regions such as saturated or misaligned areas by adaptively generating structural cues. Extensive experiments demonstrate that DDPF-PR consistently outperforms state-of-the-art methods across multiple benchmarks in both qualitative and quantitative evaluations. The code will be made publicly available.
Weiyu Zhou, Yongqing Yang, Tao Hu 0013, Pu Hui, Yu Cao 0016, Qingsen Yan, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 Boosting HDR Image Reconstruction via Semantic Knowledge Transfer
abstract
Recovering High Dynamic Range (HDR) images from multiple Standard Dynamic Range (SDR) images becomes challenging when the SDR images exhibit noticeable degradation and missing content. Leveraging scene-specific semantic priors offers a promising solution for restoring heavily degraded regions. However, these priors are typically extracted from sRGB SDR images, the domain/format gap poses a significant challenge when applying it to HDR imaging. To address this issue, we propose a general framework that transfers semantic knowledge derived from SDR domain via self-distillation to boost existing HDR reconstruction. Specifically, the proposed framework first introduces the Semantic Priors Guided Reconstruction Model (SPGRM), which leverages SDR image semantic knowledge to address ill-posed problems in the initial HDR reconstruction results. Subsequently, we leverage a self-distillation mechanism that constrains the color and content information with semantic knowledge, aligning the external outputs between the baseline and SPGRM. Furthermore, to transfer the semantic knowledge of the internal features, we utilize a Semantic Knowledge Alignment Module (SKAM) to fill the missing semantic contents with the complementary masks. Extensive experiments demonstrate that our framework significantly boosts HDR imaging quality for existing methods without altering the network architecture.
Tao Hu 0013, Longyao Wu, Wei Dong 0010, Peng Wu 0015, Jinqiu Sun, Xiaogang Xu 0002, Qingsen Yan, Yanning Zhang 0001
IEEE Trans. Image Process.7
2026 Ghost-Free HDR Imaging via Latent Low-Frequency Priors and Deformable Attention Alignment
abstract
Recovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, showing promising performance, particularly in achieving visually perceptible better results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. These priors are integrated into a carefully designed Dynamic HDR Reconstruction Network (DHRNet), which employs a regression-based approach to produce high-quality HDR images. Furthermore, we introduce the Attention-guided Deformable Alignment Module (ADAM) that utilizes correlation-driven feature matching to learn deformable receptive fields for self-attention, enabling efficient pre-alignment of LDR images by focusing on salient regions. Extensive experiments on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is $10\times $ faster than previous DM-based methods.
Tao Hu 0013, Qingsen Yan, Wei Dong 0010, Peng Wu 0015, Yuankai Qi, Weisi Lin, Yanning Zhang 0001
IEEE Trans. Image Process.2
2026 Deep Learning for Video Anomaly Detection: A Review
abstract
Video anomaly detection (VAD) aims to discover behaviors or events deviating from the normality in videos. As a long-standing task in the field of computer vision, VAD has witnessed much good progress. In the era of deep learning, with the explosion of architectures of continuously growing capability and capacity, a great variety of deep learning-based methods are constantly emerging for the VAD task, greatly improving the generalization ability of detection algorithms and broadening the application scenarios. Therefore, such a multitude of methods and a large body of literature make a comprehensive survey a pressing necessity. In this article, we present an extensive and comprehensive research review, covering the spectrum of five different categories, namely, semi-supervised, weakly supervised, fully supervised, unsupervised, and open-set supervised VAD, and we also delve into the latest VAD works based on pretrained large models and open-world learning, remedying the limitations of past reviews in terms of only focusing on semi-supervised VAD and small model-based methods. For the VAD task with different levels of supervision, we construct a well-organized taxonomy, profoundly discuss the characteristics of different types of methods, and show their performance comparisons. In addition, this review involves the public datasets, open-source codes, and evaluation metrics covering all the aforementioned VAD tasks. Finally, we provide several important research directions for the VAD community. Additional details of the survey are available on the project homepage: https://github.com/Roc-Ng/DeepVAD.
Peng Wu 0015, Chengyu Pan, Guansong Pang, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
abstract
Currently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglecting improvements to the vision encoder itself. In contrast, we propose Text Guided LLaVA (TG-LLaVA) in this paper, which optimizes VLMs by guiding the vision encoder with text, offering a new and orthogonal optimization direction. Specifically, inspired by the purpose-driven logic inherent in human behavior, we use learnable latent embeddings as a bridge to analyze textual instruction and add the analysis results to the vision encoder as guidance, refining it. Subsequently, another set of latent embeddings extracts additional detailed text-guided information from high-resolution local patches as auxiliary information. Finally, with the guidance of text, the vision encoder can extract text-related features, similar to how humans focus on the most relevant parts of an image when considering a question. This results in generating better answers. Experiments on various datasets validate the effectiveness of the proposed method. Remarkably, without the need for additional training data, our proposed method can bring more benefits to the baseline (LLaVA-1.5) compared with other concurrent methods. Furthermore, the proposed method consistently brings improvement in different settings.
Dawei Yan 0001, Hao Chen 0041, Weihua Luo, Wei Dong 0010, Qingsen Yan, Haokui Zhang, Chunhua Shen
AAAI8
2025 HVI: A New Color Space for Low-light Image Enhancement
abstract
Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable intensity. The former enforces small distances for red coordinates to remove the red artifacts, while the latter compresses the low-light regions to remove the black artifacts. To fully leverage the chromatic and intensity information, a novel Color and Intensity Decoupling Network (CIDNet) is further introduced to learn accurate photometric mapping function under different lighting conditions in the HVI space. Comprehensive results from benchmark and ablation experiments show that the proposed HVI color space with CIDNet outperforms the state-of-the-art methods on 10 datasets. The code is available at https://github.com/Fediory/HVI-CIDNet.
Qingsen Yan, Yixu Feng, Guansong Pang, Kangbiao Shi, Peng Wu 0015, Wei Dong 0010, Jinqiu Sun, Yanning Zhang 0001
CVPR1
2025 Learnable Feature Patches and Vectors for Boosting Low-Light Image Enhancement Without External Knowledge
Xiaogang Xu 0002, Jiafei Wu, Qingsen Yan, Jiequan Cui, Richang Hong, Bei Yu 0001
ICCV3
2025 Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning Strategy
abstract
A prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived through the multiplication structure of down-projection and up-projection matrices, exemplified by methods such as LoRA and Adapter. In this work, we observe an approximate orthogonality among any two row or column vectors within any weight matrix of the backbone parameters; however, this property is absent in the vectors of the down/up-projection matrices. Approximate orthogonality implies a reduction in the upper bound of the model's generalization error, signifying that the model possesses enhanced generalization capability. If the fine-tuned down/up-projection matrices were to exhibit this same property as the pre-trained backbone matrices, could the generalization capability of fine-tuned ViTs be further augmented? To address this question, we propose an Approximately Orthogonal Fine-Tuning (AOFT) strategy for representing the low-rank weight matrices. This strategy employs a single learnable vector to generate a set of approximately orthogonal vectors, which form the down/up-projection matrices, thereby aligning the properties of these matrices with those of the backbone. Extensive experimental results demonstrate that our method achieves competitive performance across a range of downstream image classification tasks, confirming the efficacy of the enhanced generalization capability embedded in the down/up-projection matrices.
Yiting Yang, Qingsen Yan, Haokui Zhang, Wei Dong 0010, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen
ICCV4
2025 Text-Visual Semantic Constrained AI-Generated Image Quality Assessment
Qingsen Yan, Haojian Huang, Peng Wu 0015, Haokui Zhang, Yanning Zhang 0001
ACM Multimedia2
2025 Click-level supervision for online action detection extended from SCOAD
Yuhan Mei, Xia Ling Lin, Genqing Bian, Qingsen Yan, Ghulam Mohiuddin, Chen Ai
Future Gener. Comput. Syst.6
2025 Enhancing the noise robustness of sparse-form patches for image denoising
Liping Qi, Yu Zhu 0004, Wei Sun 0036, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Knowl. Based Syst.6
2025 CLIP-guided continual novel class discovery
Qingsen Yan, Yiting Yang, Yutong Dai 0001, Katarzyna Wiltos, Marcin Wozniak, Wei Dong 0010, Yanning Zhang 0001
Knowl. Based Syst.1
2025 A multi-scale feature cross-dimensional interaction network for stereo image super-resolution
Yu Zhu 0004, Shengjun Peng, Axi Niu, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
Multim. Syst.5
2025 Modeling optical imaging pipeline and learning contrastive-based representation for hybrid-corrupted image restoration
Chenyuan Zhao, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Axi Niu, Yanning Zhang 0001
Multim. Syst.3
2025 Efficient Image Enhancement With a Diffusion-Based Frequency Prior
abstract
Due to the lack of appropriate priors, generating the content of dark regions remains a challenge in low-light image enhancement tasks. Currently, diffusion models employ robust image generation capabilities for enhancing low-light images. However, diffusion models require multiple iterations at the image feature level to generate details and content, which limits the speed. Moreover, the diffusion-based methods tend to generate unexpected artifacts in the degraded regions. To address these issues, we propose a Frequency Priors-guided Image Enhancement (FPIE) network, including a frequency prior generation network and an image restoration network. FPIE significantly accelerates inference by learning abstract prior with frequency domain constraints. Concretely, to learn compacted priors at the frequency domain, we introduce a joint training approach for the prior generation and restoration models to constrain the distribution of priors. Furthermore, to better utilize frequency-domain features for enhancing the network’s generation capabilities, a wavelet-based transformer block is introduced to produce intricate details and avoid the artifacts of the output. Extensive experimental results on the commonly used benchmarks demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images.
Qingsen Yan, Tao Hu 0013, Peng Wu 0015, Duwei Dai, Shuhang Gu, Wei Dong 0010, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 From Dynamic to Static: Stepwisely Generate HDR Image for Ghost Removal
abstract
Generating high-quality high dynamic range (HDR) images in dynamic scenes is particularly challenging due to the influence of large motion. Despite the effectiveness of existing deep learning methods, they still suffer from ghosting artifacts when saturation and motion coexist. Inspired by fusion on static scenes, we propose an inpainting and fusion strategy to enhance the quality of the generated HDR images. The proposed method consists of pseudo-static LDR generation and detail-guided HDR generation, which creates pseudo-static images and then generates ghost-free HDR images. Specifically, the pseudo-static LDR generation network utilizes semantic information to identify the motion regions, and employs a diffusion model-based inpainting approach to produce pseudo-static LDR images that closely resemble real scenes. In the detail-guided HDR generation network, we employ a detail enhancement module to refine diverse high-frequency features with detailed information extracted from pseudo-static LDR images, which effectively enhances the visual quality. Extensive experiments on four public datasets demonstrate the superiority of the proposed method, both quantitatively and qualitatively.
Qingsen Yan, Kangzhen Yang, Tao Hu 0013, Genggeng Chen, Kexin Dai, Peng Wu 0015, Wenqi Ren, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly Detection
abstract
The recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes full use of fine-grained associations between vision and language on the strength of CLIP and involves dual branch. One branch simply utilizes visual features for coarse-grained binary classification, while the other fully leverages the fine-grained language-image alignment. With the benefit of dual branch, VadCLIP achieves both coarse-grained and fine-grained video anomaly detection by transferring pre-trained knowledge from CLIP to WSVAD task. We conduct extensive experiments on two commonly-used benchmarks, demonstrating that VadCLIP achieves the best performance on both coarse-grained and fine-grained WSVAD, surpassing the state-of-the-art methods by a large margin. Specifically, VadCLIP achieves 84.51% AP and 88.02% AUC on XD-Violence and UCF-Crime, respectively. Code and features are released at https://github.com/nwpu-zxr/VadCLIP.
Peng Wu 0015, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001
AAAI5
2024 Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design Approach
abstract
Parameter-efficient fine-tuning for pre-trained Vision Transformers aims to adeptly tailor a model to downstream tasks by learning a minimal set of new adaptation parameters while preserving the frozen majority of pre-trained parameters. Striking a balance between retaining the generalizable representation capacity of the pre-trained model and acquiring task-specific features poses a key challenge. Currently, there is a lack of focus on guiding this delicate trade-off. In this study, we approach the problem from the perspective of Singular Value Decomposition (SVD) of pre-trained parameter matrices, providing insights into the tuning dynamics of existing methods. Building upon this understanding, we propose a Residual-based Low-Rank Rescaling (RLRR) fine-tuning strategy. This strategy not only enhances flexibility in parameter tuning but also ensures that new parameters do not deviate excessively from the pre-trained model through a residual design. Extensive experiments demonstrate that our method achieves competitive performance across various downstream image classification tasks, all while maintaining comparable new parameters. We believe this work takes a step forward in offering a unified perspective for interpreting existing methods and serves as motivation for the development of new approaches that move closer to effectively considering the crucial trade-off mentioned above. Our code is available at https://github.com/zstarN70/RLRR.git.
Wei Dong 0010, Bihui Chen, Dawei Yan 0001, Zhijun Lin, Qingsen Yan, Peng Wang 0023, Yang Yang 0002
CVPR6
2024 Generating Content for HDR Deghosting from Frequency View
abstract
Recovering ghost-free High Dynamic Range (HDR) images from multiple Low Dynamic Range (LDR) images becomes challenging when the LDR images exhibit saturation and significant motion. Recent Diffusion Models (DMs) have been introduced in HDR imaging field, demonstrating promising performance, particularly in achieving visually perceptible results compared to previous DNN-based methods. However, DMs require extensive iterations with large models to estimate entire images, resulting in inefficiency that hinders their practical application. To address this challenge, we propose the Low-Frequency aware Diffusion (LF-Diff) model for ghost-free HDR imaging. The key idea of LF-Diff is implementing the DMs in a highly compacted latent space and integrating it into a regression-based model to enhance the details of reconstructed images. Specifically, as low-frequency information is closely related to human visual perception we propose to utilize DMs to create compact low-frequency priors for the reconstruction process. In addition, to take full advantage of the above low-frequency priors, the Dynamic HDR Reconstruction Network (DHRNet) is carried out in a regression-based manner to obtain final HDR images. Extensive experiments conducted on synthetic and real-world benchmark datasets demonstrate that our LF-Diff performs favorably against several state-of-the-art methods and is 10x faster than previous DM-based methods.
Tao Hu 0013, Qingsen Yan, Yuankai Qi, Yanning Zhang 0001
CVPR2
2024 Multiple Object Tracking Based on Occlusion-Aware Embedding Consistency Learning
abstract
The Joint Detection and Embedding (JDE) framework has achieved remarkable progress for multiple object tracking. Existing methods often employ extracted embeddings to re-establish associations between new detections and previously disrupted tracks. However, the reliability of embeddings diminishes when the region of the occluded object frequently contains adjacent objects or clutters, especially in scenarios with severe occlusion. To alleviate this problem, we propose a novel multiple object tracking method based on visual embedding consistency, mainly including: 1) Occlusion Prediction Module (OPM) and 2) Occlusion-Aware Association Module (OAAM). The OPM predicts occlusion information for each true detection, facilitating the selection of valid samples for consistency learning of the track’s visual embedding. The OAAM leverages occlusion cues and visual embeddings to generate two separate embeddings for each track, guaranteeing consistency in both unoccluded and occluded detections. By integrating these two modules, our method is capable of addressing track interruptions caused by occlusion in online tracking scenarios. Extensive experimental results demonstrate that our approach achieves promising performance levels in both unoccluded and occluded tracking scenarios.
Yaoqi Hu, Axi Niu, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
ICASSP4
2024 Diffevent: Event Residual Diffusion for Image Deblurring
abstract
Traditional frame-based cameras inevitably suffer from non-uniform blur in real-world scenarios. Event cameras that record the intensity changes with high temporal resolution provide an effective solution for image deblurring. In this paper, we formulate the event-based image deblurring as an image generation problem by designing diffusion priors for the image and residual. Specifically, we propose an alternative diffusion sampling framework to jointly estimate clear and residual images to ensure the quality of the final result. In addition, to further enhance the subtle details, a pseudoinverse guidance module is leveraged to guide the prediction closer to the input with event data. Note that the proposed method can effectively handle the real unknown degradation without kernel estimation. The experiments on the benchmark event datasets demonstrate the effectiveness of our method.
Jiumei He, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
ICASSP3
2024 Efficient Content Reconstruction for High Dynamic Range Imaging
abstract
High Dynamic Range (HDR) images can be reconstructed from multiple Low Dynamic Range (LDR) images using existing deep neural network (DNN) techniques. Despite notable advancements, DNN-based methods still exhibit ghosting artifacts when handling LDR images with saturation and significant motion. Recent Diffusion models (DMs) have been introduced in HDR imaging, showcasing promising performance, especially in achieving visually perceptible results. However, DMs typically require numerous inference iterations to recover the clean image from Gaussian noise, demanding substantial computational resources. Additionally, DM only learns a probability distribution of the added noise in each step but neglects image space constraints on HDR images, limiting distortion-based metrics. To tackle these challenges, we propose an efficient network that integrates DM modules into existing regression-based models, providing reliable content reconstruction for HDR while avoiding limitations in distortion-based metrics.
Tao Hu 0013, Jiashuang He, Qingsen Yan
ICASSP4
2024 EiffHDR: An Efficient Network for Multi-Exposure High Dynamic Range Imaging
abstract
While recent progress in Multi-exposure HDR imaging is promising, the growing complexity of state-of-the-art (SOTA) methods poses challenges for their analysis and comparison. In this paper, we analyze the motivations and approaches behind previous SOTA works and introduce EiffHDR, an efficient Multi-exposure HDR imaging technique. In contrast to prior methods employing multiple branches spatial attention mechanisms, EiffHDR adopts a streamlined gating mechanism for information flow control at both spatial and channel levels, enabling implicit alignment. Subsequently, we process these features through proposed Efficient Merging Network, facilitating long-range correlations and multi-scale information perception, ultimately producing high-quality HDR images. Our experiments demonstrate that EiffHDR not only achieves outstanding performance but also significantly reduces computational complexity, making it a valuable contribution to the field.
Tao Hu 0013, Qingsen Yan
ICASSP4
2024 HL-HDR: Multi-Exposure High Dynamic Range Reconstruction with High-Low Frequency Decomposition
abstract
Generating high-quality High Dynamic Range (HDR) images in dynamic scenes is particularly challenging. Recent Transformer have been introduced in HDR imaging, demonstrating promising performance, particularly in scenarios involving large-scale motion compared to previous CNN-based methods. However, Transformer-based methods face hurdles capturing local details and come with high computational complexity, hindering further progress. In this paper, inspired by the distinct characteristics of high and low-frequency in image patterns, we propose a Frequency Decomposition Processing Block (FDPB) for ghost-free HDR imaging. In the image reconstruction process, FDPB decouples features into resolution-invariant high-frequency features and resolution-reduced low-frequency features to separately address local and global information. Specifically, considering the characteristics of different frequencies, for the high-frequency components, we design a Local Feature Extractor (LFE) based on CNN to extract local feature maps. Meanwhile, for the low-frequency components, we propose a Global Feature Extractor (GFE) that learns long-range dependencies through carefully designed Transformer modules. Importantly, the downscaled low-frequency features exploit Transformer’s remote learning capabilities while substantially reducing self-attention computational costs. By incorporating the FDPB as basic components, we further build a Low/High-Frequency Aware Network (HL-HDR), a hierarchical network to reconstruct high-quality ghost-free HDR images. Extensive experiments on four public datasets confirm the superior performance of the proposed method, both in terms of quantitative and qualitative evaluations.
Genggeng Chen, Tao Hu 0013, Kangzhen Yang, Qingsen Yan
IJCNN6
2024 Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
abstract
Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings for weakly supervised video anomaly detection and localization (WSVADL) based on pre-trained vision-language models (VLMs). Our proposed method employs a two-stream network structure, with one stream focusing on the temporal dimension and the other primarily on the spatial dimension. By leveraging the learned knowledge from pre-trained VLMs and incorporating natural motion priors from raw videos, our model learns prompt embeddings that are aligned with spatio-temporal regions of videos (e.g., patches of individual frames) for identify specific local regions of anomalies, enabling accurate video anomaly detection while mitigating the influence of background information. Without relying on detailed spatio-temporal annotations or auxiliary object detection/tracking, our method achieves state-of-the-art performance on three public benchmarks for the WSVADL task.
Peng Wu 0015, Xuerong Zhou, Guansong Pang, Zhiwei Yang 0013, Qingsen Yan, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia5
2024 Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
abstract
A common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck dimensionality being crucial for reducing the number of learnable parameters, as exemplified by prevalent methods like LoRA and Adapter. However, these low-rank strategies typically employ a fixed bottleneck dimensionality, which limits their flexibility in handling layer-wise variations. To address this limitation, we propose a novel PEFT approach inspired by Singular Value Decomposition (SVD) for representing the adaptation matrix. SVD decomposes a matrix into the product of a left unitary matrix, a diagonal matrix of scaling values, and a right unitary matrix. We utilize Householder transformations to construct orthogonal matrices that efficiently mimic the unitary matrices, requiring only a vector. The diagonal values are learned in a layer-wise manner, allowing them to flexibly capture the unique properties of each layer. This approach enables the generation of adaptation matrices with varying ranks across different layers, providing greater flexibility in adapting pre-trained models. Experiments on standard downstream vision tasks demonstrate that our method achieves promising fine-tuning performance.
Wei Dong 0010, Yiting Yang, Zhijun Lin, Qingsen Yan, Haokui Zhang, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen
NeurIPS6
2024 Take a prior from other tasks for severe blur removal
Yu Zhu 0004, Danna Xue, Qingsen Yan, Jinqiu Sun, Sung-Eui Yoon, Yanning Zhang 0001
Comput. Vis. Image Underst.4
2024 Real-time portrait image retouching extended from DualBLN
Genqing Bian, Chengzhe Lu, Sifei Wang, Ghulam Mohiuddin, Qingsen Yan
Expert Syst. Appl.7
2024 Dynamic center point learning for multiple object tracking under Severe occlusions
Yaoqi Hu, Axi Niu, Jinqiu Sun, Yu Zhu 0004, Qingsen Yan, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001
Knowl. Based Syst.5
2024 DPGS: Cross-cooperation guided dynamic points generation for scene text spotting
Wei Sun 0036, Qianzhou Wang, Xueling Chen, Qingsen Yan, Yanning Zhang 0001
Knowl. Based Syst.5
2024 I2U-Net: A dual-path U-Net with rich information interaction for medical image segmentation
Duwei Dai, Caixia Dong, Qingsen Yan, Yongheng Sun, Zongfang Li, Songhua Xu
Medical Image Anal.3
2024 GRAN: ghost residual attention network for single image super resolution
Axi Niu, Yu Zhu 0004, Jinqiu Sun, Qingsen Yan, Yanning Zhang 0001
Multim. Tools Appl.5
2024 KGSR: A kernel guided network for real-world blind super-resolution
Qingsen Yan, Axi Niu, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001
Pattern Recognit.1
2024 Uncertainty estimation in HDR imaging with Bayesian neural networks
Qingsen Yan, Haishen Wang, Yuhang Liu 0002, Wei Dong 0010, Marcin Wozniak, Yanning Zhang 0001
Pattern Recognit.1
2024 Toward High-Quality HDR Deghosting With Conditional Diffusion Models
abstract
High Dynamic Range (HDR) images can be recovered from several Low Dynamic Range (LDR) images by existing Deep Neural Networks (DNNs) techniques. Despite the remarkable progress, DNN-based methods still generate ghosting artifacts when LDR images have saturation and large motion, which hinders potential applications in real-world scenarios. To address this challenge, we formulate the HDR deghosting problem as an image generation that leverages LDR features as the diffusion model’s condition, consisting of the feature condition generator and the noise predictor. Feature condition generator employs attention and Domain Feature Alignment (DFA) layer to transform the intermediate features to avoid ghosting artifacts. With the learned features as conditions, the noise predictor leverages a stochastic iterative denoising process for diffusion models to generate an HDR image by steering the sampling process. Furthermore, to mitigate semantic confusion caused by the saturation problem of LDR images, we design a sliding window noise estimator to sample smooth noise in a patch-based manner. In addition, an image space loss is proposed to avoid the color distortion of the estimated HDR results. We empirically evaluate our model on benchmark datasets for HDR imaging. The results demonstrate that our approach achieves state-of-the-art performances and well generalization to real-world images.
Qingsen Yan, Tao Hu 0013, Hao Tang 0005, Yu Zhu 0004, Wei Dong 0010, Luc Van Gool, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 A Unified HDR Imaging Method with Pixel and Patch Level
abstract
Mapping Low Dynamic Range (LDR) images with different exposures to High Dynamic Range (HDR) remains nontrivial and challenging on dynamic scenes due to ghosting caused by object motion or camera jitting. With the success of Deep Neural Networks (DNNs), several DNNs-based methods have been proposed to alleviate ghosting, they cannot generate approving results when motion and saturation occur. To generate visually pleasing HDR images in various cases, we propose a hybrid HDR deghosting network, called HyHDRNet, to learn the complicated relationship between reference and non-reference images. The proposed HyHDRNet consists of a content alignment subnetwork and a Transformer-based fusion subnetwork. Specifically, to effectively avoid ghosting from the source, the content alignment subnetwork uses patch aggregation and ghost attention to integrate similar content from other non-reference images with patch level and suppress undesired components with pixel level. To achieve mutual guidance between patch-level and pixel-level, we leverage a gating module to sufficiently swap useful information both in ghosted and saturated regions. Furthermore, to obtain a high-quality HDR image, the Transformer-based fusion subnetwork uses a Residual Deformable Transformer Block (RDTB) to adaptively merge information for different exposed regions. We examined the proposed method on four widely used public HDR image deghosting datasets. Experiments demonstrate that HyHDRNet outperforms state-of-the-art methods both quantitatively and qualitatively, achieving appealing HDR visualization with unified textures and colors.
Qingsen Yan, Weiye Chen, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
CVPR1
2023 SMAE: Few-shot Learning for HDR Deghosting with Saturation-Aware Masked Autoencoders
abstract
Generating a high-quality High Dynamic Range (HDR) image from dynamic scenes has recently been extensively studied by exploiting Deep Neural Networks (DNNs). Most DNNs-based methods require a large amount of training data with ground truth, requiring tedious and time-consuming work. Few-shot HDR imaging aims to generate satisfactory images with limited data. However, it is difficult for modern DNNs to avoid overfitting when trained on only a few images. In this work, we propose a novel semi-supervised approach to realize few-shot HDR imaging via two stages of training, called SSHDR. Unlikely previous methods, directly recovering content and removing ghosts simultaneously, which is hard to achieve optimum, we first generate content of saturated regions with a self-supervised mechanism and then address ghosts via an iterative semi-supervised learning framework. Concretely, considering that saturated regions can be regarded as masking Low Dynamic Range (LDR) input regions, we design a Saturated Mask AutoEncoder (SMAE) to learn a robust feature representation and reconstruct a non-saturated HDR image. We also propose an adaptive pseudo-label selection strategy to pick high-quality HDR pseudo-labels in the second stage to avoid the effect of mislabeled samples. Experiments demonstrate that SSHDR outperforms state-of-the-art methods quantitatively and qualitatively within and across different datasets, achieving appealing HDR visualization with few labeled samples.
Qingsen Yan, Weiye Chen, Hao Tang 0005, Yu Zhu 0004, Jinqiu Sun, Luc Van Gool, Yanning Zhang 0001
CVPR1
2023 All-in-one Multi-degradation Image Restoration Network via Hierarchical Degradation Representation
abstract
The aim of image restoration is to recover high-quality images from distorted ones. However, current methods usually focus on a single task (e.g., denoising, deblurring or super-resolution) which cannot address the needs of real-world multi-task processing, especially on mobile devices. Thus, developing an all-in-one method that can restore images from various unknown distortions is a significant challenge. Previous works have employed contrastive learning to learn the degradation representation from observed images, but this often leads to representation drift caused by deficient positive and negative pairs. To address this issue, we propose a novel All-in-one Multi-degradation Image Restoration Network (AMIRNet) that can effectively capture and utilize accurate degradation representation for image restoration. AMIRNet learns a degradation representation for unknown degraded images by progressively constructing a tree structure through clustering, without any prior knowledge of degradation information. This tree-structured representation explicitly reflects the consistency and discrepancy of various distortions, providing a specific clue for image restoration. To further enhance the performance of the image restoration network and overcome domain gaps caused by unknown distortions, we design a feature transform block (FTB) that aligns domains and refines features with the guidance of the degradation representation. We conduct extensive experiments on multiple distorted datasets, demonstrating the effectiveness of our method and its advantages over state-of-the-art restoration methods both qualitatively and quantitatively.
Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
ACM Multimedia3
2023 An Internal-External Constrained Distillation Framework for Continual Semantic Segmentation
Qingsen Yan, Shengqiang Liu, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
PRCV (3)1
2023 Effectively fusing clinical knowledge and AI knowledge for reliable lung nodule diagnosis
Duwei Dai, Yongheng Sun, Caixia Dong, Qingsen Yan, Zongfang Li, Songhua Xu
Expert Syst. Appl.4
2023 From Distortion Manifold to Perceptual Quality: a Data Efficient Blind Image Quality Assessment Approach
Shaolin Su, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
Pattern Recognit.2
2023 3D Medical image segmentation using parallel transformers
Qingsen Yan, Shengqiang Liu, Songhua Xu, Caixia Dong, Zongfang Li, Qinfeng Shi, Yanning Zhang 0001, Duwei Dai
Pattern Recognit.1
2023 SharpFormer: Learning Local Feature Preserving Global Representations for Image Deblurring
abstract
The goal of dynamic scene deblurring is to remove the motion blur presented in a given image. To recover the details from the severe blurs, conventional convolutional neural networks (CNNs) based methods typically increase the number of convolution layers, kernel-size, or different scale images to enlarge the receptive field. However, these methods neglect the non-uniform nature of blurs, and cannot extract varied local and global information. Unlike the CNNs-based methods, we propose a Transformer-based model for image deblurring, named SharpFormer, that directly learns long-range dependencies via a novel Transformer module to overcome large blur variations. Transformer is good at learning global information but is poor at capturing local information. To overcome this issue, we design a novel Locality preserving Transformer (LTransformer) block to integrate sufficient local information into global features. In addition, to effectively apply LTransformer to the medium-resolution features, a hybrid block is introduced to capture intermediate mixed features. Furthermore, we use a dynamic convolution (DyConv) block, which aggregates multiple parallel convolution kernels to handle the non-uniform blur of inputs. We leverage a powerful two-stage attentive framework composed of the above blocks to learn the global, hybrid, and local features effectively. Extensive experiments on the GoPro and REDS datasets show that the proposed SharpFormer performs favourably against the state-of-the-art methods in blurred image restoration.
Qingsen Yan, Dong Gong, Zhen Zhang 0008, Yanning Zhang 0001, Qinfeng Shi
IEEE Trans. Image Process.1
2023 DHT-Net: Dynamic Hierarchical Transformer Network for Liver and Tumor Segmentation
abstract
Automatic segmentation of liver tumors is crucial to assist radiologists in clinical diagnosis. While various deep learningbased algorithms have been proposed, such as U-Net and its variants, the inability to explicitly model long-range dependencies in CNN limits the extraction of complex tumor features. Some researchers have applied Transformer-based 3D networks to analyze medical images. However, the previous methods focus on modeling the local information (eg. edge) or global information (eg. morphology) with fixed network weights. To learn and extract complex tumor features of varied tumor size, location, and morphology for more accurate segmentation, we propose a Dynamic Hierarchical Transformer Network, named DHT-Net. The DHT-Net mainly contains a Dynamic Hierarchical Transformer (DHTrans) structure and an Edge Aggregation Block (EAB). The DHTrans first automatically senses the tumor location by Dynamic Adaptive Convolution, which employs hierarchical operations with the different receptive field sizes to learn the features of various tumors, thus enhancing the semantic representation ability of tumor features. Then, to adequately capture the irregular morphological features in the tumor region, DHTrans aggregates global and local texture information in a complementary manner. In addition, we introduce the EAB to extract detailed edge features in the shallow fine-grained details of the network, which provides sharp boundaries of liver and tumor regions. We evaluate DHT-Net on two challenging public datasets, LiTS and 3DIRCADb. The proposed method has shown superior liver and tumor segmentation performance compared to several state-of-the-art 2D, 3D, and 2.5D hybrid models.
Longchang Xu, Kun Xie 0011, Jianfeng Song, Liang Chang 0003, Qingsen Yan
IEEE J. Biomed. Health Informatics7
2022 SCOAD: Single-Frame Click Supervision for Online Action Detection
Dawei Yan 0001, Wei Dong 0010, Qingsen Yan
ACCV (4)5
2022 DualBLN: Dual Branch LUT-Aware Network for Real-Time Image Retouching
Chengzhe Lu, Dawei Yan 0001, Wei Dong 0010, Qingsen Yan
ACCV (3)5
2022 Learning Bayesian Sparse Networks with Full Experience Replay for Continual Learning
abstract
Continual Learning (CL) methods aim to enable machine learning models to learn new tasks without catastrophic forgetting of those that have been previously mastered. Existing CL approaches often keep a buffer of previously-seen samples, perform knowledge distillation, or use regularization techniques towards this goal. Despite their performance, they still suffer from interference across tasks which leads to catastrophic forgetting. To ameliorate this problem, we propose to only activate and select sparse neurons for learning current and past tasks at any stage. More parameters space and model capacity can thus be reserved for the future tasks. This minimizes the interference between parameters for different tasks. To do so, we propose a Sparse neural Network for Continual Learning (SNCL), which employs variational Bayesian sparsity priors on the activations of the neurons in all layers. Full Experience Replay (FER) provides effective supervision in learning the sparse activations of the neurons in different layers. A loss-aware reservoir-sampling strategy is developed to maintain the memory buffer. The proposed method is agnostic as to the network structures and the task boundaries. Experiments on different datasets show that SNCL achieves state-of-the-art result for mitigating forgetting.
Qingsen Yan, Dong Gong, Yuhang Liu 0002, Anton van den Hengel, Qinfeng Shi
CVPR1
2022 Exploring and Evaluating Image Restoration Potential in Dynamic Scenes
abstract
In dynamic scenes, images often suffer from dynamic blur due to superposition of motions or low signal-noise ratio resulted from quick shutter speed when avoiding motions. Recovering sharp and clean results from the captured images heavily depends on the ability of restoration methods and the quality of the input. Although existing research on image restoration focuses on developing models for obtaining better restored results, fewer have studied to evaluate how and which input image leads to superior restored quality. In this paper, to better study an image's potential value that can be explored for restoration, we propose a novel concept, referring to image restoration potential (IRP). Specifically, We first establish a dynamic scene imaging dataset containing composite distortions and applied image restoration processes to validate the rationality of the existence to IRP. Based on this dataset, we investigate several properties of IRP and propose a novel deep model to accurately predict IRP values. By gradually distilling and selective fusing the degradation features, the proposed model shows its superiority in IRP prediction. Thanks to the proposed model, we are then able to validate how various image restoration related applications are benefited from IRP prediction. We show the potential usages of IRP as a filtering principle to select valuable frames, an auxiliary guidance to improve restoration models, and also an indicator to optimize camera settings for capturing better images under dynamic scenarios.
Shaolin Su, Yu Zhu 0004, Qingsen Yan, Jinqiu Sun, Yanning Zhang 0001
CVPR4
2022 Dual-Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001
Int. J. Comput. Vis.1
2022 Ms RED: A novel multi-scale residual encoding and decoding network for skin lesion segmentation
Duwei Dai, Caixia Dong, Songhua Xu, Qingsen Yan, Zongfang Li, Nana Luo
Medical Image Anal.4
2022 High dynamic range imaging via gradient-aware context aggregation network
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Jinqiu Sun, Yu Zhu 0004, Yanning Zhang 0001
Pattern Recognit.1
2021 A Comprehensive CT Dataset for Liver Computer Assisted Diagnosis
Qingsen Yan, Bo Wang 0011, Dong Gong, Dingwen Zhang, Yang Yang 0009, Zheng You, Yanning Zhang 0001, Qinfeng Shi
BMVC1
2021 Towards accurate HDR imaging with learning generator constraints
Qingsen Yan, Bo Wang 0011, Lei Zhang 0054, Zheng You, Qinfeng Shi, Yanning Zhang 0001
Neurocomputing1
2021 Non-uniform motion deblurring with blurry component divided guidance
Wei Sun 0036, Qingsen Yan, Axi Niu, Rui Li 0013, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
Pattern Recognit.3
2021 COVID-19 Chest CT Image Segmentation Network by Multi-Scale Fusion and Enhancement Operations
abstract
A novel coronavirus disease 2019 (COVID-19) was detected and has spread rapidly across various countries around the world since the end of the year 2019. Computed Tomography (CT) images have been used as a crucial alternative to the time-consuming RT-PCR test. However, pure manual segmentation of CT images faces a serious challenge with the increase of suspected cases, resulting in urgent requirements for accurate and automatic segmentation of COVID-19 infections. Unfortunately, since the imaging characteristics of the COVID-19 infection are diverse and similar to the backgrounds, existing medical image segmentation methods cannot achieve satisfactory performance. In this article, we try to establish a new deep convolutional neural network tailored for segmenting the chest CT images with COVID-19 infections. We first maintain a large and new chest CT image dataset consisting of 165,667 annotated chest CT images from 861 patients with confirmed COVID-19. Inspired by the observation that the boundary of the infected lung can be enhanced by adjusting the global intensity, in the proposed deep CNN, we introduce a feature variation block which adaptively adjusts the global properties of the features for segmenting COVID-19 infection. The proposed FV block can enhance the capability of feature representation effectively and adaptively for diverse cases. We fuse features at different scales by proposing Progressive Atrous Spatial Pyramid Pooling to handle the sophisticated infection areas with diverse appearance and shapes. The proposed method achieves state-of-the-art performance. Dice similarity coefficients are 0.987 and 0.726 for lung and COVID-19 segmentation, respectively. We conducted experiments on the data collected in China and Germany and show that the proposed deep CNN can produce impressive performance effectively. The proposed network enhances the segmentation ability of the COVID-19 infection, makes the connection with other techniques and contributes to the development of remedying COVID-19 infection.
Qingsen Yan, Bo Wang 0011, Dong Gong, Chuan Luo 0003, Jianhu Shen, Jingyang Ai, Qinfeng Shi, Yanning Zhang 0001, Liang Zhang 0010, Zheng You
IEEE Trans. Big Data1
2021 Attention-Guided Deep Neural Network With Multi-Scale Feature Fusion for Liver Vessel Segmentation
abstract
Liver vessel segmentation is fast becoming a key instrument in the diagnosis and surgical planning of liver diseases. In clinical practice, liver vessels are normally manual annotated by clinicians on each slice of CT images, which is extremely laborious. Several deep learning methods exist for liver vessel segmentation, however, promoting the performance of segmentation remains a major challenge due to the large variations and complex structure of liver vessels. Previous methods mainly using existing UNet architecture, but not all features of the encoder are useful for segmentation and some even cause interferences. To overcome this problem, we propose a novel deep neural network for liver vessel segmentation, called LVSNet, which employs special designs to obtain the accurate structure of the liver vessel. Specifically, we design Attention-Guided Concatenation (AGC) module to adaptively select the useful context features from low-level features guided by high-level features. The proposed AGC module focuses on capturing rich complemented information to obtain more details. In addition, we introduce an innovative multi-scale fusion block by constructing hierarchical residual-like connections within one single residual block, which is of great importance for effectively linking the local blood vessel fragments together. Furthermore, we construct a new dataset containing 40 thin thickness cases (0.625 mm) which consist of CT volumes and annotated vessels. To evaluate the effectiveness of the method with minor vessels, we also propose an automatic stratification method to split major and minor liver vessels. Extensive experimental results demonstrate that the proposed LVSNet outperforms previous methods on liver vessel segmentation datasets. Additionally, we conduct a series of ablation studies that comprehensively support the superiority of the underlying concepts.
Qingsen Yan, Bo Wang 0011, Wei Zhang 0098, Chuan Luo 0003, Wei Xu 0005, Zhengqing Xu, Yanning Zhang 0001, Qinfeng Shi, Liang Zhang 0010, Zheng You
IEEE J. Biomed. Health Informatics1
2020 Blindly Assess Image Quality in the Wild Guided by a Self-Adaptive Hyper Network
abstract
Blind image quality assessment (BIQA) for authentically distorted images has always been a challenging problem, since images captured in the wild include varies contents and diverse types of distortions. The vast majority of prior BIQA methods focus on how to predict synthetic image quality, but fail when applied to real-world distorted images. To deal with the challenge, we propose a self-adaptive hyper network architecture to blind assess image quality in the wild. We separate the IQA procedure into three stages including content understanding, perception rule learning and quality predicting. After extracting image semantics, perception rule is established adaptively by a hyper network, and then adopted by a quality prediction network. In our model, image quality can be estimated in a self-adaptive manner, thus generalizes well on diverse images captured in the wild. Experimental results verify that our approach not only outperforms the state-of-the-art methods on challenging authentic image databases but also achieves competing performances on synthetic image databases, though it is not explicitly designed for the synthetic task.
Shaolin Su, Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
CVPR2
2020 Attention-Based Network For Low-Light Image Enhancement
abstract
The captured images under low-light conditions often suffer insufficient brightness and notorious noise. Hence, low-light image enhancement is a key challenging task in computer vision. A variety of methods have been proposed for this task, but these methods often failed in an extreme low-light environment and amplified the underlying noise in the input image. To address such a difficult problem, this paper presents a novel attention-based neural network to generate high-quality enhanced low-light images from the raw sensor data. Specifically, we first employ attention strategy (i.e. spatial attention and channel attention modules) to suppress undesired chromatic aberration and noise. The spatial attention module focuses on denoising by taking advantage of the non-local correlation in the image. The channel attention module guides the network to refine redundant colour features. Furthermore, we propose a new pooling layer, called inverted shuffle layer, which adaptively selects useful information from previous features. Extensive experiments demonstrate the superiority of the proposed network in terms of suppressing the chromatic aberration and noise artifacts in enhancement, especially when the low-light image has severe noise.
Qingsen Yan, Yu Zhu 0004, Xianjun Li, Jinqiu Sun, Yanning Zhang 0001
ICME2
2020 A Benchmark Dataset for Segmenting Liver, Vasculature and Lesions from Large-scale Computed Tomography Data
abstract
How to build a high-performance liver-related computer assisted diagnosis system is an open question of great interest. However, the performance of the state-of-art algorithm is always limited by the amount of data and the quality of the label. To address this problem, we propose the biggest treatment-oriented liver cancer dataset for liver surgery and treatment planning. This dataset provides 216 cases (total about 268K frames) scanned images in contrast-enhanced computed tomography (CT). We labeled all the CT images with the liver, liver vasculature, and liver tumor segmentation ground truth for train and tune segmentation algorithms in advance. Based on that, we evaluate several recent and state-of-the-art segmentation algorithms, including 7 deep learning methods, on CT sequences. All results are compared to reference segmentations five error metrics that highlight different aspects of segmentation accuracy. In general, compared with previous datasets, our dataset is really a challenging dataset. To our knowledge, the proposed dataset and benchmark allow for the first time systematic exploration of such issues, and will be made available to allow for further research in this field.
Bo Wang 0011, Qingsen Yan, Zhengqing Xu, Jingyang Ai, Wei Xu 0005, Liang Zhang 0010, Zheng You
ICPR2
2020 Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
abstract
Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixels of a target object in all frames, given the annotation mask in the first frame. Even when such annotation is available, this remains a challenging problem because of the changing appearance and shape of the object over time. In this paper, we tackle this task by formulating it as a meta-learning problem, where the base learner grasping the semantic scene understanding for a general type of objects, and the meta learner quickly adapting the appearance of the target object with a few examples. Our proposed meta-learning method uses a closed form optimizer, the so-called "ridge regression", which has been shown to be conducive for fast and better training convergence. Moreover, we propose a mechanism, named "block splitting", to further speed up the training process as well as to reduce the number of learning parameters. In comparison with the state-of-the art methods, our proposed framework achieves significant boost up in processing speed, while having highly comparable performance compared to the best performing methods on the widely used datasets. Video demo can be found here1.
Yu Liu 0029, Lingqiao Liu, Haokui Zhang, Seyed Hamid Rezatofighi, Qingsen Yan, Ian D. Reid 0001
IROS5
2020 Ghost Removal via Channel Attention in Exposure Fusion
Qingsen Yan, Bo Wang 0011, Xianjun Li, Qinfeng Shi, Zheng You, Yu Zhu 0004, Jinqiu Sun, Yanning Zhang 0001
Comput. Vis. Image Underst.1
2020 Deep HDR Imaging via A Non-Local Network
abstract
One of the most challenging problems in reconstructing a high dynamic range (HDR) image from multiple low dynamic range (LDR) inputs is the ghosting artifacts caused by the object motion across different inputs. When the object motion is slight, most existing methods can well suppress the ghosting artifacts through aligning LDR inputs based on optical flow or detecting anomalies among them. However, they often fail to produce satisfactory results in practice, since the real object motion can be very large. In this study, we present a novel deep framework, termed NHDRRnet, which adopts an alternative direction and attempts to remove ghosting artifacts by exploiting the non-local correlation in inputs. In NHDRRnet, we first adopt an Unet architecture to fuse all inputs and map the fusion results into a low-dimensional deep feature space. Then, we feed the resultant features into a novel global non-local module which reconstructs each pixel by weighted averaging all the other pixels using the weights determined by their correspondences. By doing this, the proposed NHDRRnet is able to adaptively select the useful information (e.g., which are not corrupted by large motions or adverse lighting conditions) in the whole deep feature space to accurately reconstruct each pixel. In addition, we also incorporate a triple-pass residual module to capture more powerful local features, which proves to be effective in further boosting the performance. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed NDHRnet in terms of suppressing the ghosting artifacts in HDR reconstruction, especially when the objects have large motions.
Qingsen Yan, Lei Zhang 0054, Yu Liu 0029, Yu Zhu 0004, Jinqiu Sun, Qinfeng Shi, Yanning Zhang 0001
IEEE Trans. Image Process.1
2019 Attention-Guided Network for Ghost-Free High Dynamic Range Imaging
abstract
Ghosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results. A very recent work tries to bypass optical flows via a deep network with skip-connections, however, which still suffers from ghosting artifacts for severe movement. To avoid the ghosting from the source, we propose a novel attention-guided end-to-end deep neural network (AHDRNet) to produce high-quality ghost-free HDR images. Unlike previous methods directly stacking the LDR images or features for merging, we use attention modules to guide the merging according to the reference image. The attention modules automatically suppress undesired components caused by misalignments and saturation and enhance desirable fine details in the non-reference images. In addition to the attention model, we use dilated residual dense block (DRDB) to make full use of the hierarchical features and increase the receptive field for hallucinating the missing details. The proposed AHDRNet is a non-flow-based method, which can also avoid the artifacts generated by optical-flow estimation error. Experiments on different datasets show that the proposed AHDRNet can achieve state-of-the-art quantitative and qualitative results.
Qingsen Yan, Dong Gong, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Ian D. Reid 0001, Yanning Zhang 0001
CVPR1
2019 Multi-Scale Dense Networks for Deep High Dynamic Range Imaging
abstract
Generating a high dynamic range (HDR) image from a set of sequential exposures is a challenging task for dynamic scenes. The most common approaches are aligning the input images to a reference image before merging them into an HDR image, but artifacts often appear in cases of large scene motion. The state-of-the-art method using deep learning can solve this problem effectively. In this paper, we propose a novel deep convolutional neural network to generate HDR, which attempts to produce more vivid images. The key idea of our method is using the coarse-to-fine scheme to gradually reconstruct the HDR image with the multi-scale architecture and residual network. By learning the relative changes of inputs and ground truth, our method can produce not only artificial free image but also restore missing information. Furthermore, we compare to existing methods for HDR reconstruction, and show high-quality results from a set of low dynamic range (LDR) images. We evaluate the results in qualitative and quantitative experiments, our method consistently produces excellent results than existing state-of-the-art approaches in challenging scenes.
Qingsen Yan, Dong Gong, Qinfeng Shi, Jinqiu Sun, Ian D. Reid 0001, Yanning Zhang 0001
WACV1
2019 Robust artifact-free high dynamic range imaging of dynamic scenes
Qingsen Yan, Yu Zhu 0004, Yanning Zhang 0001
Multim. Tools Appl.1
2019 Enhancing image visuality by multi-exposure fusion
Qingsen Yan, Yu Zhu 0004, Jinqiu Sun, Lei Zhang 0054, Yanning Zhang 0001
Pattern Recognit. Lett.1
2019 Two-Stream Convolutional Networks for Blind Image Quality Assessment
abstract
Traditional image quality assessment (IQA) methods do not perform robustly due to the shallow hand-designed features. It has been demonstrated that deep neural network can learn more effective features than ever. In this paper, we describe a new deep neural network to predict the image quality accurately without relying on the reference image. To learn more effective feature representations for non-reference IQA, we propose a two-stream convolution network that includes two subcomponents for image and gradient image. The motivation for this design is using a two-stream scheme to capture different-level information of inputs and easing the difficulty of extracting features from one steam. The gradient stream focuses on extracting structure features in details, and the image stream pays more attention to the information in intensity. In addition, to consider the locally non-uniform distribution of distortion in images, we add a region-based fully convolutional layer for using the information around the center of the input image patch. The final score of the overall image is calculated by averaging of the patch scores. The proposed network performs in an end-to-end manner in both the training and testing phases. The experimental results on a series of benchmark datasets, e.g., LIVE, CISQ, IVC, TID2013, and Waterloo Exploration Database, show that the proposed algorithm outperforms the state-of-the-art methods, which verifies the effectiveness of our network architecture.
Qingsen Yan, Dong Gong, Yanning Zhang 0001
IEEE Trans. Image Process.1
2018 Blind Image Quality Assessment via Deep Recursive Convolutional Network with Skip Connection
Qingsen Yan, Jinqiu Sun, Shaolin Su, Yu Zhu 0004, Haisen Li, Yanning Zhang 0001
PRCV (2)1
2017 High dynamic range imaging by sparse representation
Qingsen Yan, Jinqiu Sun, Haisen Li, Yu Zhu 0004, Yanning Zhang 0001
Neurocomputing1
2014 Kernel sparse tracking with compressive sensing
abstract
Online tracking is a challenging task to develop effective and efficient models to account for appearance change. However, most tracking algorithms only consider the holistic or local information and do not make full use of the appearance information. In this study, a novel tracking algorithm with sparse representation is proposed and the online classifier is learned to discriminate the target from the background. To reduce visual drift problem which is encountered in object tracking, a two‐stage sparse representation method is proposed. The holistic information is used to estimate the initial tracking position, and the local information is used to determine the final tracking position. To improve the performance of the classifier and robustness of the algorithm, the kernel function is applied on the sparse representation. Moreover, the dimension of the target is reduced via compressive sensing. Besides, a simple and effective method for dictionary update is proposed. Both qualitative and quantitative evaluations on challenging image sequences demonstrate that the proposed algorithm performs favourably against several state‐of‐the‐art algorithms.
Qingsen Yan, Linsheng Li
IET Comput. Vis.1