Liyuan Pan

dblp:199/2150 · DBLP profile ↗
← Back
45ranked-venue papers
12as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 8 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-author · 21 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Oligodendrocyte-Driven Spiking Neural Model
abstract
The spiking neuron model (SNM) mimics the processing paradigm of synaptic and membrane potentials in the cerebral cortex. However, existing SNMs are limited by two issues. First, they lack spike diversity. Although a spiking neuron perceives temporally varying input currents, SNMs only use identical synaptic weights for regulation. Second, they are insensitive to weak spikes. The potential accumulation in SNMs is solely driven by external inputs, ignoring the internal dynamics of potential. Oligodendrocytes, a recent revelation in neuroscience, enhance neural signaling by forming bidirectional communication. This offers the potential to alleviate the aforementioned issues. In this paper, we first propose the mechanism of the oligodendrocyte-spiking neuron (Oli-N) model. Subsequently, using the Oli-N model, we develop our Oli-inspired spiking neural network (Oli-SNN), which broadens the diversity of spike representations and enhances neurons' firing precision through improved sparse coding to enhance weak spikes. Experiments show that our Oli-SNN achieves state-of-the-art performance in the classification task on both static and neuromorphic datasets.
Mengqiao Han, Liyuan Pan, Xiabi Liu, Hongming Zhang 0002
AAAI2
2026 DCRR++: Unsupervised Reflection Removal and Novel View Synthesis via Dual-Pixel Guided 3D Gaussian Splatting
Kailong Yu, Mina Han, Liyuan Pan, Liu Liu 0009, Miaomiao Liu 0001, Wei Liang 0008
Int. J. Comput. Vis.3
2026 A mutual information-based framework for generalized image fusion via common-unique decoupling
Liyuan Pan, Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002
Knowl. Based Syst.2
2026 D2-DETR:DETR With Dual-Domain frequency-spatial modeling for unmanned aerial vehicle imagery object detection
Xuanming Liu, Huanxin Zou, Jun Li 0020, Liyuan Pan, Shitian He, Jiangshan Li, Wanyu Chen
Knowl. Based Syst.4
2026 SZCo: Self-supervised zero-shot co-segmentation with region-text alignment learning
Xin Duan, Yan Yang 0011, Liyuan Pan, Xiabi Liu, Mingyang Gong
Pattern Recognit.3
2025 EZSR: Event-based Zero-Shot Recognition
abstract
This paper studies zero-shot object recognition using event camera data. Guided by CLIP, which is pre-trained on RGB images, existing approaches achieve zero-shot object recognition by optimizing embedding similarities between event data and RGB images respectively encoded by an event encoder and the CLIP image encoder. Alternatively, several methods learn RGB frame reconstructions from event data for the CLIP image encoder. However, they often result in suboptimal zero-shot performance.This study develops an event encoder without relying on additional reconstruction networks. We theoretically analyze the performance bottlenecks of previous approaches: the embedding optimization objectives are prone to suffer from the spatial sparsity of event data, causing semantic misalignments between the learned event embedding space and the CLIP text embedding space. To mitigate the issue, we explore a scalar-wise modulation strategy. Furthermore, to scale up the number of events and RGB data pairs for training, we also study a pipeline for synthesizing event data from static RGB images in mass.Experimentally, we demonstrate an attractive scaling property in the number of parameters and synthesized data. We achieve superior zero-shot object recognition performance on extensive standard benchmark datasets, even compared with past supervised learning approaches. For example, our model with a ViT/B-16 backbone achieves 47.84% zero-shot accuracy on the N-ImageNet dataset.
Yan Yang 0011, Liyuan Pan, Dongxu Li 0003, Liu Liu 0009
CVPR2
2025 DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
abstract
The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learning on current datasets, we observe that redundancy occurs in both repeated and answer-irrelevant frames, and the corresponding frames vary with different questions. This suggests the possibility of adopting dynamic encoding to balance detailed video information preservation with token budget reduction. To this end, we propose a dynamic cooperative network, DynFocus, for memory-efficient video encoding in this paper. Specifically, i) a Dynamic Event Prototype Estimation (DPE) module to dynamically select meaningful frames for question answering; (ii) a Compact Cooperative Encoding (CCE) module that encodes meaningful frames with detailed visual appearance and the remaining frames with sketchy perception separately. We evaluate our method on five publicly available benchmarks, and experimental results consistently demonstrate that our method achieves competitive performance. Code is available at https://github.com/Simon98-AI/ DynFocus
Qingpei Guo, Liyuan Pan, Liu Liu 0009, Yu Guan 0001, Ming Yang 0007
CVPR3
2025 GliaNet: Adaptive Neural Network Structure Learning with Glia-Driven
abstract
Neural networks derived from the M-P model have excelled in various visual tasks. However, as a simplified simulation version of the brain neural pathway, their structures are locked during training, causing over-fitting and over-parameterization. Although recent models have begun using the biomimetic concept and empirical pruning, they still result in irrational pruning, potentially affecting the accuracy of the model. In this paper, we introduce the Glia unit, composed of oligodendrocytes (Oli) and astrocytes (Ast), to emulate the exact workflow of the mammalian brain, thereby enhancing the biological plausibility of neural functions. Oli selects neurons involved in signal transmission during neural communication and, together with Ast, adaptively optimizes the neural structure. Specifically, we first construct the artificial Glia-Neuron (G-N) model, which is formulated at the instance, group, and interaction levels with adaptive and collaborative mechanisms. Then, we construct GliaNet based on our G-N model, whose structure and connections can be continuously optimized during training. Experiments show that our GliaNet advances state-of-the-art on multiple tasks while significantly reducing its parameters.
Mengqiao Han, Liyuan Pan, Xiabi Liu
CVPR2
2025 Task-Specific Gradient Adaptation for Few-Shot One-Class Classification
abstract
Optimization-based meta-learning methods for few-shot one-class classification (FS-OCC) aim to fine-tune a meta-trained model to classify the positive and negative samples using only a few positive samples by adaptation. However, recent approaches primarily focus on adjusting existing meta-learning algorithms for FS-OCC, while overlooking issues stemming from the misalignment between the cross-entropy loss and OCC tasks during adaptation. This misalignment, combined with the limited availability of one-class samples and the restricted diversity of task-specific adaptation, can significantly exacerbate the adverse effects of gradient instability and generalization. To address these challenges, we propose a novel Task-Specific Gradient Adaptation (TSGA) for FS-OCC. Without extra supervision, TSGA learns to generate appropriate, stable gradients by leveraging label prediction and feature representation details of one-class samples and refines the adaptation process by recalibrating task-specific gradients and regularization terms. We evaluate TSGA on three challenging datasets and a real-world CNC Milling Machine application and demonstrate consistent improvements over baseline methods. Furthermore, we illustrate the critical impact of gradient instability and task-agnostic adaptation. Notably, TSGA achieves state-of-the-art results by effectively addressing these issues.
Xiabi Liu, Liyuan Pan, Yuchen Ren 0003
CVPR3
2025 From Subtle Hints to Grand Expressions - Mastering Fine-grained Emotions with Dynamic Multimodal Analysis
abstract
Multimodal Emotion Analysis (MEA) plays a crucial role in extracting and understanding emotional insights from diverse data sources, including text, video, and audio. However, existing methods may overlook the key issue that multimodal components exhibit asynchronism temporally and they obtain insufficient representation of fine-grained emotional expressions. In light of this, we propose a unified emotion reasoning model, EmoChat, which enhances multimodal emotion analysis by dynamically generating emotion-related tokens and fine-grained expression information through facial action modeling. To incorporate expression semantics, we design the AU Agent, a lightweight facial expression extractor, to provide LLMs with fine-grained facial knowledge for reasoning. In addition, we propose the Correlation Aggregator to alleviate the correlation differences between acoustic features and textual content. Therefore, our method decouples both the audio and vision modalities, allowing for efficient token-level emotion cues mining in misaligned multimodal input, while maintaining semantic consistency across different languages. Experiments on public benchmark datasets have demonstrated the superiority of our proposed EmoChat over the state-of-the-art methods.
Qinfu Xu, Liyuan Pan, Shaozu Yuan, Chunlei Wu
ACM Multimedia2
2025 Enhanced Dual-Pixel Image Reflection Removal via Gaussian Splatting
abstract
Image de-reflection is a critical task in computer vision. Existing methods for de-reflection using monocular cameras face challenges due to the lack of depth cues to separate the transmission and reflection layers, particularly under strong illumination or multi-layer reflection scenarios. Although recent advances, such as 3D Gaussian Splatting (3DGS), utilize novel view-synthesis capabilities to separate transmitted and reflected layers, they still encounter difficulties in practice with monocular images. In this paper, we simplify the de-reflection task by combining dual-pixel (DP) technology with 3DGS, forming the first unsupervised de-reflection framework. Specifically, we propose the Dual-View Coordinated Reflection Removal (DCRR) Framework, which integrates depth cues from DP sensors with the rendering capabilities of 3DGS. The DCRR utilizes a dual-view approach that estimates the image transmission layer and opacity via differentiable rasterization with 3DGS and reconstructs the reflection layer through a lightweight multi-layer perceptron. We then present the Dual-Pixel-Driven Reflection Gaussian Pruning (DPRGP) to refine the separation process. By using the physical properties of DP sensors, DCRR achieves significant accuracy improvements in complex reflection scenarios. A real-world DP-based dataset that includes paired reflection/reflection-free images has been collected. Extensive experiments demonstrate our competitive performance compared to state-of-the-art de-reflection approaches.
Kailong Yu, Liyuan Pan, Liu Liu 0009, Wei Liang 0008
ACM Multimedia2
2025 User-Instructed Disparity-aware Defocus Control
abstract
In photography, an All-in-Focus (AiF) image may not always effectively convey the creator’s intent. Professional photographers manipulate Depth of Field (DoF) to control which regions appear sharp or blurred, achieving compelling artistic effects. For general users, the ability to flexibly adjust DoF enhances creative expression and image quality. In this paper, we propose UiD, a User-Instructed DoF control framework, that allows users to specify refocusing regions using text, box, or point prompts, and our UiD automatically simulates in-focus and out-of-focus (OoF) regions in the given images. However, controlling defocus blur in a single-lens camera remains challenging due to the difficulty in estimating depth-aware aberrations and the suboptimal quality of reconstructed AiF images. To address this, we leverage dual-pixel (DP) sensors, commonly found in DSLR-style and mobile cameras. DP sensors provide a small-baseline stereo pair in a single snapshot, enabling depth-aware aberration estimation. Our approach first establishes an invertible mapping between OoF and AiF images to learn spatially varying defocus kernels and the disparity features. These depth-aware kernels enable bidirectional image transformation—deblurring out-of-focus (OoF) images into all-in-focus (AiF) representations, and conversely reblurring AiF images into OoF outputs—by seamlessly switching between the kernel and its inverse form. These depth-aware kernels enable both deblurring of OoF images into AiF representations and reblurring AiF images into OoF representations by flexibly switching its original form to its inverse one. For user-guided refocusing, we first generate masks based on user prompts using SAM, which modulates disparity features in closed form, allowing dynamic kernel re-estimation for reblurring. This achieves user-controlled refocusing effects. Extensive experiments on both common datasets and the self-collected dataset demonstrate that UiD offers superior flexibility and quality in DoF manipulation imaging.
Liyuan Pan
NeurIPS4
2025 Storyboard-guided Alignment for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition can be formulated as a video–text matching problem. Previous approaches primarily rely on global video semantics to consolidate video embeddings, often leading to misaligned video–text pairs due to inaccurate atomic-level action understanding. This inaccuracy arises due to i) videos with distinct global semantics may share similar atomic actions or visual appearances, and ii) atomic actions can be momentary, gradual, or not directly aligned with overarching video semantics. Inspired by storyboarding, where a script is segmented into individual shots, we propose a multi-granularity framework, SFAR. SFAR generates fine-grained descriptions of common atomic actions for each global semantic using a large language model. Unlike existing works that refine global semantics with auxiliary video frames, SFAR introduces a filtering metric to ensure correspondence between the descriptions and the global semantics, eliminating the need for direct video involvement and thereby enabling more nuanced recognition of subtle actions. By leveraging both global semantics and fine-grained descriptions, our SFAR effectively identifies prominent frames within videos, thereby improving the accuracy of embedding aggregation. Extensive experiments on various video action recognition datasets demonstrate the competitive performance of our SFAR in supervised, few-shot, and zero-shot settings.
Enqi Liu, Liyuan Pan, Yan Yang 0011, Yiran Zhong, Zhijing Wu 0001, Xinxiao Wu, Liu Liu 0009
NeurIPS2
2025 MambaRSIS: Context-aware multi-scale feature aggregation with selective state space model for remote sensing instance segmentation
abstract
Remote sensing instance segmentation aims to detect and assign pixel-level labels to each instance in remote sensing images, which holds critical engineering significance for both civil and military applications. While existing domain-specific methods have made progress, they still struggle with three persistent challenges: ineffective context modeling in cluttered backgrounds, information loss during multi-scale feature fusion, and blurred boundaries for densely clustered small objects. To address these limitations, we propose a novel remote sensing instance segmentation framework with three artificial intelligence (AI) methodological innovations, which comprises: a Context Perception Module (CPM) for context modeling, a Context Guided Multi-Scale Feature Aggregation (CGFA) method for multi-scale feature fusion, and a Multi-Path Region Proposal Extractor (MPRPE) with boundary-refined segmentation. The CPM leverages the selective state space model (Mamba) to capture long-range contextual information, effectively addressing the issue of cluttered backgrounds in remote sensing images. The CGFA replaces standard feature pyramid network architecture which is limited by direct summation or concatenation, preserving fine-grained spatial details with context guidance. The MPRPE and boundary-aware segmentation head mitigate the challenges of missed detection of small objects and blurred edge predictions, which arise from the clustered distribution of small objects and semantic ambiguity. Extensive experiments on the challenging iSAID and NWPU VHR-10 datasets validate the proposed method’s consistent improvements across metrics while demonstrating its practical engineering impact on remote sensing interpretation systems.
Liyuan Pan, Huanxin Zou, Hao Chen 0046, Shitian He, Xuanming Liu, Jiangshan Li, Wanyu Chen
Eng. Appl. Artif. Intell.1
2025 Multimodal image generation and fusion through content-style hybrid disentanglement
abstract
• Research highlight 1: We propose a novel cross-task hybrid training methodology for multimodal images, offering a simple yet unified solution that simultaneously addresses both image generation and fusion tasks. • Research highlight 2: Building upon mutual-supervised multimodal image pairs, we innovatively integrate single-modality self-supervision to develop a hybrid-supervised decoupling framework with a dedicated loss function, achieving robust separation of content-style representations. • Research highlight 3: Extensive experiments spanning on four modalities and seven popular datasets demonstrate our method’s consistent superiority and impressive cross-task capability. Ablation studies further reveal that our framework learns generalized representations transferable across different image processing tasks. Multimodal image fusion and cross-modal translation are fundamental yet challenging tasks in computer vision, with their performance directly impacting downstream applications. Existing approaches typically treat these tasks independently, developing specialized models that fail to exploit the intrinsic relationships between different modalities. This limitation not only restricts model generalizability but also hinders further performance improvements. In this paper, we propose a joint optimization framework for image generation and fusion. Specifically, we generalize multimodal image tasks as the fusion and transformation of cross-modal features, and design a hybrid task training strategy. At the data level, we introduce a self-supervised and mutual-supervised hybrid mechanism for content-style feature decoupling, which achieves superior feature separation through stepwise training on intra-modal and cross-modal data. At the model level, we construct a triple-branch decoupling head along with fusion and transformation modules to ensure synchronous and efficient execution of dual tasks. Our method not only breaks through the single task limitation of the model, but also innovatively introduces mixed supervision into multimodal processing. We conduct comprehensive experiments covering four modalities fusion tasks on seven popular datasets. Extensive experimental results demonstrate that our method achieves superior performance on two tasks as compared of the respective state-of-the-art methods, and show impressive cross-task generalization capability.
Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002, Liyuan Pan
Knowl. Based Syst.8
2025 LCCo: Lending CLIP to co-segmentation
Xin Duan, Yan Yang 0011, Liyuan Pan, Xiabi Liu
Pattern Recognit.3
2025 L2T-DFM: Learning to Teach with Dynamic Fused Metric
Zhaoyang Hai, Liyuan Pan, Xiabi Liu, Mengqiao Han
Pattern Recognit.2
2024 MA-Net: Rethinking Neural Unit in the Light of Astrocytes
abstract
The artificial neuron (N-N) model-based networks have accomplished extraordinary success for various vision tasks. However, as a simplification of the mammal neuron model, their structure is locked during training, resulting in overfitting and over-parameters. The astrocyte, newly explored by biologists, can adaptively modulate neuronal communication by inserting itself between neurons. The communication, between the astrocyte and neuron, is bidirectionally and shows the potential to alleviate issues raised by unidirectional communication in the N-N model. In this paper, we first elaborate on the artificial Multi-Astrocyte-Neuron (MA-N) model, which enriches the functionality of the artificial neuron model. Our MA-N model is formulated at both astrocyte- and neuron-level that mimics the bidirectional communication with temporal and joint mechanisms. Then, we construct the MA-Net network with the MA-N model, whose neural connections can be continuously and adaptively modulated during training. Experiments show that our MA-Net advances new state-of-the-art on multiple tasks while significantly reducing its parameters by connection optimization.
Mengqiao Han, Liyuan Pan, Xiabi Liu
AAAI2
2024 LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network
abstract
Recovering sharp images from dual-pixel (DP) pairs with disparity-dependent blur is a challenging task. Existing blur map-based deblurring methods have demonstrated promising results. In this paper, we propose, to the best of our knowledge, the first framework that introduces the contrastive language-image pre-training framework (CLIP) to accurately estimate the blur map from a DP pair unsu-pervisedly. To achieve this, we first carefully design text prompts to enable CLIP to understand blur-related geo-metric prior knowledge from the DP pair. Then, we pro-pose a format to input a stereo DP pair to CLIP without any fine-tuning, despite the fact that CLIP is pre-trained on monocular images. Given the estimated blur map, we intro-duce a blur-prior attention block, a blur-weighting loss, and a blur-aware loss to recover the all-in-focus image. Our method achieves state-of-the-art performance in extensive experiments (see Fig. 1).
Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Richard I. Hartley, Miaomiao Liu 0001
CVPR2
2024 Language-driven All-in-one Adverse Weather Removal
abstract
All-in-one (AiO) frameworks restore various adverse weather degradations with a single set of networks jointly. To handle various weather conditions, an AiO framework is expected to adaptively learn weather-specific knowledge for different degradations and shared knowledge for common patterns. However, existing methods: 1) rely on extra su-pervision signals, which are usually unknown in real-world applications; 2) employ fixed network structures, which re-strict the diversity of weather-specific knowledge. In this paper, we propose a Language-driven Restoration frame-work (LDR) to alleviate the aforementioned issues. First, we leverage the power of pre-trained vision-language (PVL) models to enrich the diversity of weather-specific knowl-edge by reasoning about the occurrence, type, and severity of degradation, generating description-based degradation priors. Then, with the guidance of degradation prior, we sparsely select restoration experts from a candidate list dy-namically based on a Mixture-of-Experts (MoE) structure. This enables us to adaptively learn the weather-specific and shared knowledge to handle various weather conditions (e.g., unknown or mixed weather). Experiments on exten-sive restoration scenarios show our superior performance.
Hao Yang 0040, Liyuan Pan, Yan Yang 0011, Wei Liang 0008
CVPR2
2024 Event Camera Data Dense Pre-training
Yan Yang 0011, Liyuan Pan, Liu Liu 0009
ECCV (43)2
2024 A Lightweight Multi-Level Relation Network for Few-shot Action Recognition
abstract
Few-shot (FS) action recognition classifies new actions with limited training samples. Most existing works focus on the variability between actions/videos by designing for either feature extraction methods or training strategies. However, they ignore the relations for a same action at different time clips, which is crucial to improve class-specific discriminability. In this paper, we propose a lightweight multi-level relation network (MLRN) that considers the variability of an action that inner- and cross-video, based on episodic training strategies. Furthermore, a query-support similarity classifier is introduced to improve the class identifiability by enhancing the feature utilisation at different levels. Experiments on three challenging benchmarks demonstrate that the proposed MLRN outperforms state-of-the-art methods while using approximately 50% fewer trainable parameters.
Enqi Liu, Liyuan Pan
ICME2
2024 YOLOX-Drone: An Improved Object Detection Method for UAV Images
abstract
Unmanned aerial vehicles (UAV) are widely used for their small size and flexibility. However, the large number of small objects and the significant difference in object size in UAV images bring great challenges to the detection task. Therefore, we propose an object detection method for UAV images with four improvements on the strong baseline model YOLOX-S, which is robust to detect small objects and multi-scale objects. Firstly, we introduce a high-resolution feature map to retain rich detailed information about small objects. Secondly, we propose new up-sampling and down-sampling modules to reduce the feature information loss during the sampling process. Thirdly, we present the triple-scale feature fusion module (TSFFM) to fuse more abundant multi-scale features in the neck’s bottom-up feature fusion process. Finally, the parrell dilated convolution attention module (PD-CAM) is proposed to learn the multi-receptive field features. Experiment results on the VisDrone-VID2019 dataset validate the effectiveness and superiority of the proposed method.
Huanxin Zou, Shitian He, Shuo Liu 0015, Liyuan Pan
IGARSS7
2024 Event-based Few-shot Fine-grained Human Action Recognition
abstract
Few-shot fine-grained human (FGH) action recognition is crucial in the context of human-robot interaction within open-set real-world environments. Existing works mainly focus on features extracted from RGB frames. However, their performances are drastically impacted in challenging scenarios, such as high-dynamic or low lighting conditions. Event cameras can independently and sparsely capture brightness changes in a scene at microsecond resolution and high dynamic range, which offer a promising solution. However, the modality differences between events and RGB frames, and the lack of paired fine-grained data hinder the development of event-based FGH action recognition. Therefore, in this paper, we introduce the first Event Camera Fine-grained Human Action (E-FAction) dataset. This dataset comprises 3304 paired ‘event stream and RGB sequence’, covering 15 coarse action classes and 128 fine-grained actions. Then, we develop a versatile event feature extractor. Considering the spatial sparsity of event stream, we design two modules to mine the temporal motion and semantic features under the guidance of paired RGB frames, facilitating robust weight initialization for the feature extractor in few-shot FGH action recognition. We conduct extensive experiments on both published and our built synthetic and real datasets, and consistently achieve state-of-the-art performance compared to existing baselines. Code and dataset will be available at link.
Zonglin Yang 0002, Yan Yang 0011, Yuheng Shi, Hao Yang 0040, Ruikun Zhang, Liu Liu 0009, Xinxiao Wu, Liyuan Pan
IROS8
2024 SpikMamba: When SNN meets Mamba in Event-based Human Action Recognition
Yan Yang 0011, Shizhuo Deng, Da Teng, Liyuan Pan
MMAsia5
2024 LMHaze: Intensity-aware Image Dehazing with a Large-scale Multi-intensity Real Haze Dataset
abstract
Image dehazing has drawn a significant attention in recent years.Learning-based methods usually require paired hazy and corresponding ground truth (haze-free) images for training.However, it is difficult to collect real-world image pairs, which prevents developments of existing methods.Although several works partially alleviate this issue by using synthetic datasets or small-scale real datasets.The haze intensity distribution bias and scene homogeneity in existing datasets limit the generalization ability of these methods, particularly when encountering images with previously unseen haze intensities.In this work, we present LMHaze, a large-scale, high-quality real-world dataset.LMHaze comprises paired hazy and haze-free images captured in diverse indoor and outdoor environments, spanning multiple scenarios and haze intensities.It contains over 5K high-resolution image pairs, surpassing the size of the biggest existing real-world dehazing dataset by over 25 times.Meanwhile, to better handle images with different haze intensities, we propose a mixture-of-experts model based on Mamba (MoE-Mamba) for dehazing, which dynamically adjusts the model parameters according to the haze intensity.Moreover, with our proposed dataset, we conduct a new large multimodal model (LMM)-based benchmark study to simulate human perception for evaluating dehazed images.Experiments demonstrate that LMHaze dataset improves the dehazing performance in real scenarios and our dehazing method provides better results compared to state-of-the-art methods.The dataset and code are available at our project page.
Ruikun Zhang, Hao Yang 0040, Yan Yang 0011, Ying Fu 0003, Liyuan Pan
MMAsia5
2024 Convolutional Masked Image Modeling for Dense Prediction Tasks on Pathology Images
abstract
This paper studies a convolutional masked image modeling approach for boosting downstream dense prediction tasks on pathology images. Our method is self-supervised, and entails two strategies in sequence. Considering features contained in the pathology images usually have a large spatial span, e.g., glands, we insert [MASK] tokens to the masked regions after the stem layer of the convolutional network for encoding unmasked pixels, which facilitates information propagation through masked regions for reconstructing unmasked pixels. Furthermore, the pathology images contain features that are represented in diverse affine shapes and color spaces. We, therefore, enforce the network to learn the affine and color invariant embedding by imposing transformation constraints between the unmasked image-encoded embedding and reconstruction targets. Our approach is simple but effective. With extensive experiments on standard benchmark datasets, we demonstrate superior transfer learning performance on downstream tasks over past state-of-the-art approaches.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Eric A. Stone
WACV2
2024 Adaptive Hypersphere Data Description for few-shot one-class classification
Yuchen Ren 0003, Xiabi Liu, Liyuan Pan, Lijuan Niu
Appl. Intell.3
2024 Weakly-Supervised Depth Estimation and Image Deblurring via Dual-Pixel Sensors
abstract
Dual-pixel (DP) imaging sensors are getting more popularly adopted by modern cameras. A DP camera captures a pair of images in a single snapshot by splitting each pixel in half. Several previous studies show how to recover depth information by treating the DP pair as an approximate stereo pair. However, dual-pixel disparity occurs only in image regions with defocus blur which is unlike classic stereo disparity. Heavy defocus blur in DP pairs affects the performance of depth estimation approaches based on matching. Therefore, we treat the blur removal and the depth estimation as a joint problem. We investigate the formation of the DP pair, which links the blur and depth information, rather than blindly removing the blur effect. We propose a mathematical DP model that can improve depth estimation by the blur. This exploration motivated us to propose our previous work, an end-to-end DDDNet (DP-based Depth and Deblur Network), which jointly estimates depth and restores the image in a supervised fashion. However, collecting the ground-truth (GT) depth map for the DP pair is challenging and limits the depth estimation potential of the DP sensor. Therefore, we propose an extension of the DDDNet, called WDDNet (Weakly-supervised Depth and Deblur Network), which includes an efficient reblur solver that does not require GT depth maps for training. To achieve this, we convert all-in-focus images into supervisory signals for unsupervised depth estimation in our WDDNet. We jointly estimate an all-in-focus image and a disparity map, then use a Reblur and Fstack module to regularize the disparity estimation and image restoration. We conducted extensive experiments on synthetic and real data to demonstrate the competitive performance of our method when compared to state-of-the-art (SOTA) supervised approaches.
Liyuan Pan, Richard I. Hartley, Liu Liu 0009, Shah Ariful Hoque Chowdhury, Yan Yang 0011, Hongdong Li, Miaomiao Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 AstroNet: When Astrocyte Meets Artificial Neural Network
abstract
Network structure learning aims to optimize network architectures and make them more efficient without compromising performance. In this paper, we first study the astrocytes, a new mechanism to regulate connections in the classic M-P neuron. Then, with the astrocytes, we propose an AstroNet that can adaptively optimize neuron connections and therefore achieves structure learning to achieve higher accuracy and efficiency. AstroNet is based on our built Astrocyte-Neuron model, with a temporal regulation mechanism and a global connection mechanism, which is inspired by the bidirectional communication property of astrocytes. With the model, the proposed AstroNet uses a neural network (NN) for performing tasks, and an astrocyte network (AN) to continuously optimize the connections of NN, i.e., assigning weight to the neuron units in the NN adaptively. Experiments on the classification task demonstrate that our AstroNet can efficiently optimize the network structure while achieving state-of-the-art (SOTA) accuracy.
Mengqiao Han, Liyuan Pan, Xiabi Liu
CVPR2
2023 K3DN: Disparity-Aware Kernel Estimation for Dual-Pixel Defocus Deblurring
abstract
The dual-pixel (DP) sensor captures a two-view image pair in a single snapshot by splitting each pixel in half. The disparity occurs in defocus blurred regions between the two views of the DP pair, while the in-focus sharp regions have zero disparity. This motivates us to propose a K3DN framework for DP pair deblurring, and it has three modules: i) a disparity-aware deblur module. It estimates a disparity feature map, which is used to query a trainable kernel set to estimate a blur kernel that best describes the spatially-varying blur. The kernel is constrained to be symmetrical per the DP formulation. A simple Fourier transform is performed for deblurring that follows the blur model; ii) a reblurring regularization module. It reuses the blur kernel, performs a simple convolution for reblurring, and regularizes the estimated kernel and disparity feature unsupervisedly, in the training stage; iii) a sharp region preservation module. It identifies in-focus regions that correspond to areas with zero disparity between DP images, aims to avoid the introduction of noises during the deblurring process, and improves image restoration performance. Experiments on four standard DP datasets show that the proposed K3DN outperforms state-of-the-art methods, with fewer parameters and flops at the same time.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Miaomiao Liu 0001
CVPR2
2023 Event Camera Data Pre-training
abstract
This paper proposes a pre-trained neural network for handling event camera data. Our model is a self-supervised learning framework, and uses paired event camera data and natural RGB images for training. Our method contains three modules connected in a sequence: i) a family of event data augmentations, generating meaningful event images for self-supervised training; ii) a conditional masking strategy to sample informative event patches from event images, encouraging our model to capture the spatial layout of a scene and accelerating training; iii) a contrastive learning approach, enforcing the similarity of embeddings between matching event images, and between paired event and RGB images. An embedding projection loss is proposed to avoid the model collapse when enforcing the event image embedding similarities. A probability distribution alignment loss is proposed to encourage the event image to be consistent with its paired RGB image in the feature space. Transfer learning performance on downstream tasks shows the superiority of our method over state-of-the-art methods. For example, we achieve top-1 accuracy at 64.83% on the N-ImageNet dataset. Our code is available at https://github.com/Yan98/Event-Camera-Data-Pre-training.
Yan Yang 0011, Liyuan Pan, Liu Liu 0009
ICCV2
2023 L2T-DLN: Learning to Teach with Dynamic Loss Network
abstract
With the concept of teaching being introduced to the machine learning community, a teacher model start using dynamic loss functions to teach the training of a student model. The dynamic intends to set adaptive loss functions to different phases of student model learning. In existing works, the teacher model 1) merely determines the loss function based on the present states of the student model, e.g., disregards the experience of the teacher; 2) only utilizes the states of the student model, e.g., training iteration number and loss/accuracy from training/validation sets, while ignoring the states of the loss function. In this paper, we first formulate the loss adjustment as a temporal task by designing a teacher model with memory units, and, therefore, enables the student learning to be guided by the experience of the teacher model. Then, with a Dynamic Loss Network, we can additionally use the states of the loss to assist the teacher learning in enhancing the interactions between the teacher and the student model. Extensive experiments demonstrate our approach can enhance student learning and improve the performance of various deep models on real-world tasks, including classification, objective detection, and semantic segmentation scenario.
Zhaoyang Hai, Liyuan Pan, Xiabi Liu, Zhengzheng Liu, Mirna Yunita
NeurIPS2
2022 ISG: I can See Your Gene Expression
Yan Yang 0011, Liyuan Pan, Liu Liu 0009, Eric A. Stone
BMVC2
2022 Biomass Prediction with 3D Point Clouds from LiDAR
abstract
With population growth and a shrinking rural workforce, agricultural technologies have become increasingly important. Above-ground biomass (AGB) is a key trait relevant to breeding, agronomy and crop physiology field experiments. However, measuring the biomass of a cereal plot requires cutting, drying and weighing processes, which are laborious, expensive and destructive tasks. This paper proposes a non-destructive and high-throughput method to predict biomass from field samples based on Light Detection and Ranging (LiDAR). Unlike previous methods that are based on the density of a point cloud or plant height, our biomass prediction network (BioNet) additionally considers plant structure. Our BioNet contains three modules: 1) a completion module to predict missing points due to canopy occlusion; 2) a regularization module to regularize the neural representation of the whole plot; and 3) a projection module to learn the salient structures from a bird’s eye view of the point cloud. An attention-based fusion block is used to achieve final biomass predictions. In addition, the complete dataset, including hand-measured biomass and LiDAR data, is made available to the community. Experiments show that our BioNet achieves ≈ 33% improvement over current state-of-the-art methods.
Liyuan Pan, Liu Liu 0009, Anthony G. Condon, Gonzalo M. Estavillo, Robert Coe, Geoff Bull, Eric A. Stone, Lars Petersson, Vivien Rolland
WACV1
2022 High Frame Rate Video Reconstruction Based on an Event Camera
abstract
Event-based cameras measure intensity changes (called 'events') with microsecond accuracy under high-speed motion and challenging lighting conditions. With the 'active pixel sensor' (APS), the 'Dynamic and Active-pixel Vision Sensor' (DAVIS) allows the simultaneous output of intensity frames and events. However, the output images are captured at a relatively low frame rate and often suffer from motion blur. A blurred image can be regarded as the integral of a sequence of latent images, while events indicate changes between the latent images. Thus, we are able to model the blur-generation process by associating event data to a latent sharp image. Based on the abundant event data alongside a low frame rate, easily blurred images, we propose a simple yet effective approach to reconstruct high-quality and high frame rate sharp videos. Starting with a single blurred frame and its event data from DAVIS, we propose the Event-based Double Integral (EDI) model and solve it by adding regularization terms. Then, we extend it to multiple Event-based Double Integral (mEDI) model to get more smooth results based on multiple images and their events. Furthermore, we provide a new and more efficient solver to minimize the proposed energy model. By optimizing the energy function, we achieve significant improvements in removing blur and the reconstruction of a high temporal resolution video. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real datasets demonstrate the superiority of our mEDI model and optimization method compared to the state-of-the-art.
Liyuan Pan, Richard I. Hartley, Cedric Scheerlinck, Miaomiao Liu 0001, Xin Yu 0002, Yuchao Dai
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration
abstract
The dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus blur in DP pairs affects the performance of matching-based depth estimation approaches. Instead of removing the blur effect blindly, we study the formation of the DP pair which links the blur and the depth information. In this paper, we propose a mathematical DP model which can benefit depth estimation by the blur. These explorations motivate us to propose an end-to-end DDDNet (DP-based Depth and Deblur Network) to jointly estimate the depth and restore the image. Moreover, we define a re-blur loss, which reflects the relationship of the DP image formation process with depth information, to regularise our depth estimate in training. To meet the requirement of a large amount of data for learning, we propose the first DP image simulator which allows us to create datasets with DP pairs from any existing RGBD dataset. As a side contribution, we collect a real dataset for further research. Extensive experimental evaluation on both synthetic and real datasets shows that our approach achieves competitive performance compared to state-of-the-art approaches.
Liyuan Pan, Shah Chowdhury, Richard I. Hartley, Miaomiao Liu 0001, Hongdong Li
CVPR1
2021 Stereo Hybrid Event-Frame (SHEF) Cameras for 3D Perception
abstract
Stereo camera systems play an important role in robotics applications to perceive the 3D world. However, conventional cameras have drawbacks such as low dynamic range, motion blur and latency due to the underlying frame- based mechanism. Event cameras address these limitations as they report the brightness changes of each pixel independently with a fine temporal resolution, but they are unable to acquire absolute intensity information directly. Although integrated hybrid event-frame sensors (e.g., DAVIS) are available, the quality of data is compromised by coupling at the pixel level in the circuit fabrication of such cameras. This paper proposes a stereo hybrid event-frame (SHEF) camera system that offers a sensor modality with separate high-quality pure event and pure frame cameras, overcoming the limitations of each separate sensor and allowing for stereo depth estimation. We provide a SHEF dataset targeted at evaluating disparity estimation algorithms and introduce a stereo disparity estimation algorithm that uses edge information extracted from the event stream correlated with the edge detected in the frame data. Our disparity estimation outperforms the state-of-the-art stereo matching algorithm on the SHEF dataset.
Ziwei Wang 0002, Liyuan Pan, Yonhon Ng, Zheyu Zhuang, Robert E. Mahony
IROS2
2020 Single Image Optical Flow Estimation With an Event Camera
abstract
Event cameras are bio-inspired sensors that asynchronously report intensity changes in microsecond resolution. DAVIS can capture high dynamics of a scene and simultaneously output high temporal resolution events and low frame-rate intensity images. In this paper, we propose a single image (potentially blurred) and events based optical flow estimation approach. First, we demonstrate how events can be used to improve flow estimates. To this end, we encode the relation between flow and events effectively by presenting an event-based photometric consistency formulation. Then, we consider the special case of image blur caused by high dynamics in the visual environments and show that including the blur formation in our model further constrains flow estimation. This is in sharp contrast to existing works that ignore the blurred images while our formulation can naturally handle either blurred or sharp images to achieve accurate flow estimation. Finally, we reduce flow estimation, as well as image deblurring, to an alternative optimization problem of an objective function using the primal-dual algorithm. Experimental results on both synthetic and real data (with blurred and non-blurred images) show the superiority of our model in comparison to state-of-the-art approaches.
Liyuan Pan, Miaomiao Liu 0001, Richard I. Hartley
CVPR1
2020 Joint Stereo Video Deblurring, Scene Flow Estimation and Moving Object Segmentation
abstract
Stereo videos for the dynamic scenes often show unpleasant blurred effects due to the camera motion and the multiple moving objects with large depth variations. Given consecutive blurred stereo video frames, we aim to recover the latent clean images, estimate the 3D scene flow and segment the multiple moving objects. These three tasks have been previously addressed separately, which fail to exploit the internal connections among these tasks and cannot achieve optimality. In this paper, we propose to jointly solve these three tasks in a unified framework by exploiting their intrinsic connections. To this end, we represent the dynamic scenes with the piece-wise planar model, which exploits the local structure of the scene and expresses various dynamic scenes. Under our model, these three tasks are naturally connected and expressed as the parameter estimation of 3D scene structure and camera motion (structure and motion for the dynamic scenes). By exploiting the blur model constraint, the moving objects and the 3D scene structure, we reach an energy minimization formulation for joint deblurring, scene flow and segmentation. We evaluate our approach extensively on both synthetic datasets and publicly available real datasets with fast-moving objects, camera motion, uncontrolled lighting conditions and shadows. Experimental results demonstrate that our method can achieve significant improvement in stereo video deblurring, scene flow estimation and moving object segmentation, over state-of-the-art methods.
Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli, Quan Pan 0001
IEEE Trans. Image Process.1
2019 Phase-Only Image Based Kernel Estimation for Single Image Blind Deblurring
abstract
The image motion blurring process is generally modelled as the convolution of a blur kernel with a latent image. Therefore, the estimation of the blur kernel is essentially important for blind image deblurring. Unlike existing approaches which focus on approaching the problem by enforcing various priors on the blur kernel and the latent image, we are aiming at obtaining a high quality blur kernel directly by studying the problem in the frequency domain. We show that the auto-correlation of the absolute phase-only image 1 can provide faithful information about the motion (e.g., the motion direction and magnitude, we call it the motion pattern in this paper.) that caused the blur, leading to a new and efficient blur kernel estimation approach. The blur kernel is then refined and the sharp image is estimated by solving an optimization problem by enforcing a regularization on the blur kernel and the latent image. We further extend our approach to handle non-uniform blur, which involves spatially varying blur kernels. Our approach is evaluated extensively on synthetic and real data and shows good results compared to the state-of-the-art deblurring approaches.
Liyuan Pan, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai
CVPR1
2019 Bringing a Blurry Frame Alive at High Frame-Rate With an Event Camera
abstract
Event-based cameras can measure intensity changes (called ‘events’) with microsecond accuracy under high-speed motion and challenging lighting conditions. With the active pixel sensor (APS), the event camera allows simultaneous output of the intensity frames. However, the output images are captured at a relatively low frame-rate and often suffer from motion blur. A blurry image can be regarded as the integral of a sequence of latent images, while the events indicate the changes between the latent images. Therefore, we are able to model the blur-generation process by associating event data to a latent image. In this paper, we propose a simple and effective approach, the Event-based Double Integral (EDI) model, to reconstruct a high frame-rate, sharp video from a single blurry frame and its event data. The video generation is based on solving a simple non-convex optimization problem in a single scalar variable. Experimental results on both synthetic and real images demonstrate the superiority of our EDI model and optimization method in comparison to the state-of-the-art.
Liyuan Pan, Cedric Scheerlinck, Xin Yu 0002, Richard I. Hartley, Miaomiao Liu 0001, Yuchao Dai
CVPR1
2019 Single Image Deblurring and Camera Motion Estimation With Depth Map
abstract
Camera shake during exposure is a major problem in hand-held photography, as it causes image blur that destroys details in the captured images. In the real world, such blur is mainly caused by both the camera motion and the complex scene structure. While considerable existing approaches have been proposed based on various assumptions regarding the scene structure or the camera motion, few existing methods could handle the real 6 DoF camera motion. In this paper, we propose to jointly estimate the 6 DoF camera motion and remove the non-uniform blur caused by camera motion by exploiting their underlying geometric relationships, with a single blurry image and its depth map (either direct depth measurements, or a learned depth map) as input. We formulate our joint deblurring and 6 DoF camera motion estimation as an energy minimization problem which is solved in an alternative manner. Our model enables the recovery of the 6 DoF camera motion and the latent clean image, which could also achieve the goal of generating a sharp sequence from a single blurry image. Experiments on challenging real-world and synthetic datasets demonstrate that image blur from camera shake can be well addressed within our proposed framework.
Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001
WACV1
2018 Depth Map Completion by Jointly Exploiting Blurry Color Images and Sparse Depth Maps
abstract
We aim at predicting a complete and high-resolution depth map from incomplete, sparse and noisy depth measurements. Existing methods handle this problem either by exploiting various regularizations on the depth maps directly or resorting to learning based methods. When the corresponding color images are available, the correlation between the depth maps and the color images are used to improve the completion performance, assuming the color images are clean and sharp. However, in real world dynamic scenes, color images are often blurry due to the camera motion and the moving objects in the scene. In this paper, we propose to tackle the problem of depth map completion by jointly exploiting the blurry color image sequences and the sparse depth map measurements, and present an energy minimization based formulation to simultaneously complete the depth maps, estimate the scene flow and deblur the color images. Our experimental evaluations on both outdoor and indoor scenarios demonstrate the state-of-the-art performance of our approach.
Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli
WACV1
2017 Simultaneous Stereo Video Deblurring and Scene Flow Estimation
abstract
Videos for outdoor scene often show unpleasant blur effects due to the large relative motion between the camera and the dynamic objects and large depth variations. Existing works typically focus monocular video deblurring. In this paper, we propose a novel approach to deblurring from stereo videos. In particular, we exploit the piece-wise planar assumption about the scene and leverage the scene flow information to deblur the image. Unlike the existing approach [31] which used a pre-computed scene flow, we propose a single framework to jointly estimate the scene flow and deblur the image, where the motion cues from scene flow estimation and blur information could reinforce each other, and produce superior results than the conventional scene flow estimation or stereo deblurring methods. We evaluate our method extensively on two available datasets and achieve significant improvement in flow estimation and removing the blur effect over the state-of-the-art methods.
Liyuan Pan, Yuchao Dai, Miaomiao Liu 0001, Fatih Porikli
CVPR1