VLDB 2026 Research / reviewers in the wild / expert
Leidong Fan
dblp:221/9543
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-3082-695XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Monocular Depth via Cascaded Iterative Refinement in Visual-Echo ScenesabstractIn recent years, integrating multimodal information, particularly visual and echo data, has shown great promise for improving depth estimation performance. While existing works demonstrate that combining binaural echo features with image attributes can enhance depth estimation, they often use rudimentary feature alignment and fusion methods, failing to fully exploit the complementary nature of cross-modal information and limiting integration effectiveness. To address these challenges, this paper introduces an innovative multimodal fusion framework. First, the framework incorporates a combination of multi-scale self-attention and cross-attention mechanisms, establishing correlations between features and facilitating cohesive interactions between the visual and echo domains. Furthermore, we propose an incremental feature updating mechanism based on Convolutional Gated Recurrent Units (ConvGRU), which implements cascaded iterative optimization, integrating contextual features with the multi-scale fused features from both echo and image modalities. In each iteration, the framework preserves contextual information from previous steps while employing a multi-level loss function to guide result updates. This approach effectively captures spatial structural information and progressively enhances depth estimation accuracy. Comprehensive experimental evaluations on the Replica, Matterport3D and BatVision (BV1) datasets validate the effectiveness of the proposed method. Comparative analyses with state-of-the-art monocular plus echo methods underscore the superior performance achievable through this novel framework. Anjie Wang, Zhijun Fang 0001, Leidong Fan, Guibiao Liao, Siwei Ma 0001, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Unified Inverse-Tone-Mapped HDR Video Quality Assessment Method across Two HDR FormatsabstractHigh Dynamic Range Video Quality Assessment (HDR VQA) plays a pivotal role in Inverse Tone Mapping (ITM) research. Existing HDR VQA datasets and models mainly focus on a single HDR format, leading to poor generalization and limited application scope. To address this problem, this paper proposes a reference-free CONTrastive ITM-VQA (CONT-ITM-VQA) model via format transformation-based data augmentation and contrastive learning. Specifically, the format transformation-based data augmentation improves the model generalization, via applying Opto-electronic Transfer Function (OETF) transformations between the HDR formats; while the contrastive learning-based quality-related feature alignment aligns the quality features from different HDR formats of the same video to obtain more effective quality representations. It is worth noting that our method can be extended to other HDR-related quality assessment, not limited to ITM-HDR VQA. Experimental results demonstrate that our model closely mimics subjective judgments. Leidong Fan, Xiongkuo Min, Qing Li 0029, Anjie Wang |
ICME | 1 |
| 2025 | Inverse-Tone-Mapped HDR Video Quality Assessment for Broadcast Television: A Comprehensive Dataset and SDR-Referenced MethodabstractInverse-Tone-Mapped High Dynamic Range Video Quality Assessment (ITM-HDR VQA) plays a pivotal role in evaluating the visual quality of ITM-enhanced HDR videos. The research community tackles this issue from the dataset and method perspectives. However, current ITM-HDR VQA datasets exhibit three key limitations: narrow scene diversity, partial HDR format representation, and inadequate distortion coverage; existing methods face challenges in HDR and SDR domain discrepancy and insufficient ITM-induced quality feature extraction. To bridge these gaps, we introduce a comprehensive ITM HDR Video Quality Assessment dataset tailored to Broadcast Television (BT-ITM-VQA), along with a novel SDR-Referenced Bidirectional Quality Interaction (SDR-R-BQI) method. The BT-ITM-VQA dataset features rich broadcast scenes, multiple HDR-format support of Hybrid Log-Gamma (HLG) and Perceptual Quantizer (PQ), and real-world distortions induced by super-resolution and deinterlacing, providing a systematic foundation for ITM-HDR VQA model development and validation. The SDR-R-BQI method effectively mitigates HDR and SDR discrepancies through luminance dynamic range alignment and color gamut alignment, and then extracts ITM-induced quality alterations by bidirectional, cross-quality-based computation in a unified feature space. Extensive validation on four datasets demonstrates the effectiveness of our newly constructed dataset and proposed method. Leidong Fan, Qian Zhang 0096, Qing Li 0029 |
ACM Multimedia | 1 |
| 2023 | Learning a Practical SDR-to-HDRTV Up-conversion using New Dataset and Degradation ModelsabstractIn media industry, the demand of SDR-to-HDRTV upconversion arises when users possess HDR-WCG (high dynamic range-wide color gamut) TVs while most off-the-shelf footage is still in SDR (standard dynamic range). The research community has started tackling this low-level vision task by learning-based approaches. When applied to real SDR, yet, current methods tend to produce dim and desaturated result, making nearly no improvement on viewing experience. Different from other network-oriented methods, we attribute such deficiency to training set (HDR-SDR pair). Consequently, we propose new HDRTV dataset (dubbed HDRTV4K) and new HDR-to-SDR degradation models. Then, it's used to train a luminance-segmented network (LSN) consisting of a global mapping trunk, and two Transformer branches on bright and dark luminance range. We also update assessment criteria by tailored metrics and subjective experiment. Finally, ablation studies are conducted to prove the effectiveness. Our work is available at: https://github.com/AndreGuo/HDRTVDM Cheng Guo 0005, Leidong Fan, Xiuhua Jiang |
CVPR | 2 |
| 2023 | A Decoupled Kernel Prediction Network Guided by Soft Mask for Single Image HDR ReconstructionabstractRecent works on single image high dynamic range (HDR) reconstruction fail to hallucinate plausible textures, resulting in information missing and artifacts in large-scale under/over-exposed regions. In this article, a decoupled kernel prediction network is proposed to infer an HDR image from a low dynamic range (LDR) image. Specifically, we first adopt a simple module to generate a preliminary result, which can precisely estimate well-exposed HDR regions. Meanwhile, an encoder-decoder backbone network with a soft mask guidance module is presented to predict pixel-wise kernels, which is further convolved with the preliminary result to obtain the final HDR output. Instead of traditional kernels, our predicted kernels are decoupled along the spatial and channel dimensions. The advantages of our method are threefold at least. First, our model is guided by the soft mask so that it can focus on the most relevant information for under/over-exposed regions. Second, pixel-wise kernels are able to adaptively solve the different degradations for differently exposed regions. Third, decoupled kernels can avoid information redundancy across channels and reduce the solution space of our model. Thus, our method is able to hallucinate fine details in the under/over-exposed regions and renders visually pleasing results. Extensive experiments demonstrate that our model outperforms state-of-the-art ones. Gaofeng Cao, Fei Zhou 0001, Kanglin Liu, Anjie Wang, Leidong Fan |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | KPN-MFI: A Kernel Prediction Network with Multi-frame Interaction for Video Inverse Tone MappingabstractUp to now, the image-based inverse tone mapping (iTM) models have been widely investigated, while there is little research on video-based iTM methods. It would be interesting to make use of these existing image-based models in the video iTM task. However, directly transferring the imagebased iTM models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both the intra-frame quality and the inter-frame consistency of a video, this article presents a new video iTM method based on a kernel prediction network (KPN), which takes advantage of multi-frame interaction (MFI) module to capture temporal-spatial information for video data. Specifically, a basic encoder-decoder KPN, essentially designed for image iTM, is trained to guarantee the mapping quality within each frame. More importantly, the MFI module is incorporated to capture temporal-spatial context information and preserve the inter-frame consistency by exploiting the correction between adjacent frames. Notably, we can readily extend any existing image iTM models to video iTM ones by involving the proposed MFI module. Furthermore, we propose an inter-frame brightness consistency loss function based on the Gaussian pyramid to reduce the video temporal inconsistency. Extensive experiments demonstrate that our model outperforms state-ofthe-art image and video-based methods. The code is available at https://github.com/caogaofeng/KPNMFI. Gaofeng Cao, Fei Zhou 0001, Han Yan 0003, Anjie Wang, Leidong Fan |
IJCAI | 5 |