Yanlin Qian

dblp:183/6226 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Optimizing for the Shortest Path in Denoising Diffusion Model
abstract
In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Path Diffusion Model (ShortDF), treats the denoising process as a shortest-path problem aimed at minimizing reconstruction error. By optimizing the initial residuals, we improve the efficiency of the reverse diffusion process and the quality of the generated samples. Extensive experiments on multiple standard benchmarks demonstrate that ShortDF significantly reduces diffusion time (or steps) while enhancing the visual fidelity of generated samples compared to prior arts. This work, we suppose, paves the way for interactive diffusion-based applications and establishes a foundation for rapid data generation. Code is available at https://github.com/UnicomAI/ShortDF.
Xingpeng Zhang, Zhaoxiang Liu, Kai Wang 0012, Min Wang 0031, Yanlin Qian, Shiguo Lian
CVPR8
2025 Integral Fast Fourier Color Constancy
abstract
Traditional auto white balance (AWB) algorithms typically assume a single global illuminant source, which leads to color distortions in multi-illuminant scenes. While recent neural network-based methods have shown excellent accuracy in such scenarios, their high parameter count and computational demands limit their practicality for real-time video applications. The Fast Fourier Color Constancy (FFCC) algorithm was proposed for single-illuminant-source scenes, predicting a global illuminant source with high efficiency. However, it cannot be directly applied to multi-illuminant scenarios unless specifically modified. To address this, we propose Integral Fast Fourier Color Constancy (IFFCC), an extension of FFCC tailored for multi-illuminant scenes. IFFCC leverages the proposed integral UV histogram to accelerate histogram computations across all possible regions in Cartesian space and parallelizes Fourier-based convolution operations, resulting in a spatially-smooth illumination map. This approach enables high-accuracy, real-time AWB in multi-illuminant scenes. Extensive experiments show that IFFCC achieves accuracy that is on par with or surpasses that of pixel-level neural networks, while reducing the parameter count by over 400× and processing speed by 20 − 100× faster than network-based approaches.
Wenjun Wei, Yanlin Qian, Huaian Chen, Junkang Dai, Yi Jin 0002
CVPR2
2025 RGB-Event ISP: The Dataset and Benchmark
abstract
Event-guided imaging has received significant attention due to its potential to revolutionize instant imaging systems. However, the prior methods primarily focus on enhancing RGB images in a post-processing manner, neglecting the challenges of image signal processor (ISP) dealing with event sensor and the benefits events provide for reforming the ISP process. To achieve this, we conduct the first research on event-guided ISP. First, we present a new event-RAW paired dataset, collected with a novel but still confidential sensor that records pixel-level aligned events and RAW images. This dataset includes 3373 RAW images with $2248\times 3264$ resolution and their corresponding events, spanning 24 scenes with 3 exposure modes and 3 lenses. Second, we propose a convential ISP pipeline to generate good RGB frames as reference. This convential ISP pipleline performs basic ISP operations, e.g., demosaicing, white balancing, denoising and color space transforming, with a ColorChecker as reference. Third, we classify the existing learnable ISP methods into 3 classes, and select multiple methods to train and evaluate on our new dataset. Lastly, since there is no prior work for reference, we propose a simple event-guided ISP method and test it on our dataset. We further put forward key technical challenges and future directions in RGB-Event ISP. In summary, to the best of our knowledge, this is the very first research focusing on event-guided ISP, and we hope it will inspire the community.
Yunfan Lu, Yanlin Qian, Ziyang Rao, Junren Xiao
ICLR2
2024 Learning Triangular Distribution in Visual World
abstract
Convolution neural network is successful in pervasive vision tasks, including label distribution learning, which usually takes the form of learning an injection from the nonlinear visual features to the well-defined labels. However, how the discrepancy between features is mapped to the label discrepancy is ambient, and its correctness is not guaranteed. To address these problems, we study the mathematical connection between feature and its label, presenting a general and simple framework for label distribution learning. We propose a so-called Triangular Distribution Transform (TDT) to build an injective function between feature and label, guaranteeing that any symmetric feature discrepancy linearly reflects the difference between labels. The proposed TDT can be used as a plug-in in mainstream backbone networks to address different label distribution learning tasks. Experiments on Facial Age Recognition, Illumination Chromaticity Estimation, and Aesthetics assessment show that TDT achieves on-par or better results than the prior arts. Code is available at https://github.com/redcping/TDT.
Xingpeng Zhang, Chengtao Zhou, Dichao Fan, Peng Tu, Le Zhang 0001, Yanlin Qian
CVPR7
2022 Point Cloud Color Constancy
abstract
In this paper, we present Point Cloud Color Constancy, in short PCCC, an illumination chromaticity estimation algorithm exploiting a point cloud. We leverage the depth information captured by the time-of-flight (ToF) sensor mounted rigidly with the RGB sensor, and form a 6D cloud where each point contains the coordinates and RGB intensities, noted as (x,y,z, r,g, b). PCCC applies the PointNet architecture to the color constancy problem, deriving the illumination vector point-wise and then making a global decision about the global illumination chromaticity. On two popular RGB-D datasets, which we extend with illumination information, as well as on a novel benchmark, PCCC obtains lower error than the state-of-the-art algorithms. Our method is simple andfast, requiring merely 16 x 16-size input and reaching speed over 140 fps (CPU time), including the cost of building the point cloud and net inference.
Xiaoyan Xing, Yanlin Qian, Sibo Feng, Yuhan Dong, Jiri Matas
CVPR2
2022 Dual-Illumination Weighting and Estimation
abstract
Illumination estimation refers to estimating the chromaticity vector of illumination, and can be used to recover the surface color under white light. Dual-illuminant is a common scenario in computational illumination estimation tasks. A straightforward way to correct the dual-illuminant image can be estimating a spatially-varying illumination map. However, it is hindered by the lack of large-scale annotated datasets for data-driven methods. In this paper, we propose a novel approach to obtain the dominant dual-illuminant and the pixel-wise illuminant map on real dual-illuminant raw images. Our method consists of 1) dual-illuminant image generator (DIG) to synthesize dual-illuminant images from unique-illuminant datasets assuming the Lambertian model; 2) dual-illuminant estimation network (DE-Net) to estimate illuminant both globally and locally. Quantitative experiments show that with DIG synthesized dual-illuminant images, DE-Net obtains the best accuracy in dual-illumination detection and estimation on the Gehalr-shi dataset and Mutlti-Illuminant Multi-Object dataset.
Xiaoyan Xing, Sibo Feng, Yanlin Qian, Yuhan Dong
ICPR3
2021 Fast Fourier Intrinsic Network
abstract
We address the problem of decomposing an image into albedo and shading. We propose the Fast Fourier Intrinsic Network, FFI-Net in short, that operates in the spectral domain, splitting the input into several spectral bands. Weights in FFI-Net are optimized in the spectral domain, allowing faster convergence to a lower error. FFI-Net is lightweight and does not need auxiliary networks for training. The network is trained end-to-end with a novel spectral loss which measures the global distance between the network prediction and corresponding ground truth. FFI-Net achieves state-of-the-art performance on MPI-Sintel, MIT Intrinsic, and IIW datasets.
Yanlin Qian, Miaojing Shi, Joni-Kristian Kämäräinen, Jiri Matas
WACV1
2020 Cascading Convolutional Color Constancy
abstract
Regressing the illumination of a scene from the representations of object appearances is popularly adopted in computational color constancy. However, it's still challenging due to intrinsic appearance and label ambiguities caused by unknown illuminants, diverse reflection properties of materials and extrinsic imaging factors (such as different camera sensors). In this paper, we introduce a novel algorithm – Cascading Convolutional Color Constancy (in short, C4) to improve robustness of regression learning and achieve stable generalization capability across datasets (different cameras and scenes) in a unique framework. The proposed C4 method ensembles a series of dependent illumination hypotheses from each cascade stage via introducing a weighted multiply-accumulate loss function, which can inherently capture different modes of illuminations and explicitly enforce coarse-to-fine network optimization. Experimental results on the public Color Checker and NUS 8-Camera benchmarks demonstrate superior performance of the proposed algorithm in comparison with the state-of-the-art methods, especially for more difficult scenes.
Huanglin Yu, Ke Chen 0004, Kaiqi Wang, Yanlin Qian, Zhaoxiang Zhang 0001, Kui Jia
AAAI4
2020 SDE-AWB: a generic solution for 2nd International Illumination Estimation Challenge
abstract
We propose a neural network-based solution for three different tracks of 2nd International Illumination Estimation Challenge (chromaticity.iitp.ru). Our method is built on pre-trained Squeeze-Net backbone, differential 2D chroma histogram layer and a shallow MLP utilizing Exif information. By combining semantic feature, color feature and Exif metadata, the resulting method – SDE-AWB – obtains 1st place in both indoor and two-illuminant tracks and 2nd place in general track.
Yanlin Qian
ICMV1
2020 DAL: A Deep Depth-Aware Long-term Tracker
abstract
The best RGBD trackers provide high accuracy but are slow to run. On the other hand, the best RGB trackers are fast but clearly inferior on the RGBD datasets. In this work, we propose a deep depth-aware long-term tracker that achieves state-of-the-art RGBD tracking performance and is fast to run. We reformulate deep discriminative correlation filter (DCF) to embed the depth information into deep features. Moreover, the same depth-aware correlation filter is used for target redetection. Comprehensive evaluations show that the proposed tracker achieves state-of-the-art performance on the Princeton RGBD, STC, and the newly-released CDTB benchmarks and runs 20 fps.
Yanlin Qian, Alan Lukezic, Matej Kristan, Joni-Kristian Kämäräinen, Jiri Matas
ICPR1
2019 On Finding Gray Pixels
abstract
We propose a novel grayness index for finding gray pixels and demonstrate its effectiveness and efficiency in illumination estimation. The grayness index, GI in short, is derived using the Dichromatic Reflection Model and is learning-free. GI allows to estimate one or multiple illumination sources in color-biased images. On standard single-illumination and multiple-illumination estimation benchmarks, GI outperforms state-of-the-art statistical methods and many recent deep methods. GI is simple and fast, written in a few dozen lines of code, processing a 1080p image in ~0.4 seconds with a non-optimized Matlab code.
Yanlin Qian, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas
CVPR1
2019 Flash Lightens Gray Pixe
abstract
In the real world, a scene is usually cast by multiple illuminants and herein we address the problem of spatial illumination estimation. Our solution is based on detecting gray pixels with the help of flash photography. We show that flash photography significantly improves the performance of gray pixel detection without illuminant prior, training data or calibration of the flash. We also introduce a novel flash photography dataset generated from the MIT intrinsic dataset.
Yanlin Qian, Joni-Kristian Kämäräinen, Jiri Matas
ICIP1
2019 Convolutional low-resolution fine-grained classification
Dingding Cai, Ke Chen 0004, Yanlin Qian, Joni-Kristian Kämäräinen
Pattern Recognit. Lett.3
2018 Object Detection in Equirectangular Panorama
abstract
We introduce a high-resolution equirectangular panorama (aka 360-degree, virtual reality, VR) dataset for object detection and propose a multi-projection variant of the YOLO detector. The main challenges with equirectangular panorama images are i) the lack of annotated training data, ii) high-resolution imagery and iii) severe geometric distortions of objects near the panorama projection poles. In this work, we solve the challenges by I) using training examples available in the “conventional datasets” (ImageNet and COCO), II) employing only low resolution images that require only moderate GPU computing power and memory, and III) our multi-projection YOLO handles projection distortions by making multiple stereographic sub-projections. In our experiments, YOLO outperforms the other state-of-the-art detector, Faster R-CNN, and our multi-projection YOLO achieves the best accuracy with low-resolution input.
Wenyan Yang, Yanlin Qian, Joni-Kristian Kämäräinen, Francesco Cricri, Lixin Fan
ICPR2
2018 Hierarchical Sliding Slice Regression for Vehicle Viewing Angle Estimation
abstract
We propose a novel hierarchical sliding slice regression which in a coarse-to-fine manner represents global circular target space with a number of ordinally localized and overlapping subspaces. Our method is particularly suitable for visual regression problems where the regression target is circular (e.g., car viewing angle) and visual similarity inconsistent over the target space (e.g., repetitive appearance). A good application example is the camera-based car viewing angle estimation problem, where visual similarity of different views is highly inconsistent-front and back views and left and right side views are pair-wise similar, but appear at the far ends of the circular view angle space. In practice, the problem is even more complicated due to large visual variation of objects (e.g., different car models). We perform extensive experiments on the Lausanne Federal of Institute of Technology Multi-view Car and KITTI Data Sets as well as the Technische Universitat Darmstadt Multi-view Pedestrians Data Set and achieve superior performance as compared to the state-of-the-art algorithms.
Dan Yang 0012, Yanlin Qian, Ke Chen 0004, Eleni Berki, Joni-Kristian Kämäräinen
IEEE Trans. Intell. Transp. Syst.2
2017 Recurrent Color Constancy
abstract
We introduce a novel formulation of temporal color constancy which considers multiple frames preceding the frame for which illumination is estimated. We propose an end-to-end trainable recurrent color constancy network – the RCC-Net – which exploits convolutional LSTMs and a simulated sequence to learn compositional representations in space and time. We use a standard single frame color constancy benchmark, the SFU Gray Ball Dataset, which can be adapted to a temporal setting. Extensive experiments show that the proposed method consistently outperforms single-frame state-of-the-art methods and their temporal variants.
Yanlin Qian, Ke Chen 0004, Jarno Nikkanen, Joni-Kristian Kämäräinen, Jiri Matas
ICCV1
2016 Deep structured-output regression learning for computational color constancy
abstract
The color constancy problem is addressed by structured-output regression on the values of the fully-connected layers of a convolutional neural network. The AlexNet and the VGG are considered and VGG slightly outperformed AlexNet. Best results were obtained with the first fully-connected “fc6” layer and with multi-output support vector regression. Experiments on the SFU Color Checker and Indoor Dataset benchmarks demonstrate that our method achieves competitive performance, outperforming the state of the art on the SFU indoor benchmark.
Yanlin Qian, Ke Chen 0004, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas
ICPR1