Tao Yue 0003

dblp:40/7423-3 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-2952-8971ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Learned Optimal Visual Time-of-Flight Imaging With Fisher Information Guidance
abstract
Indirect time-of-flight (iToF) imaging provides absolute depth by encoding light transport delays into correlation measurements via active illumination modulation and demodulation. This technology plays a crucial role in object detection and scene understanding and has been widely adopted in applications such as robotics, automotive driving, and augmented reality. However, iToF systems often suffer from low signal-to-noise ratio (SNR), particularly under strong ambient illumination. Signal attenuation and noise-induced phase errors severely degrade depth accuracy, while reliable depth discontinuity and edge reconstruction remain challenging. In this paper, we introduce an optimal visual ToF imaging scheme, optimized through end-to-end learning of iToF coding functions and depth reconstruction, guided by discriminative Fisher information supervision. By integrating vision information from RGB images, we facilitate the convergence of the optimization for iToF coding functions and enhance robustness under noisy inputs. Moreover, we design a dual-branch visual ToF reconstruction network that estimates depth from iToF measurements while exploiting edge guidance derived from visual information to preserve fine geometric details. Extensive experiments on both synthetic and real-world datasets verify the effectiveness of the proposed approach, particularly in challenging low-SNR conditions.
Kanghui Wang, Jiaqu Li, Zhangnan Li, Xiangshun Kong, Feng Yan 0002, Tao Yue 0003
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 SAM-iToF: SAM-Guided iToF Depth Recovery of Transparent and Mirror Objects
abstract
Indirect time-of-flight (iToF) cameras fail catastrophically when imaging transparent and reflective surfaces, as the complex superposition of light paths violates their fundamental operating principle. Existing methods either rely on overly simplistic physical models that lack adaptability or on data-driven deep learning approaches that exhibit poor generalization. In this letter, we introduce a novel, physics-informed paradigm that synergizes the strengths of both approaches. We are the first to steer a learnable iToF forward model using powerful spatial priors from the Segment Anything Model (SAM). Our framework leverages SAM's guidance to parameterize a reflectance-transmittance mixture model, providing a physically-grounded representation of the scene. This representation is then processed by a dual-stream network, featuring a Dynamic Channel Gate (DCG) and Cross-Attention Fusion (CAF), to intelligently interpret the raw iToF signals and reconstruct a high-fidelity depth map. Extensive experiments on the MSD, Trans10K, and ClearGrasp benchmarks demonstrate that our method not only sets a new state-of-the-art but also shows remarkable generalization from synthetic mirrors to real-world glass objects, offering a robust solution to this long-standing challenge.
Kanghui Wang, Feng Yan 0002, Tao Yue 0003
IEEE Signal Process. Lett.3
2026 OnRef-VSR: Real-Time Online Reference-Based Video Super-Resolution via Efficient Temporal and Spatial Feature Transfers
abstract
Acquiring high-resolution video is essential in applications such as surveillance, remote sensing, and microscopy. Yet pushing spatial resolution often compromises temporal resolution due to readout and bandwidth limitations. Rather than relying solely on hardware improvements, reference-based video superre-solution (Ref-VSR) provides a promising solution to achieve HR video acquisition with off-the-shelf cameras. However, existing Ref-VSR approaches are predominantly offline, relying on future frames and computationally expensive matching, which fundamentally limits their causal real-time applicability. In addition, the inherent trade-off between reconstruction fidelity and inference speed remains a key obstacle to practical use. To overcome these challenges, we present the first online reference-based video super-resolution framework (OnRef-VSR), which enables causal real-time reconstruction through efficient spatial and temporal feature transfers. The framework incorporates three key modules: Flow-Local Linear Transform-based Feature Enhancement (FLFE) module, Hierarchical Heterogeneous Feature Transfer (HHFT) module, and Detail Refinement with Flow-guided Local nonlinear transform (DRFL) module, which in combination realize effective integration of temporal cues and reference information. Extensive experiments demonstrate the superiority of the proposed approach. Specifically, OnRefVSR achieves a 5.2× faster speed with a 0.29 dB gain over the state-of-the-art real-time method, and its offline variant yields a 3.6× speedup and a 0.84 dB PSNR gain on the standard Vid4 benchmark, confirming both the effectiveness and practicality of OnRef-VSR for real-time high-fidelity video reconstruction.
Duzhong Feng, Chenxi Qiu, Tao Yue 0003
IEEE Trans. Circuits Syst. Video Technol.3
2025 Learnable Burst-Encodable Time-of-Flight Imaging for High-Fidelity Long-Distance Depth Sensing
abstract
Long-distance depth imaging holds great promise for applications such as autonomous driving and robotics. Direct time-of-flight (dToF) imaging offers high-precision, long-distance depth sensing, yet demands ultra-short pulse light sources and high-resolution time-to-digital converters. In contrast, indirect time-of-flight (iToF) imaging often suffers from phase wrapping and low signal-to-noise ratio (SNR) as the sensing distance increases. In this paper, we introduce a novel ToF imaging paradigm, termed Burst-Encodable Time-of-Flight (BE-ToF), which facilitates high-fidelity, long-distance depth imaging. Specifically, the BE-ToF system emits light pulses in burst mode and estimates the phase delay of the reflected signal over the entire burst period, thereby effectively avoiding the phase wrapping inherent to conventional iToF systems. Moreover, to address the low SNR caused by light attenuation over increasing distances, we propose an end-to-end learnable framework that jointly optimizes the coding functions and the depth reconstruction network. A specialized double well function and first-order difference term are incorporated into the framework to ensure the hardware implementability of the coding functions. The proposed approach is rigorously validated through comprehensive simulations and real-world prototype experiments, demonstrating its effectiveness and practical applicability. The code is available at: https://github.com/ComputationalPerceptionLab/BE-ToF.
Manchao Bao, Shengjiang Fang, Tao Yue 0003
NeurIPS3
2024 iToF-Flow-Based High Frame Rate Depth Imaging
abstract
iToF is a prevalent, cost-effective technology for 3D perception. While its reliance on multi-measurement commonly leads to reduced performance in dynamic environments. Based on the analysis of the physical iToF imaging process, we propose the iToF flow, composed of crossmode transformation and uni-mode photometric correction, to model the variation of measurements caused by different measurement modes and 3D motion, respectively. We propose a local linear transform (LLT) based cross-mode transfer module (LCTM) for mode-varying and pixel shift compensation of cross-mode flow, and uni-mode photometric correct module (UPCM) for estimating the depth-wise motion caused photometric residual of uni-mode flow. The iToF flow-based depth extraction network is proposed which could facilitate the estimation of the 4-phase measurements at each individual time for high framerate and accurate depth estimation. Extensive experiments, including both simulation and real-world experiments, are conducted to demonstrate the effectiveness of the proposed methods. Compared with the SOTA method, our approach reduces the computation time by 75% while improving the performance by 38%. The code and database are available at https://github.com/ComputationalPerceptionLab/iToF_flow.
Zhou Xue, Tao Yue 0003
CVPR5
2024 Reconstruction-free Cascaded Adaptive Compressive Sensing
abstract
Scene-aware Adaptive Compressive Sensing (ACS) has constituted a persistent pursuit, holding substantial promise for the enhancement of Compressive Sensing (CS) performance. Cascaded ACS furnishes a proficient multi-stage framework for adaptively allocating the CS sampling based on previous CS measurements. However, reconstruction is commonly required for analyzing and steering the successive CS sampling, which bottlenecks the ACS speed and impedes the practical application in time-sensitive scenarios. Addressing this challenge, we propose a reconstruction-free cascaded ACS method, which requires NO reconstruction during the adaptive sampling process. A lightweight Score Network (ScoreNet) is proposed to directly determine the ACS allocation with previous CS measurements and a differentiable adaptive sampling module is proposed for end-to-end training. For image reconstruction, we propose a Multi-Grid Spatial-Attention Network (MGSANet) that could facilitate efficient multi-stage training and inferencing. By introducing the reconstruction-fidelity supervision outside the loop of the multi-stage sampling process, ACS can be efficiently optimized and achieve high imaging fidelity. The effectiveness of the proposed method is demonstrated with extensive quantitative and qualitative experiments, compared with the state-of-the-art CS algorithms.
Chenxi Qiu, Tao Yue 0003
CVPR2
2024 End-to-End Fluorescence Lifetime Imaging with Optimized Encoding and Exposure Allocation
abstract
Fluorescence lifetime imaging (FLIM) is widely used in biomedical applications as a powerful technique for resolving fluorophores and their unique molecular environments. Compared to Time-Domain FLIM (TD-FLIM), Frequency-Domain FLIM (FD-FLIM) allows for faster acquisition of fluorescence lifetime images, making it better suited for live-cell imaging and real-time applications. By illuminating a sample with modulated light and analyzing the phase shift and modulation reduction of the emitted fluorescence, the lifetimes of fluorophores in the sample can be determined. How to retrieve the optimal modulation function for FD-FLIM for noise-robust FLIM is still an open problem and the exposure time allocation for measurements of different coding functions has not been explored yet. In this paper, we propose an end-to-end learnable framework to jointly optimize the coding functions, the allocation of exposure time for measurements of different coding functions, and the reconstruction networks. Specifically, we propose a differential forward model of FD-FLIM with learnable allocation of exposure time. Besides, we propose a transformer-based FLIM reconstruction network to retrieve accurate fluorescence lifetime. The effectiveness of the proposed FD-FLIM method is extensively validated on both simulated and real-world fluorescence lifetime datasets, and a prototype imaging system has also been built to further demonstrate the proposed method.
Jiaqu Li, Kanghui Wang, Tao Yue 0003
ICCP4
2024 An improved attentive residue multi-dilated network for thermal noise removal in magnetic resonance images
Tao Yue 0003
Image Vis. Comput.2
2023 Learnable Polarization-multiplexed Modulation Imager for Depth from Defocus
abstract
Estimating depth from a single snapshot image with defocus information is still a tricky problem for the ill-posedness introduced by the limited depth cues implied in the defocus images. This paper proposes a Polarization-multiplexed Modulation Imager (PoMI) to fully utilize the multiplexed polarization channels for capturing more depth cues with a single snapshot image. The polarization-dependent modulator, i.e., Liquid Crystal Spatial Light Modulator (LC-SLM), is applied to modulate the depth information into polarization channels. A differentiable polarization-dependent modulation camera model is proposed, combined with the Polarization-Driven Attention Network, to enable the joint system optimization by end-to-end training. Extensive tests have been applied to the synthetic datasets to verify the effectiveness of the proposed method. A system prototype is built to conduct real experiments demonstrating the feasibility of the proposed method for natural scenes.
Mingyou Dai, Tao Yue 0003
ICCP3
2023 Thermal Noise Removal of Magnetic Resonance Images: A Deep Learning Approach Based on an Attentive Residue Multi-Dilated Network with Adaptive Filtering and Discrete Cosine Transform
abstract
Magnetic resonance imaging (MRI) has been applied in various fields, especially for the medical purposes. However, fine details and smooth areas of some critical patterns in an MR image polluted by common thermal noise will interfere with the diagnosis of doctors. Thermal noise in an MR image obeys Rician distribution, which is hard for conventional denoising methods based on shift invariant spatial filtering approaches to dispose. Besides, the fine detail and edge information will be inevitably damaged when smoothing the noise, which is unacceptable for medical images. In this paper, we propose two corresponding solutions. First, we design a convolutional neural network (CNN) to learn a mask that aims to directly eliminate the thermal noise in the background region, and make the noise in the MR image obey almost the same distribution. Second, we propose several improvements in terms of existing deep learning approaches for thermal noise removal. Specifically, we establish a dual-branch neural network, a frequency-domain-optimizable discrete cosine transform (DCT) module, and adopt other effective structures such as residue blocks, convolutional block attention modules (CBAM), and parallel multi-dilated GoogLeNet inception based convolutional blocks to form an attentive residue multi-dilated network (ARM-Net). We evaluate our method over the BraTS 2018 dataset at noise levels ranging from 2% to 20%. Experimental results reveal that our method achieves state-of-the-art (SOTA) performance compared with the most recent works.
Tao Yue 0003
IJCNN2
2023 An Efficient Way for Active None-Line-of-Sight: End-to-End Learned Compressed NLOS Imaging
Chen Chang, Tao Yue 0003, Siqi Ni
PRCV (6)2
2023 Semi-Blindly Enhancing Extremely Noisy Videos With Recurrent Spatio-Temporal Large-Span Network
abstract
Capturing videos under the extremely dark environment is quite challenging for the extremely large and complex noise. To accurately represent the complex noise distribution, the physics-based noise modeling and learning-based blind noise modeling methods are proposed. However, these methods suffer from either the requirement of complex calibration procedure or performance degradation in practice. In this paper, we propose a semi-blind noise modeling and enhancing method, which incorporates the physics-based noise model with a learning-based Noise Analysis Module (NAM). With NAM, self-calibration of model parameters can be realized, which enables the denoising process to be adaptive to various noise distributions of either different cameras or camera settings. Besides, we develop a recurrent Spatio-Temporal Large-span Network (STLNet), constructed with a Slow-Fast Dual-branch (SFDB) architecture and an Interframe Non-local Correlation Guidance (INCG) mechanism, to fully investigate the spatio-temporal correlation in a large span. The effectiveness and superiority of the proposed method are demonstrated with extensive experiments, both qualitatively and quantitatively.
Tao Yue 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Fisher Information Guidance for Learned Time-of-Flight Imaging
abstract
Indirect Time-of-Flight (ToF) imaging is widely applied in practice for its superiorities on cost and spatial resolution. However, lower signal-to-noise ratio (SNR) of measurement leads to larger error in ToF imaging, especially for imaging scenes with strong ambient light or long distance. In this paper, we propose a Fisher-information guided framework to jointly optimize the coding functions (light modulation and sensor demodulation functions) and the reconstruction network of iToF imaging, with the super-vision of the proposed discriminative fisher loss. By introducing the differentiable modeling of physical imaging process considering various real factors and constraints, e.g., light-falloff with distance, physical implementability of coding functions, etc., followed by a dual-branch depth reconstruction neural network, the proposed method could learn the optimal iToF imaging system in an end-to-end manner. The effectiveness of the proposed method is extensively verified with both simulations and prototype experiments.
Jiaqu Li, Tao Yue 0003, Sijie Zhao
CVPR2
2021 Controlling the Rain: From Removal to Rendering
abstract
Existing rain image editing methods focus on either removing rain from rain images or rendering rain on rain-free images. This paper proposes to realize continuous control of rain intensity bidirectionally, from clear rain-free to downpour image with a single rain image as input, without changing the scene-specific characteristics, e.g. the direction, appearance and distribution of rain. Specifically, we introduce a Rain Intensity Controlling Network (RIC-Net) that contains three sub-networks of background extraction network, high-frequency rain-streak elimination network and main controlling network, which allows to control rain image of different intensities continuously by interpolation in the deep feature space. The HOG loss and autocorrelation loss are proposed to enhance consistency in orientation and suppress repetitive rain streaks. Furthermore, a decremental learning strategy that trains the network from downpour to drizzle images sequentially is proposed to further improve the performance and speedup the convergence. Extensive experiments on both rain dataset and real rain images demonstrate the effectiveness of the proposed method.
Siqi Ni, Xueyun Cao, Tao Yue 0003
CVPR3
2021 Distribution-Aware Adaptive Multi-Bit Quantization
abstract
In this paper, we explore the compression of deep neural networks by quantizing the weights and activations into multi-bit binary networks (MBNs). A distribution-aware multi-bit quantization (DMBQ) method that incorporates the distribution prior into the optimization of quantization is proposed. Instead of solving the optimization in each iteration, DMBQ search the optimal quantization scheme over the distribution space beforehand, and select the quantization scheme during training using a fast lookup table based strategy. Based upon DMBQ, we further propose loss-guided bit-width allocation (LBA) to adaptively quantize and even prune the neural network. The first-order Taylor expansion is applied to build a metric for evaluating the loss sensitivity of the quantization of each channel, and automatically adjust the bit-width of weights and activations channel-wisely. We extend our method to image classification tasks and experimental results show that our method not only outperforms state-of-the-art quantized networks in terms of accuracy but also is more efficient in terms of training time compared with state-of-the-art MBNs, even for the extremely low bit width (below 1-bit) quantization cases.
Sijie Zhao, Tao Yue 0003
CVPR2
2021 Fast Light-field Disparity Estimation with Multi-disparity-scale Cost Aggregation
abstract
Light field images contain both angular and spatial information of captured light rays. The rich information of light fields enables straightforward disparity recovery capability but demands high computational cost as well. In this paper, we design a lightweight disparity estimation model with physical-based multi-disparity-scale cost volume aggregation for fast disparity estimation. By introducing a sub-network of edge guidance, we significantly improve the recovery of geometric details near edges and improve the overall performance. We test the proposed model extensively on both synthetic and real-captured datasets, which provide both densely and sparsely sampled light fields. Finally, we significantly reduce computation cost and GPU memory consumption, while achieving comparable performance with state-of-the-art disparity estimation methods for light fields. Our source code is available at https://github.com/zcong17huang/FastLFnet.
Zhou Xue, Weizhu Xu, Tao Yue 0003
ICCV5
2019 Hyperspectral Imaging With Random Printed Mask
abstract
Hyperspectral images can provide rich clues for various computer vision tasks. However, the requirements of professional and expensive hardware for capturing hyperspectral images impede its wide applications. In this paper, based on a simple but not widely noticed phenomenon that the color printer can print color masks with a large number of independent spectral transmission responses, we propose a simple and low-budget scheme to capture the hyperspectral images with a random mask printed by the consumer-level color printer. Specifically, we notice that the printed dots with different colors are stacked together, forming multiplicative, instead of additive, spectral transmission responses. Therefore, new spectral transmission response uncorrelated with that of the original printer dyes are generated. With the random printed color mask, hyperspectral images could be captured in a snapshot way. A convolutional neural network (CNN) based method is developed to reconstruct the hyperspectral images from the captured image. The effectiveness and accuracy of the proposed system are verified on both synthetic and real captured images.
Zhan Ma 0001, Xun Cao, Tao Yue 0003
CVPR5
2019 Spectral Reconstruction From Dispersive Blur: A Novel Light Efficient Spectral Imager
abstract
Developing high light efficiency imaging techniques to retrieve high dimensional optical signal is a long-term goal in computational photography. Multispectral imaging, which captures images of different wavelengths and boosting the abilities for revealing scene properties, has developed rapidly in the last few decades. From scanning method to snapshot imaging, the limit of light collection efficiency is kept being pushed which enables wider applications especially under the light-starved scenes. In this work, we propose a novel multispectral imaging technique, that could capture the multispectral images with a high light efficiency. Through investigating the dispersive blur caused by spectral dispersers and introducing the difference of blur (DoB) constraints, we propose a basic theory for capturing multispectral information from a single dispersive-blurred image and an additional spectrum of an arbitrary point in the scene. Based on the theory, we design a prototype system and develop an optimization algorithm to realize snapshot multispectral imaging. The effectiveness of the proposed method is verified on both the synthetic data and real captured images.
Zhan Ma 0001, Tao Yue 0003, Xun Cao
CVPR5
2019 Enhancing Low Light Videos by Exploring High Sensitivity Camera Noise
abstract
Enhancing low light videos, which consists of denoising and brightness adjustment, is an intriguing but knotty problem. Under low light condition, due to high sensitivity camera setting, commonly negligible noises become obvious and severely deteriorate the captured videos. To recover high quality videos, a mass of image/video denoising/enhancing algorithms are proposed, most of which follow a set of simple assumptions about the statistic characters of camera noise, e.g., independent and identically distributed(i.i.d.), white, additive, Gaussian, Poisson or mixture noises. However, the practical noise under high sensitivity setting in real captured videos is complex and inaccurate to model with these assumptions. In this paper, we explore the physical origins of the practical high sensitivity noise in digital cameras, model them mathematically, and propose to enhance the low light videos based on the noise model by using an LSTM-based neural network. Specifically, we generate the training data with the proposed noise model and train the network with the dark noisy video as input and clear-bright video as output. Extensive comparisons on both synthetic and real captured low light videos with the state-of-the-art methods are conducted to demonstrate the effectiveness of the proposed method.
Tao Yue 0003
ICCV6
2018 Multispectral Image Intrinsic Decomposition via Subspace Constraint
abstract
Multispectral images contain many clues of surface characteristics of the objects, thus can be used in many computer vision tasks, e.g., recolorization and segmentation. However, due to the complex geometry structure of natural scenes, the spectra curves of the same surface can look very different under different illuminations and from different angles. In this paper, a new Multispectral Image Intrinsic Decomposition model (MIID) is presented to decompose the shading and reflectance from a single multispectral image. We extend the Retinex model, which is proposed for RGB image intrinsic decomposition, for multispectral domain. Based on this, a subspace constraint is introduced to both the shading and reflectance spectral space to reduce the ill-posedness of the problem and make the problem solvable. A dataset of 22 scenes is given with the ground truth of shadings and reflectance to facilitate objective evaluations. The experiments demonstrate the effectiveness of the proposed method.
Weixin Zhu, Linsen Chen, Yao Wang 0001, Tao Yue 0003, Xun Cao
CVPR6
2017 Multispectral focal stack acquisition using a chromatic aberration enlarged camera
abstract
Capturing more information, e.g. geometry and material, using optical cameras can greatly help the perception and understanding of complex scenes. This paper proposes a novel method to capture the spectral and light field information simultaneously. By using a delicately designed chromatic aberration enlarged camera, the spectral-varying slices at different depths of the scene can be easily captured. Afterwards, the multispectral focal stack, which is composed of a stack of multispectral slice images focusing on different depths, can be recovered from the spectral-varying slices by using a Local Linear Transformation (LLT) based algorithm. The experiments verify the effectiveness of the proposed method.
Yunqian Li, Linsen Chen, Xiaoming Zhong, Jin-Li Suo, Zhan Ma 0001, Tao Yue 0003, Xun Cao
ICIP7
2017 DeepCoder: A deep neural network based video compression
abstract
Inspired by recent advances in deep learning, we present the DeepCoder - a Convolutional Neural Network (CNN) based video compression framework. We apply separate CNN nets for predictive and residual signals respectively. Scalar quantization and Huffman coding are employed to encode the quantized feature maps (fMaps) into binary stream. We use the fixed 32 × 32 block in this work to demonstrate our ideas, and performance comparison is conducted with the well-known H.264/AVC video coding standard with comparable rate-distortion performance. Here distortion is measured using Structural Similarity (SSIM) because it is more close to perceptual response.
Tong Chen 0004, Qiu Shen, Tao Yue 0003, Xun Cao, Zhan Ma 0001
VCIP4
2017 A Practical System Towards the Secure, Robust and Pervasive Mobile Workstyle
abstract
We develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
VTC Spring2
2017 The role of prior in image based 3D modeling: a survey
Hao Zhu 0004, Yongming Nie, Tao Yue 0003, Xun Cao
Frontiers Comput. Sci.3
2017 Robust multi-view stereo synthesized by various parameters model
Yongming Nie, Tao Yue 0003, Hao Zhu 0004, Sidan Du, Xun Cao
J. Vis. Commun. Image Represent.2
2017 High-resolution spectral video acquisition
abstract
Compared with conventional cameras, spectral imagers provide many more features in the spectral domain. They have been used in various fields such as material identification, remote sensing, precision agriculture, and surveillance. Traditional imaging spectrometers use generally scanning systems. They cannot meet the demands of dynamic scenarios. This limits the practical applications for spectral imaging. Recently, with the rapid development in computational photography theory and semiconductor techniques, spectral video acquisition has become feasible. This paper aims to offer a review of the state-of-the-art spectral imaging technologies, especially those capable of capturing spectral videos. Finally, we evaluate the performances of the existing spectral acquisition systems and discuss the trends for future work.
Linsen Chen, Tao Yue 0003, Xun Cao, Zhan Ma 0001, David J. Brady
Frontiers Inf. Technol. Electron. Eng.2
2017 Efficient Method for High-Quality Removal of Nonuniform Blur in the Wavelet Domain
abstract
This paper presents a novel nonuniform deblurring approach, which defines the blur model and calculates regularized nonuniform deconvolution in the wavelet domain to achieve high efficiency and high accuracy simultaneously. Targeting high computation efficiency, we derive a wavelet-domain hierarchical blur model, which can be calculated efficiently by exploiting the sparsity property of natural images in the wavelet domain. Correspondingly, the blur model is incorporated into a multilayer framework and at each layer spatially varying step sizes are introduced to further accelerate the convergence of the algorithm. In addition to the efficiency advantages, the proposed approach deals with intensely nonuniform blur with high accuracy due to the intrinsic tight supportness of wavelet basis. We conduct a series of experiments and comparisons to validate the efficiency and effectiveness of our algorithm.
Tao Yue 0003, Jin-Li Suo, Xun Cao, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2017 Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle
abstract
In this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
IEEE Trans. Multim.2
2015 Blind optical aberration correction by exploring geometric and visual priors
abstract
Optical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric and visual priors. The global rotational symmetry allows us to transform the non-uniform degeneration into several uniform ones by the proposed radial splitting and warping technique. Locally, two types of symmetry constraints, i.e. central symmetry and reflection symmetry are defined as geometric priors in central and surrounding regions, respectively. Furthermore, by investigating the visual artifacts of aberration degenerated images captured by consumer-level cameras, the non-uniform distribution of sharpness across color channels and the image lattice is exploited as visual priors, resulting in a novel strategy to utilize the guidance from the sharpest channel and local image regions to improve the overall performance and robustness. Extensive evaluation on both real and synthetic data suggests that the proposed method outperforms the state-of-the-art techniques.
Tao Yue 0003, Jin-Li Suo, Jue Wang 0001, Xun Cao, Qionghai Dai
CVPR1
2014 Hybrid Image Deblurring by Fusing Edge and Power Spectrum Information
Tao Yue 0003, Sunghyun Cho, Jue Wang 0001, Qionghai Dai
ECCV (7)1
2014 Deblur a blurred RGB image with a sharp NIR image through local linear mapping
abstract
Image acquisition in a low light environment requires long exposure to achieve acceptable signal-to-noise ratio, which however causes blurry effect. This paper addresses this problem by using a sharp near-infrared (NIR) image when the environment has sufficient NIR light. We assume that an RGB and NIR image pair has a linear mapping in a local area and that the mapping function is valid for both the blur and sharp image pairs. Using this property, we solve the sharp RGB images from a blurred RGB image and the corres ponding s harp NIR image. The effectiveness of the proposed algorithm is verified with both synthetic and real captured datasets.
Tao Yue 0003, Ming-Ting Sun, Zhengyou Zhang, Jin-Li Suo, Qionghai Dai
ICME1
2014 High-Dimensional Camera Shake Removal With Given Depth Map
abstract
Camera motion blur is drastically nonuniform for large depth-range scenes, and the nonuniformity caused by camera translation is depth dependent but not the case for camera rotations. To restore the blurry images of large-depth-range scenes deteriorated by arbitrary camera motion, we build an image blur model considering 6-degrees of freedom (DoF) of camera motion with a given scene depth map. To make this 6D depth-aware model tractable, we propose a novel parametrization strategy to reduce the number of variables and an effective method to estimate high-dimensional camera motion as well. The number of variables is reduced by temporal sampling motion function, which describes the 6-DoF camera motion by sampling the camera trajectory uniformly in time domain. To effectively estimate the high-dimensional camera motion parameters, we construct the probabilistic motion density function (PMDF) to describe the probability distribution of camera poses during exposure, and apply it as a unified constraint to guide the convergence of the iterative deblurring algorithm. Specifically, PMDF is computed through a back projection from 2D local blur kernels to 6D camera motion parameter space and robust voting. We conduct a series of experiments on both synthetic and real captured data, and validate that our method achieves better performance than existing uniform methods and nonuniform methods on large-depth-range scenes.
Tao Yue 0003, Jin-Li Suo, Qionghai Dai
IEEE Trans. Image Process.1
2013 Non-uniform image deblurring using an optical computing system
Tao Yue 0003, Jin-Li Suo, Qionghai Dai
Comput. Graph.1