Stanley H. Chan

dblp:17/6278 · DBLP profile ↗
← Back
63ranked-venue papers
17as first author
25since 2021 · last 2025
0000-0001-5876-2073ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 56 · 16 first-author · 23 since 2021Artificial intelligence and machine learning · 23 · 2 first-author · 16 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
abstract
Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the generator will not be able to interpret and generate scene-consistent images. This limitation not only hinders the adoption of generative tools in professional photography but also highlights the broader challenge of aligning data-driven models with real-world physical settings. In this paper, we introduce Generative Photography, a framework that allows controlling camera intrinsic settings during content generation. The core innovation of this work are the concepts of Dimensionality Lifting and Differential Camera Intrinsics Learning, enabling smooth and consistent transitions across different camera settings. Experimental results show that our method produces significantly more scene-consistent photorealistic images than state-of-the-art models such as Stable Diffusion 3 and FLUX. Our code and additional results are available at https://generativephotography.github.io/project.
Yichen Sheng, Prateek Chennuri, Xingguang Zhang, Stanley H. Chan
CVPR6
2025 Learning Phase Distortion with Selective State Space Models for Video Turbulence Mitigation
abstract
Atmospheric turbulence is a major source of image degradation in long-range imaging systems. Although numerous deep learning-based turbulence mitigation (TM) methods have been proposed, many are slow, memory-hungry, and do not generalize well. In the spatial domain, methods based on convolutional operators have a limited receptive field, so they cannot handle a large spatial dependency required by turbulence. In the temporal domain, methods relying on self-attention can, in theory, leverage the lucky effects of turbulence, but their quadratic complexity makes it difficult to scale to many frames. Traditional recurrent aggregation methods face parallelization challenges.In this paper, we present a new TM method based on two concepts: (1) A turbulence mitigation network based on the Selective State Space Model (MambaTM). MambaTM provides a global receptive field in each layer across spatial and temporal dimensions while maintaining linear computational complexity. (2) Learned Latent Phase Distortion (LPD). LPD guides the state space model. Unlike classical Zernike-based representations of phase distortion, the new LPD map uniquely captures the actual effects of turbulence, significantly improving the model’s capability to estimate degradation by reducing the ill-posedness. Our proposed method exceeds current state-of-the-art networks on various synthetic and real-world TM benchmarks with significantly faster inference speed. The code is available at https://github.com/xg416/MambaTM.
Xingguang Zhang, Nicholas Chimitt, Stanley H. Chan
CVPR5
2025 Bayesian-Inspired Space-Time Superpixels
Kent Gauen, Stanley H. Chan
ICCV2
2025 Quanta Diffusion
abstract
We present Quanta Diffusion (QuDi), a powerful generative video reconstruction method for single-photon imaging. QuDI is an algorithm supporting the latest Quanta Image Sensors (QIS) and Single Photon Avalanche Diodes (SPADs) for extremely low-light imaging conditions. Compared to existing methods, QuDi overcomes the difficulties of simultaneously managing the motion and the strong shot noise. The core innovation of QuDi is to inject a physics-based forward model into the diffusion algorithm, while keeping the motion estimation in the loop. QuDi demonstrates an average of 2.4 dB PSNR improvement over the best existing methods.
Prateek Chennuri, Dongdong Fu, Stanley H. Chan
ICIP3
2025 Ultrafast High-Flux Single-Photon Lidar Simulator via Neural Mapping
abstract
Efficient simulation of photon registrations in single-photon LiDAR (SPL) is essential for applications such as depth estimation under high-flux conditions, where hardware dead time significantly distorts photon measurements. However, the conventional wisdom is computationally intensive due to their inherently sequential, photon-by-photon processing. In this paper, we propose a learning-based framework that accelerates the simulation process by modeling the photon count and directly predicting the photon registration probability density function (PDF) using an autoencoder (AE). Our method achieves high accuracy in estimating both the total number of registered photons and their temporal distribution, while substantially reducing simulation time. Extensive experiments validate the effectiveness and efficiency of our approach, highlighting its potential to enable fast and accurate SPL simulations for data-intensive imaging tasks in the high-flux regime.
Hashan K. Weerasooriya, Stanley H. Chan
ICIP3
2024 Resolution Limit of Single-Photon LiDAR
abstract
Single-photon Light Detection and Ranging (LiDAR) systems are often equipped with an array of detectors for improved spatial resolution and sensing speed. However, given a fixed amount of flux produced by the laser transmitter across the scene, the per-pixel Signal-to-Noise Ratio (SNR) will decrease when more pixels are packed in a unit space. This presents a fundamental trade-off between the spatial resolution of the sensor array and the SNR received at each pixel. Theoretical characterization of this fundamental limit is explored. By deriving the photon arrival statistics and introducing a series of new approximation techniques, the Mean Squared Error (MSE) of the maximum-likelihood estimator of the time delay is derived. The theoretical predictions align well with simulations and real data.
Stanley H. Chan, Hashan K. Weerasooriya, Pamela Abshire, István Gyöngy, Robert K. Henderson
CVPR1
2024 Generative Quanta Color Imaging
abstract
The astonishing development of single-photon cameras has created an unprecedented opportunity for scientific and industrial imaging. However, the high data throughput generated by these 1-bit sensors creates a significant bottleneck for low-power applications. In this paper, we explore the possibility of generating a color image from a single binary frame of a single-photon camera. We evidently find this problem being particularly difficult to standard colorization approaches due to the substantial degree of exposure variation. The core innovation of our paper is an exposure synthesis model framed under a neural ordinary differential equation (Neural ODE) that allows us to generate a contin-uum of exposures from a single observation. This innovation ensures consistent exposure in binary images that col-orizers take on, resulting in notably enhanced colorization. We demonstrate applications of the method in single-image and burst colorization and show superior generative performance over baselines. Project website can be found at https://vishal-s-p.github.io/projects/2023/generative_quanta_color.html
Vishal Purohit, Junjie Luo 0009, Yiheng Chi, Qi Guo 0009, Stanley H. Chan, Qiang Qiu 0001
CVPR5
2024 Spatio-Temporal Turbulence Mitigation: A Translational Perspective
abstract
Recovering images distorted by atmospheric turbulence is a challenging inverse problem due to the stochastic nature of turbulence. Although numerous turbulence mitigation (TM) algorithms have been proposed, their efficiency and generalization to real-world dynamic scenarios remain severely limited. Building upon the intuitions of classical TM algorithms, we present the Deep Atmospheric TUrbulence Mitigation network (DATUM). DATUM aims to overcome major challenges when transitioning from classical to deep learning approaches. By carefully integrating the merits of classical multi-frame TM methods into a deep network structure, we demonstrate that DATUM can efficiently perform long-range temporal aggregation using a recurrent fashion, while deformable attention and temporal-channel attention seamlessly facilitate pixel registration and lucky imaging. With additional supervision, tilt and blur degradation can be Jointly mitigated. These inductive biases empower DATUM to significantly outperform existing methods while delivering a tenfold increase in processing speed. A large-scale training dataset, ATSyn, is presented as a co-invention to enable the generalization to real turbulence. Our code and datasets are available at http://xg416.github.io/DATUM
Xingguang Zhang, Nicholas Chimitt, Yiheng Chi, Zhiyuan Mao, Stanley H. Chan
CVPR5
2024 Quanta Video Restoration
Prateek Chennuri, Yiheng Chi, Enze Jiang, G. M. Dilshan Godaliyadda, Abhiram Gnanasambandam, Hamid R. Sheikh, István Gyöngy, Stanley H. Chan
ECCV (40)8
2024 Kernel Diffusion: An Alternate Approach to Blind Deconvolution
Yash Sanghvi, Yiheng Chi, Stanley H. Chan
ECCV (59)3
2024 Analysis and Improvement of Rank-Ordered Mean Algorithm in Single-Photon LiDAR
abstract
Depth estimation using a single-photon LiDAR is often solved by a matched filter. It is, however, error-prone in the presence of background noise. A commonly used technique to reject background noise is the rank-ordered mean (ROM) filter previously reported by Shin et al. (2015). ROM rejects noisy photon arrival timestamps by selecting only a small range of them around the median statistics within its local neighborhood. Despite the promising performance of ROM, its theoretical performance limit is unknown. In this paper, we theoretically characterize the ROM performance by showing that ROM fails when the reflectivity drops below a threshold predetermined by the depth and signal-to-background ratio, and its accuracy undergoes a phase transition at the cutoff. Based on our theory, we propose an improved signal extraction technique by selecting tight timestamp clusters. Experimental results show that the proposed algorithm improves depth estimation performance over ROM by 3 orders of magnitude at the same signal intensities, and achieves high image fidelity at noise levels as high as 17 times that of signal. The code for this project is made available at https://github.com/yauwilliam69/NCFforLiDAR.git.
William C. Yau, Hashan K. Weerasooriya, Stanley H. Chan
MMSP4
2024 Parametric Modeling and Estimation of Photon Registrations for 3D Imaging
abstract
In single-photon light detection and ranging (SP-LiDAR) systems, the histogram distortion due to hardware dead time fundamentally limits the precision of depth estimation. To compensate for the dead time effects, the photon registration distribution is typically modeled based on the Markov chain self-exciting process. However, this is a discrete process and it is computationally expensive, thus hindering potential neural network applications and fast simulations. In this paper, we overcome the modeling challenge by proposing a continuous parametric model. We introduce a Gaussian-uniform mixture model (GUMM) and periodic padding to address high noise floors and noise slopes respectively. By deriving and implementing a customized expectation maximization (EM) algorithm, we achieve accurate histogram matching in scenarios that were deemed difficult in the literature.
Hashan K. Weerasooriya, Prateek Chennuri, Stanley H. Chan
MMSP4
2024 Soft Superpixel Neighborhood Attention
abstract
Images contain objects with deformable boundaries, such as the contours of a human face, yet attention operators act on square windows. This mixes features from perceptually unrelated regions, which can degrade the quality of a denoiser. One can exclude pixels using an estimate of perceptual groupings, such as superpixels, but the naive use of superpixels can be theoretically and empirically worse than standard attention. Using superpixel probabilities rather than superpixel assignments, this paper proposes soft superpixel neighborhood attention (SNA), which interpolates between the existing neighborhood attention and the naive superpixel neighborhood attention. This paper presents theoretical results showing SNA is the optimal denoiser under a latent superpixel model. SNA outperforms alternative local attention modules on image denoising, and we compare the superpixels learned from denoising with those learned with supervision.
Kent Gauen, Stanley H. Chan
NeurIPS2
2024 FarSight: A Physics-Driven Whole-Body Biometric System at Large Distance and Altitude
abstract
Whole-body biometric recognition is an important area of research due to its vast applications in law enforcement, border security, and surveillance. This paper presents the end-to-end design, development and evaluation of FarSight, an innovative software system designed for whole-body (fusion of face, gait and body shape) biometric recognition. FarSight accepts videos from elevated platforms and drones as input and outputs a candidate list of identities from a gallery. The system is designed to address several challenges, including (i) low-quality imagery, (ii) large yaw and pitch angles, (iii) robust feature extraction to accommodate large intra-person variabilities and large inter-person similarities, and (iv) the large domain gap between training and test sets. FarSight combines the physics of imaging and deep learning models to enhance image restoration and biometric feature encoding. We test FarSight’s effectiveness using the newly acquired IARPA Biometric Recognition and Identification at Altitude and Range (BRIAR) dataset. Notably, FarSight demonstrated a substantial performance increase on the BRIAR dataset, with gains of +11.82% Rank-20 identification and +11.30% TAR@1% FAR.
Feng Liu 0037, Ryan Ashbaugh, Nicholas Chimitt, Najmul Hassan, Ali Hassani 0001, Ajay Jaiswal, Zhiyuan Mao, Christopher Perry, Yiyang Su, Pegah Varghaei, Kai Wang 0058, Stanley H. Chan, Arun Ross, Humphrey Shi, Zhangyang Wang, Xiaoming Liu 0002
WACV14
2023 HDR Imaging with Spatially Varying Signal-to-Noise Ratios
abstract
While today's high dynamic range (HDR) image fusion algorithms are capable of blending multiple exposures, the acquisition is often controlled so that the dynamic range within one exposure is narrow. For HDR imaging in photon-limited situations, the dynamic range can be enormous and the noise within one exposure is spatially varying. Existing image denoising algorithms and HDR fusion algorithms both fail to handle this situation, leading to severe limitations in low-light HDR imaging. This paper presents two contributions. Firstly, we identify the source of the problem. We find that the issue is associated with the co-existence of (1) spatially varying signal-to-noise ratio, especially the excessive noise due to very dark regions, and (2) a wide luminance range within each exposure. We show that while the issue can be handled by a bank of denoisers, the complexity is high. Secondly, we propose a new method called the spatially varying high dynamic range (SV-HDR) fusion network to simultaneously denoise and fuse images. We introduce a new exposure-shared block within our custom-designed multi-scale transformer framework. In a variety of testing conditions, the performance of the proposed SV-HDR is better than the existing methods.
Yiheng Chi, Xingguang Zhang, Stanley H. Chan
CVPR3
2023 Structured Kernel Estimation for Photon-Limited Deconvolution
abstract
Images taken in a low light condition with the presence of camera shake suffer from motion blur and photon shot noise. While state-of-the-art image restoration networks show promising results, they are largely limited to well-illuminated scenes and their performance drops significantly when photon shot noise is strong. In this paper, we propose a new blur estimation technique customized for photon-limited conditions. The proposed method employs a gradient-based backpropagation method to estimate the blur kernel. By modeling the blur kernel using a low-dimensional representation with the key points on the motion trajectory, we significantly reduce the search space and improve the regularity of the kernel estimation problem. When plugged into an iterative framework, our novel low-dimensional representation provides improved kernel estimates and hence significantly better deconvolution performance when compared to end-to-end trained neural networks. The source code and pretrained mdoels are available at https://github.com/sanghviyashiitb/structured-kernel-cvpr23
Yash Sanghvi, Zhiyuan Mao, Stanley H. Chan
CVPR3
2023 Physics-Driven Turbulence Image Restoration with Stochastic Refinement
abstract
Image distortion by atmospheric turbulence is a stochastic degradation, which is a critical problem in long-range optical imaging systems. A number of research has been conducted during the past decades, including model-based and emerging deep-learning solutions with the help of synthetic data. Although fast and physics-grounded simulation tools have been introduced to help the deep-learning models adapt to real-world turbulence conditions recently, the training of such models only relies on the synthetic data and ground truth pairs. This paper proposes the Physics-integrated Restoration Network (PiRN) to bring the physics-based simulator directly into the training process to help the network to disentangle the stochasticity from the degradation and the underlying image. Furthermore, to overcome the "average effect" introduced by deterministic models and the domain gap between the synthetic and real-world degradation, we further introduce PiRN with Stochastic Refinement (PiRN-SR) to boost its perceptual quality. Overall, our PiRN and PiRN-SR improve the generalization to real-world unknown turbulence conditions and provide a state-of-the-art restoration in both pixel-wise accuracy and perceptual quality. Our codes are available at https://github.com/VITA-Group/PiRN.
Ajay Jaiswal, Xingguang Zhang, Stanley H. Chan, Zhangyang Wang
ICCV3
2023 DROID: Driver-Centric Risk Object Identification
abstract
Identification of high-risk driving situations is generally approached through collision risk estimation or accident pattern recognition. In this work, we approach the problem from the perspective of subjective risk. We operationalize subjective risk assessment by predicting driver behavior changes and identifying the cause of changes. To this end, we introduce a new task called driver-centric risk object identification (DROID), which uses egocentric video to identify object(s) influencing a driver's behavior, given only the driver's response as the supervision signal. We formulate the task as a cause-effect problem and present a novel two-stage DROID framework, taking inspiration from models of situation awareness and causal inference. A subset of data constructed from the Honda Research Institute Driving Dataset (HDD) is used to evaluate DROID. We demonstrate state-of-the-art DROID performance, even compared with strong baseline models using this dataset. Additionally, we conduct extensive ablative studies to justify our design choices. Moreover, we demonstrate the applicability of DROID for risk assessment.
Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Single Frame Atmospheric Turbulence Mitigation: A Benchmark Study and a New Physics-Inspired Transformer Model
Zhiyuan Mao, Ajay Jaiswal, Zhangyang Wang, Stanley H. Chan
ECCV (19)4
2022 Photon-Limited Deblurring Using Algorithm Unrolling
abstract
Image deblurring in a photon-limited condition is ubiquitous in a variety of low-light applications such as photography, microscopy and astronomy. However, the presence of photon shot noise due to a low-illumination and/or short exposure makes the deblurring task substantially more challenging than conventional deblurring. In this paper we present an algorithm unrolling approach that unrolls a Plug-and-Play algorithm using a fixed-iteration network. By changing the conventional two-variable splitting formulation of Plug-and-Play to an alternate three-variable splitting, we obtain a differentiable and end-to-end trainable network. Our algorithm outperforms existing methods for 1dB across different illuminations. We also overcome the difficulty of acquiring real motion blur kernels at low-light by presenting a photon-limited motion deblurring image dataset.
Yash Sanghvi, Abhiram Gnanasambandan, Stanley H. Chan
ICASSP3
2022 Tilt-Then-Blur or Blur-Then-Tilt? Clarifying the Atmospheric Turbulence Model
abstract
Imaging at a long distance often requires advanced image restoration algorithms to compensate for the distortions caused by atmospheric turbulence. However, unlike many standard restoration problems such as deconvolution, the forward image formation model of the atmospheric turbulence does not have a simple expression. Thanks to the Zernike representation of the phase, one can show that the forward model is a combination of tilt (pixel shifting due to the linear phase terms) and blur (image smoothing due to the high order aberrations). Confusions then arise between the ordering of the two operators. Should the model be tilt-then-blur, or blur-then-tilt? Some papers in the literature say that the model is tilt-then-blur, whereas more papers say that it is blur-then-tilt. This paper clarifies the differences between the two and discusses why the tilt-then-blur is the correct model. Recommendations are given to the research community.
Stanley H. Chan
IEEE Signal Process. Lett.1
2022 Graph-Based Depth Denoising & Dequantization for Point Cloud Enhancement
abstract
A 3D point cloud is typically constructed from depth measurements acquired by sensors at one or more viewpoints. The measurements suffer from both quantization and noise corruption. To improve quality, previous works denoise a point cloud a posteriori after projecting the imperfect depth data onto 3D space. Instead, we enhance depth measurements directly on the sensed images a priori, before synthesizing a 3D point cloud. By enhancing near the physical sensing process, we tailor our optimization to our depth formation model before subsequent processing steps that obscure measurement errors. Specifically, we model depth formation as a combined process of signal-dependent noise addition and non-uniform log-based quantization. The designed model is validated (with parameters fitted) using collected empirical data from a representative depth sensor. To enhance each pixel row in a depth image, we first encode intra-view similarities between available row pixels as edge weights via feature graph learning. We next establish inter-view similarities with another rectified depth image via viewpoint mapping and sparse linear interpolation. This leads to a maximum a posteriori (MAP) graph filtering objective that is convex and differentiable. We minimize the objective efficiently using accelerated gradient descent (AGD), where the optimal step size is approximated via Gershgorin circle theorem (GCT). Experiments show that our method significantly outperformed recent point cloud denoising schemes and state-of-the-art image denoising schemes in two established point cloud quality metrics.
Xue Zhang 0008, Gene Cheung, Jiahao Pang, Yash Sanghvi, Abhiram Gnanasambandam, Stanley H. Chan
IEEE Trans. Image Process.6
2021 Student-Teacher Learning From Clean Inputs to Noisy Inputs
abstract
Feature-based student-teacher learning, a training method that encourages the student’s hidden features to mimic those of the teacher network, is empirically successful in transferring the knowledge from a pre-trained teacher network to the student network. Furthermore, recent empirical results demonstrate that, the teacher’s features can boost the student network’s generalization even when the student’s input sample is corrupted by noise. However, there is a lack of theoretical insights into why and when this method of transferring knowledge can be successful between such heterogeneous tasks. We analyze this method theoretically using deep linear networks, and experimentally using nonlinear networks. We identify three vital factors to the success of the method: (1) whether the student is trained to zero training loss; (2) how knowledgeable the teacher is on the clean-input problem; (3) how the teacher decomposes its knowledge in its hidden features. Lack of proper control in any of the three factors leads to failure of the student-teacher learning method.
Guanzhe Hong, Zhiyuan Mao, Xiaojun Lin 0001, Stanley H. Chan
CVPR4
2021 Graph Signal Denoising Using Nested-Structured Deep Algorithm Unrolling
abstract
In this paper, we propose a deep algorithm unrolling (DAU) based on a variant of the alternating direction method of multiplier (ADMM) called Plug-and-Play ADMM (PnP-ADMM) for denoising of signals on graphs. DAU is a trainable deep architecture realized by unrolling iterations of an existing optimization algorithm which contains trainable parameters at each layer. We also propose a nested-structured DAU: Its submodules in the unrolled iterations are also designed by DAU. Several experiments for graph signal denoising are performed on synthetic signals on a community graph and U.S. temperature data to validate the proposed approach. Our proposed method outperforms alternative optimization- and deep learning-based approaches.
Masatoshi Nagahama, Koki Yamada, Yuichi Tanaka 0001, Stanley H. Chan, Yonina C. Eldar
ICASSP4
2021 Accelerating Atmospheric Turbulence Simulation via Learned Phase-to-Space Transform
abstract
Fast and accurate simulation of imaging through atmospheric turbulence is essential for developing turbulence mitigation algorithms. Recognizing the limitations of previous approaches, we introduce a new concept known as the phase-to-space (P2S) transform to significantly speed up the simulation. P2S is built upon three ideas: (1) reformulating the spatially varying convolution as a set of invariant convolutions with basis functions, (2) learning the basis function via the known turbulence statistics models, (3) implementing the P2S transform via a light-weight network that directly converts the phase representation to spatial representation. The new simulator offers 300× – 1000× speed up compared to the mainstream split-step simulators while preserving the essential turbulence statistics.
Zhiyuan Mao, Nicholas Chimitt, Stanley H. Chan
ICCV3
2020 Dynamic Low-Light Imaging with Quanta Image Sensors
Yiheng Chi, Abhiram Gnanasambandam, Vladlen Koltun, Stanley H. Chan
ECCV (21)4
2020 Image Classification in the Dark Using Quanta Image Sensors
Abhiram Gnanasambandam, Stanley H. Chan
ECCV (8)2
2020 Simulating Anisoplanatic Turbulence by Sampling Correlated Zernike Coefficients
abstract
Simulating atmospheric turbulence is an essential task for evaluating turbulence mitigation algorithms and training learning-based methods. Advanced numerical simulators for atmospheric turbulence are available, but they require sophisticated wave propagations which are computationally very expensive. In this paper, we present a propagation-free method for simulating imaging through anisoplanatic atmospheric turbulence. The key innovation that enables this work is a new method to draw spatially correlated tilts and high-order abberations in the Zernike space. By establishing the equivalence between the angle-of-arrival correlation by Basu, McCrae and Fiorino (2015) and the multi-aperture correlation by Chanan (1992), we show that the Zernike coefficients can be drawn according to a covariance matrix defining the spatial correlations. We propose fast and scalable sampling strategies to draw these samples. The new method allows us to compress the wave propagation problem into a sampling problem, hence making the new simulator significantly faster than existing ones. Experimental results show that the simulator has an excellent match with the theory and real turbulence data.
Nicholas Chimitt, Stanley H. Chan
ICCP2
2020 One Size Fits All: Can We Train One Denoiser for All Noise Levels?
abstract
When training an estimator such as a neural network for tasks like image denoising, it is often preferred to train one estimator and apply it to all noise levels. The de facto training protocol to achieve this goal is to train the estimator with noisy samples whose noise levels are uniformly distributed across the range of interest. However, why should we allocate the samples uniformly? Can we have more training samples that are less noisy, and fewer samples that are more noisy? What is the optimal distribution? How do we obtain such a distribution? The goal of this paper is to address this training sample distribution problem from a minimax risk optimization perspective. We derive a dual ascent algorithm to determine the optimal sampling distribution of which the convergence is guaranteed as long as the set of admissible estimators is closed and convex. For estimators with non-convex admissible sets such as deep neural networks, our dual formulation converges to a solution of the convex relaxation. We discuss how the algorithm can be implemented in practice. We evaluate the algorithm on linear estimators and deep networks.
Abhiram Gnanasambandam, Stanley H. Chan
ICML2
2020 Learning 3D-aware Egocentric Spatial-Temporal Interaction via Graph Convolutional Networks
abstract
To enable intelligent automated driving systems, a promising strategy is to understand how human drives and interacts with road users in complicated driving situations. In this paper, we propose a 3D-aware egocentric spatial-temporal interaction framework for automated driving applications. Graph convolution networks (GCN) is devised for interaction modeling. We introduce three novel concepts into GCN. First, we decompose egocentric interactions into ego-thing and ego- stuff interaction, modeled by two GCNs. In both GCNs, ego nodes are introduced to encode the interaction between thing objects (e.g., car and pedestrian), and interaction between stuff objects (e.g., lane marking and traffic light). Second, objects' 3D locations are explicitly incorporated into GCN to better model egocentric interactions. Third, to implement ego-stuff interaction in GCN, we propose a MaskAlign operation to extract features for irregular objects.We validate the proposed framework on tactical driver behavior recognition. Extensive experiments are conducted using Honda Research Institute Driving Dataset, the largest dataset with diverse tactical driver behavior annotations. Our framework demonstrates substantial performance boost over baselines on the two experimental settings by 3.9% and 6.0%, respectively. Furthermore, we visualize the learned affinity matrices, which encode ego-thing and ego-stuff interactions, to showcase the proposed framework can capture interactions effectively.
Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001
ICRA3
2020 Who Make Drivers Stop? Towards Driver-centric Risk Assessment: Risk Object Identification via Causal Inference
abstract
A significant amount of people die in road accidents due to driver errors. To reduce fatalities, developing intelligent driving systems assisting drivers to identify potential risks is in an urgent need. Risky situations are generally defined based on collision prediction in the existing works. However, collision is only a source of potential risks, and a more generic definition is required. In this work, we propose a novel driver-centric definition of risk, i.e., objects influencing drivers' behavior are risky. A new task called risk object identification is introduced. We formulate the task as the cause-effect problem and present a novel two-stage risk object identification framework based on causal inference with the proposed object-level manipulable driving model. We demonstrate favorable performance on risk object identification compared with strong baselines on the Honda Research Institute Driving Dataset (HDD). Our framework achieves a substantial average performance boost over a strong baseline by 7.5%.
Chengxi Li 0006, Stanley H. Chan, Yi-Ting Chen 0001
IROS2
2020 Automatic foreground extraction from imperfect backgrounds using multi-agent consensus equilibrium
Xiran Wang, Jason Juang, Stanley H. Chan
J. Vis. Commun. Image Represent.3
2019 Interpolation and Denoising of Graph Signals Using Plug-and-play Admm
abstract
Signals defined on a network or a graph are often prone to errors due to missing data and noise. In order to restore the graph signal, interpolation and denoising are two necessary steps along with other graph signal processing procedures. However, existing graph signal interpolation and denoising methods are largely decoupled due to the opposite objectives of the two tasks and the inherent high computational complexity. The goal of this paper is to integrate graph interpolation and denoising using the Plug-and-Play (PnP) ADMM, a recently developed technique in image processing. When using the subsampling process as the forward model and graph filter as the denoiser, we show that PnP ADMM is equivalent to interpolating a bandlimited signal. Preliminary results are demonstrated via experiments, where the proposed method shows significantly better performance over existing methods.
Yoshinao Yazaki, Yuichi Tanaka 0001, Stanley H. Chan
ICASSP3
2019 Optimal Combination of Image Denoisers
abstract
Given a set of image denoisers, each having a different denoising capability, is there a provably optimal way of combining these denoisers to produce an overall better result? An answer to this question is fundamental to designing an ensemble of weak estimators for complex scenes. In this paper, we present an optimal combination scheme by leveraging the deep neural networks and the convex optimization. The proposed framework, called the Consensus Neural Network (CsNet), introduces three new concepts in image denoising: 1) a provably optimal procedure to combine the denoised outputs via convex optimization; 2) a deep neural network to estimate the mean squared error (MSE) of denoised images without needing the ground truths; and 3) an image boosting procedure using a deep neural network to improve the contrast and to recover the lost details of the combined images. Experimental results show that CsNet can consistently improve the denoising performance for both deterministic and neural network denoisers.
Joon Hee Choi, Omar A. Elgendy, Stanley H. Chan
IEEE Trans. Image Process.3
2018 Fast And Robust Recursive Filter for Image Denoising
abstract
Image denoising on mobile cameras requires low complexity, but many state-of-the-art denoising methods are computationally intensive. We present a low complexity denoising algorithm using an edge-aware recursive filter (RF). We make two contributions. First, we modify the original RF so that it is significantly more robust when estimating the gradients from noisy inputs. We extend the RF to high-order for texture and heavy noise images. Second, we introduce a SURE-based image fusion technique. We show that while individual RFs have different performance, the fused result is often better. Experimental results show that the new RF performs much faster than other denoisers while providing good quality images.
Yiheng Chi, Stanley H. Chan
ICASSP2
2018 Image Reconstruction for Quanta Image Sensors Using Deep Neural Networks
abstract
Quanta Image Sensor (QIS) is a single-photon image sensor that oversamples the light field to generate binary measurements. Its single-photon sensitivity makes it an ideal candidate for the next generation image sensor after CMOS. However, image reconstruction of the sensor remains a challenging issue. Existing image reconstruction algorithms are largely based on optimization. In this paper, we present the first deep neural network approach for QIS image reconstruction. Our deep neural network takes the binary bit stream of QIS as input, learns the nonlinear transformation and denoising simultaneously. Experimental results show that the proposed network produces significantly better reconstruction results compared to existing methods.
Joon Hee Choi, Omar A. Elgendy, Stanley H. Chan
ICASSP3
2018 Plug-and-Play Unplugged: Optimization-Free Reconstruction Using Consensus Equilibrium
abstract
Regularized inversion methods for image reconstruction are used widely due to their tractability and their ability to combine complex physical sensor models with useful regularity criteria. Such methods motivated the recently developed Plug-and-Play prior method, which provides a framework to use advanced denoising algorithms as regularizers in inversion. However, the need to formulate regularized inversion as the solution to an optimization problem limits the expressiveness of possible regularity conditions and physical sensor models. In this paper, we introduce the idea of consensus equilibrium (CE), which generalizes regularized inversion to include a much wider variety of both forward (or data fidelity) components and prior (or regularity) components without the need for either to be expressed using a cost function. CE is based on the solution of a set of equilibrium equations that balance data fit and regularity. In this framework, the problem of MAP estimation in regularized inversion is replaced by the problem of solving these equilibrium equations, which can be approached in multiple ways. The key contribution of CE is to provide a novel framework for fusing multiple heterogeneous models of physical sensors or models learned from data. We describe the derivation of the CE equations and prove that the solution of the CE equations generalizes the standard MAP estimate under appropriate circumstances. We also discuss algorithms for solving the CE equations, including a version of the Douglas--Rachford/alternating direction method of multipliers algorithm with a novel form of preconditioning and Newton's method, both standard form and a Jacobian-free form using Krylov subspaces. We give several examples to illustrate the idea of CE and the convergence properties of these algorithms and demonstrate this method on some toy problems and on a denoising example in which we use an array of convolutional neural network denoisers, none of which is tuned to match the noise level in a noisy image but which in consensus can achieve a better result than any of them individually.
Gregery T. Buzzard, Stanley H. Chan, Suhas Sreehari, Charles A. Bouman
SIAM J. Imaging Sci.2
2017 Resolution enhancement for hyperspectral images: A super-resolution and fusion approach
abstract
Many remote sensing applications require a high-resolution hyperspectral image. However, resolutions of most hyperspectral imagers are limited to tens of meters. Existing resolution enhancement techniques either acquire additional multispectral band images or use a pan band image. The former poses hardware challenges, whereas the latter has limited performance. In this paper, we present a new resolution enhancement method that only requires a color image. Our approach integrates two newly developed techniques in the area: (1) A hybrid color mapping algorithm, and (2) A Plug-and-Play algorithm for single image super-resolution. Comprehensive experiments using real hyperspectral images are conducted to validate and evaluate the proposed method.
Chiman Kwan, Joon Hee Choi, Stanley H. Chan, Jin Zhou 0005, Bence Budavari
ICASSP3
2017 Parameter-free Plug-and-Play ADMM for image restoration
abstract
Plug-and-Play ADMM is a recently developed variation of the classical ADMM algorithm that replaces one of the subproblems using an off-the-shelf image denoiser. Despite its apparently ad-hoc nature, Plug-and-Play ADMM produces surprisingly good image recovery results. However, since in Plug-and-Play ADMM the denoiser is treated as a black-box, behavior of the overall algorithm is largely unknown. In particular, the internal parameter that controls the rate of convergence of the algorithm has to be adjusted by the user, and a bad choice of the parameter can lead to severe degradation of the result. In this paper, we present a parameter-free Plug-and-Play ADMMwhere internal parameters are updated as part of the optimization. Our algorithm is derived from the generalized approximate message passing, with several essential modifications. Experimentally, we find that the new algorithm produces solutions along a reliable and fast converging path.
Xiran Wang, Stanley H. Chan
ICASSP2
2017 Understanding Symmetric Smoothing Filters: A Gaussian Mixture Model Perspective
abstract
Many patch-based image denoising algorithms can be formulated as applying a smoothing filter to the noisy image. Expressed as matrices, the smoothing filters must be row normalized, so that each row sums to unity. Surprisingly, if we apply a column normalization before the row normalization, the performance of the smoothing filter can often be significantly improved. Prior works showed that such performance gain is related to the Sinkhorn-Knopp balancing algorithm, an iterative procedure that symmetrizes a row-stochastic matrix to a doubly stochastic matrix. However, a complete understanding of the performance gain phenomenon is still lacking. In this paper, we study the performance gain phenomenon from a statistical learning perspective. We show that Sinkhorn-Knopp is equivalent to an expectation-maximization (EM) algorithm of learning a Gaussian mixture model of the image patches. By establishing the correspondence between the steps of Sinkhorn-Knopp and the EM algorithm, we provide a geometrical interpretation of the symmetrization process. This observation allows us to develop a new denoising algorithm called Gaussian mixture model symmetric smoothing filter (GSF). GSF is an extension of the Sinkhorn-Knopp and is a generalization of the original smoothing filters. Despite its simple formulation, GSF outperforms many existing smoothing filters and has a similar performance compared with several state-of-the-art denoising algorithms.
Stanley H. Chan, Todd E. Zickler, Yue M. Lu
IEEE Trans. Image Process.1
2016 Image reconstruction and threshold design for Quanta Image Sensors
abstract
Quanta Image Sensor (QIS) has been envisioned as a candidate solution for next generation image sensors. We provide two new contributions to the signal processing aspects of QIS. First, we develop an image reconstruction algorithm to recover the underlying images from the QIS data, which is a massive array of binarized Poisson random variables. The new algorithm supersedes existing methods by enabling arbitrary threshold level. Second, we present a threshold design scheme to adaptively update the threshold level for optimal image reconstruction. We discuss the existence of a phase transition in determining the optimal threshold. Experimental results on tone-mapped high dynamic range images validates the effectiveness of the threshold scheme and the image reconstruction algorithm.
Omar A. Elgendy, Stanley H. Chan
ICIP2
2016 Adaptive Image Denoising by Mixture Adaptation
abstract
We propose an adaptive learning procedure to learn patch-based image priors for image denoising. The new algorithm, called the expectation-maximization (EM) adaptation, takes a generic prior learned from a generic external database and adapts it to the noisy image to generate a specific prior. Different from existing methods that combine internal and external statistics in ad hoc ways, the proposed algorithm is rigorously derived from a Bayesian hyper-prior perspective. There are two contributions of this paper. First, we provide full derivation of the EM adaptation algorithm and demonstrate methods to improve the computational complexity. Second, in the absence of the latent clean image, we show how EM adaptation can be modified based on pre-filtering. The experimental results show that the proposed adaptation algorithm yields consistently better denoising results than the one without adaptation and is superior to several state-of-the-art algorithms.
Enming Luo, Stanley H. Chan, Truong Q. Nguyen
IEEE Trans. Image Process.2
2015 Understanding symmetric smoothing filters via Gaussian mixtures
abstract
We study a class of smoothing filters for image denoising. Expressed as matrices, these smoothing filters must be row normalized so that each row sums to unity. Surprisingly, if one applies a column normalization to the matrix before the row normalization, the denoising quality can often be significantly improved. This column-row normalization corresponds to one iteration of a symmetrization process called the Sinkhorn-Knopp balancing algorithm. However, a complete understanding of the performance gain phenomenon is lacking. In this paper, we analyze the performance gain from a Gaussian mixture model (GMM) perspective. We show that the symmetrization is equivalent to an expectation-maximization (EM) algorithm for learning the GMM. Moreover, we make modifications to the symmetrization procedure and present a new denoising algorithm. Experimental results show that the new algorithm achieves comparable denoising results to some state-of-the-art methods.
Stanley H. Chan, Todd E. Zickler, Yue M. Lu
ICIP1
2015 Depth Reconstruction From Sparse Samples: Representation, Algorithm, and Sampling
abstract
The rapid development of 3D technology and computer vision applications has motivated a thrust of methodologies for depth acquisition and estimation. However, existing hardware and software acquisition methods have limited performance due to poor depth precision, low resolution, and high computational cost. In this paper, we present a computationally efficient method to estimate dense depth maps from sparse measurements. There are three main contributions. First, we provide empirical evidence that depth maps can be encoded much more sparsely than natural images using common dictionaries, such as wavelets and contourlets. We also show that a combined wavelet-contourlet dictionary achieves better performance than using either dictionary alone. Second, we propose an alternating direction method of multipliers (ADMM) for depth map reconstruction. A multiscale warm start procedure is proposed to speed up the convergence. Third, we propose a two-stage randomized sampling scheme to optimally choose the sampling locations, thus maximizing the reconstruction performance for a given sampling budget. Experimental results show that the proposed method produces high-quality dense depth estimates, and is robust to noisy measurements. Applications to real data in stereo matching are demonstrated.
Lee-Kang Liu, Stanley H. Chan, Truong Q. Nguyen
IEEE Trans. Image Process.2
2015 Adaptive Image Denoising by Targeted Databases
abstract
We propose a data-dependent denoising procedure to restore noisy images. Different from existing denoising algorithms which search for patches from either the noisy image or a generic database, the new algorithm finds patches from a database that contains relevant patches. We formulate the denoising problem as an optimal filter design problem and make two contributions. First, we determine the basis function of the denoising filter by solving a group sparsity minimization problem. The optimization formulation generalizes existing denoising algorithms and offers systematic analysis of the performance. Improvement methods are proposed to enhance the patch search process. Second, we determine the spectral coefficients of the denoising filter by considering a localized Bayesian prior. The localized prior leverages the similarity of the targeted database, alleviates the intensive Bayesian computation, and links the new method to the classical linear minimum mean squared error estimation. We demonstrate applications of the proposed method in a variety of scenarios, including text images, multiview images, and face images. Experimental results show the superiority of the new algorithm over existing methods.
Enming Luo, Stanley H. Chan, Truong Q. Nguyen
IEEE Trans. Image Process.2
2014 Image denoising by targeted external databases
abstract
Classical image denoising algorithms based on single noisy images and generic image databases will soon reach their performance limits. In this paper, we propose to denoise images using targeted external image databases. Formulating denoising as an optimal filter design problem, we utilize the targeted databases to (1) determine the basis functions of the optimal filter by means of group sparsity; (2) determine the spectral coefficients of the optimal filter by means of localized priors. For a variety of scenarios such as text images, multiview images, and face images, we demonstrate superior denoising results over existing algorithms.
Enming Luo, Stanley H. Chan, Truong Q. Nguyen
ICASSP2
2014 A Consistent Histogram Estimator for Exchangeable Graph Models
abstract
Exchangeable graph models (ExGM) subsume a number of popular network models. The mathematical object that characterizes an ExGM is termed a graphon. Finding scalable estimators of graphons, provably consistent, remains an open issue. In this paper, we propose a histogram estimator of a graphon that is provably consistent and numerically efficient. The proposed estimator is based on a sorting-and-smoothing (SAS) algorithm, which first sorts the empirical degree of a graph, then smooths the sorted graph using total variation minimization. The consistency of the SAS algorithm is proved by leveraging sparsity concepts from compressed sensing.
Stanley H. Chan, Edoardo M. Airoldi
ICML1
2014 Monte Carlo Non-Local Means: Random Sampling for Large-Scale Image Filtering
abstract
We propose a randomized version of the nonlocal means (NLM) algorithm for large-scale image filtering. The new algorithm, called Monte Carlo nonlocal means (MCNLM), speeds up the classical NLM by computing a small subset of image patch distances, which are randomly selected according to a designed sampling pattern. We make two contributions. First, we analyze the performance of the MCNLM algorithm and show that, for large images or large external image databases, the random outcomes of MCNLM are tightly concentrated around the deterministic full NLM result. In particular, our error probability bounds show that, at any given sampling ratio, the probability for MCNLM to have a large deviation from the original NLM solution decays exponentially as the size of the image or database grows. Second, we derive explicit formulas for optimal sampling patterns that minimize the error probability bound by exploiting partial knowledge of the pairwise similarity weights. Numerical experiments show that MCNLM is competitive with other state-of-the-art fast NLM algorithms for single-image denoising. When applied to denoising images using an external database containing ten billion patches, MCNLM returns a randomized solution that is within 0.2 dB of the full NLM solution while reducing the runtime by three orders of magnitude.
Stanley H. Chan, Todd E. Zickler, Yue M. Lu
IEEE Trans. Image Process.1
2013 Fast non-local filtering by random sampling: It works, especially for large images
abstract
Non-local means (NLM) is a popular denoising scheme. Conceptually simple, the algorithm is computationally intensive for large images. We propose to speed up NLM by using random sampling. Our algorithm picks, uniformly at random, a small number of columns of the weight matrix, and uses these “representatives” to compute an approximate result. It also incorporates an extra column-normalization of the sampled columns, a form of symmetrization that often boosts the denoising performance on real images. Using statistical large deviation theory, we analyze the proposed algorithm and provide guarantees on its performance. We show that the probability of having a large approximation error decays exponentially as the image size increases. Thus, for large images, the random estimates generated by the algorithm are tightly concentrated around their limit values, even if the sampling ratio is small. Numerical results confirm our theoretical analysis: the proposed algorithm reduces the run time of NLM, and thanks to the symmetrization step, actually provides some improvement in peak signal-to-noise ratios.
Stanley H. Chan, Todd E. Zickler, Yue M. Lu
ICASSP1
2013 Adaptive non-local means for multiview image denoising: Searching for the right patches via a statistical approach
abstract
We present an adaptive non-local means (NLM) denoising method for a sequence of images captured by a multiview imaging system, where direct extensions of existing single image NLM methods are incapable of producing good results. Our proposed method consists of three major components: (1) a robust joint-view distance metric to measure the similarity of patches; (2) an adaptive procedure derived from statistical properties of the estimates to determine the optimal number of patches to be used; (3) a new NLM algorithm to denoise using only a set of similar patches. Experimental results show that the proposed method is robust to disparity estimation error, out-performs existing algorithms in multiview settings, and performs competitively in video settings.
Enming Luo, Stanley H. Chan, Shengjun Pan, Truong Q. Nguyen
ICIP2
2013 Stochastic blockmodel approximation of a graphon: Theory and consistent estimation
abstract
Given a convergent sequence of graphs, there exists a limit object called the graphon from which random graphs are generated. This nonparametric perspective of random graphs opens the door to study graphs beyond the traditional parametric models, but at the same time also poses the challenging question of how to estimate the graphon underlying observed graphs. In this paper, we propose a computationally efficient algorithm to estimate a graphon from a set of observed graphs generated from it. We show that, by approximating the graphon with stochastic block models, the graphon can be consistently estimated, that is, the estimation error vanishes as the size of the graph approaches infinity.
Edoardo M. Airoldi, Thiago B. Costa, Stanley H. Chan
NIPS3
2011 An augmented Lagrangian method for video restoration
abstract
This paper presents a fast algorithm for restoring video sequences. The proposed algorithm, as opposed to existing methods, does not consider video restoration as a sequence of image restoration problems. Rather, it treats a video sequence as a space-time volume and poses a space-time total variation regularization to enhance the smoothness of the solution. The optimization problem is solved by transforming the original unconstrained minimization problem to an equivalent constrained minimization problem. An augmented Lagrangian method is used to handle the constraints, and an alternating direction method (ADM) is used to iteratively find solutions of the subproblems. The proposed algorithm has a wide range of applications, including video deblurring and denoising, disparity map refinement, and reducing hot-air turbulence effects.
Stanley H. Chan, Ramsin Khoshabeh, Kristofor B. Gibson, Philip E. Gill, Truong Q. Nguyen
ICASSP1
2011 Spatio-temporal consistency in video disparity estimation
abstract
We present a novel stereo video disparity estimation method. The proposed method is a two-stage algorithm. During the first stage, initial disparity maps are computed in a frame by-frame basis. In the second stage, the initial estimates are treated as a space-time volume. By setting up an l1-normed minimization problem with a novel three-dimensional total variation regularization, spatial smoothness and temporal consistency are handled simultaneously. Due to our unique formulation, any existing image disparity estimation technique may utilize our method as a post-processing step to refine noisy estimates or to be extended to videos. The proposed method shows superior speed, accuracy, and consistency compared to state-of-the-art algorithms.
Ramsin Khoshabeh, Stanley H. Chan, Truong Q. Nguyen
ICASSP2
2011 Single image spatially variant out-of-focus blur removal
abstract
This paper addresses an out-of-focus blur problem in which the fore ground object is in focus whereas the background scene is out of focus. To recover the details of the background scene, a spatially variant blind deconvolution problem must be solved. However, spatially variant deconvolution is computationally intensive because Fourier based methods cannot be used to handle spatially variant convolution operators. The proposed method exploits the invariant structure of the problem by first predicting the background. Then a blind deconvolution algorithm is applied to estimate the blur kernel and a coarse estimate of the image is found as a side product. Finally, the back ground is recovered using total variation minimization, and fused with the foreground to produce the final deblurred image.
Stanley H. Chan, Truong Q. Nguyen
ICIP1
2011 An Augmented Lagrangian Method for Total Variation Video Restoration
abstract
This paper presents a fast algorithm for restoring video sequences. The proposed algorithm, as opposed to existing methods, does not consider video restoration as a sequence of image restoration problems. Rather, it treats a video sequence as a space-time volume and poses a space-time total variation regularization to enhance the smoothness of the solution. The optimization problem is solved by transforming the original unconstrained minimization problem to an equivalent constrained minimization problem. An augmented Lagrangian method is used to handle the constraints, and an alternating direction method is used to iteratively find solutions to the subproblems. The proposed algorithm has a wide range of applications, including video deblurring and denoising, video disparity refinement, and hot-air turbulence effect reduction.
Stanley H. Chan, Ramsin Khoshabeh, Kristofor B. Gibson, Philip E. Gill, Truong Q. Nguyen
IEEE Trans. Image Process.1
2011 LCD Motion Blur: Modeling, Analysis, and Algorithm
abstract
Liquid crystal display (LCD) devices are well known for their slow responses due to the physical limitations of liquid crystals. Therefore, fast moving objects in a scene are often perceived as blurred. This effect is known as the LCD motion blur. In order to reduce LCD motion blur, an accurate LCD model and an efficient deblurring algorithm are needed. However, existing LCD motion blur models are insufficient to reflect the limitation of human-eye-tracking system. Also, the spatiotemporal equivalence in LCD motion blur models has not been proven directly in the discrete 2-D spatial domain, although it is widely used. There are three main contributions of this paper: modeling, analysis, and algorithm. First, a comprehensive LCD motion blur model is presented, in which human-eye-tracking limits are taken into consideration. Second, a complete analysis of spatiotemporal equivalence is provided and verified using real video sequences. Third, an LCD motion blur reduction algorithm is proposed. The proposed algorithm solves an l(1)-norm regularized least-squares minimization problem using a subgradient projection method. Numerical results show that the proposed algorithm gives higher peak SNR, lower temporal error, and lower spatial error than motion-compensated inverse filtering and Lucy-Richardson deconvolution algorithm, which are two state-of-the-art LCD deblurring algorithms.
Stanley H. Chan, Truong Q. Nguyen
IEEE Trans. Image Process.1
2010 Subpixel motion estimation without interpolation
abstract
We propose a fast subpixel motion estimation method for motion deblurring, where conventional motion estimation algorithms used in video codings are too complex. The new algorithm is a combination of block matching and optical flow. It does not require any interpolation and it does not provide motion compensated frames. Thus it is much faster than conventional methods. Statistical results show that the new algorithm performs quickly and accurately. It also demonstrates compatible performance with the benchmarking full search algorithm, yet uses significantly less amount of time.
Stanley H. Chan, Dung Trung Vo, Truong Q. Nguyen
ICASSP1
2010 Demosaicking imageswith motion blur
abstract
In standard digital color imaging, each pixel position acquires data for only one color plane and the remaining two color planes must be inferred through a process known as demosaicking. Furthermore, the image is susceptible to blurring artifacts due to a moving camera or fast moving subject. In this work we develop a robust framework to demosaick the color filter array (CFA) image while reducing the blur corrupting the image. We begin by defining a color motion blur model that describes the motion blur artifacts affecting color images. We then integrate the motion blur model in the demosaicking algorithm to obtain a computationally efficient framework for deblurring while demosaicking.
Shay Har-Noy, Stanley H. Chan, Truong Q. Nguyen
ICASSP2
2010 Constructing a sparse convolution matrix for shift varying image restoration problems
abstract
Convolution operator is a linear operator characterized by a point spread functions (PSF). In classical image restoration problems, the blur is usually shift invariant and so the convolution operator can be characterized by one single PSF. This assumption allows one to use fast operations such as Fast Fourier Transform (FFT) to perform a matrix-vector computation efficiently. However, as in most of the video motion deblurring problems, the blur is shift variant and so the matrix-vector multiplication can be difficult to perform. In this paper, we propose an efficient method to construct the convolution matrix explicitly. We exploit the submatrix structure of the convolution matrix and systematically assigning values to the nonzero locations. For small to medium sized images, the convolution matrix gives superior speed than some state-of-art convolution operators.
Stanley H. Chan
ICIP1
2010 LCD motion blur modeling and simulation
abstract
Liquid crystal display (LCD) devices are well known to have slow response due to the physical limitations of the liquid crystals. Therefore, fast moving objects in a scene are often seen blurred on an LCD. In order to reduce motion blur, an accurate LCD model and an efficient deblurring algorithm are needed. However, existing LCD models are inadequate to reflect the human eye tracking limitation. Also, the spatial-temporal equivalence in LCD models is widely used but not proven directly in the 2D discrete spatial domain. In this paper, we study the human eye tracking limit by reviewing a number of papers in the cognitive science literature. We provide both theoretical and experimental arguments to support our findings. Also, we prove the spatial-temporal equivalence rigourously and verify the results using real video sequences.
Stanley H. Chan, Truong Q. Nguyen
ICME1
2010 Comparison of Two Frame Rate Conversion Schemes for Reducing LCD Motion Blurs
abstract
Liquid crystal display (LCD) is known to have motion blur due to the slow response and sample-hold characteristics of the liquid crystal (LC). To alleviate the LCD motion blur, improving the LC response is the most fundamental solution. However, if the response time is shortened, then more frames are needed and hence frame rate up conversion (FRUC) should be used. In this paper, we study two FRUC methods. We compare the output signal qualities by studying the temporal and spatial profile of the two methods. We use the solution of Erickson-Leslie equation to derive the step response, in contrast to the existing literature where the resistor-capacitor (RC) approximation and uniform averaging function are used. The step response we derived is able to model not only the general trend of the rising and falling edges, but also the effects of different gray level transitions. Based on the step response, we analyze the two methods by comparing the observed signal in both the spatial and temporal domain.
Stanley H. Chan, Thomas X. Wu, Truong Q. Nguyen
IEEE Signal Process. Lett.1
2009 Fast LCD motion deblurring by decimation and optimization
abstract
The LCD deblurring problem is considered as a simple bounded quadratic programming problem and is solved using conjugate gradient with early stopping criteria to avoid excessive search. A decimation and interlace interpolation method is introduced to reduce the computing time. Solutions are competitive to those the generated by conventional Lucy Richardson algorithm, but using much shorter amount of time. The method can be extended to higher decimation factors. Visual subjective tests are conducted to justify our proposed method.
Stanley H. Chan, Truong Q. Nguyen
ICASSP1
2008 Inverse image problem of designing phase shifting masks in optical lithography
abstract
The continual shrinkage of minimum feature size in integrated circuit (IC) fabrication incurs more and more serious distortion in the optical lithography process, generating circuit patterns deviating from the desired ones. Conventional resolution enhancement techniques (RETs) are facing critical challenges in compensating such increasingly severe distortion. The approach of inverse lithography, which is a branch of mask design methodology to treat the design as an inverse image problem, is adopted in this paper. We apply nonlinear optimization techniques to design masks with minimally distorted output. The output patterns so generated have high contrast and low dose sensitivity. We also propose a dynamic program-based initialization scheme to pre-assign phases to the layout.
Stanley H. Chan, Edmund Y. Lam
ICIP1