EDBT 2026 Demo / reviewers in the wild / expert
Jin-Li Suo
dblp:15/898 · also Jinli Suo
· DBLP profile ↗
61ranked-venue papers
11as first author
28since 2021 · last 2026
0000-0002-3426-1634ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 34 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COGS: A Causal Representation Learning Framework for Out-of-Distribution Generalization in Time SeriesabstractTime series analysis is crucial in various fields such as healthcare and finance. However, environmental variations and the inherent non-stationarity of time series data often lead to out-of-distribution (OOD) scenarios, consequently causing model performance degradation. Most existing OOD generalization methods primarily focus on images or text, leaving time series analysis relatively underexplored. In this paper, we propose COGS, a novel framework that incorporates causal representation learning into the OOD generalization of time series. By imposing structural priors, our method identifies latent variables and learns a causal graph to disentangle causal variables from non-causal ones. These causal variables are then used to learn domain-invariant representations for stable prediction. Moreover, to tackle the challenge of the absence of domain labels, we further introduce a prototype-based domain discovery algorithm that infers domain labels in an unsupervised manner. The entire framework is optimized in a two-phase iterative manner, resulting in robust OOD performance. Extensive experiments on multiple real-world time series datasets demonstrate that our method achieves competitive performance compared to baseline methods. Xinxin Song 0001, Yuxiao Cheng, Tingxiong Xiao, Jin-Li Suo |
AAAI | 4 |
| 2026 | Causally-informed deep learning towards explainable and generalizable outcome prediction in critical care
Yuxiao Cheng, Xinxin Song 0001, Qin Zhong, Kunlun He, Jin-Li Suo |
Artif. Intell. Medicine | 6 |
| 2026 | DarkVision: A Benchmark and Study for Low-Light Image/Video AnalysisabstractLow-light image/video analysis is essential for various applications, e.g., night surveillance and photography, high-speed imaging, and autonomous vehicles. Under such conditions, cameras suffer from low signal-to-noise ratio, which degrades image quality severely and poses challenges for downstream tasks such as object detection. Data-driven methods have achieved enormous success for normal-light image/video restoration and high-level vision tasks. However, the lack of a high-quality benchmark dataset with accurate semantic annotations for low-light images and especially videos greatly hinders research progress. In this paper, we contribute the first multi-illuminance, multi-camera, low-light dataset, DarkVision, serving both image/video enhancement and object detection applications. We provide bright and dark pairs with pixel-wise registration, in which the bright counterpart provides a reliable reference for enhancement and annotation. This dataset comprises 13,455 images of 900 static scenes with objects from 15 categories, and 89,411 frames of 32 dynamic scenes with 4 categories of objects. For each scene, images/videos were captured at 5 illuminance levels using three cameras of different quality grades; average photon numbers can be reliably estimated from the calibration curves for quantitative studies. The static images and dynamic videos respectively contain around 7344 and 320,667 object instances in total. With DarkVision, we establish baselines for image/video enhancement and object detection by representative algorithms. To demonstrate an exemplary application of DarkVision, we propose two simple yet effective approaches to improve the performance of video enhancement and object detection respectively by exploiting temporal cues. Furthermore, we study the relationship between image enhancement and object detection. We believe DarkVision can help to advance the state-of-the art in both low-light image/video enhancement and object detection, as well as benefiting cross-task studies. Bo Zhang 0109, Runzhao Yang, Zhihong Zhang 0004, Jiayi Xie, Jin-Li Suo |
Comput. Vis. Media | 6 |
| 2026 | UTA-Sign: Unsupervised thermal video augmentation via event-assisted traffic signage sketching
Yuqi Han, Songqian Zhang, Weijian Su, Jin-Li Suo, Qiang Zhang 0008 |
Pattern Recognit. | 6 |
| 2026 | Seq-IF: Sequentially Consistent Infrared-Visible Video Fusion Under Time-Varying Illumination for Perception EnhancementabstractInfrared–visible image fusion leverages the complementary strengths of both modalities to enhance visual perception in challenging environments. While image-level fusion has achieved promising results, extending it to video remains challenging due to temporal illumination variations, brightness flickering, and visual inconsistency caused by motion under non-uniform illumination. To address these issues, we propose Seq-IF, a sequential fusion framework that ensures visual consistency and structural fidelity when processing video sequences. The framework comprises a static–dynamic decoupling module for robust foreground–background separation. For the background, fusion is performed by selecting the frame with the highest contrast to ensure clarity and stability. For the dynamic objects, the pixel intensity is adaptively adjusted across consecutive frames by a lightweight MLP-based illumination-consistency fine-tuning module that performs online adaptation and dynamically optimizes brightness in response to scene changes. Later we introduce a spatial–frequency fusion module integrating multi-scale encoder and edge-guided decoder to ensure structural consistency. Extensive experiments demonstrate that the fusion results produced by Seq-IF outperform baseline methods in terms of clarity, detail preservation, and illumination stability, achieving smoother temporal transitions and enhanced perceptual quality. Furthermore, the effectiveness of Seq-IF is validated on downstream tasks, including object detection and optical flow estimation, highlighting its applicability to real-world scenarios. Yuqi Han, Zhihui Zheng, Weijian Su, Mingkai Wei, Liang Zhang 0031, Jin-Li Suo, Qiang Zhang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | A Compact Implicit Neural Representation for Efficient Storage of Massive 4D Functional Magnetic Resonance ImagingabstractFunctional Magnetic Resonance Imaging (fMRI) data is a widely used kind of four-dimensional biomedical data, which requires effective compression. However, fMRI compressing poses unique challenges due to its intricate temporal dynamics, low signal-to-noise ratio, and complicated underlying redundancies. This paper reports a novel compression paradigm specifically tailored for fMRI data based on Implicit Neural Representation (INR). The proposed approach focuses on removing the various redundancies among the time series by employing several methods, including (i) conducting spatial correlation modeling for intra-region dynamics, (ii) decomposing reusable neuronal activation patterns, and (iii) using proper initialization together with nonlinear fusion to describe the inter-region similarity. This scheme appropriately incorporates the unique features of fMRI data, and experimental results on publicly available datasets demonstrate the effectiveness of the proposed method, surpassing state-of-the-art algorithms in both conventional image quality evaluation metrics and fMRI downstream tasks. This work in this paper paves the way for sharing massive fMRI data at low bandwidth and high fidelity. Ruoran Li, Runzhao Yang, Wenxin Xiang, Yuxiao Cheng, Tingxiong Xiao, Jin-Li Suo |
AAAI | 7 |
| 2025 | FIND: A Framework for Discovering Formulas in DataabstractScientific discovery serves as the cornerstone for advances in various fields, from the fundamental laws of physics to the intricate mechanisms of biology. However, two existing mainstream methods---symbolic regression and dimensional analysis, are significantly limited in this task: the former suffers from low computational efficiency due to the vast search space and often results in formulas without physical meaning; the latter provides a useful theoretical framework but also struggles in searching in a huge space because of lacking effective analysis for the latent variables. To address this issue, here we propose a framework for efficiently discovering underlying formulas in data, named FIND. We draw inspiration from Buckingham’s Pi theorem, imposing dimensional constraints on the input and output, thereby ensuring discovered expressions possess physical meaning. Additionally, we propose a theoretical scheme for identifying the latent structure as well as a coarse-to-fine framework, significantly reducing the search space of latent variables. This framework not only improves computational efficiency but also enhances model interpretability. From comprehensive experimental validation, FIND showcases its potential to uncover meaningful scientific insights across various domains, providing a robust tool for advancing our understanding of unknown systems. Tingxiong Xiao, Yuxiao Cheng, Jin-Li Suo |
AAAI | 3 |
| 2025 | X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent AttentionabstractWe propose X-NeMo, a novel zero-shot diffusion-based portrait animation pipeline that animates a static portrait using facial movements from a driving video of a different individual. Our work first identifies the root causes of the limitations in prior approaches, such as identity leakage and difficulty in capturing subtle and extreme expressions. To address these challenges, we introduce a fully end-to-end training framework that distills a 1D identity-agnostic latent motion descriptor from driving image, effectively controlling motion through cross-attention during image generation. Our implicit motion descriptor captures expressive facial motion in fine detail, learned end-to-end from a diverse video dataset without reliance on any pre-trained motion detectors. We further disentangle motion latents from identity cues with enhanced expressiveness by supervising their learning with a dual GAN decoder, alongside spatial and color augmentations. By embedding the driving motion into a 1D latent vector and controlling motion via cross-attention instead of additive spatial guidance, our design effectively eliminates the transmission of spatial-aligned structural clues from the driving condition to the diffusion backbone, substantially mitigating identity leakage. Extensive experiments demonstrate that X-NeMo surpasses state-of-the-art baselines, producing highly expressive animations with superior identity resemblance. Our code and models will be available for research. Xiaochen Zhao, Guoxian Song, You Xie, Xiu Li 0003, Linjie Luo, Jin-Li Suo, Yebin Liu |
ICLR | 8 |
| 2025 | DVI: A Derivative-based Vision Network for INRabstractRecent advancements in computer vision have seen Implicit Neural Representations (INR) becoming a dominant representation form for data due to their compactness and expressive power. To solve various vision tasks with INR data, vision networks can either be purely INR-based, but are thereby limited by simplistic operations and performance constraints, or include raster-based methods, which then tend to lose crucial structural information of the INR during the conversion process. To address these issues, we propose DVI, a novel Derivative-based Vision network for INR, capable of handling a variety of vision tasks across various data modalities, while achieving the best performance among the existing methods by incorporating state of the art raster-based methods into a INR based architecture. DVI excels by extracting semantic information from the high order derivative map of the INR, then seamlessly fusing it into a pre-existing raster-based vision network, enhancing its performance with deeper, task-relevant semantic insights. Extensive experiments on five vision tasks across three data modalities demonstrate DVI's superiority over existing methods. Additionally, our study encompasses comprehensive ablation studies to affirm the efficacy of each element of DVI, the influence of different derivative computation techniques and the impact of derivative orders. Reproducible codes are provided in the supplementary materials. Runzhao Yang, Zhihong Zhang 0004, Fabian Zhang, Tingxiong Xiao, Zongren Li, Kunlun He, Jin-Li Suo |
ICML | 8 |
| 2025 | HeRIF: A Mixture-of-Experts Framework for Infrared and Visible Image Fusion with Heterogeneous Resolutions
Songqian Zhang, Weijian Su, Yuqi Han, Jin-Li Suo, Qiang Zhang 0008 |
PRCV (5) | 5 |
| 2025 | Lightweight High-Speed Photography Built on Coded Exposure and Implicit Neural Representation of Videos
Zhihong Zhang 0004, Runzhao Yang, Jin-Li Suo, Yuxiao Cheng, Qionghai Dai |
Int. J. Comput. Vis. | 3 |
| 2025 | Event-Enhanced Snapshot Compressive Videography at 10K FPSabstractVideo snapshot compressive imaging (SCI) encodes the target dynamic scene compactly into a snapshot and reconstructs its high-speed frame sequence afterward, greatly reducing the required data footprint and transmission bandwidth as well as enabling high-speed imaging with a low frame rate intensity camera. In implementation, high-speed dynamics are encoded via temporally varying patterns, and only frames at corresponding temporal intervals can be reconstructed, while the dynamics occurring between consecutive frames are lost. To unlock the potential of conventional snapshot compressive videography, we propose a novel hybrid "intensity event imaging scheme by incorporating an event camera into a video SCI setup. Our proposed system consists of a dual-path optical setup to record the coded intensity measurement and intermediate event signals simultaneously, which is compact and photon-efficient by collecting the half photons discarded in conventional video SCI. Correspondingly, we developed a dual-branch Transformer utilizing the reciprocal relationship between two data modes to decode dense video frames. Extensive experiments on both simulated and real-captured data demonstrate our superiority to state-of-the-art video SCI and video frame interpolation (VFI) methods. Benefiting from the new hybrid design leveraging both intrinsic redundancy in videos and the unique feature of event cameras, we achieve high-quality videography at 0.1ms time intervals with a low-cost CMOS image sensor working at 24 FPS. Bo Zhang 0109, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | CUTS+: High-Dimensional Causal Discovery from Irregular Time-SeriesabstractCausal discovery in time-series is a fundamental problem in the machine learning community, enabling causal reasoning and decision-making in complex scenarios. Recently, researchers successfully discover causality by combining neural networks with Granger causality, but their performances degrade largely when encountering high-dimensional data because of the highly redundant network design and huge causal graphs. Moreover, the missing entries in the observations further hamper the causal structural learning. To overcome these limitations, We propose CUTS+, which is built on the Granger-causality-based causal discovery method CUTS and raises the scalability by introducing a technique called Coarse-to-fine-discovery (C2FD) and leveraging a message-passing-based graph neural network (MPGNN). Compared to previous methods on simulated, quasi-real, and real datasets, we show that CUTS+ largely improves the causal discovery performance on high-dimensional data with different types of irregular sampling. Yuxiao Cheng, Lianglong Li, Tingxiong Xiao, Zongren Li, Jin-Li Suo, Kunlun He, Qionghai Dai |
AAAI | 5 |
| 2024 | SHoP: A Deep Learning Framework for Solving High-Order Partial Differential EquationsabstractSolving partial differential equations (PDEs) has been a fundamental problem in computational science and of wide applications for both scientific and engineering research. Due to its universal approximation property, neural network is widely used to approximate the solutions of PDEs. However, existing works are incapable of solving high-order PDEs due to insufficient calculation accuracy of higher-order derivatives, and the final network is a black box without explicit explanation. To address these issues, we propose a deep learning framework to solve high-order PDEs, named SHoP. Specifically, we derive the high-order derivative rule for neural network, to get the derivatives quickly and accurately; moreover, we expand the network into a Taylor series, providing an explicit solution for the PDEs. We conduct experimental validations four high-order PDEs with different dimensions, showing that we can solve high-order PDEs efficiently and accurately. The source code can be found at https://github.com/HarryPotterXTX/SHoP.git. Tingxiong Xiao, Runzhao Yang, Yuxiao Cheng, Jin-Li Suo |
AAAI | 4 |
| 2024 | A Physics-Informed Low-Rank Deep Neural Network for Blind and Universal Lens Aberration CorrectionabstractHigh-end lenses, although offering high-quality images, suffer from both insufficient affordability and bulky design, which hamper their applications in low-budget scenarios or on low-payload platforms. A flexible scheme is to tackle the optical aberration of low-end lenses computationally. However, it is highly demanded but quite challenging to build a general model capable of handling non-stationary aberrations and covering diverse lenses, especially in a blind manner. To address this issue, we propose a universal solution by extensively utilizing the physical properties of camera lenses: (i) reducing the complexity of lens aberrations, i.e., lens-specific non-stationary blur, by warping annual-ring-shaped sub-images into rectangular stripes to transform non-uniform degenerations into a uniform one, (ii) building a low-dimensional nonnegative orthogonal representation of lens blur kernels to cover diverse lenses; (iii) designing a decoupling network to decompose the input low-quality image into several components degenerated by above kernel bases, and applying corresponding pretrained deconvolution networks to reverse the degeneration. Benefiting from the proper incorporation of lenses' physical properties and unique network design, the proposed method achieves superb imaging quality, wide applicability for various lenses, high running efficiency, and is totally free of kernel calibration. These advantages bring great potential for scenarios requiring lightweight high-quality photography. Jin Gong, Runzhao Yang, Jin-Li Suo, Qionghai Dai |
CVPR | 4 |
| 2024 | CausalTime: Realistically Generated Time-series for Benchmarking of Causal DiscoveryabstractTime-series causal discovery (TSCD) is a fundamental problem of machine learning. However, existing synthetic datasets cannot properly evaluate or predict the algorithms' performance on real data. This study introduces the CausalTime pipeline to generate time-series that highly resemble the real data and with ground truth causal graphs for quantitative performance evaluation. The pipeline starts from real observations in a specific scenario and produces a matching benchmark dataset. Firstly, we harness deep neural networks along with normalizing flow to accurately capture realistic dynamics. Secondly, we extract hypothesized causal graphs by performing importance analysis on the neural network or leveraging prior knowledge. Thirdly, we derive the ground truth causal graphs by splitting the causal model into causal term, residual term, and noise term. Lastly, using the fitted network and the derived causal graph, we generate corresponding versatile time-series proper for algorithm assessment. In the experiments, we validate the fidelity of the generated data through qualitative and quantitative experiments, followed by a benchmarking of existing TSCD algorithms using these generated datasets. CausalTime offers a feasible solution to evaluating TSCD algorithms in real applications and can be generalized to a wide range of fields. For easy use of the proposed approach, we also provide a user-friendly website, hosted on www.causaltime.cc. Yuxiao Cheng, Tingxiong Xiao, Qin Zhong, Jin-Li Suo, Kunlun He |
ICLR | 5 |
| 2024 | HOPE: High-Order Polynomial Expansion of Black-Box Neural NetworksabstractDespite their remarkable performance, deep neural networks remain mostly "black boxes", suggesting inexplicability and hindering their wide applications in fields requiring making rational decisions. Here we introduce HOPE (High-order Polynomial Expansion), a method for expanding a network into a high-order Taylor polynomial on a reference input. Specifically, we derive the high-order derivative rule for composite functions and extend the rule to neural networks to obtain their high-order derivatives quickly and accurately. From these derivatives, we can then derive the Taylor polynomial of the neural network, which provides an explicit expression of the network's local interpretations. We combine the Taylor polynomials obtained under different reference inputs to obtain the global interpretation of the neural network. Numerical analysis confirms the high accuracy, low computational complexity, and good convergence of the proposed method. Moreover, we demonstrate HOPE's wide applications built on deep learning, including function discovery, fast inference, and feature selection. We compared HOPE with other XAI methods and demonstrated our advantages. Tingxiong Xiao, Yuxiao Cheng, Jin-Li Suo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | HAvatar: High-fidelity Head Avatar via Facial Model Conditioned Neural Radiance FieldabstractThe problem of modeling an animatable 3D human head avatar under lightweight setups is of significant importance but has not been well solved. Existing 3D representations either perform well in the realism of portrait images synthesis or the accuracy of expression control, but not both. To address the problem, we introduce a novel hybrid explicit-implicit 3D representation, Facial Model Conditioned Neural Radiance Field, which integrates the expressiveness of NeRF and the prior information from the parametric template. At the core of our representation, a synthetic-renderings-based condition method is proposed to fuse the prior information from the parametric model into the implicit field without constraining its topological flexibility. Besides, based on the hybrid representation, we properly overcome the inconsistent shape issue presented in existing methods and improve the animation stability. Moreover, by adopting an overall GAN-based architecture using an image-to-image translation network, we achieve high-resolution, realistic and view-consistent synthesis of dynamic head appearance. Experiments demonstrate that our method can achieve state-of-the-art performance for 3D head avatar animation compared with previous methods. Xiaochen Zhao, Lizhen Wang 0002, Jingxiang Sun, Hongwen Zhang 0001, Jin-Li Suo, Yebin Liu |
ACM Trans. Graph. | 5 |
| 2023 | SCI: A Spectrum Concentrated Implicit Neural Compression for Biomedical DataabstractMassive collection and explosive growth of biomedical data, demands effective compression for efficient storage, transmission and sharing. Readily available visual data compression techniques have been studied extensively but tailored for natural images/videos, and thus show limited performance on biomedical data which are of different features and larger diversity. Emerging implicit neural representation (INR) is gaining momentum and demonstrates high promise for fitting diverse visual data in target-data-specific manner, but a general compression scheme covering diverse biomedical data is so far absent. To address this issue, we firstly derive a mathematical explanation for INR's spectrum concentration property and an analytical insight on the design of INR based compressor. Further, we propose a Spectrum Concentrated Implicit neural compression (SCI) which adaptively partitions the complex biomedical data into blocks matching INR's concentrated spectrum envelop, and design a funnel shaped neural network capable of representing each block with a small number of parameters. Based on this design, we conduct compression via optimization under given budget and allocate the available parameters with high representation accuracy. The experiments show SCI's superior performance to state-of-the-art methods including commercial compressors, data-driven ones, and INR based counterparts on diverse biomedical data. The source code can be found at https://github.com/RichealYoung/ImplicitNeuralCompression.git. Runzhao Yang, Tingxiong Xiao, Yuxiao Cheng, Qianni Cao, Jinyuan Qu, Jin-Li Suo, Qionghai Dai |
AAAI | 6 |
| 2023 | Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality AssessmentabstractRecently, some studies have shown that semantic and distortion representations both benefit the evaluation of image quality. However, the images of existing synthetic distortion databases are annotated with subjective quality scores and distortion types, lacking labels with semantic objects. Therefore, it is virtually infeasible to learn the representations of image semantics and distortion by co-guiding with semantic and distortion labels. To address this issue, we propose a dual-perception network (DPNet) via an end-to-end multi-task learning method, where knowledge distillation is lever-aged as a semantic label-free strategy. Specifically, semantic representation derived from pre-trained ResNet152 is applied to supervise the output of DPNet, while the output is utilized to construct a distortion recognition task. In this way, image semantics and distortion can be hybridly represented in an identical feature map. Finally, image quality is regressed based on the hybrid representations. Experimental results conducted on five benchmark databases validate that the proposed method can achieve state-of-the-art performance. Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005 |
ICASSP | 4 |
| 2023 | ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality AssessmentabstractThe human vision system is highly adapted to extract structural information from the viewed scenes. The irregularity of point clouds makes the extraction of structural information containing both color and geometry an important challenge for point cloud quality assessment (PCQA). This paper proposes a point structural information (PSI) network (ψ-Net) for no-reference PCQA. Firstly, a PSI module is proposed to map the position vectors of neighboring points to weights for the calculation of color and geometric structure information. Secondly, a dual-stream network is presented to introduce distortion-related features for PCQA. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005 |
ICASSP | 4 |
| 2023 | CUTS: Neural Causal Discovery from Irregular Time-Series Data
Yuxiao Cheng, Runzhao Yang, Tingxiong Xiao, Zongren Li, Jin-Li Suo, Kunlun He, Qionghai Dai |
ICLR | 5 |
| 2023 | Computational Imaging and Artificial Intelligence: The Next Revolution of Mobile VisionabstractSignal capture is at the forefront of perceiving and understanding the environment; thus, imaging plays a pivotal role in mobile vision. Recent unprecedented progress in artificial intelligence (AI) has shown great potential in the development of advanced mobile platforms with new imaging devices. Traditional imaging systems based on the “capturing images first and processing afterward” mechanism cannot meet this explosive demand. On the other hand, computational imaging (CI) systems are designed to capture high-dimensional data in an encoded manner to provide more information for mobile vision systems. Thanks to AI, CI can now be used in real-life systems by integrating deep learning algorithms into the mobile vision platform to achieve a closed loop of intelligent acquisition, processing, and decision-making, thus leading to the next revolution of mobile vision. Starting from the history of mobile vision using digital cameras, this work first introduces the advancement of CI in diverse applications and then conducts a comprehensive review of current research topics combining CI and AI. Although new-generation mobile platforms, represented by smart mobile phones, have deeply integrated CI and AI for better image acquisition and processing, most mobile vision platforms, such as self-driving cars and drones only loosely connect CI and AI, and are calling for a closer integration. Motivated by this fact, at the end of this work, we propose some potential technologies and disciplines that aid the deep integration of CI and AI and shed light on new directions in the future generation of mobile vision platforms. Jin-Li Suo, Jin Gong, Xin Yuan 0002, David J. Brady, Qionghai Dai |
Proc. IEEE | 1 |
| 2023 | Retrieving Object Motions From Coded Shutter Snapshot in Dark EnvironmentabstractVideo object detection is a widely studied topic and has made significant progress in the past decades. However, the feature extraction and calculations in existing video object detectors demand decent imaging quality and avoidance of severe motion blur. Under extremely dark scenarios, due to limited sensor sensitivity, we have to trade off signal-to-noise ratio for motion blur compensation or vice versa, and thus suffer from performance deterioration. To address this issue, we propose to temporally multiplex a frame sequence into one snapshot and extract the cues characterizing object motion for trajectory retrieval. For effective encoding, we build a prototype for encoded capture by mounting a highly compatible programmable shutter. Correspondingly, in terms of decoding, we design an end-to-end deep network called detection from coded snapshot (DECENT) to retrieve sequential bounding boxes from the coded blurry measurements of dynamic scenes. For effective network learning, we generate quasi-real data by incorporating physically-driven noise into the temporally coded imaging model, which circumvents the unavailability of training data and with high generalization ability on real dark videos. The approach offers multiple advantages, including low bandwidth, low cost, compact setup, and high accuracy. The effectiveness of the proposed approach is experimentally validated under low illumination vision and provide a feasible way for night surveillance. Kaiming Dong, Runzhao Yang, Yuxiao Cheng, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Image Process. | 5 |
| 2023 | INFWIDE: Image and Feature Space Wiener Deconvolution Network for Non-Blind Image Deblurring in Low-Light ConditionsabstractUnder low-light environment, handheld photography suffers from severe camera shake under long exposure settings. Although existing deblurring algorithms have shown promising performance on well-exposed blurry images, they still cannot cope with low-light snapshots. Sophisticated noise and saturation regions are two dominating challenges in practical low-light deblurring: the former violates the Gaussian or Poisson assumption widely used in most existing algorithms and thus degrades their performance badly, while the latter introduces non-linearity to the classical convolution-based blurring model and makes the deblurring task even challenging. In this work, we propose a novel non-blind deblurring method dubbed image and feature space Wiener deconvolution network (INFWIDE) to tackle these problems systematically. In terms of algorithm design, INFWIDE proposes a two-branch architecture, which explicitly removes noise and hallucinates saturated regions in the image space and suppresses ringing artifacts in the feature space, and integrates the two complementary outputs with a subtle multi-scale fusion network for high quality night photograph deblurring. For effective network training, we design a set of loss functions integrating a forward imaging model and backward reconstruction to form a close-loop regularization to secure good convergence of the deep neural network. Further, to optimize INFWIDE's applicability in real low-light conditions, a physical-process-based low-light noise model is employed to synthesize realistic noisy night photographs for model training. Taking advantage of the traditional Wiener deconvolution algorithm's physically driven characteristics and deep neural network's representation ability, INFWIDE can recover fine details while suppressing the unpleasant artifacts during deblurring. Extensive experiments on synthetic data and real data demonstrate the superior performance of the proposed approach. Zhihong Zhang 0004, Yuxiao Cheng, Jin-Li Suo, Liheng Bian, Qionghai Dai |
IEEE Trans. Image Process. | 3 |
| 2022 | Plug-and-Play Algorithms for Video Snapshot Compressive ImagingabstractWe consider the reconstruction problem of video snapshot compressive imaging (SCI), which captures high-speed videos using a low-speed 2D sensor (detector). The underlying principle of SCI is to modulate sequential high-speed frames with different masks and then these encoded frames are integrated into a snapshot on the sensor and thus the sensor can be of low-speed. On one hand, video SCI enjoys the advantages of low-bandwidth, low-power and low-cost. On the other hand, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging and one of the bottlenecks lies in the reconstruction algorithm. Existing algorithms are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload. We first employ the image deep denoising priors to show that PnP can recover a UHD color video with 30 frames from a snapshot measurement. Since videos have strong temporal correlation, by employing the video deep denoising priors, we achieve a significant improvement in the results. Furthermore, we extend the proposed PnP algorithms to the color SCI system using mosaic sensors, where each pixel only captures the red, green or blue channels. A joint reconstruction and demosaicing paradigm is developed for flexible and high quality reconstruction of color video SCI systems. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm. Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Frédo Durand, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Universal and Flexible Optical Aberration Correction Using Deep-Prior Based DeconvolutionabstractHigh quality imaging usually requires bulky and expensive lenses to compensate geometric and chromatic aberrations. This poses high constraints on the optical hash or low cost applications. Although one can utilize algorithmic reconstruction to remove the artifacts of low-end lenses, the degeneration from optical aberrations is spatially varying and the computation has to trade off efficiency for performance. For example, we need to conduct patch-wise optimization or train a large set of local deep neural networks to achieve high reconstruction performance across the whole image. In this paper, we propose a PSF aware deep network, which takes the aberrant image and PSF map as input and produces the latent high quality version via incorporating deep priors, thus leading to a universal and flexible optical aberration correction method. Specifically, we pre-train a base model from a set of diverse lenses and then adapt it to a given lens by quickly refining the parameters, which largely alleviates the time and memory consumption of model learning. The approach is of high efficiency in both training and testing stages. Extensive results verify the promising applications of our proposed approach for compact low-end cameras. The code is available at https://github.com/leehsiu/UABC Xiu Li 0003, Jin-Li Suo, Xin Yuan 0002, Qionghai Dai |
ICCV | 2 |
| 2021 | Sinusoidal Sampling Enhanced Compressive Camera for High Speed ImagingabstractCompressive sensing technique allows capturing fast phenomena at a much higher frame rate than the camera sensor, by recovering a frame sequence from their encoded combination. However, most conventional compressive video sensing methods limit the achieved frame rate improvement to tenfold and only support low resolution recovery. Making use of the camera's redundant spatial resolution for further frame rate improve, here we report a novel compressive video acquisition technique termed Sinusoidal Sampling Enhanced Compressive Camera (S2EC2) to encode denser frames within a snapshot. Specifically, we decompose the dense frames into groups and apply combinational coding: random codes within each group for compressive acquisition; group specific sinusoidal codes to multiplex different groups onto the high resolution sensor. The sinusoidal codes designed for these groups would shift their frequency components by different offsets in the Fourier domain and staggered the dominant frequencies of the coded measurements of these groups. Correspondingly, the reconstruction successfully separate coded measurements of different groups and recovers frames within each group. Besides, we also solve the implementation problem of insufficient gray scale spatial light modulation speed, and build a prototype achieving 2000 fps reconstruction with a 15.6 fps camera (the actual compression ratio is 0.009). The extensive experiments validate the proposed approach. Chao Deng 0005, Yuanlong Zhang, Yifeng Mao, Jingtao Fan, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Plug-and-Play Algorithms for Large-Scale Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) aims to capture the high-dimensional (usually 3D) images using a 2D sensor (detector) in a single snapshot. Though enjoying the advantages of low-bandwidth, low-power and low-cost, applying SCI to large-scale problems (HD or UHD videos) in our daily life is still challenging. The bottleneck lies in the reconstruction algorithms; they are either too slow (iterative optimization algorithms) or not flexible to the encoding process (deep learning based end-to-end networks). In this paper, we develop fast and flexible algorithms for SCI based on the plug-and-play (PnP) framework. In addition to the widely used PnP-ADMM method, we further propose the PnP-GAP (generalized alternating projection) algorithm with a lower computational workload and prove the {global convergence} of PnP-GAP under the SCI hardware constraints. By employing deep denoising priors, we first time show that PnP can recover a UHD color video (3840×1644×48 with PNSR above 30dB) from a snapshot 2D measurement. Extensive results on both simulation and real datasets verify the superiority of our proposed algorithm. Xin Yuan 0002, Yang Liu 0146, Jin-Li Suo, Qionghai Dai |
CVPR | 3 |
| 2020 | Learning Deep Landmarks for Imbalanced ClassificationabstractWe introduce a deep imbalanced learning framework called learning DEep Landmarks in laTent spAce (DELTA). Our work is inspired by the shallow imbalanced learning approaches to rebalance imbalanced samples before feeding them to train a discriminative classifier. Our DELTA advances existing works by introducing the new concept of rebalancing samples in a deeply transformed latent space, where latent points exhibit several desired properties including compactness and separability. In general, DELTA simultaneously conducts feature learning, sample rebalancing, and discriminative learning in a joint, end-to-end framework. The framework is readily integrated with other sophisticated learning concepts including latent points oversampling and ensemble learning. More importantly, DELTA offers the possibility to conduct imbalanced learning with the assistancy of structured feature extractor. We verify the effectiveness of DELTA not only on several benchmark data sets but also on more challenging real-world tasks including click-through-rate (CTR) prediction, multi-class cell type classification, and sentiment analysis with sequential inputs. Feng Bao 0002, Yue Deng 0001, Youyong Kong, Zhiquan Ren, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Rank Minimization for Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) refers to compressive imaging systems where multiple frames are mapped into a single measurement, with video compressive imaging and hyperspectral compressive imaging as two representative applications. Though exciting results of high-speed videos and hyperspectral images have been demonstrated, the poor reconstruction quality precludes SCI from wide applications. This paper aims to boost the reconstruction quality of SCI via exploiting the high-dimensional structure in the desired signal. We build a joint model to integrate the nonlocal self-similarity of video/hyperspectral frames and the rank minimization approach with the SCI sensing process. Following this, an alternating minimization algorithm is developed to solve this non-convex problem. We further investigate the special structure of the sampling process in SCI to tackle the computational workload and memory issues in SCI reconstruction. Both simulation and real data (captured by four different SCI cameras) results demonstrate that our proposed algorithm leads to significant improvements compared with current state-of-the-art algorithms. We hope our results will encourage the researchers and engineers to pursue further in compressive imaging for real applications. Yang Liu 0146, Xin Yuan 0002, Jin-Li Suo, David J. Brady, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Probabilistic natural mapping of gene-level tests for genome-wide association studiesabstractGenome-wide association studies (GWASs) generally focus on a single marker, which limits the elucidation of the genetic architecture of complex traits. Herein, we present a new computational framework, termed probabilistic natural mapping (PALM), for performing gene-level association tests. PALM robustly reveals the inherent genomic structures of genes and generates feature representations that can be seamlessly incorporated into conventional statistic tests. Our approach substantially improves the effectiveness of uncovering associations derived from a subgroup of variants with weak effects, which represents a known challenge associated with existing methods. We applied PALM in a gastric cancer GWAS and identified two additional gastric cancer-associated susceptibility genes, NOC3L and RUNDC2A. The robust susceptibility discoveries of PALM are widely supported by existing studies from other biological perspectives. PALM will be useful for further GWAS analytical strategies that use gene-level analyses. Feng Bao 0002, Yue Deng 0001, Mulong Du, Zhiquan Ren, Yanyu Zhao, Jin-Li Suo, Meilin Wang, Qionghai Dai |
Briefings Bioinform. | 7 |
| 2017 | Multispectral focal stack acquisition using a chromatic aberration enlarged cameraabstractCapturing more information, e.g. geometry and material, using optical cameras can greatly help the perception and understanding of complex scenes. This paper proposes a novel method to capture the spectral and light field information simultaneously. By using a delicately designed chromatic aberration enlarged camera, the spectral-varying slices at different depths of the scene can be easily captured. Afterwards, the multispectral focal stack, which is composed of a stack of multispectral slice images focusing on different depths, can be recovered from the spectral-varying slices by using a Local Linear Transformation (LLT) based algorithm. The experiments verify the effectiveness of the proposed method. Yunqian Li, Linsen Chen, Xiaoming Zhong, Jin-Li Suo, Zhan Ma 0001, Tao Yue 0003, Xun Cao |
ICIP | 5 |
| 2017 | Emerging theories and technologies on computational imagingabstractComputational imaging describes the whole imaging process from the perspective of light transport and information transmission, features traditional optical computing capabilities, and assists in breaking through the limitations of visual information recording. Progress in computational imaging promotes the development of diverse basic and applied disciplines. In this review, we provide an overview of the fundamental principles and methods in computational imaging, the history of this field, and the important roles that it plays in the development of science. We review the most recent and promising advances in computational imaging, from the perspective of different dimensions of visual signals, including spatial dimension, temporal dimension, angular dimension, spectral dimension, and phase. We also discuss some topics worth studying for future developments in computational imaging. Jin-Li Suo, Qionghai Dai |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | Frequency-Domain Transient ImagingabstractA transient image is the optical impulse response of a scene, which also visualizes the propagation of light during an ultra-short time interval. In contrast to the previous transient imaging which samples in the time domain using an ultra-fast imaging system, this paper proposes transient imaging in the frequency domain using a multi-frequency time-of-flight (ToF) camera. Our analysis reveals the Fourier relationship between transient images and the measurements of a multi-frequency ToF camera, and identifies the causes of the systematic error-non-sinusoidal and frequency-varying waveforms and limited frequency range of the modulation signal. Based on the analysis we propose a novel framework of frequency-domain transient imaging. By removing the systematic error and exploiting the harmonic components inside the measurements, we achieves high quality reconstruction results. Moreover, our technique significantly reduces the computational cost of ToF camera based transient image reconstruction, especially reduces the memory usage, such that it is feasible for the reconstruction of transient images at extremely small time steps. The effectiveness of frequency-domain transient imaging is tested on synthetic data, real data from the web, and real data acquired by our prototype camera. Yebin Liu, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Efficient Method for High-Quality Removal of Nonuniform Blur in the Wavelet DomainabstractThis paper presents a novel nonuniform deblurring approach, which defines the blur model and calculates regularized nonuniform deconvolution in the wavelet domain to achieve high efficiency and high accuracy simultaneously. Targeting high computation efficiency, we derive a wavelet-domain hierarchical blur model, which can be calculated efficiently by exploiting the sparsity property of natural images in the wavelet domain. Correspondingly, the blur model is incorporated into a multilayer framework and at each layer spatially varying step sizes are introduced to further accelerate the convergence of the algorithm. In addition to the efficiency advantages, the proposed approach deals with intensely nonuniform blur with high accuracy due to the intrinsic tight supportness of wavelet basis. We conduct a series of experiments and comparisons to validate the efficiency and effectiveness of our algorithm. Tao Yue 0003, Jin-Li Suo, Xun Cao, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Signal-dependent noise removal for color videos using temporal and cross-channel priors
Jin-Li Suo, Liheng Bian, Feng Chen 0007, Qionghai Dai |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Normalized filter pool for prior modeling of nature images
Yangang Wang 0001, Jin-Li Suo, Qionghai Dai |
Mach. Vis. Appl. | 2 |
| 2016 | Fast and High Quality Highlight Removal From a Single ImageabstractSpecular reflection exists widely in photography and causes the recorded color deviating from its true value, thus, fast and high quality highlight removal from a single nature image is of great importance. In spite of the progress in the past decades in highlight removal, achieving wide applicability to the large diversity of nature scenes is quite challenging. To handle this problem, we propose an analytic solution to highlight removal based on an L2chromaticity definition and corresponding dichromatic model. Specifically, this paper derives a normalized dichromatic model for the pixels with identical diffuse color: a unit circle equation of projection coefficients in two subspaces that are orthogonal to and parallel with the illumination, respectively. In the former illumination orthogonal subspace, which is specular-free, we can conduct robust clustering with an explicit criterion to determine the cluster number adaptively. In the latter, illumination parallel subspace, a property called pure diffuse pixels distribution rule helps map each specular-influenced pixel to its diffuse component. In terms of efficiency, the proposed approach involves few complex calculation, and thus can remove highlight from high resolution images fast. Experiments show that this method is of superior performance in various challenging cases. Jin-Li Suo, Dongsheng An, Xiangyang Ji, Haoqian Wang, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2015 | Blind optical aberration correction by exploring geometric and visual priorsabstractOptical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric and visual priors. The global rotational symmetry allows us to transform the non-uniform degeneration into several uniform ones by the proposed radial splitting and warping technique. Locally, two types of symmetry constraints, i.e. central symmetry and reflection symmetry are defined as geometric priors in central and surrounding regions, respectively. Furthermore, by investigating the visual artifacts of aberration degenerated images captured by consumer-level cameras, the non-uniform distribution of sharpness across color channels and the image lattice is exploited as visual priors, resulting in a novel strategy to utilize the guidance from the sharpest channel and local image regions to improve the overall performance and robustness. Extensive evaluation on both real and synthetic data suggests that the proposed method outperforms the state-of-the-art techniques. Tao Yue 0003, Jin-Li Suo, Jue Wang 0001, Xun Cao, Qionghai Dai |
CVPR | 2 |
| 2015 | Generalized iterative phase retrieval algorithms and their applicationsabstractIt is well known that the phase contains more important information about the field in comparison with the amplitude. Therefore the imaging of phase is encountered in many branches of modern science and engineering. Direct measurement of the phase is easy in the long wavelength regime of the electromagnetic spectrum, but is difficult in the short regime such as the visible light due to the limited bandwidth of imaging sensors. One must employ computational techniques to extract the phase from the captured intensity. So far many methods have been proposed for this task. These algorithms can be basically classified into three categories: Holography, deterministic algorithms such as the transport of intensity equation, and iterative algorithms such as the Gerchberg-Saxton-Fienup-type algorithm. Each of these algorithms has its own advantages and disadvantages. This paper mainly focuses on the our previous works on iterative phase retrieval techniques, and their applications in the calculation of computer-generated holograms, microscopic imaging, and optical signal processing. Guohai Situ, Jin-Li Suo, Qionghai Dai |
INDIN | 2 |
| 2015 | Extracting Depth and Radiance From a Defocused Video PairabstractWe present a novel iterative feedback approach for the simultaneous estimation of depth and all-in-focus (AIF) videos from a defocused video pair by joint spatiotemporal optimization. Depth and AIF videos benefit each other in the iterative optimization. First, for the recovery of AIF video, the sparse prior of natural video is incorporated to ensure a high-quality defocus blur removal even under inaccurate depth estimation. Second, in depth estimation step, we feed back the spatial and temporal constraints from the high-quality AIF video and adopt a numerical solution, which is robust to the inaccuracy of AIF recovery to further boost the performance of depth from the defocus algorithm. Benefitting from the incorporation of AIF video priors and the temporal consistency constraint, the proposed framework can effectively reconstruct the depth of the textureless region and is insensitive to camera parameter changes. Our approach provides better temporal consistency and higher depth accuracy than the conventional method that applies postsmoothing to the sequential frame estimation. We not only demonstrate the feasibility of our approach via real experimentation but also provide visual and quantitative evaluation on synthetic data. Jin-Li Suo, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Transparent Object Reconstruction via Coded Transport of IntensityabstractCapturing and understanding visual signals is one of the core interests of computer vision. Much progress has been made w.r.t. many aspects of imaging, but the reconstruc-tion of refractive phenomena, such as turbulence, gas and heat flows, liquids, or transparent solids, has remained a challenging problem. In this paper, we derive an intuitive formulation of light transport in refractive media using light fields and the transport of intensity equation. We show how coded illumination in combination with pairs of recorded images allow for robust computational reconstruction of dy-namic two and three-dimensional refractive phenomena. 1. Chenguang Ma, Jin-Li Suo, Qionghai Dai, Gordon Wetzstein |
CVPR | 3 |
| 2014 | Recovering Scene Geometry under Wavy Fluid via Distortion and Defocus Analysis
Mohit Gupta 0001, Jin-Li Suo, Qionghai Dai |
ECCV (5) | 4 |
| 2014 | Automatic inpainting of linearly related video framesabstractThis paper addresses automatic inpainting of a specific but common kind of videos captured by imaging a far or planar scene with a moving camera. The projective model tells that the frames of such videos can be approximately aligned by linear mappings except for some to-be-inpainted small regions. Mathematically, we treat inpainting as a global optimization with a linear system incorporating both the temporal consistency and the priors of the inpainting regions: (i) temporally registered frames form a low rank matrix; (ii) the pixels in the given inpainting regions destroy the low rank-ness with gross sparse errors. Besides, we also use a soft mask to ensure consistent global brightness before and after inpainting. Further, we propose a numerical solution to above optimization based on Augmented Lagrangian Method. The experiment results demonstrated our advantageous in both preserving thin scene structures and the details prone to be smoothed out by previous methods. Yudong Xiao, Jin-Li Suo, Liheng Bian, Qionghai Dai |
ICIP | 2 |
| 2014 | Deblur a blurred RGB image with a sharp NIR image through local linear mappingabstractImage acquisition in a low light environment requires long exposure to achieve acceptable signal-to-noise ratio, which however causes blurry effect. This paper addresses this problem by using a sharp near-infrared (NIR) image when the environment has sufficient NIR light. We assume that an RGB and NIR image pair has a linear mapping in a local area and that the mapping function is valid for both the blur and sharp image pairs. Using this property, we solve the sharp RGB images from a blurred RGB image and the corres ponding s harp NIR image. The effectiveness of the proposed algorithm is verified with both synthetic and real captured datasets. Tao Yue 0003, Ming-Ting Sun, Zhengyou Zhang, Jin-Li Suo, Qionghai Dai |
ICME | 4 |
| 2014 | Robust Image Restoration via Reweighted Low-Rank Matrix Recovery
YiGang Peng, Jin-Li Suo, Qionghai Dai, Wenli Xu |
MMM (1) | 2 |
| 2014 | Reweighted Low-Rank Matrix Recovery and its Application in Image RestorationabstractIn this paper, we propose a reweighted low-rank matrix recovery method and demonstrate its application for robust image restoration. In the literature, principal component pursuit solves low-rank matrix recovery problem via a convex program of mixed nuclear norm and l1 norm. Inspired by reweighted l1 minimization for sparsity enhancement, we propose reweighting singular values to enhance low rank of a matrix. An efficient iterative reweighting scheme is proposed for enhancing low rank and sparsity simultaneously and the performance of low-rank matrix recovery is prompted greatly. We demonstrate the utility of the proposed method both on numerical simulations and real images/videos restoration, including single image restoration, hyperspectral image restoration, and background modeling from corrupted observations. All of these experiments give empirical evidence on significant improvements of the proposed algorithm over previous work on low-rank matrix recovery. YiGang Peng, Jin-Li Suo, Qionghai Dai, Wenli Xu |
IEEE Trans. Cybern. | 2 |
| 2014 | Joint Non-Gaussian Denoising and Superresolving of Raw High Frame Rate VideosabstractHigh frame rate cameras capture sharp videos of highly dynamic scenes by trading off signal-noise-ratio and image resolution, so combinational super-resolving and denoising is crucial for enhancing high speed videos and extending their applications. The solution is nontrivial due to the fact that two deteriorations co-occur during capturing and noise is nonlinearly dependent on signal strength. To handle this problem, we propose conducting noise separation and super resolution under a unified optimization framework, which models both spatiotemporal priors of high quality videos and signal-dependent noise. Mathematically, we align the frames along temporal axis and pursue the solution under the following three criterion: 1) the sharp noise-free image stack is low rank with some missing pixels denoting occlusions; 2) the noise follows a given nonlinear noise model; and 3) the recovered sharp image can be reconstructed well with sparse coefficients and an over complete dictionary learned from high quality natural images. In computation aspects, we propose to obtain the final result by solving a convex optimization using the modern local linearization techniques. In the experiments, we validate the proposed approach in both synthetic and real captured data. Jin-Li Suo, Yue Deng 0001, Liheng Bian, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2014 | High-Dimensional Camera Shake Removal With Given Depth MapabstractCamera motion blur is drastically nonuniform for large depth-range scenes, and the nonuniformity caused by camera translation is depth dependent but not the case for camera rotations. To restore the blurry images of large-depth-range scenes deteriorated by arbitrary camera motion, we build an image blur model considering 6-degrees of freedom (DoF) of camera motion with a given scene depth map. To make this 6D depth-aware model tractable, we propose a novel parametrization strategy to reduce the number of variables and an effective method to estimate high-dimensional camera motion as well. The number of variables is reduced by temporal sampling motion function, which describes the 6-DoF camera motion by sampling the camera trajectory uniformly in time domain. To effectively estimate the high-dimensional camera motion parameters, we construct the probabilistic motion density function (PMDF) to describe the probability distribution of camera poses during exposure, and apply it as a unified constraint to guide the convergence of the iterative deblurring algorithm. Specifically, PMDF is computed through a back projection from 2D local blur kernels to 6D camera motion parameter space and robust voting. We conduct a series of experiments on both synthetic and real captured data, and validate that our method achieves better performance than existing uniform methods and nonuniform methods on large-depth-range scenes. Tao Yue 0003, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Image Process. | 2 |
| 2013 | Coded focal stack photographyabstractWe present coded focal stack photography as a computational photography paradigm that combines a focal sweep and a coded sensor readout with novel computational algorithms. We demonstrate various applications of coded focal stacks, including photography with programmable non-planar focal surfaces and multiplexed focal stack acquisition. By leveraging sparse coding techniques, coded focal stacks can also be used to recover a full-resolution depth and all-in-focus (AIF) image from a single photograph. Coded focal stack photography is a significant step towards a computational camera architecture that facilitates high-resolution post-capture refocusing, flexible depth of field, and 3D imaging. Jin-Li Suo, Gordon Wetzstein, Qionghai Dai, Ramesh Raskar |
ICCP | 2 |
| 2013 | High-rank coded aperture projection for extended depth of fieldabstractProjectors require large apertures to maximize light throughput. Unfortunately, this leads to shallow depths of field (DOF), hence blurry images, when projecting on non-planar surfaces, such as cultural heritage sites, curved screens, or when sharing visual information in everyday environments. We introduce high-rank coded aperture projectors - a new computational display technology that combines optical designs with computational processing to overcome depth of field limitations of conventional devices. In particular, we employ high-speed spatial light modulators (SLMs) on the image plane and in the aperture of modified projectors. The patterns displayed on these SLMs are computed with a new mathematical framework that uses high-rank light field factorizations and directly exploits the limited temporal resolution and contrast sensitivity of the human visual system. With an experimental prototype projector, we demonstrate significantly increased DOF as compared to conventional technology. Chenguang Ma, Jin-Li Suo, Qionghai Dai, Ramesh Raskar, Gordon Wetzstein |
ICCP | 2 |
| 2013 | Non-uniform image deblurring using an optical computing system
Tao Yue 0003, Jin-Li Suo, Qionghai Dai |
Comput. Graph. | 2 |
| 2012 | Iterative Feedback Estimation of Depth and Radiance from Defocused Images
Jin-Li Suo, Xun Cao, Qionghai Dai |
ACCV (4) | 2 |
| 2012 | An overview of computational photography
Jin-Li Suo, Xiangyang Ji, Qionghai Dai |
Sci. China Inf. Sci. | 1 |
| 2012 | A Concatenational Graph Evolution Aging ModelabstractModeling the long-term face aging process is of great importance for face recognition and animation, but there is a lack of sufficient long-term face aging sequences for model learning. To address this problem, we propose a CONcatenational GRaph Evolution (CONGRE) aging model, which adopts decomposition strategy in both spatial and temporal aspects to learn long-term aging patterns from partially dense aging databases. In spatial aspect, we build a graphical face representation, in which a human face is decomposed into mutually interrelated subregions under anatomical guidance. In temporal aspect, the long-term evolution of the above graphical representation is then modeled by connecting sequential short-term patterns following the Markov property of aging process under smoothness constraints between neighboring short-term patterns and consistency constraints among subregions. The proposed model also considers the diversity of face aging by proposing probabilistic concatenation strategy between short-term patterns and applying scholastic sampling in aging prediction. In experiments, the aging prediction results generated by the learned aging models are evaluated both subjectively and objectively to validate the proposed model. Jin-Li Suo, Xilin Chen 0001, Shiguang Shan, Wen Gao 0001, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | High-Resolution Face Fusion for Gender ConversionabstractThis paper presents an integrated face image fusion framework, which combines a hierarchical compositional paradigm with seamless image-editing techniques, for gender conversion. In our framework a high-resolution face is represented by a probabilistic graphical model that decomposes a human face into several parts (facial components) constrained by explicit spatial configurations (relationships). Benefiting from this representation, the proposed fusion strategy is able to largely preserve the face identity of each facial component while applying gender transformation. Given a face image, the basic idea is to select reference facial components from the opposite-gender group as templates and transform the appearance of the given image toward the selected facial components. Our fusion approach decomposes a face image into two parts-sketchable and nonsketchable ones. For the sketchable regions (e.g., the contours of facial components and wrinkle lines, etc.), we use a graph-matching algorithm to find the best templates and transform the structure (shape), while for the nonsketchable regions (e.g., the texture area of facial components, skin, etc.), we learn active appearance models and transform the texture attributes in the corresponding principal component analysis space. Both objective and subjective quantitative evaluation results on 200 Asian frontal-face images selected from the public Lotus Hill Image database show that the proposed approach is able to give plausible gender conversion results. Jin-Li Suo, Liang Lin 0004, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2010 | A Compositional and Dynamic Model for Face AgingabstractIn this paper, we present a compositional and dynamic model for face aging. The compositional model represents faces in each age group by a hierarchical And-Or graph, in which And nodes decompose a face into parts to describe details (e.g., hair, wrinkles, etc.) crucial for age perception and Or nodes represent large diversity of faces by alternative selections. Then a face instance is a transverse of the And-Or graph-parse graph. Face aging is modeled as a Markov process on the parse graph representation. We learn the parameters of the dynamic model from a large annotated face data set and the stochasticity of face aging is modeled in the dynamics explicitly. Based on this model, we propose a face aging simulation and prediction algorithm. Inversely, an automatic age estimation algorithm is also developed under this representation. We study two criteria to evaluate the aging results using human perception experiments: 1) the accuracy of simulation: whether the aged faces are perceived of the intended age group, and 2) preservation of identity: whether the aged faces are perceived as the same person. Quantitative statistical analysis validates the performance of our aging model and age estimation algorithm. Jin-Li Suo, Song-Chun Zhu, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Learning long term face aging patterns from partially dense aging databasesabstractStudies on face aging are handicapped by lack of long term dense aging sequences for model training. To handle this problem, we propose a new face aging model, which learns long term face aging patterns from partially dense aging databases. The learning strategy is based on two assumptions: (i) short term face aging pattern is relatively simple and is possible to be learned from currently available databases; (ii) long term face aging is a continuous and smooth Markov process. Adopting a compositional face representation, our aging algorithm learns a function-based short term aging model from real aging sequences to infer facial parameters within a short age span. Based on the predefined smoothness criteria between two overlapping short term aging patterns, we concatenate these learned short term aging patterns to build the long term aging patterns. Both the subjective assessment and objective evaluations of synthetic aging sequences validate the effectiveness of the proposed model. Jin-Li Suo, Xilin Chen 0001, Shiguang Shan, Wen Gao 0001 |
ICCV | 1 |
| 2008 | Design sparse features for age estimation using hierarchical face modelabstractA key point in automatic age estimation is to design feature set essential to age perception. To achieve this goal, this paper builds up a hierarchical graphical face model for faces appearing at low, middle and high resolution respectively. Along the hierarchy, a face image is decomposed into detailed parts from coarse to fine. Then four types of features are extracted from this graph representation guided by the priors of aging process embedded in the graphical model: topology, geometry, photometry and configuration. On age estimation, this paper follows the popular regression formulation for mapping feature vectors to its age label. The effectiveness of the presented feature set is justified by testing results on two datasets using different kinds of regression methods. The experimental results in this paper show that designing feature set for age estimation under the guidance of hierarchical face model is a promising method and a flexible framework as well. Jin-Li Suo, Tianfu Wu 0001, Song-Chun Zhu, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
FG | 1 |
| 2007 | A Multi-Resolution Dynamic Model for Face Aging SimulationabstractIn this paper we present a dynamic model for simulating face aging process. We adopt a high resolution grammatical face model[1] and augment it with age and hair features. This model represents all face images by a multi-layer And-Or graph and integrates three most prominent aspects related to aging changes: global appearance changes in hair style and shape, deformations and aging effects of facial components, and wrinkles appearance at various facial zones. Then face aging is modeled as a dynamic Markov process on this graph representation which is learned from a large dataset. Given an input image, we firstly compute the graph representation, and then sample the graph structures over various age groups according to the learned dynamic model. Finally we generate new face images with the sampled graphs. Our approach has three novel aspects: (1) the aging model is learned from a dataset of 50,000 adult faces at different ages; (2) we explicitly model the uncertainty in face aging and can sample multiple plausible aged faces for an input image; and (3) we conduct a simple human experiment to validate the simulated aging process. Jin-Li Suo, Feng Min, Song-Chun Zhu, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |