VLDB 2026 Research / reviewers in the wild / expert
Yinqiang Zheng
dblp:79/5068
· DBLP profile ↗
130ranked-venue papers
15as first author
69since 2021 · last 2026
0000-0001-7434-5069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 113 · 15 first-author · 55 since 2021Artificial intelligence and machine learning · 101 · 14 first-author · 55 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KaoLRM: Repurposing Pre-Trained Large Reconstruction Models for Parametric 3D Face ReconstructionabstractWe propose KaoLRM to re-target the learned prior of the Large Reconstruction Model (LRM) for parametric 3D face reconstruction from single-view images. Parametric 3D Morphable Models (3DMMs) have been widely used for facial reconstruction due to their compact and interpretable parameterization, yet existing 3DMM regressors often exhibit poor consistency across varying viewpoints. To address this, we harness the pre-trained 3D prior of LRM and incorporate FLAME-based 2D Gaussian Splatting into LRM's rendering pipeline. Specifically, KaoLRM projects LRM's pre-trained triplane features into the FLAME parameter space to recover geometry, and models appearance via 2D Gaussian primitives that are tightly coupled to the FLAME mesh. The rich prior enables the FLAME regressor to be aware of the 3D structure, leading to accu-rate and robust reconstructions under self-occlusions and diverse viewpoints. Experiments on both controlled and in-the-wild benchmarks demonstrate that KaoLRM achieves superior reconstruction accuracy and cross-view consistency, while existing methods remain sensitive to viewpoint variations. The code is released at https://github.com/CyberAgentAILab/KaoLRM. Qingtian Zhu, Zhixiang Wang 0001, Yinqiang Zheng, Takafumi Taketomi |
3DV | 4 |
| 2026 | HarmoQ: Harmonized Post-Training Quantization for High-Fidelity Image Super-ResolutionabstractPost-training quantization offers an efficient pathway to deploy super-resolution models, yet existing methods treat weight and activation quantization independently, missing their critical interplay. Through controlled experiments on SwinIR, we uncover a striking asymmetry: weight quantization primarily degrades structural similarity, while activation quantization disproportionately affects pixel-level accuracy. This stems from their distinct roles—weights encode learned restoration priors for textures and edges, whereas activations carry input-specific intensity information. Building on this insight, we propose HarmoQ, a unified framework that harmonizes quantization across components through three synergistic steps: structural residual calibration proactively adjusts weights to compensate for activation-induced detail loss, harmonized scale optimization analytically balances quantization difficulty via closed-form solutions, and adaptive boundary refinement iteratively maintains this balance during optimization. Experiments show HarmoQ achieves substantial gains under aggressive compression, outperforming prior art by 0.46 dB on Set5 at 2-bit while delivering 3.2× speedup and 4× memory reduction on A100 GPUs. This work provides the first systematic analysis of weight-activation coupling in super-resolution quantization and establishes a principled solution for efficient high-quality image restoration. Hongjun Wang 0007, Jiyuan Chen, Xuan Song 0001, Yinqiang Zheng |
AAAI | 4 |
| 2025 | Adversarial Attacks on Event-Based Pedestrian Detectors: A Physical ApproachabstractEvent cameras, known for their low latency and high dynamic range, show great potential in pedestrian detection applications. However, while recent research has primarily focused on improving detection accuracy, the robustness of event-based visual models against physical adversarial attacks has received limited attention. For example, adversarial physical objects, such as specific clothing patterns or accessories, can exploit inherent vulnerabilities in these systems, leading to misdetections or misclassifications. This study is the first to explore physical adversarial attacks on event-driven pedestrian detectors, specifically investigating whether certain clothing patterns worn by pedestrians can cause these detectors to fail, effectively rendering them unable to detect the person. To address this, we developed an end-to-end adversarial framework in the digital domain, framing the design of adversarial clothing textures as a 2D texture optimization problem. By crafting an effective adversarial loss function, the framework iteratively generates optimal textures through backpropagation. Our results demonstrate that the textures identified in the digital domain possess strong adversarial properties. Furthermore, we translated these digitally optimized textures into physical clothing and tested them in real-world scenarios, successfully demonstrating that the designed textures significantly degrade the performance of event-based pedestrian detection models. This work highlights the vulnerability of such models to physical adversarial attacks. Guixu Lin, Muyao Niu, Qingtian Zhu, Zhengwei Yin, Zhuoxiao Li, Shengfeng He, Yinqiang Zheng |
AAAI | 7 |
| 2025 | ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained GuidanceabstractDiffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a novel, training-free method that empowers resolution-limited diffusion models to generate high-quality images beyond resolution restrictions. Specifically, ResMaster leverages a low-resolution reference image created by a pre-trained diffusion model to provide structural and fine-grained guidance for crafting high-resolution images on a patch-by-patch basis. To ensure a coherent structure, ResMaster meticulously aligns the low-frequency components of high-resolution patches with the low-resolution reference at each denoising step. For fine-grained guidance, tailored image prompts based on the low-resolution reference and enriched textual prompts produced by a vision-language model are incorporated. This approach could significantly mitigate local pattern distortions and improve detail refinement. Extensive experiments validate that ResMaster sets a new benchmark for high-resolution image generation. Shuwei Shi, Yuechen Zhang, Jingwen He, Biao Gong, Yinqiang Zheng |
AAAI | 6 |
| 2025 | Instruction-based Image Manipulation by Watching How Things MoveabstractThis paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects and scenes, ensuring consistent content preservation during editing. Additionally, video data captures diverse, natural dynamics—such as non-rigid subject motion and complex camera movements—that are difficult to model otherwise, making it an ideal source for scalable dataset construction. Using this approach, we create a new dataset to train InstructMove, a model capable of instruction-based complex manipulations that are difficult to achieve with synthetically generated datasets. Our model demonstrates state-of-the-art performance in tasks such as adjusting subject poses, rearranging elements, and altering camera perspectives. The project page is available here. Mingdeng Cao, Xuaner Cecilia Zhang, Yinqiang Zheng, Zhihao Xia |
CVPR | 3 |
| 2025 | MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video GenerationabstractThe image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training. Shuwei Shi, Biao Gong, Zizheng Yang, Yuyuan Li 0001, Jingwen He, Kecheng Zheng, Jingdong Chen, Ming Yang 0007, Yinqiang Zheng |
CVPR | 12 |
| 2025 | Not All Degradations are Equal: A Targeted Feature Denoising Framework for Generalizable Image Super-ResolutionabstractGeneralizable Image Super-Resolution aims to enhance model generalization capabilities under unknown degradations. To achieve this goal, the models are expected to focus only on image content-related features instead of overfitting degradations. Recently, numerous approaches such as Dropout and Feature Alignment have been proposed to suppress models' natural tendency to overfit degradations and yield promising results. Nevertheless, these works have assumed that models overfit to all degradation types (e.g., blur, noise, JPEG), while through careful investigations in this paper, we discover that models predominantly overfit to noise, largely attributable to its distinct degradation pattern compared to other degradation types. In this paper, we propose a targeted feature denoising framework, comprising noise detection and denoising modules. Our approach presents a general solution that can be seamlessly integrated with existing super-resolution models without requiring architectural modifications. Our framework demonstrates superior performance compared to previous regularization-based methods across five traditional benchmarks and datasets, encompassing both synthetic and real-world scenarios. Hongjun Wang 0007, Jiyuan Chen, Zhengwei Yin, Xuan Song 0001, Yinqiang Zheng |
ICCV | 5 |
| 2025 | Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars
Yifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma, Jiancheng Zhao, Zhihang Zhong, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ICCV | 9 |
| 2025 | Tree-NeRV: Efficient Non-Uniform Sampling for Neural Video Representation via Tree-Structured Feature Grids
Jiancheng Zhao, Yifan Zhan, Qingtian Zhu, Mingze Ma, Muyao Niu, Zunian Wan, Xiang Ji 0005, Yinqiang Zheng |
ICCV | 8 |
| 2025 | Revolutionizing EMCCD Denoising through a Novel Physics-Based Learning Framework for Noise ModelingabstractElectron-multiplying charge-coupled device (EMCCD) has been instrumental in sensitive observations under low-light situations including astronomy, material science, and biology.
Despite its ingenious designs to enhance target signals overcoming read-out circuit noises, produced images are not completely noise free, which could still cast a cloud on desired experiment outcomes, especially in fluorescence microscopy.
Existing studies on EMCCD's noise model have been focusing on statistical characteristics in theory, yet unable to incorporate latest advancements in the field of computational photography, where physics-based noise models are utilized to guide deep learning processes, creating adaptive denoising algorithms for ordinary image sensors.
Still, those models are not directly applicable to EMCCD.
In this paper, we intend to pioneer EMCCD denoising by introducing a systematic study on physics-based noise model calibration procedures for an EMCCD camera, accurately estimating statistical features of observable noise components in experiments, which are then utilized to generate substantial amount of authentic training samples for one of the most recent neural networks.
A first real-world test image dataset for EMCCD is captured, containing both images of ordinary daily scenes and those of microscopic contents.
Benchmarking upon the testset and authentic microscopic images, we demonstrate distinct advantages of our model against previous methods for EMCCD and physics-based noise modeling, forging a promising new path for EMCCD denoising. Haiyang Jiang 0002, Tetsuichi Wazawa, Imari Sato, Takeharu Nagai, Yinqiang Zheng |
ICLR | 5 |
| 2025 | Random Is All You Need: Random Noise Injection on Feature Statistics for Generalizable Deep Image DenoisingabstractRecent advancements in generalizable deep image denoising have catalyzed the development of robust noise-handling models. The current state-of-the-art, Masked Training (MT), constructs a masked swinir model which is trained exclusively on Gaussian noise ($\sigma$=15) but can achieve commendable denoising performance across various noise types (*i.e.* speckle noise, poisson noise). However, this method, while focusing on content reconstruction, often produces over-smoothed images and poses challenges in mask ratio optimization, complicating its integration with other methodologies. In response, this paper introduces RNINet, a novel architecture built on a streamlined encoder-decoder framework to enhance both efficiency and overall performance. Initially, we train a pure RNINet (only simple encoder-decoder) on individual noise types, observing that feature statistics such as mean and variance shift in response to different noise conditions. Leveraging these insights, we incorporate a noise injection block that injects random noise into feature statistics within our framework, significantly improving generalization across unseen noise types. Our framework not only simplifies the architectural complexity found in MT but also delivers superior performance. Comprehensive experimental evaluations demonstrate that our method outperforms MT in various unseen noise conditions in terms of denoising effectiveness and computational efficiency (lower MACs and GPU memory usage), achieving up to 10 times faster inference speeds and underscoring it's capability for large scale deployments. Zhengwei Yin, Hongjun Wang 0007, Guixu Lin, Weihang Ran, Yinqiang Zheng |
ICLR | 5 |
| 2025 | You Always Recognize Me (YARM): Robust Texture Synthesis Against Multi-View CorruptionabstractDamage to imaging systems and complex external environments often introduce corruption, which can impair the performance of deep learning models pretrained on high-quality image data. Previous methods have focused on restoring degraded images or fine-tuning models to adapt to out-of-distribution data. However, these approaches struggle with complex, unknown corruptions and often reduce model accuracy on high-quality data. Inspired by the use of warning colors and camouflage in the real world, we propose designing a robust appearance that can enhance model recognition of low-quality image data. Furthermore, we demonstrate that certain universal features in radiance fields can be applied across objects of the same class with different geometries. We also examine the impact of different proxy models on the transferability of robust appearances. Extensive experiments demonstrate the effectiveness of our proposed method, which outperforms existing image restoration and model fine-tuning approaches across different experimental settings, and retains effectiveness when transferred to models with different architectures. Code will be available at https://github.com/SilverRAN/YARM. Weihang Ran, Wei Yuan 0004, Yinqiang Zheng |
ICML | 3 |
| 2025 | SUICA: Learning Super-high Dimensional Sparse Implicit Neural Representations for Spatial TranscriptomicsabstractSpatial Transcriptomics (ST) is a method that captures gene expression profiles aligned with spatial coordinates. The discrete spatial distribution and the super-high dimensional sequencing results make ST data challenging to be modeled effectively. In this paper, we manage to model ST in a continuous and compact manner by the proposed tool, SUICA, empowered by the great approximation capability of Implicit Neural Representations (INRs) that can enhance both the spatial density and the gene expression. Concretely within the proposed SUICA, we incorporate a graph-augmented Autoencoder to effectively model the context information of the unstructured spots and provide informative embeddings that are structure-aware for spatial mapping. We also tackle the extremely skewed distribution in a regression-by-classification fashion and enforce classification-based loss functions for the optimization of SUICA. By extensive experiments of a wide range of common ST platforms under varying degradations, SUICA outperforms both conventional INR variants and SOTA methods regarding numerical fidelity, statistical correlation, and bio-conservation. The prediction by SUICA also showcases amplified gene signatures that enriches the bio-conservation of the raw data and benefits subsequent analysis. Qingtian Zhu, Yumin Zheng, Yuling Sang, Yifan Zhan, Ziyan Zhu, Yinqiang Zheng |
ICML | 7 |
| 2025 | Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task ConditioningabstractWe introduce Dr. RAW, a unified and tuning-efficient framework for high-level computer vision tasks directly operating on camera RAW data. Unlike previous approaches that optimize image signal processing (ISP) pipelines and fully fine-tune networks for each task, Dr. RAW achieves state-of-the-art performance with minimal parameter updates. At the input stage, we apply lightweight pre-processing modules, sensor and illumination mapping, followed by re-mosaicing, to mitigate data inconsistencies stemming from sensor variation and lighting. At the network level, we introduce task-specific adaptation through two modules: Sensor Prior Prompts (SPP) and Low-Rank Adaptation (LoRA). SPP injects sensor-aware conditioning into the network via learnable prompts derived from imaging priors, while LoRA enables efficient task-specific tuning by updating only low-rank matrices in key backbone layers. Despite minimal tuning, our method delivers superior results across four RAW-based tasks (object detection, semantic segmentation, instance segmentation, and pose estimation) on nine datasets encompassing low-light and over-exposed conditions. By harnessing the intrinsic physical cues of RAW data alongside parameter-efficient techniques, our method advances RAW-based vision systems, achieving both high accuracy and computational economy. We will release our source code. Wenjun Huang 0001, Ziteng Cui, Yinqiang Zheng, Yirui He, Tatsuya Harada, Mohsen Imani |
NeurIPS | 3 |
| 2025 | EventHDR: From Event to High-Speed HDR Videos and BeyondabstractEvent cameras are innovative neuromorphic sensors that asynchronously capture the scene dynamics. Due to the event-triggering mechanism, such cameras record event streams with much shorter response latency and higher intensity sensitivity compared to conventional cameras. On the basis of these features, previous works have attempted to reconstruct high dynamic range (HDR) videos from events, but have either suffered from unrealistic artifacts or failed to provide sufficiently high frame rates. In this paper, we present a recurrent convolutional neural network that reconstruct high-speed HDR videos from event sequences, with a key frame guidance to prevent potential error accumulation caused by the sparse event data. Additionally, to address the problem of severely limited real dataset, we develop a new optical system to collect a real-world dataset with paired high-speed HDR videos and event streams, facilitating future research in this field. Our dataset provides the first real paired dataset for event-to-HDR reconstruction, avoiding potential inaccuracies from simulation strategies. Experimental results demonstrate that our method can generate high-quality, high-speed HDR videos. We further explore the potential of our work in cross-camera reconstruction and downstream computer vision tasks, including object detection, panoramic segmentation, optical flow estimation, and monocular depth estimation under HDR scenarios. Yunhao Zou, Ying Fu 0001, Tsuyoshi Takatani, Yinqiang Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Polarization State Attention Dehazing Network With a Simulated Polar-Haze DatasetabstractImage dehazing under harsh weather conditions remains a challenging and ill-posed problem. In addition, acquiring real-time haze-free counterparts of hazy images poses difficulties. Existing approaches commonly synthesize hazy data by relying on estimated depth information, which is prone to errors due to its physical unreliability. While generative networks can transfer some hazy features to clear images, the resulting hazy images still exhibit an artificial appearance. In this paper, we introduce polarization cues to propose a haze simulation strategy to synthesize hazy data, ensuring visually pleasing results that adhere to physical laws. Leveraging on the simulated Polar-Haze dataset, we present a polarization state attention dehazing network (PSADNet), which consists of a polarization extraction module and a polarization dehazing module. The proposed polarization extraction model incorporates an attention mechanism to capture high-level image features related to polarization and chromaticity. The polarization dehazing module utilizes these features derived from the polarization analysis to enhance image dehazing capabilities while preserving the accuracy of the polarization information. Promising results are observed in both qualitative and quantitative experiments, supporting the effectiveness of the proposed PSADNet and the validity of polarization-based haze simulation strategy. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Multim. | 2 |
| 2025 | A Serial Perspective on Photometric Stereo of Filtering and Serializing Spatial InformationabstractIn this paper, we introduce a novel method of Filtering and Serializing Spatial Information to tackle uncalibrated photometric stereo tasks, termed FSSI-PS. Photometric stereo aims to recover surface normals from images with varying lighting and is crucial for tasks like 3D reconstruction and defect detection. Current methods in complex surface reconstruction are costly and inaccurate due to redundant feature representations from GCN or Transformer modules, caused by the weak global information extraction capability of GCNs or the large computational cost of Transformers. Furthermore, the trainset's lack of richness in texture complexity makes reconstruction more difficult. We address these issues by optimizing feature maps and dataset richness through serializing and filtering. First, we use Mamba-RNN to optimize feature representation by directly fusing feature maps, which reduces redundancy and uses minimal computational resources. Specifically, we treat input spatial information as a sequence and serialize it by sorting. Furthermore, we introduce the Mean Angular Variation metric to assess reconstruction difficulty by measuring texture complexity. It classifies PS-Sculpture and PS-Blobby into three categories: Difficult, Normal, and Simple. We use this to construct DNS-S+B, a photometric stereo training set with rich complexity levels. Our method is compared with state-of-the-art methods on the DiLiGenT and LUCES benchmarks to highlight effectiveness. Minzhe Xu, You Yang 0002, Yinqiang Zheng, Qiong Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Navigating Beyond Dropout: An Intriguing Solution Towards Generalizable Image Super ResolutionabstractDeep learning has led to a dramatic leap on Single Image Super-Resolution (SISR) performances in recent years. While most existing work assumes a simple and fixed degradation model (e.g., bicubic downsampling), the research of Blind SR seeks to improve model generalization ability with unknown degradation. Recently, Kong et al. [37] pioneer the investigation of a more suitable training strategy for Blind SR using Dropout [63]. Although such method indeed brings substantial generalization improvements via mitigating overfitting, we argue that Dropout simultaneously introduces undesirable side-effect that compromises model's capacity to faithfully reconstruct fine details. We show both the theoretical and experimental analyses in our paper, and furthermore, we present another easy yet effective training strategy that enhances the generalization ability of the model by simply modulating its first and second-order features statistics. Experimental results have shown that our method could serve as a model-agnostic regularization and outperforms Dropout on seven benchmark datasets including both synthetic and real-world scenarios. Hongjun Wang 0007, Jiyuan Chen, Yinqiang Zheng, Tieyong Zeng |
CVPR | 3 |
| 2024 | Rolling Shutter Correction with Intermediate Distortion Flow EstimationabstractThis paper proposes to correct the rolling shutter (RS) distorted images by estimating the distortion flow from the global shutter (GS) to RS directly. Existing methods usually perform correction using the undistortion flow from the RS to GS. They initially predict the flow from consecutive RS frames, subsequently rescaling it as the displacement fields from the RS frame to the underlying GS image using time-dependent scaling factors. Following this, RS-aware forward warping is employed to convert the RS image into its GS counterpart. Nevertheless, this strategy is prone to two shortcomings. First, the undistortion flow estimation is rendered inaccurate by merely linear scaling the flow, due to the complex non-linear motion nature. Second, RS-aware forward warping often results in unavoidable artifacts. To address these limitations, we introduce a new framework that directly estimates the distortion flow and rectifies the RS image with the backward warping operation. More specifically, we first propose a global correlation-based flow attention mechanism to estimate the initial distortion flow and GS feature jointly, which are then refined by the following coarse-to-fine decoder layers. Additionally, a multi-distortion flow prediction strategy is integrated to mitigate the issue of inaccurate flow estimation further. Experimental results validate the effectiveness of the proposed method, which outperforms state-of-the-art approaches on various benchmarks while maintaining high efficiency. The project is available at https://github.com/ljzycmd/DFRSC. Mingdeng Cao, Sidi Yang, Yujiu Yang 0001, Yinqiang Zheng |
CVPR | 4 |
| 2024 | IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame InterpolationabstractAdvanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However, they struggle to handle complex scenarios with curvilinear motions since they overlook the latent acceleration information between the input frames. Moreover, the supervision of predicted motions is tricky because ground-truth motions are not available. To this end, we propose a novel frame-work for implicit quadratic video frame interpolation (IQ-VFI), which explores latent acceleration information and accurate intermediate motions via knowledge distillation. Specifically, the proposed IQ-VFI consists of an implicit acceleration estimation network (IANet) and a VFI back-bone, the former fully leverages spatio-temporal information to explore latent acceleration priors between two input frames, which is then used to progressively modulate linear motions from the latter into quadratic motions in coarse-to-fine manner. Furthermore, to encourage both components to distill more acceleration and motion cues oriented towards VFI, we propose a knowledge distillation strategy in which implicit acceleration distillation loss and implicit motion distillation loss are employed to adaptively guide latent acceleration priors and intermediate motions learning, respectively. Extensive experiments show that our proposed IQ-VFI can achieve state-of-the-art performances on various benchmark datasets. Mengshun Hu, Kui Jiang, Zhihang Zhong, Zheng Wang 0007, Yinqiang Zheng |
CVPR | 5 |
| 2024 | Motion Blur Decomposition with Cross-shutter GuidanceabstractMotion blur is a frequently observed image artifact, especially under insufficient illumination where exposure time has to be prolonged so as to collect more photons for a bright enough image. Rather than simply removing such blurring effects, recent researches have aimed at decomposing a blurry image into multiple sharp images with spatial and temporal coherence. Since motion blur decomposition itself is highly ambiguous, priors from neighbouring frames or human annotation are usually needed for motion disambiguation. In this paper, inspired by the complementary exposure characteristics of a global shutter (GS) camera and a rolling shutter (RS) camera, we propose to utilize the ordered scanline-wise delay in a rolling shutter image to robustify motion decomposition of a single blurry image. To evaluate this novel dual imaging setting, we construct a triaxial system to collect realistic data, as well as a deep network architecture that explicitly addresses temporal and contextual information through reciprocal branches for cross-shutter motion blur decomposition. Experiment results have verified the effectiveness of our proposed algorithm, as well as the validity of our dual imaging setting. Xiang Ji 0005, Haiyang Jiang 0002, Yinqiang Zheng |
CVPR | 3 |
| 2024 | Fooling Polarization-Based Vision Using Locally Controllable Polarizing ProjectionabstractPolarization is a fundamental property of light that encodes abundant information regarding surface shape, material, illumination and viewing geometry. The computer vision community has witnessed a blossom of polarization-based vision applications, such as reflection removal, shape-from-polarization (SfP), transparent object segmentation and color constancy, partially due to the emergence of single-chip mono/color polarization sensors that make polarization data acquisition easier than ever. However, is polarization-based vision vulnerable to adversarial attacks? If so, is that possible to realize these adversarial attacks in the physical world, without being perceived by human eyes? In this paper, we warn the community of the vulnerability of polarization-based vision, which can be more serious than RGB-based vision. By adapting a commercial LCD projector, we achieve locally controllable polarizing projection, which is successfully utilized to fool state-of-the-art polarization-based vision algorithms for glass segmentation and SfP. Compared with existing physical attacks on RGB-based vision, which always suffer from the trade-off between attack efficacy and eye conceivability, the adversarial attackers based on polarizing projection are contact-free and visually imperceptible, since naked human eyes can rarely perceive the difference of viciously manipulated polarizing light and ordinary illumination. This poses unprecedented risks on polarization-based vision, for which due attentions should be paid and counter measures be considered. Zhuoxiao Li, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
CVPR | 5 |
| 2024 | Within the Dynamic Context: Inertia-Aware 3D Human Modeling with Pose Sequence
Yifan Zhan, Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ECCV (49) | 7 |
| 2024 | MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
Muyao Niu, Xiaodong Cun, Xintao Wang 0002, Yong Zhang 0034, Ying Shan, Yinqiang Zheng |
ECCV (19) | 6 |
| 2024 | RS-NeRF: Neural Radiance Fields from Rolling Shutter Images
Muyao Niu, Yifan Zhan, Zhuoxiao Li, Xiang Ji 0005, Yinqiang Zheng |
ECCV (46) | 6 |
| 2024 | KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ECCV (45) | 7 |
| 2024 | RPBG: Towards Robust Neural Point-Based Graphics in the Wild
Qingtian Zhu, Zizhuang Wei, Zhongtian Zheng, Yifan Zhan, Zhuyu Yao, Jiawang Zhang, Kejian Wu, Yinqiang Zheng |
ECCV (15) | 8 |
| 2024 | Defending Against Physical Adversarial Patch attacks On Infrared Human DetectionabstractInfrared detection is an emerging technique for safety-critical tasks owing to its remarkable anti-interference capability. However, recent studies have revealed that it is vulnerable to physically-realizable adversarial patches, posing risks in its real-world applications. To address this problem, we are the first to investigate defense strategies against adversarial patch attacks on infrared detection, especially human detection. We propose a straightforward defense strategy, patch-based occlusion-aware detection (POD), which efficiently augments training samples with random patches and subsequently detects them. POD not only robustly detects people but also identifies adversarial patch locations. Surprisingly, while being extremely computationally efficient, POD easily generalizes to state-of-the-art adversarial patch attacks that are unseen during training. Furthermore, POD improves detection precision even in a clean (i.e., no-attack) situation due to the data augmentation effect. Our evaluation demonstrates that POD is robust to adversarial patches of various shapes and sizes. The effectiveness of our baseline approach is shown to be a viable defense mechanism for real-world infrared human detection systems, paving the way for exploring future research directions. Lukas Strack, Futa Waseda, Huy H. Nguyen, Yinqiang Zheng, Isao Echizen |
ICIP | 4 |
| 2024 | FlexIR: Towards Flexible and Manipulable Image RestorationabstractThe domain of image restoration encompasses a wide array of highly effective models (e.g., SwinIR, CODE, DnCNN), each exhibiting distinct advantages in either efficiency or performance. Selecting and deploying these models necessitate careful consideration of resource limitations. While some studies have explored dynamic restoration through the integration of an auxiliary network within a unified framework, these approaches often fall short in practical applications due to the complexities involved in training, retraining, and hyperparameter adjustment, as well as limitations as being totally controlled by auxiliary network and biased by training data. To address these challenges, we introduce FlexIR: a flexible and manipulable framework for image restoration. FlexIR is distinguished by three components: a meticulously designed hierarchical branch network enabling dynamic output, an innovative progressive self-distillation process, and a channel-wise evaluation method to enhance knowledge distillation efficiency. Additionally, we propose two novel inference methodologies to fully leverage FlexIR, catering to diverse user needs and deployment contexts. Through this framework, FlexIR achieves unparalleled performance across all branches, allowing users to navigate the trade-offs between quality, cost, and efficiency during the inference phase. Crucially, FlexIR employs a dynamic mechanism powered by a non-learning metric independent of training data, ensuring that FlexIR is entirely under the direct control of the user. Comprehensive experimental evaluations validate FlexIR's flexibility, manipulability, and cost-effectiveness, showcasing its potential for straightforward adjustments and quick adaptations across a range of scenarios. Zhengwei Yin, Guixu Lin, Mengshun Hu, Hao Zhang 0168, Yinqiang Zheng |
ACM Multimedia | 5 |
| 2024 | Exploring Data Efficiency in Image Restoration: A Gaussian Denoising Case StudyabstractAmidst the prevailing trend of escalating demands for data and computational resources, the efficiency of data utilization emerges as a critical lever for enhancing the performance of deep learning models, especially in the realm of image restoration tasks. This investigation delves into the intricacies of data efficiency in the context of image restoration, with Gaussian image denoising serving as a case study. We postulate a strong correlation between the model's performance and the content information encapsulated in the training images. This hypothesis is rigorously tested through experiments conducted on synthetically blurred datasets. Building on this premise, we delve into the data efficiency within training datasets and introduce an effective and stabilized method for quantifying content information, thereby enabling the ranking of training images based on their influence. Our in-depth analysis sheds light on the impact of various subset selection strategies, informed by this ranking, on model performance. Furthermore, we examine the transferability of these efficient subsets across disparate network architectures. The findings underscore the potential to achieve comparable, if not superior, performance with a fraction of the data-highlighting instances where training IRCNN and Restormer models with only 3.89% and 2.30% of the data resulted in a negligible drop and, in some cases, a slight improvement in PSNR. This investigation offers valuable insights and methodologies to address data efficiency challenges in Gaussian denoising. Similarly, our method yields comparable conclusions in other restoration tasks. We believe this will be beneficial for future research. Zhengwei Yin, Mingze Ma, Guixu Lin, Yinqiang Zheng |
ACM Multimedia | 4 |
| 2023 | HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable DesignabstractAdversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often impractical and attention-grabbing. To address the need for a physically practical and stealthy adversarial attack, we introduce HotCold Block, a novel physical attack for infrared detectors that hide persons utilizing the wearable Warming Paste and Cooling Paste. By attaching these readily available temperature-controlled materials to the body, HotCold Block evades human eyes efficiently. Moreover, unlike existing methods that build adversarial patches with complex texture and structure features, HotCold Block utilizes an SSP-oriented adversarial optimization algorithm that enables attacks with pure color blocks and explores the influence of size, shape, and position on attack performance. Extensive experimental results in both digital and physical environments demonstrate the performance of our proposed HotCold Block. Code is available: https://github.com/weihui1308/HOTCOLDBlock. Hui Wei 0004, Zhixiang Wang 0001, Xuemei Jia, Yinqiang Zheng, Hao Tang 0005, Shin'ichi Satoh 0001, Zheng Wang 0007 |
AAAI | 4 |
| 2023 | High-fidelity Event-Radiance Recovery via Transient Event FrequencyabstractHigh-fidelity radiance recovery plays a crucial role in scene information reconstruction and understanding. Conventional cameras suffer from limited sensitivity in dynamic range, bit depth, and spectral response, etc. In this paper, we propose to use event cameras with bio-inspired silicon sensors, which are sensitive to radiance changes, to recover precise radiance values. We reveal that, under active lighting conditions, the transient frequency of event signals triggering linearly reflects the radiance value. We propose an innovative method to convert the high temporal resolution of event signals into precise radiance values. The precise radiance values yields several capabilities in image analysis. We demonstrate the feasibility of recovering radiance values solely from the transient event frequency (TEF) through multiple experiments. Jin Han 0001, Yuta Asano, Boxin Shi, Yinqiang Zheng, Imari Sato |
CVPR | 4 |
| 2023 | Polarized Color Image DenoisingabstractSingle-chip polarized color photography provides both visual textures and object surface information in one snapshot. However, the use of an additional directional polarizing filter array tends to lower photon count and SNR, when compared to conventional color imaging. As a result, such a bilayer structure usually leads to unpleasant noisy images and undermines performance of polarization analysis, especially in low-light conditions. It is a challenge for traditional image processing pipelines owing to the fact that the physical constraints exerted implicitly in the channels are excessively complicated. In this paper, we propose to tackle this issue through a noise modeling method for realistic data synthesis and a powerful network structure inspired by vision Transformer. A real-world polarized color image dataset of paired raw short-exposed noisy images and long-exposed reference images is captured for experimental evaluation, which has demonstrated the effectiveness of our approaches for data synthesis and polarized color image denoising. The code and data can be found at https://github.com/bandasyou/pcdenoise. Zhuoxiao Li, Haiyang Jiang 0002, Mingdeng Cao, Yinqiang Zheng |
CVPR | 4 |
| 2023 | Visibility Constrained Wide-Band Illumination Spectrum Design for Seeing-in-the-DarkabstractSeeing-in-the-dark is one of the most important and challenging computer vision tasks due to its wide applications and extreme complexities of in-the- wild scenarios. Existing arts can be mainly divided into two threads: 1) RGB-dependent methods restore information using degraded RGB inputs only (e.g., low-light enhancement), 2) RGB-independent methods translate images captured under auxiliary near-infrared (NIR) illuminants into RGB domain (e.g., NIR2RGB translation). The latter is very attractive since it works in complete darkness and the illuminants are visually friendly to naked eyes, but tends to be unstable due to its intrinsic ambiguities. In this paper, we try to robustify NIR2RGB translation by designing the optimal spectrum of auxiliary illumination in the wide-band VIS-NIR range, while keeping visual friendliness. Our core idea is to quantify the visibility constraint implied by the human vision system and incorporate it into the design pipeline. By modeling the formation process of images in the VIS-NIR range, the optimal multiplexing of a wide range of LEDs is automatically designed in a fully differentiable manner, within the feasible region defined by the visibility constraint. We also collect a substantially expanded VIS-NIR hyperspectral image dataset for experiments by using a customized 50-band filter wheel. Experimental results show that the task can be significantly improved by using the optimized wide-band illumination than using NIR only. Codes Available: https://github.com/MyNiuuu/VCSD. Muyao Niu, Zhuoxiao Li, Zhihang Zhong, Yinqiang Zheng |
CVPR | 4 |
| 2023 | Blur Interpolation Transformer for Real-World Motion from BlurabstractThis paper studies the challenging problem of recovering motion from blur, also known as joint deblurring and interpolation or blur temporal super-resolution. The challenges are twofold: 1) the current methods still leave considerable room for improvement in terms of visual quality even on the synthetic dataset, and 2) poor generalization to real-world data. To this end, we propose a blur interpolation transformer (BiT) to effectively unravel the underlying temporal correlation encoded in blur. Based on multi-scale residual Swin transformer blocks, we introduce dual-end temporal supervision and temporally symmetric ensembling strategies to generate effective features for time-varying motion rendering. In addition, we design a hybrid camera system to collect the first real-world dataset of one-to-many blur-sharp video pairs. Experimental results show that BiT has a significant gain over the state-of-the-art methods on the public dataset Adobe240. Besides, the proposed real-world dataset effectively helps the model generalize well to real blurry scenarios. Code and data are available at https://github.com/zzh-tech/Bi'T. Zhihang Zhong, Mingdeng Cao, Xiang Ji 0005, Yinqiang Zheng, Imari Sato |
CVPR | 4 |
| 2023 | MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingabstractDespite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex nonrigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing. Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Ying Shan, Xiaohu Qie, Yinqiang Zheng |
ICCV | 6 |
| 2023 | Single Image Deblurring with Row-dependent Blur MagnitudeabstractImage degradation often occurs during fast camera or object movements, regardless of the exposure modes: global shutter (GS) or rolling shutter (RS). Since these two exposure modes give rise to intrinsically different degradations, two restoration threads have been explored separately, i.e. motion deblurring of GS images and distortion correction of RS images, both of which are challenging restoration tasks, especially in the presence of a single input image. In this paper, we explore a novel in-between exposure mode, called global reset release (GRR) shutter, which produces GS-like blur but with row-dependent blur magnitude. We take advantage of this unique characteristic of GRR to explore the latent frames within a single image and restore a clear counterpart by only relying on these latent contexts. Specifically, we propose a residual spatially-compensated and spectrally-enhanced Transformer (RSS-T) block for row-dependent deblurring of a single GRR image. Its hierarchical positional encoding compensates global positional context of windows and enables order-awareness of the local pixel’s position, along with a novel feed-forward network that simultaneously uses spatial and spectral information for gaining mixed global context. Extensive experimental results demonstrate that our method outperforms the state-of-the-art GS deblurring and RS correction methods on single GRR input. Xiang Ji 0005, Zhixiang Wang 0001, Shin'ichi Satoh 0001, Yinqiang Zheng |
ICCV | 4 |
| 2023 | Rethinking Video Frame Interpolation from Shutter Mode Induced DegradationabstractImage restoration from various motion-related degradations, like blurry effects recorded by a global shutter (GS) and jello effects caused by a rolling shutter (RS), has been extensively studied. It has been recently recognized that such degradations encode temporal information, which can be exploited for video frame interpolation (VFI), a more challenging task than pure restoration. However, these VFI researches are mainly grounded on experiments with synthetic data, rather than real data. More fundamentally, under the same imaging condition, it remains unknown which degradation will be more effective toward VFI. In this paper, we present the first real-world dataset for learning and benchmarking degraded video frame interpolation, named RD-VFI, and further explore the performance differences of three types of degradations, including GS blur, RS distortion, and an in-between effect caused by the rolling shutter with global reset (RSGR), thanks to our novel quad-axis imaging system. Moreover, we propose a unified Progressive Mutual Boosting Network (PMBNet) model to interpolate middle frames at arbitrary time for all shutter modes. Its disentanglement strategy and dual-stream correction enable us to adaptively deal with different degradations for VFI. Experimental results demonstrate that our PMBNet is superior to the respective state-of-the-art methods on all shutter modes. Xiang Ji 0005, Zhixiang Wang 0001, Zhihang Zhong, Yinqiang Zheng |
ICCV | 4 |
| 2023 | NIR-assisted Video Enhancement via Unpaired 24-hour DataabstractLow-light video enhancement in the visible (VIS) range is important yet technically challenging, and it is likely to become more tractable by introducing near-infrared (NIR) information for assistance, which in turn arouses a new challenge on how to obtain appropriate multispectral data for model training. In this paper, we defend the feasibility and superiority of NIR-assisted low-light video enhancement results by using unpaired 24-hour data for the first time, which significantly eases data collection and improves generalization performance on in-the-wild data. By accounting for different physical characteristics between unpaired daytime and nighttime videos, we first propose to turn daytime NIR & VIS into "nighttime mode". Specifically, we design a heuristic yet physics-inspired relighting algorithm to produce realistic pseudo nighttime NIR, and use a resampling strategy followed by a noiseGAN for nighttime VIS conversion. We further devise a temporal-aware network for video enhancement that extracts and fuses bi-directional temporal streams and is trained using real daytime videos and pseudo nighttime videos. We capture multi-spectral data using a co-axial camera and contribute Fulltime Multi-Spectral Video Dataset (FMSVD), the first dataset including aligned 24-hour NIR & VIS videos. Compared to alternative methods, we achieve significantly improved video quality as well as generalization ability on in-the-wild data in terms of both evaluation metrics and visual judgment. Codes and Data Available: https://github.com/MyNiuuu/NVEU. Muyao Niu, Zhihang Zhong, Yinqiang Zheng |
ICCV | 3 |
| 2023 | NeRFrac: Neural Radiance Fields through Refractive SurfaceabstractNeural Radiance Fields (NeRF) is a popular neural representation for novel view synthesis. By querying spatial points and view directions, a multilayer perceptron (MLP) can be trained to output the volume density and radiance along a ray, which lets us render novel views of the scene. The original NeRF and its recent variants, however, are limited to opaque scenes dominated with diffuse reflection surfaces and cannot handle complex refractive surfaces well. We introduce NeRFrac to realize neural novel view synthesis of scenes captured through refractive surfaces, typically water surfaces. For each queried ray, an MLP-based Refractive Field is trained to estimate the distance from the ray origin to the refractive surface. A refracted ray at each intersection point is then computed by Snell’s Law, given the input ray and the approximated local normal. Points of the scene are sampled along the refracted ray and are sent to a Radiance Field for further radiance estimation. We show that from a sparse set of images, our model achieves accurate novel view synthesis of the scene underneath the refractive surface and simultaneously reconstructs the refractive surface. We evaluate the effectiveness of our method with synthetic and real scenes seen through water surfaces. Experimental results demonstrate the accuracy of NeRFrac for modeling scenes seen through wavy refractive surfaces. Github page: https://github.com/Yifever20002/NeRFrac. Yifan Zhan, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ICCV | 4 |
| 2023 | Event-guided Frame Interpolation and Dynamic Range Expansion of Single Rolling Shutter ImageabstractIn the presence of abrupt motion, the pushbroom scanning mechanism of a rolling shutter (RS) camera tends to bring undesirable distortion, which is recently shown to be beneficial for high-speed frame interpolation. Although promising results have been reported by using multiple consecutive RS frames, to interpolate intermediate distortion-free frames from a single RS image is still an open question, due to the existence of multiple motions that can account for the recorded distortion. Another limitation of RS cameras in complex dynamic scenarios lies in the dynamic range, since traditional ways of multiple exposure for high dynamic range (HDR) imaging will fail due to alignment issues. To deal with these two challenges simultaneously, we propose to use an event camera for assistance, which has much faster temporal response and wider dynamic range. Since there does not exist learning data for this brand new imaging setup, we first build a quad-axis imaging system to capture a realistic dataset called REG-HDR, with pairs of fully aligned RS image and its associated events, as well as their corresponding high-speed HDR GS images. We also propose a flow-based network for frame interpolation, compounded with an attention-based fusion network for dynamic range expansion. Experimental results have verified the effectiveness of our proposed algorithm and the superiority of using realistic data for this challenging dural-purpose enhancement task. Guixu Lin, Jin Han 0001, Mingdeng Cao, Zhihang Zhong, Yinqiang Zheng |
ACM Multimedia | 5 |
| 2023 | Beyond Domain Gap: Exploiting Subjectivity in Sketch-Based Person RetrievalabstractPerson re-identification (re-ID) requires densely distributed cameras. In practice, the person of interest may not be captured by cameras and therefore need to be retrieved using subjective information (e.g., sketches from witnesses). Previous research defines this case using the sketch as sketch re-identification (Sketch re-ID) and focuses on eliminating the domain gap. Actually, subjectivity is another significant challenge. We model and investigate it by posing a new dataset with multi-witness descriptions. It features two aspects. 1) Large-scale. It contains over 4,763 sketches and 32,668 photos, making it the largest Sketch re-ID dataset. 2) Multi-perspective and multi-style. Our dataset offers multiple sketches for each identity. Witnesses' subjective cognition provides multiple perspectives on the same individual, while different artists' drawing styles provide variation in sketch styles. We further have two novel designs to alleviate the challenge of subjectivity. 1) Fusing subjectivity. We propose a non-local (NL) fusion module that gathers sketches from different witnesses for the same identity. 2) Introducing objectivity. An AttrAlign module utilizes attributes as an implicit mask to align cross-domain features. To push forward the advance of Sketch re-ID, we set three benchmarks (large-scale, multi-style, cross-style). Extensive experiments demonstrate our leading performance in these benchmarks. Dataset and Codes are publicly available at: https://github.com/Lin-Kayla/subjectivity-sketch-reid Kejun Lin, Zhixiang Wang 0001, Zheng Wang 0007, Yinqiang Zheng, Shin'ichi Satoh 0001 |
ACM Multimedia | 4 |
| 2023 | Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera SystemsabstractMany surveillance cameras switch between daytime and nighttime modes based on illuminance levels. During the day, the camera records ordinary RGB images through an enabled IR-cut filter. At night, the filter is disabled to capture near-infrared (NIR) light emitted from NIR LEDs typically mounted around the lens. While the vulnerabilities of RGB-based AI algorithms have been widely reported, those of NIR-based AI have rarely been investigated. In this paper, we identify fundamental vulnerabilities in NIR-based image understanding caused by color and texture loss due to the intrinsic characteristics of clothes' reflectance and cameras' spectral sensitivity in the NIR range. We further show that the nearly co-located configuration of illuminants and cameras in existing surveillance systems facilitates concealing and fully passive attacks in the physical world. Specifically, we demonstrate how retro-reflective and insulation plastic tapes can manipulate the intensity distribution of NIR images. We showcase an attack on the YOLO-based human detector using binary patterns designed in the digital space (via black-box query and searching) and then physically realized using tapes pasted onto clothes. Our attack highlights significant reliability concerns about nighttime surveillance systems, which are intended to enhance security. Codes Available: https://github.com/MyNiuuu/AdvNIR. Muyao Niu, Zhuoxiao Li, Yifan Zhan, Huy H. Nguyen, Isao Echizen, Yinqiang Zheng |
ACM Multimedia | 6 |
| 2023 | Assessor360: Multi-sequence Network for Blind Omnidirectional Image Quality AssessmentabstractBlind Omnidirectional Image Quality Assessment (BOIQA) aims to objectively assess the human perceptual quality of omnidirectional images (ODIs) without relying on pristine-quality image information. It is becoming more significant with the increasing advancement of virtual reality (VR) technology. However, the quality assessment of ODIs is severely hampered by the fact that the existing BOIQA pipeline lacks the modeling of the observer's browsing process. To tackle this issue, we propose a novel multi-sequence network for BOIQA called Assessor360, which is derived from the realistic multi-assessor ODI quality assessment procedure. Specifically, we propose a generalized Recursive Probability Sampling (RPS) method for the BOIQA task, combining content and details information to generate multiple pseudo viewport sequences from a given starting point. Additionally, we design a Multi-scale Feature Aggregation (MFA) module with a Distortion-aware Block (DAB) to fuse distorted and semantic features of each viewport. We also devise Temporal Modeling Module (TMM) to learn the viewport transition in the temporal domain. Extensive experimental results demonstrate that Assessor360 outperforms state-of-the-art methods on multiple OIQA datasets. The code and models are available at https://github.com/TianheWu/Assessor360. Tianhe Wu, Shuwei Shi, Haoming Cai, Mingdeng Cao, Jing Xiao 0006, Yinqiang Zheng, Yujiu Yang 0001 |
NeurIPS | 6 |
| 2023 | Real-World Video Deblurring: A Benchmark Dataset and an Efficient Recurrent Neural Network
Zhihang Zhong, Yinqiang Zheng, Imari Sato |
Int. J. Comput. Vis. | 3 |
| 2023 | Reliability-Aware Restoration Framework for 4D Spectral Photoacoustic DataabstractSpectral photoacoustic imaging (PAI) is a new technology that is able to provide 3D geometric structure associated with 1D wavelength-dependent absorption information of the interior of a target in a non-invasive manner. It has potentially broad applications in clinical and medical diagnosis. Unfortunately, the usability of spectral PAI is severely affected by a time-consuming data scanning process and complex noise. Therefore in this study, we propose a reliability-aware restoration framework to recover clean 4D data from incomplete and noisy observations. To the best of our knowledge, this is the first attempt for the 4D spectral PA data restoration problem that solves data completion and denoising simultaneously. We first present a sequence of analyses, including modeling of data reliability in the depth and spectral domains, developing an adaptive correlation graph, and analyzing local patch orientation. On the basis of these analyses, we explore global sparsity and local self-similarity for restoration. We demonstrated the effectiveness of our proposed approach through experiments on real data captured from patients, where our approach outperformed the state-of-the-art methods in both objective evaluation and subjective assessment. Weihang Liao, Art Subpa-Asa, Yuta Asano, Yinqiang Zheng, Hiroki Kajita, Nobuaki Imanishi, Takayuki Yagi, Sadakazu Aiso, Kazuo Kishi, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Learning Adaptive Warping for RealWorld Rolling Shutter CorrectionabstractThis paper proposes the first real-world rolling shutter (RS) correction dataset, BS-RSC, and a corresponding model to correct the RS frames in a distorted video. Mobile devices in the consumer market with CMOS-based sensors for video capture often result in rolling shutter effects when relative movements occur during the video acquisition process, calling for RS effect removal techniques. However, current state-of-the-art RS correction methods often fail to remove RS effects in real scenarios since the motions are various and hard to model. To address this issue, we propose a real-world RS correction dataset BS-RSC. Real distorted videos with corresponding ground truth are recorded simultaneously via a well-designed beam-splitter-based acquisition system. BS-RSC contains various motions of both camera and objects in dynamic scenes. Further, an RS correction model with adaptive warping is proposed. Our model can warp the learned RS features into global shutter counterparts adaptively with predicted multiple displacement fields. These warped features are aggregated and then reconstructed into high-quality global shutter frames in a coarse-to-fine strategy. Experimental results demonstrate the effectiveness of the proposed method, and our dataset can improve the model's ability to remove the RS effects in the real world. The project is available at https://github.com/ljzycmd/BSRSC. Mingdeng Cao, Zhihang Zhong, Jiahao Wang 0005, Yinqiang Zheng, Yujiu Yang 0001 |
CVPR | 4 |
| 2022 | Optimal LED Spectral Multiplexing for NIR2RGB TranslationabstractThe industry practice for night video surveillance is to use auxiliary near-infrared (NIR) LEDs, usually centered at 850nm or 940nm, for scene illumination. NIR LEDs are used to save power consumption while hiding the surveillance coverage area from naked human eyes. The captured images are almost monochromatic, and visual color and texture tend to disappear, which hinders human and machine perception. A few existing studies have tried to convert such NIR images to RGB images through deep learning, which can not provide satisfying results, nor generalize well beyond the training dataset. In this paper, we aim to break the fundamental restrictions on reliable NIR-to-RGB (NIR2RGB) translation by examining the imaging mechanism of single-chip silicon-based RGB cameras under NIR illuminations, and propose to retrieve the optimal LED multiplexing via deep learning. Experimental results show that this translation task can be significantly improved by properly multiplexing NIR LEDs close to the visible spectral range than using 850nm and 940nm LEDs. Yuze Chen, Junchi Yan, Yinqiang Zheng |
CVPR | 4 |
| 2022 | Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene UnderstandingabstractAlthough considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hinder the progress of this field. Considering the success in semantic scene understanding under clear weather, we think it is reasonable to transfer knowledge learned from clear images to the foggy domain. As such, the problem becomes to bridge the domain gap between clear images and foggy images. Unlike previous methods that mainly focus on closing the domain gap caused by fog - defogging the foggy images or fogging the clear images, we propose to alleviate the domain gap by considering fog influence and style variation simultaneously. The motivation is based on our finding that the style-related gap and the fog-related gap can be divided and closed respectively, by adding an intermediate domain. Thus, we propose a new pipeline to cumulatively adapt style, fog and the dual-factor (style and fog). Specifically, we devise a unified framework to disentangle the style factor and the fog factor separately, and then the dual-factor from images in different domains. Furthermore, we collaborate the disentanglement of three factors with a novel cumulative loss to thoroughly disentangle these three factors. Our method achieves the state-of-the-art performance on three benchmarks and shows generalization ability in rainy and snowy scenes. Xianzheng Ma, Zhixiang Wang 0001, Yacheng Zhan, Yinqiang Zheng, Zheng Wang 0007, Dengxin Dai, Chia-Wen Lin |
CVPR | 4 |
| 2022 | Neural Global Shutter: Learn to Restore Video from a Rolling Shutter Camera with Global Reset FeatureabstractMost computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However, most of the existing work relies heavily on the prior assumptions of scenes or motions. Besides, the motion estimation steps are either oversimplified or computationally inefficient due to the heavy flow warping, limiting their applicability. In this paper, we investigate using rolling shutter with a global reset feature (RSGR) to restore clean global shutter (GS) videos. This feature enables us to turn the rectification problem into a deblur-like one, getting rid of inaccurate and costly explicit motion estimation. First, we build an optic system that captures paired RSGR/GS videos. Second, we develop a novel algorithm incorporating spatial and temporal designs to correct the spatial-varying RSGR distortion. Third, we demonstrate that existing image-to-image translation algorithms can recover clean GS videos from distorted RSGR inputs, yet our algorithm achieves the best performance with the specific designs. Our rendered results are not only visually appealing but also beneficial to downstream tasks. Compared to the state-of-the-art RS solution, our RSGR solution is superior in both effectiveness and efficiency. Considering it is easy to realize without changing the hardware, we believe our RSGR solution can potentially replace the RS solution in taking distortion-free videos with low noise and low budget. Zhixiang Wang 0001, Xiang Ji 0005, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Yinqiang Zheng |
CVPR | 6 |
| 2022 | Efficient Video Deblurring Guided by Motion Magnitude
Yusheng Wang 0001, Yunfan Lu, Lin Wang 0025, Zhihang Zhong, Yinqiang Zheng, Atsushi Yamashita |
ECCV (19) | 6 |
| 2022 | Bringing Rolling Shutter Images Alive with Dual Reversed Distortion
Zhihang Zhong, Mingdeng Cao, Xiao Sun 0001, Zhirong Wu, Zhongyi Zhou, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (7) | 6 |
| 2022 | Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance
Zhihang Zhong, Xiao Sun 0001, Zhirong Wu, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (19) | 4 |
| 2022 | Graph-Based Compression of Incomplete 3D Photoacoustic Data
Weihang Liao, Yinqiang Zheng, Hiroki Kajita, Kazuo Kishi, Imari Sato |
MICCAI (6) | 2 |
| 2022 | Event-guided Video Clip Generation from Blurry ImagesabstractDynamic and active pixel vision sensors (DAVIS) can simultaneously produce streams of asynchronous events captured by the dynamic vision sensor (DVS) and intensity frames from the active pixel sensor (APS). Event sequences show high temporal resolution and high dynamic range, while intensity images easily suffer from motion blur due to the low frame rate of APS. In this paper, we present an end-to-end convolutional neural network based method under the local and global constraints of events to restore clear, sharp intensity frames through collaborative learning from a blurry image and its associated event streams. Specifically, we first learn a function of the relationship between the sharp intensity frame and the corresponding blurry image with its event data. Then we propose a generation module to realize it with a supervision module to constrain the restoration in the motion process. We also capture the first realistic dataset with paired blurry frame/events and sharp frames by synchronizing a DAVIS camera and a high-speed camera. Experimental results show that our method can reconstruct high-quality sharp video clips, and outperform the state-of-the-art on both simulated and real-world data. Tsuyoshi Takatani, Zhongyuan Wang 0001, Ying Fu 0001, Yinqiang Zheng |
ACM Multimedia | 5 |
| 2022 | Joint Camera Spectral Response Selection and Hyperspectral Image RecoveryabstractHyperspectral image (HSI) recovery from a single RGB image has attracted much attention, whose performance has recently been shown to be sensitive to the camera spectral response (CSR). In this paper, we present an efficient convolutional neural network (CNN) based method, which can jointly select the optimal CSR from a candidate dataset and learn a mapping to recover HSI from a single RGB image captured with this algorithmically selected camera under multi-chip or single-chip setups. Given a specific CSR, we first present a HSI recovery network, which accounts for the underlying characteristics of the HSI, including spectral nonlinear mapping and spatial similarity. Later, we append a CSR selection layer onto the recovery network, and the optimal CSR under both multi-chip and single-chip setups can thus be automatically determined from the network weights under the nonnegative sparse constraint. Experimental results on three hyperspectral datasets and two camera spectral response datasets demonstrate that our HSI recovery network outperforms state-of-the-art methods in terms of both quantitative metrics and perceptive quality, and the selection layer always returns a CSR consistent to the best one determined by exhaustive search. Finally, we show that our method can also perform well in the real capture system, and collect a hyperspectral flower dataset to evaluate the effect from HSI recovery on classification problem. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Physics-Based Noise Modeling for Extreme Low-Light PhotographyabstractEnhancing the visibility in extreme low-light environments is a challenging task. Under nearly lightless condition, existing image denoising methods could easily break down due to significantly low SNR. In this paper, we systematically study the noise statistics in the imaging pipeline of CMOS photosensors, and formulate a comprehensive noise model that can accurately characterize the real noise structures. Our novel model considers the noise sources caused by digital camera electronics which are largely overlooked by existing methods yet have significant influence on raw measurement in the dark. It provides a way to decouple the intricate noise structure into different statistical distributions with physical interpretations. Moreover, our noise model can be used to synthesize realistic training data for learning-based low-light denoising algorithms. In this regard, although promising results have been shown recently with deep convolutional neural networks, the success heavily depends on abundant noisy-clean image pairs for training, which are tremendously difficult to obtain in practice. Generalizing their trained models to images from new devices is also problematic. Extensive experiments on multiple low-light denoising datasets - including a newly collected one in this work covering various devices - show that a deep neural network trained with our proposed noise formation model can reach surprisingly-high accuracy. The results are on par with or sometimes even outperform training with paired real data, opening a new door to real-world extreme low-light photography. Kaixuan Wei, Ying Fu 0001, Yinqiang Zheng, Jiaolong Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Optical Flow in the DarkabstractOptical flow estimation in low-light conditions is a challenging task for existing methods and current optical flow datasets lack low-light samples. Even if the dark images are enhanced before estimation, which could achieve great visual perception, it still leads to suboptimal optical flow results because information like motion consistency may be broken during the enhancement. We propose to apply a novel training policy to learn optical flow directly from new synthetic and real low-light images. Specifically, first, we design a method to collect a new optical flow dataset in multiple exposures with shared optical flow pseudo labels. Then we apply a two-step process to create a synthetic low-light optical flow dataset, based on an existing bright one, by simulating low-light raw features from the multi-exposure raw images we collected. To extend the data diversity, we also include published low-light raw videos without optical flow labels. In our training pipeline, with the three datasets, we create two teacher-student pairs to progressively obtain optical flow labels for all data. Finally, we apply a mix-up training policy with our diversified datasets to produce low-light-robust optical flow models for release. The experiments show that our method can relatively maintain the optical flow accuracy as the image exposure descends and the generalization ability of our method is tested with different cameras in multiple practical scenes. Mingfang Zhang 0002, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Hyperspectral Image Reconstruction Using Multi-scale Fusion LearningabstractHyperspectral imaging is a promising imaging modality that simultaneously captures several images for the same scene on narrow spectral bands, and it has made considerable progress in different fields, such as agriculture, astronomy, and surveillance. However, the existing hyperspectral (HS) cameras sacrifice the spatial resolution for providing the detail spectral distribution of the imaged scene, which leads to low-resolution (LR) HS images compared with the common red-green-blue (RGB) images. Generating a high-resolution HS (HR-HS) image via fusing an observed LR-HS image with the corresponding HR-RGB image has been actively studied. Existing methods for this fusing task generally investigate hand-crafted priors to model the inherent structure of the latent HR-HS image, and they employ optimization approaches for solving it. However, proper priors for different scenes can possibly be diverse, and to figure it out for a specific scene is difficult. This study investigates a deep convolutional neural network (DCNN)-based method for automatic prior learning, and it proposes a novel fusion DCNN model with multi-scale spatial and spectral learning for effectively merging an HR-RGB and LR-HS images. Specifically, we construct an U-shape network architecture for gradually reducing the feature sizes of the HR-RGB image (Encoder-side) and increasing the feature sizes of the LR-HS image (Decoder-side), and we fuse the HR spatial structure and the detail spectral attribute in multiple scales for tackling the large resolution difference in spatial domain of the observed HR-RGB and LR-HS images. Then, we employ multi-level cost functions for the proposed multi-scale learning network to alleviate the gradient vanish problem in long-propagation procedure. In addition, for further improving the reconstruction performance of the HR-HS image, we refine the predicted HR-HS image using an alternating back-projection method for minimizing the reconstruction errors of the observed LR-HS and HR-RGB images. Experiments on three benchmark HS image datasets demonstrate the superiority of the proposed method in both quantitative values and visual qualities. Xianhua Han, Yinqiang Zheng, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Multi-View 3D Reconstruction of a Texture-Less Smooth Surface of Unknown Generic ReflectanceabstractRecovering the 3D geometry of a purely texture-less object with generally unknown surface reflectance (e.g. non-Lambertian) is regarded as a challenging task in multi-view reconstruction. The major obstacle revolves around establishing cross-view correspondences where photometric constancy is violated. This paper proposes a simple and practical solution to overcome this challenge based on a co-located camera-light scanner device. Unlike existing solutions, we do not explicitly solve for correspondence. Instead, we argue the problem is generally well-posed by multi-view geometrical and photometric constraints, and can be solved from a small number of input views. We formulate the reconstruction task as a joint energy minimization over the surface geometry and reflectance. Despite this energy is highly non-convex, we develop an optimization algorithm that robustly recovers globally optimal shape and reflectance even from a random initialization. Extensive experiments on both simulated and real data have validated our method, and possible future extensions are discussed. Ziang Cheng, Hongdong Li, Yuta Asano, Yinqiang Zheng, Imari Sato |
CVPR | 4 |
| 2021 | 4D Hyperspectral Photoacoustic Data Restoration With Reliability AnalysisabstractHyperspectral photoacoustic (HSPA) spectroscopy is an emerging bi-modal imaging technology that is able to show the wavelength-dependent absorption distribution of the interior of a 3D volume. However, HSPA devices have to scan an object exhaustively in the spatial and spectral domains; and the acquired data tend to suffer from complex noise. This time-consuming scanning process and noise severely affects the usability of HSPA. It is therefore critical to examine the feasibility of 4D HSPA data restoration from an in-complete and noisy observation. In this work, we present a data reliability analysis for the depth and spectral domain. On the basis of this analysis, we explore the inherent data correlations and develop a restoration algorithm to recover 4D HSPA cubes. Experiments on real data verify that the proposed method achieves satisfactory restoration results. Weihang Liao, Art Subpa-Asa, Yinqiang Zheng, Imari Sato |
CVPR | 3 |
| 2021 | Tuning IR-Cut Filter for Illumination-Aware Spectral Reconstruction From RGBabstractTo reconstruct spectral signals from multi-channel observations, in particular trichromatic RGBs, has recently emerged as a promising alternative to traditional scanning-based spectral imager. It has been proven that the reconstruction accuracy relies heavily on the spectral response of the RGB camera in use. To improve accuracy, data-driven algorithms have been proposed to retrieve the best response curves of existing RGB cameras, or even to design brand new three-channel response curves. Instead, this paper explores the filter-array based color imaging mechanism of existing RGB cameras, and proposes to design the IR-cut filter properly for improved spectral recovery, which stands out as an in-between solution with better trade-off between reconstruction accuracy and implementation complexity. We further propose a deep learning based spectral reconstruction method, which allows to recover the illumination spectrum as well. Experiment results with both synthetic and real images under daylight illumination have shown the benefits of our IR-cut filter tuning method and our illumination-aware spectral reconstruction method. Junchi Yan, Yinqiang Zheng |
CVPR | 4 |
| 2021 | Event-Based Bispectral Photometry Using Temporally Modulated IlluminationabstractAnalysis of bispectral difference plays a critical role in various applications that involve rays propagating in a light absorbing medium. In general, the bispectral difference is obtained by subtracting signals at two individual wave-lengths captured by ordinary digital cameras, which tends to inherit the drawbacks of conventional cameras in dynamic range, response speed and quantization precision. In this paper, we propose a novel method to obtain a bispectral difference image using an event camera with temporally modulated illumination. Our method is rooted in a key observation on the analogy between the bispectral photometry principle of the participating medium and the event generating mechanism in an event camera. By carefully modulating the bispectral illumination, our method allows to read out the bispectral difference directly from triggered events. Experiments using a prototype imaging system have verified the feasibility of this novel usage of event cameras in photometry based vision tasks, such as 3D shape reconstruction in water. Tsuyoshi Takatani, Yuzuha Ito, Ayaka Ebisu, Yinqiang Zheng, Takahito Aoto 0002 |
CVPR | 4 |
| 2021 | Towards Rolling Shutter Correction and Deblurring in Dynamic ScenesabstractJoint rolling shutter correction and deblurring (RSCD) techniques are critical for the prevalent CMOS cameras. However, current approaches are still based on conventional energy optimization and are developed for static scenes. To enable learning-based approaches to address real-world RSCD problem, we contribute the first dataset, BS-RSCD, which includes both ego-motion and object-motion in dynamic scenes. Real distorted and blurry videos with corresponding ground truth are recorded simultaneously via a beam-splitter-based acquisition system.Since direct application of existing individual rolling shutter correction (RSC) or global shutter deblurring (GSD) methods on RSCD leads to undesirable results due to inherent flaws in the network architecture, we further present the first learning-based model (JCD) for RSCD. The key idea is that we adopt bi-directional warping streams for displacement compensation, while also preserving the non-warped deblurring stream for details restoration. The experimental results demonstrate that JCD achieves state-of-the-art performance on the realistic RSCD dataset (BS-RSCD) and the synthetic RSC dataset (Fastec-RS). The dataset and code are available at https://github.com/zzh-tech/RSCD. Zhihang Zhong, Yinqiang Zheng, Imari Sato |
CVPR | 2 |
| 2021 | Learning To Reconstruct High Speed and High Dynamic Range Videos From EventsabstractEvent cameras are novel sensors that capture the dynamics of a scene asynchronously. Such cameras record event streams with much shorter response latency than images captured by conventional cameras, and are also highly sensitive to intensity change, which is brought by the triggering mechanism of events. On the basis of these two features, previous works attempt to reconstruct high speed and high dynamic range (HDR) videos from events. However, these works either suffer from unrealistic artifacts, or cannot provide sufficiently high frame rate. In this paper, we present a convolutional recurrent neural network which takes a sequence of neighboring events to reconstruct high speed HDR videos, and temporal consistency is well considered to facilitate the training process. In addition, we setup a prototype optical system to collect a real-world dataset with paired high speed HDR videos and event streams, which will be made publicly accessible for future researches in this field. Experimental results on both simulated and real scenes verify that our method can generate high speed HDR videos with high quality, and outperform the state-of-the-art reconstruction methods. Yunhao Zou, Yinqiang Zheng, Tsuyoshi Takatani, Ying Fu 0001 |
CVPR | 2 |
| 2021 | Depth Sensing by Near-Infrared Light Absorption in WaterabstractThis paper introduces a novel depth recovery method based on light absorption in water. Water absorbs light at almost all wavelengths whose absorption coefficient is related to the wavelength. Based on the Beer-Lambert model, we introduce a bispectral depth recovery method that leverages the light absorption difference between two near-infrared wavelengths captured with a distant point source and orthographic cameras. Through extensive analysis, we show that accurate depth can be recovered irrespective of the surface texture and reflectance, and introduce algorithms to correct for nonidealities of a practical implementation including tilted light source and camera placement, nonideal bandpass filters and the perspective effect of the camera with a diverging point light source. We construct a coaxial bispectral depth imaging system using low-cost off-the-shelf hardware and demonstrate its use for recovering the shapes of complex and dynamic objects in water. We also present a trispectral variant to further improve robustness to extremely challenging surface reflectance. Experimental results validate the theory and practical implementation of this novel depth recovery paradigm, which we refer to as shape from water. Yuta Asano, Yinqiang Zheng, Ko Nishino, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | A Microfacet-Based Model for Photometric Stereo with General Isotropic ReflectanceabstractThis paper presents a precise, stable, and invertible reflectance model for photometric stereo. This microfacet-based model is applicable to all types of isotropic surface reflectance, covering cases from diffusion to specular reflections. We introduce a single variable to physically quantify the surface smoothness, and by monotonically sliding this variable between 0 and 1, our model enables a versatile representation that can smoothly transform between an ellipsoid of revolution and the equation for Lambertian reflectance. In the inverse domain, this model offers a compact and physically interpretable formulation, for which we introduce a fast and lightweight solver that allows accurate estimations for both surface smoothness and surface shape. Finally, extensive experiments on the appearances of synthesized and real objects evidence that this model is state-of-the-art in our off-the-shelf solution. Lixiong Chen, Yinqiang Zheng, Boxin Shi, Art Subpa-Asa, Imari Sato |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | A Sparse Representation Based Joint Demosaicing Method for Single-Chip Polarized Color SensorabstractThe emergence of the single-chip polarized color sensor now allows for simultaneously capturing chromatic and polarimetric information of the scene on a monochromatic image plane. However, unlike the usual camera with an embedded demosaicing method, the latest polarized color camera is not delivered with an in-built demosaicing tool. For demosaicing, the users have to down-sample the captured images or to use traditional interpolation techniques. Neither of them can perform well since the polarization and color are interdependent. Therefore, joint chromatic and polarimetric demosaicing is the key to obtaining high-quality polarized color images. In this paper, we propose a joint chromatic and polarimetric demosaicing model to address this challenging problem. Instead of mechanically demosaicing for the multi-channel polarized color image, we further present a sparse representation-based optimization strategy that utilizes chromatic information and polarimetric information to jointly optimize the model. To avoid the interaction between color and polarization during demosaicing, we separately construct the corresponding dictionaries. We also build an optical data acquisition system to collect a dataset, which contains various sources of polarization, such as illumination, reflectance and birefringence. Results of both qualitative and quantitative experiments have shown that our method is capable of faithfully recovering full RGB information of four polarization angles for each pixel from a single mosaic input image. Moreover, the proposed method can perform well not only on the synthetic data but the real captured data. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Image Process. | 2 |
| 2021 | Polarization Guided Specular Reflection SeparationabstractSince specular reflection often exists in the real captured images and causes deviation between the recorded color and intrinsic color, specular reflection separation can bring advantages to multiple applications that require consistent object surface appearance. However, due to the color of an object is significantly influenced by the color of the illumination, the existing researches still suffer from the near-duplicate challenge, that is, the separation becomes unstable when the illumination color is close to the surface color. In this paper, we derive a polarization guided model to incorporate the polarization information into a designed iteration optimization separation strategy to separate the specular reflection. Based on the analysis of polarization, we propose a polarization guided model to generate a polarization chromaticity image, which is able to reveal the geometrical profile of the input image in complex scenarios, e.g., diversity of illumination. The polarization chromaticity image can accurately cluster the pixels with similar diffuse color. We further use the specular separation of all these clusters as an implicit prior to ensure that the diffuse component will not be mistakenly separated as the specular component. With the polarization guided model, we reformulate the specular reflection separation into a unified optimization function which can be solved by the ADMM strategy. The specular reflection will be detected and separated jointly by RGB and polarimetric information. Both qualitative and quantitative experimental results have shown that our method can faithfully separate the specular reflection, especially in some challenging scenarios. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Image Process. | 2 |
| 2020 | Underwater Scene Recovery Using Wavelength-Dependent Refraction of LightabstractThis paper proposes a method of underwater depth estimation from an orthographic multispectral image. In accordance with Snell's law, incoming light is refracted when it enters the water surface, and its directions are determined by the refractive index and the normals of the water surface. The refractive index is wavelength-dependent, and this leads to some disparity between images taken at different wavelengths. Given the camera orientation and the refractive index of a medium such as water, our approach can reconstruct the underwater scene with unknown water surface from the disparity observed in images taken at different wavelengths. We verified the effectiveness of our method through simulations and real experiments on various scenes. Shin Ishihara, Yuta Asano, Yinqiang Zheng, Imari Sato |
3DV | 3 |
| 2020 | An Integrated Enhancement Solution for 24-Hour Colorful ImagingabstractThe current industry practice for 24-hour outdoor imaging is to use a silicon camera supplemented with near-infrared (NIR) illumination. This will result in color images with poor contrast at daytime and absence of chrominance at nighttime. For this dilemma, all existing solutions try to capture RGB and NIR images separately. However, they need additional hardware support and suffer from various drawbacks, including short service life, high price, specific usage scenario, etc. In this paper, we propose a novel and integrated enhancement solution that produces clear color images, whether at abundant sunlight daytime or extremely low-light nighttime. Our key idea is to separate the VIS and NIR information from mixed signals, and enhance the VIS signal adaptively with the NIR signal as assistance. To this end, we build an optical system to collect a new VIS-NIR-MIX dataset and present a physically meaningful image processing algorithm based on CNN. Extensive experiments show outstanding results, which demonstrate the effectiveness of our solution. Feifan Lv, Yinqiang Zheng, Feng Lu 0005 |
AAAI | 2 |
| 2020 | Optical Flow in the DarkabstractMany successful optical flow estimation methods have been proposed, but they become invalid when tested in dark scenes because low-light scenarios are not considered when they are designed and current optical flow benchmark datasets lack low-light samples. Even if we preprocess to enhance the dark images, which achieves great visual perception, it still leads to poor optical flow results or even worse ones, because information like motion consistency may be broken while enhancing. We propose an end-to-end data-driven method that avoids error accumulation and learns optical flow directly from low-light noisy images. Specifically, we develop a method to synthesize large-scale low-light optical flow datasets by simulating the noise model on dark raw images. We also collect a new optical flow dataset in raw format with a large range of exposure to be used as a benchmark. The models trained on our synthetic dataset can relatively maintain optical flow accuracy as the image brightness descends and they outperform the existing methods greatly on low-light images. Yinqiang Zheng, Mingfang Zhang 0002, Feng Lu 0005 |
CVPR | 1 |
| 2020 | Layered Neighborhood Expansion for Incremental Multiple Graph Matching
Zhihui Xie 0002, Junchi Yan, Yinqiang Zheng, Xiaokang Yang 0001 |
ECCV (10) | 4 |
| 2020 | Learn to Recover Visible Color for Video Surveillance in a Day
Guangming Wu, Yinqiang Zheng, Zhiling Guo, Zekun Cai, Xiaodan Shi, Yifei Huang 0002, Ryosuke Shibasaki |
ECCV (1) | 2 |
| 2020 | Efficient Spatio-Temporal Recurrent Neural Network for Video Deblurring
Zhihang Zhong, Yinqiang Zheng |
ECCV (6) | 3 |
| 2020 | Beyond Intra-modality: A Survey of Heterogeneous Person Re-identificationabstractAn efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous person re-identification (Hetero-ReID). In this paper, we provide a comprehensive review of state-of-the-art Hetero-ReID methods that address the challenge of inter-modality discrepancies. According to the application scenario, we classify the methods into four categories --- low-resolution, infrared, sketch, and text. We begin with an introduction of ReID, and make a comparison between Homogeneous ReID (Homo-ReID) and Hetero-ReID tasks. Then, we describe and compare existing datasets for performing evaluations, and survey the models that have been widely employed in Hetero-ReID. We also summarize and compare the representative approaches from two perspectives, i.e., the application scenario and the learning pipeline. We conclude by a discussion of some future research directions. Follow-up updates are available at https://github.com/lightChaserX/Awesome-Hetero-reID Zheng Wang 0007, Zhixiang Wang 0001, Yinqiang Zheng, Yang Wu 0001, Wenjun Zeng 0001, Shin'ichi Satoh 0001 |
IJCAI | 3 |
| 2020 | Simultaneous hyperspectral image super-resolution and geometric alignment with a hybrid camera system
Ying Fu 0001, Yongrong Zheng, Yinqiang Zheng, Hua Huang 0001 |
Neurocomputing | 4 |
| 2020 | Illumination-Adaptive Person Re-IdentificationabstractMost person re-identification (ReID) approaches assume that person images are captured under relatively similar illumination conditions. In reality, long-term person retrieval is common, and person images are often captured under different illumination conditions at different times across a day. In this situation, the performances of existing ReID models often degrade dramatically. This paper addresses the ReID problem with illumination variations and names it as Illumination-Adaptive Person Re-identification (IA-ReID). We propose an Illumination-Identity Disentanglement (IID) network to dispel different scales of illuminations away while preserving individuals' identity information. To demonstrate the illumination issue and to evaluate our model, we construct two large-scale simulated datasets with a wide range of illumination variations. Experimental results on the simulated datasets and real-world images demonstrate the effectiveness of the proposed framework. Zelong Zeng, Zhixiang Wang 0001, Zheng Wang 0007, Yinqiang Zheng, Yung-Yu Chuang, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Hyperspectral Reconstruction with Redundant Camera Spectral Sensitivity FunctionsabstractHigh-resolution hyperspectral (HS) reconstruction has recently achieved significantly progress, among which the method based on the fusion of the RGB and HS images of the same scene can greatly improve the reconstruction performance compared with those based on the individually spectral or spatial enhancement. It is well known that the HS image is obtained only via the costly hypersoectral sensor, whereas the RGB images can be provided by low-price RGB cameras and the spectral sensitivity (SS) functions of RGB cameras are usually different. Thus, this study proposes a HS reconstruction, which fuses merely two RGB images with redundant spectral responses. In this work, we design a new RGB camera via shifting the SS of an existed RGB camera, which can provide similar strength of spectral response with different spectral centers of SS, and fuse the new achieved color image with an existed RGB image by a deep ResNet. Experiments validate that fusion of two existed RGB images can provide impressive HS reconstruction performance and further improvement can be achieved by integrating the color image of the simulated SS with the RGB image. Xianhua Han, Yinqiang Zheng, Jiande Sun 0001, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Hyperspectral Image Super-Resolution With Optimized RGB GuidanceabstractTo overcome the limitations of existing hyperspectral cameras on spatial/temporal resolution, fusing a low resolution hyperspectral image (HSI) with a high resolution RGB (or multispectral) image into a high resolution HSI has been prevalent. Previous methods for this fusion task usually employ hand-crafted priors to model the underlying structure of the latent high resolution HSI, and the effect of the camera spectral response (CSR) of the RGB camera on super-resolution accuracy has rarely been investigated. In this paper, we first present a simple and efficient convolutional neural network (CNN) based method for HSI super-resolution in an unsupervised way, without any prior training. Later, we append a CSR optimization layer onto the HSI super-resolution network, either to automatically select the best CSR in a given CSR dataset, or to design the optimal CSR under some physical restrictions. Experimental results show our method outperforms the state-of-the-arts, and the CSR optimization can further boost the accuracy of HSI super-resolution. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
CVPR | 3 |
| 2019 | Turn a Silicon Camera Into an InGaAs CameraabstractShort-wave infrared (SWIR) imaging has a wide range of applications for both industry and civilian. However, the InGaAs sensors commonly used for SWIR imaging suffer from a variety of drawbacks, including high price, low resolution, unstable quality, and so on. In this paper, we propose a novel solution for SWIR imaging using a common Silicon sensor, which has cheaper price, higher resolution and better technical maturity compared with the specialized InGaAs sensor. Our key idea is to approximate the response of the InGaAs sensor by exploiting the largely ignored sensitivity of a Silicon sensor, weak as it is, in the SWIR range. To this end, we build a multi-channel optical system to collect a new SWIR dataset and present a physically meaningful three-stage image processing algorithm on the basis of CNN. Both qualitative and quantitative experiments show promising experimental results, which demonstrate the effectiveness of the proposed method. Feifan Lv, Yinqiang Zheng, Feng Lu 0005 |
CVPR | 2 |
| 2019 | Learning to Reduce Dual-Level Discrepancy for Infrared-Visible Person Re-IdentificationabstractInfrared-Visible person RE-IDentification (IV-REID) is a rising task. Compared to conventional person re-identification (re-ID), IV-REID concerns the additional modality discrepancy originated from the different imaging processes of spectrum cameras, in addition to the person's appearance discrepancy caused by viewpoint changes, pose variations and deformations presented in the conventional re-ID task. The co-existed discrepancies make IV-REID more difficult to solve. Previous methods attempt to reduce the appearance and modality discrepancies simultaneously using feature-level constraints. It is however difficult to eliminate the mixed discrepancies using only feature-level constraints. To address the problem, this paper introduces a novel Dual-level Discrepancy Reduction Learning (D$^2$RL) scheme which handles the two discrepancies separately. For reducing the modality discrepancy, an image-level sub-network is trained to translate an infrared image into its visible counterpart and a visible image to its infrared version. With the image-level sub-network, we can unify the representations for images with different modalities. With the help of the unified multi-spectral images, a feature-level sub-network is trained to reduce the remaining appearance discrepancy through feature embedding. By cascading the two sub-networks and training them jointly, the dual-level reductions take their responsibilities cooperatively and attentively. Extensive experiments demonstrate the proposed approach outperforms the state-of-the-art methods. Zhixiang Wang 0001, Zheng Wang 0007, Yinqiang Zheng, Yung-Yu Chuang, Shin'ichi Satoh 0001 |
CVPR | 3 |
| 2019 | Polarimetric Camera Calibration Using an LCD MonitorabstractIt is crucial for polarimetric imaging to accurately calibrate the polarizer angles and the camera response function (CRF) of a polarizing camera. When this polarizing camera is used in a setting of multiview geometric imaging, it is often required to calibrate its intrinsic and extrinsic parameters as well, for which Zhang's calibration method is the most widely used with either a physical checker board, or more conveniently a virtual checker pattern displayed on a monitor. In this paper, we propose to jointly calibrate the polarizer angles and the inverse CRF (ICRF) using a slightly adapted checker pattern displayed on a liquid crystal display (LCD) monitor. Thanks to the lighting principles and the industry standards of the LCD monitors, the polarimetric and radiometric calibration can be significantly simplified, when assisted by the extrinsic parameters estimated from the checker pattern. We present a simple linear method for polarizer angle calibration and a convex method for radiometric calibration, both of which can be jointly refined in a process similar to bundle adjustment. Experiments have verified the feasibility and accuracy of the proposed calibration method. Zhixiang Wang 0001, Yinqiang Zheng, Yung-Yu Chuang |
CVPR | 2 |
| 2019 | Non-Local Intrinsic Decomposition With Near-Infrared PriorsabstractIntrinsic image decomposition is a highly under-constrained problem that has been extensively studied by computer vision researchers. Previous methods impose additional constraints by exploiting either empirical or data-driven priors. In this paper, we revisit intrinsic image decomposition with the aid of near-infrared (NIR) imagery. We show that NIR band is considerably less sensitive to textures and can be exploited to reduce ambiguity caused by reflectance variation, promoting a simple yet powerful prior for shading smoothness. With this observation, we formulate intrinsic decomposition as an energy minimisation problem. Unlike existing methods, our energy formulation decouples reflectance and shading estimation, into a convex local shading component based on NIR-RGB image pair, and a reflectance component that encourages reflectance homogeneity both locally and globally. We further show the minimisation process can be approached by a series of multi-dimensional kernel convolutions, each within linear time complexity. To validate the proposed algorithm, a NIR-RGB dataset is captured over real-world objects, where our NIR-assisted approach demonstrates clear superiority over RGB methods. Ziang Cheng, Yinqiang Zheng, Shaodi You, Imari Sato |
ICCV | 2 |
| 2019 | Learning to See Moving Objects in the DarkabstractVideo surveillance systems have wide range of utilities, yet easily suffer from great quality degeneration under dim light circumstances. Industrial solutions mainly use extra near-infrared illuminations, even though it doesn't preserve color and texture information. A variety of researches enhanced low-light videos shot by visible light cameras, while they either relied on task specific preconditions or trained with synthetic datasets. We propose a novel optical system to capture bright and dark videos of the exact same scenes, generating training and groud truth pairs for authentic low-light video dataset. A fully convolutional network with 3D and 2D miscellaneous operations is utilized to learn an enhancement mapping with proper spatial-temporal transformation from raw camera sensor data to bright RGB videos. Experiments show promising results by our method, and it outperforms state-of-the-art low-light image/video enhancement algorithms. Haiyang Jiang 0002, Yinqiang Zheng |
ICCV | 2 |
| 2019 | Image restoration from patch-based compressed sensing measurement
Hua Huang 0001, Guangtao Nie, Yinqiang Zheng, Ying Fu 0001 |
Neurocomputing | 3 |
| 2019 | Editorial
Junchi Yan, Minsu Cho, Francesc Serratosa, Gui-Song Xia, Yinqiang Zheng |
Pattern Recognit. Lett. | 5 |
| 2018 | Camera Pose Estimation With Unknown Principal PointabstractTo estimate the 6-DoF extrinsic pose of a pinhole camera with partially unknown intrinsic parameters is a critical sub-problem in structure-from-motion and camera localization. In most of existing camera pose estimation solvers, the principal point is assumed to be in the image center. Unfortunately, this assumption is not always true, especially for asymmetrically cropped images. In this paper, we develop the first exactly minimal solver for the case of unknown principal point and focal length by using four and a half point correspondences (P4.5Pfuv). We also present an extremely fast solver for the case of unknown aspect ratio (P5Pfuva). The new solvers outperform the previous state-of-the-art in terms of stability and speed. Finally, we explore the extremely challenging case of both unknown principal point and radial distortion, and develop the first practical non-minimal solver by using seven point correspondences (P7Pfruv). Experimental results on both simulated data and real Internet images demonstrate the usefulness of our new solvers. Viktor Larsson, Zuzana Kukelova, Yinqiang Zheng |
CVPR | 3 |
| 2018 | Deeply Learned Filter Response Functions for Hyperspectral ReconstructionabstractHyperspectral reconstruction from RGB imaging has recently achieved significant progress via sparse coding and deep learning. However, a largely ignored fact is that existing RGB cameras are tuned to mimic human trichromatic perception, thus their spectral responses are not necessarily optimal for hyperspectral reconstruction. In this paper, rather than use RGB spectral responses, we simultaneously learn optimized camera spectral response functions (to be implemented in hardware) and a mapping for spectral reconstruction by using an end-to-end network. Our core idea is that since camera spectral filters act in effect like the convolution layer, their response functions could be optimized by training standard neural networks. We propose two types of designed filters: a three-chip setup without spatial mosaicing and a single-chip setup with a Bayer-style 2x2 filter array. Numerical simulations verify the advantages of deeply learned spectral responses compared to existing RGB cameras. More interestingly, by considering physical restrictions in the design process, we are able to realize the deeply learned spectral response functions by using modern film filter production technologies, and thus construct data-inspired multispectral cameras for snapshot hyperspectral imaging. Shijie Nie, Lin Gu 0003, Yinqiang Zheng, Antony Lam, Nobutaka Ono, Imari Sato |
CVPR | 3 |
| 2018 | Self-Calibrating Polarising Radiometric CalibrationabstractWe present a self-calibrating polarising radiometric calibration method. From a set of images taken from a single viewpoint under different unknown polarising angles, we recover the inverse camera response function and the polarising angles relative to the first angle. The problem is solved in an integrated manner, recovering both of the unknowns simultaneously. The method exploits the fact that the intensity of polarised light should vary sinusoidally as the polarising filter is rotated, provided that the response is linear. It offers the first solution to demonstrate the possibility of radiometric calibration through polarisation. We evaluate the accuracy of our proposed method using synthetic data and real world objects captured using different cameras. The self-calibrated results were found to be comparable with those from multiple exposure sequence. Daniel Teo, Boxin Shi, Yinqiang Zheng, Sai-Kit Yeung |
CVPR | 3 |
| 2018 | Coded Illumination and Imaging for Fluorescence Based Classification
Yuta Asano, Misaki Meguro, Antony Lam, Yinqiang Zheng, Takahiro Okabe, Imari Sato |
ECCV (8) | 5 |
| 2018 | Polarimetric Three-View Geometry
Lixiong Chen, Yinqiang Zheng, Art Subpa-Asa, Imari Sato |
ECCV (16) | 2 |
| 2018 | Joint Camera Spectral Sensitivity Selection and Hyperspectral Image Recovery
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
ECCV (3) | 3 |
| 2018 | Simultaneous 3D Reconstruction for Water Surface and Underwater Scene
Yiming Qian, Yinqiang Zheng, Minglun Gong, Yee-Hong Yang |
ECCV (3) | 2 |
| 2018 | Stereo Relative Pose from Line and Point Feature Triplets
Alexander Vakhitov, Victor S. Lempitsky, Yinqiang Zheng |
ECCV (8) | 3 |
| 2018 | SSF-CNN: Spatial and Spectral Fusion with CNN for Hyperspectral Image Super-ResolutionabstractFusing a low-resolution hyperspectral image with the corresponding high-resolution RGB image to obtain a high-resolution hyperspectral image is usually solved as an optimization problem with prior-knowledge such as sparsity representation and spectral physical properties as constraints, which have limited applicability. Deep convolutional neural network extracts more comprehensive features and is proved to be effective in upsampling RGB images. However, directly applying CNNs to upsample either the spatial or spectral dimension alone may not produce pleasing results due to the neglect of complementary information from both low resolution hyper spectral and high resolution RGB images. This paper proposes two types of novel CNN architectures to take advantages of spatial and spectral fusion for hyperspectral image superresolution. Experiment results on benchmark datasets validate that the proposed spatial and spectral fusion CNNs outperforms the state-of-the-art methods and baseline CNN architectures in both quantitative values and visual qualities. Xianhua Han, Boxin Shi, Yinqiang Zheng |
ICIP | 3 |
| 2018 | Residual HSRCNN: Residual Hyper-Spectral Reconstruction CNN from an RGB ImageabstractHyper-spectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyper-spectral and high-resolution multi-spectral or RGB images can be captured at video rate. This study aims to generate a hyper-spectral image via enhancing spectral resolution of an RGB image, which might be easily obtained by a commodity camera. Motivated by the success of deep convolutional neural network (DCNN) for spatial resolution enhancement of natural images, we explore a spectral reconstruction CNN for spectral super-resolution with an available RGB image, which predicts the high-frequency content of the fine spectral wavelength in narrow band interval. Since the lost high-frequency content can not be perfectly recovered, by leveraging on the baseline CNN, we further propose a novel residual hyper-spectral reconstruction CNN framework to estimate the non-recovered high-frequency content (Residual) from the output of the baseline CNN. Experiments on benchmark hyper-spectral datasets validate that the proposed method achieves promising performances compared with the existing state-of-the-art methods. Xianhua Han, Boxin Shi, Yinqiang Zheng |
ICPR | 3 |
| 2018 | Hyperspectral Image Super-Resolution With a Mosaic RGB ImageabstractRecently, many hyperspectral (HS) image superresolution methods that merge a low spatial resolution HS image and a high spatial resolution three-channel RGB image have been proposed in spectral imaging. A largely ignored fact is that most existing commercial RGB cameras capture high resolution images by a single CCD/CMOS sensor equipped with a color filter array (CFA). In this paper, we account for the common imaging mechanism of commercial RGB cameras, and propose to use a mosaic RGB image for HS image super-resolution, which prevents demosaicing error and thus its propagation into the HS image super-resolution results. We design a proper nonlocal low-rank regularization to exploit the intrinsic properties - rich self-repeating patterns and high correlation across spectra - within HS images of natural scenes, and formulate the HS image super-resolution task into a variational optimization problem, which can be efficiently solved via the alternating direction method of multipliers (ADMM). The effectiveness of the proposed method has been evaluated on two benchmark datasets, demonstrating that the proposed method can provide substantial improvement over the current state-of-the-art HS image superresolution methods without considering the mosaicing effect. Finally, we show that our method can also perform well in the real capture system. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001, Imari Sato, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Self-Similarity Constrained Sparse Representation for Hyperspectral Image Super-ResolutionabstractFusing a low-resolution hyperspectral image with the corresponding high-resolution multispectral image to obtain a high-resolution hyperspectral image is an important technique for capturing comprehensive scene information in both spatial and spectral domains. Existing approaches adopt sparsity promoting strategy, and encode the spectral information of each pixel independently, which results in noisy sparse representation. We propose a novel hyperspectral image super-resolution method via a self-similarity constrained sparse representation. We explore the similar patch structures across the whole image and the pixels with close appearance in local regions to create globalstructure groups and local-spectral super-pixels. By forcing the similarity of the sparse representations for pixels belonging to the same group and super-pixel, we alleviate the effect of the outliers in the learned sparse coding. Experiment results on benchmark datasets validate that the proposed method outperforms the stateof- the-art methods in both quantitative metrics and visual effect. Xianhua Han, Boxin Shi, Yinqiang Zheng |
IEEE Trans. Image Process. | 3 |
| 2017 | A Microfacet-Based Reflectance Model for Photometric Stereo with Highly Specular SurfacesabstractA precise, stable and invertible model for surface reflectance is the key to the success of photometric stereo with real world materials. Recent developments in the field have enabled shape recovery techniques for surfaces of various types, but an effective solution to directly estimating the surface normal in the presence of highly specular reflectance remains elusive. In this paper, we derive an analytical isotropic microfacet-based reflectance model, based on which a physically interpretable approximate is tailored for highly specular surfaces. With this approximate, we identify the equivalence between the surface recovery problem and the ellipsoid of revolution fitting problem, where the latter can be described as a system of polynomials. Additionally, we devise a fast, non-iterative and globally optimal solver for this problem. Experimental results on both synthetic and real images validate our model and demonstrate that our solution can stably deliver superior performance in its targeted application domain. Lixiong Chen, Yinqiang Zheng, Boxin Shi, Art Subpa-Asa, Imari Sato |
ICCV | 2 |
| 2017 | From RGB to Spectrum for Natural Scenes via Manifold-Based MappingabstractSpectral analysis of natural scenes can provide much more detailed information about the scene than an ordinary RGB camera. The richer information provided by hyperspectral images has been beneficial to numerous applications, such as understanding natural environmental changes and classifying plants and soils in agriculture based on their spectral properties. In this paper, we present an efficient manifold learning based method for accurately reconstructing a hyperspectral image from a single RGB image captured by a commercial camera with known spectral response. By applying a nonlinear dimensionality reduction technique to a large set of natural spectra, we show that the spectra of natural scenes lie on an intrinsically low dimensional manifold. This allows us to map an RGB vector to its corresponding hyperspectral vector accurately via our proposed novel manifold-based reconstruction pipeline. Experiments using both synthesized RGB images using hyperspectral datasets and real world data demonstrate our method outperforms the state-of-the-art. Yan Jia 0005, Yinqiang Zheng, Lin Gu 0003, Art Subpa-Asa, Antony Lam, Yoichi Sato 0001, Imari Sato |
ICCV | 2 |
| 2017 | Making Minimal Solvers for Absolute Pose Estimation Compact and RobustabstractIn this paper we present new techniques for constructing compact and robust minimal solvers for absolute pose estimation. We focus on the P4Pfr problem, but the methods we propose are applicable to a more general setting. Previous approaches to P4Pfr suffer from artificial degeneracies which come from their formulation and not the geometry of the original problem. In this paper we show how to avoid these false degeneracies to create more robust solvers. Combined with recently published techniques for Gröbner basis solvers we are also able to construct solvers which are significantly smaller. We evaluate our solvers on both real and synthetic data, and show improved performance compared to competing solvers. Finally we show that our techniques can be directly applied to the P3.5Pf problem to get a non-degenerate solver, which is competitive with the current state-of-the-art. Viktor Larsson, Zuzana Kukelova, Yinqiang Zheng |
ICCV | 3 |
| 2017 | Visibility enhancement of fluorescent substance under ambient illumination using flash photographyabstractMany natural and manmade objects contain fluorescent substance. To visualize the distribution of fluorescence emitting substance is of great importance for food freshness examination, molecular dynamics analysis and so on. Unfortunately, the presence of fluorescent substance is usually imperceptible under strong ambient illumination, since fluorescent emission is relatively weak compared with surface reflectance. Even assuming that surface reflectance could be somehow blocked out, shading effect on fluorescent emission that relates to surface geometry would still interfere with visibility of fluorescent substance in the scene. In this paper, we propose a visibility enhancement method to better visualize the distribution of fluorescent substance under unknown and uncontrolled ambient illumination. By using an image pair captured with UV and visible flash illumination, we obtain a shading-free luminance image that visualizes the distribution of fluorescent emission. We further replace the luminance of the RGB image under ambient illumination by using this fluorescent emission luminance, so as to obtain a full colored image. The effectiveness of our method has been verified when used to visualize weak fluorescence from bacteria on rotting cheese and meat. Misaki Meguro, Yuta Asano, Yinqiang Zheng, Imari Sato |
ICIP | 3 |
| 2017 | Light transport component decomposition using multi-frequency illuminationabstractScene appearance is a mixture of light transport phenomena ranging from direct reflection to complicated effect such as inter-reflection and subsurface scattering. To decompose scene appearance into meaningful photometric components is very helpful in scene understanding and image editing. However, it has proven to be a difficult task. In this paper, we explore the difference of direct components obtained by multi-frequency illumination for light transport component decomposition. We apply independent vector analysis (IVA) to this task with no fixed constraints. Experiment results have verified the effectiveness of our method and its applicability to generic scenes. Art Subpa-Asa, Yinqiang Zheng, Nobutaka Ono, Imari Sato |
ICIP | 2 |
| 2017 | Camera spectral sensitivity, illumination and spectral reflectance estimation for a hybrid hyperspectral image capture systemabstractA variety of methods have been proposed to restore high resolution hyperspectral image (HSI) from a hybrid camera system, which captures high spatial resolution RGB images and low spatial resolution HSI. They focused unanimously on HSI super-resolution via fusion, yet did not explore the potential of this kind of system for camera spectral sensitivity (CSS), illumination spectrum, and high spatial resolution spectral reflectance recovery. In this paper, we present a sparse representation based method to estimate the CSS of the RGB camera under unknown illumination for the hybrid camera system. Furthermore, the illumination and high spatial resolution spectral reflectance are simultaneously recovered. Experimental results show the effectiveness of the proposed methods on camera spectral sensitivity, illumination spectrum and spectral reflectance recovery. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001 |
ICIP | 3 |
| 2017 | Hyper-spectral Image Super-resolution Using Non-negative Spectral Representation with Data-Guided SparsityabstractHyperspectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyperspectral and high-resolution multi-spectral images can be available using existing imaging techniques. This study aims to generate a high-resolution hyperspectral image via fusion of the available LR-HS and HR-MS images. We propose a novel hyperspectral image superresolution method via non-negative sparse representation of reflectance spectral with adaptive sparsity constraint. By analyzing local content similarity of a focused pixel in the available high-resolution multi-spectral image, which can measure pixel material purity according to surrounding pixels, we generate a sparsity map for guiding non-negative sparse coding optimization procedure of the spectral representation called non-negative spectral representation with data-guided sparsity. Since the proposed method adaptively adjust the sparsity in the spectral representation based on the local content of the available high-resolution multi-spectral image, it can produce more robust spectral representation for recovering the target high-resolution hyper-spectral image. Comprehensive experiments on two public hyperspectral datasets validate that the proposed method achieves promising performances compared with the existing state of the art methods. Xianhua Han, Jan Wang, Boxin Shi, Yinqiang Zheng, Yen-Wei Chen 0001 |
ISM | 4 |
| 2017 | Semi-supervised Learning for Biomedical Image Segmentation via Forest Oriented Super Pixels(Voxels)
Lin Gu 0003, Yinqiang Zheng, Ryoma Bise, Imari Sato, Nobuaki Imanishi, Sadakazu Aiso |
MICCAI (1) | 2 |
| 2017 | Separation of Transmitted Light and Scattering Components in Transmitted Microscopy
Mihoko Shimano, Ryoma Bise, Yinqiang Zheng, Imari Sato |
MICCAI (2) | 3 |
| 2016 | Spectral Reflectance Recovery with Interreflection Using a Hyperspectral Image
Hiroki Okawa, Yinqiang Zheng, Antony Lam, Imari Sato |
ACCV (4) | 2 |
| 2016 | Direct and Global Component Separation from a Single Image Using Basis Representation
Art Subpa-Asa, Ying Fu 0001, Yinqiang Zheng, Toshiyuki Amano, Imari Sato |
ACCV (3) | 3 |
| 2016 | Exploiting Spectral-Spatial Correlation for Coded Hyperspectral Image RestorationabstractConventional scanning and multiplexing techniques for hyperspectral imaging suffer from limited temporal and/or spatial resolution. To resolve this issue, coding techniques are becoming increasingly popular in developing snapshot systems for high-resolution hyperspectral imaging. For such systems, it is a critical task to accurately restore the 3D hyperspectral image from its corresponding coded 2D image. In this paper, we propose an effective method for coded hyperspectral image restoration, which exploits extensive structure sparsity in the hyperspectral image. Specifically, we simultaneously explore spectral and spatial correlation via low-rank regularizations, and formulate the restoration problem into a variational optimization model, which can be solved via an iterative numerical algorithm. Experimental results using both synthetic data and real images show that the proposed method can significantly outperform the state-of-the-art methods on several popular coding-based hyperspectral imaging systems. Ying Fu 0001, Yinqiang Zheng, Imari Sato, Yoichi Sato 0001 |
CVPR | 2 |
| 2016 | A Direct Least-Squares Solution to the PnP Problem with Unknown Focal LengthabstractIn this work, we propose a direct least-squares solution to the perspective-n-point (PnP) pose estimation problem of a partially uncalibrated camera, whose intrinsic parameters except the focal length are known. The basic idea is to construct a proper objective function with respect to the target variables and extract all its stationary points so as to find the global minimum. The advantages of our proposed solution over existing ones are that (i) the objective function is directly built upon the imaging equation, such that all the 3D-to-2D correspondences contribute equally to the minimized error, and that (ii) the proposed solution is noniterative, in the sense that the stationary points are retrieved by means of eigenvalue factorization and the common iterative refinement step is not needed. In addition, the proposed solution has O(n) complexity, and can be used to handle both planar and nonplanar 3D points. Experimental results show that the proposed solution is much more accurate than the existing state-of-the-art solutions, and is even comparable to the maximum likelihood estimation by minimizing the reprojection error. Yinqiang Zheng, Laurent Kneip |
CVPR | 1 |
| 2016 | Shape from Water: Bispectral Light Absorption for Depth Recovery
Yuta Asano, Yinqiang Zheng, Ko Nishino, Imari Sato |
ECCV (6) | 2 |
| 2016 | Simultaneous linear separation and unmixing of fluorescent and reflective components from a single hyperspectral imageabstractRecently an algorithm to separate fluorescent and reflective components from a hyperspectral image has been reported, in which the important task of spectral unmixing of multiple fluorescent components was left unresolved. In this paper, we present the algorithm to simultaneously separate those components and unmix fluorophores (SSUF: Simultaneous Separation and Unmixing of Fluorescent components) from a single hyperspectral image. Two variants are introduced for the cases when fluorophore spectra are known and unknown. Experimental results confirm the validity of the proposed method. Naoyuki Ohara, Yinqiang Zheng, Imari Sato, Tomoya Nakamura, Masahiro Yamaguchi 0002 |
ICIP | 2 |
| 2015 | Illumination and reflectance spectra separation of a hyperspectral image meets low-rank matrix factorizationabstractThis paper addresses the illumination and reflectance spectra separation (IRSS) problem of a hyperspectral image captured under general spectral illumination. The huge amount of pixels in a hypersepctral image poses tremendous challenges on computational efficiency, yet in turn offers greater color variety that might be utilized to improve separation accuracy and relax the restrictive subspace illumination assumption in existing works. We show that this IRSS problem can be modeled into a low-rank matrix factorization problem, and prove that the separation is unique up to an unknown scale under the standard low-dimensionality assumption of reflectance. We also develop a scalable algorithm for this separation task that works in the presence of model error and image noise. Experiments on both synthetic data and real images have demonstrated that our separation results are sufficiently accurate, and can benefit some important applications, such as spectra relighting and illumination swapping. Yinqiang Zheng, Imari Sato, Yoichi Sato 0001 |
CVPR | 1 |
| 2015 | Separating Fluorescent and Reflective Components by Using a Single Hyperspectral ImageabstractThis paper introduces a novel method to separate fluorescent and reflective components in the spectral domain. In contrast to existing methods, which require to capture two or more images under varying illuminations, we aim to achieve this separation task by using a single hyperspectral image. After identifying the critical hurdle in single-image component separation, we mathematically design the optimal illumination spectrum, which is shown to contain substantial high-frequency components in the frequency domain. This observation, in turn, leads us to recognize a key difference between reflectance and fluorescence in response to the frequency modulation effect of illumination, which fundamentally explains the feasibility of our method. On the practical side, we successfully find an off-the-shelf lamp as the light source, which is strong in irradiance intensity and cheap in cost. A fast linear separation algorithm is developed as well. Experiments using both synthetic data and real images have confirmed the validity of the selected illuminant and the accuracy of our separation algorithm. Yinqiang Zheng, Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001 |
ICCV | 1 |
| 2014 | Partial Symmetry in Polynomial Systems and Its Applications in Computer VisionabstractAlgorithms for solving systems of polynomial equations are key components for solving geometry problems in computer vision. Fast and stable polynomial solvers are essential for numerous applications e.g. minimal problems or finding for all stationary points of certain algebraic errors. Recently, full symmetry in the polynomial systems has been utilized to simplify and speed up state-of-the-art polynomial solvers based on Gröbner basis method. In this paper, we further explore partial symmetry (i.e. where the symmetry lies in a subset of the variables) in the polynomial systems. We develop novel numerical schemes to utilize such partial symmetry. We then demonstrate the advantage of our schemes in several computer vision problems. In both synthetic and real experiments, we show that utilizing partial symmetry allow us to obtain faster and more accurate polynomial solvers than the general solvers. Yubin Kuang, Yinqiang Zheng, Kalle Åström |
CVPR | 2 |
| 2014 | A General and Simple Method for Camera Pose and Focal Length DeterminationabstractIn this paper, we revisit the pose determination problem of a partially calibrated camera with unknown focal length, hereafter referred to as the PnPf problem, by using n(n ≥ 4) 3D-to-2D point correspondences. Our core contribution is to introduce the angle constraint and derive a compact bivariate polynomial equation for each point triplet. Based on this polynomial equation, we propose a truly general method for the PnPf problem, which is suited both to the minimal 4-point based RANSAC application, and also to large scale scenarios with thousands of points, irrespective of the 3D point configuration. In addition, by solving bivariate polynomial systems via the Sylvester resultant, our method is very simple and easy to implement. Its simplicity is especially obvious when one needs to develop a fast solver for the 4-point case on the basis of the characteristic polynomial technique. Experiment results have also demonstrated its superiority in accuracy and efficiency when compared with the existing state-of-the-art solutions. Yinqiang Zheng, Shigeki Sugimoto, Imari Sato, Masatoshi Okutomi |
CVPR | 1 |
| 2014 | Spectra Estimation of Fluorescent and Reflective Scenes by Using Ordinary Illuminants
Yinqiang Zheng, Imari Sato, Yoichi Sato 0001 |
ECCV (5) | 1 |
| 2013 | A Practical Rank-Constrained Eight-Point Algorithm for Fundamental Matrix EstimationabstractDue to its simplicity, the eight-point algorithm has been widely used in fundamental matrix estimation. Unfortunately, the rank-2 constraint of a fundamental matrix is enforced via a posterior rank correction step, thus leading to non-optimal solutions to the original problem. To address this drawback, existing algorithms need to solve either a very high order polynomial or a sequence of convex relaxation problems, both of which are computationally ineffective and numerically unstable. In this work, we present a new rank-2 constrained eight-point algorithm, which directly incorporates the rank-2 constraint in the minimization process. To avoid singularities, we propose to solve seven sub problems and retrieve their globally optimal solutions by using tailored polynomial system solvers. Our proposed method is noniterative, computationally efficient and numerically stable. Experiment results have verified its superiority over existing algebraic error based algorithms in terms of accuracy, as well as its advantages when used to initialize geometric error based algorithms. Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi |
CVPR | 1 |
| 2013 | Revisiting the PnP Problem: A Fast, General and Optimal SolutionabstractIn this paper, we revisit the classical perspective-n-point (PnP) problem, and propose the first non-iterative O(n) solution that is fast, generally applicable and globally optimal. Our basic idea is to formulate the PnP problem into a functional minimization problem and retrieve all its stationary points by using the Gr"obner basis technique. The novelty lies in a non-unit quaternion representation to parameterize the rotation and a simple but elegant formulation of the PnP problem into an unconstrained optimization problem. Interestingly, the polynomial system arising from its first-order optimality condition assumes two-fold symmetry, a nice property that can be utilized to improve speed and numerical stability of a Grobner basis solver. Experiment results have demonstrated that, in terms of accuracy, our proposed solution is definitely better than the state-of-the-art O(n) methods, and even comparable with the reprojection error minimization method. Yinqiang Zheng, Yubin Kuang, Shigeki Sugimoto, Kalle Åström, Masatoshi Okutomi |
ICCV | 1 |
| 2012 | Practical low-rank matrix approximation under robust L1-normabstractA great variety of computer vision tasks, such as rigid/nonrigid structure from motion and photometric stereo, can be unified into the problem of approximating a low-rank data matrix in the presence of missing data and outliers. To improve robustness, the L1-norm measurement has long been recommended. Unfortunately, existing methods usually fail to minimize the L1-based nonconvex objective function sufficiently. In this work, we propose to add a convex trace-norm regularization term to improve convergence, without introducing too much heterogenous information. We also customize a scalable first-order optimization algorithm to solve the regularized formulation on the basis of the augmented Lagrange multiplier (ALM) method. Extensive experimental results verify that our regularized formulation is reasonable, and the solving algorithm is very efficient, insensitive to initialization and robust to high percentage of missing data and/or outliers1. Yinqiang Zheng, Guangcan Liu, Shigeki Sugimoto, Shuicheng Yan, Masatoshi Okutomi |
CVPR | 1 |
| 2012 | Generalizing Wiberg algorithm for rigid and nonrigid factorizations with missing components and metric constraintsabstractIn spite of intensive endeavor over decades, rigid and nonrigid factorizations under metric constraints, possibly in the presence of missing components, remain to be very challenging. In this work, we try to break the hard nut by generalizing to these problems the Wiberg algorithm, one of the most successful solutions for unconstrained bilinear factorization. To properly handle missing components, we advocate a bilinear factorization formulation with an extra mean vector. In spirit of the Wiberg algorithm, we first propose an efficient and initialization-insensitive algorithm for unconstrained factorization, posterior correction of whose solution offers reasonable initialization for metric upgrade. For factorization with metric constraints, we reformulate it into an unconstrained problem through quaternion parametrization, which merges elegantly into our unconstrained factorization algorithm. Extensive experiment results verify that our proposed methods are fast, accurate and robust to high percentage of missing components. Yinqiang Zheng, Shigeki Sugimoto, Shuicheng Yan, Masatoshi Okutomi |
CVPR | 1 |
| 2011 | Deterministically maximizing feasible subsystem for robust model fitting with unit norm constraintabstractMany computer vision problems can be accounted for or properly approximated by linearity, and the robust model fitting (parameter estimation) problem in presence of outliers is actually to find the Maximum Feasible Subsystem (MaxFS) of a set of infeasible linear constraints. We propose a deterministic branch and bound method to solve the MaxFS problem with guaranteed global optimality. It can be used in a wide class of computer vision problems, in which the model variables are subject to the unit norm constraint. In contrast to the convex and concave relaxations in existing works, we introduce a piecewise linear relaxation to build very tight under- and over-estimators for square terms by partitioning variable bounds into smaller segments. Based on this novel relaxation technique, our branch and bound method can converge in a few iterations. For homogeneous linear systems, which correspond to some quasi-convex problems based on L∞-L∞-norm, our method is non-iterative and certainly reaches the globally optimal solution at the root node by partitioning each variable range into two segments with equal length. Throughout this work, we rely on the so-called Big-M method, and successfully avoid potential numerical problems by exploiting proper parametrization and problem structure. Experimental results demonstrate the stability and efficiency of our proposed method. Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi |
CVPR | 1 |
| 2011 | A branch and contract algorithm for globally optimal fundamental matrix estimationabstractWe propose a unified branch and contract method to estimate the fundamental matrix with guaranteed global optimality, by minimizing either the Sampson error or the point to epipolar line distance, and explicitly handling the rank-2 constraint and scale ambiguity. Based on a novel denominator linearization strategy, the fundamental matrix estimation problem can be transformed into an equivalent problem that involves 9 squared univariate, 12 bilinear and 6 trilin-ear terms. We build tight convex and concave relaxations for these nonconvex terms and solve the problem deterministically under the branch and bound framework. For acceleration, a bound contraction mechanism is introduced to reduce the size of the branching region at the root node. Given high-quality correspondences and proper data normalization, our experiments show that the state-of-the-art locally optimal methods generally converge to the globally optimal solution. However, they indeed have the risk of being trapped into local minimum in case of noise. As another important experimental result, we also demonstrate, from the viewpoint of global optimization, that the point to epipolar line distance is slightly inferior to the Sampson error in case of drastically varying object scales across two views. Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi |
CVPR | 1 |
| 2010 | 3D Structure Refinement of Nonrigid Surfaces through Efficient Image Alignment
Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi |
ACCV (4) | 1 |
| 2009 | Using silhouette for pose estimation of object with surface of revolutionabstractA novel approach of pose estimation is proposed for the object with surface of revolution(SOR). The silhouette of the object is the only information necessary for this method and no cross section circle (latitude circle) is needed. In this article, we explain the property of tangent circle and use it to establish constraint between two images of object with different poses. Such constraint can help to solve the pose of object in both images. We test our method with a simulation experiment and use it to estimate the pose for both rigid body and articulated object. Yinqiang Zheng, Yuncai Liu |
ICIP | 2 |
| 2008 | Another way of looking at monocular circle pose estimationabstractIt is well known that there are generally two possible sets of pose parameters from one calibrated perspective view of a circle. What is the relation between these two possible sets? Where does this ambiguity arise from? In which case can this ambiguity be resolved? Considering these questions, we suggest a novel viewpoint toward circle pose estimation from a single view. Different from existing methods on the basis of analytical geometry, we originally develop the projective equation of a circle, based on which a closed form solution is developed and a brand new geometric explanation for the ambiguity of solutions is presented. Experimental results verify the correctness of our proposed method. Yinqiang Zheng, Wenjuan Ma, Yuncai Liu |
ICIP | 1 |
| 2008 | Deformable surface stereo tracking-by-detection using Second Order Cone ProgrammingabstractWe present a method for tracking deformable surfaces in 3D using a stereo rig. Different from traditional recursive tracking approaches that provide a strong prior on the pose for each new frame, the proposed method tracks deformable surfaces by detecting them in individual frames. In our method, the model of the surface is represented by a triangulated mesh. The constraints for model to image keypoint correspondences, together with the constraints that preserve the lengths of mesh edges, are formulated as Second Order Cone Programming (SOCP) constraints, leading this tracking-by-detection method to be an SOCP problem that can be effectively solved. Experiments on a piece of deformed paper demonstrate the capability of the proposed tracking-by-detection method. Shuhan Shen, Yinqiang Zheng, Yuncai Liu |
ICPR | 2 |
| 2008 | The projective equation of a circle and its application in camera calibrationabstractIn this article, we present the projective equation of a circle in a perspective view, which naturally encodes such important geometric entities as the projected circle center, the vanishing point of the normal direction of the circle’s supporting plane and the degenerate conic envelope spanned by the image of circular points (ICPs). Based on this projective equation, we propose an easy technique to calibrate the focal length and the extrinsic parameters of a camera merely by using one perspective view of two arbitrary coplanar circles. Unlike existing optimization algorithm, our method offers a closed form solution through simple matrix manipulation. Experimental results verify the correctness and efficiency of our proposed technique. Yinqiang Zheng, Yuncai Liu |
ICPR | 1 |