Boxin Shi

dblp:69/783 · DBLP profile ↗
← Back
247ranked-venue papers
14as first author
154since 2021 · last 2026
0000-0001-6749-0364ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 189 · 9 first-author · 130 since 2021Graphics, computer vision, multimedia, augmented reality and games · 171 · 11 first-author · 95 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Event-Based Multi-Range Radiance Separation and 3D Reconstruction via Line-Scan Pseudo-Square Illumination
abstract
Decomposing scene radiance into physically meaningful components, including direct reflection, interreflection, and scattering, enables a deeper understanding of scene appearance. In this paper, we propose the first method to perform multi-range radiance component separation using only events captured by an event camera, without requiring any additional frame-based measurements. Our approach scans the scene by swiping line-shaped illumination across it, while exploiting the event camera's high temporal resolution and wide dynamic range to recover both direct and multiple global components corresponding to different light propagation distances. To address the noise inherent in event-integration-based radiance recovery, we present a pixel-wise calibration strategy that leverages the reproducibility of per-pixel noise patterns. We demonstrate that this calibration is highly effective in suppressing noise, enabling stable recovery from subtle signals. Moreover, we show that by detecting the timing at which the scanning line passes each pixel, the same line-scan event data can be exploited for coarse 3D reconstruction. Experimental results on real scenes show that our event-based approach achieves faster and finer component separation, while also enabling coarse depth estimation without the exposure control required by frame-based cameras.
Ryuji Hashimoto, Yuta Asano, Shin Ishihara, Bohan Yu, Chu Zhou, Boxin Shi, Imari Sato
3DV6
2026 ReContraster: Making Your Posters Stand Out with Regional Contrast
abstract
Effective poster design requires rapidly capturing attention and clearly conveying messages.Inspired by the "contrast effects" principle, we propose ReContraster, the first training-free model to leverage regional contrast to make posters stand out.By emulating the cognitive behaviors of a poster designer, ReContraster introduces the compositional multi-agent system to identify elements, organize layout, and evaluate generated poster candidates.To further ensure harmonious transitions across region boundaries, ReContraster integrates the hybrid denoising strategy during the diffusion process.We additionally contribute a new benchmark dataset for comprehensive evaluation.Seven quantitative metrics and four user studies confirm its superiority over relevant state-of-the-art methods, producing visually striking and aesthetically appealing posters.
Peixuan Zhang, Zijian Jia, Ziqi Cai, Shuchen Weng, Si Li 0001, Boxin Shi
ACL (1)6
2026 L-VOCAL: Language-based Video Colorization with Audio Alignment
Shuchen Weng, Huan Ouyang, Yuchen Hong, Lihan Lin, Si Li 0001, Boxin Shi
Int. J. Comput. Vis.7
2026 Affective Image Editing: Shaping Emotional Factors via Text Descriptions
Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li 0001, Boxin Shi
Int. J. Comput. Vis.7
2026 L-C4: Language-based video colorization for creative and consistent color
Shuchen Weng, Huan Ouyang, Lihan Lin, Yu Li 0003, Si Li 0001, Boxin Shi
Neurocomputing7
2026 Visual-in-Visual: A Unified and Efficient Baseline for Image Restoration
abstract
Recent years have witnessed remarkable progress in image restoration, yet achieving both high performance and efficiency remains a persistent challenge. To address this issue, we present VIVNet, a strong and efficient unified baseline designed to balance accuracy and practicality. Drawing inspiration from the high efficiency of the human visual system, VIVNet embeds a biologically inspired micro visual module into each block of a macro U-shaped vision architecture. This module mimics key perceptual processes such as retinal encoding, lateral inhibition, and high-order processing by combining lightweight depth-wise convolutions for multi-receptive-field feature extraction, a similarity-aware weighting mechanism to emphasize informative signals, and high-order interactions implemented via iterative element-wise multiplication to capture complex dependencies. This design enhances the model's representational capacity while maintaining computational efficiency. Unlike most existing methods that are limited to narrow task settings, we evaluate VIVNet across a wide range of scenarios, including general, all-in-one, and composite degradation tasks, as well as ultra-high-definition (UHD), underwater, medical, and remote sensing datasets. Extensive experiments show that VIVNet delivers competitive performance with high efficiency.
Yuning Cui 0001, Wenqi Ren, Boxin Shi, Alois C. Knoll
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Coded Event Focal Stack for Continuous Refocusing in Dynamic Scene
abstract
Traditional cameras face limitations in maintaining focus across dynamic scenes, especially during rapid motion, due to the constraints of their lenses. Post-capture refocusing techniques, including deep learning-based methods and light field cameras, have been explored to mitigate these challenges. However, these approaches frequently struggle with temporal consistency or experience a trade-off in spatial resolution. In this paper, we introduce the coded event focal stack, a novel approach that captures both motion and depth information through event streams recorded during a modulated focal sweep. Our coded event focal stack enables the generation of full-time intermediate frames refocused at arbitrary focal distances. Extensive experiments on both synthetic and real-world datasets demonstrate the superior refocusing capability of our method over state-of-the-art techniques, particularly in dynamic scenes with complex motion and depth variations.
Minggui Teng, Suhang Xuan, Zhiang Yan, Hanyue Lou, Bin Fan 0002, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Efficient 3D Surface Super-Resolution via Normal-Based Multimodal Restoration
abstract
High-fidelity 3D surface is essential for vision tasks across various domains such as medical imaging, cultural heritage preservation, quality inspection, virtual reality, and autonomous navigation. However, the intricate nature of 3D data representations poses significant challenges in restoring diverse 3D surfaces while capturing fine-grained geometric details at a low cost. This paper introduces an efficient multimodal normal-based 3D surface super-resolution (mn3DSSR) framework, designed to address the challenges of microgeometry enhancement and computational overhead. Specifically, we have constructed one of the largest normal-based multimodal dataset, ensuring superior data quality and diversity through meticulous subjective selection. Furthermore, we explore a new two-branch multimodal alignment approach along with a multimodal split fusion module to mitigate computational complexity while improving restoration performances. To address the limitations associated with normal-based multimodal learning, we develop novel normal-induced loss functions that facilitate geometric consistency and improve feature alignment. Extensive experiments conducted on seven benchmark datasets across four different 3D data representations demonstrate that mn3DSSR consistently outperforms state-of-the-art super-resolution methods in terms of restoration accuracy with high computational efficiency.
Miaohui Wang, Yunheng Liu, Wuyuan Xie, Boxin Shi, Jianmin Jiang
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Toward Deeper Emotional Reflection: Crafting Affective Image Filters With Generative Priors
abstract
Social media platforms enable users to express emotions by posting text with accompanying images. In this paper, we propose the Affective Image Filter (AIF) task, which aims to reflect visually-abstract emotionsfrom text into visually-concrete images, thereby creating emotionally compelling results. We first introduce the AIF dataset and the formulation of the AIF models. Then, we present AIF-B as an initial attempt based on a multi-modal transformer architecture. After that, we propose AIF-D as an extension of AIF-B towards deeper emotional reflection, effectively leveraging generative priors from pre-trained large-scale diffusion models. Quantitative and qualitative experiments demonstrate that AIF models achieve superior performance for both content consistency and emotional fidelity compared to state-of-the-art methods. Extensive user study experiments demonstrate that AIF models are significantly more effective at evoking specific emotions. Based on the presented results, we comprehensively discuss the value and potential of AIF models.
Peixuan Zhang, Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Toward a Unified Complementary Fusion Framework for Robust Polarimetric Imaging
abstract
Polarization, as an intrinsic property of light alongside amplitude and phase, has demonstrated great potential in a variety of downstream applications by providing valuable physical cues encoded in the degree of polarization (DoP) and the angle of polarization (AoP). Polarimetric imaging aims to acquire these polarimetric parameters by capturing polarized snapshots. However, compared to conventional imaging, it faces greater difficulties due to the presence of polarizers, which attenuate light intensity in a spatially variant manner. Such attenuation complicates exposure control: a short exposure leads to low signal-to-noise ratio and color distortion, whereas a relatively long exposure increases the risk of motion blur and saturation. To address these challenges, this work proposes PolFusion+, a unified framework that robustly produces clean and sharp polarized snapshots by complementarily fusing a degraded pair of short-exposed noisy and long-exposed blurry inputs. Building upon a polarization-aware three-phase fusion scheme, PolFusion+ introduces two key advancements. First, to handle saturation in the blurry snapshot, the irradiance restoration phase extracts and rectifies color information from both inputs, effectively mitigating saturation-induced degradation. Second, to ensure physically faithful polarization reconstruction, the framework explicitly models the individual characteristics and interdependencies of the DoP and AoP, enabling their joint restoration. These improvements are supported by a degradation-oriented neural network tailored to the fusion scheme. Experimental results demonstrate that PolFusion+ achieves state-of-the-art performance, effectively benefiting downstream applications.
Chu Zhou, Minggui Teng, Chao Xu 0006, Boxin Shi, Imari Sato
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 CDIR: LoRA-Inspired Attention for Efficient Composite Degradation Image Restoration
abstract
Specialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability.
Yuning Cui 0001, Wenqi Ren, Boxin Shi, Jianhou Gan, Alois C. Knoll
IEEE Trans. Image Process.3
2026 Dark-EvGS: Event Camera as an Eye for Radiance Field in the Dark
abstract
In low-light environments, conventional cameras often struggle to capture clear multi-view images of objects due to dynamic range limitations and motion blur caused by long exposure. Event cameras, with their high-dynamic range and high-speed properties, have the potential to mitigate these issues. Additionally, 3D Gaussian Splatting (GS) enables radiance field reconstruction, facilitating bright frame synthesis from multiple viewpoints in low-light conditions. However, naively using an event-assisted 3D GS approach still faced challenges because, in low lights, events are noisy, frames lack quality, and the color tone may be inconsistent. To address these issues, we propose Dark-EvGS, the first event-assisted 3D GS framework that enables the reconstruction of bright frames from arbitrary viewpoints along the camera trajectory. Triplet-level supervision is proposed to gain holistic knowledge, granular details, and sharp scene rendering. The color tone matching block is proposed to guarantee the color consistency of the rendered frames. Furthermore, we introduce the first real-captured dataset for the event-guided bright frame synthesis task via 3D GS-based radiance field reconstruction. Experiments demonstrate that our method achieves better results than existing methods, conquering radiance field reconstruction under challenging low-light conditions. The code and sample data are included in the supplementary material.
Jingqian Wu, Peiqi Duan 0002, Zongqiang Wang, Changwei Wang 0001, Boxin Shi, Edmund Y. Lam
IEEE Trans. Image Process.5
2025 Zero-Shot Low-Light Image Enhancement via Latent Diffusion Models
abstract
Low-light image enhancement (LLIE) aims to improve visibility and signal-to-noise ratio in images captured under poor lighting conditions. While deep learning has shown promise in this domain, current approaches require extensive paired training data, limiting their practical utility. We present a novel framework that reformulates low-light image enhancement as a zero-shot inference problem using pre-trained latent diffusion models (LDMs), eliminating the need for task-specific training data. Our key insight is that the rich natural image priors encoded in LDMs can be leveraged to recover well-lit images through a carefully designed optimization process. To address the ill-posed nature of low-light degradation and the complexity of latent space optimization, our framework introduces an exposure-aware degradation module that adaptively models illumination variations and a principled latent regularization scheme with adaptive guidance that ensures both enhancement quality and natural image statistics. Experimental results demonstrate that our framework outperforms existing zero-shot methods across diverse real-world scenarios.
Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Yuhui Quan, Boxin Shi, Yong Xu 0007
AAAI5
2025 PlaNet: Learning to Mitigate Atmospheric Turbulence in Planetary Images
abstract
Obtaining planetary images with good visual quality is not an easy task since they are usually degenerated by atmospheric turbulence during the imaging procedure. Existing atmospheric turbulence mitigation methods designed for conventional images cannot be applied to planetary images, since the objects on the Earth have totally different degeneration patterns to planets. Besides, in planetary imaging, photographers often capture as many frames as possible to reduce the noise level of planetary images, which requires the method designed for planetary images to support an arbitrary number of input frames. In this paper, we propose a vertical distance-aware turbulence simulation pipeline to synthesize realistic planetary images in accordance with their unique degeneration patterns at a large scale with affordable computational cost, and design a neural network to mitigate the turbulence with flexible input frames by adopting an edge-based supervision strategy to handle the background scarcity issue. Experimental results show that our method achieves state-of-the-art performance on both synthetic and real-world images.
Chu Zhou, Chengxuan Zhu, Boxin Shi
AAAI5
2025 Polarization Guided Mask-Free Shadow Removal
abstract
Shadow is a phenomenon that degenerates image quality and decreases the performance of downstream vision algorithms. Despite the fact that current image shadow removal methods have achieved promising progress, many of them require an externally obtained shadow mask as a necessary part of the input data, which not only introduces additional workload but also leads to degenerated performance near the shadow boundary due to the inaccuracy of the mask. Some of them do not require the shadow mask, however, they need to simultaneously consider the restoration of the brightness and color information along with the preservation of the texture and structure information inside the shadow region without external clues, which poses highly ill-posedness and makes the results prone to artifacts. In this paper, we propose Pol-ShaRe, the first Polarization-guided image Shadow Removal solution, to remove shadow in a mask-free manner with fewer artifacts. Specifically, it consists of a two-stage pipeline to relieve the ill-posedness and a neural network tailored to the pipeline to suppress the artifacts. Experimental results show that our Pol-ShaRe achieves state-of-the-art performance on both synthetic and real-world images.
Chu Zhou, Boxin Shi
AAAI3
2025 PhyS-EdiT: Physics-aware Semantic Image Editing with Text Description
abstract
Achieving joint control over material properties, lighting, and high-level semantics in images is essential for applications in digital media, advertising, and interactive design. Existing methods often isolate these properties, lacking a cohesive approach to manipulating them simultaneously. We introduce PhyS-EdiT, a novel diffusion-based model that enables precise control over four critical material properties: roughness, metallicity, albedo, and transparency while integrating lighting and semantic adjustments within a single framework. To facilitate this disentangled control, we present PR-TIPS, a large and diverse synthetic dataset designed to improve the disentanglement of material and lighting effects. PhyS-EdiT incorporates a dual-network architecture and robust training strategies to balance low-level physical realism with high-level semantic coherence, supporting localized and continuous property adjustments. Extensive experiments demonstrate the superiority of PhyS-EdiT in editing both synthetic and real-world images, achieving state-of-the-art performance on material, lighting, and semantic editing tasks.
Ziqi Cai, Shuchen Weng, Boxin Shi
CVPR4
2025 Unified Reconstruction of Static and Dynamic Scenes from Events
abstract
This paper addresses the challenge that current event-based video reconstruction methods cannot produce static background information. Recent research has uncovered the potential of event cameras in capturing static scenes. Nonetheless, image quality deteriorates due to noise interference and detail loss, failing to provide reliable background information. We propose a two-stage reconstruction strategy to address these challenges and reconstruct static scene images comparable to frame cameras. Building on this, we introduce the URSEE framework designed for reconstructing motion videos with static backgrounds. This framework includes a parallel channel that can simultaneously process static and dynamic events, and a network module designed to reconstruct videos encompassing both static and dynamic scenes in an end-to-end manner. We also collect a real-captured dataset for static reconstruction, containing both indoor and outdoor scenes. Comparison results indicate that the proposed method achieves state-of-the-art performance on both synthetic and real data.
Qiyao Gao, Peiqi Duan 0002, Hanyue Lou, Minggui Teng, Ziqi Cai, Xu Chen 0001, Boxin Shi
CVPR7
2025 SpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS), a recently emerged multi-view 3D reconstruction technique, has shown significant advantages in real-time rendering and explicit editing. However, 3DGS encounters challenges in the accurate modeling of both high-frequency view-dependent appearances and global illumination effects, including inter-reflection. This paper introduces SpecTRe-GS, which addresses these challenges and models highly Specular surfaces that reflect nearby objects through Tracing Rays in 3D Gaussian Splatting. SpecTRe-GS separately models reflections from highly specular and rough surfaces to leverage the distinctions between their reflective properties and integrates an efficient ray tracer within the 3DGS framework for querying secondary rays, thus achieving fast and accurate rendering. Also, it incorporates normal prior guidance and joint geometry optimization at various stages of the training process to enhance geometry reconstruction for undistorted reflections. Experiments on both synthetic and real-world scenes demonstrate the superiority of SpecTRe-GS compared to existing 3DGS-based methods in capturing highly specular inter-reflections and also showcase its editing applications.
Jiajun Tang 0001, Zhihao Li 0002, Shiyong Liu, Youyu Chen, Binxiao Huang, Boxin Shi
CVPR10
2025 VIRES: Video Instance Repainting via Sketch and Text Guided Generation
abstract
We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generative priors of text-to-video models to maintain temporal consistency and produce visually pleasing results. We propose the Sequential ControlNet with the standardized self-scaling, which effectively extracts structure layouts and adaptively captures high-contrast sketch details. We further augment the diffusion transformer backbone with the sketch attention to interpret and inject fine-grained sketch semantics. A sketch-aware encoder ensures that repainted results are aligned with the provided sketch sequence. Additionally, we contribute the VIRESET, a dataset with detailed annotations tailored for training and evaluating video instance editing methods. Experimental results demonstrate the effectiveness of VIRES, which outperforms state-of-the-art methods in visual quality, temporal consistency, condition alignment, and human ratings. The code, dataset and pretrained models are available at: https://hjzheng.net/projects/VIRES.
Shuchen Weng, Haojie Zheng, Peixuan Zhang, Yuchen Hong, Si Li 0001, Boxin Shi
CVPR7
2025 EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event Camera
abstract
Simultaneously acquisition of the surface normal and reflectance parameters is a crucial but challenging technique in the field of computer vision and graphics. It requires capturing multiple high dynamic range (HDR) images in existing methods using frame-based cameras. In this paper, we propose EventPSR, the first work to recover surface normal and reflectance parameters (e.g., metallic and roughness) simultaneously using an event camera. Compared with the existing methods based on photometric stereo or neural radiance fields, EventPSR is a robust and efficient approach that works consistently with different materials. Thanks to the extremely high temporal resolution and high dynamic range coverage of event cameras, EventPSR can recover accurate surface normal and reflectance of objects with various materials in 10 seconds. Extensive experiments on both synthetic data and real objects show that compared with existing methods using more than 100 HDR images, EventPSR recovers comparable surface normal and reflectance parameters with only about 30% of the data rate.
Bohan Yu, Jin Han 0001, Boxin Shi, Imari Sato
CVPR3
2025 Active Hyperspectral Imaging Using an Event Camera
abstract
Hyperspectral imaging plays a critical role in numerous scientific and industrial fields. Conventional hyperspectral imaging systems often struggle with the trade-off between capture speed, spectral resolution, and bandwidth, particularly in dynamic environments. In this work, we present a novel event-based active hyperspectral imaging system designed for real-time capture with low bandwidth in dynamic scenes. By combining an event camera with a dynamic illumination strategy, our system achieves unprecedented temporal resolution while maintaining high spectral fidelity, all at a fraction of the bandwidth requirements of traditional systems. Unlike basis-based methods that sacrifice spectral resolution for efficiency, our approach enables continuous spectral sampling through an innovative "sweeping rainbow" illumination pattern synchronized with a rotating mirror array. The key insight is leveraging the sparse, asynchronous nature of event cameras to encode spectral variations as temporal contrasts, effectively transforming the spectral reconstruction problem into a series of geometric constraints. Extensive evaluations of both synthetic and real data demonstrate that our system outperforms state-of-the-art methods in temporal resolution while maintaining competitive spectral reconstruction quality.
Bohan Yu, Jinxiu Liang, Zhuofeng Wang, Bin Fan 0002, Art Subpa-Asa, Boxin Shi, Imari Sato
CVPR6
2025 PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution
abstract
Polarization cameras can capture multiple polarized images with different polarizer angles in a single shot, bringing convenience to polarization-based downstream tasks. However, their direct outputs are color-polarization filter array (CPFA) raw images, requiring demosaicing to reconstruct full-resolution, full-color polarized images; unfortunately, this necessary step introduces artifacts that make polarization-related parameters such as the degree of polarization (DoP) and angle of polarization (AoP) prone to error. Besides, limited by the hardware design, the resolution of a polarization camera is often much lower than that of a conventional RGB camera. Existing polarized image demosaicing (PID) methods are limited in that they cannot enhance resolution, while polarized image super-resolution (PISR) methods, though designed to obtain high-resolution (HR) polarized images from the demosaicing results, tend to retain or even amplify errors in the DoP and AoP introduced by demosaicing artifacts. In this paper, we propose PIDSR, a joint framework that performs complementary Polarized Image Demosaicing and Super-Resolution, showing the ability to robustly obtain high-quality HR polarized images with more accurate DoP and AoP from a CPFA raw image in a direct manner. Experiments show our PIDSR not only achieves state-of-the-art performance on both synthetic and real data, but also facilitates downstream tasks.
Shuangfan Zhou, Chu Zhou, Youwei Lyu, Heng Guo 0003, Zhanyu Ma, Boxin Shi, Imari Sato
CVPR6
2025 No Redundancy, No Stall: Lightweight Streaming 3D Gaussian Splatting for Real-time Rendering
abstract
3D Gaussian Splatting (3DGS) enables high-quality rendering of 3D scenes and is getting increasing adoption in domains like autonomous driving and embodied intelligence. However, 3DGS still faces major efficiency challenges when faced with high frame rate requirements and resource-constrained edge deployment. To enable efficient 3DGS, in this paper, we propose LS-Gaussian, an algorithm/hardware co-design framework for lightweight streaming 3D rendering. LS-Gaussian is motivated by the core observation that 3DGS suffers from substantial computation redundancy and stalls. On one hand, in practical scenarios, high-frame-rate 3DGS is often applied in settings where a camera observes and renders the same scene continuously but from slightly different viewpoints. Therefore, instead of rendering each frame separately, LS-Gaussian proposes a viewpoint transformation algorithm that leverages inter-frame continuity for efficient sparse rendering. On the other hand, as different tiles within an image are rendered in parallel but have imbalanced workloads, frequent hardware stalls also slow down the rendering process. LS-Gaussian predicts the workload for each tile based on viewpoint transformation to enable more balanced parallel computation and co-designs a customized 3DGS accelerator to support the workload-aware mapping in real-time. Experimental results demonstrate that LS-Gaussian achieves 5.41× speedup over the edge GPU baseline on average and up to 17.3× speedup with the customized accelerator, while incurring only minimal visual quality degradation.
Linye Wei, Jiajun Tang 0001, Boxin Shi, Runsheng Wang
ICCAD4
2025 PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction
Yufei Han 0002, Bowen Tie, Heng Guo 0003, Youwei Lyu, Si Li 0001, Boxin Shi, Zhanyu Ma
ICCV6
2025 EventUPS: Uncalibrated Photometric Stereo Using an Event Camera
Jinxiu Liang, Bohan Yu, Haotian Zhuang, Jieji Ren, Peiqi Duan 0002, Boxin Shi
ICCV7
2025 Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset
abstract
With more event datasets being released online, safeguarding the event dataset against unauthorized usage has become a serious concern for data owners. Unlearnable Examples are proposed to prevent the unauthorized exploitation of image datasets. However, it's unclear how to create unlearnable asynchronous event streams to prevent event misuse. In this work, we propose the first unlearnable event stream generation method to prevent unauthorized training from event datasets. A new form of asynchronous event error-minimizing noise is proposed to perturb event streams, tricking the unauthorized model into learning embedded noise instead of realistic features. To be compatible with the sparse event, a projection strategy is presented to sparsify the noise to render our unlearnable event streams (UEvs). Extensive experiments demonstrate that our method effectively protects event data from unauthorized exploitation, while preserving their utility for legitimate use. We hope our UEvs contribute to the advancement of secure and trustworthy event dataset sharing. Code is available at: https://github.com/rfww/uevs.
Ruofei Wang, Peiqi Duan 0002, Boxin Shi, Renjie Wan
ICCV3
2025 AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic Scenes
abstract
Mainstream high dynamic range imaging techniques typically rely on fusing multiple images captured with different exposure setups (shutter speed and ISO). A good balance between shutter speed and ISO is crucial for achieving high-quality HDR, as high ISO values introduce significant noise, while long shutter speeds can lead to noticeable motion blur. However, existing methods often overlook the complex interaction between shutter speed and ISO and fail to account for motion blur effects in dynamic scenes. In this work, we propose AdaptiveAE, a reinforcement learning-based method that optimizes the selection of shutter speed and ISO combinations to maximize HDR reconstruction quality in dynamic environments. AdaptiveAE integrates an image synthesis pipeline that incorporates motion blur and noise simulation into our training procedure, leveraging semantic information and exposure histograms. It can adaptively select optimal ISO and shutter speed sequences based on a user-defined exposure time budget, and find a better exposure schedule than traditional solutions. Experimental results across multiple datasets demonstrate that it achieves the state-of-the-art performance.
Boxin Shi, Tianfan Xue, Yujin Wang 0001
ICCV3
2025 SpikeDiff: Zero-Shot High-Quality Video Reconstruction from Chromatic Spike Camera and Sub-Millisecond Spike Streams
Jinxiu Liang, Zhaojun Huang, Yeliduosi Xiaokaiti, Yakun Chang, Zhaofei Yu, Boxin Shi
ICCV7
2025 Event-Guided HDR Reconstruction with Diffusion Priors
Yixin Yang 0008, Jiawei Zhang 0002, Yunxuan Wei, Dongqing Zou, Jimmy S. J. Ren, Boxin Shi
ICCV7
2025 PolarAnything: Diffusion-based Polarimetric Image Synthesis
abstract
Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a parametric polarization image formation model and requires extensive 3D assets covering shape and PBR materials, preventing it from generating large-scale photorealistic images. To address this problem, we propose PolarAnything, capable of synthesizing polarization images from a single RGB input with both photorealism and physical accuracy, eliminating the dependency on 3D asset collections. Drawing inspiration from the zero-shot performance of pretrained diffusion models, we introduce a diffusion-based generative framework with an effective representation strategy that preserves the fidelity of polarization properties. Experiments show that our model generates high-quality polarization images and supports downstream tasks like shape from polarization.
Kailong Zhang, Youwei Lyu, Heng Guo 0003, Si Li 0001, Zhanyu Ma, Boxin Shi
ICCV6
2025 Event-Based Visual Vibrometry
Peiqi Duan 0002, Yeliduosi Xiaokaiti, Chao Xu 0006, Boxin Shi
ICCV5
2025 Polarimetric Neural Field via Unified Complex-Valued Wave Representation
Chu Zhou, Yixin Yang 0008, Junda Liao, Heng Guo 0003, Boxin Shi, Imari Sato
ICCV5
2025 BokehDiff: Neural Lens Blur with One-Step Diffusion
abstract
We introduce BokehDiff, a novel lens blur rendering method that achieves physically accurate and visually appealing outcomes, with the help of generative diffusion prior. Previous methods are bounded by the accuracy of depth estimation, generating artifacts in depth discontinuities. Our method employs a physics-inspired self-attention module that aligns with the image formation process, incorporating depth-dependent circle of confusion constraint and self-occlusion effects. We adapt the diffusion model to the one-step inference scheme without introducing additional noise, and achieve results of high quality and fidelity. To address the lack of scalable paired data, we propose to synthesize photorealistic foregrounds with transparency with diffusion models, balancing authenticity and scene diversity.
Chengxuan Zhu, Qingnan Fan, Qi Zhang 0066, Huaqi Zhang, Boxin Shi
ICCV7
2025 EDeF-Net: Spatio-temporal Association Network for Flicker Removal in Event Streams
abstract
Event cameras with bio-inspired neuromorphic sensors are highly sensitive to brightness changes. When there are moving objects in a scene under constant lighting, event cameras only record motion information and output a sequence of events asynchronously. However, the common flickering light sources, such as fluorescent or LED lamps powered by alternating current exist in various real-world scenarios. When operating under a flickering light source, event cameras output numerous redundant event signals that are triggered by the flickering effect, which overwhelm the useful signals that encode motion information. In this paper, we propose EDeF-Net, an Event streams DeFlickering Network that effectively leverages the spatio-temporal correlation of event streams by modeling both the inter-channel temporal attention and inter-patch spatial attention. To facilitate network training and evaluation, we synthesize the first dataset containing paired flickering and flicker-free event streams. Moreover, we demonstrate that event streams filtered by EDeF-Net yield performance improvements on down-stream applications such as event-based optical flow estimation and object tracking.
Jin Han 0001, Yixin Yang 0008, Zhan Zhan, Boxin Shi, Imari Sato
ACM Multimedia4
2025 Dense Metric Depth Estimation via Event-based Differential Focus Volume Prompting
abstract
Dense metric depth estimation has witnessed great developments in recent years. While single-image-based methods have demonstrated commendable performance in certain circumstances, they may encounter challenges regarding scale ambiguities and visual illusions in real world. Traditional depth-from-focus methods are constrained by low sampling rates during data acquisition. In this paper, we introduce a novel approach to enhance dense metric depth estimation by fusing events with image foundation models via a prompting approach. Specifically, we build Event-based Differential Focus Volumes (EDFV) using events triggered through focus sweeping, which are subsequently transformed into sparse metric depth maps. These maps are then utilized for prompting dense depth estimation via our proposed Event-based Depth Prompting Network. We further construct synthetic and real-captured datasets to facilitate the training and evaluation of both frame-based and event-based methods. Quantitative and qualitative results, including both in-domain and zero-shot experiments, demonstrate the superior performance of our method compared to existing approaches. Code and data will be available at https://github.com/liboyu02/EDFV/.
Peiqi Duan 0002, Zhaojun Huang, Boxin Shi
NeurIPS6
2025 V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
abstract
Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150× reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours—an order of magnitude larger than existing event datasets, yielding substantial improvements.
Hanyue Lou, Jinxiu Liang, Minggui Teng, Boxin Shi
NeurIPS5
2025 Audio-Sync Video Generation with Multi-Stream Temporal Control
abstract
Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio narratives (e.g., Podcasts or historical recordings). However, existing approaches fall short in generating high-quality videos with precise audio-visual synchronization, especially across diverse and complex audio types. In this work, we introduce MTV, a versatile framework for audio-sync video generation. MTV explicitly separates audios into speech, effects, and music tracks, enabling disentangled control over lip motion, event timing, and visual mood, respectively—resulting in fine-grained and semantically aligned video generation. To support the framework, we additionally present DEMIX, a dataset comprising high-quality cinematic videos and demixed audio tracks. DEMIX is structured into five overlapped subsets, enabling scalable multi-stage training for diverse generation scenarios. Extensive experiments demonstrate that MTV achieves state-of-the-art performance across six standard metrics spanning video quality, text-video consistency, and audio-video alignment.
Shuchen Weng, Haojie Zheng, Si Li 0001, Boxin Shi
NeurIPS5
2025 PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
Shuchen Weng, Jingqi Liu, Chengxuan Zhu, Minggui Teng, Zijian Jia, Boxin Shi
NeurIPS9
2025 Spk2ImgMamba: Spiking Camera Image Reconstruction with Multi-Scale State Space Models
abstract
As a bio-inspired vision sensor, the spiking camera has showcased remarkable capability in high-speed imaging with a sampling rate of 40,000 Hz. Reconstructing clear images from continuous spike streams, which is obtained by each photosensor continuously detecting photons and firing them asynchronously, has garnered significant attention. Despite promising results, existing spike-to-image reconstruction methods face challenges in balancing global receptive fields and efficient computation due to the inherent limitations of their backbones. Recently, due to powerful long-range modeling and linear complexity, the state space model (SSM) has emerged as a competitive alternative to CNNs and Transformers. In this paper, we propose a lightweight spike-to-image reconstruction network that harnesses Mamba as the backbone. Our approach sequentially executes three core modules: temporal information integration, spatial feature enhancement, and progressive image reconstruction. The former accumulates cues across diverse temporal windows to explore both long-term and short-term contexts. Subsequently, to model global dependencies while heightening local detail perception, we develop a multi-scale SSM block characterized by multi-scale multi-direction scanning, which effectively boosts spatial feature representations. Finally, intensity images are decoded progressively from the enhanced light-intensity features. Extensive experiments on both synthetic and real-captured data demonstrate that our approach achieves state-of-the-art performance, with only 10% of the network parameters and nearly two orders of magnitude less computational effort. The code will be available at https://github.com/interstellarH/Spk2ImgMamba.
Jiaoyang Yin, Bin Fan 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
WACV5
2025 Self-supervised Shutter Unrolling with Events
Mingyuan Lin, Yangguang Wang, Xiang Zhang 0022, Boxin Shi, Wen Yang 0001, Chu He, Gui-Song Xia, Lei Yu 0006
Int. J. Comput. Vis.4
2025 Learning to Deblur Polarized Images
Chu Zhou, Minggui Teng, Chao Xu 0006, Imari Sato, Boxin Shi
Int. J. Comput. Vis.6
2025 EventAid: Benchmarking Event-Aided Image/Video Enhancement Algorithms With Real-Captured Hybrid Dataset
abstract
Event cameras are emerging imaging technology that offer advantages over conventional frame-based imaging sensors in dynamic range and sensing speed. Complementing the rich texture and color perception of traditional image frames, the hybrid camera system of event and frame-based cameras enables high-performance imaging. With the assistance of event cameras, high-quality image/video enhancement methods make it possible to break the limits of traditional frame-based cameras, especially exposure time, resolution, dynamic range, and frame rate limits. This paper focuses on five event-aided image and video enhancement tasks (i.e., event-based video reconstruction, event-aided high frame rate video reconstruction, image deblurring, image super-resolution, and high dynamic range image reconstruction), provides an analysis of the effects of different event properties, a real-captured and ground truth labeled benchmark dataset, a unified benchmarking of state-of-the-art methods, and an evaluation for two mainstream event simulators. In detail, this paper collects a real-captured evaluation dataset EventAid for five event-aided image/video enhancement tasks, by using "Event-RGB" multi-camera hybrid system, taking into account scene diversity and spatiotemporal synchronization. We further perform quantitative and visual comparisons for state-of-the-art algorithms, provide a controlled experiment to analyze the performance limit of event-aided image deblurring methods, and discuss open problems to inspire future research.
Peiqi Duan 0002, Yixin Yang 0008, Hanyue Lou, Minggui Teng, Yi Ma 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier Embedding
abstract
This paper introduces a one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), for non-Lambertian objects under unknown light directions. It departs from traditional two-stage methods that first explicitly learn lighting information and then estimate surface normals. Two-stage methods were deployed because the interplay of lighting with shading cues presents challenges for directly estimating surface normals without explicit lighting information. However, these two-stage networks are disjointed and separately trained so that the error in explicit light calibration will propagate to the second stage and cannot be eliminated. In contrast, the proposed FUPS-Net utilizes an embedded Fourier transform network to implicitly learn lighting features by decomposing inputs, rather than employing a disjointed light estimation network. Our approach is motivated from observations in the Fourier domain of photometric stereo images: lighting information is mainly encoded in amplitudes, while geometry information is mainly associated with phases. Leveraging this property, our method "decomposes" geometry and lighting in the Fourier domain as guidance, via the proposed Fourier Embedding Extraction (FEE) block and Fourier Embedding Aggregation (FEA) block, which generate lighting and geometry features for the FUPS-Net to implicitly resolve the geometry-lighting ambiguity. Furthermore, we propose a Frequency-Spatial Weighted (FSW) block that assigns weights to combine features extracted from the frequency domain and those from the spatial domain for enhancing surface reconstructions. FUPS-Net overcomes the limitations of two-stage UPS methods, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks. Experimental results on synthetic and real datasets demonstrate the superior performance of our approach, and its simpler training setup, potentially paving the way for a new strategy in deep learning-based UPS methods.
Yakun Ju, Boxin Shi, Bihan Wen, Kin-Man Lam 0001, Xudong Jiang 0001, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 The NeRF Signature: Codebook-Aided Watermarking for Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based creations, the need for copyright protection has emerged as a critical issue. Although some approaches have been proposed to embed digital watermarks into NeRF, they often neglect essential model-level considerations and incur substantial time overheads, resulting in reduced imperceptibility and robustness, along with user inconvenience. In this paper, we extend the previous criteria for image watermarking to the model level and propose NeRF Signature, a novel watermarking method for NeRF. We employ a Codebook-aided Signature Embedding (CSE) that does not alter the model structure, thereby maintaining imperceptibility and enhancing robustness at the model level. Furthermore, after optimization, any desired signatures can be embedded through the CSE, and no fine-tuning is required when NeRF owners want to use new binary signatures. Then, we introduce a joint pose-patch encryption watermarking strategy to hide signatures into patches rendered from a specific viewpoint for higher robustness. In addition, we explore a Complexity-Aware Key Selection (CAKS) scheme to embed signatures in high visual complexity patches to enhance imperceptibility. The experimental results demonstrate that our method outperforms other baseline methods in terms of imperceptibility and robustness.
Ziyuan Luo, Anderson Rocha 0001, Boxin Shi, Qing Guo 0005, Haoliang Li, Renjie Wan
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Guest Editorial: Introduction to the Special Section on Computational Photography
Boxin Shi, Ashok Veeraraghavan, Roarke Horstmeyer, Wolfgang Heidrich
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Revisiting Supervised Learning-Based Photometric Stereo Networks
abstract
Deep learning has significantly propelled the development of photometric stereo by handling the challenges posed by unknown reflectance and global illumination effects. However, how supervised learning-based photometric stereo networks resolve these challenges remains to be elucidated. In this paper, we aim to reveal how existing methods address these challenges by revisiting their deep features, deep feature encoding strategies, and network architectures. Based on the insights gained from our analysis, we propose ESSENCE-Net, which effectively encodes deep shading features with an easy-first-encoding strategy, enhances shading features with shading supervision, and accurately decodes normal with spatial context-aware attention. The experimental results verify that the proposed method outperforms state-of-the-art methods on three benchmark datasets, whether with dense or sparse inputs.
Xiaoyao Wei, Zongrui Li 0001, Binjie Ding, Boxin Shi, Xudong Jiang 0001, Gang Pan 0001, Yanlong Cao
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 OpenCIR: Conditional Image Repainting With Open Condition Mixture
abstract
In this paper, we introduce OpenCIR, a fully-functional Conditional Image Repainting (CIR) model designed for local image editing. Given an image and a combination of conditions related to geometry, texture, and color, CIR models are required to repaint instances and seamlessly composite them with the original images. Previous CIR models suffer from limited object categories, restricted condition modalities, and demanded geometry precision. In contrast, leveraging the generative priors from pre-trained models, OpenCIR could repaint open object categories. Equipped with redesigned condition injection modules and the condition extension strategy, OpenCIR is able to understand open condition modalities. Adopting the contour refinement strategy, OpenCIR allows users to specify instances with open geometry precision. In addition, we contribute the Open-CIR dataset, which includes detailed annotations, tailored for the comprehensive training and evaluation of the OpenCIR model. Extensive experiments demonstrate that OpenCIR outperforms relevant state-of-the-art methods, achieving superior visual quality, and more favorable results by human evaluators.
Shuchen Weng, Xiaocheng Gong, Haojie Zheng, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Self-Supervised Learning for Rolling Shutter Temporal Super-Resolution
abstract
Most cameras on portable devices adopt a rolling shutter (RS) mechanism, encoding sufficient temporal dynamic information through sequential readouts. This advantage can be exploited to recover a temporal sequence of latent global shutter (GS) images. Existing methods rely on fully supervised learning, necessitating specialized optical devices to collect paired RS-GS images as ground-truth, which is too costly to scale. In this paper, we propose a self-supervised learning framework for the first time to produce a high frame rate GS video from two consecutive RS images, unleashing the potential of RS cameras. Specifically, we first develop the unified warping model of RS2GS and GS2RS, enabling the complement conversions of RS2GS and GS2RS to be incorporated into a uniform network model. Then, based on the cycle consistency constraint, given a triplet of consecutive RS frames, we minimize the discrepancy between the input middle RS frame and its cycle reconstruction, generated by interpolating back from the predicted two intermediate GS frames. Experiments on various benchmarks show that our approach achieves comparable or better performance than state-of-the-art supervised methods while enjoying stronger generalization capabilities. Moreover, our approach makes it possible to recover smooth and distortion-free videos from two adjacent RS frames in the real-world BS-RSC dataset, surpassing prior limitations.
Bin Fan 0002, Ying Guo 0014, Yuchao Dai, Chao Xu 0006, Boxin Shi
IEEE Trans. Circuits Syst. Video Technol.5
2025 Detail-Preserving Diffusion Models for Low-Light Image Enhancement
abstract
Existing diffusion models for low-light image enhancement typically incrementally remove noise introduced during the forward diffusion process using a denoising loss, with the process being conditioned on input low-light images. While these models demonstrate remarkable abilities in generating realistic high-frequency details, they often struggle to restore fine details that are faithful to the input. To address this, we present a novel detail-preserving diffusion model for realistic and faithful low-light image enhancement. Our approach integrates a size-agnostic diffusion process with a reverse process reconstruction loss, significantly enhancing the fidelity of enhanced images to their low-light counterparts and enabling more accurate recovery of fine details. To ensure the preservation of region- and content-aware details, we employ an efficient noise estimation network with a simplified channel-spatial attention mechanism. Additionally, we propose a multiscale ensemble scheme to maintain detail fidelity across diverse illumination regions. Comprehensive experiments on eight benchmark datasets demonstrate that our method achieves state-of-the-art results compared to over twenty existing methods in terms of both perceptual quality (LPIPS) and distortion metrics (PSNR and SSIM). The code is available at:https://github.com/CSYanH/DePDiff.
Yan Huang 0031, Xiaoshan Liao, Jinxiu Liang, Boxin Shi, Yong Xu 0007, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.4
2025 SweepEvGS: Event-Based 3D Gaussian Splatting for Macro and Micro Radiance Field Rendering From a Single Sweep
abstract
Recent advancements in 3D Gaussian Splatting (3D-GS) have demonstrated the potential of using 3D Gaussian primitives for high-speed, high-fidelity, and cost-efficient novel view synthesis from continuously calibrated input views. However, conventional methods require high-frame-rate dense and high-quality sharp images, which are time-consuming and inefficient to capture, especially in dynamic environments. Event cameras, with their high temporal resolution and ability to capture asynchronous brightness changes, offer a promising alternative for more reliable scene reconstruction without motion blur. In this paper, we propose SweepEvGS, a novel hardware-integrated method that leverages event cameras for robust and accurate novel view synthesis across various imaging settings from a single sweep. SweepEvGS utilizes the initial static frame with dense event streams captured during a single camera sweep to effectively reconstruct detailed scene views. We also introduce different real-world hardware imaging systems for real-world data collection and evaluation for future research. We validate the robustness and efficiency of SweepEvGS through experiments in three different imaging settings: synthetic objects, real-world macro-level, and real-world micro-level view synthesis. Our results demonstrate that SweepEvGS surpasses existing methods in visual rendering quality, rendering speed, and computational efficiency, highlighting its potential for dynamic practical applications.
Jingqian Wu, Chutian Wang, Boxin Shi, Edmund Y. Lam
IEEE Trans. Circuits Syst. Video Technol.4
2024 Colorizing Monochromatic Radiance Fields
abstract
Though Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fields becomes crucial. To achieve this goal, instead of manipulating the monochromatic radiance fields directly, we consider it as a representation-prediction task in the Lab color space. By first constructing the luminance and density representation using monochromatic images, our prediction stage can recreate color representation on the basis of an image colorization module. We then reproduce a colorful implicit model through the representation of luminance, density, and color. Extensive experiments have been conducted to validate the effectiveness of our approaches. Our project page: https://liquidammonia.github.io/color-nerf.
Yean Cheng, Renjie Wan, Shuchen Weng, Chengxuan Zhu, Yakun Chang, Boxin Shi
AAAI6
2024 Pano-NeRF: Synthesizing High Dynamic Range Novel Views with Geometry from Sparse Low Dynamic Range Panoramic Images
abstract
Panoramic imaging research on geometry recovery and High Dynamic Range (HDR) reconstruction becomes a trend with the development of Extended Reality (XR). Neural Radiance Fields (NeRF) provide a promising scene representation for both tasks without requiring extensive prior data. How- ever, in the case of inputting sparse Low Dynamic Range (LDR) panoramic images, NeRF often degrades with under-constrained geometry and is unable to reconstruct HDR radiance from LDR inputs. We observe that the radiance from each pixel in panoramic images can be modeled as both a signal to convey scene lighting information and a light source to illuminate other pixels. Hence, we propose the irradiance fields from sparse LDR panoramic images, which increases the observation counts for faithful geometry recovery and leverages the irradiance-radiance attenuation for HDR reconstruction. Extensive experiments demonstrate that the irradiance fields outperform state-of-the-art methods on both geometry recovery and HDR reconstruction and validate their effectiveness. Furthermore, we show a promising byproduct of spatially-varying lighting estimation. The code is available at https://github.com/Lu-Zhan/Pano-NeRF.
Zhan Lu, Boxin Shi, Xudong Jiang 0001
AAAI3
2024 DiLiGenRT: A Photometric Stereo Dataset with Quantified Roughness and Translucency
abstract
Photometric stereo faces challenges from non-Lambertian reflectance in real-world scenarios. Systematically measuring the reliability of photometric stereo methods in handling such complex reflectance necessitates a real-world dataset with quantitatively controlled reflectances. This paper introduces DiLiGenRT, the first real-world dataset for evaluating photometric stereo methods under quantified reflectances by manufacturing 54 hemispheres with varying degrees of two reflectance properties: Roughness and Transluency, Unlike qualitative and semantic labels, such as “diffuse” and “specular,” that have been used in previous datasets, our quantified dataset allows comprehensive and systematic benchmark evaluations. In addition, it facilitates selecting best-fit photometric stereo methods based on the quantitative reflectance properties. Our dataset and benchmark results are available at https://photometricstereo.github.io/diligentrt.html.
Heng Guo 0003, Jieji Ren, Feishi Wang, Boxin Shi, Ming Jun Ren, Yasuyuki Matsushita
CVPR4
2024 Real-Time 3D-Aware Portrait Video Relighting
abstract
Synthesizing realistic videos of talking faces under custom lighting conditions and viewing angles benefits various downstream applications like video conferencing. However, most existing relighting methods are either time-consuming or unable to adjust the viewpoints. In this paper, we present the first real-time 3D-aware method for relighting in-the-wild videos of talking faces based on Neural Radiance Fields (NeRF). Given an input portrait video, our method can synthesize talking faces under both novel views and novel lighting conditions with a photo-realistic and disentangled 3D representation. Specifically, we infer an albedo tri-plane, as well as a shading tri-plane based on a desired lighting condition for each video frame with fast dual-encoders. We also leverage a temporal consistency network to ensure smooth transitions and reduce flickering artifacts. Our method runs at 32.98 fps on consumer-level hardware and achieves state-of-the-art results in terms of reconstruction quality, lighting error, lighting instability, temporal consistency and inference speed. We demonstrate the effectiveness and interactivity of our method on various portrait videos with diverse lighting and viewing conditions.
Ziqi Cai, Yukun Lai, Hongbo Fu 0001, Boxin Shi, Lin Gao 0004
CVPR6
2024 Towards HDR and HFR Video from Rolling-Mixed-Bit Spikings
abstract
The spiking cameras offer the benefits of high dynamic range (HDR), high temporal resolution, and low data redundancy. However, reconstructing HDR videos in high-speed conditions using single-bit spikings presents challenges due to the limited bit depth. Increasing the bit depth of the spikings is advantageous for boosting HDR performance, but the readout efficiency will be decreased, which is unfavorable for achieving a high frame rate (HFR) video. To address these challenges, we propose a readout mechanism to obtain rolling-mixed-bit (RMB) spikings, which involves inter-leaving multi-bit spikings within the single-bit spikings in a rolling manner, thereby combining the characteristics of high bit depth and efficient readout. Furthermore, we introduce RMB-Net for reconstructing HDR and HFR videos. RMB-Net comprises a cross-bit attention block for fusing mixed-bit spikings and a cross-time attention block for achieving temporal fusion. Extensive experiments conducted on synthetic and real-synthetic data demonstrate the superiority of our method. For instance, pure 3 -bit spikings result in 3 times of data volume, whereas our method achieves comparable performance with less than 2% increase in data volume.
Yakun Chang, Yeliduosi Xiaokaiti, Yujia Liu 0005, Bin Fan 0002, Zhaojun Huang, Tiejun Huang 0001, Boxin Shi
CVPR7
2024 VMINer: Versatile Multi-view Inverse Rendering with Near-and Far-field Light Sources
abstract
This paper introduces a versatile multi-view inverse rendering framework with near-and far-field light sources. Tackling the fundamental challenge of inherent ambiguity in inverse rendering, our framework adopts a lightweight yet inclusive lighting model for different near-and far-field lights, thus is able to make use of input images under varied lighting conditions available during capture. It leverages observations under each lighting to disentangle the intrinsic geometry and material from the external lighting, using both neural radiance field rendering and physically-based surface rendering on the 3D implicit fields. After training, the reconstructed scene is extracted to a textured triangle mesh for seamless integration into industrial rendering soft-ware for various applications. Quantitatively and qualitatively tested on synthetic and real-world scenes, our method shows superiority to state-of-the-art multi-view inverse rendering methods in both speed and quality.
Jiajun Tang 0001, Ping Tan 0002, Boxin Shi
CVPR4
2024 NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images
abstract
We present NeRSP, a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand, sparse image inputs, as a practical capture setting, commonly cause incomplete or distorted results due to the lack of correspondence matching. This paper jointly han-dles the challenges from sparse inputs and reflective surfaces by leveraging polarized images. We derive photomet-ric and geometric cues from the polarimetric image formation model and multiview azimuth consistency, which jointly optimize the surface geometry modeled via implicit neural representation. Based on the experiments on our synthetic and real datasets, we achieve the state-of-the-art surface reconstruction results with only 6 views as input.
Yufei Han 0002, Heng Guo 0003, Koki Fukai, Hiroaki Santo, Boxin Shi, Fumio Okura, Zhanyu Ma
CVPR5
2024 Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction
abstract
Reliable hand mesh reconstruction (HMR) from commonly-used color and depth sensors is challenging es-pecially under scenarios with varied illuminations and fast motions. Event camera is a highly promising alternative for its high dynamic range and dense temporal resolution properties, but it lacks salient texture appearance for hand mesh reconstruction. In this paper, we propose EvRGBHand - the first approach for 3D hand mesh reconstruction with an event camera and an RGB camera compensating for each other. By fusing two modalities of data across time, space, and information dimensions, EvRGBHand can tackle overexposure and motion blur issues in RGB-based HMR and foreground scarcity as well as background overflow issues in event-based HMR. We further propose EvRGBDegrader, which allows our model to generalize effectively in challenging scenes, even when trained solely on standard scenes, thus reducing data acquisition costs. Experiments on real-world data demonstrate that EvRGBHand can effectively solve the challenging issues when using either type of camera alone via retaining the merits of both, and shows the potential of generalization to outdoor scenes and another type of event camera. For code, models, and dataset, please refer to https://alanjiang98.github.io/.evrgbhand.github.io/.
Bingxuan Wang, Xiaoming Deng 0001, Boxin Shi
CVPR6
2024 Spin-UP: Spin Light for Natural Light Uncalibrated Photometric Stereo
abstract
Natural Light Uncalibrated Photometric Stereo (NaUPS) relieves the strict environment and light assumptions in classical Uncalibrated Photometric Stereo (UPS) methods. However, due to the intrinsic ill-posedness and high-dimensional ambiguities, addressing NaUPS is still an open question. Existing works impose strong assumptions on the environment lights and objects' material, restricting the effectiveness in more general scenarios. Alternatively, some methods leverage supervised learning with intricate models while lacking interpretability, resulting in a biased estimation. In this work, we propose Spin Light Uncalibrated Photometric Stereo (Spin-UP), an unsupervised method to tackle NaUPS in various environment lights and objects. The proposed method uses a novel setup that captures the object's images on a rotatable platform, which mitigates NaUPS's ill-posedness by reducing unknowns and provides reliable priors to alleviate NaUPS's ambiguities. Leveraging neural inverse rendering and the proposed training strategies, Spin-UP recovers surface normals, environment light, and isotropic reflectance under complex natural light with low computational cost. Experiments have shown that Spin-UP outperforms other supervised / unsupervised NaUPS meth-ods and achieves state-of-the-art performance on synthetic and real-world datasets. Codes and data are available at https://github.com/LMozart/CVPR2024-SpinUP.
Zongrui Li 0001, Zhan Lu, Haojie Yan, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001
CVPR4
2024 Neural Underwater Scene Representation
abstract
Among the numerous efforts towards digitally recovering the physical world, Neural Radiance Fields (NeRFs) have proved effective in most cases. However, underwater scene introduces unique challenges due to the absorbing water medium, the local change in lighting and the dynamic contents in the scene. We aim at developing a neural under-water scene representation for these challenges, modeling the complex process of attenuation, unstable in-scattering and moving objects during light transport. The proposed method can reconstruct the scenes from both established datasets and in-the-wild videos with outstanding fidelity.
Yunkai Tang, Chengxuan Zhu, Renjie Wan, Boxin Shi
CVPR5
2024 NB-GTR: Narrow-Band Guided Turbulence Removal
abstract
The removal of atmospheric turbulence is crucial for long-distance imaging. Leveraging the stochastic nature of atmospheric turbulence, numerous algorithms have been developed that employ multi-frame input to mitigate the tur-bulence. However, when limited to a single frame, existing algorithms face substantial performance drops, partic-ularly in diverse real-world scenes. In this paper, we propose a robust solution to turbulence removal from an RGB image under the guidance of an additional narrow-band image, broadening the applicability of turbulence mitigation techniques in real-world imaging scenarios. Our approach exhibits a substantial suppression in the magnitude of tur-bulence artifacts by using only a pair of images, thereby enhancing the clarity and fidelity of the captured scene.
Chu Zhou, Chengxuan Zhu, Minggui Teng, Boxin Shi
CVPR6
2024 Latency Correction for Event-Guided Deblurring and Frame Interpolation
abstract
Event cameras, with their high temporal resolution, dynamic range, and low power consumption, are particu-larly good at time-sensitive applications like deblurring and frame interpolation. However, their performance is hindered by latency variability, especially under low-light conditions and with fast-moving objects. This paper addresses the challenge of latency in event cameras - the temporal discrepancy between the actual occurrence of changes in the corresponding timestamp assigned by the sensor. Focusing on event-guided deblurring and frame interpolation tasks, we propose a latency correction method based on a parameterized latency model. To enable data-driven learning, we develop an event-based temporal fidelity to describe the sharpness of latent images reconstructed from events and the corresponding blurry images, and reformulate the event-based double integral model differentiable to latency. The proposed method is validated using synthetic and real-world datasets, demonstrating the benefits of latency correction for deblurring and interpolation across different lighting conditions.
Yixin Yang 0008, Jinxiu Liang, Bohan Yu, Jimmy S. J. Ren, Boxin Shi
CVPR6
2024 EventPS: Real-Time Photometric Stereo Using an Event Camera
abstract
Photometric stereo is a well-established technique to es-timate the surface normal of an object. However, the re-quirement of capturing multiple high dynamic range images under different illumination conditions limits the speed and real-time applications. This paper introduces EventPS, a novel approach to real-time photometric stereo using an event camera. Capitalizing on the exceptional temporal resolution, dynamic range, and low bandwidth character-istics of event cameras, EventPS estimates surface nor-mal only from the radiance changes, significantly enhancing data efficiency. EventPS seamlessly integrates with both optimization-based and deep-learning-based photo-metric stereo techniques to offer a robust solution for non-Lambertian surfaces. Extensive experiments validate the effectiveness and efficiency of EventPS compared to frame-based counterparts. Our algorithm runs at over 30 fps in real-world scenarios, unleashing the potential of EventPS in time-sensitive and high-speed downstream applications.11Code available: https://codeberg.org/ybh1998/EventPS
Bohan Yu, Jieji Ren, Jin Han 0001, Feishi Wang, Jinxiu Liang, Boxin Shi
CVPR6
2024 Language-guided Image Reflection Separation
abstract
This paper studies the problem of language-guided re-flection separation, which aims at addressing the ill-posed reflection separation problem by introducing language de-scriptions to provide layer content. We propose a unified framework to solve this problem, which leverages the cross-attention mechanism with contrastive learning strategies to construct the correspondence between language descriptions and image layers. A gated network design and a ran-domized training strategy are employed to tackle the rec-ognizable layer ambiguity. The effectiveness of the pro-posed method is validated by the significant performance advantage over existing reflection separation methods on both quantitative and qualitative comparisons.
Haofeng Zhong, Yuchen Hong, Shuchen Weng, Jinxiu Liang, Boxin Shi
CVPR5
2024 EvDiG: Event-guided Direct and Global Components Separation
abstract
Separating the direct and global components of a scene aids in shape recovery and basic material understanding. Conventional methods capture multiple frames under high frequency illumination patterns or shadows, requiring the scene to keep stationary during the image acquisition process. Single-frame methods simplify the capture procedure but yield lower-quality separation results. In this paper, we leverage the event camera to facilitate the separation of direct and global components, enabling video-rate separation of high quality. In detail, we adopt an event camera to record rapid illumination changes caused by the shadow of a line occluder sweeping over the scene, and reconstruct the coarse separation results through event accumulation. We then design a network to resolve the noise in the coarse sep-aration results and restore color information. A real-world dataset is collected using a hybrid camera system for network training and evaluation. Experimental results show superior performance over state-of-the-art methods.
Peiqi Duan 0002, Chu Zhou, Chao Xu 0006, Boxin Shi
CVPR6
2024 L-DiffER: Single Image Reflection Removal with Language-Based Diffusion Model
Yuchen Hong, Haofeng Zhong, Shuchen Weng, Jinxiu Liang, Boxin Shi
ECCV (20)5
2024 Imaging Interiors: An Implicit Solution to Electromagnetic Inverse Scattering Problems
Ziyuan Luo, Boxin Shi, Haoliang Li, Renjie Wan
ECCV (7)2
2024 Real-Data-Driven 2000 FPS Color Video from Mosaicked Chromatic Spikes
Zhaojun Huang, Yakun Chang, Bin Fan 0002, Zhaofei Yu, Boxin Shi
ECCV (12)6
2024 Color4E: Event Demosaicing for Full-color Event Guided Image Deblurring
abstract
Neuromorphic event sensors are novel visual cameras that feature high-speed illumination-variation sensing and have found widespread application in guiding frame-based imaging enhancement. This paper focuses on color restoration in the event-guided image deblurring task, we fuse blurry images with mosaic color events instead of mono events to avoid artifacts such as color bleeding. The challenges associated with this approach include demosaicing color events for reconstructing full-resolution sampled signals and fusing bimodal signals to achieve image deblurring. To meet these challenges, we propose a novel network called Color4E to enhance the color restoration quality for the image deblurring task. Color4E leverages an event demosaicing module to upsample the spatial resolution of mosaic color events and a cross-encoding image deblurring module for fusing bimodal signals, a refinement module is designed to fuse full-color events and refine initial deblurred images. Furthermore, to avoid the real-simulated gap of events, we implement a display-filter-camera system that enables mosaic and full-color event data captured synchronously, to collect a real-captured dataset used for network training and validation. The results on the public dataset and our collected dataset show that Color4E enables high-quality event-based image deblurring compared to state-of-the-art methods.
Yi Ma 0001, Peiqi Duan 0002, Yuchen Hong, Chu Zhou, Yu Zhang 0035, Jimmy S. J. Ren, Boxin Shi
ACM Multimedia7
2024 Spatio-Temporal Interactive Learning for Efficient Image Reconstruction of Spiking Cameras
abstract
The spiking camera is an emerging neuromorphic vision sensor that records high-speed motion scenes by asynchronously firing continuous binary spike streams. Prevailing image reconstruction methods, generating intermediate frames from these spike streams, often rely on complex step-by-step network architectures that overlook the intrinsic collaboration of spatio-temporal complementary information. In this paper, we propose an efficient spatio-temporal interactive reconstruction network to jointly perform inter-frame feature alignment and intra-frame feature filtering in a coarse-to-fine manner. Specifically, it starts by extracting hierarchical features from a concise hybrid spike representation, then refines the motion fields and target frames scale-by-scale, ultimately obtaining a full-resolution output. Meanwhile, we introduce a symmetric interactive attention block and a multi-motion field estimation block to further enhance the interaction capability of the overall network. Experiments on synthetic and real-captured data show that our approach exhibits excellent performance while maintaining low model complexity.
Bin Fan 0002, Jiaoyang Yin, Yuchao Dai, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
NeurIPS6
2024 Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image Domain
abstract
Event-intensity asymmetric stereo systems have emerged as a promising approach for robust 3D perception in dynamic and challenging environments by integrating event cameras with frame-based sensors in different views. However, existing methods often suffer from overfitting and poor generalization due to limited dataset sizes and lack of scene diversity in the event domain. To address these issues, we propose a zero-shot framework that utilizes monocular depth estimation and stereo matching models pretrained on diverse image datasets. Our approach introduces a visual prompting technique to align the representations of frames and events, allowing the use of off-the-shelf stereo models without additional training. Furthermore, we introduce a monocular cue-guided disparity refinement module to improve robustness across static and dynamic regions by incorporating monocular depth information from foundation models. Extensive experiments on real-world datasets demonstrate the superior zero-shot evaluation performance and enhanced generalization ability of our method compared to existing approaches.
Hanyue Lou, Jinxiu Liang, Minggui Teng, Bin Fan 0002, Yong Xu 0007, Boxin Shi
NeurIPS6
2024 SfPUEL: Shape from Polarization under Unknown Environment Light
abstract
Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment light. To handle this challenging light condition, we design a transformer-based framework for enhancing the perception of global context features. We further propose to integrate photometric stereo (PS) priors from pretrained models to enrich extracted features for high-quality normal predictions. As metallic and dielectric materials exhibit different BRDFs, SfPUEL additionally predicts dielectric and metallic material segmentation to further boost performance. Experimental results on synthetic and our collected real-world dataset demonstrate that SfPUEL significantly outperforms existing SfP and single-shot normal estimation methods. The code and dataset is available at https://github.com/YouweiLyu/SfPUEL.
Youwei Lyu, Heng Guo 0003, Kailong Zhang, Si Li 0001, Boxin Shi
NeurIPS5
2024 Quality-Improved and Property-Preserved Polarimetric Imaging via Complementarily Fusing
abstract
Polarimetric imaging is a challenging problem in the field of polarization-based vision, since setting a short exposure time reduces the signal-to-noise ratio, making the degree of polarization (DoP) and the angle of polarization (AoP) severely degenerated, while if setting a relatively long exposure time, the DoP and AoP would tend to be over-smoothed due to the frequently-occurring motion blur. This work proposes a polarimetric imaging framework that can produce clean and clear polarized snapshots by complementarily fusing a degraded pair of noisy and blurry ones. By adopting a neural network-based three-phase fusing scheme with specially-designed modules tailored to each phase, our framework can not only improve the image quality but also preserve the polarization properties. Experimental results show that our framework achieves state-of-the-art performance.
Chu Zhou, Boxin Shi
NeurIPS4
2024 Light Flickering Guided Reflection Removal
Yuchen Hong, Yakun Chang, Jinxiu Liang, Lei Ma 0008, Tiejun Huang 0001, Boxin Shi
Int. J. Comput. Vis.6
2024 SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face Intrinsics
abstract
This paper proposes a novel pipeline to estimate a non-parametric environment map with high dynamic range from a single human face image. Lighting-independent and -dependent intrinsic images of the face are first estimated separately in a cascaded network. The influence of face geometry on the two lighting-dependent intrinsics, diffuse shading and specular reflection, are further eliminated by distributing the intrinsics pixel-wise onto spherical representations using the surface normal as indices. This results in two representations simulating images of a diffuse sphere and a glossy sphere under the input scene lighting. Taking into account the distinctive nature of light sources and ambient terms, we further introduce a two-stage lighting estimator to predict both accurate and realistic lighting from these two representations. Our model is trained supervisedly on a large-scale and high-quality synthetic face image dataset. We demonstrate that our method allows accurate and detailed lighting estimation and intrinsic decomposition, outperforming state-of-the-art methods both qualitatively and quantitatively on real face images.
Yean Cheng, Yongjie Zhu, Si Li 0001, Gang Pan 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 EvHandPose: Event-Based 3D Hand Pose Estimation With Sparse Supervision
abstract
Event camera shows great potential in 3D hand pose estimation, especially addressing the challenges of fast motion and high dynamic range in a low-power way. However, due to the asynchronous differential imaging mechanism, it is challenging to design event representation to encode hand motion information especially when the hands are not moving (causing motion ambiguity), and it is infeasible to fully annotate the temporally dense event stream. In this paper, we propose EvHandPose with novel hand flow representations in Event-to-Pose module for accurate hand pose estimation and alleviating the motion ambiguity issue. To solve the problem under sparse annotation, we design contrast maximization and hand-edge constraints in Pose-to-IWE (Image with Warped Events) module and formulate EvHandPose in a weakly-supervision framework. We further build EvRealHands, the first large-scale real-world event-based hand pose dataset on several challenging scenes to bridge the real-synthetic domain gap. Experiments on EvRealHands demonstrate that EvHandPose outperforms previous event-based methods under all evaluation scenes, achieves accurate and stable hand pose estimation with high temporal resolution in fast motion and strong light scenes compared with RGB-based methods, generalizes well to outdoor scenes and another type of event camera, and shows the potential for the hand gesture recognition task.
Jiahe Li 0006, Baowen Zhang, Xiaoming Deng 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Deep Learning Methods for Calibrated Photometric Stereo and Beyond
abstract
Photometric stereo recovers the surface normals of an object from multiple images with varying shading cues, i.e., modeling the relationship between surface orientation and intensity at each pixel. Photometric stereo prevails in superior per-pixel resolution and fine reconstruction details. However, it is a complicated problem because of the non-linear relationship caused by non-Lambertian surface reflectance. Recently, various deep learning methods have shown a powerful ability in the context of photometric stereo against non-Lambertian surfaces. This paper provides a comprehensive review of existing deep learning-based calibrated photometric stereo methods utilizing orthographic cameras and directional light sources. We first analyze these methods from different perspectives, including input processing, supervision, and network architecture. We summarize the performance of deep learning photometric stereo models on the most widely-used benchmark data set. This demonstrates the advanced performance of deep learning-based photometric stereo methods. Finally, we give suggestions and propose future research trends based on the limitations of existing models.
Yakun Ju, Kin-Man Lam 0001, Wuyuan Xie, Huiyu Zhou 0001, Junyu Dong, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Hybrid All-in-Focus Imaging From Neuromorphic Focal Stack
abstract
Creating an image focal stack requires multiple shots, which captures images at different depths within the same scene. Such methods are not suitable for scenes undergoing continuous changes. Achieving an all-in-focus image from a single shot poses significant challenges, due to the highly ill-posed nature of rectifying defocus and deblurring from a single image. In this paper, to restore an all-in-focus image, we introduce the neuromorphic focal stack, which is defined as neuromorphic signal streams captured by an event/ a spike camera during a continuous focal sweep, aiming to restore an all-in-focus image. Given an RGB image focused at any distance, we harness the high temporal resolution of neuromorphic signal streams. From neuromorphic signal streams, we automatically select refocusing timestamps and reconstruct corresponding refocused images to form a focal stack. Guided by the neuromorphic signal around the selected timestamps, we can merge the focal stack using proper weights and restore a sharp all-in-focus image. We test our method on two distinct neuromorphic cameras. Experimental results from both synthetic and real datasets demonstrate a marked improvement over existing State-of-the-Art methods.
Minggui Teng, Hanyue Lou, Yixin Yang 0008, Tiejun Huang 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Conditional Image Repainting
abstract
A number of advanced image editing technologies have demonstrated impressive performance in synthesizing visually pleasing results in accordance with user instructions. In this paper, we further extend the practicalities of image editing technology by proposing the conditional image repainting (CIR) task, which requires the model to synthesize realistic visual content based on multiple cross-modality conditions provided by the user. We first define condition inputs and formulate two-phased CIR models as the baseline. After that, we further design unified CIR models with novel condition fusion modules to improve the performance. For allowing users to express their intent more freely, our CIR models support both attributes and language to represent colors of repainted visual content. We demonstrate the effectiveness of CIR models by collecting and processing four datasets. Finally, we present a number of practical application scenarios of CIR models to demonstrate its usability.
Shuchen Weng, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Unified Video Reconstruction for Rolling Shutter and Global Shutter Cameras
abstract
Currently, the general domain of video reconstruction (VR) is fragmented into different shutters spanning global shutter and rolling shutter cameras. Despite rapid progress in the state-of-the-art, existing methods overwhelmingly follow shutter-specific paradigms and cannot conceptually generalize to other shutter types, hindering the uniformity of VR models. In this paper, we propose UniVR, a versatile framework to handle various shutters through unified modeling and shared parameters. Specifically, UniVR encodes diverse shutter types into a unified space via a tractable shutter adapter, which is parameter-free and thus can be seamlessly delivered to current well-established VR architectures for cross-shutter transfer. To demonstrate its effectiveness, we conceptualize UniVR as three shutter-generic VR methods, namely Uni-SoftSplat, Uni-SuperSloMo, and Uni-RIFE. Extensive experimental results demonstrate that the pre-trained model without any fine-tuning can achieve reasonable performance even on novel shutters. After fine-tuning, new state-of-the-art performances are established that go beyond shutter-specific methods and enjoy strong generalization. The code is available at https://github.com/GitCVfb/UniVR.
Bin Fan 0002, Zhexiong Wan, Boxin Shi, Chao Xu 0006, Yuchao Dai
IEEE Trans. Image Process.3
2024 MOUNT: Learning 6DoF Motion Prediction Based on Uncertainty Estimation for Delayed AR Rendering
abstract
The delay of rendering on AR devices requires prediction of head motion using sensor data acquired tens of even one hundred milliseconds ago to avoid misalignment between the virtual content and the physical world, where the misalignment will lead to a sense of time latency and dizziness for users. To solve the problem, we propose a method for the 6DoF motion prediction to compensate for the time latency. Compared with traditional hand-crafted methods, our method is based on deep learning, which has better motion prediction ability to deal with complex human motion. In particular, we propose a MOtion UNcerTainty encode decode network (MOUNT) that estimates the uncertainty of input data and predicts the uncertainty of output motion to improve the prediction accuracy and smoothness. Experiments on the EuRoC and our collected dataset demonstrate that our method significantly outperforms the traditional method and greatly improves AR visual effects.
Haoran Chen 0010, Lantian Wei, Haomin Liu, Boxin Shi, Guofeng Zhang 0001, Hongbin Zha
IEEE Trans. Vis. Comput. Graph.4
2024 GR-PSN: Learning to Estimate Surface Normal and Reconstruct Photometric Stereo Images
abstract
In this paper, we propose a novel method, namely GR-PSN, which learns surface normals from photometric stereo images and generates the photometric images under distant illumination from different lighting directions and surface materials. The framework is composed of two subnetworks, named GeometryNet and ReconstructNet, which are cascaded to perform shape reconstruction and image rendering in an end-to-end manner. ReconstructNet introduces additional supervision for surface-normal recovery, forming a closed-loop structure with GeometryNet. We also encode lighting and surface reflectance in ReconstructNet, to achieve arbitrary rendering. In training, we set up a parallel framework to simultaneously learn two arbitrary materials for an object, providing an additional transform loss. Therefore, our method is trained based on the supervision by three different loss functions, namely the surface-normal loss, reconstruction loss, and transform loss. We alternately input the predicted surface-normal map and the ground-truth into ReconstructNet, to achieve stable training for ReconstructNet. Experiments show that our method can accurately recover the surface normals of an object with an arbitrary number of inputs, and can re-render images of the object with arbitrary surface materials. Extensive experimental results show that our proposed method outperforms those methods based on a single surface recovery network and shows realistic rendering results on 100 different materials.
Yakun Ju, Boxin Shi, Yang Chen 0036, Huiyu Zhou 0001, Junyu Dong, Kin-Man Lam 0001
IEEE Trans. Vis. Comput. Graph.2
2023 Polarization-Aware Low-Light Image Enhancement
abstract
Polarization-based vision algorithms have found uses in various applications since polarization provides additional physical constraints. However, in low-light conditions, their performance would be severely degenerated since the captured polarized images could be noisy, leading to noticeable degradation in the degree of polarization (DoP) and the angle of polarization (AoP). Existing low-light image enhancement methods cannot handle the polarized images well since they operate in the intensity domain, without effectively exploiting the information provided by polarization. In this paper, we propose a Stokes-domain enhancement pipeline along with a dual-branch neural network to handle the problem in a polarization-aware manner. Two application scenarios (reflection removal and shape from polarization) are presented to show how our enhancement can improve their results.
Chu Zhou, Minggui Teng, Youwei Lyu, Si Li 0001, Chao Xu 0006, Boxin Shi
AAAI6
2023 L-CoIns: Language-based Colorization With Instance Awareness
abstract
Language-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the same object words. In this paper, we propose a transformer-based framework to automatically aggregate similar image patches and achieve instance awareness without any additional knowledge. By applying our presented luminance augmentation and counter-color loss to break down the statistical correlation between luminance and color words, our model is driven to synthesize colors with better descriptive consistency. We further collect a dataset to provide distinctive visual characteristics and detailed language descriptions for multiple instances in the same image. Extensive experiments demonstrate our advantages of synthesizing visually pleasing and description-consistent results of instance-aware colorization.
Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi
CVPR6
2023 1000 FPS HDR Video with a Spike-RGB Hybrid Camera
abstract
Capturing high frame rate and high dynamic range (HFR&HDR) color videos in high-speed scenes with conventional frame-based cameras is very challenging. The increasing frame rate is usually guaranteed by using shorter exposure time so that the captured video is severely interfered by noise. Alternating exposures can alleviate the noise issue but sacrifice frame rate due to involving long-exposure frames. The neuromorphic spiking camera records high-speed scenes of high dynamic range without colors using a completely different sensing mechanism and visual representation. We introduce a hybrid camera system composed of a spiking and an alternating-exposure RGB camera to capture HFR&HDR scenes with high fidelity. Our insight is to bring each camera's superiority into full play. The spike frames, with accurate fast motion information encoded, are firstly reconstructed for motion representation, from which the spike-based optical flows guide the recovery of missing temporal information for long-exposure RGB images while retaining their reliable color appearances. With the strong temporal constraint estimated from spike trains, both missing and distorted colors cross RGB frames are recovered to generate time-consistent and HFR color frames. We collect a new Spike-RGB dataset that contains 300 sequences of synthetic data and 20 groups of real-world data to demonstrate 1000 FPS HDR videos outperforming HDR video reconstruction methods and commercial high-speed cameras.
Yakun Chang, Chu Zhou, Yuchen Hong, Liwen Hu 0002, Chao Xu 0002, Tiejun Huang 0001, Boxin Shi
CVPR7
2023 High-fidelity Event-Radiance Recovery via Transient Event Frequency
abstract
High-fidelity radiance recovery plays a crucial role in scene information reconstruction and understanding. Conventional cameras suffer from limited sensitivity in dynamic range, bit depth, and spectral response, etc. In this paper, we propose to use event cameras with bio-inspired silicon sensors, which are sensitive to radiance changes, to recover precise radiance values. We reveal that, under active lighting conditions, the transient frequency of event signals triggering linearly reflects the radiance value. We propose an innovative method to convert the high temporal resolution of event signals into precise radiance values. The precise radiance values yields several capabilities in image analysis. We demonstrate the feasibility of recovering radiance values solely from the transient event frequency (TEF) through multiple experiments.
Jin Han 0001, Yuta Asano, Boxin Shi, Yinqiang Zheng, Imari Sato
CVPR3
2023 DANI-Net: Uncalibrated Photometric Stereo by Differentiable Shadow Handling, Anisotropic Reflectance Modeling, and Neural Inverse Rendering
abstract
Uncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by the unknown light. Although the ambiguity is alleviated on non-Lambertian objects, the problem is still difficult to solve for more general objects with complex shapes introducing irregular shadows and general materials with complex reflectance like anisotropic reflectance. To exploit cues from shadow and reflectance to solve UPS and improve performance on general materials, we propose DANI-Net, an inverse rendering framework with differentiable shadow handling and anisotropic reflectance modeling. Unlike most previous methods that use non-differentiable shadow maps and assume isotropic material, our network benefits from cues of shadow and anisotropic reflectance through two differentiable paths. Experiments on multiple real-world datasets demonstrate our superior and robust performance.
Zongrui Li 0001, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001
CVPR3
2023 All-in-Focus Imaging from Event Focal Stack
abstract
Traditional focal stack methods require multiple shots to capture images focused at different distances of the same scene, which cannot be applied to dynamic scenes well. Generating a high-quality all-in-focus image from a single shot is challenging, due to the highly ill-posed nature of the single-image defocus and deblurring problem. In this paper, to restore an all-in-focus image, we propose the event focal stack which is defined as event streams captured during a continuous focal sweep. Given an RGB image focused at an arbitrary distance, we explore the high temporal resolution of event streams, from which we automatically select refocusing timestamps and reconstruct corresponding refocused images with events to form a focal stack. Guided by the neighbouring events around the selected timestamps, we can merge the focal stack with proper weights and restore a sharp all-in-focus image. Experimental results on both synthetic and real datasets show superior performance over state-of-the-art methods.
Hanyue Lou, Minggui Teng, Yixin Yang 0008, Boxin Shi
CVPR4
2023 Complementary Intrinsics from Neural Radiance Fields and CNNs for Outdoor Scene Relighting
abstract
Relighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multiview stereo. With neural radiance fields (NeRF), editing the appearance code could produce more realistic results without interpreting the outdoor scene image formation explicitly. This paper proposes to complement the intrinsic estimation from volume rendering using NeRF and from inversing the photometric image formation model using convolutional neural networks (CNNs). The former produces richer and more reliable pseudo labels (cast shadows and sky appearances in addition to albedo and normal) for training the latter to predict interpretable and editable lighting parameters via a single-image prediction pipeline. We demonstrate the advantages of our method for both intrinsic image decomposition and relighting for various real outdoor scenes.
Xuanning Cui, Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Zhaofei Yu, Boxin Shi
CVPR7
2023 Learning Event Guided High Dynamic Range Video Reconstruction
abstract
Limited by the trade-off between frame rate and exposure time when capturing moving scenes with conventional cameras, frame based HDR video reconstruction suffers from scene-dependent exposure ratio balancing and ghosting artifacts. Event cameras provide an alternative visual representation with a much higher dynamic range and temporal resolution free from the above issues, which could be an effective guidance for HDR imaging from LDR videos. In this paper, we propose a multimodal learning framework for event guided HDR video reconstruction. In order to better leverage the knowledge of the same scene from the two modalities of visual signals, a multimodal representation alignment strategy to learn a shared latent space and a fusion module tailored to complementing two types of signals for different dynamic ranges in different regions are proposed. Temporal correlations are utilized recurrently to suppress the flickering effects in the reconstructed HDR video. The proposed HDRev-Net demonstrates state-of-the-art performance quantitatively and qualitatively for both synthetic and real-world data.
Yixin Yang 0008, Jin Han 0001, Jinxiu Liang, Imari Sato, Boxin Shi
CVPR5
2023 Occlusion-Free Scene Recovery via Neural Radiance Fields
abstract
Our everyday lives are filled with occlusions that we strive to see through. By aggregating desired background information from different viewpoints, we can easily eliminate such occlusions without any external occlusion-free supervision. Though several occlusion removal methods have been proposed to empower machine vision systems with such ability, their performances are still unsatisfactory due to reliance on external supervision. We propose a novel method for occlusion removal by directly building a mapping between position and viewing angles and the corresponding occlusion-free scene details leveraging Neural Radiance Fields (NeRF). We also develop an effective scheme to jointly optimize camera parameters and scene reconstruction when occlusions are present. An additional depth constraint is applied to supervise the entire optimizaion without labeled external data for training. The experimental results on existing and newly collected datasets validate the effectiveness of our method. Our project page: https://freebutuselesssoul.github.io/occnerf.
Chengxuan Zhu, Renjie Wan, Yunkai Tang, Boxin Shi
CVPR4
2023 ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric Stereo
abstract
Illumination planning in photometric stereo aims to find a balance between surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illumination, and the choice of photometric stereo backbones, which are too complex to be handled by existing methods based on handcrafted illumination planning rules. This paper proposes a learning-based illumination planning method that jointly considers these factors via integrating a neural network and a generalized image formation model. As it is impractical to supervise illumination planning due to the enormous search space for ground truth light configurations, we formulate illumination planning using reinforcement learning, which explores the light space in a photometric stereo-aware and reward-driven manner. Experiments on synthetic and real-world datasets demonstrate that photometric stereo under the 20-light configurations from our method is comparable to, or even surpasses that of using lights from all available directions.
Jun Hoong Chan, Bohan Yu, Heng Guo 0003, Jieji Ren, Zongqing Lu 0002, Boxin Shi
ICCV6
2023 Coherent Event Guided Low-Light Video Enhancement
abstract
With frame-based cameras, capturing fast-moving scenes without suffering from blur often comes at the cost of low SNR and low contrast. Worse still, the photometric constancy that enhancement techniques heavily relied on is fragile for frames with short exposure. Event cameras can record brightness changes at an extremely high temporal resolution. For low-light videos, event data are not only suitable to help capture temporal correspondences but also provide alternative observations in the form of intensity ratios between consecutive frames and exposure-invariant information. Motivated by this, we propose a low-light video enhancement method with hybrid inputs of events and frames. Specifically, a neural network is trained to establish spatiotemporal coherence between visual signals with different modalities and resolutions by constructing correlation volume across space and time. Experimental results on synthetic and real data demonstrate the superiority of the proposed method compared to the state-of-the-art methods.
Jinxiu Liang, Yixin Yang 0008, Peiqi Duan 0002, Yong Xu 0007, Boxin Shi
ICCV6
2023 DiLiGenT-Π: Photometric Stereo for Planar Surfaces with Rich Details - Benchmark Dataset and Beyond
abstract
Photometric stereo aims to recover detailed surface shapes from images captured under varying illuminations. However, existing real-world datasets primarily focus on evaluating photometric stereo for general non-Lambertian reflectances and feature bulgy shapes that have a certain height. As shape detail recovery is the key strength of photometric stereo over other 3D reconstruction techniques, and the near-planar surfaces widely exist in cultural relics and manufacturing workpieces, we present a new real-world dataset DiLiGenT-Π containing 30 near-planar scenes with rich surface details. This dataset enables us to evaluate recent photometric stereo methods specifically for their ability to estimate shape details under diverse materials and to identify open problems such as near-planar surface normal estimation from uncalibrated photometric stereo and surface detail recovery for translucent materials. To inspire future research, this dataset will open soruced at https://photometricstereo.github.io/diligentpi.html.
Feishi Wang, Jieji Ren, Heng Guo 0003, Ming Jun Ren, Boxin Shi
ICCV5
2023 Affective Image Filter: Reflecting Emotions from Text to Images
abstract
Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the visually-abstract emotions from the text and reflect them to visually-concrete images with appropriate colors and textures. We build our model based on the multi-modal transformer architecture, which unifies both images and texts into tokens and encodes the emotional prior knowledge. Various loss functions are proposed to understand complex emotions and produce appropriate visualization. In addition, we collect and contribute a new dataset with abundant aesthetic images and emotional texts for training and evaluating the AIF model. We carefully design four quantitative metrics and conduct a user study to comprehensively evaluate the performance, which demonstrates our AIF model outperforms state-of-the-art methods and could evoke specific emotional responses from human observers.
Shuchen Weng, Peixuan Zhang, Si Li 0001, Boxin Shi
ICCV6
2023 Non-Lambertian Multispectral Photometric Stereo via Spectral Reflectance Decomposition
abstract
Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multispectral image captured under multispectral illuminations. Existing MPS methods adopt the Lambertian reflectance model to make the problem tractable, but it greatly limits their application to real-world surfaces. In this paper, we propose a deep neural network named NeuralMPS to solve the MPS problem under non-Lambertian spectral reflectances. Specifically, we present a spectral reflectance decomposition model to disentangle the spectral reflectance into a geometric component and a spectral component. With this decomposition, we show that the MPS problem for surfaces with a uniform material is equivalent to the conventional photometric stereo (CPS) with unknown light intensities. In this way, NeuralMPS reduces the difficulty of the non-Lambertian MPS problem by leveraging the well-studied non-Lambertian CPS methods. Experiments on both synthetic and real-world scenes demonstrate the effectiveness of our method.
Jipeng Lv, Heng Guo 0003, Guanying Chen, Jinxiu Liang, Boxin Shi
IJCAI5
2023 L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors
abstract
Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal performance. In this paper, we propose a unified model to perform language-based colorization with any-level descriptions. We leverage the pretrained cross-modality generative model for its robust language understanding and rich color priors to handle the inherent ambiguity of any-level descriptions. We further design modules to align with input conditions to preserve local spatial structures and prevent the ghosting effect. With the proposed novel sampling strategy, our model achieves instance-aware colorization in diverse and complex scenarios. Extensive experimental results demonstrate our advantages of effectively handling any-level descriptions and outperforming both language-based and automatic colorization methods. The code and pretrained models are available at: https://github.com/changzheng123/L-CAD.
Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi
NeurIPS6
2023 Slow and Weak Attractor Computation Embedded in Fast and Strong E-I Balanced Neural Dynamics
abstract
Attractor networks require neuronal connections to be highly structured in order to maintain attractor states that represent information, while excitation and inhibition balanced networks (E-INNs) require neuronal connections to be random and sparse to generate irregular neuronal firings. Despite being regarded as canonical models of neural circuits, both types of networks are usually studied in isolation, and it remains unclear how they coexist in the brain, given their very different structural demands. In this study, we investigate the compatibility of continuous attractor neural networks (CANNs) and E-INNs. In line with recent experimental data, we find that a neural circuit can exhibit both the traits of CANNs and E-INNs if the neuronal synapses consist of two sets: one set is strong and fast for irregular firing, and the other set is weak and slow for attractor dynamics. Our results from simulations and theoretical analysis reveal that the network also exhibits enhanced performance compared to the case of using only one set of synapses, with accelerated convergence of attractor states and retained E-I balanced condition for localized input. We also apply the network model to solve a real-world tracking problem and demonstrate that it can track fast-moving objects well. We hope that this study provides insight into how structured neural computations are realized by irregular firings of neurons.
Xiaohan Lin, Liyuan Li, Boxin Shi, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001
NeurIPS3
2023 LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic Generation
abstract
We present the ilLumination-Aware conditional Image Repainting (LuminAIRe) task to address the unrealistic lighting effects in recent conditional image repainting (CIR) methods. The environment lighting and 3D geometry conditions are explicitly estimated from given background images and parsing masks using a parametric lighting representation and learning-based priors. These 3D conditions are then converted into illumination images through the proposed physically-based illumination rendering and illumination attention module. With the injection of illumination images, physically-correct lighting information is fed into the lighting-realistic generation process and repainted images with harmonized lighting effects in both foreground and background regions can be acquired, whose superiority over the results of state-of-the-art methods is confirmed through extensive experiments. For facilitating and validating the LuminAIRe task, a new dataset Car-LuminAIRe with lighting annotations and rich appearance variants is collected.
Jiajun Tang 0001, Haofeng Zhong, Shuchen Weng, Boxin Shi
NeurIPS4
2023 Deblurring Low-Light Images with Events
Chu Zhou, Minggui Teng, Jin Han 0001, Jinxiu Liang, Chao Xu 0006, Boxin Shi
Int. J. Comput. Vis.7
2023 A Survey on In-Vehicle Time-Sensitive Networking
abstract
With the continued evolution of autonomous driving, the communication demand for in-vehicle networks (IVNs) has dramatically increased, and the traditional IVN solutions have become ineligible for the needs of new in-vehicle applications. Time-sensitive network (TSN) with the characteristics of determinism, low latency, high reliability, large bandwidth, and open industry standards, has been considered as the most promising technology for the next-generation IVNs. This article summarizes the current state and methodologies of in-vehicle TSN research through the following five aspects: 1) Quality of Service (QoS) strategy; 2) reliability; 3) clock synchronization; 4) network planning; and 5) network management. The corresponding research priorities and technical challenges are also addressed, respectively.
Yifei Peng, Boxin Shi, Tigang Jiang, Xiaodong Tu, Du Xu, Kun Hua
IEEE Internet Things J.2
2023 NeuroZoom: Denoising and Super Resolving Neuromorphic Events and Spikes
abstract
Neuromorphic cameras are emerging imaging technology that has advantages over conventional imaging sensors in several aspects including dynamic range, sensing latency, and power consumption. However, the signal-to-noise level and the spatial resolution still fall behind the state of conventional imaging sensors. In this article, we address the denoising and super-resolution problem for modern neuromorphic cameras. We employ 3D U-Net as the backbone neural architecture for such a task. The networks are trained and tested on two types of neuromorphic cameras: a dynamic vision sensor and a spike camera. Their pixels generate signals asynchronously, the former is based on perceived light changes and the latter is based on accumulated light intensity. To collect the datasets for training such networks, we design a display-camera system to record high frame-rate videos at multiple resolutions, providing supervision for denoising and super-resolution. The networks are trained in a noise-to-noise fashion, where the two ends of the network are unfiltered noisy data. The output of the networks has been tested for downstream applications including event-based visual object tracking and image reconstruction. Experimental results demonstrate the effectiveness of improving the quality of neuromorphic events and spikes, and the corresponding improvement to downstream applications with state-of-the-art performance.
Peiqi Duan 0002, Yi Ma 0001, Xinyu Shi 0004, Zihao W. Wang, Tiejun Huang 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Hybrid High Dynamic Range Imaging fusing Neuromorphic and Conventional Images
abstract
Reconstruction of high dynamic range image from a single low dynamic range image captured by a conventional RGB camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras like event camera and spike camera can record high dynamic range scenes in the form of intensity maps, but with much lower spatial resolution and no color information. In this article, we propose a hybrid imaging system (denoted as NeurImg) that captures and fuses the visual information from a neuromorphic camera and ordinary images from an RGB camera to reconstruct high-quality high dynamic range images and videos. The proposed NeurImg-HDR+ network consists of specially designed modules, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images to reconstruct high-resolution, high dynamic range images and videos. We capture a test dataset of hybrid signals on various HDR scenes using the hybrid camera, and analyze the advantages of the proposed fusing strategy by comparing it to state-of-the-art inverse tone mapping methods and merging two low dynamic range images approaches. Quantitative and qualitative experiments on both synthetic data and real-world scenarios demonstrate the effectiveness of the proposed hybrid high dynamic range imaging system. Code and dataset can be found at: https://github.com/hjynwa/NeurImg-HDR.
Jin Han 0001, Yixin Yang 0008, Peiqi Duan 0002, Chu Zhou, Lei Ma 0008, Chao Xu 0006, Tiejun Huang 0001, Imari Sato, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.9
2023 PAR$^{2}$2Net: End-to-End Panoramic Image Reflection Removal
abstract
In this article, we investigate the problem of panoramic image reflection removal to relieve the content ambiguity between the reflection layer and the transmission scene. Although a partial view of the reflection scene is attainable in the panoramic image and provides additional information for reflection removal, it is not trivial to directly apply this for getting rid of undesired reflections due to its misalignment with the reflection-contaminated image. We propose an end-to-end framework to tackle this problem. By resolving misalignment issues with adaptive modules, the high-fidelity recovery of reflection layer and transmission scenes is accomplished. We further propose a new data generation approach that considers the physics-based formation model of mixture images and the in-camera dynamic range clipping to diminish the domain gap between synthetic and real data. Experimental results demonstrate the effectiveness of the proposed method and its applicability for mobile devices and industrial applications.
Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Physics-Guided Reflection Separation From a Pair of Unpolarized and Polarized Images
abstract
Undesirable reflections contained in photos taken in front of glass windows or doors often degrade visual quality of the image. Separating two layers apart benefits both human and machine perception. The polarization status of the light changes after refraction or reflection, providing more observations of the scene, which can benefit the reflection separation. Different from previous works that take three or more polarization images as input, we propose to exploit physical constraints from a pair of unpolarized and polarized images to separate reflection and transmission layers in this paper. Due to the simplified capturing setup, the system is more under-determined compared to the existing polarization-based works. In order to solve this problem, we propose to estimate the semi-reflector orientation first to make the physical image formation well-posed, and then learn to reliably separate two layers using additional networks based on both physical and numerical analysis. In addition, a motion estimation network is introduced to handle the misalignment of paired input. Quantitative and qualitative experimental results show our approach performs favorably over existing polarization and single image based solutions.
Youwei Lyu, Zhaopeng Cui, Si Li 0001, Marc Pollefeys, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Shape From Polarization With Distant Lighting Estimation
abstract
This article presents a new approach for surface normal recovery from polarization images under an unknown distant light. Polarization provides rich cues of object geometry and material, but it is also influenced by different lighting conditions. Different from previous Shape-from-Polarization (SfP) methods, which rely on handcrafted or data-driven priors, we analytically investigate the benefits of estimating distant lighting for resolving the ambiguity in normal estimation from SfP using the polarimetric Bidirectional Reflectance Distribution Function (pBRDF) based image formation model. We then propose a two-stage learning framework that first effectively exploits polarization and shading cues to estimate the reflectance and lighting information and then optimizes the initial normal as the geometric prior. Leveraging the normal prior with the polarization cues from the input images, our network further generates the surface normal with more details in the second stage. We also present a data generation pipeline derived from the pBRDF model enabling model training and create a real dataset for evaluation of SfP approaches. Extensive ablation studies show the effectiveness of our designed architecture, and our approach outperforms existing methods in quantitative and qualitative experiments on real data.
Youwei Lyu, Lingran Zhao, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Benchmarking Single-Image Reflection Removal Algorithms
abstract
Reflection removal has been discussed for more than decades. This paper aims to provide the analysis for different reflection properties and factors that influence image formation, an up-to-date taxonomy for existing methods, a benchmark dataset, and the unified benchmarking evaluations for state-of-the-art (especially learning-based) methods. Specifically, this paper presents a SIngle-image Reflection Removal Plus dataset “SIR$^{2+}$” with the new consideration for in-the-wild scenarios and glass with diverse color and unplanar shapes. We further perform quantitative and visual quality comparisons for state-of-the-art single-image reflection removal algorithms. Open problems for improving reflection removal algorithms are discussed at the end. Our dataset and follow-up update can be found athttps://reflectionremoval.github.io/sir2data/.
Renjie Wan, Boxin Shi, Haoliang Li, Yuchen Hong, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Coarse-to-fine Disentangling Demoiréing Framework for Recaptured Screen Images
abstract
Removing the undesired moiré patterns from images capturing the contents displayed on screens is of increasing research interest, as the need for recording and sharing the instant information conveyed by the screens is growing. Previous demoiréing methods provide limited investigations into the formation process of moiré patterns to exploit moiré-specific priors for guiding the learning of demoiréing models. In this paper, we investigate the moiré pattern formation process from the perspective of signal aliasing, and correspondingly propose a coarse-to-fine disentangling demoiréing framework. In this framework, we first disentangle the moiré pattern layer and the clean image with alleviated ill-posedness based on the derivation of our moiré image formation model. Then we refine the demoiréing results exploiting both the frequency domain features and edge attention, considering moiré patterns' property on spectrum distribution and edge intensity revealed in our aliasing based analysis. Experiments on several datasets show that the proposed method performs favorably against state-of-the-art methods. Besides, the proposed method is validated to adapt well to different data sources and scales, especially on the high-resolution moiré images.
Ce Wang 0007, Shengsen Wu, Renjie Wan, Boxin Shi, Ling-Yu Duan
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Surface Geometry Processing: An Efficient Normal-Based Detail Representation
abstract
With the rapid development of high-resolution 3D vision applications, the traditional way of manipulating surface detail requires considerable memory and computing time. To address these problems, we introduce an efficient surface detail processing framework in 2D normal domain, which extracts new normal feature representations as the carrier of micro geometry structures that are illustrated both theoretically and empirically in this article. Compared with the existing state of the arts, we verify and demonstrate that the proposed normal-based representation has three important properties, including detail separability, detail transferability and detail idempotence. Finally, three new schemes are further designed for geometric surface detail processing applications, including geometric texture synthesis, geometry detail transfer, and 3D surface super-resolution. Theoretical analysis and experimental results on the latest benchmark dataset verify the effectiveness and versatility of our normal-based representation, which accepts 30 times of the input surface vertices but at the same time only takes 6.5% memory cost and 14.0% running time in comparison with existing competing algorithms.
Wuyuan Xie, Miaohui Wang, Di Lin 0002, Boxin Shi, Jianmin Jiang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 MILO: Multi-Bounce Inverse Rendering for Indoor Scene With Light-Emitting Objects
abstract
Recently, many advances in inverse rendering are achieved by high-dimensional lighting representations and differentiable rendering. However, multi-bounce lighting effects can hardly be handled correctly in scene editing using high-dimensional lighting representations, and light source model deviation and ambiguities exist in differentiable rendering methods. These problems limit the applications of inverse rendering. In this paper, we present a multi-bounce inverse rendering method based on Monte Carlo path tracing, to enable correct complex multi-bounce lighting effects rendering in scene editing. We propose a novel light source model that is more suitable for light source editing in indoor scenes, and design a specific neural network with corresponding disambiguation constraints to alleviate ambiguities during the inverse rendering. We evaluate our method on both synthetic and real indoor scenes through virtual object insertion, material editing, relighting tasks, and so on. The results demonstrate that our method achieves better photo-realistic quality.
Bohan Yu, Xuanning Cui, Siyan Dong, Baoquan Chen, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Polarization Guided HDR Reconstruction via Pixel-Wise Depolarization
abstract
Taking photos with digital cameras often accompanies saturated pixels due to their limited dynamic range, and it is far too ill-posed to restore them. Capturing multiple low dynamic range images with bracketed exposures can make the problem less ill-posed, however, it is prone to ghosting artifacts caused by spatial misalignment among images. A polarization camera can capture four spatially-aligned and temporally-synchronized polarized images with different polarizer angles in a single shot, which can be used for ghost-free high dynamic range (HDR) reconstruction. However, real-world scenarios are still challenging since existing polarization-based HDR reconstruction methods treat all pixels in the same manner and only utilize the spatially-variant exposures of the polarized images (without fully exploiting the degree of polarization (DoP) and the angle of polarization (AoP) of the incoming light to the sensor, which encode abundant structural and contextual information of the scene) to handle the problem still in an ill-posed manner. In this paper, we propose a pixel-wise depolarization strategy to solve the polarization guided HDR reconstruction problem, by classifying the pixels based on their levels of ill-posedness in HDR reconstruction procedure and applying different solutions to different classes. To utilize the strategy with better generalization ability and higher robustness, we propose a network-physics-hybrid polarization-based HDR reconstruction pipeline along with a neural network tailored to it, fully exploiting the DoP and AoP. Experimental results show that our approach achieves state-of-the-art performance on both synthetic and real-world images.
Chu Zhou, Yufei Han 0002, Minggui Teng, Jin Han 0001, Si Li 0001, Chao Xu 0006, Boxin Shi
IEEE Trans. Image Process.7
2023 A Residual Learning Approach to Deblur and Generate High Frame Rate Video With an Event Camera
abstract
Event cameras are bio-inspired cameras that can measure the intensity change asynchronously with high temporal resolution. One of the advantages of event cameras is that they suffer less from motion blur than traditional frame cameras when recording daily scenes with fast-moving objects. In this paper, we formulate the deblurring task on traditional cameras directed by events to be a residual learning one, and propose corresponding network architectures for effective learning of deblurring and high frame rate video generation tasks. We first train a modified U-Net network to restore a sharp image from a blurry image using the corresponding events. Then we train another similar network by replacing the downsampling blocks with blocks of the convolutional long short-term memory (Conv-LSTM) to recurrently generate high frame rate video using the restored sharp image and part of the events. Benefitting from the blur-free events and the proposed learning strategy, the experimental results show that the proposed method outperforms state-of-the-art methods for generating sharp images and high frame rate videos.
Minggui Teng, Boxin Shi, Yizhou Wang 0001, Tiejun Huang 0001
IEEE Trans. Multim.3
2023 Reflection Removal With NIR and RGB Image Feature Fusion
abstract
Removing undesirable reflections in photographs benefits both human perceptions and downstream computer vision tasks, but it is a highly ill-posed problem based on a single RGB image. Different from RGB images, near-infrared (NIR) images captured by an active NIR camera are less likely to be affected by reflections when glass and camera planes form certain angles, while textures on objects could “vanish” in some situations. Based on this observation, we propose a cascaded reflection removal network with an image feature fusion strategy to utilize auxiliary information in active NIR images. To tackle the insufficiency of training data, we propose a data generation pipeline to approximate perceptual properties and the reflection-suppressing nature of active NIR images. We further build a dataset with synthetic and real images to facilitate the research. Experimental results show that the proposed method outperforms state-of-the-art reflection removal methods in both quantitative metrics and visual quality.
Yuchen Hong, Youwei Lyu, Si Li 0001, Boxin Shi
IEEE Trans. Multim.5
2023 Purifying Low-Light Images via Near-Infrared Enlightened Image
abstract
Cameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results.
Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.2
2023 Background Scene Recovery From an Image Looking Through Colored Glass
abstract
Colored glass, which is commonly seen in modern city life, often degrades images taken through it with co-occurring reflection and color bias due to its optical property of simultaneous transmission, reflection, and wavelength-selective absorption. Recovering the clean background behind colored glass is inherently challenging due to the mutual interference of two degradations within a single mixture observation, and has barely been specifically considered by existing image restoration methods. In this paper, we aim at realizing faithful background scene recovery for an image taken in front of colored glass. We first analyze the formation model of mixed degradations caused by colored glass, and propose a cooperative framework to address the mutual interference problem, featuring a novel glass color invariant loss and progressive refinement. Besides, we propose a data synthesis strategy for network training. Experimental results on our newly collected real-world dataset show that our proposed method achieves state-of-the-art performance.
Ce Wang 0007, Dejia Xu, Renjie Wan, Boxin Shi, Ling-Yu Duan
IEEE Trans. Multim.5
2022 L-CoDe: Language-Based Colorization Using Color-Object Decoupled Conditions
abstract
Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to color difficult. In this paper, we propose L-CoDe, a Language-based Colorization network using color-object Decoupled conditions. A predictor for object-color corresponding matrix (OCCM) and a novel attention transfer module (ATM) are introduced to solve the color-object coupling problem. To deal with color-object mismatch that results in incorrect color-object correspondence, we adopt a soft-gated injection module (SIM). We further present a new dataset containing annotated color-object pairs to provide supervisory signals for resolving the coupling problem. Experimental results show that our approach outperforms state-of-the-art methods conditioned on captions.
Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi
AAAI6
2022 Optical Flow Estimation for Spiking Camera
abstract
As a bio-inspired sensor with high temporal resolution, the spiking camera has an enormous potential in real applications, especially for motion estimation in high-speed scenes. However, frame-based and event-based methods are not well suited to spike streams from the spiking camera due to the different data modalities. To this end, we present, SCFlow, a tailored deep learning pipeline to estimate optical flow in high-speed scenes from spike streams. Importantly, a novel input representation is introduced which can adaptively remove the motion blur in spike streams according to the prior motion. Further, for training SCFlow, we synthesize two sets of optical flow data for the spiking camera, SPIkingly Flying Things and Photo-realistic Highspeed Motion, denoted as SPIFT and PHM respectively, corresponding to random high-speed and well-designed scenes. Experimental results show that the SCFlow can predict optical flow from spike streams in different high-speed scenes. Moreover, SCFlow shows promising generalization on real spike streams. Codes and datasets refer to https://github.com/Acnext/Optical-Flow-For-Spiking-Camera.
Liwen Hu 0002, Rui Zhao 0010, Ziluo Ding, Lei Ma 0008, Boxin Shi, Ruiqin Xiong, Tiejun Huang 0001
CVPR5
2022 DiLiGenT102: A Photometric Stereo Benchmark Dataset with Controlled Shape and Material Variation
abstract
Evaluating photometric stereo using real-world dataset is important yet difficult. Existing datasets are insufficient due to their limited scale and random distributions in shape and material. This paper presents a new real-world photometric stereo dataset with “ground truth” normal maps, which is 10 times larger than the widely adopted one. More importantly, we propose to control the shape and material variations by fabricating objects from CAD models with carefully selected materials, covering typical aspects of reflectance properties that are distinctive for evaluating photometric stereo methods. By benchmarking recent photometric stereo methods using these 100 sets of images, with a special focus on recent learning based solutions, a 10x 10 shape-material error distribution matrix is visualized to depict a “portrait” for each evaluated method. From such comprehensive analysis, open problems in this field are discussed. To inspire future research, this dataset is available at https://photometricstereo.github.io.
Jieji Ren, Feishi Wang, Ming Jun Ren, Boxin Shi
CVPR6
2022 UniCoRN: A Unified Conditional Image Repainting Network
abstract
Conditional image repainting (CIR) is an advanced image editing task, which requires the model to generate visual content in user-specified regions conditioned on multiple cross-modality constraints, and composite the visual content with the provided background seamlessly. Existing methods based on two-phase architecture design assume dependency between phases and cause color-image incongruity. To solve these problems, we propose a novel Unified Conditional image Repainting Network (UniCoRN). We break the two-phase assumption in the CIR task by constructing the interaction and dependency relationship between background and other conditions. We further introduce the hierarchical structure into cross-modality similarity model to capture feature patterns at different levels and bridge the gap between visual content and color condition. A new Landscape-CIR dataset is collected and annotated to expand the application scenarios of the CIR task. Experiments show that UniCoRN achieves higher synthetic quality, better condition consistency, and more realistic compositing effect.
Jimeng Sun 0002, Shuchen Weng, Si Li 0001, Boxin Shi
CVPR5
2022 EvUnroll: Neuromorphic Events based Rolling Shutter Image Correction
abstract
This paper proposes to use neuromorphic events for correcting rolling shutter (RS) images as consecutive global shutter (GS) frames. RS effect introduces edge distortion and region occlusion into images caused by row-wise read-out of CMOS sensors. We introduce a novel computational imaging setup consisting of an RS sensor and an event sensor, and propose a neural network called EvUnroll to solve this problem by exploring the high-temporal-resolution property of events. We use events to bridge a spatio-temporal connection between RS and GS, establish a flow estimation module to correct edge distortions, and design a synthesis-based restoration module to restore occluded regions. The results of two branches are fused through a refining module to generate corrected GS images. We further propose datasets captured by a high-speed camera and an RS-Event hybrid camera system for training and testing our network. Experimental results on both public and proposed datasets show a systematic performance improvement compared to state-of-the-art methods.
Peiqi Duan 0002, Yi Ma 0001, Boxin Shi
CVPR4
2022 Bilateral Normal Integration
Hiroaki Santo, Boxin Shi, Fumio Okura, Yasuyuki Matsushita
ECCV (1)3
2022 L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer
Shuchen Weng, Yu Li 0003, Si Li 0001, Boxin Shi
ECCV (18)5
2022 Real-Time Intermediate Flow Estimation for Video Frame Interpolation
Zhewei Huang, Wen Heng, Boxin Shi, Shuchang Zhou 0001
ECCV (14)4
2022 Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation
Jiajun Tang 0001, Yongjie Zhu, Jun Hoong Chan, Si Li 0001, Boxin Shi
ECCV (6)6
2022 NEST: Neural Event Stack for Event-Based Image Enhancement
Minggui Teng, Chu Zhou, Hanyue Lou, Boxin Shi
ECCV (6)4
2022 CT2: Colorization Transformer via Color Tokens
Shuchen Weng, Jimeng Sun 0002, Yu Li 0003, Si Li 0001, Boxin Shi
ECCV (7)5
2022 Instance Contour Adjustment via Structure-Driven CNN
Shuchen Weng, Ming-Ching Chang, Boxin Shi
ECCV (7)4
2022 Data Association Between Event Streams and Intensity Frames Under Diverse Baselines
Dehao Zhang, Qiankun Ding, Peiqi Duan 0002, Chu Zhou, Boxin Shi
ECCV (7)5
2022 AdsCVLR: Commercial Visual-Linguistic Representation Modeling in Sponsored Search
abstract
Sponsored search advertisements (ads) appear next to search results when consumers look for products and services on search engines. As the fundamental basis of search ads, relevance modeling has attracted increasing attention due to the significant research challenges and tremendous practical value. In this paper, we address the problem of multi-modal modeling in sponsored search, which models the relevance between user query and commercial ads with multi-modal structured information. To solve this problem, we propose a transformer architecture with Ads data on Commercial Visual-Linguistic Representation (AdsCVLR) with contrastive learning that naturally extends the transformer encoder with the complementary multi-modal inputs, serving as a strong aggregator of image-text features. We also make a public advertising dataset, which includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on. Empirically, we evaluate the AdsCVLR model over the large industry dataset, and the experimental results of online/offline tests show the superiority of our method.
Yongjie Zhu, Chunhui Han, Yuefeng Zhan, Bochen Pang, Zhaoju Li, Hao Sun 0015, Si Li 0001, Boxin Shi, Nan Duan 0001, Ruofei Zhang, Liangjie Zhang, Qi Zhang 0066
ACM Multimedia8
2022 Neural Transmitted Radiance Fields
abstract
Neural radiance fields (NeRF) have brought tremendous progress to novel view synthesis. Though NeRF enables the rendering of subtle details in a scene by learning from a dense set of images, it also reconstructs the undesired reflections when we capture images through glass. As a commonly observed interference, the reflection would undermine the visibility of the desired transmitted scene behind glass by occluding the transmitted light rays. In this paper, we aim at addressing the problem of rendering novel transmitted views given a set of reflection-corrupted images. By introducing the transmission encoder and recurring edge constraints as guidance, our neural transmitted radiance fields can resist such reflection interference during rendering and reconstruct high-fidelity results even under sparse views. The proposed method achieves superior performance from the experiments on a newly collected dataset compared with state-of-the-art methods.
Chengxuan Zhu, Renjie Wan, Boxin Shi
NeurIPS3
2022 Shape and Albedo Recovery by Your Phone using Stereoscopic Flash and No-Flash Photography
abstract
Recovering shape and albedo for the immense number of existing cultural heritage artifacts is challenging. Accurate 3D reconstruction systems are typically expensive and thus inaccessible to many and cheaper off-the-shelf 3D sensors often generate results of unsatisfactory quality. This paper presents a high-fidelity shape and albedo recovery method that only requires a stereo camera and a flashlight, a typical camera setup equipped in many off-the-shelf smartphones. The stereo camera allows us to infer rough shape from a pair of no-flash images, and a flash image is further captured for shape refinement based on our flash/no-flash image formation model. We verify the effectiveness of our method on real-world artifacts in indoor and outdoor conditions using smartphones with different camera/flashlight configurations. Comparison results demonstrate that our stereoscopic flash and no-flash photography benefits the high-fidelity shape and albedo recovery on a smartphone. Using our method, people can immediately turn their phones into high-fidelity 3D scanners, facilitating the digitization of cultural heritage artifacts.
Michael Waechter, Boxin Shi, Fumio Okura, Yasuyuki Matsushita
Int. J. Comput. Vis.3
2022 Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances
Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita
Int. J. Comput. Vis.3
2022 NormAttention-PSN: A High-frequency Region Enhanced Photometric Stereo Network with Normalized Attention
Yakun Ju, Boxin Shi, Muwei Jian, Lin Qi 0004, Junyu Dong, Kin-Man Lam 0001
Int. J. Comput. Vis.2
2022 Deep Photometric Stereo for Non-Lambertian Surfaces
abstract
This paper addresses the problem of photometric stereo, in both calibrated and uncalibrated scenarios, for non-Lambertian surfaces based on deep learning. We first introduce a fully convolutional deep network for calibrated photometric stereo, which we call PS-FCN. Unlike traditional approaches that adopt simplified reflectance models to make the problem tractable, our method directly learns the mapping from reflectance observations to surface normal, and is able to handle surfaces with general and unknown isotropic reflectance. At test time, PS-FCN takes an arbitrary number of images and their associated light directions as input and predicts a surface normal map of the scene in a fast feed-forward pass. To deal with the uncalibrated scenario where light directions are unknown, we introduce a new convolutional network, named LCNet, to estimate light directions from input images. The estimated light directions and the input images are then fed to PS-FCN to determine the surface normals. Our method does not require a pre-defined set of light directions and can handle multiple images in an order-agnostic manner. Thorough evaluation of our approach on both synthetic and real datasets shows that it outperforms state-of-the-art methods in both calibrated and uncalibrated scenarios.
Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Guided Event Filtering: Synergy Between Intensity Images and Neuromorphic Events for High Performance Imaging
abstract
Many visual and robotics tasks in real-world scenarios rely on robust handling of high speed motion and high dynamic range (HDR) with effectively high spatial resolution and low noise. Such stringent requirements, however, cannot be directly satisfied by a single imager or imaging modality, rather by multi-modal sensors with complementary advantages. In this paper, we address high performance imaging by exploring the synergy between traditional frame-based sensors with high spatial resolution and low sensor noise, and emerging event-based sensors with high speed and high dynamic range. We introduce a novel computational framework, termed Guided Event Filtering (GEF), to process these two streams of input data and output a stream of super-resolved yet noise-reduced events. To generate high quality events, GEF first registers the captured noisy events onto the guidance image plane according to our flow model. it then performs joint image filtering that inherits the mutual structure from both inputs. Lastly, GEF re-distributes the filtered event frame in the space-time volume while preserving the statistical characteristics of the original events. When the guidance images under-perform, GEF incorporates an event self-guiding mechanism that resorts to neighbor events for guidance. We demonstrate the benefits of GEF by applying the output high quality events to existing event-based algorithms across diverse application categories, including high speed object tracking, depth estimation, high frame-rate video synthesis, and super resolution/HDR/color image restoration.
Peiqi Duan 0002, Zihao W. Wang, Boxin Shi, Oliver Cossairt, Tiejun Huang 0001, Aggelos K. Katsaggelos
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Patch-Based Uncalibrated Photometric Stereo Under Natural Illumination
abstract
This paper presents a photometric stereo method that works with unknown natural illumination without any calibration objects or initial guess of the target shape. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary orthogonal ambiguity. We further build the patch connections by extracting consistent surface normal pairs via spatial overlaps among patches and intensity profiles. Guided by these connections, the local ambiguities are unified to a global orthogonal one through Markov Random Field optimization and rotation averaging. After applying the integrability constraint, our solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods.
Heng Guo 0003, Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Ping Tan 0002, Yasuyuki Matsushita
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Optimizing Latent Distributions for Non-Adversarial Generative Networks
abstract
The generator in generative adversarial networks (GANs) is driven by a discriminator to produce high-quality images through an adversarial game. At the same time, the difficulty of reaching a stable generator has been increased. This paper focuses on non-adversarial generative networks that are trained in a plain manner without adversarial loss. The given limited number of real images could be insufficient to fully represent the real data distribution. We therefore investigate a set of distributions in a Wasserstein ball centred on the distribution induced by the training data and propose to optimize the generator over this Wasserstein ball. We theoretically discuss the solvability of the newly defined objective function and develop a tractable reformulation to learn the generator. The connections and differences between the proposed non-adversarial generative networks and GANs are analyzed. Experimental results on real-world datasets demonstrate that the proposed algorithm can effectively learn image generators in a non-adversarial approach, and the generated images are of comparable quality with those from GANs.
Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Deep Photometric Stereo Networks for Determining Surface Normal and Reflectances
abstract
This article presents a photometric stereo method based on deep learning. One of the major difficulties in photometric stereo is designing an appropriate reflectance model that is both capable of representing real-world reflectances and computationally tractable for deriving surface normal. Unlike previous photometric stereo methods that rely on a simplified parametric image formation model, such as the Lambert's model, the proposed method aims at establishing a flexible mapping between complex reflectance observations and surface normal using a deep neural network. In addition, the proposed method predicts the reflectance, which allows us to understand surface materials and to render the scene under arbitrary lighting conditions. As a result, we propose a deep photometric stereo network (DPSN) that takes reflectance observations under varying light directions and infers the surface normal and reflectance in a per-pixel manner. To make the DPSN applicable to real-world scenes, a dataset of measured BRDFs (MERL BRDF dataset) has been used for training the network. Evaluation using simulation and real-world scenes shows the effectiveness of the proposed approach in estimating both surface normal and reflectances.
Hiroaki Santo, Masaki Samejima, Yusuke Sugano, Boxin Shi, Yasuyuki Matsushita
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Hybrid Face Reflectance, Illumination, and Shape From a Single Image
abstract
We propose HyFRIS-Net to jointly estimate the hybrid reflectance and illumination models, as well as the refined face shape from a single unconstrained face image in a pre-defined texture space. The proposed hybrid reflectance and illumination representation ensure photometric face appearance modeling in both parametric and non-parametric spaces for efficient learning. While forcing the reflectance consistency constraint for the same person and face identity constraint for different persons, our approach recovers an occlusion-free face albedo with disambiguated color from the illumination color. Our network is trained in a self-evolving manner to achieve general applicability on real-world data. We conduct comprehensive qualitative and quantitative evaluations with state-of-the-art methods to demonstrate the advantages of HyFRIS-Net in modeling photo-realistic face albedo, illumination, and shape.
Yongjie Zhu, Chen Li 0031, Si Li 0001, Boxin Shi, Yu-Wing Tai
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Self-Supervised Low-Light Image Enhancement Using Discrepant Untrained Network Priors
abstract
This paper proposes a deep learning method for low-light image enhancement, which exploits the generation capability of Neural Networks (NNs) while requiring no training samples except the input image itself. Based on the Retinex decomposition model, the reflectance and illumination of a low-light image are parameterized by two untrained NNs. The ambiguity between the two layers is resolved by the discrepancy between the two NNs in terms of architecture and capacity, while the complex noise with spatially-varying characteristics is handled by an illumination-adaptive self-supervised denoising module. The enhancement is done by jointly optimizing the Retinex decomposition and the illumination adjustment. Extensive experiments show that the proposed method not only outperforms existing non-learning-based and unsupervised-learning-based methods, but also competes favorably with some supervised-learning-based methods in extreme low-light conditions.
Jinxiu Liang, Yong Xu 0007, Yuhui Quan, Boxin Shi, Hui Ji 0002
IEEE Trans. Circuits Syst. Video Technol.4
2022 $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference Prediction
abstract
With the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework.
Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.4
2021 Normal Integration via Inverse Plane Fitting With Minimum Point-to-Plane Distance
abstract
This paper presents a surface normal integration method that solves an inverse problem of local plane fitting. Surface reconstruction from normal maps is essential in photometric shape reconstruction. To this end, we formulate normal integration in the camera coordinates and jointly solve for 3D point positions and local plane displacements. Unlike existing methods that consider the vertical distances between 3D points, we minimize the sum of squared point-to-plane distances. Our method can deal with both orthographic or perspective normal maps with arbitrary boundaries. Compared to existing normal integration methods, our method avoids the checkerboard artifact and performs more robustly against natural boundaries, sharp features, and outliers. We further provide a geometric analysis of the source of artifacts that appear in previous methods based on our plane fitting formulation. Experimental results on analytically computed, synthetic, and real-world surfaces show that our method yields accurate and stable reconstruction for both orthographic and perspective normal maps1.
Boxin Shi, Fumio Okura, Yasuyuki Matsushita
CVPR2
2021 EventZoom: Learning To Denoise and Super Resolve Neuromorphic Events
abstract
We address the problem of jointly denoising and super resolving neuromorphic events, a novel visual signal that represents thresholded temporal gradients in a space-time window. The challenge for event signal processing is that they are asynchronously generated, and do not carry absolute intensity but only binary signs informing temporal variations. To study event signal formation and degradation, we implement a display-camera system which enables multi-resolution event recording. We further propose Event- Zoom, a deep neural framework with a backbone architecture of 3D U-Net. EventZoom is trained in a noise-to-noise fashion where the two ends of the network are unfiltered noisy events, enforcing noise-free event restoration. For resolution enhancement, EventZoom incorporates an event-to- image module supervised by high resolution images. Our results showed that EventZoom achieves at least 40 × temporal efficiency compared to state-of-the-art (SOTA) event denoisers. Additionally, we demonstrate that EventZoom enables performance improvements on applications including event-based visual object tracking and image reconstruction. EventZoom achieves SOTA super resolution image reconstruction results while being 10× faster.
Peiqi Duan 0002, Zihao W. Wang, Yi Ma 0001, Boxin Shi
CVPR5
2021 Multispectral Photometric Stereo for Spatially-Varying Spectral Reflectances: A Well Posed Problem?
abstract
Multispectral photometric stereo (MPS) aims at recovering the surface normal of a scene from a single-shot multi-spectral image, which is known as an ill-posed problem. To make the problem well-posed, existing MPS methods rely on restrictive assumptions, such as shape prior, surfaces having a monochromatic with uniform albedo. This paper alleviates the restrictive assumptions in existing methods. We show that the problem becomes well-posed for a surface with a uniform chromaticity but spatially-varying albedos based on our new formulation. Specifically, if at least three (or two) scene points share the same chromaticity, the proposed method uniquely recovers their surface normals and spectral reflectance with the illumination of more than or equal to four (or five) spectral lights. Besides, our method can be made robust by having many (i.e., 4 or more) spectral bands using robust estimation techniques for conventional photometric stereo. Experiments on both synthetic and real-world scenes demonstrate the effectiveness of our method. Our data and result can be found at https://github.com/GH-HOME/MultispectralPS.git.
Heng Guo 0003, Fumio Okura, Boxin Shi, Takuya Funatomi, Yasuhiro Mukaigawa, Yasuyuki Matsushita
CVPR3
2021 Panoramic Image Reflection Removal
abstract
This paper studies the problem of panoramic image reflection removal, aiming at reliving the content ambiguity between reflection and transmission scenes. Although a partial view of the reflection scene is included in the panoramic image, it cannot be utilized directly due to its misalignment with the reflection-contaminated image. We propose a two-step approach to solve this problem, by first accomplishing geometric and photometric alignment for the reflection scene via a coarse-to-fine strategy, and then restoring the transmission scene via a recovery network. The proposed method is trained with a synthetic dataset and verified quantitatively with a real panoramic image dataset. The effectiveness of the proposed method is validated by the significant performance advantage over single image-based reflection removal methods and generalization capacity to limited-FoV scenarios captured by conventional camera or mobile phone users.
Yuchen Hong, Lingran Zhao, Xudong Jiang 0001, Alex Chichung Kot, Boxin Shi
CVPR6
2021 Single Image Reflection Removal With Absorption Effect
abstract
In this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solution that explicitly takes the absorption effect into account. The first step estimates the absorption effect from a reflection-contaminated image, while the second step recovers the transmission image by taking a reflection-contaminated image and the estimated absorption effect as the input. Experimental results on four public datasets show that our two-step solution not only successfully removes reflection artifact, but also faithfully restores the intensity distortion caused by the absorption effect. Our ablation studies further demonstrate that our method achieves superior performance on the recovery of overall intensity and has good model generalization capacity. The code is available at https://github.com/q-zh/absorption.
Boxin Shi, Jinnan Chen, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
CVPR2
2021 High-Speed Image Reconstruction Through Short-Term Plasticity for Spiking Cameras
abstract
Fovea, located in the centre of the retina, is specialized for high-acuity vision. Mimicking the sampling mechanism of the fovea, a retina-inspired camera, named spiking camera, is developed to record the external information with a sampling rate of 40,000 Hz, and outputs asynchronous binary spike streams. Although the temporal resolution of visual information is improved, how to reconstruct the scenes is still a challenging problem. In this paper, we present a novel high-speed image reconstruction model through the short-term plasticity (STP) mechanism of the brain. We derive the relationship between postsynaptic potential regulated by STP and the firing frequency of each pixel. By setting up the STP model at each pixel of the spiking camera, we can infer the scene radiance with the temporal regularity of the spike stream. Moreover, we show that STP can be used to distinguish the static and motion areas and further enhance the reconstruction results. The experimental results show that our methods achieve state-of-the-art performance in both image quality and computing time.
Yajing Zheng, Lingxiao Zheng, Zhaofei Yu, Boxin Shi, Yonghong Tian 0001, Tiejun Huang 0001
CVPR4
2021 Spatially-Varying Outdoor Lighting Estimation From Intrinsics
abstract
We present SOLID-Net, a neural network for spatially- varying outdoor lighting estimation from a single outdoor image for any 2D pixel location. Previous work has used a unified sky environment map to represent outdoor lighting. Instead, we generate spatially-varying local lighting environment maps by combining global sky environment map with warped image information according to geometric information estimated from intrinsics. As no outdoor dataset with image and local lighting ground truth is readily available, we introduce the SOLID-Img dataset with physically- based rendered images and their corresponding intrinsic and lighting information. We train a deep neural network to regress intrinsic cues with physically-based constraints and use them to conduct global and local lightings estimation. Experiments on both synthetic and real datasets show that SOLID-Net significantly outperforms previous methods.
Yongjie Zhu, Yinda Zhang 0001, Si Li 0001, Boxin Shi
CVPR4
2021 DeRenderNet: Intrinsic Image Decomposition of Urban Scenes with Shape-(In)dependent Shading Rendering
abstract
We propose DeRenderNet, a deep neural network to decompose the albedo and latent lighting, and render shape-(in)dependent shadings, given a single image of an outdoor urban scene, trained in a self-supervised manner. To achieve this goal, we propose to use the albedo maps extracted from scenes in videogames as direct supervision and pre-compute the normal and shadow prior maps based on the depth maps provided as indirect supervision. Compared with state-of-the-art intrinsic image decomposition methods, DeRenderNet produces shadow-free albedo maps with clean details and an accurate prediction of shadows in the shape-independent shading, which is shown to be effective in re-rendering and improving the accuracy of high-level vision tasks for urban scenes.
Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Boxin Shi
ICCP4
2021 EvIntSR-Net: Event Guided Multiple Latent Frames Reconstruction and Super-resolution
abstract
An event camera detects the scene radiance changes and sends a sequence of asynchronous event streams with high dynamic range, high temporal resolution, and low latency. However, the spatial resolution of event cameras is limited as a trade-off for these outstanding properties. To reconstruct high-resolution intensity images from event data, we propose EvIntSR-Net that converts Event data to multiple latent Intensity frames to achieve Super-Resolution on intensity images in this paper. EvIntSR-Net bridges the domain gap between event streams and intensity frames and learns to merge a sequence of latent intensity frames in a recurrent updating manner. Experimental results show that EvIntSR-Net can reconstruct SR intensity images with higher dynamic range and fewer blurry artifacts by fusing events with intensity frames for both simulated and real-world data. Furthermore, the proposed EvIntSR-Net is able to generate high-frame-rate videos with super-resolved frames.
Jin Han 0001, Yixin Yang 0008, Chu Zhou, Chao Xu 0006, Boxin Shi
ICCV5
2021 Learning to dehaze with polarization
abstract
Haze, a common kind of bad weather caused by atmospheric scattering, decreases the visibility of scenes and degenerates the performance of computer vision algorithms. Single-image dehazing methods have shown their effectiveness in a large variety of scenes, however, they are based on handcrafted priors or learned features, which do not generalize well to real-world images. Polarization information can be used to relieve its ill-posedness, however, real-world images are still challenging since existing polarization-based methods usually assume that the transmitted light is not significantly polarized, and they require specific clues to estimate necessary physical parameters. In this paper, we propose a generalized physical formation model of hazy images and a robust polarization-based dehazing pipeline without the above assumption or requirement, along with a neural network tailored to the pipeline. Experimental results show that our approach achieves state-of-the-art performance on both synthetic data and real-world hazy images.
Chu Zhou, Minggui Teng, Yufei Han 0002, Chao Xu 0006, Boxin Shi
NeurIPS5
2021 Face Image Reflection Removal
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot
Int. J. Comput. Vis.2
2021 A Microfacet-Based Model for Photometric Stereo with General Isotropic Reflectance
abstract
This paper presents a precise, stable, and invertible reflectance model for photometric stereo. This microfacet-based model is applicable to all types of isotropic surface reflectance, covering cases from diffusion to specular reflections. We introduce a single variable to physically quantify the surface smoothness, and by monotonically sliding this variable between 0 and 1, our model enables a versatile representation that can smoothly transform between an ellipsoid of revolution and the equation for Lambertian reflectance. In the inverse domain, this model offers a compact and physically interpretable formulation, for which we introduce a fast and lightweight solver that allows accurate estimations for both surface smoothness and surface shape. Finally, extensive experiments on the appearances of synthesized and real objects evidence that this model is state-of-the-art in our off-the-shelf solution.
Lixiong Chen, Yinqiang Zheng, Boxin Shi, Art Subpa-Asa, Imari Sato
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Pose-Normalized and Appearance-Preserved Street-to-Shop Clothing Image Generation and Feature Learning
abstract
We tackle the task of street-to-shop clothing image synthesis. Given a daily person image with a particular clothing item captured in the street scenario, we aim to synthesize the frontal facing view of that item in the shop scenario. This problem has the following challenges: 1) the distinct visual discrepancy between the street and shop scenario; 2) the severe shape deformation of clothing in the presence of an arbitrary human pose; 3) the preservation of fine-grained details during the process of clothing image generation. In this paper, we jointly solve these difficulties by proposing a Pose-Normalized and Appearance-Preserved Generative Adversarial Network (PNAP-GAN). More specifically, conditioned on the clothing-agnostic representation (i.e., clothing landmarks and semantic parsing map), we disentangle the shape and appearance synthesis in a coarse-to-fine framework. Moreover, a semantic embedding loss is introduced to guide the domain transfer in the semantic level (i.e., keeping the clothing attributes). With the synthesized frontal shop image, a pose-normalized representation in complementary to the domain-invariant feature learnt from the original street image are integrated to facilitate the problem of street-to-shop clothing retrieval. Extensive experiments conducted demonstrate the effectiveness of the proposed PNAP-GAN on generating high quality frontal-view images and the excellence of the learnt pose-normalized features on the retrieval task than existing methods. In addition, we demonstrate that the pose-normalized retrieval feature benefits the cross-scenario (i.e., street-to-shop) clothing image generation in a semantic-preserved manner.
Huijing Zhan, Chenyu Yi, Boxin Shi, Jie Lin 0001, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.3
2020 Dense Point Diffusion for 3D Object Detection
abstract
The backbone network adopted in state-of-the-art 3D object detectors lacks a good balance between high point resolution and large receptive field, both of which are desirable for object detection on point clouds. This work proposes Dense Point Diffusion module, a novel backbone network that solves these issues. It adopts dilated point convolution as a building block to enlarge the receptive field and retain the point resolution at the same time. Further, a number of such layers are densely connected, giving rise to large receptive field and multi-scale feature fusion, which are effective for object detection task. Comprehensive experiments verify the efficacy of our approach. The source code1has been released to facilitate the reproduction of the results.
Jiayan Cao, Qianqian Bi, Boxin Shi
3DV5
2020 Distilling Portable Generative Adversarial Networks for Image Translation
abstract
Despite Generative Adversarial Networks (GANs) have been widely used in various image-to-image translation tasks, they can be hardly applied on mobile devices due to their heavy computation and storage cost. Traditional network compression methods focus on visually recognition tasks, but never deal with generation tasks. Inspired by knowledge distillation, a student generator of fewer parameters is trained by inheriting the low-level and high-level information from the original heavy teacher generator. To promote the capability of student generator, we include a student discriminator to measure the distances between real images, and images generated by student and teacher generators. An adversarial learning process is therefore established to optimize student generator and student discriminator. Qualitative and quantitative analysis by conducting experiments on benchmark datasets demonstrate that the proposed method can learn portable generative models with strong performance.
Hanting Chen, Yunhe Wang 0001, Han Shu, Changyuan Wen, Chunjing Xu, Boxin Shi, Chao Xu 0006, Chang Xu 0002
AAAI6
2020 Beyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks
abstract
Deep neural networks often consist of a great number of trainable parameters for extracting powerful features from given datasets. One one hand, massive trainable parameters significantly enhance the performance of these deep networks. One the other hand, they bring the problem of over-fitting. To this end, dropout based methods disable some elements in the output feature maps during the training phase for reducing the co-adaptation of neurons. Although the generalization ability of the resulting models can be enhanced by these approaches, the conventional binary dropout is not the optimal solution. Therefore, we investigate the empirical Rademacher complexity related to intermediate layers of deep neural networks and propose a feature distortion method for addressing the aforementioned problem. In the training period, randomly selected elements in the feature maps will be replaced with specific values by exploiting the generalization error bound. The superiority of the proposed feature map distortion for producing deep neural network with higher testing performance is analyzed and demonstrated on several benchmark image datasets.
Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Boxin Shi, Chao Xu 0006, Chunjing Xu, Chang Xu 0002
AAAI4
2020 Reborn Filters: Pruning Convolutional Neural Networks with Limited Data
abstract
Channel pruning is effective in compressing the pretrained CNNs for their deployment on low-end edge devices. Most existing methods independently prune some of the original channels and need the complete original dataset to fix the performance drop after pruning. However, due to commercial protection or data privacy, users may only have access to a tiny portion of training examples, which could be insufficient for the performance recovery. In this paper, for pruning with limited data, we propose to use all original filters to directly develop new compact filters, named reborn filters, so that all useful structure priors in the original filters can be well preserved into the pruned networks, alleviating the performance drop accordingly. During training, reborn filters can be easily implemented via 1×1 convolutional layers and then be fused in the inference stage for acceleration. Based on reborn filters, the proposed channel pruning algorithm shows its effectiveness and superiority on extensive experiments.
Yehui Tang 0001, Shan You, Chang Xu 0002, Jin Han 0001, Chen Qian 0006, Boxin Shi, Chao Xu 0006, Changshui Zhang
AAAI6
2020 Stereoscopic Flash and No-Flash Photography for Shape and Albedo Recovery
abstract
We present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pair. From the stereo image pair, we recover a rough shape that captures low-frequency shape variation without high-frequency details. From the flash/no-flash pair, we derive an image formation model for Lambertian objects under natural lighting, based on which a fine normal map is obtained and fused with the rough shape. Further, we use the flash/no-flash pair for cast shadow detection and albedo canceling, making the shape recovery robust against shadows and albedo variation. We verify the effectiveness of our approach on both synthetic and real-world data.
Michael Waechter, Boxin Shi, Yasuyuki Matsushita
CVPR3
2020 Frequency Domain Compact 3D Convolutional Neural Networks
abstract
This paper studies the compression and acceleration of 3-dimensional convolutional neural networks (3D CNNs). To reduce the memory cost and computational complexity of deep neural networks, a number of algorithms have been explored by discovering redundant parameters in pre-trained networks. However, most of existing methods are designed for processing neural networks consisting of 2-dimensional convolution filters (i.e. image classification and detection) and cannot be straightforwardly applied for 3-dimensional filters (i.e. time series data). In this paper, we develop a novel approach for eliminating redundancy in the time dimensionality of 3D convolution filters by converting them into the frequency domain through a series of learned optimal transforms with extremely fewer parameters. Moreover, these transforms are forced to be orthogonal, and the calculation of feature maps can be accomplished in the frequency domain to achieve considerable speed-up rates. Experimental results on benchmark 3D CNN models and datasets demonstrate that the proposed Frequency Domain Compact 3D CNNs (FDC3D) can achieve the state-of-the-art performance, \eg a 2x speed-up ratio on the 3D-ResNet-18 without obviously affecting its accuracy.
Hanting Chen, Yunhe Wang 0001, Han Shu, Yehui Tang 0001, Chunjing Xu, Boxin Shi, Chao Xu 0006, Qi Tian 0001, Chang Xu 0002
CVPR6
2020 AdderNet: Do We Really Need Multiplications in Deep Learning?
abstract
Compared with cheap addition operation, multiplication operation is of much higher computation complexity. The widely-used convolutions in deep neural networks are exactly cross-correlation to measure the similarity between input feature and convolution filters, which involves massive multiplications between float values. In this paper, we present adder networks (AdderNets) to trade these massive multiplications in deep neural networks, especially convolutional neural networks (CNNs), for much cheaper additions to reduce computation costs. In AdderNets, we take the L1-norm distance between filters and input feature as the output response. The influence of this new similarity measure on the optimization of neural network have been thoroughly analyzed. To achieve a better performance, we develop a special back-propagation approach for AdderNets by investigating the full-precision gradient. We then propose an adaptive learning rate strategy to enhance the training procedure of AdderNets according to the magnitude of each neuron's gradient. As a result, the proposed AdderNets can achieve 74.9% Top-1 accuracy 91.7% Top-5 accuracy using ResNet-50 on the ImageNet dataset without any multiplication in convolutional layer. The codes are publicly available at: (https://github.com/huaweinoah/AdderNet).
Hanting Chen, Yunhe Wang 0001, Chunjing Xu, Boxin Shi, Chao Xu 0006, Qi Tian 0001, Chang Xu 0002
CVPR4
2020 On Positive-Unlabeled Classification in GAN
abstract
This paper defines a positive and unlabeled classification problem for standard GANs, which then leads to a novel technique to stabilize the training of the discriminator in GANs. Traditionally, real data are taken as positive while generated data are negative. This positive-negative classification criterion was kept fixed all through the learning process of the discriminator without considering the gradually improved quality of generated data, even if they could be more realistic than real data at times. In contrast, it is more reasonable to treat the generated data as unlabeled, which could be positive or negative according to their quality. The discriminator is thus a classifier for this positive and unlabeled classification problem, and we derive a new Positive-Unlabeled GAN (PUGAN). We theoretically discuss the global optimality the proposed model will achieve and the equivalent optimization goal. Empirically, we find that PUGAN can achieve comparable or even better performance than those sophisticated discriminator stabilization methods.
Tianyu Guo 0001, Chang Xu 0002, Yunhe Wang 0001, Boxin Shi, Chao Xu 0006, Dacheng Tao
CVPR5
2020 Neuromorphic Camera Guided High Dynamic Range Imaging
abstract
Reconstruction of high dynamic range image from a single low dynamic range image captured by a frame-based conventional camera, which suffers from over- or under-exposure, is an ill-posed problem. In contrast, recent neuromorphic cameras are able to record high dynamic range scenes in the form of an intensity map, with much lower spatial resolution, and without color. In this paper, we propose a neuromorphic camera guided high dynamic range imaging pipeline, and a network consisting of specially designed modules according to each step in the pipeline, which bridges the domain gaps on resolution, dynamic range, and color representation between two types of sensors and images. A hybrid camera system has been built to validate that the proposed method is able to reconstruct quantitatively and qualitatively high-quality high dynamic range images by successfully fusing the images and intensity maps for various real-world scenarios.
Jin Han 0001, Chu Zhou, Peiqi Duan 0002, Yehui Tang 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
CVPR8
2020 DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing
abstract
We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is particularly problematic when the function is represented as a neural network. We optimize both the forward and backward pass of our rendering layer to make it run efficiently with affordable memory consumption on a commodity graphics card. Our rendering method is fully differentiable such that losses can be directly computed on the rendered 2D observations, and the gradients can be propagated backward to optimize the 3D geometry. We show that our rendering method can effectively reconstruct accurate 3D shapes from various inputs, such as sparse depth and multi-view images, through inverse optimization. With the geometry based reasoning, our 3D shape prediction methods show excellent generalization capability and robustness against various noises.
Shaohui Liu, Yinda Zhang 0001, Songyou Peng, Boxin Shi, Marc Pollefeys, Zhaopeng Cui
CVPR4
2020 A Semi-Supervised Assessor of Neural Architectures
abstract
Neural architecture search (NAS) aims to automatically design deep neural networks of satisfactory performance. Wherein, architecture performance predictor is critical to efficiently value an intermediate neural architecture. But for the training of this predictor, a number of neural architectures and their corresponding real performance often have to be collected. In contrast with classical performance predictor optimized in a fully supervised way, this paper suggests a semi-supervised assessor of neural architectures. We employ an auto-encoder to discover meaningful representations of neural architectures. Taking each neural architecture as an individual instance in the search space, we construct a graph to capture their intrinsic similarities, where both labeled and unlabeled architectures are involved. A graph convolutional neural network is introduced to predict the performance of architectures based on the learned representations and their relation modeled by the graph. Extensive experimental results on the NAS-Benchmark-101 dataset demonstrated that our method is able to make a significant reduction on the required fully trained architectures for finding efficient architectures.
Yehui Tang 0001, Yunhe Wang 0001, Yixing Xu, Hanting Chen, Boxin Shi, Chao Xu 0006, Chunjing Xu, Qi Tian 0001, Chang Xu 0002
CVPR5
2020 Reflection Scene Separation From a Single Image
abstract
For images taken through glass, existing methods focus on the restoration of the background scene by regarding the reflection components as noise. However, the scene reflected by glass surface also contains important information to be recovered, especially for the surveillance or criminal investigations. In this paper, instead of removing reflection components from the mixture image, we aim at recovering reflection scenes from the mixture image. We first propose a strategy to obtain such ground truth and its corresponding input images. Then, we propose a two-stage framework to obtain the visible reflection scene from the mixture image. Specifically, we train the network with a shift-invariant loss which is robust to misalignment between the input and output images. The experimental results show that our proposed method achieves promising results.
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot
CVPR2
2020 Joint Filtering of Intensity Images and Neuromorphic Events for High-Resolution Noise-Robust Imaging
abstract
We present a novel computational imaging system with high resolution and low noise. Our system consists of a traditional video camera which captures high-resolution intensity images, and an event camera which encodes high-speed motion as a stream of asynchronous binary events. To process the hybrid input, we propose a unifying framework that first bridges the two sensing modalities via a noise-robust motion compensation model, and then performs joint image filtering. The filtered output represents the temporal gradient of the captured space-time volume, which can be viewed as motion-compensated event frames with high resolution and low noise. Therefore, the output can be widely applied to many existing event-based algorithms that are highly dependent on spatial resolution and noise robustness. In experimental results performed on both publicly available datasets as well as our contributing RGB-DAVIS dataset, we show systematic performance improvement in applications such as high frame-rate video synthesis, feature/corner detection and tracking, as well as high dynamic range image reconstruction.
Zihao W. Wang, Peiqi Duan 0002, Oliver Cossairt, Aggelos K. Katsaggelos, Tiejun Huang 0001, Boxin Shi
CVPR6
2020 MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image Synthesis
abstract
In this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named ``MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by leveraging the inter-condition correlations. For the image compositing, we theoretically prove the weaknesses of the cutting-edge methods, and make it more robust by removing the spatially-invariance constraint, and enabling the bounding mechanism and the spatial adaptability. We show the effectiveness of our method on the Video Instance-level Parsing dataset, and demonstrate the robustness through controllability tests.
Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi
CVPR5
2020 CARS: Continuous Evolution for Efficient Neural Architecture Search
abstract
Searching techniques in most of existing neural architecture search (NAS) algorithms are mainly dominated by differentiable methods for the efficiency reason. In contrast, we develop an efficient continuous evolutionary approach for searching neural networks. Architectures in the population that share parameters within one SuperNet in the latest generation will be tuned over the training dataset with a few epochs. The searching in the next evolution generation will directly inherit both the SuperNet and the population, which accelerates the optimal network generation. The non-dominated sorting strategy is further applied to preserve only results on the Pareto front for accurately updating the SuperNet. Several neural networks with different model sizes and performances will be produced after the continuous search with only 0.4 GPU days. As a result, our framework provides a series of networks with the number of parameters ranging from 3.7M to 5.1M under mobile settings. These networks surpass those produced by the state-of-the-art methods on the benchmark ImageNet dataset.
Zhaohui Yang 0003, Yunhe Wang 0001, Xinghao Chen 0001, Boxin Shi, Chao Xu 0006, Chunjing Xu, Qi Tian 0001, Chang Xu 0002
CVPR4
2020 What Does Plate Glass Reveal About Camera Calibration?
abstract
This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image contents. To reduce the impact of a noisy calibration cue estimated from a reflection-contaminated image, we propose two strategies: an optimization-based method that imposes part of though reliable entries on the map and a learning-based method that fully exploits all entries. We collect a dataset containing 320 samples as well as their camera parameters for evaluation. We demonstrate that our method not only facilitates a general single image camera calibration method that leverages image contents but also contributes to improving the performance of single image reflection removal. Furthermore, we show our byproduct output helps alleviate the ill-posed problem of estimating the panorama from a single image.
Jinnan Chen, Zhan Lu, Boxin Shi, Xudong Jiang 0001, Kim-Hui Yap, Ling-Yu Duan, Alex Chichung Kot
CVPR4
2020 Deep Shape from Polarization
Yunhao Ba, Alex Gilbert, Franklin Wang, Jinfa Yang, Rui Chen 0006, Boxin Shi, Achuta Kadambi
ECCV (24)8
2020 What Is Learned in Deep Uncalibrated Photometric Stereo?
Guanying Chen, Michael Waechter, Boxin Shi, Kwan-Yee Kenneth Wong, Yasuyuki Matsushita
ECCV (14)3
2020 FHDe2Net: Full High Definition Demoireing Network
Ce Wang 0007, Boxin Shi, Ling-Yu Duan
ECCV (22)3
2020 Conditional Image Repainting via Semantic Bridge and Piecewise Value Function
Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi
ECCV (9)5
2020 Near-Infrared Image Guided Reflection Removal
abstract
Removing reflections from a single RGB image is a highly ill-posed problem. Unlike RGB images, near-infrared (NIR) images obtained through an active NIR camera are less likely to be affected by reflections when glass and camera planes form certain angles, while textures on objects could “vanish” under certain circumstances. Based on this observation, we propose a two-stream neural network to remove undesired reflections in an RGB image with the guidance of an NIR image. To tackle the insufficiency of training data, we propose a synthetic data generation pipeline that simulates the reflection-suppressing nature of the active NIR imaging and build a dataset mixed with synthetic and real data. Experimental results show that the proposed method outperforms state-of-the-art reflection removal methods in both quantitative metrics and visual quality.
Yuchen Hong, Youwei Lyu, Si Li 0001, Boxin Shi
ICME4
2020 Group Contextual Encoding for 3D Point Clouds
abstract
Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characterize the global semantic context, and then based on these code words, the method learns a global contextual descriptor to reweight the featuremaps accordingly. Moreover, compared to 2D scenarios, data sparsity becomes a major issue in 3D point cloud scenarios, and the performance of contextual encoding quickly saturates when the number of code words increases. To mitigate this problem, we further proposed a group contextual encoding method, which divides the channel into groups and then performs encoding on group-divided feature vectors. This method facilitates learning of global context in grouped subspace for 3D point clouds. We evaluate the effectiveness and generalizability of our method on three widely-studied 3D point cloud tasks. Experimental results have shown that the proposed method outperformed the VoteNet remarkably with 3 mAP on the benchmark of SUN-RGBD, with the metrics of mAP@ 0.25, and a much greater margin of 6.57 mAP on ScanNet with the metrics of mAP@ 0.5. Compared to the baseline of PointNet++, the proposed method leads to an accuracy of 86 %, outperforming the baseline by 1.5 %. Our proposed method have outperformed the non-grouping baseline methods across the board and establishes new state-of-the-art on these benchmarks.
Xu Liu 0017, Chengtao Li, Jian Wang 0100, Boxin Shi, Xiaodong He 0001
NeurIPS5
2020 GPS-Net: Graph-based Photometric Stereo Network
abstract
Learning-based photometric stereo methods predict the surface normal either in a per-pixel or an all-pixel manner. Per-pixel methods explore the inter-image intensity variation of each pixel but ignore features from the intra-image spatial domain. All-pixel methods explore the intra-image intensity variation of each input image but pay less attention to the inter-image lighting variation. In this paper, we present a Graph-based Photometric Stereo Network, which unifies per-pixel and all-pixel processings to explore both inter-image and intra-image information. For per-pixel operation, we propose the Unstructured Feature Extraction Layer to connect an arbitrary number of input image-light pairs into graph structures, and introduce Structure-aware Graph Convolution filters to balance the input data by appropriately weighting shadows and specular highlights. For all-pixel operation, we propose the Normal Regression Network to make efficient use of the intra-image spatial information for predicting a surface normal map with rich details. Experimental results on the real-world benchmark show that our method achieves excellent performance under both sparse and dense lighting distributions.
Zhuokun Yao, Kun Li 0001, Ying Fu 0001, Haofeng Hu, Boxin Shi
NeurIPS5
2020 UnModNet: Learning to Unwrap a Modulo Image for High Dynamic Range Imaging
abstract
A conventional camera often suffers from over- or under-exposure when recording a real-world scene with a very high dynamic range (HDR). In contrast, a modulo camera with a Markov random field (MRF) based unwrapping algorithm can theoretically accomplish unbounded dynamic range but shows degenerate performances when there are modulus-intensity ambiguity, strong local contrast, and color misalignment. In this paper, we reformulate the modulo image unwrapping problem into a series of binary labeling problems and propose a modulo edge-aware model, named as UnModNet, to iteratively estimate the binary rollover masks of the modulo image for unwrapping. Experimental results show that our approach can generate 12-bit HDR images from 8-bit modulo images reliably, and runs much faster than the previous MRF-based algorithm thanks to the GPU acceleration.
Chu Zhou, Hang Zhao 0021, Jin Han 0001, Chang Xu 0002, Chao Xu 0006, Tiejun Huang 0001, Boxin Shi
NeurIPS7
2020 Ambiguity-Free Radiometric Calibration for Internet Photo Collections
abstract
Radiometrically calibrating nonlinear images from Internet photo collections makes photometric analysis applicable not only to lab data but also to big image data in the wild. However, conventional calibration methods cannot be directly applied to such photo collections. This paper presents a method to jointly perform radiometric calibration for a set of nonlinear images in Internet photo collections. By incorporating the consistency of scene reflectance of corresponding pixels across nonlinear images, the proposed method first estimates radiometric response functions of all the nonlinear images up to a unique exponential ambiguity using a rank minimization framework. The ambiguity is then resolved using the linear edge color blending constraint. Quantitative evaluation using both synthetic and real-world data shows the effectiveness of the proposed method.
Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 CoRRN: Cooperative Reflection Removal Network
abstract
Removing the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose a network with the feature-sharing strategy to tackle this problem in a cooperative and unified framework, by integrating image context information and the multi-scale gradient information. To remove the strong reflections existed in some local regions, we propose a statistic loss by considering the gradient level statistics between the background and reflections. Our network is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods.
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Multi-View Photometric Stereo: A Robust Solution and Benchmark Dataset for Spatially Varying Isotropic Materials
abstract
We present a method to capture both 3D shape and spatially varying reflectance with a multi-view photometric stereo (MVPS) technique that works for general isotropic materials. Our algorithm is suitable for perspective cameras and nearby point light sources. Our data capture setup is simple, which consists of only a digital camera, some LED lights, and an optional automatic turntable. From a single viewpoint, we use a set of photometric stereo images to identify surface points with the same distance to the camera. We collect this information from multiple viewpoints and combine it with structure-from-motion to obtain a precise reconstruction of the complete 3D shape. The spatially varying isotropic bidirectional reflectance distribution function (BRDF) is captured by simultaneously inferring a set of basis BRDFs and their mixing weights at each surface point. In experiments, we demonstrate our algorithm with two different setups: a studio setup for highest precision and a desktop setup for best usability. According to our experiments, under the studio setting, the captured shapes are accurate to 0.5 millimeters and the captured reflectance has a relative root-mean-square error (RMSE) of 9%. We also quantitatively evaluate state-of-the-art MVPS on a newly collected benchmark dataset, which is publicly available for inspiring future research.
Min Li 0049, Zhenglong Zhou, Boxin Shi, Changyu Diao, Ping Tan 0002
IEEE Trans. Image Process.4
2020 Robust Student Network Learning
abstract
Deep neural networks bring in impressive accuracy in various applications, but the success often relies on heavy network architectures. Taking well-trained heavy networks as teachers, classical teacher-student learning paradigm aims to learn a student network that is lightweight yet accurate. In this way, a portable student network with significantly fewer parameters can achieve considerable accuracy, which is comparable to that of a teacher network. However, beyond accuracy, the robustness of the learned student network against perturbation is also essential for practical uses. Existing teacher-student learning frameworks mainly focus on accuracy and compression ratios, but ignore the robustness. In this paper, we make the student network produce more confident predictions with the help of the teacher network, and analyze the lower bound of the perturbation that will destroy the confidence of the student network. Two important objectives regarding prediction scores and gradients of examples are developed to maximize this lower bound, to enhance the robustness of the student network without sacrificing the performance. Experiments on benchmark data sets demonstrate the efficiency of the proposed approach to learning robust student networks that have satisfying accuracy and compact sizes.
Tianyu Guo 0001, Chang Xu 0002, Shiyi He, Boxin Shi, Chao Xu 0006, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.4
2020 Multirobot Object Transport via Robust Caging
abstract
In this paper, we propose a control algorithm to collectively transport an object using a group of relatively low-cost robots. We address this problem using the robust caging, which features reliable object closure with minimum number of robots, and requires no high-precision control capability on the individual robot. Given a 2-D convex object, the proposed method uses the quality of complete robustness to first optimize the number of robots in the initial formation, and then reorient and move the formation. The method is free of force analysis, and therefore less prone to sensor errors and failures. Compared with state-of-the-art multirobot object transport approaches, which require more robots and rely heavily on high-precision control, such as force and torque feedback control, our method uses fewer robots and has high tolerance to control noises. We performed both simulation and real-time experiments to demonstrate the performance of our method. We conclude that the proposed robust caging is promising under reduced number of robots and a certain level of control noises in multirobot object transport tasks.
Weiwei Wan, Boxin Shi, Zijian Wang 0003, Rui Fukui
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Summary study of data-driven photometric stereo methods
abstract
A photometric stereo method aims to recover the surface normal of a 3D object observed under varying light directions. It is an ill-defined problem because the general reflectance properties of the surface are unknown. This paper reviews existing data-driven methods, with a focus on their technical insights into the photometric stereo problem. We divide these methods into two categories, per-pixel and all-pixel, according to how they process an image. We discuss the differences and relationships between these methods from the perspective of inputs, networks, and data, which are key factors in designing a deep learning approach. We demonstrate the performance of the models using a popular benchmark dataset. Data-driven photometric stereo methods have shown that they possess a superior performance advantage over traditional methods. However, these methods suffer from various limitations, such as limited generalization capability. Finally, this study suggests directions for future research.
Boxin Shi, Gang Pan 0001
Virtual Real. Intell. Hardw.2
2019 Smooth Deep Image Generator from Noises
abstract
Generative Adversarial Networks (GANs) have demonstrated a strong ability to fit complex distributions since they were presented, especially in the field of generating natural images. Linear interpolation in the noise space produces a continuously changing in the image space, which is an impressive property of GANs. However, there is no special consideration on this property in the objective function of GANs or its derived models. This paper analyzes the perturbation on the input of the generator and its influence on the generated images. A smooth generator is then developed by investigating the tolerable input perturbation. We further integrate this smooth generator with a gradient penalized discriminator, and design smooth GAN that generates stable and high-quality images. Experiments on real-world image datasets demonstrate the necessity of studying smooth generator and the effectiveness of the proposed algorithm.
Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao
AAAI3
2019 Self-Calibrating Deep Photometric Stereo Networks
abstract
This paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a scene with unknown arbitrary reflectances observed under unknown varying light directions. To achieve this goal, we propose a two-stage deep learning architecture, called SDPS-Net, which can effectively take advantage of intermediate supervision, resulting in reduced learning difficulty compared to a single-stage model. Experiments on both synthetic and real datasets show that our proposed approach significantly outperforms previous uncalibrated photometric stereo methods.
Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong
CVPR3
2019 Data-Free Learning of Student Networks
abstract
Learning portable neural networks is very essential for computer vision for the purpose that pre-trained heavy deep models can be well applied on edge devices such as mobile phones and micro sensors. Most existing deep neural network compression and speed-up methods are very effective for training compact deep models, when we can directly access the training dataset. However, training data for the given deep network are often unavailable due to some practice problems (\eg privacy, legal issue, and transmission), and the architecture of the given network are also unknown except some interfaces. To this end, we propose a novel framework for training efficient deep neural networks by exploiting generative adversarial networks (GANs). To be specific, the pre-trained teacher networks are regarded as a fixed discriminator and the generator is utilized for derivating training samples which can obtain the maximum response on the discriminator. Then, an efficient network with smaller model size and computational complexity is trained using the generated data and the teacher network, simultaneously. Efficient student networks learned using the proposed Data-Free Learning (DFL) method achieve 92.22% and 74.47% accuracies without any training data on the CIFAR-10 and CIFAR-100 datasets, respectively. Meanwhile, our student network obtains an 80.56% accuracy on the CelebA benchmark.
Hanting Chen, Yunhe Wang 0001, Chang Xu 0002, Zhaohui Yang 0003, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu 0006, Qi Tian 0001
ICCV6
2019 Mop Moiré Patterns Using MopNet
abstract
Moiré pattern is a common image quality degradation caused by frequency aliasing between monitors and cameras when taking screen-shot photos. The complex frequency distribution, imbalanced magnitude in colour channels, and diverse appearance attributes of moiré pattern make its removal a challenging problem. In this paper, we propose a Moiré pattern Removal Neural Network (MopNet) to solve this problem. All core components of MopNet are specially designed for unique properties of moire patterns, including the multi-scale feature aggregation addressing complex frequency, the channel-wise target edge predictor to exploit imbalanced magnitude among colour channels, and the attribute-aware classifier to characterize the diverse appearance for better modelling Moiré patterns. Quantitative and qualitative experimental comparison validate the state-of-the-art performance of MopNet.
Ce Wang 0007, Boxin Shi, Ling-Yu Duan
ICCV3
2019 Learning to Jointly Generate and Separate Reflections
abstract
Existing learning-based single image reflection removal methods using paired training data have fundamental limitations about the generalization capability on real-world reflections due to the limited variations in training pairs. In this work, we propose to jointly generate and separate reflections within a weakly-supervised learning framework, aiming to model the reflection image formation more comprehensively with abundant unpaired supervision. By imposing the adversarial losses and combinable mapping mechanism in a multi-task structure, the proposed framework elegantly integrates the two separate stages of reflection generation and separation into a unified model. The gradient constraint is incorporated into the concurrent training process of the multi-task learning as well. In particular, we built up an unpaired reflection dataset with 4,027 images, which is useful for facilitating the weakly-supervised learning of reflection removal model. Extensive experiments on a public benchmark dataset show that our framework performs favorably against state-of-the-art methods and consistently produces visually appealing results.
Daiqian Ma, Renjie Wan, Boxin Shi, Alex Chichung Kot, Ling-Yu Duan
ICCV3
2019 SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation Networks
abstract
This paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estimation network to estimate surface normals. Both networks are jointly constrained by the proposed symmetric and asymmetric loss functions to enforce isotropic constrain and perform outlier rejection of global illumination effects. SPLINE-Net is verified to outperform existing methods for photometric stereo of general BRDFs by using only ten images of different lights instead of using nearly one hundred images.
Yiming Jia, Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
ICCV3
2019 Fashion Recommendation on Street Images
abstract
Learning the compatibility relationship is of vital importance to a fashion recommendation system, while existing works achieve this merely on product images but not on street images in the complex daily life scenario. In this paper, we propose a novel fashion recommendation system: Given a query item of interest in the street scenario, the system can return the compatible items. More specifically, a two-stage curriculum learning scheme is developed to transfer the semantics from the product to street outfit images. We also propose a domain-specific missing item imputation method based on style and color similarity to handle the incomplete outfits. To support the training of deep recommendation model, we collect a large dataset with street outfit images. The experiments on the dataset demonstrate the advantages of the proposed method over the state-of-the-art approaches on both the street images and the product images.
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
ICIP2
2019 Denoising Adversarial Networks for Rain Removal and Reflection Removal
abstract
This paper presents a novel adversarial scheme to perform image denoising for the tasks of rain streak removal and reflection removal. Similar to several previous works, the proposed method first estimates a prior image and then uses it to guide the inference of noise-free image. The novelty of our approach is to jointly learn the gradient and noise-free image based on an adversarial scheme. More specifically, we use the gradient map as the prior image. The inferred noise-free image guided by an estimated gradient is regarded as a negative sample, while the noise-free image guided by the ground truth of a gradient is taken as a positive sample. With the anchor defined by the ground truth of noise-free image, we play a min-max game to jointly train two optimizers for the estimation of the gradient and the inference of noise-free images. We show that both prior image and noise-free image can be accurately obtained under this adversarial scheme. Our state-of-the-art performance achieved on two public benchmark datasets validate the effectiveness of our approach.
Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot
ICIP2
2019 Learning to Remove Reflections for Text Images
abstract
Text images taken behind a piece of glass in the wild are largely contaminated by reflections. Directly applying existing reflection removal methods on text images with reflections cannot recover clear and correct text contents due to the ignorance of special characteristics of texts. This paper proposes a stacked framework to solve the text image reflection removal problem by specifically considering the regional properties of reflection and embedding the specific text priors into the estimation process in a unified manner. Experiment results on a newly collected dataset demonstrate that the proposed method outperforms state-of-the-art methods in recovering visually pleasant reflection-free images and recognizable text features.
Ce Wang 0007, Renjie Wan, Feng Gao 0014, Boxin Shi, Ling-Yu Duan
ICME4
2019 From Market to Dish: Multi-ingredient Image Recognition for Personalized Recipe Recommendation
abstract
Recognition of food ingredients enables applications on recipe recommendation for developing a healthier eating habit. Existing ingredients recognition methods largely rely on ideal images captured in a controlled environment, while ingredients are usually displayed unorderly in a complex environment in the market. We propose the multi-ingredient recognition problem in the market and develop a Spatial Regularization Network (SRN) based method to solve it by using a newly collected multiple vegetable image dataset captured in the market. We further use the recognition result to develop a recipe recommendation system to satisfy the daily nutrition requirements and individual preference of each user. Experiments show that our multi-ingredient recognition outperforms previous methods over 14% in mAP and recommendation model shows an improvement of over 23% in HR@10.
Lin Zhang 0014, Jianbo Zhao 0002, Si Li 0001, Boxin Shi, Ling-Yu Duan
ICME4
2019 LegoNet: Efficient Convolutional Neural Networks with Lego Filters
abstract
This paper aims to build efficient convolutional neural networks using a set of Lego filters. Many successful building blocks, e.g., inception and residual modules, have been designed to refresh state-of-the-art records of CNNs on visual recognition tasks. Beyond these high-level modules, we suggest that an ordinary filter in the neural network can be upgraded to a sophisticated module as well. Filter modules are established by assembling a shared set of Lego filters that are often of much lower dimensions. Weights in Lego filters and binary masks to stack Lego filters for these filter modules can be simultaneously optimized in an end-to-end manner as usual. Inspired by network engineering, we develop a split-transform-merge strategy for an efficient convolution by exploiting intermediate Lego feature maps. The compression and acceleration achieved by Lego Networks using the proposed Lego filters have been theoretically discussed. Experimental results on benchmark datasets and deep models demonstrate the advantages of the proposed Lego filters and their potential real-world applications on mobile devices.
Zhaohui Yang 0003, Yunhe Wang 0001, Chuanjian Liu, Hanting Chen, Chunjing Xu, Boxin Shi, Chao Xu 0006, Chang Xu 0002
ICML6
2019 See Through the Windshield from Surveillance Camera
abstract
This paper attempts to address the challenging task of seeing through the windshield images captured by surveillance cameras in the wild. Such images usually have very low visibility due to heterogeneous degradations caused by blur, haze, reflection, noise etc., which makes existing image enhancing methods inapplicable. We propose a windshield image restoration generative adversarial network (WIRE-GAN) to restore and enhance the visibility of windshield images. We adopt the weakly supervised framework based on the generative model, which has effectively released the request of paired training data for a specific type of degradation. To generate more semantically consistent results even in extreme lighting conditions, we introduce a novel content-preserving strategy into the proposed weakly-supervised framework. To make the image restoration more reliable, the WIRE-GAN network constructs a sort of content-aware embedding space and enforces the constraint of the restored windshield images being closer to the original input in the embedding space. Moreover, we collect a large-scale windshield image dataset (WIRE dataset) to validate the advantage of our method in improving the image quality, and further evaluate the impact of windshield restoration on the vehicle ReID performance.
Daiqian Ma, Renjie Wan, Ce Wang 0007, Boxin Shi, Ling-Yu Duan
ACM Multimedia5
2019 Learning from Bad Data via Generation
abstract
Bad training data would challenge the learning model from understanding the underlying data-generating scheme, which then increases the difficulty in achieving satisfactory performance on unseen test data. We suppose the real data distribution lies in a distribution set supported by the empirical distribution of bad data. A worst-case formulation can be developed over this distribution set, and then be interpreted as a generation task in an adversarial manner. The connections and differences between GANs and our framework have been thoroughly discussed. We further theoretically show the influence of this generation task on learning from bad data and reveal its connection with a data-dependent regularization. Given different distance measures (\eg, Wasserstein distance or JS divergence) of distributions, we can derive different objective functions for the problem. Experimental results on different kinds of bad training data demonstrate the necessity and effectiveness of the proposed method.
Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao
NeurIPS3
2019 Reflection Separation using a Pair of Unpolarized and Polarized Images
abstract
When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical constraints from a pair of unpolarized and polarized images to separate reflection and transmission layers. Due to the simplified capturing setup, the system becomes more underdetermined compared with existing polarization based solutions that take three or more images as input. We propose to solve semireflector orientation estimation first to make the physical image formation well-posed and then learn to reliably separate two layers using a refinement network with gradient loss. Quantitative and qualitative experimental results show our approach performs favorably over existing polarization and single image based solutions.
Youwei Lyu, Zhaopeng Cui, Si Li 0001, Marc Pollefeys, Boxin Shi
NeurIPS5
2019 Data-driven photometric 3D modeling
abstract
course Share on Data-driven photometric 3D modeling Author: Boxin Shi Peking University Peking UniversityView Profile Authors Info & Claims SA '19: SIGGRAPH Asia 2019 CoursesNovember 2019 Article No.: 129Pages 1–139https://doi.org/10.1145/3355047.3359422Published:17 November 2019Publication History 1citation249DownloadsMetricsTotal Citations1Total Downloads249Last 12 Months26Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Boxin Shi
SIGGRAPH Asia1
2019 DeepShoe: An improved Multi-Task View-invariant CNN for street-to-shop shoe retrieval
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
Comput. Vis. Image Underst.2
2019 A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo
abstract
Classic photometric stereo is often extended to deal with real-world materials and work with unknown lighting conditions for practicability. To quantitatively evaluate non-Lambertian and uncalibrated photometric stereo, a photometric stereo image dataset containing objects of various shapes with complex reflectance properties and high-quality ground truth normals is still missing. In this paper, we introduce the 'DiLiGenT' dataset with calibrated Directional Lightings, objects of General reflectance with different shininess, and 'ground Truth' normals from high-precision laser scanning. We use our dataset to quantitatively evaluate state-of-the-art photometric stereo methods for general materials and unknown lighting conditions, selected from a newly proposed photometric stereo taxonomy emphasizing non-Lambertian and uncalibrated methods. The dataset and evaluation results are made publicly available, and we hope it can serve as a benchmark platform that inspires future research.
Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Learning to remove reflections from windshield images
Ce Wang 0007, Boxin Shi, Ling-Yu Duan
Signal Process. Image Commun.2
2019 Numerical Reflectance Compensation for Non-Lambertian Photometric Stereo
abstract
The surface normal estimation from photometric stereo becomes less reliable when the surface reflectance deviates from the Lambertian assumption. The non-Lambertian effect can be explicitly addressed by physics modeling to the reflectance function, at the cost of introducing highly nonlinear optimization. This paper proposes a numerical compensation scheme that attempts to minimize the angular error to address the non-Lambertian photometric stereo problem. Due to the multifaceted influence in the modeling of non-Lambertian reflectance in photometric stereo, directly minimizing the angular errors of surface normal is a highly complex problem. We introduce an alternating strategy, in which the estimated reflectance can be temporarily regarded as a known variable, to simplify the formulation of angular error. To reduce the impact of inaccurately estimated reflectance in this simplification, we propose a numerical compensation scheme whose compensation weight is formulated to reflect the reliability of estimated reflectance. Finally, the solution for the proposed numerical compensation scheme is efficiently computed by using cosine difference to approximate the angular difference. The experimental results show that our method can significantly improve the performance of the state-of-the-art methods on both synthetic data and real data with small additive costs. Moreover, our method initialized by results from the baseline method (least-square-based) achieves the state-of-the-art performance on both synthetic data and real data with significantly smaller overall computation, i.e., about eight times faster compared with the state-of-the-art methods.
Ajay Kumar 0001, Boxin Shi, Gang Pan 0001
IEEE Trans. Image Process.3
2019 A self-calibrated photo-geometric depth camera
Liang Xie 0008, Yuhua Xu 0003, Chenpeng Tong, Boxin Shi
Vis. Comput.6
2018 Depth-Aware Stereo Video Retargeting
abstract
As compared with traditional video retargeting, stereo video retargeting poses new challenges because stereo video contains the depth information of salient objects and its time dynamics. In this work, we propose a depth-aware stereo video retargeting method by imposing the depth fidelity constraint. The proposed depth-aware retargeting method reconstructs the 3D scene to obtain the depth information of salient objects. We cast it as a constrained optimization problem, where the total cost function includes the shape, temporal and depth distortions of salient objects. As a result, the solution can preserve the shape, temporal and depth fidelity of salient objects simultaneously. It is demonstrated by experimental results that the depth-aware retargeting method achieves higher retargeting quality and provides better user experience.
Bing Li 0024, Chia-Wen Lin, Boxin Shi, Tiejun Huang 0001, Wen Gao 0001, C.-C. Jay Kuo
CVPR3
2018 Uncalibrated Photometric Stereo Under Natural Illumination
abstract
This paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary rotation ambiguity. Our method connects the resulting patches and unifies the local ambiguities to a global rotation one through angular distance propagation defined over the whole surface. After applying the integrability constraint, our final solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods.
Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Yasuyuki Matsushita
CVPR2
2018 Self-Calibrating Polarising Radiometric Calibration
abstract
We present a self-calibrating polarising radiometric calibration method. From a set of images taken from a single viewpoint under different unknown polarising angles, we recover the inverse camera response function and the polarising angles relative to the first angle. The problem is solved in an integrated manner, recovering both of the unknowns simultaneously. The method exploits the fact that the intensity of polarised light should vary sinusoidally as the polarising filter is rotated, provided that the response is linear. It offers the first solution to demonstrate the possibility of radiometric calibration through polarisation. We evaluate the accuracy of our proposed method using synthetic data and real world objects captured using different cameras. The self-calibrated results were found to be comparable with those from multiple exposure sequence.
Daniel Teo, Boxin Shi, Yinqiang Zheng, Sai-Kit Yeung
CVPR2
2018 CRRN: Multi-Scale Guided Concurrent Reflection Removal Network
abstract
Removing the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose the Concurrent Reflection Removal Network (CRRN) to tackle this problem in a unified framework. Our proposed network integrates image appearance information and multi-scale gradient information with human perception inspired loss function, and is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Extensive experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods.
Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot
CVPR2
2018 SSF-CNN: Spatial and Spectral Fusion with CNN for Hyperspectral Image Super-Resolution
abstract
Fusing a low-resolution hyperspectral image with the corresponding high-resolution RGB image to obtain a high-resolution hyperspectral image is usually solved as an optimization problem with prior-knowledge such as sparsity representation and spectral physical properties as constraints, which have limited applicability. Deep convolutional neural network extracts more comprehensive features and is proved to be effective in upsampling RGB images. However, directly applying CNNs to upsample either the spatial or spectral dimension alone may not produce pleasing results due to the neglect of complementary information from both low resolution hyper spectral and high resolution RGB images. This paper proposes two types of novel CNN architectures to take advantages of spatial and spectral fusion for hyperspectral image superresolution. Experiment results on benchmark datasets validate that the proposed spatial and spectral fusion CNNs outperforms the state-of-the-art methods and baseline CNN architectures in both quantitative values and visual qualities.
Xianhua Han, Boxin Shi, Yinqiang Zheng
ICIP2
2018 Residual HSRCNN: Residual Hyper-Spectral Reconstruction CNN from an RGB Image
abstract
Hyper-spectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyper-spectral and high-resolution multi-spectral or RGB images can be captured at video rate. This study aims to generate a hyper-spectral image via enhancing spectral resolution of an RGB image, which might be easily obtained by a commodity camera. Motivated by the success of deep convolutional neural network (DCNN) for spatial resolution enhancement of natural images, we explore a spectral reconstruction CNN for spectral super-resolution with an available RGB image, which predicts the high-frequency content of the fine spectral wavelength in narrow band interval. Since the lost high-frequency content can not be perfectly recovered, by leveraging on the baseline CNN, we further propose a novel residual hyper-spectral reconstruction CNN framework to estimate the non-recovered high-frequency content (Residual) from the output of the baseline CNN. Experiments on benchmark hyper-spectral datasets validate that the proposed method achieves promising performances compared with the existing state-of-the-art methods.
Xianhua Han, Boxin Shi, Yinqiang Zheng
ICPR2
2018 ChipGAN: A Generative Adversarial Network for Chinese Ink Wash Painting Style Transfer
abstract
Style transfer has been successfully applied on photos to generate realistic western paintings. However, because of the inherently different painting techniques adopted by Chinese and western paintings, directly applying existing methods cannot generate satisfactory results for Chinese ink wash painting style transfer. This paper proposes ChipGAN, an end-to-end Generative Adversarial Network based architecture for photo to Chinese ink wash painting style transfer. The core modules of ChipGAN enforce three constraints -- voids, brush strokes, and ink wash tone and diffusion -- to address three key techniques commonly adopted in Chinese ink wash painting. We conduct stylization perceptual study to score the similarity of generated paintings to real paintings by consulting with professional artists based on the newly built Chinese ink wash photo and image dataset. The advantages in visual quality compared with state-of-the-art networks and high stylization perceptual study scores show the effectiveness of the proposed method.
Feng Gao 0014, Daiqian Ma, Boxin Shi, Ling-Yu Duan
ACM Multimedia4
2018 Self-Similarity Constrained Sparse Representation for Hyperspectral Image Super-Resolution
abstract
Fusing a low-resolution hyperspectral image with the corresponding high-resolution multispectral image to obtain a high-resolution hyperspectral image is an important technique for capturing comprehensive scene information in both spatial and spectral domains. Existing approaches adopt sparsity promoting strategy, and encode the spectral information of each pixel independently, which results in noisy sparse representation. We propose a novel hyperspectral image super-resolution method via a self-similarity constrained sparse representation. We explore the similar patch structures across the whole image and the pixels with close appearance in local regions to create globalstructure groups and local-spectral super-pixels. By forcing the similarity of the sparse representations for pixels belonging to the same group and super-pixel, we alleviate the effect of the outliers in the learned sparse coding. Experiment results on benchmark datasets validate that the proposed method outperforms the stateof- the-art methods in both quantitative metrics and visual effect.
Xianhua Han, Boxin Shi, Yinqiang Zheng
IEEE Trans. Image Process.2
2018 Region-Aware Reflection Removal With Unified Content and Gradient Priors
abstract
Removing the undesired reflections in images taken through the glass is of broad application to various image processing and computer vision tasks. Existing single image based solutions heavily rely on scene priors such as separable sparse gradients caused by different levels of blur, and they are fragile when such priors are not observed. In this paper, we notice that strong reflections usually dominant a limited region in the whole image, and propose a Region-aware Reflection Removal (R3) approach by automatically detecting and heterogeneously processing regions with and without reflections. We integrate content and gradient priors to jointly achieve missing contents restoration as well as background and reflection separation in a unified optimization framework. Extensive validation using 50 sets of real data shows that the proposed method outperforms state-of-the-art on both quantitative metrics and visual qualities.
Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Wen Gao 0001, Alex Chichung Kot
IEEE Trans. Image Process.2
2017 DeepShoe: A Multi-Task View-Invariant CNN for Street-to-Shop Shoe Retrieval
Huijing Zhan, Boxin Shi, Alex Chichung Kot
BMVC2
2017 Polarimetric Multi-view Stereo
abstract
Multi-view stereo relies on feature correspondences for 3D reconstruction, and thus is fundamentally flawed in dealing with featureless scenes. In this paper, we propose polarimetric multi-view stereo, which combines per-pixel photometric information from polarization with epipolar constraints from multiple views for 3D reconstruction. Polarization reveals surface normal information, and is thus helpful to propagate depth to featureless regions. Polarimetric multi-view stereo is completely passive and can be applied outdoors in uncontrolled illumination, since the data capture can be done simply with either a polarizer or a polarization camera. Unlike previous work on shape-from-polarization which is limited to either diffuse polarization or specular polarization only, we propose a novel polarization imaging model that can handle real-world objects with mixed polarization. We prove there are exactly two types of ambiguities on estimating surface azimuth angles from polarization, and we resolve them with graph optimization and iso-depth contour tracing. This step significantly improves the initial depth map estimate, which are later fused together for complete 3D reconstruction. Extensive experimental results demonstrate high-quality 3D reconstruction and better performance than state-of-the-art multi-view stereo methods, especially on featureless 3D objects, such as ceramic tiles, office room with white walls, and highly reflective cars in the outdoors.
Zhaopeng Cui, Jinwei Gu, Boxin Shi, Ping Tan 0002, Jan Kautz
CVPR3
2017 Radiometric Calibration for Internet Photo Collections
abstract
Radiometrically calibrating the images from Internet photo collections brings photometric analysis from lab data to big image data in the wild, but conventional calibration methods cannot be directly applied to such image data. This paper presents a method to jointly perform radiometric calibration for a set of images in an Internet photo collection. By incorporating the consistency of scene reflectance for corresponding pixels in multiple images, the proposed method estimates radiometric response functions of all the images using a rank minimization framework. Our calibration aligns all response functions in an image set up to the same exponential ambiguity in a robust manner. Quantitative results using both synthetic and real data show the effectiveness of the proposed method.
Zhipeng Mo, Boxin Shi, Sai-Kit Yeung, Yasuyuki Matsushita
CVPR2
2017 A Microfacet-Based Reflectance Model for Photometric Stereo with Highly Specular Surfaces
abstract
A precise, stable and invertible model for surface reflectance is the key to the success of photometric stereo with real world materials. Recent developments in the field have enabled shape recovery techniques for surfaces of various types, but an effective solution to directly estimating the surface normal in the presence of highly specular reflectance remains elusive. In this paper, we derive an analytical isotropic microfacet-based reflectance model, based on which a physically interpretable approximate is tailored for highly specular surfaces. With this approximate, we identify the equivalence between the surface recovery problem and the ellipsoid of revolution fitting problem, where the latter can be described as a system of polynomials. Additionally, we devise a fast, non-iterative and globally optimal solver for this problem. Experimental results on both synthetic and real images validate our model and demonstrate that our solution can stably deliver superior performance in its targeted application domain.
Lixiong Chen, Yinqiang Zheng, Boxin Shi, Art Subpa-Asa, Imari Sato
ICCV3
2017 Benchmarking Single-Image Reflection Removal Algorithms
abstract
Removing undesired reflections from a photo taken in front of a glass is of great importance for enhancing the efficiency of visual computing systems. Various approaches have been proposed and shown to be visually plausible on small datasets collected by their authors. A quantitative comparison of existing approaches using the same dataset has never been conducted due to the lack of suitable benchmark data with ground truth. This paper presents the first captured Single-image Reflection Removal dataset ‘SIR2’ with 40 controlled and 100 wild scenes, ground truth of background and reflection. For each controlled scene, we further provide ten sets of images under varying aperture settings and glass thicknesses. We perform quantitative and visual quality comparisons for four state-of-the-art single-image reflection removal algorithms using four error metrics. Open problems for improving reflection removal algorithms are discussed at the end.
Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot
ICCV2
2017 Saliency detection by forward and backward cues in deep-CNN
abstract
As prior knowledge of objects or object features helps us make relations for similar objects on attentional tasks, pre-trained deep convolutional neural networks (CNNs) can be used to detect salient objects on images regardless of the object class is in the network knowledge or not. In this paper, we propose a top-down saliency model using CNN, a weakly supervised CNN model trained for 1000 object labelling task from RGB images. The model detects attentive regions based on their objectness scores predicted by selected features from CNNs. To estimate the salient objects effectively, we combine both forward and backward features, while demonstrating that partially-guided backpropagation will provide sufficient information for selecting the features from forward run of CNN model. Finally, these top-down cues are enhanced with a state-of-the-art bottom-up model as complementing the overall saliency. As the proposed model is an effective integration of forward and backward cues through objectness without any supervision or regression to ground truth data, it gives promising results compared to state-of-the-art models in two different datasets.
Nevrez Imamoglu, Chi Zhang 0027, Wataru Shimoda, Yuming Fang 0001, Boxin Shi
ICIP5
2017 Street-to-shop shoe retrieval with multi-scale viewpoint invariant triplet network
abstract
In this paper we aim to find exactly the same shoes given a daily shoe photo (street scenario) that matches the online shop shoe photo (shop scenario). There are large visual differences between the street and shop scenario shoe images. To handle the discrepancy of different scenarios, we learn a feature embedding for shoes via a viewpoint-invariant triplet network, the feature activations of which reflect the inherent similarity between any two shoe images. Specifically, we propose a new loss function that minimizes the distances between images of the same shoes captured from different viewpoints. Moreover, we train the proposed triplet network at two different scales so that the representation of shoes incorporates different levels of invariance at different scales. To support training the multi-scale triplet networks, we collect a large dataset with shoe images from the daily life and online shopping websites. Experiments on the dataset show excellence over state-of-the-art approaches, which demonstrate the effectiveness of our proposed method.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ICIP2
2017 Sparsity based reflection removal using external patch search
abstract
Reflection removal aims at separating the mixture of the desired background scenes and the undesired reflections, when the photos are taken through the glass. It has both aesthetic and practical applications which can largely improve the performance of many multimedia tasks. Existing reflection removal approaches heavily rely on scene priors such as separable sparse gradients brought by different levels of blur, and they easily fail when such priors are not observed in many real scenes. Sparse representation models and nonlocal image priors have shown their effectiveness in image restoration with self similarity. In this work, we propose a reflection removal method benefited from the sparsity and nonlocal image prior as a unified optimization framework. We leverage the retrieved image patch from an external database to overcome the limited prior information in the input mixture image and self similarity search. The experimental results show that our proposed model performs better than the existing state-of-the-art reflection removal method for both objective and subjective image qualities.
Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot
ICME2
2017 Fashion analysis with a subordinate attribute classification network
abstract
In this paper we deal with two image-based object search tasks in the fashion domain, clothing attribute prediction and cross-domain shoe retrieval. Clothing attribute prediction is about describing the appearances of clothes via semantic attributes and cross-domain shoe retrieval aims at retrieving the same shoe items from online stores given a daily life shoe photo. We jointly solve these two problems by a novel Subordinate Attribute Convolutional Neural Network (SA-CNN), with the newly designed loss function that systematically merges semantic attributes of closer visual appearance to prevent images with obvious visual differences being confused with each other. A three-level feature representation is further developed based on SA-CNN for shoes from different domains. The experimental results demonstrate that the clothing attribute prediction using the proposed SA-CNN achieves better performance than that using traditional features and fine-tuned conventional CNN. Moreover, for the task of cross-domain shoe retrieval, the top-20 retrieval accuracy with deep features extracted from SA-CNN has a significant improvement of 43% compared to that with the pretrained CNN features.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ICME2
2017 Cross-domain shoe retrieval using a three-level deep feature representation
abstract
In this paper, we address the problem of matching the shoes from the daily life photos to exactly the same shoes from online shops. The problem is extremely challenging because of the significant visual differences between street domain images (shoe images captured in the daily life scenario) and online domain photos (images from online shops taken in the controlled environment). This paper presents a semantic Shoe Attribute-Guided Convolutional Neural Network (SAG-CNN) to extract the deep features. Moreover, we develop a three-level feature representation based on SAG-CNN. The deep features extracted from the image, region and part levels effectively match the images across different domains. We collect a novel shoe dataset, which consists of 8021 street domain and 5821 online domain images. The experimental results on our dataset show that the top-20 retrieval accuracy of our approach improves over that using the pre-trained CNN features by about 40%.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ISCAS2
2017 Hyper-spectral Image Super-resolution Using Non-negative Spectral Representation with Data-Guided Sparsity
abstract
Hyperspectral imaging has great potential for understanding the characteristics of different materials in many applications ranging from remote sensing to medical imaging. However, due to various hardware limitations, only low-resolution hyperspectral and high-resolution multi-spectral images can be available using existing imaging techniques. This study aims to generate a high-resolution hyperspectral image via fusion of the available LR-HS and HR-MS images. We propose a novel hyperspectral image superresolution method via non-negative sparse representation of reflectance spectral with adaptive sparsity constraint. By analyzing local content similarity of a focused pixel in the available high-resolution multi-spectral image, which can measure pixel material purity according to surrounding pixels, we generate a sparsity map for guiding non-negative sparse coding optimization procedure of the spectral representation called non-negative spectral representation with data-guided sparsity. Since the proposed method adaptively adjust the sparsity in the spectral representation based on the local content of the available high-resolution multi-spectral image, it can produce more robust spectral representation for recovering the target high-resolution hyper-spectral image. Comprehensive experiments on two public hyperspectral datasets validate that the proposed method achieves promising performances compared with the existing state of the art methods.
Xianhua Han, Jan Wang, Boxin Shi, Yinqiang Zheng, Yen-Wei Chen 0001
ISM3
2017 Depth Sensing Using Geometrically Constrained Polarization Normals
Achuta Kadambi, Vage Taamazyan, Boxin Shi, Ramesh Raskar
Int. J. Comput. Vis.3
2017 Cross-Domain Shoe Retrieval With a Semantic Hierarchy of Attribute Classification Network
abstract
Cross-domain shoe image retrieval is a challenging problem, because the query photo from the street domain (daily life scenario) and the reference photo in the online domain (online shop images) have significant visual differences due to the viewpoint and scale variation, self-occlusion, and cluttered background. This paper proposes the semantic hierarchy of attribute convolutional neural network (SHOE-CNN) with a three-level feature representation for discriminative shoe feature expression and efficient retrieval. The SHOE-CNN with its newly designed loss function systematically merges semantic attributes of closer visual appearances to prevent shoe images with the obvious visual differences being confused with each other; the features extracted from image, region, and part levels effectively match the shoe images across different domains. We collect a large-scale shoe data set composed of 14341 street domain and 12652 corresponding online domain images with fine-grained attributes to train our network and evaluate our system. The top-20 retrieval accuracy improves significantly over the solution with the pre-trained CNN features.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
IEEE Trans. Image Process.2
2016 A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo
abstract
Recent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we first survey and categorize existing methods using a photometric stereo taxonomy emphasizing on non-Lambertian and uncalibrated methods. We then introduce the 'DiLiGenT' photometric stereo image dataset with calibrated Directional Lightings, objects of General reflectance, and 'ground Truth' shapes (normals). Based on our dataset, we quantitatively evaluate state-of-the-art photometric stereo methods for general non-Lambertian materials and unknown lightings to analyze their strengths and limitations.
Boxin Shi, Zhipeng Mo, Dinglong Duan, Sai-Kit Yeung, Ping Tan 0002
CVPR1
2016 Depth of field guided reflection removal
abstract
Reflection removal aims at separating the mixture of the desired scene and the undesired reflections. Locating reflection and background edges is a key step for reflection removal. In this paper, we present a visual depth guided method to remove reflections. Our idea is to use Depth of Field (DoF) to label the background and reflection edges. We propose a DoF confidence map where pixels with higher DoF values are assumed to belong to the desired background components. Moreover, we observe that images with different resolutions show different properties in the DoF map. Thus, we introduce a multi-scale DoF computing strategy to classify edge pixels more efficiently. Based on the results of edge classification, the background and reflection layers can be separated. Experimental results validate the effectiveness of our method using real-world photos.
Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot
ICIP2
2016 Occluded Imaging with Time-of-Flight Sensors
abstract
We explore the question of whether phase-based time-of-flight (TOF) range cameras can be used for looking around corners and through scattering diffusers. By connecting TOF measurements with theory from array signal processing, we conclude that performance depends on two primary factors: camera modulation frequency and the width of the specular lobe (“shininess”) of the wall. For purely Lambertian walls, commodity TOF sensors achieve resolution on the order of meters between targets. For seemingly diffuse walls, such as posterboard, the resolution is drastically reduced, to the order of 10cm. In particular, we find that the relationship between reflectance and resolution is nonlinear—a slight amount of shininess can lead to a dramatic improvement in resolution. Since many realistic scenes exhibit a slight amount of shininess, we believe that off-the-shelf TOF cameras can look around corners.
Achuta Kadambi, Hang Zhao 0021, Boxin Shi, Ramesh Raskar
ACM Trans. Graph.3
2015 SpecTrans: Versatile Material Classification for Interaction with Textureless, Specular and Transparent Surfaces
abstract
Surface and object recognition is of significant importance in ubiquitous and wearable computing. While various techniques exist to infer context from material properties and appearance, they are typically neither designed for real-time applications nor for optically complex surfaces that may be specular, textureless, and even transparent. These materials are, however, becoming increasingly relevant in HCI for transparent displays, interactive surfaces, and ubiquitous computing. We present SpecTrans, a new sensing technology for surface classification of exotic materials, such as glass, transparent plastic, and metal. The proposed technique extracts optical features by employing laser and multi-directional, multi-spectral LED illumination that leverages the material's optical properties. The sensor hardware is small in size, and the proposed classification method requires significantly lower computational cost than conventional image-based methods, which use texture features or reflectance analysis, thereby providing real-time performance for ubiquitous computing. Our evaluation of the sensing technique for nine different transparent materials, including air, shows a promising recognition rate of 99.0%. We demonstrate a variety of possible applications using SpecTrans' capabilities.
Munehiko Sato, Shigeo Yoshida, Alex Olwal, Boxin Shi, Atsushi Hiyama, Tomohiro Tanikawa, Michitaka Hirose, Ramesh Raskar
CHI4
2015 Unbounded High Dynamic Range Photography Using a Modulo Camera
abstract
This paper presents a novel framework to extend the dynamic range of images called Unbounded High Dynamic Range (UHDR) photography with a modulo camera. A modulo camera could theoretically take unbounded radiance levels by keeping only the least significant bits. We show that with limited bit depth, very high radiance levels can be recovered from a single modulus image with our newly proposed unwrapping algorithm for natural images. We can also obtain an HDR image with details equally well preserved for all radiance levels by merging the least number of modulus images. Synthetic experiment and experiment with a real modulo camera show the effectiveness of the proposed approach.
Hang Zhao 0021, Boxin Shi, Christy Fernandez-Cull, Sai-Kit Yeung, Ramesh Raskar
ICCP2
2015 Polarized 3D: High-Quality Depth Sensing with Polarization Cues
abstract
Coarse depth maps can be enhanced by using the shape information from polarization cues. We propose a framework to combine surface normals from polarization (hereafter polarization normals) with an aligned depth map. Polarization normals have not been used for depth enhancement before. This is because polarization normals suffer from physics-based artifacts, such as azimuthal ambiguity, refractive distortion and fronto-parallel signal degradation. We propose a framework to overcome these key challenges, allowing the benefits of polarization to be used to enhance depth maps. Our results demonstrate improvement with respect to state-of-the-art 3D reconstruction techniques.
Achuta Kadambi, Vage Taamazyan, Boxin Shi, Ramesh Raskar
ICCV3
2015 Photometric Stereo with Small Angular Variations
abstract
Most existing successful photometric stereo setups require large angular variations in illumination directions, which results in acquisition rigs that have large spatial extent. For many applications, especially involving mobile devices, it is important that the device be spatially compact. This naturally implies smaller angular variations in the illumination directions. This paper studies the effect of small angular variations in illumination directions to photometric stereo. We explore both theoretical justification and practical issues in the design of a compact and portable photometric stereo device on which a camera is surrounded by a ring of point light sources. We first derive the relationship between the estimation error of surface normal and the baseline of the point light sources. Armed with this theoretical insight, we develop a small baseline photometric stereo prototype to experimentally examine the theory and its practicality.
Jian Wang 0100, Yasuyuki Matsushita, Boxin Shi, Aswin C. Sankaranarayanan
ICCV3
2014 Photometric Stereo Using Internet Images
abstract
Photometric stereo using unorganized Internet images is very challenging, because the input images are captured under unknown general illuminations, with uncontrolled cameras. We propose to solve this difficult problem by a simple yet effective approach that makes use of a coarse shape prior. The shape prior is obtained from multi-view stereo and will be useful in twofold: resolving the shape-light ambiguity in uncalibrated photometric stereo and guiding the estimated normals to produce the high quality 3D surface. By assuming the surface albedo is not highly contrasted, we also propose a novel linear approximation of the nonlinear camera responses with our normal estimation algorithm. We evaluate our method using synthetic data and demonstrate the surface improvement on real data over multi-view stereo results.
Boxin Shi, Kenji Inose, Yasuyuki Matsushita, Ping Tan 0002, Sai-Kit Yeung, Katsushi Ikeuchi
3DV1
2014 Sub-pixel Layout for Super-Resolution with Images in the Octic Group
Boxin Shi, Hang Zhao 0021, Moshe Ben-Ezra, Sai-Kit Yeung, Christy Fernandez-Cull, R. Hamilton Shepard, Christopher Barsi, Ramesh Raskar
ECCV (1)1
2014 Bi-Polynomial Modeling of Low-Frequency Reflectances
abstract
We present a bi-polynomial reflectance model that can precisely represent the low-frequency component of reflectance. Most existing reflectance models aim at accurately representing the complete reflectance domain for photo-realistic rendering purposes. In contrast, our bi-polynomial model is developed for the purpose of accurately solving inverse problems by effectively discarding the high-frequency component while retaining nonlinear variations in the low-frequency part. The bi-polynomial reflectance model is useful for estimating reflectance and shape of an object. Experimental evaluation in comparison with other parametric reflectance models demonstrates that the proposed model achieves better performance in reflectometry and photometric stereo applications.
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Radiometric Calibration by Rank Minimization
abstract
We present a robust radiometric calibration framework that capitalizes on the transform invariant low-rank structure in the various types of observations, such as sensor irradiances recorded from a static scene with different exposure times, or linear structure of irradiance color mixtures around edges. We show that various radiometric calibration problems can be treated in a principled framework that uses a rank minimization approach. This framework provides a principled way of solving radiometric calibration problems in various settings. The proposed approach is evaluated using both simulation and real-world datasets and shows superior performance to previous approaches.
Joon-Young Lee, Yasuyuki Matsushita, Boxin Shi, In-So Kweon, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 A biquadratic reflectance model for radiometric image analysis
abstract
Radiometric image analysis methods heavily rely on reflectance models. Due to the complexity of real materials, methods based on simple models such as the Lambertian model often suffer from inaccuracy. On the other hand, more advanced models such as the Cook-Torrance model severely complicate the analysis problem. We tackle this dilemma by focusing on the low-frequency component of the reflectance. We propose a compact biquadratic reflectance model to represent the reflectance of a broad class of materials precisely in the low-frequency domain. We validate our model by fitting to both existing parametric models and non-parametric measured data, and show that our model outperforms existing parametric diffuse models. We show applications of reflectometry using general diffuse surfaces and photometric stereo for general isotropic materials. Experimental results show the effectiveness of our biquadratic model and its usefulness in radiometric image analysis.
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
CVPR1
2012 Elevation Angle from Reflectance Monotonicity: Photometric Stereo for General Isotropic Reflectances
Boxin Shi, Ping Tan 0002, Yasuyuki Matsushita, Katsushi Ikeuchi
ECCV (3)1
2011 Radiometric calibration by transform invariant low-rank structure
abstract
We present a robust radiometric calibration method that capitalizes on the transform invariant low-rank structure of sensor irradiances recorded from a static scene with different exposure times. We formulate the radiometric calibration problem as a rank minimization problem. Unlike previous approaches, our method naturally avoids over-fitting problem; therefore, it is robust against biased distribution of the input data, which is common in practice. When the exposure times are completely unknown, the proposed method can robustly estimate the response function up to an exponential ambiguity. The method is evaluated using both simulation and real-world datasets and shows a superior performance than previous approaches.
Joon-Young Lee, Boxin Shi, Yasuyuki Matsushita, In-So Kweon, Katsushi Ikeuchi
CVPR2
2010 Robust Photometric Stereo via Low-Rank Matrix Completion and Recovery
Lun Wu, Arvind Ganesh, Boxin Shi, Yasuyuki Matsushita, Yongtian Wang, Yi Ma 0001
ACCV (3)3
2010 Self-calibrating photometric stereo
abstract
We present a self-calibrating photometric stereo method. From a set of images taken from a fixed viewpoint under different and unknown lighting conditions, our method automatically determines a radiometric response function and resolves the generalized bas-relief ambiguity for estimating accurate surface normals and albedos. We show that color and intensity profiles, which are obtained from registered pixels across images, serve as effective cues for addressing these two calibration problems. As a result, we develop a complete auto-calibration method for photometric stereo. The proposed method is useful in many practical scenarios where calibrations are difficult. Experimental results validate the accuracy of the proposed method using various real-world scenes.
Boxin Shi, Yasuyuki Matsushita, Chao Xu 0006, Ping Tan 0002
CVPR1
2009 Color Correction and Compression for Multi-view Video Using H.264 Features
Boxin Shi, Yangxi Li, Chao Xu 0006
ACCV (3)1
2009 Integrating Color Constancy into Multi-view Video Coding
abstract
Color constancy is the ability to remove the dependency of illuminant and show the intrinsic color of objects. A novel multi-view video coding scheme integrated with color constancy algorithms is introduced in this paper. Color constancy algorithms based on gray-edge hypothesis are selected. The new scheme takes color constancy as a preprocessing to make the sequence independent of the illuminant conditions. Based on the examination of the scene change in the sequence, the key frames are picked out and the parameters of color constancy are determined. The illuminant information is sent into the encoder along with the illuminant independent sequence, and then it is used to solve the color cast problem of multi-view coding at the decoder. Both visual quality and coding performance prove that this novel scheme can lower the bitrates and gain better color quality.
Yangxi Li, Boxin Shi, Chao Xu 0006
ICIG2
2009 Intrinsic Image Decomposition Using Color Invariant Edge
abstract
The intrinsic image composed of reflectance and shading images plays important roles in various computer vision applications. This paper focuses on solving the problem of intrinsic image decomposition. Based on the assumption that the image derivatives can be classified into either reflectance-related or shading-related, the reflectance and shading image can be restored from the classified derivatives. We improve the classification result using only color information by introducing the color invariant edge. Considering some color invariant properties in the image, the color invariant edge can provide more useful information in guiding the classification and producing more robust decomposition result as it is shown in the experiment.
Boxin Shi, Yangxi Li, Chao Xu 0006
ICIG1
2009 Block-based color correction algorithm for multi-view video coding
abstract
The color variations among different viewpoints in multiview video sequences may deteriorate the visual quality and coding efficiency. Various color correction methods have been proposed, however, the color appearance and histogram of corrected target frames are not similar enough to the reference frames in details. Focusing on restoring more similar color, a block-based color correction algorithm is proposed. The blocks in reference frames are matched into target frames through spatial prediction, and the colorization scheme is then adopted to expand color as a coarse correction. Finally the mixture with global color transfer result yields the fine correction. The experiment results show this novel method can provide better visual effect in detail and also provide the corrected frames with histograms more similar to reference histograms.
Boxin Shi, Yangxi Li, Chao Xu 0006
ICME1
2008 Comparison between JPEG2000 and H.264 for digital cinema
abstract
JPEG2000 and H.264 are the latest image and video coding standards respectively. Digital cinema is a new kind of application for super high definition video. The DCI (Digital Cinema Initiative) specification published in 2005 has selected JPEG2000 instead of H.264 as the video coding standard for digital cinema. It shows that JPEG2000 has a better performance in the field of super high definition video coding. Until now, only a few basic tests have been done to compare between JPEG2000 and H.264. Moreover, the test resolutions and characteristics of input sequences are very limited. In this paper, based on JPEG2000, H.264 intra and inter-frame coding, we compare the coding efficiency and subjective image quality on multiple series of test sequences from low to super high resolutions. The experiment results demonstrate some regularity for JPEG2000 and H.264 video coding, and reveal JPEG2000 is more suitable for digital cinema.
Boxin Shi, Chao Xu 0006
ICME1