VLDB 2026 Research / reviewers in the wild / expert
Bingyao Huang
dblp:155/4775 · also BingYao Huang
· DBLP profile ↗
24ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-8647-5730ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text Recovery Attacks and Defenses on Physical EnvelopesabstractRecovering text from sealed envelopes poses a severe security threat to physical privacy. While recent imaging advances can see through paper layers to recover general textures, they fail to reconstruct legible text due to the high-frequency detail and semantic structure required for OCR. To address this, we present the first unified pipeline for Physical Envelope Text Recovery (PETR), covering both attack and defense mechanisms. On the attack side, we design two entirely different recovery architectures: ETNN and ETFormer. In particular, we propose a Text-Aware Loss that injects text-specific guidance into the restoration process, where constraints on both text content and text structure jointly drive gradient back-propagation to enhance structural clarity. We also introduce an LLM-based evaluation protocol to measure the semantic readability of recovered content. On the defense side, we analyze these attack vectors to design secure envelopes using localized disruption patterns that impede recovery without affecting human readability after opening. Extensive experiments show our models outperform state-of-the-art baselines in both fidelity (PSNR/SSIM) and OCR accuracy. The source code and dataset are available on the project page https://github.com/YongfTao/PETR. Yongfei Tao, Yi Zhang 0181, Runze Liao, Jingwei Qu, Bingyao Huang |
ICMR | 5 |
| 2026 | ProCap: Projection-Aware Captioning for Spatial Augmented RealityabstractSpatial augmented reality (SAR) directly projects digital content onto physical scenes using projectors, creating immersive experience without head-mounted displays. However, for SAR to support intelligent interaction, such as reasoning about the scene or answering user queries, it must semantically distinguish between the physical scene and the projected content. Standard Vision Language Models (VLMs) struggle with this virtual-physical ambiguity, often confusing the two contexts. To address this issue, we introduce ProCap, a novel framework that explicitly decouples projected content from physical scenes. ProCap employs a two-stage pipeline: first it visually isolates virtual and physical layers via automated segmentation; then it uses region-aware retrieval to avoid ambiguous semantic context due to projection distortion. To support this, we present RGBP (RGB + Projections), the first large-scale SAR semantic benchmark dataset, featuring 65 diverse physical scenes and over 180,000 projections with dense, decoupled annotations. Finally, we establish a dual-captioning evaluation protocol using task-specific tokens to assess physical scene and projection descriptions independently. Our experiments show that ProCap provides a robust semantic foundation for future SAR research. The source code, pre-trained models and the RGBP dataset are available on the project page: https://ZimoCao.github.io/ProCap/. Zimo Cao, Yuchen Deng, Haibin Ling, Bingyao Huang |
VR | 4 |
| 2026 | E2SL: Efficient Depth Sensing from Event-Based Structured LightabstractStructured light (SL) is a popular approach for 3D reconstruction. Most SL techniques rely on frame-based cameras and are often not robust in high-speed dynamic scenes. Recently, event cameras have sparked growing interest in high-speed SL imaging, due to their high temporal resolution. The event-based SL enjoys the high-speed data acquisition, however, most existing methods tend to pursue the reconstruction accuracy but sacrificing the computational efficiency, limiting the applicability in real-world scenarios. To this end, we propose E2SL, an Efficient deep network tailored for monocular Event-based SL. Specifically, E2SL comprises three key components: binary embedding lookup table (BE-LUT), spatial context enhancement (SCE), and geometric-prior regression (GPR). Given the input event frame, BE-LUT, which is precomputed and stored, first retrieves the features efficiently. Then, SCE extends the receptive field of the features and captures the spatial context. Finally, GPR conducts the geometric-prior-based tree classification for fast and robust depth estimation. To support training and evaluation, we contribute an event-based SL simulator, which generates a large-scale and diverse synthetic dataset. Besides, we develop an event-based SL prototype and collect a dataset with accurate ground truth for real-world evaluation. Extensive experiments demonstrate that our method achieves state-of-the-art accuracy while maintaining a per-frame reconstruction time of 7.7 ms, meeting the demands of high-speed depth sensing. The code and dataset are available on the project page https://dongxin000.github.io/E2SL/. Jiacheng Fu, Wenming Weng, Yueyi Zhang 0001, Bingyao Huang, Zhiwei Xiong |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | Setup-Independent Full Projector CompensationabstractProjector compensation seeks to correct geometric and photometric distortions that occur when images are projected onto nonplanar or textured surfaces. However, most existing methods are highly setup-dependent, requiring fine-tuning or retraining whenever the surface, lighting, or projector-camera pose changes. Progress has been limited by two key challenges: (1) the absence of large, diverse training datasets and (2) existing geometric correction models are typically constrained by specific spatial setups; without further retraining or fine-tuning, they often fail to generalize directly to novel geometric configurations. We introduce SIComp, the first Setup-Independent framework for full projector Compensation, capable of generalizing to unseen setups without fine-tuning or retraining. To enable this, we construct a large-scale real-world dataset spanning 277 distinct projector-camera setups. SIComp adopts a co-adaptive design that decouples geometry and photometry: A carefully tailored optical flow module performs online geometric correction, while a novel photometric network handles photometric compensation. To further enhance robustness under varying illumination, we integrate intensity-varying surface priors into the network design. Extensive experiments demonstrate that SIComp consistently produces high-quality compensation across diverse unseen setups, substantially outperforming existing methods in terms of generalization ability and establishing the first generalizable solution to projector compensation. The code and dataset are available on our project page: https://hai-bo-li.github.io/SIComp/. Qingyue Deng, Jijiang Li, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | DiffPC: Diffusion-Based Projector Photometric CompensationabstractProjector photometric compensation corrects color distortions introduced by surface texture, reflection, and ambient lighting. Existing deep learning-based methods usually require professional scene-specific data collection and lack consideration for perceptual quality. To address this limitation, we present a diffusion-based photometric compensation method that reconstructs compensation images under photometric and content-aware guidance. Specifically, we fi rst mo del th e ph otometric distortions introduced during projection as environment-dependent additive noise, thereby reformulating the photometric compensation problem as a denoising task with physical constraints. Next, we introduce a diffusion model, which generates compensation images by following an additive trajectory to iteratively remove the noise. Finally, to accurately estimate the noise at each timestep, by analyzing the factors that contribute to distortions in the physical process of projection and capturing, we design a noise estimation network that incorporates features of both photometry-aware and content conditions. Experiments show that our method achieves superior visual performance in unknown scenarios, thereby exhibiting significant practical advantages over prior art. Our source code is available at https://github.com/cyxwang/DiffPC. Yuxi Wang 0002, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Mixture of Cluster-Guided Experts for Retrieval-Augmented Label PlacementabstractText labels are widely used to convey auxiliary information in visualization and graphic design. The substantial variability in the categories and structures of labeled objects leads to diverse label layouts. Recent single-model learning-based solutions in label placement struggle to capture fine-grained differences between these layouts, which in turn limits their performance. In addition, although human designers often consult previous works to gain design insights, existing label layouts typically serve merely as training data, limiting the extent to which embedded design knowledge can be exploited. To address these challenges, we propose a mixture of cluster-guided experts (MoCE) solution for label placement. In this design, multiple experts jointly refine layout features, with each expert responsible for a specific cluster of layouts. A cluster-based gating function assigns input samples to experts based on representation clustering. We implement this idea through the Label Placement Cluster-guided Experts (LPCE) model, in which a MoCE layer integrates multiple feed-forward networks (FFNs), with each expert composed of a pair of FFNs. Furthermore, we introduce a retrieval augmentation strategy into LPCE, which retrieves and encodes reference layouts for each input sample to enrich its representations. Extensive experiments demonstrate that LPCE achieves superior performance in label placement, both quantitatively and qualitatively, surpassing a range of state-of-the-art baselines. Our algorithm is available at https://github.com/PingshunZhang/LPCE. Pingshun Zhang, Enyu Che, Bingyao Huang, Haibin Ling, Jingwei Qu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | CAPAA: Classifier-Agnostic Projector-Based Adversarial AttackabstractProjector-Based adversarial attack aims to project carefully designed light patterns (i.e., adversarial projections) onto scenes to deceive deep image classifiers. It has potential applications in privacy protection and the development of more robust classifiers. However, existing approaches primarily focus on individual classifiers and fixed camera poses, often neglecting the complexities of multi-classifier systems and scenarios with varying camera poses. This limitation reduces their effectiveness when introducing new classifiers or camera poses. In this paper, we introduce Classifier-Agnostic Projector-Based Adversarial Attack (CAPAA) to address these issues. First, we develop a novel classifier-agnostic adversarial loss and optimization framework that aggregates adversarial and stealthiness loss gradients from multiple classifiers. Then, we propose an attention-based gradient weighting mechanism that concentrates perturbations on regions of high classification activation, thereby improving the robustness of adversarial projections when applied to scenes with varying camera poses. Our extensive experimental evaluations demonstrate that CAPAA achieves both a higher attack success rate and greater stealthiness compared to existing baselines. Codes are available at: https://github.com/ZhanLiQxQ/CAPAA. Haibin Ling, Bingyao Huang |
ICME | 5 |
| 2025 | MAD-paint: Mask-Aware Diffusion Sampling for Image InpaintingabstractImage inpainting aims to repair digital images with defects such as holes and scratches at both semantic and textural levels. Diffusion models have shown great success in image inpainting, delivering high-quality results. However, existing diffusion-based methods often overlook the shape of defective regions/masks, applying a uniform sampling strategy across varying shapes. This oversight may lead to low-quality or semantically inappropriate restored images. In this paper, we propose MAD-paint (Mask-Aware Diffusion sampling for inpainting), and show that applying different noise types tailored to specific defect regions/mask shapes during the reverse diffusion process can significantly improve the inpainting quality. We begin by introducing a metric for mask uncertainty to assess the impact of different masks on inpainting quality. Using this metric, we propose a mask-aware sampling approach that automatically adjusts its sampling strategy according to different mask shapes, as indicated by the mask uncertainty. In addition, leveraging the known image texture consistency, we propose a known region-guided iterative refinement mechanism to condition texture restoration. The experimental results demonstrate the advantages of our method over other diffusion-based inpainting methods. Shipeng Jiang, Jingwei Qu, Bingyao Huang |
ICMR | 3 |
| 2025 | NeuroPump: Simultaneous Geometric and Color Rectification for Underwater ImagesabstractUnderwater image restoration aims to remove geometric and color distortions due to water refraction, absorption, and scattering. Previous studies focus on restoring either color or geometry, but to our best knowledge, not both. However, in practice it may be cumbersome to address the two rectifications one by one. In this paper, we propose NeuroPump, a self-supervised method to simultaneously optimize and rectify underwater geometry and color as if water were pumped out. The key idea is to explicitly model refraction, absorption, and scattering in Neural Radiance Field (NeRF) pipeline, such that it not only performs simultaneous geometric and color rectification, but also enables to synthesize novel views and optical effects by controlling the decoupled parameters. In addition, to address the lack of real paired ground truth images, we propose an underwater 360 benchmark dataset that has real paired (i.e., with and without water) images. Our method clearly outperforms other baselines both quantitatively and qualitatively. Our code and dataset is available at https://ygswu.github.io/NeuroPump.github.io/. Haoxiang Liao, Haibin Ling, Bingyao Huang |
ACM Multimedia | 4 |
| 2025 | LAPIG: Language Guided Projector Image Generation with Surface Adaptation and StylizationabstractWe propose LAPIG, a language guided projector image generation method with surface adaptation and stylization. LAPIG consists of a projector-camera system and a target textured projection surface. LAPIG takes the user text prompt as input and aims to transform the surface style using the projector. LAPIG's key challenge is that due to the projector's physical brightness limitation and the surface texture, the viewer's perceived projection may suffer from color saturation and artifacts in both dark and bright regions, such that even with the state-of-the-art projector compensation techniques, the viewer may see clear surface texture-related artifacts. Therefore, how to generate a projector image that follows the user's instruction while also displaying minimum surface artifacts is an open problem. To address this issue, we propose projection surface adaptation (PSA) that can generate compensable surface stylization. We first train two networks to simulate the projector compensation and project-and-capture processes, this allows us to find a satisfactory projector image without real project-and-capture and utilize gradient descent for fast convergence. Then, we design content and saturation losses to guide the projector image generation, such that the generated image shows no clearly perceivable artifacts when projected. Finally, the generated image is projected for visually pleasing surface style morphing effects. The source code and more results are available on the project page: https://Yu-chen-Deng.github.io/LAPIG/. Yuchen Deng, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | GS-ProCams: Gaussian Splatting-Based Projector-Camera SystemsabstractWe present GS-ProCams, the first Gaussian Splatting-based framework for projector-camera systems (ProCams). GS-ProCams is not only view-agnostic but also significantly enhances the efficiency of projection mapping (PM) that requires establishing geometric and radiometric mappings between the projector and the camera. Previous CNN-based ProCams are constrained to a specific viewpoint, limiting their applicability to novel perspectives. In contrast, NeRF-based ProCams support view-agnostic projection mapping, however, they require an additional co-located light source and demand significant computational and memory resources. To address this issue, we propose GS-ProCams that employs 2D Gaussian for scene representations, and enables efficient view-agnostic ProCams applications. In particular, we explicitly model the complex geometric and photometric mappings of ProCams using projector responses, the projection surface's geometry and materials represented by Gaussians, and the global illumination component. Then, we employ differentiable physically-based rendering to jointly estimate them from captured multi-view projections. Compared to state-of-the-art NeRF-based methods, our GS-ProCams eliminates the need for additional devices, achieving superior ProCams simulation quality. It also uses only 1/10 of the GPU memory for training and is 900 times faster in inference speed. Please refer to our project page for the code and dataset: https://realqingyue.github.io/GS-ProCams/. Qingyue Deng, Jijiang Li, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | DPCS: Path Tracing-Based Differentiable Projector-Camera SystemsabstractProjector-camera systems (ProCams) simulation aims to model the physical project-and-capture process and associated scene parameters of a ProCams, and is crucial for spatial augmented reality (SAR) applications such as ProCams relighting and projector compensation. Recent advances use an end-to-end neural network to learn the project-and-capture process. However, these neural network-based methods often implicitly encapsulate scene parameters, such as surface material, gamma, and white balance in the network parameters, and are less interpretable and hard for novel scene simulation. Moreover, neural networks usually learn the indirect illumination implicitly in an image-to-image translation way which leads to poor performance in simulating complex projection effects such as soft-shadow and interreflection. In this paper, we introduce a novel path tracing-based differentiable projector-camera systems (DPCS), offering a differentiable ProCams simulation method that explicitly integrates multi-bounce path tracing. Our DPCS models the physical project-and-capture process using differentiable physically-based rendering (PBR), enabling the scene parameters to be explicitly decoupled and learned using much fewer samples. Moreover, our physically-based method not only enables high-quality downstream ProCams tasks, such as ProCams relighting and projector compensation, but also allows novel scene simulation using the learned scene parameters. In experiments, DPCS demonstrates clear advantages over previous approaches in ProCams simulation, offering better interpretability, more efficient handling of complex interreflection and shadow, and requiring fewer training samples. The code and dataset are available on the project page: https://jijiangli.github.io/DPCS/. Jijiang Li, Qingyue Deng, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | ViComp: Video Compensation for Projector-Camera SystemsabstractProjector video compensation aims to cancel the geometric and photometric distortions caused by non-ideal projection surfaces and environments when projecting videos. Most existing projector compensation methods start by projecting and capturing a set of sampling images, followed by an offline compensation model training step. Thus, abundant user effort is required before the users can watch the video. Moreover, the sampling images have little prior knowledge of the video content and may lead to suboptimal results. To address these issues, this paper builds a video compensation system that can online adapt the compensation parameters. Our approach consists of five threads and can perform compensation, projection, capturing, and short-term and long-term model updates in parallel. Due to the parallel mechanism, rather than projecting and capturing hundreds of sampling images and training the model offline, we can directly use the projected and captured video frames for model updates on the fly. To quickly apply to the new environment, we introduce a deep learning-based compensation model that integrates a fixed transformer-based method and a novel CNN-based network. Moreover, for fast convergence and to reduce error accumulation during fine-tuning, we present a strategy that cooperates with short-term and long-term memory model updates. Experiments show that it significantly outperforms state-of-the-art baselines. Yuxi Wang 0002, Haibin Ling, Bingyao Huang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Adaptive Color Structured Light for Calibration and Shape ReconstructionabstractColor structured light (SL) plays an important role in spatial augmented reality and shape reconstruction. Compared to traditional non-color multi-shot SL, it has the advantage of fewer projections, and can even achieve single-shot. However, distortions caused by ambient light and imaging devices limit color SL’s applicability and accuracy. A common solution is to apply color adaptation techniques to cancel the disturbances. Previous studies focus on either robust fixed color patterns or adaptation approaches that may require preliminary geometric calibrations. In this paper, we propose an approach that can efficiently adapt color SL to arbitrary ambient light and imaging devices’ color responses, without device response function calibration or geometric calibration. First, we design a novel algorithm to quickly find the most distinct colors that are easily separable under a new environment and device setup. Then, we design a maximum a posteriori (MAP)-based color detection algorithm that can utilize ambient light and device priors to robustly detect the SL colors. In experiments, our adaptive color SL outperforms previous methods in both calibration and shape reconstruction tasks across a variety of setups. Haibin Ling, Bingyao Huang |
ISMAR | 3 |
| 2023 | CompenHR: Efficient Full Compensation for High-resolution ProjectorabstractFull projector compensation is a practical task of projector-camera systems. It aims to find a projector input image, named compensation image, such that when projected it cancels the geometric and photometric distortions due to the physical environment and hardware. State-of-the-art methods use deep learning to address this problem and show promising performance for low-resolution setups. However, directly applying deep learning to high-resolution setups is impractical due to the long training time and high memory cost. To address this issue, this paper proposes a practical full compensation solution. Firstly, we design an attention-based grid refinement network to improve geometric correction quality. Secondly, we integrate a novel sampling scheme into an end-to-end compensation network to alleviate computation and introduce attention blocks to preserve key features. Finally, we construct a benchmark dataset for high-resolution projector full compensation. In experiments, our method demonstrates clear advantages in both efficiency and quality. Yuxi Wang 0002, Haibin Ling, Bingyao Huang |
VR | 3 |
| 2022 | SPAA: Stealthy Projector-based Adversarial Attacks on Deep Image ClassifiersabstractLight-based adversarial attacks use spatial augmented reality (SAR) techniques to fool image classifiers by altering the physical light condition with a controllable light source, e.g., a projector. Compared with physical attacks that place hand-crafted adversarial objects, projector-based ones obviate modifying the physical entities, and can be performed transiently and dynamically by altering the projection pattern. However, subtle light perturbations are insufficient to fool image classifiers, due to the complex environment and project-and-capture process. Thus, existing approaches focus on projecting clearly perceptible adversarial patterns, while the more interesting yet challenging goal, stealthy projector-based attack, remains open. In this paper, for the first time, we formulate this problem as an end-to-end differentiable process and propose a Stealthy Projector-based Adversarial Attack (SPAA) solution. In SPAA, we approximate the real Project-and-Capture process using a deep neural network named PCNet, then we include PCNet in the optimization of projector-based attacks such that the generated adversarial projection is physically plausible. Finally, to generate both robust and stealthy adversarial projections, we propose an algorithm that uses minimum perturbation and adversarial confidence thresholds to alternate between the adversarial loss and stealthiness loss optimization. Our experimental evaluations show that SPAA clearly outperforms other methods by achieving higher attack success rates and meanwhile being stealthier, for both targeted and untargeted attacks. Bingyao Huang, Haibin Ling |
VR | 1 |
| 2022 | End-to-End Full Projector CompensationabstractFull projector compensation aims to modify a projector input image to compensate for both geometric and photometric disturbance of the projection surface. Traditional methods usually solve the two parts separately and may suffer from suboptimal solutions. In this paper, we propose the first end-to-end differentiable solution, named CompenNeSt++, to solve the two problems jointly. First, we propose a novel geometric correction subnet, named WarpingNet, which is designed with a cascaded coarse-to-fine structure to learn the sampling grid directly from sampling images. Second, we propose a novel photometric compensation subnet, named CompenNeSt, which is designed with a siamese architecture to capture the photometric interactions between the projection surface and the projected images, and to use such information to compensate the geometrically corrected images. By concatenating WarpingNet with CompenNeSt, CompenNeSt++ accomplishes full projector compensation and is end-to-end trainable. Third, to improve practicability, we propose a novel synthetic data-based pre-training strategy to significantly reduce the number of training images and training time. Moreover, we construct the first setup-independent full compensation benchmark to facilitate future studies. In thorough experiments, our method shows clear advantages over prior art with promising compensation quality and meanwhile being practically convenient. Bingyao Huang, Tao Sun 0009, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Modeling Deep Learning Based Privacy Attacks on Physical Mail
Bingyao Huang, Ruyi Lian, Dimitris Samaras, Haibin Ling |
AAAI | 1 |
| 2021 | A Fast and Flexible Projector-Camera Calibration SystemabstractExisting projector-camera calibration methods typically warp keypoints from a camera image to a projector image using estimated homographies and often suffer from errors in camera parameters and noises due to imperfect planarity of the calibration target. This article proposes a practical and robust projector-camera calibration system that explicitly deals with these challenges. First, a graph-theory-based correspondence algorithm is built on top of a color-coded spatial structured light (SL) pattern. Such SL correspondences are then used for a coarse projector-camera calibration. To gain more robustness against noises from an imperfect planar calibration board, we develop a bundle adjustment algorithm to jointly optimize the estimated projector-camera parameters and the correspondences’ coordinates. Moreover, our system requires only one shot of an SL pattern for each calibration board pose, which is much more practical than multishot solutions. Comprehensive experimental validation is conducted on both synthetic and real data sets, and our method clearly outperforms the existing methods in all experiments. For the benefit of the society, a practical open-source software with graphical user interface (GUI) of the developed system is publicly available athttps://github.com/bingyaohuang/single-shot-pro-cam-calib.Note to Practitioners—The proposed method is motivated by two challenges in industrial structured light (SL) system calibration: 1) robustness against imperfect planarity of the calibration target and 2) the number of SL projections per pose. In many industrial SL-based 3-D reconstruction systems, the calibration accuracy greatly affects the reconstruction reliability. Our SL calibration system explicitly deals with calibration target’s imperfect planarity and thus outperforms the existing methods in terms of system accuracy. Another advantage of our SL calibration system is single-shot-per-pose, allowing fast recalibration and reducing the decoding error due to slight pattern misalignment in multishot methods[37]. In addition, we release the open-source calibration software with a graphical user interface (GUI), with which calibration and sparse 3-D reconstruction can be easily performed without any further instructions. Moreover, considering the complex calibration environment and setup, we make the camera and projector imaging parameters, such as exposure, brightness, and contrast, adjustable through widgets and preview. Finally, a limitation of our color-coded SL system is its sensitivity to environment lighting and target texture. This problem may be solved by projector photometric compensation[16],[18],[19],[39]. Bingyao Huang, Ying Tang 0001, Samed Ozdemir, Haibin Ling |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2021 | DeProCams: Simultaneous Relighting, Compensation and Shape Reconstruction for Projector-Camera SystemsabstractImage-based relighting, projector compensation and depth/normal reconstruction are three important tasks of projector-camera systems (ProCams) and spatial augmented reality (SAR). Although they share a similar pipeline of finding projector-camera image mappings, in tradition, they are addressed independently, sometimes with different prerequisites, devices and sampling images. In practice, this may be cumbersome for SAR applications to address them one-by-one. In this paper, we propose a novel end-to-end trainable model named DeProCams to explicitly learn the photometric and geometric mappings of ProCams, and once trained, DeProCams can be applied simultaneously to the three tasks. DeProCams explicitly decomposes the projector-camera image mappings into three subprocesses: shading attributes estimation, rough direct light estimation and photorealistic neural rendering. A particular challenge addressed by DeProCams is occlusion, for which we exploit epipolar constraint and propose a novel differentiable projector direct light mask. Thus, it can be learned end-to-end along with the other modules. Afterwards, to improve convergence, we apply photometric and geometric constraints such that the intermediate results are plausible. In our experiments, DeProCams shows clear advantages over previous arts with promising quality and meanwhile being fully differentiable. Moreover, by solving the three tasks in a unified model, DeProCams waives the need for additional optical devices, radiometric calibrations and structured light. Bingyao Huang, Haibin Ling |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | End-To-End Projector Photometric CompensationabstractProjector photometric compensation aims to modify a projector input image such that it can compensate for disturbance from the appearance of projection surface. In this paper, for the first time, we formulate the compensation problem as an end-to-end learning problem and propose a convolutional neural network, named CompenNet, to implicitly learn the complex compensation function. CompenNet consists of a UNet-like backbone network and an autoencoder subnet. Such architecture encourages rich multi-level interactions between the camera-captured projection surface image and the input image, and thus captures both photometric and environment information of the projection surface. In addition, the visual details and interaction information are carried to deeper layers along the multi-level skip convolution layers. The architecture is of particular importance for the projector compensation task, for which only a small training dataset is allowed in practice. Another contribution we make is a novel evaluation benchmark, which is independent of system setup and thus quantitatively verifiable. Such benchmark is not previously available, to our best knowledge, due to the fact that conventional evaluation requests the hardware system to actually project the final results. Our key idea, motivated from our end-to-end problem formulation, is to use a reasonable surrogate to avoid such projection process so as to be setup-independent. Our method is evaluated carefully on the benchmark, and the results show that our end-to-end learning solution outperforms state-of-the-arts both qualitatively and quantitatively by a significant margin. Bingyao Huang, Haibin Ling |
CVPR | 1 |
| 2019 | CompenNet++: End-to-End Full Projector CompensationabstractFull projector compensation aims to modify a projector input image such that it can compensate for both geometric and photometric disturbance of the projection surface. Traditional methods usually solve the two parts separately, although they are known to correlate with each other. In this paper, we propose the first end-to-end solution, named CompenNet++, to solve the two problems jointly. Our work non-trivially extends CompenNet, which was recently proposed for photometric compensation with promising performance. First, we propose a novel geometric correction subnet, which is designed with a cascaded coarse-to-fine structure to learn the sampling grid directly from photometric sampling images. Second, by concatenating the geometric correction subset with CompenNet, CompenNet++ accomplishes full projector compensation and is end-to-end trainable. Third, after training, we significantly simplify both geometric and photometric compensation parts, and hence largely improves the running time efficiency. Moreover, we construct the first setup-independent full compensation benchmark to facilitate the study on this topic. In our thorough experiments, our method shows clear advantages over previous arts with promising compensation quality and meanwhile being practically convenient. Bingyao Huang, Haibin Ling |
ICCV | 1 |
| 2016 | Dynamic Behavior of Artificial Hodgkin-Huxley Neuron Model Subject to Additive NoiseabstractMotivated by neuroscience discoveries during the last few years, many studies consider pulse-coupled neural networks with spike-timing as an essential component in information processing by the brain. There also exists some technical challenges while simulating the networks of artificial spiking neurons. The existing studies use a Hodgkin-Huxley (H-H) model to describe spiking dynamics and neuro-computational properties of each neuron. But they fail to address the effect of specific non-Gaussian noise on an artificial H-H neuron system. This paper aims to analyze how an artificial H-H neuron responds to add different types of noise using an electrical current and subunit noise model. The spiking and bursting behavior of this neuron is also investigated through numerical simulations. In addition, through statistic analysis, the intensity of different kinds of noise distributions is discussed to obtain their relationship with the mean firing rate, interspike intervals, and stochastic resonance. Qi Kang 0001, Bingyao Huang, MengChu Zhou |
IEEE Trans. Cybern. | 2 |
| 2014 | Fast 3D reconstruction using one-shot spatial structured lightabstractStructured light gains its popularity in 3D reconstruction applications due to its robustness against outliers. In the last few decades, a number of high-accuracy temporally encoded structured light emerged to solve the 3D reconstruction problem. However their applications are mainly limited to scanning stationary objects. When dealing with dynamic scenes and real-time data acquisition, one-shot spatial multiplexed structured light has the speed advantage. In this paper, we propose a fast 3D reconstruction method using one-shot special structured light. It works by projecting a static two-dimensional 8-color De Bruijn spatial grid pattern onto the scene, analyzing the deformation of the observed light pattern with respect to the projected one, and identifying their correspondence. Several local optimization strategies are used to offer a confident solution, including special vote majority for color detection and correction, and De Bruijn-based Hamming distance minimization to improve intersection neighborhood information. The effectiveness of the proposed method is verified through 3D reconstruction of a complicated bust. Bingyao Huang, Ying Tang 0001 |
SMC | 1 |