EDBT 2026 Demo / reviewers in the wild / expert
Gaochang Wu
dblp:156/2555
· DBLP profile ↗
23ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-5149-2995ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HCLIP-AD: Calibrating text-image foundation models with hierarchical semantic alignment for zero-shot anomaly detection
Gaochang Wu, Jinliang Ding |
Pattern Recognit. | 2 |
| 2026 | Prescribed-Time Fuzzy Adaptive Consensus Control for Photovoltaic Systems With Dead-Zone Input and Actuator FaultsabstractThis article presents a novel prescribed-time fuzzy adaptive consensus control scheme for nonlinear photovoltaic (PV) energy systems with dead-zone and actuator faults. The PV panels are controlled to track the maximum power point so that the considered systems can efficiently operate under different complex conditions. To quickly regulate the voltage to a desired reference, a novel prescribed-time performance function is presented to guarantee systems states convergence within a specified time. Besides, dead zones and faults are serious nonlinearity constraints in practical grid-connected PV systems. Thus, a finite-time tracking controller and the adaptive laws are designed to compensate for the effect of unknown nonlinear constraints. By applying fuzzy logic systems, the difficulty of dealing with unknown nonlinear dynamics is overcome. The stability of the closed-loop system is rigorously proven using Lyapunov stability theory, demonstrating the uniformly ultimately boundedness of all signals. Finally, extensive simulation experiments validate the effectiveness and superiority of the proposed approach. Zilong Tan, Gaochang Wu |
IEEE Trans. Cybern. | 2 |
| 2026 | XMatchAD: A Cross-Modal Matching Perspective on Reconstruction-Based Anomaly DetectionabstractThe remarkable success of reconstruction-based methods in Unsupervised Anomaly Detection (UAD) lies in their ability to identify and localize anomalies by modeling discrepancies between input images and their reconstructed counterparts. However, these approaches often struggle to capture subtle anomalies and tend to produce blurred anomaly boundaries, which significantly limits their effectiveness, particularly in complex multi-class scenarios. To address these issues, we present XMatchAD, a novel UAD framework that reinterprets the task from a pseudo cross-modal matching perspective. Specifically, the input and reconstructed images are treated as two complementary modalities and their matching relationships are precisely exploited for anomaly detection. First, a pre-trained feature extractor is employed to encode discriminative representations. Second, an attention-guided cross-modal matching mechanism is introduced to match local inter-modal anomaly-related patterns while mutually refining the features. This enhances the sensitivity to anomalies with diverse shapes and subtle deviations and significantly improves the precision of anomaly detection and localization. Third, we design an adaptive frequency-aware fusion module that further delineates sharp anomaly boundaries through the coupling of high-frequency components from cross-modal multi-scale representations. Comprehensive evaluations on MVTec-AD, VisA, and MPDD benchmarks demonstrate that our method consistently achieves superior performance, outperforming state-of-the-art methods in multi-class anomaly detection and localization. The code will be released at https://github.com/Mingxiu-Cai/XMatchAD. Mingxiu Cai, Gaochang Wu, Tianyou Chai |
IEEE Trans. Image Process. | 3 |
| 2026 | DuaDiff: Dual-Conditional Diffusion Model for Guided Thermal Image Super-ResolutionabstractThermal imaging offers valuable properties, but suffers from inherently low spatial resolution, which can be enhanced using a high-resolution (HR) visible image as guidance. However, the substantial modality differences between thermal and visible images, coupled with significant resolution gaps, pose challenges to existing guided super-resolution (SR) approaches. In this article, we present dual-conditional diffusion (DuaDiff), an innovative diffusion model featuring a dual-conditioning mechanism to enhance guided thermal image SR. Unlike typical conditional diffusion models, DuaDiff integrates a learnable Laplacian pyramid to extract high-frequency details from the visible image, serving as one of the conditioning inputs. By capturing multiscale high-frequency components, DuaDiff effectively focuses on intricate textures and edges in the HR visible images, significantly enhancing thermal image fidelity. Furthermore, we project both thermal and visible images into a semantic latent space, constructing another conditioning input. Leveraging these complementary conditions, DuaDiff employs a multimodal latent feature cross-attention module to facilitate effective interaction between noise, thermal, and visible latent representations. Extensive experiments on the FLIR-ADAS and CATS datasets for $4\times $ and $8\times $ guided SR demonstrate that combining learnable Laplacian conditioning with semantic latent conditioning enables DuaDiff to surpass state-of-the-art methods in both visual quality and metric evaluation, particularly in scenarios with a large resolution gap. Besides, the applications to downstream tasks further confirm the capability of DuaDiff to recover high-fidelity semantic information. The code will be released. Linrui Shi, Gaochang Wu, Yingqian Wang 0002, Yebin Liu, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | CostFilter-AD: Enhancing Anomaly Detection through Matching Cost FilteringabstractUnsupervised anomaly detection (UAD) seeks to localize the anomaly mask of an input image with respect to normal samples. Either by reconstructing normal counterparts (reconstruction-based) or by learning an image feature embedding space (embedding-based), existing approaches fundamentally rely on image-level or feature-level matching to derive anomaly scores. Often, such a matching process is inaccurate yet overlooked, leading to sub-optimal detection. To address this issue, we introduce the concept of cost filtering, borrowed from classical matching tasks, such as depth and flow estimation, into the UAD problem. We call this approach CostFilter-AD. Specifically, we first construct a matching cost volume between the input and normal samples, comprising two spatial dimensions and one matching dimension that encodes potential matches. To refine this, we propose a cost volume filtering network, guided by the input observation as an attention query across multiple feature layers, which effectively suppresses matching noise while preserving edge structures and capturing subtle anomalies. Designed as a generic post-processing plug-in, CostFilter-AD can be integrated with either reconstruction-based or embedding-based methods. Extensive experiments on MVTec-AD and VisA benchmarks validate the generic benefits of CostFilter-AD for both single- and multi-class UAD tasks. Code and models will be released at https://github.com/ZHE-SAPI/CostFilter-AD. Mingxiu Cai, Gaochang Wu, Tianyou Chai, Xiatian Zhu |
ICML | 4 |
| 2025 | Lightweight Steel Surface Defect Detection via Knowledge DistillationabstractDeploying steel plate surface defect detection models to edge devices necessitates lightweight and real-time capabilities. To address this, we propose an Ultra Lite Residual Module with a smaller parameter count, which we then use to construct an Ultra Lite ResNet (UL-ResNet). Compared to ResNet-50, our UL-ResNet achieves a 7.3× inference speedup and a 142.6× reduction in model parameters. To ensure that this highly simplified model can still effectively learn feature extraction and construction methods, we employ model distillation. This technique allows the feature extraction and classification capabilities of existing steel plate surface defect detection models to be transferred to our smaller model. Recognizing the disparities in feature maps and hierarchical feature levels between the teacher and student models, we also utilize a trainable attention module to facilitate knowledge transfer from the teacher model. Experimental results demonstrate that our student model effectively learns from the teacher model, achieving a classification accuracy of 99.44%. Gaochang Wu |
MMSP | 2 |
| 2025 | Geo-NI: Geometry-Aware Neural Interpolation for Light Field RenderingabstractWe present a novel Geometry-aware Neural Interpolation (Geo-NI) framework for light field rendering. Previous learning-based approaches either perform direct interpolation via neural networks, which we dubbed Neural Interpolation (NI), or explore scene geometry for novel view synthesis, also known as Depth Image-Based Rendering (DIBR). Both kinds of approaches have their own strengths and weaknesses in addressing non-Lambert effect and large disparity problems. In this paper, we incorporate the ideas behind these two kinds of approaches by launching the NI within a specific DIBR pipeline. Specifically, a DIBR network in the proposed Geo-NI serves to construct a novel reconstruction cost volume for neural interpolated light fields sheared by different depth hypotheses. The reconstruction cost can be interpreted as an indicator reflecting the reconstruction quality under a certain depth hypothesis, and is further applied to guide the rendering of the final high angular resolution light field. To implement the Geo-NI framework more practically, we further propose an efficient modeling strategy to encode high-dimensional cost volumes using a lower-dimension network. By combining the superiorities of NI and DIBR, the proposed Geo-NI is able to render views with large disparities with the help of scene geometry while also reconstructing the non-Lambertian effect when depth is prone to be ambiguous. Extensive experiments on various datasets demonstrate the superior performance of the proposed geometry-aware light field rendering framework. Gaochang Wu, Yuemei Zhou, Lu Fang 0001, Yebin Liu, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Unified Domain Adaptive Semantic SegmentationabstractUnsupervised Domain Adaptive Semantic Segmentation (UDA-SS) aims to transfer the supervision from a labeled source domain to an unlabeled and shifted target domain. The majority of existing UDA-SS works typically consider images whilst recent attempts have extended further to tackle videos by modeling the temporal dimension. Although two lines of research share the major challenges - overcoming the underlying domain distribution shift, their studies are largely independent. It causes several issues: (1) The insights gained from each line of research remain fragmented, leading to a lack of holistic understanding of the problem and potential solutions. (2) Preventing the unification of methods and best practices across two scenarios (images and videos) will lead to redundant efforts and missed opportunities for cross-pollination of ideas. (3) Without a unified approach, the knowledge and advancements made in one scenario may not be effectively transferred to the other, leading to suboptimal performance and slower progress. Under this observation, we advocate unifying the study of UDA-SS across video and image scenarios, enabling a more comprehensive understanding, synergistic advancements, and efficient knowledge sharing. To that end, we explore the unified UDA-SS from a general domain augmentation perspective, serving as a unifying framework, enabling improved generalization, and potential for cross-pollination, ultimately contributing to the practical impact and overall progress. Specifically, we propose a Quad-directional Mixup (QuadMix) method, characterized by tackling intra-domain discontinuity, fragmented gap bridging, and feature inconsistencies through four-directional paths designed for intra- and inter-domain mixing within an explicit feature space. To deal with temporal shifts within videos, we incorporate optical flow-guided feature aggregation across spatial and temporal dimensions for fine-grained domain alignment, which is extendable to image scenarios. Extensive experiments show that QuadMix outperforms the state-of-the-art works by large margins on four challenging UDA-SS benchmarks. Gaochang Wu, Jing Zhang 0037, Xiatian Zhu, Dacheng Tao, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Cross-Modal Learning for Anomaly Detection in Complex Industrial Process: Methodology and BenchmarkabstractAnomaly detection in complex industrial processes plays a pivotal role in ensuring efficient, stable, and secure operation. Existing anomaly detection methods primarily focus on analyzing dominant anomalies using the process variables (such as arc current) or constructing neural networks based on abnormal visual features, while overlooking the intrinsic correlation of cross-modal information. This paper proposes a cross-modal Transformer (dubbed FmFormer), designed to facilitate anomaly detection by exploring the correlation between visual features (video) and process variables (current) in the context of the fused magnesium smelting process. Our approach introduces a novel tokenization paradigm to effectively bridge the substantial dimensionality gap between the 3D video modality and the 1D current modality in a multiscale manner, enabling a hierarchical reconstruction of pixel-level anomaly detection. Subsequently, the FmFormer leverages self-attention to learn internal features within each modality and bidirectional cross-attention to capture correlations across modalities. By decoding the bidirectional correlation features, we obtain the final detection result and even locate the specific anomaly region. To validate the effectiveness of the proposed method, we also present a pioneering cross-modal benchmark of the fused magnesium smelting process, featuring synchronously acquired video and current data for over 2.2 million samples. Leveraging cross-modal learning, the proposed FmFormer achieves state-of-the-art performance in detecting anomalies, particularly under extreme interferences such as current fluctuations and visual occlusion caused by heavy water mist. The presented methodology and benchmark may be applicable to other industrial applications with some amendments. The benchmark will be released athttps://github.com/GaochangWu/FMF-Benchmark. Gaochang Wu, Lan Deng, Jingxin Zhang 0001, Tianyou Chai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | ProbIBR: Fast Image-Based Rendering With Learned Probability-Guided SamplingabstractWe present a general, fast, and practical solution for interpolating novel views of diverse real-world scenes given a sparse set of nearby views. Existing generic novel view synthesis methods rely on time-consuming scene geometry pre-computation or redundant sampling of the entire space for neural volumetric rendering, limiting the overall efficiency. Instead, we incorporate learned MVS priors into the neural volume rendering pipeline while improving the rendering efficiency by reducing sampling points under the guidance of depth probability distributions. Specifically, fewer but important points are sampled under the guidance of depth probability distributions extracted from the learned MVS architecture. Based on the learned probability-guided sampling, we develop a sophisticated neural volume rendering module that effectively integrates source view information with the learned scene structures. We further propose confidence-aware refinement to improve the rendering results in uncertain, occluded, and unreferenced regions. Moreover, we build a four-view camera system for holographic display and provide a real-time version of our framework for free-viewpoint experience, where novel view images of a spatial resolution of 512×512 can be rendered at around 20 fps on a single GTX 3090 GPU. Experiments show that our method achieves 15 to 40 times faster rendering compared to state-of-the-art baselines, with strong generalization capacity and comparable high-quality novel view synthesis performance. Yuemei Zhou, Tao Yu 0007, Zerong Zheng, Gaochang Wu, Guihua Zhao, Ying Fu 0001, Yebin Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Disentangling Light Fields for Super-Resolution and Disparity EstimationabstractLight field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been proposed for various LF image processing tasks. However, it is challenging for CNNs to effectively process LF images since the spatial and angular information are highly inter-twined with varying disparities. In this paper, we propose a generic mechanism to disentangle these coupled information for LF image processing. Specifically, we first design a class of domain-specific convolutions to disentangle LFs from different dimensions, and then leverage these disentangled features by designing task-specific modules. Our disentangling mechanism can well incorporate the LF structure prior and effectively handle 4D LF data. Based on the proposed mechanism, we develop three networks (i.e., DistgSSR, DistgASR and DistgDisp) for spatial super-resolution, angular super-resolution and disparity estimation. Experimental results show that our networks achieve state-of-the-art performance on all these three tasks, which demonstrates the effectiveness, efficiency, and generality of our disentangling mechanism. Project page: https://yingqianwang.github.io/DistgLF/. Yingqian Wang 0002, Longguang Wang, Gaochang Wu, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Kernel Fisher Dictionary Transfer LearningabstractDictionary learning is an efficient knowledge representation method that can learn the essential features of data. Traditional dictionary learning methods are difficult to obtain nonlinear information when processing large-scale and high-dimensional datasets. While most dictionary learning algorithms are based on the assumption that the training data and test data have the same feature distribution, which is not always true in practical applications. To address the above problems, we propose the Kernel Fisher Dictionary Transfer Learning (KFDTL) algorithm. First, we map each sample to high-dimensional space through kernel mapping and use any dictionary learning algorithm to learn the essential features. Then, the feature-based transfer learning method is performed to predict the labels of the target samples. This method includes three main contributions: (1) KFDTL constructs a discriminative Fisher embedding model to make the same class samples have similar coding coefficients; (2) Based on the relationship between profiles and atoms, KFDTL constructs an adaptive model that adapts source domain samples to target domain samples; (3) The kernel method is used to efficiently solve nonlinear problems. Experiments on a large number of public image datasets have proved the effectiveness of the proposed method. The source code of the proposed method is available at https://github.com/zzfan3/KFDTL . Linrui Shi, Zheng Zhang 0006, Zizhu Fan, Chao Xi, Gaochang Wu |
ACM Trans. Knowl. Discov. Data | 6 |
| 2022 | Revisiting Light Field Rendering With Deep Anti-Aliasing Neural NetworkabstractThe light field (LF) reconstruction is mainly confronted with two challenges, large disparity and the non-Lambertian effect. Typical approaches either address the large disparity challenge using depth estimation followed by view synthesis or eschew explicit depth information to enable non-Lambertian rendering, but rarely solve both challenges in a unified framework. In this paper, we revisit the classic LF rendering framework to address both challenges by incorporating it with advanced deep learning techniques. First, we analytically show that the essential issue behind the large disparity and non-Lambertian challenges is the aliasing problem. Classic LF rendering approaches typically mitigate the aliasing with a reconstruction filter in the Fourier domain, which is, however, intractable to implement within a deep learning pipeline. Instead, we introduce an alternative framework to perform anti-aliasing reconstruction in the image domain and analytically show comparable efficacy on the aliasing issue. To explore the full potential, we then embed the anti-aliasing framework into a deep neural network through the design of an integrated architecture and trainable parameters. The network is trained through end-to-end optimization using a peculiar training set, including regular LFs and unstructured LFs. The proposed deep learning pipeline shows a substantial superiority in solving both the large disparity and the non-Lambertian challenges compared with other state-of-the-art approaches. In addition to the view interpolation for an LF, we also show that the proposed pipeline also benefits light field view extrapolation. Gaochang Wu, Yebin Liu, Lu Fang 0001, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Cross-MPI: Cross-Scale Stereo for Image Super-Resolution Using Multiplane ImagesabstractVarious combinations of cameras enrich computational photography, among which reference-based super-resolution (RefSR) plays a critical role in multiscale imaging systems. However, existing RefSR approaches fail to accomplish high-fidelity super-resolution under a large resolution gap, e.g., 8× upscaling, due to the lower consideration of the underlying scene structure. In this paper, we aim to solve the RefSR problem in actual multiscale camera systems inspired by multiplane image (MPI) representation. Specifically, we propose Cross-MPI, an end-to-end RefSR network composed of a novel plane-aware attention-based MPI mechanism, a multiscale guided upsampling module as well as a super-resolution (SR) synthesis and fusion module. Instead of using a direct and exhaustive matching between the cross-scale stereo, the proposed plane-aware attention mechanism fully utilizes the concealed scene structure for efficient attention-based correspondence searching. Further combined with a gentle coarse-to-fine guided upsampling strategy, the proposed Cross-MPI can achieve a robust and accurate detail transmission. Experimental results on both digitally synthesized and optical zoom cross-scale data show that the Cross-MPI framework can achieve superior performance against the existing RefSR methods and is a real fit for actual multiscale camera systems even with large-scale differences. Yuemei Zhou, Gaochang Wu, Ying Fu 0001, Kun Li 0001, Yebin Liu |
CVPR | 2 |
| 2021 | LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography EstimationabstractCross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondences between them, which leads to degraded accuracy in cross-resolution challenges. In this paper, we consider the cross-resolution homography estimation as a multimodal problem, and propose a local transformer network embedded within a multiscale structure to explicitly learn correspondences between the multimodal inputs, namely, input images with different resolutions. The proposed local transformer adopts a local attention map specifically for each position in the feature. By combining the local transformer with the multiscale structure, the network is able to capture long-short range correspondences efficiently and accurately. Experiments on both the MS-COCO dataset and the real-captured cross-resolution dataset show that the proposed network outperforms existing state-of-the-art feature-based and deep-learning-based homography estimation methods, and is able to accurately align images under 10× resolution gap. Ruizhi Shao, Gaochang Wu, Yuemei Zhou, Ying Fu 0001, Lu Fang 0001, Yebin Liu |
ICCV | 2 |
| 2021 | Boosting Single Image Super-Resolution Learnt From Implicit Multi-Image PriorabstractLearning-based single image super-resolution (SISR) aims to learn a versatile mapping from low resolution (LR) image to its high resolution (HR) version. The critical challenge is to bias the network training towards continuous and sharp edges. For the first time in this work, we propose an implicit boundary prior learnt from multi-view observations to significantly mitigate the challenge in SISR we outline. Specifically, the multi-image prior that encodes both disparity information and boundary structure of the scene supervise a SISR network for edge-preserving. For simplicity, in the training procedure of our framework, light field (LF) serves as an effective multi-image prior, and a hybrid loss function jointly considers the content, structure, variance as well as disparity information from 4D LF data. Consequently, for inference, such a general training scheme boosts the performance of various SISR networks, especially for the regions along edges. Extensive experiments on representative backbone SISR architectures constantly show the effectiveness of the proposed method, leading to around 0.6 dB gain without modifying the network architecture. Dingjian Jin, Mengqi Ji, Lan Xu 0003, Gaochang Wu, Lu Fang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Spatial-Angular Attention Network for Light Field ReconstructionabstractTypical learning-based light field reconstruction methods demand in constructing a large receptive field by deepening their networks to capture correspondences between input views. In this paper, we propose a spatial-angular attention network to perceive non-local correspondences in the light field, and reconstruct high angular resolution light field in an end-to-end manner. Motivated by the non-local attention mechanism (Wang et al., 2018; Zhang et al., 2019), a spatial-angular attention module specifically for the high-dimensional light field data is introduced to compute the response of each query pixel from all the positions on the epipolar plane, and generate an attention map that captures correspondences along the angular dimension. Then a multi-scale reconstruction structure is proposed to efficiently implement the non-local attention in the low resolution feature space, while also preserving the high frequency components in the high-resolution feature space. Extensive experiments demonstrate the superior performance of the proposed spatial-angular attention network for reconstructing sparsely-sampled light fields with Non-Lambertian effects. Gaochang Wu, Yingqian Wang 0002, Yebin Liu, Lu Fang 0001, Tianyou Chai |
IEEE Trans. Image Process. | 1 |
| 2020 | All-in-depth via Cross-baseline Light Field CameraabstractLight-field (LF) camera holds great promise for passive/general depth estimation benefited from high angular resolution, yet suffering small baseline for distanced region. While stereo solution with large baseline is superior to handle distant scenarios, the problem of limited angular resolution becomes bothering for near objects. Aiming for all-in-depth solution, we propose a cross-baseline LF camera using a commercial LF camera and a monocular camera, which naturally form a 'stereo camera' enabling compensated baseline for LF camera. The idea is simple yet non-trivial, due to the significant angular resolution gap and baseline gap between LF and stereo cameras. Dingjian Jin, Anke Zhang, Gaochang Wu, Haoqian Wang, Lu Fang 0001 |
ACM Multimedia | 4 |
| 2019 | Joint view synthesis and disparity refinement for stereo matching
Gaochang Wu, Yuanhao Huang, Yebin Liu |
Frontiers Comput. Sci. | 1 |
| 2019 | Light Field Reconstruction Using Convolutional Network on EPI and Extended ApplicationsabstractIn this paper, a novel convolutional neural network (CNN)-based framework is developed for light field reconstruction from a sparse set of views. We indicate that the reconstruction can be efficiently modeled as angular restoration on an epipolar plane image (EPI). The main problem in direct reconstruction on the EPI involves an information asymmetry between the spatial and angular dimensions, where the detailed portion in the angular dimensions is damaged by undersampling. Directly upsampling or super-resolving the light field in the angular dimensions causes ghosting effects. To suppress these ghosting effects, we contribute a novel "blur-restoration-deblur" framework. First, the "blur" step is applied to extract the low-frequency components of the light field in the spatial dimensions by convolving each EPI slice with a selected blur kernel. Then, the "restoration" step is implemented by a CNN, which is trained to restore the angular details of the EPI. Finally, we use a non-blind "deblur" operation to recover the spatial high frequencies suppressed by the EPI blur. We evaluate our approach on several datasets, including synthetic scenes, real-world scenes and challenging microscope light field data. We demonstrate the high performance and robustness of the proposed framework compared with state-of-the-art algorithms. We further show extended applications, including depth enhancement and interpolation for unstructured input. More importantly, a novel rendering approach is presented by combining the proposed framework and depth information to handle large disparities. Gaochang Wu, Yebin Liu, Lu Fang 0001, Qionghai Dai, Tianyou Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Learning Sheared EPI Structure for Light Field ReconstructionabstractResearch in light field reconstruction focuses on synthesizing novel views with the assistance of depth information. In this paper, we present a learning-based light field reconstruction approach by fusing a set of sheared epipolar plane images (EPIs). We start by showing that a patch in a sheared EPI will exhibit a clear structure when the sheared value equals the depth of that patch. By taking advantage of this pattern, a convolutional neural network (CNN) is then trained to evaluate the sheared EPIs, and output a reference score for fusing the sheared EPIs. The proposed CNN is elaborately designed to learn the similarity degree between the input sheared EPI and the ground truth EPI. Therefore, no depth information is required for network training and reasoning. We demonstrate the high performance of the proposed method through evaluations on synthetic scenes, real-world scenes, and challenging microscope light fields. We also show a further application of our proposed network for depth inference. Gaochang Wu, Yebin Liu, Qionghai Dai, Tianyou Chai |
IEEE Trans. Image Process. | 1 |
| 2017 | Light Field Reconstruction Using Deep Convolutional Network on EPIabstractIn this paper, we take advantage of the clear texture structure of the epipolar plane image (EPI) in the light field data and model the problem of light field reconstruction from a sparse set of views as a CNN-based angular detail restoration on EPI. We indicate that one of the main challenges in sparsely sampled light field reconstruction is the information asymmetry between the spatial and angular domain, where the detail portion in the angular domain is damaged by undersampling. To balance the spatial and angular information, the spatial high frequency components of an EPI is removed using EPI blur, before feeding to the network. Finally, a non-blind deblur operation is used to recover the spatial detail suppressed by the EPI blur. We evaluate our approach on several datasets including synthetic scenes, real-world scenes and challenging microscope light field data. We demonstrate the high performance and robustness of the proposed framework compared with the state-of-the-arts algorithms. We also show a further application for depth enhancement by using the reconstructed light field. Gaochang Wu, Mandan Zhao, Liangyong Wang, Qionghai Dai, Tianyou Chai, Yebin Liu |
CVPR | 1 |
| 2015 | Automatic determination of cutoff frequency for filter design using neuro-fuzzy systems
Yongfu Wang 0001, Gaochang Wu, Gang (Sheng) Chen |
Neurocomputing | 2 |