EDBT 2026 Demo / reviewers in the wild / expert
Jie Chen 0026
dblp:92/6289-26
· DBLP profile ↗
61ranked-venue papers
14as first author
36since 2021 · last 2026
0000-0001-8419-4620ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 10 first-author · 28 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 14 since 2021Systems, architecture and hardware · 7 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing ControlabstractTraditional text-to-motion frameworks often lack precise control, and existing approaches based on joint keyframe locations provide only positional guidance, making it challenging and unintuitive to specify body part orientations and motion timing. To address these limitations, we introduce the Salient Orientation Symbolic (SOS) script, a programmable symbolic framework for specifying body part orientations and motion timing at keyframes. We further propose an automatic SOS extraction pipeline that employs temporally-constrained agglomerative clustering for frame saliency detection and a Saliency-based Masking Scheme (SMS) to generate sparse, interpretable SOS scripts directly from motion data. Moreover, we present the SOSControl framework, which treats the available orientation symbols in the sparse SOS script as salient and prioritizes satisfying these constraints during motion generation. By incorporating SMS-based data augmentation and gradient-based iterative optimization, the framework enhances alignment with user-specified constraints. Additionally, it employs a ControlNet-based ACTOR-PAE Decoder to ensure smooth and natural motion outputs. Extensive experiments demonstrate that the SOS extraction pipeline generates human-interpretable scripts with symbolic annotations at salient keyframes, while the SOSControl framework outperforms existing baselines in motion quality, controllability, and generalizability with respect to motion timing and body part orientation control. Ho Yin Au, Junkun Jiang, Jie Chen 0026 |
AAAI | 3 |
| 2026 | SAOT: An Enhanced Locality-Aware Spectral Transformer for Solving PDEsabstractNeural operators have shown great potential in solving a family of Partial Differential Equations (PDEs) by modeling the mappings between input and output functions. Fourier Neural Operator (FNO) implements global convolutions via parameterizing the integral operators in Fourier space. However, it often results in over-smoothing solutions and fails to capture local details and high-frequency components. To address these limitations, we investigate incorporating the spatial-frequency localization property of Wavelet transforms into the Transformer architecture. We propose a novel Wavelet Attention (WA) module with linear computational complexity to efficiently learn locality-aware features. Building upon WA, we further develop the Spectral Attention Operator Transformer (SAOT), a hybrid spectral Transformer framework that integrates WA’s localized focus with the global receptive field of Fourier-based Attention (FA) through a gated fusion block. Experimental results demonstrate that WA significantly mitigates the limitations of FA and outperforms existing Wavelet-based neural operators by a large margin. By integrating the locality-aware and global spectral representations, SAOT achieves state-of-the-art performance on six operator learning benchmarks and exhibits strong discretization-invariant ability. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang |
AAAI | 2 |
| 2025 | SphereFusion: Efficient Panorama Depth Estimation via Gated FusionabstractDue to the rapid development of panorama cameras, the task of estimating panorama depth has attracted significant attention from the computer vision community, especially in applications such as robot sensing and autonomous driving. However, existing methods relying on different projection formats often encounter challenges, either struggling with distortion and discontinuity in the case of equirectangular, cubemap, and tangent projections, or experiencing a loss of texture details with the spherical projection. To tackle these concerns, we present SphereFusion, an end-toend framework that combines the strengths of various projection methods. Specifically, SphereFusion initially employs$2 D$image convolution and mesh operations to extract two distinct types of features from the panorama image in both equirectangular and spherical projection domains. These features are then projected onto the spherical domain, where a gate fusion module selects the most reliable features for fusion. Finally, SphereFusion estimates panorama depth within the spherical domain. Meanwhile, SphereFusion employs a cache strategy to improve the efficiency of mesh operation. Extensive experiments on three public panorama datasets demonstrate that SphereFusion achieves competitive results with other state-of-theart methods, while presenting the fastest inference speed at only 17 ms on a$512 \times 1024$panorama image. Qingsong Yan, Qiang Wang 0022, Kaiyong Zhao, Jie Chen 0026, Bo Li 0001, Xiaowen Chu 0001 |
3DV | 4 |
| 2025 | NCDI-Diffusion: Neural Contextual and Directional Inversion for Novel View Synthesis through Diffusion ModelsabstractNovel view synthesis typically requires a comprehensive set of multi-view images for either image-based rendering or scene representation-based optimization. However, achieving high-fidelity novel view rendering often demands a large number of images. To address this limitation, we propose NCDI-Diffusion, a novel diffusion-based view synthesis method that reduces the number of required images by leveraging the prior knowledge embedded in pre-trained diffusion models. Specifically, NCDI-Diffusion encapsulates both the contextual and directional information of a scene by utilizing neural descriptors, which are inversely derived from a limited set of positioned multi-view training images. These descriptors guide the diffusion model's image synthesis process, enabling the generation of high-quality novel views. Empirical results on the Forward-facing Dataset demonstrate the effectiveness of our approach to novel view synthesis. Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Changting Lin |
ICASSP | 2 |
| 2025 | Fast and Physically-based Neural Explicit Surface for Relightable Human AvatarsabstractEfficiently modeling relightable human avatars from sparse-view videos is crucial for AR/VR applications. Current methods use neural implicit representations to capture dynamic geometry and reflectance, which incur high costs due to the need for dense sampling in volume rendering. To overcome these challenges, we introduce Physically-based Neural Explicit Surface (PhyNES), which employs compact neural material maps based on the Neural Explicit Surface (NES) representation. PhyNES organizes human models in a compact 2D space, enhancing material disentanglement efficiency. By connecting Signed Distance Fields to explicit surfaces, PhyNES enables efficient geometry inference around a parameterized human shape model. This approach models dynamic geometry, texture, and material maps as 2D neural representations, enabling efficient rasterization. PhyNES effectively captures physical surface attributes under varying illumination, enabling real-time physically-based rendering. Experiments show that PhyNES achieves relighting quality comparable to SOTA methods while significantly improving rendering speed, memory efficiency, and reconstruction quality. Jie Chen 0026, Hui Zhang 0062 |
ICME | 3 |
| 2025 | Learning Physics-Informed Color-Aware Transforms for Low-Light Image EnhancementabstractImage decomposition offers deep insights into the imaging factors of visual data and significantly enhances various advanced computer vision tasks. In this work, we introduce a novel approach to low-light image enhancement based on decomposed physics-informed priors. Existing methods that directly map low-light to normal-light images in the sRGB color space suffer from inconsistent color predictions and high sensitivity to spectral power distribution (SPD) variations, resulting in unstable performance under diverse lighting conditions. To address these challenges, we introduce a Physics-informed Color-aware Transform (PiCat), a learning-based framework that converts low-light images from the sRGB color space into deep illumination-invariant descriptors via our proposed Color-aware Transform (CAT). This transformation enables robust handling of complex lighting and SPD variations. Complementing this, we propose the Content-Noise Decomposition Network (CNDN), which refines the descriptor distributions to better align with well-lit conditions by mitigating noise and other distortions, thereby effectively restoring content representations to low-light images. The CAT and the CNDN collectively act as a physical prior, guiding the transformation process from low-light to normal-light domains. Our proposed PiCat framework demonstrates superior performance compared to state-of-the-art methods across five benchmark datasets. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
ICME | 2 |
| 2025 | Dual-Balancing for Physics-Informed Neural NetworksabstractPhysics-informed neural networks (PINNs) have emerged as a new learning paradigm for solving partial differential equations (PDEs) by enforcing the constraints of physical equations, boundary conditions (BCs), and initial conditions (ICs) into the loss function. Despite their successes, vanilla PINNs still suffer from poor accuracy and slow convergence due to the intractable multi-objective optimization issue. In this paper, we propose a novel Dual-Balanced PINN (DB-PINN), which dynamically adjusts loss weights by integrating inter-balancing and intra-balancing to alleviate two imbalance issues in PINNs. Inter-balancing aims to mitigate the gradient imbalance between PDE residual loss and condition-fitting losses by determining an aggregated weight that offsets their gradient distribution discrepancies. Intra-balancing acts on condition-fitting losses to tackle the imbalance in fitting difficulty across diverse conditions. By evaluating the fitting difficulty based on the loss records, intra-balancing can allocate the aggregated weight proportionally to each condition loss according to its fitting difficulty level. We further introduce a robust weight update strategy to prevent abrupt spikes and arithmetic overflow in instantaneous weight values caused by large loss variances, enabling smooth weight updating and stable training. Extensive experiments demonstrate that DB-PINN achieves significantly superior performance than those popular gradient-based weighting methods in terms of convergence speed and prediction accuracy. Our code and supplementary material are available at https://github.com/chenhong-zhou/DualBalanced-PINNs. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang, Ching Eng Png |
IJCAI | 2 |
| 2025 | RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation under Complex TrajectoriesabstractNeural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have emerged as powerful tools for 3D reconstruction and SLAM tasks. However, their performance depends heavily on accurate camera pose priors. Existing approaches attempt to address this issue by introducing external constraints but fall short of achieving satisfactory accuracy, particularly when camera trajectories are complex. In this paper, we propose a novel method, RA-NeRF, capable of predicting highly accurate camera poses even with complex camera trajectories. Following the incremental pipeline, RA-NeRF reconstructs the scene using NeRF with photometric consistency and incorporates flow-driven pose regulation to enhance robustness during initialization and localization. Additionally, RA-NeRF employs an implicit pose filter to capture the camera movement pattern and eliminate the noise for pose estimation. To validate our method, we conduct extensive experiments on the Tanks&Temple dataset for standard evaluation, as well as the NeRFBuster dataset, which presents challenging camera pose trajectories. On both datasets, RA-NeRF achieves state-of-the-art results in both camera pose estimation and visual quality, demonstrating its effectiveness and robustness in scene reconstruction under complex pose trajectories. Qingsong Yan, Qiang Wang 0022, Kaiyong Zhao, Jie Chen 0026, Bo Li 0001, Xiaowen Chu 0001 |
IROS | 4 |
| 2025 | Efficient Real-Time Fine-Grained Action Recognition over a Progressive and Hierarchical Classification FrameworkabstractReal-Time fine-grained action recognition (AR) presents significant challenges in resource-constrained environments with strict accuracy requirements. This paper proposes an efficient real-time AR system that utilizes a progressive hierarchical classification framework to achieve high accuracy while minimizing computational demands. The system utilizes the YOLO model for initial single-frame classification, enabling precise identification of alarming actions with a high recall rate to facilitate timely alerts. Subsequently, a second-tier recognizer that relies on spatiotemporal features is applied for fine-grained AR of identified alarming actions. To enhance recognition accuracy, we introduce a hierarchical classification model where actions are grouped based on semantic and kinematic similarity, followed by further classification within each group. Additionally, we implement a multi-threaded scheduling pipeline that ensures prompt alarms with reasonable loading time for precise AR. Experimental results demonstrate that our system effectively balances computational efficiency with recognition accuracy, making it suitable for real-time deployment in resource-constrained settings. Shuwen Niu, Junkun Jiang, Jie Chen 0026 |
ISCAS | 3 |
| 2025 | Deep Compositional Phase Diffusion for Long Motion Sequence GenerationabstractRecent research on motion generation has shown significant progress in generating semantically aligned motion with singular semantics. However, when employing these models to create composite sequences containing multiple semantically generated motion clips, they often struggle to preserve the continuity of motion dynamics at the transition boundaries between clips, resulting in awkward transitions and abrupt artifacts. To address these challenges, we present Compositional Phase Diffusion, which leverages the Semantic Phase Diffusion Module (SPDM) and Transitional Phase Diffusion Module (TPDM) to progressively incorporate semantic guidance and phase details from adjacent motion clips into the diffusion process. Specifically, SPDM and TPDM operate within the latent motion frequency domain established by the pre-trained Action-Centric Motion Phase Autoencoder (ACT-PAE). This allows them to learn semantically important and transition-aware phase information from variable-length motion clips during training. Experimental results demonstrate the competitive performance of our proposed framework in generating compositional motion sequences that align semantically with the input conditions, while preserving phase transitional continuity between preceding and succeeding motion clips. Additionally, motion inbetweening task is made possible by keeping the phase parameter of the input motion sequences fixed throughout the diffusion process, showcasing the potential for extending the proposed framework to accommodate various application scenarios. Codes are available at
https://github.com/asdryau/TransPhase. Ho Yin Au, Jie Chen 0026, Junkun Jiang, Jingyu Xiang |
NeurIPS | 2 |
| 2025 | Unraveling Metameric Dilemma for Spectral Reconstruction: A High-Fidelity Approach via Semi-Supervised LearningabstractSpectral reconstruction from RGB images often suffers from a metameric dilemma, where distinct spectral distributions map to nearly identical RGB values, making them indistinguishable to current models and leading to unreliable reconstructions.
In this paper, we present Diff-Spectra that integrates supervised physics-aware spectral estimation and unsupervised high-fidelity spectral regularization for HSI reconstruction.
We first introduce an Adaptive illumiChroma Decoupling (AICD) module to decouple illumination and chrominance information, which learns intrinsic and distinctive feature distributions, thereby mitigating the metameric issue.
Then, we incorporate the AICD into a learnable spectral response function (SRF) guided hyperspectral initial estimation mechanism to mimic the physical image formation and thus inject physics-aware reasoning into neural networks, turning an ill-posed problem into a constrained, interpretable task.
We also introduce a metameric spectra augmentation method to synthesize comprehensive hyperspectral data to pre-train a Spectral Diffusion Module (SDM), which internalizes the statistical properties of real-world HSI data, enforcing unsupervised high-fidelity regularization on the spectral transitions via inner-loop optimization during inference.
Extensive experimental evaluations demonstrate that our Diff-Spectra achieves SOTA performance on both Spectral reconstruction and downstream HSI classification. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
NeurIPS | 2 |
| 2025 | Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint CloudabstractMulti-person motion capture over sparse angular observations is a challenging problem under interference from both self- and mutual-occlusions. Existing works produce accurate 2D joint detection, however, when these are triangulated and lifted into 3D, available solutions all struggle in selecting the most accurate candidates and associating them to the correct joint type and target identity. As such, in order to fully utilize all accurate 2D joint location information, we propose to independently triangulate between all same-typed 2D joints from all camera views regardless of their target ID, forming the Joint Cloud. Joint Cloud consist of both valid joints lifted from the same joint type and target ID, as well as falsely constructed ones that are from different 2D sources. These redundant and inaccurate candidates are processed over the proposed Joint Cloud Selection and Aggregation Transformer (JCSAT) involving three cascaded encoders which deeply explore the trajectile, skeletal structural, and view-dependent correlations among all 3D point candidates in the cross-embedding space. An Optimal Token Attention Path (OTAP) module is proposed which subsequently selects and aggregates informative features from these redundant observations for the final prediction of human motion. To demonstrate the effectiveness of JCSAT, we build and publish a new multi-person motion capture dataset BUMocap-X with complex interactions and severe occlusions. Comprehensive experiments over the newly presented as well as benchmark datasets validate the effectiveness of the proposed framework, which outperforms all existing state-of-the-art methods, especially under challenging occlusion scenarios. Junkun Jiang, Jie Chen 0026, Ho Yin Au, Wei Xue 0002, Yike Guo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | CF-NeRF: Camera Parameter Free Neural Radiance Fields with Incremental LearningabstractNeural Radiance Fields have demonstrated impressive performance in novel view synthesis. However, NeRF and most of its variants still rely on traditional complex pipelines to provide extrinsic and intrinsic camera parameters, such as COLMAP. Recent works, like NeRFmm, BARF, and L2G-NeRF, directly treat camera parameters as learnable and estimate them through differential volume rendering. However, these methods work for forward-looking scenes with slight motions and fail to tackle the rotation scenario in practice. To overcome this limitation, we propose a novel camera parameter free neural radiance field (CF-NeRF), which incrementally reconstructs 3D representations and recovers the camera parameters inspired by incremental structure from motion. Given a sequence of images, CF-NeRF estimates camera parameters of images one by one and reconstructs the scene through initialization, implicit localization, and implicit optimization. To evaluate our method, we use a challenging real-world dataset, NeRFBuster, which provides 12 scenes under complex trajectories. Results demonstrate that CF-NeRF is robust to rotation and achieves state-of-the-art results without providing prior information and constraints. Qingsong Yan, Qiang Wang 0022, Kaiyong Zhao, Jie Chen 0026, Bo Li 0001, Xiaowen Chu 0001 |
AAAI | 4 |
| 2024 | Hyperspectral Image Reconstruction via Combinatorial Embedding of Cross-Channel Spatio-Spectral CluesabstractExisting learning-based hyperspectral reconstruction methods show limitations in fully exploiting the information among the hyperspectral bands. As such, we propose to investigate the chromatic inter-dependencies in their respective hyperspectral embedding space. These embedded features can be fully exploited by querying the inter-channel correlations in a combinatorial manner, with the unique and complementary information efficiently fused into the final prediction. We found such independent modeling and combinatorial excavation mechanisms are extremely beneficial to uncover marginal spectral features, especially in the long wavelength bands. In addition, we have proposed a spatio-spectral attention block and a spectrum-fusion attention module, which greatly facilitates the excavation and fusion of information at both semantically long-range levels and fine-grained pixel levels across all dimensions. Extensive quantitative and qualitative experiments show that our method (dubbed CESST) achieves SOTA performance. Code for this project is at: https://github.com/AlexYangxx/CESST. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
AAAI | 2 |
| 2024 | Exploring Latent Cross-Channel Embedding for Accurate 3d Human Pose Reconstruction in a Diffusion FrameworkabstractMonocular 3D human pose estimation poses significant challenges due to the inherent depth ambiguities that arise during the reprojection process from 2D to 3D. Conventional approaches that rely on estimating an over-fit projection matrix struggle to effectively address these challenges and often result in noisy outputs. Recent advancements in diffusion models have shown promise in incorporating structural priors to address reprojection ambiguities. However, there is still ample room for improvement as these methods often overlook the exploration of correlation between the 2D and 3D jointlevel features. In this study, we propose a novel cross-channel embedding framework that aims to fully explore the correlation between joint-level features of 3D coordinates and their 2D projections. In addition, we introduce a context guidance mechanism to facilitate the propagation of joint graph attention across latent channels during the iterative diffusion process. To evaluate the effectiveness of our proposed method, we conduct experiments on two benchmark datasets, namely Human3.6M and MPI-INF-3DHP. Our results demonstrate a significant improvement in terms of reconstruction accuracy compared to state-of-the-art methods. The code for our method will be made available online for further reference. Junkun Jiang, Jie Chen 0026 |
ICASSP | 2 |
| 2024 | Motion Part-Level Interpolation and Manipulation over Automatic Symbolic Labanotation AnnotationabstractMotion sequencing is a crucial process in creating smooth and natural animations by arranging individual motion sequences based on desired action scripts. Existing methods either rely on carefully engineered key-frame libraries or implicitly encoded latent phase manifolds for sequential interpolation and manipulation. However, ensuring smooth and natural transitions becomes challenging when dealing with complex and diverse actions, and the manipulation flexibility is limited to the frame level. In this study, we introduce a novel motion sequencing framework centered around Labanotation. The framework leverages automatically annotated Labanotation for explicit representation of motion elements to the body-part level. The proposed Laban Masked Autoencoder (LBN-MAE) is able to directly complete, interpolate and translate Laban symbols into natural 3D trajectories. Our framework offers a compact and descriptive representation of motion, enabling precise motion control and reediting. Comparative evaluations against both conventional and state-of-the-art learning-based methods validate the effectiveness of our proposed framework. Junkun Jiang, Ho Yin Au, Jie Chen 0026, Jingyu Xiang |
IJCNN | 3 |
| 2024 | Enhanced Physics-Informed Neural Networks with Optimized Sensor Placement via Multi-Criteria Adaptive SamplingabstractPhysics-informed neural networks (PINNs) have emerged as promising and powerful tools for solving partial differential equations (PDEs). To enforce PINNs that yield accurate solutions to PDEs, a set of scattered spatiotemporal points, known as collocation points, are typically sampled within the computational domain. The choices of collocation points significantly impact the performance of PINNs. However, existing sampling methods primarily rely on the PDE residual, which is insufficient for solutions with steep gradients. To enhance the accuracy of PINNs, we propose a novel multi-criteria adaptive sampling (MCAS) approach to optimally select appropriate collocation points. The MCAS approach integrates three sampling criteria: PDE’s residual, the gradient of residual, and the gradient of solutions, enabling us to capture the PDE violations and the sharpness of solutions. Experimental results demonstrate that the proposed MCAS approach is not only applicable for collocation points but also for optimizing sensor placement, consistently outperforming the residual-based sampling methods. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang, Alexander Matyasko, Ching Eng Png |
IJCNN | 2 |
| 2024 | Mesh-Centric Gaussian Splatting for Human Avatar Modelling with Real-time Dynamic Mesh ReconstructionabstractReal-time mesh reconstruction is highly demanded for integrating human avatar in modern computer graphics applications. Current methods typically use coordinate-based MLP to represent 3D scene as Signed Distance Field (SDF) and optimize it through volumetric rendering, relying on Marching Cubes for mesh extraction. However, volumetric rendering lacks training and rendering efficiency, and the dependence on Marching Cubes significantly impacts mesh extraction efficiency. This study introduces a novel approach, Mesh-Centric Gaussian Splatting (MCGS), which introduces a unique representation Mesh-Centric SDF and optimizes it using high-efficiency Gaussian Splatting. The primary innovation introduces Mesh-Centric SDF, a thin layer of SDF enveloping the underlying mesh, and could be efficiently derived from mesh. This derivation of SDF from mesh allows for mesh optimization through SDF, providing mesh as 0 iso-surface, and eliminating the need for slow Marching Cubes. The secondary innovation focuses on optimizing Mesh-Centric SDF with high-efficiency Gaussian Splatting. By dispersing the underlying mesh of Mesh-Centric SDF into multiple layers and generating Mesh-Constrained Gaussians on them, we create Multi-Layer Gaussians. These Mesh-Constrained Gaussians confine Gaussians within a 2D surface space defined by mesh, ensuring an accurate correspondence between Gaussian rendering and mesh geometry. The Multi-Layer Gaussians serve as sampling layers of Mesh-Centric SDF and can be optimized with Gaussian Splatting, which would further optimize Mesh-Centric SDF and its underlying mesh. As a result, our method can directly optimize the underlying mesh through Gaussian Splatting, providing fast training and rendering speeds derived from Gaussian Splatting, as well as precise surface learning of SDF. Experiments demonstrate that our method achieves dynamic mesh reconstruction at over 30 FPS. In contrast, SDF-based methods using Marching Cubes achieve less than 1 FPS, and concurrent 3D Gaussian Splatting-based methods cannot extract reasonable mesh. Jie Chen 0026 |
ACM Multimedia | 2 |
| 2023 | CasTensoRF: Cascaded Tensorial Radiance Fields for Novel View SynthesisabstractNovel views synthesized from Neural Radiance Fields (NeRF) have reached remarkable rendering quality. However, a 5D radiance field volume is too large to be stored or directly rendered. In order to efficiently reconstruct and manipulate such a high-order tensor, we leverage inspirations from previous tensor decomposition methods, e.g. Tensorial Radiance Fields (TensoRF) and Hierarchical Tucker decomposition. And we propose a Hierarchical Vector-Matrix decomposition (HVMD) framework to learn a sparse approximation of high-order tensors. The proposed HVMD takes advantage of tensor separation and factorization properties and builds a hierarchical scheme that enables a better approximation of the high-order tensor with a very limited number of parameters. Our method achieves better-rendering quality than TensoRF in the NeRF-synthetic dataset given the same model size. The advantage gets more significant when the network parameter number becomes extremely small. Wenpeng Xing, Jie Chen 0026 |
ICME | 2 |
| 2023 | IRCasTRF: Inverse Rendering by Optimizing Cascaded Tensorial Radiance Fields, Lighting, and Materials From Multi-view ImagesabstractWe propose an inverse rendering pipeline that simultaneously reconstructs scene geometry, lighting, and spatially-varying material from a set of multi-view images. Specifically, the proposed pipeline involves volume and physics-based rendering, which are performed separately in two steps: exploration and exploitation. During the exploration step, our method utilizes the compactness of neural radiance fields and a flexible differentiable volume rendering technique to learn an initial volumetric field. Here, we introduce a novel cascaded tensorial radiance field method on top of the Canonical Polyadic (CP) decomposition to boost model compactness beyond conventional methods. In the exploitation step, a shading pass that incorporates a differentiable physics-based shading method is applied to jointly optimize the scene's geometry, spatially-varying materials, and lighting, using image reconstruction loss. Experimental results demonstrate that our proposed inverse rendering pipeline, IRCasTRF, outperforms prior works in inverse rendering quality. The final output is highly compatible with downstream applications like scene editing and advanced simulations. Further details are available on the project page: https://ircasrf.github.io/. Wenpeng Xing, Jie Chen 0026, Ka Chun Cheung, Simon See |
ACM Multimedia | 2 |
| 2023 | Cooperative Colorization: Exploring Latent Cross-Domain Priors for NIR Image Spectrum TranslationabstractNear-infrared (NIR) image spectrum translation is a challenging problem with many promising applications. Existing methods struggle with the mapping ambiguity between the NIR and the RGB domains, and generalize poorly due to the limitations of models' learning capabilities and the unavailability of sufficient NIR-RGB image pairs for training. To address these challenges, we propose a cooperative learning paradigm that colorizes NIR images in parallel with another proxy grayscale colorization task by exploring latent cross-domain priors (i.e., latent spectrum context priors and task domain priors), dubbed CoColor. The complementary statistical and semantic spectrum information from these two task domains -- in the forms of pre-trained colorization networks -- are brought in as task domain priors. A bilateral domain translation module is subsequently designed, in which intermittent NIR images are generated from grayscale and colorized in parallel with authentic NIR images; and vice versa for the grayscale images. These intermittent transformations act as latent spectrum context priors for efficient domain knowledge exchange. We progressively fine-tune and fuse these modules with a series of pixel-level and feature-level consistency constraints. Experiments show that our proposed cooperative learning framework produces satisfactory spectrum translation outputs with diverse colors and rich textures, and outperforms state-of-the-art counterparts by 3.95dB and 4.66dB in terms of PNSR for the NIR and grayscale colorization tasks, respectively. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
ACM Multimedia | 2 |
| 2023 | CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelabstractDenoising diffusion probabilistic models (DDPMs) have shown promising performance for speech synthesis. However, a large number of iterative steps are required to achieve high sample quality, which restricts the inference speed. Maintaining sample quality while increasing sampling speed has become a challenging task. In this paper, we propose a Consistency Model-based Speech synthesis method, CoMoSpeech, which achieve speech synthesis through a single diffusion sampling step while achieving high audio quality. The consistency constraint is applied to distill a consistency model from a well-designed diffusion-based teacher model, which ultimately yields superior performances in the distilled CoMoSpeech. Our experiments show that by generating audio recordings by a single sampling step, the CoMoSpeech achieves an inference speed more than 150 times faster than real-time on a single NVIDIA A100 GPU, which is comparable to FastSpeech2, making diffusion-sampling based speech synthesis truly practical. Meanwhile, objective and subjective evaluations on text-to-speech and singing voice synthesis show that the proposed teacher models yield the best audio quality, and the one-step sampling based CoMoSpeech achieves the best inference speed with better or comparable audio quality to other conventional multi-step diffusion model baselines. Audio samples and codes are available at https://comospeech.github. https://comospeech.github.io/. Zhen Ye 0006, Wei Xue 0002, Xu Tan 0003, Jie Chen 0026, Yike Guo |
ACM Multimedia | 4 |
| 2023 | Explicifying Neural Implicit Fields for Efficient Dynamic Human Avatar Modeling via a Neural Explicit SurfaceabstractThis paper proposes a technique for efficiently modeling dynamic humans by explicifying the implicit neural fields via a Neural Explicit Surface (NES). Implicit neural fields have advantages over traditional explicit representations in modeling dynamic 3D content from sparse observations and effectively representing complex geometries and appearances. Implicit neural fields defined in 3D space, however, are expensive to render due to the need for dense sampling during volumetric rendering. Moreover, their memory efficiency can be further optimized when modeling sparse 3D space. To overcome these issues, the paper proposes utilizing Neural Explicit Surface (NES) to explicitly represent implicit neural fields, facilitating memory and computational efficiency. To achieve this, the paper creates a fully differentiable conversion between the implicit neural fields and the explicit rendering interface of NES, leveraging the strengths of both implicit and explicit approaches. This conversion enables effective training of the hybrid representation using implicit methods and efficient rendering by integrating the explicit rendering interface with a newly proposed rasterization-based neural renderer that only incurs a texture color query once for the initial ray interaction with the explicit surface, resulting in improved inference efficiency. NES describes dynamic human geometries with pose-dependent neural implicit surface deformation fields and their dynamic neural textures both in 2D space, which is a more memory-efficient alternative to traditional 3D methods, reducing redundancy and computational load. The comprehensive experiments show that NES performs similarly to previous 3D approaches, with greatly improved rendering speed and reduced memory cost. Jie Chen 0026, Qiang Wang 0022 |
ACM Multimedia | 2 |
| 2023 | Multi-scale Progressive Feature Embedding for Accurate NIR-to-RGB Spectral Domain TranslationabstractNIR-to-RGB spectral domain translation is a challenging task due to the mapping ambiguities and existing methods show limited learning capacities. To address these challenges, we propose to colorize NIR images via a multi-scale progressive feature embedding network (MPFNet), with the guidance of grayscale image colorization. Specifically, we first introduce a domain translation module that translates NIR source images into the grayscale target domain. By incorporating a progressive training strategy, the statistical and semantic knowledge from both task domains are efficiently aligned with a series of pixel-/feature-level consistency constraints. Besides, a multi-scale progressive feature embedding network is designed to improve learning capabilities. Experiments show that our MPFNet outperforms state-of-the-art counterparts by 2.55dB in the NIR-to-RGB spectral domain translation task in terms of PSNR. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
VCIP | 2 |
| 2022 | Temporal-MPI: Enabling Multi-plane Images for Dynamic Scene Modelling via Temporal Basis Learning
Wenpeng Xing, Jie Chen 0026 |
ECCV (15) | 2 |
| 2022 | NDF: Neural Deformable Fields for Dynamic Human Modelling
Jie Chen 0026 |
ECCV (32) | 2 |
| 2022 | NEX+: Novel View Synthesis with Neural Regularisation Over Multi-Plane ImagesabstractWe propose Nex+, a neural Multi-Plane Image (MPI) representation with alpha denoising for the task of novel view synthesis (NVS). Overfitting to training data is a common challenge for all learning-based models. We propose a novel solution for resolving such issue in the context of NVS with signal denoising-motivated operations over the alpha coefficients of the MPI, without any additional requirements for supervision. Nex+contains a novel 5D Alpha Neural Regulariser (ANR), which favors low-frequency components in the angular domain, i.e., the alpha coefficients’ signal sub-space indicating various viewing directions. ANR’s angular low-frequency property derives from its small number of angular encoding levels and output basis. The regularised alpha in Nex+can model the scene geometry more accurately than Nex, and outperforms other state-of-the-art methods on public datasets for the task of NVS. Wenpeng Xing, Jie Chen 0026 |
ICASSP | 2 |
| 2022 | ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic GraphabstractTo generate dance that temporally and aesthetically matches the music is a challenging problem, as the following factors need to be considered. First, the aesthetic styles and messages conveyed by the motion and music should be consistent. Second, the beats of the generated motion should be locally aligned to the musical features. And finally, basic choreomusical rules should be observed, and the motion generated should be diverse. To address these challenges, we propose ChoreoGraph, which choreographs high-quality dance motion for a given piece of music over a Dynamic Graph. A data-driven learning strategy is proposed to evaluate the aesthetic style and rhythmic connections between music and motion in a progressively learned cross-modality embedding space. The motion sequences will be beats-aligned based on the music segments and then incorporated as nodes of a Dynamic Motion Graph. Compatibility factors such as the style and tempo consistency, motion context connection, action completeness, and transition smoothness are comprehensively evaluated to determine the node transition in the graph. We demonstrate that our repertoire-based framework can generate motions with aesthetic consistency and robustly extensible in diversity. Both quantitative and qualitative experiment results show that our proposed model outperforms other baseline models. Ho Yin Au, Jie Chen 0026, Junkun Jiang, Yike Guo |
ACM Multimedia | 2 |
| 2022 | A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionabstractMulti-person motion capture can be challenging due to ambiguities caused by severe occlusion, fast body movement, and complex interactions. Existing frameworks build on 2D pose estimations and triangulate to 3D coordinates via reasoning the appearance, trajectory, and geometric consistencies among multi-camera observations. However, 2D joint detection is usually incomplete and with wrong identity assignments due to limited observation angle, which leads to noisy 3D triangulation results. To overcome this issue, we propose to explore the short-range autoregressive characteristics of skeletal motion using transformer. First, we propose an adaptive, identity-aware triangulation module to reconstruct 3D joints and identify the missing joints for each identity. To generate complete 3D skeletal motion, we then propose a Dual-Masked Auto-Encoder (D-MAE) which encodes the joint status with both skeletal-structural and temporal position encoding for trajectory completion. D-MAE's flexible masking and encoding mechanism enable arbitrary skeleton definitions to be conveniently deployed under the same framework. In order to demonstrate the proposed model's capability in dealing with severe data loss scenarios, we contribute a high-accuracy and challenging motion capture dataset of multi-person interactions with severe occlusion. Evaluations on both benchmark and our new dataset demonstrate the efficiency of our proposed model, as well as its advantage against the other state-of-the-art methods. Junkun Jiang, Jie Chen 0026, Yike Guo |
ACM Multimedia | 2 |
| 2022 | MVSPlenOctree: Fast and Generic Reconstruction of Radiance Fields in PlenOctree from Multi-view StereoabstractWe present MVSPlenOctree, a novel approach that can efficiently reconstruct radiance fields for view synthesis. Unlike previous scene-specific radiance fields reconstruction methods, we present a generic pipeline that can efficiently reconstruct 360-degree-renderable radiance fields via multi-view stereo (MVS) inference from tens of sparse-spread out images. Our approach leverages variance-based statistic features for MVS inference, and combines this with image based rendering and volume rendering for radiance field reconstruction. We first train a MVS Machine for reasoning scene's density and appearance. Then, based on the spatial hierarchy of the PlenOctree and coarse-to-fine dense sampling mechanism, we design a robust and efficient sampling strategy for PlenOctree reconstruction, which handles occlusion robustly. A 360-degree-renderable radiance fields can be reconstructed in PlenOctree from MVS Machine in an efficient single forward pass. We trained our method on real-world DTU, LLFF datasets, and synthetic datasets. We validate its generalizability by evaluating on the test set of DTU dataset which are unseen in training. In summary, our radiance field reconstruction method is both efficient and generic, a coarse 360-degree-renderable radiance field can be reconstructed in seconds and a dense one within minutes. Please visit the project page for more details: https://derry-xing.github.io/projects/MVSPlenOctree. Wenpeng Xing, Jie Chen 0026 |
ACM Multimedia | 2 |
| 2022 | Deep Spatial-Angular Regularization for Light Field Imaging, Denoising, and Super-ResolutionabstractCoded aperture is a promising approach for capturing the 4-D light field (LF), in which the 4-D data are compressively modulated into 2-D coded measurements that are further decoded by reconstruction algorithms. The bottleneck lies in the reconstruction algorithms, resulting in rather limited reconstruction quality. To tackle this challenge, we propose a novel learning-based framework for the reconstruction of high-quality LFs from acquisitions via learned coded apertures. The proposed method incorporates the measurement observation into the deep learning framework elegantly to avoid relying entirely on data-driven priors for LF reconstruction. Specifically, we first formulate the compressive LF reconstruction as an inverse problem with an implicit regularization term. Then, we construct the regularization term with a deep efficient spatial-angular separable convolutional sub-network in the form of local and global residual learning to comprehensively explore the signal distribution free from the limited representation ability and inefficiency of deterministic mathematical modeling. Furthermore, we extend this pipeline to LF denoising and spatial super-resolution, which could be considered as variants of coded aperture imaging equipped with different degradation matrices. Extensive experimental results demonstrate that the proposed methods outperform state-of-the-art approaches to a significant extent both quantitatively and qualitatively, i.e., the reconstructed LFs not only achieve much higher PSNR/SSIM but also preserve the LF parallax structure better on both real and synthetic LF benchmarks. The code will be publicly available at https://github.com/MantangGuo/DRLF. Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Deep Coarse-to-Fine Dense Light Field Reconstruction With Flexible Sampling and Geometry-Aware FusionabstractA densely-sampled light field (LF) is highly desirable in various applications, such as 3-D reconstruction, post-capture refocusing and virtual reality. However, it is costly to acquire such data. Although many computational methods have been proposed to reconstruct a densely-sampled LF from a sparsely-sampled one, they still suffer from either low reconstruction quality, low computational efficiency, or the restriction on the regularity of the sampling pattern. To this end, we propose a novel learning-based method, which accepts sparsely-sampled LFs with irregular structures, and produces densely-sampled LFs with arbitrary angular resolution accurately and efficiently. We also propose a simple yet effective method for optimizing the sampling pattern. Our proposed method, an end-to-end trainable network, reconstructs a densely-sampled LF in a coarse-to-fine manner. Specifically, the coarse sub-aperture image (SAI) synthesis module first explores the scene geometry from an unstructured sparsely-sampled LF and leverages it to independently synthesize novel SAIs, in which a confidence-based blending strategy is proposed to fuse the information from different input SAIs, giving an intermediate densely-sampled LF. Then, the efficient LF refinement module learns the angular relationship within the intermediate result to recover the LF parallax structure. Comprehensive experimental evaluations demonstrate the superiority of our method on both real-world and synthetic LF images when compared with state-of-the-art methods. In addition, we illustrate the benefits and advantages of the proposed approach when applied in various LF-based applications, including image-based rendering and depth estimation enhancement. The code is available at https://github.com/jingjin25/LFASR-FS-GAF. Jing Jin 0006, Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Sam Kwong, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Attention-Guided Progressive Neural Texture Fusion for High Dynamic Range Image RestorationabstractHigh Dynamic Range (HDR) imaging via multi-exposure fusion is an important task for most modern imaging platforms. In spite of recent developments in both hardware and algorithm innovations, challenges remain over content association ambiguities caused by saturation, motion, and various artifacts introduced during multi-exposure fusion such as ghosting, noise, and blur. In this work, we propose an Attention-guided Progressive Neural Texture Fusion (APNT-Fusion) HDR restoration model which aims to address these issues within one framework. An efficient two-stream structure is proposed which separately focuses on texture feature transfer over saturated regions and multi-exposure tonal and texture feature fusion. A neural feature transfer mechanism is proposed which establishes spatial correspondence between different exposures based on multi-scale VGG features in the masked saturated HDR domain for discriminative contextual clues over the ambiguous image areas. A progressive texture blending module is designed to blend the encoded two-stream features in a multi-scale and progressive manner. In addition, we introduce several novel attention mechanisms, i.e., the motion attention module detects and suppresses the content discrepancies among the reference images; the saturation attention module facilitates differentiating the misalignment caused by saturation from those caused by motion; and the scale attention module ensures texture blending consistency between different coder/decoder scales. We carry out comprehensive qualitative and quantitative evaluations and ablation studies, which validate that these novel modules work coherently under the same framework and outperform state-of-the-art methods. Jie Chen 0026, Zaifeng Yang, Tsz Nam Chan, Hui Li 0029, Junhui Hou, Lap-Pui Chau |
IEEE Trans. Image Process. | 1 |
| 2022 | Scale-Consistent Fusion: From Heterogeneous Local Sampling to Global Immersive RenderingabstractImage-based geometric modeling and novel view synthesis based on sparse large-baseline samplings are challenging but important tasks for emerging multimedia applications such as virtual reality and immersive telepresence. Existing methods fail to produce satisfactory results due to the limitation on inferring reliable depth information over such challenging reference conditions. With the popularization of commercial light field (LF) cameras, capturing LF images (LFIs) is as convenient as taking regular photos, and geometry information can be reliably inferred. This inspires us to use a sparse set of LF captures to render high-quality novel views globally. However, the fusion of LF captures from multiple angles is challenging due to the scale inconsistency caused by various capture settings. To overcome this challenge, we propose a novel scale-consistent volume rescaling algorithm that robustly aligns the disparity probability volumes (DPV) among different captures for scale-consistent global geometry fusion. Based on the fused DPV projected to the target camera frustum, novel learning-based modules (i.e., the attention-guided multi-scale residual fusion module, and the disparity field-guided deep re-regularization module), which comprehensively regularize noisy observations from heterogeneous captures for high-quality rendering of novel LFIs, have been proposed. Both quantitative and qualitative experiments over the Stanford Lytro Multi-view LF dataset show that the proposed method outperforms state-of-the-art methods significantly under different experiment settings for disparity inference and LF synthesis. Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Qiang Wang 0022, Yike Guo |
IEEE Trans. Image Process. | 2 |
| 2021 | A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced DataabstractIn this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the first learning stage, pre-processed volumetric CT data together with the segmented lung masks are fed into a 3D ResNet module, and an initial classification result can be obtained. However, due to categorical data imbalance, we observe large differences in sensitivity between COVID-19 and CAP cases. In the second stage, five learning models are independently trained over data with only COVID-19 and CAP cases, and are then ensembled to further discriminate the two classes. The final classification results are obtained by combining the predictions from both stages. Based on the validation dataset, we have evaluated our method and compared it with up-to-date methods in terms of overall accuracy and sensitivity for each class. The validation results validate the accuracy of the proposed multi-stage learning strategy. The overall accuracy of the validation dataset is 88.8%, and the sensitivities are 0.873, 0.789 and 1 for COVID-19, CAP and normal cases, respectively. Zaifeng Yang, Yubo Hou, Zhenghua Chen, Le Zhang 0001, Jie Chen 0026 |
ICASSP | 5 |
| 2021 | Hyperspectral Image Super-Resolution via Deep Progressive Zero-Centric Residual LearningabstractThis paper explores the problem of hyperspectral image (HSI) super-resolution that merges a low resolution HSI (LR-HSI) and a high resolution multispectral image (HR-MSI). The cross-modality distribution of the spatial and spectral information makes the problem challenging. Inspired by the classic wavelet decomposition-based image fusion, we propose a novel lightweight deep neural network-based framework, namely progressive zero-centric residual network (PZRes-Net), to address this problem efficiently and effectively. Specifically, PZRes-Net learns a high resolution and zero-centric residual image, which contains high-frequency spatial details of the scene across all spectral bands, from both inputs in a progressive fashion along the spectral dimension. And the resulting residual image is then superimposed onto the up-sampled LR-HSI in a mean-value invariant manner, leading to a coarse HR-HSI, which is further refined by exploring the coherence across all spectral bands simultaneously. To learn the residual image efficiently and effectively, we employ spectral-spatial separable convolution with dense connections. In addition, we propose zero-mean normalization implemented on the feature maps of each layer to realize the zero-mean characteristic of the residual image. Extensive experiments over both real and synthetic benchmark datasets demonstrate that our PZRes-Net outperforms state-of-the-art methods to a significant extent in terms of both 4 quantitative metrics and visual quality, e.g., our PZRes-Net improves the PSNR more than 3dB, while saving 2.3× parameters and consuming 15× less FLOPs. The code is publicly available at https://github.com/zbzhzhy/PZRes-Net. Junhui Hou, Jie Chen 0026, Huanqiang Zeng, Jiantao Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Light Field Spatial Super-Resolution via Deep Combinatorial Geometry Embedding and Structural Consistency RegularizationabstractLight field (LF) images acquired by hand-held devices usually suffer from low spatial resolution as the limited sampling resources have to be shared with the angular dimension. LF spatial super-resolution (SR) thus becomes an indispensable part of the LF camera processing pipeline. The high-dimensionality characteristic and complex geometrical structure of LF images makes the problem more challenging than traditional single-image SR. The performance of existing methods are still limited as they fail to thoroughly explore the coherence among LF views and are insufficient in accurately preserving the parallax structure of the scene. In this paper, we propose a novel learning-based LF spatial SR framework, in which each view of an LF image is first individually super-resolved by exploring the complementary information among views with combinatorial geometry embedding. For accurate preservation of the parallax structure among the reconstructed views, a regularization network trained over a structure-aware loss function is subsequently appended to enforce correct parallax relationships over the intermediate estimation. Our proposed approach is evaluated over datasets with a large number of testing images including both synthetic and real-world scenes. Experimental results demonstrate the advantage of our approach over state-of-the-art methods, i.e., our method not only improves the average PSNR by more than 1.0 dB but also preserves more accurate parallax details, at a lower computation cost. Jing Jin 0006, Junhui Hou, Jie Chen 0026, Sam Kwong |
CVPR | 3 |
| 2020 | Deep Spatial-Angular Regularization for Compressive Light Field Reconstruction over Coded Apertures
Mantang Guo, Junhui Hou, Jing Jin 0006, Jie Chen 0026, Lap-Pui Chau |
ECCV (2) | 4 |
| 2020 | Surface Consistent Light Field Extrapolation Over Stratified Disparity And Spatial GranularitiesabstractThe light field captures both the spatial and angular configurations of the scene, which facilitates a wide range of imaging possibilities. In this work, we propose an LF view extrapolation algorithm which renders high quality novel LF views far outside the range of given angular baselines. A stratified synthesis strategy is adopted which projects the scene content based on stratified disparity layers and across varying scales of spatial granularities. Such a stratified methodology proves to help preserve scene structures over large angular shifts, and provide informative clues for inferring the contents of occluded regions. A generative-adversarial network model is further adopted for parallax correction and occlusion completion conditioned on surface consistent feature. Experiments show that our proposed model can provide more reliable novel view extrapolation quality at large baseline extension ratios compared with state-of-the-art LF synthesis algorithms. Jie Chen 0026, Lap-Pui Chau, Junhui Hou |
ICME | 1 |
| 2020 | Accurate Light Field Depth Estimation via an Occlusion-Aware NetworkabstractDepth estimation is a fundamental problem for light field based applications. Although recent learning-based methods have proven to be effective for light field depth estimation, they still have troubles when handling occlusion regions. In this paper, by leveraging the explicitly learned occlusion map, we propose an occlusion-aware network, which is capable of estimating accurate depth maps with sharp edges. Our main idea is to separate the depth estimation on non-occlusion and occlusion regions, as they contain different properties with respect to the light field structure, i.e., obeying and violating the angular photo consistency constraint. To this end, three modules are involved in our network: the occlusion region detection network (ORDNet), the coarse depth estimation network (CDENet), and the refined depth estimation network (RDENet). Specifically, ORDNet predicts the occlusion map as a mask, while under the guidance of the resulting occlusion map, CDENet and REDNet focus on the depth estimation on non-occlusion and occlusion areas, respectively. Experimental results show that our method achieves better performance on 4D light field benchmark, especially in occlusion regions, when compared with current state-of-the-art light-field depth estimation algorithms. Chunle Guo, Jing Jin 0006, Junhui Hou, Jie Chen 0026 |
ICME | 4 |
| 2020 | Haze Removal with Fusion of Local and Non-Local StatisticsabstractMost of the outdoor images suffer from contrast degradation caused by fog and haze. Two statistical frameworks have been proposed in recent years that exploit local (dark channel prior) and non-local (haze-lines) characteristics of hazy images for the estimation of scene configurations and the restoration of scene albedo. Both frameworks show intrinsic limitations due to the basic assumptions they rely on. In this paper we propose a novel dehazing method that combines the advantages of local and non-local dehazing methods. Exploiting their complementary statistical properties, we use the local features to regulate the estimation of non-local haze-lines for a better final restoration at challenging regions. Both quantitative and qualitative results validate the effectiveness of our proposed method over state-of-the-art frameworks. Jie Chen 0026, Cheen-Hau Tan, Lap-Pui Chau |
ISCAS | 1 |
| 2020 | Light Field Super-resolution via Attention-Guided Fusion of Hybrid LensesabstractThis paper explores the problem of reconstructing high-resolution light field (LF) images from hybrid lenses, including a high-resolution camera surrounded by multiple low-resolution cameras. To tackle this challenge, we propose a novel end-to-end learning-based approach, which can comprehensively utilize the specific characteristics of the input from two complementary and parallel perspectives. Specifically, one module regresses a spatially consistent intermediate estimation by learning a deep multidimensional and cross-domain feature representation; the other one constructs another intermediate estimation, which maintains the high-frequency textures, by propagating the information of the high-resolution view. We finally leverage the advantages of the two intermediate estimations via the learned attention maps, leading to the final high-resolution LF image. Extensive experiments demonstrate the significant superiority of our approach over state-of-the-art ones. That is, our method not only improves the PSNR by more than 2 dB, but also preserves the LF structure much better. To the best of our knowledge, this is the first end-to-end deep learning method for reconstructing a high-resolution LF image with a hybrid input. We believe our framework could potentially decrease the cost of high-resolution LF data acquisition and also be beneficial to LF data storage and transmission. The code is available at https://github.com/jingjin25/LFhybridSR-Fusion. Jing Jin 0006, Junhui Hou, Jie Chen 0026, Sam Kwong, Jingyi Yu 0001 |
ACM Multimedia | 3 |
| 2019 | Light Field Image Compression Based on Bi-Level View Compensation With Rate-Distortion OptimizationabstractCompared with conventional color images, light field images (LFIs) contain richer scene information, which allows a wide range of interesting applications. However, such additional information is obtained at the cost of generating substantially more data, which poses challenges to both data storage and transmission. In this paper, we propose a new hybrid framework for effective compression of LFIs. The proposed framework takes the particular characteristics of LFIs into account so that the inter- and intra-view correlations of LFIs can be more efficiently exploited to produce better compression performance. Specifically, the proposed scheme partitions sub-aperture images (SAIs) of an LFI into two groups, namely, key SAIs and non-key SAIs. Bi-level view compensation is proposed to exploit the inter-view correlation: first, based on the group of selected key SAIs, learning-based angular super-resolution is performed to compensate non-key SAIs in pixel-wise, during which heterogeneous inter-view correlation between the non-key SAIs is efficiently removed; second, the two groups of SAIs are respectively reorganized as pseudo-sequences, and block-wise motion compensation is carried out with a standard video encoder, during which the homogeneous inter-view correlation is subsequently exploited. The video encoder also helps to remove the intra-view correlation of the SAIs and finally generates the encoded bitstream. Moreover, the bits allocated to each group are optimally determined via model-based rate distortion optimization. Extensive experimental evaluations and comparisons demonstrate the advantage of the proposed framework over existing methods in terms of rate-distortion performance. Junhui Hou, Jie Chen 0026, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Light Field Spatial Super-Resolution Using Deep Efficient Spatial-Angular Separable ConvolutionabstractLight field (LF) photography is an emerging paradigm for capturing more immersive representations of the real-world. However, arising from the inherent trade-off between the angular and spatial dimensions, the spatial resolution of LF images captured by commercial micro-lens based LF cameras are significantly constrained. In this paper, we propose effective and efficient end-to-end convolutional neural network models for spatially super-resolving LF images. Specifically, the proposed models have an hourglass shape, which allows feature extraction to be performed at the low resolution level to save both computational and memory costs. To fully make use of the four-dimensional (4-D) structure information of LF data in both spatial and angular domains, we propose to use 4-D convolution to characterize the relationship among pixels. Moreover, as an approximation of 4-D convolution, we also propose to use spatialangular separable (SAS) convolutions for more computationallyand memory-efficient extraction of spatial-angular joint features. Extensive experimental results on 57 test LF images with various challenging natural scenes show significant advantages from the proposed models over state-of-the-art methods. That is, an average PSNR gain of more than 3.0 dB and better visual quality are achieved, and our methods preserve the LF structure of the super-resolved LF images better, which is highly desirable for subsequent applications. In addition, the SAS convolutionbased model can achieve 3× speed up with only negligible reconstruction quality decrease when compared with the 4-D convolution-based one. The source code of our method is online available at https://github.com/spatialsr/DeepLightFieldSSR. Henry Wing Fung Yeung, Junhui Hou, Xiaoming Chen 0006, Jie Chen 0026, Zhibo Chen 0001, Vera Chung |
IEEE Trans. Image Process. | 4 |
| 2018 | Robust Video Content Alignment and Compensation for Rain Removal in a CNN FrameworkabstractRain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which applies superpixel (SP) segmentation to decompose the scene into depth consistent units. Alignment of scene contents are done at the SP level, which proves to be robust towards rain occlusion and fast camera motion. Two alignment output tensors, i.e., optimal temporal match tensor and sorted spatial-temporal match tensor, provide informative clues for rain streak location and occluded background contents to generate an intermediate derain output. These tensors will be subsequently prepared as input features for a convolutional neural network to restore high frequency details to the intermediate output for compensation of mis-alignment blur. Extensive evaluations show that up to 5dB reconstruction PSNR advantage is achieved over state-of-the-art methods. Visual inspection shows that much cleaner rain removal is achieved especially for highly dynamic scenes with heavy and opaque rainfall from a fast moving camera. Jie Chen 0026, Cheen-Hau Tan, Junhui Hou, Lap-Pui Chau |
CVPR | 1 |
| 2018 | Fast Light Field Reconstruction with Deep Coarse-to-Fine Modeling of Spatial-Angular Clues
Henry Wing Fung Yeung, Junhui Hou, Jie Chen 0026, Vera Chung, Xiaoming Chen 0006 |
ECCV (6) | 3 |
| 2018 | Light Field Denoising via Anisotropic Parallax Analysis in a CNN FrameworkabstractLight field (LF) cameras provide perspective information of scenes by taking directional measurements of the focusing light rays. The raw outputs are usually dark with additive camera noise, which impedes subsequent processing and applications. We propose a novel LF denoising framework based on anisotropic parallax analysis (APA). Two convolutional neural networks are jointly designed for the task: first, the structural parallax synthesis network predicts the parallax details for the entire LF based on a set of anisotropic parallax features. These novel features can efficiently capture the high-frequency perspective components of a LF from noisy observations. Second, the view-dependent detail compensation network restores non-Lambertian variation to each LF view by involving view-specific spatial energies. Extensive experiments show that the proposed APA LF denoiser provides a much better denoising performance than state-of-the-art methods in terms of visual quality and in preservation of parallax details. Jie Chen 0026, Junhui Hou, Lap-Pui Chau |
IEEE Signal Process. Lett. | 1 |
| 2018 | Simultaneous Spatial and Spectral Low-Rank Representation of Hyperspectral Images for ClassificationabstractArising from various environmental and atmos- pheric conditions and sensor interference, spectral variations are inevitable during hyperspectral remote sensing, which degrade the subsequent hyperspectral image analysis significantly. In this paper, we propose simultaneous spatial and spectral low-rank representation (S3LRR) that can effectively suppress the within-class spectral variations for classification purposes. The S3LRR recovers an intrinsic component with the same dimension as the original image, in which both spatial and spectral low-rank priors are adopted to regularize the intrinsic component simultaneously and compensate to each other, together with robust modeling of spectral variations. Compared with existing methods that explore only the spectral low-rank prior, the novel spatial low-rank prior (i.e., low-rank prior in band-wise) can take the spatial structure information of hyperspectral images into account, which has demonstrated to be very useful. Technically, we formulate S3LRR as a constrained convex optimization problem, and solve it using the efficient inexact augmented Lagrangian multiplier method. The resulting intrinsic component is less interfered by within-class spectral variations, and more discriminatory to offer higher classification accuracy. Comprehensive experiments on benchmark data sets demonstrate that the proposed S3LRR improves classification accuracy significantly, which outperforms state-of-the-art methods. Shaohui Mei, Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Light Field Compression With Disparity-Guided Sparse Coding Based on Structural Key ViewsabstractRecent imaging technologies are rapidly evolving for sampling richer and more immersive representations of the 3D world. One of the emerging technologies is light field (LF) cameras based on micro-lens arrays. To record the directional information of the light rays, a much larger storage space and transmission bandwidth are required by an LF image as compared with a conventional 2D image of similar spatial dimension. Hence, the compression of LF data becomes a vital part of its application. In this paper, we propose an LF codec with disparity guided Sparse Coding over a learned perspective-shifted LF dictionary based on selected Structural Key Views (SC-SKV). The sparse coding is based on a limited number of optimally selected SKVs; yet the entire LF can be recovered from the coding coefficients. By keeping the approximation identical between encoder and decoder, only the residuals of the non-key views, disparity map, and the SKVs need to be compressed into the bit stream. An optimized SKV selection method is proposed such that most LF spatial information can be preserved. To achieve optimum dictionary efficiency, the LF is divided into several coding regions, over which the reconstruction works individually. Experiments and comparisons have been carried out over benchmark LF data set, which show that the proposed SC-SKV codec produces convincing compression results in terms of both rate-distortion performance and visual quality compared with Joint Exploration Model: with 37.9% BD-rate reduction and 1.17-dB BD-PSNR improvement achieved on average, especially with up to 6-dB improvement for low bit rate scenarios. Jie Chen 0026, Junhui Hou, Lap-Pui Chau |
IEEE Trans. Image Process. | 1 |
| 2018 | Accurate Light Field Depth Estimation With Superpixel Regularization Over Partially Occluded RegionsabstractDepth estimation is a fundamental problem for light field photography applications. Numerous methods have been proposed in recent years, which either focus on crafting cost terms for more robust matching, or on analyzing the geometry of scene structures embedded in the epipolar-plane images. Significant improvements have been made in terms of overall depth estimation error; however, current state-of-the-art methods still show limitations in handling intricate occluding structures and complex scenes with multiple occlusions. To address these challenging issues, we propose a very effective depth estimation framework which focuses on regularizing the initial label confidence map and edge strength weights. Specifically, we first detect partially occluded boundary regions (POBR) via superpixel-based regularization. Series of shrinkage/reinforcement operations are then applied on the label confidence map and edge strength weights over the POBR. We show that after weight manipulations, even a low-complexity weighted least squares model can produce much better depth estimation than the state-of-the-art methods in terms of average disparity error rate, occlusion boundary precision-recall rate, and the preservation of intricate visual features. Jie Chen 0026, Junhui Hou, Yun Ni, Lap-Pui Chau |
IEEE Trans. Image Process. | 1 |
| 2017 | Reflection removal based on single light field captureabstractPhotography through reflective surfaces suffers from the obstruction of reflections, which deteriorates the visibility of background targets and causes challenges for subsequent computer vision applications. In this paper, we propose a novel reflection removal algorithm using light field (LF) cameras. Unlike conventional cameras, LF cameras capture extra directional information of incoming rays which enable our algorithm to remove reflections with only a single shot. We analyze the optical geometry of the background and reflection imagery in a LF camera, and generalize a set of rules that could facilitate our algorithm to differentiate the edges from different optical sources. We show that the proposed method produces significantly better reflection removal results based on the LF data as compared to traditional methods based on multiple-shot image sequences as input. Yun Ni, Jie Chen 0026, Lap-Pui Chau |
ISCAS | 2 |
| 2017 | Light Field Compressed Sensing Over a Disparity-Aware DictionaryabstractLight field (LF) acquisition faces the challenge of extremely bulky data. Available hardware solutions usually compromise the sensor resource between spatial and angular resolutions. In this paper, a compressed sensing framework is proposed for the sampling and reconstruction of a high-resolution LF based on a coded aperture camera. First, an LF dictionary based on perspective shifting is proposed for the sparse representation of the highly correlated LF. Then, two separate methods, i.e., subaperture scan and normalized fluctuation, are proposed to acquire/calculate the scene disparity, which will be used during the LF reconstruction with the proposed disparity-aware dictionary. At last, a hardware implementation of the proposed LF acquisition/reconstruction scheme is carried out. Both quantitative and qualitative evaluation show that the proposed methods produce the state-of-the-art performance in both reconstruction quality and computation efficiency. Jie Chen 0026, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Sparse two-dimensional singular value decompositionabstractIn this paper, we propose a new data-driven transform, called sparse two-dimensional singular value decomposition (S2DSVD). By leveraging the advantages of discrete cosine transform and the conventional 2D SVD, we decompose a set of matrices into transform coefficient matrices with sparse and orthogonal basis functions. Such sparsity characteristic can significantly reduce their overhead, hence being beneficial to data compression. We formulate S2DSVD as a constrained optimization problem and solve it via alternative iteration. We demonstrate the efficacy of S2DSVD on image and video datasets, and observe that it can produce results with error comparable to 2D SVD whereas its space complexity is significantly smaller than 2D SVD. Junhui Hou, Jie Chen 0026, Lap-Pui Chau, Ying He 0001 |
ICME | 2 |
| 2015 | Heavy haze removal in a learning frameworkabstractExtreme weather hazards happens more often these days due to climate changes and increased human industrial activities, and one of most notorious of them is haze. State-of-the-art haze removal methods generally work well with light haze conditions, however when haze gets heavier, the physical model tend to produce over-shadowed, noisy, and color distorted restorations. A new physical model has been proposed in this paper for heavy haze weathers. An airlight vector map has been proposed to address the problem caused by uneven aerosol distribution w.r.t. altitude variation. A Random Decision Forest model has been adopted to deal with the additional light attenuation and transmission map underestimation problem caused by heavy haze. Experiment shows the proposed model produces much better visual restoration for heavy haze weathers compared to state-of-the-art methods in terms of colour fidelity, noise reduction, and overall contrast. Jie Chen 0026, Lap-Pui Chau |
ISCAS | 1 |
| 2015 | Multiscale Dictionary Learning via Cross-Scale Cooperative Learning and Atom Clustering for Visual Signal ProcessingabstractFor sparse signal representation, the sparsity across the scales is a promising yet underinvestigated direction. In this paper, we aim to design a multiscale sparse representation scheme to explore such potential. A multiscale dictionary (MD) structure is designed. A cross-scale matching pursuit algorithm is proposed for multiscale sparse coding. Two dictionary learning methods, cross-scale cooperative learning and cross-scale atom clustering, are proposed each focusing on one of the two important attributes of an efficient MD: the similarity and uniqueness of corresponding atoms in different scales. We analyze and compare their different advantages in the application of image denoising under different noise levels, where both methods produce state-of-the-art denoising results. Jie Chen 0026, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | A fast adaptive guided filtering algorithm for light field depth interpolationabstractLight field camera provides 4D information of the light rays, from which the scene depth information can be inferred. The disparity/depth maps calculated from light field data are always noisy with missing and false entries in homogeneous regions or areas where view-dependant effects are present. In this paper we proposed an adaptive guided filtering (AGF) algorithm to get an optimized output disparity/depth map. A guidance image is used to provide the image contour and texture information, the filter is able to preserve the disparity edges, smooth the regions without influence of the image texture, and reject the data entries with low confidence during coefficients regression. Experiment shows AGF is much faster in implementation as compared to other variational or hierarchical based optimization algorithms, and produces competitive visual results. Jie Chen 0026, Lap-Pui Chau |
ISCAS | 1 |
| 2014 | A Rain Pixel Recovery Algorithm for Videos With Highly Dynamic ScenesabstractRain removal is a very useful and important technique in applications such as security surveillance and movie editing. Several rain removal algorithms have been proposed these years, where photometric, chromatic, and probabilistic properties of the rain have been exploited to detect and remove the rainy effect. Current methods generally work well with light rain and relatively static scenes, when dealing with heavier rainfall in dynamic scenes, these methods give very poor visual results. The proposed algorithm is based on motion segmentation of dynamic scene. After applying photometric and chromatic constraints for rain detection, rain removal filters are applied on pixels such that their dynamic property as well as motion occlusion clue are considered; both spatial and temporal informations are then adaptively exploited during rain pixel recovery. Results show that the proposed algorithm has a much better performance for rainy scenes with large motion than existing algorithms. Jie Chen 0026, Lap-Pui Chau |
IEEE Trans. Image Process. | 1 |
| 2013 | An enhanced window-variant dark channel prior for depth estimation using single foggy imageabstractThe dark channel prior is a simple yet efficient way to estimate the scene depth information using one single foggy image. However the prior fails for pixels with low colour saturation. Based on the observation that areas with dramatic colour changes tend to belong to similar depth, a window variation mechanism is proposed in this paper based on the neighbourhood scene complexity and colour saturation rate to achieve an ideal compromise between depth resolution and precision. The proposed method greatly alleviates the intrinsic drawbacks of the original dark channel prior. Experiments show the proposed method produces more accurate depth estimation in most of the scenes than the original prior. Jie Chen 0026, Lap-Pui Chau |
ICIP | 1 |
| 2013 | Human motion capture data recovery via trajectory-based sparse representationabstractMotion capture is widely used in sports, entertainment and medical applications. An important issue is to recover motion capture data that has been corrupted by noise and missing data entries during acquisition. In this paper, we propose a new method to recover corrupted motion capture data through trajectory-based sparse representation. The data is firstly represented as trajectories with fixed length and high correlation. Then, based on the sparse representation theory, the original trajectories can be recovered by solving the sparse representation of the incomplete trajectories through the OMP algorithm using a dictionary learned by K-SVD. Experimental results show that the proposed algorithm achieves much better performance, especially when significant portions of data is missing, than the existing algorithms. Junhui Hou, Lap-Pui Chau, Ying He 0001, Jie Chen 0026, Nadia Magnenat-Thalmann |
ICIP | 4 |
| 2013 | A novel SVD-based image quality assessment metricabstractImage distortion can be categorized into two aspects: content-dependent degradation and content-independent one. An existing full-reference image quality assessment (IQA) metric cannot deal with these two different impacts well. Singular value decomposition (SVD) as a useful mathematical tool has been used in various image processing applications. In this paper, SVD is employed to separate the structural (content-dependent) and the content-independent components. For each portion, we design a specific assessment model to tailor for its corresponding distortion properties. The proposed models are then fused to obtain the final quality score. Experimental results with the TID database demonstrate that the proposed metric achieves better performance in comparison with the relevant state-of-the-art quality metrics. Shuigen Wang, Chenwei Deng, Weisi Lin, Baojun Zhao, Jie Chen 0026 |
ICIP | 5 |
| 2013 | Rain removal from dynamic scene based on motion segmentationabstractRain removal technique has been intensively studied over these years, the photometric, chromatic, and probabilistic properties of the rain have been exploited to remove the rainy effect. However, current available algorithms only work well with light rain and static scenes, when dealing with heavier rain fall in dynamic scenes, obvious visual degradation will occur especially in motion intensive areas. The proposed algorithm is based on motion segmentation of dynamic scenes. Photometric and chromatic constraints are used for rain detection, motion occlusion information are involved in the adaptive prediction of the rain pixels' original value, using both spatial and temporal neighbor information. Results show the proposed algorithm has a much better performance for rainy scenes with large motion than existing algorithms. Jie Chen 0026, Lap-Pui Chau |
ISCAS | 1 |