EDBT 2026 Demo / reviewers in the wild / expert
Zaifeng Yang
dblp:183/5576
· DBLP profile ↗
17ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-7667-5309ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAOT: An Enhanced Locality-Aware Spectral Transformer for Solving PDEsabstractNeural operators have shown great potential in solving a family of Partial Differential Equations (PDEs) by modeling the mappings between input and output functions. Fourier Neural Operator (FNO) implements global convolutions via parameterizing the integral operators in Fourier space. However, it often results in over-smoothing solutions and fails to capture local details and high-frequency components. To address these limitations, we investigate incorporating the spatial-frequency localization property of Wavelet transforms into the Transformer architecture. We propose a novel Wavelet Attention (WA) module with linear computational complexity to efficiently learn locality-aware features. Building upon WA, we further develop the Spectral Attention Operator Transformer (SAOT), a hybrid spectral Transformer framework that integrates WA’s localized focus with the global receptive field of Fourier-based Attention (FA) through a gated fusion block. Experimental results demonstrate that WA significantly mitigates the limitations of FA and outperforms existing Wavelet-based neural operators by a large margin. By integrating the locality-aware and global spectral representations, SAOT achieves state-of-the-art performance on six operator learning benchmarks and exhibits strong discretization-invariant ability. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang |
AAAI | 3 |
| 2025 | NCDI-Diffusion: Neural Contextual and Directional Inversion for Novel View Synthesis through Diffusion ModelsabstractNovel view synthesis typically requires a comprehensive set of multi-view images for either image-based rendering or scene representation-based optimization. However, achieving high-fidelity novel view rendering often demands a large number of images. To address this limitation, we propose NCDI-Diffusion, a novel diffusion-based view synthesis method that reduces the number of required images by leveraging the prior knowledge embedded in pre-trained diffusion models. Specifically, NCDI-Diffusion encapsulates both the contextual and directional information of a scene by utilizing neural descriptors, which are inversely derived from a limited set of positioned multi-view training images. These descriptors guide the diffusion model's image synthesis process, enabling the generation of high-quality novel views. Empirical results on the Forward-facing Dataset demonstrate the effectiveness of our approach to novel view synthesis. Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Changting Lin |
ICASSP | 3 |
| 2025 | Learning Physics-Informed Color-Aware Transforms for Low-Light Image EnhancementabstractImage decomposition offers deep insights into the imaging factors of visual data and significantly enhances various advanced computer vision tasks. In this work, we introduce a novel approach to low-light image enhancement based on decomposed physics-informed priors. Existing methods that directly map low-light to normal-light images in the sRGB color space suffer from inconsistent color predictions and high sensitivity to spectral power distribution (SPD) variations, resulting in unstable performance under diverse lighting conditions. To address these challenges, we introduce a Physics-informed Color-aware Transform (PiCat), a learning-based framework that converts low-light images from the sRGB color space into deep illumination-invariant descriptors via our proposed Color-aware Transform (CAT). This transformation enables robust handling of complex lighting and SPD variations. Complementing this, we propose the Content-Noise Decomposition Network (CNDN), which refines the descriptor distributions to better align with well-lit conditions by mitigating noise and other distortions, thereby effectively restoring content representations to low-light images. The CAT and the CNDN collectively act as a physical prior, guiding the transformation process from low-light to normal-light domains. Our proposed PiCat framework demonstrates superior performance compared to state-of-the-art methods across five benchmark datasets. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
ICME | 3 |
| 2025 | Dual-Balancing for Physics-Informed Neural NetworksabstractPhysics-informed neural networks (PINNs) have emerged as a new learning paradigm for solving partial differential equations (PDEs) by enforcing the constraints of physical equations, boundary conditions (BCs), and initial conditions (ICs) into the loss function. Despite their successes, vanilla PINNs still suffer from poor accuracy and slow convergence due to the intractable multi-objective optimization issue. In this paper, we propose a novel Dual-Balanced PINN (DB-PINN), which dynamically adjusts loss weights by integrating inter-balancing and intra-balancing to alleviate two imbalance issues in PINNs. Inter-balancing aims to mitigate the gradient imbalance between PDE residual loss and condition-fitting losses by determining an aggregated weight that offsets their gradient distribution discrepancies. Intra-balancing acts on condition-fitting losses to tackle the imbalance in fitting difficulty across diverse conditions. By evaluating the fitting difficulty based on the loss records, intra-balancing can allocate the aggregated weight proportionally to each condition loss according to its fitting difficulty level. We further introduce a robust weight update strategy to prevent abrupt spikes and arithmetic overflow in instantaneous weight values caused by large loss variances, enabling smooth weight updating and stable training. Extensive experiments demonstrate that DB-PINN achieves significantly superior performance than those popular gradient-based weighting methods in terms of convergence speed and prediction accuracy. Our code and supplementary material are available at https://github.com/chenhong-zhou/DualBalanced-PINNs. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang, Ching Eng Png |
IJCAI | 3 |
| 2025 | Unraveling Metameric Dilemma for Spectral Reconstruction: A High-Fidelity Approach via Semi-Supervised LearningabstractSpectral reconstruction from RGB images often suffers from a metameric dilemma, where distinct spectral distributions map to nearly identical RGB values, making them indistinguishable to current models and leading to unreliable reconstructions.
In this paper, we present Diff-Spectra that integrates supervised physics-aware spectral estimation and unsupervised high-fidelity spectral regularization for HSI reconstruction.
We first introduce an Adaptive illumiChroma Decoupling (AICD) module to decouple illumination and chrominance information, which learns intrinsic and distinctive feature distributions, thereby mitigating the metameric issue.
Then, we incorporate the AICD into a learnable spectral response function (SRF) guided hyperspectral initial estimation mechanism to mimic the physical image formation and thus inject physics-aware reasoning into neural networks, turning an ill-posed problem into a constrained, interpretable task.
We also introduce a metameric spectra augmentation method to synthesize comprehensive hyperspectral data to pre-train a Spectral Diffusion Module (SDM), which internalizes the statistical properties of real-world HSI data, enforcing unsupervised high-fidelity regularization on the spectral transitions via inner-loop optimization during inference.
Extensive experimental evaluations demonstrate that our Diff-Spectra achieves SOTA performance on both Spectral reconstruction and downstream HSI classification. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
NeurIPS | 3 |
| 2024 | Hyperspectral Image Reconstruction via Combinatorial Embedding of Cross-Channel Spatio-Spectral CluesabstractExisting learning-based hyperspectral reconstruction methods show limitations in fully exploiting the information among the hyperspectral bands. As such, we propose to investigate the chromatic inter-dependencies in their respective hyperspectral embedding space. These embedded features can be fully exploited by querying the inter-channel correlations in a combinatorial manner, with the unique and complementary information efficiently fused into the final prediction. We found such independent modeling and combinatorial excavation mechanisms are extremely beneficial to uncover marginal spectral features, especially in the long wavelength bands. In addition, we have proposed a spatio-spectral attention block and a spectrum-fusion attention module, which greatly facilitates the excavation and fusion of information at both semantically long-range levels and fine-grained pixel levels across all dimensions. Extensive quantitative and qualitative experiments show that our method (dubbed CESST) achieves SOTA performance. Code for this project is at: https://github.com/AlexYangxx/CESST. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
AAAI | 3 |
| 2024 | Enhanced Physics-Informed Neural Networks with Optimized Sensor Placement via Multi-Criteria Adaptive SamplingabstractPhysics-informed neural networks (PINNs) have emerged as promising and powerful tools for solving partial differential equations (PDEs). To enforce PINNs that yield accurate solutions to PDEs, a set of scattered spatiotemporal points, known as collocation points, are typically sampled within the computational domain. The choices of collocation points significantly impact the performance of PINNs. However, existing sampling methods primarily rely on the PDE residual, which is insufficient for solutions with steep gradients. To enhance the accuracy of PINNs, we propose a novel multi-criteria adaptive sampling (MCAS) approach to optimally select appropriate collocation points. The MCAS approach integrates three sampling criteria: PDE’s residual, the gradient of residual, and the gradient of solutions, enabling us to capture the PDE violations and the sharpness of solutions. Experimental results demonstrate that the proposed MCAS approach is not only applicable for collocation points but also for optimizing sensor placement, consistently outperforming the residual-based sampling methods. Chenhong Zhou, Jie Chen 0026, Zaifeng Yang, Alexander Matyasko, Ching Eng Png |
IJCNN | 3 |
| 2024 | Design and Analysis of Variable Airgap Permanent Magnet Motor for Electric DrivetrainabstractThis paper presents performance analysis of a variable flux permanent magnet motor. The proposed machine can flexibly control the air-gap flux by adjusting the separation distance between the rotor and stator, allowing for variation in electromagnetic coupling based on operating conditions. Simulation results are provided to verify the effectiveness of the self-adjusted variable air gap motor, which maximizes power transfer by controlling the electromagnetic coupling. H. N. Phyu, Zaifeng Yang, Jonathan Hey |
TENCON | 2 |
| 2024 | Knowledge-Based Extraction for Inverse Design of Optical Thin-FilmsabstractIn this paper, we've developed an inverse design framework for optical thin-films assisted by Generative Pre-trained Transformers (GPT). As a Large Language Models (LLMs)-powered tool, this creates an integrated platform for users to perform various functions related to inverse design on thin-films using an accessible and unique natural language interface via the assistant tool. With the consultant tool, users can query important design information and receive relevant answers to enhance the design knowledge of optical thin-films. This platform proves capable of speeding up design of thin-film structures compared to traditional ways and can also save many hours for repeated knowledge querying. The developed platform significantly reduces the resources and time (~3 x) for designers in developing optical thin-films to reach an optimal design. This will allow the thin-films' applications to be brought to the market faster and with increased efficiency. Ze Dong Saw, Ethan Jun Qi Wang, Zaifeng Yang, Viet Phuong Bui, Ching Eng Png |
TENCON | 3 |
| 2024 | Generative Pixelated Microstrip Millimeter-Wave Circuit Components DesignabstractIn this paper, we present a novel method for the generative design of pixelated microstrip Millimeter-Wave (mm-wave) circuit components. Our approach utilizes a primary meandering transmission line, which is connected to the input and output/terminated ports, with random small pixelated patches surrounding the primary transmission line as perturbations. A 2D Convolutional Neural Network (CNN) incorporating a Convolutional Block Attention Module (CBAM) is trained and employed as a forward solver to predict the input impedance of the circuit component. Additionally, a Variational Autoencoder (VAE) is trained to generate new pixelated microstrip mm-wave circuit component patterns by sampling from the latent space, thereby extending beyond the original training dataset. By leveraging the CNN-based forward solver for impedance prediction and the VAE for pattern generation, we can inversely generate the mm-wave circuit component design that meets specific input impedance requirements. Zaifeng Yang, Chuan Ge, Kevin Tshun Chuan Chai, Viet Phuong Bui, Ching Eng Png |
TENCON | 1 |
| 2023 | Sodium Niobate Slitted Ultrasonic Transducer Model with Simulated 172× Figure-of-Merit EnhancementabstractPiezoelectric ultrasound transducers have been commercialized by industry. However, the trade-off between acoustic pressure output and the sensitivity of resonant frequencies to the non-uniformity of residual stresses in the piezoelectric layer (known as stress sensitivity) has not been fully addressed. We present a novel combination of our patented slitted membrane design with our Sodium Niobate piezoelectric material. Our combination not only reduces the stress sensitivity by 4.9×, but also increases the acoustic pressure output by 35×. This brings our figure-of-merit to ($4.9\times 35$) 172× improvement over fully clamped Aluminum Nitride membranes. Xing Haw Marvin Tan, Khuong Phuong Ong, Zaifeng Yang, Viet Phuong Bui, Ching Eng Png, Hong Son Chu, Huajun Liu |
IECON | 3 |
| 2023 | Cooperative Colorization: Exploring Latent Cross-Domain Priors for NIR Image Spectrum TranslationabstractNear-infrared (NIR) image spectrum translation is a challenging problem with many promising applications. Existing methods struggle with the mapping ambiguity between the NIR and the RGB domains, and generalize poorly due to the limitations of models' learning capabilities and the unavailability of sufficient NIR-RGB image pairs for training. To address these challenges, we propose a cooperative learning paradigm that colorizes NIR images in parallel with another proxy grayscale colorization task by exploring latent cross-domain priors (i.e., latent spectrum context priors and task domain priors), dubbed CoColor. The complementary statistical and semantic spectrum information from these two task domains -- in the forms of pre-trained colorization networks -- are brought in as task domain priors. A bilateral domain translation module is subsequently designed, in which intermittent NIR images are generated from grayscale and colorized in parallel with authentic NIR images; and vice versa for the grayscale images. These intermittent transformations act as latent spectrum context priors for efficient domain knowledge exchange. We progressively fine-tune and fuse these modules with a series of pixel-level and feature-level consistency constraints. Experiments show that our proposed cooperative learning framework produces satisfactory spectrum translation outputs with diverse colors and rich textures, and outperforms state-of-the-art counterparts by 3.95dB and 4.66dB in terms of PNSR for the NIR and grayscale colorization tasks, respectively. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
ACM Multimedia | 3 |
| 2023 | Multi-scale Progressive Feature Embedding for Accurate NIR-to-RGB Spectral Domain TranslationabstractNIR-to-RGB spectral domain translation is a challenging task due to the mapping ambiguities and existing methods show limited learning capacities. To address these challenges, we propose to colorize NIR images via a multi-scale progressive feature embedding network (MPFNet), with the guidance of grayscale image colorization. Specifically, we first introduce a domain translation module that translates NIR source images into the grayscale target domain. By incorporating a progressive training strategy, the statistical and semantic knowledge from both task domains are efficiently aligned with a series of pixel-/feature-level consistency constraints. Besides, a multi-scale progressive feature embedding network is designed to improve learning capabilities. Experiments show that our MPFNet outperforms state-of-the-art counterparts by 2.55dB in the NIR-to-RGB spectral domain translation task in terms of PSNR. Xingxing Yang 0002, Jie Chen 0026, Zaifeng Yang |
VCIP | 3 |
| 2022 | Attention-Guided Progressive Neural Texture Fusion for High Dynamic Range Image RestorationabstractHigh Dynamic Range (HDR) imaging via multi-exposure fusion is an important task for most modern imaging platforms. In spite of recent developments in both hardware and algorithm innovations, challenges remain over content association ambiguities caused by saturation, motion, and various artifacts introduced during multi-exposure fusion such as ghosting, noise, and blur. In this work, we propose an Attention-guided Progressive Neural Texture Fusion (APNT-Fusion) HDR restoration model which aims to address these issues within one framework. An efficient two-stream structure is proposed which separately focuses on texture feature transfer over saturated regions and multi-exposure tonal and texture feature fusion. A neural feature transfer mechanism is proposed which establishes spatial correspondence between different exposures based on multi-scale VGG features in the masked saturated HDR domain for discriminative contextual clues over the ambiguous image areas. A progressive texture blending module is designed to blend the encoded two-stream features in a multi-scale and progressive manner. In addition, we introduce several novel attention mechanisms, i.e., the motion attention module detects and suppresses the content discrepancies among the reference images; the saturation attention module facilitates differentiating the misalignment caused by saturation from those caused by motion; and the scale attention module ensures texture blending consistency between different coder/decoder scales. We carry out comprehensive qualitative and quantitative evaluations and ablation studies, which validate that these novel modules work coherently under the same framework and outperform state-of-the-art methods. Jie Chen 0026, Zaifeng Yang, Tsz Nam Chan, Hui Li 0029, Junhui Hou, Lap-Pui Chau |
IEEE Trans. Image Process. | 2 |
| 2022 | Scale-Consistent Fusion: From Heterogeneous Local Sampling to Global Immersive RenderingabstractImage-based geometric modeling and novel view synthesis based on sparse large-baseline samplings are challenging but important tasks for emerging multimedia applications such as virtual reality and immersive telepresence. Existing methods fail to produce satisfactory results due to the limitation on inferring reliable depth information over such challenging reference conditions. With the popularization of commercial light field (LF) cameras, capturing LF images (LFIs) is as convenient as taking regular photos, and geometry information can be reliably inferred. This inspires us to use a sparse set of LF captures to render high-quality novel views globally. However, the fusion of LF captures from multiple angles is challenging due to the scale inconsistency caused by various capture settings. To overcome this challenge, we propose a novel scale-consistent volume rescaling algorithm that robustly aligns the disparity probability volumes (DPV) among different captures for scale-consistent global geometry fusion. Based on the fused DPV projected to the target camera frustum, novel learning-based modules (i.e., the attention-guided multi-scale residual fusion module, and the disparity field-guided deep re-regularization module), which comprehensively regularize noisy observations from heterogeneous captures for high-quality rendering of novel LFIs, have been proposed. Both quantitative and qualitative experiments over the Stanford Lytro Multi-view LF dataset show that the proposed method outperforms state-of-the-art methods significantly under different experiment settings for disparity inference and LF synthesis. Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Qiang Wang 0022, Yike Guo |
IEEE Trans. Image Process. | 3 |
| 2021 | A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced DataabstractIn this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the first learning stage, pre-processed volumetric CT data together with the segmented lung masks are fed into a 3D ResNet module, and an initial classification result can be obtained. However, due to categorical data imbalance, we observe large differences in sensitivity between COVID-19 and CAP cases. In the second stage, five learning models are independently trained over data with only COVID-19 and CAP cases, and are then ensembled to further discriminate the two classes. The final classification results are obtained by combining the predictions from both stages. Based on the validation dataset, we have evaluated our method and compared it with up-to-date methods in terms of overall accuracy and sensitivity for each class. The validation results validate the accuracy of the proposed multi-stage learning strategy. The overall accuracy of the validation dataset is 88.8%, and the sensitivities are 0.873, 0.789 and 1 for COVID-19, CAP and normal cases, respectively. Zaifeng Yang, Yubo Hou, Zhenghua Chen, Le Zhang 0001, Jie Chen 0026 |
ICASSP | 1 |
| 2020 | Learning From Paired and Unpaired Data: Alternately Trained CycleGAN for Near Infrared Image ColorizationabstractThis paper presents a novel near infrared (NIR) image colorization approach for the Grand Challenge held by 2020 IEEE International Conference on Visual Communications and Image Processing (VCIP). A Cycle-Consistent Generative Adversarial Network (CycleGAN) with cross-scale dense connections is developed to learn the color translation from the NIR domain to the RGB domain based on both paired and unpaired data. Due to the limited number of paired NIR-RGB images, data augmentation via cropping, scaling, contrast and mirroring operations have been adopted to increase the variations of the NIR domain. An alternating training strategy has been designed, such that CycleGAN can efficiently and alternately learn the explicit pixel-level mappings from the paired NIR-RGB data, as well as the implicit domain mappings from the unpaired ones. Based on the validation data, we have evaluated our method and compared it with conventional CycleGAN method in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and angular error (AE). The experimental results validate the proposed colorization framework. Zaifeng Yang, Zhenghua Chen |
VCIP | 1 |