Zixiang Zhao

dblp:65/5420 · DBLP profile ↗
← Back
41ranked-venue papers
15as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 11 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Nussbaum-Based Adaptive Neural Network Event-Triggered Control for Hyperbolic PDE Systems With Uncertain Actuator Dynamics and Deception Attacks
abstract
This paper explores the security control problem of a class of hyperbolic partial differential equation (PDE) systems described by a set of nonlinear ordinary differential equation (ODE) under deception attacks. Compared with the control design process of traditional single ODE systems, the strong coupling characteristics of partial differential equation-ordinary differential equation (PDE-ODE) systems make the control design under deception attacks more difficult. By applying the infinite-dimensional backstepping transformation and its inverse transformation, the original PDE subsystem is transformed into a more tractable target system, thus effectively achieving control performance. The neural network approximation algorithm is adopted to separate the coupling effects caused by attacks, focusing on addressing the unknown nonlinear dynamics of the system. Deception attacks are modeled as time-varying weights with unknown control directions, and the Nussbaum function is employed to address the problem of unknown control directions. In addition, a dynamic event-triggered mechanism with a novel switching threshold is designed to alleviate the communication burden. The proposed control algorithm ensures that both the closed-loop system states and the actuator states are bounded. Finally, this conclusion is verified through numerical simulation.
Liang Zhang 0039, Zixiang Zhao, Ning Zhao 0002, Ben Niu 0003
IEEE Internet Things J.2
2025 Task-driven Image Fusion with Learnable Fusion Loss
abstract
Multi-modal image fusion aggregates information from multiple sensor sources, achieving superior visual quality and perceptual features compared to single-source images, often improving downstream tasks. However, current fusion methods for downstream tasks still use predefined fusion objectives that potentially mismatch the downstream tasks, limiting adaptive guidance and reducing model flexibility. To address this, we propose Task-driven Image Fusion (TDFusion), a fusion framework incorporating a learnable fusion loss guided by task loss. Specifically, our fusion loss includes learnable parameters modeled by a neural network called the loss generation module. This module is supervised by the downstream task loss in a meta-learning manner. The learning objective is to minimize the task loss of fused images after optimizing the fusion module with the fusion loss. Iterative updates between the fusion module and the loss module ensure that the fusion network evolves toward minimizing task loss, guiding the fusion process toward the task objectives. TDFusion’s training relies entirely on the downstream task loss, making it adaptable to any specific task. It can be applied to any architecture of fusion and task networks. Experiments demonstrate TDFusion’s performance through fusion experiments conducted on four different datasets, in addition to evaluations on semantic segmentation and object detection tasks. The code is available at https://github.com/HaowenBai/TDFusion.
Haowen Bai, Jiangshe Zhang 0001, Zixiang Zhao, Lilun Deng, Yukun Cui
CVPR3
2025 Retinex-MEF: Retinex-Based Glare Effects Aware Unsupervised Multi-Exposure Image Fusion
Haowen Bai, Jiangshe Zhang 0001, Zixiang Zhao, Lilun Deng, Yukun Cui
ICCV3
2025 Hipandas: Hyperspectral Image Joint Denoising and Super-Resolution by Image Fusion with the Panchromatic Image
abstract
Hyperspectral images (HSIs) are frequently noisy and of low resolution due to the constraints of imaging devices. Recently launched satellites can concurrently acquire HSIs and panchromatic (PAN) images, enabling the restoration of HSIs to generate clean and high-resolution imagery through fusing PAN images for denoising and super-resolution. However, previous studies treated these two tasks as independent processes, resulting in accumulated errors. This paper introduces \textbf{H}yperspectral \textbf{I}mage Joint \textbf{Pand}enoising \textbf{a}nd Pan\textbf{s}harpening (Hipandas), a novel learning paradigm that reconstructs HRHS images from noisy low-resolution HSIs (LRHS) and high-resolution PAN images. The proposed zero-shot Hipandas framework consists of a guided denoising network, a guided super-resolution network, and a PAN reconstruction network, utilizing an HSI low-rank prior and a newly introduced detail-oriented low-rank prior. The interconnection of these networks complicates the training process, necessitating a two-stage training strategy to ensure effective training. Experimental results on both simulated and real-world datasets indicate that the proposed method surpasses state-of-the-art algorithms, yielding more accurate and visually pleasing HRHS images.
Zixiang Zhao, Haowen Bai, Jiangjun Peng, Xiangyong Cao, Deyu Meng
ICCV2
2025 BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
abstract
With the advancement of diffusion models (DMs) and the substantially increased computational requirements, quantization emerges as a practical solution to obtain compact and efficient low-bit DMs. However, the highly discrete representation leads to severe accuracy degradation, hindering the quantization of diffusion models to ultra-low bit-widths. This paper proposes a novel weight binarization approach for DMs, namely BinaryDM, pushing binarized DMs to be accurate and efficient by improving the representation and optimization. From the representation perspective, we present an Evolvable-Basis Binarizer (EBB) to enable a smooth evolution of DMs from full-precision to accurately binarized. EBB enhances information representation in the initial stage through the flexible combination of multiple binary bases and applies regularization to evolve into efficient single-basis binarization. The evolution only occurs in the head and tail of the DM architecture to retain the stability of training. From the optimization perspective, a Low-rank Representation Mimicking (LRM) is applied to assist the optimization of binarized DMs. The LRM mimics the representations of full-precision DMs in low-rank space, alleviating the direction ambiguity of the optimization process caused by fine-grained alignment. Comprehensive experiments demonstrate that BinaryDM achieves significant accuracy and efficiency gains compared to SOTA quantization methods of DMs under ultra-low bit-widths. With 1-bit weight and 4-bit activation (W1A4), BinaryDM achieves as low as 7.74 FID and saves the performance from collapse (baseline FID 10.87). As the first binarization method for diffusion models, W1A4 BinaryDM achieves impressive 15.2x OPs and 29.2x model size savings, showcasing its substantial potential for edge deployment.
Xingyu Zheng, Xianglong Liu 0001, Haotong Qin, Xudong Ma, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo 0002, Michele Magno
ICLR8
2025 Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
abstract
Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters. Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present **Q-VDiT**, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the *Token aware Quantization Estimator* (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce *Temporal Maintenance Distillation* (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency score of 23.40, setting a new benchmark and outperforming the current state-of-the-art quantization methods by **1.9$\times$**.
Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Zhulin An, Libo Huang 0001, Boyu Diao, Zixiang Zhao, Yongjun Xu 0001, Michele Magno
ICML9
2025 LLplace: Embodied 3D Indoor Layout Synthesis Framework with Large Language Model
abstract
Designing 3D indoor layouts is a crucial task with significant applications in embodied robot intelligence, virtual reality, and interior design. Existing methods for 3D layout design either rely on diffusion models, which utilize spatial relationship priors, or heavily leverage the inferential capabilities of proprietary Large Language Models (LLMs) , which require extensive prompt engineering and in-context exemplars via black-box trials. These methods often face limitations in generalization and dynamic scene editing. In this paper, we introduce LLplace, a novel 3D indoor scene layout designer based on lightweight, fine-tuned, open-source LLM Llama3. LLplace circumvents the need for spatial relationship priors and in-context exemplars, enabling efficient and credible room layout generation based solely on user inputs specifying the room type and desired objects. We curated a new dialogue dataset based on the 3D-Front dataset, expanding the original data volume and incorporating dialogue data for adding and removing objects. This dataset can enhance the LLM’s spatial understanding. Furthermore, through dialogue, LLplace activates the LLM’s capability to understand 3D layouts and perform dynamic scene editing, enabling the addition and removal of objects. Our approach demonstrates that LLplace can effectively generate and edit 3D indoor layouts interactively and outperform existing methods in delivering high-quality 3D design solutions.
Junru Lu, Zixiang Zhao, Wanxi Dong, Victor Sanchez, Feng Zheng 0001
IROS3
2025 A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking
abstract
The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion (UniVF), a novel and unified framework for video fusion that leverages multi-frame learning and optical flow-based feature warping for informative, temporally coherent video fusion. To support its development, we also introduce Video Fusion Benchmark (VF-Bench), the first comprehensive benchmark covering four video fusion tasks: multi-exposure, multi-focus, infrared-visible, and medical fusion. VF-Bench provides high-quality, well-aligned video pairs obtained through synthetic data generation and rigorous curation from existing datasets, with a unified evaluation protocol that jointly assesses the spatial quality and temporal consistency of video fusion. Extensive experiments show that UniVF achieves state-of-the-art results across all tasks on VF-Bench. Project page: [vfbench.github.io](https://vfbench.github.io).
Zixiang Zhao, Haowen Bai, Bingxin Ke, Yukun Cui, Lilun Deng, Yulun Zhang 0001, Kai Zhang 0008, Konrad Schindler
NeurIPS1
2025 ReFusion: Learning Image Fusion from Reconstruction with Learnable Loss Via Meta-Learning
Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Lilun Deng, Yukun Cui, Baisong Jiang
Int. J. Comput. Vis.2
2025 Parameterized Low-Rank Regularizer for High-dimensional Visual Data
Zixiang Zhao, Xiangyong Cao, Jiangjun Peng, Xi-Le Zhao, Deyu Meng, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
Int. J. Comput. Vis.2
2025 Deep Unfolding Multi-Modal Image Fusion Network via Attribution Analysis
abstract
Multi-modal image fusion synthesizes information from multiple sources into a single image, facilitating downstream tasks such as semantic segmentation. Current approaches primarily focus on acquiring informative fusion images at the visual display stratum through intricate mappings. Although some approaches attempt to jointly optimize image fusion and downstream tasks, these efforts often lack direct guidance or interaction, serving only to assist with a predefined fusion loss. To address this, we propose an “Unfolding Attribution Analysis Fusion network” (UAAFusion), using attribution analysis to tailor fused images more effectively for semantic segmentation, enhancing the interaction between the fusion and segmentation. Specifically, we utilize attribution analysis techniques to explore the contributions of semantic regions in the source images to task discrimination. At the same time, our fusion algorithm incorporates more beneficial features from the source images, thereby allowing the segmentation to guide the fusion process. Our method constructs a model-driven unfolding network that uses optimization objectives derived from attribution analysis, with an attribution fusion loss calculated from the current state of the segmentation network. We also develop a new pathway function for attribution analysis, specifically tailored to the fusion tasks in our unfolding network. An attribution attention mechanism is integrated at each network stage, allowing the fusion network to prioritize areas and pixels crucial for high-level recognition tasks. Additionally, to mitigate the information loss in traditional unfolding networks, a memory augmentation module is incorporated into our network to improve the information flow across various network layers. Extensive experiments demonstrate our method’s superiority in image fusion and applicability to semantic segmentation. The code is available athttps://github.com/HaowenBai/UAAFusion.
Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Baisong Jiang, Lilun Deng, Yukun Cui, Chunxia Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 An Automotive Onboard Self-Heating Method Based on Reconfigurable Battery System in Cold Climates
abstract
Lithium-ion batteries in cold climates suffer from significant performance degradation, such as reduced available power and life cycle deterioration. To address this problem, a reconfigurable battery system (RBS) based self-heating method is proposed in this article. This innovative approach leverages a switch array integrated within the RBS, achieving self-heating with a fast temperature rise and low energy loss. Additionally, employing ac heating current with high frequency further mitigates damage to the battery. The modularized three-switch reconfigurable topology is proposed, and the ac-heating principle is designed. Also, based on the frequency-dependent characteristics of the battery impedance, the heating strategy is developed to accurately control the heating current. The experimental results demonstrate the efficient heating capability: the battery can be heated from −30 °C to 0 °C within 237 s by consuming only 6.27% of nominal capacity, and the battery capacity fade rate is only 0.47% after 200 heating cycles.
Zixiang Zhao, Jun Xu 0018, Zhaohuan Liu, Zhongyue Zou, Xuesong Mei
IEEE Trans. Ind. Informatics1
2025 RefComp: A Reference-Guided Unified Framework for Unpaired Point Cloud Completion
Zixiang Zhao, Victor Sanchez, Feng Zheng 0001
IEEE Trans. Multim.3
2025 CACNN: Capsule Attention Convolutional Neural Networks for 3D Object Recognition
abstract
Recently, view-based approaches, which recognize a 3D object through its projected 2-D images, have been extensively studied and have achieved considerable success in 3D object recognition. Nevertheless, most of them use a pooling operation to aggregate viewwise features, which usually leads to the visual information loss. To tackle this problem, we propose a novel layer called capsule attention layer (CAL) by using attention mechanism to fuse the features expressed by capsules. In detail, instead of dynamic routing algorithm, we use an attention module to transmit information from the lower level capsules to higher level capsules, which obviously improves the speed of capsule networks. In particular, the view pooling layer of multiview convolutional neural network (MVCNN) becomes a special case of our CAL when the trainable weights are chosen on some certain values. Furthermore, based on CAL, we propose a capsule attention convolutional neural network (CACNN) for 3D object recognition. Extensive experimental results on three benchmark datasets demonstrate the efficiency of our CACNN and show that it outperforms many state-of-the-art methods.
Kai Sun 0007, Jiangshe Zhang 0001, Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Junying Hu
IEEE Trans. Neural Networks Learn. Syst.4
2024 Equivariant Multi-Modality Image Fusion
abstract
Multi-modality image fusion is a technique that combines information from different sensors or modalities, en-abling the fused image to retain complementary features from each modality, such as functional highlights and texture details. However, effective training of such fusion models is challenging due to the scarcity of ground truth fusion data. To tackle this issue, we propose the Equivariant Multi-Modality imAge fusion (EMMA) paradigm for end-to-end self-supervised learning. Our approach is rooted in the prior knowledge that natural imaging responses are equiv-ariant to certain transformations. Consequently, we introduce a novel training paradigm that encompasses a fusion module, a pseudo-sensing module, and an equivariant fusion module. These components enable the net training to follow the principles of the natural sensing-imaging process while satisfying the equivariant imaging prior. Extensive experiments confirm that EMMA yields high-quality fusion results for infraredvisible and medical images, concurrently facilitating downstream multi-modal segmentation and detection tasks. The code is available at https://github.com/Zhaozixiang1228/MMIF-EMMA.
Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Radu Timofte, Luc Van Gool
CVPR1
2024 Flexible Residual Binarization for Image Super-Resolution
abstract
Binarized image super-resolution (SR) has attracted much research attention due to its potential to drastically reduce parameters and operations. However, most binary SR works binarize network weights directly, which hinders high-frequency information extraction. Furthermore, as a pixel-wise reconstruction task, binarization often results in heavy representation content distortion. To address these issues, we propose a flexible residual binarization (FRB) method for image SR. We first propose a second-order residual binarization (SRB), to counter the information loss caused by binarization. In addition to the primary weight binarization, we also binarize the reconstruction error, which is added as a residual term in the prediction. Furthermore, to narrow the representation content gap between the binarized and full-precision networks, we propose Distillation-guided Binarization Training (DBT). We uniformly align the contents of different bit widths by constructing a normalized attention form. Finally, we generalize our method by applying our FRB to binarize convolution and Transformer-based SR networks, resulting in two binary baselines: FRBC and FRBT. We conduct extensive experiments and comparisons with recent leading binarization methods. Our proposed baselines, FRBC and FRBT, achieve superior performance both quantitatively and visually. The code and model will be released.
Yulun Zhang 0001, Haotong Qin, Zixiang Zhao, Xianglong Liu 0001, Martin Danelljan, Fisher Yu 0001
ICML3
2024 Image Fusion via Vision-Language Model
abstract
Image fusion integrates essential information from multiple images into a single composite, enhancing structures, textures, and refining imperfections. Existing methods predominantly focus on pixel-level and semantic visual features for recognition, but often overlook the deeper text-level semantic information beyond vision. Therefore, we introduce a novel fusion paradigm named image Fusion via vIsion-Language Model (FILM), for the first time, utilizing explicit textual information from source images to guide the fusion process. Specifically, FILM generates semantic prompts from images and inputs them into ChatGPT for comprehensive textual descriptions. These descriptions are fused within the textual domain and guide the visual information fusion, enhancing feature extraction and contextual understanding, directed by textual semantic information via cross-attention. FILM has shown promising results in four image fusion tasks: infrared-visible, medical, multi-exposure, and multi-focus image fusion. We also propose a vision-language dataset containing ChatGPT-generated paragraph descriptions for the eight image fusion datasets across four fusion tasks, facilitating future research in vision-language model-based image fusion. Code and dataset are available at https://github.com/Zhaozixiang1228/IF-FILM.
Zixiang Zhao, Lilun Deng, Haowen Bai, Yukun Cui, Yulun Zhang 0001, Haotong Qin, Jiangshe Zhang 0001, Luc Van Gool
ICML1
2024 Make Continual Learning Stronger via C-Flat
abstract
How to balance the learning ’sensitivity-stability’ upon new task training and memory preserving is critical in CL to resolve catastrophic forgetting. Improving model generalization ability within each learning phase is one solution to help CL learning overcome the gap in the joint knowledge space. Zeroth-order loss landscape sharpness-aware minimization is a strong training regime improving model generalization in transfer learning compared with optimizer like SGD. It has also been introduced into CL to improve memory representation or learning efficiency. However, zeroth-order sharpness alone could favors sharper over flatter minima in certain scenarios, leading to a rather sensitive minima rather than a global optima. To further enhance learning stability, we propose a Continual Flatness (C-Flat) method featuring a flatter loss landscape tailored for CL. C-Flat could be easily called with only one line of code and is plug-and-play to any CL methods. A general framework of C-Flat applied to all CL categories and a thorough comparison with loss minima optimizer and flat minima based CL approaches is presented in this paper, showing that our method can boost CL performance in almost all cases. Code is available at https://github.com/WanNaa/C-Flat.
Ang Bian, Wei Li 0313, Hangjie Yuan, Chengrong Yu, Mang Wang 0001, Zixiang Zhao, Aojun Lu, Pengliang Ji, Tao Feng 0014
NeurIPS6
2024 Prescribed performance adaptive neural event-triggered control for switched nonlinear cyber-physical systems under deception attacks
Liang Zhang 0039, Zixiang Zhao, Ning Zhao 0002
Neural Networks2
2024 Simultaneous Automatic Picking and Manual Picking Refinement for First-Break
abstract
First-break picking is a pivotal procedure in processing microseismic data for geophysics and resource exploration. Recent advancements in deep learning have catalyzed the evolution of automated methods for identifying first-break. Nevertheless, the complexity of seismic data acquisition and the requirement for detailed, expert-driven labeling often result in outliers and potential mislabeling within manually labeled datasets. These issues can negatively affect the training of neural networks, necessitating algorithms that handle outliers or mislabeled data effectively. We introduce the Simultaneous Picking and Refinement (SPR) algorithm, designed to handle datasets plagued by outlier samples or even noisy labels. Unlike conventional approaches that regard manual picks as ground truth, our method treats the true first-break as a latent variable within a probabilistic model that includes a first-break labeling prior. SPR aims to uncover this variable, enabling dynamic adjustments and improved accuracy across the dataset. This strategy mitigates the impact of outliers or inaccuracies in manual labels. Intra-site picking experiments and cross-site generalization experiments on publicly available data confirm our method’s performance in identifying first-break and its generalization across different sites. Additionally, our investigations into noisy signals and labels underscore SPR’s resilience to both types of noise and its capability to refine misaligned manual annotations. Moreover, the flexibility of SPR, not being limited to any single network architecture, enhances its adaptability across various deep learning-based picking methods. Focusing on learning from data that may contain outliers or partial inaccuracies, SPR provides a robust solution to some of the principal obstacles in automatic first-break picking.
Haowen Bai, Zixiang Zhao, Jiangshe Zhang 0001, Yukun Cui, Chunxia Zhang 0002, Zhenbo Guo
IEEE Trans. Geosci. Remote. Sens.2
2024 Pan-Denoising: Guided Hyperspectral Image Denoising via Weighted Represent Coefficient Total Variation
abstract
This article introduces a novel paradigm for hyperspectral image (HSI) denoising, which is termed pan-denoising. In a given scene, panchromatic (PAN) images capture similar structures and textures to HSIs but with less noise. This enables the utilization of PAN images to guide the HSI denoising process. Consequently, pan-denoising, which incorporates an additional prior, has the potential to uncover underlying structures and details beyond the internal information modeling of traditional HSI denoising methods. However, the proper modeling of this additional prior poses a significant challenge. To alleviate this issue, the article proposes a novel regularization term, panchromatic weighted representation coefficient total variation (PWRCTV). It employs the gradient maps of PAN images to automatically assign different weights of total variation (TV) regularization for each pixel, resulting in larger weights for smooth areas and smaller weights for edges. This regularization forms the basis of a pan-denoising model, which is solved using the alternating direction method of multipliers (ADMM). Extensive experiments on synthetic and real-world datasets demonstrate that PWRCTV outperforms several state-of-the-art methods in terms of metrics and visual quality. Furthermore, an HSI classification experiment confirms that PWRCTV, as a preprocessing method, can enhance the performance of downstream classification tasks. The code and data are available athttps://github.com/shuangxu96/PWRCTV.
Qiao Ke, Jiangjun Peng, Xiangyong Cao, Zixiang Zhao
IEEE Trans. Geosci. Remote. Sens.5
2024 Reconfigurable Battery System-Based Hybrid Self-Heating Method for Low Temperature Applications
abstract
Battery performance is significantly reduced at low temperatures, posing a challenge. To overcome this issue, the reconfigurable battery system (RBS) based hybrid self-heating (HSH) method is proposed in this article. This innovative approach leverages the flexible mode-switching characteristics of the RBS, achieving HSH with a high temperature rise rate and minimal energy loss. Additionally, employing square ac heating current with high frequency and low amplitude further mitigates damage to the battery. The physical configuration of the RBS- based HSH method is designed, and the modularized three-switch reconfigurable topology is proposed. Furthermore, the heating strategy is developed to further reduce battery fading. The experimental results demonstrate the efficient heating capability of this approach: the battery can be rapidly heated from –20 °C to 10 °C in just 239 s, consuming only 6.29% of the nominal capacity.
Zixiang Zhao, Jun Xu 0018, Zhaohuan Liu, Xianggong Zhang, Xuesong Mei
IEEE Trans. Ind. Informatics1
2023 CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion
abstract
Multi-modality (MM) image fusion aims to render fused images that maintain the merits of different modalities, e.g., functional highlight and detailed textures. To tackle the challenge in modeling cross-modality features and decomposing desirable modality-specific and modality-shared features, we propose a novel Correlation-Driven feature Decomposition Fusion (CDDFuse) network. Firstly, CDDFuse uses Restormer blocks to extract cross-modality shallow features. We then introduce a dual-branch Transformer-CNN feature extractor with Lite Transformer (LT) blocks leveraging long-range attention to handle low-frequency global features and Invertible Neural Networks (INN) blocks focusing on extracting high-frequency local information. A correlation-driven loss is further proposed to make the low-frequency features correlated while the high-frequency features uncorrelated based on the embedded information. Then, the LT-based global fusion and INN-based local fusion layers output the fused image. Extensive experiments demonstrate that our CDDFuse achieves promising results in multiple fusion tasks, including infrared-visible image fusion and medical image fusion. We also show that CDDFuse can boost the performance in downstream infrared-visible semantic segmentation and object detection in a unified benchmark. The code is available at https://github.om/haozixiang1228/MMIF-CDDFuse.
Zixiang Zhao, Haowen Bai, Jiangshe Zhang 0001, Yulun Zhang 0001, Zudi Lin, Radu Timofte, Luc Van Gool
CVPR1
2023 Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution
abstract
Guided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The critical step of this task is to effectively extract domain-shared and domain-private RGB/depth features. In addition, three detailed issues, namely blurry edges, noisy surfaces, and over-transferred RGB texture, need to be addressed. In this paper, we propose the Spherical Space feature Decomposition Network (SSDNet) to solve the above issues. To better model cross-modality features, Restormer block-based RGB/depth encoders are employed for extracting local-global features. Then, the extracted features are mapped to the spherical space to complete the separation of private features and the alignment of shared features. Shared features of RGB are fused with the depth features to complete the GDSR task. Subsequently, a spherical contrast refinement (SCR) module is proposed to further address the detail issues. Patches that are classified according to imperfect categories are input into the SCR module, where the patch features are pulled closer to the ground truth and pushed away from the corresponding imperfect samples in the spherical feature space via contrastive learning. Extensive experiments demonstrate that our method can achieve state-of-the-art results on four test datasets, as well as successfully generalize to real-world scenes. The code is available at https://github.com/Zhaozixiang1228/GDSR-SSDNet.
Zixiang Zhao, Jiangshe Zhang 0001, Chengli Tan, Yulun Zhang 0001, Radu Timofte, Luc Van Gool
ICCV1
2023 DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion
abstract
Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of interpretability for GAN-based generative methods, we propose a novel fusion algorithm based on the denoising diffusion probabilistic model (DDPM). The fusion task is formulated as a conditional generation problem under the DDPM sampling framework, which is further divided into an unconditional generation subproblem and a maximum likelihood subproblem. The latter is modeled in a hierarchical Bayesian manner with latent variables and inferred by the expectation-maximization (EM) algorithm. By integrating the inference solution into the diffusion sampling iteration, our method can generate high-quality fused images with natural image generative priors and cross-modality information from source images. Note that all we required is an unconditional pre-trained generative model, and no fine-tuning is needed. Our extensive experiments indicate that our approach yields promising fusion results in infrared-visible image fusion and medical image fusion. The code is available at https://github.com/Zhaozixiang1228/MMIF-DDFM.
Zixiang Zhao, Haowen Bai, Yuanzhi Zhu 0001, Jiangshe Zhang 0001, Yulun Zhang 0001, Kai Zhang 0008, Deyu Meng, Radu Timofte, Luc Van Gool
ICCV1
2023 AR assistance for efficient dynamic target search
abstract
When searching for a dynamic target in an unknown real world scene, search efficiency is greatly reduced if users lack information about the spatial structure of the scene. Most target search studies, especially in robotics, focus on determining either the shortest path when the target’s position is known, or a strategy to find the target as quickly as possible when the target’s position is unknown. However, the target’s position is often known intermittently in the real world, e.g., in the case of using surveillance cameras. Our goal is to help user find a dynamic target efficiently in the real world when the target’s position is intermittently known. In order to achieve this purpose, we have designed an AR guidance assistance system to provide optimal current directional guidance to users, based on searching a prediction graph. We assume that a certain number of depth cameras are fixed in a real scene to obtain dynamic target’s position. The system automatically analyzes all possible meetings between the user and the target, and generates optimal directional guidance to help the user catch up with the target. A user study was used to evaluate our method, and its results showed that compared to free search and a top-view method, our method significantly improves target search efficiency.
Zixiang Zhao, Jian Wu 0033, Lili Wang 0006
Comput. Vis. Media1
2022 Discrete Cosine Transform Network for Guided Depth Map Super-Resolution
abstract
Guided depth super-resolution (GDSR) is an essential topic in multi-modal image processing, which reconstructs high-resolution (HR) depth maps from low-resolution ones collected with suboptimal conditions with the help of HR RGB images of the same scene. To solve the challenges in interpreting the working mechanism, extracting cross-modal features and RGB texture over-transferred, we propose a novel Discrete Cosine Transform Network (DCTNet) to alleviate the problems from three aspects. First, the Discrete Cosine Transform (DCT) module reconstructs the multi-channel HR depth features by using DCT to solve the channel-wise optimization problem derived from the image domain. Second, we introduce a semi-coupled feature extraction module that uses shared convolutional kernels to extract common information and private kernels to extract modality-specific information. Third, we employ an edge attention mechanism to highlight the contours informative for guided upsampling. Extensive quantitative and qualitative evaluations demonstrate the effectiveness of our DCTNet, which outperforms previous state-of-the-art methods with a relatively small number of parameters. The code is available at https://github.com/Zhaozixiang1228/GDSR-DCTNet.
Zixiang Zhao, Jiangshe Zhang 0001, Zudi Lin, Hanspeter Pfister
CVPR1
2022 Optimization Algorithm Unfolding Deep Networks of Detail Injection Model for Pansharpening
abstract
Pansharpening aims at integrating a high-spatial-resolution panchromatic (PAN) image with a low-spatial-resolution multispectral (MS) image to generate a high-resolution MS (HRMS) image. It is a fundamental and significant task in the field of remotely sensed images. Classic and convolutional neural network (CNN)-based algorithms have been developed, over the last decades, for pansharpening based on the spatial detail injection model. However, these algorithms have difficulties in extracting sufficient details or lack interpretability. In this letter, we present an algorithm unfolding pansharpening (AUP) for this task. In the proposed AUP, a two-step optimization model is first designed based on the spatial detail decomposition model. Then, the iteration processes induced by an optimization model are mapped to several detailed convolution (dc) blocks to solve the detail injection by a trainable neural network. Finally, the desired MS details are obtained in end-to-end manners through a decoder. The superiority of the proposed AUP is demonstrated by extensive experiments on datasets acquired by two different kinds of satellites. Each module of the AUP is interpretable, and its fused results are with fewer spectral and spatial distortions.
Yunqiao Feng, Junmin Liu, Zixiang Zhao
IEEE Geosci. Remote. Sens. Lett.5
2022 Efficient and Model-Based Infrared and Visible Image Fusion via Algorithm Unrolling
abstract
Infrared and visible image fusion (IVIF) expects to obtain images that retain thermal radiation information from infrared images and texture details from visible images. In this paper, a model-based convolutional neural network (CNN) model, referred to as Algorithm Unrolling Image Fusion (AUIF), is proposed to overcome the shortcomings of traditional CNN-based IVIF models. The proposed AUIF model starts with the iterative formulas of two traditional optimization models, which are established to accomplish two-scale decomposition, i.e., separating low-frequency base information and high-frequency detail information from source images. Then the algorithm unrolling is implemented where each iteration is mapped to a CNN layer and each optimization model is transformed into a trainable neural network. Compared with the general network architectures, the proposed framework combines the model-based prior information and is designed more reasonably. After the unrolling operation, our model contains two decomposers (encoders) and an additional reconstructor (decoder). In the training phase, this network is trained to reconstruct the input image. While in the test phase, the base (or detail) decomposed feature maps of infrared/visible images are merged respectively by an extra fusion layer, and then the decoder outputs the fusion image. Qualitative and quantitative comparisons demonstrate the superiority of our model, which can robustly generate fusion images containing highlight targets and legible details, exceeding the state-of-the-art methods. Furthermore, our network has fewer weights and faster speed.
Zixiang Zhao, Jiangshe Zhang 0001, Chengyang Liang, Chunxia Zhang 0002, Junmin Liu
IEEE Trans. Circuits Syst. Video Technol.1
2022 Automatic Velocity Picking Using a Multi-Information Fusion Deep Semantic Segmentation Network
abstract
Velocity picking, a critical step in seismic data processing, has been studied for decades. Although manual picking can produce accurate normal moveout (NMO) velocities from the velocity spectra of prestack gathers, it is time-consuming and becomes infeasible with the emergence of a large amount of seismic data. Numerous automatic velocity picking methods have thus been developed. In recent years, deep learning (DL) methods have produced good results on the seismic data with medium and high signal-to-noise ratios (SNR). Unfortunately, it still lacks a picking method to automatically generate accurate velocities in situations of low SNR. In this paper, we propose a multi-information fusion network (MIFN) to estimate stacking velocity from the fusion information of velocity spectra and stack gather segments (SGS). In particular, we transform the velocity picking problem into a semantic segmentation problem based on the velocity spectrum images. Meanwhile, the information provided by SGS is used as a prior in the network to assist segmentation. The experimental results on two field datasets show that the picking results of MIFN are stable and accurate for the scenarios with medium and high SNR, and it also performs well in low SNR scenarios. Code is made publicly available at https://github.com/newbee-ML/MIFN-Velocity-Picking.
Jiangshe Zhang 0001, Zixiang Zhao, Chunxia Zhang 0002, Li Long, Weifeng Geng
IEEE Trans. Geosci. Remote. Sens.3
2022 Hybrid Loss-Guided Coarse-to-Fine Model for Seismic Data Consecutively Missing Trace Reconstruction
abstract
Seismic data are generally sampled irregularly and sparsely along spatial coordinates because economic costs and obstacles hinder the regular arrangement of geophones in the field. Thus, the sampled seismic data often contain missing traces which result in difficulties for later processing steps. To alleviate this issue, versatile interpolation methods have been developed to interpolate the missing traces. However, the existing models for recovering seismic data with consecutively missing traces in a large amplitude range tend to produce artifacts and blurred signal details. We propose in this paper a hybrid loss guided coarse-to-fine model which consists of a coarse network and a refinement network to allow different regions of seismic data to be recovered in different stages. The coarse network is designed to reconstruct the strong signals and the refinement network is implemented subsequently to recover the weak signals. In addition, the refinement network focuses its attention on the areas which are not well recovered by the coarse network via a weight-masked mechanism. By resorting to the hybrid loss function L1+SSIM+Relativistic Average Least-Square Generative Adversarial Network (RaLSGAN), our model enables more accurate and realistic signal details to be reconstructed. Experiments with synthetic and field data demonstrate that our model is superior to the existing mainstream approaches and the role of the key components is also investigated through ablation studies.
Xiao-Li Wei, Chunxia Zhang 0002, Zixiang Zhao, Xiong Deng, Jiangshe Zhang 0001, Sang-Woon Kim
IEEE Trans. Geosci. Remote. Sens.4
2021 Deep Gradient Projection Networks for Pan-sharpening
abstract
Pan-sharpening is an important technique for remote sensing imaging systems to obtain high resolution multi-spectral images. Recently, deep learning has become the most popular tool for pan-sharpening. This paper develops a model-based deep pan-sharpening approach. Specifically, two optimization problems regularized by the deep prior are formulated, and they are separately responsible for the generative models for panchromatic images and low resolution multispectral images. Then, the two problems are solved by a gradient projection algorithm, and the iterative steps are generalized into two network blocks. By alternatively stacking the two blocks, a novel network, called gradient projection based pan-sharpening neural network, is constructed. The experimental results on different kinds of satellite datasets demonstrate that the new network out-performs state-of-the-art methods both visually and quantitatively. The codes are available at https://github.com/xsxjtu/GPPNN.
Jiangshe Zhang 0001, Zixiang Zhao, Kai Sun 0007, Junmin Liu, Chunxia Zhang 0002
CVPR3
2021 FGF-GAN: A Lightweight Generative Adversarial Network for Pansharpening via Fast Guided Filter
abstract
Pansharpening is a widely used image enhancement technique for remote sensing. Its principle is to fuse the input high-resolution single-channel panchromatic (PAN) image and low-resolution multi-spectral image and to obtain a high-resolution multi-spectral (HRMS) image. The existing deep learning pansharpening method has two shortcomings. First, features of two input images need to be concatenated along the channel dimension to reconstruct the HRMS image, which makes the importance of PAN images not prominent, and also leads to high computational cost. Second, the implicit information of features is difficult to extract through the manually designed loss function. To this end, we propose a generative adversarial network via the fast guided filter (FGF) for pansharpening. In generator, traditional channel concatenation is replaced by FGF to better retain the spatial information while reducing the number of parameters. Meanwhile, the fusion objects can be highlighted by the spatial attention module. In addition, the latent information of features can be preserved effectively through adversarial training. Numerous experiments illustrate that our network generates high-quality HRMS images that can surpass existing methods, and with fewer parameters.
Zixiang Zhao, Jiangshe Zhang 0001, Kai Sun 0007, Junmin Liu, Chunxia Zhang 0002
ICME1
2021 Deep Convolutional Sparse Coding Network For Pansharpening With Guidance Of Side Information
abstract
Pansharpening is a fundamental issue in remote sensing field. This paper proposes a side information partially guided convolutional sparse coding (SCSC) model for pansharpening. The key idea is to split the low resolution multispectral image into a panchromatic image related feature map and a panchromatic image irrelated feature map, where the former one is regularized by the side information from panchromatic images. With the principle of algorithm unrolling techniques, the proposed model is generalized as a deep neural network, called as SCSC pansharpening neural network (SCSC-PNN). Compared with 13 classic and state-of-the-art methods on three satellites, the numerical experiments show that SCSC-PNN is superior to others. The codes are available at https://github.com/xsxjtu/SCSC-PNN.
Jiangshe Zhang 0001, Kai Sun 0007, Zixiang Zhao, Junmin Liu, Chunxia Zhang 0002
ICME4
2021 MFIF-GAN: A new generative adversarial network for multi-focus image fusion
Junmin Liu, Zixiang Zhao, Chunxia Zhang 0002, Jiangshe Zhang 0001
Signal Process. Image Commun.4
2021 Dynamic targets searching assistance based on virtual camera priority
abstract
When a user walks freely in an unknown virtual scene and searches for multiple dynamic targets, the lack of a comprehensive understanding of the environment may have a negative impact on the execution of virtual reality tasks. Previous studies can help users with auxiliary tools, such as top view maps or trails, and exploration guidance, for example, automatically generated paths according to the user location and important static spots in virtual scenes. However, in some virtual reality applications, when the scene has complex occlusions, and the user cannot obtain any real-time position information of the dynamic target, the above assistance cannot help the user complete the task more effectively. We design a virtual camera priority-based assistance to help the user search dynamic targets efficiently. Instead of forcing users to go to destinations, we provide an optimized instant path to guide them to places where they are more likely to find dynamic targets when they ask for help. We assume that a certain number of virtual cameras are fixed in virtual scenes to obtain extra depth maps, which capture the depth information of the scene and the locations of the dynamic targets. Our methodautomatically analyzes the priority of these virtual cameras, chooses the destination, and generates an instant path to assist the user in finding the dynamic targets. Our method is suitable for various virtual reality applications that do not require manual supervision or input. A user study is designed to evaluate the proposed method. The results indicate that compared with three conventional navigation methods, such as the top-view method, our method can help users find dynamic targets more efficiently. The advantages include reducing the task completion time, reducing the number of resets, increasing the average distance between resets, and reducing user task load. We presented a method for improving dynamic target searching efficiency in virtual scenes by virtual camera priority-based path guidance. Compared with three conventional navigation methods, such as the top-view method, this method can help users find dynamic targets more effectively.
Zixiang Zhao, Quanwei Zhou, Xiaoguang Han 0001, Lili Wang 0006
Virtual Real. Intell. Hardw.1
2020 DIDFuse: Deep Image Decomposition for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion, a hot topic in the field of image processing, aims at obtaining fused images keeping the advantages of source images. This paper proposes a novel auto-encoder (AE) based fusion network. The core idea is that the encoder decomposes an image into background and detail feature maps with low- and high-frequency information, respectively, and that the decoder recovers the original image. To this end, the loss function makes the background/detail feature maps of source images similar/dissimilar. In the test phase, background and detail feature maps are respectively merged via a fusion module, and the fused image is recovered by the decoder. Qualitative and quantitative results illustrate that our method can generate fusion images containing highlighted targets and abundant detail texture information with strong reproducibility and meanwhile surpass state-of-the-art (SOTA) approaches.
Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Jiangshe Zhang 0001
IJCAI1
2020 Bayesian fusion for infrared and visible images
Zixiang Zhao, Chunxia Zhang 0002, Junmin Liu, Jiangshe Zhang 0001
Signal Process.1
2019 On the using of Rényi's quadratic entropy for physical layer key generation
Furui Zhan, Zixiang Zhao, Nianmin Yao
Comput. Commun.2
2008 Modeling and Simulation of Self-similar Storage I/O
Zixiang Zhao, Yunhan Jiang
GPC3
2008 SVD based Kalman particle filter for robust visual tracking
abstract
Object tracking is one of the most important tasks in computer vision. The unscented particle filter algorithm has been extensively used to tackle this problem and achieved a great success, because it uses the UKF (unscented Kalman filter) to generate a sophisticated proposal distributions which incorporates the newest observations into the state transition distribution and thus overcomes the sample impoverishment problem suffered by the particle filter. However, UKF often encounters the ill-conditioned problem when solving the square root of the covariance matrix in practice. In this paper, we propose a novel Kalman particle filter based on SVD (singular value decomposition), and apply it for visual tracking. Experimental results demonstrate that, compared with the particle filter and the unscented particle filter, the proposed algorithm is more robust in tracking performance.
Xiaoqin Zhang 0002, Weiming Hu 0004, Zixiang Zhao, Yanguo Wang, Xi Li 0001, Qingdi Wei
ICPR3