Zhengjue Wang

dblp:203/3538 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-1846-495XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning transferable representations by topic guided graph adversarial network
Zhengjue Wang, Zhihui Xin, Chiyu Chen, Hao Zhang 0050, Yunsong Li 0001, Hongwei Liu 0001, Bo Chen 0001
Signal Process.1
2025 Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image information, leading to two main drawbacks: i) they often produce spurious visual-concept relations, hence decreasing model reliability; and ii) though CBMs could explain the importance of every concept to the final prediction, it is still challenging to tell which visual region produces the prediction. To solve these problems, this paper proposes a Disentangled Optimal Transport CBM (DOT-CBM) framework to explore fine-grained visual-concept relations between local image patches and concepts. Specifically, we model the concept prediction process as a transportation problem between the patches and concepts, thereby achieving explicit fine-grained feature alignment. We also incorporate orthogonal projection losses within the modality to enhance local feature disentanglement. To further address the shortcut issues caused by statistical biases in the data, we utilize the visual saliency map and concept label statistics as transportation priors. Thus, DOT-CBM can visualize inversion heatmaps, provide more reliable concept predictions, and produce more accurate class predictions. Comprehensive experiments demonstrate that our proposed DOT-CBM achieves SOTA performance on several tasks, including image classification, local part detection and out-of-distribution generalization. Codes are available in supplementary material.
Zequn Zeng, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001
CVPR6
2025 Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification
abstract
Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE.
Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma
CVPR6
2025 HyLiOSR: Staged Progressive Learning for Joint Open-Set Recognition of Hyperspectral and LiDAR Data
abstract
The joint classification of hyperspectral images (HSIs) and light detection and ranging (LiDAR) data have seen significant advancements in recent research. However, it would be more practical if we could simultaneously detect the unknown classes in a more realistic open-set scenario. In this article, we introduce a novel open-set recognition (OSR) method for HSI and LiDAR data, termed HyLiOSR, which devises a staged progressive learning strategy to effectively bridge the gap between closed-set and open-set feature distributions within an autoencoder framework. Specifically, for the first stage, the reconstruction-based network is dedicated to accurately modeling each known category by learning multiple Gaussian prototypes, which facilitates OSR by disentangling the distribution of known classes. In the second stage, we actively synthesize samples of unknown classes during the feature extraction phase and create a virtual unknown classifier, enabling the network to effectively differentiate between known and unknown class samples. This approach establishes a distinct separation between known and unknown classes in the latent feature space, thereby enhancing the capability of the frameworks to distinguish between them. Comprehensive experiments conducted on three benchmark datasets demonstrate that the proposed HyLiOSR outperforms existing state-of-the-art methods. The source code will be accessible athttps://github.com/B-Xi/TGRS_2025_HyLiOSR.
Bobo Xi, Mingshuo Cai, Jiaojiao Li 0001, Zhengjue Wang, Shou Feng, Yunsong Li 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 MeaCap: Memory-Augmented Zero-shot Image Captioning
abstract
Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pre-trained vision-language models like CLIP for image-text similarity evaluation and a pre-trained language model (LM) for caption generation. The main difference be-tween them is whether using a textual corpus to train the LM. Though achieving attractive performance w.r.t. some metrics, existing methods often exhibit some com-mon drawbacks. Training-free methods tend to produce hallucinations, while text-only-training often lose gener-alization capability. To move forward, in this paper, we propose a novel Memory-Augmented zero-shot image Captioning framework (MeaCap). Specifically, equipped with a textual memory, we introduce a retrieve-then-filter module to get key concepts that are highly related to the image. By deploying our proposed memory-augmented visual-related fusion score in a keywords-to-sentence LM, MeaCap can generate concept-centered captions that keep high consistency with the image with fewer hallucinations and more world-knowledge. The framework of Mea-Cap achieves the state-of-the-art performance on a se-ries of zero-shot IC settings. Our code is available at https://github.com/joeyzOz/MeaCap.
Zequn Zeng, Hao Zhang 0050, Chiyu Chen, Bo Chen 0001, Zhengjue Wang
CVPR6
2024 HICEScore: A Hierarchical Metric for Image Captioning Evaluation
abstract
Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE.
Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001
ACM Multimedia7
2024 Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2
abstract
Text generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations.
Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.3
2023 ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based Polishing
abstract
Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searches every word in the caption using the knowledge of large-scale pre-trained models. Though effective, its autoregressive generation and gradient-directed searching mechanism limit the diversity of captions and inference speed, respectively. Moreover, ZeroCap does not consider the controllability issue of zero-shot IC. To move forward, we propose a framework for Controllable Zero-shot IC, named ConZIC. The core of ConZIC is a novel sampling-based non-autoregressive language model named Gibbs-BERT, which can generate and continuously polish every word. Extensive quantitative and qualitative results demonstrate the superior performance of our proposed ConZIC for both zero-shot IC and controllable zero-shot IC. Especially, ConZIC achieves about$5\times$generation speed than ZeroCap, and about$1.5\times$diversity scores, with accurate generation given different control signals. Our code is available at https://github.com/joeyz0z/ConZIC.
Zequn Zeng, Hao Zhang 0050, Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Zhengjue Wang
CVPR6
2023 Recurrent Neural Networks for Snapshot Compressive Imaging
abstract
Conventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT.
Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multiscale Visual-Attribute Co-Attention for Zero-Shot Image Recognition
abstract
Zero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis.
Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Learning Hierarchical Document Graphs From Multilevel Sentence Relations
abstract
Organizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching.
Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.3
2022 Multimodal Weibull Variational Autoencoder for Jointly Modeling Image-Text Data
abstract
For multimodal representation learning, traditional black-box approaches often fall short of extracting interpretable multilayer hidden structures, which contribute to visualize the connections between different modalities at multiple semantic levels. To extract interpretable multimodal latent representations and visualize the hierarchial semantic relationships between different modalities, based on deep topic models, we develop a novel multimodal Poisson gamma belief network (mPGBN) that tightly couples the observations of different modalities via imposing sparse connections between their modality-specific hidden layers. To alleviate the time-consuming Gibbs sampler adopted by traditional topic models in the testing stage, we construct a Weibull-based variational inference network (encoder) to directly map the observations to their latent representations, and further combine it with the mPGBN (decoder), resulting in a novel multimodal Weibull variational autoencoder (MWVAE), which is fast in out-of-sample prediction and can handle large-scale multimodal datasets. Qualitative evaluations on bimodal data consisting of image-text pairs show that the developed MWVAE can successfully extract expressive multimodal latent representations for downstream tasks like missing modality imputation and multimodal retrieval. Further extensive quantitative results demonstrate that both MWVAE and its supervised extension sMWVAE achieve state-of-the-art performance on various multimodal benchmarks.
Chaojie Wang 0001, Bo Chen 0001, Sucheng Xiao, Zhengjue Wang, Hao Zhang 0050, Ning Han 0004, Mingyuan Zhou
IEEE Trans. Cybern.4
2022 Infinite Bayesian Max-Margin Discriminant Projection
abstract
In this article, considering the supervised dimensionality reduction, we first propose a model, called infinite Bayesian max-margin linear discriminant projection (iMMLDP), by assembling a set of local regions, where we make use of Bayesian nonparametric priors to handle the model selection problem, for example, the underlying number of local regions. In each local region, our model jointly learns a discriminative subspace and the corresponding classifier. Under this framework, iMMLDP combines dimensionality reduction, clustering, and classification in a principled way. Moreover, to deal with more complex data, for example, a local nonlinear separable structure, we extend the linear projection to a nonlinear case based on the kernel trick and develop an infinite kernel max-margin discriminant projection (iKMMDP) model. Thanks to the conjugate property, the parameters in these two models can be inferred efficiently via the Gibbs sampler. Finally, we implement our models on synthesized and real-world data, including multimodally distributed datasets and measured radar image data, to validate their efficiency and effectiveness.
Bo Chen 0001, Xuefei Cao, Xuefeng Zhang 0003, Zhengjue Wang, Hongwei Liu 0001
IEEE Trans. Cybern.5
2022 Unsupervised Hyperspectral and Multispectral Images Fusion Based on Nonlinear Variational Probabilistic Generative Model
abstract
Due to hardware limitations, it is challenging for sensors to acquire images of high resolution in both spatial and spectral domains, which arouses a trend that utilizing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to fuse an HR-HSI in an unsupervised manner. Considering the fact that most existing methods are restricted by using linear spectral unmixing, we propose a nonlinear variational probabilistic generative model (NVPGM) for the unsupervised fusion task based on nonlinear unmixing. We model the joint full likelihood of the observed pixels in an LR-HSI and an HR-MSI, both of which are assumed to be generated from the corresponding latent representations, i.e., the abundance vectors. The sufficient statistics of the generative conditional distributions are nonlinear functions with respect to the latent variable, realized by neural networks, which results in a nonlinear spectral mixture model. For scalability and efficiency, we construct two recognition models to infer the latent representations, which are parameterized by neural networks as well. Simultaneously inferring the latent representations and optimizing the parameters are achieved using stochastic gradient variational inference, after which the target HR-HSI is retrieved via feedforward mapping. Though without supervised information about the HR-HSI, NVPGM still can be trained based on extra LR-HSI and HR-MSI data sets in advance unsupervisedly and processes the images at the test phase in real time. Three commonly used data sets are used to evaluate the effectiveness and efficiency of NVPGM, illustrating the outperformance of NVPGM in the unsupervised LR-HSI and HR-MSI fusion task.
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering
abstract
Zhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Bo Chen 0001, Mingyuan Zhou
ACL/IJCNLP (1)4
2021 Memory-Efficient Network for Large-Scale Video Compressive Sensing
abstract
Video snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net.
Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002
CVPR6
2021 MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive Sensing
abstract
To capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-speed frames, where the state-of-the-art results are achieved by deep learning networks. However, these networks are usually trained for specific small-scale masks and often have high demands of training time and GPU memory, which are hence not flexible to i) a new mask with the same size and ii) a larger-scale mask. We address these challenges by developing a Meta Modulated Convolutional Network for SCI reconstruction, dubbed MetaSCI. MetaSCI is composed of a shared backbone for different masks, and light-weight meta-modulation parameters to evolve to different modulation parameters for each mask, thus having the properties of fast adaptation to new masks (or systems) and ready to scale to large data. Extensive simulation and real data results demonstrate the superior performance of our proposed approach. Our code is available at https://github.com/xyvirtualgroup/MetaSCI-CVPR2021.
Zhengjue Wang, Hao Zhang 0050, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002
CVPR1
2020 Learning Dynamic Hierarchical Topic Graph with Graph Convolutional Network for Document Classification
abstract
Constructing a graph with graph convolutional network (GCN) to explore the relational structure of the data has attracted lots of interests in various tasks. However, for document classification, existing graph based methods often focus on the straightforward word-word and word-document relations, ignoring the hierarchical semantics. Besides, the graph construction is often independent from the task-specific GCN learning. To address these constrains, we integrate a probabilistic deep topic model into graph construction, and propose a novel trainable hierarchical topic graph (HTG), including word-level, hierarchical topic-level and document-level nodes, exhibiting semantic variation from fine-grained to coarse. Regarding the document classification as a document-node label generation task, HTG can be dynamically evolved with GCN by performing variational inference, which leads to an end-to-end document classification method, named dynamic HTG (DHTG). Besides achieving state-of-the-art classification results, our model learns an interpretable document graph with meaningful node embeddings and semantic edges.
Zhengjue Wang, Chaojie Wang 0001, Hao Zhang 0050, Zhibin Duan, Mingyuan Zhou, Bo Chen 0001
AISTATS1
2020 BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002
ECCV (24)3
2020 Friendly Topic Assistant for Transformer Based Abstractive Summarization
abstract
ive document summarization is a comprehensive task including document understanding and summary generation, in which area Transformer-based models have achieved the state-of-the-art performance. Compared with Transformers, topic models are better at learning explicit document semantics, and hence could be integrated into Transformers to further boost their performance. To this end, we rearrange and explore the semantics learned by a topic model, and then propose a topic assistant (TA) including three modules. TA is compatible with various Transformer-based models and user-friendly since i) TA is a plug-and-play model that does not break any structure of the original Transformer network, making users easily fine-tune Transformer+TA based on a well pre-trained model; ii) TA only introduces a small number of extra parameters. Experimental results on three datasets demonstrate that TA is able to improve the performance of several Transformer-based models.
Zhengjue Wang, Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou
EMNLP (1)1
2020 Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Mingyuan Zhou
ICLR4
2020 Deep Relational Topic Modeling via Graph Poisson Gamma Belief Network
abstract
To analyze a collection of interconnected documents, relational topic models (RTMs) have been developed to describe both the link structure and document content, exploring their underlying relationships via a single-layer latent representation with limited expressive capability. To better utilize the document network, we first propose graph Poisson factor analysis (GPFA) that constructs a probabilistic model for interconnected documents and also provides closed-form Gibbs sampling update equations, moving beyond sophisticated approximate assumptions of existing RTMs. Extending GPFA, we develop a novel hierarchical RTM named graph Poisson gamma belief network (GPGBN), and further introduce two different Weibull distribution based variational graph auto-encoders for efficient model inference and effective network information aggregation. Experimental results demonstrate that our models extract high-quality hierarchical latent document representations, leading to improved performance over baselines on various graph analytic tasks.
Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Dongsheng Wang 0003, Zhengjue Wang, Mingyuan Zhou
NeurIPS5
2020 FusionNet: An Unsupervised Convolutional Variational Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to hardware limitations of the imaging sensors, it is challenging to acquire images of high resolution in both spatial and spectral domains. Fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to obtain an HR-HSI in an unsupervised manner has drawn considerable attention. Though effective, most existing fusion methods are limited due to the use of linear parametric modeling for the spectral mixture process, and even the deep learning-based methods only focus on deterministic fully-connected networks without exploiting the spatial correlation and local spectral structures of the images. In this paper, we propose a novel variational probabilistic autoencoder framework implemented by convolutional neural networks, in order to fuse the spatial and spectral information contained in the LR-HSI and HR-MSI, called FusionNet. The FusionNet consists of a spectral generative network, a spatial-dependent prior network, and a spatial-spectral variational inference network, which are jointly optimized in an unsupervised manner, leading to an end-to-end fusion system. Further, for fast adaptation to different observation scenes, we give a meta-learning explanation to the fusion problem, and combine the FusionNet with meta-learning in a synergistic manner. Effectiveness and efficiency of the proposed method are evaluated based on several publicly available datasets, demonstrating that the proposed FusionNet outperforms the state-of-the-art fusion methods.
Zhengjue Wang, Bo Chen 0001, Ruiying Lu, Hao Zhang 0050, Hongwei Liu 0001, Pramod K. Varshney
IEEE Trans. Image Process.1
2019 Variational probabilistic generative framework for single image super-resolution
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001
Signal Process.1
2019 Deep Max-Margin Discriminant Projection
abstract
In this paper, a unified Bayesian max-margin discriminant projection framework is proposed, which is able to jointly learn the discriminant feature space and the max-margin classifier with different relationships between the latent representations and observations. We assume that the latent representation follows a normal distribution whose sufficient statistics are functions of the observations. The function can be flexibly realized through either shallow or deep structures. The shallow structure includes linear, nonlinear kernel-based functions, and even the convolutional projection, which can be further trained layerwisely to build a multilayered convolutional feature learning model. To take the advantage of the deep neural networks, especially their highly expressive ability and efficient parameter learning, we integrate Bayesian modeling and the popular neural networks, for example, mltilayer perceptron and convolutional neural network, to build an end-to-end Bayesian deep discriminant projection under the proposed framework, which degenerated into the existing shallow linear or convolutional projection with the single-layer structure. Moreover, efficient scalable inferences for the realizations with different functions are derived to handle large-scale data via a stochastic gradient Markov chain Monte Carlo. Finally, we demonstrate the effectiveness and efficiency of the proposed models by the experiments on real-world data, including four image benchmarks (MNIST, CIFAR-10, STL-10, and SVHN) and one measured radar high-resolution range profile dataset, with the detailed analysis about the parameters and computational complexity.
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Hongwei Liu 0001
IEEE Trans. Cybern.3
2017 Structured Kernel Dictionary Learning With Correlation Constraint for Object Recognition
abstract
In this paper, we propose a new discriminative non-linear dictionary learning approach, called correlation constrained structured kernel KSVD, for object recognition. The objective function for dictionary learning contains a reconstructive term and a discriminative term. In the reconstructive term, signals are implicitly non-linearly mapped into a space, where a structured kernel dictionary, each sub-dictionary of which lies in the span of the mapped signals from the corresponding class, is established. In the discriminative term, by analyzing the classification mechanism, the correlation constraint is proposed in kernel form, constraining the correlations between different discriminative codes, and restricting the coefficient vectors to be transformed into a feature space, where the features are highly correlated inner-class and nearly independent between-classes. The objective function is optimized by the proposed structured kernel KSVD. During the classification stage, the specific form of the discriminative feature is needless to be known, while the inner product of the discriminative feature with kernel matrix embedded is available, and is suitable for a linear SVM classifier. Experimental results demonstrate that the proposed approach outperforms many state-of-the-art dictionary learning approaches for face, scene, and synthetic aperture radar vehicle target recognition.
Zhengjue Wang, Yinghua Wang, Hongwei Liu 0001, Hao Zhang 0050
IEEE Trans. Image Process.1