Hao Zhang 0050

dblp:55/2270-50 · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-2928-2692ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Non-Negative Deep VAE: The Generalized Gamma Belief Network
abstract
Gamma belief network (GBN), widely viewed as deep probabilistic topic models, has demonstrated its potential for uncovering multi-layer interpretable latent representations from text corpora. Its notable performance in document modeling largely arises from the expressive nature of gamma-distributed latent variables, which naturally capture sparsity, nonnegativity, skewness, heavy-tailed pattens, and from their seamless extension to multi-layer hierarchical structures. However, existing GBN and its variations are constrained by linear generative model, thereby limiting their expressiveness and applicability. To address this limitation, we introduce Generalized Gamma Belief Network (Generalized GBN), which extends original linear generative model to a more expressive non-linear generative model. Since parameters of Generalized GBN no longer possess an analytic conditional posterior, we further propose an upward-downward Weibull inference network to approximate posterior distribution of latent variables. The parameters of both generative model and inference network are jointly trained within variational inference framework. In addition, we provide theoretical analyses that demonstrate the effectiveness of Generalized GBN in modeling data variability and achieving disentangled representations. The former benefit arises from its hierarchical latent-variable structure, while the latter stems from its inherent ability to model sparsity. Finally, we conduct comprehensive experiments on both expressivity and disentangled representation learning tasks to evaluate the performance of Generalized GBN against Gaussian variational autoencoders serving as strong baseline models.
Zhibin Duan, Tiansheng Wen, Muyao Wang, Hao Zhang 0050, Bo Chen 0001, Hongwei Liu 0001, Mingyuan Zhou
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Learning transferable representations by topic guided graph adversarial network
Zhengjue Wang, Zhihui Xin, Chiyu Chen, Hao Zhang 0050, Yunsong Li 0001, Hongwei Liu 0001, Bo Chen 0001
Signal Process.4
2025 Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image information, leading to two main drawbacks: i) they often produce spurious visual-concept relations, hence decreasing model reliability; and ii) though CBMs could explain the importance of every concept to the final prediction, it is still challenging to tell which visual region produces the prediction. To solve these problems, this paper proposes a Disentangled Optimal Transport CBM (DOT-CBM) framework to explore fine-grained visual-concept relations between local image patches and concepts. Specifically, we model the concept prediction process as a transportation problem between the patches and concepts, thereby achieving explicit fine-grained feature alignment. We also incorporate orthogonal projection losses within the modality to enhance local feature disentanglement. To further address the shortcut issues caused by statistical biases in the data, we utilize the visual saliency map and concept label statistics as transportation priors. Thus, DOT-CBM can visualize inversion heatmaps, provide more reliable concept predictions, and produce more accurate class predictions. Comprehensive experiments demonstrate that our proposed DOT-CBM achieves SOTA performance on several tasks, including image classification, local part detection and out-of-distribution generalization. Codes are available in supplementary material.
Zequn Zeng, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001
CVPR3
2025 Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification
abstract
Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE.
Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma
CVPR5
2025 Manifoldron: Direct Space Partition via Manifold Discovery
abstract
A neural network (NN) with the widely-used ReLU activation has been shown to partition the sample space into many convex polytopes for prediction. However, the parametric way a NN and other machine learning models use to partition the space has imperfections, e.g., the compromised interpretability for complex models, the inflexibility in decision boundary construction due to the generic character of the model, and the risk of being trapped into shortcut solutions. In contrast, although the nonparameterized models can adorably avoid or downplay these issues, they are usually insufficiently powerful either due to over-simplification or the failure to accommodate the manifold structures of data. In this context, we first propose a new type of machine learning models referred to as Manifoldron that directly derives decision boundaries from data and partitions the space via manifold structure discovery. Then, we systematically analyze the key characteristics of the Manifoldron such as manifold characterization capability and its link to NNs. The experimental results on four synthetic examples, 20 public benchmark datasets, and one real-world application demonstrate that the proposed Manifoldron performs competitively compared to the mainstream machine learning models. We have shared our code in https://github.com/wdayang/Manifoldron for free download and evaluation.
Dayang Wang, Fenglei Fan, Bojian Hou, Hao Zhang 0050, Boce Zhang, Rongjie Lai, Hengyong Yu, Fei Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 MeaCap: Memory-Augmented Zero-shot Image Captioning
abstract
Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pre-trained vision-language models like CLIP for image-text similarity evaluation and a pre-trained language model (LM) for caption generation. The main difference be-tween them is whether using a textual corpus to train the LM. Though achieving attractive performance w.r.t. some metrics, existing methods often exhibit some com-mon drawbacks. Training-free methods tend to produce hallucinations, while text-only-training often lose gener-alization capability. To move forward, in this paper, we propose a novel Memory-Augmented zero-shot image Captioning framework (MeaCap). Specifically, equipped with a textual memory, we introduce a retrieve-then-filter module to get key concepts that are highly related to the image. By deploying our proposed memory-augmented visual-related fusion score in a keywords-to-sentence LM, MeaCap can generate concept-centered captions that keep high consistency with the image with fewer hallucinations and more world-knowledge. The framework of Mea-Cap achieves the state-of-the-art performance on a se-ries of zero-shot IC settings. Our code is available at https://github.com/joeyzOz/MeaCap.
Zequn Zeng, Hao Zhang 0050, Chiyu Chen, Bo Chen 0001, Zhengjue Wang
CVPR3
2024 HICEScore: A Hierarchical Metric for Image Captioning Evaluation
abstract
Image captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE.
Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001
ACM Multimedia3
2024 Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2
abstract
Text generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations.
Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.1
2023 ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based Polishing
abstract
Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searches every word in the caption using the knowledge of large-scale pre-trained models. Though effective, its autoregressive generation and gradient-directed searching mechanism limit the diversity of captions and inference speed, respectively. Moreover, ZeroCap does not consider the controllability issue of zero-shot IC. To move forward, we propose a framework for Controllable Zero-shot IC, named ConZIC. The core of ConZIC is a novel sampling-based non-autoregressive language model named Gibbs-BERT, which can generate and continuously polish every word. Extensive quantitative and qualitative results demonstrate the superior performance of our proposed ConZIC for both zero-shot IC and controllable zero-shot IC. Especially, ConZIC achieves about$5\times$generation speed than ZeroCap, and about$1.5\times$diversity scores, with accurate generation given different control signals. Our code is available at https://github.com/joeyz0z/ConZIC.
Zequn Zeng, Hao Zhang 0050, Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Zhengjue Wang
CVPR2
2023 Recurrent Neural Networks for Snapshot Compressive Imaging
abstract
Conventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT.
Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Generative Text Convolutional Neural Network for Hierarchical Document Representation Learning
abstract
For document analysis, existing methods often resort to the document representation that either discards the word order information or projects each word into a low-dimensional dense embedding vector. However, confined by the data's sparsity and high-dimensionality, limited effort has been made to explore the semantic structures underlying the document representation that formulates each document as a sequence of one-hot vectors, especially in the probabilistic modeling literature. To construct a probabilistic generative model for this type of document representation, we first develop convolutional Poisson factor analysis (CPFA) that not only utilizes the sparse property of data but also enables model parallelism. Through interleaving probabilistic Dirichlet-gamma pooling layers with learnable parameters, we extend the shallow CPFA into a generative text convolutional neural network (GTCNN), which captures richer semantic information with multiple probabilistic convolutional layers and can be coupled with existing deep topic models to alleviate their loss of word order. For efficient and scalable model inference, we not only develop both a parallel upward-downward Gibbs sampler and SG-MCMC based algorithm for training GTCNN, but also construct a hierarchical Weibull convolutional inference network for fast out-of-sample prediction. Experimental results on document representation learning tasks demonstrate the effectiveness of the proposed methods.
Chaojie Wang 0001, Bo Chen 0001, Zhibin Duan, Hao Zhang 0050, Mingyuan Zhou
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Multiscale Visual-Attribute Co-Attention for Zero-Shot Image Recognition
abstract
Zero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis.
Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Learning Hierarchical Document Graphs From Multilevel Sentence Relations
abstract
Organizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching.
Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin
IEEE Trans. Neural Networks Learn. Syst.1
2022 Optimal Transport with a New Preprocessing for Deep-Learning Full Waveform Inversion
abstract
Full waveform inversion (FWI) has been implemented using deep learning techniques as an analogue recurrent neural network for geophysics. However, the cycle-skipping issue, from which the conventional FWI suffers, troubles the deeplearning aided FWI as well if the least-square loss function is used to measure the misfit between observed and synthetic data. We propose to use a Wasserstein distance loss function combined with a newly designed preprocessing transform, named integration affine scaling, for the inversion. This transform transfers the seismograms into probability densities, and significantly improves the inversion results. Numerical results show that the proposed method outperforms its counterparts in mitigating cycle-skipping, in comparison with other loss functions including the least-square, the absolute, and the quadratic Wasserstein distance losses.
Hao Zhang 0050, Jianwei Ma 0006
ICIP1
2022 A Variational Edge Partition Model for Supervised Graph Representation Learning
abstract
Graph neural networks (GNNs), which propagate the node features through the edges and learn how to transform the aggregated features under label supervision, have achieved great success in supervised feature extraction for both node-level and graph-level classification tasks. However, GNNs typically treat the graph structure as given and ignore how the edges are formed. This paper introduces a graph generative process to model how the observed edges are generated by aggregating the node interactions over a set of overlapping node communities, each of which contributes to the edges via a logical OR mechanism. Based on this generative model, we partition each edge into the summation of multiple community-specific weighted edges and use them to define community-specific GNNs. A variational inference framework is proposed to jointly learn a GNN-based inference network that partitions the edges into different communities, these community-specific GNNs, and a GNN-based predictor that combines community-specific GNNs for the end classification task. Extensive evaluations on real-world graph datasets have verified the effectiveness of the proposed method in learning discriminative representations for both node-level and graph-level classification tasks.
Yilin He, Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou
NeurIPS3
2022 Multi-scale visual attention for attribute disambiguation in zero-shot learning
Bo Chen 0001, Hao Zhang 0050, Ning Han 0004, Yuanwei Chen, Hongwei Liu 0001
Signal Process. Image Commun.4
2022 Max-Margin Deep Diverse Latent Dirichlet Allocation With Continual Learning
abstract
Deep probabilistic aspect models are widely utilized in document analysis to extract the semantic information and obtain descriptive topics. However, there are two problems that may affect their applications. One is that common words shared among all documents with low representational meaning may reduce the representation ability of learned topics. The other is introducing supervision information to hierarchical topic models to fully utilize the side information of documents that is difficult. To address these problems, in this article, we first propose deep diverse latent Dirichlet allocation (DDLDA), a deep hierarchical topic model that can yield more meaningful semantic topics with less common and meaningless words by introducing shared topics. Moreover, we develop a variational inference network for DDLDA, which helps us to further generalize DDLDA to a supervised deep topic model called max-margin DDLDA (mmDDLDA) by employing max-margin principle as the classification criterion. Compared to DDLDA, mmDDLDA can discover more discriminative topical representations. In addition, a continual hybrid method with stochastic-gradient MCMC and variational inference is put forward for deep latent Dirichlet allocation (DLDA)-based models to make them more practical in real-world applications. The experimental results demonstrate that DDLDA and mmDDLDA are more efficient than existing unsupervised and supervised topic models in discovering highly discriminative topic representations and achieving higher classification accuracy. Meanwhile, DLDA and our proposed models trained by the proposed continual learning approach cannot only show good performance on preventing catastrophic forgetting but also fit the evolving new tasks well.
Bo Chen 0001, Yingqi Liu, Xuefei Cao, Qianru Zhao, Hao Zhang 0050
IEEE Trans. Cybern.6
2022 Multimodal Weibull Variational Autoencoder for Jointly Modeling Image-Text Data
abstract
For multimodal representation learning, traditional black-box approaches often fall short of extracting interpretable multilayer hidden structures, which contribute to visualize the connections between different modalities at multiple semantic levels. To extract interpretable multimodal latent representations and visualize the hierarchial semantic relationships between different modalities, based on deep topic models, we develop a novel multimodal Poisson gamma belief network (mPGBN) that tightly couples the observations of different modalities via imposing sparse connections between their modality-specific hidden layers. To alleviate the time-consuming Gibbs sampler adopted by traditional topic models in the testing stage, we construct a Weibull-based variational inference network (encoder) to directly map the observations to their latent representations, and further combine it with the mPGBN (decoder), resulting in a novel multimodal Weibull variational autoencoder (MWVAE), which is fast in out-of-sample prediction and can handle large-scale multimodal datasets. Qualitative evaluations on bimodal data consisting of image-text pairs show that the developed MWVAE can successfully extract expressive multimodal latent representations for downstream tasks like missing modality imputation and multimodal retrieval. Further extensive quantitative results demonstrate that both MWVAE and its supervised extension sMWVAE achieve state-of-the-art performance on various multimodal benchmarks.
Chaojie Wang 0001, Bo Chen 0001, Sucheng Xiao, Zhengjue Wang, Hao Zhang 0050, Ning Han 0004, Mingyuan Zhou
IEEE Trans. Cybern.5
2022 Unsupervised Hyperspectral and Multispectral Images Fusion Based on Nonlinear Variational Probabilistic Generative Model
abstract
Due to hardware limitations, it is challenging for sensors to acquire images of high resolution in both spatial and spectral domains, which arouses a trend that utilizing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to fuse an HR-HSI in an unsupervised manner. Considering the fact that most existing methods are restricted by using linear spectral unmixing, we propose a nonlinear variational probabilistic generative model (NVPGM) for the unsupervised fusion task based on nonlinear unmixing. We model the joint full likelihood of the observed pixels in an LR-HSI and an HR-MSI, both of which are assumed to be generated from the corresponding latent representations, i.e., the abundance vectors. The sufficient statistics of the generative conditional distributions are nonlinear functions with respect to the latent variable, realized by neural networks, which results in a nonlinear spectral mixture model. For scalability and efficiency, we construct two recognition models to infer the latent representations, which are parameterized by neural networks as well. Simultaneously inferring the latent representations and optimizing the parameters are achieved using stochastic gradient variational inference, after which the target HR-HSI is retrieved via feedforward mapping. Though without supervised information about the HR-HSI, NVPGM still can be trained based on extra LR-HSI and HR-MSI data sets in advance unsupervisedly and processes the images at the test phase in real time. Three commonly used data sets are used to evaluate the effectiveness and efficiency of NVPGM, illustrating the outperformance of NVPGM in the unsupervised LR-HSI and HR-MSI fusion task.
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering
abstract
Zhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Bo Chen 0001, Mingyuan Zhou
ACL/IJCNLP (1)2
2021 Memory-Efficient Network for Large-Scale Video Compressive Sensing
abstract
Video snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net.
Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002
CVPR4
2021 MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive Sensing
abstract
To capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-speed frames, where the state-of-the-art results are achieved by deep learning networks. However, these networks are usually trained for specific small-scale masks and often have high demands of training time and GPU memory, which are hence not flexible to i) a new mask with the same size and ii) a larger-scale mask. We address these challenges by developing a Meta Modulated Convolutional Network for SCI reconstruction, dubbed MetaSCI. MetaSCI is composed of a shared backbone for different masks, and light-weight meta-modulation parameters to evolve to different modulation parameters for each mask, thus having the properties of fast adaptation to new masks (or systems) and ready to scale to large data. Extensive simulation and real data results demonstrate the superior performance of our proposed approach. Our code is available at https://github.com/xyvirtualgroup/MetaSCI-CVPR2021.
Zhengjue Wang, Hao Zhang 0050, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002
CVPR2
2021 A Prototype-Oriented Framework for Unsupervised Domain Adaptation
abstract
Existing methods for unsupervised domain adaptation often rely on minimizing some statistical distance between the source and target samples in the latent space. To avoid the sampling variability, class imbalance, and data-privacy concerns that often plague these methods, we instead provide a memory and computation-efficient probabilistic framework to extract class prototypes and align the target features with them. We demonstrate the general applicability of our method on a wide range of scenarios, including single-source, multi-source, class-imbalance, and source-private domain adaptation. Requiring no additional model parameters and having a moderate increase in computation over the source model alone, the proposed method achieves competitive performance with state-of-the-art methods.
Korawat Tanwisuth, Xinjie Fan, Huangjie Zheng, Shujian Zhang, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou
NeurIPS5
2021 Deep Autoencoding Topic Model With Scalable Hybrid Bayesian Inference
abstract
To build a flexible and interpretable model for document analysis, we develop deep autoencoding topic model (DATM) that uses a hierarchy of gamma distributions to construct its multi-stochastic-layer generative network. In order to provide scalable posterior inference for the parameters of the generative network, we develop topic-layer-adaptive stochastic gradient Riemannian MCMC that jointly learns simplex-constrained global parameters across all layers and topics, with topic and layer specific learning rates. Given a posterior sample of the global parameters, in order to efficiently infer the local latent representations of a document under DATM across all stochastic layers, we propose a Weibull upward-downward variational encoder that deterministically propagates information upward via a deep neural network, followed by a Weibull distribution based stochastic downward generative model. To jointly model documents and their associated labels, we further propose supervised DATM that enhances the discriminative power of its latent representations. The efficacy and scalability of our models are demonstrated on both unsupervised and supervised learning tasks on big corpora.
Hao Zhang 0050, Bo Chen 0001, Yulai Cong, Dandan Guo, Hongwei Liu 0001, Mingyuan Zhou
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Learning Dynamic Hierarchical Topic Graph with Graph Convolutional Network for Document Classification
abstract
Constructing a graph with graph convolutional network (GCN) to explore the relational structure of the data has attracted lots of interests in various tasks. However, for document classification, existing graph based methods often focus on the straightforward word-word and word-document relations, ignoring the hierarchical semantics. Besides, the graph construction is often independent from the task-specific GCN learning. To address these constrains, we integrate a probabilistic deep topic model into graph construction, and propose a novel trainable hierarchical topic graph (HTG), including word-level, hierarchical topic-level and document-level nodes, exhibiting semantic variation from fine-grained to coarse. Regarding the document classification as a document-node label generation task, HTG can be dynamically evolved with GCN by performing variational inference, which leads to an end-to-end document classification method, named dynamic HTG (DHTG). Besides achieving state-of-the-art classification results, our model learns an interpretable document graph with meaningful node embeddings and semantic edges.
Zhengjue Wang, Chaojie Wang 0001, Hao Zhang 0050, Zhibin Duan, Mingyuan Zhou, Bo Chen 0001
AISTATS3
2020 BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002
ECCV (24)4
2020 Friendly Topic Assistant for Transformer Based Abstractive Summarization
abstract
ive document summarization is a comprehensive task including document understanding and summary generation, in which area Transformer-based models have achieved the state-of-the-art performance. Compared with Transformers, topic models are better at learning explicit document semantics, and hence could be integrated into Transformers to further boost their performance. To this end, we rearrange and explore the semantics learned by a topic model, and then propose a topic assistant (TA) including three modules. TA is compatible with various Transformer-based models and user-friendly since i) TA is a plug-and-play model that does not break any structure of the original Transformer network, making users easily fine-tune Transformer+TA based on a well pre-trained model; ii) TA only introduces a small number of extra parameters. Experimental results on three datasets demonstrate that TA is able to improve the performance of several Transformer-based models.
Zhengjue Wang, Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou
EMNLP (1)3
2020 Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Mingyuan Zhou
ICLR1
2020 Bidirectional Convolutional Poisson Gamma Dynamical Systems
abstract
Incorporating the natural document-sentence-word structure into hierarchical Bayesian modeling, we propose convolutional Poisson gamma dynamical systems (PGDS) that introduce not only word-level probabilistic convolutions, but also sentence-level stochastic temporal transitions. With word-level convolutions capturing phrase-level topics and sentence-level transitions capturing how the topic usages evolve over consecutive sentences, we aggregate the topic proportions of all sentences of a document as its feature representation. To consider not only forward but also backward sentence-level information transmissions, we further develop a bidirectional convolutional PGDS to incorporate the full contextual information to represent each sentence. For efficient inference, we construct a convolutional-recurrent inference network, which provides both sentence-level and document-level representations, and introduce a hybrid Bayesian inference scheme combining stochastic-gradient MCMC and amortized variational inference. Experimental results on a variety of document corpora demonstrate that the proposed models can extract expressive multi-level latent representations, including interpretable phrase-level topics and sentence-level temporal transitions as well as discriminative document-level features, achieving state-of-the-art document categorization performance while being memory and computation efficient.
Chaojie Wang 0001, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou
NeurIPS5
2020 Deep Relational Topic Modeling via Graph Poisson Gamma Belief Network
abstract
To analyze a collection of interconnected documents, relational topic models (RTMs) have been developed to describe both the link structure and document content, exploring their underlying relationships via a single-layer latent representation with limited expressive capability. To better utilize the document network, we first propose graph Poisson factor analysis (GPFA) that constructs a probabilistic model for interconnected documents and also provides closed-form Gibbs sampling update equations, moving beyond sophisticated approximate assumptions of existing RTMs. Extending GPFA, we develop a novel hierarchical RTM named graph Poisson gamma belief network (GPGBN), and further introduce two different Weibull distribution based variational graph auto-encoders for efficient model inference and effective network information aggregation. Experimental results demonstrate that our models extract high-quality hierarchical latent document representations, leading to improved performance over baselines on various graph analytic tasks.
Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Dongsheng Wang 0003, Zhengjue Wang, Mingyuan Zhou
NeurIPS2
2020 FusionNet: An Unsupervised Convolutional Variational Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to hardware limitations of the imaging sensors, it is challenging to acquire images of high resolution in both spatial and spectral domains. Fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to obtain an HR-HSI in an unsupervised manner has drawn considerable attention. Though effective, most existing fusion methods are limited due to the use of linear parametric modeling for the spectral mixture process, and even the deep learning-based methods only focus on deterministic fully-connected networks without exploiting the spatial correlation and local spectral structures of the images. In this paper, we propose a novel variational probabilistic autoencoder framework implemented by convolutional neural networks, in order to fuse the spatial and spectral information contained in the LR-HSI and HR-MSI, called FusionNet. The FusionNet consists of a spectral generative network, a spatial-dependent prior network, and a spatial-spectral variational inference network, which are jointly optimized in an unsupervised manner, leading to an end-to-end fusion system. Further, for fast adaptation to different observation scenes, we give a meta-learning explanation to the fusion problem, and combine the FusionNet with meta-learning in a synergistic manner. Effectiveness and efficiency of the proposed method are evaluated based on several publicly available datasets, demonstrating that the proposed FusionNet outperforms the state-of-the-art fusion methods.
Zhengjue Wang, Bo Chen 0001, Ruiying Lu, Hao Zhang 0050, Hongwei Liu 0001, Pramod K. Varshney
IEEE Trans. Image Process.4
2019 Variational probabilistic generative framework for single image super-resolution
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001
Signal Process.3
2019 Deep Max-Margin Discriminant Projection
abstract
In this paper, a unified Bayesian max-margin discriminant projection framework is proposed, which is able to jointly learn the discriminant feature space and the max-margin classifier with different relationships between the latent representations and observations. We assume that the latent representation follows a normal distribution whose sufficient statistics are functions of the observations. The function can be flexibly realized through either shallow or deep structures. The shallow structure includes linear, nonlinear kernel-based functions, and even the convolutional projection, which can be further trained layerwisely to build a multilayered convolutional feature learning model. To take the advantage of the deep neural networks, especially their highly expressive ability and efficient parameter learning, we integrate Bayesian modeling and the popular neural networks, for example, mltilayer perceptron and convolutional neural network, to build an end-to-end Bayesian deep discriminant projection under the proposed framework, which degenerated into the existing shallow linear or convolutional projection with the single-layer structure. Moreover, efficient scalable inferences for the realizations with different functions are derived to handle large-scale data via a stochastic gradient Markov chain Monte Carlo. Finally, we demonstrate the effectiveness and efficiency of the proposed models by the experiments on real-world data, including four image benchmarks (MNIST, CIFAR-10, STL-10, and SVHN) and one measured radar high-resolution range profile dataset, with the detailed analysis about the parameters and computational complexity.
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Hongwei Liu 0001
IEEE Trans. Cybern.1
2018 WHAI: Weibull Hybrid Autoencoding Inference for Deep Topic Modeling
Hao Zhang 0050, Bo Chen 0001, Dandan Guo, Mingyuan Zhou
ICLR (Poster)1
2018 Deep Poisson gamma dynamical systems
abstract
We develop deep Poisson-gamma dynamical systems (DPGDS) to model sequentially observed multivariate count data, improving previously proposed models by not only mining deep hierarchical latent structure from the data, but also capturing both first-order and long-range temporal dependencies. Using sophisticated but simple-to-implement data augmentation techniques, we derived closed-form Gibbs sampling update equations by first backward and upward propagating auxiliary latent counts, and then forward and downward sampling latent variables. Moreover, we develop stochastic gradient MCMC inference that is scalable to very long multivariate count time series. Experiments on both synthetic and a variety of real-world data demonstrate that the proposed model not only has excellent predictive performance, but also provides highly interpretable multilayer latent structure to represent hierarchical and temporal information propagation.
Dandan Guo, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou
NeurIPS3
2017 Structured Kernel Dictionary Learning With Correlation Constraint for Object Recognition
abstract
In this paper, we propose a new discriminative non-linear dictionary learning approach, called correlation constrained structured kernel KSVD, for object recognition. The objective function for dictionary learning contains a reconstructive term and a discriminative term. In the reconstructive term, signals are implicitly non-linearly mapped into a space, where a structured kernel dictionary, each sub-dictionary of which lies in the span of the mapped signals from the corresponding class, is established. In the discriminative term, by analyzing the classification mechanism, the correlation constraint is proposed in kernel form, constraining the correlations between different discriminative codes, and restricting the coefficient vectors to be transformed into a feature space, where the features are highly correlated inner-class and nearly independent between-classes. The objective function is optimized by the proposed structured kernel KSVD. During the classification stage, the specific form of the discriminative feature is needless to be known, while the inner product of the discriminative feature with kernel matrix embedded is available, and is suitable for a linear SVM classifier. Experimental results demonstrate that the proposed approach outperforms many state-of-the-art dictionary learning approaches for face, scene, and synthetic aperture radar vehicle target recognition.
Zhengjue Wang, Yinghua Wang, Hongwei Liu 0001, Hao Zhang 0050
IEEE Trans. Image Process.4
2015 Max-Margin Discriminant Projection via Data Augmentation
abstract
In this paper, we introduce a new max-margin discriminant projection method, which takes advantage of the latent variable representation for support vector machine (SVM) as the classification criterion. Specifically, the proposed model jointly learns the discriminative subspace and classifier in a Bayesian framework by conditioning on augmented variables. Moreover, an extended nonlinear model is developed based on the kernel trick, where the similar model can be used in this setting with few modifications. To explore the sparsity in the kernel expansion, we use the spike-and-slab prior to seek basis vectors (BVs) from the corresponding candidates. Unlike existing methods, which employ BVs to approximate the original feature space, in our method BVs are sought to associate the final classification task. Thanks to the conditionally conjugate property, the parameters in our models can be inferred via the simple and efficient Gibbs sampler. Finally, we test our methods on synthesized and real-world data, including large-scale data sets to demonstrate their efficiency and effectiveness.
Bo Chen 0001, Hao Zhang 0050, Xuefeng Zhang 0003, Hongwei Liu 0001, Jun Liu 0004
IEEE Trans. Knowl. Data Eng.2