VLDB 2026 Research / reviewers in the wild / expert
Hao Zhang 0050
dblp:55/2270-50
· DBLP profile ↗
37ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-2928-2692ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Non-Negative Deep VAE: The Generalized Gamma Belief NetworkabstractGamma belief network (GBN), widely viewed as deep probabilistic topic models, has demonstrated its potential for uncovering multi-layer interpretable latent representations from text corpora. Its notable performance in document modeling largely arises from the expressive nature of gamma-distributed latent variables, which naturally capture sparsity, nonnegativity, skewness, heavy-tailed pattens, and from their seamless extension to multi-layer hierarchical structures. However, existing GBN and its variations are constrained by linear generative model, thereby limiting their expressiveness and applicability. To address this limitation, we introduce Generalized Gamma Belief Network (Generalized GBN), which extends original linear generative model to a more expressive non-linear generative model. Since parameters of Generalized GBN no longer possess an analytic conditional posterior, we further propose an upward-downward Weibull inference network to approximate posterior distribution of latent variables. The parameters of both generative model and inference network are jointly trained within variational inference framework. In addition, we provide theoretical analyses that demonstrate the effectiveness of Generalized GBN in modeling data variability and achieving disentangled representations. The former benefit arises from its hierarchical latent-variable structure, while the latter stems from its inherent ability to model sparsity. Finally, we conduct comprehensive experiments on both expressivity and disentangled representation learning tasks to evaluate the performance of Generalized GBN against Gaussian variational autoencoders serving as strong baseline models. Zhibin Duan, Tiansheng Wen, Muyao Wang, Hao Zhang 0050, Bo Chen 0001, Hongwei Liu 0001, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Learning transferable representations by topic guided graph adversarial network
Zhengjue Wang, Zhihui Xin, Chiyu Chen, Hao Zhang 0050, Yunsong Li 0001, Hongwei Liu 0001, Bo Chen 0001 |
Signal Process. | 4 |
| 2025 | Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck ModelsabstractConcept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image information, leading to two main drawbacks: i) they often produce spurious visual-concept relations, hence decreasing model reliability; and ii) though CBMs could explain the importance of every concept to the final prediction, it is still challenging to tell which visual region produces the prediction. To solve these problems, this paper proposes a Disentangled Optimal Transport CBM (DOT-CBM) framework to explore fine-grained visual-concept relations between local image patches and concepts. Specifically, we model the concept prediction process as a transportation problem between the patches and concepts, thereby achieving explicit fine-grained feature alignment. We also incorporate orthogonal projection losses within the modality to enhance local feature disentanglement. To further address the shortcut issues caused by statistical biases in the data, we utilize the visual saliency map and concept label statistics as transportation priors. Thus, DOT-CBM can visualize inversion heatmaps, provide more reliable concept predictions, and produce more accurate class predictions. Comprehensive experiments demonstrate that our proposed DOT-CBM achieves SOTA performance on several tasks, including image classification, local part detection and out-of-distribution generalization. Codes are available in supplementary material. Zequn Zeng, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001 |
CVPR | 3 |
| 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationabstractConcept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE. Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma |
CVPR | 5 |
| 2025 | Manifoldron: Direct Space Partition via Manifold DiscoveryabstractA neural network (NN) with the widely-used ReLU activation has been shown to partition the sample space into many convex polytopes for prediction. However, the parametric way a NN and other machine learning models use to partition the space has imperfections, e.g., the compromised interpretability for complex models, the inflexibility in decision boundary construction due to the generic character of the model, and the risk of being trapped into shortcut solutions. In contrast, although the nonparameterized models can adorably avoid or downplay these issues, they are usually insufficiently powerful either due to over-simplification or the failure to accommodate the manifold structures of data. In this context, we first propose a new type of machine learning models referred to as Manifoldron that directly derives decision boundaries from data and partitions the space via manifold structure discovery. Then, we systematically analyze the key characteristics of the Manifoldron such as manifold characterization capability and its link to NNs. The experimental results on four synthetic examples, 20 public benchmark datasets, and one real-world application demonstrate that the proposed Manifoldron performs competitively compared to the mainstream machine learning models. We have shared our code in https://github.com/wdayang/Manifoldron for free download and evaluation. Dayang Wang, Fenglei Fan, Bojian Hou, Hao Zhang 0050, Boce Zhang, Rongjie Lai, Hengyong Yu, Fei Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | MeaCap: Memory-Augmented Zero-shot Image CaptioningabstractZero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pre-trained vision-language models like CLIP for image-text similarity evaluation and a pre-trained language model (LM) for caption generation. The main difference be-tween them is whether using a textual corpus to train the LM. Though achieving attractive performance w.r.t. some metrics, existing methods often exhibit some com-mon drawbacks. Training-free methods tend to produce hallucinations, while text-only-training often lose gener-alization capability. To move forward, in this paper, we propose a novel Memory-Augmented zero-shot image Captioning framework (MeaCap). Specifically, equipped with a textual memory, we introduce a retrieve-then-filter module to get key concepts that are highly related to the image. By deploying our proposed memory-augmented visual-related fusion score in a keywords-to-sentence LM, MeaCap can generate concept-centered captions that keep high consistency with the image with fewer hallucinations and more world-knowledge. The framework of Mea-Cap achieves the state-of-the-art performance on a se-ries of zero-shot IC settings. Our code is available at https://github.com/joeyzOz/MeaCap. Zequn Zeng, Hao Zhang 0050, Chiyu Chen, Bo Chen 0001, Zhengjue Wang |
CVPR | 3 |
| 2024 | HICEScore: A Hierarchical Metric for Image Captioning EvaluationabstractImage captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE. Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001 |
ACM Multimedia | 3 |
| 2024 | Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2abstractText generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations. Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based PolishingabstractZero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searches every word in the caption using the knowledge of large-scale pre-trained models. Though effective, its autoregressive generation and gradient-directed searching mechanism limit the diversity of captions and inference speed, respectively. Moreover, ZeroCap does not consider the controllability issue of zero-shot IC. To move forward, we propose a framework for Controllable Zero-shot IC, named ConZIC. The core of ConZIC is a novel sampling-based non-autoregressive language model named Gibbs-BERT, which can generate and continuously polish every word. Extensive quantitative and qualitative results demonstrate the superior performance of our proposed ConZIC for both zero-shot IC and controllable zero-shot IC. Especially, ConZIC achieves about$5\times$generation speed than ZeroCap, and about$1.5\times$diversity scores, with accurate generation given different control signals. Our code is available at https://github.com/joeyz0z/ConZIC. Zequn Zeng, Hao Zhang 0050, Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Zhengjue Wang |
CVPR | 2 |
| 2023 | Recurrent Neural Networks for Snapshot Compressive ImagingabstractConventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT. Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Generative Text Convolutional Neural Network for Hierarchical Document Representation LearningabstractFor document analysis, existing methods often resort to the document representation that either discards the word order information or projects each word into a low-dimensional dense embedding vector. However, confined by the data's sparsity and high-dimensionality, limited effort has been made to explore the semantic structures underlying the document representation that formulates each document as a sequence of one-hot vectors, especially in the probabilistic modeling literature. To construct a probabilistic generative model for this type of document representation, we first develop convolutional Poisson factor analysis (CPFA) that not only utilizes the sparse property of data but also enables model parallelism. Through interleaving probabilistic Dirichlet-gamma pooling layers with learnable parameters, we extend the shallow CPFA into a generative text convolutional neural network (GTCNN), which captures richer semantic information with multiple probabilistic convolutional layers and can be coupled with existing deep topic models to alleviate their loss of word order. For efficient and scalable model inference, we not only develop both a parallel upward-downward Gibbs sampler and SG-MCMC based algorithm for training GTCNN, but also construct a hierarchical Weibull convolutional inference network for fast out-of-sample prediction. Experimental results on document representation learning tasks demonstrate the effectiveness of the proposed methods. Chaojie Wang 0001, Bo Chen 0001, Zhibin Duan, Hao Zhang 0050, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Multiscale Visual-Attribute Co-Attention for Zero-Shot Image RecognitionabstractZero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis. Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Learning Hierarchical Document Graphs From Multilevel Sentence RelationsabstractOrganizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching. Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Optimal Transport with a New Preprocessing for Deep-Learning Full Waveform InversionabstractFull waveform inversion (FWI) has been implemented using deep learning techniques as an analogue recurrent neural network for geophysics. However, the cycle-skipping issue, from which the conventional FWI suffers, troubles the deeplearning aided FWI as well if the least-square loss function is used to measure the misfit between observed and synthetic data. We propose to use a Wasserstein distance loss function combined with a newly designed preprocessing transform, named integration affine scaling, for the inversion. This transform transfers the seismograms into probability densities, and significantly improves the inversion results. Numerical results show that the proposed method outperforms its counterparts in mitigating cycle-skipping, in comparison with other loss functions including the least-square, the absolute, and the quadratic Wasserstein distance losses. Hao Zhang 0050, Jianwei Ma 0006 |
ICIP | 1 |
| 2022 | A Variational Edge Partition Model for Supervised Graph Representation LearningabstractGraph neural networks (GNNs), which propagate the node features through the edges and learn how to transform the aggregated features under label supervision, have achieved great success in supervised feature extraction for both node-level and graph-level classification tasks. However, GNNs typically treat the graph structure as given and ignore how the edges are formed. This paper introduces a graph generative process to model how the observed edges are generated by aggregating the node interactions over a set of overlapping node communities, each of which contributes to the edges via a logical OR mechanism. Based on this generative model, we partition each edge into the summation of multiple community-specific weighted edges and use them to define community-specific GNNs. A variational inference framework is proposed to jointly learn a GNN-based inference network that partitions the edges into different communities, these community-specific GNNs, and a GNN-based predictor that combines community-specific GNNs for the end classification task. Extensive evaluations on real-world graph datasets have verified the effectiveness of the proposed method in learning discriminative representations for both node-level and graph-level classification tasks. Yilin He, Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 3 |
| 2022 | Multi-scale visual attention for attribute disambiguation in zero-shot learning
Bo Chen 0001, Hao Zhang 0050, Ning Han 0004, Yuanwei Chen, Hongwei Liu 0001 |
Signal Process. Image Commun. | 4 |
| 2022 | Max-Margin Deep Diverse Latent Dirichlet Allocation With Continual LearningabstractDeep probabilistic aspect models are widely utilized in document analysis to extract the semantic information and obtain descriptive topics. However, there are two problems that may affect their applications. One is that common words shared among all documents with low representational meaning may reduce the representation ability of learned topics. The other is introducing supervision information to hierarchical topic models to fully utilize the side information of documents that is difficult. To address these problems, in this article, we first propose deep diverse latent Dirichlet allocation (DDLDA), a deep hierarchical topic model that can yield more meaningful semantic topics with less common and meaningless words by introducing shared topics. Moreover, we develop a variational inference network for DDLDA, which helps us to further generalize DDLDA to a supervised deep topic model called max-margin DDLDA (mmDDLDA) by employing max-margin principle as the classification criterion. Compared to DDLDA, mmDDLDA can discover more discriminative topical representations. In addition, a continual hybrid method with stochastic-gradient MCMC and variational inference is put forward for deep latent Dirichlet allocation (DLDA)-based models to make them more practical in real-world applications. The experimental results demonstrate that DDLDA and mmDDLDA are more efficient than existing unsupervised and supervised topic models in discovering highly discriminative topic representations and achieving higher classification accuracy. Meanwhile, DLDA and our proposed models trained by the proposed continual learning approach cannot only show good performance on preventing catastrophic forgetting but also fit the evolving new tasks well. Bo Chen 0001, Yingqi Liu, Xuefei Cao, Qianru Zhao, Hao Zhang 0050 |
IEEE Trans. Cybern. | 6 |
| 2022 | Multimodal Weibull Variational Autoencoder for Jointly Modeling Image-Text DataabstractFor multimodal representation learning, traditional black-box approaches often fall short of extracting interpretable multilayer hidden structures, which contribute to visualize the connections between different modalities at multiple semantic levels. To extract interpretable multimodal latent representations and visualize the hierarchial semantic relationships between different modalities, based on deep topic models, we develop a novel multimodal Poisson gamma belief network (mPGBN) that tightly couples the observations of different modalities via imposing sparse connections between their modality-specific hidden layers. To alleviate the time-consuming Gibbs sampler adopted by traditional topic models in the testing stage, we construct a Weibull-based variational inference network (encoder) to directly map the observations to their latent representations, and further combine it with the mPGBN (decoder), resulting in a novel multimodal Weibull variational autoencoder (MWVAE), which is fast in out-of-sample prediction and can handle large-scale multimodal datasets. Qualitative evaluations on bimodal data consisting of image-text pairs show that the developed MWVAE can successfully extract expressive multimodal latent representations for downstream tasks like missing modality imputation and multimodal retrieval. Further extensive quantitative results demonstrate that both MWVAE and its supervised extension sMWVAE achieve state-of-the-art performance on various multimodal benchmarks. Chaojie Wang 0001, Bo Chen 0001, Sucheng Xiao, Zhengjue Wang, Hao Zhang 0050, Ning Han 0004, Mingyuan Zhou |
IEEE Trans. Cybern. | 5 |
| 2022 | Unsupervised Hyperspectral and Multispectral Images Fusion Based on Nonlinear Variational Probabilistic Generative ModelabstractDue to hardware limitations, it is challenging for sensors to acquire images of high resolution in both spatial and spectral domains, which arouses a trend that utilizing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to fuse an HR-HSI in an unsupervised manner. Considering the fact that most existing methods are restricted by using linear spectral unmixing, we propose a nonlinear variational probabilistic generative model (NVPGM) for the unsupervised fusion task based on nonlinear unmixing. We model the joint full likelihood of the observed pixels in an LR-HSI and an HR-MSI, both of which are assumed to be generated from the corresponding latent representations, i.e., the abundance vectors. The sufficient statistics of the generative conditional distributions are nonlinear functions with respect to the latent variable, realized by neural networks, which results in a nonlinear spectral mixture model. For scalability and efficiency, we construct two recognition models to infer the latent representations, which are parameterized by neural networks as well. Simultaneously inferring the latent representations and optimizing the parameters are achieved using stochastic gradient variational inference, after which the target HR-HSI is retrieved via feedforward mapping. Though without supervised information about the HR-HSI, NVPGM still can be trained based on extra LR-HSI and HR-MSI data sets in advance unsupervisedly and processes the images at the test phase in real time. Three commonly used data sets are used to evaluate the effectiveness and efficiency of NVPGM, illustrating the outperformance of NVPGM in the unsupervised LR-HSI and HR-MSI fusion task. Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | EnsLM: Ensemble Language Model for Data Diversity by Semantic ClusteringabstractZhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Bo Chen 0001, Mingyuan Zhou |
ACL/IJCNLP (1) | 2 |
| 2021 | Memory-Efficient Network for Large-Scale Video Compressive SensingabstractVideo snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net. Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002 |
CVPR | 4 |
| 2021 | MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive SensingabstractTo capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-speed frames, where the state-of-the-art results are achieved by deep learning networks. However, these networks are usually trained for specific small-scale masks and often have high demands of training time and GPU memory, which are hence not flexible to i) a new mask with the same size and ii) a larger-scale mask. We address these challenges by developing a Meta Modulated Convolutional Network for SCI reconstruction, dubbed MetaSCI. MetaSCI is composed of a shared backbone for different masks, and light-weight meta-modulation parameters to evolve to different modulation parameters for each mask, thus having the properties of fast adaptation to new masks (or systems) and ready to scale to large data. Extensive simulation and real data results demonstrate the superior performance of our proposed approach. Our code is available at https://github.com/xyvirtualgroup/MetaSCI-CVPR2021. Zhengjue Wang, Hao Zhang 0050, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002 |
CVPR | 2 |
| 2021 | A Prototype-Oriented Framework for Unsupervised Domain AdaptationabstractExisting methods for unsupervised domain adaptation often rely on minimizing some statistical distance between the source and target samples in the latent space. To avoid the sampling variability, class imbalance, and data-privacy concerns that often plague these methods, we instead provide a memory and computation-efficient probabilistic framework to extract class prototypes and align the target features with them. We demonstrate the general applicability of our method on a wide range of scenarios, including single-source, multi-source, class-imbalance, and source-private domain adaptation. Requiring no additional model parameters and having a moderate increase in computation over the source model alone, the proposed method achieves competitive performance with state-of-the-art methods. Korawat Tanwisuth, Xinjie Fan, Huangjie Zheng, Shujian Zhang, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 5 |
| 2021 | Deep Autoencoding Topic Model With Scalable Hybrid Bayesian InferenceabstractTo build a flexible and interpretable model for document analysis, we develop deep autoencoding topic model (DATM) that uses a hierarchy of gamma distributions to construct its multi-stochastic-layer generative network. In order to provide scalable posterior inference for the parameters of the generative network, we develop topic-layer-adaptive stochastic gradient Riemannian MCMC that jointly learns simplex-constrained global parameters across all layers and topics, with topic and layer specific learning rates. Given a posterior sample of the global parameters, in order to efficiently infer the local latent representations of a document under DATM across all stochastic layers, we propose a Weibull upward-downward variational encoder that deterministically propagates information upward via a deep neural network, followed by a Weibull distribution based stochastic downward generative model. To jointly model documents and their associated labels, we further propose supervised DATM that enhances the discriminative power of its latent representations. The efficacy and scalability of our models are demonstrated on both unsupervised and supervised learning tasks on big corpora. Hao Zhang 0050, Bo Chen 0001, Yulai Cong, Dandan Guo, Hongwei Liu 0001, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Learning Dynamic Hierarchical Topic Graph with Graph Convolutional Network for Document ClassificationabstractConstructing a graph with graph convolutional network (GCN) to explore the relational structure of the data has attracted lots of interests in various tasks. However, for document classification, existing graph based methods often focus on the straightforward word-word and word-document relations, ignoring the hierarchical semantics. Besides, the graph construction is often independent from the task-specific GCN learning. To address these constrains, we integrate a probabilistic deep topic model into graph construction, and propose a novel trainable hierarchical topic graph (HTG), including word-level, hierarchical topic-level and document-level nodes, exhibiting semantic variation from fine-grained to coarse. Regarding the document classification as a document-node label generation task, HTG can be dynamically evolved with GCN by performing variational inference, which leads to an end-to-end document classification method, named dynamic HTG (DHTG). Besides achieving state-of-the-art classification results, our model learns an interpretable document graph with meaningful node embeddings and semantic edges. Zhengjue Wang, Chaojie Wang 0001, Hao Zhang 0050, Zhibin Duan, Mingyuan Zhou, Bo Chen 0001 |
AISTATS | 3 |
| 2020 | BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002 |
ECCV (24) | 4 |
| 2020 | Friendly Topic Assistant for Transformer Based Abstractive Summarizationabstractive document summarization is a comprehensive task including document understanding and summary generation, in which area Transformer-based models have achieved the state-of-the-art performance. Compared with Transformers, topic models are better at learning explicit document semantics, and hence could be integrated into Transformers to further boost their performance. To this end, we rearrange and explore the semantics learned by a topic model, and then propose a topic assistant (TA) including three modules. TA is compatible with various Transformer-based models and user-friendly since i) TA is a plug-and-play model that does not break any structure of the original Transformer network, making users easily fine-tune Transformer+TA based on a well pre-trained model; ii) TA only introduces a small number of extra parameters. Experimental results on three datasets demonstrate that TA is able to improve the performance of several Transformer-based models. Zhengjue Wang, Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou |
EMNLP (1) | 3 |
| 2020 | Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Mingyuan Zhou |
ICLR | 1 |
| 2020 | Bidirectional Convolutional Poisson Gamma Dynamical SystemsabstractIncorporating the natural document-sentence-word structure into hierarchical Bayesian modeling, we propose convolutional Poisson gamma dynamical systems (PGDS) that introduce not only word-level probabilistic convolutions, but also sentence-level stochastic temporal transitions. With word-level convolutions capturing phrase-level topics and sentence-level transitions capturing how the topic usages evolve over consecutive sentences, we aggregate the topic proportions of all sentences of a document as its feature representation. To consider not only forward but also backward sentence-level information transmissions, we further develop a bidirectional convolutional PGDS to incorporate the full contextual information to represent each sentence. For efficient inference, we construct a convolutional-recurrent inference network, which provides both sentence-level and document-level representations, and introduce a hybrid Bayesian inference scheme combining stochastic-gradient MCMC and amortized variational inference. Experimental results on a variety of document corpora demonstrate that the proposed models can extract expressive multi-level latent representations, including interpretable phrase-level topics and sentence-level temporal transitions as well as discriminative document-level features, achieving state-of-the-art document categorization performance while being memory and computation efficient. Chaojie Wang 0001, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou |
NeurIPS | 5 |
| 2020 | Deep Relational Topic Modeling via Graph Poisson Gamma Belief NetworkabstractTo analyze a collection of interconnected documents, relational topic models (RTMs) have been developed to describe both the link structure and document content, exploring their underlying relationships via a single-layer latent representation with limited expressive capability. To better utilize the document network, we first propose graph Poisson factor analysis (GPFA) that constructs a probabilistic model for interconnected documents and also provides closed-form Gibbs sampling update equations, moving beyond sophisticated approximate assumptions of existing RTMs. Extending GPFA, we develop a novel hierarchical RTM named graph Poisson gamma belief network (GPGBN), and further introduce two different Weibull distribution based variational graph auto-encoders for efficient model inference and effective network information aggregation. Experimental results demonstrate that our models extract high-quality hierarchical latent document representations, leading to improved performance over baselines on various graph analytic tasks. Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Dongsheng Wang 0003, Zhengjue Wang, Mingyuan Zhou |
NeurIPS | 2 |
| 2020 | FusionNet: An Unsupervised Convolutional Variational Network for Hyperspectral and Multispectral Image FusionabstractDue to hardware limitations of the imaging sensors, it is challenging to acquire images of high resolution in both spatial and spectral domains. Fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to obtain an HR-HSI in an unsupervised manner has drawn considerable attention. Though effective, most existing fusion methods are limited due to the use of linear parametric modeling for the spectral mixture process, and even the deep learning-based methods only focus on deterministic fully-connected networks without exploiting the spatial correlation and local spectral structures of the images. In this paper, we propose a novel variational probabilistic autoencoder framework implemented by convolutional neural networks, in order to fuse the spatial and spectral information contained in the LR-HSI and HR-MSI, called FusionNet. The FusionNet consists of a spectral generative network, a spatial-dependent prior network, and a spatial-spectral variational inference network, which are jointly optimized in an unsupervised manner, leading to an end-to-end fusion system. Further, for fast adaptation to different observation scenes, we give a meta-learning explanation to the fusion problem, and combine the FusionNet with meta-learning in a synergistic manner. Effectiveness and efficiency of the proposed method are evaluated based on several publicly available datasets, demonstrating that the proposed FusionNet outperforms the state-of-the-art fusion methods. Zhengjue Wang, Bo Chen 0001, Ruiying Lu, Hao Zhang 0050, Hongwei Liu 0001, Pramod K. Varshney |
IEEE Trans. Image Process. | 4 |
| 2019 | Variational probabilistic generative framework for single image super-resolution
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001 |
Signal Process. | 3 |
| 2019 | Deep Max-Margin Discriminant ProjectionabstractIn this paper, a unified Bayesian max-margin discriminant projection framework is proposed, which is able to jointly learn the discriminant feature space and the max-margin classifier with different relationships between the latent representations and observations. We assume that the latent representation follows a normal distribution whose sufficient statistics are functions of the observations. The function can be flexibly realized through either shallow or deep structures. The shallow structure includes linear, nonlinear kernel-based functions, and even the convolutional projection, which can be further trained layerwisely to build a multilayered convolutional feature learning model. To take the advantage of the deep neural networks, especially their highly expressive ability and efficient parameter learning, we integrate Bayesian modeling and the popular neural networks, for example, mltilayer perceptron and convolutional neural network, to build an end-to-end Bayesian deep discriminant projection under the proposed framework, which degenerated into the existing shallow linear or convolutional projection with the single-layer structure. Moreover, efficient scalable inferences for the realizations with different functions are derived to handle large-scale data via a stochastic gradient Markov chain Monte Carlo. Finally, we demonstrate the effectiveness and efficiency of the proposed models by the experiments on real-world data, including four image benchmarks (MNIST, CIFAR-10, STL-10, and SVHN) and one measured radar high-resolution range profile dataset, with the detailed analysis about the parameters and computational complexity. Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Hongwei Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | WHAI: Weibull Hybrid Autoencoding Inference for Deep Topic Modeling
Hao Zhang 0050, Bo Chen 0001, Dandan Guo, Mingyuan Zhou |
ICLR (Poster) | 1 |
| 2018 | Deep Poisson gamma dynamical systemsabstractWe develop deep Poisson-gamma dynamical systems (DPGDS) to model sequentially observed multivariate count data, improving previously proposed models by not only mining deep hierarchical latent structure from the data, but also capturing both first-order and long-range temporal dependencies. Using sophisticated but simple-to-implement data augmentation techniques, we derived closed-form Gibbs sampling update equations by first backward and upward propagating auxiliary latent counts, and then forward and downward sampling latent variables. Moreover, we develop stochastic gradient MCMC inference that is scalable to very long multivariate count time series. Experiments on both synthetic and a variety of real-world data demonstrate that the proposed model not only has excellent predictive performance, but also provides highly interpretable multilayer latent structure to represent hierarchical and temporal information propagation. Dandan Guo, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou |
NeurIPS | 3 |
| 2017 | Structured Kernel Dictionary Learning With Correlation Constraint for Object RecognitionabstractIn this paper, we propose a new discriminative non-linear dictionary learning approach, called correlation constrained structured kernel KSVD, for object recognition. The objective function for dictionary learning contains a reconstructive term and a discriminative term. In the reconstructive term, signals are implicitly non-linearly mapped into a space, where a structured kernel dictionary, each sub-dictionary of which lies in the span of the mapped signals from the corresponding class, is established. In the discriminative term, by analyzing the classification mechanism, the correlation constraint is proposed in kernel form, constraining the correlations between different discriminative codes, and restricting the coefficient vectors to be transformed into a feature space, where the features are highly correlated inner-class and nearly independent between-classes. The objective function is optimized by the proposed structured kernel KSVD. During the classification stage, the specific form of the discriminative feature is needless to be known, while the inner product of the discriminative feature with kernel matrix embedded is available, and is suitable for a linear SVM classifier. Experimental results demonstrate that the proposed approach outperforms many state-of-the-art dictionary learning approaches for face, scene, and synthetic aperture radar vehicle target recognition. Zhengjue Wang, Yinghua Wang, Hongwei Liu 0001, Hao Zhang 0050 |
IEEE Trans. Image Process. | 4 |
| 2015 | Max-Margin Discriminant Projection via Data AugmentationabstractIn this paper, we introduce a new max-margin discriminant projection method, which takes advantage of the latent variable representation for support vector machine (SVM) as the classification criterion. Specifically, the proposed model jointly learns the discriminative subspace and classifier in a Bayesian framework by conditioning on augmented variables. Moreover, an extended nonlinear model is developed based on the kernel trick, where the similar model can be used in this setting with few modifications. To explore the sparsity in the kernel expansion, we use the spike-and-slab prior to seek basis vectors (BVs) from the corresponding candidates. Unlike existing methods, which employ BVs to approximate the original feature space, in our method BVs are sought to associate the final classification task. Thanks to the conditionally conjugate property, the parameters in our models can be inferred via the simple and efficient Gibbs sampler. Finally, we test our methods on synthesized and real-world data, including large-scale data sets to demonstrate their efficiency and effectiveness. Bo Chen 0001, Hao Zhang 0050, Xuefeng Zhang 0003, Hongwei Liu 0001, Jun Liu 0004 |
IEEE Trans. Knowl. Data Eng. | 2 |