VLDB 2026 Research / reviewers in the wild / expert
Bo Chen 0001
dblp:89/5615-1
· DBLP profile ↗
136ranked-venue papers
15as first author
84since 2021 · last 2026
0000-0001-5151-9388ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 10 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 1 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantically guided dynamic visual prototype refinement for compositional zero-shot learning
Zhong Peng, Yishi Xu, Gerong Wang, Jing Zhang 0151, Bo Chen 0001, Hongwei Liu 0001 |
Neurocomputing | 6 |
| 2026 | A hybrid SNN-ANN co-training paradigm for SAR ships with heterogeneous views and auxiliary flow
Mengqi Shen, Yinghua Wang, Zhiyong Suo, Hongwei Liu 0001, Bo Chen 0001 |
Neurocomputing | 6 |
| 2026 | A PID-like neural network control method for a 5PUS-RPUR parallel robot considering force coupling errors
Zesheng Wang 0009, Xuanyin Wang, Yanbiao Li 0002, Bo Chen 0001, Xianyong Dai |
Neurocomputing | 5 |
| 2026 | Complex distribution-aware calibration for imbalanced classification via optimal transport
Guanliang Liu, Kun Qin, Bo Chen 0001 |
Knowl. Based Syst. | 6 |
| 2026 | Alignment and disentanglement with adaptive-weighted optimal transport for compositional zero-shot learning
Zhong Peng, Gerong Wang, Dongsheng Wang 0003, Jing Zhang 0151, Bo Chen 0001 |
Knowl. Based Syst. | 6 |
| 2026 | A Non-Negative Deep VAE: The Generalized Gamma Belief NetworkabstractGamma belief network (GBN), widely viewed as deep probabilistic topic models, has demonstrated its potential for uncovering multi-layer interpretable latent representations from text corpora. Its notable performance in document modeling largely arises from the expressive nature of gamma-distributed latent variables, which naturally capture sparsity, nonnegativity, skewness, heavy-tailed pattens, and from their seamless extension to multi-layer hierarchical structures. However, existing GBN and its variations are constrained by linear generative model, thereby limiting their expressiveness and applicability. To address this limitation, we introduce Generalized Gamma Belief Network (Generalized GBN), which extends original linear generative model to a more expressive non-linear generative model. Since parameters of Generalized GBN no longer possess an analytic conditional posterior, we further propose an upward-downward Weibull inference network to approximate posterior distribution of latent variables. The parameters of both generative model and inference network are jointly trained within variational inference framework. In addition, we provide theoretical analyses that demonstrate the effectiveness of Generalized GBN in modeling data variability and achieving disentangled representations. The former benefit arises from its hierarchical latent-variable structure, while the latter stems from its inherent ability to model sparsity. Finally, we conduct comprehensive experiments on both expressivity and disentangled representation learning tasks to evaluate the performance of Generalized GBN against Gaussian variational autoencoders serving as strong baseline models. Zhibin Duan, Tiansheng Wen, Muyao Wang, Hao Zhang 0050, Bo Chen 0001, Hongwei Liu 0001, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Adaptive distribution calibration via optimal transport for imbalanced SAR automatic target recognition
Guanliang Liu, Bo Chen 0001, Zhiqiang Wan, Hongwei Liu 0001 |
Signal Process. | 3 |
| 2026 | Masked variational transformer for complex clutter modeling and target detection
Lixing Shi, Xueling Liang, Yaoqiang Liu, Kun Qin, Bo Chen 0001, Hongwei Liu 0001 |
Signal Process. | 7 |
| 2026 | Learning transferable representations by topic guided graph adversarial network
Zhengjue Wang, Zhihui Xin, Chiyu Chen, Hao Zhang 0050, Yunsong Li 0001, Hongwei Liu 0001, Bo Chen 0001 |
Signal Process. | 8 |
| 2026 | CAIR-Net: Reliability-Aware Information Routing for Robust Multimodal Object Detection Under Modality DegradationabstractMultimodal remote sensing combines optical and synthetic aperture radar (SAR) imagery to improve perception, yet real deployments face spatially varying degradations (e.g., clouds, low light, sensor interference) that can corrupt fusion. To make robustness measurable, we introduce a controlled mixed-severity setting in which only the optical stream is synthetically cloud-degraded while SAR remains intact, providing a standardized testbed for evaluating multimodal detection under modality imbalance. We further present CAIR-Net, a reliability–aware information routing network that follows adenoise-then-fuseprinciple: a Local Reliability Modulation (LRM) module learns soft, spatial reliability maps to suppress degraded regionsbeforecross-modal interaction, and a Global Information Selection Mechanism (GISM) performs confidence-aware expert routing across optical, fused, and SAR experts. On the mixed-severity benchmark, CAIR-Net consistently outperforms strong unimodal and fusion baselines and exhibits a substantially smaller performance drop under severe clouds (only a 7.3% AP reduction versus drops exceeding 25% for representative alternatives). These results indicate that explicit reliability modeling and quality-guided routing provide a practical path toward robust multimodal detection when one modality is partially or nearly completely occluded. Yudi Su, Jialei Ni, Tiansheng Wen, Hongwei Liu 0001, Hongtao Su, Bo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | MePAT: Meta-Prior Aided Transformer for Adverse Weather Condition RestorationabstractImage restoration under adverse weather conditions is critical for real-world applications. However, existing approaches mainly suffer from two fundamental limitations,i) the impractical requirement of prior degradation knowledge for task-specific model selection andii) performance degradation when handling with in-the-wild corruptions. To address the above issues, in this paper, we propose a novelMeta-prior Aided Transformerrestoration framework, MePAT, to synergize dynamic feature modulation with optimal transport (OT) theory. Specifically, we first architect an efficient attention mechanism,rectified self-channel attention(RSCA) to capture long-range associations along the channel dimension. Then, to adaptively tackle different conditions, we design atask-shared prior learning network(TPLN) to generate content-adaptive weather embeddings and serve as feature modulators to direct a more flexible and robust restoration process. In addition to learn discriminative task features, we propose an weakly-supervised OT-driven contrastive loss to measure the discrepancy between different weather corruptions. During the inference process, through the shared TPLN, we derive image-oriented vectors for unseen corruptions and then perform image restoration. The superior experimental results on three synthetic benchmarks demonstrate the effectiveness of MePAT. We also conduct experiments on real-world applications to verify the generalization ability and robustness. The code and pre-trained models will be made available. Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002, Chunhui Qu, Hongwei Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Rethinking Topic Modeling With Information Bottleneck Principle
Zhibin Duan, Bo Chen 0001, Chaojie Wang 0001, Xuefei Cao, Mingyuan Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck ModelsabstractConcept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image information, leading to two main drawbacks: i) they often produce spurious visual-concept relations, hence decreasing model reliability; and ii) though CBMs could explain the importance of every concept to the final prediction, it is still challenging to tell which visual region produces the prediction. To solve these problems, this paper proposes a Disentangled Optimal Transport CBM (DOT-CBM) framework to explore fine-grained visual-concept relations between local image patches and concepts. Specifically, we model the concept prediction process as a transportation problem between the patches and concepts, thereby achieving explicit fine-grained feature alignment. We also incorporate orthogonal projection losses within the modality to enhance local feature disentanglement. To further address the shortcut issues caused by statistical biases in the data, we utilize the visual saliency map and concept label statistics as transportation priors. Thus, DOT-CBM can visualize inversion heatmaps, provide more reliable concept predictions, and produce more accurate class predictions. Comprehensive experiments demonstrate that our proposed DOT-CBM achieves SOTA performance on several tasks, including image classification, local part detection and out-of-distribution generalization. Codes are available in supplementary material. Zequn Zeng, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001 |
CVPR | 7 |
| 2025 | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationabstractConcept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE. Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma |
CVPR | 7 |
| 2025 | Enhancing Uncertainty Estimation and Interpretability with Bayesian Non-negative Decision LayerabstractAlthough deep neural networks have demonstrated significant success due to their
powerful expressiveness, most models struggle to meet practical requirements for
uncertainty estimation. Concurrently, the entangled nature of deep neural net-
works leads to a multifaceted problem, where various localized explanation tech-
niques reveal that multiple unrelated features influence the decisions, thereby un-
dermining interpretability. To address these challenges, we develop a Bayesian
Nonnegative Decision Layer (BNDL), which reformulates deep neural networks
as a conditional Bayesian non-negative factor analysis. By leveraging stochastic
latent variables, the BNDL can model complex dependencies and provide robust
uncertainty estimation. Moreover, the sparsity and non-negativity of the latent
variables encourage the model to learn disentangled representations and decision
layers, thereby improving interpretability. We also offer theoretical guarantees
that BNDL can achieve effective disentangled learning. In addition, we developed
a corresponding variational inference method utilizing a Weibull variational in-
ference network to approximate the posterior distribution of the latent variables.
Our experimental results demonstrate that with enhanced disentanglement capa-
bilities, BNDL not only improves the model’s accuracy but also provides reliable
uncertainty estimation and improved interpretability. Zhibin Duan, Bo Chen 0001, Mingyuan Zhou |
ICLR | 3 |
| 2025 | OmiAD: One-Step Adaptive Masked Diffusion Model for Multi-class Anomaly Detection via Adversarial DistillationabstractDiffusion models have demonstrated outstanding performance in industrial anomaly detection. However, their iterative denoising nature results in slow inference speed, limiting their practicality for real-time industrial deployment. To address this challenge, we propose OmiAD, a one-step masked diffusion model for multi-class anomaly detection, derived from a well-designed multi-step Adaptive Masked Diffusion Model (AMDM) and compressed using Adversarial Score Distillation (ASD). OmiAD first introduces AMDM, equipped with an adaptive masking strategy that dynamically adjusts masking patterns based on noise levels and encourages the model to reconstruct anomalies as normal counterparts by leveraging broader context, to reduce the pixel-level shortcut reliance. Then, ASD is developed to compress the multi-step diffusion process into a single-step generator by score distillation and incorporating a shared-weight discriminator effectively reusing parameters while significantly improving both inference efficiency and detection performance. The effectiveness of OmiAD is validated on four diverse datasets, achieving state-of-the-art performance across seven metrics while delivering a remarkable inference speedup. Yaoxuan Feng, Yuxin Li 0003, Bo Chen 0001, Yubiao Wang, Hongwei Liu 0001, Mingyuan Zhou |
ICML | 4 |
| 2025 | Beyond Matryoshka: Revisiting Sparse Coding for Adaptive RepresentationabstractMany large-scale systems rely on high-quality deep representations (embeddings) to facilitate tasks like retrieval, search, and generative modeling. Matryoshka Representation Learning (MRL) recently emerged as a solution for adaptive embedding lengths, but it requires full model retraining and suffers from noticeable performance degradations at short lengths. In this paper, we show that sparse coding offers a compelling alternative for achieving adaptive representation with minimal overhead and higher fidelity. We propose Contrastive Sparse Representation (CSR), a method that specifies pre-trained embeddings into a high-dimensional but selectively activated feature space. By leveraging lightweight autoencoding and task-aware contrastive objectives, CSR preserves semantic quality while allowing flexible, cost-effective inference at different sparsity levels. Extensive experiments on image, text, and multimodal benchmarks demonstrate that CSR consistently outperforms MRL in terms of both accuracy and retrieval speed—often by large margins—while also cutting training time to a fraction of that required by MRL. Our results establish sparse coding as a powerful paradigm for adaptive representation learning in real-world applications where efficiency and fidelity are both paramount. Code is available at this URL. Tiansheng Wen, Yifei Wang 0001, Zequn Zeng, Zhong Peng, Yudi Su, Bo Chen 0001, Hongwei Liu 0001, Stefanie Jegelka, Chenyu You |
ICML | 7 |
| 2025 | Channel Matters: Estimating Channel Influence for Multivariate Time SeriesabstractThe influence function serves as an efficient post-hoc interpretability tool that quantifies the impact of training data modifications on model parameters, enabling enhanced model performance, improved generalization, and interpretability insights without the need for expensive retraining processes. Recently, Multivariate Time Series (MTS) analysis has become an important yet challenging task, attracting significant attention. While channel extremely matters to MTS tasks, channel-centric methods are still largely under-explored for MTS. Particularly, no previous work studied the effects of channel information of MTS in order to explore counterfactual effects between these channels and model performance. To fill this gap, we propose a novel Channel-wise Influence (ChInf) method that is the first to estimate the influence of different channels in MTS. Based on ChInf, we naturally derived two channel-wise algorithms by incorporating ChInf into classic MTS tasks. Extensive experiments demonstrate the effectiveness of ChInf and ChInf-based methods in critical MTS analysis tasks, such as MTS anomaly detection and MTS data pruning. Specifically, our ChInf-based methods rank top-1 among all methods for comparison, while previous influence functions do not perform well on MTS anomaly detection tasks and MTS data pruning problem. This fully supports the superiority and necessity of ChInf. Muyao Wang, Zeke Xie, Bo Chen 0001, James T. Kwok |
NeurIPS | 3 |
| 2025 | A Prompter Guided Distribution Matching for Zero-Shot SAR ATRabstractZero-shot learning (ZSL) in SAR automatic target recognition (ATR) aims to recognize targets of the unseen classes. ZSL methods have been widely studied in optical images; however, most of them suffer from lacking shared semantics space and domain shift problems in SAR images. To address these, we propose an optimizable visual prompter (OVP) that automatically extracts semantic features from casually selected optical images, improving flexibility compared to strict alignment methods. We also introduce an optimal transport (OT)-based distribution matching method to align unseen SAR images with learned category-agnostic semantics, addressing domain shift issues. The OT method is theoretically optimal for matching samples from two distributions, making it particularly effective for zero-shot SAR ATR. Experimental results on real-world SAR datasets validate the efficiency and accuracy of our approach for small- and large-scale zero-shot tasks. Chunhui Qu, Yujia Gan, Yaoqing Li, Mengqi Shen, Bo Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2025 | Robust Imaging of Aerial Targets With Maneuvering Trajectory Based on Multioptimization Strategies Under a Spaceborne SAR/ISAR Hybrid ModeabstractThe aerial targets with maneuvering trajectories can result in severe defocusing in inverse synthetic aperture radar (ISAR) images, while the low signal-to-noise ratio (SNR) from remote observations can also pose challenges to imaging. However, due to the stable motion state of satellites and their higher speeds compared to aerial targets, the accumulation time required for imaging is relatively short. This makes the impact of aerial target trajectory maneuvers on imaging effects relatively minor, providing a significant advantage for robust imaging of aerial targets with maneuvering trajectory under a spaceborne SAR/ISAR hybrid mode. This letter proposes an algorithm for robust imaging of aerial targets with maneuvering trajectory based on multioptimization strategies (OSs) under a spaceborne SAR/ISAR hybrid mode. In this algorithm, robust imaging is divided into two parts: 1) multistep optimal imaging time interval selection method based on OSs (MS-OITI-OSs). Through multiple optimization steps based on the target’s aerial trajectory and image information, the method achieves robust selection of the OITI by utilizing an “attitude first, quality later” optimization strategy, reducing computational complexity and increasing imaging success rates; and 2) joint translational motion compensation method based on OSs (JTMC-OSs). Characterizing motion parameters using a polynomial model, the method optimizes the motion parameters using image entropy as the objective function through the Grasshopper optimization algorithm (GOA). During the optimization process, a “high-order first, low-order later” optimization strategy is employed based on the impact of motion parameters on imaging quality to achieve robust translational motion compensation. The proposed algorithm enables robust imaging of aerial targets with maneuvering trajectory based on multi-OSs under a spaceborne SAR/ISAR hybrid mode. Extensive experimental validation confirms the effectiveness and robustness of the proposed method. Zhiqiang Wan, Shuai Shao 0011, Jiabo Fan, Gang Xu 0002, Bo Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Continual learning with Bayesian compression for shared and private latent representations
Yang Yang 0072, Dandan Guo, Bo Chen 0001, Dexiu Hu |
Neural Networks | 3 |
| 2025 | Supervised contrastive deep Q-Network for imbalanced radar automatic target recognition
Guanliang Liu, Bo Chen 0001, Bo Feng 0012, Hongwei Liu 0001 |
Pattern Recognit. | 3 |
| 2025 | Advancing Hyperspectral and Multispectral Image Fusion: An Information-Aware Transformer-Based Unfolding NetworkabstractIn hyperspectral image (HSI) processing, the fusion of the high-resolution multispectral image (HR-MSI) and the low-resolution HSI (LR-HSI) on the same scene, known as MSI-HSI fusion, is a crucial step in obtaining the desired high-resolution HSI (HR-HSI). With the powerful representation ability, convolutional neural network (CNN)-based deep unfolding methods have demonstrated promising performances. However, limited receptive fields of CNN often lead to inaccurate long-range spatial features, and inherent input and output images for each stage in unfolding networks restrict the feature transmission, thus limiting the overall performance. To this end, we propose a novel and efficient information-aware transformer-based unfolding network (ITU-Net) to model the long-range dependencies and transfer more information across the stages. Specifically, we employ a customized transformer block to learn representations from both the spatial and frequency domains as well as avoid the quadratic complexity with respect to the input length. For spatial feature extractions, we develop an information transfer guided linearized attention (ITLA), which transmits high-throughput information between adjacent stages and extracts contextual features along the spatial dimension in linear complexity. Moreover, we introduce frequency domain learning in the feedforward network (FFN) to capture token variations of the image and narrow the frequency gap. Via integrating our proposed transformer blocks with the unfolding framework, our ITU-Net achieves state-of-the-art (SOTA) performance on both synthetic and real hyperspectral datasets. Bo Chen 0001, Ruiying Lu, Ziheng Cheng 0001, Chunhui Qu, Xin Yuan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Considering Nonstationary within Multivariate Time Series with Variational Hierarchical Transformer for ForecastingabstractThe forecasting of Multivariate Time Series (MTS) has long been an important but challenging task. Due to the non-stationary problem across long-distance time steps, previous studies primarily adopt stationarization method to attenuate the non-stationary problem of original series for better predictability. However, existed methods always adopt the stationarized series, which ignore the inherent non-stationarity, and have difficulty in modeling MTS with complex distributions due to the lack of stochasticity. To tackle these problems, we first develop a powerful hierarchical probabilistic generative module to consider the non-stationarity and stochastity characteristics within MTS, and then combine it with transformer for a well-defined variational generative dynamic model named Hierarchical Time series Variational Transformer (HTV-Trans), which recovers the intrinsic non-stationary information into temporal dependencies. Being an powerful probabilistic model, HTV-Trans is utilized to learn expressive representations of MTS and applied to the forecasting tasks. Extensive experiments on diverse datasets show the efficiency of HTV-Trans on MTS forecasting tasks. Muyao Wang, Bo Chen 0001 |
AAAI | 3 |
| 2024 | MeaCap: Memory-Augmented Zero-shot Image CaptioningabstractZero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pre-trained vision-language models like CLIP for image-text similarity evaluation and a pre-trained language model (LM) for caption generation. The main difference be-tween them is whether using a textual corpus to train the LM. Though achieving attractive performance w.r.t. some metrics, existing methods often exhibit some com-mon drawbacks. Training-free methods tend to produce hallucinations, while text-only-training often lose gener-alization capability. To move forward, in this paper, we propose a novel Memory-Augmented zero-shot image Captioning framework (MeaCap). Specifically, equipped with a textual memory, we introduce a retrieve-then-filter module to get key concepts that are highly related to the image. By deploying our proposed memory-augmented visual-related fusion score in a keywords-to-sentence LM, MeaCap can generate concept-centered captions that keep high consistency with the image with fewer hallucinations and more world-knowledge. The framework of Mea-Cap achieves the state-of-the-art performance on a se-ries of zero-shot IC settings. Our code is available at https://github.com/joeyzOz/MeaCap. Zequn Zeng, Hao Zhang 0050, Chiyu Chen, Bo Chen 0001, Zhengjue Wang |
CVPR | 5 |
| 2024 | Instruction Tuning-Free Visual Token Complement for Multimodal LLMs
Dongsheng Wang 0003, Jiequan Cui, Miaoge Li, Bo Chen 0001, Hanwang Zhang |
ECCV (81) | 5 |
| 2024 | Transformer-Modulated Diffusion Models for Probabilistic Multivariate Time Series ForecastingabstractTransformers have gained widespread usage in multivariate time series (MTS) forecasting, delivering impressive performance. Nonetheless, these existing transformer-based methods often neglect an essential aspect: the incorporation of uncertainty into the predicted series, which holds significant value in decision-making. In this paper, we introduce a Transformer-Modulated Diffusion Model (TMDM), uniting conditional diffusion generative process with transformers into a unified framework to enable precise distribution forecasting for MTS. TMDM harnesses the power of transformers to extract essential insights from historical time series data. This information is then utilized as prior knowledge, capturing covariate-dependence in both the forward and reverse processes within the diffusion model. Furthermore, we seamlessly integrate well-designed transformer-based forecasting methods into TMDM to enhance its overall performance. Additionally, we introduce two novel metrics for evaluating uncertainty estimation performance. Through extensive experiments on six datasets using four evaluation metrics, we establish the effectiveness of TMDM in probabilistic MTS forecasting. Yuxin Li 0003, Bo Chen 0001, Baolin Sun, Mingyuan Zhou |
ICLR | 4 |
| 2024 | Vague Prototype-Oriented Diffusion Model for Multi-Class Anomaly DetectionabstractMulti-class unsupervised anomaly detection aims to create a unified model for identifying anomalies in objects from multiple classes when only normal data is available. In such a challenging setting, widely used reconstruction-based networks persistently grapple with the "identical shortcut" problem, wherein the infiltration of abnormal information from the condition biases the output towards an anomalous distribution. In response to this critical challenge, we introduce a Vague Prototype-Oriented Diffusion Model (VPDM) that extracts only fundamental information from the condition to prevent the occurrence of the "identical shortcut" problem from the input layer. This model leverages prototypes that contain only vague information about the target as the initial condition. Subsequently, a novel conditional diffusion model is introduced to incrementally enhance details based on vague conditions. Finally, a Vague Prototype-Oriented Optimal Transport (VPOT) method is proposed to provide more accurate information about conditions. All these components are seamlessly integrated into a unified optimization objective. The effectiveness of our approach is demonstrated across diverse datasets, including the MVTec, VisA, and MPDD benchmarks, achieving state-of-the-art results. Yuxin Li 0003, Yaoxuan Feng, Bo Chen 0001, Yubiao Wang, Baolin Sun, Chunhui Qu, Mingyuan Zhou |
ICML | 3 |
| 2024 | HICEScore: A Hierarchical Metric for Image Captioning EvaluationabstractImage captioning evaluation metrics can be divided into two categories, reference-based metrics and reference-free metrics. However, reference-based approaches may struggle to evaluate descriptive captions with abundant visual details produced by advanced multimodal large language models, due to their heavy reliance on limited human-annotated references. In contrast, previous reference-free metrics have been proven effective via CLIP cross-modality similarity. Nonetheless, CLIP-based metrics, constrained by their solution of global image-text compatibility, often have a deficiency in detecting local textual hallucinations and are insensitive to small visual objects. Besides, their single-scale designs are unable to provide an interpretable evaluation process such as pinpointing the position of caption mistakes and identifying visual regions that have not been described. To move forward, we propose a novel reference-free metric for image captioning evaluation, dubbed Hierarchical Image Captioning Evaluation Score (HICE-S). By detecting local visual regions and textual phrases, HICE-S builds an interpretable hierarchical scoring mechanism, breaking through the barriers of the single-scale structure of existing reference-free metrics. Comprehensive experiments indicate that our proposed metric achieves the SOTA performance on several benchmarks, outperforming existing reference-free metrics like CLIP-S and PAC-S, and reference-based metrics like METEOR and CIDEr. Moreover, several case studies reveal that the assessment process of HICE-S on detailed captions closely resembles interpretable human judgments.Our code is available at https://github.com/joeyz0z/HICE. Zequn Zeng, Hao Zhang 0050, Tiansheng Wen, Yudi Su, Zhengjue Wang, Bo Chen 0001 |
ACM Multimedia | 8 |
| 2024 | Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language ModelsabstractFor downstream applications of vision-language pre-trained models, there has been significant interest in constructing effective prompts. Existing works on prompt engineering, which either require laborious manual designs or optimize the prompt tuning as a point estimation problem, may fail to describe diverse characteristics of categories and limit their applications. We introduce a Bayesian probabilistic resolution to prompt tuning, where the label-specific stochastic prompts are generated hierarchically by first sampling a latent vector from an underlying distribution and then employing a lightweight generative model. Importantly, we semantically regularize the tuning process by minimizing the statistic distance between the visual patches and linguistic prompts, which pushes the stochastic label representations to faithfully capture diverse visual concepts, instead of overfitting the training categories. We evaluate the effectiveness of our approach on four tasks: few-shot image recognition, base-to-new generalization, dataset transfer learning, and domain shifts. Extensive results on over 15 datasets show promising transferability and generalization performance of our proposed model, both quantitatively and qualitatively. Dongsheng Wang 0003, Bowei Fang, Miaoge Li, Yishi Xu, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou |
UAI | 7 |
| 2024 | Discriminative shapelet learning via temporal clustering and matrix factorization
Bo Chen 0001, Guizhi Wang |
Appl. Intell. | 1 |
| 2024 | An adaptive class prototype generation framework for partial label learning
Haixiang Li, Xiao Li 0008, Bo Chen 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Selective-generative feature representations for generalized zero-shot open-set classification by learning a tightly clustered space
Xiao Li 0008, Haikun Li, Bo Chen 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Denoising matrix factorization for high-dimensional time series forecasting
Bo Chen 0001, Xiao Li 0008 |
Neural Comput. Appl. | 1 |
| 2024 | Motion-Aware Dynamic Graph Neural Network for Video Compressive SensingabstractVideo snapshot compressive imaging (SCI) utilizes a 2D detector to capture sequential video frames and compress them into a single measurement. Various reconstruction methods have been developed to recover the high-speed video frames from the snapshot measurement. However, most existing reconstruction methods are incapable of efficiently capturing long-range spatial and temporal dependencies, which are critical for video processing. In this paper, we propose a flexible and robust approach based on the graph neural network (GNN) to efficiently model non-local interactions between pixels in space and time regardless of the distance. Specifically, we develop a motion-aware dynamic GNN for better video representation, i.e., represent each node as the aggregation of relative neighbors under the guidance of frame-by-frame motions, which consists of motion-aware dynamic sampling, cross-scale node sampling, global knowledge integration, and graph aggregation. Extensive results on both simulation and real data demonstrate both the effectiveness and efficiency of the proposed approach, and the visualization illustrates the intrinsic dynamic sampling operations of our proposed model for boosting the video SCI reconstruction results. The code and model will be released. Ruiying Lu, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Hierarchical Topic-Aware Contextualized TransformersabstractTraining on disjoint fixed-length segments, Transformers convert static word embeddings into contextualized word representations. However, they often restrict the context of a token to the segment it resides in and hence neglect the contextual information across segments, failing to capture longer-term dependencies beyond the predefined segment length. This article uses a probabilistic deep topic model to provide hierarchical contextualized embeddings at both the token and segment levels, and integrate topic information through a constrained attention mechanism. The proposed method not only injects contextualized topic information into Transformers, but also controls languages generation guided by specific topics, styles, and sentiments. Three plug-and-play modules are proposed, including the contextual topical token embedding, the segment embedding, and the multi-head topic attention mechanism. We aim to capture the semantic coherence and word concurrence patterns at the global level, and also enrich the representation of each token by adapting to its local context, with negligible increased memory footprint and computational time. Experiments on various corpora show that by adding marginal extra parameters, the proposed hierarchical topic-aware contextualized Transformers consistently outperform their conventional counterparts, and generate sentences and paragraphs according to human preferences. Ruiying Lu, Bo Chen 0001, Dandan Guo, Dongsheng Wang 0003, Mingyuan Zhou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based PolishingabstractZero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC), ZeroCap abandons supervised training and sequentially searches every word in the caption using the knowledge of large-scale pre-trained models. Though effective, its autoregressive generation and gradient-directed searching mechanism limit the diversity of captions and inference speed, respectively. Moreover, ZeroCap does not consider the controllability issue of zero-shot IC. To move forward, we propose a framework for Controllable Zero-shot IC, named ConZIC. The core of ConZIC is a novel sampling-based non-autoregressive language model named Gibbs-BERT, which can generate and continuously polish every word. Extensive quantitative and qualitative results demonstrate the superior performance of our proposed ConZIC for both zero-shot IC and controllable zero-shot IC. Especially, ConZIC achieves about$5\times$generation speed than ZeroCap, and about$1.5\times$diversity scores, with accurate generation given different control signals. Our code is available at https://github.com/joeyz0z/ConZIC. Zequn Zeng, Hao Zhang 0050, Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Zhengjue Wang |
CVPR | 5 |
| 2023 | PatchCT: Aligning Patch Set and Label Set with Conditional Transport for Multi-Label Image ClassificationabstractMulti-label image classification is a prediction task that aims to identify more than one label from a given image. This paper considers the semantic consistency of the latent space between the visual patch and linguistic label domains and introduces the conditional transport (CT) theory to bridge the acknowledged gap. While recent cross-modal attention-based studies have attempted to align such two representations and achieved impressive performance, they required carefully-designed alignment modules and extra complex operations in the attention computation. We find that by formulating the multi-label classification as a CT problem, we can exploit the interactions between the image and label efficiently by minimizing the bidirectional CT cost. Specifically, after feeding the images and textual labels into the modality-specific encoders, we view each image as a mixture of patch embeddings and a mixture of label embeddings, which capture the local region features and the class prototypes, respectively. CT is then employed to learn and align those two semantic sets by defining the forward and backward navigators. Importantly, the defined navigators in CT distance model the similarities between patches and labels, which provides an interpretable tool to visualize the learned prototypes. Extensive experiments on three public image benchmarks show that the proposed model consistently outperforms the previous methods. Miaoge Li, Dongsheng Wang 0003, Zequn Zeng, Ruiying Lu, Bo Chen 0001, Mingyuan Zhou |
ICCV | 6 |
| 2023 | Prototypes-oriented Transductive Few-shot Learning with Conditional TransportabstractTransductive Few-Shot Learning (TFSL) has recently attracted increasing attention since it typically outperforms its inductive peer by leveraging statistics of query samples. However, previous TFSL methods usually encode uniform prior that all the classes within query samples are equally likely, which is biased in imbalanced TFSL and causes severe performance degradation. Given this pivotal issue, in this work, we propose a novel Conditional Transport (CT) based imbalanced TFSL model called Prototypes-oriented Unbiased Transfer Model (PUTM) to fully exploit unbiased statistics of imbalanced query samples, which employs forward and backward navigators as transport matrices to balance the prior of query samples per class between uniform and adaptive data-driven distributions. For efficiently transferring statistics learned by CT, we further derive a closed form solution to refine prototypes based on MAP given the learned navigators. The above two steps of discovering and transferring unbiased statistics follow an iterative manner, formulating our EM-based solver. Experimental results on four standard benchmarks including miniImageNet, tiered-ImageNet, CUB, and CIFAR-FS demonstrate superiority of our model in class-imbalanced generalization1. Jingyi Feng, Xiaoqiang Chai, Bo Chen 0001 |
ICCV | 7 |
| 2023 | Bayesian Progressive Deep Topic Model with Knowledge Informed Textual Data Coarsening ProcessabstractDeep topic models have shown an impressive ability to extract multi-layer document latent representations and discover hierarchical semantically meaningful topics.However, most deep topic models are limited to the single-step generative process, despite the fact that the progressive generative process has achieved impressive performance in modeling image data. To this end, in this paper, we propose a novel progressive deep topic model that consists of a knowledge-informed textural data coarsening process and a corresponding progressive generative model. The former is used to build multi-level observations ranging from concrete to abstract, while the latter is used to generate more concrete observations gradually. Additionally, we incorporate a graph-enhanced decoder to capture the semantic relationships among words at different levels of observation. Furthermore, we perform a theoretical analysis of the proposed model based on the principle of information theory and show how it can alleviate the well-known "latent variable collapse" problem. Finally, extensive experiments demonstrate that our proposed model effectively improves the ability of deep topic models, resulting in higher-quality latent document representations and topics. Zhibin Duan, Yudi Su, Yishi Xu, Bo Chen 0001, Mingyuan Zhou |
ICML | 5 |
| 2023 | Prototype-oriented unsupervised anomaly detection for multivariate time seriesabstractUnsupervised anomaly detection (UAD) of multivariate time series (MTS) aims to learn robust representations of normal multivariate temporal patterns. Existing UAD methods try to learn a fixed set of mappings for each MTS, entailing expensive computation and limited model adaptation. To address this pivotal issue, we propose a prototype-oriented UAD (PUAD) method under a probabilistic framework. Specifically, instead of learning the mappings for each MTS, the proposed PUAD views multiple MTSs as the distribution over a group of prototypes, which are extracted to represent a diverse set of normal patterns. To learn and regulate the prototypes, PUAD introduces a reconstruction-based unsupervised anomaly detection approach, which incorporates a prototype-oriented optimal transport method into a Transformer-powered probabilistic dynamical generative framework. Leveraging meta-learned transferable prototypes, PUAD can achieve high model adaptation capacity for new MTSs. Experiments on five public MTS datasets all verify the effectiveness of the proposed UAD method. Yuxin Li 0003, Bo Chen 0001, Dongsheng Wang 0003, Mingyuan Zhou |
ICML | 3 |
| 2023 | Tuning Multi-mode Token-level Prompt Alignment across ModalitiesabstractAdvancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily focus on single-mode (only one prompt for each modality) and holistic level (image or sentence) semantic alignment, which fails to capture the sample diversity, leading to sub-optimal prompt discovery. To address the limitation, we propose a multi-mode token-level tuning framework that leverages the optimal transportation to learn and align a set of prompt tokens across modalities. Specifically, we rely on two essential factors: 1) multi-mode prompts discovery, which guarantees diverse semantic representations, and 2) token-level alignment, which helps explore fine-grained similarity. Consequently, the similarity can be calculated as a hierarchical transportation problem between the modality-specific sets. Extensive experiments on popular image recognition benchmarks show the superior generalization and few-shot abilities of our approach. The qualitative analysis demonstrates that the learned prompt tokens have the ability to capture diverse visual concepts. Dongsheng Wang 0003, Miaoge Li, Mingsheng Xu, Bo Chen 0001, Hanwang Zhang |
NeurIPS | 5 |
| 2023 | Few-shot Generation via Recalling Brain-Inspired Episodic-Semantic MemoryabstractAimed at adapting a generative model to a novel generation task with only a few given data samples, the capability of few-shot generation is crucial for many real-world applications with limited data, \emph{e.g.}, artistic domains.
Instead of training from scratch, recent works tend to leverage the prior knowledge stored in previous datasets, which is quite similar to the memory mechanism of human intelligence, but few of these works directly imitate the memory-recall mechanism that humans make good use of in accomplishing creative tasks, \emph{e.g.}, painting and writing.
Inspired by the memory mechanism of human brain, in this work, we carefully design a variational structured memory module (VSM), which can simultaneously store both episodic and semantic memories to assist existing generative models efficiently recall these memories during sample generation.
Meanwhile, we introduce a bionic memory updating strategy for the conversion between episodic and semantic memories, which can also model the uncertainty during conversion.
Then, we combine the developed VSM with various generative models under the Bayesian framework, and evaluate these memory-augmented generative models with few-shot generation tasks, demonstrating the effectiveness of our methods. Zhibin Duan, Zhiyi Lv, Chaojie Wang 0001, Bo Chen 0001, Bo An 0001, Mingyuan Zhou |
NeurIPS | 4 |
| 2023 | Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly DetectionabstractUnsupervised image Anomaly Detection (UAD) aims to learn robust and discriminative representations of normal samples. While separate solutions per class endow expensive computation and limited generalizability, this paper focuses on building a unified framework for multiple classes. Under such a challenging setting, popular reconstruction-based networks with continuous latent representation assumption always suffer from the "identical shortcut" issue, where both normal and abnormal samples can be well recovered and difficult to distinguish. To address this pivotal issue, we propose a hierarchical vector quantized prototype-oriented Transformer under a probabilistic framework. First, instead of learning the continuous representations, we preserve the typical normal patterns as discrete iconic prototypes, and confirm the importance of Vector Quantization in preventing the model from falling into the shortcut. The vector quantized iconic prototypes are integrated into the Transformer for reconstruction, such that the abnormal data point is flipped to a normal data point. Second, we investigate an exquisite hierarchical framework to relieve the codebook collapse issue and replenish frail normal patterns. Third, a prototype-oriented optimal transport method is proposed to better regulate the prototypes and hierarchically evaluate the abnormal score. By evaluating on MVTec-AD and VisA datasets, our model surpasses the state-of-the-art alternatives and possesses good interpretability. The code is available at https://github.com/RuiyingLu/HVQ-Trans. Ruiying Lu, Dongsheng Wang 0003, Bo Chen 0001, Ruimin Hu |
NeurIPS | 5 |
| 2023 | Context-guided Embedding Adaptation for Effective Topic Modeling in Low-Resource RegimesabstractEmbedding-based neural topic models have turned out to be a superior option for low-resourced topic modeling. However, current approaches consider static word embeddings learnt from source tasks as general knowledge that can be transferred directly to the target task, discounting the dynamically changing nature of word meanings in different contexts, thus typically leading to sub-optimal results when adapting to new tasks with unfamiliar contexts. To settle this issue, we provide an effective method that centers on adaptively generating semantically tailored word embeddings for each task by fully exploiting contextual information. Specifically, we first condense the contextual syntactic dependencies of words into a semantic graph for each task, which is then modeled by a Variational Graph Auto-Encoder to produce task-specific word representations. On this basis, we further impose a learnable Gaussian mixture prior on the latent space of words to efficiently learn topic representations from a clustering perspective, which contributes to diverse topic discovery and fast adaptation to novel tasks. We have conducted a wealth of quantitative and qualitative experiments, and the results show that our approach comprehensively outperforms established topic models. Yishi Xu, Yudi Su, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 6 |
| 2023 | Recurrent Neural Networks for Snapshot Compressive ImagingabstractConventional high-speed and spectral imaging systems are expensive and they usually consume a significant amount of memory and bandwidth to save and transmit the high-dimensional data. By contrast, snapshot compressive imaging (SCI), where multiple sequential frames are coded by different masks and then summed to a single measurement, is a promising idea to use a 2-dimensional camera to capture 3-dimensional scenes. In this paper, we consider the reconstruction problem in SCI, i.e., recovering a series of scenes from a compressed measurement. Specifically, the measurement and modulation masks are fed into our proposed network, dubbed BIdirectional Recurrent Neural networks with Adversarial Training (BIRNAT) to reconstruct the desired frames. BIRNAT employs a deep convolutional neural network with residual blocks and self-attention to reconstruct the first frame, based on which a bidirectional recurrent neural network is utilized to sequentially reconstruct the following frames. Moreover, we build an extended BIRNAT-color algorithm for color videos aiming at joint reconstruction and demosaicing. Extensive results on both video and spectral, simulation and real data from three SCI cameras demonstrate the superior performance of BIRNAT. Ziheng Cheng 0001, Bo Chen 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Ziyi Meng 0001, Xin Yuan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Generative Text Convolutional Neural Network for Hierarchical Document Representation LearningabstractFor document analysis, existing methods often resort to the document representation that either discards the word order information or projects each word into a low-dimensional dense embedding vector. However, confined by the data's sparsity and high-dimensionality, limited effort has been made to explore the semantic structures underlying the document representation that formulates each document as a sequence of one-hot vectors, especially in the probabilistic modeling literature. To construct a probabilistic generative model for this type of document representation, we first develop convolutional Poisson factor analysis (CPFA) that not only utilizes the sparse property of data but also enables model parallelism. Through interleaving probabilistic Dirichlet-gamma pooling layers with learnable parameters, we extend the shallow CPFA into a generative text convolutional neural network (GTCNN), which captures richer semantic information with multiple probabilistic convolutional layers and can be coupled with existing deep topic models to alleviate their loss of word order. For efficient and scalable model inference, we not only develop both a parallel upward-downward Gibbs sampler and SG-MCMC based algorithm for training GTCNN, but also construct a hierarchical Weibull convolutional inference network for fast out-of-sample prediction. Experimental results on document representation learning tasks demonstrate the effectiveness of the proposed methods. Chaojie Wang 0001, Bo Chen 0001, Zhibin Duan, Hao Zhang 0050, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Data association for maneuvering targets through a combined siamese network and XGBoost model
Chang Gao 0004, Junkun Yan, Bo Chen 0001, Pramod K. Varshney, Tianyi Jia, Hongwei Liu 0001 |
Signal Process. | 3 |
| 2023 | Suspicious Object Detection for Millimeter-Wave Images With Multi-View Fusion Siamese NetworkabstractMillimeter-wave (MMW) imaging techniques have been widely used in the public security industries for their under-controlled privacy concerns and no health hazards. However, since MMW images are low resolution and most objects are small, reflection-weak, diverse, suspicious object detection in the MMW images is a very challenging task. This paper develops a robust suspicious object detector for the MMW images based on the Siamese network integrated with the pose estimation and image segmentation, which estimates the coordinates of human joints and segments the complete human images into symmetrical body part images. Unlike most existing detectors, which detect and recognize suspicious objects in MMW images and require a complete training set with correct annotations, our proposed model aims to learn the similarity between two symmetrical human body part images segmented from the complete MMW images. Furthermore, to decrease the misdetection caused by the restricted field of view, we further fuse the multi-view MMW images observed from the same person by designing a decision-level fusion strategy and feature-level fusion strategy based on the attention mechanism. Experimental results on the measured MMW images show that our proposed models have favorable detection accuracy and speed in practical application and thus prove their effectiveness. Dandan Guo, Chuan Du, Bo Chen 0001, Lei Zhang 0019 |
IEEE Trans. Image Process. | 5 |
| 2023 | Multiscale Visual-Attribute Co-Attention for Zero-Shot Image RecognitionabstractZero-shot image recognition aims to classify data from unseen classes, by exploring the association between visual features and the semantic representations of each class. Most existing approaches focus on learning a shared single-scale embedding space (often at the output layer of the network) for both visual and semantic features, ignoring a fact that different-scale visual features exhibit different semantics. In this article, we propose a multi-scale visual-attribute co-attention (mVACA) model, considering both visual-semantic alignment and visual discrimination at multiple scales. At each scale, a hybrid visual attention is realized by attribute-related attention and visual self-attention. The attribute-related attention is guided by a pseudo attribute vector inferred via a mutual information regularization (MIR). The visual self-attentive features further influence the attribute attention to emphasize visual-associated attributes. Leveraging multiscale visual discrimination, mVACA unifies standard zero-shot learning (ZSL) and generalized ZSL tasks in one framework, achieving state-of-the-art or competitive performance on several commonly used benchmarks of both setups. To better understand the interaction between images and attributes in mVACA, we also provide visualized analysis. Hao Zhang 0050, Zhengjue Wang, Yishi Xu, Pengyu Cheng, Ke Bai 0001, Bo Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Learning Hierarchical Document Graphs From Multilevel Sentence RelationsabstractOrganizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching. Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Representing Mixtures of Word Embeddings with Mixtures of Topic Embeddings
Dongsheng Wang 0003, Dandan Guo, He Zhao 0001, Huangjie Zheng, Korawat Tanwisuth, Bo Chen 0001, Mingyuan Zhou |
ICLR | 6 |
| 2022 | Deep Variational Graph Convolutional Recurrent Network for Multivariate Time Series Anomaly DetectionabstractAnomaly detection within multivariate time series (MTS) is an essential task in both data mining and service quality management. Many recent works on anomaly detection focus on designing unsupervised probabilistic models to extract robust normal patterns of MTS. In this paper, we model sensor dependency and stochasticity within MTS by developing an embedding-guided probabilistic generative network. We combine it with adaptive variational graph convolutional recurrent network %and get variational GCRN (VGCRN) to model both spatial and temporal fine-grained correlations in MTS. To explore hierarchical latent representations, we further extend VGCRN into a deep variational network, which captures multilevel information at different layers and is robust to noisy time series. Moreover, we develop an upward-downward variational inference scheme that considers both forecasting-based and reconstruction-based losses, achieving an accurate posterior approximation of latent variables with better MTS representations. The experiments verify the superiority of the proposed method over current state-of-the-art methods. Bo Chen 0001, Zhibin Duan, Mingyuan Zhou |
ICML | 3 |
| 2022 | Bayesian Deep Embedding Topic Meta-LearnerabstractExisting deep topic models are effective in capturing the latent semantic structures in textual data but usually rely on a plethora of documents. This is less than satisfactory in practical applications when only a limited amount of data is available. In this paper, we propose a novel framework that efficiently solves the problem of topic modeling under the small data regime. Specifically, the framework involves two innovations: a bi-level generative model that aims to exploit the task information to guide the document generation, and a topic meta-learner that strives to learn a group of global topic embeddings so that fast adaptation to the task-specific topic embeddings can be achieved with a few examples. We apply the proposed framework to a hierarchical embedded topic model and achieve better performance than various baseline models on diverse experiments, including few-shot topic discovery and few-shot document classification. Zhibin Duan, Yishi Xu, Bo Chen 0001, Chaojie Wang 0001, Mingyuan Zhou |
ICML | 4 |
| 2022 | Switching Gaussian Mixture Variational RNN for Anomaly Detection of Diverse CDN WebsitesabstractTo conduct service quality management of industry devices or Internet infrastructures, various deep learning approaches have been used for extracting the normal patterns of multivariate Key Performance Indicators (KPIs) for unsupervised anomaly detection. However, in the scenario of Content Delivery Networks (CDN), KPIs that belong to diverse websites usually exhibit various structures at different timesteps and show the non-stationary sequential relationship between them, which is extremely difficult for the existing deep learning approaches to characterize and identify anomalies. To address this issue, we propose a switching Gaussian mixture variational recurrent neural network (SGmVRNN) suitable for multivariate CDN KPIs. Specifically, SGmVRNN introduces the variational recurrent structure and assigns its latent variables into a mixture Gaussian distribution to model complex KPI time series and capture the diversely structural and dynamical characteristics within them, while in the next step it incorporates a switching mechanism to characterize these diversities, thus learning richer representations of KPIs. For efficient inference, we develop an upward-downward autoencoding inference method which combines the bottom-up likelihood and up-bottom prior information of the parameters for accurate posterior approximation. Extensive experiments on real-world data show that SGmVRNN significantly outperforms the state-of-the-art approaches according to F1-score on CDN KPIs from diverse websites. Yanwei Liu 0001, Antonios Argyriou, Tao Lin 0001, Zhen Xu 0009, Bo Chen 0001 |
INFOCOM | 9 |
| 2022 | Knowledge-Aware Bayesian Deep Topic ModelabstractWe propose a Bayesian generative model for incorporating prior domain knowledge into hierarchical topic modeling. Although embedded topic models (ETMs) and its variants have gained promising performance in text analysis, they mainly focus on mining word co-occurrence patterns, ignoring potentially easy-to-obtain prior topic hierarchies that could help enhance topic coherence. While several knowledge-based topic models have recently been proposed, they are either only applicable to shallow hierarchies or sensitive to the quality of the provided prior knowledge. To this end, we develop a novel deep ETM that jointly models the documents and the given prior knowledge by embedding the words and topics into the same space. Guided by the provided domain knowledge, the proposed model tends to discover topic hierarchies that are organized into interpretable taxonomies. Moreover, with a technique for adapting a given graph, our extended version allows the structure of the prior knowledge to be fine-tuned to match the target corpus. Extensive experiments show that our proposed model efficiently integrates the prior knowledge and improves both hierarchical topic discovery and document representation. Dongsheng Wang 0003, Yishi Xu, Miaoge Li, Zhibin Duan, Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 6 |
| 2022 | A Variational Edge Partition Model for Supervised Graph Representation LearningabstractGraph neural networks (GNNs), which propagate the node features through the edges and learn how to transform the aggregated features under label supervision, have achieved great success in supervised feature extraction for both node-level and graph-level classification tasks. However, GNNs typically treat the graph structure as given and ignore how the edges are formed. This paper introduces a graph generative process to model how the observed edges are generated by aggregating the node interactions over a set of overlapping node communities, each of which contributes to the edges via a logical OR mechanism. Based on this generative model, we partition each edge into the summation of multiple community-specific weighted edges and use them to define community-specific GNNs. A variational inference framework is proposed to jointly learn a GNN-based inference network that partitions the edges into different communities, these community-specific GNNs, and a GNN-based predictor that combines community-specific GNNs for the end classification task. Extensive evaluations on real-world graph datasets have verified the effectiveness of the proposed method in learning discriminative representations for both node-level and graph-level classification tasks. Yilin He, Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 4 |
| 2022 | Alleviating "Posterior Collapse" in Deep Topic Models via Policy GradientabstractDeep topic models have been proven as a promising way to extract hierarchical latent representations from documents represented as high-dimensional bag-of-words vectors.However, the representation capability of existing deep topic models is still limited by the phenomenon of "posterior collapse", which has been widely criticized in deep generative models, resulting in the higher-level latent representations exhibiting similar or meaningless patterns.To this end, in this paper, we first develop a novel deep-coupling generative process for existing deep topic models, which incorporates skip connections into the generation of documents, enforcing strong links between the document and its multi-layer latent representations.After that, utilizing data augmentation techniques, we reformulate the deep-coupling generative process as a Markov decision process and develop a corresponding Policy Gradient (PG) based training algorithm, which can further alleviate the information reduction at higher layers.Extensive experiments demonstrate that our developed methods can effectively alleviate "posterior collapse" in deep topic models, contributing to providing higher-quality latent document representations. Yewen Li, Chaojie Wang 0001, Zhibin Duan, Dongsheng Wang 0003, Bo Chen 0001, Bo An 0001, Mingyuan Zhou |
NeurIPS | 5 |
| 2022 | HyperMiner: Topic Taxonomy Mining with Hyperbolic EmbeddingabstractEmbedded topic models are able to learn interpretable topics even with large and heavy-tailed vocabularies. However, they generally hold the Euclidean embedding space assumption, leading to a basic limitation in capturing hierarchical relations. To this end, we present a novel framework that introduces hyperbolic embeddings to represent words and topics. With the tree-likeness property of hyperbolic space, the underlying semantic hierarchy among words and topics can be better exploited to mine more interpretable topics. Furthermore, due to the superiority of hyperbolic geometry in representing hierarchical data, tree-structure knowledge can also be naturally injected to guide the learning of a topic hierarchy. Therefore, we further develop a regularization term based on the idea of contrastive learning to inject prior structural knowledge efficiently. Experiments on both topic taxonomy discovery and document representation demonstrate that the proposed framework achieves improved performance against existing embedded topic models. Yishi Xu, Dongsheng Wang 0003, Bo Chen 0001, Ruiying Lu, Zhibin Duan, Mingyuan Zhou |
NeurIPS | 3 |
| 2022 | Matching Visual Features to Hierarchical Semantic Topics for Image Paragraph Captioning
Dandan Guo, Ruiying Lu, Bo Chen 0001, Zequn Zeng, Mingyuan Zhou |
Int. J. Comput. Vis. | 3 |
| 2022 | Intelligent multiframe detection aided by Doppler information and a deep neural network
Chang Gao 0004, Junkun Yan, Xiaojun Peng, Bo Chen 0001, Hongwei Liu 0001 |
Inf. Sci. | 4 |
| 2022 | Fast C&W: A Fast Adversarial Attack Algorithm to Fool SAR Target Recognition With Deep Convolutional Neural NetworksabstractIn recent years, deep convolutional neural networks (CNNs) pose superior synthetic aperture radar target recognition (SAR-TR) performance. However, CNN-based SAR classifiers would be vulnerable to adversarial attack (AA) when strong nonlinearity of CNN is contrapuntally utilized by AA. The AA can cause a CNN classifier to produce erroneous predictions with extremely high confidence by injecting a tiny adversarial perturbation to the input SAR images. In this letter, an accelerated SAR-TR AA algorithm is proposed named Fast C&W. We introduce a well-trained deep encoder network to replace the process of searching for the optimal perturbation of the input SAR image iteratively in the vanilla C&W algorithm. In this way, an adversarial perturbation can be generated much faster through the rapid forward mapping during an attack. Meanwhile, as a feature extraction network, the encoder network can learn the separable data region by optimizing the attack loss function. Through the encoder network, the added perturbation energy can be mainly concentrated on a region of target instead of background clutter area. This property would be of advantages in the perturbation location control in an SAR image. In the experiments, we use the proposed AA algorithm to interfere with the deep CNN-based high-accuracy SAR-TR model trained on the moving and stationary target acquisition and recognition (MSTAR) data set, which demonstrates its excellent effectiveness and thousands of times of efficiency improvement. Chuan Du, Chaoying Huo, Lei Zhang 0019, Bo Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Generalized zero-shot domain adaptation with target unseen class prototype learning
Xiao Li 0008, Bo Chen 0001 |
Neural Comput. Appl. | 3 |
| 2022 | Bayesian compression for dynamically expandable networks
Yang Yang 0072, Bo Chen 0001, Hongwei Liu 0001 |
Pattern Recognit. | 2 |
| 2022 | Open set HRRP recognition with few samples based on multi-modality prototypical networks
Bo Chen 0001, Zekun Guo, Chuan Du, Hongwei Liu 0001 |
Signal Process. | 2 |
| 2022 | Multi-scale visual attention for attribute disambiguation in zero-shot learning
Bo Chen 0001, Hao Zhang 0050, Ning Han 0004, Yuanwei Chen, Hongwei Liu 0001 |
Signal Process. Image Commun. | 2 |
| 2022 | Max-Margin Deep Diverse Latent Dirichlet Allocation With Continual LearningabstractDeep probabilistic aspect models are widely utilized in document analysis to extract the semantic information and obtain descriptive topics. However, there are two problems that may affect their applications. One is that common words shared among all documents with low representational meaning may reduce the representation ability of learned topics. The other is introducing supervision information to hierarchical topic models to fully utilize the side information of documents that is difficult. To address these problems, in this article, we first propose deep diverse latent Dirichlet allocation (DDLDA), a deep hierarchical topic model that can yield more meaningful semantic topics with less common and meaningless words by introducing shared topics. Moreover, we develop a variational inference network for DDLDA, which helps us to further generalize DDLDA to a supervised deep topic model called max-margin DDLDA (mmDDLDA) by employing max-margin principle as the classification criterion. Compared to DDLDA, mmDDLDA can discover more discriminative topical representations. In addition, a continual hybrid method with stochastic-gradient MCMC and variational inference is put forward for deep latent Dirichlet allocation (DLDA)-based models to make them more practical in real-world applications. The experimental results demonstrate that DDLDA and mmDDLDA are more efficient than existing unsupervised and supervised topic models in discovering highly discriminative topic representations and achieving higher classification accuracy. Meanwhile, DLDA and our proposed models trained by the proposed continual learning approach cannot only show good performance on preventing catastrophic forgetting but also fit the evolving new tasks well. Bo Chen 0001, Yingqi Liu, Xuefei Cao, Qianru Zhao, Hao Zhang 0050 |
IEEE Trans. Cybern. | 2 |
| 2022 | Multimodal Weibull Variational Autoencoder for Jointly Modeling Image-Text DataabstractFor multimodal representation learning, traditional black-box approaches often fall short of extracting interpretable multilayer hidden structures, which contribute to visualize the connections between different modalities at multiple semantic levels. To extract interpretable multimodal latent representations and visualize the hierarchial semantic relationships between different modalities, based on deep topic models, we develop a novel multimodal Poisson gamma belief network (mPGBN) that tightly couples the observations of different modalities via imposing sparse connections between their modality-specific hidden layers. To alleviate the time-consuming Gibbs sampler adopted by traditional topic models in the testing stage, we construct a Weibull-based variational inference network (encoder) to directly map the observations to their latent representations, and further combine it with the mPGBN (decoder), resulting in a novel multimodal Weibull variational autoencoder (MWVAE), which is fast in out-of-sample prediction and can handle large-scale multimodal datasets. Qualitative evaluations on bimodal data consisting of image-text pairs show that the developed MWVAE can successfully extract expressive multimodal latent representations for downstream tasks like missing modality imputation and multimodal retrieval. Further extensive quantitative results demonstrate that both MWVAE and its supervised extension sMWVAE achieve state-of-the-art performance on various multimodal benchmarks. Chaojie Wang 0001, Bo Chen 0001, Sucheng Xiao, Zhengjue Wang, Hao Zhang 0050, Ning Han 0004, Mingyuan Zhou |
IEEE Trans. Cybern. | 2 |
| 2022 | Infinite Bayesian Max-Margin Discriminant ProjectionabstractIn this article, considering the supervised dimensionality reduction, we first propose a model, called infinite Bayesian max-margin linear discriminant projection (iMMLDP), by assembling a set of local regions, where we make use of Bayesian nonparametric priors to handle the model selection problem, for example, the underlying number of local regions. In each local region, our model jointly learns a discriminative subspace and the corresponding classifier. Under this framework, iMMLDP combines dimensionality reduction, clustering, and classification in a principled way. Moreover, to deal with more complex data, for example, a local nonlinear separable structure, we extend the linear projection to a nonlinear case based on the kernel trick and develop an infinite kernel max-margin discriminant projection (iKMMDP) model. Thanks to the conjugate property, the parameters in these two models can be inferred efficiently via the Gibbs sampler. Finally, we implement our models on synthesized and real-world data, including multimodally distributed datasets and measured radar image data, to validate their efficiency and effectiveness. Bo Chen 0001, Xuefei Cao, Xuefeng Zhang 0003, Zhengjue Wang, Hongwei Liu 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Attitude and Size Estimation of Satellite Targets Based on ISAR Image InterpretationabstractThe attitude and size of satellite targets are essential information for their activity analysis. This article proposes a novel approach to estimate the absolute attitude and size of satellite targets in the 3-D stable coordinates based on inverse synthetic aperture radar (ISAR) image interpretation. In an ISAR image of a satellite, the satellite’s body is chosen as an individual structure segmented from each ISAR image by employing pix2pix generative adversarial network (Pix2pixGAN). By exploring the shape feature of the satellite body with principal component analysis (PCA), the satellite attitude and size are estimated jointly through solving an optimization based on the gradient iteration method. The optimization is established by bridging range-Doppler (RD) images and the target feature parameters (attitude and size) with the accommodation of target trajectory information and the ISAR geometric projection model. In the experiments, the simulation data are generated from real satellite orbital parameters and the computer-aided-design (CAD) models of two satellite targets: TianGong (TG) and KeyHole (KH). Compared with the factorization-based reconstruction method, the proposed method can estimate the attitude and size of the satellite simultaneously and has a higher size estimation accuracy. Lan Du 0001, Yachao Li 0001, Guoxin Lyu, Bo Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Unsupervised Hyperspectral and Multispectral Images Fusion Based on Nonlinear Variational Probabilistic Generative ModelabstractDue to hardware limitations, it is challenging for sensors to acquire images of high resolution in both spatial and spectral domains, which arouses a trend that utilizing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to fuse an HR-HSI in an unsupervised manner. Considering the fact that most existing methods are restricted by using linear spectral unmixing, we propose a nonlinear variational probabilistic generative model (NVPGM) for the unsupervised fusion task based on nonlinear unmixing. We model the joint full likelihood of the observed pixels in an LR-HSI and an HR-MSI, both of which are assumed to be generated from the corresponding latent representations, i.e., the abundance vectors. The sufficient statistics of the generative conditional distributions are nonlinear functions with respect to the latent variable, realized by neural networks, which results in a nonlinear spectral mixture model. For scalability and efficiency, we construct two recognition models to infer the latent representations, which are parameterized by neural networks as well. Simultaneously inferring the latent representations and optimizing the parameters are achieved using stochastic gradient variational inference, after which the target HR-HSI is retrieved via feedforward mapping. Though without supervised information about the HR-HSI, NVPGM still can be trained based on extra LR-HSI and HR-MSI data sets in advance unsupervisedly and processes the images at the test phase in real time. Three commonly used data sets are used to evaluate the effectiveness and efficiency of NVPGM, illustrating the outperformance of NVPGM in the unsupervised LR-HSI and HR-MSI fusion task. Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | EnsLM: Ensemble Language Model for Data Diversity by Semantic ClusteringabstractZhibin Duan, Hao Zhang, Chaojie Wang, Zhengjue Wang, Bo Chen, Mingyuan Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Bo Chen 0001, Mingyuan Zhou |
ACL/IJCNLP (1) | 5 |
| 2021 | Memory-Efficient Network for Large-Scale Video Compressive SensingabstractVideo snapshot compressive imaging (SCI) captures a sequence of video frames in a single shot using a 2D detector. The underlying principle is that during one exposure time, different masks are imposed on the high-speed scene to form a compressed measurement. With the knowledge of masks, optimization algorithms or deep learning methods are employed to reconstruct the desired high-speed video frames from this snapshot measurement. Unfortunately, though these methods can achieve decent results, the long running time of optimization algorithms or huge training memory occupation of deep networks still preclude them in practical applications. In this paper, we develop a memory-efficient network for large-scale video SCI based on multi-group reversible 3D convolutional neural networks. In addition to the basic model for the grayscale SCI system, we take one step further to combine demosaicing and SCI reconstruction to directly recover color video from Bayer measurements. Extensive results on both simulation and real data captured by SCI cameras demonstrate that our proposed model outperforms previous state-of-the-art with less memory and thus can be used in large-scale problems. The code is at https: //github.com/BoChenGroup/RevSCI-net. Ziheng Cheng 0001, Bo Chen 0001, Guanliang Liu, Hao Zhang 0050, Ruiying Lu, Zhengjue Wang, Xin Yuan 0002 |
CVPR | 2 |
| 2021 | MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive SensingabstractTo capture high-speed videos using a two-dimensional detector, video snapshot compressive imaging (SCI) is a promising system, where the video frames are coded by different masks and then compressed to a snapshot measurement. Following this, efficient algorithms are desired to reconstruct the high-speed frames, where the state-of-the-art results are achieved by deep learning networks. However, these networks are usually trained for specific small-scale masks and often have high demands of training time and GPU memory, which are hence not flexible to i) a new mask with the same size and ii) a larger-scale mask. We address these challenges by developing a Meta Modulated Convolutional Network for SCI reconstruction, dubbed MetaSCI. MetaSCI is composed of a shared backbone for different masks, and light-weight meta-modulation parameters to evolve to different modulation parameters for each mask, thus having the properties of fast adaptation to new masks (or systems) and ready to scale to large data. Extensive simulation and real data results demonstrate the superior performance of our proposed approach. Our code is available at https://github.com/xyvirtualgroup/MetaSCI-CVPR2021. Zhengjue Wang, Hao Zhang 0050, Ziheng Cheng 0001, Bo Chen 0001, Xin Yuan 0002 |
CVPR | 4 |
| 2021 | Sawtooth Factorial Topic Embeddings Guided Gamma Belief NetworkabstractHierarchical topic models such as the gamma belief network (GBN) have delivered promising results in mining multi-layer document representations and discovering interpretable topic taxonomies. However, they often assume in the prior that the topics at each layer are independently drawn from the Dirichlet distribution, ignoring the dependencies between the topics both at the same layer and across different layers. To relax this assumption, we propose sawtooth factorial topic embedding guided GBN, a deep generative model of documents that captures the dependencies and semantic similarities between the topics in the embedding space. Specifically, both the words and topics are represented as embedding vectors of the same dimension. The topic matrix at a layer is factorized into the product of a factor loading matrix and a topic embedding matrix, the transpose of which is set as the factor loading matrix of the layer above. Repeating this particular type of factorization, which shares components between adjacent layers, leads to a structure referred to as sawtooth factorization. An auto-encoding variational inference network is constructed to optimize the model parameter via stochastic gradient descent. Experiments on big corpora show that our models outperform other neural topic models on extracting deeper interpretable topics and deriving better document representations. Zhibin Duan, Dongsheng Wang 0003, Bo Chen 0001, Chaojie Wang 0001, Yewen Li, Mingyuan Zhou |
ICML | 3 |
| 2021 | Bayesian Attention Belief NetworksabstractAttention-based neural networks have achieved state-of-the-art results on a wide range of tasks. Most such models use deterministic attention while stochastic attention is less explored due to the optimization difficulties or complicated model design. This paper introduces Bayesian attention belief networks, which construct a decoder network by modeling unnormalized attention weights with a hierarchy of gamma distributions, and an encoder network by stacking Weibull distributions with a deterministic-upward-stochastic-downward structure to approximate the posterior. The resulting auto-encoding networks can be optimized in a differentiable way with a variational lower bound. It is simple to convert any models with deterministic attention, including pretrained ones, to the proposed Bayesian attention belief networks. On a variety of language understanding tasks, we show that our method outperforms deterministic attention and state-of-the-art stochastic attention in accuracy, uncertainty estimation, generalization across domains, and robustness to adversarial attacks. We further demonstrate the general applicability of our method on neural machine translation and visual question answering, showing great potential of incorporating our method into various attention-related tasks. Shujian Zhang, Xinjie Fan, Bo Chen 0001, Mingyuan Zhou |
ICML | 3 |
| 2021 | TopicNet: Semantic Graph-Guided Topic DiscoveryabstractExisting deep hierarchical topic models are able to extract semantically meaningful topics from a text corpus in an unsupervised manner and automatically organize them into a topic hierarchy. However, it is unclear how to incorporate prior belief such as knowledge graph to guide the learning of the topic hierarchy. To address this issue, we introduce TopicNet as a deep hierarchical topic model that can inject prior structural knowledge as inductive bias to influence the learning. TopicNet represents each topic as a Gaussian-distributed embedding vector, projects the topics of all layers into a shared embedding space, and explores both the symmetric and asymmetric similarities between Gaussian embedding vectors to incorporate prior semantic hierarchies. With a variational auto-encoding inference network, the model parameters are optimized by minimizing the evidence lower bound and supervised loss via stochastic gradient descent. Experiments on widely used benchmark show that TopicNet outperforms related deep topic models on discovering deeper interpretable topics and mining better document representations. Zhibin Duan, Yishi Xu, Bo Chen 0001, Dongsheng Wang 0003, Chaojie Wang 0001, Mingyuan Zhou |
NeurIPS | 3 |
| 2021 | A Prototype-Oriented Framework for Unsupervised Domain AdaptationabstractExisting methods for unsupervised domain adaptation often rely on minimizing some statistical distance between the source and target samples in the latent space. To avoid the sampling variability, class imbalance, and data-privacy concerns that often plague these methods, we instead provide a memory and computation-efficient probabilistic framework to extract class prototypes and align the target features with them. We demonstrate the general applicability of our method on a wide range of scenarios, including single-source, multi-source, class-imbalance, and source-private domain adaptation. Requiring no additional model parameters and having a moderate increase in computation over the source model alone, the proposed method achieves competitive performance with state-of-the-art methods. Korawat Tanwisuth, Xinjie Fan, Huangjie Zheng, Shujian Zhang, Hao Zhang 0050, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 6 |
| 2021 | Dual-view Snapshot Compressive Imaging via Optical Flow Aided Recurrent Neural Network
Ruiying Lu, Bo Chen 0001, Guanliang Liu, Ziheng Cheng 0001, Xin Yuan 0002 |
Int. J. Comput. Vis. | 2 |
| 2021 | Deep Autoencoding Topic Model With Scalable Hybrid Bayesian InferenceabstractTo build a flexible and interpretable model for document analysis, we develop deep autoencoding topic model (DATM) that uses a hierarchy of gamma distributions to construct its multi-stochastic-layer generative network. In order to provide scalable posterior inference for the parameters of the generative network, we develop topic-layer-adaptive stochastic gradient Riemannian MCMC that jointly learns simplex-constrained global parameters across all layers and topics, with topic and layer specific learning rates. Given a posterior sample of the global parameters, in order to efficiently infer the local latent representations of a document under DATM across all stochastic layers, we propose a Weibull upward-downward variational encoder that deterministically propagates information upward via a deep neural network, followed by a Weibull distribution based stochastic downward generative model. To jointly model documents and their associated labels, we further propose supervised DATM that enhances the discriminative power of its latent representations. The efficacy and scalability of our models are demonstrated on both unsupervised and supervised learning tasks on big corpora. Hao Zhang 0050, Bo Chen 0001, Yulai Cong, Dandan Guo, Hongwei Liu 0001, Mingyuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Bidirectional recurrent gamma belief network for HRRP target recognition
Bo Chen 0001, Xiaojun Peng, Haoyang Fan, Fangxu Yu, Hongwei Liu 0001 |
Signal Process. | 2 |
| 2021 | Region-factorized recurrent attentional network with deep clustering for radar HRRP target recognition
Chuan Du, Bo Chen 0001, Lei Zhang 0019, Hongwei Liu 0001 |
Signal Process. | 3 |
| 2021 | Domain-aware meta network for radar HRRP target recognition with missing aspects
Bo Chen 0001, Yishi Xu, Hongwei Liu 0001 |
Signal Process. | 2 |
| 2021 | Generalized zero-shot classification via iteratively generating and selecting unseen samples
Xiao Li 0008, Bo Chen 0001 |
Signal Process. Image Commun. | 3 |
| 2020 | Learning Dynamic Hierarchical Topic Graph with Graph Convolutional Network for Document ClassificationabstractConstructing a graph with graph convolutional network (GCN) to explore the relational structure of the data has attracted lots of interests in various tasks. However, for document classification, existing graph based methods often focus on the straightforward word-word and word-document relations, ignoring the hierarchical semantics. Besides, the graph construction is often independent from the task-specific GCN learning. To address these constrains, we integrate a probabilistic deep topic model into graph construction, and propose a novel trainable hierarchical topic graph (HTG), including word-level, hierarchical topic-level and document-level nodes, exhibiting semantic variation from fine-grained to coarse. Regarding the document classification as a document-node label generation task, HTG can be dynamically evolved with GCN by performing variational inference, which leads to an end-to-end document classification method, named dynamic HTG (DHTG). Besides achieving state-of-the-art classification results, our model learns an interpretable document graph with meaningful node embeddings and semantic edges. Zhengjue Wang, Chaojie Wang 0001, Hao Zhang 0050, Zhibin Duan, Mingyuan Zhou, Bo Chen 0001 |
AISTATS | 6 |
| 2020 | BIRNAT: Bidirectional Recurrent Neural Networks with Adversarial Training for Video Snapshot Compressive Imaging
Ziheng Cheng 0001, Ruiying Lu, Zhengjue Wang, Hao Zhang 0050, Bo Chen 0001, Ziyi Meng 0001, Xin Yuan 0002 |
ECCV (24) | 5 |
| 2020 | Friendly Topic Assistant for Transformer Based Abstractive Summarizationabstractive document summarization is a comprehensive task including document understanding and summary generation, in which area Transformer-based models have achieved the state-of-the-art performance. Compared with Transformers, topic models are better at learning explicit document semantics, and hence could be integrated into Transformers to further boost their performance. To this end, we rearrange and explore the semantics learned by a topic model, and then propose a topic assistant (TA) including three modules. TA is compatible with various Transformer-based models and user-friendly since i) TA is a plug-and-play model that does not break any structure of the original Transformer network, making users easily fine-tune Transformer+TA based on a well pre-trained model; ii) TA only introduces a small number of extra parameters. Experimental results on three datasets demonstrate that TA is able to improve the performance of several Transformer-based models. Zhengjue Wang, Zhibin Duan, Hao Zhang 0050, Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou |
EMNLP (1) | 6 |
| 2020 | Variational Hetero-Encoder Randomized GANs for Joint Image-Text Modeling
Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Mingyuan Zhou |
ICLR | 2 |
| 2020 | Recurrent Hierarchical Topic-Guided RNN for Language GenerationabstractTo simultaneously capture syntax and global semantics from a text corpus, we propose a new larger-context recurrent neural network (RNN) based language model, which extracts recurrent hierarchical semantic structure via a dynamic deep topic model to guide natural language generation. Moving beyond a conventional RNN-based language model that ignores long-range word dependencies and sentence order, the proposed model captures not only intra-sentence word dependencies, but also temporal transitions between sentences and inter-sentence topic dependencies. For inference, we develop a hybrid of stochastic-gradient Markov chain Monte Carlo and recurrent autoencoding variational Bayes. Experimental results on a variety of real-world text corpora demonstrate that the proposed model not only outperforms larger-context RNN-based language models, but also learns interpretable recurrent multilayer topics and generates diverse sentences and paragraphs that are syntactically correct and semantically coherent. Dandan Guo, Bo Chen 0001, Ruiying Lu, Mingyuan Zhou |
ICML | 2 |
| 2020 | Meta Network for Radar HRRP Noncooperative Target Recognition with Missing AspectsabstractWe propose a meta network (MNet) for the problem of target-aspect missing in radar high-resolution range profile (HRRP)-based noncooperative target recognition, where a classifier must be generalized to new aspects not seen in the training set, given only a small number of HRRP data of each new aspect. The MNet is a time domain convolutional neural network (TCNN) that is built based upon recent progress in meta-learning. In effect, it learns a model that is easy and fast to fine-tune, allowing the adaptation to happen in the right space for fast learning. Besides, we construct a new controllable HRRP dataset suitable for the scenario of noncooperative target-aspect missing using electromagnetic simulation. Compared with the traditional methods, the MNet is more efficient and could achieve better performance. Extensive experiments on the simulated HRRP dataset are conducted to illustrate the effectiveness of the proposed method. Bo Chen 0001, Chuan Du, Hongwei Liu 0001 |
IGARSS | 2 |
| 2020 | Switching Poisson Gamma Dynamical SystemsabstractWe propose Switching Poisson gamma dynamical systems (SPGDS) to model sequentially observed multivariate count data. Different from previous models, SPGDS assigns its latent variables into mixture of gamma distributed parameters to model complex sequences and describe the nonlinear dynamics, meanwhile, capture various temporal dependencies. For efficient inference, we develop a scalable hybrid stochastic gradient-MCMC and switching recurrent autoencoding variational inference, which is scalable to large scale sequences and fast in out-of-sample prediction. Experiments on both unsupervised and supervised tasks demonstrate that the proposed model not only has excellent fitting and prediction performance on complex dynamic sequences, but also separates different dynamical patterns within them. Bo Chen 0001, Qianru Zhao, Mingyuan Zhou |
IJCAI | 2 |
| 2020 | Bidirectional Convolutional Poisson Gamma Dynamical SystemsabstractIncorporating the natural document-sentence-word structure into hierarchical Bayesian modeling, we propose convolutional Poisson gamma dynamical systems (PGDS) that introduce not only word-level probabilistic convolutions, but also sentence-level stochastic temporal transitions. With word-level convolutions capturing phrase-level topics and sentence-level transitions capturing how the topic usages evolve over consecutive sentences, we aggregate the topic proportions of all sentences of a document as its feature representation. To consider not only forward but also backward sentence-level information transmissions, we further develop a bidirectional convolutional PGDS to incorporate the full contextual information to represent each sentence. For efficient inference, we construct a convolutional-recurrent inference network, which provides both sentence-level and document-level representations, and introduce a hybrid Bayesian inference scheme combining stochastic-gradient MCMC and amortized variational inference. Experimental results on a variety of document corpora demonstrate that the proposed models can extract expressive multi-level latent representations, including interpretable phrase-level topics and sentence-level temporal transitions as well as discriminative document-level features, achieving state-of-the-art document categorization performance while being memory and computation efficient. Chaojie Wang 0001, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou |
NeurIPS | 3 |
| 2020 | Bayesian Attention ModulesabstractAttention modules, as simple and effective tools, have not only enabled deep neural networks to achieve state-of-the-art results in many domains, but also enhanced their interpretability. Most current models use deterministic attention modules due to their simplicity and ease of optimization. Stochastic counterparts, on the other hand, are less popular despite their potential benefits. The main reason is that stochastic attention often introduces optimization issues or requires significant model changes. In this paper, we propose a scalable stochastic version of attention that is easy to implement and optimize. We construct simplex-constrained attention distributions by normalizing reparameterizable distributions, making the training process differentiable. We learn their parameters in a Bayesian framework where a data-dependent prior is introduced for regularization. We apply the proposed stochastic attention modules to various attention-based models, with applications to graph node classification, visual question answering, image captioning, machine translation, and language understanding. Our experiments show the proposed method brings consistent improvements over the corresponding baselines. Xinjie Fan, Shujian Zhang, Bo Chen 0001, Mingyuan Zhou |
NeurIPS | 3 |
| 2020 | Deep Relational Topic Modeling via Graph Poisson Gamma Belief NetworkabstractTo analyze a collection of interconnected documents, relational topic models (RTMs) have been developed to describe both the link structure and document content, exploring their underlying relationships via a single-layer latent representation with limited expressive capability. To better utilize the document network, we first propose graph Poisson factor analysis (GPFA) that constructs a probabilistic model for interconnected documents and also provides closed-form Gibbs sampling update equations, moving beyond sophisticated approximate assumptions of existing RTMs. Extending GPFA, we develop a novel hierarchical RTM named graph Poisson gamma belief network (GPGBN), and further introduce two different Weibull distribution based variational graph auto-encoders for efficient model inference and effective network information aggregation. Experimental results demonstrate that our models extract high-quality hierarchical latent document representations, leading to improved performance over baselines on various graph analytic tasks. Chaojie Wang 0001, Hao Zhang 0050, Bo Chen 0001, Dongsheng Wang 0003, Zhengjue Wang, Mingyuan Zhou |
NeurIPS | 3 |
| 2020 | Packet-Based Intrusion Detection Using Bayesian Topic Models in Mobile Edge ComputingabstractIn this paper, a network intrusion detection system is proposed using Bayesian topic model latent Dirichlet allocation (LDA) for mobile edge computing (MEC). The method employs tcpdump packets and extracts multiple features from the packet headers. The tcpdump packets are transferred into documents based on the features. A topic model is trained using only attack-free traffic in order to learn the behavior patterns of normal traffic. Then, the test traffic is analyzed against the learned behavior patterns to measure the extent to which the test traffic resembles the normal traffic. A threshold is defined in the training phase as the minimum likelihood of a host. In the test phase, when a host’s test traffic has a likelihood lower than the host’s threshold, the traffic is labeled as an intrusion. The intrusion detection system is validated using DARPA 1999 dataset. Experiment shows that our method is suitable to protect the security of MEC. Xuefei Cao, Bo Chen 0001 |
Secur. Commun. Networks | 3 |
| 2020 | RAFnet: Recurrent attention fusion network of hyperspectral and multispectral images
Ruiying Lu, Bo Chen 0001, Ziheng Cheng 0001 |
Signal Process. | 2 |
| 2020 | FusionNet: An Unsupervised Convolutional Variational Network for Hyperspectral and Multispectral Image FusionabstractDue to hardware limitations of the imaging sensors, it is challenging to acquire images of high resolution in both spatial and spectral domains. Fusing a low-resolution hyperspectral image (LR-HSI) and a high-resolution multispectral image (HR-MSI) to obtain an HR-HSI in an unsupervised manner has drawn considerable attention. Though effective, most existing fusion methods are limited due to the use of linear parametric modeling for the spectral mixture process, and even the deep learning-based methods only focus on deterministic fully-connected networks without exploiting the spatial correlation and local spectral structures of the images. In this paper, we propose a novel variational probabilistic autoencoder framework implemented by convolutional neural networks, in order to fuse the spatial and spectral information contained in the LR-HSI and HR-MSI, called FusionNet. The FusionNet consists of a spectral generative network, a spatial-dependent prior network, and a spatial-spectral variational inference network, which are jointly optimized in an unsupervised manner, leading to an end-to-end fusion system. Further, for fast adaptation to different observation scenes, we give a meta-learning explanation to the fusion problem, and combine the FusionNet with meta-learning in a synergistic manner. Effectiveness and efficiency of the proposed method are evaluated based on several publicly available datasets, demonstrating that the proposed FusionNet outperforms the state-of-the-art fusion methods. Zhengjue Wang, Bo Chen 0001, Ruiying Lu, Hao Zhang 0050, Hongwei Liu 0001, Pramod K. Varshney |
IEEE Trans. Image Process. | 2 |
| 2019 | Convolutional Poisson Gamma Belief NetworkabstractFor text analysis, one often resorts to a lossy representation that either completely ignores word order or embeds each word as a low-dimensional dense feature vector. In this paper, we propose convolutional Poisson factor analysis (CPFA) that directly operates on a lossless representation that processes the words in each document as a sequence of high-dimensional one-hot vectors. To boost its performance, we further propose the convolutional Poisson gamma belief network (CPGBN) that couples CPFA with the gamma belief network via a novel probabilistic pooling layer. CPFA forms words into phrases and captures very specific phrase-level topics, and CPGBN further builds a hierarchy of increasingly more general phrase-level topics. For efficient inference, we develop both a Gibbs sampler and a Weibull distribution based convolutional variational auto-encoder. Experimental results demonstrate that CPGBN can extract high-quality text latent representations that capture the word order information, and hence can be leveraged as a building block to enrich a wide variety of existing latent variable models that ignore word order. Chaojie Wang 0001, Bo Chen 0001, Sucheng Xiao, Mingyuan Zhou |
ICML | 2 |
| 2019 | Factorized discriminative conditional variational auto-encoder for radar HRRP target recognition
Chuan Du, Bo Chen 0001, Dandan Guo, Hongwei Liu 0001 |
Signal Process. | 2 |
| 2019 | Long short-term memory-based recurrent neural networks for nonlinear target tracking
Chang Gao 0004, Junkun Yan, Shenghua Zhou, Bo Chen 0001, Hongwei Liu 0001 |
Signal Process. | 4 |
| 2019 | Variational probabilistic generative framework for single image super-resolution
Zhengjue Wang, Bo Chen 0001, Hao Zhang 0050, Hongwei Liu 0001 |
Signal Process. | 2 |
| 2019 | Target-Aware Recurrent Attentional Network for Radar HRRP Target Recognition
Bo Chen 0001, Jinwei Wan, Hongwei Liu 0001, Lin Jin |
Signal Process. | 2 |
| 2019 | Deep Max-Margin Discriminant ProjectionabstractIn this paper, a unified Bayesian max-margin discriminant projection framework is proposed, which is able to jointly learn the discriminant feature space and the max-margin classifier with different relationships between the latent representations and observations. We assume that the latent representation follows a normal distribution whose sufficient statistics are functions of the observations. The function can be flexibly realized through either shallow or deep structures. The shallow structure includes linear, nonlinear kernel-based functions, and even the convolutional projection, which can be further trained layerwisely to build a multilayered convolutional feature learning model. To take the advantage of the deep neural networks, especially their highly expressive ability and efficient parameter learning, we integrate Bayesian modeling and the popular neural networks, for example, mltilayer perceptron and convolutional neural network, to build an end-to-end Bayesian deep discriminant projection under the proposed framework, which degenerated into the existing shallow linear or convolutional projection with the single-layer structure. Moreover, efficient scalable inferences for the realizations with different functions are derived to handle large-scale data via a stochastic gradient Markov chain Monte Carlo. Finally, we demonstrate the effectiveness and efficiency of the proposed models by the experiments on real-world data, including four image benchmarks (MNIST, CIFAR-10, STL-10, and SVHN) and one measured radar high-resolution range profile dataset, with the detailed analysis about the parameters and computational complexity. Hao Zhang 0050, Bo Chen 0001, Zhengjue Wang, Hongwei Liu 0001 |
IEEE Trans. Cybern. | 2 |
| 2019 | Noise-Robust Motion Compensation for Aerial Maneuvering Target ISAR Imaging by Parametric Minimum Entropy OptimizationabstractWhen a target is involved in maneuvering motion, the nonuniform 3-D rotation motion will cause a continuous change of image projection plane (IPP), which would induce 2-D spatial-variant phase errors. In this case, the inverse synthetic aperture (ISAR) image would be seriously blurred when using the traditional compensation methods. On the other hand, strong noise has been always challenging the conventional methods in motion parameters estimation and phase error compensation. In this paper, we propose a noise-robust compensation method to compensate the 2-D spatial-variant phase errors of the maneuvering target via using tracking information and parametric minimum entropy optimization. First, the maneuvering signal model is developed based on a 2-D spatial-variant model and a 3-D rotation motion model. Based on the developed signal model, a parametric entropy minimum optimization is established to estimate the rotation motion parameters. A gradient-based solver of this optimization is then adopted to iteratively find the global optimum. Meanwhile, in order to increase the robustness of this optimization under low SNR, an extended Kalman filter is adopted here for coarse motion estimation via using tracking information. By treating these estimated motion parameters as initial values, we can effectively prevent this optimization from trapping into a local optimum. Finally, the 2-D spatial-variant phase error can be iteratively compensated, and a well-focused ISAR image can be obtained. The proposed method has three main contributions: 1) it is applicable in the case of changing IPP; 2) it gives the exact expression of chip parameters; and 3) it can efficiently compensate the 2-D spatial-variant phase errors under low SNR. Experiments based on the simulated data and the real measured data prove the effectiveness and robustness of the proposed method. Lei Zhang 0019, Lan Du 0001, Dongwen Yang, Bo Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Multimodal Poisson Gamma Belief NetworkabstractTo learn a deep generative model of multimodal data, we propose a multimodal Poisson gamma belief network (mPGBN) that tightly couple the data of different modalities at multiple hidden layers. The mPGBN unsupervisedly extracts a nonnegative latent representation using an upward-downward Gibbs sampler. It imposes sparse connections between different layers, making it simple to visualize the generative process and the relationships between the latent features of different modalities. Our experimental results on bi-modal data consisting of images and tags show that the mPGBN can easily impute a missing modality and hence is useful for both image annotation and retrieval. We further demonstrate that the mPGBN achieves state-of-the-art results on unsupervisedly extracting latent features from multimodal data. Chaojie Wang 0001, Bo Chen 0001, Mingyuan Zhou |
AAAI | 2 |
| 2018 | WHAI: Weibull Hybrid Autoencoding Inference for Deep Topic Modeling
Hao Zhang 0050, Bo Chen 0001, Dandan Guo, Mingyuan Zhou |
ICLR (Poster) | 2 |
| 2018 | Deep Poisson gamma dynamical systemsabstractWe develop deep Poisson-gamma dynamical systems (DPGDS) to model sequentially observed multivariate count data, improving previously proposed models by not only mining deep hierarchical latent structure from the data, but also capturing both first-order and long-range temporal dependencies. Using sophisticated but simple-to-implement data augmentation techniques, we derived closed-form Gibbs sampling update equations by first backward and upward propagating auxiliary latent counts, and then forward and downward sampling latent variables. Moreover, we develop stochastic gradient MCMC inference that is scalable to very long multivariate count time series. Experiments on both synthetic and a variety of real-world data demonstrate that the proposed model not only has excellent predictive performance, but also provides highly interpretable multilayer latent structure to represent hierarchical and temporal information propagation. Dandan Guo, Bo Chen 0001, Hao Zhang 0050, Mingyuan Zhou |
NeurIPS | 2 |
| 2017 | Deep Latent Dirichlet Allocation with Topic-Layer-Adaptive Stochastic Gradient Riemannian MCMCabstractIt is challenging to develop stochastic gradient based scalable inference for deep discrete latent variable models (LVMs), due to the difficulties in not only computing the gradients, but also adapting the step sizes to different latent factors and hidden layers. For the Poisson gamma belief network (PGBN), a recently proposed deep discrete LVM, we derive an alternative representation that is referred to as deep latent Dirichlet allocation (DLDA). Exploiting data augmentation and marginalization techniques, we derive a block-diagonal Fisher information matrix and its inverse for the simplex-constrained global model parameters of DLDA. Exploiting that Fisher information matrix with stochastic gradient MCMC, we present topic-layer-adaptive stochastic gradient Riemannian (TLASGR) MCMC that jointly learns simplex-constrained global parameters across all layers and topics, with topic and layer specific learning rates. State-of-the-art results are demonstrated on big data sets. Yulai Cong, Bo Chen 0001, Hongwei Liu 0001, Mingyuan Zhou |
ICML | 2 |
| 2017 | Multiple-Instance feature extraction at the bag and instance levels using the maximum trace-difference criterion
Jing Chai, Bo Chen 0001, Xinghao Ding |
Inf. Sci. | 2 |
| 2017 | Radar HRRP target recognition with deep networks
Bo Feng 0012, Bo Chen 0001, Hongwei Liu 0001 |
Pattern Recognit. | 2 |
| 2016 | Performance analysis of a modified Rao test for adaptive subspace detectionabstractThe problem of detecting a subspace signal is studied in colored Gaussian noise with an unknown covariance matrix. In the subspace model, the target signal belongs to a known subspace, but with unknown coordinates. We propose a modified Rao test (MRT) by introducing a tunable parameter. The MRT is more general, which includes the Rao test and the generalized likelihood ratio test as special cases. Moreover, closed-form expressions for the probabilities of false alarm and detection of the MRT are derived. Numerical results demonstrate that the MRT can offer the flexibility of being adjustable in the mismatched case where the target signal deviates from the presumed signal subspace. In particular, the MRT provides better mismatch rejection capacities as the tunable parameter increases. Jun Liu 0004, Bo Chen 0001, Hongwei Liu 0001, Weijian Liu 0001 |
ICASSP | 2 |
| 2016 | Multiple scattering effects on the localization of two point scatterersabstractMultiple scattering effects are commonly ignored in the detection and estimation of scatterers in signal processing research, because the energy of the first-order scattering is much larger than that of higher-order components. Although multiple scattering can significantly increase the estimation precision of point scatterers, it does not always lead to an improvement. Identifying conditions under which multiple scattering is beneficial or detrimental to estimation in a general setup is still an open problem. In this paper, we consider the effects of multiple scattering on the localization of two point scatterers. By comparing the Fisher information matrix on location parameters when multiple scattering exists and does not exist, we show analytically that information on ranges can benefit estimating directions of arrival via multiple scattering when the two scatterers are in far-field and well resolved. Arye Nehorai, Hongwei Liu 0001, Bo Chen 0001, Yuehai Wang |
ICASSP | 4 |
| 2016 | Augmentable Gamma Belief NetworksabstractTo infer multilayer deep representations of high-dimensional discrete and nonnegative real vectors, we propose an augmentable gamma belief network (GBN) that factorizes each of its hidden layers into the product of a sparse connection weight matrix and the nonnegative real hidden units of the next layer. The GBN's hidden layers are jointly trained with an upward-downward Gibbs sampler that solves each layer with the same subroutine. The gamma-negative binomial process combined with a layer-wise training strategy allows inferring the width of each layer given a fixed budget on the width of the first layer. Example results illustrate interesting relationships between the width of the first layer and the inferred network structure, and demonstrate that the GBN can add more layers to improve its performance in both unsupervisedly extracting features and predicting heldout data. For exploratory data analysis, we extract trees and subnetworks from the learned deep network to visualize how the very specific factors discovered at the first hidden layer and the increasingly more general factors discovered at deeper hidden layers are related to each other, and we generate synthetic data by propagating random variables through the deep network from the top hidden layer back to the bottom data layer. Mingyuan Zhou, Yulai Cong, Bo Chen 0001 |
J. Mach. Learn. Res. | 3 |
| 2016 | Convolutional Neural Network With Data Augmentation for SAR Target RecognitionabstractMany methods have been proposed to improve the performance of synthetic aperture radar (SAR) target recognition but seldom consider the issues in real-world recognition systems, such as the invariance under target translation, the invariance under speckle variation in different observations, and the tolerance of pose missing in training data. In this letter, we investigate the capability of a deep convolutional neural network (CNN) combined with three types of data augmentation operations in SAR target recognition. Experimental results demonstrate the effectiveness and efficiency of the proposed method. The best performance is obtained by using the CNN trained by all types of augmentation operations, showing that it is a practical approach for target recognition in challenging conditions of target translation, random speckle noise, and missing pose. Bo Chen 0001, Hongwei Liu 0001, Mengyuan Huang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Infinite max-margin factor analysis via data augmentation
Xuefeng Zhang 0003, Bo Chen 0001, Hongwei Liu 0001, Bo Feng 0012 |
Pattern Recognit. | 2 |
| 2016 | Performance of the SMI beamformer with signal steering vector errors in heterogeneous environments
Jun Liu 0004, Weijian Liu 0001, Hongwei Liu 0001, Zi-Jing Zhang, Bo Chen 0001 |
Signal Process. | 5 |
| 2015 | The Poisson Gamma Belief NetworkabstractTo infer a multilayer representation of high-dimensional count vectors, we propose the Poisson gamma belief network (PGBN) that factorizes each of its layers into the product of a connection weight matrix and the nonnegative real hidden units of the next layer. The PGBN's hidden layers are jointly trained with an upward-downward Gibbs sampler, each iteration of which upward samples Dirichlet distributed connection weight vectors starting from the first layer (bottom data layer), and then downward samples gamma distributed hidden units starting from the top hidden layer. The gamma-negative binomial process combined with a layer-wise training strategy allows the PGBN to infer the width of each layer given a fixed budget on the width of the first layer. The PGBN with a single hidden layer reduces to Poisson factor analysis. Example results on text analysis illustrate interesting relationships between the width of the first layer and the inferred network structure, and demonstrate that the PGBN, whose hidden units are imposed with correlated gamma priors, can add more layers to increase its performance gains over Poisson factor analysis, given the same limit on the width of the first layer. Mingyuan Zhou, Yulai Cong, Bo Chen 0001 |
NIPS | 3 |
| 2015 | A Three-Component Fisher-Based Feature Weighting Method for Supervised PolSAR Image ClassificationabstractThis letter presents a feature weighting method for polarimetric synthetic aperture radar (PolSAR) image classification. Appropriate feature weighting is essential for obtaining accurate classifications but so far has remained an open research problem. We propose in this letter a supervised three-component feature weighting method based on the Fisher linear discriminant. Fisher linear discriminant method is used to calculate a coefficient for each feature. Then, these coefficients are modified according to a three-component scattering power decomposition model, combining both physical and statistical scattering characteristics to adapt them for the particular scattering mechanisms inherent in PolSAR data and assigned to the coherency matrix to enhance the discriminating ability of the features. Freeman decomposition and Wishart classifier are used to classify the PolSAR image. The effectiveness of the proposed method is demonstrated by experiments NASA/JPL AIRSAR L-band and CSA Radarsat-2 C-band PolSAR images of the San Francisco area. Bo Chen 0001, Shuang Wang 0001, Licheng Jiao, Rustam Stolkin, Hongying Liu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | Weighted classifier ensemble based on quadratic form
Shasha Mao, Licheng Jiao, Shuiping Gou, Bo Chen 0001, Sai-Kit Yeung |
Pattern Recognit. | 5 |
| 2015 | Detection Probability of a CFAR Matched Filter with Signal Steering Vector ErrorsabstractOur aim in this work is to analyze the detection performance of a constant false alarm rata matched filter (CFAR-MF) which was developed for the detection problem in white Gaussian noise with unknown noise power. An exact expression for the detection probability of the CFAR-MF is derived in the mismatched case where mismatch exists between the actual signal steering vector and the nominal one. This theoretical expression can be used to facilitate the performance evaluation of the CFAR-MF in real-world scenarios when signal mismatch cannot be neglected. Jun Liu 0004, Weijian Liu 0001, Bo Chen 0001, Hongwei Liu 0001, Hongbin Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2015 | Max-Margin Discriminant Projection via Data AugmentationabstractIn this paper, we introduce a new max-margin discriminant projection method, which takes advantage of the latent variable representation for support vector machine (SVM) as the classification criterion. Specifically, the proposed model jointly learns the discriminative subspace and classifier in a Bayesian framework by conditioning on augmented variables. Moreover, an extended nonlinear model is developed based on the kernel trick, where the similar model can be used in this setting with few modifications. To explore the sparsity in the kernel expansion, we use the spike-and-slab prior to seek basis vectors (BVs) from the corresponding candidates. Unlike existing methods, which employ BVs to approximate the original feature space, in our method BVs are sought to associate the final classification task. Thanks to the conditionally conjugate property, the parameters in our models can be inferred via the simple and efficient Gibbs sampler. Finally, we test our methods on synthesized and real-world data, including large-scale data sets to demonstrate their efficiency and effectiveness. Bo Chen 0001, Hao Zhang 0050, Xuefeng Zhang 0003, Hongwei Liu 0001, Jun Liu 0004 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Classification Method for Fully PolSAR Data Based on Three Novel ParametersabstractIn this letter, a new classification method for fully polarimetric synthetic aperture radar (PolSAR) data based on three novel parameters is presented. The three parameters are derived from the eigenspace of the coherency matrix as linear combinations of its three eigenvalues. In the proposed classification method, the maximum value out of the three parameters is determined to assign a label to each image pixel, and the PolSAR image is classified into three classes accordingly. Experimental results based on NASA/JPL AIRSAR L-band data and CSA RADARSAT-2 C-band data illustrate the validity and efficacy of the procedure. Shuang Wang 0001, Bo Chen 0001, Shasha Mao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Deep Learning with Hierarchical Convolutional Factor AnalysisabstractUnsupervised multilayered (“deep”) models are considered for imagery. The model is represented using a hierarchical convolutional factor-analysis construction, with sparse factor loadings and scores. The computation of layer-dependent model parameters is implemented within a Bayesian setting, employing a Gibbs sampler and variational Bayesian (VB) analysis that explicitly exploit the convolutional nature of the expansion. To address large-scale and streaming data, an online version of VB is also developed. The number of dictionary elements at each layer is inferred from the data, based on a beta-Bernoulli implementation of the Indian buffet process. Example results are presented for several image-processing applications, with comparisons to related models in the literature. Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David M. Blei, David B. Dunson, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | A Large Scene Deceptive Jamming Method for Space-Borne SARabstractBased on the synthetic aperture radar (SAR) geometric model, a novel, fast algorithm of large scene deceptive jamming against the space-borne SAR is proposed. First, we divide the jamming scene template into sub-templates according to the depth of focus in the range dimension. Next, each sub-template is decomposed into the slow-time-dependent and slow-time-independent terms in the range frequency-azimuth time domain. The slow-time-independent terms are generated off-line while the slow-time-dependent terms are generated by real-time 1-D frequency modulation. Then, the sub-templates are convolved with the intercepted SAR signals simultaneously. Finally, fast deceptive jamming is achieved by incorporating all the sub-templates together. In the proposed method, the two-step realization of the sub-templates and the parallel sub-block processing improves the algorithm efficiency. The simulation results prove the validity of the proposed algorithm. Feng Zhou 0001, Bo Zhao 0006, Mingliang Tao, Xueru Bai, Bo Chen 0001, Guangcai Sun |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2011 | The Hierarchical Beta Process for Convolutional Factor Analysis and Deep Learning
Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David B. Dunson, Lawrence Carin |
ICML | 1 |
| 2011 | Unsupervised classification of POLSAR data based on the polarimetric decomposition and the co-polarization ratioabstractIn this paper, a new classification scheme for polarimetric SAR data sets is presented. The proposed method mainly involves two concepts: Freeman-Durden decomposition and co-polarization ratio. The core concept of Freeman-Durden decomposition is to decompose the covariance matrix into three scattering mechanisms: surface scattering, double bounce scattering and volume scattering; the co-polarization radio represents the proportion of horizontal polarization and vertical polarization. The combination of these two characters can distinguish different vegetation types effectively. The proposed method mainly consists of three steps: first, apply Freeman-Durden decomposition to get the scattering characters; second, combine the scattering powers and co-polarization radio to divide the images into corresponding initial clusters; finally, improve the representation of each class, the data sets of which are classified by an iterative algorithm based on a complex Wishart density function. The effectiveness of this algorithm is demonstrated by two sets of data: NASA/JPL AIRSAR L-band data of San Francisco and Flevoland in The Netherlands by NASA/JPL AIRSAR in 1989. Shuang Wang 0001, Jingjing Pei, Kun Liu 0011, Bo Chen 0001 |
IGARSS | 5 |
| 2011 | On the Analysis of Multi-Channel Neural Spike DataabstractNonparametric Bayesian methods are developed for analysis of multi-channel spike-train data, with the feature learning and spike sorting performed jointly. The feature learning and sorting are performed simultaneously across all channels. Dictionary learning is implemented via the beta-Bernoulli process, with spike sorting performed via the dynamic hierarchical Dirichlet process (dHDP), with these two models coupled. The dHDP is augmented to eliminate refractoryperiod violations, it allows the “appearance” and “disappearance” of neurons over time, and it models smooth variation in the spike statistics. Bo Chen 0001, David E. Carlson, Lawrence Carin |
NIPS | 1 |
| 2010 | Sparse linear regression with beta process priorsabstractA Bayesian approximation to finding the minimum ℓ0norm solution for an underdetermined linear system is proposed that is based on the beta process prior. The beta process linear regression (BP-LR) model finds sparse solutions to the underdetermined model y = Φx + ϵ, by modeling the vector x as an element-wise product of a non-sparse weight vector, w, and a sparse binary vector, z, that is drawn from the beta process prior. The hierarchical model is fully conjugate and therefore is amenable to fast inference methods. We demonstrate the model on a compressive sensing problem and on a correlated-feature problem, where we show the ability of the BP-LR to selectively remove the irrelevant features, while preserving the relevant groups of correlated features. Bo Chen 0001, John W. Paisley, Lawrence Carin |
ICASSP | 1 |
| 2010 | Bayesian Inference of the Number of Factors in Gene-Expression Analysis: Application to Human Virus Challenge StudiesabstractBACKGROUND: Nonparametric Bayesian techniques have been developed recently to extend the sophistication of factor models, allowing one to infer the number of appropriate factors from the observed data. We consider such techniques for sparse factor analysis, with application to gene-expression data from three virus challenge studies. Particular attention is placed on employing the Beta Process (BP), the Indian Buffet Process (IBP), and related sparseness-promoting techniques to infer a proper number of factors. The posterior density function on the model parameters is computed using Gibbs sampling and variational Bayesian (VB) analysis. RESULTS: Time-evolving gene-expression data are considered for respiratory syncytial virus (RSV), Rhino virus, and influenza, using blood samples from healthy human subjects. These data were acquired in three challenge studies, each executed after receiving institutional review board (IRB) approval from Duke University. Comparisons are made between several alternative means of per-forming nonparametric factor analysis on these data, with comparisons as well to sparse-PCA and Penalized Matrix Decomposition (PMD), closely related non-Bayesian approaches. CONCLUSIONS: Applying the Beta Process to the factor scores, or to the singular values of a pseudo-SVD construction, the proposed algorithms infer the number of factors in gene-expression data. For real data the "true" number of factors is unknown; in our simulations we consider a range of noise variances, and the proposed Bayesian models inferred the number of factors accurately relative to other methods in the literature, such as sparse-PCA and PMD. We have also identified a "pan-viral" factor of importance for each of the three viruses considered in this study. We have identified a set of genes associated with this pan-viral factor, of interest for early detection of such viruses based upon the host response, as quantified via gene-expression data. Bo Chen 0001, Minhua Chen, John W. Paisley, Aimee K. Zaas, Christopher W. Woods, Geoffrey S. Ginsburg, Alfred O. Hero III, Joseph E. Lucas, David B. Dunson, Lawrence Carin |
BMC Bioinform. | 1 |
| 2010 | Large margin nearest local mean classifier
Jing Chai, Hongwei Liu 0001, Bo Chen 0001, Zheng Bao 0001 |
Signal Process. | 3 |
| 2009 | Large Margin Feature Weighting Method via Linear ProgrammingabstractThe problem of feature selection is a difficult combinatorial task in machine learning and of high practical relevance. In this paper, we consider feature selection method for multimodally distributed data, and present a large margin feature weighting method for k-nearest neighbor (kNN) classifiers. The method learns the feature weighting factors by minimizing a cost function, which aims at separating different classes by large local margins and pulling closer together points from the same class, based on using as few features as possible. The consequent optimization problem can be efficiently solved by linear programming. Finally, the proposed approach is assessed through a series of experiments with UCI and microarray data sets, as well as a more specific and challenging task, namely, radar high-resolution range profiles (HRRP) automatic target recognition (ATR). The experimental results demonstrate the effectiveness of the proposed algorithms. Bo Chen 0001, Hongwei Liu 0001, Jing Chai, Zheng Bao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | A kernel optimization method based on the localized kernel Fisher criterion
Bo Chen 0001, Hongwei Liu 0001, Zheng Bao 0001 |
Pattern Recognit. | 1 |
| 2008 | Optimizing the data-dependent kernel under a unified kernel optimization framework
Bo Chen 0001, Hongwei Liu 0001, Zheng Bao 0001 |
Pattern Recognit. | 1 |
| 2007 | Kernel subclass discriminant analysis
Bo Chen 0001, Hongwei Liu 0001, Zheng Bao 0001 |
Neurocomputing | 1 |
| 2006 | Speeding Up SVM in Test Phase: Application to Radar HRRP ATR
Bo Chen 0001, Hongwei Liu 0001, Zheng Bao 0001 |
ICONIP (1) | 1 |
| 2006 | A Kernel Optimization Method Based on the Localized Kernel Fisher Criterion
Bo Chen 0001, Hongwei Liu 0001, Zheng Bao 0001 |
ISNN (1) | 1 |