Wenbo Hu 0001

dblp:95/7076-1 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-0639-2012ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
abstract
Chain of Thought (CoT) reasoning has demonstrated remarkable deep reasoning capabilities in both large language models (LLMs) and multimodal large language models (MLLMs). However, its reliability is often undermined by the accumulation of errors in intermediate steps. This paper proposes a novel approach to calibrating CoT reasoning accuracy by leveraging the model’s internal cognition of truthfulness. Our findings suggest that the model implicitly tracks the evolving veracity of intermediate steps throughout the dynamic, progressive reasoning process. We train a confidence predictor to quantify the model’s internal cognition of truthfulness at each reasoning step, enabling dynamic selection of the most plausible reasoning path through beam search. Experimental results demonstrate that our method significantly outperforms the state-of-the-art baselines (e.g., Self-Consistency, and PRM Guided Search) across the mathematical, symbolic, and commonsense reasoning tasks, exhibiting superior accuracy and reliability in both unimodal and multimodal settings. This study proposes a novel path toward improving the reliability of CoT reasoning, demonstrating strong potential for wide-ranging applications.
Zijun Chen 0001, Wenbo Hu 0001, Richang Hong
AAAI2
2026 Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting
abstract
Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in time series remains a major challenge. Existing methods often adopt time-windows of varying sizes, which may introduce noisy or irrelevant representations when mismatched with the underlying temporal patterns, potentially leading to overfitting. In this paper, we propose Sparse-Scale Transformer (SSformer) with Bidirectional Awareness for Time Series Forecasting to enhance the multi-scale modeling for time series. Specifically, we propose a novel Sparse-Scale Convolution (SSC) block that imposes sparsity on scales to obtain the informative representations by evaluating the intra-scale segment similarity of time series, and utilizes scale-specific convolutions to extract local patterns. Furthermore, we design a Bidirectional-Scale Interaction (BSI) block to explicitly model scale correlations in both coarse-to-fine and fine-to-coarse directions. Finally, scale predictions are ensembled to fully exploit the complementary forecasting capabilities across scales. Extensive experiments on various real-world datasets demonstrate that SSformer achieves state-of-the-art performance with superior efficiency.
Ying Liu 0096, Bo Liu 0005, Sheng Huang 0001, Wenbo Hu 0001, Meng Wang 0001, Richang Hong
AAAI5
2026 Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding
abstract
Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy.
Youze Wang, Zijun Chen 0001, Shishen Gu, Wenbo Hu 0001, Yinpeng Dong, Hang Su 0006, Jun Zhu 0001, Meng Wang 0001, Richang Hong
AAAI5
2025 Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) combine visual and textual data for tasks like image captioning and visual question answering. Proper uncertainty calibration is crucial but challenging for reliable use in areas like healthcare and autonomous driving. This paper investigates several MLLMs, focusing on their calibration across various scenarios, including before and after visual fine-tuning as well as before and after multimodal training of the base LLMs. We observed miscalibration in their performance, and at the same time, no significant differences in calibration across these scenarios. We also highlight differences in uncertainty between text and the impact of the integration of these two types of information in uncertainty. To better understand MLLMs’ miscalibration and their ability to self-assess uncertainty, we developed the IDK (I don’t know) dataset, which is key for evaluating how they handle unknowns. Our findings reveal that MLLMs tend to give answers rather than admit uncertainty, but this self-assessment improves with prompt adjustments. Finally, to calibrate MLLMs and enhance model reliability, we propose techniques such as temperature scaling and iterative prompt optimization. Our results provide insights into improving MLLMs for effective and responsible deployment in multimodal applications.
Zijun Chen 0001, Wenbo Hu 0001, Guande He, Zhijie Deng, Zheng Zhang 0006, Richang Hong
COLING2
2025 SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information.However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak attacks.This paper first analyzes the impact of inserting safety reasoning prompt on various aspects of the model.We find that this external method can help the model resist jailbreak attacks to some extent, but the model still fails to distinguish specific semantic scenarios, resulting in a significantly increased refusal rate for benign queries.Inspired by this, we propose a novel training framework, SURE (Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models), designed to help models internalize chain-of-thought-based safety decision-making capabilities.Extensive experiments demonstrate that SURE significantly improves model safety while effectively avoiding over-defense, achieving a good balance between safety and generality.Finally, we create a large-scale multimodal safety reasoning dataset, MLLM-SCoT-Plus, to facilitate research on safety alignment in multimodal models.Our code and the dataset are publicly available at https://github.com/hfutml/SURE.Warning: This paper contains offensive and harmful examples.
Yuxin Gou, Xiaoning Dong, Shishen Gu, Richang Hong, Wenbo Hu 0001
EMNLP6
2025 Towards V2X HD Mapping for Autonomous Driving: A Concise Review
abstract
High-definition (HD) maps are fundamental components to autonomous driving systems, providing essential centimeter-level accuracy and lane-level semantic information. While traditional HD mapping methods have evolved into online learning approaches, current solutions face significant challenges due to sensor limitations and environmental constraints. This paper presents a systematic review of HD map construction methods, tracking their evolution from conventional techniques to advanced Vehicle-to-Everything (V2X) cooperative mapping enabled by edge computing and communication technologies. Through comprehensive analysis of methodologies, algorithms, and datasets, we identify critical challenges in current HD mapping systems. Our review encompasses three key domains: traditional mapping methods, online learning approaches, and V2X cooperative construction of HD maps. We evaluate existing solutions against standardized metrics, compare their effective-ness, and outline promising directions for future research. This work provides researchers and practitioners with a structured understanding of the HD mapping landscape and highlights opportunities for advancing autonomous driving systems.
Suhui Yang, Shengtong Xu, Xiangzeng Liu, Wenbo Hu 0001, Haoyi Xiong
IV6
2025 A Concise Survey on Lane Topology Reasoning for HD Mapping
abstract
Lane topology reasoning techniques play a crucial role in high-definition (HD) mapping and autonomous driving applications. While recent years have witnessed significant advances in this field, there has been limited effort to consolidate these works into a comprehensive overview. This survey systematically reviews the evolution and current state of lane topology reasoning methods, categorizing them into three major paradigms: procedural modeling-based methods, aerial imagery-based methods, and onboard sensors-based methods. We analyze the progression from early rule-based approaches to modern learning-based solutions utilizing transformers, graph neural networks (GNNs), and other deep learning architectures. The paper examines standardized evaluation metrics, including road-level measures (APLS and TLTS score), and lane-level metrics (DET and TOP score), along with performance comparisons on benchmark datasets such as OpenLane- V2. We identify key technical challenges, including dataset availability and model efficiency, and outline promising directions for future research. This comprehensive review provides researchers and practitioners with insights into the theoretical frameworks, practical implementations, and emerging trends in lane topology reasoning for HD mapping applications.
Shengtong Xu, Haoyi Xiong, Xiangzeng Liu, Wenbo Hu 0001, Wenbing Huang 0001
IV6
2025 Joint Adversarial Purification: Mitigating the Threat of Multimodal Adversarial Examples
abstract
Vision-language pre-training (VLP) models exhibit exceptional generalization capabilities across diverse vision-language (V+L) tasks. However, studies reveal their vulnerability to carefully crafted multimodal adversarial examples. Notably, state-of-the-art attacks like Co-Attack (white-box) and SGA (transfer-based) demonstrate alarming success rates, posing critical security threats to VLP models. Current defense mechanisms against such multimodal attacks remain insufficiently explored. To address this challenge, we propose Joint Adversarial Purification (JAP), a novel defense framework that synergistically eliminates adversarial perturbations across modalities through cross-modal interaction. Our approach harnesses cross-modal semantic synergy to jointly purify adversarial perturbations: Generative denoising establishes visual-semantic anchors through diffusion processes, while purified linguistic cues conversely enhance visual perturbation filtering, forming a self-reinforcing defense cycle. Extensive experiments demonstrate that JAP effectively mitigates adversarial threats from both white-box Co-Attack and transfer-based SGA, significantly outperforming existing unimodal defense baselines. This work establishes a new paradigm for securing VLP models against multimodal adversarial attacks.
Youze Wang, Wenbo Hu 0001, Richang Hong
ICMR3
2025 RATE: Robust Adversarial Training and Temperature-scaled Ensemble Framework for Trustworthy Misinformation Detection
abstract
Social media platforms have become major sources of misinformation, primarily manifesting as rumors and fake news. While existing detection methods demonstrate good accuracy, they still exhibit numerous issues related to confidence calibration. Existing methods often lack robustness due to insufficient consideration of adversarial attacks, which compromises their security and introduces greater uncertainty in predictions. Moreover, the predominant focus on improving model performance has led to the neglect of prediction risk, resulting in numerous overconfident and unconfident predictions. To address the above issues, we present a novel temperature-scaled ensemble frame with robust adversarial training(RATE) for calibrating misinformation detection models. We employed the fast gradient sign method for adversarial training to smooth the predictive distribution, defined an ensemble framework with independently trained backbone networks to reduce uncertainty, and applied temperature scaling during validation to align confidence outputs with the true distribution. Through extensive evaluation of context-based rumor detection and content-based fake news detection models across six diverse datasets, we reveal critical issues in the models' predictive confidence calibration. Experimental results demonstrate that RATE effectively addresses confidence bias while maintaining predictive accuracy, offering a universal solution for constructing reliable misinformation detection systems.
Wenbo Hu 0001, Qiang Liu 0006, Richang Hong
ICMR2
2025 Making Strides Security in Multimodal Fake News Detection Models: A Comprehensive Analysis of Adversarial Attacks
Jiahua Si, Youze Wang, Wenbo Hu 0001, Qiang Liu 0006, Richang Hong
MMM (2)3
2025 Deep sub-ensembles meets quantile regression: uncertainty-aware imputation for time series
Ying Liu 0096, Peng Cui 0007, Wenbo Hu 0001, Richang Hong
Mach. Learn.3
2025 Revealing Security Flaws in Cross-Modal Retrieval Models Through Video Poisoning
abstract
Video-text cross-modal retrieval is widely studied to improve retrieval accuracy. However, the security of video-text cross-modal retrieval models receives little attention. If attackers exploit the security vulnerabilities in these models, it poses a significant threat to the retrieval models. Thus, identifying security flaws in video-text cross-modal retrieval models becomes the focus of our research. We are the first to design a video poisoning model to uncover security vulnerabilities in retrieval models. Existing poisoning models have certain limitations when it comes to exploiting vulnerabilities in retrieval models. These include failing to comprehensively embed malicious information into the original video and being unable to maintain visual consistency between the original and poisoned videos. These limitations can result in unsuccessful attacks on retrieval models and an inability to effectively identify security flaws within them. To address these shortcomings, we design an efficient poisoning model that embeds malicious information thoroughly into the original clean data to attack video-text cross-modal retrieval models. We are the first to use a poisoning model to attack retrieval models, thereby uncovering their security vulnerabilities. Second, we introduce a bi-level poisoning module to ensure that malicious information is thoroughly embedded into the original video, thereby enhancing the attack capability of the poisoning model. Finally, we design an adversarial module to improve visual consistency between the original and poisoned videos, thus enhancing the concealment of malicious information within the training data of retrieval models. Our poisoning model can identify security flaws in video-text cross-modal retrieval models, providing insights into improving the security of retrieval models. The effectiveness of our model is validated on the MSR-VTT, LSMDC, and MSVD datasets.
Ming Jin 0007, Wenbo Hu 0001, Richang Hong, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 Align Is Not Enough: Multimodal Universal Jailbreak Attack Against Multimodal Large Language Models
abstract
Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the nature of human intelligence, which processes a variety of data forms beyond just text. Despite advancements, the undesirable generation of these models remains a critical concern, particularly due to vulnerabilities exposed by text-based jailbreak attacks, which have represented a significant threat by challenging existing safety protocols. Motivated by the unique security risks posed by the integration of new and old modalities for MLLMs, we propose a unified multimodal universal jailbreak attack framework that leverages iterative image-text interactions and transfer-based strategy to generate a universal adversarial suffix and image. Our work not only highlights the interaction of image-text modalities can be used as a critical vulnerability but also validates that multimodal universal jailbreak attacks can bring higher-quality undesirable generations across different MLLMs. We evaluate the undesirable context generation of MLLMs like LLaVA, Yi-VL, MiniGPT4, MiniGPT-v2, and InstructBLIP, and reveal significant multimodal safety alignment issues, highlighting the inadequacy of current safety mechanisms against sophisticated multimodal attacks. This study underscores the urgent need for robust safety measures in MLLMs, advocating for a comprehensive review and enhancement of security protocols to mitigate potential risks associated with multimodal capabilities.
Youze Wang, Wenbo Hu 0001, Yinpeng Dong, Jing Liu 0001, Hanwang Zhang, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.2
2025 SDE-HNN: Accurate and Well-Calibrated Forecasting Using Stochastic Differential Equations
abstract
It is crucial yet challenging for deep learning models to properly characterize uncertainty that is pervasive in real-world environments. Heteroscedastic neural networks (HNNs) are promising methods that capture data uncertainty for forecasting problems while existing HNNs have difficulties in conjoining calibrated uncertainty estimation and satisfactory predictive performance due to the failure to construct an explicit interaction between the prediction and its associated uncertainty. This article develops SDE-HNN, an improved HNN equipped with stochastic differential equations (SDE), to characterize the interaction between the predictive mean and variance inside HNNs for accurate and reliable forecasting. The existence and uniqueness of the solution to the devised neural SDE are guaranteed. Moreover, based on the bias-variance tradeoff for the optimization in SDE-HNN, we design an enhanced numerical SDE solver to improve learning stability. Finally, we present two new diagnostic uncertainty metrics to systematically evaluate the predictive uncertainty. Experiments on various challenging datasets show that our method significantly outperforms state-of-the-art baselines on both predictive performance and uncertainty quantification, delivering well-calibrated and sharp prediction intervals in time-series forecasting.
Peng Cui 0007, Zhijie Deng, Wenbo Hu 0001, Jun Zhu 0001
ACM Trans. Knowl. Discov. Data3
2025 Uncertainty Calibration for Counterfactual Propensity Estimation in Recommendation
abstract
Post-click conversion rate (CVR) is a reliable indicator of online customers' preferences, making it crucial for developing recommender systems. A major challenge in predicting CVR is severe selection bias, arising from users' inherent self-selection behavior and the system's item selection process. To mitigate this issue, the inverse propensity score (IPS) is employed to weight the prediction error of each observed instance. However, current propensity score estimations are unreliable due to the lack of a quality measure. To address this, we evaluate the quality of propensity scores from the perspective of uncertainty calibration, proposing the use of Expected Calibration Error (ECE) as a measure of propensity-score quality, which quantifies the extent to which predicted probabilities are overconfident by assessing the difference between predicted probabilities and actual observed frequencies. Miscalibrated propensity scores can lead to distorted IPS weights, thereby compromising the debiasing process in CVR prediction. In this paper, we introduce a model-agnostic calibration framework for propensity-based debiasing of CVR predictions. Theoretical analysis on bias and generalization bounds demonstrates the superiority of calibrated propensity estimates over uncalibrated ones. Experiments conducted on the Coat, Yahoo and KuaiRand datasets show improved uncertainty calibration, as evidenced by lower ECE values, leading to enhanced CVR prediction outcomes.
Wenbo Hu 0001, Qiang Liu 0006, Le Wu 0001, Liang Wang 0001
IEEE Trans. Knowl. Data Eng.1
2025 Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-Training Models With Contrastive Learning
abstract
The integration of visual and textual data in Vision-Language Pre-training (VLP) models is crucial for enhancing vision-language understanding. However, the adversarial robustness of these models, especially in the alignment of image-text features, has not yet been sufficiently explored. In this paper, we introduce a novel gradient-based multimodal adversarial attack method, underpinned by contrastive learning, to improve the transferability of multimodal adversarial samples in VLP models. This method concurrently generates adversarial texts and images within imperceptive perturbation, employing both image-text and intra-modal contrastive loss. We evaluate the effectiveness of our approach on image-text retrieval and visual entailment tasks, using publicly available datasets in a black-box setting. Extensive experiments indicate a significant advancement over existing single-modal transfer-based adversarial attack methods and current multimodal adversarial attack approaches.
Youze Wang, Wenbo Hu 0001, Yinpeng Dong, Hanwang Zhang, Hang Su 0006, Richang Hong
IEEE Trans. Multim.2
2024 AtomTool: Empowering Large Language Models with Tool Utilization Skills
Yongle Li, Zheng Zhang 0006, Wenbo Hu 0001, Yongyu Wu, Richang Hong
PRCV (1)4
2024 Based on Spatial and Temporal Implicit Semantic Relational Inference for Cross-Modal Retrieval
abstract
To meet users’ demands for video retrieval, text-video cross-modal retrieval technology continues to evolve. Methods based on pre-trained models and transfer learning are widely employed in designing cross-modal retrieval models, significantly enhancing the accuracy of video retrieval. However, these methods exhibit shortcomings when it comes to studying the relationships between video frames, preventing the model from fully establishing the hidden semantic relationships within video features. To further deduce the implicit semantic relationships among video frames, we propose a cross-modal retrieval model based on graph convolutional networks (GCN) and visual semantic inference (GVSI). The GCN is utilized to establish relationships between video frame features, facilitating the mining of hidden semantic information across video frames. In order to use text semantic features to help the model to infer temporal and implicit semantic information between video frames, we introduce a semantic mining and temporal space (SM&TS) inference module. Additionally, we design semantic alignment modules (SA_M) to align explicit and implicit object features present in both video and text. Finally, we analyze and validate the effectiveness of the model using MSR-VTT, MSVD, and LSMDC datasets.
Ming Jin 0007, Wenbo Hu 0001, Lei Zhu 0002, Xiang Wang 0010, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Iterative Adversarial Attack on Image-Guided Story Ending Generation
abstract
Multimodal learning involves developing models that can integrate information from various sources like images and texts. In this field, multimodal text generation is a crucial aspect that involves processing data from multiple modalities and outputting text. The image-guided story ending generation (IgSEG) is a particularly significant task, targeting on an understanding of complex relationships between text and image data with a complete story text ending. Unfortunately, deep neural networks, which are the backbone of recent IgSEG models, are vulnerable to adversarial samples. Current adversarial attack methods mainly focus on single-modality data and do not analyze adversarial attacks for multimodal text generation tasks that use cross-modal information. To this end, we propose an iterative adversarial attack method (Iterative-attack) that fuses image and text modality attacks, allowing for an attack search for adversarial text and image in a more effective iterative way. Experimental results demonstrate that the proposed method outperforms existing single-modal and non-iterative multimodal attack methods, indicating the potential for improving the adversarial robustness of multimodal text generation models, such as multimodal machine translation, multimodal question answering, etc.
Youze Wang, Wenbo Hu 0001, Richang Hong
IEEE Trans. Multim.2
2023 Physics-Guided Discovery of Highly Nonlinear Parametric Partial Differential Equations
abstract
Partial differential equations (PDEs) that fit scientific data can represent physical laws with explainable mechanisms for various mathematically-oriented subjects, such as physics and finance. The data-driven discovery of PDEs from scientific data thrives as a new attempt to model complex phenomena in nature, but the effectiveness of current practice is typically limited by the scarcity of data and the complexity of phenomena. Especially, the discovery of PDEs with highly nonlinear coefficients from low-quality data remains largely under-addressed. To deal with this challenge, we propose a novel physics-guided learning method, which can not only encode observation knowledge such as initial and boundary conditions but also incorporate the basic physical principles and laws to guide the model optimization. We theoretically show that our proposed method strictly reduces the coefficient estimation error of existing baselines, and is also robust against noise. Extensive experiments show that the proposed method is more robust against data noise, and can reduce the estimation error by a large margin. Moreover, all the PDEs in the experiments are correctly discovered, and for the first time we are able to discover three-dimensional PDEs with highly nonlinear coefficients.
Yingtao Luo, Qiang Liu 0006, Yuntian Chen, Wenbo Hu 0001, Tian Tian 0001, Jun Zhu 0001
KDD4
2021 Two Birds with One Stone: Series Saliency for Accurate and Interpretable Multivariate Time Series Forecasting
abstract
It is important yet challenging to perform accurate and interpretable time series forecasting. Though deep learning methods can boost forecasting accuracy, they often sacrifice interpretability. In this paper, we present a new scheme of series saliency to boost both accuracy and interpretability. By extracting series images from sliding windows of the time series, we design series saliency as a mixup strategy with a learnable mask between the series images and their perturbed versions. Series saliency is model agnostic and performs as an adaptive data augmentation method for training deep models. Moreover, by slightly changing the objective, we optimize series saliency to find a mask for interpretable forecasting in both feature and time dimensions. Experimental results on several real datasets demonstrate that series saliency is effective to produce accurate time-series forecasting results as well as generate temporal interpretations.
Qingyi Pan, Wenbo Hu 0001, Ning Chen 0002
IJCAI2
2020 Calibrated Reliable Regression using Maximum Mean Discrepancy
abstract
Accurate quantification of uncertainty is crucial for real-world applications of machine learning. However, modern deep neural networks still produce unreliable predictive uncertainty, often yielding over-confident predictions. In this paper, we are concerned with getting well-calibrated predictions in regression tasks. We propose the calibrated regression method using the maximum mean discrepancy by minimizing the kernel embedding measure. Theoretically, the calibration error of our method asymptotically converges to zero when the sample size is large enough. Experiments on non-trivial real datasets show that our method can produce well-calibrated and sharp prediction intervals, which outperforms the related state-of-the-art methods.
Peng Cui 0007, Wenbo Hu 0001, Jun Zhu 0001
NeurIPS2
2017 Semi-supervised Max-margin Topic Model with Manifold Posterior Regularization
abstract
Supervised topic models leverage label information to learn discriminative latent topic representations. As collecting a fully labeled dataset is often time-consuming, semi-supervised learning is of high interest. In this paper, we present an effective semi-supervised max-margin topic model by naturally introducing manifold posterior regularization to a regularized Bayesian topic model, named LapMedLDA. The model jointly learns latent topics and a related classifier with only a small fraction of labeled documents. To perform the approximate inference, we derive an efficient stochastic gradient MCMC method. Unlike the previous semi-supervised topic models, our model adopts a tight coupling between the generative topic model and the discriminative classifier. Extensive experiments demonstrate that such tight coupling brings significant benefits in quantitative and qualitative performance.
Wenbo Hu 0001, Jun Zhu 0001, Hang Su 0006, Jingwei Zhuo, Bo Zhang 0010
IJCAI1
2017 Fast sampling methods for Bayesian max-margin models
Wenbo Hu 0001, Jun Zhu 0001, Bo Zhang 0010
Expert Syst. Appl.1
2012 Neural-network-based cooperative adaptive identification of nonlinear systems
abstract
This paper considers the problem of cooperative adaptive identification for a class of nonlinear systems via neural networks. The proposed adaptive laws of neural network weights are distributed, and the interconnection topologies are established among identification models in order to share their data on-line. It is proved that if the interconnection topologies are undirected and connected, then all adaptive laws of neural network weights for the same system function can converge to a small neighborhood around their optimal values over a union of sets consisting of system trajectories. Thus, the learned system model has the better generalization capability. A simulation example are provided to verify the effectiveness and advantages of the algorithms proposed in this paper.
Weisheng Chen, Shaoyong Hua, Wenlong Ren, Wenbo Hu 0001
ICARCV4