Feng Hou

dblp:65/8064 · DBLP profile ↗
← Back
9ranked-venue papers in the field
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 4Data Mining & Knowledge Discovery · 3 (2 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 A diversified and heterogeneous ensemble learning framework with predictive uncertainty calibration for explainable anomaly analysis
abstract
Abstract Ensemble methods have been the norm for anomaly detection. However, existing ensemble methods for anomaly detection have three main issues: (1) Lack of diversity, the base classifiers are of the same algorithm with different initializations, for example, decision trees or neural networks, achieving only sub-optimal results. (2) The predictive uncertainty is not well calibrated for reliable and explainable anomaly detection, where overconfident predictions for both correct and erroneous classifications are made, making the results unreliable and hard to explain. (3) Traditional ensemble methods (e.g., bagging) cannot effectively capture the distinct distributions of predictions from diverse base models. In this paper, we propose a D iversified and H eterogeneous E nsemble learning framework with C alibrated predictive U ncertainty estimation (DHE-CU). We utilize a multi-layer perceptron (MLP) as the meta-classifier to combine the confidence of diverse base models, thereby achieving more explainable anomaly detection. We devise a global diversity loss that considers a global measure of diversity for the selection and pruning of the base models. The MLP meta-classifier can capture the diverse and distinct distributions of predictions from base classifiers. We use a simple yet effective method to quantify the predictive uncertainty of the meta-classifier. We propose a weighted accuracy-uncertainty calibration loss for class-imbalanced data to effectively calibrate predictive uncertainties. Various datasets are used to perform experimental evaluation extensively. The proposed DHE-CU framework demonstrates strong ensemble learning ability, achieving an average improvement of 8.8% in classification accuracy across 27 UCR time series datasets and other anomaly detection benchmark datasets.
Feng Hou, Ruili Wang 0001
Knowl. Inf. Syst.2
2024 Incorporating Pre-ordering Representations for Low-resource Neural Machine Translation
Yuan Gao 0057, Feng Hou, Ruili Wang 0001
MMAsia2
2024 Multimodal Energy Prompting for Video Salient Object Detection
Feng Hou, Yi Wang 0037
MMAsia2
2024 Mix-fine-tune: An Alternate Fine-tuning Strategy for Domain Adaptation and Generalization of Low-resource ASR
abstract
Self-supervised Learning (SSL) using extensive unlabeled speech data has significantly improved the performance of ASR models on datasets like LibriSpeech.However, few studies have addressed the issue of domain mismatch between the data used to pre-train and fine-tune ASR models.Moreover, the Empirical Risk Minimization (ERM) principle, commonly used to train deep learning models, often causes the trained models to exhibit undesirable behaviors such as memorizing training data and being sensitive to adversarial examples.Thus, in this paper, we propose an alternate fine-tuning strategy, called Mix-fine-tune, to address domain mismatch in ASR systems and the limitations of the ERM training principle.Mix-finetune use a data-driven weighted sum of two speech sequences as input and the corresponding text sequences are used to calculate a weighted audio-text alignment Connectionist Temporal Classification (CTC) loss for fine-tuning a pre-trained model.Additionally, Mix-fine-tune incorporates the masked Contrastive Predictive Coding (CPC) loss, previously used exclusively for pre-training, into the fine-tuning process.Our novel strategy alternates between minimizing the CTC loss and the CPC loss to address the domain mismatch between pre-training and fine-tuning.We validate our method by fine-tuning different sizes of the Wav2Vec model using the public Air Traffic Control (ATC) corpus.The experiments show that Mix-fine-tune efficiently adapts the models pre-trained on general speech corpora like LibriSpeech to a specific domain (e.g., the air traffic control domain) by fine-turning.
Chengxi Lei, Satwinder Singh, Feng Hou, Ruili Wang 0001
MMAsia3
2024 Structured Bipartite Graph Ensemble Clustering
Chen Wang 0108, Feng Hou, Yi Wang 0037, Ruili Wang 0001
MMAsia2
2023 Learning and integration of adaptive hybrid graph structures for multivariate time series forecasting
abstract
Recent status-of-the-art methods for multivariate time series forecasting can be categorized into graph-based approach and global-local approach. The former approach uses graphs to represent the dependencies among variables and apply graph neural networks to the forecasting problem. The latter approach decomposes the matrix of multivariate time series into global components and local components to capture the shared information across variables. However, both approaches cannot capture the propagation delay of the dependencies among individual variables of a multivariate time series, for example, the congestion at intersection A has a delayed effects on the neighbouring intersection B. In addition, graph-based forecasting methods cannot capture the shared global tendency across the variables of a multivariate time series; and global-local forecasting methods cannot reflect the nonlinear inter-dependencies among variables of a multivariate time series. In this paper, we propose to combine the advantages of both approaches by integrating Adaptive Global-Local Graph Structure Learning with Gated Recurrent Units (AGLG-GRU). We learn a global graph to represent the shared information across variables. And we learn dynamic local graphs to capture the local randomness and nonlinear dependencies among variables. We apply diffusion convolution and graph convolution operations to global and dynamic local graphs to integrate the information of graphs and update gated recurrent unit for multivariate time series forecasting. The experimental results on seven representative real-world datasets demonstrate that our approach outperform various existing methods.
Feng Hou, Xiaoyun Jia, Ruili Wang 0001
Inf. Sci.2
2023 Exploiting anonymous entity mentions for named entity linking
Feng Hou, Ruili Wang 0001, See-Kiong Ng, Michael Witbrock, Fangyi Zhu, Xiaoyun Jia
Knowl. Inf. Syst.1
2023 Fine-Grained Entity Typing With a Type Taxonomy: A Systematic Review
abstract
Fine-grained entity typing (FGET) is an important natural language processing task. It is to assign fine-grained semantic types of a type taxonomy (e.g., Person/artist/actor) to entity mentions. Fine-grained entity semantic types have been successfully applied in many natural language processing (NLP) applications, such as relation extraction, entity linking and question answering. The key challenge for FGET is how to deal with label noises that disperse in the corpora since the corpora are normally automatically annotated. Various type taxonomies, typing methods and representation learning approaches for FGET have been proposed and developed in the past two decades. This paper systematically categorizes and reviews these various typing methods and representation learning approaches to provide a reference for future studies on FGET. We identify the current trends in FGET research: (i) Learning embedded feature representations to address the challenges posed by label noises, tail types and new entities; (ii) Tackling FGET jointly with other entity analysis sub-tasks (e.g., entity linking and coreference resolution) is also a promising direction. We also present a comprehensive review of type taxonomies, resources, applications for FGET and methods for automatically generating FGET training corpora.
Ruili Wang 0001, Feng Hou, Steven F. Cahan, Lily Chen, Xiaoyun Jia, Wanting Ji
IEEE Trans. Knowl. Data Eng.2
2021 Transfer learning for fine-grained entity typing
Feng Hou, Ruili Wang 0001
Knowl. Inf. Syst.1