Feng Hou

dblp:65/8064 · DBLP profile ↗
← Back
38ranked-venue papers
7as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A diversified and heterogeneous ensemble learning framework with predictive uncertainty calibration for explainable anomaly analysis
abstract
Abstract Ensemble methods have been the norm for anomaly detection. However, existing ensemble methods for anomaly detection have three main issues: (1) Lack of diversity, the base classifiers are of the same algorithm with different initializations, for example, decision trees or neural networks, achieving only sub-optimal results. (2) The predictive uncertainty is not well calibrated for reliable and explainable anomaly detection, where overconfident predictions for both correct and erroneous classifications are made, making the results unreliable and hard to explain. (3) Traditional ensemble methods (e.g., bagging) cannot effectively capture the distinct distributions of predictions from diverse base models. In this paper, we propose a D iversified and H eterogeneous E nsemble learning framework with C alibrated predictive U ncertainty estimation (DHE-CU). We utilize a multi-layer perceptron (MLP) as the meta-classifier to combine the confidence of diverse base models, thereby achieving more explainable anomaly detection. We devise a global diversity loss that considers a global measure of diversity for the selection and pruning of the base models. The MLP meta-classifier can capture the diverse and distinct distributions of predictions from base classifiers. We use a simple yet effective method to quantify the predictive uncertainty of the meta-classifier. We propose a weighted accuracy-uncertainty calibration loss for class-imbalanced data to effectively calibrate predictive uncertainties. Various datasets are used to perform experimental evaluation extensively. The proposed DHE-CU framework demonstrates strong ensemble learning ability, achieving an average improvement of 8.8% in classification accuracy across 27 UCR time series datasets and other anomaly detection benchmark datasets.
Feng Hou, Ruili Wang 0001
Knowl. Inf. Syst.2
2026 Pre-ordering representations improve low-resource neural machine translation and application in the Māori language
abstract
Abstract Pre-ordering is a data pre-processing technique that reorganizes the word order in a source language sentence to better align with the syntax of the target language, and it has been shown to significantly impact various tasks. While previous studies have primarily leveraged token position embeddings in pre-ordered sentences to enhance machine translation, they have not directly explored learning contextualized representations that encapsulate richer semantic information and also reflect the word position information in the pre-ordered sentences. In this work, we introduce a novel pre-ordering-aware neural network that explicitly integrates the representations of pre-ordered sentences into the training process. Our network utilizes a pre-ordering encoder, which works alongside the original Transformer encoder to process the pre-ordered source sentence. We propose a Cross-Encoder Consistency (CEC) block, which encourages the original encoder to produce representations that better reflect the word order of the target language by closing the gap between representations of the original and pre-ordered sentence. Additionally, we incorporate a Sentence Encoding Consistency (SEC) block to help the model preserve the semantic integrity of the source sentence. We conduct extensive experiments on multiple low-resource neural machine translation benchmarks, including an endangered language, the Māori language. Our method delivers substantial enhancements in translation quality, achieving improvements of up to +1.46 BLEU in Tr $$\rightarrow$$ En translation.
Yuan Gao 0057, Feng Hou, Huia Jahnke
Multim. Tools Appl.2
2025 DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes
abstract
Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance.
Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
AAAI2
2025 Empowering Māori Automatic Speech Recognition through EMD-Based Augmentation
Chengxi Lei, Sheng Li 0010, Satwinder Singh, Feng Hou, Huia Jahnke, Ruili Wang 0001
PRICAI4
2025 Parameter-Efficient Personalized Speech Synthesis via EMD-based Speaker Modeling
abstract
Personalized speech synthesis has attracted increasing attention in the field of robotics. Compared to traditional speech synthesis, it faces two primary challenges: the limited availability of adaptation data and the necessity for highly efficient adaptation method using compact parameters to reduce both training time and memory consumption. To address both the challenges, this paper proposes a personalized speech synthesis approach that incorporates an Empirical Mode Decomposition (EMD)-based speaker modeling method alongside a novel decoder structure with masked inputs, which improves the model’s ability to extract speaker-specific features accurately. Furthermore, we introduce a parameter-efficient fine-tuning technique, Attention-based Speaker-Text Scaling and Shifting Feature (AST-SSF), to enhance adaptation efficiency. We validate our approach using the MAGICDATA Corpus. The results indicate that our proposed approach outperforms the baseline in both naturalness and similarity, demonstrating its effectiveness. Moreover, although the proposed adaptation method substantially reduces the number of parameters, it exhibits only minimal performance degradation compared to full and partial fine-tuning strategies.
Chengxi Lei, Feng Hou, Huia Jahnke, Ruili Wang 0001
RO-MAN2
2025 Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
Multim. Syst.1
2025 Enhancing video salient object detection via SAM-based multimodal energy prompting
Yi Wang 0037, Feng Hou, Li-li Liu
Pattern Anal. Appl.3
2025 Knowledge-sharing hierarchical memory fusion network for scribble-supervised video salient object detection
Feng Hou, Yi Wang 0037, Guangzhu Chen, Ruili Wang 0001
Pattern Recognit. Lett.2
2025 DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain Generalization
abstract
Traditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods.
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui
IEEE Trans. Multim.1
2024 Incorporating Pre-ordering Representations for Low-resource Neural Machine Translation
Yuan Gao 0057, Feng Hou, Ruili Wang 0001
MMAsia2
2024 Multimodal Energy Prompting for Video Salient Object Detection
Feng Hou, Yi Wang 0037
MMAsia2
2024 Mix-fine-tune: An Alternate Fine-tuning Strategy for Domain Adaptation and Generalization of Low-resource ASR
abstract
Self-supervised Learning (SSL) using extensive unlabeled speech data has significantly improved the performance of ASR models on datasets like LibriSpeech.However, few studies have addressed the issue of domain mismatch between the data used to pre-train and fine-tune ASR models.Moreover, the Empirical Risk Minimization (ERM) principle, commonly used to train deep learning models, often causes the trained models to exhibit undesirable behaviors such as memorizing training data and being sensitive to adversarial examples.Thus, in this paper, we propose an alternate fine-tuning strategy, called Mix-fine-tune, to address domain mismatch in ASR systems and the limitations of the ERM training principle.Mix-finetune use a data-driven weighted sum of two speech sequences as input and the corresponding text sequences are used to calculate a weighted audio-text alignment Connectionist Temporal Classification (CTC) loss for fine-tuning a pre-trained model.Additionally, Mix-fine-tune incorporates the masked Contrastive Predictive Coding (CPC) loss, previously used exclusively for pre-training, into the fine-tuning process.Our novel strategy alternates between minimizing the CTC loss and the CPC loss to address the domain mismatch between pre-training and fine-tuning.We validate our method by fine-tuning different sizes of the Wav2Vec model using the public Air Traffic Control (ATC) corpus.The experiments show that Mix-fine-tune efficiently adapts the models pre-trained on general speech corpora like LibriSpeech to a specific domain (e.g., the air traffic control domain) by fine-turning.
Chengxi Lei, Satwinder Singh, Feng Hou, Ruili Wang 0001
MMAsia3
2024 Structured Bipartite Graph Ensemble Clustering
Chen Wang 0108, Feng Hou, Yi Wang 0037, Ruili Wang 0001
MMAsia2
2024 IENet: inheritance enhancement network for video salient object detection
Yi Wang 0037, Feng Hou, Ruili Wang 0001
Multim. Tools Appl.3
2024 Anisotropic span embeddings and the negative impact of higher-order inference for coreference resolution: An empirical analysis
abstract
Abstract Coreference resolution is the task of identifying and clustering mentions that refer to the same entity in a document. Based on state-of-the-art deep learning approaches, end-to-end coreference resolution considers all spans as candidate mentions and tackles mention detection and coreference resolution simultaneously. Recently, researchers have attempted to incorporate document-level context using higher-order inference (HOI) to improve end-to-end coreference resolution. However, HOI methods have been shown to have marginal or even negative impact on coreference resolution. In this paper, we reveal the reasons for the negative impact of HOI coreference resolution. Contextualized representations (e.g., those produced by BERT) for building span embeddings have been shown to be highly anisotropic. We show that HOI actually increases and thus worsens the anisotropy of span embeddings and makes it difficult to distinguish between related but distinct entities (e.g., pilots and flight attendants ). Instead of using HOI, we propose two methods, Less-Anisotropic Internal Representations (LAIR) and Data Augmentation with Document Synthesis and Mention Swap (DSMS), to learn less-anisotropic span embeddings for coreference resolution. LAIR uses a linear aggregation of the first layer and the topmost layer of contextualized embeddings. DSMS generates more diversified examples of related but distinct entities by synthesizing documents and by mention swapping. Our experiments show that less-anisotropic span embeddings improve the performance significantly (+2.8 F1 gain on the OntoNotes benchmark) reaching new state-of-the-art performance on the GAP dataset.
Feng Hou, Ruili Wang 0001, See-Kiong Ng, Fangyi Zhu, Michael Witbrock, Steven F. Cahan, Lily Chen, Xiaoyun Jia
Nat. Lang. Eng.1
2024 Domain-Aware Graph Network for Bridging Multi-Source Domain Adaptation
abstract
Domain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG.
Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui
IEEE Trans. Multim.2
2024 A Survey of Visual Transformers
abstract
Transformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers.
Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
IEEE Trans. Neural Networks Learn. Syst.4
2023 Learning How to Learn Domain-Invariant Parameters for Domain Generalization
abstract
Due to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability.
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
ICASSP1
2023 A Novel Self-training Approach for Low-resource Speech Recognition
Satwinder Singh, Feng Hou, Ruili Wang 0001
INTERSPEECH2
2023 CLRL-Tuning: A Novel Continual Learning Approach for Automatic Speech Recognition
Zhihan Wang, Feng Hou, Ruili Wang 0001
INTERSPEECH2
2023 Data Augmentation with Diversified Rephrasing for Low-Resource Neural Machine Translation
abstract
Data augmentation is an effective way to enhance the performance of neural machine translation models, especially for low-resource languages. Existing data augmentation methods are either at a token level or a sentence level. The data augmented using token level methods lack syntactic diversity and may alter original meanings. Sentence level methods usually generate low-quality source sentences that are not semantically paired with the original target sentences. In this paper, we propose a novel data augmentation method to generate diverse, high-quality and meaning-preserved new instances. Our method leverages high-quality translation models trained with high-resource languages to rephrase an original sentence by translating it into an intermediate language and then back to the original language. Through this process, the high-performing translation models guarantee the quality of the rephrased sentences, and the syntactic knowledge from the intermediate language can bring syntactic diversity to the rephrased sentences. Experimental results show our method can enhance the performance in various low-resource machine translation tasks. Moreover, by combining our method with other techniques that facilitate NMT, we can yield even better results.
Yuan Gao 0057, Feng Hou, Huia Jahnke, Ruili Wang 0001
MTSummit (1)2
2023 Learning and integration of adaptive hybrid graph structures for multivariate time series forecasting
abstract
Recent status-of-the-art methods for multivariate time series forecasting can be categorized into graph-based approach and global-local approach. The former approach uses graphs to represent the dependencies among variables and apply graph neural networks to the forecasting problem. The latter approach decomposes the matrix of multivariate time series into global components and local components to capture the shared information across variables. However, both approaches cannot capture the propagation delay of the dependencies among individual variables of a multivariate time series, for example, the congestion at intersection A has a delayed effects on the neighbouring intersection B. In addition, graph-based forecasting methods cannot capture the shared global tendency across the variables of a multivariate time series; and global-local forecasting methods cannot reflect the nonlinear inter-dependencies among variables of a multivariate time series. In this paper, we propose to combine the advantages of both approaches by integrating Adaptive Global-Local Graph Structure Learning with Gated Recurrent Units (AGLG-GRU). We learn a global graph to represent the shared information across variables. And we learn dynamic local graphs to capture the local randomness and nonlinear dependencies among variables. We apply diffusion convolution and graph convolution operations to global and dynamic local graphs to integrate the information of graphs and update gated recurrent unit for multivariate time series forecasting. The experimental results on seven representative real-world datasets demonstrate that our approach outperform various existing methods.
Feng Hou, Xiaoyun Jia, Ruili Wang 0001
Inf. Sci.2
2023 Exploiting anonymous entity mentions for named entity linking
Feng Hou, Ruili Wang 0001, See-Kiong Ng, Michael Witbrock, Fangyi Zhu, Xiaoyun Jia
Knowl. Inf. Syst.1
2023 Fine-Grained Entity Typing With a Type Taxonomy: A Systematic Review
abstract
Fine-grained entity typing (FGET) is an important natural language processing task. It is to assign fine-grained semantic types of a type taxonomy (e.g., Person/artist/actor) to entity mentions. Fine-grained entity semantic types have been successfully applied in many natural language processing (NLP) applications, such as relation extraction, entity linking and question answering. The key challenge for FGET is how to deal with label noises that disperse in the corpora since the corpora are normally automatically annotated. Various type taxonomies, typing methods and representation learning approaches for FGET have been proposed and developed in the past two decades. This paper systematically categorizes and reviews these various typing methods and representation learning approaches to provide a reference for future studies on FGET. We identify the current trends in FGET research: (i) Learning embedded feature representations to address the challenges posed by label noises, tail types and new entities; (ii) Tackling FGET jointly with other entity analysis sub-tasks (e.g., entity linking and coreference resolution) is also a promising direction. We also present a comprehensive review of type taxonomies, resources, applications for FGET and methods for automatically generating FGET training corpora.
Ruili Wang 0001, Feng Hou, Steven F. Cahan, Lily Chen, Xiaoyun Jia, Wanting Ji
IEEE Trans. Knowl. Data Eng.2
2022 Determining the best Acoustic Features for Smoker Identification
abstract
Speech-based automatic smoker identification (also known as smoker/non-smoker classification) aims to identify speakers’ smoking status from their speech. In the COVID-19 pandemic, speech-based automatic smoker identification approaches have received more attention in smoking cessation research due to low cost and contactless sample collection. This study focuses on determining the best acoustic features for smoker identification. In this paper, we investigate the performance of four acoustic feature sets/representations extracted using three feature extraction/learning approaches: (i) hand-crafted feature sets including the extended Geneva Minimalistic Acoustic Parameter Set and the Computational Paralinguistics Challenge Set, (ii) the Bag-of-Audio-Words representations, (iii) the neural representations extracted from raw waveform signals by SincNet. Experimental results show that: (i) SincNet feature representations are the most effective for smoker identification and outperform the MFCC baseline features by 16% in absolute accuracy; (ii) the performance of hand-crafted feature sets and the Bag-of-Audio-Words representations rely on the scale of the dimensions of feature vectors.
Zhizhong Ma, Yuanhang Qiu, Feng Hou, Ruili Wang 0001, Joanna Ting Wai Chu, Chris Bullen
ICASSP3
2022 Improved Meta Learning for Low Resource Speech Recognition
abstract
We propose a new meta learning based framework for low resource speech recognition that improves the previous model agnostic meta learning (MAML) approach. The MAML is a simple yet powerful meta learning approach. However, the MAML presents some core deficiencies such as training instabilities and slower convergence speed. To address these issues, we adopt multi-step loss (MSL). The MSL aims to calculate losses at every step of the inner loop of MAML and then combines them with a weighted importance vector. The importance vector ensures that the loss at the last step has more importance than the previous steps. Our empirical evaluation shows that MSL significantly improves the stability of the training procedure and it thus also improves the accuracy of the overall system. Our proposed system outperforms MAML based low resource ASR system on various languages in terms of character error rates and stable training behavior.
Satwinder Singh, Ruili Wang 0001, Feng Hou
ICASSP3
2022 CyclicAugment: Speech Data Random Augmentation with Cosine Annealing Scheduler for Auotmatic Speech Recognition
Zhihan Wang, Feng Hou, Yuanhang Qiu, Zhizhong Ma, Satwinder Singh, Ruili Wang 0001
INTERSPEECH2
2022 Self-Supervised Graph Neural Network for Multi-Source Domain Adaptation
abstract
Domain adaptation (DA) tries to tackle the scenarios when the test data does not fully follow the same distribution of the training data, and multi-source domain adaptation (MSDA) is very attractive for real world applications. By learning from large-scale unlabeled samples, self-supervised learning has now become a new trend in deep learning. It is worth noting that both self-supervised learning and multi-source domain adaptation share a similar goal: they both aim to leverage unlabeled data to learn more expressive representations. Unfortunately, traditional multi-task self-supervised learning faces two challenges: (1) the pretext task may not strongly relate to the downstream task, thus it could be difficult to learn useful knowledge being shared from the pretext task to the target task; (2) when the same feature extractor is shared between the pretext task and the downstream one and only different prediction heads are used, it is ineffective to enable inter-task information exchange and knowledge sharing. To address these issues, we propose a novel Self-Supervised Graph Neural Network (SSG), where a graph neural network is used as the bridge to enable more effective inter-task information exchange and knowledge sharing. More expressive representation is learned by adopting a mask token strategy to mask some domain information. Our extensive experiments have demonstrated that our proposed SSG method has achieved state-of-the-art results over four multi-source domain adaptation datasets, which have shown the effectiveness of our proposed SSG method from different aspects.
Feng Hou, Yangzhou Du, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui
ACM Multimedia2
2022 FADA: A Cloud-Fog-Edge Architecture and Ontology for Data Acquisition
abstract
Large and complex machines are the backbone of manufacturing and will remain key to industrialization for the foreseeable future as will leveraging essential technologies, such as the Internet of Things. However, efforts to make this machinery as smart as it could be are falling behind the curve. Data quality is often too low or too heterogeneous for useful analytics making maintenance troublesome. Additionally, many machines are still dumb with no uniform way to extract data or monitor their operation. In part, the success of future concepts like intelligent manufacturing, cyber physical systems, and industry 4.0 depends on solving these problems. Hence, this article presents FADA the groundwork for an ontology and 3-layer cloud-fog-edge architecture for large and complex machines that places data acquisition and IoT at the forefront. The ontology provides a flexible framework for standardizing data. The fog nodes acquire data directly from smart machines, while the edge nodes harvest data from dumb equipment through a recognition model. The fog nodes are flexible and multi-threaded to provide faster higher-performance computing power. To evaluate the proposed architecture and concepts, we implemented FADA in two factory-based testbeds: one with IoT-enabled equipment, the other with mostly dumb machines. The response times and influence rates recorded are promising and indicate that the system is highly adaptable to many different scenarios. We also conducted comparative experiments between FADA and a conventional data acquisition system to compare the occupied disk space, processing time, and data uploading time, which show that the FADA can save 2.9 TB of disk space per day, and reduce the server’s processing time by 184.8 ms per time over the conventional data acquisition system(CDAS), when 20000 fog nodes simultaneously access the server. The results show improvements by FADA in all metrics.
Huifeng Wu, Baiping Chen, Feng Hou, Danfeng Sun
IEEE Trans. Cloud Comput.4
2021 Self-Supervised Learning Based Phone-Fortified Speech Enhancement
Yuanhang Qiu, Ruili Wang 0001, Satwinder Singh, Zhizhong Ma, Feng Hou
Interspeech5
2021 Transfer learning for fine-grained entity typing
Feng Hou, Ruili Wang 0001
Knowl. Inf. Syst.1
2021 The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge
Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xiaoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan 0002, Guangrui Mu, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiawei Yang 0002, Guangwei Xiong, Jiang Tian, Christopher J. Weight
Medical Image Anal.15
2021 VerSe: A Vertebrae labelling and segmentation benchmark for multi-detector CT images
Anjany Sekuboyina, Malek El Husseini, Amirhossein Bayat, Maximilian Löffler, Hans Liebl, Hongwei Li 0004, Giles Tetteh, Jan Kukacka, Christian Payer, Darko Stern, Martin Urschler, Maodong Chen, Dalong Cheng, Nikolas Leßmann, Yujin Hu, Tianfu Wang 0001, Dong Yang 0005, Daguang Xu, Felix Ambellan, Tamaz Amiranashvili, Moritz Ehlke, Hans Lamecker, Sebastian Lehnert, Marilia Lirio, Nicolás Pérez de Olaguer, Heiko Ramm, Manish Sahu, Alexander Tack, Stefan Zachow, Xinjun Ma, Christoph Angerman, Xin Wang 0113, Alexandre Kirszenberg, Élodie Puybareau, Yiwei Bai, Brandon H. Rapazzo, Timyoas Yeah, Amber Zhang, Shangliang Xu, Feng Hou, Zhiqiang He 0002, Chan Zeng, Zheng Xiangshang, Xu Liming, Tucker J. Netherton, Raymond P. Mumme, Laurence E. Court, Zixun Huang, Chenhang He, Li-Wen Wang, Sai-Ho Ling, Lê Duy Huynh, Nicolas Boutry, Roman Jakubícek, Jirí Chmelík, Supriti Mulay, Mohanasankar Sivaprakasam, Johannes C. Paetzold, Suprosanna Shit, Ivan Ezhov, Benedikt Wiestler, Ben Glocker, Alexander Valentinitsch, Markus Rempfler, Bjoern Menze, Jan Kirschke
Medical Image Anal.43
2020 Improving Entity Linking through Semantic Reinforced Entity Embeddings
abstract
Entity embeddings, which represent different aspects of each entity with a single vector like word embeddings, are a key component of neural entity linking models.Existing entity embeddings are learned from canonical Wikipedia articles and local contexts surrounding target entities.Such entity embeddings are effective, but too distinctive for linking models to learn contextual commonality.We propose a simple yet effective method, FGS2EE, to inject fine-grained semantic information into entity embeddings to reduce the distinctiveness and facilitate the learning of contextual commonality.FGS2EE first uses the embeddings of semantic type words to generate semantic embeddings, and then combines them with existing entity embeddings through linear aggregation.Extensive experiments show the effectiveness of such embeddings.Based on our entity embeddings, we achieved new state-ofthe-art performance on entity linking.
Feng Hou, Ruili Wang 0001, Jun He 0001
ACL1
2020 Analysis of electronic health records based on long short-term memory
abstract
Summary There is a large amount of historical data of the patient's hospitalization named the electronic health records (EHRs), but the data are not fully utilized for great challenges as poor quality, high dimension, and so on. Previous studies have primarily used machine learning methods that rely heavily on manual extraction of features. Recently, many deep learning approaches are applied to predictive model of EHRs. Recurrent neural networks (RNN) are often used to model EHR data, but RNN performance degrades in the face of large sequence lengths. To solve these challenges, we develop a long short‐term memory with attention mechanism for mortality prediction. The dataset used in this article is the Medical Information Mart for Intensive Care III, which contains comprehensive clinical data for the patients. The experimental results demonstrate that the predicted results can be effectively interpreted using the attention mechanism. Compared with other baseline models, our model improves the accuracy of prediction, and helps doctors reduce the average diagnostic time.
Peiying Shi, Feng Hou, Xiangwei Zheng 0001
Concurr. Comput. Pract. Exp.2
2015 Adaptive differential evolution with directional strategy and cloud model
Jin Gou, Wang-Ping Guo, Feng Hou, Cheng Wang 0020, Yiqiao Cai
Appl. Intell.3
2015 Improving Wang-Mendel method performance in fuzzy rules generation using the fuzzy C-means clustering algorithm
Jin Gou, Feng Hou, Cheng Wang 0020
Neurocomputing2
2008 A hybrid approach combining an improved genetic algorithm and optimization strategies for the asymmetric traveling salesman problem
Ling-Ning Xing, Ying-Wu Chen 0001, Ke-Wei Yang 0001, Feng Hou, Xue-Shi Shen, Huai-Ping Cai
Eng. Appl. Artif. Intell.4