Yu Huang 0018

dblp:39/6301-18 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
13since 2021 · last 2025
0000-0001-7373-4716ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 From image to report: automating lung cancer screening interpretation and reporting with vision-language models
Tien-Yu Chang, Qinglin Gou, Leyi Zhao, Tiancheng Zhou, Dong Yang 0005, Huiwen Ju, Kaleb E. Smith, Chengkun Sun, Jinqian Pan, Yu Huang 0018, Xing He 0003, Xuhong Zhang 0001, Daguang Xu, Jie Xu 0012, Jiang Bian 0001, Aokun Chen
J. Biomed. Informatics11
2025 Variational temporal deconfounder network for individualized treatment effect estimation with longitudinal observational data
Yu Huang 0018, Yuxi Liu 0003, Xing He 0003, Jingchuan Guo, Mattia Prosperi, Jiang Bian 0001
J. Biomed. Informatics2
2025 Identifying progression subphenotypes of Alzheimer's disease from large-scale electronic health records with machine learning
Manqi Zhou, Alice S. Tang, Alison M. C. Ke, Chang Su 0002, Yu Huang 0018, William G. Mantyh, Michael Jaffee, Katherine P. Rankin, Steven DeKosky, Jiang Bian 0001, Marina Sirota, Fei Wang 0001
J. Biomed. Informatics7
2024 Graph contrastive learning as a versatile foundation for advanced scRNA-seq data analysis
abstract
Single-cell RNA sequencing (scRNA-seq) offers unprecedented insights into transcriptome-wide gene expression at the single-cell level. Cell clustering has been long established in the analysis of scRNA-seq data to identify the groups of cells with similar expression profiles. However, cell clustering is technically challenging, as raw scRNA-seq data have various analytical issues, including high dimensionality and dropout values. Existing research has developed deep learning models, such as graph machine learning models and contrastive learning-based models, for cell clustering using scRNA-seq data and has summarized the unsupervised learning of cell clustering into a human-interpretable format. While advances in cell clustering have been profound, we are no closer to finding a simple yet effective framework for learning high-quality representations necessary for robust clustering. In this study, we propose scSimGCL, a novel framework based on the graph contrastive learning paradigm for self-supervised pretraining of graph neural networks. This framework facilitates the generation of high-quality representations crucial for cell clustering. Our scSimGCL incorporates cell-cell graph structure and contrastive learning to enhance the performance of cell clustering. Extensive experimental results on simulated and real scRNA-seq datasets suggest the superiority of the proposed scSimGCL. Moreover, clustering assignment analysis confirms the general applicability of scSimGCL, including state-of-the-art clustering algorithms. Further, ablation study and hyperparameter analysis suggest the efficacy of our network architecture with the robustness of decisions in the self-supervised learning setting. The proposed scSimGCL can serve as a robust framework for practitioners developing tools for cell clustering. The source code of scSimGCL is publicly available at https://github.com/zhangzh1328/scSimGCL.
Yuxi Liu 0003, Meichen Xiao, Yu Huang 0018, Jiang Bian 0001, Ruolin Yang 0003, Fuyi Li
Briefings Bioinform.5
2024 A scoping review of fair machine learning techniques when using real-world data
abstract
OBJECTIVE: The integration of artificial intelligence (AI) and machine learning (ML) in health care to aid clinical decisions is widespread. However, as AI and ML take important roles in health care, there are concerns about AI and ML associated fairness and bias. That is, an AI tool may have a disparate impact, with its benefits and drawbacks unevenly distributed across societal strata and subpopulations, potentially exacerbating existing health inequities. Thus, the objectives of this scoping review were to summarize existing literature and identify gaps in the topic of tackling algorithmic bias and optimizing fairness in AI/ML models using real-world data (RWD) in health care domains. METHODS: We conducted a thorough review of techniques for assessing and optimizing AI/ML model fairness in health care when using RWD in health care domains. The focus lies on appraising different quantification metrics for accessing fairness, publicly accessible datasets for ML fairness research, and bias mitigation approaches. RESULTS: We identified 11 papers that are focused on optimizing model fairness in health care applications. The current research on mitigating bias issues in RWD is limited, both in terms of disease variety and health care applications, as well as the accessibility of public datasets for ML fairness research. Existing studies often indicate positive outcomes when using pre-processing techniques to address algorithmic bias. There remain unresolved questions within the field that require further research, which includes pinpointing the root causes of bias in ML models, broadening fairness research in AI/ML with the use of RWD and exploring its implications in healthcare settings, and evaluating and addressing bias in multi-modal data. CONCLUSION: This paper provides useful reference material and insights to researchers regarding AI/ML fairness in real-world health care data and reveals the gaps in the field. Fair AI/ML in health care is a burgeoning field that requires a heightened research focus to cover diverse applications and different types of RWD.
Yu Huang 0018, Jingchuan Guo, Wei-Han William Chen, Hsin-Yueh Lin, Huilin Tang, Fei Wang 0001, Hua Xu 0001, Jiang Bian 0001
J. Biomed. Informatics1
2024 Snippet Policy Network V2: Knee-Guided Neuroevolution for Multi-Lead ECG Early Classification
abstract
Early time series classification predicts the class label of a given time series before it is completely observed. In time-critical applications, such as arrhythmia monitoring in ICU, early treatment contributes to the patient's fast recovery, and early warning could even save lives. Hence, in these cases, it is worthy of trading, to some extent, classification accuracy in favor of earlier decisions when the time series data are collected over time. In this article, we propose a novel deep reinforcement learning-based framework, snippet policy network V2 (SPN-V2), for long and varied-length multi-lead electrocardiogram (ECG) early classification. The proposed SNP-V2 contains two main components: snippet representation learning (SRL) and early classification timing learning (ECTL). The SRL is proposed to encode inner-snippet spatial correlations and inter-snippet temporal correlations into the hidden representations of the subsegment (snippet) of the input ECG. ECTL aims to learn a decision agent to classify the time series early and accurately. To optimize the proposed framework, we design a novel knee-guided neuroevolution algorithm (KGNA) to solve cardiovascular diseases' early classification problem, automatically optimizing the proposed SPN-V2 regarding the tradeoff between accuracy and earliness. In addition, we conduct a series of experiments on two real-world ECG datasets. The experimental results show the superiority of the proposed algorithm over the state-of-the-art competing methods.
Yu Huang 0018, Gary G. Yen, Vincent S. Tseng
IEEE Trans. Neural Networks Learn. Syst.1
2023 Continual learning with attentive recurrent neural networks for temporal data classification
Shao-Yu Yin, Yu Huang 0018, Tien-Yu Chang, Shih-Fang Chang, Vincent S. Tseng
Neural Networks2
2023 Snippet Policy Network for Multi-Class Varied-Length ECG Early Classification
abstract
Arrhythmia detection from ECG is an important research subject in the prevention and diagnosis of cardiovascular diseases. The prevailing studies formulate arrhythmia detection from ECG as a time series classification problem. Meanwhile, early detection of arrhythmia presents a real-world demand for early prevention and diagnosis. In this paper, we address a problem of cardiovascular diseases early classification, which is a varied-length and long-length time series early classification problem as well. For solving this problem, we propose a deep reinforcement learning-based framework, namely Snippet Policy Network (SPN), consisting of four modules, snippet generator, backbone network, controlling agent, and discriminator. Comparing to the existing approaches, the proposed framework features flexible input length, solves the dual-optimization solution of the earliness and accuracy goals. Experimental results demonstrate that SPN achieves an excellent performance of over 80% in terms of accuracy. Compared to the state-of-the-art methods, at least 7% improvement on different metrics, including the precision, recall, F1-score, and harmonic mean, is delivered by the proposed SPN. To the best of our knowledge, this is the first work focusing on solving the cardiovascular early classification problem based on varied-length ECG data. Based on these excellent features from SPN, it offers a good exemplification for addressing all kinds of varied-length time series early classification problems.
Yu Huang 0018, Gary G. Yen, Vincent S. Tseng
IEEE Trans. Knowl. Data Eng.1
2022 Periodic Attention-based Stacked Sequence to Sequence framework for long-term travel time prediction
Yu Huang 0018, Vincent S. Tseng
Knowl. Based Syst.1
2022 A Novel Constraint-Based Knee- Guided Neuroevolutionary Algorithm for Context-Specific ECG Early Classification
abstract
Cardiovascular diseases (CVDs) are considered the greatest threat to human life according to World Health Organization. Early classification of CVDs and the appropriate follow-up treatment are crucial for preventing sudden deaths. Electrocardiogram (ECG) is one of the most common non-invasive tools used to evaluate the state of the heart, which can be exploited to automatically diagnose as well. However, the importance of diagnosing CVDs is varying in different context-specific scenarios. For example, ST-segment elevation (STE) is an acute myocardial infarction indicator for patients associated with chest pain and cardiac biomarker. In in-hospital healthcare, STE should be diagnosed with a higher priority than the other phenotypes of ECG. Hence, the context-specific requirements should be considered in ECG early classification problems. We formalize the ECG early classification problem as the context-specific time series classification problem. We propose a novel Constraint-based Knee-guided Neuroevolutionary Algorithm (CKNA) based on the Snippet Policy Networks V2 to solve this problem. To validate the proposed method, we perform a series of experiments on two public ECG datasets under various context-specific simulated scenarios after consulting with physicians specializing in the area. Experimental results show that CKNA significantly improves the average recall of disease classification by 5.5% compared to the competing baseline under user-specified requirements. Moreover, experimental results prove that CKNA presents a feasible solution for the early classifying of cardiac arrhythmias under different user-specified scenarios.
Yu Huang 0018, Gary G. Yen, Vincent S. Tseng
IEEE J. Biomed. Health Informatics1
2021 Stable High Utility Itemset Mining
abstract
High Utility Itemset Mining (HUIM) aims at finding all sets of items that have high importance in a database, as measured by a utility function. Although HUIM has many applications, a key limitation is that the discovered patterns often have an unstable utility over time. For example, while a set of products may yield a high utility (profit) over a year, that utility may fluctuate from weeks to weeks. To discover patterns that have a stable utility and hence that are more suitable for decision-making, this paper redefines HUIM as the task of discovering Stable High Utility Itemsets (StableHUI). An efficient tree-based and pattern-growth algorithm named Stable-Growth is proposed to extract all the StableHUI. Several experiments on two real-world datasets and two synthetic datasets show that Stable-Growth is up to 60% faster than a baseline and that it can filter out numerous unstable HUI.
Acquah Hackman, Yu Huang 0018, Philippe Fournier-Viger, Vincent S. Tseng
iiWAS2
2021 Spatio-attention embedded recurrent neural network for air quality prediction
Yu Huang 0018, Jia-Ching Ying, Vincent S. Tseng
Knowl. Based Syst.1
2021 Dynamic Graph Mining for Multi-weight Multi-destination Route Planning with Deadlines Constraints
abstract
Route planning satisfied multiple requests is an emerging branch in the route planning field and has attracted significant attention from the research community in recent years. The prevailing studies focus only on seeking a route by minimizing a single kind of Travel Cost, such as trip time or distance, among others. In reality, most users would like to choose an appropriate route, neither fastest nor shortest route. Usually, a user may have multiple requirements, and an appropriate route would satisfy all requirements requested by the user. In fact, planning an appropriate route could be formulated as a problem of Multi-weight Multi-destination Route Planning with Deadlines Constraints (MWMDRP-DC). In this article, we propose a framework, namely, MWMD-Router, which addresses the MWMDRP-DC problem comprehensively. To consider the travel costs with time-variation, we propose not only four novel dynamic graph miner to extract travel costs that reveal users’ requirements but also two new algorithms, namely, Basic MWMD Route Planning and Advanced MWMD Route Planning , to plan a route that satisfies deadline requirements and optimizes another criterion like travel cost with time-variation efficiently. To the best of our knowledge, this is the first work on route planning that considers handling multiple deadlines for multi-destination planning as well as optimizing multiple travel costs with time-variation simultaneously. Experimental results demonstrate that our proposed algorithms deliver excellent performance with respect to efficiency and effectiveness.
Yu Huang 0018, Jia-Ching Ying, Philip S. Yu, Vincent S. Tseng
ACM Trans. Knowl. Discov. Data1
2019 Mining Emerging High Utility Itemsets over Streaming Database
Acquah Hackman, Yu Huang 0018, Philip S. Yu, Vincent S. Tseng
ADMA2
2019 DeepIdentifier: A Deep Learning-Based Lightweight Approach for User Identity Recognition
Meng-Chieh Lee, Yu Huang 0018, Jia-Ching Ying, Chien Chen, Vincent S. Tseng
ADMA2
2019 Robust Sensor-based Human Activity Recognition with Snippet Consensus Neural Networks
abstract
Sensor-based human activity recognition is an important problem in pervasive computing, which has attracted lots of attention from the research community in the past few years. The existing relevant studies focused on using handcrafted features or machine learning-based methods to tackle this problem. However, these methods are usually limited to specific datasets, such that the generality is limited. Some methods are also limited to strict experimental environments, which do not take stability into consideration. In this paper, we propose a robust and novel deep learning-based framework, named Snippet Consensus Neural Networks (SCNet), which aims to conquer these challenges. Through a series of experiments, the proposed framework is verified to outperform seven state-of-the-art methods on five datasets in terms of not only accuracy but also generality and stability, averagely improving 10% on mean accuracy.
Yu Huang 0018, Meng-Chieh Lee, Vincent S. Tseng, Ching-Jui Hsiao, Chi-Chiang Huang
BSN1
2019 Long-Term Traffic Time Prediction Using Deep Learning with Integration of Weather Effect
Chih-Hsin Chou, Yu Huang 0018, Chian-Yun Huang, Vincent S. Tseng
PAKDD (2)2
2018 Mining Trending High Utility Itemsets from Temporal Transaction Databases
Acquah Hackman, Yu Huang 0018, Vincent S. Tseng
DEXA (2)2
2017 Efficient Multi-Destinations Route Planning with Deadlines and Cost Constraints
abstract
In recent years, multi-destinations route planning has been the topic of much research, which is an emerging branch of the route planning problem. The existing works have been focusing on how to find routes that minimize a single kind of trip cost, such as trip time or distance, amongst others. In fact, users may have multiple requirements in real-life multi-destinations route planning applications, including for personal or business purposes (e.g., express delivery). We observed the fact that (i) there may exist a respective deadline in reaching each of the destinations, (ii) users may consider to reduce further kinds of trip costs, such as fuel, in addition to the deadline constraint. In this paper, we address a novel route planning problem named Multi-Destinations Route Planning with Deadlines and Cost Constraints and propose two approaches, namely BMDC (Basic Multi-Destinations Route Computation) and AMDC (Advanced Multi-Destinations Route Computation) to efficiently plan a route that satisfies deadline requirements and optimizes another criterion such as trip cost. To the best of our knowledge, this is the first work on route planning that considers multiple deadlines for multi-destinations as well as optimizing trip cost, simultaneously. Experimental results demonstrate that our proposed algorithms deliver excellent performance in terms of efficiency and effectiveness.
Yu Huang 0018, Bo-Hau Lin, Vincent S. Tseng
MDM1