Haoyan Wang

dblp:145/4324 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable Hardware
abstract
We present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.
Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck
ACM Trans. Reconfigurable Technol. Syst.15
2025 Mkdban-Tei: a Multi-Level Knowledge Distillation-Based Deep Learning Architecture for Predicting T Cell Receptor-Epitope Binding Specificity
abstract
Understanding the underlying mechanisms of TCR-epitope interactions is crucial for studying the adaptive immune system and promoting the field of immunotherapy. Given the high cost of traditional experimental methods, it is urgent to develop computational methods to predict TCRepitope binding. With the advancement of experimental technology, an increasing number of TCR-epitope binding pairs have been archived in public databases, creating opportunities for the advancement of computational methods. In this study, we propose a novel framework called MKDBAN-TEI for predicting TCR-epitope binding. We encode TCR and epitope sequences using a learnable residue embedding matrix and employ CNN layers to extract features. An interpretable bilinear attention network is then used to capture the interaction patterns between TCR and epitope. To improve the model's performance and generalization capability, we introduce a multi-level knowledge distillation framework: first, we cluster epitopes in the training set based on sequence similarity to define distinct domains; second, for each domain, we integrate three types of protein sequence features (protein language embeddings, physicochemical information, and evolutionary information) to train domain-specific teacher models via internal multi-feature knowledge distillation, capturing domain-specific binding patterns; finally, we distill knowledge from all domain-specific teachers to a universal student model through inter-domain knowledge distillation, enhancing generalization to unseen epitopes. Compared to several state-of-the-art models, MKDBAN-TEI demonstrates superior performance and generalization capability. Further experiments illustrate the effectiveness of the model in realworld scenarios. Visualizing the attention maps learned by MKDBAN-TEI provides new insights into TCR-epitope interactions. The code is available at: https://github.com/X/MKDBAN-TEI
Haoyan Wang, Tianyi Zang
BIBM1
2025 CKG-TPI: integrating collaborative knowledge graph with sequence interactions for TCR-peptide binding specificity
abstract
Accurately identifying interactions between T-cell receptors (TCRs) and peptides is a fundamental challenge in immunology, with significant implications for vaccine design and immunotherapy. While computational methods offer efficient alternatives to labor-intensive experimental screening, achieving robust and accurate TCR-peptide binding prediction remains a challenging task. To address this, we propose collaborative knowledge graph (CKG-TPI), a novel prediction framework based on graph neural networks that integrates both interaction patterns between TCR and peptide sequences and their higher-order biological context through a constructed collaborative knowledge graph. Experimental results on multiple publicly available independent datasets demonstrate that CKG-TPI consistently outperforms state-of-the-art models. Specifically, it achieves a 9.89% improvement in area under the ROC curve compared to the strongest baseline model UnifyImmun, and a 23.93% increase in area under the precision-recall curve over the leading baseline method. Moreover, attention weight visualization and peptide-specific TCR screening validate the model's effectiveness, underscoring its potential as a powerful tool for immunological research and therapeutic discovery.
Yue Liu 0034, Haoyan Wang, Guohua Wang 0001, Yadong Liu 0001, Tao Jiang 0021, Yadong Wang 0001
Briefings Bioinform.2
2025 SAIR-YOLO: An Improved YOLOv8 Network for Sea-Air Background IR Small-Object Detection
abstract
The performance of IR small-object detection algorithms determines the detection capability and reliability of electrooptic tracking devices in complex environments. For IR images with sea-air backgrounds, reflections on surface waves increase background thermal noise that might occlude or alter small objects, complicating their detection. Additionally, clouds and waves complicate background texture, also reducing object detectability. In this study, we propose a new sea-air background IR (SAIR) detection model on the basis of the YOLOv8 network, called SAIR-YOLO, with three major improvements. First, an asymptotic multiscale feature fusion network gradually integrates different-scale features to mitigate the semantic gap between nonadjacent features while reducing the influence of sea-air background noise on feature representation. Second, a strengthened detection head discriminates irrelevant background features and focuses network attention on the objects. Third, a hybrid intersection-over-union (IoU) loss function improves detection performance, by focusing on shape similarities, and expands the effective regression range. Experimental results yield average SAIR-YOLO precisions of 80.2%, 84.4%, and 96.4% for three distinct datasets: a custom dataset and the SIRST-V2 and NUDT-SIRST publicly available datasets. This represents improvements of 7.0%, 4.9%, and 0.7%, respectively, on the YOLOv8 model.
Yue Yang 0020, Haoyan Wang, Peijie Pang
IEEE Geosci. Remote. Sens. Lett.2
2024 MFTEP: A Multimodal Fusion Deep Learning Framework for T Cell Receptor-epitope Interaction Prediction
abstract
Accurately predicting immunogenic peptides recognized by T cell receptors (TCR) is a crucial step toward personalized immunotherapy. However, prediction of TCR-epitope interactions is still a challenging task. Early works, like molecular dynamics simulation-based methods, suffer from slow speed and poor generalization capabilities. It is necessary to develop novel computational methods to predict TCR-epitope interactions precisely. With the development of high-throughput sequencing technologies, more and more TCR-epitope interaction data have been recorded in public databases. With the help of these databases, many in silico predictive methods have shown promising performance. However, current methods still perform poorly on unseen TCRs and epitopes. Moreover, most current models still accept single-modal information about TCRs and epitopes, such as sequences or physicochemical information. Effectively utilizing the multimodal information of TCRs and epitopes, such as molecular graphs and 3D structure, may enhance the model’s prediction performance. To address the above issues, we presented MFTEP, a multimodal fusion method to predict TCR-epitope interactions by fusing the sequence features, molecular graph features, and 3D structure features of TCRs and epitopes. The ablation study highlights the importance of the multi-modal fusion module in enhancing the model’s performance. Several datasets were collected and utilized to evaluate the generality and robustness of the proposed model. According to the experimental results, MFTEP performs better than other state-of-the-art methods, indicating its high predictive power. Overall, the results demonstrate that MFTEP can learn general TCR-epitope interaction patterns and is a powerful prediction tool to apply to real-world scenarios. All the data and code are available at: https://github.com/skybluewhy/MFTEP
Haoyan Wang, Tianyi Zang
BIBM1
2024 TDLM: A Diffusion Language Model for TCR Sequence Exploration and Generation
abstract
The adaptive immune response relies on the ability of T-cell receptors (TCRs) to recognize specific antigens. The vast diversity of TCRs allows T-cells to recognize a broad spectrum of antigens, but this complexity also poses challenges for understanding and predicting TCR-antigen binding specificity. Despite the development of various machine learning and deep learning methods for prediction and clustering, there remains a need for a versatile and effective TCR language framework that can be flexibly applied to various downstream tasks, including sequence generation. Here we present TDLM, a T-cell Receptor (TCR) diffusion language model, designed to decode complex patterns within TCR sequences and apply them across various downstream tasks. Firstly, TDLM can be trained on unlabeled TCR sequence data, enabling it to utilize vast datasets to generate comprehensive embeddings. When compared to other embedding methods, TDLM embeddings enhance TCR-antigen binding prediction accuracy and enable effective TCR sequence clustering and similarity analysis, helping identify TCRs with shared antigen specificity. Furthermore, as a diffusion-based generative model, TDLM can generate highly diverse and specific TCR sequences. This ability is invaluable for the rapid screening and optimization of TCRs with target antigen specificities, offering significant potential in disease diagnosis, personalized immunotherapy, and vaccine research. The code is available at: https://github.com/skybluewhy/TDLM
Haoyan Wang, Tianyi Zang, Yadong Liu 0001
BIBM1
2023 TetraCVD: A Temporal-Textual Transformer based Model for Cardiovascular Disease Diagnosis
abstract
Cardiovascular disease (CVD) is one of the leading causes of death globally. There is considerable clinical significance and an emerging need of assisting doctors to diagnose cardiovascular disease and identify the subtype of it, from which doctors can provide different treatments and medications to increase the cure rate. The goal of this paper is to develop a deep learning model to predict cardiovascular disease and classify its subtype, which by handling data from two modalities of time-series vital signs and text report. We propose a temporal-textual transformer based model for cardiovascular disease diagnosis, TetraCVD, to address the challenges of irregular temporal feature extraction and medical long-text feature extraction respectively. TetraCVD is a multimodal deep learning model, consisting of two networks, cvdGNN and cvdHierBERT, as its time-series and language backbones, which leverage knowledge from temporal vital signs and text reports of the individuals respectively. Our results show that TetraCVD achieves promising performance in predicting subtypes of cardiovascular disease using the P18-ECER dataset and obtains state-of-the-art results. This study is among the first efforts that use both time-series vital signs and text report data to predict cardiovascular disease and its subtype. We argue that our approach can be generalized to predict and diagnose other diseases easily, and it can potentially play a significant role in the domain of general disease diagnosis in the future.
Kailong Lu, Penghuan Gu, Haoyan Wang, Tianyi Zang
BIBM4
2023 A Path Increment Map Matching Method for High-Frequency Trajectory
abstract
Aiming at the problems of low matching accuracy and slow matching speed of high-frequency trajectory data in complex urban road networks, this paper proposes a matching method based on path increment. This method consists of two parts: combined filtering and incremental matching. Firstly, the road network is simplified through combined filtering, and then the incremental matching is carried out by taking the paths as increments. In the matching procedure, a comprehensive evaluation scheme of similarity based on distance factor and curvature is adopted. The above measures effectively reduce the impact of complex road segments on the matching results, while the path increment method enables the matching process to be executed more rapidly and accurately. The experiments were conducted using the Geolife datasets. The results show that our algorithm has obvious advantages over similar algorithms in terms of matching accuracy and efficiency, and shows good stability in road matching tests with different complexity.
Haoyan Wang, Yuangang Liu, Bo Liang 0010, Zongyi He
IEEE Trans. Intell. Transp. Syst.1
2022 MHCRoBERTa: pan-specific peptide-MHC class I binding prediction through transfer learning with label-agnostic protein sequences
abstract
Predicting the binding of peptide and major histocompatibility complex (MHC) plays a vital role in immunotherapy for cancer. The success of Alphafold of applying natural language processing (NLP) algorithms in protein secondary struction prediction has inspired us to explore the possibility of NLP methods in predicting peptide-MHC class I binding. Based on the above motivations, we propose the MHCRoBERTa method, RoBERTa pre-training approach, for predicting the binding affinity between type I MHC and peptides. Analysis of the results on benchmark dataset demonstrates that MHCRoBERTa can outperform other state-of-art prediction methods with an increase of the Spearman rank correlation coefficient (SRCC) value. Notably, our model gave a significant improvement on IC50 value. Our method has achieved SRCC value and AUC value as 0.785 and 0.817, respectively. Our SRCC value is 14.3% higher than NetMHCpan3.0 (the second highest SRCC value on pan-specific) and is 3% higher than MHCflurry (the second highest SRCC value on all methods). The AUC value is also better than any other pan-specific methods. Moreover, we visualize the multi-head self-attention for the token representation across the layers and heads by this method. Through the analysis of the representation of each layer and head, we can show whether the model has learned the syntax and semantics necessary to perform the prediction task well. All these results demonstrate that our model can accurately predict the peptide-MHC class I binding affinity and that MHCRoBERTa is a powerful tool for screening potential neoantigens for cancer immunotherapy. MHCRoBERTa is available as an open source software at github (https://github.com/FuxuWang/MHCRoBERTa).
Fuxu Wang, Haoyan Wang, Lizhuang Wang, Haoyu Lu, Shizheng Qiu, Tianyi Zang, Xinjun Zhang, Yang Hu 0008
Briefings Bioinform.2