Ziwei Yang 0002

dblp:95/11411-2 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-9846-840XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Targeted Pathway Inference for Biological Knowledge Bases via Graph Learning and Explanation
abstract
Retrieving targeted pathways in biological knowledge bases, particularly when incorporating wet-lab experimental data, remains a challenging task and often requires downstream analyses and specialized expertise. In this paper, we frame this challenge as a solvable graph learning and explaining task and propose a novel subgraph inference framework, ExPath, that explicitly integrates experimental data to classify various graphs (bio-networks) in biological databases. The links (representing pathways) that contribute more to classification can be considered as targeted pathways. Our framework can seamlessly integrate biological foundation models to encode the experimental molecular data. We propose ML-oriented biological evaluations and a new metric. The experiments involving 301 bio-networks evaluations demonstrate that pathways inferred by ExPath are biologically meaningful, achieving up to 4.5× higher Fidelity+ (necessity) and 14× lower Fidelity- (sufficiency) than explainer baselines, while preserving signaling chains up to 4× longer.
Rikuto Kotoge, Ziwei Yang 0002, Zheng Chen 0012, Yushun Dong, Yasuko Matsubara, Jimeng Sun 0001, Yasushi Sakurai
AAAI2
2025 GeSubNet: Gene Interaction Inference for Disease Subtype Network Generation
abstract
Retrieving gene functional networks from knowledge databases presents a challenge due to the mismatch between disease networks and subtype-specific variations. Current solutions, including statistical and deep learning methods, often fail to effectively integrate gene interaction knowledge from databases or explicitly learn subtype-specific interactions. To address this mismatch, we propose GeSubNet, which learns a unified representation capable of predicting gene interactions while distinguishing between different disease subtypes. Graphs generated by such representations can be considered subtype-specific networks. GeSubNet is a multi-step representation learning framework with three modules: First, a deep generative model learns distinct disease subtypes from patient gene expression profiles. Second, a graph neural network captures representations of prior gene networks from knowledge databases, ensuring accurate physical gene interactions. Finally, we integrate these two representations using an inference loss that leverages graph generation capabilities, conditioned on the patient separation loss, to refine subtype-specific information in the learned representation. GeSubNet consistently outperforms traditional methods, with average improvements of 30.6%, 21.0%, 20.1%, and 56.6% across four graph evaluation metrics, averaged over four cancer datasets. Particularly, we conduct a biological simulation experiment to assess how the behavior of selected genes from over 11,000 candidates affects subtypes or patient distributions. The results show that the generated network has the potential to identify subtype-specific genes with an 83% likelihood of impacting patient distribution shifts.
Ziwei Yang 0002, Zheng Chen 0012, Xin Liu 0020, Rikuto Kotoge, Peng Chen 0035, Yasuko Matsubara, Yasushi Sakurai, Jimeng Sun 0001
ICLR1
2025 DBgDel: Database-Enhanced Gene Deletion Framework for Growth-Coupled Production in Genome-Scale Metabolic Models
abstract
When simulating metabolite productions with genome-scale constraint-based metabolic models, gene deletion strategies are necessary to achieve growth-coupled production, which means cell growth and target metabolite production occur simultaneously. Since obtaining gene deletion strategies for large genome-scale models suffers from significant computational time, it is necessary to develop methods to mitigate this computational burden. In this study, we introduce a novel framework for computing gene deletion strategies. The proposed framework first mines related databases to extract prior information about gene deletions for growth-coupled production. It then integrates the extracted information with downstream algorithms to narrow down the algorithmic search space, resulting in highly efficient calculations on genome-scale models. Computational experiment results demonstrated that our framework can compute stoichiometrically feasible gene deletion strategies for numerous target metabolites, showcasing a noteworthy improvement in computational efficiency. Specifically, our framework achieves an average 6.1-fold acceleration in computational speed compared to existing methods while maintaining a respectable success rate.
Ziwei Yang 0002, Takeyuki Tamura
IEEE Trans. Comput. Biol. Bioinform.1
2025 DeepGDel: Deep Learning-Based Gene Deletion Prediction Framework for Growth-Coupled Production in Genome-Scale Metabolic Models
abstract
In genome-scale constraint-based metabolic models, gene deletion strategies are crucial for achieving growth-coupled production, where cell growth and target metabolite production are simultaneously achieved. While computational methods for calculating gene deletions have been widely explored and have contributed to gene deletion strategy databases, current approaches remain computationally demanding and have yet to fully leverage emerging data-driven paradigms, such as machine learning, for more efficient strain design. Therefore, it is necessary to propose a fundamental framework for this objective. In this study, we first formulate the problem of gene deletion strategy prediction and then propose a framework for predicting gene deletion strategies for growth-coupled production in genome-scale metabolic models. The proposed framework leverages deep learning algorithms to learn and integrate sequential gene and metabolite data representation, enabling the automatic gene deletion strategy prediction. Computational experiment results demonstrate the feasibility of the proposed framework, showing substantial improvements over baseline methods. Specifically, the proposed framework achieves a 14.69%, 22.52%, and 13.03% increase in overall accuracy across three metabolic models of different scales under study, while maintaining balanced precision and recall in predicting gene deletion statuses.
Ziwei Yang 0002, Takeyuki Tamura
IEEE Trans. Comput. Biol. Bioinform.1
2023 MFHCC: Multi-View Feature Hierarchical Contrastive Clustering Model for Multi-Omics Data
abstract
Comprehensive analysis of multi-omics data has now garnered significant attention. However, due to the diversity of multi-omics data, integrating multi-omics information presents a formidable challenge for researchers. Moreover, the high dimensionality and sparsity characteristics in omics data further complicate multi-omics data analysis. To address these challenges and obtain high-quality representations suitable for downstream tasks, we propose a self-supervised clustering learning framework called the Multi-view Feature Hierarchical Contrastive Clustering model (MFHCC) to extract multi-level features. Firstly, the proposed model considers multi-omics as multi-modality and employs an autoencoder for each modality to integrate diverse omics information simultaneously. Secondly, it utilizes a multilevel feature extraction framework with contrastive learning methods to mitigate the impact of redundant information and null values on representation quality while capturing semantic information embedded in the data. Additionally, the model incorporates a deep clustering module to guide the representation toward downstream tasks while integrating high-level features for guidance. Through extensive experiments conducted on pan-cancer datasets, we validate the effectiveness of MFHCC. For instance, the model achieves an accuracy exceeding 76% by omics types, thus confirming its superior performance.
Zongli Jiang, Ziwei Yang 0002, Jinli Zhang, Zheng Chen 0012
BIBM3
2023 MoCLIM: Towards Accurate Cancer Subtyping via Multi-Omics Contrastive Learning with Omics-Inference Modeling
Ziwei Yang 0002, Zheng Chen 0012, Yasuko Matsubara, Yasushi Sakurai
CIKM1
2022 Automatic Sleep Staging via Frequency-Wise Spiking Neural Networks
abstract
Identifying sleep stages is a fundamental step for both early detection of disease and neuroscientific exploration. Automatic sleep staging is classic research for replacing the time-consuming gold-standard manual staging procedure. Recently, promising results have been achieved on automatic staging by extracting spatio-temporal features via deep neural networks from electroencephalogram (EEG). However, such methods fail to consistently yield good performance due to a missing piece in data representation: the dynamic fluctuations of neurons on top of EEG features that is non-trivial for automatic sleep staging task. This paper introduces a biomimicry spiking neural network (SNN) to map the aforementioned features serving for automatic sleep staging. Such SNNs are designed as an array of encoders that converts frequency-specific features into long-term spiking coding and then a popular ANN model is used as the staging machine by absorbing the stage-dependent spiking representation. For proof-of-concept, the performance of the proposed framework is demonstrated by introducing multiple sleep datasets. The experimental results showed that the proposed method achieved a competitive stage scoring performance, especially for Wake, N2, and N3, with higher Precision of 0.94, 0.87, and 0.86. Moreover, the ablation studies prove the SNN has the potential for extracting the neuron’s variation features.
Haohui Jia, Ziwei Yang 0002, Pei Gao, Man Wu, Chen Li 0027, Yirong Kan
BIBM2
2022 Hierarchical Categorical Generative Modeling for Multi-omics Cancer Subtyping
abstract
Identifying a specific cancer subtype from a variety of candidates is vital for precise and effective treatment. However, cancer subtyping is highly non-trivial as a result of cancer heterogeneity. While significant efforts have been put into understanding the mechanism of cancer subtypes via studying the omics data, existing methods run the risk of presenting biased analyses resulted from overfitting the high-dimensional and scarce omics data. In this paper, we propose a novel generative model that directly models the cancer data distribution by which downstream tasks can circumvent the curse of overfitting and achieve better performance. Unlike conventional generative modeling schemes, the proposed method underlines hierarchical categorical latent spaces to extract global features and local details respectively from transcriptomics and genomics profiles, which is the first to be considered in the cancer subtyping literature. By extensive experiments we verify that the proposed architecture achieves more clearly separated subtypes, as well as medically significant insights into real subtyping.
Ziwei Yang 0002, Lingwei Zhu, Chen Li 0027, Zheng Chen 0012, Naoki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM1
2022 Multi-Tier Platform for Cognizing Massive Electroencephalogram
abstract
An end-to-end platform assembling multiple tiers is built for precisely cognizing brain activities. Being fed massive electroencephalogram (EEG) data, the time-frequency spectrograms are conventionally projected into the episode-wise feature matrices (seen as tier-1). A spiking neural network (SNN) based tier is designed to distill the principle information in terms of spike-streams from the rare features, which maintains the temporal implication in the nature of EEGs. The proposed tier-3 transposes time- and space-domain of spike patterns from the SNN; and feeds the transposed pattern-matrices into an artificial neural network (ANN, Transformer specifically) known as tier-4, where a special spanning topology is proposed to match the two-dimensional input form. In this manner, cognition such as classification is conducted with high accuracy. For proof-of-concept, the sleep stage scoring problem is demonstrated by introducing multiple EEG datasets with the largest comprising 42,560 hours recorded from 5,793 subjects. From experiment results, our platform achieves the general cognition overall accuracy of 87% by leveraging sole EEG, which is 2% superior to the state-of-the-art. Moreover, our developed multi-tier methodology offers visible and graphical interpretations of the temporal characteristics of EEG by identifying the critical episodes, which is demanded in neurodynamics but hardly appears in conventional cognition scenarios.
Zheng Chen 0012, Lingwei Zhu, Ziwei Yang 0002
IJCAI3
2022 Automated Cancer Subtyping via Vector Quantization Mutual Information Maximization
Zheng Chen 0012, Lingwei Zhu, Ziwei Yang 0002, Takashi Matsubara 0001
ECML/PKDD (1)3
2021 An End-to-End Sleep Staging Simulator Based on Mixed Deep Neural Networks
abstract
Sleep screening is not only a major tool in the assessment of pathophysiology, but also a bridge between the central neuronal systems and behaviour/cognition. Automatic sleep staging is an alternative for the time-consuming gold standard manual scoring procedure. Most of the existing works designed such procedure by using deep neural networks without considering the medical criterion of the sleep staging task. We argue that capturing the stage-specific features which meet the criterion is of significant importance for the automatic sleep staging alternative. In this work we propose an end-to-end sleep staging simulator based on mixed neural networks, i.e., CNN, LSTM, and Transformer. The framework consists of two subnetworks: stage dependent feature mapping network which is constructed by the idea of physiological sleep nature, and an attention-based parallel staging network. Moreover, we adopt a mixed precision training strategy to quantize the model for exploring feasible usage in the clinical settings. Through an experiment with a large EEG database (Sleep Heart Health Study), the proposed method has a competitive stage scoring performance, especially in stages Wake, N2, and N3, with higher precision of 0.92, 0.85, and 0.86, respectively. Our study proves that the quantized model has potential capability for further application in the clinical staging task.
Zheng Chen 0012, Ziwei Yang 0002, Dong Wang 0044, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM2
2021 Exploring Feasibility of Truth-Involved Automatic Sleep Staging Combined with Transformer
abstract
Recently, deep learning-based methods have been successfully proposed for electrophysiology signal-based sleep staging with promising results. Most existing methods use convolutional layers and recurrent-based architectures to implement a model structure from feature extraction to sequence signal classification. In this study, we propose a method of segmenting electroencephalogram (EEG) and electrooculogram (EOG) data according to frequency bands and construct a Transformer based automatic sleep classification model on top of it. The results show that the classifications of the stage Wake, N3, and REM outperform the state-of-art works, with the Fl-scores of 0.92, 0.85 and 0.91. Our work is the first attempt to explore the feasibility of a truth-involved Transformer-based model with a large-scale sleep database.
Ziwei Yang 0002, Dong Wang 0044, Zheng Chen 0012, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM1