Mingzhu Yin

dblp:239/6470 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A cell-interacting and multi-correcting method for automatic circulating tumor cells detection
Rensheng Lai, Ling Bai, Jianxin Ji, Ruihao Qin, Lihong Jiang, Xiang Kui, Liuchao Zhang, Dimin Ning, Liuying Wang, Yujiang Chen, Xinling Wang, Menglei Hua, Yuanning Wang, Chenjing Ma, Yanyan Dai, Yongzhen Song, Hesong Wang, Lijun Fan, Mingzhu Yin
Artif. Intell. Medicine30
2024 Noise Shaping Enabled Signal Generation with Low-resolution Digital-to-Analog Converter in Optical Interconnects
abstract
Digital-to-analog converter (DAC) plays an important role in a communication system, which determines the quality of the generated signal. A high-resolution DAC is necessary to guarantee the system performance. However, the application of a high-resolution DAC results in high system cost and power consumption, which is incompatible with the future low-cost data centers. To be cost-effective and energy-efficient, a system with low-resolution DAC and negligible quantization noise induced performance degradation is highly desirable. Noise shaping (NS) technique is proposed to effectively suppress the quantization noise induced by low-resolution DAC. In this paper, we review the principle of the NS technique, and its effectiveness is investigated experimentally in 100 Gb/s intensity modulation and direct detection (IM/DD) PAM-4 system with low-resolution DAC. The experimental results show that the system with 4-bit DAC coupled with NS technique shows almost the same receiver sensitivity as the system with 8-bit DAC without NS technique over 2-km and 10-km single mode fiber (SMF) transmission in O-band. It indicates that significant performance improvement can be achieved by using NS technique in low-resolution DAC systems, and it is a promising solution for future low-cost data centers.
Wei Wang 0213, Fan Li 0011, Mingzhu Yin, Weihao Ni
GLOBECOM3
2024 Electrical Dispersion Pre-Compensation Based on Multi-Output Neural Network for 200 Gbps O-band IM-DD Short Reach Optical Systems
abstract
A non-iterative electrical dispersion pre-compensation (pre-EDC) scheme based on a multi-output neural network (NN) is proposed. The scheme is evaluated in an O-band simulation system utilizing the 1270nm optical carrier with the dispersion coefficient of -5 ps/(nm.km). The simulation results show that the multi-output NN pre-EDC realizes the same performance as the traditional iterative modified GS pre-EDC while surpassing previously proposed non-iterative pre-EDC. Additionally, the proposed multi-output NN-based pre-EDC can reduce the time consumption of the pre-EDC procedure by 50%.
Weihao Ni, Mingzhu Yin, Fan Li 0011
GLOBECOM3
2024 stAA: adversarial graph autoencoder for spatial clustering task of spatially resolved transcriptomics
abstract
With the development of spatially resolved transcriptomics technologies, it is now possible to explore the gene expression profiles of single cells while preserving their spatial context. Spatial clustering plays a key role in spatial transcriptome data analysis. In the past 2 years, several graph neural network-based methods have emerged, which significantly improved the accuracy of spatial clustering. However, accurately identifying the boundaries of spatial domains remains a challenging task. In this article, we propose stAA, an adversarial variational graph autoencoder, to identify spatial domain. stAA generates cell embedding by leveraging gene expression and spatial information using graph neural networks and enforces the distribution of cell embeddings to a prior distribution through Wasserstein distance. The adversarial training process can make cell embeddings better capture spatial domain information and more robust. Moreover, stAA incorporates global graph information into cell embeddings using labels generated by pre-clustering. Our experimental results show that stAA outperforms the state-of-the-art methods and achieves better clustering results across different profiling platforms and various resolutions. We also conducted numerous biological analyses and found that stAA can identify fine-grained structures in tissues, recognize different functional subtypes within tumors and accurately identify developmental trajectories.
Zhaoyu Fang, Ruiqing Zheng, Jin A, Mingzhu Yin, Min Li 0007
Briefings Bioinform.5
2024 Assembling spatial clustering framework for heterogeneous spatial transcriptomics data with GRAPHDeep
abstract
MOTIVATION: Spatial clustering is essential and challenging for spatial transcriptomics' data analysis to unravel tissue microenvironment and biological function. Graph neural networks are promising to address gene expression profiles and spatial location information in spatial transcriptomics to generate latent representations. However, choosing an appropriate graph deep learning module and graph neural network necessitates further exploration and investigation. RESULTS: In this article, we present GRAPHDeep to assemble a spatial clustering framework for heterogeneous spatial transcriptomics data. Through integrating 2 graph deep learning modules and 20 graph neural networks, the most appropriate combination is decided for each dataset. The constructed spatial clustering method is compared with state-of-the-art algorithms to demonstrate its effectiveness and superiority. The significant new findings include: (i) the number of genes or proteins of spatial omics data is quite crucial in spatial clustering algorithms; (ii) the variational graph autoencoder is more suitable for spatial clustering tasks than deep graph infomax module; (iii) UniMP, SAGE, SuperGAT, GATv2, GCN, and TAG are the recommended graph neural networks for spatial clustering tasks; and (iv) the used graph neural network in the existent spatial clustering frameworks is not the best candidate. This study could be regarded as desirable guidance for choosing an appropriate graph neural network for spatial clustering. AVAILABILITY AND IMPLEMENTATION: The source code of GRAPHDeep is available at https://github.com/narutoten520/GRAPHDeep. The studied spatial omics data are available at https://zenodo.org/record/8141084.
Zhaoyu Fang, Lining Zhang, Dong-Sheng Cao 0001, Min Li 0007, Mingzhu Yin
Bioinform.7
2023 Graph deep learning enabled spatial domains identification for spatial transcriptomics
abstract
Advancing spatially resolved transcriptomics (ST) technologies help biologists comprehensively understand organ function and tissue microenvironment. Accurate spatial domain identification is the foundation for delineating genome heterogeneity and cellular interaction. Motivated by this perspective, a graph deep learning (GDL) based spatial clustering approach is constructed in this paper. First, the deep graph infomax module embedded with residual gated graph convolutional neural network is leveraged to address the gene expression profiles and spatial positions in ST. Then, the Bayesian Gaussian mixture model is applied to handle the latent embeddings to generate spatial domains. Designed experiments certify that the presented method is superior to other state-of-the-art GDL-enabled techniques on multiple ST datasets. The codes and dataset used in this manuscript are summarized at https://github.com/narutoten520/SCGDL.
Zhaoyu Fang, Li-Ning Zhang, Dong-Sheng Cao 0001, Mingzhu Yin
Briefings Bioinform.6
2021 QSAR-assisted-MMPA to expand chemical transformation space for lead optimization
abstract
Matched molecular pairs analysis (MMPA) has become a powerful tool for automatically and systematically identifying medicinal chemistry transformations from compound/property datasets. However, accurate determination of matched molecular pair (MMP) transformations largely depend on the size and quality of existing experimental data. Lack of high-quality experimental data heavily hampers the extraction of more effective medicinal chemistry knowledge. Here, we developed a new strategy called quantitative structure-activity relationship (QSAR)-assisted-MMPA to expand the number of chemical transformations and took the logD7.4 property endpoint as an example to demonstrate the reliability of the new method. A reliable logD7.4 consensus prediction model was firstly established, and its applicability domain was strictly assessed. By applying the reliable logD7.4 prediction model to screen two chemical databases, we obtained more high-quality logD7.4 data by defining a strict applicability domain threshold. Then, MMPA was performed on the predicted data and experimental data to derive more chemical rules. To validate the reliability of the chemical rules, we compared the magnitude and directionality of the property changes of the predicted rules with those of the measured rules. Then, we compared the novel chemical rules generated by our proposed approach with the published chemical rules, and found that the magnitude and directionality of the property changes were consistent, indicating that the proposed QSAR-assisted-MMPA approach has the potential to enrich the collection of rule types or even identify completely novel rules. Finally, we found that the number of the MMP rules derived from the experimental data could be amplified by the predicted data, which is helpful for us to analyze the medicinal chemical rules in local chemical environment. In summary, the proposed QSAR-assisted-MMPA approach could be regarded as a very promising strategy to expand the chemical transformation space for lead optimization, especially when no enough experimental data can support MMPA.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.4
2021 ChemFLuo: a web-server for structure analysis and identification of fluorescent compounds
abstract
BACKGROUND: Fluorescent detection methods are indispensable tools for chemical biology. However, the frequent appearance of potential fluorescent compound has greatly interfered with the recognition of compounds with genuine activity. Such fluorescence interference is especially difficult to identify as it is reproducible and possesses concentration-dependent characteristic. Therefore, the development of a credible screening tool to detect fluorescent compounds from chemical libraries is urgently needed in early stages of drug discovery. RESULTS: In this study, we developed a webserver ChemFLuo for fluorescent compound detection, based on two large and high-quality training datasets containing 4906 blue and 8632 green fluorescent compounds. These molecules were used to construct a group of prediction models based on the combination of three machine learning algorithms and seven types of molecular representations. The best blue fluorescence prediction model achieved with balanced accuracy (BA) = 0.858 and area under the receiver operating characteristic curve (AUC) = 0.931 for the validation set, and BA = 0.823 and AUC = 0.903 for the test set. The best green fluorescence prediction model achieved the prediction accuracy with BA = 0.810 and AUC = 0.887 for the validation set, and BA = 0.771 and AUC = 0.852 for the test set. Besides prediction model, 22 blue and 16 green representative fluorescent substructures were summarized for the screening of potential fluorescent compounds. The comparison with other fluorescence detection tools and theapplication to external validation sets and large molecule libraries have demonstrated the reliability of prediction model for fluorescent compound detection. CONCLUSION: ChemFLuo is a public webserver to filter out compounds with undesirable fluorescent properties, which will benefit the design of high-quality chemical libraries for drug discovery. It is freely available at http://admet.scbdd.com/chemfluo/index/.
Zhi-Jiang Yang, Mingzhu Yin, Hong-Li Jiang, Aiping Lu, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.4
2021 PySmash: Python package and individual executable program for representative substructure generation and application
abstract
BACKGROUND: Substructure screening is widely applied to evaluate the molecular potency and ADMET properties of compounds in drug discovery pipelines, and it can also be used to interpret QSAR models for the design of new compounds with desirable physicochemical and biological properties. With the continuous accumulation of more experimental data, data-driven computational systems which can derive representative substructures from large chemical libraries attract more attention. Therefore, the development of an integrated and convenient tool to generate and implement representative substructures is urgently needed. RESULTS: In this study, PySmash, a user-friendly and powerful tool to generate different types of representative substructures, was developed. The current version of PySmash provides both a Python package and an individual executable program, which achieves ease of operation and pipeline integration. Three types of substructure generation algorithms, including circular, path-based and functional group-based algorithms, are provided. Users can conveniently customize their own requirements for substructure size, accuracy and coverage, statistical significance and parallel computation during execution. Besides, PySmash provides the function for external data screening. CONCLUSION: PySmash, a user-friendly and integrated tool for the automatic generation and implementation of representative substructures, is presented. Three screening examples, including toxicophore derivation, privileged motif detection and the integration of substructures with machine learning (ML) models, are provided to illustrate the utility of PySmash in safety profile evaluation, therapeutic activity exploration and molecular optimization, respectively. Its executable program and Python package are available at https://github.com/kotori-y/pySmash.
Zhi-Jiang Yang, Mingzhu Yin, Aiping Lu, Shao Liu 0002, Tingjun Hou, Dong-Sheng Cao 0001
Briefings Bioinform.4