Yanli Zhao

dblp:05/6435 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LoKA: Low-Precision Kernel Applications for Recommendation Models at Scale
Yinbin Ma, Quanyu Zhu, Vasiliy Kuznetsov, Yuxin Chen 0001, Jiecao Yu, Buyun Zhang, Tongyi Tang, Xiaohan Wei, Yanli Zhao, Zeliang Chen, Yuchen Hao, Venkatesh Ranganathan, Sandeep Parab, Yantao Yao, Maxim Naumov, Chunzhi Yang, Ellie Wen, Chunqiang Tang
ISCA11
2026 SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
abstract
Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from non-negligible costs and missing co-design opportunities. Such inefficiency makes them difficult to support complex model architectures, such as learned similarities and multi-task retrieval. In this paper, we present SilverTorch, a model-based serving system that brings all components into one unified model. It unifies model serving by replacing standalone indexing and filtering services with model layers. We propose a model-based GPU Bloom index for feature filtering and a fused Int8 ANN kernel for nearest neighbor search. Through co-design of the ANN search and feature filtering, we reduce GPU memory usage and eliminate computation. Benefiting from this design, we scale up retrieval by introducing an OverArch scoring layer and a multi-task retrieval with a Value Model to aggregate scores. These advancements improve the retrieval accuracy and enable future studies for serving more complex models. Our evaluation on industry-scale datasets shows that SilverTorch achieves up to 23.7× higher throughput compared to the state-of-the-art approaches. We also demonstrate that SilverTorch's solution is 13.35× more cost-efficient than CPU-based solution while improving accuracy via serving more complex models.
Bi Xue, Xiaoheng Mao, Xialu Li, Rui Jian, Yanli Zhao, Yanzun Huang, Yijie Deng, Harry Tran, Ryan Chang, Eric Dong, Jiazhou Wang, Keke Zhai, Hongzhang Yin, Pawel Garbacki, Zheng Fang 0009, Yiyi Pan, Min Ni
SIGIR14
2025 Multi-Channel Learning Framework Based on Relation-Aware Transformer For Drug Synergy Prediction
abstract
Combined therapeutic strategies have demonstrated their importance in addressing complex diseases, particularly among patient populations where monotherapy is less effective. Compared to single drug treatments, the use of drug combinations can reduce the occurrence of drug resistance and enhance the efficacy of cancer treatments. Therefore, developing effective combination therapies through clinical trials holds significant importance for researchers and society as a whole. However, facing a vast library of compounds, conducting high-throughput drug combination screening is not only challenging but also costly. To overcome these difficulties, researchers have developed various computational methods that utilize biomedical data related to drugs to efficiently identify potential drug combinations. This study has developed a novel multi-channel representation learning framework based on knowledge graphs and Transformers models, aimed at predicting drug synergy. Unlike full-graph analysis, we employ a node-centric sampling strategy to extract subgraphs for learning node representations. Furthermore, we utilize a relation-based self-attention mechanism to handle complex relations between nodes. Through the multi-channel network, we input drug molecular structures, cell line information, and biomedical knowledge graphs into different channels, and fuse the feature representations of these channels to enhance the accuracy of drug synergy prediction. Our model benefits from both the structural features of drugs and rich biomedical background information. Extensive experiments on two representative databases have validated the effectiveness of our model.
Kaiyuan Zhang 0006, Tianyi Zang, Yanli Zhao
IJCNN5
2025 A novel enhanced Bayesian classifier with multiple smoothing parameters
Yanli Zhao, Guang Yang 0005, Huihui Wei
Knowl. Inf. Syst.1
2025 DECK: Experiences on Delta Checkpointing for Industrial Recommendation Systems
abstract
In large-scale industrial recommendation systems, model checkpoints are instrumental in maintaining training goodput and numerical correctness during system failures and job preemptions. The increasing prevalence of multi-terabyte models has rendered frequent regular model checkpoints impractical, resulting in substantial lost progress when recovering from failures. As model sizes continue to grow, researchers and practitioners are compelled to investigate more efficient and scalable solutions. This paper presents DECK, a novel approach to delta model checkpointing designed for real-world industrial systems. Specifically, DECK focuses on extracting delta states with near-zero overhead, staging and streaming delta checkpoints without interrupting the training process, and merging delta checkpoints in an optimal and decoupled manner. Experimental results demonstrate that DECK achieves a 12-fold increase in checkpoint frequency while maintaining negligible impact on training throughput, thereby attaining state-of-the-art (SOTA) production performance.
Sibasish Acharya, Sihui Han, Yongxiong Ren, Yanli Zhao, Chucheng Wang, Pradeep Fernando, Siqi Yan, Yicong Du, Elzbieta Krepska, Intaik Park, Min Ni, Qunshu Zhang
Proc. VLDB Endow.5
2024 Wukong: Towards a Scaling Law for Large-Scale Recommendation
abstract
Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong’s unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong’s scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
Buyun Zhang, Yuxin Chen 0001, Jade Nie, Xi Liu 0011, Yanli Zhao, Yuchen Hao, Yantao Yao, Ellie Wen, Jongsoo Park, Maxim Naumov
ICML7
2024 BI²Net: Graph-Based Boundary-Interior Interaction Network for Raft Aquaculture Area Extraction From Remote Sensing Images
abstract
Accurate monitoring of raft aquaculture areas (RAAs) is particularly important for the protection of marine ecosystems. However, existing semantic segmentation methods are often degraded by severe shrinkage when extracting inapparent RAAs caused by natural factors such as tide level changes and human activities including laver harvesting. In this letter, a graph-based boundary-interior interaction network (BI2Net) is proposed for laver RAA extraction. Graph convolution based on soft clustering is introduced to capture the global distribution patterns among RAAs. For the network structure, we designed the RAA-boundary branch and RAA-interior branch, one for locating the boundaries and the other for extracting the RAAs. Importantly, unlike traditional dual-branch methods that only use additional information from the auxiliary tasks to assist with the primary task, BI2Net introduces a graph interaction module (GIM) that implements reasoning about the relationship between differ-ent distribution patterns, which sufficiently and comprehensively exploits mutual benefits between RAA extraction and boundary detection. After the addition of GIM, precision, recall, F1 and IoU were improved by 2.0%, 8.2%, 0.060 and 8.2% respectively.Extensive experiments have shown that our proposed BI2Net outperforms existing methods in terms of the consistency and completeness of RAA extraction, with precision, recall, F1 and IoU reached 93.0%, 90.6%, 0.916 and 84.7%, respectively.
Yan Lu 0014, Yuchao Zhao, Mingkai Yang, Yanli Zhao, Binge Cui
IEEE Geosci. Remote. Sens. Lett.4
2023 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
abstract
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.
Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews
Proc. VLDB Endow.1
2022 An Imbalance Modified Convolutional Neural Network With Incremental Learning for Chemical Fault Diagnosis
abstract
Fault diagnosis that identifies the root of the abnormal status is of great importance to eliminate faults in the complex chemical processes. Many data-driven fault diagnosis models ignore different faults that occur with varied frequencies in chemical plants, and they need a complete retraining process with the arrival of new fault modes. In this article, a novel incremental imbalance modified convolutional neural network is proposed to solve the aforementioned issues. The proposed method employs an imbalance modified method to extract the valuable information from the imbalance data, and generate new samples. After that, a local hyperplane-based dynamic Relief is designed to reduce the dimension of the chemical data and simplify the complex learning process. Finally, for the arrival of new fault modes, the proposed method is prompted in an incremental hierarchical way. Unlike the traditional models that are trained on static data, the proposed method inherits the existing knowledge and updates itself to include new coming fault classes. The proposed method is utilized in a simulated process and a real industrial process. Experimental results illustrate that the proposed method is better than the existing methods and has significant robustness and reliability in chemical fault diagnosis.
Xiaohua Gu, Yanli Zhao, Guang Yang 0005, Lusi Li
IEEE Trans. Ind. Informatics2
2020 PyTorch Distributed: Experiences on Accelerating Data Parallel Training
abstract
This paper presents the design, implementation, and evaluation of the PyTorch distributed data parallel module. Py-Torch is a widely-adopted scientific computing package used in deep learning research and applications. Recent advances in deep learning argue for the value of large datasets and large models, which necessitates the ability to scale out model training to more computational resources. Data parallelism has emerged as a popular solution for distributed training thanks to its straightforward principle and broad applicability. In general, the technique of distributed data parallelism replicates the model on every computational resource to generate gradients independently and then communicates those gradients at each iteration to keep model replicas consistent. Despite the conceptual simplicity of the technique, the subtle dependencies between computation and communication make it non-trivial to optimize the distributed training efficiency. As of v1.5, PyTorch natively provides several techniques to accelerate distributed data parallel, including bucketing gradients, overlapping computation with communication, and skipping gradient synchronization. Evaluations show that, when configured appropriately, the PyTorch distributed data parallel module attains near-linear scalability using 256 GPUs.
Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li 0009, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, Soumith Chintala
Proc. VLDB Endow.2
2015 Detecting temporal protein complexes based on Neighbor Closeness and time course protein interaction networks
abstract
The detection of temporal protein complexes would be a great aid in furthering our knowledge of the dynamic features and molecular mechanism in cell life activities. Inspired by the idea of that the tighter a protein's neighbors inside a module connect, the greater the possibility that the protein belongs to the module, we propose a novel clustering algorithm CNC (Clustering based on Neighbor Closeness) and apply it to the time course protein interaction networks (TCPINs) to detect temporal protein complexes. Our novel algorithm has better performance on identifying protein complexes than five state-of-the-art algorithms—Hunter, MCODE, CFinder, SPICI, and ClusterONE—in terms of matching degree and accuracy metric, meanwhile it obtains many protein complexes with strong biological significance.
Xianjun Shen, Xingpeng Jiang, Yanli Zhao, Tingting He 0003, Jincai Yang
BIBM4
2014 An efficient protein complex mining algorithm based on Multistage Kernel Extension
abstract
BACKGROUND: In recent years, many protein complex mining algorithms, such as classical clique percolation (CPM) method and markov clustering (MCL) algorithm, have developed for protein-protein interaction network. However, most of the available algorithms primarily concentrate on mining dense protein subgraphs as protein complexes, failing to take into account the inherent organizational structure within protein complexes. Thus, there is a critical need to study the possibility of mining protein complexes using the topological information hidden in edges. Moreover, the recent massive experimental analyses reveal that protein complexes have their own intrinsic organization. METHODS: Inspired by the formation process of cliques of the complex social network and the centrality-lethality rule, we propose a new protein complex mining algorithm called Multistage Kernel Extension (MKE) algorithm, integrating the idea of critical proteins recognition in the Protein- Protein Interaction (PPI) network,. MKE first recognizes the nodes with high degree as the first level kernel of protein complex, and then adds the weighted best neighbour node of the first level kernel into the current kernel to form the second level kernel of the protein complex. This process is repeated, extending the current kernel to form protein complex. In the end, overlapped protein complexes are merged to form the final protein complex set. RESULTS: Here MKE has better accuracy compared with the classical clique percolation method and markov clustering algorithm. MKE also performs better than the classical clique percolation method both on Gene Ontology semantic similarity and co-localization enrichment and can effectively identify protein complexes with biological significance in the PPI network.
Xianjun Shen, Yanli Zhao, Tingting He 0003, Jincai Yang, Xiaohua Hu 0001
BMC Bioinform.2
2013 An efficient protein complex mining algorithm based on multistage kernel extension
abstract
Inspired by the formation process of cliques of the complex social network and the centrality-lethality rule, and integrating the idea of critical proteins recognition in the Protein-Protein Interaction (PPI) network, we propose a new protein complex mining algorithm called MKE (Multistage Kernel Extension). MKE first recognizes the nodes with high degree as the first level kernel of protein complex, then adds the weighted best neighbor node of the first level kernel into the current kernel to form the second level kernel of the protein complex, this process is repeated, extending the current kernel to form protein complex. Overlapped protein complexes are merged to form the final protein complex set. The results show that MKE has better accuracy compared with the classical clique percolation method. MKE also performs better than markov clustering algorithm on Gene Ontology semantic similarity and co-localization enrichment and can effectively identify protein complexes with biological significance.
Xianjun Shen, Yanli Zhao, Jincai Yang
BIBM2
2013 An integrated approach to identify protein complex based on best neighbor and modularity increment
abstract
In order to overcome the limitations of global modularity and the deficiency of local modularity, we introduce a hybrid modularity measure LGQ (Local-Global Quantification) which adopts a suitable modularity adjustable parameter to control the balance of global detecting capability and local search capability in Protein-Protein Interaction (PPI) network. On the other hand, a new protein complex mining algorithm called BN-LGQ has been proposed, which integrates the definitions of best neighbor node and the modularity increment. And by comparison with other known algorithms, the experimental results show BN-LGQ performs a better accuracy on predicting protein complexes and has a higher match with the reference protein complexes. Moreover, it can identify protein complexes with better biological significance in PPI network.
Xianjun Shen, Yanli Zhao, Jincai Yang, Tingting He 0003, Xiaohua Hu 0001
BIBM2
2013 A parallel computing approach to viewshed analysis of large terrain data using graphics processing units
abstract
Viewshed analysis, often supported by geographic information system, is widely used in many application domains. However, as terrain data continue to become increasingly large and available at high resolutions, data-intensive viewshed analysis poses significant computational challenges. General-purpose computation on graphics processing units (GPUs) provides a promising means to address such challenges. This article describes a parallel computing approach to data-intensive viewshed analysis of large terrain data using GPUs. Our approach exploits the high-bandwidth memory of GPUs and the parallelism of massive spatial data to enable memory-intensive and computation-intensive tasks while central processing units are used to achieve efficient input/output (I/O) management. Furthermore, a two-level spatial domain decomposition strategy has been developed to mitigate a performance bottleneck caused by data transfer in the memory hierarchy of GPU-based architecture. Computational experiments were designed to evaluate computational performance of the approach. The experiments demonstrate significant performance improvement over a well-known sequential computing method, and an enhanced ability of analyzing sizable datasets that the sequential computing method cannot handle.
Yanli Zhao, Anand Padmanabhan, Shaowen Wang 0001
Int. J. Geogr. Inf. Sci.1
2008 The Measurement on the Dielectric Properties of Fresh-Water Ice with Rectangular Waveguide at 2.6GHz-3.9GHz
abstract
The complex dielectric permittivity of pure ice is measured using the transmission/reflection method in the rectangular waveguide at frequencies between 2.6GHz and 3.9GHz and over the temperature range from -25 to -2.5°C, in order to extract the influence of temperature and frequency on the dielectric properties of ice quantitatively. The S band of microwave frequency is particularly investigated because of its importance and specialty. The experiments in this study show that the real part of the complex permittivity of pure ice is around 3.16, independent of frequency. The imaginary part changes over the frequency with the nonlinear function and the special frequency point 3.3GHz is found. Two polynomial functions of temperature are selected for the real part and the imaginary part, respectively. The real part increases firstly and then decreases with the increasing temperature, while the imaginary part just monotonously increases with the increasing temperature.
Yanli Zhao, Yan Chen 0003, Ling Tong 0001, Mingquan Jia
IGARSS (4)1
2003 Modified Nonparametric Approaches to Detecting Differentially Expressed Genes in Replicated Microarray Experiments
abstract
MOTIVATION: An important goal in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Various parametric tests, such as the two-sample t-test, have been used, but their possibly too strong parametric assumptions or large sample justifications may not hold in practice. As alternatives, a class of three nonparametric statistical methods, including the empirical Bayes method of Efron et al. (2001), the significance analysis of microarray (SAM) method of Tusher et al. (2001) and the mixture model method (MMM) of Pan et al. (2001), have been proposed. All the three methods depend on constructing a test statistic and a so-called null statistic such that the null statistic's distribution can be used to approximate the null distribution of the test statistic. However, relatively little effort has been directed toward assessment of the performance or the underlying assumptions of the methods in constructing such test and null statistics. RESULTS: We point out a problem of a current method to construct the test and null statistics, which may lead to largely inflated Type I errors (i.e. false positives). We also propose two modifications that overcome the problem. In the context of MMM, the improved performance of the modified methods is demonstrated using simulated data. In addition, our numerical results also provide evidence to support the utility and effectiveness of MMM.
Yanli Zhao, Wei Pan 0011
Bioinform.1