Yan Zhang 0100

dblp:04/3348-100 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1729-3976ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiGen: Active Lipid Generation via a Molecular Language Model
abstract
Lipid nanoparticles (LNPs) can deliver cargos to both tumor and immune cells, playing a crucial role in biomedicine.Traditional approaches rely on experimental screening and expert knowledge, which can be costly and time-consuming.Recent methods based on language models have accelerated this process using deep learning.Although these methods can retrieve molecules for fusion or rank candidates from existing libraries, they are still limited by the scope of known formulations.In this work, we propose LiGen to generate lipid molecules efficiently and actively, facilitating the discovery of high-performing LNP formulations.We first train a lipid-specific molecular language model, LiCore, to learn hidden representations of lipid molecules.We then explore the learned latent space to generate improved candidate formulations.This process is guided by a trained predictor, which evaluates delivery efficiency and provides directional signals.In reconstruction task, LiCore achieves near-perfect reconstruction performance output with a low invalid ratio on both the LNP-Virtual900k and LNP-Exp12k datasets.The predictor consistently improves ranking-oriented metrics across multiple cell lines, with our method outperforming the best baselines by an average of 4.1%, 10.8%, and 8.1% in Top-50, Top-10, and Top-5 identification accuracy, respectively.Guided by predictor, LiGen generates novel lipid candidates that achieve a 30.7% relative improvement over baseline methods in predicted delivery efficiency, with some candidates exceeding 50% improvement.
Ying Zhan 0001, Xiuqi Tang, Yan Zhang 0100, Xiao Tan 0005, Dian Shen, Beilun Wang
ACL (1)3
2026 Identification of Influential Node Group in Attributed Graph through Explaining Graph Neural Network
abstract
Identification of influential groups of nodes in attributed graphs has applications in a wide range of real-world problems, for instance, collecting important proceedings in citation networks, or identifying essential genes for diagnosing disease in Protein-Protein Interaction networks. Previous approaches for influence maximization manipulated on the graph structure, despite their proliferation, neglect the node attribute information containing additional knowledge. In this work, we introduce Global Graph UNderstanding (GGUN), a perturbation-based framework leveraging the explanatory power of Graph Neural Networks. It takes into account the entire graph structure and node attributes simultaneously and fuses knowledge through GNN layers. Following the perturbation-based explanation, GGUN fills the gap between Deep Neural Network gradient-based feature importance analysis and discrete structure in the graph, which is formulated as a combinatorial optimization problem. Moreover, GGUN obtains an efficient solution by relaxing the infeasible combinatorial optimization problem with performance guaranteed. Evaluations of synthetic and real-world datasets show that GGUN outperforms baselines on both quantitative metrics and human-intelligible analysis.
Xiao Tan 0005, Tongtong Su, Yan Zhang 0100, Binghui Xu, Dian Shen, Meng Wang 0009, Beilun Wang
WWW4
2025 Information-Agnostic Model Poisoning Attacks Against Byzantine-Robust Federated Learning
Yan Zhang 0100, Yueyao Chen, Xiao Tan 0005, Dian Shen, Meng Wang 0009, Beilun Wang
DASFAA (4)1
2025 NoTeNet: Normalized Mutual Information-Driven Tuning-free Dynamic Dependence Network Inference Method for Multimodal Data
abstract
Dynamic Dependence Network (DDN) inference is crucial for understanding evolving relationships in multimodal time series web data, with broad applications in fields like medical and financial network analysis. The inherent dynamic nature, temporal continuity, and heterogeneous data sources in multimodal time series data pose three fundamental challenges: computational efficiency, prediction stability and robustness, and modality quality disparity. Previous methods, generally lacking utilization of multiple modalities, either struggle with computational efficiency due to the time-intensive manual hyperparameter tuning, or compromise prediction stability and robustness by neglecting temporal coherence. To address these challenges, we propose a Normalized mutual information-driven Tuning-free Dynamic Dependence Network inference method for multimodal data, namely NoTeNet. NoTeNet provides a promising paradigm that can integrate two different data modalities to enhance prediction accuracy. It uses normalized mutual information transforms noisy auxiliary data into relationship matrices and employs a kernel function for smooth temporal estimation. Additionally, NoTeNet significantly reduces the need for manual hyperparameter adjustments, offering a tuning-free approach with theoretical guarantees. On various synthetic datasets and real-world data, NoTeNet demonstrates superior prediction accuracy and efficiency without the need for hyperparameter tuning, making it potential for a wide range of web data applications.
Xiao Tan 0005, Yangyang Shen, Yan Zhang 0100, Jingwen Shao, Dian Shen, Meng Wang 0009, Beilun Wang
WWW3
2023 A Robust Framework for Fixing The Vulnerability of Compressed Distributed Learning
abstract
Nowadays, as a prevailing paradigm for large-scale machine learning, distributed learning has been faced with two challenges, communication bottleneck and limited robustness. For the communication challenge, compression is widely-used as a solution. For the robustness challenge, some robust defence methods have been proposed. However, previous works that simultaneously consider these two challenges are limited. Through an experiment on the w8a dataset, we found that compressed distributed learning with rand-K is vulnerable to poisoning attacks. Therefore, in this paper, we propose a robust compressed distributed framework for distributed learning settings. Experimental evaluations on a9a and w8a datasets have shown the effectiveness of our proposed framework, which markedly decreases the average optimality gap from 1.47E −2 and 2.15E −2 to 3.98E −4 and 4.33E −4 respectively.
Yueyao Chen, Beilun Wang, Yan Zhang 0100, Jingyu Kuang
CSCWD3
2023 Take CARE: Improving Inherent Robustness of Spiking Neural Networks with Channel-wise Activation Recalibration Module
abstract
Spiking Neural Networks (SNNs) are considered the next generation of deep neural networks for their computation efficiency and biological plausibility. Still, SNN models can be fooled with adversarial perturbations and noises. There is an urgent need for building a robust SNN model that can be deployed in safety-critical domains. Recent works successfully proposed some defense methods inspired by those designed for traditional deep neural network models. However, these methods neglect the inherent robustness of SNN models, which has been proven by previous studies. In this paper, we dedicate ourselves to improving the inherent robustness of SNN without additional training. To do that, we unveil that the success of most attacks relies on obfuscating the model activation. Inspired by this phenomenon, we propose a spiking neural network framework Channel-wise Activation Recalibration (CARE) to improve SNN inherent robustness, which is named CARENet. By analyzing the model activation pattern, we prove that the CARE module has a strong capability of activation preservation. We evaluate our method on three benchmarks. Under diverse attacks, including hybrid attacks using multiple attacks, our method shows significant accuracy gains compared to baselines. Furthermore, our framework achieves competitive performance on natural benchmarks.
Yan Zhang 0100, Dian Shen, Meng Wang 0009, Beilun Wang
ICDM1
2021 Scalable Estimator for Multi-task Gaussian Graphical Models Based in an IoT Network
abstract
Recently, the Internet of Things (IoT) receives significant interest due to its rapid development. But IoT applications still face two challenges: heterogeneity and large scale of IoT data. Therefore, how to efficiently integrate and process these complicated data becomes an essential problem. In this article, we focus on the problem that analyzing variable dependencies of data collected from different edge devices in the IoT network. Because data from different devices are heterogeneous and the variable dependencies can be characterized into a graphical model, we can focus on the problem that jointly estimating multiple, high-dimensional, and sparse Gaussian Graphical Models for many related tasks (edge devices). This is an important goal in many fields. Many IoT networks have collected massive multi-task data and require the analysis of heterogeneous data in many scenarios. Past works on the joint estimation are non-distributed and involve computationally expensive and complex non-smooth optimizations. To address these problems, we propose a novel approach: Multi-FST. Multi-FST can be efficiently implemented on a cloud-server-based IoT network. The cloud server has a low computational load and IoT devices use asynchronous communication with the server, leading to efficiency. Multi-FST shows significant improvement, over baselines, when tested on various datasets.
Beilun Wang, Jiaqi Zhang 0005, Yan Zhang 0100, Meng Wang 0009, Sen Wang 0001
ACM Trans. Sens. Networks3