Yanlin Zhang

dblp:121/6745 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
abstract
Weicai Long, Yusen Hou, Junning Feng, Houcheng su, Shuo Yang, Donglin Xie, Yanlin Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Weicai Long, Yusen Hou, Junning Feng 0001, Houcheng Su, Donglin Xie, Yanlin Zhang
ACL (1)7
2026 Porcine MutBERT: a family of lightweight genomic foundation models for functional element prediction in pigs
abstract
The pig (Sus scrofa) is both an economically important livestock species and a valuable biomedical model . Its genome bears regulatory features shaped by domestication and selection that are often poorly captured by genomic language models (gLMs) trained on human or model organism data. To address these challenges, we developed Porcine MutBERT, a suite of lightweight gLMs with 86 million parameters that employs a probabilistic masking strategy targeting evolutionarily informative single-nucleotide polymorphisms. This design captures population-specific variation while reducing computational cost. We further propose PorcineBench, a benchmark that evaluates gLM performance across porcine functional genomics tasks, including chromatin accessibility (ATAC-seq), CTCF binding, and histone modifications (H3K27ac, H3K4me1, and H3K27me3). Results show that Porcine MutBERT family achieves highly competitive performance on PorcineBench relative to substantially larger models, while providing an explicitly porcine-adapted alternative for downstream functional genomics in pigs. These findings underscore the advantages of species-adapted, efficient architectures in agricultural genomics and demonstrate that compact gLMs can expand accessibility and impact in resource-constrained settings. The code and data are available at https://github.com/ai4nucleome/pigmutbert.
Weicai Long, Wenkang Wei, Xiaoai Zhang, Yanlin Zhang, Zishuai Wang
Briefings Bioinform.7
2025 MutBERT: probabilistic genome representation improves genomics foundation models
abstract
MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.
Weicai Long, Houcheng Su, Jiaqi Xiong, Yanlin Zhang
Bioinform.4
2025 Fast fixed-/preassigned-time synchronization of Clifford-valued neural networks for medical image encryption
Yanlin Zhang, Kit Ian Kou
Neurocomputing1
2025 Exploiting explicit item-item correlations from knowledge graphs for enhanced sequential recommendation
abstract
In recent years, the research of employing knowledge graphs (KGs) in sequential recommendation (SR) has received a lot of attention, since the side information extracted from KGs, especially the information of the correlations between items, indeed helps the SR models achieve better performance. However, many previous KG-based SR models tend to introduce some noise information when learning item embeddings, or insufficiently fuse item–item correlations into their sequential modeling, thus limiting their performance improvements . In this paper, we propose a D istance- A ware K nowledge-based S equential R ecommendation model ( DAKSR ), which exploits the explicit item–item correlations from KGs to achieve enhanced SR. Specifically, as one critical component in our DAKSR, the distance score matrix (DSM) is first obtained to indicate the correlations between items, and then leveraged in the following three major modules of DAKSR. First, in the Item-Set Embedding layer (ISE) all item embeddings are learned based on DSM, in which the noise information is eliminated effectively. Meanwhile, the Knowledge-Infused Transformer (KIT) incorporates DSM into its attention mechanism to improve the feature extraction. Furthermore, the Knowledge Contrastive Learning module (KCL) also leverages the item–item correlations presented in DSM to generate two credible sequence views, which are used to refine sample representations through a contrastive learning strategy, and thus improve the model’s robustness. Our extensive experiments on three SR benchmarks obviously demonstrate our DAKSR’s superior performance over the state-of-the-art (SOTA) KG-based recommendation models. The implementation of our DAKSR is available at https://github.com/Easonsi/DAKSR for reproducing our experiment results conveniently.
Yanlin Zhang, Deqing Yang, Xiaodong Gu 0001
Inf. Syst.1
2025 Advancements in exponential synchronization and encryption techniques: Quaternion-Valued Artificial Neural Networks with two-sided coefficients
Kit Ian Kou, Yanlin Zhang, Yang Liu 0040
Neural Networks3
2025 Randomized quaternion tensor UTV decompositions for color image and color video processing
Liqiao Yang, Jifei Miao, Tai-Xiang Jiang, Yanlin Zhang, Kit Ian Kou
Pattern Recognit.4
2024 ELF-Gym: Evaluating Large Language Models Generated Features for Tabular Prediction
abstract
Crafting effective features is a crucial yet labor-intensive and domain-specific task within machine learning pipelines. Fortunately, recent advancements in Large Language Models (LLMs) have shown promise in automating various data science tasks, including feature engineering. But despite this potential, evaluations thus far are primarily based on the end performance of a complete ML pipeline, providing limited insight into precisely how LLMs behave relative to human experts in feature engineering. To address this gap, we propose ELF-Gym, a framework for Evaluating LLM-generated Features. We curated a new dataset from historical Kaggle competitions, including 251 golden features used by top-performing teams. ELF-Gym then quantitatively evaluates LLM-generated features by measuring their impact on downstream model performance as well as their alignment with expert-crafted features through semantic and functional similarity assessments. This approach provides a more comprehensive evaluation of disparities between LLMs and human experts, while offering valuable insights into specific areas where LLMs may have room for improvement. For example, using ELF-Gym we empirically demonstrate that, in the best-case scenario, LLMs can semantically capture approximately 56% of the golden features, but at the more demanding implementation level this overlap drops to 13%. Moreover, in other cases LLMs may fail completely, particularly on datasets that require complex features, indicating broad potential pathways for improvement.
Yanlin Zhang, Ning Li 0029, Weinan Zhang 0001, David P. Wipf
CIKM1
2024 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs
abstract
Given a relational database (RDB), how can we predict missing column values in some target table of interest? Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably falls well behind advances in other domains such as computer vision or natural language processing. This deficit stems, at least in part, from the lack of established/public RDB benchmarks as needed for training and evaluation purposes. As a result, related model development thus far often defaults to tabular approaches trained on ubiquitous single-table benchmarks, or on the relational side, graph-based alternatives such as GNNs applied to a completely different set of graph datasets devoid of tabular characteristics. To more precisely target RDBs lying at the nexus of these two complementary regimes, we explore a broad class of baseline models predicated on: (i) converting multi-table datasets into graphs using various strategies equipped with efficient subsampling, while preserving tabular characteristics; and (ii) trainable models with well-matched inductive biases that output predictions based on these input subgraphs. Then, to address the dearth of suitable public benchmarks and reduce siloed comparisons, we assemble a diverse collection of (i) large-scale RDB datasets and (ii) coincident predictive tasks. From a delivery standpoint, we operationalize the above four dimensions (4D) of exploration within a unified, scalable open-source toolbox called 4DBInfer; please see https://github.com/awslabs/multi-table-benchmark .
David P. Wipf, Zheng Zhang 0001, Christos Faloutsos, Weinan Zhang 0001, Muhan Zhang, Zhenkun Cai, Jiahang Li 0002, Zunyao Mao, Yakun Song, Yanlin Zhang, Chuan Lei, Xiao Qin 0003, Ning Li 0029, Han Zhang 0057
NeurIPS13
2024 ARGV: 3D genome structure exploration using augmented reality
abstract
Over the past two decades, scientists have increasingly realized the importance of the three-dimensional (3D) genome organization in regulating cellular activity. Hi-C and related experiments yield 2D contact matrices that can be used to infer 3D models of chromosome structure. Visualizing and analyzing genomes in 3D space remains challenging. Here, we present ARGV, an augmented reality 3D Genome Viewer. ARGV contains more than 350 pre-computed and annotated genome structures inferred from Hi-C and imaging data. It offers interactive and collaborative visualization of genomes in 3D space, using standard mobile phones or tablets. A user study comparing ARGV to existing tools demonstrates its benefits.
Chrisostomos Drogaris, Yanlin Zhang, Elena Nazarova, Roman Sarrazin-Gendron, Sélik Wilhelm-Landry, Yan Cyr, Jacek Majewski, Mathieu Blanchette, Jérôme Waldispühl
BMC Bioinform.2
2024 Synchronization of fractional-order quaternion-valued neural networks with image encryption via event-triggered impulsive control
Yanlin Zhang, Liqiao Yang, Kit Ian Kou, Yang Liu 0040
Knowl. Based Syst.1
2024 Enhancing user and item representation with collaborative signals for KG-based recommendation
Yanlin Zhang, Xiaodong Gu 0001
Neural Comput. Appl.1
2024 Identification and Evolutionary Analysis of User Collusion Behavior in Blockchain Online Social Media
abstract
Blockchain technology has given rise to a series of new blockchain online social media (BOSMs), of which Steemit is representative. Such communities are based on a token reward system and attempt to engross users in the knowledge activities of the community through knowledge payment. Studies have found that the reward system of such communities has been abused (e.g., collusion for profit), but few studies have performed an in-depth analysis for this phenomenon. Consequently, real data for Steemit are used as a case study herein to examine the collusion of users in BOSMs. Two user collusion behaviors (group-voting and vote-buying) are defined and measured. On this basis, an identification and evolutionary survival analysis of the two collusion behaviors are conducted for colluding users and colluding groups, and the behavior patterns of user collusion under the token system are deconstructed. The results of this study improve stakeholders’ understanding of user participation behavior in new online communities, and serve as a reference for decision-making in community governance and token design.
Hongting Tang, Jian Ni, Yanlin Zhang
IEEE Trans. Comput. Soc. Syst.3
2023 Reference panel-guided super-resolution inference of Hi-C data
abstract
MOTIVATION: Accurately assessing contacts between DNA fragments inside the nucleus with Hi-C experiment is crucial for understanding the role of 3D genome organization in gene regulation. This challenging task is due in part to the high sequencing depth of Hi-C libraries required to support high-resolution analyses. Most existing Hi-C data are collected with limited sequencing coverage, leading to poor chromatin interaction frequency estimation. Current computational approaches to enhance Hi-C signals focus on the analysis of individual Hi-C datasets of interest, without taking advantage of the facts that (i) several hundred Hi-C contact maps are publicly available and (ii) the vast majority of local spatial organizations are conserved across multiple cell types. RESULTS: Here, we present RefHiC-SR, an attention-based deep learning framework that uses a reference panel of Hi-C datasets to facilitate the enhancement of Hi-C data resolution of a given study sample. We compare RefHiC-SR against tools that do not use reference samples and find that RefHiC-SR outperforms other programs across different cell types, and sequencing depths. It also enables high-accuracy mapping of structures such as loops and topologically associating domains. AVAILABILITY AND IMPLEMENTATION: https://github.com/BlanchetteLab/RefHiC.
Yanlin Zhang, Mathieu Blanchette
Bioinform.1
2023 Fixed-time synchronization for quaternion-valued memristor-based neural networks with mixed delays
abstract
In this paper, the fixed-time synchronization (FXTSYN) of unilateral coefficients quaternion-valued memristor-based neural networks (UCQVMNNs) with mixed delays is investigated. A direct analytical approach is suggested to obtain FXTSYN of UCQVMNNs utilizing one-norm smoothness in place of decomposition. When dealing with drive-response system discontinuity issues, use the set-valued map and the differential inclusion theorem. To accomplish the control objective, innovative nonlinear controllers and the Lyapunov functions are designed. Furthermore, some criteria of FXTSYN for UCQVMNNs are given using inequality techniques and the novel FXTSYN theory. And the accurate settling time is obtained explicitly. Finally, in order to show that the obtained theoretical results are accurate, useful, and applicable, numerical simulations are presented at the conclusion.
Yanlin Zhang, Liqiao Yang, Kit Ian Kou, Yang Liu 0040
Neural Networks1
2021 Subspace Constraint for Single Image Super-Resolution
Yanlin Zhang, Ding Qin, Xiaodong Gu 0001
ICANN (3)1
2020 User Recruitment with Budget Redistribution in Edge-Aided Mobile Crowdsensing
Yanlin Zhang, Peng Li 0046, Tao Zhang 0043
ICA3PP (2)1
2020 Novel design of Hardware Trojan: A generic approach for defeating testability based detection
abstract
Hardware design, especially the very large scale integration(VLSI) and systems on chip design(SOC), utilizes many codes from third-party intellectual property (IP) providers and former designers. Hardware Trojans (HTs) are easily inserted in this process. Recently researchers have proposed many HTs detection techniques targeting the design codes. State-of-art detections are based on the testability including Controllability and Observability, which are effective to all HTs from TrustHub, and advanced HTs like DeTrust. Meanwhile, testability based detections have advantages in the timing complexity and can be easily integrated into recently industrial verification. Undoubtedly, the adversaries will upgrade their designs accordingly to evade these detection techniques. Designing a variety of complex trojans is a significant way to perfect the existing detection, therefore, we present a novel design of HTs to defeat the testability based detection methods, namely DeTest. Our approach is simple and straight forward, yet it proves to be effective at adding some logic. Without changing HTs malicious function, DeTest decreases controllability and observability values to about 10% of the original, which invalidates distinguishers like clustering and support vector machines (SVM). As shown in our practical attack results, adversaries can easily use DeTest to upgrade their HTs to evade testability based detections. Combined with advanced HTs design techniques like DeTrust, DeTest can evade previous detecions, like UCI, VeriTrust and FANCI. We further discuss how to extend existing solutions to reduce the threat posed by DeTest.
Zhiqiang Lv, Yanlin Zhang, Weiqing Huang
TrustCom3
2020 Fixed-Time Synchronization of Complex-Valued Memristor-Based Neural Networks with Impulsive Effects
Yanlin Zhang, Shengfu Deng
Neural Process. Lett.1
2012 A multi-sources data assimilation system for catchment scale research
abstract
A multi-sources data assimilation system has been developed for the catchment scale land water and energy cycle researches. The land surface model, microwave radiative transfer model and ensemble Kalman filter have been coupled in this system, the high performance computing was also considered. This system is being used in the data assimilation of soil moisture, soil temperature, microwave brightness temperature and snow water equivalent at catchment scale.
Xujun Han, Xin Li 0029, Yanlin Zhang, Jian Kang 0004
IGARSS3