Yanfen Zheng

dblp:335/2734 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An end-to-end DNA storage coding method based on a low-complexity multiple biological constraints loss and RL-inspired differentiable solver
Yanfen Zheng, Xue Li 0019, Bin Wang 0005, Shihua Zhou, Ben Cao, Pan Zheng 0001
Expert Syst. Appl.2
2025 DNA Sequence Clustering in High Error Rates via Hash Sketches Fuzzy Clustering for Efficient Stored Data Reconstruction
Yanfen Zheng, Ben Cao, Zhenlu Liu, Bin Wang 0005, Shihua Zhou, Pan Zheng 0001
PAKDD (3)2
2025 Efficient DNA Fragment Assembly Based on Discrete Slime Mould Algorithm
Shuqing Si, Ben Cao, Yanfen Zheng
PAKDD (7)5
2025 DNA sequence design model for multi-scene fusion
Yanfen Zheng, Yaqing Hou, Qiang Zhang 0008, Xiaopeng Wei
Neural Comput. Appl.2
2025 Stable DNA Storage Encoding Scheme Based on Repeating Substring Tree
abstract
DNA storage is considered to be a promising storage media in the current era of data explosion. DNA encoding is the beginning of the DNA storage process and lays the foundation for subsequent processes. However, many encoding methods suffer from low encoding rate, do not satisfy important constraints, or have insufficient sequence stability. To address these issues and improved sequences stability, this paper proposes a novel approach called the Repeating Substring Tree Encoding (RSTE) method. The method begins by applying the Longest Substring Backtracking Method (LSBM) to identify the longest repeated substrings within the binary file. These substrings are then encoded into compact DNA motifs using Huffman encoding. In contrast to the ideal coding density of 2 bits per nucleotide (2 bit/nt) targeted by previous studies, RSTE enhances the encoding rate by 13% through efficient utilization of repeated substrings. Furthermore, the DNA sequences generated by the RSTE method successfully meet three biological constraints: run-length limitation, GC content balance and end constraints. The experimental results of minimum free energy and melting temperature indicate that the stability of the sequences encoded by RSTE is also greatly improved. A series of experiments showed that the sequences encoded by RSTE have a higher coding rate, satisfy constraints, and are more stable.
Jieqiong Wu, Penghao Wang 0001, Yanfen Zheng, Bin Wang 0005, Qiang Zhang 0008, Pan Zheng 0001
IEEE Trans. Comput. Biol. Bioinform.3
2024 MFN: Explainable DNA triple helixes Stabilized Design based on mCGR and flow network
abstract
DNA triple helix structure, as a highly specific gene targeting tool, enable gene regulation by precisely identifying and binding to target DNA sequences. However, the limits of design quality and efficiency affect their wide application in gene therapy. Therefore, in this paper, we propose an antiparallel DNA triple helixes design method-MFN based on matrix chaotic game representation (mCGR) and flow network. Leveraging the structural characteristics of DNA sequences, this method employs the mCGR algorithm to construct an initial matrix, generating a set of DNA sequences that conform to foundational constraints. Then, these sequences are mapped as stream network nodes to screen crosstalk structures by path search, and the whole process is observable and interpretable. Experimental results show that MFN significantly improves the design efficiency of triple helix structure and reduces crosstalk phenomenon. Wet experiments further verify the effectiveness of the method. In summary, MFN achieves an efficient and high-quality design of DNA triple helixes and provides a new idea for targeted gene therapy.
Xiaoru Wen, Yanfen Zheng, Ben Cao, Bin Wang 0005
BIBM4
2024 Application of Static Virus Spread Algorithm in Base-Balanced DNA Fragment Optimization
abstract
DNA has found applications in a diverse array of fields such as computing, medical diagnosis, and circuits. Designing DNA fragments that meet specific requirements is crucial for ensuring the smooth execution of tasks in these domains. However, the conventional approach of combining constraints with evolutionary algorithms often encounters challenges like orthogonality and thermodynamic instability. This poses risks to the control of reaction processes. To tackle these challenges, we have undertaken two tasks in this study: expanding constraint sets and innovating evolutionary algorithms. Firstly, we propose a base balance strategy aimed at enhancing the thermodynamic properties within DNA fragment groups by diversifying neighboring base combinations. The addition of this strategy achieves a minimum variance of 0.03 for the melting temperature while ensuring orthogonality. It represents a significant improvement compared to previous results and reduces the complexity of controlling reactions. Secondly, our static virus spread algorithm optimizes target DNA fragments by base mutations and virus amplification. Simultaneously,it demonstrates good performance across 23 benchmark functions, highlighting its optimization potential. This work is expected to further refine theoretical shortcomings and offer a convenient tool for DNA fragment optimization.
Yanfen Zheng, Xin Liu 0120, Bin Wang 0005, Qiang Zhang 0008
BIBM2
2023 High Net Information Density DNA Data Storage by the MOPE Encoding Algorithm
abstract
DNA has recently been recognized as an attractive storage medium due to its high reliability, capacity, and durability. However, encoding algorithms that simply map binary data to DNA sequences have the disadvantages of low net information density and high synthesis cost. Therefore, this paper proposes an efficient, feasible, and highly robust encoding algorithm called MOPE (Modified Barnacles Mating Optimizer and Payload Encoding). The Modified Barnacles Mating Optimizer (MBMO) algorithm is used to construct the non-payload coding set, and the Payload Encoding (PE) algorithm is used to encode the payload. The results show that the lower bound of the non-payload coding set constructed by the MBMO algorithm is 3%-18% higher than the optimal result of previous work, and theoretical analysis shows that the designed PE algorithm has a net information density of 1.90 bits/nt, which is close to the ideal information capacity of 2 bits per nucleotide. The proposed MOPE encoding algorithm with high net information density and satisfying constraints can not only effectively reduce the cost of DNA synthesis and sequencing but also reduce the occurrence of errors during DNA storage.
Yanfen Zheng, Ben Cao, Jieqiong Wu, Bin Wang 0005, Qiang Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 Design of Constraint Coding Sets for Archive DNA Storage
abstract
With the advent of the era of massive data, the increase of storage demand has far exceeded current storage capacity. DNA molecules provide a reliable solution for big data storage by virtue of their large capacity, high density, and long-term stability. To reduce errors in storing procedures, constructing a sufficient set of constraint encoding is critical for achieving DNA storage. A new version of the Marine Predator algorithm (called QRSS-MPA) is proposed in this paper to increase the lower bound of the coding set while satisfying the specific combination of constraints. In order to demonstrate the effectiveness of the improvement, the classical CEC-05 test function is used to test and compare the mean, variance, scalability, and significance. In terms of storage, the lower bound of construction is compared with previous works, and the result is found to be significantly improved. In order to prevent the emergence of a secondary structure that leads to sequencing failure, we give a more stringent lower bound for the constraint coding set, which is of great significance for reducing the error rate of DNA storage amidst its rapid development.
Qiang Yin 0006, Yanfen Zheng, Bin Wang 0005, Qiang Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.2