Yiran Cheng

dblp:285/1876 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Vercation: Precise Vulnerable Open-Source Software Version Identification Based on Static Analysis and LLM
abstract
Open-source software (OSS) has experienced a surge in popularity, attributed to its collaborative development model and cost-effective nature. However, the adoption of specific software versions in development projects may introduce security risks when these versions bring along vulnerabilities. Current methods of identifying vulnerable versions typically analyze and extract the code features involved in vulnerability patches using static analysis with pre-defined rules. They then use code clone detection to identify the vulnerable versions. These methods are hindered by imprecision due to (1) the exclusion of vulnerability-irrelevant code in the analysis and (2) the inadequacy of code clone detection. This paper presents VERCATION, an approach designed to identify vulnerable versions of OSS written in C/C++. VERCATION combines program slicing with a Large Language Model (LLM) to identify vulnerability-relevant code from vulnerability patches. It then backtracks historical commits to gather previous modifications of identified vulnerability-relevant code. We propose code clone detection based on expanded and normalized ASTs to compare the differences between pre-modification and post-modification code, thereby locating the vulnerability-introducing commit (vic) and enabling the identification of the vulnerable versions between the vulnerability-fixing commit and thevic. We curate a dataset linking 122 OSS vulnerabilities and 1,211 versions to evaluate VERCATION. On this dataset, our approach achieves an F1 score of 93.1%, outperforming current state-of-the-art methods. More importantly, VERCATION detected 202 incorrect vulnerable OSS versions in NVD reports.
Yiran Cheng, Ting Zhang 0011, Lwin Khin Shar, Shouguo Yang, Chaopeng Dong, David Lo 0001, Shichao Lv, Zhiqiang Shi, Limin Sun 0001
IEEE Trans. Software Eng.1
2025 RapidAVA: Efficient All-vs-All Overlap Detection in HiFi Reads
abstract
The High-Fidelity (HiFi) sequencing technology produces long, highly accurate reads, leading to higher-quality results in genome assembly. All-vs-all overlap detection is an important task for de novo assembly and other genomics applications. However, existing methods remain inefficient for HiFi reads, as they lack optimizations for long read length and high accuracy. To improve efficiency, we introduce RapidAVA, a new method for all-vs-all overlap detection in HiFi reads. First, we identify a set of representative subsequences, or minimizers, from all reads as indices for overlap detection; we keep the total number of minimizers small while ensuring their coverage of the reads. Second, we develop a fast seed-chaining algorithm to link the common minimizers (seeds) between a pair of reads. Last, to align a pair of reads to identify overlap, we adopt the wavefront alignment (WFA) algorithm for its lower complexity than other dynamic programming algorithms, and improve the algorithm with heuristics suitable for HiFi reads. Our experiments show that RapidAVA outperforms existing tools, achieving speedups of$4.7 \times$and$2.3 \times$over Minimap 2 and Hifiasm, respectively, with its peak memory consumption at 13 - 50 % of the other systems.
Yiran Cheng, Zhuoyang Chen, Qiong Luo 0001
BIBM1
2025 Generalized Rate Splitting for Enhanced Max-Min Fairness in Weak-User RSMA Systems
abstract
Rate splitting multiple access (RSMA) is a powerful multiple access technology that enables communication systems to achieve both reliable and fair data transmission by splitting and encoding user messages into common and private streams. This capability is particularly critical for space-air-ground-sea (SAGS) integrated networks, where heterogeneous nodes (e.g., satellites, UAVs, and underwater sensors) coexist with significant channel quality disparities. The common stream is formed by consolidating the diverse common messages that all users can decode, allowing the system to balance resource allocation and maintain reliable connectivity even for users with weaker channel conditions. Nevertheless, the performance of the common stream is often limited by the user with the weakest channel strength, a prevalent challenge in SAGS integrated networks with mixed near-far field communications and dynamic topology. To address this issue, this study proposes a generalized RSMA strategy to mitigate the rate limitation of common streams in RSMA systems, thereby enhancing system fairness and overall performance. The solution holds potential for crossdomain applications where strong and weak maritime/aerial users share spectrum resources. Furthermore, an algorithm is designed to optimize Max-min fairness (MMF) rate among all users. This is formulated as a non-convex optimization problem, which poses significant challenges for direct solution. To tackle this challenge, we design a low-complexity suboptimal iterative algorithm employing the successive convex approximation (SCA) method. Simulations demonstrate that the proposed generalized RSMA system outperforms traditional one-layer RSMA system and other existing counterparts, particularly in systems with weak users, by effectively enhancing the MMF rate and overall system fairness. This improvement suggests broader applicability for future integrated networks requiring unified management of heterogeneous links.
Junji Pan, Chang Liu 0008, Zheng Xue, Zhong Zheng 0001, Yiran Cheng, Muhammad Umar Farooq 0002, Guojun Han
VTC2025-Spring6
2025 UPicker: a semi-supervised particle picking transformer method for cryo-EM micrographs
abstract
Automatic single particle picking is a critical step in the data processing pipeline of cryo-electron microscopy structure reconstruction. In recent years, several deep learning-based algorithms have been developed, demonstrating their potential to solve this challenge. However, current methods highly depend on manually labeled training data, which is labor-intensive and prone to biases especially for high-noise and low-contrast micrographs, resulting in suboptimal precision and recall. To address these problems, we propose UPicker, a semi-supervised transformer-based particle-picking method with a two-stage training process: unsupervised pretraining and supervised fine-tuning. During the unsupervised pretraining, an Adaptive Laplacian of Gaussian region proposal generator is proposed to obtain pseudo-labels from unlabeled data for initial feature learning. For the supervised fine-tuning, UPicker only needs a small amount of labeled data to achieve high accuracy in particle picking. To further enhance model performance, UPicker employs a contrastive denoising training strategy to reduce redundant detections and accelerate convergence, along with a hybrid data augmentation strategy to deal with limited labeled data. Comprehensive experiments on both simulated and experimental datasets demonstrate that UPicker outperforms state-of-the-art particle-picking methods in terms of accuracy and robustness while requiring fewer labeled data than other transformer-based models. Furthermore, ablation studies demonstrate the effectiveness and necessity of each component of UPicker. The source code and data are available at https://github.com/JachyLikeCoding/UPicker.
Chi Zhang 0110, Yiran Cheng, Kaiwen Feng, Fa Zhang 0001, Renmin Han, Jieqing Feng
Briefings Bioinform.2
2024 RapidGKC: GPU-Accelerated K-Mer Counting
abstract
Many bioinformatics applications, e.g., genome assembly, genome profiling, and sequence alignment, break biological sequences into k-mers, or length-k substrings, for sub-sequent processing. In these applications, counting the number of occurrences of distinct k-mers is a common but expensive step due to the data and computation intensity. As such, prior work proposed to parallelize this task and utilize GPUs for further acceleration. However, these solutions under-utilize the GPU parallelism because the encoding format of intermediate data forces sequential decoding. To address this problem, we design a new encoding scheme for variable-length genomic data to support parallel encoding and decoding. Furthermore, we propose a novel rule to select common substrings among k-mers for partitioning, reducing the space cost as well as facilitating efficient parallel processing. Finally, we parallelize the entire workflow of partitioning and counting through pipelining, CPU-GPU co-processing, and work stealing. As a result, RapidGKC, our end-to-end GPU-accelerated k-mer counting system, outperforms state-of-the-art CPU-based and GPU-accelerated methods on real-world datasets.
Yiran Cheng, Xibo Sun, Qiong Luo 0001
ICDE1
2024 Serial Section Microscopy Image Inpainting Guided by Axial Optical Flow
abstract
Volume electron microscopy (vEM) is becoming a prominent technique in three-dimensional (3D) cellular visualization. vEM collects a series of two-dimensional (2D) images and reconstructs ultrastructures at the nanometer scale by rational axial interpolation between neighboring sections. However, section damage inevitably occurs in the sample preparation and imaging process, suffering from manual operational errors or occasional mechanical failures. The damaged regions present blurry and contaminated structure information, even local blank holes. Despite significant progress in single-image inpainting, it is still a great challenge to recover missing biological structures, that satisfy 3D structural continuity among sections. In this paper, we propose an optical flow-based serial section inpainting architecture to effectively combine the 3D structure information from neighboring sections and 2D image features from surrounding regions. We design a two-stage reference generation strategy to predict a rational and detailed intermediate state image from coarse to fine. Then, a GAN-based inpainting network is adopted to integrate all reference information and guide the restoration of missing structures, while ensuring consistent distribution of pixel values across the 2D image. Extensive experimental results well demonstrate the superiority of our method over existing inpainting tools. Our code is available at https://github.com/chengyr1999/FlowInpaint/.
Yiran Cheng, Bintao He, Fa Zhang 0001, Renmin Han
ACM Multimedia1
2024 Asteria-Pro: Enhancing Deep Learning-based Binary Code Similarity Detection by Incorporating Domain Knowledge
abstract
Widespread code reuse allows vulnerabilities to proliferate among a vast variety of firmware. There is an urgent need to detect these vulnerable codes effectively and efficiently. By measuring code similarities, AI-based binary code similarity detection is applied to detecting vulnerable code at scale. Existing studies have proposed various function features to capture the commonality for similarity detection. Nevertheless, the significant code syntactic variability induced by the diversity of IoT hardware architectures diminishes the accuracy of binary code similarity detection. In our earlier study and the tool Asteria , we adopted a Tree-LSTM network to summarize function semantics as function commonality, and the evaluation result indicates an advanced performance. However, it still has utility concerns due to excessive time costs and inadequate precision while searching for large-scale firmware bugs. To this end, we propose a novel deep learning-enhancement architecture by incorporating domain knowledge-based pre-filtration and re-ranking modules, and we develop a prototype named Asteria-Pro based on Asteria . The pre-filtration module eliminates dissimilar functions, thus reducing the subsequent deep learning-model calculations. The re-ranking module boosts the rankings of vulnerable functions among candidates generated by the deep learning model. Our evaluation indicates that the pre-filtration module cuts the calculation time by 96.9%, and the re-ranking module improves MRR and Recall by 23.71% and 36.4%, respectively. By incorporating these modules, Asteria-Pro outperforms existing state-of-the-art approaches in the bug search task by a significant margin. Furthermore, our evaluation shows that embedding baseline methods with pre-filtration and re-ranking modules significantly improves their precision. We conduct a large-scale real-world firmware bug search, and Asteria-Pro manages to detect 1,482 vulnerable functions with a high precision 91.65%.
Shouguo Yang, Chaopeng Dong, Yang Xiao 0011, Yiran Cheng, Zhiqiang Shi, Zhi Li 0018, Limin Sun 0001
ACM Trans. Softw. Eng. Methodol.4
2023 VERI: A Large-scale Open-Source Components Vulnerability Detection in IoT Firmware
Yiran Cheng, Shouguo Yang, Zhe Lang, Zhiqiang Shi, Limin Sun 0001
Comput. Secur.1
2023 Time-Capturing Dynamic Graph Embedding for Temporal Linkage Evolution
abstract
Dynamic graph embedding learns representation vectors for vertices and edges in a graph that evolves over time. We aim to capture and embed the evolution of vertices' temporal connectivity. Existing work studies the vertices' dynamic connection changes but neglects the time it takes for edges to evolve, failing to embed temporal linkage information into the evolution of the graph. To capture vertices' temporal linkage evolution, we model dynamic graphs as a sequence of snapshot graphs, appending the respective timespans of edges (ToE). We co-train a linear regressor to embed ToE while inferring a common latent space for all snapshot graphs by a matrix-factorization-based model to embed vertices' dynamic connection changes. Vertices' temporal linkage evolution is captured as their moving trajectories within the common latent representation space. Our embedding algorithm converges quickly with our proposed training methods, which is very time efficient and scalable. Extensive evaluations on several datasets show that our model can achieve significant performance improvements, i.e. 22.98% on average across all datasets, over the state-of-the-art baselines in the tasks of vertex classification, static and time-aware link prediction, and ToE prediction.
Yu Yang 0012, Jiannong Cao 0001, Milos Stojmenovic, Senzhang Wang, Yiran Cheng, Chun Lum, Zhetao Li
IEEE Trans. Knowl. Data Eng.5
2022 Effective Attribute Selection for Multi-dimensional Root Cause Analysis
abstract
Using large-scale multi-dimensional data for root cause analysis (MDRCA) is vitally important for online software services. It helps operators narrow down the scope of anomalies and failures quickly and localize the root cause to a finer granularity. However, most existing MDRCA algorithms can only solve low-dimensional problems. When dealing with high-dimensional data, the complexity of these algorithms would significantly increase, and even some algorithms would no longer work. Intuitively, passing only a subset of attributes rather than full attributes can improve the performance of these MDRCA algorithms. However, it is challenging due to data imbalance and novel root cause attributes. To better understand the problem of root-cause-oriented attribute selection (RCOAS), we conduct a preliminary study based on real-world data. We find that there exist several straightforward rules to filter out some attributes. In addition, we reveal that existing approaches do not fit the requirements of RCOAS. Motivated by the study, we propose an RCOAS approach, RC-LIR, to select a subset of attributes for downstream algorithms. RC-LIR first performs rule-based selection. Then it improves a feature selection algorithm by two strategies, i.e., scaling up imbalanced data and considering the redundant cost. Experiments on 1000 real-world fault cases demonstrate that RC-LIR can achieve an F1-score of 0.88, outper-forming the baseline approaches by at least 0.15. Furthermore, our experiments with four widely adopted MDRCA algorithms show that integrating RC-LIR can lead to more effective and efficient MDRCA.
Yiran Cheng, Pengxiang Jin, Yongqian Sun, Xiaohui Nie, Nengwen Zhao, Shenglin Zhang, Dan Pei
ISSRE1
2021 PMatch: Semantic-based Patch Detection for Binary Programs
abstract
Binary function matching has been proposed to detect the known vulnerabilities. However, the high similarity between the vulnerable and patched versions leads to a large of false positives. Patch detection is proposed to improve the accuracy of function matching by identifying the patched functions from matching results. However, the accuracy of existing methods decreases significantly due to the function changes introduced by high compiler optimization levels.In this paper, we propose PMatch, a method based on code semantic similarity to detect the patched binary functions. Firstly, PMatch extracts patch-affected code snippets from the patched binary function. Secondly, PMatch leverages a novel unsupervised sentence embedding technique in Natural Language Processing (NLP) to generate the semantic representations of binary code. Finally, PMatch matches the patch-affected code snippets with target blocks obtained by function diffing. To evaluate PMatch, we collect 101 CVEs and compile 304 binary programs with 4 different optimization levels. PMatch achieves an 86.43% average accuracy in detecting the patched functions, which outperforms the state-of-the-art work, and costs only 65.14ms per function. Besides, at the O3 high optimization level, PMatch achieves an accuracy improvement of over 20%.
Zhe Lang, Shouguo Yang, Yiran Cheng, Xiaoling Zhang 0009, Zhiqiang Shi, Limin Sun 0001
IPCCC3