Rui Mei

dblp:34/6478 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5Security and privacy · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
abstract
The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing automated jailbreak generation approaches e.g. AutoDAN, suffer from limited mutation diversity, shallow fitness evaluation, and fragile keyword-based detection. To address these limitations, we propose ForgeDAN, a novel evolutionary framework for generating semantically coherent and highly effective adversarial prompts against aligned LLMs. First, ForgeDAN introduces multi-strategy textual perturbations across character, word, and sentence-level operations to enhance attack diversity; then we employ interpretable semantic fitness evaluation based on a text similarity model to guide the evolutionary process toward semantically relevant and harmful outputs; finally, ForgeDAN integrates dual-dimensional jailbreak judgment, leveraging an LLM-based classifier to jointly assess model compliance and output harmfulness, thereby reducing false positives and improving detection effectiveness. Our evaluation demonstrates ForgeDAN achieves high jailbreaking success rates while maintaining naturalness and stealth, outperforming existing SOTA solutions.
Siyang Cheng, Gaotian Liu, Rui Mei, Kaishuo Wei, Yuqi Yu, Weiping Wen
TrustCom3
2022 APT Attribution for Malware Based on Time Series Shapelets
abstract
To discover and defend against APT attacks more efficiently, we need to conduct binary analysis and source tracing research on APT malicious codes. This paper attributes APT groups for malicious codes from the perspective of binary similarity. First, we innovatively select the local features of the binary functions for classification and apply time series mining techniques to the mining of sequences of basic blocks (called paths). The Shapelet model selects path shapelets, which are path fragments that can best represent paths and are used to distinguish paths. Path shapelets can provide path-level interpretability for classification. Second, we use API calls to filter functions and generate paths of interest to reduce resource consumption. To evaluate the proposed method, we collect APT malicious codes based on publicly available threat intelligence reports. Our method filters 92.82% of functions and generates an average of 1.37 paths per function. The classification effect has obvious advantages over other methods.
Qinqin Wang, Rui Mei, Zhihui Han
TrustCom4
2022 Measurement of Malware Family Classification on a Large-Scale Real-World Dataset
abstract
There are many review articles on malware analysis, which provide a comprehensive summary of the features, methods, and challenges of malware analysis. But these are limited to a theoretical overview. The purpose of this paper is to take malware family classification as an example, to restore and display the malware analysis in real scenarios.In this paper, the measurement of malware family classification is carried out on a large-scale dataset in the real world. We use the BODMAS dataset, which contains a total of 57,293 malware samples, with carefully curated family information (581 families). Referring to the common features (including static features and dynamic features) and machine learning methods mentioned in the review articles, we conduct feature extraction and classification experiments. Then, we summarize the classification results. Static features can efficiently classify a large number of samples, even with 10% packed samples. In real scenarios, the family distribution is extremely unbalanced, and the family classification results with a large number of samples are better. Different evaluation methods have different interpretations of the classification results. These measurement results provide a basis for further malware analysis in real applications.
Qinqin Wang, Rui Mei, Zhihui Han
TrustCom4
2022 Meltdown-type attacks are still feasible in the wall of kernel page-Table isolation
Yueqiang Cheng, Zhi Zhang 0001, Yansong Gao 0001, Zhaofeng Chen, Shengjian Guo, Qifei Zhang 0001, Rui Mei, Surya Nepal, Yang Xiang 0001
Comput. Secur.7
2021 CTSCOPY: Hunting Cyber Threats within Enterprise via Provenance Graph-based Analysis
abstract
In recent years, the security community has been working on detecting increasingly sophisticated cyber threats effectively and responding efficiently. A large body of approaches has been proposed and deployed for discovering malicious behaviors within enterprise IT environments. However, combating two main types of attacks, namely outsider Advanced Persistent Threats (APTs) and insider employee's malicious activities, is still a long-lasting confrontation. We propose a novel provenance graph-based approach for detecting these two major threats. First, we collect system event logs of endpoints in the enterprise IT environment and generate whole-system provenance graph and corresponding correlation graph. Then, we extract most uncommon or abnormal causality subgraphs for further graph embedding. Finally, we adopt an anomaly detection model to separate malicious parts from a mass of benign parts in the correlation graph for analysts' decision. We implement a prototype of CTSCOPY. Our evaluation demonstrates that it outperforms state-of-the-art approaches in various attack scenarios.
Rui Mei, Zhihui Han, Jian-Chun Jiang
QRS1
2006 CARAT: A novel method for allelic detection of DNA copy number changes using high density oligonucleotide arrays
abstract
BACKGROUND: DNA copy number alterations are one of the main characteristics of the cancer cell karyotype and can contribute to the complex phenotype of these cells. These alterations can lead to gains in cellular oncogenes as well as losses in tumor suppressor genes and can span small intervals as well as involve entire chromosomes. The ability to accurately detect these changes is central to understanding how they impact the biology of the cell. RESULTS: We describe a novel algorithm called CARAT (Copy Number Analysis with Regression And Tree) that uses probe intensity information to infer copy number in an allele-specific manner from high density DNA oligonuceotide arrays designed to genotype over 100,000 SNPs. Total and allele-specific copy number estimations using CARAT are independently evaluated for a subset of SNPs using quantitative PCR and allelic TaqMan reactions with several human breast cancer cell lines. The sensitivity and specificity of the algorithm are characterized using DNA samples containing differing numbers of X chromosomes as well as a test set of normal individuals. Results from the algorithm show a high degree of agreement with results from independent verification methods. CONCLUSION: Overall, CARAT automatically detects regions with copy number variations and assigns a significance score to each alteration as well as generating allele-specific output. When coupled with SNP genotype calls from the same array, CARAT provides additional detail into the structure of genome wide alterations that can contribute to allelic imbalance.
Jing Huang 0024, Joyce Chen, Jane Zhang, Xiaojun Di, Rui Mei, Shumpei Ishikawa, Hiroyuki Aburatani, Keith W. Jones, Michael H. Shapero
BMC Bioinform.7
2005 Dynamic model based algorithms for screening and genotyping over 100K SNPs on oligonucleotide microarrays
abstract
MOTIVATION: A high density of single nucleotide polymorphism (SNP) coverage on the genome is desirable and often an essential requirement for population genetics studies. Region-specific or chromosome-specific linkage studies also benefit from the availability of as many high quality SNPs as possible. The availability of millions of SNPs from both Perlegen and the public domain and the development of an efficient microarray-based assay for genotyping SNPs has brought up some interesting analytical challenges. Effective methods for the selection of optimal subsets of SNPs spanning the genome and methods for accurately calling genotypes from probe hybridization patterns have enabled the development of a new microarray-based system for robustly genotyping over 100,000 SNPs per sample. RESULTS: We introduce a new dynamic model-based algorithm (DM) for screening over 3 million SNPs and genotyping over 100,000 SNPs. The model is based on four possible underlying states: Null, A, AB and B for each probe quartet. We calculate a probe-level log likelihood for each model and then select between the four competing models with an SNP-level statistical aggregation across multiple probe quartets to provide a high-quality genotype call along with a quality measure of the call. We assess performance with HapMap reference genotypes, informative Mendelian inheritance relationship in families, and consistency between DM and another genotype classification method. At a call rate of 95.91% the concordance with reference genotypes from the HapMap Project is 99.81% based on over 1.5 million genotypes, the Mendelian error rate is 0.018% based on 10 trios, and the consistency between DM and MPAM is 99.90% at a comparable rate of 97.18%. We also develop methods for SNP selection and optimal probe selection. AVAILABILITY: The DM algorithm is available in Affymetrix's Genotyping Tools software package and in Affymetrix's GDAS software package. See http://www.affymetrix.com for further information. 10 K and 100 K mapping array data are available on the Affymetrix website.
Xiaojun Di, Hajime Matsuzaki, Teresa A. Webster, Earl Hubbell, Shoulian Dong, Dan Bartell, Jing Huang 0024, Richard Chiles, Geoffrey Yang, Mei-mei Shen, David Kulp, Giulia C. Kennedy, Rui Mei, Keith W. Jones, Simon Cawley
Bioinform.14
2003 Algorithms for large-scale genotyping microarrays
abstract
MOTIVATION: Analysis of many thousands of single nucleotide polymorphisms (SNPs) across whole genome is crucial to efficiently map disease genes and understanding susceptibility to diseases, drug efficacy and side effects for different populations and individuals. High density oligonucleotide microarrays provide the possibility for such analysis with reasonable cost. Such analysis requires accurate, reliable methods for feature extraction, classification, statistical modeling and filtering. RESULTS: We propose the modified partitioning around medoids as a classification method for relative allele signals. We use the average silhouette width, separation and other quantities as quality measures for genotyping classification. We form robust statistical models based on the classification results and use these models to make genotype calls and calculate quality measures of calls. We apply our algorithms to several different genotyping microarrays. We use reference types, informative Mendelian relationship in families, and leave-one-out cross validation to verify our results. The concordance rates with the single base extension reference types are 99.36% for the SNPs on autosomes and 99.64% for the SNPs on sex chromosomes. The concordance of the leave-one-out test is over 99.5% and is 99.9% higher for AA, AB and BB cells. We also provide a method to determine the gender of a sample based on the heterozygous call rate of SNPs on the X chromosome. See http://www.affymetrix.com for further information. The microarray data will also be available from the Affymetrix web site. AVAILABILITY: The algorithms will be available commercially in the Affymetrix software package.
Xiaojun Di, Geoffrey Yang, Hajime Matsuzaki, Jing Huang 0024, Rui Mei, Thomas B. Ryder, Teresa A. Webster, Shoulian Dong, Keith W. Jones, Giulia C. Kennedy, David Kulp
Bioinform.6
2002 Robust estimators for expression analysis
abstract
Abstract Motivation: We consider the problem of estimating values associated with gene expression from oligonucleotide arrays. Such estimates should linearly track concentration, yield non-negative results, have statistical guarantees of robustness against outliers, and allow estimates of significance and variance. Results: A hierarchy of simple models is used to design robust estimators meeting these goals for both stand alone and comparative experiments. This algorithm has been validated against an extensive panel of known spike experiments, and shows comparable performance to existing standards. Availability: Algorithms available commercially as part of the MAS 5.0 software package. Data sets available from the Affymetrix website. Contact: [email protected] Supplemental Information: http://www.affymetrix.com/community/publications/affymetrix/index.affx * To whom correspondence should be addressed.
Earl Hubbell, Rui Mei
Bioinform.3
2002 Analysis of high density expression microarrays with signed-rank call algorithms
abstract
MOTIVATION: We consider the detection of expressed genes and the comparison of them in different experiments with the high-density oligonucleotide microarrays. The results are summarized as the detection calls and comparison calls, and they should be robust against data outliers over a wide target concentration range. It is also helpful to provide parameters that can be adjusted by the user to balance specificity and sensitivity under various experimental conditions. RESULTS: We present rank-based algorithms for making detection and comparison calls on expression microarrays. The detection call algorithm utilizes the discrimination scores. The comparison call algorithm utilizes intensity differences. Both algorithms are based on Wilcoxon's signed-rank test. Several parameters in the algorithms can be adjusted by the user to alter levels of specificity and sensitivity. The algorithms were developed and analyzed using spiked-in genes arrayed in a Latin square format. In the call process, p-values are calculated to give a confidence level for the pertinent hypotheses. For comparison calls made between two arrays, two primary normalization factors are defined. To overcome the difficulty that constant normalization factors do not fit all probe sets, we perturb these primary normalization factors and make increasing or decreasing calls only if all resulting p-values fall within a defined critical region. Our algorithms also automatically handle scanner saturation.
Rui Mei, Xiaojun Di, Thomas B. Ryder, Earl Hubbell, S. Dee, Teresa A. Webster, C. A. Harrington, Ming-Hsiu Ho, J. Baid, S. P. Smeekens
Bioinform.2