Runqing Yang

dblp:03/7065 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 3 since 2021Security and privacy · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Mixture-of-Experts Distillation for Cross-Satellite Generalizable Incremental Remote Sensing Scene Classification
abstract
Incremental learning aims to continuously acquire new knowledge from data streams while maintaining previously learned knowledge. Existing incremental learning methods typically assume that the training (source domain) and testing (target domain) data are identically distributed. However, differences in sensor parameters and imaging conditions inevitably lead to distribution gaps between data collected from different satellites (domains). The ensuing domain shift problem substantially impairs the generalization of continuously learned knowledge from source domains to unseen ones. To tackle this problem, we propose adaptive mixture-of-experts distillation (AMoED) for cross-satellite generalizable incremental remote sensing scene classification (CSGIRSSC). Specifically, AMoED adopts a high-level semantic learning pipeline, in which new knowledge is acquired through the coordinated guidance of multiple domain-specific experts, rather than directly from raw data. This pipeline prevents the model from being exposed to large volumes of newly emerging data, thereby alleviating the erasure of previous knowledge when adapting to new data distributions. Besides, the adaptive mixture of domain-specific experts facilitates the formation of universal class concepts, which exhibit strong generalizability across different domains. During the learning process, an equi-partite subset is constructed for knowledge acquisition and consolidation, accompanied by a shallow style-mixing operation to mitigate the interference of domain discrepancies. Extensive experiments are conducted on four remote sensing scene classification datasets, and the proposed method consistently achieves state-of-the-art performance across various scenarios and settings. The code is released at https://github.com/fuyimin96/AMoED.
Yimin Fu, Runqing Yang, Zhunga Liu, Michael Kwok-Po Ng
IEEE Trans. Circuits Syst. Video Technol.2
2022 Optimizing genomic control in mixed model associations with binary diseases
abstract
Complex computation and approximate solution hinder the application of generalized linear mixed models (GLMM) into genome-wide association studies. We extended GRAMMAR to handle binary diseases by considering genomic breeding values (GBVs) estimated in advance as a known predictor in genomic logit regression, and then reduced polygenic effects by regulating downward genomic heritability to control false negative errors produced in the association tests. Using simulations and case analyses, we showed in optimizing GRAMMAR, polygenic effects and genomic controls could be evaluated using the fewer sampling markers, which extremely simplified GLMM-based association analysis in large-scale data. Further, joint association analysis for quantitative trait nucleotide (QTN) candidates chosen by multiple testing offered significant improved statistical power to detect QTNs over existing methods.
Zhiyu Hao, Runqing Yang, Pao Xu
Briefings Bioinform.5
2022 Conan: A Practical Real-Time APT Detection System With High Accuracy and Efficiency
abstract
Advanced Persistent Threat (APT) attacks have caused serious security threats and financial losses worldwide. Various real-time detection mechanisms that combine context information and provenance graphs have been proposed to defend against APT attacks. However, existing real-time APT detection mechanisms suffer from accuracy and efficiency issues due to inaccurate detection models and the growing size of provenance graphs. To address the accuracy issue, we propose a novel and accurate APT detection model that removes unnecessary phases and focuses on the remaining ones with improved definitions. To address the efficiency issue, we propose a state-based framework in which events are consumed as streams and each entity is represented in an FSA-like structure without storing historic data. Additionally, we reconstruct attack scenarios by storing just one in a thousand events in a database. Finally, we implement our design, calledConan, on Windows and conduct comprehensive experiments under real-world scenarios to show thatConancan accurately and efficiently detect all attacks within our evaluation. The memory usage and CPU efficiency ofConanremain constant over time (1-10 MB of memory and hundreds of times faster than data generation), makingConana practical design for detecting both known and unknown APT attacks in real-world scenarios.
Chun-lin Xiong, Tiantian Zhu 0001, Weihao Dong, Linqi Ruan, Runqing Yang, Yueqiang Cheng, Yan Chen 0004, Xutong Chen
IEEE Trans. Dependable Secur. Comput.5
2022 RATScope: Recording and Reconstructing Missing RAT Semantic Behaviors for Forensic Analysis on Windows
abstract
Remote Access Trojan (RAT) attacks have become an extensively prevailing and serious threat to enterprise security. A forensic system targeting RAT attacks is needed to record and reconstruct fine-grained semantic behaviors of RATs. However, existing forensic systems suffer from various issues such as intrusive instrumentation, nontrivial recording overhead, and RAT behavior blindness. In this article, we first conduct a large-scale study of a representative set of real-world RAT families active from 1999 to 2016. This is the first study to understand the landscape of RATs in the literature. Based on the study, we then proposeRATScope, an instrumentation-free RAT forensic system targeting Windows platform. Specifically,RATScopeoffers an audit logging module to efficiently record system logs by leveraging Event Tracing for Windows (ETW), and provides a novel program behavior modeling technique to reconstruct semantic behaviors of RATs accurately. We implement a prototype ofRATScopeand evaluate the recording overhead and the behavior identification accuracy. The results show that the audit logging module only incurs 3.7 percent runtime overhead on average. Our system can achieve around 90 percent true positive rate in the cross-family experiment, around 80 percent true positive rate in the two-year spanning temporal experiment, and nearzerofalse positive rate.
Runqing Yang, Xutong Chen, Haitao Xu 0002, Yueqiang Cheng, Chun-lin Xiong, Linqi Ruan, Mohammad Kavousi, Zhenyuan Li, Liheng Xu, Yan Chen 0004
IEEE Trans. Dependable Secur. Comput.1
2021 SemFlow: Accurate Semantic Identification from Low-Level System Data
Mohammad Kavousi, Runqing Yang, Shiqing Ma, Yan Chen 0004
SecureComm (1)2
2021 Genome-wide hierarchical mixed model association analysis
abstract
In genome-wide mixed model association analysis, we stratified the genomic mixed model into two hierarchies to estimate genomic breeding values (GBVs) using the genomic best linear unbiased prediction and statistically infer the association of GBVs with each SNP using the generalized least square. The hierarchical mixed model (Hi-LMM) can correct confounders effectively with polygenic effects as residuals for association tests, preventing potential false-negative errors produced with genome-wide rapid association using mixed model and regression or an efficient mixed-model association expedited (EMMAX). Meanwhile, the Hi-LMM performs the same statistical power as the exact mixed model association and the same computing efficiency as EMMAX. When the GBVs have been estimated precisely, the Hi-LMM can detect more quantitative trait nucleotides (QTNs) than existing methods. Especially under the Hi-LMM framework, joint association analysis can be made straightforward to improve the statistical power of detecting QTNs.
Zhiyu Hao, Runqing Yang
Briefings Bioinform.4
2021 Hierarchical mixed-model expedites genome-wide longitudinal association analysis
abstract
A hierarchical random regression model (Hi-RRM) was extended into a genome-wide association analysis for longitudinal data, which significantly reduced the dimensionality of repeated measurements. The Hi-RRM first modeled the phenotypic trajectory of each individual using a RRM and then associated phenotypic regressions with genetic markers using a multivariate mixed model (mvLMM). By spectral decomposition of genomic relationship and regression covariance matrices, the mvLMM was transformed into a multiple linear regression, which improved computing efficiency while implementing mvLMM associations in efficient mixed-model association expedited (EMMAX). Compared with the existing RRM-based association analyses, the statistical utility of Hi-RRM was demonstrated by simulation experiments. The method proposed here was also applied to find the quantitative trait nucleotides controlling the growth pattern of egg weights in poultry data.
Runqing Yang
Briefings Bioinform.6
2021 Threat detection and investigation with system-level provenance graphs: A survey
Zhenyuan Li, Qi Alfred Chen, Runqing Yang, Yan Chen 0004
Comput. Secur.3
2020 UIScope: Accurate, Instrumentation-free, and Visible Attack Investigation for GUI Applications
Runqing Yang, Shiqing Ma, Haitao Xu 0002, Xiangyu Zhang 0001, Yan Chen 0004
NDSS1
2019 Quick approximation of threshold values for genome-wide association studies
abstract
Standard normal statistics, chi-squared statistics, Student's t statistics and F statistics are used to map quantitative trait nucleotides for both small and large sample sizes. In genome-wide association studies (GWASs) of single-nucleotide polymorphisms (SNPs), the statistical distributions depend on both genetic effects and SNPs but are independent of SNPs under the null hypothesis of no genetic effects. Therefore, hypothesis testing when a nuisance parameter is present only under the alternative was introduced to quickly approximate the critical thresholds of these test statistics for GWASs. When only the statistical probabilities are available for high-throughput SNPs, the approximate critical thresholds can be estimated with chi-squared statistics, formulated by statistical probabilities with a degree of freedom of two. High similarities in the critical thresholds between the accurate and approximate estimations were demonstrated by extensive simulations and real data analysis.
Zhiyu Hao, Jinhua Ye, Jingli Zhao, Shuling Li, Runqing Yang
Briefings Bioinform.7
2018 Automatic Benchmark Generation Framework for Malware Detection
abstract
To address emerging security threats, various malware detection methods have been proposed every year. Therefore, a small but representative set of malware samples are usually needed for detection model, especially for machine-learning-based malware detection models. However, current manual selection of representative samples from large unknown file collection is labor intensive and not scalable. In this paper, we firstly propose a framework that can automatically generate a small data set for malware detection. With this framework, we extract behavior features from a large initial data set and then use a hierarchical clustering technique to identify different types of malware. An improved genetic algorithm based on roulette wheel sampling is implemented to generate final test data set. The final data set is only one-eighteenth the volume of the initial data set, and evaluations show that the data set selected by the proposed framework is much smaller than the original one but does not lose nearly any semantics.
Guanghui Liang, Jianmin Pang, Zheng Shan, Runqing Yang
Secur. Commun. Networks4
2015 Vetting SSL Usage in Applications with SSLINT
abstract
Secure Sockets Layer (SSL) and Transport Layer Security (TLS) protocols have become the security backbone of the Web and Internet today. Many systems including mobile and desktop applications are protected by SSL/TLS protocols against network attacks. However, many vulnerabilities caused by incorrect use of SSL/TLS APIs have been uncovered in recent years. Such vulnerabilities, many of which are caused due to poor API design and inexperience of application developers, often lead to confidential data leakage or man-in-the-middle attacks. In this paper, to guarantee code quality and logic correctness of SSL/TLS applications, we design and implement SSLINT, a scalable, automated, static analysis system for detecting incorrect use of SSL/TLS APIs. SSLINT is capable of performing automatic logic verification with high efficiency and good accuracy. To demonstrate it, we apply SSLINT to one of the most popular Linux distributions -- Ubuntu. We find 27 previously unknown SSL/TLS vulnerabilities in Ubuntu applications, most of which are also distributed with other Linux distributions.
Boyuan He, Vaibhav Rastogi, Yinzhi Cao, Yan Chen 0004, V. N. Venkatakrishnan, Runqing Yang, Zhenrui Zhang
IEEE Symposium on Security and Privacy6
2014 Forward LASSO analysis for high-order interactions in genome-wide association study
abstract
Previous genome-wide association study (GWAS) focused on low-order interactions between pairwise single-nucleotide polymorphisms (SNPs) with significant main effects. Little is known how high-order interactions effect, especially one among the SNPs without main effects regulates quantitative traits. Within the frameworks of linear model and generalized linear model, the LASSO with coordinate descent step can be used to simultaneously analyze thousands and thousands of SNPs for normal and discrete traits. With consideration of high-order interactions among SNPs, a huge number of genetic effects make the LASSO failing to work under the presented condition of computation. Forward LASSO analysis is, therefore, proposed to shrink most of genetic effects to be zeros stage by stage. Simulation demonstrates that our proposed method could be used instead of the LASSO method for full model in mapping high-order interactions. Application of forward LASSO method is provided to GWAS for carcass traits and meat quality traits in beef cattle.
Huijiang Gao, Yang Wu 0004, Jiahan Li, Hongwang Li, Junya Li, Runqing Yang
Briefings Bioinform.6
2014 Iteratively reweighted LASSO for mapping multiple quantitative trait loci
abstract
The iteratively reweighted least square (IRLS) method is mostly identical to maximum likelihood (ML) method in terms of parameter estimation and power of quantitative trait locus (QTL) detection. But the IRLS is greatly superior to ML in terms of computing speed and the robustness of parameter estimation. In conjunction with the priors of parameters, ML can analyze multiple QTL model based on Bayesian theory, whereas under a single QTL model, IRLS has very limited statistical power to detect multiple QTLs. In this study, we proposed the iteratively reweighted least absolute shrinkage and selection operator (IRLASSO) for extending IRLS to simultaneously map multiple QTLs. The LASSO with coordinate descent step is employed to efficiently estimate non-zero genetic effect of each locus scanned over entire genome. Simulations demonstrate that IRLASSO has a higher precision of parameter estimation and power to detect QTL than IRLS, and is able to estimate residual variance more accurately than the unweighted LASSO based on LS. Especially, IRLASSO is very fast, usually taking less than five iterations to converge. The barley dataset from the North American Barley Genome Mapping Project is reanalyzed by our proposed method.
Tianfu Yang, Hongwang Li, Runqing Yang
Briefings Bioinform.4
2014 An efficient approach to large-scale genotype-phenotype association analyses
abstract
Modern molecular biotechnology generates a great deal of intermediate information, such as transcriptional and metabolic products in bridging DNA and complex traits. In genome-wide linkage analysis and genome-wide association study, regression analysis for large-scale correlated phenotypes is applied to map genes for those by-products that are regarded as quantitative traits. For a single trait, least absolute shrinkage and selection operator with coordinate descent step can be employed to efficiently shrink sparse non-zero genetic effects of quantitative trait loci (QTLs). However, regression analyses in a trait-by-trait basis do not take account of the correlations among the analyzed traits. In this study, conditional phenotype of each trait is defined, given other traits. Large-scale genotype-phenotype association analyses are therefore transformed to separate genotype-conditional phenotype ones. Meanwhile, the correlation architecture between each trait and other traits can also be provided by shrinkage estimation for each conditional phenotype. Simulation demonstrates that the proposed conditional mapping method is generally identical to joint mapping method based on multivariate analysis in terms of statistical detection power and parameter estimation. Application of the method is provided to locate eQTL in yeast.
Runqing Yang, Hongwang Li, Lina Fu
Briefings Bioinform.1
2012 Bayesian inference for genomic imprinting underlying developmental characteristics
abstract
The identification of imprinted genes is becoming a standard procedure in searching for quantitative trait loci (QTL) underlying complex traits. When a developmental characteristic such as growth or drug response is observed at multiple time points, understanding the dynamics of gene function governing the underlying feature should provide more biological information regarding the genetic control of an organism. Recognizing that differential imprinting can be development-specific, mapping imprinted genes considering the dynamic imprinting effect can provide additional biological insights into the epigenetic control of a complex trait. In this study, we proposed a Bayesian imprinted QTL (iQTL) mapping framework considering the dynamics of imprinting effects and model multiple iQTLs with an efficient Bayesian model selection procedure. The method overcomes the limitation of likelihood-based mapping procedure, and can simultaneously identify multiple iQTLs with different gene action modes across the whole genome with high computational efficiency. An inference procedure using Bayes factors to distinguish different imprinting patterns of iQTL was proposed. Monte Carlo simulations were conducted to evaluate the performance of the method. The utility of the approach was illustrated through an analysis of a body weight growth data set in an F(2) family derived from LG/J and SM/J mouse stains. The proposed Bayesian mapping method provides an efficient and computationally feasible framework for genome-wide multiple iQTL inference with complex developmental traits.
Runqing Yang, Yuehua Cui
Briefings Bioinform.1
2010 Bayesian model selection for characterizing genomic imprinting effects and patterns
abstract
MOTIVATION: Although imprinted genes have been ubiquitously observed in nature, statistical methodology still has not been systematically developed for jointly characterizing genomic imprinting effects and patterns. To detect imprinting genes influencing quantitative traits, the least square and maximum likelihood approaches for fitting a single quantitative trait loci (QTL) and Bayesian method for simultaneously modeling multiple QTLs have been adopted in various studies. RESULTS: In a widely used F(2) reciprocal mating population for mapping imprinting genes, we herein propose a genomic imprinting model which describes additive, dominance and imprinting effects of multiple imprinted quantitative trait loci (iQTL) for traits of interest. Depending upon the estimates of the above genetic effects, we categorized imprinting patterns into seven types, which provides a complete classification scheme for describing imprinting patterns. Bayesian model selection was employed to identify iQTL along with many genetic parameters in a computationally efficient manner. To make statistical inference on the imprinting types of iQTL detected, a set of Bayes factors were formulated using the posterior probabilities for the genetic effects being compared. We demonstrated the performance of the proposed method by computer simulation experiments and then applied this method to two real datasets. Our approach can be generally used to identify inheritance modes and determine the contribution of major genes for quantitative variations.
Runqing Yang, Zeyuan Wu, Daniel R. Prows
Bioinform.1
2009 Bayesian robust analysis for genetic architecture of quantitative traits
abstract
MOTIVATION: In most quantitative trait locus (QTL) mapping studies, phenotypes are assumed to follow normal distributions. Deviations from this assumption may affect the accuracy of QTL detection and lead to detection of spurious QTLs. To improve the robustness of QTL mapping methods, we replaced the normal distribution for residuals in multiple interacting QTL models with the normal/independent distributions that are a class of symmetric and long-tailed distributions and are able to accommodate residual outliers. Subsequently, we developed a Bayesian robust analysis strategy for dissecting genetic architecture of quantitative traits and for mapping genome-wide interacting QTLs in line crosses. RESULTS: Through computer simulations, we showed that our strategy had a similar power for QTL detection compared with traditional methods assuming normal-distributed traits, but had a substantially increased power for non-normal phenotypes. When this strategy was applied to a group of traits associated with physical/chemical characteristics and quality in rice, more main and epistatic QTLs were detected than traditional Bayesian model analyses under the normal assumption.
Runqing Yang, Jian Li 0012, Hong-Wen Deng
Bioinform.1