Qi Long

dblp:47/7320 · DBLP profile ↗
← Back
10ranked-venue papers in the field
0as first author
4since 2021 · last 2024
0000-0003-0660-5230ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 6Big Data, Cloud & Distributed Data Systems · 4
YearPublicationVenuePosition
2024 SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text
abstract
techniques that often evade existing detection methods. To address this challenge, we present a novel semantic-enhanced framework for detecting LLM-generated text (SEFD) that leverages a retrieval-based mechanism to fully utilize text semantics. Our framework improves upon existing detection methods by systematically integrating retrieval-based techniques with traditional detectors, employing a carefully curated retrieval mechanism that strikes a balance between comprehensive coverage and computational efficiency. We showcase the effectiveness of our approach in sequential text scenarios common in real-world applications, such as online forums and Q&A platforms. Through comprehensive experiments across various LLM-generated texts and detection methods, we demonstrate that our framework substantially enhances detection accuracy in paraphrasing scenarios while maintaining robustness for standard LLM-generated content. This work contributes significantly to ongoing efforts to safeguard information integrity in an era where AI-generated content is increasingly prevalent.
Weiqing He, Bojian Hou, Tianqi Shang, D. Ataee Tarzanagh, Qi Long, Li Shen 0001
IEEE Big Data5
2023 MISNN: Multiple Imputation via Semi-parametric Neural Networks
Zhiqi Bu, Zongyu Dai, Yiliang Zhang, Qi Long
PAKDD (1)4
2022 Differentially Private Bayesian Neural Networks on Accuracy, Privacy and Reliability
Qiyiwen Zhang, Zhiqi Bu, Qi Long
ECML/PKDD (4)4
2021 Graph-guided Bayesian SVM with Adaptive Structured Shrinkage Prior for High-dimensional Data
abstract
Support vector machine (SVM) is a popular classification method for the analysis of a wide range of data including big biomedical data. Many SVM methods with feature selection have been developed under the frequentist regularization or Bayesian shrinkage frameworks. On the other hand, the value of incorporating a priori known biological knowledge, such as those from functional genomics and functional proteomics, into statistical analysis of -omic data has been recognized in recent years. Such biological information is often represented by graphs. We propose a novel method that assigns Laplace priors to the regression coefficients and incorporates the underlying graph information via a hyper-prior for the shrinkage parameters in the Laplace priors. This enables smoothing of shrinkage parameters for connected variables in the graph and conditional independence between shrinkage parameters for disconnected variables. Extensive simulations demonstrate that our proposed methods achieve the best performance compared to the other existing SVM methods in terms of prediction accuracy. The proposed method are also illustrated in analysis of genomic data from cancer studies, demonstrating its advantage in generating biologically meaningful results and identifying potentially important features.
Wenli Sun, Changgee Chang, Qi Long
IEEE BigData3
2020 Joint Bayesian Variable Selection and Graph Estimation for Non-linear SVM with Application to Genomics Data
abstract
Support vector machine (SVM) is a powerful classification tool for analysis of high dimensional data such as genomics. Regularized linear and nonlinear SVM methods with feature selection have been developed. On the other hand, there is a growing body of literature showing that incorporating prior biological knowledge such as functional genomics, which are typically represented by graphs, into the analysis of genomic data can improve feature selection and prediction. In practice, however, such biological knowledge can often be inaccurate or unavailable. To attack this problem, we propose a Bayesian modeling approach which enables us to learn the graph structure among features and perform feature selection simultaneously. Our approach employs a Gaussian graphical model for inferring the graphical information and exploits the inferred graph to guide feature selection for SVM. An efficient MCMC algorithm is developed and our numerical analysis demonstrates that the proposed method has advantages over existing methods in feature selection and prediction via simulations and an application to the analysis of glioblastoma patient data.
Wenli Sun, Changgee Chang, Qi Long
DSAA3
2020 GRIA: Graphical Regularization for Integrative Analysis
abstract
Integrative analysis jointly analyzes multiple data sets to overcome curse of dimensionality. It can detect important but weak signals by jointly selecting features for all data sets, but unfortunately the sets of important features are not always the same for all data sets. Variations which allows heterogeneous sparsity structure-a subset of data sets can have a zero coefficient for a selected feature-have been proposed, but it compromises the effect of integrative analysis recalling the problem of losing weak important signals. We propose a new integrative analysis approach which not only aggregates weak important signals well in homogeneity setting but also substantially alleviates the problem of losing weak important signals in heterogeneity setting. Our approach exploits a priori known graphical structure of features by forcing joint selection of adjacent features, and integrating such information over multiple data sets can increase the power while taking into account the heterogeneity across data sets. We confirm the problem of existing approaches and demonstrate the superiority of our method through a simulation study and an application to gene expression data from ADNI.
Changgee Chang, Jihwan Oh, Qi Long
SDM3
2019 Bayesian Non-linear Support Vector Machine for High-Dimensional Data with Incorporation of Graph Information on Features
abstract
Support vector machine (SVM) is a popular classification method for analysis of high dimensional data such as genomics data. Recently a number of linear SVM methods have been developed to achieve feature selection through either frequentist regularization or Bayesian shrinkage, but the linear assumption may not be plausible for many real applications. In addition, recent work has demonstrated that incorporating known biological knowledge, such as those from functional genomics, into the statistical analysis of genomic data offers great promise of improved predictive accuracy and feature selection. Such biological knowledge can often be represented by graphs. In this article, we propose a novel knowledge-guided nonlinear Bayesian SVM approach for analysis of high-dimensional data. Our model uses graph information that represents the relationship among the features to guide feature selection. To achieve knowledge-guided feature selection, we assign an Ising prior to the indicators representing inclusion/exclusion of the features in the model. An efficient MCMC algorithm is developed for posterior inference. The performance of our method is evaluated and compared with several penalized linear SVM and the standard kernel SVM method in terms of prediction and feature selection in extensive simulation studies. Also, analyses of genomic data from a cancer study show that our method yields a more accurate prediction model for patient survival and reveals biologically more meaningful results than the existing methods.
Wenli Sun, Changgee Chang, Qi Long
IEEE BigData3
2018 Knowledge-Guided Bayesian Support Vector Machine for High-Dimensional Data with Application to Analysis of Genomics Data
abstract
Support vector machine (SVM) is a popular classification method for the analysis of wide range of data including big data. Many SVM methods with feature selection have been developed under frequentist regularization or Bayesian shrinkage frameworks. On the other hand, the importance of incorporating a priori known biological knowledge, such as gene pathway information which stems from the gene regulatory network, into the statistical analysis of genomic data has been recognized in recent years. In this article, we propose a new Bayesian SVM approach that enables the feature selection to be guided by the knowledge on the graphical structure among predictors. The proposed method uses the spike-and-slab prior for feature selection, combined with the Ising prior that encourages group-wise selection of the predictors adjacent to each other on the known graph. Gibbs sampling algorithm is used for Bayesian inference. The performance of our method is evaluated and compared with existing SVM methods in terms of prediction and feature selection in extensive simulation settings. In addition, our method is illustrated in the analysis of genomic data from a cancer study, demonstrating its advantage in generating biologically meaningful results and identifying potentially important features.
Wenli Sun, Changgee Chang, Yize Zhao, Qi Long
IEEE BigData4
2018 Generalized Bayesian Factor Analysis for Integrative Clustering with Applications to Multi-Omics Data
abstract
Integrative clustering is a clustering approach for multiple datasets, which provide different views of a common group of subjects. It enables analyzing multi-omics data jointly to, for example, identify the subtypes of diseases, cells, and so on, capturing the complex underlying biological processes more precisely. On the other hand, there has been a great deal of interest in incorporating the prior structural knowledge on the features into statistical analyses over the past decade. The knowledge on the gene regulatory network (pathways) can potentially be incorporated into many genomic studies. In this paper, we propose a novel integrative clustering method which can incorporate the prior graph knowledge. We first develop a generalized Bayesian factor analysis (GBFA) framework, a sparse Bayesian factor analysis which can take into account the graph information. Our GBFA framework employs the spike and slab lasso (SSL) prior to impose sparsity on the factor loadings and the Markov random field (MRF) prior to encourage smoothing over the adjacent factor loadings, which establishes a unified shrinkage adaptive to the loading size and the graph structure. Then, we use the framework to extend iCluster+, a factor analysis based integrative clustering approach. A novel variational EM algorithm is proposed to efficiently estimate the MAP estimator for the factor loadings. Extensive simulation studies and the application to the NCI60 cell line dataset demonstrate that the propose method is superior and delivers more biologically meaningful outcomes.
Eun Jeong Min, Changgee Chang, Qi Long
DSAA3
2016 Sparse Linear Discriminant Analysis in Structured Covariates Space
abstract
Classification with high dimensional variables is a popular goal in many modern statistical studies. Fisher's linear discriminant analysis (LDA) is a common and effective tool for classifying entities into existing groups. It is well known that classification using Fisher's discriminant for high dimensional data is as bad as random guessing due to the many noise features that increases misclassification rate. Recently, it is being acknowledged that complex biological mechanisms occur through multiple features working together, though individually these features may contribute to noise accumulation in the data. In view of these, it is important to perform classification with discriminant vectors that use a subset of important variables, while also utilizing prior biological relationships among features. We tackle this problem in this article and propose methods that incorporate variable selection into the classification problem, for the identification of important biomarkers. Furthermore, we incorporate into the LDA problem prior information on the relationships among variables using undirected graphs in order to identify functionally meaningful biomarkers. We compare our methods to existing sparse LDA approaches via simulation studies and real data analysis.
Sandra Safo, Qi Long
DSAA2