VLDB 2026 Research / reviewers in the wild / expert
Qian Xu 0005
dblp:81/5941-5
· DBLP profile ↗
23ranked-venue papers
4as first author
12since 2021 · last 2024
0009-0004-7726-5928ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FedCORE: Federated Learning for Cross-Organization Recommendation EcosystemabstractA recommendation system is of vital importance in delivering personalization services, which often brings continuous dual improvement in user experience and organization revenue. However, the data of one single organization may not be enough to build an accurate recommendation model for inactive or new cold-start users. Moreover, due to the recent regulatory restrictions on user privacy and data security, as well as the commercial conflicts, the raw data in different organizations cannot be merged to alleviate the scarcity issue in training a model. In order to learn users’ preferences from such cross-silo data of different organizations and then provide recommendations to the cold-start users, we propose a novel federated learning framework, i.e., federated cross-organization recommendation ecosystem (FedCORE). Specifically, we first focus on the ecosystem problem of cross-organization federated recommendation, including cooperation patterns and privacy protection. For the former, we propose a privacy-aware collaborative training and inference algorithm. For the latter, we define four levels of privacy leakage and propose some methods for protecting the privacy. We then conduct extensive experiments on three real-world datasets and two seminal recommendation models to study the impact of cooperation in our proposed ecosystem and the effectiveness of privacy protection. Zhitao Li 0005, Xueyang Wu 0001, Weike Pan, Youlong Ding, Zeheng Wu, Shengqi Tan, Qian Xu 0005, Qiang Yang 0001, Zhong Ming 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2023 | FedNP: Towards Non-IID Federated Learning via Federated Neural PropagationabstractTraditional federated learning (FL) algorithms, such as FedAvg, fail to handle non-i.i.d data because they learn a global model by simply averaging biased local models that are trained on non-i.i.d local data, therefore failing to model the global data distribution. In this paper, we present a novel Bayesian FL algorithm that successfully handles such a non-i.i.d FL setting by enhancing the local training task with an auxiliary task that explicitly estimates the global data distribution. One key challenge in estimating the global data distribution is that the data are partitioned in FL, and therefore the ground-truth global data distribution is inaccessible. To address this challenge, we propose an expectation-propagation-inspired probabilistic neural network, dubbed federated neural propagation (FedNP), which efficiently estimates the global data distribution given non-i.i.d data partitions. Our algorithm is sampling-free and end-to-end differentiable, can be applied with any conventional FL frameworks and learns richer global data representation. Experiments on both image classification tasks with synthetic non-i.i.d image data partitions and real-world non-i.i.d speech recognition tasks demonstrate that our framework effectively alleviates the performance deterioration caused by non-i.i.d data. Xueyang Wu 0001, Hengguan Huang, Youlong Ding, Hao Wang 0014, Ye Wang 0007, Qian Xu 0005 |
AAAI | 6 |
| 2022 | Contribution-Aware Federated Learning for Smart HealthcareabstractArtificial intelligence (AI) is a promising technology to transform the healthcare industry. Due to the highly sensitive nature of patient data, federated learning (FL) is often leveraged to build models for smart healthcare applications. Existing deployed FL frameworks cannot address the key issues of varying data quality and heterogeneous data distributions across multiple institutions in this sector. In this paper, we report our experience developing and deploying the Contribution-Aware Federated Learning (CAFL) framework for smart healthcare. It provides an efficient and accurate approach to fairly evaluate FL participants' contribution to model performance without exposing their private data, and improves the FL model training protocol to allow the best performing intermediate models to be distributed to participants for FL training. Since its deployment in Yidu Cloud Technology Inc. in March 2021, CAFL has served 8 well-established medical institutions in China to build healthcare decision support models. It can perform contribution evaluations 2.84 times faster than the best existing approach, and has improved the average accuracy of the resulting models by 2.62% compared to the previous system (which is significant in industrial settings). To our knowledge, it is the first contribution-aware federated learning successfully deployed in the healthcare industry. Zelei Liu, Yuanyuan Chen 0012, Yansong Zhao, Han Yu 0001, Yang Liu 0165, Renyi Bao, Jinpeng Jiang, Zaiqing Nie, Qian Xu 0005, Qiang Yang 0001 |
AAAI | 9 |
| 2022 | A Platform for Deploying the TFE Ecosystem of Automatic Speech RecognitionabstractSince data regulations such as the European Union's General Data Protection Regulation (GDPR) have taken effect, the traditional two-step Automatic Speech Recognition (ASR) optimization strategy (i.e., training a one-size-fits-all model with vendor's centralized data and fine-tuning the model with clients' private data) has become infeasible. To meet these privacy requirements, TFE, a novel GDPR-compliant ASR ecosystem, has been proposed by us to incorporate transfer learning, federated learning, and evolutionary learning towards effective ASR model optimization. In this demonstration, we further design and implement a novel platform to promote the deployment and applicability of TFE. Our proposed platform allows enterprises to easily conduct the ASR optimization task using TFE across organizations. Yuanfeng Song, Rongzhong Lian, Di Jiang 0004, Xuefang Zhao, Conghui Tan, Qian Xu 0005, Raymond Chi-Wing Wong |
ACM Multimedia | 7 |
| 2021 | A Health-friendly Speaker Verification System Supporting Mask WearingabstractWe demonstrate a health-friendly speaker verification system for voice-based identity verification on mobile devices. The system is built upon a speech processing module, a ResNet-based local acoustic feature extractor and a multi-head attention-based embedding layer, and is optimized under an additive margin softmax loss for discriminative speaker verification. It is shown that the system achieves superior performance no matter whether there is mask wearing or not. This characteristic is important for speaker verification services operating in regions affected by the raging coronavirus pneumonia. With this demonstration, the audience will have an in-depth experience of how the accuracy of bio-metric verification and the personal health are simultaneously ensured. We wish that this demonstration would boost the development of next-generation bio-metric verification technologies. Chaotao Chen, Di Jiang 0004, Jinhua Peng, Rongzhong Lian, Chen Zhang 0013, Qian Xu 0005, Lixin Fan, Qiang Yang 0001 |
AAAI | 6 |
| 2021 | L2RS: A Learning-to-Rescore Mechanism for Hybrid Speech RecognitionabstractThis paper aims to advance the performance of industrial ASR systems by exploring a more effective method for N-best rescoring, a critical step that greatly affects the final recognition accuracy. Existing rescoring approaches suffer the following issues: (i) limited performance since they optimize an unnecessarily harder problem, namely predicting accurate grammatical legitimacy scores of the N-best hypotheses rather than directly predicting their partial orders regarding a specific acoustic input; (ii) hard to incorporate various information by advanced natural language processing (NLP) models such as BERT to achieve a comprehensive evaluation of each N-best candidate. To relieve the above drawbacks, we propose a simple yet effective mechanism, Learning-to-Rescore (L2RS), to empower ASR systems with state-of-the-art information retrieval (IR) techniques. Specifically, L2RS utilizes a wide range of textual information from the state-of-the-art NLP models and automatically deciding their weights to directly learn the ranking order of each N-best hypothesis with respect to a specific acoustic input. We incorporate various features including BERT sentence embeddings, the topic vectors, and perplexity scores produced by an n-gram language model (LM), topic modeling LM, BERT, and RNNLM to train the rescoring model. Experimental results on a public dataset show that L2RS outperforms not only traditional rescoring methods but also its deep neural network counterparts by a substantial margin of 20.85% in terms of [email protected] The L2RS toolkit has been successfully deployed for many online commercial services in WeBank Co., Ltd, China's leading digital bank. The efficacy and applicability of L2RS are validated by real-life online customer datasets. Yuanfeng Song, Di Jiang 0004, Xuefang Zhao, Qian Xu 0005, Raymond Chi-Wing Wong, Lixin Fan, Qiang Yang 0001 |
ACM Multimedia | 4 |
| 2021 | SmartMeeting: Automatic Meeting Transcription and Summarization for In-Person ConversationsabstractMeetings are a necessary part of the operations of any institution, whether they are held online or in-person. However, meeting transcription and summarization are always painful requirements since they involve tedious human effort. This drives the need for automatic meeting transcription and summarization (AMTS) systems. A successful AMTS system relies on systematic integration of multiple natural language processing (NLP) techniques, such as automatic speech recognition, speaker identification, and meeting summarization, which are traditionally developed separately and validated offline with standard datasets. In this demonstration, we provide a novel productive meeting tool named SmartMeeting, which enables users to automatically record, transcribe, summarize, and manage the information in an in-person meeting. SmartMeeting transcribes every word on the fly, enriches the transcript with speaker identification and voice separation, and extracts essential decisions and crucial insights automatically. In our demonstration, the audience can experience the great potential of the state-of-the-art NLP techniques in this real-life application. Yuanfeng Song, Di Jiang 0004, Xuefang Zhao, Qian Xu 0005, Raymond Chi-Wing Wong, Qiang Yang 0001 |
ACM Multimedia | 5 |
| 2021 | SmartSales: An AI-Powered Telemarketing Coaching System in FinTechabstractTelemarketing is a primary and mature method for enterprises to solicit prospective customers to buy products or services. However, training telesales representatives is always a pain point for enterprises since it is usually conducted manually and costs great effort and time. In this demonstration, we propose a telemarketing coaching system named SmartSales to help enterprises develop better salespeople. Powered by artificial intelligence (AI), SmartSales aims to accumulate the experienced sales pitch from customer-sales dialogues and use it to coach junior salespersons. To the best of our knowledge, this is the first practice of an AI telemarketing coaching system in the domain of Chinese FinTech in the literature. SmartSales has been successfully deployed in the WeBank's telemarketing team. We expect that SmartSales will inspire more research on AI assistant systems. Yuanfeng Song, Xuefang Zhao, Di Jiang 0004, Qian Xu 0005, Raymond Chi-Wing Wong, Qiang Yang 0001 |
ACM Multimedia | 6 |
| 2021 | Memetic Federated Learning for Biomedical Natural Language Processing
Xinya Zhou, Conghui Tan, Di Jiang 0004, Bosen Zhang, Si Li 0001, Qian Xu 0005, Sheng Gao 0001 |
NLPCC (2) | 7 |
| 2021 | FATE: An Industrial Grade Platform for Collaborative Learning With Data ProtectionabstractCollaborative and federated learning has become an emerging solution to many industrial applications where data values from different sites are exploit jointly with privacy protection. We introduce FATE, an industrial-grade project that supports enterprises and institutions to build machine learning models collaboratively at large-scale in a distributed manner. FATE supports a variety of secure computation protocols and machine learning algorithms, and features out-of-box usability with end-to-end building modules and visualization tools. Documentations are available at https://github.com/FederatedAI/FATE. Case studies and other information are available at https://www.fedai.org. Yang Liu 0165, Tao Fan 0002, Tianjian Chen, Qian Xu 0005, Qiang Yang 0001 |
J. Mach. Learn. Res. | 4 |
| 2021 | A GDPR-compliant Ecosystem for Speech Recognition with Transfer, Federated, and Evolutionary LearningabstractAutomatic Speech Recognition (ASR) is playing a vital role in a wide range of real-world applications. However, Commercial ASR solutions are typically “one-size-fits-all” products and clients are inevitably faced with the risk of severe performance degradation in field test. Meanwhile, with new data regulations such as the European Union’s General Data Protection Regulation (GDPR) coming into force, ASR vendors, which traditionally utilize the speech training data in a centralized approach, are becoming increasingly helpless to solve this problem, since accessing clients’ speech data is prohibited. Here, we show that by seamlessly integrating three machine learning paradigms (i.e., T ransfer learning, F ederated learning, and E volutionary learning (TFE)), we can successfully build a win-win ecosystem for ASR clients and vendors and solve all the aforementioned problems plaguing them. Through large-scale quantitative experiments, we show that with TFE, the clients can enjoy far better ASR solutions than the “one-size-fits-all” counterpart, and the vendors can exploit the abundance of clients’ data to effectively refine their own ASR products. Di Jiang 0004, Conghui Tan, Jinhua Peng, Chaotao Chen, Xueyang Wu 0001, Yuanfeng Song, Yongxin Tong, Chang Liu 0069, Qian Xu 0005, Qiang Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 10 |
| 2021 | Industrial Federated Topic ModelingabstractProbabilistic topic modeling has been applied in a variety of industrial applications. Training a high-quality model usually requires a massive amount of data to provide comprehensive co-occurrence information for the model to learn. However, industrial data such as medical or financial records are often proprietary or sensitive, which precludes uploading to data centers. Hence, training topic models in industrial scenarios using conventional approaches faces a dilemma: A party (i.e., a company or institute) has to either tolerate data scarcity or sacrifice data privacy. In this article, we propose a framework named Industrial Federated Topic Modeling (iFTM), in which multiple parties collaboratively train a high-quality topic model by simultaneously alleviating data scarcity and maintaining immunity to privacy adversaries. iFTM is inspired by federated learning, supports two representative topic models (i.e., Latent Dirichlet Allocation and SentenceLDA) in industrial applications, and consists of novel techniques such as private Metropolis-Hastings, topic-wise normalization, and heterogeneous model integration. We conduct quantitative evaluations to verify the effectiveness of iFTM and deploy iFTM in two real-life applications to demonstrate its utility. Experimental results verify iFTM’s superiority over conventional topic modeling. Di Jiang 0004, Yongxin Tong, Yuanfeng Song, Xueyang Wu 0001, Jinhua Peng, Rongzhong Lian, Qian Xu 0005, Qiang Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2020 | Federated Acoustic Model Optimization for Automatic Speech Recognition
Conghui Tan, Di Jiang 0004, Huaxiao Mo, Jinhua Peng, Yongxin Tong, Chaotao Chen, Rongzhong Lian, Yuanfeng Song, Qian Xu 0005 |
DASFAA (3) | 10 |
| 2020 | A De Novo Divide-and-Merge Paradigm for Acoustic Model Optimization in Automatic Speech RecognitionabstractDue to the rising awareness of privacy protection and the voluminous scale of speech data, it is becoming infeasible for Automatic Speech Recognition (ASR) system developers to train the acoustic model with complete data as before. In this paper, we propose a novel Divide-and-Merge paradigm to solve salient problems plaguing the ASR field. In the Divide phase, multiple acoustic models are trained based upon different subsets of the complete speech data, while in the Merge phase two novel algorithms are utilized to generate a high-quality acoustic model based upon those trained on data subsets. We first propose the Genetic Merge Algorithm (GMA), which is a highly specialized algorithm for optimizing acoustic models but suffers from low efficiency. We further propose the SGD-Based Optimizational Merge Algorithm (SOMA), which effectively alleviates the efficiency bottleneck of GMA and maintains superior performance. Extensive experiments on public data show that the proposed methods can significantly outperform the state-of-the-art. Conghui Tan, Di Jiang 0004, Jinhua Peng, Xueyang Wu 0001, Qian Xu 0005, Qiang Yang 0001 |
IJCAI | 5 |
| 2020 | GoldenRetriever: A Speech Recognition System Powered by Modern Information RetrievalabstractExisting Automatic Speech Recognition (ASR) systems usually generate the N-best hypotheses list first, and then rescore them with the language model score and the acoustic model score to find the best one. This procedure is essentially analogous to the working mechanism of modern Information Retrieval (IR) systems, which retrieve a relatively large amount of relevant candidates first, re-rank them, and output the top-N list. Exploiting their commonality, this demonstration proposes a novel system named GoldenRetriever that marries IR with ASR. GoldenRetriever transforms the problem of N-best hypotheses rescoring as a Learning-to-Rescore (L2RS) problem and utilizes a wide range of features beyond the language model score and the acoustic model score. In this demonstration, the audience can experience the great potential of marrying IR with ASR for the first time. GoldenRetriever should inspire more research on transferring the state-of-the-art IR techniques to ASR. Yuanfeng Song, Di Jiang 0004, Yawen Li 0001, Qian Xu 0005, Raymond Chi-Wing Wong, Qiang Yang 0001 |
ACM Multimedia | 5 |
| 2020 | Graph Random Neural Networks for Semi-Supervised Learning on GraphsabstractWe study the problem of semi-supervised learning on graphs, for which graph neural networks (GNNs) have been extensively explored. However, most existing GNNs inherently suffer from the limitations of over-smoothing, non-robustness, and weak-generalization when labeled nodes are scarce. In this paper, we propose a simple yet effective framework—GRAPH RANDOM NEURAL NETWORKS (GRAND)—to address these issues. In GRAND, we first design a random propagation strategy to perform graph data augmentation. Then we leverage consistency regularization to optimize the prediction consistency of unlabeled nodes across different data augmentations. Extensive experiments on graph benchmark datasets suggest that GRAND significantly outperforms state-of- the-art GNN baselines on semi-supervised node classification. Finally, we show that GRAND mitigates the issues of over-smoothing and non-robustness, exhibiting better generalization behavior than existing GNNs. The source code of GRAND is publicly available at https://github.com/Grand20/grand. Wenzheng Feng, Jie Zhang 0078, Yuxiao Dong, Yu Han 0001, Huan-Bo Luan, Qian Xu 0005, Qiang Yang 0001, Evgeny Kharlamov, Jie Tang 0001 |
NeurIPS | 6 |
| 2019 | Federated Topic ModelingabstractTopic modeling has been widely applied in a variety of industrial applications. Training a high-quality model usually requires massive amount of in-domain data, in order to provide comprehensive co-occurrence information for the model to learn. However, industrial data such as medical or financial records are often proprietary or sensitive, which precludes uploading to data centers. Hence training topic models in industrial scenarios using conventional approaches faces a dilemma: a party (i.e., a company or institute) has to either tolerate data scarcity or sacrifice data privacy. In this paper, we propose a novel framework named Federated Topic Modeling (FTM), in which multiple parties collaboratively train a high-quality topic model by simultaneously alleviating data scarcity and maintaining immune to privacy adversaries. FTM is inspired by federated learning and consists of novel techniques such as private Metropolis Hastings, topic-wise normalization and heterogeneous model integration. We conduct a series of quantitative evaluations to verify the effectiveness of FTM and deploy FTM in an Automatic Speech Recognition (ASR) system to demonstrate its utility in real-life applications. Experimental results verify FTM's superiority over conventional topic modeling. Di Jiang 0004, Yuanfeng Song, Yongxin Tong, Xueyang Wu 0001, Qian Xu 0005, Qiang Yang 0001 |
CIKM | 6 |
| 2019 | Topic-Aware Dialogue Speech Recognition with Transfer LearningabstractDialogue speech widely exists in scenarios such as chitchat, meeting and customer service. General-purpose speech recognition systems usually neglect the topic information in the context of dialogue speech, which has great potential for improving the performance of speech recognition. In this paper, we propose a transfer learning mechanism to conduct topic-aware recognition for dialogue speech. We first propose a new probabilistic topic model named Dialogue Speech Topic Model (DSTM) that is specialized for modeling the context of dialogue speech. We further propose a novel transfer learning mechanism for DSTM to significantly reduce its training cost while preserving its effectiveness for accurate topic inference. The experiment results demonstrate that proposed techniques in language model adaptation effectively improve the performance of the state-of-the-art Automatic Speech Recognition (ASR) system. Copyright © 2019 ISCA Yuanfeng Song, Di Jiang 0004, Xueyang Wu 0001, Qian Xu 0005, Raymond Chi-Wing Wong, Qiang Yang 0001 |
INTERSPEECH | 4 |
| 2011 | Multitask Learning for Protein Subcellular Location PredictionabstractProtein subcellular localization is concerned with predicting the location of a protein within a cell using computational methods. The location information can indicate key functionalities of proteins. Thus, accurate prediction of subcellular localizations of proteins can help the prediction of protein functions and genome annotations, as well as the identification of drug targets. Machine learning methods such as Support Vector Machines (SVMs) have been used in the past for the problem of protein subcellular localization, but have been shown to suffer from a lack of annotated training data in each species under study. To overcome this data sparsity problem, we observe that because some of the organisms may be related to each other, there may be some commonalities across different organisms that can be discovered and used to help boost the data in each localization task. In this paper, we formulate protein subcellular localization problem as one of multitask learning across different organisms. We adapt and compare two specializations of the multitask learning algorithms on 20 different organisms. Our experimental results show that multitask learning performs much better than the traditional single-task methods. Among the different multitask learning methods, we found that the multitask kernels and supertype kernels under multitask learning that share parameters perform slightly better than multitask learning by sharing latent features. The most significant improvement in terms of localization accuracy is about 25 percent. We find that if the organisms are very different or are remotely related from a biological point of view, then jointly training the multiple models cannot lead to significant improvement. However, if they are closely related biologically, the multitask learning can do much better than individual learning. Qian Xu 0005, Sinno Jialin Pan, Hong Xue 0001, Qiang Yang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | Protein-protein interaction prediction via Collective Matrix FactorizationabstractProtein-protein interactions (PPI) play an important role in cellular processes and metabolic processes within a cell. An important task is to determine the existence of interactions among proteins. Unfortunately, existing biological experimental techniques are expensive, time-consuming and labor-intensive. The network structures of many such networks are sparse, incomplete and noisy, containing many false positive and false negatives. Thus, state-of-the-art methods for link prediction in these networks often cannot give satisfactory prediction results, especially when some networks are extremely sparse. Noticing that we typically have more than one PPI network available, we naturally wonder whether it is possible to 'transfer' the linkage knowledge from some existing, relatively dense networks to a sparse network, to improve the prediction performance. Noticing that a network structure can be modeled using a matrix model, in this paper, we introduce the well-known Collective Matrix Factorization (CMF) technique to 'transfer' usable linkage knowledge from relatively dense interaction network to a sparse target network. Our approach is to establish the correspondence between a source and a target network via network similarities. We test this method on two real protein-protein interaction networks, Helicobacter pylori (as a target network) and Human (as a source network). Our experimental results show that our method can achieve higher and more robust performance as compared to some baseline methods. Qian Xu 0005, Evan Wei Xiang, Qiang Yang 0001 |
BIBM | 1 |
| 2010 | Predicting chemical activities from structures by attributed molecular graph classificationabstractDesigning Quantitative Structure-Activity Relationship (QSAR) models has been a recurrent research interest for biologists and computer scientists. An example is to predict the toxicity of chemical compounds using their structural properties as features represented by graphs. A popular method to classify these graphs is to exploit classifiers such as support vector machines (SVMs) and graph kernels to incorporate the sequential, structural and chemical information. Previous works have focused on designing specific graph kernels for this task, amongst which graph alignment kernels are one of the most popular approach. Graph alignment kernels align the nodes of one graph to the nodes of the second graph so that the total overall similarity is maximized with respect to all possible alignments. However, taking both vertex and edge similarities into account makes the problem NP-Hard. In this paper, we present a novel general graph-matching based method for QSAR. We view the problem of calculating optimal assignments of two attributed graphs from a different perspective. Instead of first designing an atom kernel function and a bond kernel function, we first provide a training set of pairs of graphs with their corresponding matchings. We then try to learn the compatibility function over atoms and use only the atom kernel function to compute graph matchings. Our algorithm has the advantage of being more general and yet efficient than previous approaches for the QSAR problem. We evaluate our method on a set of chemical structure-activity prediction benchmark datasets, and show that our algorithm can achieve better or comparable accuracies over the optimal assignment kernel method. Qian Xu 0005, Derek Hao Hu, Hong Xue 0001, Qiang Yang 0001 |
CIBCB | 1 |
| 2010 | Multi-task learning for cross-platform siRNA efficacy prediction: an in-silico studyabstractBACKGROUND: Gene silencing using exogenous small interfering RNAs (siRNAs) is now a widespread molecular tool for gene functional study and new-drug target identification. The key mechanism in this technique is to design efficient siRNAs that incorporated into the RNA-induced silencing complexes (RISC) to bind and interact with the mRNA targets to repress their translations to proteins. Although considerable progress has been made in the computational analysis of siRNA binding efficacy, few joint analysis of different RNAi experiments conducted under different experimental scenarios has been done in research so far, while the joint analysis is an important issue in cross-platform siRNA efficacy prediction. A collective analysis of RNAi mechanisms for different datasets and experimental conditions can often provide new clues on the design of potent siRNAs. RESULTS: An elegant multi-task learning paradigm for cross-platform siRNA efficacy prediction is proposed. Experimental studies were performed on a large dataset of siRNA sequences which encompass several RNAi experiments recently conducted by different research groups. By using our multi-task learning method, the synergy among different experiments is exploited and an efficient multi-task predictor for siRNA efficacy prediction is obtained. The 19 most popular biological features for siRNA according to their jointly importance in multi-task learning were ranked. Furthermore, the hypothesis is validated out that the siRNA binding efficacy on different messenger RNAs(mRNAs) have different conditional distribution, thus the multi-task learning can be conducted by viewing tasks at an "mRNA"-level rather than at the "experiment"-level. Such distribution diversity derived from siRNAs bound to different mRNAs help indicate that the properties of target mRNA have important implications on the siRNA binding efficacy. CONCLUSIONS: The knowledge gained from our study provides useful insights on how to analyze various cross-platform RNAi data for uncovering of their complex mechanism. Qi Liu 0019, Qian Xu 0005, Vincent Wenchen Zheng, Hong Xue 0001, Qiang Yang 0001 |
BMC Bioinform. | 2 |
| 2009 | Semi-supervised protein subcellular localizationabstractBACKGROUND: Protein subcellular localization is concerned with predicting the location of a protein within a cell using computational method. The location information can indicate key functionalities of proteins. Accurate predictions of subcellular localizations of protein can aid the prediction of protein function and genome annotation, as well as the identification of drug targets. Computational methods based on machine learning, such as support vector machine approaches, have already been widely used in the prediction of protein subcellular localization. However, a major drawback of these machine learning-based approaches is that a large amount of data should be labeled in order to let the prediction system learn a classifier of good generalization ability. However, in real world cases, it is laborious, expensive and time-consuming to experimentally determine the subcellular localization of a protein and prepare instances of labeled data. RESULTS: In this paper, we present an approach based on a new learning framework, semi-supervised learning, which can use much fewer labeled instances to construct a high quality prediction model. We construct an initial classifier using a small set of labeled examples first, and then use unlabeled instances to refine the classifier for future predictions. CONCLUSION: Experimental results show that our methods can effectively reduce the workload for labeling data using the unlabeled data. Our method is shown to enhance the state-of-the-art prediction results of SVM classifiers by more than 10%. Qian Xu 0005, Derek Hao Hu, Hong Xue 0001, Weichuan Yu, Qiang Yang 0001 |
BMC Bioinform. | 1 |