Feng Luo 0001

dblp:l/FengLuo · DBLP profile ↗
← Back
61ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0002-4813-2403ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-authorComputer networks · 3 · 1 since 2021Security and privacy · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Guest Editorial for the 21th Asia Pacific Bioinformatics Conference
Min Li 0007, Feng Luo 0001, Yi-Ping Phoebe Chen
IEEE Trans. Comput. Biol. Bioinform.2
2025 Few-Shot Generalized Category Discovery With Retrieval-Guided Decision Boundary Enhancement
Yunhan Ren, Feng Luo 0001, Siyu Huang
ICMR2
2024 Effectively Prompting Small-Sized Language Models for Cross-Lingual Tasks via Winning Tickets
abstract
Current soft prompt methods yield limited performance when applied to small-sized models (fewer than a billion parameters). Deep prompt-tuning, which entails prepending parameters in each layer for enhanced efficacy, presents a solution for prompting small-sized models, albeit requiring carefully designed implementation. In this paper, we introduce the Lottery Ticket Prompt-learning (LTP) framework that integrates winning tickets with soft prompts. The LTP offers a simpler implementation and requires only a one-time execution. We demonstrate LTP on cross-lingual tasks, where prior works rely on external tools like human-designed multilingual templates and bilingual dictionaries, which may not be feasible in a low-resource regime. Specifically, we select a subset of parameters that have been changed the most during the fine-tuning with the Masked Language Modeling objective. Then, we prepend soft prompts to the original pretrained language model and only update the selected parameters together with prompt-related parameters when adapting to the downstream tasks. We verify the effectiveness of our LTP framework on cross-lingual tasks, specifically targeting low-resource languages. Our approach outperforms the baselines by only updating 20% of the original parameters.
Feng Luo 0001
ICMLA2
2024 RNA m6A detection using raw current signals and basecalling errors from Nanopore direct RNA sequencing reads
abstract
MOTIVATION: Nanopore direct RNA sequencing (DRS) enables the detection of RNA N6-methyladenosine (m6A) without extra laboratory techniques. A number of supervised or comparative approaches have been developed to identify m6A from Nanopore DRS reads. However, existing methods typically utilize either statistical features of the current signals or basecalling-error features, ignoring the richer information of the raw signals of DRS reads. RESULTS: Here, we propose RedNano, a deep-learning method designed to detect m6A from Nanopore DRS reads by utilizing both raw signals and basecalling errors. RedNano processes the raw-signal feature and basecalling-error feature through residual networks. We validated the effectiveness of RedNano using synthesized, Arabidopsis, and human DRS data. The results demonstrate that RedNano surpasses existing methods by achieving higher area under the ROC curve (AUC) and area under the precision-recall curve (AUPRs) in all three datasets. Furthermore, RedNano performs better in cross-species validation, demonstrating its robustness. Additionally, when detecting m6A from an independent dataset of Populus trichocarpa, RedNano achieves the highest AUC and AUPR, which are 3.8%-9.9% and 5.5%-13.8% higher than other methods, respectively. AVAILABILITY AND IMPLEMENTATION: The source code of RedNano is freely available at https://github.com/Derryxu/RedNano.
Jinrui Xu, Zeyu Zhong, Feng Luo 0001, Jianxin Wang 0001
Bioinform.4
2024 Phasing nanopore genome assembly by integrating heterozygous variations and Hi-C data
abstract
MOTIVATION: Haplotype-resolved genome assemblies serve as vital resources in various research domains, including genomics, medicine, and pangenomics. Algorithms employing Hi-C data to generate haplotype-resolved assemblies are particularly advantageous due to its ready availability. Existing methods primarily depend on mapping quality to filter out uninformative Hi-C alignments which may be susceptible to sequencing errors. Setting a high mapping quality threshold filters out numerous informative Hi-C alignments, whereas a low mapping quality threshold compromises the accuracy of Hi-C alignments. Maintaining high accuracy while retaining a maximum number of Hi-C alignments can be challenging. RESULTS: In our experiments, heterozygous variations play an important role in filtering uninformative Hi-C alignments. Here, we introduce Diphase, a novel phasing tool that harnesses heterozygous variations to accurately identify the informative Hi-C alignments for phasing and to extend primary/alternate assemblies. Diphase leverages mapping quality and heterozygous variations to filter uninformative Hi-C alignments, thereby enhancing the accuracy of phasing and the detection of switches. To validate its performance, we conducted a comparative analysis of Diphase, FALCON-Phase, and GFAse on various human datasets. The results demonstrate that Diphase achieves a longer phased block N50 and exhibits higher phasing accuracy while maintaining a lower hamming error rate. AVAILABILITY AND IMPLEMENTATION: The source code of Diphase is available at https://github.com/zhangjuncsu/Diphase.
Fan Nie, Feng Luo 0001, Jianxin Wang 0001
Bioinform.3
2024 Dual-Level Knowledge Distillation via Knowledge Alignment and Correlation
abstract
Knowledge distillation (KD) has become a widely used technique for model compression and knowledge transfer. We find that the standard KD method performs the knowledge alignment on an individual sample indirectly via class prototypes and neglects the structural knowledge between different samples, namely, knowledge correlation. Although recent contrastive learning-based distillation methods can be decomposed into knowledge alignment and correlation, their correlation objectives undesirably push apart representations of samples from the same class, leading to inferior distillation results. To improve the distillation performance, in this work, we propose a novel knowledge correlation objective and introduce the dual-level knowledge distillation (DLKD), which explicitly combines knowledge alignment and correlation together instead of using one single contrastive objective. We show that both knowledge alignment and correlation are necessary to improve the distillation performance. In particular, knowledge correlation can serve as an effective regularization to learn generalized representations. The proposed DLKD is task-agnostic and model-agnostic, and enables effective knowledge transfer from supervised or self-supervised pretrained teachers to students. Experiments show that DLKD outperforms other state-of-the-art methods on a large number of experimental settings including: 1) pretraining strategies; 2) network architectures; 3) datasets; and 4) tasks.
Yin Yang 0002, Hongxin Hu, Venkat N. Krovi, Feng Luo 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Is GPT Powerful Enough to Analyze the Emotions of Memes?
abstract
Large Language Models (LLMs), representing a significant achievement in artificial intelligence (AI) research, have demonstrated their ability in a multitude of tasks. This project aims to explore the capabilities of GPT-3.5, a leading example of LLMs, in processing the sentiment analysis of Internet memes. Memes, which include both verbal and visual aspects, act as a powerful yet complex tool for expressing ideas and sentiments, demanding an understanding of societal norms and cultural contexts. Notably, the detection and moderation of hateful memes pose a significant challenge due to their implicit offensive nature. This project investigates GPT's proficiency in such subjective tasks, revealing its strengths and potential limitations. The tasks include the classification of meme sentiment, determination of humor type, and detection of implicit hate in memes. The performance evaluation, using datasets from SemEval-2020 Task 8 and Facebook hateful memes, offers a comparative understanding of GPT responses against human annotations. Despite GPT's remarkable progress, our findings underscore the challenges faced by these models in handling subjective tasks, which are rooted in their inherent limitations including contextual understanding, interpretation of implicit meanings, and data biases. This research contributes to the broader discourse on the applicability of AI in handling complex, context-dependent tasks, and offers valuable insights for future advancements.
Jingjing Wang 0001, Joshua Luo, Grace Yang, Allen Hong, Feng Luo 0001
ICMLA5
2023 Analysis of COVID-19 Offensive Tweets and Their Targets
abstract
During the global COVID-19 pandemic, people utilized social media platforms, especially Twitter, to spread and express opinions about the pandemic. Such discussions also drove the rise in COVID-related offensive speech. In this work, focusing on Twitter, we present a comprehensive analysis of COVID-related offensive tweets and their targets. We collected a COVID-19 dataset with over 747 million tweets for 30 months and fine-tuned a BERT classifier to detect offensive tweets. Our offensive tweets analysis shows that the ebb and flow of COVID-related offensive tweets potentially reflect events in the physical world. We then studied the targets of these offensive tweets. There was a large number of offensive tweets with abusive words, which could negatively affect the targeted groups or individuals. We also conducted a user network analysis, and found that offensive users interact more with other offensive users and that the pandemic had a lasting impact on some offensive users. Our study offers novel insights into the persistence and evolution of COVID-related offensive tweets during the pandemic
Song Liao, Ebuka Okpala, Long Cheng 0005, Nishant Vishwamitra, Hongxin Hu, Feng Luo 0001, Matthew Costello
KDD7
2023 NanoSNP: a progressive and haplotype-aware SNP caller on low-coverage nanopore sequencing data
abstract
MOTIVATION: Oxford Nanopore sequencing has great potential and advantages in population-scale studies. Due to the cost of sequencing, the depth of whole-genome sequencing for per individual sample must be small. However, the existing single nucleotide polymorphism (SNP) callers are aimed at high-coverage Nanopore sequencing reads. Detecting the SNP variants on low-coverage Nanopore sequencing data is still a challenging problem. RESULTS: We developed a novel deep learning-based SNP calling method, NanoSNP, to identify the SNP sites (excluding short indels) based on low-coverage Nanopore sequencing reads. In this method, we design a multi-step, multi-scale and haplotype-aware SNP detection pipeline. First, the pileup model in NanoSNP utilizes the naive pileup feature to predict a subset of SNP sites with a Bi-long short-term memory (LSTM) network. These SNP sites are phased and used to divide the low-coverage Nanopore reads into different haplotypes. Finally, the long-range haplotype feature and short-range pileup feature are extracted from each haplotype. The haplotype model combines two features and predicts the genotype for the candidate site using a Bi-LSTM network. To evaluate the performance of NanoSNP, we compared NanoSNP with Clair, Clair3, Pepper-DeepVariant and NanoCaller on the low-coverage (∼16×) Nanopore sequencing reads. We also performed cross-genome testing on six human genomes HG002-HG007, respectively. Comprehensive experiments demonstrate that NanoSNP outperforms Clair, Pepper-DeepVariant and NanoCaller in identifying SNPs on low-coverage Nanopore sequencing data, including the difficult-to-map regions and major histocompatibility complex regions in the human genome. NanoSNP is comparable to Clair3 when the coverage exceeds 16×. AVAILABILITY AND IMPLEMENTATION: https://github.com/huangnengCSU/NanoSNP.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Minghua Xu 0005, Fan Nie, Chuan-Le Xiao, Feng Luo 0001, Jianxin Wang 0001
Bioinform.6
2022 HoD-Net: High-Order Differentiable Deep Neural Networks and Applications
abstract
We introduce a deep architecture named HoD-Net to enable high-order differentiability for deep learning. HoD-Net is based on and generalizes the complex-step finite difference (CSFD) method. While similar to classic finite difference, CSFD approaches the derivative of a function from a higher-dimension complex domain, leading to highly accurate and robust differentiation computation without numerical stability issues. This method can be coupled with backpropagation and adjoint perturbation methods for an efficient calculation of high-order derivatives. We show how this numerical scheme can be leveraged in challenging deep learning problems, such as high-order network training, deep learning-based physics simulation, and neural differential equations.
Tianjia Shao, Kun Zhou 0001, Chenfanfu Jiang, Feng Luo 0001, Yin Yang 0002
AAAI5
2022 Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model
abstract
Pre-trained multilingual language models play an important role in cross-lingual natural language understanding tasks.However, existing methods did not focus on learning the semantic structure of representation, and thus could not optimize their performance.In this paper, we propose Multi-level Multilingual Knowledge Distillation (MMKD), a novel method for improving multilingual language models.Specifically, we employ a teacher-student framework to adopt rich semantic representation knowledge in English BERT.We propose token-, word-, sentence-, and structure-level alignment objectives to encourage multiple levels of consistency between source-target pairs and correlation similarity between teacher and student models.We conduct experiments on crosslingual evaluation benchmarks including XNLI, PAWS-X, and XQuAD.Experimental results show that MMKD outperforms other baseline models of similar size on XNLI and XQuAD and obtains comparable performance on PAWS-X.Especially, MMKD obtains significant performance gains on low-resource languages.
Long Cheng 0005, Hongxin Hu, Feng Luo 0001
EMNLP6
2022 Clustering by Directly Disentangling Latent Space
abstract
To overcome the high dimensionality problem of data, learning feature representations for clustering has been widely studied. In this work, we propose Disentangling Latent Space Clustering (DLS-Clustering), a new clustering framework that directly learns cluster assignments from disentangled latent spacing without additional clustering methods. We enforce the encoder and the generator of GAN to form an encoder-generator pair in addition to the generator-encoder pair. We train the encoder-generator pair using real data, which can implicitly estimate the real conditional distribution. Meanwhile, this framework enforces the outputs of the encoder to match the inputs of GAN and the prior noise distribution, which disentangles latent space into two parts: one-hot discrete and continuous latent variables. The former can be directly expressed as clusters and the latter represents remaining unspecified factors. Our experiments show that the proposed method achieves the optimal disentanglement performance and outperforms existing generative model-based clustering methods.
Yin Yang 0002, Feng Luo 0001
ICIP3
2022 AAEBERT: Debiasing BERT-based Hate Speech Detection Models via Adversarial Learning
abstract
Hate speech datasets contain bias which machine learning models propagate. When these models classify tweets written in African American English (AAE), they predict AAE tweets as hate/abusive at a higher rate than tweets written in Standard American English (SAE). This paper assesses bias in language models fine-tuned for hate speech detection and the effectiveness of adversarial learning in reducing such bias. We introduce AAEBERT, a pre-trained language model for African American English obtained by re-training BERT-base on AAE tweets. AAEBERT is used to extract the representation of each tweet in the various hate speech datasets and to classify tweets into two classes - AAE dialect and non-AAE dialect. A three-layer feedforward neural network that takes the representation from AAEBERT and a dialect label as input is used as the adversarial network for debiasing. We evaluate bias in language models fine-tuned for hate speech detection. Then assess the effectiveness of adversarial debiasing in these models by comparing results before and after adversarial debiasing is applied. Analysis reveals that the fine-tuned models are biased towards AAE, and adversarial debiasing is effective in reducing bias.
Ebuka Okpala, Long Cheng 0005, Nicodemus Msafiri John Mbwambo, Feng Luo 0001
ICMLA4
2022 BlockPolish: accurate polishing of long-read assembly via block divide-and-conquer
abstract
Long-read sequencing technology enables significant progress in de novo genome assembly. However, the high error rate and the wide error distribution of raw reads result in a large number of errors in the assembly. Polishing is a procedure to fix errors in the draft assembly and improve the reliability of genomic analysis. However, existing methods treat all the regions of the assembly equally while there are fundamental differences between the error distributions of these regions. How to achieve very high accuracy in genome assembly is still a challenging problem. Motivated by the uneven errors in different regions of the assembly, we propose a novel polishing workflow named BlockPolish. In this method, we divide contigs into blocks with low complexity and high complexity according to statistics of aligned nucleotide bases. Multiple sequence alignment is applied to realign raw reads in complex blocks and optimize the alignment result. Due to the different distributions of error rates in trivial and complex blocks, two multitask bidirectional Long short-term memory (LSTM) networks are proposed to predict the consensus sequences. In the whole-genome assemblies of NA12878 assembled by Wtdbg2 and Flye using Nanopore data, BlockPolish has a higher polishing accuracy than other state-of-the-arts including Racon, Medaka and MarginPolish & HELEN. In all assemblies, errors are predominantly indels and BlockPolish has a good performance in correcting them. In addition to the Nanopore assemblies, we further demonstrate that BlockPolish can also reduce the errors in the PacBio assemblies. The source code of BlockPolish is freely available on Github (https://github.com/huangnengCSU/BlockPolish).
Fan Nie, Xin Gao 0001, Feng Luo 0001, Jianxin Wang 0001
Briefings Bioinform.5
2022 Fec: a fast error correction method based on two-rounds overlapping and caching
abstract
The third-generation sequencing technology has advanced genome analysis with long-read length, but the reads need error correction due to the high error rate. Error correction is a time-consuming process especially when the sequencing coverage is high. Generally, for a pair of overlapping reads A and B, the existing error correction methods perform a base-level alignment from B to A when correcting the read A. And another base-level alignment from A to B is performed when correcting the read B. However, based on our observation, the base-level alignment information can be reused. In this article, we present a fast error correction tool Fec, using two-rounds overlapping and caching. Fec can be used independently or as an error correction step in an assembly pipeline. In the first round, Fec uses a large window size (20) to quickly find enough overlaps to correct most of the reads. In the second round, a small window size (5) is used to find more overlaps for the reads with insufficient overlaps in the first round. When performing base-level alignment, Fec searches the cache first. If the alignment exists in the cache, Fec takes this alignment out and deduces the second alignment from it. Otherwise, Fec performs base-level alignment and stores the alignment in the cache. We test Fec on nine datasets, and the results show that Fec has 1.24-38.56 times speed-up compared to MECAT, CANU and MINICNS on five PacBio datasets and 1.16-27.8 times speed-up compared to NECAT and CANU on four nanopore datasets. AVAILABILITY AND IMPLEMENTATION: Fec is available at https://github.com/zhangjuncsu/Fec. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fan Nie, Feng Luo 0001, Jianxin Wang 0001
Bioinform.5
2022 CatCharger: Deploying In-Motion Wireless Chargers in a Metropolitan Road Network via Categorization and Clustering of Vehicle Traffic
abstract
In metropolitan areas with heavy transit demands, electric vehicles (EVs) are expected to be continuously driving without recharging downtime. Wireless power transfer (WPT) provides a promising solution for in-motion EV charging. Nevertheless, previous works are not directly applicable for the deployment of in-motion wireless chargers due to their different charging characteristics. The challenge of deploying in-motion wireless chargers to support the continuous driving of EVs in a metropolitan road network with the minimum cost remains unsolved. We proposeCatChargerto tackle this challenge. By analyzing a metropolitan-scale data set, we found that traffic attributes like vehicle passing speed, daily visit frequency at intersections (i.e., landmarks), and their variances are diverse, and these attributes are critical to in-motion wireless charging performance. Driven by these observations, we first group landmarks with similar attribute values using the entropy minimization clustering method, and select candidate landmarks from the groups with suitable attribute values. Then, we use the kernel density estimator (KDE) to deduce the expected vehicle residual energy at each candidate landmark and consider EV drivers’ routing choice behavior in charger deployment. Finally, we determine the deployment locations by formulating and solving a multiobjective optimization problem, which maximizes vehicle traffic flow at charger deployment positions while guaranteeing the continuous driving of EVs at each landmark. Trace-driven experiments demonstrate thatCatChargerincreases the ratio of driving EVs at the end of a day by 12.5% under the same deployment cost.
Li Yan 0004, Haiying Shen, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001, Feng Luo 0001, Chenxi Qiu, Zhe Zhang 0048, Shohaib Mahmud
IEEE Internet Things J.5
2022 SACall: A Neural Network Basecaller for Oxford Nanopore Sequencing Data Based on Self-Attention Mechanism
abstract
Highly portable Oxford Nanopore sequencer producing long reads in real-time at low cost has made many breakthroughs in genomics studies. However, a major limitation of nanopore sequencing is its high errors when deciphering DNA sequences from noisy and complex raw data. In this paper, we developed an end-to-end basecaller, SACall, based on convolution layers, transformer self-attention layers and a CTC decoder. In SACall, the convolution layers are used to downsample the signals and capture the local patterns. To achieve the contextual relevance of signals, self-attention layers are adopted to calculate the similarity of the signals at any two positions in the raw signal sequence. Finally, the CTC decoder generates the DNA sequence by a beam search algorithm. We use a benchmark consisting of nine isolated genomes to test the quality of different basecallers including SACall, Albacore, and Guppy. The performances of basecallers are evaluated from the perspective of read accuracy, assembly quality, and consensus accuracy. Among most of the genomes in the test benchmark, the reads basecalled by SACall have fewer errors than the reads basecalled by other basecallers. When assembling the basecalled reads of each genome, the assembly from SACall basecalled reads achieves a higher assembly identity. In addition, there are fewer errors in the polished assembly from reads basecalled by SACall compared to those basecalled by Albacore and Guppy. In general, SACall outperforms the Nanopore official basecallers Albacore and Guppy in the benchmark. Moreover, SACall is an open-source and freely available basecaller, which gives a chance for researchers to train their own basecalling models on specific data and basecall Nanopore reads.
Fan Nie, Feng Luo 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 COVID-HateBERT: a Pre-trained Language Model for COVID-19 related Hate Speech Detection
abstract
With the dramatic growth of hate speech on social media during the COVID-19 pandemic, there is an urgent need to detect various hate speech effectively. Existing methods only achieve high performance when the training and testing data come from the same data distribution. The models trained on the traditional hateful dataset cannot fit well on COVID-19 related dataset. Meanwhile, manually annotating the hate speech dataset for supervised learning is time-consuming. Here, we propose COVID-HateBERT, a pre-trained language model to detect hate speech on English Tweets to address this problem. We collect 200M English tweets based on COVID-19 related hateful keywords and hashtags. Then, we use a classifier to extract the 1.27M potential hateful tweets to re-train BERT-base. We evaluate our COVID-HateBERT on four benchmark datasets. The COVID-HateBERT achieves a 14.8%-23.8% higher macro average F1 score on traditional hate speech detection comparing to baseline methods and a 2.6%-6.73% higher macro average F1 score on COVID-19 related hate speech detection comparing to classifiers using BERT and BERTweet, which shows that COVID-HateBERT can generalize well on different datasets.
Song Liao, Ebuka Okpala, Max Tong, Matthew Costello, Long Cheng 0005, Hongxin Hu, Feng Luo 0001
ICMLA8
2021 Towards Understanding and Detecting Cyberbullying in Real-world Images
Nishant Vishwamitra, Hongxin Hu, Feng Luo 0001, Long Cheng 0005
NDSS3
2021 NeuralPolish: a novel Nanopore polishing method based on alignment matrix construction and orthogonal Bi-GRU Networks
abstract
MOTIVATION: Oxford Nanopore sequencing producing long reads at low cost has made many breakthroughs in genomics studies. However, the large number of errors in Nanopore genome assembly affect the accuracy of genome analysis. Polishing is a procedure to correct the errors in genome assembly and can improve the reliability of the downstream analysis. However, the performances of the existing polishing methods are still not satisfactory. RESULTS: We developed a novel polishing method, NeuralPolish, to correct the errors in assemblies based on alignment matrix construction and orthogonal Bi-GRU networks. In this method, we designed an alignment feature matrix for representing read-to-assembly alignment. Each row of the matrix represents a read, and each column represents the aligned bases at each position of the contig. In the network architecture, a bi-directional GRU network is used to extract the sequence information inside each read by processing the alignment matrix row by row. After that, the feature matrix is processed by another bi-directional GRU network column by column to calculate the probability distribution. Finally, a CTC decoder generates a polished sequence with a greedy algorithm. We used five real datasets and three assembly tools including Wtdbg2, Flye and Canu for testing, and compared the results of different polishing methods including NeuralPolish, Racon, MarginPolish, HELEN and Medaka. Comprehensive experiments demonstrate that NeuralPolish achieves more accurate assembly with fewer errors than other polishing methods and can improve the accuracy of assembly obtained by different assemblers. AVAILABILITY AND IMPLEMENTATION: https://github.com/huangnengCSU/NeuralPolish.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fan Nie, Feng Luo 0001, Xin Gao 0001, Jianxin Wang 0001
Bioinform.4
2021 EPGA-SC : A Framework for de novo Assembly of Single-Cell Sequencing Reads
abstract
Assembling genomes from single-cell sequencing data is essential for single-cell studies. However, single-cell assemblies are challenging due to (i) the highly non-uniform read coverage and (ii) the elevated levels of sequencing errors and chimeric reads. Although several assemblers for single-cell data have been proposed in recent years, most of them fail to construct correct long contigs. In this study, we present a new framework called EPGA-SC for de novo assembly of single-cell sequencing reads. The EPGA assembler has designed strategies to solve the problems caused by sequencing errors, sequencing biases, and repetitive regions. However, the extremely unbalanced and richer error types prevent EPGA to achieve high performance in single-cell sequencing data. In this study, we designed EPGA-SC based on EPGA. The main innovations of EPGA-SC are as follows: (i) classifying reads to reduce the proportion of false reads; (ii) using multiple sets of high precision paired-end reads generated from the high precision assemblies produced by other assembler such as SPAdes to overcome the impact of sequencing biases and repetitive regions; and (iii) developing novel algorithms for removing chimeric errors and extending contigs. We test EPGA-SC with seven datasets. The experimental results show that EPGA-SC can generate better assemblies than most current tools in most time in term of MAX contig, N50, NG50, NA50, and NGA50.
Xingyu Liao, Min Li 0007, You Zou, Fang-Xiang Wu, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2021 A Gene Rank Based Approach for Single Cell Similarity Assessment and Clustering
abstract
Single-cell RNA sequencing (scRNA-seq) technology provides quantitative gene expression profiles at single-cell resolution. As a result, researchers have established new ways to explore cell population heterogeneity and genetic variability of cells. One of the current research directions for scRNA-seq data is to identify different cell types accurately through unsupervised clustering methods. However, scRNA-seq data analysis is challenging because of their high noise level, high dimensionality and sparsity. Moreover, the impact of multiple latent factors on gene expression heterogeneity and on the ability to accurately identify cell types remains unclear. How to overcome these challenges to reveal the biological difference between cell types has become the key to analyze scRNA-seq data. For these reasons, the unsupervised learning for cell population discovery based on scRNA-seq data analysis has become an important research area. A cell similarity assessment method plays a significant role in cell clustering. Here, we present BioRank, a new cell similarity assessment method based on annotated gene sets and gene ranks. To evaluate the performances, we cluster cells by two classical clustering algorithms based on the similarity between cells obtained by BioRank. In addition, BioRank can be used by any clustering algorithm that requires a similarity matrix. Applying BioRank to 12 public scRNA-seq datasets, we show that it is better than or at least as well as several popular similarity assessment methods for single cell clustering.
Yunpei Xu, Hong-Dong Li, Yi Pan 0001, Feng Luo 0001, Fang-Xiang Wu, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 DeepPower: Non-intrusive and Deep Learning-based Detection of IoT Malware Using Power Side Channels
abstract
The vulnerability of Internet of Things (IoT) devices to malware attacks poses huge challenges to current Internet security. The IoT malware attacks are usually composed of three stages: intrusion, infection and monetization. Existing approaches for IoT malware detection cannot effectively identify the executed malicious activities at intrusion and infection stages, and thus cannot help stop potential attacks timely. In this paper, we present DeepPower, a non-intrusive approach to infer malicious activities of IoT malware via analyzing power side-channel signals using deep learning. DeepPower first filters raw power signals of IoT devices to obtain suspicious signals, and then performs a fine-grained analysis on these signals to infer corresponding executed activities inside the devices. DeepPower determines whether there exists an ongoing malware infection by conducting a correlation analysis on these identified activities. We implement a prototype of DeepPower leveraging low-cost sensors and devices and evaluate the effectiveness of DeepPower against real-world IoT malware using commodity IoT devices. Our experimental results demonstrate that DeepPower is able to detect infection activities of different IoT malware with a high accuracy without any changes to the monitored devices.
Hongda Li 0002, Feng Luo 0001, Hongxin Hu, Long Cheng 0005, Hai Xiao, Rong Ge 0002
AsiaCCS3
2020 On the Impact of Word Representation in Hate Speech and Offensive Language Detection and Explanation
abstract
Online hate speech and offensive language have been widely recognized as critical social problems. To defend against this problem, several recent works have emerged that focus on the detection and explanation of hate speech and offensive language using machine learning approaches. Although these approaches are quite effective in the detection and explanation of hate speech and offensive language samples, they do not explore the impact of the representation of such samples. In this work, we introduce a novel, pronunciation-based representation of hate speech and offensive language samples to enable its detection with high accuracy. To demonstrate the effectiveness of our pronunciation-based representation, we extend an existing hate-speech and offensive language defense model based on deep Long Short-term Memory (LSTM) neural networks by using our pronunciation-based representation of hate speech and offensive language samples to train this model. Our work finds that the pronunciation-based presentation significantly reduces noise in the datasets and enhances the overall performance of the existing model.
Ruijia (Roger) Hu, Wyatt Dorris, Nishant Vishwamitra, Feng Luo 0001, Matthew Costello
CODASPY4
2020 On Analyzing COVID-19-related Hate Speech Using BERT Attention
abstract
The emergence of COVID-19 has engendered a new wave of online hate speech in social media platforms such as Twitter. Its widespread effects range from acts of cyber-harassment towards certain ethnic communities (e.g., the Asian community), to targeting older people belonging to age groups correlated with higher mortality rates (termed infamously as "Boomer Remover"). Thus, an urgent need arises for a timely mitigation of this new wave of online hate speech. In this work, we aim to discover the hate-related keywords linked to COVID-19 in hateful tweets posted on Twitter so that users posting such keywords can be asked to reconsider posting them. We first collect a new dataset of tweets targeting older people supplementing with a dataset targeting the Asian community. Then, we develop an approach to analyze the datasets with BERT (a transformer-based model) attention mechanism and discover 186 novel keywords targeting the Asian community and 100 keywords targeting older people. Based on our study, we then propose a control mechanism wherein a user can be asked to reconsider using certain sensitive words identified by our approach. We further perform an exploratory analysis of BERT attention mechanism and find that the most high-impact, long distance attentions are learned in the earlier or later layers of the model depending on the underlying data distribution. Our study indicates that the BERT model in some cases uses a hate keyword and an associated group or individual to make predictions, a finding that is inline with existing hate-speech research, which suggests that hate-speech is often aimed at certain groups or individuals.
Nishant Vishwamitra, Ruijia (Roger) Hu, Feng Luo 0001, Long Cheng 0005, Matthew Costello, Yin Yang 0002
ICMLA3
2020 MultiGuideScan: a multi-processing tool for designing CRISPR guide RNA libraries
abstract
SUMMARY: The recent advance in genome engineering technologies based on CRISPR/Cas9 system is enabling people to systematically understand genomic functions. A short RNA string (the CRISPR guide RNA) can guide the Cas9 endonuclease to specific locations in complex genomes to cut DNA double-strands. The CRISPR guide RNA is essential for gene editing systems. Recently, the GuideScan software is developed to design CRISPR guide RNA libraries, which can be used for genome editing of coding and non-coding genomic regions effectively. However, GuideScan is a serial program and computationally expensive for designing CRISPR guide RNA libraries from large genomes. Here, we present an efficient guide RNA library designing tool (MultiGuideScan) by implementing multiple processes of GuideScan. MultiGuideScan speeds up the guide RNA library designing about 9-12 times on a 32-process mode comparing to GuideScan. MultiGuideScan makes it possible to design guide RNA libraries from large genomes. AVAILABILITY AND IMPLEMENTATION: MULTIGUIDESCAN IS AVAILABLE AT GITHUB: https://github.com/bioinfomaticsCSU/MultiGuideScan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Li 0033, Shaokai Wang, Feng Luo 0001, Fang-Xiang Wu, Jianxin Wang 0001
Bioinform.3
2020 MultiMotifMaker: A Multi-Thread Tool for Identifying DNA Methylation Motifs from Pacbio Reads
abstract
The methylation of DNA is an important mechanism to control biological processes. Recently, the Pacbio SMRT technology provides a new way to identify base methylation in the genome. MotifMaker is a tool developed by Pacbio for discovering DNA methylation motifs from methylated DNA sequences. However, MotifMaker is single-threaded and computational expensive for identifying methylation motifs from large genomes. Here, we present an efficient motif finding algorithm (MultiMotifMaker) by implementing multi threads of the MotifMaker. The MultiMotifMaker speeds up the motif search about 8-9 times on a 32 core computer comparing to MotifMaker. MultiMotifMaker makes it possible to identify methylation motifs from Pacbio reads for large genomes.
Tao Li 0033, Xiankai Zhang, Feng Luo 0001, Fang-Xiang Wu, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Improving de novo Assembly Based on Read Classification
abstract
Due to sequencing bias, sequencing error, and repeat problems, the genome assemblies usually contain misarrangements and gaps. When tackling these problems, current assemblers commonly consider the read libraries as a whole and adopt the same strategy to deal with them. However, if we can divide reads into different categories and take different assembly strategies for different read categories, we expect to reduce the mutual effects on problems in genome assembly and facilitate to produce satisfactory assemblies. In this paper, we present a new pipeline for genome assembly based on read classification (ARC). ARC classifies reads into three categories according to the frequencies of k-mers they contain. The three categories refer to (1) low depth reads, which contain a certain low frequency k-mers and are often caused by sequencing errors or bias; (2) high depth reads, which contain a certain high frequency k-mers and usually come from repetitive regions; and (3) normal depth reads, which are the rest of reads. After read classification, an existing assembler is used to assemble different read categories separately, which is beneficial to resolve problems in the genome assembly. ARC adopts loose assembly parameters for low depth reads, and strict assembly parameters for normal depth and high depth reads. We test ARC using five datasets. The experimental results show that, assemblers combining with ARC can generate better assemblies in terms of NA50, NGA50, and genome fraction.
Xingyu Liao, Min Li 0007, You Zou, Fang-Xiang Wu, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2019 An attention-based neural network basecaller for Oxford Nanopore sequencing data
abstract
Highly portable Oxford Nanopore sequencer producing long reads in real time at low cost has made many breakthroughts in genomics studies. However, a major limitation of nanopore sequencing is its high errors when deciphering DNA sequences from noisy and complex raw data. Here we develops SACall, an end-to-end basecaller based on convolution layers, transformer self-attention layers and CTC decoder. From the perspective of read accuracy, SACall yields better performance in the benchmark than ONT official basecaller Guppy and Albacore. SACall is an open-source, freely available basecaller, which gives a chance for researchers to train new basecalling models on specific data and basecall Nanopore reads.
Fan Nie, Feng Luo 0001, Jianxin Wang 0001
BIBM4
2019 DeepSignal: detecting DNA methylation state from Nanopore sequencing reads using deep-learning
abstract
MOTIVATION: The Oxford Nanopore sequencing enables to directly detect methylation states of bases in DNA from reads without extra laboratory techniques. Novel computational methods are required to improve the accuracy and robustness of DNA methylation state prediction using Nanopore reads. RESULTS: In this study, we develop DeepSignal, a deep learning method to detect DNA methylation states from Nanopore sequencing reads. Testing on Nanopore reads of Homo sapiens (H. sapiens), Escherichia coli (E. coli) and pUC19 shows that DeepSignal can achieve higher performance at both read level and genome level on detecting 6 mA and 5mC methylation states comparing to previous hidden Markov model (HMM) based methods. DeepSignal achieves similar performance cross different DNA methylation bases, different DNA methylation motifs and both singleton and mixed DNA CpG. Moreover, DeepSignal requires much lower coverage than those required by HMM and statistics based methods. DeepSignal can achieve 90% above accuracy for detecting 5mC and 6 mA using only 2× coverage of reads. Furthermore, for DNA CpG methylation state prediction, DeepSignal achieves 90% correlation with bisulfite sequencing using just 20× coverage of reads, which is much better than HMM based methods. Especially, DeepSignal can predict methylation states of 5% more DNA CpGs that previously cannot be predicted by bisulfite sequencing. DeepSignal can be a robust and accurate method for detecting methylation states of DNA bases. AVAILABILITY AND IMPLEMENTATION: DeepSignal is publicly available at https://github.com/bioinfomaticsCSU/deepsignal. SUPPLEMENTARY INFORMATION: Supplementary data are available at bioinformatics online.
De-Peng Wang, Chuan-Le Xiao, Feng Luo 0001, Jianxin Wang 0001
Bioinform.8
2019 Stability of Cloud-Based UAV Systems Supporting Big Data Acquisition and Processing
abstract
Unmanned Aerial Vehicle (UAV) technology has been widely applied in both military and civilian applications. Recent researches on UAV systems feature in the dramatic augment of the variety and number of equipped sensors, which results in such an issue that multiple UAVs cannot afford to handle the big data generated by a range of sensors in the air. Considering this practical problem, in this paper, we propose a cloud-based UAV system which incorporates the computing capability of the terrestrial cloud into the UAV systems. Relying on proposed cloud-based UAV system, one critical theoretic issue is how to acquire the big data generated by the sensors while guaranteeing a stable operation state of the system. First, we analyze the cloud-based system's on-demand service ability as well as its impact on UAVs' control procedure. Second, the UAV cloud control system is modeled as a network control system. Moreover, the stable condition of the UAV cloud control system is derived, which reveals the relationship between the acquisition rate of sensor data and the stability of the cloud-based UAV system. Finally, simulations are conducted to verify the effectiveness of our theoretical analysis.
Feng Luo 0001, Chunxiao Jiang, Shui Yu 0001, Jingjing Wang 0001, Yong Ren 0001
IEEE Trans. Cloud Comput.1
2018 BioRank: A Similarity Assessment Method for Single Cell Clustering
Yunpei Xu, Hong-Dong Li, Yi Pan 0001, Feng Luo 0001, Jianxin Wang 0001
BIBM4
2018 CrescendoNet: A New Deep Convolutional Neural Network with Ensemble Behavior
abstract
We introduce a new deep convolutional neural network, CrescendoNet, by stacking simple building blocks without residual connections. Each Crescendo block contains independent convolution paths with increased depths. The numbers of convolution layers and parameters are only increased linearly in Crescendo blocks. In experiments, CrescendoNet with only 15 layers outperforms almost all networks without residual connections on benchmark datasets, CIFAR10, CIFAR100, and SVHN. Given sufficient amount of data as in SVHN dataset, CrescendoNet with 15 layers and 4.1M parameters can match the performance of DenseNet-BC with 250 layers and 15.3M parameters. CrescendoNet provides a new way to construct high performance deep convolutional neural networks with simple network architecture. Moreover, by investigating a various combination of subnetworks in CrescendoNet, we note that the high performance of CrescendoNet may come from its implicit ensemble behavior, which gives CrescendoNet an anytime classification property. Furthermore, the independence between paths in CrescendoNet allows us to introduce a new path-wise training procedure, which can reduce the memory needed for training.
Nishant Vishwamitra, Hongxin Hu, Feng Luo 0001
ICMLA4
2018 Prediction of lncRNA-disease associations based on inductive matrix completion
abstract
Motivation: Accumulating evidences indicate that long non-coding RNAs (lncRNAs) play pivotal roles in various biological processes. Mutations and dysregulations of lncRNAs are implicated in miscellaneous human diseases. Predicting lncRNA-disease associations is beneficial to disease diagnosis as well as treatment. Although many computational methods have been developed, precisely identifying lncRNA-disease associations, especially for novel lncRNAs, remains challenging. Results: In this study, we propose a method (named SIMCLDA) for predicting potential lncRNA-disease associations based on inductive matrix completion. We compute Gaussian interaction profile kernel of lncRNAs from known lncRNA-disease interactions and functional similarity of diseases based on disease-gene and gene-gene onotology associations. Then, we extract primary feature vectors from Gaussian interaction profile kernel of lncRNAs and functional similarity of diseases by principal component analysis, respectively. For a new lncRNA, we calculate the interaction profile according to the interaction profiles of its neighbors. At last, we complete the association matrix based on the inductive matrix completion framework using the primary feature vectors from the constructed feature matrices. Computational results show that SIMCLDA can effectively predict lncRNA-disease associations with higher accuracy compared with previous methods. Furthermore, case studies show that SIMCLDA can effectively predict candidate lncRNAs for renal cancer, gastric cancer and prostate cancer. Availability and implementation: https://github.com//bioinfomaticsCSU/SIMCLDA. Supplementary information: Supplementary data are available at Bioinformatics online.
Chengqian Lu, Mengyun Yang, Feng Luo 0001, Fang-Xiang Wu, Min Li 0007, Yi Pan 0001, Yaohang Li, Jianxin Wang 0001
Bioinform.3
2017 Dynamic Management of In-memory Storage for Efficiently Integrating Compute- and Data-intensive Computing on HPC Systems
abstract
In order to boost the performance of data-intensive computing on HPC systems, in-memory computing frameworks, such as Apache Spark and Flink, use local DRAM for data storage. Optimizing the memory allocation to data storage is critical to delivering performance to traditional HPC compute jobs and throughput to data-intensive applications sharing the HPC resources. Current practices that statically configure in-memory storage may leave inadequate space for compute jobs or miss the opportunity to utilize available space for data-intensive applications. In this paper, we explore techniques to dynamically adjust in-memory storage allocation and provide optimum memory to compute jobs. We have developed a dynamic in-memory storage controller, DynIMS, which monitors memory demands of compute tasks in real time and employs a feedback-based control mechanism to adapt the allocation of in-memory storage. We test DynIMS using HPCC and Spark workloads on a HPC cluster. Experimental results show that DynIMS can achieve up to 5X performance improvement compared to systems with static memory allocations.
Pengfei Xuan, Feng Luo 0001, Rong Ge 0002, Pradip K. Srimani
CCGrid2
2017 CatCharger: Deploying wireless charging lanes in a metropolitan road network through categorization and clustering of vehicle traffic
abstract
The future generation of transportation system will be featured by electrified public transportation. To fulfill metropolitan transit demands, electric vehicles (EVs) must be continuously operable without recharging downtime. Wireless Power Transfer (WPT) techniques for in-motion EV charging is a solution. It however brings up a challenge: how to deploy charging lanes in a metropolitan road network to minimize the deployment cost while enabling EVs' continuous operability. In this paper, we propose CatCharger, which is the first work that handles this challenge. From a metropolitan-scale dataset collected from multiple sources of vehicles, we observe the diversity of vehicle passing speed and daily visit frequency (called traffic attributes) at intersections (i.e., landmarks), which are important factors for charging lane deployment. To select landmarks for deployment, we first group landmarks with similar traffic attribute values using the entropy minimization clustering method, and choose better candidate landmarks from each group suitable for deployment. To determine the deployment locations from the candidate landmarks, we infer the expected vehicle residual energy at each landmark using a Kernel Density Estimator fed by the vehicles' mobility, and formulate and solve an optimization problem to minimize the total deployment cost while ensuring a certain level of expected residual energy of EVs at each landmark. Our trace-driven experiments demonstrate the superior performance of CatCharger over other methods.
Li Yan 0004, Haiying Shen, Juanjuan Zhao 0001, Cheng-Zhong Xu 0001, Feng Luo 0001, Chenxi Qiu
INFOCOM5
2017 Accelerating big data analytics on HPC clusters using two-level storage
Pengfei Xuan, Walter B. Ligon III, Pradip K. Srimani, Rong Ge 0002, Feng Luo 0001
Parallel Comput.5
2016 A de novo genome assembler based on MapReduce and bi-directed de Bruijn graph
abstract
The next generation sequencing (NGS) techniques have enabled biologists to generate large DNA sequences in a high-throughput and low-cost way. Assembly of NGS reads still face great challenges due to the short reads and enormous high volume. In this paper, we presented a new assembler, called GAMR, which is based on bi-directed de Bruijn graph and implemented using MapReduce framework. We designed distributed algorithm for each step in GAMR, making it scalable in assembling large-scale genomes. We evaluated GAMR using GAGE's data and compared it against other NGS assemblers. The results showed GAMR assembled contigs and scaffolds with better accuracy and longer N50 values.
Yuehua Zhang, Pengfei Xuan, Pradip K. Srimani, Feng Luo 0001
BIBM5
2016 Cyberbullying Detection with a Pronunciation Based Convolutional Neural Network
abstract
Cyberbullying can have a deep and long lasting impact on its victims, who are often adolescents. Accurately detecting cyberbullying helps prevent it. However, the noise and errors in social media posts and messages make detecting cyberbullying very challenging. In this paper, we propose a novel pronunciation based convolutional neural network (PCNN) to address this challenge. Upon observing that the pronunciation of misspelled words in informal online conversations is often unchanged, we used the phoneme codes of the text as the features for a convolutional neural network. This procedure corrects spelling errors that did not alter the pronunciation, thereby alleviating the problem of noise and bullying data sparsity. To overcome class imbalance, a common problem in cyberbullying datasets, we implement three techniques that include threshold-moving, cost function adjusting, and a hybrid solution in our model. We evaluate the performance of our models using two cyberbullying datasets collected from Twitter and Formspring.me. The results of our experiment show that PCNN can achieve improved recall and precision compared to baseline convolutional neural networks.
Jonathan Tong, Nishant Vishwamitra, Elizabeth Whittaker, Joseph P. Mazer, Robin M. Kowalski, Hongxin Hu, Feng Luo 0001, Edward Dillon 0002
ICMLA8
2016 CatCharge: Deploying wireless charging lane in metropolitan scale through categorization and clustering of vehicle mobility
abstract
The future generation transportation system will be featured by electrified public transportation. To fulfill metropolitan transit demands, electric vehicles (EVs) must be continuously operable without recharging downtime. Wireless Power Transfer (WPT) techniques for in-motion EV charging is a solution [1], [2]. It however brings up a challenge: how to deploy charging lanes in a metropolitan road network to minimize the deployment cost while enabling EVs' continuous operability.
Li Yan 0004, Juanjuan Zhao 0001, Haiying Shen, Cheng-Zhong Xu 0001, Feng Luo 0001
ICNP5
2015 Predicting protein phosphorylation from gene expression: top methods from the IMPROVER Species Translation Challenge
abstract
MOTIVATION: Using gene expression to infer changes in protein phosphorylation levels induced in cells by various stimuli is an outstanding problem. The intra-species protein phosphorylation challenge organized by the IMPROVER consortium provided the framework to identify the best approaches to address this issue. RESULTS: Rat lung epithelial cells were treated with 52 stimuli, and gene expression and phosphorylation levels were measured. Competing teams used gene expression data from 26 stimuli to develop protein phosphorylation prediction models and were ranked based on prediction performance for the remaining 26 stimuli. Three teams were tied in first place in this challenge achieving a balanced accuracy of about 70%, indicating that gene expression is only moderately predictive of protein phosphorylation. In spite of the similar performance, the approaches used by these three teams, described in detail in this article, were different, with the average number of predictor genes per phosphoprotein used by the teams ranging from 3 to 124. However, a significant overlap of gene signatures between teams was observed for the majority of the proteins considered, while Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways were enriched in the union of the predictor genes of the three teams for multiple proteins. AVAILABILITY AND IMPLEMENTATION: Gene expression and protein phosphorylation data are available from ArrayExpress (E-MTAB-2091). Software implementation of the approach of Teams 49 and 75 are available at http://bioinformaticsprb.med.wayne.edu and http://people.cs.clemson.edu/∼luofeng/sbv.rar, respectively. CONTACT: [email protected] or [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Adel Dayarian, Roberto Romero, Michael Biehl, Erhan Bilal, Sahand Hormoz, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Gyan Bhanot, Feng Luo 0001, Adi L. Tarca
Bioinform.11
2015 Guest Editorial for Special Section on BIBM 2013
abstract
The articles in this special section include a selection of seven papers presented at the IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2013) held in Shanghai, China, on December 18-21, 2013.
Feng Luo 0001, Xintao Wu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 Combining Hadoop and GPU to preprocess large Affymetrix microarray data
abstract
High density oligonucleotide array (microarray) from Affymetrix has been widely used for the measurements of gene expressions. Currently, public data repositories, such as Gene Expression Omnibus (GEO) of the National Center for Biotechnology Information (NCBI), have accumulated large amounts of microarray data. Efficient integrative analysis of those microarray data will provide significant knowledge about biological systems. None of the existing microarray preprocessing and quality assessment tools can handle very large microarray datasets with tens of thousands of experiments. The preprocessing and quality assessment of microarray datasets contain both data-intensive and compute-intensive tasks. In this paper, we develop a new set of tools using a mix of the Hadoop (for data intensive tasks) and the General-Purpose Graphics Processing Units (GPGPUs) (for compute intensive tasks) to efficiently process large microarray data. Evaluation of our new tools on large microarray datasets with ten thousands of experiments showed promising superior performance. We demonstrate that the combination of Hadoop and GPGPU computation is effective for complex scientific applications that contain both data-intensive and compute-intensive tasks. Our new tool set will make it possible to utilize valuable large microarray data in the public repositories.
Sufeng Niu, Nilim Sarma, Pengfei Xuan, Melissa C. Smith, Pradip K. Srimani, Feng Luo 0001
IEEE BigData7
2014 Node Energy Consumption Analysis in Wireless Sensor Networks
abstract
The limited sensor node energy and the large number of nodes with dynamic network topology information have always been the important design concerns in Wireless Sensor Networks (WSN). Node clustering is an effective way to tackle with the two issues by grouping the nodes into hierarchies in order to reduce communication distance and the amount of message. This paper mainly focuses on the unification of the node energy consumption in WSN. The distributions of the energy consumption for various scenarios in the hierarchical network are analyzed for the first time and two main reasons are found leading to the asymmetry of the energy consumption among nodes. One is the energy consumption from the communications between nodes and base station, and the other is that from the cluster head for receiving data from other nodes. It is concluded that the probability of the node acting as cluster head should depend on the distribution of the head's energy consumption, and a variable sampling space oriented to the potential number of cluster heads is established thereafter. Furthermore, a new clustering algorithm, the Segment Equalization Clustering based on Cluster Head Energy Consumption (SECHEC) algorithm is proposed, which can effectively improve the network lifetime and ensure the availability of the system within its entire lifespan.
Feng Luo 0001, Chunxiao Jiang, Haijun Zhang 0001, Xuexia Wang, Yong Ren 0001
VTC Fall1
2013 Sesame: A new bioinformatics semantic workflow design system
abstract
Biologists have become increasingly dependent on bioinformatics tools to analyze and interpret their datasets. The number, variety and complexity of these bioinformatics tools have increased dramatically and they have become more and more computationally complex, expensive and resource intensive. Powerful workflow design systems have been developed to automate the execution of set of tools for a specific task. However, designing a complex executable workflow using such tools still requires considerable computational expertise or the help from a bioinformatics expert. In this paper, we developed Sesame, a bioinformatics semantic workflow design system. We have designed a new ontology for bioinformatics tools and services (OBTS) and proposed an ontology driven semantic workflow design mechanism, using this new OBTS. Compared to an executable workflow, the semantic workflow is at the level of biological concepts that are closer to the scientific research. Biologists will greatly benefit from the decoupling of semantic workflow design from executable workflow design with computational implementation details. Currently, a prototype version of Sesame system has been implemented and deployed. Sesame will allow biologists to efficiently perform complex data analysis to address scientific questions.
Lu Zhang 0021, Pengfei Xuan, Alexander Duvall, Jonathan Lowe, Arvind Subramanian, Pradip K. Srimani, Feng Luo 0001, Yongping Duan
BIBM9
2012 O-linked Glycosylation Site Prediction Using Ensemble of Graphical Models
abstract
Prediction of O-linked glycosylation sites in proteins is a challenging problem. In this paper, we introduced a new method to predict glycosylation sites in proteins. First, we built a Markov random field (MRF) to represent the sequence position relationship and model the underlying distribution of glycosylation sites. We then considered glycosylation site prediction as a class imbalance problem and employed the AdaBoost algorithm to improve the predictive performance of the classifier. We applied our method to two types of proteins: the transmembrane (TM) proteins and the non-transmembrane (non-TM) proteins. We showed that for both datasets, our methods outperform existing methods. We also showed that the performance of the system was improved significantly with the help of AdaBoost.
Aditya Sriram, Feng Luo 0001
ICMLA (2)2
2012 Molecular ecological network analyses
abstract
BACKGROUND: Understanding the interaction among different species within a community and their responses to environmental changes is a central goal in ecology. However, defining the network structure in a microbial community is very challenging due to their extremely high diversity and as-yet uncultivated status. Although recent advance of metagenomic technologies, such as high throughout sequencing and functional gene arrays, provide revolutionary tools for analyzing microbial community structure, it is still difficult to examine network interactions in a microbial community based on high-throughput metagenomics data. RESULTS: Here, we describe a novel mathematical and bioinformatics framework to construct ecological association networks named molecular ecological networks (MENs) through Random Matrix Theory (RMT)-based methods. Compared to other network construction methods, this approach is remarkable in that the network is automatically defined and robust to noise, thus providing excellent solutions to several common issues associated with high-throughput metagenomics data. We applied it to determine the network structure of microbial communities subjected to long-term experimental warming based on pyrosequencing data of 16 S rRNA genes. We showed that the constructed MENs under both warming and unwarming conditions exhibited topological features of scale free, small world and modularity, which were consistent with previously described molecular ecological networks. Eigengene analysis indicated that the eigengenes represented the module profiles relatively well. In consistency with many other studies, several major environmental traits including temperature and soil pH were found to be important in determining network interactions in the microbial communities examined. To facilitate its application by the scientific community, all these methods and statistical tools have been integrated into a comprehensive Molecular Ecological Network Analysis Pipeline (MENAP), which is open-accessible now (http://ieg2.ou.edu/MENA). CONCLUSIONS: The RMT-based molecular ecological network analysis provides powerful tools to elucidate network interactions in microbial communities and their responses to environmental changes, which are fundamentally important for research in microbial ecology and environmental microbiology.
Ye Deng 0001, Yi-Huei Jiang, Yunfeng Yang, Feng Luo 0001, Jizhong Zhou
BMC Bioinform.5
2011 Predicing Yeast Synthetic Lethal Genetic Interactions Using Short Polypeptide Clusters
abstract
Synthetic lethal genetic interactions (SLGI) among proteins have been widely used to define functional relationships between proteins and pathways. However, the molecular mechanism of synthetic lethal genetic interactions is still unclear. In this study we used the clusters of short polypeptide sequences, which are typically shorter than the classically defined protein domains, to characterize the functionalities of proteins. We developed a framework to identify significant short polypeptide clusters from yeast protein sequences. We then used these short polypeptide clusters as features to predict SLGIs. Both cross-validation and evaluation on experimental data sets showed that the short polypeptide clusters based approach is superior to the previous protein domain based approach. The short polypeptide clusters based approach provides significantly higher coverage for predicting SLGIs. Moreover, the short polypeptide clusters based approach produced less false positive predictions.
Bo Li 0036, Yuehua Zhang, Pradip K. Srimani, Feng Luo 0001
BIBM4
2010 A non-parameter Ising model for network-based identification of differentially expressed genes in recurrent breast cancer patients
abstract
Identification of genes and pathways involving in diseases and physiological conditions is a major task in systems biology. In this study, we develop a new non-parameter Ising model to integrate protein-protein interaction network and microarray data for identifying differentially expressed (DE) genes. We also propose a simulated annealing algorithm to find the optimal configuration of the Ising model. We test the Ising model to two breast cancer microarray data sets. The results show that more cancer related differentially expressed subnetworks and genes are identified by the Ising model than by the Markov random filed (MRF) model.
Xumeng Li, F. Alex Feltus, Xiaoqian Sun, James Zijun Wang, Feng Luo 0001
BIBM5
2009 Predicting Yeast Synthetic Lethal Genetic Interactions Using Protein Domains
abstract
Synthetic lethal genetic interactions are of interest as they can be used to predict function of unknown proteins and find drug target or drug combinations. In this study, we applied support vector machine (SVM) classifier to predict synthetic lethal genetic interactions in Saccharomyces cerevisiae based on domain information in proteins. We found that our method can predict synthetic lethal genetic interactions with high sensitivity (88.35%) and specificity (82.00%). To the best of our knowledge, the work reported in this paper is the first domain-based model for the prediction of genetic interactions. Our study indicates that there is strong correlation between protein domain relationship and synthetic lethal genetic interactions.
Bo Li 0036, Feng Luo 0001
BIBM2
2009 Core and periphery structures in protein interaction networks
abstract
BACKGROUND: Characterizing the structural properties of protein interaction networks will help illuminate the organizational and functional relationships among elements in biological systems. RESULTS: In this paper, we present a systematic exploration of the core/periphery structures in protein interaction networks (PINs). First, the concepts of cores and peripheries in PINs are defined. Then, computational methods are proposed to identify two types of cores, k-plex cores and star cores, from PINs. Application of these methods to a yeast protein interaction network has identified 110 k-plex cores and 109 star cores. We find that the k-plex cores consist of either "party" proteins, "date" proteins, or both. We also reveal that there are two classes of 1-peripheral proteins: "party" peripheries, which are more likely to be part of protein complex, and "connector" peripheries, which are more likely connected to different proteins or protein complexes. Our results also show that, besides connectivity, other variations in structural properties are related to the variation in biological properties. Furthermore, the negative correlation between evolutionary rate and connectivity are shown toysis. Moreover, the core/periphery structures help to reveal the existence of multiple levels of protein expression dynamics. CONCLUSION: Our results show that both the structure and connectivity can be used to characterize topological properties in protein interaction networks, illuminating the functional organization of cellular systems.
Feng Luo 0001, Bo Li 0036, Xiu-Feng Wan, Richard H. Scheuermann
BMC Bioinform.1
2008 Exploring Core/Periphery Structures in Protein Interaction Networks Provides Structure-Property Relation Insights
abstract
In this study, we present a systematic exploration of the core/periphery structures in protein interaction networks (PINs). First, the concepts of cores and peripheries in PINs are defined. Then, computational methods are proposed to identify cores from PINs. Application of these methods to a combined yeast PIN has identified 110 k-plex cores and 138 star cores. Based on more precise structural characteristics, our studies reveal new prospects of principles and roles of proteins. Our results show that, aside from connectivity, the structural variations between different types of proteins are also related to the variation in biological properties. Two classes of 1-peripheral proteins have been identified: party peripheries, which are more likely to be part of protein complex, and connector peripheries, which are more likely connected to different complexes or individual proteins. This study may facilitate the understanding of the topological characteristics of proteins in interaction networks and thus help elucidate the organization of cellular systems.
Thomas Grindinger, Feng Luo 0001, Xiu-Feng Wan, Richard H. Scheuermann
BIBM2
2008 Exploring local community structures in large networks
abstract
In this paper, three new algorithms, a greedy algorithm, a KL-like algorithm, and an add-all algorithm, are proposed to find local optimal community structures in large networks starting from a given source vertex. The time complexity for finding a l
Feng Luo 0001, James Zijun Wang, Eric Promislow
Web Intell. Agent Syst.1
2007 Modular organization of protein interaction networks
abstract
MOTIVATION: Accumulating evidence suggests that biological systems are composed of interacting, separable, functional modules. Identifying these modules is essential to understand the organization of biological systems. RESULT: In this paper, we present a framework to identify modules within biological networks. In this approach, the concept of degree is extended from the single vertex to the sub-graph, and a formal definition of module in a network is used. A new agglomerative algorithm was developed to identify modules from the network by combining the new module definition with the relative edge order generated by the Girvan-Newman (G-N) algorithm. A JAVA program, MoNet, was developed to implement the algorithm. Applying MoNet to the yeast core protein interaction network from the database of interacting proteins (DIP) identified 86 simple modules with sizes larger than three proteins. The modules obtained are significantly enriched in proteins with related biological process Gene Ontology terms. A comparison between the MoNet modules and modules defined by Radicchi et al. (2004) indicates that MoNet modules show stronger co-clustering of related genes and are more robust to ties in betweenness values. Further, the MoNet output retains the adjacent relationships between modules and allows the construction of an interaction web of modules providing insight regarding the relationships between different functional modules. Thus, MoNet provides an objective approach to understand the organization and interactions of biological processes in cellular systems. AVAILABILITY: MoNet is available upon request from the authors.
Feng Luo 0001, Yunfeng Yang, Chin-Fu Chen, Roger L. Chang, Jizhong Zhou, Richard H. Scheuermann
Bioinform.1
2007 Modular organization of protein interaction networks
Feng Luo 0001, Yunfeng Yang, Chin-Fu Chen, Roger L. Chang, Jizhong Zhou, Richard H. Scheuermann
Bioinform.1
2007 A quantitative genotype algorithm reflecting H5N1 Avian influenza niches
abstract
MOTIVATION: Computational genotyping analyses are critical for characterizing molecular evolutionary footprints, thus providing important information for designing the strategies of influenza prevention and control. Most of the current methods that are available are based on multiple sequence alignment and phylogenetic tree construction, which are time consuming and limited by the number of taxa. Arbitrarily defining genotypes further complicates the interpretation of genotyping results. METHODS: In this study, we describe a quantitative influenza genotyping algorithm based on the theory of quasispecies. First, the complete composition vector (CCV) was utilized to calculate the pairwise evolutionary distance between genotypes. Next, Hierarchical Bayesian Modeling using the Gibbs Sampling algorithm was applied to identify the segment genotype threshold, which is used to identify influenza segment genotype through a modularity calculation. The viral genotype was defined by combining eight segment genotypes based on the genetic reassortment feature of influenza A viruses. RESULTS: We applied this method for H5N1 avian influenza viruses and identified 107 niches among 283 viruses with a complete genome set. The diversity of viral genotypes, and their correlation with geographic locations suggests that these viruses form local niches after being introduced to a new ecological environment through poultry trade or bird migration. This novel method allows us to define genotypes in a robust, quantitative as well as hierarchical manner. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiu-Feng Wan, Feng Luo 0001, Michael Emch, Ruben O. Donis
Bioinform.3
2007 Constructing gene co-expression networks and predicting functions of unknown genes by random matrix theory
abstract
BACKGROUND: Large-scale sequencing of entire genomes has ushered in a new age in biology. One of the next grand challenges is to dissect the cellular networks consisting of many individual functional modules. Defining co-expression networks without ambiguity based on genome-wide microarray data is difficult and current methods are not robust and consistent with different data sets. This is particularly problematic for little understood organisms since not much existing biological knowledge can be exploited for determining the threshold to differentiate true correlation from random noise. Random matrix theory (RMT), which has been widely and successfully used in physics, is a powerful approach to distinguish system-specific, non-random properties embedded in complex systems from random noise. Here, we have hypothesized that the universal predictions of RMT are also applicable to biological systems and the correlation threshold can be determined by characterizing the correlation matrix of microarray profiles using random matrix theory. RESULTS: Application of random matrix theory to microarray data of S. oneidensis, E. coli, yeast, A. thaliana, Drosophila, mouse and human indicates that there is a sharp transition of nearest neighbour spacing distribution (NNSD) of correlation matrix after gradually removing certain elements insider the matrix. Testing on an in silico modular model has demonstrated that this transition can be used to determine the correlation threshold for revealing modular co-expression networks. The co-expression network derived from yeast cell cycling microarray data is supported by gene annotation. The topological properties of the resulting co-expression network agree well with the general properties of biological networks. Computational evaluations have showed that RMT approach is sensitive and robust. Furthermore, evaluation on sampled expression data of an in silico modular gene system has showed that under-sampled expressions do not affect the recovery of gene co-expression network. Moreover, the cellular roles of 215 functionally unknown genes from yeast, E. coli and S. oneidensis are predicted by the gene co-expression networks using guilt-by-association principle, many of which are supported by existing information or our experimental verification, further demonstrating the reliability of this approach for gene function prediction. CONCLUSION: Our rigorous analysis of gene expression microarray profiles using RMT has showed that the transition of NNSD of correlation matrix of microarray profile provides a profound theoretical criterion to determine the correlation threshold for identifying gene co-expression networks.
Feng Luo 0001, Yunfeng Yang, Jianxin Zhong, Haichun Gao, Latifur Khan, Dorothea K. Thompson, Jizhong Zhou
BMC Bioinform.1
2006 Exploring Local Community Structures in Large Networks
abstract
In this paper, we extend the concept of degree from single vertex to sub-graph, and present a formal definition of module/community in a network based on this extension. A new locally optimized algorithm is designed to find the module for a given source vertex in a network. Our analysis shows that the complexity of this algorithm is O(K2d) where K is the number of vertices to be explored in the sub-graph and d is the average degree of the vertices in the sub-graph. Based on this algorithm, we implement a JAVA tool, MoNet, for exploring local community structures in large networks. Using this tool to analyze a co-purchase network from Amazon shows that there are local community structures in this network. Further analyses on these local community structures demonstrate that media items are much easier to form compact local modules than book items do, indicating that recommending digital media items to customers based on co-purchasing information in the online store is more efficient than recommending books
Feng Luo 0001, James Zijun Wang, Eric Promislow
Web Intelligence1
2004 A dynamically growing self-organizing tree (DGSOT) for hierarchical clustering gene expression profiles
abstract
MOTIVATION: The increasing use of microarray technologies is generating large amounts of data that must be processed in order to extract useful and rational fundamental patterns of gene expression. Hierarchical clustering technology is one method used to analyze gene expression data, but traditional hierarchical clustering algorithms suffer from several drawbacks (e.g. fixed topology structure; mis-clustered data which cannot be reevaluated). In this paper, we introduce a new hierarchical clustering algorithm that overcomes some of these drawbacks. RESULT: We propose a new tree-structure self-organizing neural network, called dynamically growing self-organizing tree (DGSOT) algorithm for hierarchical clustering. The DGSOT constructs a hierarchy from top to bottom by division. At each hierarchical level, the DGSOT optimizes the number of clusters, from which the proper hierarchical structure of the underlying dataset can be found. In addition, we propose a new cluster validation criterion based on the geometric property of the Voronoi partition of the dataset in order to find the proper number of clusters at each hierarchical level. This criterion uses the Minimum Spanning Tree (MST) concept of graph theory and is computationally inexpensive for large datasets. A K-level up distribution (KLD) mechanism, which increases the scope of data distribution in the hierarchy construction, was used to improve the clustering accuracy. The KLD mechanism allows the data misclustered in the early stages to be reevaluated at a later stage and increases the accuracy of the final clustering result. The clustering result of the DGSOT is easily displayed as a dendrogram for visualization. Based on a yeast cell cycle microarray expression dataset, we found that our algorithm extracts gene expression patterns at different levels. Furthermore, the biological functionality enrichment in the clusters is considerably high and the hierarchical structure of the clusters is more reasonable. AVAILABILITY: DGSOT is available upon request from the authors.
Feng Luo 0001, Latifur Khan, Farokh B. Bastani, I-Ling Yen, Jizhong Zhou
Bioinform.1
2003 Hierarchical Clustering of Gene Expression Data
abstract
Rapid development of biological technologies generates a huge amount of data, which provides a processing and global view of the gene expression levels across different conditions and over multiple stages. Analyzation and interpretation of these massive data is a challenging task. One of the most important steps is to extract useful and rational fundamental patterns of gene expression inherent in these huge data. Clustering technology is one of the useful and popular methods to obtain these patterns. In this paper we propose a new hierarchical clustering algorithm to obtain gene expression patterns. This algorithm constructs a hierarchy from top to bottom based on a self-organizing tree. It dynamically finds the number of clusters at each level. We compare our algorithm with the traditional hierarchical agglomerative clustering (HAC) algorithm. We apply our algorithm to an existing 112 rat central nervous system gene expression data. We observe that our algorithm extracts patterns with different levels of abstraction. Furthermore, our approach is useful on recognizing features in complex gene expression data.
Feng Luo 0001, Latifur Khan
BIBE1
2002 Ontology Construction for Information Selection
abstract
Technology in the field of digital media generates huge amounts of non-textual information, audio, video, and images, along with more familiar textual information. The potential for exchange and retrieval of information is vast and daunting. The key problem in achieving efficient and user-friendly retrieval is the development of a search mechanism to guarantee delivery of minimal irrelevant information (high precision) while ensuring relevant information is not overlooked (high recall). The traditional solution employs keyword-based search. The only documents retrieved are those containing user specified keywords. But many documents convey desired semantic information without containing these keywords. One can overcome this problem by indexing documents according to meanings rather than words, although this will entail a way of converting words to meanings and the creation of ontology. We have solved the problem of an index structure through the design and implementation of a concept-based model using domain-dependent ontology. Ontology is a collection of concepts and their interrelationships, which provide an abstract view of an application domain. We propose a new mechanism that can generate ontology automatically in order to make our approach scalable. For this we modify the existing self-organizing tree algorithm (SOTA) that constructs a hierarchy from top to bottom. Furthermore, in order to find an appropriate concept for each node in the hierarchy we propose an automatic concept selection algorithm from WordNet called linguistic ontology. To illustrate the effectiveness of our automatic ontology construction method, we have explored our ontology construction in text documents. The Reuters21578 text document corpus has been used. We have observed that our modified SOTA outperforms hierarchical agglomerative clustering (HAC).
Latifur Khan, Feng Luo 0001
ICTAI2