Noorul Amin

dblp:307/3408 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0001-7747-7956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis › pattern matching
approximate string matching
0.812024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.812024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis
0.212024
Flexiplex: a versatile demultiplexer and search tool for omics data · Bioinform. 2024

Methods — techniques the papers use, named apart from their topics

levenshtein distance · 0.8
YearPublicationVenuePosition
2024 Flexiplex: a versatile demultiplexer and search tool for omics data
abstract
MOTIVATION: The process of analyzing high throughput sequencing data often requires the identification and extraction of specific target sequences. This could include tasks, such as identifying cellular barcodes and UMIs in single-cell data, and specific genetic variants for genotyping. However, existing tools, which perform these functions are often task-specific, such as only demultiplexing barcodes for a dedicated type of experiment, or are not tolerant to noise in the sequencing data. RESULTS: To overcome these limitations, we developed Flexiplex, a versatile and fast sequence searching and demultiplexing tool for omics data, which is based on the Levenshtein distance and thus allows imperfect matches. We demonstrate Flexiplex's application on three use cases, identifying cell-line-specific sequences in Illumina short-read single-cell data, and discovering and demultiplexing cellular barcodes from noisy long-read single-cell RNA-seq data. We show that Flexiplex achieves an excellent balance of accuracy and computational efficiency compared to leading task-specific tools. AVAILABILITY AND IMPLEMENTATION: Flexiplex is available at https://davidsongroup.github.io/flexiplex/.
Oliver Cheng, Min Hao Ling, Shuyi Wu, Matthew E. Ritchie, Jonathan Göke, Noorul Amin, Nadia M. Davidson
Bioinform.7
2021 FexRNA: Exploratory Data Analysis and Feature Selection of Non-Coding RNA
abstract
Non-coding RNA (ncRNA) is involved in many biological processes and diseases in all species. Many ncRNA datasets exist that provide ncRNA data in FASTA format which is well suited for biomedical purposes. However, for ncRNA analysis and classification, statistical learning methods require hidden numerical features from the data. Furthermore, in the literature, a wealth of sequence intrinsic features has been proposed for ncRNA identification. The extraction of hidden features, their analysis, and usage of a suitable set of features is crucial for the performance of any statistical learning method. To alleviate the posed challenges, we generated 96 feature datasets from ncRNA widely used features. The feature datasets are based on RNACentral and consist of species, ncRNA types, and expert databases that are available on the FexRNA platform. Additionally, the feature datasets are explored and analysed to provide statistical information, univariate, and bivariate analysis. We sought to determine which of these 17 features would be most appropriate to use in developing ncRNA classification approaches. For feature selection (FS), a two-phase hierarchical FS framework based on correlation and majority voting is proposed and evaluated on 5 species. The FexRNA platform provides information about ncRNA feature analysis and selection.
Noorul Amin, Annette McGrath, Yi-Ping Phoebe Chen
IEEE ACM Trans. Comput. Biol. Bioinform.1