Aritra Bose

dblp:250/0857 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 InDeepFake: A novel multimodal multilingual indian deepfake video dataset
Arnab Kumar Das, Aritra Bose, Priya Manohar, Anurag Dutta, Ruchira Naskar, Rajat Subhra Chakraborty
Pattern Recognit. Lett.2
2024 MaSk-LMM: A Matrix Sketching Framework for Linear Mixed Models in Association Studies
Myson C. Burch, Aritra Bose, Gregory Dexter, Laxmi Parida, Petros Drineas
RECOMB2
2024 Epidemiological topology data analysis links severe COVID-19 to RAAS and hyperlipidemia associated metabolic syndrome conditions
abstract
MOTIVATION: The emergence of COVID-19 (C19) created incredible worldwide challenges but offers unique opportunities to understand the physiology of its risk factors and their interactions with complex disease conditions, such as metabolic syndrome. To address the challenges of discovering clinically relevant interactions, we employed a unique approach for epidemiological analysis powered by redescription-based topological data analysis (RTDA). RESULTS: Here, RTDA was applied to Explorys data to discover associations among severe C19 and metabolic syndrome. This approach was able to further explore the probative value of drug prescriptions to capture the involvement of RAAS and hypertension with C19, as well as modification of risk factor impact by hyperlipidemia (HL) on severe C19. RTDA found higher-order relationships between RAAS pathway and severe C19 along with demographic variables of age, gender, and comorbidities such as obesity, statin prescriptions, HL, chronic kidney failure, and disproportionately affecting Black individuals. RTDA combined with CuNA (cumulant-based network analysis) yielded a higher-order interaction network derived from cumulants that furthered supported the central role that RAAS plays. TDA techniques can provide a novel outlook beyond typical logistic regressions in epidemiology. From an observational cohort of electronic medical records, it can find out how RAAS drugs interact with comorbidities, such as hypertension and HL, of patients with severe bouts of C19. Where single variable association tests with outcome can struggle, TDA's higher-order interaction network between different variables enables the discovery of the comorbidities of a disease such as C19 work in concert. AVAILABILITY AND IMPLEMENTATION: Code for performing TDA/RTDA is available in https://github.com/IBM/Matilda and code for CuNA can be found in https://github.com/BiomedSciAI/Geno4SD/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel E. Platt, Aritra Bose, Kahn Rhrissorrakrai, Chaya Levovitz, Laxmi Parida
Bioinform.2
2023 Structure-informed clustering for population stratification in association studies
abstract
BACKGROUND: Identifying variants associated with complex traits is a challenging task in genetic association studies due to linkage disequilibrium (LD) between genetic variants and population stratification, unrelated to the disease risk. Existing methods of population structure correction use principal component analysis or linear mixed models with a random effect when modeling associations between a trait of interest and genetic markers. However, due to stringent significance thresholds and latent interactions between the markers, these methods often fail to detect genuinely associated variants. RESULTS: To overcome this, we propose CluStrat, which corrects for complex arbitrarily structured populations while leveraging the linkage disequilibrium induced distances between genetic markers. It performs an agglomerative hierarchical clustering using the Mahalanobis distance covariance matrix of the markers. In simulation studies, we show that our method outperforms existing methods in detecting true causal variants. Applying CluStrat on WTCCC2 and UK Biobank cohorts, we found biologically relevant associations in Schizophrenia and Myocardial Infarction. CluStrat was also able to correct for population structure in polygenic adaptation of height in Europeans. CONCLUSIONS: CluStrat highlights the advantages of biologically relevant distance metrics, such as the Mahalanobis distance, which captures the cryptic interactions within populations in the presence of LD better than the Euclidean distance.
Aritra Bose, Myson C. Burch, Agniva Chowdhury, Peristera Paschou, Petros Drineas
BMC Bioinform.1
2022 Epidemiological topology data analysis links severe COVID-19 to RAAS and hyperlipidemia associated metabolic syndrome conditions
Daniel E. Platt, Aritra Bose, Chaya Levovitz, Kahn Rhrissorrakrai, Laxmi Parida
AMIA2
2022 A Fast, Provably Accurate Approximation Algorithm for Sparse Principal Component Analysis Reveals Human Genetic Variation Across the World
Agniva Chowdhury, Aritra Bose, Samson Zhou, David P. Woodruff, Petros Drineas
RECOMB2
2021 Impact of Clinical and Genomic Factors on COVID-19 Disease Severity
Sanjoy Dey, Aritra Bose, Subrata Saha, Prithwish Chakraborty, Mohamed F. Ghalwash, Filippo Utro, Aldo Guzmán-Sáenz, Kenney Ng, Jianying Hu, Laxmi Parida, Daby M. Sow
AMIA2
2020 CluStrat: A Structure Informed Clustering Strategy for Population Stratification
Aritra Bose, Myson C. Burch, Agniva Chowdhury, Peristera Paschou, Petros Drineas
RECOMB1
2019 TeraPCA: a fast and scalable software package to study genetic variation in tera-scale genotypes
abstract
MOTIVATION: Principal Component Analysis is a key tool in the study of population structure in human genetics. As modern datasets become increasingly larger in size, traditional approaches based on loading the entire dataset in the system memory (Random Access Memory) become impractical and out-of-core implementations are the only viable alternative. RESULTS: We present TeraPCA, a C++ implementation of the Randomized Subspace Iteration method to perform Principal Component Analysis of large-scale datasets. TeraPCA can be applied both in-core and out-of-core and is able to successfully operate even on commodity hardware with a system memory of just a few gigabytes. Moreover, TeraPCA has minimal dependencies on external libraries and only requires a working installation of the BLAS and LAPACK libraries. When applied to a dataset containing a million individuals genotyped on a million markers, TeraPCA requires <5 h (in multi-threaded mode) to accurately compute the 10 leading principal components. An extensive experimental analysis shows that TeraPCA is both fast and accurate and is competitive with current state-of-the-art software for the same task. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are both available at https://github.com/aritra90/TeraPCA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aritra Bose, Vassilis Kalantzis, Eugenia-Maria Kontopoulou, Mai Elkady, Peristera Paschou, Petros Drineas
Bioinform.1