Mohammadsaleh Refahi

dblp:428/5510 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Generative modeling · 50%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › gene regulation › regulatory element discovery
regulatory element prediction
0.312025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

transition-matrix loss · 1.7transformer · 1.7n-gram statistics · 1.7
YearPublicationVenuePosition
2025 Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis
abstract
Transformers have revolutionized nucleotide sequence analysis, yet capturing long‑range dependencies remains challenging. Recent studies show that autoregressive transformers often exhibit Markovian behavior by relying on fixed-length context windows for next-token prediction. However, standard self-attention mechanisms are computationally inefficient for long sequences due to their quadratic complexity and do not explicitly enforce global transition consistency. We introduce CARMANIA (Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis), a self-supervised pretraining framework that augments next-token (NT) prediction with a transition-matrix (TM) loss. The TM loss aligns predicted token transitions with empirically derived n-gram statistics from each input sequence, encouraging the model to capture higher-order dependencies beyond local context. This integration enables CARMANIA to learn organism-specific sequence structures that reflect both evolutionary constraints and functional organization. We evaluate CARMANIA across diverse genomic tasks, including regulatory element prediction, functional gene classification, taxonomic inference, antimicrobial resistance detection, and biosynthetic gene cluster classification. CARMANIA outperforms the previous best long-context model by at least 7\%, matches state-of-the-art on shorter sequences (exceeding prior results on 20/40 tasks while running $\sim$2.5$\times$ faster), and shows particularly strong improvements on enhancer and housekeeping gene classification tasks—including up to a 34\% absolute gain in Matthews correlation coefficient (MCC) for enhancer prediction. The TM loss boosts accuracy in 33 of 40 tasks, especially where local motifs or regulatory patterns drive prediction. This enables more effective modeling of sequence-dependent biological features while maintaining robustness across non-coding and low-signal regions. Code available at https://github.com/EESI/carmania.
Mohammadsaleh Refahi, Mahdi Abavisani, Bahrad A. Sokhansanj, James R. Brown, Gail L. Rosen
NeurIPS1