Emaad Khwaja

dblp:369/5959 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
2 papers
Deep learning architectures and training · 54% Generative modeling · 46%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › foundation model
time series foundation model
0.912025
This Time is Different: An Observability Perspective on Time Series Foundation Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
CELL-E: A Text-to-Image Transformer for Protein Image Prediction · RECOMB 2024
Bioinformatics and computational biology
protein structure prediction
0.812024
CELL-E: A Text-to-Image Transformer for Protein Image Prediction · RECOMB 2024
Bioinformatics and computational biology › protein design
de novo protein design
0.712023
CELLE-2: Translating Proteins to Pictures and Back with a Bidirectional Text-to-Image Transformer · NeurIPS 2023
Bioinformatics and computational biology
protein design
0.712023
CELLE-2: Translating Proteins to Pictures and Back with a Bidirectional Text-to-Image Transformer · NeurIPS 2023
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction
0.712023
CELLE-2: Translating Proteins to Pictures and Back with a Bidirectional Text-to-Image Transformer · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

text-to-image generation · 2.2synthetic data generation · 1.7pre-training · 1.7transformer · 1.5bidirectional transformer · 0.7
YearPublicationVenuePosition
2025 This Time is Different: An Observability Perspective on Time Series Foundation Models
abstract
We introduce Toto, a time series forecasting foundation model with 151 million parameters. Toto uses a modern decoder-only architecture coupled with architectural innovations designed to account for specific challenges found in multivariate observability time series data. Toto's pre-training corpus is a mixture of observability data, open datasets, and synthetic data, and is 4-10$\times$ larger than those of leading time series foundation models. Additionally, we introduce BOOM, a large-scale benchmark consisting of 350 million observations across 2,807 real-world time series. For both Toto and BOOM, we source observability data exclusively from our own telemetry and internal observability metrics. Extensive evaluations demonstrate that Toto achieves state-of-the-art performance on both BOOM and on established general purpose time series forecasting benchmarks. Toto's model weights, inference code, and evaluation scripts, as well as BOOM's data and evaluation code, are all available as open source under the Apache 2.0 License.
Ben Cohen, Emaad Khwaja, Youssef Doubli, Salahidine Lemaachi, Chris Lettieri, Charles Masson, Hugo Miccinilli, Elise Ramé, Qiqi Ren, Afshin Rostamizadeh, Jean Ogier du Terrail, Anna-Monica Toon, Stephan Xie, Zongzhe Xu, Viktoriya Zhukova, David Asker, Ameet Talwalkar, Othmane Abou-Amal
NeurIPS2
2024 CELL-E: A Text-to-Image Transformer for Protein Image Prediction
Emaad Khwaja, Yun S. Song
RECOMB1
2023 CELLE-2: Translating Proteins to Pictures and Back with a Bidirectional Text-to-Image Transformer
abstract
We present CELL-E 2, a novel bidirectional transformer that can generate images depicting protein subcellular localization from the amino acid sequences (and vice versa). Protein localization is a challenging problem that requires integrating sequence and image information, which most existing methods ignore. CELL-E 2 extends the work of CELL-E, not only capturing the spatial complexity of protein localization and produce probability estimates of localization atop a nucleus image, but also being able to generate sequences from images, enabling de novo protein design. We train and finetune CELL-E 2 on two large-scale datasets of human proteins. We also demonstrate how to use CELL-E 2 to create hundreds of novel nuclear localization signals (NLS). Results and interactive demos are featured at https://bohuanglab.github.io/CELL-E_2/.
Emaad Khwaja, Yun Song, Aaron Agarunov
NeurIPS1