VLDB 2026 Research / reviewers in the wild / expert
Aleem Khan
dblp:292/7529
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 23% Deep learning architectures and training · 23% Probabilistic and Bayesian machine learning · 23% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.9 | 1 | 2025 | Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.9 | 1 | 2025 | Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025 |
Digital forensics and information hiding › synthetic media detection
machine-generated text detection |
0.8 | 1 | 2024 | Few-Shot Detection of Machine-Generated Text using Style Representations · ICLR 2024 |
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
authorship attribution |
0.2 | 1 | 2024 | Few-Shot Detection of Machine-Generated Text using Style Representations · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
style representations · 1.5few-shot learning · 1.5autoregressive modeling · 0.9MCMC · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Extrapolative Sequence Transformations from Markov ChainsabstractMost successful applications of deep learning involve similar training and test conditions. However, tasks such as biological sequence design involve searching for sequences that improve desirable properties beyond previously known values, which requires novel hypotheses that \emph{extrapolate} beyond training data. In these settings, extrapolation may be achieved by using random search methods such as Markov chain Monte Carlo (MCMC), which, given an initial state, sample local transformations to approximate a target density that rewards states with the desired properties. However, even with a well-designed proposal, MCMC may struggle to explore large structured state spaces efficiently. Rather than relying on stochastic search, it would be desirable to have a model that greedily optimizes the properties of interest, successfully extrapolating in as few steps as possible. We propose to learn such a model from the Markov chains resulting from MCMC search. Specifically, our approach uses selected states from Markov chains as a source of training data for an autoregressive model, which is then able to efficiently generate novel sequences that extrapolate along the sequence-level properties of interest. The proposed approach is validated on three problems: protein sequence design, text sentiment control, and text anonymization. We find that the autoregressive model can extrapolate as well or better than MCMC, but with the additional benefits of scalability and significantly higher sample efficiency. Sophia Hager, Aleem Khan, Nicholas Andrews |
ICML | 2 |
| 2024 | Few-Shot Detection of Machine-Generated Text using Style RepresentationsabstractThe advent of instruction-tuned language models that convincingly mimic human writing poses a significant risk of abuse. For example, such models could be used for plagiarism, disinformation, spam, or phishing. However, such abuse may be counteracted with the ability to detect whether a piece of text was composed by a language model rather than a human. Some previous approaches to this problem have relied on supervised methods trained on corpora of confirmed human and machine-written documents. Unfortunately, model under-specification poses an unavoidable challenge for such detectors, making them brittle in the face of data shifts, such as the release of further language models producing still more fluent text than the models used to train the detectors. Other previous approaches require access to the models that generated the text to be detected at inference or detection time, which is often impractical. In light of these challenge, we pursue a fundamentally different approach not relying on samples from language models of concern at training time. Instead, we propose to leverage representations of writing style estimated from human-authored text. Indeed, we find that features effective at distinguishing among human authors are also effective at distinguishing human from machine authors, including state of the art large language models like Llama 2, ChatGPT, and GPT-4. Furthermore, given handfuls of examples composed by each of several specific language models of interest, our approach affords the ability to predict which model specifically generated a given document. Rafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen, Marcus Bishop, Nicholas Andrews |
ICLR | 3 |
| 2021 | Learning Universal Authorship RepresentationsabstractRafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Rafael A. Rivera Soto, Olivia Elizabeth Miano, Juanita Ordoñez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews |
EMNLP (1) | 5 |
| 2021 | A Deep Metric Learning Approach to Account LinkingabstractWe consider the task of linking social media accounts that belong to the same author in an automated fashion on the basis of the content and meta-data of the corresponding document streams. We focus on learning an embedding that maps variable-sized samples of user activity–ranging from single posts to entire months of activity–to a vector space, where samples by the same author map to nearby points. Our approach does not require human-annotated data for training purposes, which allows us to leverage large amounts of social media content. The proposed model outperforms several competitive baselines under a novel evaluation framework modeled after established recognition benchmarks in other domains. Our method achieves high linking accuracy, even with small samples from accounts not seen at training time, a prerequisite for practical applications of the proposed linking framework. Aleem Khan, Elizabeth Fleming, Noah Schofield, Marcus Bishop, Nicholas Andrews |
NAACL-HLT | 1 |