Aleem Khan

dblp:292/7529 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 23% Deep learning architectures and training · 23% Probabilistic and Bayesian machine learning · 23%
Network and information security
1 paper
Digital forensics and information hiding · 100%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
0.912025
Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.912025
Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.912025
Learning Extrapolative Sequence Transformations from Markov Chains · ICML 2025
Digital forensics and information hiding › synthetic media detection
machine-generated text detection
0.812024
Few-Shot Detection of Machine-Generated Text using Style Representations · ICLR 2024
Natural language and speech › Information extraction and text analysis › text mining › authorship analysis
authorship attribution
0.212024
Few-Shot Detection of Machine-Generated Text using Style Representations · ICLR 2024

Methods — techniques the papers use, named apart from their topics

style representations · 1.5few-shot learning · 1.5autoregressive modeling · 0.9MCMC · 0.9
YearPublicationVenuePosition
2025 Learning Extrapolative Sequence Transformations from Markov Chains
abstract
Most successful applications of deep learning involve similar training and test conditions. However, tasks such as biological sequence design involve searching for sequences that improve desirable properties beyond previously known values, which requires novel hypotheses that \emph{extrapolate} beyond training data. In these settings, extrapolation may be achieved by using random search methods such as Markov chain Monte Carlo (MCMC), which, given an initial state, sample local transformations to approximate a target density that rewards states with the desired properties. However, even with a well-designed proposal, MCMC may struggle to explore large structured state spaces efficiently. Rather than relying on stochastic search, it would be desirable to have a model that greedily optimizes the properties of interest, successfully extrapolating in as few steps as possible. We propose to learn such a model from the Markov chains resulting from MCMC search. Specifically, our approach uses selected states from Markov chains as a source of training data for an autoregressive model, which is then able to efficiently generate novel sequences that extrapolate along the sequence-level properties of interest. The proposed approach is validated on three problems: protein sequence design, text sentiment control, and text anonymization. We find that the autoregressive model can extrapolate as well or better than MCMC, but with the additional benefits of scalability and significantly higher sample efficiency.
Sophia Hager, Aleem Khan, Nicholas Andrews
ICML2
2024 Few-Shot Detection of Machine-Generated Text using Style Representations
abstract
The advent of instruction-tuned language models that convincingly mimic human writing poses a significant risk of abuse. For example, such models could be used for plagiarism, disinformation, spam, or phishing. However, such abuse may be counteracted with the ability to detect whether a piece of text was composed by a language model rather than a human. Some previous approaches to this problem have relied on supervised methods trained on corpora of confirmed human and machine-written documents. Unfortunately, model under-specification poses an unavoidable challenge for such detectors, making them brittle in the face of data shifts, such as the release of further language models producing still more fluent text than the models used to train the detectors. Other previous approaches require access to the models that generated the text to be detected at inference or detection time, which is often impractical. In light of these challenge, we pursue a fundamentally different approach not relying on samples from language models of concern at training time. Instead, we propose to leverage representations of writing style estimated from human-authored text. Indeed, we find that features effective at distinguishing among human authors are also effective at distinguishing human from machine authors, including state of the art large language models like Llama 2, ChatGPT, and GPT-4. Furthermore, given handfuls of examples composed by each of several specific language models of interest, our approach affords the ability to predict which model specifically generated a given document.
Rafael A. Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen, Marcus Bishop, Nicholas Andrews
ICLR3
2021 Learning Universal Authorship Representations
abstract
Rafael A. Rivera-Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Rafael A. Rivera Soto, Olivia Elizabeth Miano, Juanita Ordoñez, Barry Y. Chen, Aleem Khan, Marcus Bishop, Nicholas Andrews
EMNLP (1)5
2021 A Deep Metric Learning Approach to Account Linking
abstract
We consider the task of linking social media accounts that belong to the same author in an automated fashion on the basis of the content and meta-data of the corresponding document streams. We focus on learning an embedding that maps variable-sized samples of user activity–ranging from single posts to entire months of activity–to a vector space, where samples by the same author map to nearby points. Our approach does not require human-annotated data for training purposes, which allows us to leverage large amounts of social media content. The proposed model outperforms several competitive baselines under a novel evaluation framework modeled after established recognition benchmarks in other domains. Our method achieves high linking accuracy, even with small samples from accounts not seen at training time, a prerequisite for practical applications of the proposed linking framework.
Aleem Khan, Elizabeth Fleming, Noah Schofield, Marcus Bishop, Nicholas Andrews
NAACL-HLT1