VLDB 2026 Research / reviewers in the wild / expert
Philippe Chlenski
dblp:250/6177
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2951-4385ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 42% Representation and self-supervised learning · 29% Kernel, tree and ensemble methods · 25% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
embedding space |
0.9 | 1 | 2025 | Mixed-curvature decision trees and random forests · ICML 2025 |
Machine learning › Representation and self-supervised learning
product manifold |
0.9 | 1 | 2025 | Mixed-curvature decision trees and random forests · ICML 2025 |
Bioinformatics and computational biology › sequence analysis › sequence feature extraction
DNA sequence representation |
0.9 | 1 | 2025 | Hyperbolic Genome Embeddings · ICLR 2025 |
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling |
0.9 | 1 | 2025 | Hyperbolic Genome Embeddings · ICLR 2025 |
Data mining › predictive modeling
classification |
0.9 | 1 | 2025 | Mixed-curvature decision trees and random forests · ICML 2025 |
Data mining › predictive modeling › classification
decision tree learning |
0.9 | 1 | 2025 | Mixed-curvature decision trees and random forests · ICML 2025 |
Data mining › predictive modeling › classification › ensemble learning
random forest |
0.9 | 1 | 2025 | Mixed-curvature decision trees and random forests · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis |
0.8 | 1 | 2024 | Transcoders find interpretable LLM feature circuits · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods
decision tree |
0.8 | 1 | 2024 | Fast Hyperboloid Decision Tree Algorithms · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Transcoders find interpretable LLM feature circuits · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.8 | 1 | 2024 | Transcoders find interpretable LLM feature circuits · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest |
0.8 | 1 | 2024 | Fast Hyperboloid Decision Tree Algorithms · ICLR 2024 |
Computer vision › 3D vision › geometric deep learning
hyperbolic neural networks |
0.3 | 1 | 2025 | Hyperbolic Genome Embeddings · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability |
0.2 | 1 | 2024 | Transcoders find interpretable LLM feature circuits · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
hyperbolic geometry · 2.5hyperspherical geometry · 1.7hyperbolic CNNs · 1.7contrastive learning · 1.7transcoder · 0.8sparse autoencoder · 0.8inner product · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Variational Combinatorial Sequential Monte Carlo for Bayesian Phylogenetics in Hyperbolic SpaceabstractHyperbolic space naturally encodes hierarchical structures such as phylogenies (binary trees), where inward-bending geodesics reflect paths through least common ancestors, and the exponential growth of neighborhoods mirrors the super-exponential scaling of topologies. This scaling challenge limits the efficiency of Euclidean-based approximate Bayesian inference methods. Motivated by the geometric connections between trees and hyperbolic space, we develop novel hyperbolic extensions of two sequential search algorithms: Combinatorial and Nested Combinatorial Sequential Monte Carlo (\textsc{Csmc} and \textsc{Ncsmc}). Our approach introduces consistent and unbiased estimators, along with variational inference methods (\textsc{H-Vcsmc} and \textsc{H-Vncsmc}), which outperform their Euclidean counterparts. Empirical results demonstrate improved speed, scalability and performance in high-dimensional Bayesian phylogenetic inference tasks. Alex Chen, Philippe Chlenski, Kenneth Munyuza, Antonio Khalil Moretti, Christian A. Naesseth, Itsik Pe'er |
AISTATS | 2 |
| 2025 | Hyperbolic Genome EmbeddingsabstractCurrent approaches to genomic sequence modeling often struggle to align the inductive biases of machine learning models with the evolutionarily-informed structure of biological systems. To this end, we formulate a novel application of hyperbolic CNNs that exploits this structure, enabling more expressive DNA sequence representations. Our strategy circumvents the need for explicit phylogenetic mapping while discerning key properties of sequences pertaining to core functional and regulatory behavior. Across 37 out of 42 genome interpretation benchmark datasets, our hyperbolic models outperform their Euclidean equivalents. Notably, our approach even surpasses state-of-the-art performance on seven GUE benchmark datasets, consistently outperforming many DNA language models while using orders of magnitude fewer parameters and avoiding pretraining. Our results include a novel set of benchmark datasets---the Transposable Elements Benchmark---which explores a major but understudied component of the genome with deep evolutionary significance. We further motivate our work by exploring how our hyperbolic models recognize genomic signal under various data-generating conditions and by constructing an empirical method for interpreting the hyperbolicity of dataset embeddings. Throughout these assessments, we find persistent evidence highlighting the potential of our hyperbolic framework as a robust paradigm for genome representation learning. Our code and benchmark datasets are available at https://github.com/rrkhan/HGE. Raiyan R. Khan, Philippe Chlenski, Itsik Pe'er |
ICLR | 2 |
| 2025 | Mixed-curvature decision trees and random forestsabstractDecision trees (DTs) and their random forest (RF) extensions are workhorses of classification and regression in Euclidean spaces. However, algorithms for learning in non-Euclidean spaces are still limited. We extend DT and RF algorithms to product manifolds: Cartesian products of several hyperbolic, hyperspherical, or Euclidean components. Such manifolds handle heterogeneous curvature while still factorizing neatly into simpler components, making them compelling embedding spaces for complex datasets. Our novel angular reformulation respects manifold geometry while preserving the algorithmic properties that make decision trees effective. In the special cases of single-component manifolds, our method simplifies to its Euclidean or hyperbolic counterparts, or introduces hyperspherical DT algorithms, depending on the curvature. In benchmarks on a diverse suite of 57 classification, regression, and link prediction tasks, our product RFs ranked first on 29 tasks and came in the top 2 for 41. This highlights the value of product RFs as straightforward yet powerful new tools for data analysis in product manifolds. Code for our method is available at https://github.com/pchlenski/manify. Philippe Chlenski, Quentin Chu, Raiyan R. Khan, Kaizhu Du, Antonio Khalil Moretti, Itsik Pe'er |
ICML | 1 |
| 2024 | Fast Hyperboloid Decision Tree AlgorithmsabstractHyperbolic geometry is gaining traction in machine learning due to its capacity to effectively capture hierarchical structures in real-world data. Hyperbolic spaces, where neighborhoods grow exponentially, offer substantial advantages and have consistently delivered state-of-the-art results across diverse applications. However, hyperbolic classifiers often grapple with computational challenges. Methods reliant on Riemannian optimization frequently exhibit sluggishness, stemming from the increased computational demands of operations on Riemannian manifolds. In response to these challenges, we present HyperDT, a novel extension of decision tree algorithms into hyperbolic space. Crucially, HyperDT eliminates the need for computationally intensive Riemannian optimization, numerically unstable exponential and logarithmic maps, or pairwise comparisons between points by leveraging inner products to adapt Euclidean decision tree algorithms to hyperbolic space. Our approach is conceptually straightforward and maintains constant-time decision complexity while mitigating the scalability issues inherent in high-dimensional Euclidean spaces. Building upon HyperDT, we introduce HyperRF, a hyperbolic random forest model. Extensive benchmarking across diverse datasets underscores the superior performance of these models, providing a swift, precise, accurate, and user-friendly toolkit for hyperbolic data analysis. Philippe Chlenski, Ethan Turok, Antonio Khalil Moretti, Itsik Pe'er |
ICLR | 1 |
| 2024 | Transcoders find interpretable LLM feature circuitsabstractA key goal in mechanistic interpretability is circuit analysis: finding sparse subgraphs of models corresponding to specific behaviors or capabilities. However, MLP sublayers make fine-grained circuit analysis on transformer-based language models difficult. In particular, interpretable features—such as those found by sparse autoencoders (SAEs)—are typically linear combinations of extremely many neurons, each with its own nonlinearity to account for. Circuit analysis in this setting thus either yields intractably large circuits or fails to disentangle local and global behavior. To address this we explore **transcoders**, which seek to faithfully approximate a densely activating MLP layer with a wider, sparsely-activating MLP layer. We introduce a novel method for using transcoders to perform weights-based circuit analysis through MLP sublayers. The resulting circuits neatly factorize into input-dependent and input-invariant terms. We then successfully train transcoders on language models with 120M, 410M, and 1.4B parameters, and find them to perform at least on par with SAEs in terms of sparsity, faithfulness, and human-interpretability. Finally, we apply transcoders to reverse-engineer unknown circuits in the model, and we obtain novel insights regarding the "greater-than circuit" in GPT2-small. Our results suggest that transcoders can prove effective in decomposing model computations involving MLPs into interpretable circuits. Code is available at https://github.com/jacobdunefsky/transcoder_circuits/ Jacob Dunefsky, Philippe Chlenski, Neel Nanda |
NeurIPS | 2 |
| 2019 | A machine learning-based service for estimating quality of genomes using PATRICabstractBACKGROUND: Recent advances in high-volume sequencing technology and mining of genomes from metagenomic samples call for rapid and reliable genome quality evaluation. The current release of the PATRIC database contains over 220,000 genomes, and current metagenomic technology supports assemblies of many draft-quality genomes from a single sample, most of which will be novel. DESCRIPTION: We have added two quality assessment tools to the PATRIC annotation pipeline. EvalCon uses supervised machine learning to calculate an annotation consistency score. EvalG implements a variant of the CheckM algorithm to estimate contamination and completeness of an annotated genome.We report on the performance of these tools and the potential utility of the consistency score. Additionally, we provide contamination, completeness, and consistency measures for all genomes in PATRIC and in a recent set of metagenomic assemblies. CONCLUSION: EvalG and EvalCon facilitate the rapid quality control and exploration of PATRIC-annotated draft genomes. Bruce D. Parrello, Rory Butler, Philippe Chlenski, Robert Olson, Jamie C. Overbeek, Gordon D. Pusch, Veronika Vonstein, Ross A. Overbeek |
BMC Bioinform. | 3 |