Philippe Chlenski

dblp:250/6177 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2951-4385ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 42% Representation and self-supervised learning · 29% Kernel, tree and ensemble methods · 25%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
embedding space
0.912025
Mixed-curvature decision trees and random forests · ICML 2025
Machine learning › Representation and self-supervised learning
product manifold
0.912025
Mixed-curvature decision trees and random forests · ICML 2025
Bioinformatics and computational biology › sequence analysis › sequence feature extraction
DNA sequence representation
0.912025
Hyperbolic Genome Embeddings · ICLR 2025
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling
0.912025
Hyperbolic Genome Embeddings · ICLR 2025
Data mining › predictive modeling
classification
0.912025
Mixed-curvature decision trees and random forests · ICML 2025
Data mining › predictive modeling › classification
decision tree learning
0.912025
Mixed-curvature decision trees and random forests · ICML 2025
Data mining › predictive modeling › classification › ensemble learning
random forest
0.912025
Mixed-curvature decision trees and random forests · ICML 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.812024
Transcoders find interpretable LLM feature circuits · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods
decision tree
0.812024
Fast Hyperboloid Decision Tree Algorithms · ICLR 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Transcoders find interpretable LLM feature circuits · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.812024
Transcoders find interpretable LLM feature circuits · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.812024
Fast Hyperboloid Decision Tree Algorithms · ICLR 2024
Computer vision › 3D vision › geometric deep learning
hyperbolic neural networks
0.312025
Hyperbolic Genome Embeddings · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.212024
Transcoders find interpretable LLM feature circuits · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

hyperbolic geometry · 2.5hyperspherical geometry · 1.7hyperbolic CNNs · 1.7contrastive learning · 1.7transcoder · 0.8sparse autoencoder · 0.8inner product · 0.8
YearPublicationVenuePosition
2025 Variational Combinatorial Sequential Monte Carlo for Bayesian Phylogenetics in Hyperbolic Space
abstract
Hyperbolic space naturally encodes hierarchical structures such as phylogenies (binary trees), where inward-bending geodesics reflect paths through least common ancestors, and the exponential growth of neighborhoods mirrors the super-exponential scaling of topologies. This scaling challenge limits the efficiency of Euclidean-based approximate Bayesian inference methods. Motivated by the geometric connections between trees and hyperbolic space, we develop novel hyperbolic extensions of two sequential search algorithms: Combinatorial and Nested Combinatorial Sequential Monte Carlo (\textsc{Csmc} and \textsc{Ncsmc}). Our approach introduces consistent and unbiased estimators, along with variational inference methods (\textsc{H-Vcsmc} and \textsc{H-Vncsmc}), which outperform their Euclidean counterparts. Empirical results demonstrate improved speed, scalability and performance in high-dimensional Bayesian phylogenetic inference tasks.
Alex Chen, Philippe Chlenski, Kenneth Munyuza, Antonio Khalil Moretti, Christian A. Naesseth, Itsik Pe'er
AISTATS2
2025 Hyperbolic Genome Embeddings
abstract
Current approaches to genomic sequence modeling often struggle to align the inductive biases of machine learning models with the evolutionarily-informed structure of biological systems. To this end, we formulate a novel application of hyperbolic CNNs that exploits this structure, enabling more expressive DNA sequence representations. Our strategy circumvents the need for explicit phylogenetic mapping while discerning key properties of sequences pertaining to core functional and regulatory behavior. Across 37 out of 42 genome interpretation benchmark datasets, our hyperbolic models outperform their Euclidean equivalents. Notably, our approach even surpasses state-of-the-art performance on seven GUE benchmark datasets, consistently outperforming many DNA language models while using orders of magnitude fewer parameters and avoiding pretraining. Our results include a novel set of benchmark datasets---the Transposable Elements Benchmark---which explores a major but understudied component of the genome with deep evolutionary significance. We further motivate our work by exploring how our hyperbolic models recognize genomic signal under various data-generating conditions and by constructing an empirical method for interpreting the hyperbolicity of dataset embeddings. Throughout these assessments, we find persistent evidence highlighting the potential of our hyperbolic framework as a robust paradigm for genome representation learning. Our code and benchmark datasets are available at https://github.com/rrkhan/HGE.
Raiyan R. Khan, Philippe Chlenski, Itsik Pe'er
ICLR2
2025 Mixed-curvature decision trees and random forests
abstract
Decision trees (DTs) and their random forest (RF) extensions are workhorses of classification and regression in Euclidean spaces. However, algorithms for learning in non-Euclidean spaces are still limited. We extend DT and RF algorithms to product manifolds: Cartesian products of several hyperbolic, hyperspherical, or Euclidean components. Such manifolds handle heterogeneous curvature while still factorizing neatly into simpler components, making them compelling embedding spaces for complex datasets. Our novel angular reformulation respects manifold geometry while preserving the algorithmic properties that make decision trees effective. In the special cases of single-component manifolds, our method simplifies to its Euclidean or hyperbolic counterparts, or introduces hyperspherical DT algorithms, depending on the curvature. In benchmarks on a diverse suite of 57 classification, regression, and link prediction tasks, our product RFs ranked first on 29 tasks and came in the top 2 for 41. This highlights the value of product RFs as straightforward yet powerful new tools for data analysis in product manifolds. Code for our method is available at https://github.com/pchlenski/manify.
Philippe Chlenski, Quentin Chu, Raiyan R. Khan, Kaizhu Du, Antonio Khalil Moretti, Itsik Pe'er
ICML1
2024 Fast Hyperboloid Decision Tree Algorithms
abstract
Hyperbolic geometry is gaining traction in machine learning due to its capacity to effectively capture hierarchical structures in real-world data. Hyperbolic spaces, where neighborhoods grow exponentially, offer substantial advantages and have consistently delivered state-of-the-art results across diverse applications. However, hyperbolic classifiers often grapple with computational challenges. Methods reliant on Riemannian optimization frequently exhibit sluggishness, stemming from the increased computational demands of operations on Riemannian manifolds. In response to these challenges, we present HyperDT, a novel extension of decision tree algorithms into hyperbolic space. Crucially, HyperDT eliminates the need for computationally intensive Riemannian optimization, numerically unstable exponential and logarithmic maps, or pairwise comparisons between points by leveraging inner products to adapt Euclidean decision tree algorithms to hyperbolic space. Our approach is conceptually straightforward and maintains constant-time decision complexity while mitigating the scalability issues inherent in high-dimensional Euclidean spaces. Building upon HyperDT, we introduce HyperRF, a hyperbolic random forest model. Extensive benchmarking across diverse datasets underscores the superior performance of these models, providing a swift, precise, accurate, and user-friendly toolkit for hyperbolic data analysis.
Philippe Chlenski, Ethan Turok, Antonio Khalil Moretti, Itsik Pe'er
ICLR1
2024 Transcoders find interpretable LLM feature circuits
abstract
A key goal in mechanistic interpretability is circuit analysis: finding sparse subgraphs of models corresponding to specific behaviors or capabilities. However, MLP sublayers make fine-grained circuit analysis on transformer-based language models difficult. In particular, interpretable features—such as those found by sparse autoencoders (SAEs)—are typically linear combinations of extremely many neurons, each with its own nonlinearity to account for. Circuit analysis in this setting thus either yields intractably large circuits or fails to disentangle local and global behavior. To address this we explore **transcoders**, which seek to faithfully approximate a densely activating MLP layer with a wider, sparsely-activating MLP layer. We introduce a novel method for using transcoders to perform weights-based circuit analysis through MLP sublayers. The resulting circuits neatly factorize into input-dependent and input-invariant terms. We then successfully train transcoders on language models with 120M, 410M, and 1.4B parameters, and find them to perform at least on par with SAEs in terms of sparsity, faithfulness, and human-interpretability. Finally, we apply transcoders to reverse-engineer unknown circuits in the model, and we obtain novel insights regarding the "greater-than circuit" in GPT2-small. Our results suggest that transcoders can prove effective in decomposing model computations involving MLPs into interpretable circuits. Code is available at https://github.com/jacobdunefsky/transcoder_circuits/
Jacob Dunefsky, Philippe Chlenski, Neel Nanda
NeurIPS2
2019 A machine learning-based service for estimating quality of genomes using PATRIC
abstract
BACKGROUND: Recent advances in high-volume sequencing technology and mining of genomes from metagenomic samples call for rapid and reliable genome quality evaluation. The current release of the PATRIC database contains over 220,000 genomes, and current metagenomic technology supports assemblies of many draft-quality genomes from a single sample, most of which will be novel. DESCRIPTION: We have added two quality assessment tools to the PATRIC annotation pipeline. EvalCon uses supervised machine learning to calculate an annotation consistency score. EvalG implements a variant of the CheckM algorithm to estimate contamination and completeness of an annotated genome.We report on the performance of these tools and the potential utility of the consistency score. Additionally, we provide contamination, completeness, and consistency measures for all genomes in PATRIC and in a recent set of metagenomic assemblies. CONCLUSION: EvalG and EvalCon facilitate the rapid quality control and exploration of PATRIC-annotated draft genomes.
Bruce D. Parrello, Rory Butler, Philippe Chlenski, Robert Olson, Jamie C. Overbeek, Gordon D. Pusch, Veronika Vonstein, Ross A. Overbeek
BMC Bioinform.3