Paul Kahlmeyer

dblp:299/5248 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 44% Knowledge representation and reasoning · 34% Information extraction and text analysis · 22%
Theoretical computer science
2 papers
Algorithms and data structures · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 50% Information retrieval · 50%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression
0.812024
Scaling Up Unbiased Search-based Symbolic Regression · IJCAI 2024
Machine learning and data management › model evaluation
embedding evaluation
0.612022
Leveraging the Wikipedia Graph for Evaluating Word Embeddings · IJCAI 2022
Information retrieval › similarity measure
graph-based similarity
0.612022
Leveraging the Wikipedia Graph for Evaluating Word Embeddings · IJCAI 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.512021
Method of Moments for Topic Models with Mixed Discrete and Continuous Features · IJCAI 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
method of moments
0.512021
Method of Moments for Topic Models with Mixed Discrete and Continuous Features · IJCAI 2021
Natural language and speech › Information extraction and text analysis
topic model
0.512021
Method of Moments for Topic Models with Mixed Discrete and Continuous Features · IJCAI 2021

Methods — techniques the papers use, named apart from their topics

unbiased search · 1.5genetic programming · 1.5symbolic regression · 0.9functional dependence test · 0.9word embeddings · 0.6pearson's method of moments · 0.5method of moments · 0.5
YearPublicationVenuePosition
2025 Dimension Reduction for Symbolic Regression
abstract
Solutions of symbolic regression problems are expressions that are composed of input variables and operators from a finite set of function symbols. One measure for evaluating symbolic regression algorithms is their ability to recover formulae, up to symbolic equivalence, from finite samples. Not unexpectedly, the recovery problem becomes harder when the formula gets more complex, that is, when the number of variables and operators gets larger. Variables in naturally occurring symbolic formulas often appear only in fixed combinations. This can be exploited in symbolic regression by substituting one new variable for the combination, effectively reducing the number of variables. However, finding valid substitutions is challenging. Here, we address this challenge by searching over the expression space of small substitutions and testing for validity. The validity test is reduced to a test of functional dependence. The resulting iterative dimension reduction procedure can be used with any symbolic regression approach. We show that it reliably identifies valid substitutions and significantly boosts the performance of different types of state-of-the-art symbolic regression algorithms.
Paul Kahlmeyer, Markus Fischer 0005, Joachim Giesen
AAAI1
2025 Discovering Symmetries of ODEs by Symbolic Regression
abstract
Solving systems of ordinary differential equations (ODEs) is essential when it comes to understanding the behavior of dynamical systems. Yet, automated solving remains challenging, in particular for nonlinear systems. Computer algebra systems (CASs) provide support for solving ODEs by first simplifying them, in particular through the use of Lie point symmetries. Finding these symmetries is, however, itself a difficult problem for CASs. Recent works in symbolic regression have shown promising results for recovering symbolic expressions from data. Here, we adapt search-based symbolic regression to the task of finding generators of Lie point symmetries. With this approach, we can find symmetries of ODEs that existing CASs cannot find.
Paul Kahlmeyer, Niklas Merk, Joachim Giesen
AAAI1
2024 Scaling Up Unbiased Search-based Symbolic Regression
Paul Kahlmeyer, Joachim Giesen, Michael Habeck, Henrik Voigt
IJCAI1
2022 Leveraging the Wikipedia Graph for Evaluating Word Embeddings
abstract
Deep learning models for different NLP tasks often rely on pre-trained word embeddings, that is, vector representations of words. Therefore, it is crucial to evaluate pre-trained word embeddings independently of downstream tasks. Such evaluations try to assess whether the geometry induced by a word embedding captures connections made in natural language, such as, analogies, clustering of words, or word similarities. Here, traditionally, similarity is measured by comparison to human judgment. However, explicitly annotating word pairs with similarity scores by surveying humans is expensive. We tackle this problem by formulating a similarity measure that is based on an agent for routing the Wikipedia hyperlink graph. In this graph, word similarities are implicitly encoded by edges between articles. We show on the English Wikipedia that our measure correlates well with a large group of traditional similarity measures, while covering a much larger proportion of words and avoiding explicit human labeling. Moreover, since Wikipedia is available in more than 300 languages, our measure can easily be adapted to other languages, in contrast to traditional similarity measures.
Joachim Giesen, Paul Kahlmeyer, Frank Nussbaum, Sina Zarrieß
IJCAI2
2021 Method of Moments for Topic Models with Mixed Discrete and Continuous Features
abstract
Topic models are characterized by a latent class variable that represents the different topics. Traditionally, their observable variables are modeled as discrete variables like, for instance, in the prototypical latent Dirichlet allocation (LDA) topic model. In LDA, words in text documents are encoded by discrete count vectors with respect to some dictionary. The classical approach for learning topic models optimizes a likelihood function that is non-concave due to the presence of the latent variable. Hence, this approach mostly boils down to using search heuristics like the EM algorithm for parameter estimation. Recently, it was shown that topic models can be learned with strong algorithmic and statistical guarantees through Pearson's method of moments. Here, we extend this line of work to topic models that feature discrete as well as continuous observable variables (features). Moving beyond discrete variables as in LDA allows for more sophisticated features and a natural extension of topic models to other modalities than text, like, for instance, images. We provide algorithmic and statistical guarantees for the method of moments applied to the extended topic model that we corroborate experimentally on synthetic data. We also demonstrate the applicability of our model on real-world document data with embedded images that we preprocess into continuous state-of-the-art feature vectors.
Joachim Giesen, Paul Kahlmeyer, Sören Laue, Matthias Mitterreiter, Frank Nussbaum, Christoph Staudt, Sina Zarrieß
IJCAI2