Ulf Brefeld

dblp:99/122 · DBLP profile ↗
← Back
49ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-9600-6463ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 21 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Self-improvement for Computerized Adaptive Testing
Yannick Rudolph, Kai Neubauer, Ulf Brefeld
ECML/PKDD (2)3
2025 Interactive sequential generative models for team sports
abstract
Abstract Understanding spatiotemporal coordination of players in team sports is key to movement models, pattern detection, and computational tactics. Existing generative models propose to capture all stochasticity by a single latent variable and may suffer from entangled representations, or aim to uncover interaction structures of players but then do not focus on their generative ability. As a remedy, we propose a hierarchical latent variable model for predicting trajectories of multiple players. In the generative model, both, discrete role assignments and a latent interaction graph are sampled to allow for different models in subsequent node updates and message passing operations between nodes, where standard Gaussian latent variables are employed per agent and timestep. We cast our approach as a variational autoencoder that provides a disentangled latent space to capture variability in team sport movements and propose a neural architecture for its optimization. We empirically evaluate our approach on tracking data from basketball and soccer and observe that our contribution outperforms the state-of-art in all experiments.
Dennis Fassmeyer, Moritz Cordes, Ulf Brefeld
Mach. Learn.3
2025 Masked autoencoder for multiagent trajectories
abstract
Abstract Automatically labeling trajectories of multiple agents is key to behavioral analyses but usually requires a large amount of manual annotations. This also applies to the domain of team sport analyses. In this paper, we specifically show how pretraining transformer models improves the classification performance on tracking data from professional soccer. For this purpose, we propose a novel self-supervised masked autoencoder for multiagent trajectories to effectively learn from only a few labeled sequences. Our approach builds upon a factorized transformer architecture for multiagent trajectory data and employs a masking scheme on the level of individual agent trajectories. As a result, our model allows for a reconstruction of masked trajectory segments while being permutation equivariant with respect to the agent trajectories. In addition to experiments on soccer, we demonstrate the usefulness of the proposed pretraining approach on multiagent pose data from entomology. In contrast to related work, our approach is conceptually much simpler, does not require handcrafted features and naturally allows for permutation invariance in downstream tasks.
Yannick Rudolph, Ulf Brefeld
Mach. Learn.2
2023 Hands in Focus: Sign Language Recognition Via Top-Down Attention
abstract
In this paper, we propose a novel Sign Language Recognition (SLR) model that leverages the task-specific knowledge to incorporate Top-Down (TD) attention to focus the processing of the network on the most relevant parts of the input video sequence. For SLR, this includes information about the hands’ shape, orientation and positions, and motion trajectory. Our model consists of three streams that process RGB, optical flow and TD attention data. For the TD attention, we generate pixel-precise attention maps focusing on both hands, thereby retaining valuable hand information, while eliminating distracting background information. Our proposed method outperforms state-of-the-art on a challenging large-scale dataset by over 2%, and achieves strong results with a much simpler architecture compared to other systems on the newly released AUTSL dataset [1].
Noha A. Sarhan, Christian Wilms, Vanessa Closius, Ulf Brefeld, Simone Frintrop
ICIP4
2023 User Authentication via Multifaceted Mouse Movements and Outlier Exposure
Jennifer Jorina Matthiesen, Hanne Hastedt, Ulf Brefeld
IDA3
2022 Modeling Conditional Dependencies in Multiagent Trajectories
abstract
We study modeling joint densities over sets of random variables (next-step movements of multiple agents) which are conditioned on aligned observations (past trajectories). For this setting, we propose an autoregressive approach to model intra-timestep dependencies, where distributions over joint movements are represented by autoregressive factorizations. In our approach, factors are randomly ordered and estimated with a graph neural network to account for permutation equivariance, while a recurrent neural network encodes past trajectories. We further propose a conditional two-stream attention mechanism, to allow for efficient training of random factorizations. We experiment on trajectory data from professional soccer matches and find that we model low frequency trajectories better than variational approaches.
Yannick Rudolph, Ulf Brefeld
AISTATS2
2022 Semi-Supervised Generative Models for Multiagent Trajectories
abstract
Analyzing the spatiotemporal behavior of multiple agents is of great interest to many communities. Existing probabilistic models in this realm are formalized either in an unsupervised framework, where the latent space is described by discrete or continuous variables, or in a supervised framework, where weakly preserved labels add explicit information to continuous latent representations. To overcome inherent limitations, we propose a novel objective function for processing multi-agent trajectories based on semi-supervised variational autoencoders, where equivariance and interaction of agents are captured via customized graph networks. The resulting architecture disentangles discrete and continuous latent effects and provides a natural solution for injecting expensive domain knowledge into interactive sequential systems. Empirically, our model not only outperforms various state-of-the-art baselines in trajectory forecasting, but also learns to effectively leverage unsupervised multi-agent sequences for classification tasks on interactive real-world sports datasets.
Dennis Fassmeyer, Pascal Fassmeyer, Ulf Brefeld
NeurIPS3
2022 Who can receive the pass? A computational model for quantifying availability in soccer
abstract
Abstract The paper presents a computational approach to Availability of soccer players. Availability is defined as the probability that a pass reaches the target player without being intercepted by opponents. Clearly, a computational model for this probability grounds on models for ball dynamics, player movements, and technical skills of the pass giver. Our approach aggregates these quantities for all possible passes to the target player to compute a single Availability value. Empirically, our approach outperforms state-of-the-art competitors using data from 58 professional soccer matches. Moreover, our experiments indicate that the model can even outperform soccer coaches in assessing the availability of soccer players from static images.
Uwe Dick, Daniel Link 0002, Ulf Brefeld
Data Min. Knowl. Discov.3
2021 Principled Interpolation in Normalizing Flows
Samuel G. Fadel, Sebastian Mair 0001, Ricardo da Silva Torres, Ulf Brefeld
ECML/PKDD (2)4
2021 Joint optimization of an autoencoder for clustering and embedding
abstract
Abstract Deep embedded clustering has become a dominating approach to unsupervised categorization of objects with deep neural networks. The optimization of the most popular methods alternates between the training of a deep autoencoder and a k -means clustering of the autoencoder’s embedding. The diachronic setting, however, prevents the former to benefit from valuable information acquired by the latter. In this paper, we present an alternative where the autoencoder and the clustering are learned simultaneously. This is achieved by providing novel theoretical insight, where we show that the objective function of a certain class of Gaussian mixture models (GMM’s) can naturally be rephrased as the loss function of a one-hidden layer autoencoder thus inheriting the built-in clustering capabilities of the GMM. That simple neural network, referred to as the clustering module, can be integrated into a deep autoencoder resulting in a deep clustering model able to jointly learn a clustering and an embedding. Experiments confirm the equivalence between the clustering module and Gaussian mixture models. Further evaluations affirm the empirical relevance of our deep architecture as it outperforms related baselines on several data sets.
Ahcène Boubekki, Michael Kampffmeyer, Ulf Brefeld, Robert Jenssen
Mach. Learn.3
2019 Coresets for Archetypal Analysis
abstract
Archetypal analysis represents instances as linear mixtures of prototypes (the archetypes) that lie on the boundary of the convex hull of the data. Archetypes are thus often better interpretable than factors computed by other matrix factorization techniques. However, the interpretability comes with high computational cost due to additional convexity-preserving constraints. In this paper, we propose efficient coresets for archetypal analysis. Theoretical guarantees are derived by showing that quantization errors of k-means upper bound archetypal analysis; the computation of a provable absolute-coreset can be performed in only two passes over the data. Empirically, we show that the coresets lead to improved performance on several data sets.
Sebastian Mair 0001, Ulf Brefeld
NeurIPS2
2019 Probabilistic movement models and zones of control
Ulf Brefeld, Jan Lasek, Sebastian Mair 0001
Mach. Learn.1
2018 Mining User Trajectories in Electronic Text Books
Ahcène Boubekki, Shailee Jain, Ulf Brefeld
EDM3
2018 MDP-based Itinerary Recommendation using Geo-Tagged Social Media
Radhika Gaonkar, Maryam Tavakol, Ulf Brefeld
IDA3
2018 Frame-Based Optimal Design
Sebastian Mair 0001, Yannick Rudolph, Vanessa Closius, Ulf Brefeld
ECML/PKDD (2)4
2018 Distributed robust Gaussian Process regression
Sebastian Mair 0001, Ulf Brefeld
Knowl. Inf. Syst.2
2017 Frame-based Data Factorizations
abstract
Archetypal Analysis is the method of choice to compute interpretable matrix factorizations. Every data point is represented as a convex combination of factors, i.e., points on the boundary of the convex hull of the data. This renders computation inefficient. In this paper, we show that the set of vertices of a convex hull, the so-called frame, can be efficiently computed by a quadratic program. We provide theoretical and empirical results for our proposed approach and make use of the frame to accelerate Archetypal Analysis. The novel method yields similar reconstruction errors as baseline competitors but is much faster to compute.
Sebastian Mair 0001, Ahcène Boubekki, Ulf Brefeld
ICML3
2017 A Unified Contextual Bandit Framework for Long- and Short-Term Recommendations
Maryam Tavakol, Ulf Brefeld
ECML/PKDD (2)2
2017 Guest editorial: Special issue on sports analytics
Ulf Brefeld, Albrecht Zimmermann
Data Min. Knowl. Discov.1
2016 Spatio-temporal convolution kernels
Konstantin Knauf, Daniel Memmert, Ulf Brefeld
Mach. Learn.3
2015 Generalising IRT to Discriminate Between Examinees
Ahcène Boubekki, Ulf Brefeld, Thomas Delacroix
EDM2
2015 Toward Data-Driven Analyses of Electronic Text Books
Ahcène Boubekki, Ulf Kröhne, Frank Goldhammer, Waltraud Schreiber, Ulf Brefeld
EDM5
2014 An Off-the-shelf Approach to Authorship Attribution
Jamal Abdul Nasir, Nico Görnitz, Ulf Brefeld
COLING3
2014 Learning to Summarise Related Sentences
Emmanouil Tzouridis, Jamal Abdul Nasir, Ulf Brefeld
COLING3
2014 Computer-based Adaptive Speed Tests
Daniel Bengs, Ulf Brefeld
EDM2
2014 Factored MDPs for detecting topics of user sessions
abstract
Recommender systems aim to capture interests of users to provide tailored recommendations. User interests are however often unique and depend on many unobservable factors including a user's mood and the local weather. We take a contextual session-based approach and propose a sequential framework using factored Markov decision processes (fMDPs) to detect the user's goal (the topic) of a session. We show that an independence assumption on the attributes of items leads to a set of independent models that can be optimised efficiently. Our approach results in interpretable topics that can be effectively turned into recommendations. Empirical results on a real world click log from a large e-commerce company exhibit highly accurate topic prediction rates of about 90%. Translating our approach into a topic-driven recommender system outperforms several baseline competitors.
Maryam Tavakol, Ulf Brefeld
RecSys2
2013 Toward Supervised Anomaly Detection
abstract
Anomaly detection is being regarded as an unsupervised learning task as anomalies stem from adversarial or unlikely events with unknown distributions. However, the predictive performance of purely unsupervised anomaly detection often fails to match the required detection rates in many tasks and there exists a need for labeled data to guide the model generation. Our first contribution shows that classical semi-supervised approaches, originating from a supervised classifier, are inappropriate and hardly detect new and unknown anomalies. We argue that semi-supervised anomaly detection needs to ground on the unsupervised learning paradigm and devise a novel algorithm that meets this requirement. Although being intrinsically non-convex, we further show that the optimization problem has a convex equivalent under relatively mild assumptions. Additionally, we propose an active learning strategy to automatically filter candidates for labeling. In an empirical study on network intrusion detection data, we observe that the proposed learning methodology requires much less labeled data than the state-of-the-art, while achieving higher detection accuracies.
Nico Görnitz, Marius Kloft, Konrad Rieck, Ulf Brefeld
J. Artif. Intell. Res.4
2012 Discriminative clustering for market segmentation
abstract
We study discriminative clustering for market segmentation tasks. The underlying problem setting resembles discriminative clustering, however, existing approaches focus on the prediction of univariate cluster labels. By contrast, market segments encode complex (future) behavior of the individuals which cannot be represented by a single variable. In this paper, we generalize discriminative clustering to structured and complex output variables that can be represented as graphical models. We devise two novel methods to jointly learn the classifier and the clustering using alternating optimization and collapsed inference, respectively. The two approaches jointly learn a discriminative segmentation of the input space and a generative output prediction model for each segment. We evaluate our methods on segmenting user navigation sequences from Yahoo! News. The proposed collapsed algorithm is observed to outperform baseline approaches such as mixture of experts. We showcase exemplary projections of the resulting segments to display the interpretability of the solutions.
Peter Haider, Luca Chiarandini, Ulf Brefeld
KDD3
2011 Hybrid models for future event prediction
abstract
We present a hybrid method to turn off-the-shelf information retrieval (IR) systems into future event predictors. Given a query, a time series model is trained on the publication dates of the retrieved documents to capture trends and periodicity of the associated events. The periodicity of historic data is used to estimate a probabilistic model to predict future bursts. Finally, a hybrid model is obtained by intertwining the probabilistic and the time-series model. Our empirical results on the New York Times corpus show that autocorrelation functions of time-series suffice to classify queries accurately and that our hybrid models lead to more accurate future event predictions than baseline competitors.
Giuseppe Amodeo, Roi Blanco, Ulf Brefeld
CIKM3
2011 Learning to rank user intent
abstract
Personalized retrieval models aim at capturing user interests to provide personalized results that are tailored to the respective information needs. User interests are however widely spread, subject to change, and cannot always be captured well, thus rendering the deployment of personalized models challenging. We take a different approach and study ranking models for user intent. We exploit user feedback in terms of click data to cluster ranking models for historic queries according to user behavior and intent. Each cluster is finally represented by a single ranking model that captures the contained search interests expressed by users. Once new queries are issued, these are mapped to the clustering and the retrieval process diversifies possible intents by combining relevant ranking functions. Empirical evidence shows that our approach significantly outperforms baseline approaches on a large corporate query log.
Giorgos Giannopoulos, Ulf Brefeld, Theodore Dalamagas 0001, Timos K. Sellis
CIKM2
2011 Learning from Partially Annotated Sequences
Eraldo Rezende Fernandes, Ulf Brefeld
ECML/PKDD (1)2
2011 Document assignment in multi-site search engines
abstract
Assigning documents accurately to sites is critical for the performance of multi-site Web search engines. In such settings, sites crawl only documents they index and forward queries to obtain best-matching documents from other sites. Inaccurate assignments may lead to inefficiencies when crawling Web pages or processing user queries. In this work, we propose a machine-learned document assignment strategy that uses the locality of document views in search results to decide upon assignments. We evaluate the performance of our strategy using various document features extracted from a large Web collection. Our experimental setup uses query logs from a number of search front-ends spread across different geographic locations and uses these logs to learn the document access patterns. We compare our technique against baselines such as region- and language-based document assignment and observe that our technique achieves substantial performance improvements with respect to recall. With our technique, we are able to obtain a small query forwarding rate (0.04) requiring roughly 45% less replication of documents compared to replicating all documents across all sites.
Ulf Brefeld, Berkant Barla Cambazoglu, Flavio Paiva Junqueira
WSDM1
2011 lp-Norm Multiple Kernel Learning
Marius Kloft, Ulf Brefeld, Sören Sonnenburg, Alexander Zien
J. Mach. Learn. Res.2
2010 Approximate Tree Kernels
Konrad Rieck, Tammo Krueger, Ulf Brefeld, Klaus-Robert Müller
J. Mach. Learn. Res.3
2009 Efficient Classification of Images with Taxonomies
Alexander Binder, Motoaki Kawanabe, Ulf Brefeld
ACCV (3)3
2009 Efficient and Accurate Lp-Norm Multiple Kernel Learning
abstract
Learning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations and hence support interpretability. Unfortunately, L1-norm MKL is hardly observed to outperform trivial baselines in practical applications. To allow for robust kernel mixtures, we generalize MKL to arbitrary Lp-norms. We devise new insights on the connection between several existing MKL formulations and develop two efficient interleaved optimization strategies for arbitrary p>1. Empirically, we demonstrate that the interleaved optimization strategies are much faster compared to the traditionally used wrapper approaches. Finally, we apply Lp-norm MKL to real-world problems from computational biology, showing that non-sparse MKL achieves accuracies that go beyond the state-of-the-art.
Marius Kloft, Ulf Brefeld, Sören Sonnenburg, Pavel Laskov, Klaus-Robert Müller, Alexander Zien
NIPS2
2009 Active and Semi-supervised Data Domain Description
Nico Görnitz, Marius Kloft, Ulf Brefeld
ECML/PKDD (1)3
2009 Feature Selection for Density Level-Sets
Marius Kloft, Shinichi Nakajima, Ulf Brefeld
ECML/PKDD (1)3
2008 Exact and Approximate Inference for Annotating Graphs with Structural SVMs
Thoralf Klein, Ulf Brefeld, Tobias Scheffer
ECML/PKDD (1)2
2007 Supervised clustering of streaming data for email batch detection
abstract
We address the problem of detecting batches of emails that have been created according to the same template. This problem is motivated by the desire to filter spam more effectively by exploiting collective information about entire batches of jointly generated messages. The application matches the problem setting of supervised clustering, because examples of correct clusterings can be collected. Known decoding procedures for supervised clustering are cubic in the number of instances. When decisions cannot be reconsidered once they have been made --- owing to the streaming nature of the data --- then the decoding problem can be solved in linear time. We devise a sequential decoding procedure and derive the corresponding optimization problem of supervised clustering. We study the impact of collective attributes of email batches on the effectiveness of recognizing spam emails.
Peter Haider, Ulf Brefeld, Tobias Scheffer
ICML2
2007 Transductive support vector machines for structured variables
abstract
We study the problem of learning kernel machines transductively for structured output variables. Transductive learning can be reduced to combinatorial optimization problems over all possible labelings of the unlabeled data. In order to scale transductive learning to structured variables, we transform the corresponding non-convex, combinatorial, constrained optimization problems into continuous, unconstrained optimization problems. The discrete optimization parameters are eliminated and the resulting differentiable problems can be optimized efficiently. We study the effectiveness of the generalized TSVM on multiclass classification and label-sequence learning problems empirically.
Alexander Zien, Ulf Brefeld, Tobias Scheffer
ICML2
2006 Efficient co-regularised least squares regression
abstract
In many applications, unlabelled examples are inexpensive and easy to obtain. Semi-supervised approaches try to utilise such examples to reduce the predictive error. In this paper, we investigate a semi-supervised least squares regression algorithm based on the co-learning approach. Similar to other semi-supervised algorithms, our base algorithm has cubic runtime complexity in the number of unlabelled examples. To be able to handle larger sets of unlabelled examples, we devise a semi-parametric variant that scales linearly in the number of unlabelled examples. Experiments show a significant error reduction by co-regularisation and a large runtime improvement for the semi-parametric approximation. Last but not least, we propose a distributed procedure that can be applied without collecting all data at a single site.
Ulf Brefeld, Thomas Gärtner 0001, Tobias Scheffer, Stefan Wrobel
ICML1
2006 Semi-supervised learning for structured output variables
abstract
The problem of learning a mapping between input and structured, interdependent output variables covers sequential, spatial, and relational learning as well as predicting recursive structures. Joint feature representations of the input and output variables have paved the way to leveraging discriminative learners such as SVMs to this class of problems. We address the problem of semi-supervised learning in joint input output spaces. The co-training approach is based on the principle of maximizing the consensus among multiple independent hypotheses; we develop this principle into a semi-supervised support vector learning algorithm for joint input output spaces and arbitrary loss functions. Experiments investigate the benefit of semi-supervised structured models in terms of accuracy and F1 score.
Ulf Brefeld, Tobias Scheffer
ICML1
2005 Multi-view Discriminative Sequential Learning
Ulf Brefeld, Christoph Büscher, Tobias Scheffer
ECML1
2005 Systematic feature evaluation for gene name recognition
abstract
In task 1A of the BioCreAtIvE evaluation, systems had to be devised that recognize words and phrases forming gene or protein names in natural language sentences. We approach this problem by building a word classification system based on a sliding window approach with a Support Vector Machine, combined with a pattern-based post-processing for the recognition of phrases. The performance of such a system crucially depends on the type of features chosen for consideration by the classification method, such as pre- or postfixes, character n-grams, patterns of capitalization, or classification of preceding or following words. We present a systematic approach to evaluate the performance of different feature sets based on recursive feature elimination, RFE. Based on a systematic reduction of the number of features used by the system, we can quantify the impact of different feature sets on the results of the word classification problem. This helps us to identify descriptive features, to learn about the structure of the problem, and to design systems that are faster and easier to understand. We observe that the SVM is robust to redundant features. RFE improves the performance by 0.7%, compared to using the complete set of attributes. Moreover, a performance that is only 2.3% below this maximum can be obtained using fewer than 5% of the features.
Jörg Hakenberg, Steffen Bickel, Conrad Plake, Ulf Brefeld, Hagen Zahn, Lukas C. Faulstich, Ulf Leser, Tobias Scheffer
BMC Bioinform.4
2004 Co-EM support vector learning
abstract
Multi-view algorithms, such as co-training and co-EM, utilize unlabeled data when the available attributes can be split into independent and compatible subsets. Co-EM outperforms co-training for many problems, but it requires the underlying learner to estimate class probabilities, and to learn from probabilistically labeled data. Therefore, co-EM has so far only been studied with naive Bayesian learners. We cast linear classifiers into a probabilistic framework and develop a co-EM version of the Support Vector Machine. We conduct experiments on text classification problems and compare the family of semi-supervised support vector algorithms under different conditions, including violations of the assumptions underlying multi-view learning. For some problems, such as course web page classification, we observe the most accurate results reported so far.
Ulf Brefeld, Tobias Scheffer
ICML1
2004 Perceptron and SVM learning with generalized cost models
Peter Geibel, Ulf Brefeld, Fritz Wysotzki
Intell. Data Anal.2
2003 Support Vector Machines with Example Dependent Costs
Ulf Brefeld, Peter Geibel, Fritz Wysotzki
ECML1
2003 Learning Linear Classifiers Sensitive to Example Dependent and Noisy Costs
Peter Geibel, Ulf Brefeld, Fritz Wysotzki
IDA2