Tobias Scheffer

dblp:s/TobiasScheffer · DBLP profile ↗
← Back
90ranked-venue papers
15as first author
13since 2021 · last 2025
0000-0003-4405-7925ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 14 first-author · 6 since 2021Databases, data management, data science and information retrieval · 31 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Detection of Alcohol Inebriation from Eye Movements using Remote and Wearable Eye Trackers
abstract
This OSF contains the data for the paper 'Detection of Alcohol Inebriation from Eye Movements using Remote and Wearable Eye Trackers'. The raw data can be found in the folder raw_data (each zip file contains recorded data up- /downsampled to 1,000 Hz as csv-files). - Each csv file contains the recording (remote and wearable) for one subject for one PVT trial. - Each csv file contains the following columns: trial_id: trial-id for current recording block_id: block-id for current recording x_pix_eyelink: x-pixel coordinates using eyelink remote eye-tracker y_pix_eyelink: y-pixel coordinates using eyelink remote eye-tracker eyelink_timestamp: timestamp or recording in ms x_pix_pupilcore_interpolated: x-pixel coordinates using pupil-core eye-tracker upsampled to 1,000 Hz y_pix_pupilcore_interpolated: y-pixel coordinates using pupil-core eye-tracker upsampled to 1,000 Hz pupil_size_eyelink: pupil-size of pupil using eyelink remote eye-tracker target_distance: distance to eyelink remote eye-tracker (screen) in mm pupil_size_pupilcore_interpolated: pupil-size of pupil pupil-core eye-tracker upsampled to 1,000 Hz pupil_confidence_interpolated: pupil detection confidence of pupil pupil-core eye-tracker upsampled to 1,000 Hz time_to_prev_bac: elapsed time from previous BAC testing in ms time_to_next_bac: remaining time for next BAC testing in ms prev_bac: previous BAC concentration next_bac: next BAC concentration For more details see: https://github.com/aeye-lab/etra-potsdam-binge-pvt
Paul Prasse, David R. Reich, Jakob Chwastek, Silvia Makowski, Lena A. Jäger, Tobias Scheffer
ETRA6
2025 PanThera: predictive analysis of higher-order combination therapies using deep neural networks
abstract
This paper develops a deep neural network that accepts cell descriptors and molecules of multiple administered drugs and predicts the joint dose-response hypersurface of the combinatorial treatment. Since the dose-response hypersurface over several concentration dimensions fully characterizes the interaction dynamics of the administered drugs, the model is a computational tool that guides the discovery of synergistic treatments. The neural network is a biochemistry-informed universal approximator; it can estimate any shape of a dose-response hypersurface and has desirable invariances built into its architecture. The model excels at interpolating and extrapolating dose-response surfaces; its predictions align well with known mechanisms of action (MOA). It is the first model that can estimate joint dose-response hypersurfaces of arbitrarily many drugs, including untried combinations, in the presence of arbitrary, potentially nonlinear interactions between drugs. We release the model itself as well as a database of likely synergistic drug triplets. Our code is available at https://github.com/alonsocampana/PanThera/; the database of likely synergistic drug triplets at https://zenodo.org/records/14001717.
Pedro Alonso Campana, Paul Prasse, Ralf Herwig, Tobias Scheffer
Briefings Bioinform.4
2024 Predicting Dose-Response Curves with Deep Neural Networks
abstract
Dose-response curves characterize the relationship between the concentration of drugs and their inhibitory effect on the growth of specific types of cells. The predominant Hill-equation model of an ideal enzymatic inhibition unduly simplifies the biochemical reality of many drugs; and for these drugs the widely-used drug performance indicator of the half-inhibitory concentration $IC_{50}$ can lead to poor therapeutic recommendations and poor selections of promising drug candidates. We develop a neural model that uses an embedding of the interaction between drug molecules and the tissue transcriptome to estimate the entire dose-response curve rather than a scalar aggregate. We find that, compared to the prior state of the art, this model excels at interpolating and extrapolating the inhibitory effect of untried concentrations. Unlike prevalent parametric models, it it able to accurately predict dose-response curves of drugs on previously unseen tumor tissues as well as of previously untested drug molecules on established tumor cell lines.
Pedro Alonso Campana, Paul Prasse, Tobias Scheffer
ICML3
2024 Improving cognitive-state analysis from eye gaze with synthetic eye-movement data
abstract
Eye movements can be used to analyze a viewer’s cognitive capacities or mental state. Neural networks that process the raw eye-tracking signal can outperform methods that operate on scan paths preprocessed into fixations and saccades. However, the scarcity of such data poses a major challenge. We therefore develop SP-EyeGAN, a neural network that generates synthetic raw eye-tracking data. SP-EyeGAN consists of Generative Adversarial Networks; it produces a sequence of gaze angles indistinguishable from human ocular micro- and macro-movements. We explore the use of these synthetic eye movements for pre-training neural networks using contrastive learning. We find that pre-training on synthetic data does not help for biometric identification, while results are inconclusive for the detection of ADHD and gender classification. However, for the eye movement-based assessment of higher-level cognitive skills such general reading comprehension, text comprehension, and the distinction of native from non-native readers, pre-training on synthetic eye-gaze data improves the models’ performance and even advances the state-of-the-art for reading comprehension. The SP-EyeGAN model, pre-trained on GazeBase, along with the code for developing your own raw eye-tracking machine learning model with contrastive learning, is available at https://github.com/aeye-lab/sp-eyegan.
Paul Prasse, David R. Reich, Silvia Makowski, Tobias Scheffer, Lena A. Jäger
Comput. Graph.4
2023 Pre-Trained Language Models Augmented with Synthetic Scanpaths for Natural Language Understanding
abstract
Human gaze data offer cognitive information that reflects natural language comprehension.Indeed, augmenting language models with human scanpaths has proven beneficial for a range of NLP tasks, including language understanding.However, the applicability of this approach is hampered because the abundance of text corpora is contrasted by a scarcity of gaze data.Although models for the generation of humanlike scanpaths during reading have been developed, the potential of synthetic gaze data across NLP tasks remains largely unexplored.We develop a model that integrates synthetic scanpath generation with a scanpath-augmented language model, eliminating the need for human gaze data.Since the model's error gradient can be propagated throughout all parts of the model, the scanpath generator can be fine-tuned to downstream tasks.We find that the proposed model not only outperforms the underlying language model, but achieves a performance that is comparable to a language model augmented with real human gaze data.Our code is publicly available.1
Shuwen Deng, Paul Prasse, David R. Reich, Tobias Scheffer, Lena A. Jäger
EMNLP4
2023 Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence Models
abstract
Recent work in XAI for eye tracking data has evaluated the suitability of feature attribution methods to explain the output of deep neural sequence models for the task of oculomotric biometric identification. These methods provide saliency maps to highlight important input features of a specific eye gaze sequence. However, to date, its localization analysis has been lacking a quantitative approach across entire datasets. In this work, we employ established gaze event detection algorithms for fixations and saccades and quantitatively evaluate the impact of these events by determining their concept influence. Input features that belong to saccades are shown to be substantially more important than features that belong to fixations. By dissecting saccade events into sub-events, we are able to show that gaze samples that are close to the saccadic peak velocity are most influential. We further investigate the effect of event properties like saccadic amplitude or fixational dispersion on the resulting concept influence.
Daniel Krakowczyk, Paul Prasse, David R. Reich, Sebastian Lapuschkin, Tobias Scheffer, Lena A. Jäger
ETRA5
2023 SP-EyeGAN: Generating Synthetic Eye Movement Data with Generative Adversarial Networks
abstract
Neural networks that process the raw eye-tracking signal can outperform traditional methods that operate on scanpaths preprocessed into fixations and saccades. However, the scarcity of such data poses a major challenge. We, therefore, present SP-EyeGAN, a neural network that generates synthetic raw eye-tracking data. SP-EyeGAN consists of Generative Adversarial Networks; it produces a sequence of gaze angles indistinguishable from human micro- and macro-movements. We demonstrate how the generated synthetic data can be used to pre-train a model using contrastive learning. This model is fine-tuned on labeled human data for the task of interest. We show that for the task of predicting reading comprehension from eye movements, this approach outperforms the previous state-of-the-art.
Paul Prasse, David R. Reich, Silvia Makowski, Seoyoung Ahn, Tobias Scheffer, Lena A. Jäger
ETRA5
2023 Detection of Alcohol Inebriation from Eye Movements
abstract
In this repository we provide the Binge / PVT data set, extracted gaze features for the data set and the code to replicate the results presented in the paper "Detection of Alcohol Inebriation from Eye Movements ". The raw data set contains binocular gaze recordings and eye closure signals of 44 subjects, aged 18 to 47 with mean age 24. Each participant is recorded over 3 experimental sessions, with a time lag of at least one week between two sessions. [Entire raw data set will be added upon acceptance.]
Silvia Makowski, Annika Bätz, Paul Prasse, Lena A. Jäger, Tobias Scheffer
KES5
2023 Eyettention: An Attention-based Dual-Sequence Model for Predicting Human Scanpaths during Reading
abstract
Eye movements during reading offer insights into both the reader's cognitive processes and the characteristics of the text that is being read. Hence, the analysis of scanpaths in reading have attracted increasing attention across fields, ranging from cognitive science over linguistics to computer science. In particular, eye-tracking-while-reading data has been argued to bear the potential to make machine-learning-based language models exhibit a more human-like linguistic behavior. However, one of the main challenges in modeling human scanpaths in reading is their dual-sequence nature: the words are ordered following the grammatical rules of the language, whereas the fixations are chronologically ordered. As humans do not strictly read from left-to-right, but rather skip or refixate words and regress to previous words, the alignment of the linguistic and the temporal sequence is non-trivial. In this paper, we develop Eyettention, the first dual-sequence model that simultaneously processes the sequence of words and the chronological sequence of fixations. The alignment of the two sequences is achieved by a cross-sequence attention mechanism. We show that Eyettention outperforms state-of-the-art models in predicting scanpaths. We provide an extensive within- and across-data set evaluation on different languages. An ablation study and qualitative analysis support an in-depth understanding of the model's behavior.
Shuwen Deng, David R. Reich, Paul Prasse, Patrick Haller 0001, Tobias Scheffer, Lena A. Jäger
Proc. ACM Hum. Comput. Interact.5
2022 Fairness in Oculomotoric Biometric Identification
abstract
Gaze patterns are known to be highly individual, and therefore eye movements can serve as a biometric characteristic. We explore aspects of the fairness of biometric identification based on gaze patterns. We find that while oculomotoric identification does not favor any particular gender and does not significantly favor by age range, it is unfair with respect to ethnicity. Moreover, fairness concerning ethnicity cannot be achieved by balancing the training data for the best-performing model.
Paul Prasse, David R. Reich, Silvia Makowski, Lena A. Jäger, Tobias Scheffer
ETRA5
2022 Oculomotoric Biometric Identification under the Influence of Alcohol and Fatigue
abstract
Patterns of micro- and macro-movements of the eyes are highly individual and can serve as a biometric characteristic. It is also known that both alcohol inebriation and fatigue can reduce saccadic velocity and accuracy. This prompts the question of whether changes of gaze patterns caused by alcohol consumption and fatigue impact the accuracy of oculomotoric biometric identification. We collect an eye tracking data set from 66 participants in sober, fatigued and alcohol-intoxicated states. We find that after enrollment in a rested and sober state, identity verification based on a deep neural embedding of gaze sequences is significantly less accurate when probe sequences are taken in either an inebriated or a fatigued state. Moreover, we find that fatigue and intoxication appear to randomize gaze patterns: when the model is fine-tuned for invariance with respect to inebriation and fatigue, and even when it is trained exclusively on inebriated training person, the model still performs significantly better for sober than for sleep-deprived or intoxicated subjects.
Silvia Makowski, Paul Prasse, Lena A. Jäger, Tobias Scheffer
IJCB4
2022 Detection of ADHD Based on Eye Movements During Natural Viewing
Shuwen Deng, Paul Prasse, David R. Reich, Sabine Dziemian, Maja Stegenwallner-Schütz, Daniel Krakowczyk, Silvia Makowski, Nicolas Langer, Tobias Scheffer, Lena A. Jäger
ECML/PKDD (6)9
2021 Learning Explainable Representations of Malware Behavior
Paul Prasse, Jan Brabec, Jan Kohout, Martin Kopp, Lukás Bajer, Tobias Scheffer
ECML/PKDD (4)6
2020 Biometric Identification and Presentation-Attack Detection using Micro- and Macro-Movements of the Eyes
abstract
We study involuntary micro-movements of both eyes, in addition to saccadic macro-movements, as biometric characteristic. We develop a deep convolutional neural network that processes binocular oculomotoric signals and identifies the viewer. In order to be able to detect presentation attacks, we develop a model in which the movements are a response to a controlled stimulus. The model detects replay attacks by processing both the controlled but randomized stimulus and the ocular response to this stimulus. We acquire eye movement data from 150 participants, with 4 sessions per participant. We observe that the model detects replay attacks reliably; compared to prior work, the model attains substantially lower error rates.
Silvia Makowski, Lena A. Jäger, Paul Prasse, Tobias Scheffer
IJCB4
2020 Discriminative Viewer Identification using Generative Models of Eye Gaze
abstract
We study the problem of identifying viewers of arbitrary images based on their eye gaze. Psychological research has derived generative stochastic models of eye movements. In order to exploit this background knowledge within a discriminatively trained classification model, we derive Fisher kernels from different generative models of eye gaze. Experimentally, we find that the performance of the classifier strongly depends on the underlying generative model. Using an SVM with Fisher kernel improves the classification performance over the underlying generative model.
Silvia Makowski, Lena A. Jäger, Lisa Schwetlick, Hans Trukenbrod, Ralf Engbert, Tobias Scheffer
KES6
2020 On the Relationship between Eye Tracking Resolution and Performance of Oculomotoric Biometric Identification
abstract
Distributional properties of fixations and saccades are known to constitute biometric characteristics. Additionally, high-frequency micro-movements of the eyes have recently been found to constitute biometric characteristics that allow for faster and more robust biometric identification than just macro-movements. Micro-movements of the eyes occur on scales that are very close to the precision of currently available eye trackers. This study therefore characterizes the relationship between the temporal and spatial resolution of eye tracking recordings on one hand and the performance of a biometric identification method that processes micro-and macro-movements via a deep convolutional network. We find that that the deteriorating effects of decreasing both, the temporal and spatial resolution are not cumulative. We observe that on low-resolution data, the network reaches performance levels above chance and outperforms statistical approaches.
Paul Prasse, Lena A. Jäger, Silvia Makowski, Moritz Feuerpfeil, Tobias Scheffer
KES5
2019 Deep Eyedentification: Biometric Identification Using Micro-movements of the Eye
Lena A. Jäger, Silvia Makowski, Paul Prasse, Sascha Liehr, Maximilian Seidler, Tobias Scheffer
ECML/PKDD (2)6
2019 Joint detection of malicious domains and infected clients
Paul Prasse, René Knaebel, Lukás Machlica, Tomás Pevný, Tobias Scheffer
Mach. Learn.5
2018 Detecting Autism by Analyzing a Simulated Social Interaction
Hanna Drimalla, Niels Landwehr, Irina Baskow, Behnoush Behnia, Stefan Roepke, Isabel Dziobek, Tobias Scheffer
ECML/PKDD (1)7
2018 A Discriminative Model for Identifying Readers and Assessing Text Comprehension from Eye Movements
Silvia Makowski, Lena A. Jäger, Ahmed AbdelWahab, Niels Landwehr, Tobias Scheffer
ECML/PKDD (1)5
2017 Malware Detection by Analysing Encrypted Network Traffic with Neural Networks
Paul Prasse, Lukás Machlica, Tomás Pevný, Jirí Havelka, Tobias Scheffer
ECML/PKDD (2)5
2017 Varying-coefficient models for geospatial transfer learning
Matthias Bussas, Christoph Sawade, Nicolas Kühn, Tobias Scheffer, Niels Landwehr
Mach. Learn.4
2016 Huber-Norm Regularization for Linear Prediction Models
Oleksandr Zadorozhnyi, Gunthard Benecke, Stephan Mandt, Tobias Scheffer, Marius Kloft
ECML/PKDD (1)4
2016 Learning to control a structured-prediction decoder for detection of HTTP-layer DDoS attackers
Uwe Dick, Tobias Scheffer
Mach. Learn.2
2015 Solving Prediction Games with Parallel Batch Gradient Descent
Michael Großhans, Tobias Scheffer
ECML/PKDD (1)2
2015 Learning to identify concise regular expressions that describe email campaigns
Paul Prasse, Christoph Sawade, Niels Landwehr, Tobias Scheffer
J. Mach. Learn. Res.4
2014 A Model of Individual Differences in Gaze Control During Reading
abstract
We develop a statistical model of saccadic eye movements during reading of isolated sentences.The model is focused on representing individual differences between readers and supports the inference of the most likely reader for a novel set of eye movement patterns.We empirically study the model for biometric reader identification using eye-tracking data collected from 20 individuals and observe that the model distinguishes between 20 readers with an accuracy of up to 98%.
Niels Landwehr, Sebastian Arzt, Tobias Scheffer, Reinhold Kliegl
EMNLP3
2014 Joint Prediction of Topics in a URL Hierarchy
Michael Großhans, Christoph Sawade, Tobias Scheffer, Niels Landwehr
ECML/PKDD (1)3
2013 Lagrangian Strain Tensor Computation with Higher Order Variational Models
abstract
The reliable estimation of the Lagrangian stress tensor from an image sequence is a challenging problem in mechanical engineering. Since this tensor involves first order motion derivatives, it appears tempting to estimate the optical flow field with a highly accurate variational model and compute its derivatives afterwards. In this paper we explain why this idea is inappropriate due to lower order smoothness assumptions and the ill-posedness of differentiation. As a remedy, we propose a variational framework that performs higher order regularisation of the optical flow field and directly computes the Lagrangian stress tensor from the image measurements. Due to its recursive structure, this framework is very generic. It can incorporate smoothness assumptions of arbitrary high order and allows to compute derivatives of any desired order in a stable way. With a biaxial tensile experiment with an elastomer we demonstrate that our novel approach gives substantially better results for the Lagrangian stress tensor than computing derivatives of the optical flow field. Moreover, it also outperforms a frequently used commercial software that marks the state-of-the-art for Lagrangian stress tensor computation.
Alexander Hewer, Joachim Weickert, Henning Seibert, Tobias Scheffer, Stefan Diebels
BMVC4
2013 Bayesian Games for Adversarial Regression Problems
abstract
We study regression problems in which an adversary can exercise some control over the data generation process. Learner and adversary have conflicting but not necessarily perfectly antagonistic objectives. We study the case in which the learner is not fully informed about the adversary’s objective; instead, any knowledge of the learner about parameters of the adversary’s goal may be reflected in a Bayesian prior. We model this problem as a Bayesian game, and characterize conditions under which a unique Bayesian equilibrium point exists. We experimentally compare the Bayesian equilibrium strategy to the Nash equilibrium strategy, the minimax strategy, and regular linear regression.
Michael Großhans, Christoph Sawade, Michael Brückner, Tobias Scheffer
ICML (3)4
2013 Active Evaluation of Ranking Functions Based on Graded Relevance (Extended Abstract)
Christoph Sawade, Steffen Bickel, Timo von Oertzen, Tobias Scheffer, Niels Landwehr
IJCAI4
2013 Active evaluation of ranking functions based on graded relevance
Christoph Sawade, Steffen Bickel, Timo von Oertzen, Tobias Scheffer, Niels Landwehr
Mach. Learn.4
2012 Finding Botnets Using Minimal Graph Clusterings
Peter Haider, Tobias Scheffer
ICML2
2012 Learning to Identify Regular Expressions that Describe Email Campaigns
Paul Prasse, Christoph Sawade, Niels Landwehr, Tobias Scheffer
ICML4
2012 Active Comparison of Prediction Models
abstract
We address the problem of comparing the risks of two given predictive models - for instance, a baseline model and a challenger - as confidently as possible on a fixed labeling budget. This problem occurs whenever models cannot be compared on held-out training data, possibly because the training data are unavailable or do not reflect the desired test distribution. In this case, new test instances have to be drawn and labeled at a cost. We devise an active comparison method that selects instances according to an instrumental sampling distribution. We derive the sampling distribution that maximizes the power of a statistical test applied to the observed empirical risks, and thereby minimizes the likelihood of choosing the inferior model. Empirically, we investigate model selection problems on several classification and regression tasks and study the accuracy of the resulting p-values.
Christoph Sawade, Niels Landwehr, Tobias Scheffer
NIPS3
2012 Active Evaluation of Ranking Functions Based on Graded Relevance
Christoph Sawade, Steffen Bickel, Timo von Oertzen, Tobias Scheffer, Niels Landwehr
ECML/PKDD (2)4
2012 Static prediction games for adversarial learning problems
Michael Brückner, Christian Kanzow, Tobias Scheffer
J. Mach. Learn. Res.3
2011 Stackelberg games for adversarial prediction problems
abstract
The standard assumption of identically distributed training and test data is violated when test data are generated in response to a predictive model. This becomes apparent, for example, in the context of email spam filtering, where an email service provider employs a spam filter and the spam sender can take this filter into account when generating new emails. We model the interaction between learner and data generator as a Stackelberg competition in which the learner plays the role of the leader and the data generator may react on the leader's move. We derive an optimization problem to determine the solution of this game and present several instances of the Stackelberg prediction game. We show that the Stackelberg prediction game generalizes existing prediction models. Finally, we explore properties of the discussed models empirically in the context of email spam filtering.
Michael Brückner, Tobias Scheffer
KDD2
2010 Active Risk Estimation
Christoph Sawade, Niels Landwehr, Steffen Bickel, Tobias Scheffer
ICML4
2010 Throttling Poisson Processes
abstract
We study a setting in which Poisson processes generate sequences of decision-making events. The optimization goal is allowed to depend on the rate of decision outcomes; the rate may depend on a potentially long backlog of events and decisions. We model the problem as a Poisson process with a throttling policy that enforces a data-dependent rate limit and reduce the learning problem to a convex optimization problem that can be solved efficiently. This problem setting matches applications in which damage caused by an attacker grows as a function of the rate of unsuppressed hostile events. We report on experiments on abuse detection for an email service.
Uwe Dick, Peter Haider, Thomas Vanck, Michael Brückner, Tobias Scheffer
NIPS5
2010 Active Estimation of F-Measures
abstract
We address the problem of estimating the F-measure of a given model as accurately as possible on a fixed labeling budget. This problem occurs whenever an estimate cannot be obtained from held-out training data; for instance, when data that have been used to train the model are held back for reasons of privacy or do not reflect the test distribution. In this case, new test instances have to be drawn and labeled at a cost. An active estimation procedure selects instances according to an instrumental sampling distribution. An analysis of the sources of estimation error leads to an optimal sampling distribution that minimizes estimator variance. We explore conditions under which active estimates of F-measures are more accurate than estimates based on instances sampled from the test distribution.
Christoph Sawade, Niels Landwehr, Tobias Scheffer
NIPS3
2009 Bayesian clustering for email campaign detection
abstract
We discuss the problem of clustering elements according to the sources that have generated them. For elements that are characterized by independent binary attributes, a closed-form Bayesian solution exists. We derive a solution for the case of dependent attributes that is based on a transformation of the instances into a space of independent feature functions. We derive an optimization problem that produces a mapping into a space of independent binary feature vectors; the features can reflect arbitrary dependencies in the input space. This problem setting is motivated by the application of spam filtering for email service providers. Spam traps deliver a real-time stream of messages known to be spam. If elements of the same campaign can be recognized reliably, entire spam and phishing campaigns can be contained. We present a case study that evaluates Bayesian clustering for this application.
Peter Haider, Tobias Scheffer
ICML2
2009 Nash Equilibria of Static Prediction Games
abstract
The standard assumption of identically distributed training and test data can be violated when an adversary can exercise some control over the generation of the test data. In a prediction game, a learner produces a predictive model while an adversary may alter the distribution of input data. We study single-shot prediction games in which the cost functions of learner and adversary are not necessarily antagonistic. We identify conditions under which the prediction game has a unique Nash equilibrium, and derive algorithms that will find the equilibrial prediction models. In a case study, we explore properties of Nash-equilibrial prediction models for email spam filtering empirically.
Michael Brückner, Tobias Scheffer
NIPS2
2009 Localizing Bugs in Program Executions with Graphical Models
abstract
We devise a graphical model that supports the process of debugging software by guiding developers to code that is likely to contain defects. The model is trained using execution traces of passing test runs; it reflects the distribution over transitional patterns of code positions. Given a failing test case, the model determines the least likely transitional pattern in the execution trace. The model is designed such that Bayesian inference has a closed-form solution. We evaluate the Bernoulli graph model on data of the software projects AspectJ and Rhino.
Laura Dietz, Valentin Dallmeier, Andreas Zeller, Tobias Scheffer
NIPS4
2009 Scalable pattern mining with Bayesian networks as background knowledge
abstract
We study a discovery framework in which background knowledge on variables and their relations within a discourse area is available in the form of a graphical model. Starting from an initial, hand-crafted or possibly empty graphical model, the network evolves in an interactive process of discovery. We focus on the central step of this process: given a graphical model and a database, we address the problem of finding the most interesting attribute sets. We formalize the concept of interestingness of attribute sets as the divergence between their behavior as observed in the data, and the behavior that can be explained given the current model. We derive an exact algorithm that finds all attribute sets whose interestingness exceeds a given threshold. We then consider the case of a very large network that renders exact inference unfeasible, and a very large database or data stream. We devise an algorithm that efficiently finds the most interesting attribute sets with prescribed approximation bound and confidence probability, even for very large networks and infinite streams. We study the scalability of the methods in controlled experiments; a case-study sheds light on the practical usefulness of the approach.
Szymon Jaroszewicz, Tobias Scheffer, Dan A. Simovici
Data Min. Knowl. Discov.2
2009 Discriminative Learning Under Covariate Shift
Steffen Bickel, Michael Brückner, Tobias Scheffer
J. Mach. Learn. Res.3
2008 Multi-task learning for HIV therapy screening
abstract
We address the problem of learning classifiers for a large number of tasks. We derive a solution that produces resampling weights which match the pool of all examples to the target distribution of any given task. Our work is motivated by the problem of predicting the outcome of a therapy attempt for a patient who carries an HIV virus with a set of observed genetic properties. Such predictions need to be made for hundreds of possible combinations of drugs, some of which use similar biochemical mechanisms. Multi-task learning enables us to make predictions even for drug combinations with few or no training examples and substantially improves the overall prediction accuracy.
Steffen Bickel, Jasmina Bogojeska, Thomas Lengauer, Tobias Scheffer
ICML4
2008 Learning from incomplete data with infinite imputations
abstract
We address the problem of learning decision functions from training data in which some attribute values are unobserved. This problem can arise, for instance, when training data is aggregated from multiple sources, and some sources record only a subset of attributes. We derive a generic joint optimization problem in which the distribution governing the missing values is a free parameter. We show that the optimal solution concentrates the density mass on finitely many imputations, and provide a corresponding algorithm for learning from incomplete data. We report on empirical results on benchmark data, and on the email spam application that motivates our work. 1.
Uwe Dick, Peter Haider, Tobias Scheffer
ICML3
2008 Transfer Learning by Distribution Matching for Targeted Advertising
abstract
We address the problem of learning classifiers for several related tasks that may differ in their joint distribution of input and output variables. For each task, small - possibly even empty - labeled samples and large unlabeled samples are available. While the unlabeled samples reflect the target distribution, the labeled samples may be biased. We derive a solution that produces resampling weights which match the pool of all examples to the target distribution of any given task. Our work is motivated by the problem of predicting sociodemographic features for users of web portals, based on the content which they have accessed. Here, questionnaires offered to a small portion of each portal's users produce biased samples. Transfer learning enables us to make predictions even for new portals with few or no training data and improves the overall prediction accuracy.
Steffen Bickel, Christoph Sawade, Tobias Scheffer
NIPS3
2008 Exact and Approximate Inference for Annotating Graphs with Structural SVMs
Thoralf Klein, Ulf Brefeld, Tobias Scheffer
ECML/PKDD (1)3
2008 Schema matching on streams with accuracy guarantees
Szymon Jaroszewicz, Lenka Ivantysynova, Tobias Scheffer
Intell. Data Anal.3
2007 Discriminative learning for differing training and test distributions
abstract
We address classification problems for which the training instances are governed by a distribution that is allowed to differ arbitrarily from the test distribution---problems also referred to as classification under covariate shift. We derive a solution that is purely discriminative: neither training nor test distribution are modeled explicitly. We formulate the general problem of learning under covariate shift as an integrated optimization problem. We derive a kernel logistic regression classifier for differing training and test distributions.
Steffen Bickel, Michael Brückner, Tobias Scheffer
ICML3
2007 Unsupervised prediction of citation influences
abstract
Abstract Publication repositories contain an abundance of information about the evolution of scientific research areas. We address the problem of creating a visualization of a research area that describes the flow of topics between papers, quantifies the impact that papers have on each other, and helps to identify key contributions. To this end, we devise a probabilistic topic model that explains the generation of documents; the model incorporates the aspects of topical innovation and topical inheritance via citations. We evaluate the model's ability to predict the strength of influence of citations against manually rated citations.
Laura Dietz, Steffen Bickel, Tobias Scheffer
ICML3
2007 Supervised clustering of streaming data for email batch detection
abstract
We address the problem of detecting batches of emails that have been created according to the same template. This problem is motivated by the desire to filter spam more effectively by exploiting collective information about entire batches of jointly generated messages. The application matches the problem setting of supervised clustering, because examples of correct clusterings can be collected. Known decoding procedures for supervised clustering are cubic in the number of instances. When decisions cannot be reconsidered once they have been made --- owing to the streaming nature of the data --- then the decoding problem can be solved in linear time. We devise a sequential decoding procedure and derive the corresponding optimization problem of supervised clustering. We study the impact of collective attributes of email batches on the effectiveness of recognizing spam emails.
Peter Haider, Ulf Brefeld, Tobias Scheffer
ICML3
2007 Transductive support vector machines for structured variables
abstract
We study the problem of learning kernel machines transductively for structured output variables. Transductive learning can be reduced to combinatorial optimization problems over all possible labelings of the unlabeled data. In order to scale transductive learning to structured variables, we transform the corresponding non-convex, combinatorial, constrained optimization problems into continuous, unconstrained optimization problems. The discrete optimization parameters are eliminated and the resulting differentiable problems can be optimized efficiently. We study the effectiveness of the generalized TSVM on multiclass classification and label-sequence learning problems empirically.
Alexander Zien, Ulf Brefeld, Tobias Scheffer
ICML3
2007 Scalable look-ahead linear regression trees
abstract
Most decision tree algorithms base their splitting decisions on a piecewise constant model. Often these splitting algorithms are extrapolated to trees with non-constant models at the leaf nodes. The motivation behind Look-ahead Linear Regression Trees (LLRT) is that out of all the methods proposed to date, there has been no scalable approach to exhaustively evaluate all possible models in the leaf nodes in order to obtain an optimal split. Using several optimizations, LLRT is able to generate and evaluate thousands of linear regression models per second. This allows for a near-exhaustive evaluation of all possible splits in a node, based on the quality of fit of linear regression models in the resulting branches. We decompose the calculation of the Residual Sum of Squares in such a way that a large part of it is pre-computed. The resulting method is highly scalable. We observe it to obtain high predictive accuracy for problems with strong mutual dependencies between attributes. We report on experiments with two simulated and seven real data sets.
David S. Vogel, Ognian Asparouhov, Tobias Scheffer
KDD3
2006 Efficient co-regularised least squares regression
abstract
In many applications, unlabelled examples are inexpensive and easy to obtain. Semi-supervised approaches try to utilise such examples to reduce the predictive error. In this paper, we investigate a semi-supervised least squares regression algorithm based on the co-learning approach. Similar to other semi-supervised algorithms, our base algorithm has cubic runtime complexity in the number of unlabelled examples. To be able to handle larger sets of unlabelled examples, we devise a semi-parametric variant that scales linearly in the number of unlabelled examples. Experiments show a significant error reduction by co-regularisation and a large runtime improvement for the semi-parametric approximation. Last but not least, we propose a distributed procedure that can be applied without collecting all data at a single site.
Ulf Brefeld, Thomas Gärtner 0001, Tobias Scheffer, Stefan Wrobel
ICML3
2006 Semi-supervised learning for structured output variables
abstract
The problem of learning a mapping between input and structured, interdependent output variables covers sequential, spatial, and relational learning as well as predicting recursive structures. Joint feature representations of the input and output variables have paved the way to leveraging discriminative learners such as SVMs to this class of problems. We address the problem of semi-supervised learning in joint input output spaces. The co-training approach is based on the principle of maximizing the consensus among multiple independent hypotheses; we develop this principle into a semi-supervised support vector learning algorithm for joint input output spaces and arbitrary loss functions. Experiments investigate the benefit of semi-supervised structured models in terms of accuracy and F1 score.
Ulf Brefeld, Tobias Scheffer
ICML2
2006 Dirichlet-Enhanced Spam Filtering based on Biased Samples
abstract
We study a setting that is motivated by the problem of filtering spam messages for many users. Each user receives messages according to an individual, unknown distribution, reflected only in the unlabeled inbox. The spam filter for a user is required to perform well with respect to this distribution. Labeled messages from publicly available sources can be utilized, but they are governed by a distinct distribution, not adequately representing most inboxes. We devise a method that minimizes a loss function with respect to a user's personal distribution based on the available biased sample. A nonparametric hierarchical Bayesian model furthermore generalizes across users by learning a common prior which is imposed on new email accounts. Empirically, we observe that bias-corrected learning outperforms naive reliance on the assumption of independent and identically distributed data; Dirichlet-enhanced generalization across users outperforms a single ("one size fits all") filter as well as independent filters for all users.
Steffen Bickel, Tobias Scheffer
NIPS2
2005 Learning to Complete Sentences
Steffen Bickel, Peter Haider, Tobias Scheffer
ECML3
2005 Estimation of Mixture Models Using Co-EM
Steffen Bickel, Tobias Scheffer
ECML2
2005 Multi-view Discriminative Sequential Learning
Ulf Brefeld, Christoph Büscher, Tobias Scheffer
ECML3
2005 Thwarting the Nigritude Ultramarine: Learning to Identify Link Spam
Isabel Drost-Fromm, Tobias Scheffer
ECML2
2005 Fast discovery of unexpected patterns in data, relative to a Bayesian network
abstract
We consider a model in which background knowledge on a given domain of interest is available in terms of a Bayesian network, in addition to a large database. The mining problem is to discover unexpected patterns: our goal is to find the strongest discrepancies between network and database. This problem is intrinsically difficult because it requires inference in a Bayesian network and processing the entire, potentially very large, database. A sampling-based method that we introduce is efficient and yet provably finds the approximately most interesting unexpected patterns. We give a rigorous proof of the method's correctness. Experiments shed light on its efficiency and practicality for large-scale Bayesian networks and databases.
Szymon Jaroszewicz, Tobias Scheffer
KDD2
2005 Systematic feature evaluation for gene name recognition
abstract
In task 1A of the BioCreAtIvE evaluation, systems had to be devised that recognize words and phrases forming gene or protein names in natural language sentences. We approach this problem by building a word classification system based on a sliding window approach with a Support Vector Machine, combined with a pattern-based post-processing for the recognition of phrases. The performance of such a system crucially depends on the type of features chosen for consideration by the classification method, such as pre- or postfixes, character n-grams, patterns of capitalization, or classification of preceding or following words. We present a systematic approach to evaluate the performance of different feature sets based on recursive feature elimination, RFE. Based on a systematic reduction of the number of features used by the system, we can quantify the impact of different feature sets on the results of the word classification problem. This helps us to identify descriptive features, to learn about the structure of the problem, and to design systems that are faster and easier to understand. We observe that the SVM is robust to redundant features. RFE improves the performance by 0.7%, compared to using the complete set of attributes. Moreover, a performance that is only 2.3% below this maximum can be obtained using fewer than 5% of the features.
Jörg Hakenberg, Steffen Bickel, Conrad Plake, Ulf Brefeld, Hagen Zahn, Lukas C. Faulstich, Ulf Leser, Tobias Scheffer
BMC Bioinform.8
2005 Finding association rules that trade support optimally against confidence
Tobias Scheffer
Intell. Data Anal.1
2004 Learning from Message Pairs for Automatic Email Answering
Steffen Bickel, Tobias Scheffer
ECML2
2004 Multi-View Clustering
abstract
We consider clustering problems in which the available attributes can be split into two independent subsets, such that either subset suffices for learning. Example applications of this multi-view setting include clustering of Web pages which have an intrinsic view (the pages themselves) and an extrinsic view (e.g., anchor texts of inbound hyperlinks); multi-view learning has so far been studied in the context of classification. We develop and study partitioning and agglomerative, hierarchical multi-view clustering algorithms for text data. We find empirically that the multi-view versions of k-means and EM greatly improve on their single-view counterparts. By contrast, we obtain negative results for agglomerative hierarchical multi-view clustering. Our analysis explains this surprising phenomenon.
Steffen Bickel, Tobias Scheffer
ICDM2
2004 Co-EM support vector learning
abstract
Multi-view algorithms, such as co-training and co-EM, utilize unlabeled data when the available attributes can be split into independent and compatible subsets. Co-EM outperforms co-training for many problems, but it requires the underlying learner to estimate class probabilities, and to learn from probabilistically labeled data. Therefore, co-EM has so far only been studied with naive Bayesian learners. We cast linear classifiers into a probabilistic framework and develop a co-EM version of the Support Vector Machine. We conduct experiments on text classification problems and compare the family of semi-supervised support vector algorithms under different conditions, including violations of the assumptions underlying multi-view learning. For some problems, such as course web page classification, we observe the most accurate results reported so far.
Ulf Brefeld, Tobias Scheffer
ICML2
2004 Sentence completion
abstract
We discuss a retrieval model in which the task is to complete a sentence, given an initial fragment, and given an application specific document collection. This model is motivated by administrative and call center environments, in which users have to write documents with a certain repetitiveness. We formulate the problem setting and discuss appropriate performance metrics. We present an index-based retrieval algorithm and a cluster-based approach, and evaluate our algorithms using collections of emails that have been written by two distinct service centers.
Korinna Bade, Tobias Scheffer
SIGIR2
2004 Email answering assistance by semi-supervised text classification
Tobias Scheffer
Intell. Data Anal.1
2004 Multi-Relational Learning, Text Mining, and Semi-Supervised Learning for Functional Genomics
Mark-A. Krogel, Tobias Scheffer
Mach. Learn.2
2003 Effectiveness of Information Extraction, Multi-Relational, and Semi-Supervised Learning for Predicting Functional Properties of Genes
abstract
We focus on the problem of predicting functional properties of the proteins corresponding to genes in the yeast genome. Our goal is to study the effectiveness of approaches that utilize all data sources that are available in this problem setting, including unlabeled and relational data, and abstracts of research papers. We study transduction and co-training for using unlabeled data. We investigate a propositionalization approach which uses relational gene interaction data. We study the benefit of information extraction for utilizing a collection of scientific abstracts. The studied tasks are KDD Cup tasks of 2001 and 2002. The solutions which we describe achieved the highest score for task 2 in 2001, the fourth rank for task 3 in 2001, the highest score for one of the two subtasks and the third place for the overall task 2 in 2002.
Mark-A. Krogel, Tobias Scheffer
ICDM2
2003 Learning to Answer Emails
Michael Kockelkorn, Andreas Lüneburg, Tobias Scheffer
IDA3
2003 Using Transduction and Multi-view Learning to Answer Emails
Michael Kockelkorn, Andreas Lüneburg, Tobias Scheffer
PKDD3
2002 A Scalable Constant-Memory Sampling Algorithm for Pattern Discovery in Large Databases
Tobias Scheffer, Stefan Wrobel
PKDD1
2002 Finding the Most Interesting Patterns in a Database Quickly by Using Sequential Sampling
Tobias Scheffer, Stefan Wrobel
J. Mach. Learn. Res.1
2001 Clipping and Analyzing News Using Machine Learning Techniques
Hans Gründel, Tino Naphtali, Christian Wiech, Jan-Marian Gluba, Maiken Rohdenburg, Tobias Scheffer
Discovery Science6
2001 Mining the Web with Active Hidden Markov Models
abstract
Given the enormous amounts of information available only in unstructured or semi-structured textual documents, tools for information extraction (IE) have become enormously important. IE tools identify the relevant information in such documents and convert it into a structured format such as a database or an XML document. While first IE algorithms were hand-crafted sets of rules, researchers soon turned to learning extraction rules from hand-labeled documents. Unfortunately, rule-based approaches sometimes fail to provide the necessary robustness against the inherent variability of document, structure, which has led to the recent interest in using hidden Markov models (HMMs). By using additional unlabeled documents as they are usually readily available in most applications, we can perform active learning of HMMs. The idea of active learning algorithms is to identify unlabeled observations that would be most useful when labeled by the user. Such algorithms are known for classification, clustering, and regression; we present the first algorithm for active learning of hidden Markov models.
Tobias Scheffer, Christian Decomain, Stefan Wrobel
ICDM1
2001 Incremental Maximization of Non-Instance-Averaging Utility Functions with Applications to Knowledge Discovery Problems
Tobias Scheffer, Stefan Wrobel
ICML1
2001 Active Hidden Markov Models for Information Extraction
Tobias Scheffer, Christian Decomain, Stefan Wrobel
IDA1
2001 Finding Association Rules That Trade Support Optimally against Confidence
Tobias Scheffer
PKDD1
2000 Average-Case Analysis of Classification Algorithms for Boolean Functions and Decision Trees
Tobias Scheffer
ALT1
2000 Nonparametric Regularization of Decision Trees
Tobias Scheffer
ECML1
2000 Predicting the Generalization Performance of Cross Validatory Model Selection Criteria
Tobias Scheffer
ICML1
2000 A sequential sampling algorithm for a general class of utility criteria
abstract
Many discovery problems, e.g., subgroup or association rule discovery, can naturally be cast as n-best hypothesis problems where the goal is to nd the n hypotheses from a given hypothesis space that score best according to a given utility function. We present a sampling algorithm that solves this problem by issuing a small number of database queries while guaranteeing precise bounds on condence and quality of solutions. Known sampling algorithms assume that the utility be the average (over the examples) of some function, which is not the case for many frequently used utility functions. We show that our algorithm works for all utilities that can be estimated with bounded error. We provide such error bounds and resulting worst-case sample bounds for some of the most frequently used utilities, and prove that there is no sampling algorithm for another popular class of utility functions. The algorithm is sequential in the sense that it starts to return (or discard) hypotheses that already...
Tobias Scheffer, Stefan Wrobel
KDD1
1999 The VC-Dimension of Subclasses of Pattern
Andrew R. Mitchell, Tobias Scheffer, Arun Sharma 0001, Frank Stephan 0001
ALT2
1999 Expected Error Analysis for Model Selection
Tobias Scheffer, Thorsten Joachims
ICML1
1997 Why Experimentation can be better than "Perfect Guidance"
Tobias Scheffer, Russell Greiner, Christian J. Darken
ICML1
1997 Unbiased Assesment of Learning Algorithms
Tobias Scheffer, Ralf Herbrich
IJCAI (2)1