Benjamin Schiller

dblp:55/8733 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Computer networks · 3 · 2 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 54% Language models and text generation · 23% Optimization for machine learning · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
argument mining
1.532024
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets · EMNLP 2024
Classification and Clustering of Arguments with Contextualized Word Embeddings · ACL (1) 2019
Cross-topic Argument Mining from Heterogeneous Sources · EMNLP 2018
Natural language and speech › Language models and text generation › text summarization
argument summarization
0.912025
Argument Summarization and its Evaluation in the Era of Large Language Models · EMNLP 2025
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.912025
BoFire: Bayesian Optimization Framework Intended for Real Experiments · J. Mach. Learn. Res. 2025
Natural language and speech › Information extraction and text analysis
stance detection
0.812024
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets · EMNLP 2024
Natural language and speech › Information extraction and text analysis › argument mining
argument classification
0.412019
Classification and Clustering of Arguments with Contextualized Word Embeddings · ACL (1) 2019
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
Argument Summarization and its Evaluation in the Era of Large Language Models · EMNLP 2025
Computational science and engineering
chemistry
0.312025
BoFire: Bayesian Optimization Framework Intended for Real Experiments · J. Mach. Learn. Res. 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.212024
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

design of experiments · 1.7bayesian optimization · 1.7large language model · 0.9fine-tuning · 0.8pre-training · 0.4contextualized word embeddings · 0.4ELMo · 0.4BERT · 0.4multi-task learning · 0.3bidirectional LSTM · 0.3
YearPublicationVenuePosition
2025 Argument Summarization and its Evaluation in the Era of Large Language Models
abstract
Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Yanran Chen, Tim Altendorf, Philipp Cimiano, Benjamin Schiller. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Moritz Altemeyer, Steffen Eger, Johannes Daxenberger, Yanran Chen, Tim Altendorf, Philipp Cimiano, Benjamin Schiller
EMNLP7
2025 BoFire: Bayesian Optimization Framework Intended for Real Experiments
abstract
Our open-source Python package BoFire combines Bayesian Optimization (BO) with other design of experiments (DoE) strategies focusing on developing and optimizing new chemistry. Previous BO implementations, for example as they exist in the literature or software, require substantial adaptation for effective real-world deployment in chemical industry. BoFire provides a rich feature-set with extensive configurability and realizes our vision of fast-tracking research contributions into industrial use via maintainable open-source software. Owing to quality-of-life features like JSON-serializability of problem formulations, BoFire enables seamless integration of BO into RESTful APIs, a common architecture component for both self-driving laboratories and human-in-the-loop setups. This paper discusses the differences between BoFire and other BO implementations and outlines ways that BO research needs to be adapted for real-world use in a chemistry setting.
Johannes P. Dürholt, Thomas S. Asche, Johanna Kleinekorte, Gabriel Mancino-Ball, Benjamin Schiller, Simon Sung, Julian Keupp, Aaron Osburg, Toby Boyne, Ruth Misener, Rosona Eldred, Chrysoula Kappatou, Robert M. Lee, Dominik Linzner, Wagner Steuer Costa, David Walz, Niklas Wulkow, Behrang Shafei
J. Mach. Learn. Res.5
2024 Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets
abstract
Topic-Dependent Argument Mining (TDAM), that is extracting and classifying argument components for a specific topic from large document sources, is an inherently difficult task for machine learning models and humans alike, as large TDAM datasets are rare and recognition of argument components requires expert knowledge.The task becomes even more difficult if it also involves stance detection of retrieved arguments.In this work, we investigate the effect of TDAM dataset composition in few-and zeroshot settings.Our findings show that, while fine-tuning is mandatory to achieve acceptable model performance, using carefully composed training samples and reducing the training sample size by up to almost 90% can still yield 95% of the maximum performance.This gain is consistent across three TDAM tasks on three different datasets.We also publish a new dataset 1 and code 2 for future benchmarking.
Benjamin Schiller, Johannes Daxenberger, Andreas Waldis, Iryna Gurevych
EMNLP1
2021 Aspect-Controlled Neural Argument Generation
abstract
Benjamin Schiller, Johannes Daxenberger, Iryna Gurevych. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Benjamin Schiller, Johannes Daxenberger, Iryna Gurevych
NAACL-HLT1
2019 Classification and Clustering of Arguments with Contextualized Word Embeddings
abstract
We experiment with two recent contextualized word embedding methods (ELMo and BERT) in the context of open-domain argument search.For the first time, we show how to leverage the power of contextualized word embeddings to classify and cluster topic-dependent arguments, achieving impressive results on both tasks and across multiple datasets.For argument classification, we improve the state-of-the-art for the UKP Sentential Argument Mining Corpus by 20.8 percentage points and for the IBM Debater -Evidence Sentences dataset by 7.4 percentage points.For the understudied task of argument clustering, we propose a pre-training step which improves by 7.8 percentage points over strong baselines on a novel dataset, and by 12.3 percentage points for the Argument Facet Similarity (AFS) Corpus. 1
Nils Reimers 0001, Benjamin Schiller, Tilman Beck, Johannes Daxenberger, Christian Stab, Iryna Gurevych
ACL (1)2
2018 A Retrospective Analysis of the Fake News Challenge Stance-Detection Task
abstract
The 2017 Fake News Challenge Stage 1 (FNC-1) shared task addressed a stance classification task as a crucial first step towards detecting fake news. To date, there is no in-depth analysis paper to critically discuss FNC-1’s experimental setup, reproduce the results, and draw conclusions for next-generation stance classification methods. In this paper, we provide such an in-depth analysis for the three top-performing systems. We first find that FNC-1’s proposed evaluation metric favors the majority class, which can be easily classified, and thus overestimates the true discriminative power of the methods. Therefore, we propose a new F1-based metric yielding a changed system ranking. Next, we compare the features and architectures used, which leads to a novel feature-rich stacked LSTM model that performs on par with the best systems, but is superior in predicting minority classes. To understand the methods’ ability to generalize, we derive a new dataset and perform both in-domain and cross-domain experiments. Our qualitative and quantitative study helps interpreting the original FNC-1 scores and understand which features help improving performance and why. Our new dataset and all source code used during the reproduction study are publicly available for future research.
Andreas Hanselowski, P. V. S. Avinesh, Benjamin Schiller, Felix Caspelherr, Debanjan Chaudhuri, Christian M. Meyer, Iryna Gurevych
COLING3
2018 Cross-topic Argument Mining from Heterogeneous Sources
abstract
Argument mining is a core technology for automating argument search in large document collections. Despite its usefulness for this task, most current approaches are designed for use only with specific text types and fall short when applied to heterogeneous texts. In this paper, we propose a new sentential annotation scheme that is reliably applicable by crowd workers to arbitrary Web texts. We source annotations for over 25,000 instances covering eight controversial topics. We show that integrating topic information into bidirectional long short-term memory networks outperforms vanilla BiLSTMs by more than 3 percentage points in F1 in two- and three-label cross-topic settings. We also show that these results can be further improved by leveraging additional data for topic relevance using multi-task learning.
Christian Stab, Tristan Miller, Benjamin Schiller, Pranav Rai, Iryna Gurevych
EMNLP3
2016 StreAM- T_g : Algorithms for Analyzing Coarse Grained RNA Dynamics Based on Markov Models of Connectivity-Graphs
Sven Jager, Benjamin Schiller, Thorsten Strufe, Kay Hamacher
WABI2
2016 SWAP: Protecting pull-based P2P video streaming systems from inference attacks
abstract
In pull-based Peer-to-Peer video streaming systems, peers exchange buffer maps to reveal the availability of video chunks in their buffer. When collecting these buffer maps, a malicious party can infer the system's overlay structure and even identify head nodes, the direct communication partners of the stream's source. Attacking these head nodes can isolate peers from the source resulting in a disruption of the video dissemination for most peers in the system. We introduce a lightweight SWAP scheme, which allows peers to proactively change their partners, to reduce the chance of head nodes to be identified by such an inference attacker. Extensive simulation studies demonstrate that our scheme effectively undermines the attack's accuracy in identifying head nodes. So, SWAP lowers the chunk miss ratio while causing only a slight increase in signaling overhead.
Giang T. Nguyen 0002, Stefanie Roos, Benjamin Schiller, Thorsten Strufe
WoWMoM3
2015 Growing a Web of Trust
abstract
Web of Trust (WoT) graphs represent trust relations between people. They are used for research and analysis in various domains. Real-world instances are only available in sizes of up to 55k vertices. This renders the analysis of larger systems based on realistic input graphs impossible. To close this gap, we develop a growth model to generate WoT graphs of arbitrary size. New edges are formed based on realistic assumptions about trust establishment. We analyze the growth of a real-world WoT and perform a parameter study of our model. We compare both with many existing models and show that ours is the only one that matches the properties of the real-world WoT.
Benjamin Schiller, Thorsten Strufe, Dirk Kohlweyer, Jan Seedorf
LCN1
2014 Measuring Freenet in the Wild: Censorship-Resilience under Observation
Stefanie Roos, Benjamin Schiller, Stefan Hacker, Thorsten Strufe
Privacy Enhancing Technologies2
2013 Resilient tree-based live streaming in reality
abstract
Our main contribution in this work is a deployable multitree-push system for P2P-based live streaming. It runs on both desktop PCs and Android-based mobile devices. Additionally, it provides controlling, monitoring, and measurement functionalities which help with debugging in the development phase, visualize the topology during a demonstration, and support the deployment of test scenarios in a distributed setting. Besides, the generic architecture of the system also allows for the extension to other classes of streaming systems.
Benjamin Schiller, Giang T. Nguyen 0002, Thorsten Strufe
P2P1
2011 Native support of multi-tenancy in RDBMS for software as a service
abstract
Software as a Service (SaaS) facilitates acquiring a huge number of small tenants by providing low service fees. To achieve low service fees, it is essential to reduce costs per tenant. For this, consolidating multiple tenants onto a single relational schema instance turned out beneficial because of low overheads per tenant and scalable manageability. This approach implements data isolation between tenants, per-tenant schema extension and further tenant-centric data management features in application logic. This is complex, disables some optimization opportunities in the RDBMS and represents a conceptual misstep with Separation of Concerns in mind.
Oliver Schiller, Benjamin Schiller, Andreas Brodt, Bernhard Mitschang
EDBT2