Björn Buchhold

dblp:117/3870 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5Artificial intelligence and machine learning · 4 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Information retrieval · 55% Knowledge graphs · 25% Web and social media mining · 21%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines › semantic search
entity retrieval
0.312017
WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017
Knowledge graphs
knowledge graph quality
0.312017
WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017
Web and social media mining › content moderation
vandalism detection
0.312017
WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017
Information retrieval › ranking › relevance estimation
relevance scoring
0.212015
Relevance Scores for Triples from Type-Like Relations · SIGIR 2015
Information retrieval › search engines
semantic search
0.212014
Semantic full-text search with broccoli · SIGIR 2014
Information retrieval › search engines › semantic search › entity retrieval
entity ranking
0.112015
Relevance Scores for Triples from Type-Like Relations · SIGIR 2015
Knowledge graphs
knowledge graph querying
0.112014
Semantic full-text search with broccoli · SIGIR 2014

Methods — techniques the papers use, named apart from their topics

crowdsourcing · 0.2
YearPublicationVenuePosition
2024 XAI-Attack: Utilizing Explainable AI to Find Incorrectly Learned Patterns for Black-Box Adversarial Example Creation
abstract
Adversarial examples, capable of misleading machine learning models into making erroneous predictions, pose significant risks in safety-critical domains such as crisis informatics, medicine, and autonomous driving. To counter this, we introduce a novel textual adversarial example method that identifies falsely learned word indicators by leveraging explainable AI methods as importance functions on incorrectly predicted instances, thus revealing and understanding the weaknesses of a model. To evaluate the effectiveness of our approach, we conduct a human and a transfer evaluation and propose a novel adversarial training evaluation setting for better robustness assessment. While outperforming current adversarial example and training methods, the results also show our method’s potential in facilitating the development of more resilient transformer models by detecting and rectifying biases and patterns in training data, showing baseline improvements of up to 23 percentage points in accuracy on adversarial tasks. The code of our approach is freely available for further exploration and use.
Markus Bayer, Markus Neiczer, Maximilian Samsinger, Björn Buchhold, Christian Reuter 0001
LREC/COLING4
2017 QLever: A Query Engine for Efficient SPARQL+Text Search
abstract
We present QLever, a query engine for efficient combined search on a knowledge base and a text corpus, in which named entities from the knowledge base have been identified (that is, recognized and disambiguated). The query language is SPARQL extended by two QLever-specific predicates ql:contains-entity and ql:contains-word, which can express the occurrence of an entity or word (the object of the predicate) in a text record (the subject of the predicate). We evaluate QLever on two large datasets, including FACC (the ClueWeb12 corpus linked to Freebase). We compare against three state-of-the-art query engines for knowledge bases with varying support for text search: RDF-3X, Virtuoso, Broccoli. Query times are competitive and often faster on the pure SPARQL queries, and several orders of magnitude faster on the SPARQL+Text queries. Index size is larger for pure SPARQL queries, but smaller for SPARQL+Text queries.
Hannah Bast, Björn Buchhold
CIKM2
2017 WSDM Cup 2017: Vandalism Detection and Triple Scoring
abstract
The WSDM Cup 2017 was a data mining challenge held in conjunction with the 10th International Conference on Web Search and Data Mining (WSDM). It addressed key challenges of knowledge bases today: quality assurance and entity search. For quality assurance, we tackle the task of vandalism detection, based on a dataset of more than 82 million user-contributed revisions of the Wikidata knowledge base, all of which annotated with regard to whether or not they are vandalism. For entity search, we tackle the task of triple scoring, using a dataset that comprises relevance scores for triples from type-like relations including occupation and country of citizenship, based on about 10,000 human relevance judgments. For reproducibility sake, participants were asked to submit their software on TIRA, a cloud-based evaluation platform, and they were incentivized to share their approaches open source.
Stefan Heindorf, Martin Potthast, Hannah Bast, Björn Buchhold, Elmar Haussmann
WSDM4
2015 Relevance Scores for Triples from Type-Like Relations
abstract
We compute and evaluate relevance scores for knowledge-base triples from type-like relations. Such a score measures the degree to which an entity "belongs" to a type. For example, Quentin Tarantino has various professions, including Film Director, Screenwriter, and Actor. The first two would get a high score in our setting, because those are his main professions. The third would get a low score, because he mostly had cameo appearances in his own movies. Such scores are essential in the ranking for entity queries, e.g. "American actors" or "Quentin Tarantino professions". These scores are different from scores for "correctness" or "accuracy" (all three professions above are correct and accurate). We propose a variety of algorithms to compute these scores. For our evaluation we designed a new benchmark, which includes a ground truth based on about 14K human judgments obtained via crowdsourcing. Inter-judge agreement is slightly over 90%. Existing approaches from the literature give results far from the optimum. Our best algorithms achieve an agreement of about 80% with the ground truth.
Hannah Bast, Björn Buchhold, Elmar Haussmann
SIGIR2
2014 Semantic full-text search with broccoli
abstract
We combine search in triple stores with full-text search into what we call \emph{semantic full-text search}. We provide a fully functional web application that allows the incremental construction of complex queries on the English Wikipedia combined with the facts from Freebase. The user is guided by context-sensitive suggestions of matching words, instances, classes, and relations after each keystroke. We also provide a powerful API, which may be used for research tasks or as a back end, e.g., for a question answering system. Our web application and public API are available under \url{http://broccoli.cs.uni-freiburg.de}.
Hannah Bast, Florian Bäurle, Björn Buchhold, Elmar Haussmann
SIGIR3
2013 An index for efficient semantic full-text search
abstract
In this paper we present a novel index data structure tailored towards semantic full-text search. Semantic full-text search, as we call it, deeply integrates keyword-based full-text search with structured search in ontologies. Queries are SPARQL-like, with additional relations for specifying word-entity co-occurrences. In order to build such queries the user needs to be guided. We believe that incremental query construction with context-sensitive suggestions in every step serves that purpose well. Our index has to answer queries and provide such suggestions in real time. We achieve this through a novel kind of posting lists and query processing, avoiding very long (intermediate) result lists and expensive (non-local) operations on these lists. In an evaluation of 8000 queries on the full English Wikipedia (40 GB XML dump) and the YAGO ontology (26.6 million facts), we achieve average query and suggestion times of around 150ms.
Hannah Bast, Björn Buchhold
CIKM2