EDBT 2026 Demo / reviewers in the wild / expert
Elmar Haussmann
dblp:81/8216 · also Elmar Haußmann
· DBLP profile ↗
10ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 55% Knowledge graphs · 25% Web and social media mining · 21% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › search engines › semantic search
entity retrieval |
0.3 | 1 | 2017 | WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017 |
Knowledge graphs
knowledge graph quality |
0.3 | 1 | 2017 | WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017 |
Web and social media mining › content moderation
vandalism detection |
0.3 | 1 | 2017 | WSDM Cup 2017: Vandalism Detection and Triple Scoring · WSDM 2017 |
Information retrieval › ranking › relevance estimation
relevance scoring |
0.2 | 1 | 2015 | Relevance Scores for Triples from Type-Like Relations · SIGIR 2015 |
Information retrieval › search engines
semantic search |
0.2 | 1 | 2014 | Semantic full-text search with broccoli · SIGIR 2014 |
Information retrieval › search engines › semantic search › entity retrieval
entity ranking |
0.1 | 1 | 2015 | Relevance Scores for Triples from Type-Like Relations · SIGIR 2015 |
Knowledge graphs
knowledge graph querying |
0.1 | 1 | 2014 | Semantic full-text search with broccoli · SIGIR 2014 |
Methods — techniques the papers use, named apart from their topics
crowdsourcing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Object-Level Targeted Selection via Deep Template MatchingabstractRetrieving images with objects that are semantically similar to objects of interest (OOI) in a query image has many practical use cases. A few examples include fixing failures like false negatives/positives of a learned model or mitigating class imbalance in a dataset. The targeted selection task requires finding the relevant data from a large-scale pool of unlabeled data. Manual mining at this scale is infeasible. Further, the OOI are often small and occupy less than 1% of image area, are occluded, and co-exist with many semantically different objects in cluttered scenes. Existing semantic image retrieval methods often focus on mining for larger sized geographical landmarks, and/or require extra labeled data, such as images/image-pairs with similar objects, for mining images with generic objects. We propose a fast and robust template matching algorithm in the DNN feature space, that retrieves semantically similar images at the object-level from a large unlabeled pool of data. We project the region(s) around the OOI in the query image to the DNN feature space for use as the template. This enables our method to focus on the semantics of the OOI without requiring extra labeled data. In the context of autonomous driving, we evaluate our system for targeted selection by using failure cases of object detectors as OOI. We demonstrate its efficacy on a large unlabeled dataset with 2.2M images and show high recall in mining for images with small-sized OOI. We compare our method against a well-known semantic image retrieval method, which also does not require extra labeled data. Lastly, we show that our method is flexible and retrieves images with one or more semantically different co-occurring OOI seamlessly. Suraj Kothawade, Donna Roy, Michele Fenzi, Elmar Haussmann, José M. Álvarez 0004, Christoph Angerer |
IV | 4 |
| 2022 | Training Data Subset Search With Ensemble Active LearningabstractDeep Neural Networks (DNNs) often rely on vast datasets for training. Given the large size of such datasets, it is conceivable that they contain specific samples that either do not contribute or negatively impact the DNN’s optimization. Modifying the training distribution to exclude such samples could provide an effective solution to improve performance and reduce training time. This paper proposes to scale up ensemble Active Learning (AL) methods to perform acquisition at a large scale (10k to 500k samples at a time). We do this with ensembles of hundreds of models, obtained at a minimal computational cost by reusing intermediate training checkpoints. This allows us to automatically and efficiently perform a training data subset search for large labeled datasets. We observe that our approach obtains favorable subsets of training data, which can be used to train more accurate DNNs than training with the entire dataset. We perform an extensive experimental study of this phenomenon on three image classification benchmarks (CIFAR-10, CIFAR-100, and ImageNet), as well as an internal object detection benchmark for prototyping perception models for autonomous driving. Unlike existing studies, our experiments on object detection are at the scale required for production-ready autonomous driving systems. We provide insights on the impact of different initialization schemes, acquisition functions, and ensemble configurations at this scale. Our results provide strong empirical evidence that optimizing the training data distribution can significantly benefit large-scale vision tasks. Kashyap Chitta, José M. Álvarez 0004, Elmar Haussmann, Clément Farabet |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Scalable Active Learning for Object DetectionabstractDeep Neural Networks trained in a fully supervised fashion are the dominant technology in perception-based autonomous driving systems. While collecting large amounts of unlabeled data is already a major undertaking, only a subset of it can be labeled by humans due to the effort needed for high-quality annotation. Therefore, finding the right data to label has become a key challenge. Active learning is a powerful technique to improve data efficiency for supervised learning methods, as it aims at selecting the smallest possible training set to reach a required performance. We have built a scalable production system for active learning in the domain of autonomous driving. In this paper, we describe the resulting high-level design, sketch some of the challenges and their solutions, present our current results at scale, and briefly describe the open problems and future directions. Elmar Haussmann, Michele Fenzi, Kashyap Chitta, Jan Ivanecky, Hanson Xu, Donna Roy, Akshita Mittel, Nicolas Koumchatzky, Clément Farabet, José M. Álvarez 0004 |
IV | 1 |
| 2017 | WSDM Cup 2017: Vandalism Detection and Triple ScoringabstractThe WSDM Cup 2017 was a data mining challenge held in conjunction with the 10th International Conference on Web Search and Data Mining (WSDM). It addressed key challenges of knowledge bases today: quality assurance and entity search. For quality assurance, we tackle the task of vandalism detection, based on a dataset of more than 82 million user-contributed revisions of the Wikidata knowledge base, all of which annotated with regard to whether or not they are vandalism. For entity search, we tackle the task of triple scoring, using a dataset that comprises relevance scores for triples from type-like relations including occupation and country of citizenship, based on about 10,000 human relevance judgments. For reproducibility sake, participants were asked to submit their software on TIRA, a cloud-based evaluation platform, and they were incentivized to share their approaches open source. Stefan Heindorf, Martin Potthast, Hannah Bast, Björn Buchhold, Elmar Haussmann |
WSDM | 5 |
| 2015 | More Accurate Question Answering on FreebaseabstractReal-world factoid or list questions often have a simple structure, yet are hard to match to facts in a given knowledge base due to high representational and linguistic variability. For example, to answer "who is the ceo of apple" on Freebase requires a match to an abstract "leadership" entity with three relations "role", "organization" and "person", and two other entities "apple inc" and "managing director". Recent years have seen a surge of research activity on learning-based solutions for this method. We further advance the state of the art by adopting learning-to-rank methodology and by fully addressing the inherent entity recognition problem, which was neglected in recent works. Hannah Bast, Elmar Haussmann |
CIKM | 2 |
| 2015 | Relevance Scores for Triples from Type-Like RelationsabstractWe compute and evaluate relevance scores for knowledge-base triples from type-like relations. Such a score measures the degree to which an entity "belongs" to a type. For example, Quentin Tarantino has various professions, including Film Director, Screenwriter, and Actor. The first two would get a high score in our setting, because those are his main professions. The third would get a low score, because he mostly had cameo appearances in his own movies. Such scores are essential in the ranking for entity queries, e.g. "American actors" or "Quentin Tarantino professions". These scores are different from scores for "correctness" or "accuracy" (all three professions above are correct and accurate). We propose a variety of algorithms to compute these scores. For our evaluation we designed a new benchmark, which includes a ground truth based on about 14K human judgments obtained via crowdsourcing. Inter-judge agreement is slightly over 90%. Existing approaches from the literature give results far from the optimum. Our best algorithms achieve an agreement of about 80% with the ground truth. Hannah Bast, Björn Buchhold, Elmar Haussmann |
SIGIR | 3 |
| 2014 | More Informative Open Information Extraction via Simple Inference
Hannah Bast, Elmar Haussmann |
ECIR | 2 |
| 2014 | Semantic full-text search with broccoliabstractWe combine search in triple stores with full-text search into what we call \emph{semantic full-text search}. We provide a fully functional web application that allows the incremental construction of complex queries on the English Wikipedia combined with the facts from Freebase. The user is guided by context-sensitive suggestions of matching words, instances, classes, and relations after each keystroke. We also provide a powerful API, which may be used for research tasks or as a back end, e.g., for a question answering system. Our web application and public API are available under \url{http://broccoli.cs.uni-freiburg.de}. Hannah Bast, Florian Bäurle, Björn Buchhold, Elmar Haussmann |
SIGIR | 4 |
| 2012 | Rethinking Java call stack design for tiny embedded devicesabstractThe ability of tiny embedded devices to run large feature-rich programs is typically constrained by the amount of memory installed on such devices. Furthermore, the useful operation of these devices in wireless sensor applications is limited by their battery life. This paper presents a call stack redesign targeted at an efficient use of RAM storage and CPU cycles by a Java program running on a wireless sensor mote. Without compromising the application programs, our call stack redesign saves 30% of RAM, on average, evaluated over a large number of benchmarks. On the same set of bench-marks, our design also avoids frequent RAM allocations and deallocations, resulting in average 80% fewer memory operations and 23% faster program execution. These may be critical improvements for tiny embedded devices that are equipped with small amount of RAM and limited battery life. However, our call stack redesign is equally effective for any complex multi-threaded object oriented program developed for desktop computers. We describe the redesign, measure its performance and report the resulting savings in RAM and execution time for a wide variety of programs. Faisal Aslam, Ghufran Baig, Mubashir Adnan Qureshi, Zartash Afzal Uzmi, Luminous Fennell, Peter Thiemann 0001, Christian Schindelhauer, Elmar Haussmann |
LCTES | 8 |
| 2010 | Optimized Java Binary and Virtual Machine for Tiny Motes
Faisal Aslam, Luminous Fennell, Christian Schindelhauer, Peter Thiemann 0001, Gidon Ernst, Elmar Haussmann, Stefan Rührup, Zartash Afzal Uzmi |
DCOSS | 6 |