EDBT 2026 Demo / reviewers in the wild / expert
Hinrich Schütze
dblp:s/HinrichSchutze
· DBLP profile ↗
22ranked-venue papers in the field
5as first author
7since 2021 · last 2026
0000-0001-9514-7934ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 18 (5 first)Data Mining & Knowledge Discovery · 2Other / Interdisciplinary · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GlotWeb: Web Indexing for Minority LanguagesabstractInternational audience Abdullah Al Sefat, Amir Hossein Kargaran, François Yvon, Hinrich Schütze |
WWW | 4 |
| 2024 | A Unified Data Augmentation Framework for Low-Resource Multi-domain Dialogue Generation
Yongkang Liu 0002, Ercong Nie, Shi Feng 0001, Zifeng Ding, Daling Wang, Yifei Zhang 0003, Hinrich Schütze |
ECML/PKDD (2) | 8 |
| 2023 | GIRT-Data: Sampling GitHub Issue Report TemplatesabstractGitHub’s issue reports provide developers with valuable information that is essential to the evolution of a software development project. Contributors can use these reports to perform software engineering tasks like submitting bugs, requesting features, and collaborating on ideas. In the initial versions of issue reports, there was no standard way of using them. As a result, the quality of issue reports varied widely. To improve the quality of issue reports, GitHub introduced issue report templates (IRTs), which pre-fill issue descriptions when a new issue is opened. An IRT usually contains greeting contributors, describing project guidelines, and collecting relevant information. However, despite of effectiveness of this feature which was introduced in 2016, only nearly 5% of GitHub repositories (with more than 10 stars) utilize it. There are currently few articles on IRTs, and the available ones only consider a small number of repositories.In this work, we introduce GIRT-DATA, the first and largest dataset of IRTs in both YAML and Markdown format. This dataset and its corresponding open-source crawler tool are intended to support research in this area and to encourage more developers to use IRTs in their repositories. The stable version of the dataset contains 1,084,300 repositories and 50,032 of them support IRTs. The stable version of the dataset and crawler is available here: https://github.com/kargaranamir/girt-data Nafiseh Nikeghbal, Amir Hossein Kargaran, Abbas Heydarnoori, Hinrich Schütze |
MSR | 4 |
| 2022 | The Reddit Politosphere: A Large-Scale Text and Network Resource of Online Political Discourse
Valentin Hofmann, Hinrich Schütze, Janet B. Pierrehumbert |
ICWSM | 2 |
| 2022 | Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts
Lutfi Kerem Senel, Furkan Sahinuç, Veysel Yücesoy, Hinrich Schütze, Tolga Çukur, Aykut Koç |
Inf. Process. Manag. | 4 |
| 2021 | Data Centric Domain Adaptation for Historical Text with OCR Errors
Luisa März, Stefan Schweter, Nina Pörner, Benjamin Roth 0001, Hinrich Schütze |
ICDAR (2) | 5 |
| 2021 | Semantic Text Segment Classification of Structured Technical Content
Julian Höllig, Philipp Dufter, Michaela Geierhos, Wolfgang Ziegler, Hinrich Schütze |
NLDB | 5 |
| 2019 | SMAPH: A Piggyback Approach for Entity-Linking in Web QueriesabstractWe study the problem of linking the terms of a web-search query to a semantic representation given by the set of entities (a.k.a. concepts) mentioned in it. We introduce SMAPH, a system that performs this task using the information coming from a web search engine, an approach we call “piggybacking.” We employ search engines to alleviate the noise and irregularities that characterize the language of queries. Snippets returned as search results also provide a context for the query that makes it easier to disambiguate the meaning of the query. From the search results, SMAPH builds a set of candidate entities with high coverage. This set is filtered by linking back the candidate entities to the terms occurring in the input query, ensuring high precision. A greedy disambiguation algorithm performs this filtering; it maximizes the coherence of the solution by iteratively discovering the pertinent entities mentioned in the query. We propose three versions of SMAPH that outperform state-of-the-art solutions on the known benchmarks and on the GERDAQ dataset, a novel dataset that we have built specifically for this problem via crowd-sourcing and that we make publicly available. Marco Cornolti, Paolo Ferragina, Massimiliano Ciaramita, Stefan Rüd, Hinrich Schütze |
ACM Trans. Inf. Syst. | 5 |
| 2016 | A Piggyback System for Joint Entity Mention Detection and Linking in Web QueriesabstractIn this paper we study the problem of linking open-domain web-search queries towards entities drawn from the full entity inventory of Wikipedia articles. We introduce SMAPH-2, a second-order approach that, by piggybacking on a web search engine, alleviates the noise and irregularities that characterize the language of queries and puts queries in a larger context in which it is easier to make sense of them. The key algorithmic idea underlying SMAPH-2 is to first discover a candidate set of entities and then link-back those entities to their mentions occurring in the input query. This allows us to confine the possible concepts pertinent to the query to only the ones really mentioned in it. The link-back is implemented via a collective disambiguation step based upon a supervised ranking model that makes one joint prediction for the annotation of the complete query optimizing directly the F1 measure. We evaluate both known features, such as word embeddings and semantic relatedness among entities, and several novel features such as an approximate distance between mentions and entities (which can handle spelling errors). We demonstrate that SMAPH-2 achieves state-of-the-art performance on the [email protected] benchmark. We also publish GERDAQ (General Entity Recognition, Disambiguation and Annotation in Queries), a novel, public dataset built specifically for web-query entity linking via a crowdsourcing effort. SMAPH-2 outperforms the benchmarks by comparable margins also on GERDAQ. Marco Cornolti, Paolo Ferragina, Massimiliano Ciaramita, Stefan Rüd, Hinrich Schütze |
WWW | 5 |
| 2012 | Crosslingual distant supervision for extracting relations of different complexityabstractWe propose crosslingual distant supervision (crosslingual DS) for relation extraction, an approach that automatically extracts labels from a pivot language for labeling one or more target languages. The approach has two benefits compared to standard DS: (i) increased coverage if target language labels are not available; and (ii) higher accuracy of automatically generated labels because noisy labels are eliminated in crosslingual filtering. An evaluation for two relations of different complexity shows that crosslingual DS increases the accuracy of relation extraction. Our approach is language independent; we successfully apply it to four different languages: Chinese, English, French and German. André Blessing, Hinrich Schütze |
CIKM | 2 |
| 2012 | Preliminary study of technical terminology for the retrieval of scientific book metadata recordsabstractBooks only represented by brief metadata (book records) are particularly hard to retrieve. One way of improving their retrieval is by extracting retrieval enhancing features from them. This work focusses on scientific (physics) book records. We ask if their technical terminology can be used as a retrieval enhancing feature. A study of 18,443 book records shows a strong correlation between their technical terminology and their likelihood of relevance. Using this finding for retrieval yields >+5% precision and recall gains. Birger Larsen, Christina Lioma, Ingo Frommholz, Hinrich Schütze |
SIGIR | 4 |
| 2011 | Sense discrimination for physics retrievalabstractInformation Retrieval in technical domains like physics is characterised by long and precise queries, whose meaning is strongly influenced by term context and domain. We treat this as a disambiguation problem, and present initial findings of a retrieval model that posits a higher probability of relevance for documents matching disambiguated query terms. Preliminary evaluation on a real-life physics test collection shows promising performance improvement. Christina Lioma, Alok Kothari, Hinrich Schütze |
SIGIR | 3 |
| 2010 | Relational feature engineering of natural language processingabstractWe present a new framework for feature engineering of natural language processing that is based on a relational data model of text. It includes fast and flexible methods for implementing and extracting new features and thereby reduces the effort of creating an NLP system for a particular task. Hamidreza Kobdani, Hinrich Schütze, Andre Burkovski, Wiltrud Kessler, Gunther Heidemann |
CIKM | 2 |
| 2010 | IR, NLP, and Visualization
Hinrich Schütze |
ECIR | 1 |
| 2008 | Automatic acquisition of vernacular placesabstractThis paper delineates our approach to augment geospatial datasets by named regions which are not typically handled by surveyors. Such regions are very important in vernacular speech and play a significant role in context-aware systems. Three different region types can be distinguished: (i) functional regions which are named after their purpose (e.g., financial district, business quarter), (ii) regions which have vernacular names (e.g. Bohnenviertel) and (iii) regions which are named after spatial attributes (Stuttgart-Süd). To acquire such regions we use a web-based approach. Our first implementation is used as a proof of concept and provides promising results. We describe several improvements which will be implemented in the future. Finally a possible scenario for a rigorous evaluation is introduced. André Blessing, Hinrich Schütze |
iiWAS | 2 |
| 2008 | Disorder inequality: a combinatorial approach to nearest neighbor searchabstractWe say that an algorithm for nearest neighbor search is combinatorial if only direct comparisons between two pairwise similarity values are allowed. Combinatorial algorithms for nearest neighbor search have two important advantages: (1) they do not map similarity values to artificial distance values and do not use the triangle inequality for the latter, and (2) they work for arbitrarily complicated data representations and similarity functions. Navin Goyal, Yury Lifshits, Hinrich Schütze |
WSDM | 3 |
| 2007 | Improving active learning recall via disjunctive boolean constraintsabstractActive learning efficiently hones in on the decision boundary between relevant and irrelevant documents, but in the process can miss entire clusters of relevant documents, yielding classifiers with low recall. In this paper, we propose a method to increase active learning recall by constraining sampling to a document subset rich in relevant examples. Emre Velipasaoglu, Hinrich Schütze, Jan O. Pedersen 0001 |
SIGIR | 2 |
| 2006 | Performance thresholding in practical text classificationabstractIn practical classification, there is often a mix of learnable and unlearnable classes and only a classifier above a minimum performance threshold can be deployed. This problem is exacerbated if the training set is created by active learning. The bias of actively learned training sets makes it hard to determine whether a class has been learned. We give evidence that there is no general and efficient method for reducing the bias and correctly identifying classes that have been learned. However, we characterize a number of scenarios where active learning can succeed despite these difficulties. Hinrich Schütze, Emre Velipasaoglu, Jan O. Pedersen 0001 |
CIKM | 1 |
| 1997 | Projections for Efficient Document ClusteringabstractClustering is increasing in importance, but linear-and even constant-time clustering algorithms are often too slow for real-time applications.A simple way to speed up clustering is to speed up the distance calculations at the heart of clustering routines.We study two techniques for improving the cost ofdistance calculations, LSI and trrmcation, and determine both how much these techniques speed up clustering and how much they affect the quality of the resulting clusters.We find that the speed increase is significant whilesurprisingly -the quality of clustering is not adversely affected.We conclude that truncation yields clusters as good as those produced by full-profile clustering while offering a significant speed advantage. Hinrich Schütze, Craig Silverstein |
SIGIR | 1 |
| 1997 | A Cooccurrence-Based Thesaurus and Two Applications to Information Retrieval
Hinrich Schütze, Jan O. Pedersen 0001 |
Inf. Process. Manag. | 1 |
| 1996 | Method Combination For Document FilteringabstractThere is strong empirical and theoretic evidence that combination of retrieval methods can improve performance.In this paper, we systematically compare combination strategies in the context of document filtering, using queries from the Tipster reference corpus.We find that simple averaging strategies do indeed improve performance, but that direet averaging of probability estimates is not the correet approach.Instead, the probabiJit y estimates must be renormalized using logistic regression on the known relevance judgments.We examine more complex combination strat~ gies but find them less successful due to the high correlations among our filtering methods which are optimized over the same training data and employ similar document represerttations.1 David A. Hull, Jan O. Pedersen 0001, Hinrich Schütze |
SIGIR | 3 |
| 1995 | A Comparison of Classifiers and Document Representations for the Routing ProblemabstractIn this paper, we compare learning techniques based on statistical classification to traditional methods of relevance feedback for the document routing problem.We consider three classification techniques which have decision rules that are derived via explicit error minimization linear discriminant analysis, logistic regression, and neuraf networks.We demonstrate that the classifiers perform 10-15% better than relevance feedback via Rocchio expansion for the TREC-2 and TREC-3 routing tasks.Error minimization is difficult in high-dimensional feature spaces because the convergence process is slow and the models ~e prone to overfitting.We use two different strategies, latent semantic indexing and optimaJ term selection, to reduce the number of features.Our results indicate that features based on latent semantic indexing are more effective for techniques such as linear discriminant anafysis and logistic regression, which have no way to protect against overfitting.Neural networks perform equally well with either set of features and can take advantage of the additional information available when both feature sets are used as input. Hinrich Schütze, David A. Hull, Jan O. Pedersen 0001 |
SIGIR | 1 |