Peter Mika

dblp:m/PMika · DBLP profile ↗
← Back
40ranked-venue papers
7as first author
0since 2021 · last 2017
0009-0007-8902-519XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 37 · 7 first-authorArtificial intelligence and machine learning · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Information retrieval · 72% Knowledge graphs · 12% Query processing and optimization · 7%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › query suggestion
query auto-completion
0.322016
Term-by-Term Query Auto-Completion for Mobile Search · WSDM 2016
An evaluation of entity and frequency based query completion methods · SIGIR 2009
Information retrieval › interactive information retrieval › search tasks
re-finding
0.312017
Re-Finding Behaviour in Vertical Domains · ACM Trans. Inf. Syst. 2017
Query processing and optimization
ad-hoc query
0.212015
Online News Tracking for Ad-Hoc Queries · SIGIR 2015
Information retrieval
query processing
0.212015
Online News Tracking for Ad-Hoc Queries · SIGIR 2015
Information retrieval › text analysis › topic analysis
topic detection and tracking
0.212015
Online News Tracking for Ad-Hoc Queries · SIGIR 2015
Data integration and cleaning
entity disambiguation
0.212013
WOO: A Scalable and Multi-tenant Platform for Continuous Knowledge Base Synthesis · Proc. VLDB Endow. 2013
Knowledge graphs
knowledge graph construction
0.212013
WOO: A Scalable and Multi-tenant Platform for Continuous Knowledge Base Synthesis · Proc. VLDB Endow. 2013
Web and social media mining
web usage mining
0.212013
Web usage mining with semantic analysis · WWW 2013
Information retrieval › search engines › semantic search › entity retrieval
ad-hoc object retrieval
0.122011
Ad-hoc object retrieval in the web of data · WWW 2010
Repeatable and reliable search system evaluation using crowdsourcing · SIGIR 2011
Information retrieval › evaluation › evaluation methodology
crowdsourced evaluation
0.112011
Repeatable and reliable search system evaluation using crowdsourcing · SIGIR 2011
Information retrieval
retrieval evaluation
0.112011
Repeatable and reliable search system evaluation using crowdsourcing · SIGIR 2011
Information retrieval
web search
0.112011
Enhanced results for web search · SIGIR 2011
Information retrieval › retrieval models
ad-hoc retrieval
0.112010
Ad-hoc object retrieval in the web of data · WWW 2010
Information retrieval
retrieval models
0.112010
Ad-hoc object retrieval in the web of data · WWW 2010
Information retrieval › search engines
semantic search
0.112010
Ad-hoc object retrieval in the web of data · WWW 2010
Information retrieval › user behavior
search session analysis
0.112017
Re-Finding Behaviour in Vertical Domains · ACM Trans. Inf. Syst. 2017
Information retrieval › web search
mobile search
0.112016
Term-by-Term Query Auto-Completion for Mobile Search · WSDM 2016
Information retrieval
query log analysis
0.112016
Term-by-Term Query Auto-Completion for Mobile Search · WSDM 2016
Information retrieval
text summarization
0.112015
Online News Tracking for Ad-Hoc Queries · SIGIR 2015
Information retrieval › text summarization › temporal summarization
update summarization
0.112015
Online News Tracking for Ad-Hoc Queries · SIGIR 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.012004
Foundations for service ontologies: aligning OWL-S to dolce · WWW 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology matching
0.012004
Foundations for service ontologies: aligning OWL-S to dolce · WWW 2004
Information retrieval › image retrieval
object retrieval
0.012011
Repeatable and reliable search system evaluation using crowdsourcing · SIGIR 2011
Information retrieval › search interfaces
search result presentation
0.012011
Enhanced results for web search · SIGIR 2011
Knowledge graphs › semantic web
linked data
0.012010
Ad-hoc object retrieval in the web of data · WWW 2010

Methods — techniques the papers use, named apart from their topics

entity enrichment · 0.3disambiguation · 0.3machine-learned models · 0.3feature engineering · 0.3user model · 0.2predictive keyboard · 0.2online tracking · 0.2semantic analysis · 0.2pattern mining · 0.2crowdsourcing · 0.1foundational ontology alignment · 0.0OWL · 0.0
YearPublicationVenuePosition
2017 Re-Finding Behaviour in Vertical Domains
abstract
Re-finding is the process of searching for information that a user has previously encountered and is a common activity carried out with information retrieval systems. In this work, we investigate re-finding in the context of vertical search, differentiating and modeling user re-finding behavior within different media and topic domains, including images, news, reference material, and movies. We distinguish the re-finding behavior in vertical domains from re-finding in a general search context and engineer features that are effective in differentiating re-finding across the domains. The features are then used to build machine-learned models, achieving an accuracy of re-finding detection in verticals of 85.7% on average. Our results demonstrate that detecting re-finding in specific verticals is more difficult than examining re-finding for general search tasks. We then investigate the effectiveness of differentiating re-finding behavior in two restricted contexts: We consider the case where the history of a searcher’s interactions with the search system is not available. In this scenario, our features and models achieve an average accuracy of 77.5% across the domains. We then examine the detection of re-finding during the early part of a search session. Both of these restrictions represent potential real-world search scenarios, where a system is attempting to learn about a user but may have limited information available. Finally, we investigate in which types of domains re-finding is most difficult. Here, it would appear that re-finding images is particularly challenging for users. This research has implications for search engine design, in terms of adapting search results by predicting the type of user tasks and potentially enabling the presentation of vertical-specific results when re-finding is identified. To the best of our knowledge, this is the first work to investigate the issue of vertical re-finding.
Seyedeh Sargol Sadeghi, Roi Blanco, Peter Mika, Mark Sanderson, Falk Scholer, David Vallet
ACM Trans. Inf. Syst.3
2016 Enriching Product Ads with Metadata from HTML Annotations
Petar Ristoski, Peter Mika
ESWC2
2016 Term-by-Term Query Auto-Completion for Mobile Search
abstract
With the ever increasing usage of mobile search, where text input is typically slow and error-prone, assisting users to formulate their queries contributes to a more satisfactory search experience. Query auto-completion (QAC) techniques, which predict possible completions for user queries, are the archetypal example of query assistance and are present in most search engines. We argue, however, that classic QAC, which operates by suggesting whole-query completions, may be sub-optimal for the case of mobile search as the available screen real estate to show suggestions is limited and editing is typically slower than in desktop search. In this paper we propose the idea of term-by-term QAC, which is a new technique inspired by predictive keyboards that suggests to the user one term at a time, instead of whole-query completions. We describe an efficient mechanism to implement this technique and an adaptation of a prior user model to evaluate the effectiveness of both standard and term-by-term QAC approaches using query log data. Our experiments with a mobile query log from a commercial search engine show the validity of our approach according to this user model with respect to saved characters, saved terms and examination effort. Finally, a user study provides further insights about our term-by-term technique compared with standard QAC with respect to the variables analyzed in the query log-based evaluation and additional variables related to the successfulness, the speed of the interactions and the properties of the submitted queries.
Saul Vargas, Roi Blanco, Peter Mika
WSDM3
2015 Predicting Re-finding Activity and Difficulty
Seyedeh Sargol Sadeghi, Roi Blanco, Peter Mika, Mark Sanderson, Falk Scholer, David Vallet
ECIR3
2015 Timely Semantics: A Study of a Stream-Based Ranking System for Entity Relationships
Lorenz Fischer, Roi Blanco, Peter Mika, Abraham Bernstein
ISWC (2)3
2015 Online News Tracking for Ad-Hoc Queries
abstract
Following news about a specific event can be a difficult task as new information is often scattered across web pages. An up-to-date summary of the event would help to inform users and allow them to navigate to articles that are likely to contain relevant and novel details. We demonstrate an approach that is feasible for online tracking of news that is relevant to a user's ad-hoc query.
Jeroen B. P. Vuurens, Arjen P. de Vries, Roi Blanco, Peter Mika
SIGIR4
2015 Ranking of daily deals with concept expansion
Roi Blanco, Michael Matthews, Peter Mika
Inf. Process. Manag.3
2014 Focused Crawling for Structured Data
abstract
The Web is rapidly transforming from a pure document collection to the largest connected public data space. Semantic annotations of web pages make it notably easier to extract and reuse data and are increasingly used by both search engines and social media sites to provide better search experiences through rich snippets, faceted search, task completion, etc. In our work, we study the novel problem of crawling structured data embedded inside HTML pages. We describe Anthelion, the first focused crawler addressing this task. We propose new methods of focused crawling specifically designed for collecting data-rich pages with greater efficiency. In particular, we propose a novel combination of online learning and bandit-based explore/exploit approaches to predict data-rich web pages based on the context of the page as well as using feedback from the extraction of metadata from previously seen pages. We show that these techniques significantly outperform state-of-the-art approaches for focused crawling, measured as the ratio of relevant pages and non-relevant pages collected within a given budget.
Robert Meusel, Peter Mika, Roi Blanco
CIKM2
2014 The PARLANCE mobile application for interactive search in English and Mandarin
abstract
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014.
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu
SIGDIAL Conference12
2013 Learning Relevance of Web Resources across Domains to Make Recommendations
abstract
Most traditional recommender systems focus on the objective of improving the accuracy of recommendations in a single domain. However, preferences of users may extend over multiple domains, especially in the Web where users often have browsing preferences that span across different sites, while being unaware of relevant resources on other sites. This work tackles the problem of recommending resources from various domains by exploiting the semantic content of these resources in combination with patterns of user browsing behavior. We overcome the lack of overlaps between domains by deriving connections based on the explored semantic content of Web resources. We present an approach that applies Support Vector Machines for learning the relevance of resources and predicting which ones are the most relevant to recommend to a user, given that the user is currently viewing a certain page. In real-world datasets of semantically-enriched logs of user browsing behavior at multiple Web sites, we study the impact of structure in generating accurate recommendations and conduct experiments that demonstrate the effectiveness of our approach.
Julia Hoxha, Peter Mika, Roi Blanco
ICMLA (2)2
2013 Entity Recommendations in Web Search
Roi Blanco, Berkant Barla Cambazoglu, Peter Mika, Nicolas Torzec
ISWC (2)3
2013 Federated Entity Search Using On-the-Fly Consolidation
Daniel M. Herzig, Peter Mika, Roi Blanco, Thanh Tran 0001
ISWC (1)2
2013 Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay
SIGDIAL Conference10
2013 Web usage mining with semantic analysis
abstract
Web usage mining has traditionally focused on the individual queries or query words leading to a web site or web page visit, mining patterns in such data. In our work, we aim to characterize websites in terms of the semantics of the queries that lead to them by linking queries to large knowledge bases on the Web. We demonstrate how to exploit such links for more effective pattern mining on query log data. We also show how such patterns can be used to qualitatively describe the differences between competing websites in the same domain and to quantitatively predict website abandonment.
Laura Hollink, Peter Mika, Roi Blanco
WWW2
2013 WOO: A Scalable and Multi-tenant Platform for Continuous Knowledge Base Synthesis
abstract
Search, exploration and social experience on the Web has recently undergone tremendous changes with search engines, web portals and social networks offering a different perspective on information discovery and consumption. This new perspective is aimed at capturing user intents, and providing richer and highly connected experiences. The new battleground revolves around technologies for the ingestion, disambiguation and enrichment of entities from a variety of structured and unstructured data sources - we refer to this process as knowledge base synthesis. This paper presents the design, implementation and production deployment of the Web Of Objects (WOO) system, a Hadoop-based platform tackling such challenges. WOO has been designed and implemented to enable various products in Yahoo! to synthesize knowledge bases (KBs) of entities relevant to their domains. Currently, the implementation of WOO we describe is used by various Yahoo! properties such as Intonow, Yahoo! Local, Yahoo! Events and Yahoo! Search. This paper highlights: (i) challenges that arise in designing, building and operating a platform that handles multi-domain, multi-version, and multi-tenant disambiguation of web-scale knowledge bases (hundreds of millions of entities), (ii) the architecture and technical solutions we devised, and (iii) an evaluation on real-world production datasets.
Kedar Bellare, Carlo Curino, Ashwin Machanavajjhala, Peter Mika, Mandar Rahurkar, Aamod Sane
Proc. VLDB Endow.4
2013 Repeatable and reliable semantic search evaluation
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001
J. Web Semant.4
2012 Fifth workshop on exploiting semantic annotations in information retrieval: ESAIR"12)
abstract
There is an increasing amount of structure on the Web as a result of modern Web languages, user tagging and annotation, emerging robust NLP tools, and an ever growing volume of linked data. These meaningful, semantic, annotations hold the promise to significantly enhance information access, by enhancing the depth of analysis of today's systems. Currently, we have only started exploring the possibilities and only begin to understand how these valuable semantic cues can be put to fruitful use. To complicate matters, standard text search excels at shallow information needs expressed by short keyword queries, and here semantic annotation contributes very little, if anything. The main questions for the workshop are how to leverage the rich context currently available, especially in a mobile search scenario, giving powerful new handles to exploit semantic annotations. And how can we fruitfully combine information retrieval and semantic web approaches, and for the first time work actively toward a unified view on exploiting semantic annotations.
Jaap Kamps, Jussi Karlgren, Peter Mika, Vanessa Murdock 0001
CIKM3
2012 Measuring website similarity using an entity-aware click graph
abstract
Query logs record the actual usage of search systems and their analysis has proven critical to improving search engine functionality. Yet, despite the deluge of information, query log analysis often suffers from the sparsity of the query space. Based on the observation that most queries pivot around a single entity that represents the main focus of the user's need, we propose a new model for query log data called the entity-aware click graph. In this representation, we decompose queries into entities and modifiers, and measure their association with clicked pages. We demonstrate the benefits of this approach on the crucial task of understanding which websites fulfill similar user needs, showing that using this representation we can achieve a higher precision than other query log-based approaches.
Pablo N. Mendes, Peter Mika, Hugo Zaragoza, Roi Blanco
CIKM2
2012 Robust Runtime Optimization and Skew-Resistant Execution of Analytical SPARQL Queries on Pig
Spyros Kotoulas, Jacopo Urbani, Peter Boncz, Peter Mika
ISWC (1)4
2011 Coreference aware web object retrieval
abstract
As user demands become increasingly sophisticated, search engines today are competing in more than just returning document results from the Web. One area of competition is providing web object results from structured data extracted from a multitude of information sources. We address the problem of performing keyword retrieval over a collection of objects containing a large degree of duplication as different Web-based information sources provide descriptions of the same object. We develop a method for coreference aware retrieval that performs topic-specific coreference resolution on retrieved objects in order to improve object search results. Our results demonstrate that coreference has a significant impact on the effectiveness of retrieval in the domain of local search. Our results show that a coreference aware system outperforms naive object retrieval by more than 20% in P5 and P10.
Jeff Dalton 0001, Roi Blanco, Peter Mika
CIKM3
2011 Effective and Efficient Entity Search in RDF Data
Roi Blanco, Peter Mika, Sebastiano Vigna
ISWC (1)2
2011 Repeatable and reliable search system evaluation using crowdsourcing
abstract
The primary problem confronting any new kind of search task is how to boot-strap a reliable and repeatable evaluation campaign, and a crowd-sourcing approach provides many advantages. However, can these crowd-sourced evaluations be repeated over long periods of time in a reliable manner? To demonstrate, we investigate creating an evaluation campaign for the semantic search task of keyword-based ad-hoc object retrieval. In contrast to traditional search over web-pages, object search aims at the retrieval of information from factual assertions about real-world objects rather than searching over web-pages with textual descriptions. Using the first large-scale evaluation campaign that specifically targets the task of ad-hoc Web object retrieval over a number of deployed systems, we demonstrate that crowd-sourced evaluation campaigns can be repeated over time and still maintain reliable results. Furthermore, we show how these results are comparable to expert judges when ranking systems and that the results hold over different evaluation and relevance metrics. This work provides empirical support for scalable, reliable, and repeatable search system evaluation using crowdsourcing.
Roi Blanco, Harry Halpin, Daniel M. Herzig, Peter Mika, Jeffrey Pound, Henry S. Thompson, Thanh Tran 0001
SIGIR4
2011 Enhanced results for web search
abstract
"Ten blue links" have defined web search results for the last fifteen years -- snippets of text combined with document titles and URLs. In this paper, we establish the notion of enhanced search results that extend web search results to include multimedia objects such as images and video, intent-specific key value pairs, and elements that allow the user to interact with the contents of a web page directly from the search results page. We show that users express a preference for enhanced results both explicitly, and when observed in their search behavior. We also demonstrate the effectiveness of enhanced results in helping users to assess the relevance of search results. Lastly, we show that we can efficiently generate enhanced results to cover a significant fraction of search result pages.
Kevin Haas, Peter Mika, Paul Tarjan, Roi Blanco
SIGIR2
2010 Making Sense of Twitter
David Laniado, Peter Mika
ISWC (1)2
2010 Ad-hoc object retrieval in the web of data
abstract
Semantic Search refers to a loose set of concepts, challenges and techniques having to do with harnessing the information of the growing Web of Data (WoD) for Web search. Here we propose a formal model of one specific semantic search task: ad-hoc object retrieval. We show that this task provides a solid framework to study some of the semantic search problems currently tackled by commercial Web search engines. We connect this task to the traditional ad-hoc document retrieval and discuss appropriate evaluation metrics. Finally, we carry out a realistic evaluation of this task in the context of a Web search application.
Jeffrey Pound, Peter Mika, Hugo Zaragoza
WWW2
2010 The Semantic Web Challenge, 2009
Christian Bizer, Peter Mika
J. Web Semant.2
2009 Investigating the Semantic Gap through Query Log Analysis
Peter Mika, Edgar Meij, Hugo Zaragoza
ISWC1
2009 An evaluation of entity and frequency based query completion methods
abstract
We present a semantic approach to suggesting query completions which leverages entity and type information. When compared to a frequency-based approach, we show that such information mostly helps rare queries.
Edgar Meij, Peter Mika, Hugo Zaragoza
SIGIR2
2009 The Semantic Web challenge, 2008
Peter Mika, James A. Hendler
J. Web Semant.1
2008 Towards Semantic Search
Ricardo Baeza-Yates, Massimiliano Ciaramita, Peter Mika, Hugo Zaragoza
NLDB3
2008 Introduction to the special issue on the Semantic Web Challenge 2006 and 2007
Jennifer Golbeck, Peter Mika, Michael Uschold
J. Web Semant.2
2008 Semantic Web and Web 2.0
Mark Greaves, Peter Mika
J. Web Semant.2
2007 Ranking very many typed entities on wikipedia
abstract
We discuss the problem of ranking very many entities of different types. In particular we deal with a heterogeneous set of types, some being very generic and some very specific. We discuss two approaches for this problem: i) exploiting the entity containment graph and ii) using a Web search engine to compute entity relevance. We evaluate these approaches on the real task of ranking Wikipedia entities typed with a state-of-the-art named-entity tagger. Results show that both approaches can greatly increase the performance of methods based only on passage retrieval.
Hugo Zaragoza, Henning Rode, Peter Mika, Jordi Atserias Batalla, Massimiliano Ciaramita, Giuseppe Attardi
CIKM3
2007 Ontologies are us: A unified model of social networks and semantics
Peter Mika
J. Web Semant.1
2005 Ontologies Are Us: A Unified Model of Social Networks and Semantics
Peter Mika
ISWC1
2005 Flink: Semantic Web technology for the extraction and analysis of social networks
Peter Mika
J. Web Semant.1
2004 Bibster - A Semantics-Based Bibliographic Peer-to-Peer System
Peter Haase 0001, Jeen Broekstra, Marc Ehrig, Maarten Menken, Peter Mika, Mariusz Olko, Michal Plechawski, Pawel Pyszlak, Björn Schnizler, Ronny Siebes, Steffen Staab, Christoph Tempich
ISWC5
2004 Social Networks and the Semantic Web
abstract
A formal, web-based representation of social networks is both a necessity in terms of infrastructure as well as a prominent application for the Semantic Web. In this paper we present three advances in exploiting the opportunity of semantically-enriched network data: (1) an ontology for the representation of social networks and relationships (2) a hybrid system for online data acquisition that combines traditional web mining techniques with the collection of Semantic Web data (2) a case study highlighting some of the possible analysis of this data using methods from Social Network Analysis, the branch of sociology concerned with relational data.
Peter Mika
Web Intelligence1
2004 Foundations for service ontologies: aligning OWL-S to dolce
abstract
Clarity in semantics and a rich formalization of this semantics are important requirements for ontologies designed to be deployed in large-scale, open, distributed systems such as the envisioned Semantic Web This is especially important for the description of Web Services, which should enable complex tasks involving multiple agents. As one of the first initiatives of the Semantic Webcommunity for describing Web Services, OWL-S attracts a lot of interest even though it is still under development. We identify problematic aspects of OWL-S and suggest enhancements through alignment to a foundational ontology. Another contribution of ourwork is the Core Ontology of Services that tries to fill the epistemological gap between the foundational ontology and OWL-S. It can be reused to align other Web Service description languages as well. Finally, we demonstrate the applicability of our work byaligning OWL-S' standard example called CongoBuy.
Peter Mika, Daniel Oberle, Aldo Gangemi, Marta Sabou
WWW1
2004 Bibster - a semantics-based bibliographic Peer-to-Peer system
Peter Haase 0001, Björn Schnizler, Jeen Broekstra, Marc Ehrig, Frank van Harmelen, Maarten Menken, Peter Mika, Michal Plechawski, Pawel Pyszlak, Ronny Siebes, Steffen Staab, Christoph Tempich
J. Web Semant.7