Vassilis Plachouras

dblp:59/5877 · also Vasileios Plachouras · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 28 · 7 first-authorArtificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
11 papers
Information retrieval · 84% Web and social media mining · 10% Query processing and optimization · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 57% Distributed systems · 43%

Topics — the 30 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
fact-checking
0.412019
ROME 2019: Workshop on Reducing Online Misinformation Exposure · SIGIR 2019
Web and social media mining
misinformation detection
0.412019
ROME 2019: Workshop on Reducing Online Misinformation Exposure · SIGIR 2019
Information retrieval › document retrieval › domain-specific retrieval
financial information retrieval
0.212016
Interacting with Financial Data using Natural Language · SIGIR 2016
Information retrieval › query formulation
natural language querying
0.212016
Interacting with Financial Data using Natural Language · SIGIR 2016
Information retrieval
question answering
0.212016
Interacting with Financial Data using Natural Language · SIGIR 2016
Information retrieval
search interfaces
0.212016
Interacting with Financial Data using Natural Language · SIGIR 2016
Information retrieval › search engines
search engine architecture
0.222010
A refreshing perspective of search engine caching · WWW 2010
Challenges on Distributed Web Retrieval · ICDE 2007
Information retrieval › search engines
search engine caching
0.222010
A refreshing perspective of search engine caching · WWW 2010
The impact of caching on search engines · SIGIR 2007
Information retrieval › distributed information retrieval
distributed web search
0.222009
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Challenges on Distributed Web Retrieval · ICDE 2007
Query processing and optimization
query result caching
0.222008
ResIn: a combination of results caching and index pruning for high-performance web search engines · SIGIR 2008
The impact of caching on search engines · SIGIR 2007
Information retrieval
query processing
0.122009
On efficient posting list intersection with multicore processors · SIGIR 2009
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Information retrieval
retrieval models
0.122007
Incorporating term dependency in the dfr framework · SIGIR 2007
Usefulness of hyperlink structure for query-biased topic distillation · SIGIR 2004
Information retrieval › evaluation
benchmark
0.112019
ROME 2019: Workshop on Reducing Online Misinformation Exposure · SIGIR 2019
Information retrieval
evaluation
0.112019
ROME 2019: Workshop on Reducing Online Misinformation Exposure · SIGIR 2019
Information retrieval
distributed information retrieval
0.112009
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Information retrieval › query processing
list intersection
0.112009
On efficient posting list intersection with multicore processors · SIGIR 2009
Information retrieval › query log analysis
clickthrough data
0.112008
Online learning from click data for sponsored search · WWW 2008
Recommender systems
click-through rate prediction
0.112008
Online learning from click data for sponsored search · WWW 2008
Information retrieval › ranking
learning to rank
0.112008
Online learning from click data for sponsored search · WWW 2008
Information retrieval › online advertising
sponsored search
0.112008
Online learning from click data for sponsored search · WWW 2008
Information retrieval › distributed information retrieval
distributed search
0.112007
Challenges on Distributed Web Retrieval · ICDE 2007
Information retrieval › retrieval models
term dependency
0.112007
Incorporating term dependency in the dfr framework · SIGIR 2007
Information retrieval
retrieval evaluation
0.012004
Usefulness of hyperlink structure for query-biased topic distillation · SIGIR 2004
Information retrieval
search engines
0.012009
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Information retrieval › search engines
web crawling
0.012009
Quantifying performance and quality gains in distributed web search engines · SIGIR 2009
Processor architecture and microarchitecture
chip multiprocessor
0.012009
On efficient posting list intersection with multicore processors · SIGIR 2009
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron
0.012008
Online learning from click data for sponsored search · WWW 2008
Machine learning › Learning theory
online learning
0.012008
Online learning from click data for sponsored search · WWW 2008
Information retrieval › indexing › index compression
index pruning
0.012008
ResIn: a combination of results caching and index pruning for high-performance web search engines · SIGIR 2008
Information retrieval
query log analysis
0.012007
The impact of caching on search engines · SIGIR 2007

Methods — techniques the papers use, named apart from their topics

semantic web · 0.4natural language processing · 0.4complex networks · 0.4natural language generation · 0.2named entity recognition · 0.2time-to-live · 0.1refresh heuristic · 0.1simulation · 0.1cost model · 0.1pairwise preference learning · 0.1online multilayer perceptron · 0.1caching · 0.1survey · 0.1
YearPublicationVenuePosition
2021 Concept Matching for Low-Resource Classification
abstract
In many applications that rely on machine learning, the availability of labelled data is a matter of primary importance. However, when tackling new tasks, labels are usually missing and must be collected from scratch by the users. In this work, we address the problem of learning classifiers when the amount of labels is very scarce. We do so by learning multiple vectors, called prototypes, that represent relevant semantic concepts for the task at hand. We propose a theoretically inspired mechanism that computes probabilities of matching between the prototypes and the input elements, and we combine these probabilities to increase the expressiveness of the classifier. Moreover, by leveraging low-cost extra annotations in the training data, a simple error-boosting technique guides the learning process and provides substantial performance improvements. Empirical results confirm the benefits of the proposed approach in both balanced and unbalanced datasets. Our methodology is thus of practical use when gathering and labelling new examples is more expensive than annotating what we already have.
Federico Errica, Fabrizio Silvestri, Bora Edizel, Ludovic Denoyer, Fabio Petroni, Vassilis Plachouras, Sebastian Riedel 0001
IJCNN6
2021 KILT: a Benchmark for Knowledge Intensive Language Tasks
abstract
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick S. H. Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, Sebastian Riedel 0001
NAACL-HLT11
2019 ROME 2019: Workshop on Reducing Online Misinformation Exposure
abstract
The spread of misinformation online is a challenge that may have an impact on society by misleading and undermining the trust of people in domains such as politics or public health. While fact-checking is one way to identify misinformation, it is a slow process and requires significant effort. Improving the efficiency of fact-checking by automating parts of the process or defining new processes to validate claims is a challenging task with a need for expertise from multiple disciplines. The aim of ROME 2019 is to bring together researchers from various fields such as Information Retrieval, Natural Language Processing, Semantic Web and Complex Networks to discuss these problems and define new directions in the area of automated fact-checking.
Guillaume Bouchard, Guido Caldarelli, Vassilis Plachouras
SIGIR3
2018 attr2vec: Jointly Learning Word and Contextual Attribute Embeddings with Factorization Machines
abstract
Fabio Petroni, Vassilis Plachouras, Timothy Nugent, Jochen L. Leidner. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Fabio Petroni, Vassilis Plachouras, Timothy Nugent, Jochen L. Leidner
NAACL-HLT2
2016 When to Plummet and When to Soar: Corpus Based Verb Selection for Natural Language Generation
abstract
For data-to-text tasks in Natural Language Generation (NLG), researchers are often faced with choices about the right words to express phenomena seen in the data.One common phenomenon centers around the description of trends between two data points and selecting the appropriate verb to express both the direction and intensity of movement.Our research shows that rather than simply selecting the same verbs again and again, variation and naturalness can be achieved by quantifying writers' patterns of usage around verbs.
Charese Smiley, Vassilis Plachouras, Frank Schilder, Hiroko Bretz, Jochen L. Leidner, Dezhao Song
INLG2
2016 Interacting with Financial Data using Natural Language
abstract
Financial and economic data are typically available in the form of tables and comprise mostly of monetary amounts, numeric and other domain-specific fields. They can be very hard to search and they are often made available out of context, or in forms which cannot be integrated with systems where text is required, such as voice-enabled devices. This work presents a novel system that enables both experts in the finance domain and non-expert users to search financial data with both keyword and natural language queries. Our system answers the queries with an automatically generated textual description using Natural Language Generation (NLG). The answers are further enriched with derived information, not explicitly asked in the user query, to provide the context of the answer. The system is designed to be flexible in order to accommodate new use cases without significant development effort, thus allowing fast integration of new datasets.
Vassilis Plachouras, Charese Smiley, Hiroko Bretz, Ola Taylor, Jochen L. Leidner, Dezhao Song, Frank Schilder
SIGIR1
2015 Information Extraction of Regulatory Enforcement Actions: From Anti-Money Laundering Compliance to Countering Terrorism Finance
abstract
Financial fines imposed by regulatory bodies to penalize illegal activities and violations against regulations (cases of non-compliance) have recently become more common, and the sizes of fines have increased. This development coincides with the ongoing increase of complexity of regulatory rules. Huge fines have been imposed on banks for financial fraud and regulations have been made more stringent after 9/11 to curb funding of terrorist groups. Market players would also like to have available a database of fine events for a range of applications, such as to benchmark their competitors performance, or to use it as an early warning system for detecting shifts in regulators' enforcement behavior. To this end, we introduce the task of extracting fines from regulatory enforcement actions and we present a method to extract such fine event instances from timeline-like descriptions of regulatory investigation activities authored by legal professionals for a commercial product. We evaluate how well a rule-based method can extract information about fine events and we compare its performance to a machine-learning baseline. To the best of our knowledge, this work is the first one addressing this task.
Vassilis Plachouras, Jochen L. Leidner
ASONAM1
2012 Named Entity Recognition and Identification for Finding the Owner of a Home Page
Vassilis Plachouras, Matthieu Rivière, Michalis Vazirgiannis
PAKDD (1)1
2011 Large-scale information retrieval experimentation with terrier
abstract
This tutorial aims to provide a practical introduction to conducting large-scale information retrieval (IR) experiments, using Terrier (http://terrier.org) as an experimentation platform. Written in Java, Terrier provides an open-source, feature-rich, flexible, and robust environment for large-scale IR experimentation. This tutorial will cover the experimentation process end-to-end, from configuring Terrier to a particular experimental setting, to efficiently indexing a document corpus and retrieving from it, and to evaluating the outcome. Moreover, it will describe how to use and extend the platform to one's own needs, and will be illustrated by practical research-driven examples. As a half-day tutorial, it will be split into two major sessions, with each session comprising both background information and practical demonstrations. In the first session, we will provide an overview of several aspects of large-scale IR experimentation, spanning areas such as indexing, data structures, query languages, and advanced retrieval models, and how these are implemented within Terrier. In the second session, we will discuss how to extend Terrier to conduct one's own experiments in a large-scale setting, including how to facilitate the evaluation of non-standard IR tasks through crowdsourcing. The practical demonstrations will cover recent use cases identified from Terrier's online discussion forum, so as to provide attendees with concrete examples of what can be done within Terrier.
Rodrygo L. T. Santos, Richard McCreadie, Vassilis Plachouras
CIKM3
2010 A refreshing perspective of search engine caching
abstract
Commercial Web search engines have to process user queries over huge Web indexes under tight latency constraints. In practice, to achieve low latency, large result caches are employed and a portion of the query traffic is served using previously computed results. Moreover, search engines need to update their indexes frequently to incorporate changes to the Web. After every index update, however, the content of cache entries may become stale, thus decreasing the freshness of served results. In this work, we first argue that the real problem in today's caching for large-scale search engines is not eviction policies, but the ability to cope with changes to the index, i.e., cache freshness. We then introduce a novel algorithm that uses a time-to-live value to set cache entries to expire and selectively refreshes cached results by issuing refresh queries to back-end search clusters. The algorithm prioritizes the entries to refresh according to a heuristic that combines the frequency of access with the age of an entry in the cache. In addition, for setting the rate at which refresh queries are issued, we present a mechanism that takes into account idle cycles of back-end servers. Evaluation using a real workload shows that our algorithm can achieve hit rate improvements as well as reduction in average hit ages. An implementation of this algorithm is currently in production use at Yahoo!.
Berkant Barla Cambazoglu, Flavio Paiva Junqueira, Vassilis Plachouras, Scott A. Banachowski, Baoqiu Cui, Swee Lim, Bill Bridge
WWW3
2009 On the feasibility of multi-site web search engines
abstract
Web search engines are often implemented as centralized systems. Designing and implementing a Web search engine in a distributed environment is a challenging engineering task that encompasses many interesting research questions. However, distributing a search engine across multiple sites has several advantages, such as utilizing less compute resources and exploiting data locality. In this paper we investigate the cost-effectiveness of building a distributed Web search engine. We propose a model for assessing the total cost of a distributed Web search engine that includes the computational costs and the communication cost among all distributed sites. We then present a query-processing algorithm that maximizes the amount of queries answered locally, without sacrificing the quality of the results compared to a centralized search engine. We simulate the algorithm on real document collections and query workloads to measure the actual parameters needed for our cost model, and we show that a distributed search engine can be competitive compared to a centralized architecture with respect to real cost.
Ricardo Baeza-Yates, Aristides Gionis, Flavio Paiva Junqueira, Vassilis Plachouras, Luca Telloli
CIKM4
2009 A Study of the Impact of Index Updates on Distributed Query Processing for Web Search
Charalampos Sarigiannis, Vassilis Plachouras, Ricardo Baeza-Yates
ECIR2
2009 Quantifying performance and quality gains in distributed web search engines
abstract
Distributed search engines based on geographical partitioning of a central Web index emerge as a feasible solution to the immense growth of the Web, user bases, and query traffic. However, there is still lack of research in quantifying the performance and quality gains that can be achieved by such architectures. In this paper, we develop various cost models to evaluate the performance benefits of a geographically distributed search engine architecture based on partial index replication and query forwarding. Specifically, we focus on possible performance gains due to the distributed nature of query processing and Web crawling processes. We show that any response time gain achieved by distributed query processing can be utilized to improve search relevance as the use of complex but more accurate algorithms can now be enabled for document ranking. We also show that distributed Web crawling leads to better Web coverage and try to see if this improves the search quality. We verify the validity of our claims over large, real-life datasets via simulations.
Berkant Barla Cambazoglu, Vassilis Plachouras, Ricardo Baeza-Yates
SIGIR2
2009 On efficient posting list intersection with multicore processors
abstract
No abstract available.
Shirish Tatikonda, Flavio Paiva Junqueira, Berkant Barla Cambazoglu, Vassilis Plachouras
SIGIR4
2008 To swing or not to swing: learning when (not) to advertise
abstract
Web textual advertising can be interpreted as a search problem over the corpus of ads available for display in a particular context. In contrast to conventional information retrieval systems, which always return results if the corpus contains any documents lexically related to the query, in Web advertising it is acceptable, and occasionally even desirable, not to show any results. When no ads are relevant to the user's interests, then showing irrelevant ads should be avoided since they annoy the user and produce no economic benefit. In this paper we pose a decision problem to swing, that is, whether or not to show any of the ads for the incoming request. We propose two methods for addressing this problem, a simple thresholding approach and a machine learning approach, which collectively analyzes the set of candidate ads augmented with external knowledge. Our experimental evaluation, based on over 28,000 editorial judgments, shows that we are able to predict, with high accuracy, when to swing for both content match and sponsored search advertising.
Andrei Z. Broder, Massimiliano Ciaramita, Marcus Fontoura, Evgeniy Gabrilovich, Vanja Josifovski, Donald Metzler, Vanessa Murdock 0001, Vassilis Plachouras
CIKM8
2008 ResIn: a combination of results caching and index pruning for high-performance web search engines
abstract
Results caching is an efficient technique for reducing the query processing load, hence it is commonly used in real search engines. This technique, however, bounds the maximum hit rate due to the large fraction of singleton queries, which is an important limitation. In this paper we propose ResIn - an architecture that uses a combination of results caching and index pruning to overcome this limitation.
Gleb Skobeltsyn, Flavio Paiva Junqueira, Vassilis Plachouras, Ricardo Baeza-Yates
SIGIR3
2008 Online learning from click data for sponsored search
abstract
Sponsored search is one of the enabling technologies for today's Web search engines. It corresponds to matching and showing ads related to the user query on the search engine results page. Users are likely to click on topically related ads and the advertisers pay only when a user clicks on their ad. Hence, it is important to be able to predict if an ad is likely to be clicked, and maximize the number of clicks. We investigate the sponsored search problem from a machine learning perspective with respect to three main sub-problems: how to use click data for training and evaluation, which learning framework is more suitable for the task, and which features are useful for existing models. We perform a large scale evaluation based on data from a commercial Web search engine. Results show that it is possible to learn and evaluate directly and exclusively on click data encoding pairwise preferences following simple and conservative assumptions. We find that online multilayer perceptron learning, based on a small set of features representing content similarity of different kinds, significantly outperforms an information retrieval baseline and other learning models, providing a suitable framework for the sponsored search task.
Massimiliano Ciaramita, Vanessa Murdock 0001, Vassilis Plachouras
WWW3
2008 Design trade-offs for search engine caching
abstract
In this article we study the trade-offs in designing efficient caching systems for Web search engines. We explore the impact of different approaches, such as static vs. dynamic caching, and caching query results vs. caching posting lists. Using a query log spanning a whole year, we explore the limitations of caching and we demonstrate that caching posting lists can achieve higher hit rates than caching query answers. We propose a new algorithm for static caching of posting lists, which outperforms previous methods. We also study the problem of finding the optimal way to split the static cache between answers and posting lists. Finally, we measure how the changes in the query log influence the effectiveness of static caching, given our observation that the distribution of the queries changes slowly over time. Our results and observations are applicable to different levels of the data-access hierarchy, for instance, for a memory/disk layer or a broker/remote server layer.
Ricardo Baeza-Yates, Aristides Gionis, Flavio Paiva Junqueira, Vanessa Murdock 0001, Vassilis Plachouras, Fabrizio Silvestri
ACM Trans. Web5
2007 Performance Comparison of Clustered and Replicated Information Retrieval Systems
Fidel Cacheda, Victor Carneiro, Vassilis Plachouras, Iadh Ounis
ECIR3
2007 Multinomial Randomness Models for Retrieval with Document Fields
Vassilis Plachouras, Iadh Ounis
ECIR1
2007 Challenges on Distributed Web Retrieval
abstract
In the ocean of Web data, Web search engines are the primary way to access content. As the data is on the order of petabytes, current search engines are very large centralized systems based on replicated clusters. Web data, however, is always evolving. The number of Web sites continues to grow rapidly and there are currently more than 20 billion indexed pages. In the near future, centralized systems are likely to become ineffective against such a load, thus suggesting the need of fully distributed search engines. Such engines need to achieve the following goals: high quality answers, fast response time, high query throughput, and scalability. In this paper we survey and organize recent research results, outlining the main challenges of designing a distributed Web retrieval system.
Ricardo Baeza-Yates, Carlos Castillo 0001, Flavio Paiva Junqueira, Vassilis Plachouras, Fabrizio Silvestri
ICDE4
2007 The impact of caching on search engines
abstract
In this paper we study the trade-offs in designing efficient caching systems for Web search engines. We explore the impact of different approaches, such as static vs. dynamic caching, and caching query results vs.caching posting lists. Using a query log spanning a whole year we explore the limitations of caching and we demonstrate that caching posting lists can achieve higher hit rates than caching query answers. We propose a new algorithm for static caching of posting lists, which outperforms previous methods. We also study the problem of finding the optimal way to split the static cache between answers and posting lists. Finally, we measure how the changes in the query log affect the effectiveness of static caching, given our observation that the distribution of the queries changes slowly over time. Our results and observations are applicable to different levels of the data-access hierarchy, for instance, for a memory/disk layer or a broker/remote server layer.
Ricardo Baeza-Yates, Aristides Gionis, Flavio Paiva Junqueira, Vanessa Murdock 0001, Vassilis Plachouras, Fabrizio Silvestri
SIGIR5
2007 Incorporating term dependency in the dfr framework
abstract
Term dependency, or co-occurrence, has been studied in language modelling, for instance by Metzler & Croft who showed that retrieval performance could be significantlyenhanced using term dependency information. In this work, weshow how term dependency can be modelled within the Divergence From Randomness (DFR) framework. We evaluate our term dependency model on the two adhoc retrieval tasks using the TREC .GOV2 Terabyte collection. Furthermore, we examine the effect of varying the term dependency window size on the retrieval performance of the proposed model. Our experiments show that term dependency can indeed besuccessfully incorporated within the DFR framework.
Jie Peng 0003, Craig Macdonald, Vassilis Plachouras, Iadh Ounis
SIGIR4
2007 Admission Policies for Caches of Search Engine Results
Ricardo Baeza-Yates, Flavio Paiva Junqueira, Vassilis Plachouras, Hans Friedrich Witschel
SPIRE3
2007 Performance analysis of distributed information retrieval architectures using an improved network simulation model
Fidel Cacheda, Victor Carneiro, Vassilis Plachouras, Iadh Ounis
Inf. Process. Manag.3
2006 A decision mechanism for the selective combination of evidence in topic distillation
Vassilis Plachouras, Fidel Cacheda, Iadh Ounis
Inf. Retr.1
2005 Network Analysis for Distributed Information Retrieval Architectures
Fidel Cacheda, Victor Carneiro, Vassilis Plachouras, Iadh Ounis
ECIR3
2005 Terrier Information Retrieval Platform
Iadh Ounis, Gianni Amati, Vassilis Plachouras, Craig Macdonald, Douglas Johnson
ECIR3
2005 A case study of distributed information retrieval architectures to index one terabyte of text
Fidel Cacheda, Vassilis Plachouras, Iadh Ounis
Inf. Process. Manag.2
2005 Dempster-Shafer Theory for a Query-Biased Combination of Evidence on the Web
Vassilis Plachouras, Iadh Ounis
Inf. Retr.1
2005 The Static Absorbing Model for the Web
Vassilis Plachouras, Iadh Ounis, Gianni Amati
J. Web Eng.1
2004 Performance Analysis of Distributed Architectures to Index One Terabyte of Text
Fidel Cacheda, Vassilis Plachouras, Iadh Ounis
ECIR2
2004 Usefulness of hyperlink structure for query-biased topic distillation
abstract
In this paper, we introduce an information theoretic method for estimating the usefulness of the hyperlink structure induced from the set of retrieved documents. We evaluate the effectiveness of this method in the context of an optimal Bayesian decision mechanism, which selects the most appropriate retrieval approaches on a per-query basis for two TREC tasks. The estimation of the hyperlink structure's usefulness is stable when we use different weighting schemes, or when we employ sampling of documents to reduce the computational overhead. Next, we evaluate the effectiveness of the hyperlink structure's usefulness in a realistic setting, by setting the thresholds of a decision mechanism automatically. Our results show that improvements over the baselines are obtained.
Vassilis Plachouras, Iadh Ounis
SIGIR1