Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gleb Skobeltsyn

dblp:82/4088 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 84% Query processing and optimization · 12% Indexing and storage engines · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
distributed information retrieval
0.232008
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Query-driven indexing for peer-to-peer text retrieval · WWW 2007
Web text retrieval with a P2P query-driven index · SIGIR 2007
Information retrieval › distributed information retrieval
peer-to-peer search
0.232008
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Query-driven indexing for peer-to-peer text retrieval · WWW 2007
Web text retrieval with a P2P query-driven index · SIGIR 2007
Distributed systems
peer-to-peer systems
0.122008
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Query-driven indexing for peer-to-peer text retrieval · WWW 2007
Information retrieval
indexing
0.122008
Query-driven indexing for peer-to-peer text retrieval · WWW 2007
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Query processing and optimization
query result caching
0.112008
ResIn: a combination of results caching and index pruning for high-performance web search engines · SIGIR 2008
Distributed systems › peer-to-peer systems › overlay networks
structured overlay
0.112008
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Indexing and storage engines
distributed indexing
0.012008
AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network · Proc. VLDB Endow. 2008
Information retrieval › indexing › index compression
index pruning
0.012008
ResIn: a combination of results caching and index pruning for high-performance web search engines · SIGIR 2008

Methods — techniques the papers use, named apart from their topics

overlay network optimization · 0.2indexing mechanisms · 0.2term set indexing · 0.1posting list truncation · 0.1index pruning · 0.1caching · 0.1query log analysis · 0.1
YearPublicationVenuePosition
2016 Contextual Prediction Models for Speech Recognition
Yoni Halpern, Keith B. Hall, Vlad Schogol, Michael Riley 0001, Brian Roark, Gleb Skobeltsyn, Martin Bäuml
INTERSPEECH6
2009 Entity Search with NECESSITY
Ekaterini Ioannou, Saket Sathe 0001, Nicolas Bonvin, Anshul Jain, Srikanth Bondalapati, Gleb Skobeltsyn, Claudia Niederée, Zoltán Miklós 0001
WebDB6
2009 Query-driven indexing for scalable peer-to-peer text retrieval
Gleb Skobeltsyn, Toan Luu, Ivana Podnar Zarko, Martin Rajman, Karl Aberer
Future Gener. Comput. Syst.1
2008 ResIn: a combination of results caching and index pruning for high-performance web search engines
abstract
Results caching is an efficient technique for reducing the query processing load, hence it is commonly used in real search engines. This technique, however, bounds the maximum hit rate due to the large fraction of singleton queries, which is an important limitation. In this paper we propose ResIn - an architecture that uses a combination of results caching and index pruning to overcome this limitation.
Gleb Skobeltsyn, Flavio Paiva Junqueira, Vassilis Plachouras, Ricardo Baeza-Yates
SIGIR1
2008 AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network
abstract
In this paper we present the AlvisP2P IR engine, which enables efficient retrieval with multi-keyword queries from a global document collection available in a P2P network. In such a network, each peer publishes its local index and invests a part of its local computing resources (storage, CPU, bandwidth) to maintain a fraction of a global P2P index. This investment is rewarded by the network-wide accessibility of the local documents via the global search facility. The AlvisP2P engine uses an optimized overlay network and relies on novel indexing/retrieval mechanisms that ensure low bandwidth consumption, thus enabling unlimited network growth. Our demonstration shows how an easy-to-install AlvisP2P client can be used to join an existing P2P network, index local (text or even multimedia) documents with collection-specific indexing mechanisms, and control access rights to them.
Toan Luu, Gleb Skobeltsyn, Fabius Klemm, Maroje Puh, Ivana Podnar Zarko, Martin Rajman, Karl Aberer
Proc. VLDB Endow.2
2007 Web text retrieval with a P2P query-driven index
abstract
In this paper, we present a query-driven indexing/retrieval strategy for efficient full text retrieval from large document collections distributed within a structured P2P network. Our indexing strategy is based on two important properties: (1) the generated distributed index stores posting lists for carefully chosen indexing term combinations, and (2) the posting lists containing too many document references are truncated to a bounded number of their top-ranked elements. These two properties guarantee acceptable storage and bandwidth requirements, essentially because the number of indexing term combinations remains scalable and the transmitted posting lists never exceed a constant size. However, as the number of generated term combinations can still become quite large, we also use term statistics extracted from available query logs to index only such combinations that are frequently present in user queries. Thus, by avoiding the generation of superfluous indexing term combinations, we achieve an additional substantial reduction in bandwidth and storage consumption. As a result, the generated distributed index corresponds to a constantly evolving query-driven indexing structure that efficiently follows current information needs of the users. More precisely, our theoretical analysis and experimental results indicate that, at the price of a marginal loss in retrieval quality for rare queries, the generated index size and network traffic remain manageable even for web-size document collections. Furthermore, our experiments show that at the same time the achieved retrieval quality is fully comparable to the one obtained with a state-of-the-art centralized query engine.
Gleb Skobeltsyn, Toan Luu, Ivana Podnar Zarko, Martin Rajman, Karl Aberer
SIGIR1
2007 Query-driven indexing for peer-to-peer text retrieval
abstract
We describe a query-driven indexing framework for scalable text retrieval over structured P2P networks. To cope with the bandwidth consumption problem that has been identified as the major obstacle for full-text retrieval in P2P networks, we truncate posting lists associated with indexing features to a constant size storing only top-k ranked document references. To compensate for the loss of information caused by the truncation, we extend the set of indexing features with carefully chosen term sets. Indexing term sets are selected based on the query statistics extracted from query logs, thus we index only such combinations that are a) frequently present in user queries and b) non-redundant w.r.t the rest of the index. The distributed index is compact and efficient as it constantly evolves adapting to the current query popularity distribution. Moreover, it is possible to control the tradeoff between the storage/bandwidth requirements and the quality of query answering by tuning the indexing parameters. Our theoretical analysis and experimental results indicate that we can indeed achieve scalable P2P text retrieval for very large document collections and deliver good retrieval performance.
Gleb Skobeltsyn, Toan Luu, Karl Aberer, Martin Rajman, Ivana Podnar Zarko
WWW1