VLDB 2026 Research / reviewers in the wild / expert
Vasileios Kandylas
dblp:07/6901 · also Vasilis Kandylas
· DBLP profile ↗
16ranked-venue papers
5as first author
0since 2021 · last 2020
0009-0000-0369-4034ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 55% Data mining · 32% Machine learning and data management · 14% | |
| Artificial intelligence
1 paper |
Learning theory · 50% Reinforcement learning · 50% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
crowdsourcing |
0.2 | 1 | 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016 |
Machine learning and data management › data annotation
label collection |
0.2 | 1 | 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016 |
Information retrieval › information filtering › technology-assisted review
stopping criteria |
0.2 | 1 | 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.2 | 1 | 2013 | Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013 |
Machine learning › Learning theory
online learning |
0.2 | 1 | 2013 | Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013 |
Collaborative and social computing
crowdsourcing |
0.2 | 1 | 2013 | Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013 |
Information retrieval › interactive information retrieval
search and browsing |
0.1 | 1 | 2019 | Social Knowledge Graph Explorer · SIGIR 2019 |
Data mining › anomaly detection
spam detection |
0.1 | 1 | 2010 | The utility of tweeted URLs for web search · WWW 2010 |
Information retrieval
web search |
0.1 | 1 | 2010 | The utility of tweeted URLs for web search · WWW 2010 |
Information retrieval
evaluation |
0.1 | 1 | 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016 |
Information retrieval › evaluation › test collection
ground truth creation |
0.1 | 1 | 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016 |
Data mining
clustering |
0.1 | 1 | 2007 | Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007 |
Data mining › structured data mining › graph mining
community detection |
0.1 | 1 | 2007 | Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007 |
Data mining › structured data mining
graph mining |
0.1 | 1 | 2007 | Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007 |
Methods — techniques the papers use, named apart from their topics
bandit algorithms · 0.3adaptive sampling · 0.3worker quality scores · 0.2adaptive exploration · 0.2URL feature analysis · 0.1predictive modeling · 0.1citation analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Answering recreational web searches with relevant things to do results
Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay, Stewart Whiting |
Inf. Process. Manag. | 2 |
| 2019 | Social Knowledge Graph ExplorerabstractWe present SKG Explorer, an application for querying and browsing a social knowledge graph derived from Twitter that contains relationships between entities, links, and topics. A temporal dimension is also added for generating timelines for well-known events that allows the construction of stories in a wiki-like style. In this paper we describe the main components of the system and showcase some examples. Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay |
SIGIR | 2 |
| 2018 | Automatic Story Evolution Wikification from Social Data
Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay |
ICWSM | 2 |
| 2018 | Urban Maps of Social Activity
Stewart Whiting, Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay |
ICWSM | 3 |
| 2016 | e#: Sharper Expertise Detection from MicroblogsabstractMicroblogging platforms such as Twitter provide low cost access to an immense reserve of authoritative professionals, opinion leaders and hobbyists for a wide range of topics. Yet, as microposts are short and incredibly diverse, many of these experts are hidden. In this paper, we present e#, a system to retrieve experts automatically for a given set of keywords. Our design targets exhaustivity: e# can detect previously undetectable experts. The core idea is to enhance a state-ofthe-art expert detection algorithm with a graph of expertise domains. Our system produces this graph from hundreds of Gigabytes of Web search query logs and behavioral data, processed in a distributed, parallel fashion. We provide a detailed description of our architecture, including an original SQL-based community detection algorithm. We then benchmark our system with 750 queries, using crowdsourcing. We observe that e# finds many more experts than a state-of-the-art baseline. Thibault Sellam, Martin Hentschel 0001, Vasileios Kandylas, Omar Alonso |
EDBT | 3 |
| 2016 | How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality LabelsabstractCrowdsourcing has been part of the IR toolbox as a cheap and fast mechanism to obtain labels for system development and evaluation. Successful deployment of crowdsourcing at scale involves adjusting many variables, a very important one being the number of workers needed per human intelligence task (HIT). We consider the crowdsourcing task of learning the answer to simple multiple-choice HITs, which are representative of many relevance experiments. In order to provide statistically significant results, one often needs to ask multiple workers to answer the same HIT. A stopping rule is an algorithm that, given a HIT, decides for any given set of worker answers to stop and output an answer or iterate and ask one more worker. In contrast to other solutions that try to estimate worker performance and answer at the same time, our approach assumes the historical performance of a worker is known and tries to estimate the HIT difficulty and answer at the same time. The difficulty of the HIT decides how much weight to give to each worker's answer. In this paper we investigate how to devise better stopping rules given workers' performance quality scores. We suggest adaptive exploration as a promising approach for scalable and automatic creation of ground truth. We conduct a data analysis on an industrial crowdsourcing platform, and use the observations from this analysis to design new stopping rules that use the workers' quality scores in a non-trivial manner. We then perform a number of experiments using real-world datasets and simulated data, showing that our algorithm performs better than other approaches. Ittai Abraham, Omar Alonso, Vasileios Kandylas, Rajesh Patel, Steven Shelford, Aleksandrs Slivkins |
SIGIR | 3 |
| 2015 | CrowdSTAR: A Social Task Routing Framework for Online Communities
Besmira Nushi, Omar Alonso, Martin Hentschel 0001, Vasileios Kandylas |
ICWE | 4 |
| 2014 | Using Worker Quality Scores to Improve Stopping RulesabstractWe consider the crowdsourcing task of learning the answer to simple multiple-choice microtasks. In order to provide statistically significant results, one often needs to ask multiple workers to answer the same microtask. A stopping rule is an algorithm that for a given microtask decides for any given set of worker answers if the system should stop and output an answer or iterate and ask one more worker. A quality score for a worker is a score that reflects the historic performance of that worker. In this paper we investigate how to devise better stopping rules given such quality scores. We conduct a data analysis on a large-scale industrial crowdsourcing platform, and use the observations from this analysis to design new stopping rules that use the workers’ quality scores in a non-trivial manner. We then conduct a simulation based on a real-world workload, showing that our algorithm performs better than the more naive approaches. Ittai Abraham, Omar Alonso, Vasileios Kandylas, Rajesh Patel, Steven Shelford, Aleksandrs Slivkins |
HCOMP | 3 |
| 2014 | Finding Users we Trust: Scaling up Verified Twitter Users Using their Communication Patterns
Martin Hentschel 0001, Omar Alonso, Scott Counts, Vasileios Kandylas |
ICWSM | 4 |
| 2013 | Adaptive Crowdsourcing Algorithms for the Bandit Survey ProblemabstractVery recently crowdsourcing has become the de facto platform for distributing and collecting human computation for a wide range of tasks and applications such as information retrieval, natural language processing and machine learning. Current crowdsourcing platforms have some limitations in the area of quality control. Most of the effort to ensure good quality has to be done by the experimenter who has to manage the number of workers needed to reach good results.We propose a simple model for adaptive quality control in crowdsourced multiple-choice tasks which we call the “bandit survey problem”. This model is related to, but technically different from the well-known multi-armed bandit problem. We present several algorithms for this problem, and support them with analysis and simulations.Our approach is based in our experience conducting relevance evaluation for a large commercial search engine. Ittai Abraham, Omar Alonso, Vasileios Kandylas, Aleksandrs Slivkins |
COLT | 3 |
| 2010 | Improving web search relevance and freshness with content previewsabstractTraditional web search engines find it challenging to achieve good search quality for recency-sensitive queries, as they are prone to delays in discovering, indexing and ranking new web pages. In this paper we introduce PreGen, an adaptive preview generation system, which is run as part of a web search engine to improve search result quality for recency-sensitive queries. PreGen uses a machine learning algorithm to classify and select live web feeds, and generates "previews" of new web pages based on the link descriptions available in these feeds. The search engine can then index and present relevant page previews as part of its search results before the pages are fetched from the web, thereby reducing end-to-end delays. Our experiments show that PreGen improves the search relevance of a state-of-the-art search engine for recency-sensitive queries by 3% and reduces the average latencies of affected documents by 50%. Siva Gurumurthy, Vasileios Kandylas, Vidhyashankar Venkataraman |
CIKM | 3 |
| 2010 | The utility of tweeted URLs for web searchabstractMicroblogging as introduced by Twitter is becoming a source of tracking real-time news. Although identifying the highest quality or most useful posts or tweets from Twitter for breaking news is still an open problem, major web search engines seem convinced of the value of such posts and have already started allocating part of their search results pages to them. In this paper, we study a different aspect of the problem for a search engine: instead of the value of the posts, we study the value of the (shortened) URLs referenced in these posts. Our results indicate that unlike frequently bookmarked URLs, which are generally of high quality, frequently tweeted URLs tend to fall in two opposite categories: they are either high in quality, or they are spam. Identifying the quality category of a URL is not trivial, but the combination of characteristics can reveal some trends. Vasileios Kandylas, Ali Dasdan |
WWW | 1 |
| 2010 | Analyzing knowledge communities using foreground and background clustersabstractInsight into the growth (or shrinkage) of “knowledge communities” of authors that build on each other's work can be gained by studying the evolution over time of clusters of documents. We cluster documents based on the documents they cite in common using the Streemer clustering method, which finds cohesive foreground clusters (the knowledge communities) embedded in a diffuse background. We build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors and use these models to study the drivers of community growth and the predictors of how widely a paper will be cited. We find that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and use narrow vocabulary and that papers that lie on the periphery of a community have the highest impact, while those not in any community have the lowest impact. Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
ACM Trans. Knowl. Discov. Data | 1 |
| 2008 | Multiway Clustering for Creating Biomedical Term SetsabstractWe present an EM-based clustering method that can be used for constructing or augmenting ontologies such as MeSH. Our algorithm simultaneously clusters verbs and nouns using both verb-noun and noun-noun co-occurrence pairs. This strategy provides greater coverage of words than using either set of pairs alone, since not all words appear in both datasets. We demonstrate it on data extracted from Medline and evaluate the results using MeSH and Wordnet. Vasileios Kandylas, Lyle H. Ungar, Ted Sandler, Shane T. Jensen |
BIBM | 1 |
| 2008 | Finding cohesive clusters for analyzing knowledge communities
Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
Knowl. Inf. Syst. | 1 |
| 2007 | Finding Cohesive Clusters for Analyzing Knowledge CommunitiesabstractDocuments and authors can be clustered into "knowledge communities" based on the overlap in the papers they cite. We introduce a new clustering algorithm, Streemer, which finds cohesive foreground clusters embedded in a diffuse background, and use it to identify knowledge communities as foreground clusters of papers which share common citations. To analyze the evolution of these communities over time, we build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors. Findings include that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and if they use a narrow vocabulary. Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar |
ICDM | 1 |