Vasileios Kandylas

dblp:07/6901 · also Vasilis Kandylas · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
0since 2021 · last 2020
0009-0000-0369-4034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 55% Data mining · 32% Machine learning and data management · 14%
Artificial intelligence
1 paper
Learning theory · 50% Reinforcement learning · 50%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
crowdsourcing
0.212016
How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016
Machine learning and data management › data annotation
label collection
0.212016
How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016
Information retrieval › information filtering › technology-assisted review
stopping criteria
0.212016
How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016
Machine learning › Reinforcement learning
multi-armed bandit
0.212013
Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013
Machine learning › Learning theory
online learning
0.212013
Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013
Collaborative and social computing
crowdsourcing
0.212013
Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem · COLT 2013
Information retrieval › interactive information retrieval
search and browsing
0.112019
Social Knowledge Graph Explorer · SIGIR 2019
Data mining › anomaly detection
spam detection
0.112010
The utility of tweeted URLs for web search · WWW 2010
Information retrieval
web search
0.112010
The utility of tweeted URLs for web search · WWW 2010
Information retrieval
evaluation
0.112016
How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016
Information retrieval › evaluation › test collection
ground truth creation
0.112016
How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels · SIGIR 2016
Data mining
clustering
0.112007
Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007
Data mining › structured data mining › graph mining
community detection
0.112007
Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007
Data mining › structured data mining
graph mining
0.112007
Finding Cohesive Clusters for Analyzing Knowledge Communities · ICDM 2007

Methods — techniques the papers use, named apart from their topics

bandit algorithms · 0.3adaptive sampling · 0.3worker quality scores · 0.2adaptive exploration · 0.2URL feature analysis · 0.1predictive modeling · 0.1citation analysis · 0.1
YearPublicationVenuePosition
2020 Answering recreational web searches with relevant things to do results
Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay, Stewart Whiting
Inf. Process. Manag.2
2019 Social Knowledge Graph Explorer
abstract
We present SKG Explorer, an application for querying and browsing a social knowledge graph derived from Twitter that contains relationships between entities, links, and topics. A temporal dimension is also added for generating timelines for well-known events that allows the construction of stories in a wiki-like style. In this paper we describe the main components of the system and showcase some examples.
Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay
SIGIR2
2018 Automatic Story Evolution Wikification from Social Data
Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay
ICWSM2
2018 Urban Maps of Social Activity
Stewart Whiting, Omar Alonso, Vasileios Kandylas, Serge-Eric Tremblay
ICWSM3
2016 e#: Sharper Expertise Detection from Microblogs
abstract
Microblogging platforms such as Twitter provide low cost access to an immense reserve of authoritative professionals, opinion leaders and hobbyists for a wide range of topics. Yet, as microposts are short and incredibly diverse, many of these experts are hidden. In this paper, we present e#, a system to retrieve experts automatically for a given set of keywords. Our design targets exhaustivity: e# can detect previously undetectable experts. The core idea is to enhance a state-ofthe-art expert detection algorithm with a graph of expertise domains. Our system produces this graph from hundreds of Gigabytes of Web search query logs and behavioral data, processed in a distributed, parallel fashion. We provide a detailed description of our architecture, including an original SQL-based community detection algorithm. We then benchmark our system with 750 queries, using crowdsourcing. We observe that e# finds many more experts than a state-of-the-art baseline.
Thibault Sellam, Martin Hentschel 0001, Vasileios Kandylas, Omar Alonso
EDBT3
2016 How Many Workers to Ask?: Adaptive Exploration for Collecting High Quality Labels
abstract
Crowdsourcing has been part of the IR toolbox as a cheap and fast mechanism to obtain labels for system development and evaluation. Successful deployment of crowdsourcing at scale involves adjusting many variables, a very important one being the number of workers needed per human intelligence task (HIT). We consider the crowdsourcing task of learning the answer to simple multiple-choice HITs, which are representative of many relevance experiments. In order to provide statistically significant results, one often needs to ask multiple workers to answer the same HIT. A stopping rule is an algorithm that, given a HIT, decides for any given set of worker answers to stop and output an answer or iterate and ask one more worker. In contrast to other solutions that try to estimate worker performance and answer at the same time, our approach assumes the historical performance of a worker is known and tries to estimate the HIT difficulty and answer at the same time. The difficulty of the HIT decides how much weight to give to each worker's answer. In this paper we investigate how to devise better stopping rules given workers' performance quality scores. We suggest adaptive exploration as a promising approach for scalable and automatic creation of ground truth. We conduct a data analysis on an industrial crowdsourcing platform, and use the observations from this analysis to design new stopping rules that use the workers' quality scores in a non-trivial manner. We then perform a number of experiments using real-world datasets and simulated data, showing that our algorithm performs better than other approaches.
Ittai Abraham, Omar Alonso, Vasileios Kandylas, Rajesh Patel, Steven Shelford, Aleksandrs Slivkins
SIGIR3
2015 CrowdSTAR: A Social Task Routing Framework for Online Communities
Besmira Nushi, Omar Alonso, Martin Hentschel 0001, Vasileios Kandylas
ICWE4
2014 Using Worker Quality Scores to Improve Stopping Rules
abstract
We consider the crowdsourcing task of learning the answer to simple multiple-choice microtasks. In order to provide statistically significant results, one often needs to ask multiple workers to answer the same microtask. A stopping rule is an algorithm that for a given microtask decides for any given set of worker answers if the system should stop and output an answer or iterate and ask one more worker. A quality score for a worker is a score that reflects the historic performance of that worker. In this paper we investigate how to devise better stopping rules given such quality scores. We conduct a data analysis on a large-scale industrial crowdsourcing platform, and use the observations from this analysis to design new stopping rules that use the workers’ quality scores in a non-trivial manner. We then conduct a simulation based on a real-world workload, showing that our algorithm performs better than the more naive approaches.
Ittai Abraham, Omar Alonso, Vasileios Kandylas, Rajesh Patel, Steven Shelford, Aleksandrs Slivkins
HCOMP3
2014 Finding Users we Trust: Scaling up Verified Twitter Users Using their Communication Patterns
Martin Hentschel 0001, Omar Alonso, Scott Counts, Vasileios Kandylas
ICWSM4
2013 Adaptive Crowdsourcing Algorithms for the Bandit Survey Problem
abstract
Very recently crowdsourcing has become the de facto platform for distributing and collecting human computation for a wide range of tasks and applications such as information retrieval, natural language processing and machine learning. Current crowdsourcing platforms have some limitations in the area of quality control. Most of the effort to ensure good quality has to be done by the experimenter who has to manage the number of workers needed to reach good results.We propose a simple model for adaptive quality control in crowdsourced multiple-choice tasks which we call the “bandit survey problem”. This model is related to, but technically different from the well-known multi-armed bandit problem. We present several algorithms for this problem, and support them with analysis and simulations.Our approach is based in our experience conducting relevance evaluation for a large commercial search engine.
Ittai Abraham, Omar Alonso, Vasileios Kandylas, Aleksandrs Slivkins
COLT3
2010 Improving web search relevance and freshness with content previews
abstract
Traditional web search engines find it challenging to achieve good search quality for recency-sensitive queries, as they are prone to delays in discovering, indexing and ranking new web pages. In this paper we introduce PreGen, an adaptive preview generation system, which is run as part of a web search engine to improve search result quality for recency-sensitive queries. PreGen uses a machine learning algorithm to classify and select live web feeds, and generates "previews" of new web pages based on the link descriptions available in these feeds. The search engine can then index and present relevant page previews as part of its search results before the pages are fetched from the web, thereby reducing end-to-end delays. Our experiments show that PreGen improves the search relevance of a state-of-the-art search engine for recency-sensitive queries by 3% and reduces the average latencies of affected documents by 50%.
Siva Gurumurthy, Vasileios Kandylas, Vidhyashankar Venkataraman
CIKM3
2010 The utility of tweeted URLs for web search
abstract
Microblogging as introduced by Twitter is becoming a source of tracking real-time news. Although identifying the highest quality or most useful posts or tweets from Twitter for breaking news is still an open problem, major web search engines seem convinced of the value of such posts and have already started allocating part of their search results pages to them. In this paper, we study a different aspect of the problem for a search engine: instead of the value of the posts, we study the value of the (shortened) URLs referenced in these posts. Our results indicate that unlike frequently bookmarked URLs, which are generally of high quality, frequently tweeted URLs tend to fall in two opposite categories: they are either high in quality, or they are spam. Identifying the quality category of a URL is not trivial, but the combination of characteristics can reveal some trends.
Vasileios Kandylas, Ali Dasdan
WWW1
2010 Analyzing knowledge communities using foreground and background clusters
abstract
Insight into the growth (or shrinkage) of “knowledge communities” of authors that build on each other's work can be gained by studying the evolution over time of clusters of documents. We cluster documents based on the documents they cite in common using the Streemer clustering method, which finds cohesive foreground clusters (the knowledge communities) embedded in a diffuse background. We build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors and use these models to study the drivers of community growth and the predictors of how widely a paper will be cited. We find that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and use narrow vocabulary and that papers that lie on the periphery of a community have the highest impact, while those not in any community have the lowest impact.
Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar
ACM Trans. Knowl. Discov. Data1
2008 Multiway Clustering for Creating Biomedical Term Sets
abstract
We present an EM-based clustering method that can be used for constructing or augmenting ontologies such as MeSH. Our algorithm simultaneously clusters verbs and nouns using both verb-noun and noun-noun co-occurrence pairs. This strategy provides greater coverage of words than using either set of pairs alone, since not all words appear in both datasets. We demonstrate it on data extracted from Medline and evaluate the results using MeSH and Wordnet.
Vasileios Kandylas, Lyle H. Ungar, Ted Sandler, Shane T. Jensen
BIBM1
2008 Finding cohesive clusters for analyzing knowledge communities
Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar
Knowl. Inf. Syst.1
2007 Finding Cohesive Clusters for Analyzing Knowledge Communities
abstract
Documents and authors can be clustered into "knowledge communities" based on the overlap in the papers they cite. We introduce a new clustering algorithm, Streemer, which finds cohesive foreground clusters embedded in a diffuse background, and use it to identify knowledge communities as foreground clusters of papers which share common citations. To analyze the evolution of these communities over time, we build predictive models with features based on the citation structure, the vocabulary of the papers, and the affiliations and prestige of the authors. Findings include that scientific knowledge communities tend to grow more rapidly if their publications build on diverse information and if they use a narrow vocabulary.
Vasileios Kandylas, S. Phineas Upham, Lyle H. Ungar
ICDM1