VLDB 2026 Research / reviewers in the wild / expert
Kostas Tsioutsiouliklis
dblp:18/2174
· DBLP profile ↗
19ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0002-2505-653XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Information retrieval · 52% Web and social media mining · 30% Data mining · 12% | |
| Theoretical computer science
5 papers |
Graph algorithms and graph theory · 48% Algorithms and data structures · 28% Approximation and online algorithms · 13% | |
| Artificial intelligence
4 papers |
Information extraction and text analysis · 46% Representation and self-supervised learning · 40% Transfer learning and domain adaptation · 8% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › document retrieval › temporal information retrieval
new event detection |
0.5 | 1 | 2021 | A General Framework for First Story Detection Utilizing Entities and Their Relations · IEEE Trans. Knowl. Data Eng. 2021 |
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification |
0.4 | 1 | 2019 | Hierarchical Transfer Learning for Multi-label Text Classification · ACL (1) 2019 |
Graph algorithms and graph theory › graph connectivity
connected components |
0.3 | 1 | 2018 | Shortcutting Label Propagation for Distributed Connected Components · WSDM 2018 |
Graph algorithms and graph theory
graph mining |
0.3 | 1 | 2018 | Shortcutting Label Propagation for Distributed Connected Components · WSDM 2018 |
Machine learning › Representation and self-supervised learning › contrastive learning
negative sampling |
0.3 | 1 | 2017 | Distributed Negative Sampling for Word Embeddings · AAAI 2017 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.3 | 1 | 2017 | Distributed Negative Sampling for Word Embeddings · AAAI 2017 |
Web and social media mining › misinformation detection
clickbait detection |
0.2 | 1 | 2016 | "8 Amazing Secrets for Getting More Clicks": Detecting Clickbaits in News Streams Using Article Informality · AAAI 2016 |
Algorithms and data structures › parallel algorithms
mapreduce algorithms |
0.2 | 1 | 2015 | Set Cover at Web Scale · KDD 2015 |
Algorithms and data structures
parallel algorithms |
0.2 | 1 | 2015 | Set Cover at Web Scale · KDD 2015 |
Approximation and online algorithms
set cover |
0.2 | 1 | 2015 | Set Cover at Web Scale · KDD 2015 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.1 | 1 | 2021 | A General Framework for First Story Detection Utilizing Entities and Their Relations · IEEE Trans. Knowl. Data Eng. 2021 |
Web and social media mining › location-based social network
geographical topic discovery |
0.1 | 1 | 2012 | Discovering geographical topics in the twitter stream · WWW 2012 |
Web and social media mining › location-based social network
geotagged social media analysis |
0.1 | 1 | 2012 | Discovering geographical topics in the twitter stream · WWW 2012 |
Web and social media mining
location estimation |
0.1 | 1 | 2012 | Discovering geographical topics in the twitter stream · WWW 2012 |
Recommender systems
user profiling |
0.1 | 1 | 2012 | Discovering geographical topics in the twitter stream · WWW 2012 |
Computational social science and digital humanities
social media analysis |
0.1 | 1 | 2011 | Linguistic Redundancy in Twitter · EMNLP 2011 |
Data mining
text mining |
0.1 | 1 | 2011 | A time-dependent topic model for multiple text streams · KDD 2011 |
Data mining › text mining
topic modeling |
0.1 | 1 | 2011 | A time-dependent topic model for multiple text streams · KDD 2011 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
hierarchical transfer learning |
0.1 | 1 | 2019 | Hierarchical Transfer Learning for Multi-label Text Classification · ACL (1) 2019 |
Information retrieval › evaluation › effectiveness metrics
relevance measure |
0.1 | 1 | 2010 | Web search engine metrics: (direct metrics to measure user satisfaction) · WWW 2010 |
Information retrieval
retrieval evaluation |
0.1 | 1 | 2010 | Web search engine metrics: (direct metrics to measure user satisfaction) · WWW 2010 |
Information retrieval
search engines |
0.1 | 1 | 2010 | Optimal Web-Scale Tiering as a Flow Problem · NIPS 2010 |
Information retrieval › retrieval evaluation
user satisfaction metrics |
0.1 | 1 | 2010 | Web search engine metrics: (direct metrics to measure user satisfaction) · WWW 2010 |
Graph algorithms and graph theory › graph algorithms › network flow
maximum flow |
0.1 | 1 | 2010 | Optimal Web-Scale Tiering as a Flow Problem · NIPS 2010 |
Information retrieval
web search |
0.1 | 2 | 2015 | Set Cover at Web Scale · KDD 2015 Using web structure for classifying and describing web pages · WWW 2002 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2017 | Distributed Negative Sampling for Word Embeddings · AAAI 2017 |
Information retrieval › web search › web information retrieval
web content quality |
0.1 | 1 | 2016 | "8 Amazing Secrets for Getting More Clicks": Detecting Clickbaits in News Streams Using Article Informality · AAAI 2016 |
Information retrieval › search engines › web crawling
focused crawling |
0.0 | 1 | 2003 | Evolving Strategies for Focused Web Crawling · ICML 2003 |
Information retrieval › search engines
web crawling |
0.0 | 1 | 2003 | Evolving Strategies for Focused Web Crawling · ICML 2003 |
Data mining › text mining › text classification
web page classification |
0.0 | 1 | 2002 | Using web structure for classifying and describing web pages · WWW 2002 |
Methods — techniques the papers use, named apart from their topics
relation extraction · 1.0named entity recognition · 1.0mapreduce · 0.4greedy algorithm · 0.4transfer learning · 0.4attention · 0.4GRU · 0.4shortcutting · 0.3label propagation · 0.3distributed algorithm · 0.3machine learning · 0.2feature engineering · 0.2corpus analysis · 0.2stochastic gradient descent · 0.2lagrangian relaxation · 0.2integer linear programming · 0.2sparse factorial coding · 0.1markov model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Leveraging Large Language Models for Improving Keyphrase Generation for Contextual TargetingabstractGenerating a set of keyphrases that convey the main concepts discussed in a document has been applied to improve various applications including document retrieval and online advertising. The state-of-the-art approaches mostly rely on the neural sequence-to-sequence framework to generate keyphrases. However, training such deep neural networks either requires a significant amount of human efforts in obtaining ground truth keyphrases or suffers from lower quality training data derived from weakly supervised signals. More recently, pre-trained language models are fine-tuned to build more data-efficient keyphrase generation models. Yet, the documents often need to be truncated to adapt to the pre-trained context window. On the other hand, large language models (LLMs) have demonstrated impressive abilities in understanding very long text and generating answers for a wide range of natural language processing tasks, making them great candidates for improving keyphrase generation. There however is a lack of a systematic study on how to use LLMs, especially in an industrial setting that requires low generation latency. In this work, we present an empirical study to facilitate a more informed use of LLMs for keyphrase generation. We compare zero-shot and few-shot in-context learning with parameter efficient fine-tuning using a number of open-source LLMs. We show that using only a handful of well selected human annotated samples, the LLMs already outperform the fine-tuned language model baselines. When thousands of human labeled samples are available, fine-tuned large language models significantly improve the amount and the quality of the generated keyphrases. To enable efficient keyphrase generation at scale, we distill the knowledge from LLMs to a base-size language model. Our evaluation shows significant increase in user reach when the generated keyphrases are used for contextual targeting at Yahoo. Xiao Bai 0002, Ivan Stojkovic, Kostas Tsioutsiouliklis |
CIKM | 4 |
| 2023 | Precision/Recall on Imbalanced Test DataabstractIn this paper we study the problem of estimating accurately the precision and recall for binary classification when the classes are imbalanced and only a limited number of human labels are available. One common strategy is to over-sample the small positive class predicted by the classifier. Rather than random sampling where the values in a confusion matrix are observations coming from a multinomial distribution, we over-sample the minority positive class predicted by the classifier, resulting in two independent binomial distributions. But how much should we over-sample? And what confidence/credible intervals can we deduce based on our over-sampling? We provide formulas for (1) the confidence intervals of the adjusted precision/recall after over-sampling; (2) Bayesian credible intervals of adjusted precision/recall. For precision, the higher the over-sampling rate, the narrower the confidence/credible interval. For recall, there exists an optimal over-sampling ratio, which minimizes the width of the confidence/credible interval. Also, we present experiments on synthetic data and real data to demonstrate the capability of our method to construct accurate intervals. Finally, we demonstrate how we can apply our techniques to Yahoo mail’s quality monitoring system. Hongwei Shang 0001, Jean-Marc Langlois, Kostas Tsioutsiouliklis, Changsung Kang |
AISTATS | 3 |
| 2021 | A General Framework for First Story Detection Utilizing Entities and Their RelationsabstractNews portals, such as Yahoo News or Google News, collect large amounts of news articles from a variety of sources on a daily basis. Only a small portion of these documents can be selected and displayed on the homepage. Thus, there is a strong preference for major, recent events. In this work, we propose a scalable First Story Detection (FSD) pipeline that identifies fresh news. This pipeline is used in order to instantiate a variety of FSD approaches. In addition we suggest a novel FSD technique that in comparison to existing systems, relies on relation extraction algorithms and exploits the named entities and their relations in order to decide about the freshness of an article. We evaluate our technique by instantiating existing state of art FSD techniques within our generic pipeline. As ground truth we use multiple datasets that cover different categories. Experimental results demonstrate that our FSD method in many cases provides an improvement over state-of-the-art techniques. In addition, we show using a large synthetic dataset that our general FSD pipeline has constant space and time requirements and is suitable for very high volume streams. Nikolaos Panagiotou, Cem Akkaya, Kostas Tsioutsiouliklis, Vana Kalogeraki, Dimitrios Gunopulos |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Hierarchical Transfer Learning for Multi-label Text ClassificationabstractMulti-Label Hierarchical Text Classification (MLHTC) is the task of categorizing documents into one or more topics organized in an hierarchical taxonomy.MLHTC can be formulated by combining multiple binary classification problems with an independent classifier for each category.We propose a novel transfer learning based strategy, HTrans, where binary classifiers at lower levels in the hierarchy are initialized using parameters of the parent classifier and fine-tuned on the child category classification task.In HTrans, we use a Gated Recurrent Unit (GRU)-based deep learning architecture coupled with attention.Compared to binary classifiers trained from scratch, our HTrans approach results in significant improvements of 1% on micro-F1 and 3% on macro-F1 on the RCV1 dataset.Our experiments also show that binary classifiers trained from scratch are significantly better than single multi-label models. Siddhartha Banerjee, Cem Akkaya, Francisco Perez-Sorrosal, Kostas Tsioutsiouliklis |
ACL (1) | 4 |
| 2018 | Identifying Domain Independent Update Intents in Task Based DialogsabstractOne important problem in task-based conversations is that of effectively updating the belief estimates of user-mentioned slot-value pairs.Given a user utterance, the intent of a slot-value pair is captured using dialog acts (DA) expressed in that utterance.However, in certain cases, DA's fail to capture the actual update intent of the user.In this paper, we describe such cases and propose a new type of semantic class for user intents.This new type, Update Intents (UI), is directly related to the type of update a user intends to perform for a slot-value pair.We define five types of UI's, which are independent of the domain of the conversation.We build a multi-class classification model using LSTM's to identify the type of UI in user utterances in the Restaurant and Shopping domains.Experimental results show that our models achieve strong classification performance in terms of F-1 score. Prakhar Biyani, Cem Akkaya, Kostas Tsioutsiouliklis |
SIGDIAL Conference | 3 |
| 2018 | Shortcutting Label Propagation for Distributed Connected ComponentsabstractConnected Components is a fundamental graph mining problem that has been studied for the PRAM, MapReduce and BSP models. We present a simple CC algorithm for BSP that does not mutate the graph, converges in O(log n) supersteps and scales to graphs of trillions of edges. Stergios Stergiou, Dipen Rughwani, Kostas Tsioutsiouliklis |
WSDM | 3 |
| 2017 | Distributed Negative Sampling for Word EmbeddingsabstractWord2Vec recently popularized dense vector word representations as fixed-length features for machine learning algorithms and is in widespread use today. In this paper we investigate one of its core components, Negative Sampling, and propose efficient distributed algorithms that allow us to scale to vocabulary sizes of more than 1 billion unique words and corpus sizes of more than 1 trillion words. Stergios Stergiou, Zygimantas Straznickas, Rolina Wu, Kostas Tsioutsiouliklis |
AAAI | 4 |
| 2016 | "8 Amazing Secrets for Getting More Clicks": Detecting Clickbaits in News Streams Using Article InformalityabstractClickbaits are articles with misleading titles, exaggerating the content on the landing page. Their goal is to entice users to click on the title in order to monetize the landing page. The content on the landing page is usually of low quality. Their presence in user homepage stream of news aggregator sites (e.g., Yahoo news, Google news) may adversely impact user experience. Hence, it is important to identify and demote or block them on homepages. In this paper, we present a machine-learning model to detect clickbaits. We use a variety of features and show that the degree of informality of a webpage (as measured by different metrics) is a strong indicator of it being a clickbait. We conduct extensive experiments to evaluate our approach and analyze properties of clickbait and non-clickbait articles. Our model achieves high performance (74.9% F-1 score) in predicting clickbaits. Prakhar Biyani, Kostas Tsioutsiouliklis, John Blackmer |
AAAI | 2 |
| 2016 | First Story Detection using Entities and RelationsabstractNews portals, such as Yahoo News or Google News, collect large amounts of documents from a variety of sources on a daily basis. Only a small portion of these documents can be selected and displayed on the homepage. Thus, there is a strong preference for major, recent events. In this work, we propose a scalable and accurate First Story Detection (FSD) pipeline that identifies fresh news. In comparison to other FSD systems, our method relies on relation extraction methods exploiting entities and their relations. We evaluate our pipeline using two distinct datasets from Yahoo News and Google News. Experimental results demonstrate that our method improves over the state-of-the-art systems on both datasets with constant space and time requirements. Nikolaos Panagiotou, Cem Akkaya, Kostas Tsioutsiouliklis, Vana Kalogeraki, Dimitrios Gunopulos |
COLING | 3 |
| 2015 | Set Cover at Web ScaleabstractThe classic Set Cover problem requires selecting a minimum size subset A ⊆ F from a family of finite subsets F Of U such that the elements covered by A are the ones covered by F. It naturally occurs in many settings in web search, web mining and web advertising. The greedy algorithm that iteratively selects a set in F that covers the most uncovered elements, yields an optimum (1+ln |U|)-approximation but is inherently sequential. In this work we give the first MapReduce Set Cover algorithm that scales to problem sizes of ∼ 1 trillion elements and runs in logp Δ iterations for a nearly optimum approximation ratio of p ln Δ, where Δ is the cardinality of the largest set in F Stergios Stergiou, Kostas Tsioutsiouliklis |
KDD | 2 |
| 2012 | Discovering geographical topics in the twitter streamabstractMicro-blogging services have become indispensable communication tools for online users for disseminating breaking news, eyewitness accounts, individual expression, and protest groups. Recently, Twitter, along with other online social networking services such as Foursquare, Gowalla, Facebook and Yelp, have started supporting location services in their messages, either explicitly, by letting users choose their places, or implicitly, by enabling geo-tagging, which is to associate messages with latitudes and longitudes. This functionality allows researchers to address an exciting set of questions: 1) How is information created and shared across geographical locations, 2) How do spatial and linguistic characteristics of people vary across regions, and 3) How to model human mobility. Although many attempts have been made for tackling these problems, previous methods are either complicated to be implemented or oversimplified that cannot yield reasonable performance. It is a challenge task to discover topics and identify users' interests from these geo-tagged messages due to the sheer amount of data and diversity of language variations used on these location sharing services. In this paper we focus on Twitter and present an algorithm by modeling diversity in tweets based on topical diversity, geographical diversity, and an interest distribution of the user. Furthermore, we take the Markovian nature of a user's location into account. Our model exploits sparse factorial coding of the attributes, thus allowing us to deal with a large and diverse set of covariates efficiently. Our approach is vital for applications such as user profiling, content recommendation and topic tracking. We show high accuracy in location estimation based on our model. Moreover, the algorithm identifies interesting topics based on location and language. Liangjie Hong, Amr Ahmed 0001, Siva Gurumurthy, Alexander J. Smola, Kostas Tsioutsiouliklis |
WWW | 5 |
| 2011 | Linguistic Redundancy in Twitter
Fabio Massimo Zanzotto, Marco Pennacchiotti, Kostas Tsioutsiouliklis |
EMNLP | 3 |
| 2011 | A time-dependent topic model for multiple text streamsabstractIn recent years social media have become indispensable tools for information dissemination, operating in tandem with traditional media outlets such as newspapers, and it has become critical to understand the interaction between the new and old sources of news. Although social media as well as traditional media have attracted attention from several research communities, most of the prior work has been limited to a single medium. In addition temporal analysis of these sources can provide an understanding of how information spreads and evolves. Modeling temporal dynamics while considering multiple sources is a challenging research problem. In this paper we address the problem of modeling text streams from two news sources - Twitter and Yahoo! News. Our analysis addresses both their individual properties (including temporal dynamics) and their inter-relationships. This work extends standard topic models by allowing each text stream to have both local topics and shared topics. For temporal modeling we associate each topic with a time-dependent function that characterizes its popularity over time. By integrating the two models, we effectively model the temporal dynamics of multiple correlated text streams in a unified framework. We evaluate our model on a large-scale dataset, consisting of text streams from both Twitter and news feeds from Yahoo! News. Besides overcoming the limitations of existing models, we show that our work achieves better perplexity on unseen data and identifies more coherent topics. We also provide analysis of finding real-world events from the topics obtained by our model. Liangjie Hong, Byron Dom, Siva Gurumurthy, Kostas Tsioutsiouliklis |
KDD | 4 |
| 2010 | Optimal Web-Scale Tiering as a Flow ProblemabstractWe present a fast online solver for large scale maximum-flow problems as they occur in portfolio optimization, inventory management, computer vision, and logistics. Our algorithm solves an integer linear program in an online fashion. It exploits total unimodularity of the constraint matrix and a Lagrangian relaxation to solve the problem as a convex online game. The algorithm generates approximate solutions of max-flow problems by performing stochastic gradient descent on a set of flows. We apply the algorithm to optimize tier arrangement of over 80 Million web pages on a layered set of caches to serve an incoming query stream optimally. We provide an empirical demonstration of the effectiveness of our method on real query-pages data. Gilbert Leung, Novi Quadrianto, Alexander J. Smola, Kostas Tsioutsiouliklis |
NIPS | 4 |
| 2010 | Web search engine metrics: (direct metrics to measure user satisfaction)abstractSearch engines are important resources for finding information on the Web. They are also important for publishers and advertisers to present their content to users. Thus, user satisfaction is key and must be quantified. In this tutorial, we give a practical review of web search metrics from a user satisfaction point of view. We cover metrics for relevance, comprehensiveness, coverage, diversity, discovery freshness, content freshness, and presentation. We will also describe how these metrics can be mapped to proxy metrics for the stages of a generic search engine pipeline. The practitioners can apply these metrics readily and the researchers can get motivation for new problems to work on, especially in formalizing and refining metrics. Ali Dasdan, Kostas Tsioutsiouliklis, Emre Velipasaoglu |
WWW | 2 |
| 2003 | Evolving Strategies for Focused Web Crawling
Judy Johnson, Kostas Tsioutsiouliklis, C. Lee Giles |
ICML | 2 |
| 2002 | Using web structure for classifying and describing web pagesabstractThe structure of the web is increasingly being used to improve organization, search, and analysis of information on the web. For example, Google uses the text in citing documents (documents that link to the target document) for search. We analyze the relative utility of document text, and the text in citing documents near the citation, for classification and description. Results show that the text in citing documents, when available, often has greater discriminative and descriptive power than the text in the target document itself. The combination of evidence from a document and citing documents can improve on either information source alone. Moreover, by ranking words and phrases in the citing documents according to expected entropy loss, we are able to accurately name clusters of web pages, even with very few positive examples. Our results confirm, quantify, and extend previous research using web structure in these areas, introducing new methods for classification and description of pages. Eric J. Glover, Kostas Tsioutsiouliklis, Steve Lawrence, David M. Pennock, Gary William Flake |
WWW | 2 |
| 2001 | Faster kinetic heaps and their use in broadcast scheduling
Haim Kaplan, Robert E. Tarjan, Kostas Tsioutsiouliklis |
SODA | 3 |
| 1999 | Cut Tree Algorithms
Andrew V. Goldberg, Kostas Tsioutsiouliklis |
SODA | 2 |