VLDB 2026 Research / reviewers in the wild / expert
Ismini Lourentzou
dblp:136/7883
· DBLP profile ↗
12ranked-venue papers in the field
4as first author
7since 2021 · last 2023
0000-0002-1238-772XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 8 (1 first)Big Data, Cloud & Distributed Data Systems · 3 (2 first)Data Mining & Knowledge Discovery · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | MArBLE: Hierarchical Multi-Armed Bandits for Human-in-the-Loop Set ExpansionabstractThe modern-day research community has an embarrassment of riches regarding pre-trained AI models. Even for a simple task such as lexicon set expansion, where an AI model suggests new entities to add to a predefined seed set of entities, thousands of models are available. However, deciding which model to use for a given set expansion task is non-trivial. In hindsight, some models can be 'off topic' for specific set expansion tasks, while others might work well initially but quickly exhaust what they have to offer. Additionally, certain models may require more careful priming in the form of samples or feedback before being finetuned to the task at hand. In this work, we frame this model selection as a sequential non-stationary problem, where there exist a large number of diverse pre-trained models that may or may not fit a task at hand, and an expert is shown one suggestion at a time to include in the set or not, i.e., accept or reject the suggestion. The goal is to expand the list with the most entities as quickly as possible. We introduce MArBLE, a hierarchical multi-armed bandit method for this task, and two strategies designed to address cold-start problems. Experimental results on three set expansion tasks demonstrate MArBLE's effectiveness compared to baselines. Muntasir Wahed, Daniel Gruhl, Ismini Lourentzou |
CIKM | 3 |
| 2023 | Multi-view Graph-Based Text Representations for Imbalanced Classification
Ola Karajeh, Ismini Lourentzou, Edward A. Fox |
TPDL | 2 |
| 2023 | Sedition Hunters: A Quantitative Study of the Crowdsourced Investigation into the 2021 U.S. Capitol AttackabstractSocial media platforms have enabled extremists to organize violent events, such as the 2021 U.S. Capitol Attack. Simultaneously, these platforms enable professional investigators and amateur sleuths to collaboratively collect and identify imagery of suspects with the goal of holding them accountable for their actions. Through a case study of Sedition Hunters, a Twitter community whose goal is to identify individuals who participated in the 2021 U.S. Capitol Attack, we explore what are the main topics or targets of the community, who participates in the community, and how. Using topic modeling, we find that information sharing is the main focus of the community. We also note an increase in awareness of privacy concerns. Furthermore, using social network analysis, we show how some participants played important roles in the community. Finally, we discuss implications for the content and structure of online crowdsourced investigations. Tianjiao Yu, Sukrit Venkatagiri, Ismini Lourentzou, Kurt Luther |
WWW | 3 |
| 2022 | Task-Driven Privacy-Preserving Data-Sharing Framework for the Industrial InternetabstractIndustrial Internet provides a collaborative computational platform for participating enterprises, allowing the collection of big data for machine learning tasks. Despite the promise of training and deployment acceleration, and the potential to optimize decision-making processes through data-sharing, the adoption of such technologies is impacted by the increasing concerns about information privacy. As enterprises prefer to keep data private, this limits interoperability. While prior work has largely explored privacy-preserving mechanisms, the proposed methods naively average or randomly sample data shared from all participants instead of selecting the most well-suited subsets for a particular downstream learning task. Motivated by the lack of effective data-sharing mechanisms for heterogeneous machine learning tasks in Industrial Internet, we propose PriED, a task-driven data-sharing framework that selectively fuses shared data and local data from participants to improve supervised learning performance. PriED utilizes privacy-preserving data distillation to facilitate data exchange, and dynamic data selection to optimize downstream machine learning tasks. We demonstrate performance improvements on a real semiconductor manufacturing case study. Parshin Shojaee, Yingyan Zeng, Muntasir Wahed, Avi Seth, Ismini Lourentzou |
IEEE Big Data | 6 |
| 2022 | Drink Bleach or Do What Now? COVID-HeRA: A Study of Risk-Informed Health Decision Making in the Presence of COVID-19 Misinformation
Arkin Dharawat, Ismini Lourentzou, Alex Morales, ChengXiang Zhai |
ICWSM | 2 |
| 2021 | SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus ExpansionabstractRecent advances in text representation have shown that training on large amounts of text is crucial for natural language understanding. However, models trained without predefined notions of topical interest typically require careful fine-tuning when transferred to specialized domains. When a sufficient amount of within-domain text may not be available, expanding a seed corpus of relevant documents from large-scale web data poses several challenges. First, corpus expansion requires scoring and ranking each document in the collection, an operation that can quickly become computationally expensive as the web corpora size grows. Relying on dense vector spaces and pairwise similarity adds to the computational expense. Secondly, as the domain concept becomes more nuanced, capturing the long tail of domain-specific rare terms becomes non-trivial, especially under limited seed corpora scenarios. Muntasir Wahed, Daniel Gruhl, Alfredo Alba, Anna Lisa Gentile, Petar Ristoski, Chad DeLuca, Steve Welch, Ismini Lourentzou |
CIKM | 8 |
| 2021 | DeepQAMVS: Query-Aware Hierarchical Pointer Networks for Multi-Video SummarizationabstractThe recent growth of web video sharing platforms has increased the demand for systems that can efficiently browse, retrieve and summarize video content. Query-aware multi-video summarization is a promising technique that caters to this demand. In this work, we introduce a novel Query-Aware Hierarchical Pointer Network for Multi-Video Summarization, termed DeepQAMVS, that jointly optimizes multiple criteria: (1) conciseness, (2) representativeness of important query-relevant events and (3) chronological soundness. We design a hierarchical attention model that factorizes over three distributions, each collecting evidence from a different modality, followed by a pointer network that selects frames to include in the summary. DeepQAMVS is trained with reinforcement learning, incorporating rewards that capture representativeness, diversity, query-adaptability and temporal coherence. We achieve state-of-the-art results on the MVS1K dataset, with inference time scaling linearly with the number of input video frames. Safa Messaoud, Ismini Lourentzou, Assma Boughoula, Mona Zehni, Zhizhen Zhao 0001, ChengXiang Zhai, Alexander G. Schwing |
SIGIR | 2 |
| 2019 | Adapting Sequence to Sequence Models for Text Normalization in Social Media
Ismini Lourentzou, Kabir Manghnani, ChengXiang Zhai |
ICWSM | 1 |
| 2018 | Mining Relations from Unstructured Content
Ismini Lourentzou, Alfredo Alba, Anni Coden, Anna Lisa Gentile, Daniel Gruhl, Steve Welch |
PAKDD (2) | 1 |
| 2017 | Text-based geolocation prediction of social media users with neural networksabstractInferring the location of a user has been a valuable step for many applications that leverage social media, such as marketing, security monitoring and recommendation systems. Motivated by the recent success of Deep Learning techniques for many other tasks such as computer vision, speech recognition, and natural language processing, we study the application of neural networks to the problem of geolocation prediction and experiment with multiple techniques to improve neural networks for geolocation inference based solely on text. Experimental results on three Twitter datasets suggest that choosing appropriate network architecture, activation function, and performing Batch Normalization, can all increase performance on this task. Ismini Lourentzou, Alex Morales, ChengXiang Zhai |
IEEE BigData | 1 |
| 2015 | Hotspots of news articles: Joint mining of news text & social media to discover controversial points in newsabstractWe propose and study a novel problem of mining news text and social media jointly to discover controversial points in news, which enables many applications such as highlighting controversial points in news articles for readers, revealing controversies in news and their trends over time, and quantifying the controversy of a news source. We design a controversy scoring function to discover the most controversial sentences in a news article by leveraging relevant comments in Twitter and comments on news web sites to assess the controversy of opinions about an issue mentioned in the news article. Multiple scoring strategies based on sentiment analysis and linguistic cues are proposed and studied. Experimental results show that the proposed algorithms can effectively discover controversial parts in news articles. Ismini Lourentzou, Graham Dyer, Abhishek Sharma 0015, ChengXiang Zhai |
IEEE BigData | 1 |
| 2013 | Automated snippet generation for online advertisingabstractProducts, services or brands can be advertised alongside the search results in major search engines, while recently smaller displays on devices like tablets and smartphones have imposed the need for smaller ad texts. In this paper, we propose a method that produces in an automated manner compact text ads (promotional text snippets), given as input a product description webpage (landing page). The challenge is to produce a small comprehensive ad while maintaining at the same time relevance, clarity, and attractiveness. Our method includes the following phases. Initially, it extracts relevant and important n-grams (keywords) given the landing page. The keywords reserved must have a positive meaning in order to have a call-to-action style, thus we attempt sentiment analysis on them. Next, we build an Advertising Language Model to evaluate phrases in terms of their marketing appeal. We experiment with two variations of our method and we show that they outperform all the baseline approaches. Stamatina Thomaidou, Ismini Lourentzou, Panagiotis Katsivelis-Perakis, Michalis Vazirgiannis |
CIKM | 2 |