EDBT 2026 Demo / reviewers in the wild / expert
Shady Elbassuoni
dblp:82/4089
· DBLP profile ↗
30ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-3491-6311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 25 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 2Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Factify: An Automated Fact-Checker for Web InformationabstractAutomated fact-checking has received widespread attention in recent years. In this paper, we present Factify, a transformer-based fact-checking approach that retrieves textual evidence from a trusted source for a given claim and subsequently uses the retrieved evidence to judge the validity of the claim. Our retrieval model is based on extractive question-answering and to train the model, we fine-tune various pre-trained BERT models using the FEVER fact-checking dataset. Our claim verification model is also BERT-based and is again trained using the FEVER dataset. Our evaluation using the test data of the FEVER dataset demonstrates the effectiveness of our approach in retrieving relevant textual evidences for given claims from trusted sources, and for determining whether claims are true or false based on the retrieved evidences, compared to various state-of-the-art automated fact-checking approaches. Ademola Adesokan, Shady Elbassuoni |
IEEE Big Data | 2 |
| 2023 | A Framework to Maximize Group Fairness for Workers on Online Labor PlatformsabstractAbstract As the number of online labor platforms and the diversity of jobs on these platforms increase, ensuring group fairness for workers needs to be the focus of job-matching services. Risk of discrimination against workers occurs in two different job-matching services: when someone is looking for a job (i.e., a job seeker) and when someone wants to deploy jobs (i.e., a job provider). To maximize their chances of getting hired, job seekers submit their profiles on different platforms. Similarly, job providers publish their job offers on multiple platforms with the goal of reaching a wide and diverse workforce. In this paper, we propose a theoretical framework to maximize group fairness for workers 1) when job seekers are looking for jobs on multiple platforms, and 2) when jobs are being deployed by job providers on multiple platforms. We formulate each goal as different optimization problems with different constraints, prove most of them are computationally hard to solve and propose various efficient algorithms to solve all of them in reasonable time. We then design a series of experiments that rely on synthetic and semi-synthetic data generated from a real-world online labor platform to evaluate our framework. Anis El Rabaa, Shady Elbassuoni, Jihad Hanna, Amer E. Mouawad, Ayham Olleik, Sihem Amer-Yahia |
Data Sci. Eng. | 2 |
| 2021 | An RDF Data Management System for Conflict CasualtiesabstractIn a world embroiled in armed conflicts, documenting conflict casualties is an important goal for many NGOs. Most of such documented records of casualties are however managed through internal databases, spreadsheets or Web forms. As such, exploring and querying such data becomes extremely chaotic. In this paper, we demonstrate CasualtIS, an RDF data management system for conflict casualties. Our system models conflict casualties data as RDF graphs and allows users to query such data using a SPARQL endpoint. Our system also includes a template-based natural-language querying interface to support non-expert users. Our system can be used for various purposes by end users, such as fact-checking certain claims about conflict casualties, aggregating casualties over time and location, and finding contextual information about casualties, such as the cause of death, actors involved, and other similar critical information. We demonstrate our system using two case studies, one related to casualties in the Iraqi war and the other related to casualties in the Syrian war. Yad Fatah, Mark Nourallah, Lynn Wahab, Fatima K. Abu Salem, Shady Elbassuoni |
CIKM | 5 |
| 2021 | Quantifying and Addressing Ranking Disparity in Human-Powered Data AcquisitionabstractAlgorithmic bias has been identified as a key challenge in many AI applications. One major source of bias is the data used to build these applications. For instance, many AI applications rely on human users to generate training data. The generated data might be biased if the data acquisition process is skewed towards certain groups of people based on say gender, ethnicity or location. This typically happens as a result of a hidden association between the people's qualifications for data acquisition and the people's protected attributes. In this paper, we study how to unveil and address disparity in data acquisition. We focus on the case where the data acquisition process involves ranking of people and we define disparity as the unbalanced targeting of people by the data acquisition process. To quantify disparity, we formulate an optimization problem that partitions people on their protected attributes, computes the qualifications of people in each partition, and finds the partitioning that exhibits the highest disparity in qualifications. Due to the combinatorial nature of our problem, we devise heuristics to navigate the space of partitions. We also discuss how to address disparity between partitions. We conduct a series of experiments on real and simulated datasets that demonstrate that our proposed approach is successful in quantifying and addressing ranking disparity in human-powered data acquisition. Sihem Amer-Yahia, Shady Elbassuoni, Ahmad Ghizzawi, Anas Hosami |
KDD | 2 |
| 2020 | Fairness in Online Jobs: A Case Study on TaskRabbit and GoogleabstractInternational audience Sihem Amer-Yahia, Shady Elbassuoni, Ahmad Ghizzawi, Ria Mae Borromeo, Emilie Hoareau, Philippe Mulhem |
EDBT | 2 |
| 2020 | Time-Aware Word Embeddings for Three Lebanese News ArchivesabstractWord embeddings have proven to be an effective method for capturing semantic relations among distinct terms within a large corpus. In this paper, we present a set of word embeddings learnt from three large Lebanese news archives, which collectively consist of 609,386 scanned newspaper images and spanning a total of 151 years, ranging from 1933 till 2011. The diversified ideological nature of the news archives alongside the temporal variability of the embeddings offer a rare glimpse onto the variation of word representation across the left-right political spectrum. To train the word embeddings, Google’s Tesseract 4.0 OCR engine was employed to transcribe the scanned news archives, and various archive-level as well as decade-level word embeddings were learnt. To evaluate the accuracy of the learnt word embeddings, a benchmark of analogy tasks was used. Finally, we demonstrate an interactive system that allows the end user to visualize for a given word of interest, the variation of the top-k closest words in the embedding space as a function of time and across news archives using an animated scatter plot. Jad Doughman, Fatima K. Abu Salem, Shady Elbassuoni |
LREC | 3 |
| 2019 | GroupTravel: Customizing Travel Packages for GroupsabstractInternational audience Sihem Amer-Yahia, Shady Elbassuoni, Behrooz Omidvar-Tehrani, Ria Mae Borromeo, Mehrdad Farokhnejad |
EDBT | 2 |
| 2019 | Exploring Fairness of Ranking in Online Job MarketplacesabstractInternational audience Shady Elbassuoni, Sihem Amer-Yahia, Christine El Atie, Ahmad Ghizzawi, Bilel Oualha |
EDBT | 1 |
| 2019 | FaiRank: An Interactive System to Explore Fairness of Ranking in Online Job MarketplacesabstractInternational audience Ahmad Ghizzawi, Julien Marinescu, Shady Elbassuoni, Sihem Amer-Yahia, Gilles Bisson |
EDBT | 3 |
| 2019 | Retrieving Textual Evidence for Knowledge Graph FactsabstractKnowledge graphs have become vital resources for semantic search and provide users with precise answers to their information needs. Knowledge graphs often consist of billions of facts, typically encoded in the form of RDF triples. In most cases, these facts are extracted automatically and can thus be susceptible to errors. For many applications, it can therefore be very useful to complement knowledge graph facts with textual evidence. For instance, it can help users make informed decisions about the validity of the facts that are returned as part of an answer to a query. In this paper, we therefore propose , an approach that given a knowledge graph and a text corpus, retrieves the top-k most relevant textual passages for a given set of facts. Since our goal is to retrieve short passages, we develop a set of IR models combining exact matching through the Okapi BM25 model with semantic matching using word embeddings. To evaluate our approach, we built an extensive benchmark consisting of facts extracted from YAGO and text passages retrieved from Wikipedia. Our experimental results demonstrate the effectiveness of our approach in retrieving textual evidence for knowledge graph facts. Gönenç Ercan, Shady Elbassuoni, Katja Hose |
ESWC | 2 |
| 2019 | FA-KES: A Fake News Dataset around the Syrian War
Fatima K. Abu Salem, Roaa Al Feel, Shady Elbassuoni, Mohamad Jaber 0001, May Farah |
ICWSM | 3 |
| 2018 | Effective searching of RDF knowledge graphs
Hiba Arnaout, Shady Elbassuoni |
J. Web Semant. | 2 |
| 2017 | Calories Prediction from Food Images
Manal Chokr, Shady Elbassuoni |
AAAI | 2 |
| 2017 | Website Navigation Behavior Analysis for Bot DetectionabstractDetecting bots is an important goal for most website admins. In this paper, we propose a novel machine learning bot detection approach based on local website navigation behavior. While machine learning has been used before for bot detection, most existing approaches rely on general hypotheses based on statistical analysis over multiple websites and are thus easy to counter. In our work, we build a website-specific hypothesis or classifier based on the actual navigation data of the website. The advantages of our approach is that it can be generally used to detect any type of bots and is difficult to counter unless website-specific bots are designed as well. Our classifier uses a Two-Class Boosted Decision Tree classification model and can be periodically re-trained to learn new hypotheses as bots evolve. We tested our approach on two real-world websites and achieved an accuracy of around 83%, outperforming the state-of-the-art machine-learning-based bot detection techniques by almost 14%. We also show that our approach can successfully distinguish between various classes of bots and we show how it can be deployed as a real-world application by any website to automatically detect bots as they navigate the website. Rabih Haidar, Shady Elbassuoni |
DSAA | 2 |
| 2017 | Customizing Travel Packages with Interactive Composite ItemsabstractWe examine the applicability of Composite Items (CIs) for generating customized travel packages consisting of Points of Interest (POIs) in a given city. CIs have been shown to serve complex information needs such as selecting books for a reading club, identifying a set of products for a promotion, or planning a city tour. In the travel domain, a synthesized view of travel options in a city can be provided with a set of cohesive CIs, each of which is covering a different region in the city. In this paper, we attempt to understand the benefit of letting users customize travel packages, and examine the relationship between customization and personalization. For personalization, we gather user preferences on POI features when available or on latent topics extracted from POI tags. For customization, we develop a framework within which a user interacts with proposed travel packages and the system suggests new CIs according to refined user preferences. Our experiments reveal a tension between personalization and the cohesiveness of items forming each CI. As a result, customization is necessary to find a balance between POI personalization and CI cohesiveness. We also show that the refined user preferences obtained from customization in one city help build better travel packages in another city. Manish Singh 0002, Ria Mae Borromeo, Anas Hosami, Sihem Amer-Yahia, Shady Elbassuoni |
DSAA | 5 |
| 2016 | Arabic Corpora for Credibility Analysis
Ayman Al Zaatari, Rim El Ballouli, Shady Elbassuoni, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban, Nizar Habash, Emad Yahya |
LREC | 3 |
| 2016 | Crowdsourcing Reliable Ratings for Underexposed ItemsabstractInternational audience Beatrice Valeri, Shady Elbassuoni, Sihem Amer-Yahia |
WEBIST (2) | 2 |
| 2015 | Acquiring Reliable Ratings from the CrowdabstractWe address the problem of acquiring reliable ratings of items such as restaurants or movies from the crowd. We propose a crowdsourcing platform that takes into consideration the workers’ skills with respect to the items being rated and assigns workers the best items to rate. Our platform focuses on acquiring ratings from skilled workers and for items that only have a few ratings. We evaluate the effectiveness of our system using a real-world dataset about restaurants. Beatrice Valeri, Shady Elbassuoni, Sihem Amer-Yahia |
HCOMP | 2 |
| 2013 | Robust question answering over the web of linked dataabstractKnowledge bases and the Web of Linked Data have become important assets for search, recommendation, and analytics. Natural-language questions are a user-friendly mode of tapping this wealth of knowledge and data. However, question answering technology does not work robustly in this setting as questions have to be translated into structured queries and users have to be careful in phrasing their questions. This paper advocates a new approach that allows questions to be partially translated into relaxed queries, covering the essential but not necessarily all aspects of the user's input. To compensate for the omissions, we exploit textual sources associated with entities and relational facts. Our system translates user questions into an extended form of structured SPARQL queries, with text predicates attached to triple patterns. Our solution is based on a novel optimization model, cast into an integer linear program, for joint decomposition and disambiguation of the user question. We demonstrate the quality of our methods through experiments with the QALD benchmark. Mohamed Yahya 0001, Klaus Berberich, Shady Elbassuoni, Gerhard Weikum |
CIKM | 3 |
| 2012 | Natural Language Questions for the Web of Data
Mohamed Yahya 0001, Klaus Berberich, Shady Elbassuoni, Maya Ramanath, Volker Tresp, Gerhard Weikum |
EMNLP-CoNLL | 3 |
| 2012 | RDF Xpress: a flexible expressive RDF search engineabstractWe demonstrate RDF Xpress, a search engine that enables users to effectively retrieve information from large RDF knowledge bases or Linked Data Sources. RDF Xpress provides a search interface where users can combine triple patterns with keywords to form queries. Moreover, RDF Xpress supports automatic query relaxation and returns a ranked list of diverse query results. Shady Elbassuoni, Maya Ramanath, Gerhard Weikum |
SIGIR | 1 |
| 2011 | Keyword search over RDF graphsabstractLarge knowledge bases consisting of entities and relationships between them have become vital sources of information for many applications. Most of these knowledge bases adopt the Semantic-Web data model RDF as a representation model. Querying these knowledge bases is typically done using structured queries utilizing graph-pattern languages such as SPARQL. However, such structured queries require some expertise from users which limits the accessibility to such data sources. To overcome this, keyword search must be supported. In this paper, we propose a retrieval model for keyword queries over RDF graphs. Our model retrieves a set of subgraphs that match the query keywords, and ranks them based on statistical language models. We show that our retrieval model outperforms the-state-of-the-art IR and DB models for keyword search over structured data using experiments over two real-world datasets. Shady Elbassuoni, Roi Blanco |
CIKM | 1 |
| 2011 | S3K: seeking statement-supporting top-K witnessesabstractTraditional information retrieval techniques based on keyword search help to identify a ranked set of relevant documents, which often contains many documents in the top ranks that do not meet the user's intention. By considering the semantics of the keywords and their relationships, both precision and recall can be improved. Using an ontology and mapping keywords to entities/concepts and identifying the relationship between them that the user is interested in, allows for retrieving documents that actually meet the user's intention. In this paper, we present a framework that enables semantic-aware document retrieval. User queries are mapped to semantic statements based on entities and their relationships. The framework searches for documents expressing these statements in different variations, e.g., synonymous names for entities or different textual expressions for relations between them. The size of potential result sets makes ranking documents according to their relevance to the user an essential component of such a system. The ranking model proposed in this paper is based on statistical language-models and considers aspects such as the authority of a document and the confidence in the textual pattern representing the queried information. Steffen Metzger, Shady Elbassuoni, Katja Hose, Ralf Schenkel |
CIKM | 2 |
| 2011 | Query Relaxation for Entity-Relationship Search
Shady Elbassuoni, Maya Ramanath, Gerhard Weikum |
ESWC (2) | 1 |
| 2010 | ROXXI: Reviving witness dOcuments to eXplore eXtracted InformationabstractIn recent years, there has been considerable research on information extraction and constructing RDF knowledge bases. In general, the goal is to extract all relevant information from a corpus of documents, store it into an ontology, and answer future queries based only on the created knowledge base. Thus, the original documents become dispensable. On the one hand, an ontology is a convenient and non-redundant structured source of information, based on which specific queries can be answered efficiently. On the other hand, many users doubt the correctness of facts and ontology subgraphs presented to them as query results without proof. Instead, users often wish to verify the obtained facts or subgraphs by reading about them in context, i.e., in a document relating the facts and providing background information. In this demo, we present ROXXI, a system operating on top of an existing knowledge base and reviving the abandoned witness documents. In doing so, it goes the opposite way of information extraction approaches -- starting with ontological facts and tracing their way back to the documents they were extracted from. ROXXI offers interfaces for expert users (SPARQL) as well as for non-experts (ontology browser) and provides a ranked list of documents each associated with a content snippet highlighting the queried facts in context. At the demonstration site, we will show the advantages of this novel approach towards document retrieval and illustrate the benefits of reviving the documents that information extraction approaches neglect. Shady Elbassuoni, Katja Hose, Steffen Metzger, Ralf Schenkel |
Proc. VLDB Endow. | 1 |
| 2009 | Language-model-based ranking for queries on RDF-graphsabstractThe success of knowledge-sharing communities like Wikipedia and the advances in automatic information extraction from textual and Web sources have made it possible to build large "knowledge repositories" such as DBpedia, Freebase, and YAGO. These collections can be viewed as graphs of entities and relationships (ER graphs) and can be represented as a set of subject-property-object (SPO) triples in the Semantic-Web data model RDF. Queries can be expressed in the W3C-endorsed SPARQL language or by similarly designed graph-pattern search. However, exact-match query semantics often fall short of satisfying the users' needs by returning too many or too few results. Therefore, IR-style ranking models are crucially needed. Shady Elbassuoni, Maya Ramanath, Ralf Schenkel, Marcin Sydow, Gerhard Weikum |
CIKM | 1 |
| 2009 | MING: mining informative entity relationship subgraphsabstractMany modern applications are faced with the task of knowledge discovery in entity-relationship graphs, such as domain-specific knowledge bases or social networks. Mining an "informative" subgraph that can explain the relations between k(>= 2) given entities of interest is a frequent knowledge discovery scenario on such graphs. We present MING, a principled method for extracting an informative subgraph for given query nodes. MING builds on a new notion of informativeness of nodes. This is used in a random-walk-with-restarts process to compute the informativeness of entire subgraphs. Gjergji Kasneci, Shady Elbassuoni, Gerhard Weikum |
CIKM | 2 |
| 2008 | Matching task profiles and user needs in personalized web searchabstractPersonalization has been deemed one of the major challenges in information retrieval with a significant potential for providing better search experience to individual users. Especially, the need for enhanced user models better capturing elements such as users' goals, tasks, and contexts has been identified. In this paper, we introduce a statistical language model for user tasks representing different granularity levels of a user profile, ranging from very specific search goals to broad topics. We propose a personalization framework that selectively matches the actual user information need with relevant past user tasks, and allows to dynamically switch the course of personalization from re-finding very precise information to biasing results to general user interests. In the extreme, our model is able to detect when the user's search and browse history is not appropriate for aiding the user in satisfying her current information quest. Instead of blindly applying personalization to all user queries, our approach refrains from undue actions in these cases, accounting for the user's desire of discovering new topics, and changing interests over time. The effectiveness of our method is demonstrated by an empirical user study. Julia Luxenburger, Shady Elbassuoni, Gerhard Weikum |
CIKM | 2 |
| 2008 | Task-aware search personalizationabstractSearch personalization has been pursued in many ways, in order to\nprovide better result rankings and better overall search experience\nto individual users.\nHowever, blindly applying personalization to all user queries, for example,\nby a background model derived from the user's long-term query-and-click\nhistory, is not always appropriate for aiding the user in accomplishing her \nactual task.\nUser interests change over time, a user sometimes works on very different \ncategories of tasks\nwithin a short timespan, and history-based personalization\nmay impede a user's desire of discovering new topics.\nIn this paper we propose a personalization framework that is\nselective in a twofold sense. First, it selectively employs\npersonalization techniques for queries that are expected to benefit from prior \nhistory\ninformation, while refraining from undue actions otherwise.\nSecond, we introduce the notion of tasks representing\ndifferent granularity levels of a user profile, ranging from very\nspecific search goals to broad topics, and base our reasoning selectively\non query-relevant user tasks.\nThese considerations are cast into a statistical language model for tasks, \nqueries, and\ndocuments, supporting both judicious query expansion and result re-ranking.\nThe effectiveness of our method is demonstrated by an empirical user study. Julia Luxenburger, Shady Elbassuoni, Gerhard Weikum |
SIGIR | 2 |
| 2008 | NAGA: harvesting, searching and ranking knowledgeabstractThe presence of encyclopedic Web sources, such as Wikipedia, the Internet Movie \nDatabase (IMDB), World Factbook, etc. calls for new querying techniques that \nare simple and yet more expressive than those provided by standard \nkeyword-based search engines. Searching for explicit knowledge needs to \nconsider inherent semantic structures involving entities and relationships.\n\nIn this demonstration proposal, we describe a semantic search system named \nNAGA. NAGA operates on a knowledge graph, which contains millions of entities \nand relationships derived from various encyclopedic Web sources, such as the \nones above. NAGA's graph-based query language is geared towards expressing \nqueries with additional semantic information. Its scoring model is based on the \nprinciples of generative language models, and formalizes several desiderata \nsuch as confidence, informativeness and compactness of answers.\n\nWe propose a demonstration of NAGA which will allow users to browse the \nknowledge base through a user interface, enter queries in NAGA's query language \nand tune the ranking parameters to test various ranking aspects. Gjergji Kasneci, Fabian M. Suchanek, Georgiana Ifrim, Shady Elbassuoni, Maya Ramanath, Gerhard Weikum |
SIGMOD Conference | 4 |