VLDB 2026 Research / reviewers in the wild / expert
Omar Sharif
dblp:270/0002
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-1971-6522ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Small Language Models on FPGAsabstractAttention is a major bottleneck when mapping Transformer-like models to FPGAs, as its matrix multiplications and normalisation stages exhibit differing numerical requirements and are highly sensitive to accumulation error. In this work, we propose operator-wise mixed-precision schemes and configurable accumulation strategies for attention-like pipelines based on shared-exponent low-bit, block floating-point style formats. By combining custom arithmetic with FPGA-specific design optimisations, our approach improves the trade-off between model quality and hardware cost, enabling more efficient deployment of small language models on reconfigurable hardware. Filip Wojcicki, Omar Sharif, Ebby Samson, Paul H. J. Kelly, George A. Constantinides, Christos-Savvas Bouganis, Wayne Luk |
FCCM | 2 |
| 2025 | Document-Level Event-Argument Data Augmentation for Challenging Role TypesabstractEvent Argument Extraction (EAE) is a daunting information extraction problem -with significant limitations in few-shot cross-domain (FSCD) settings.A common solution to FSCD modeling is data augmentation.Unfortunately, existing augmentation methods are not wellsuited to a variety of real-world EAE contexts, including (i) modeling long documents (documents with over 10 sentences), and (ii) modeling challenging role types (i.e., event roles with little to no training data and semantically outlying roles).We introduce two novel LLMpowered data augmentation methods for generating extractive document-level EAE samples using zero in-domain training data.We validate the generalizability of our approach on four datasets -showing significant performance increases in low-resource settings.Our highest performing models provide a 13-pt increase in F1 score on zero-shot role extraction in FSCD evaluation. Joseph Gatto, Omar Sharif, Parker Seegmiller, Sarah Masud Preum |
ACL (1) | 2 |
| 2025 | A Resource-Aware Residual-Based Gaussian Belief Propagation Accelerator Toolflow
Omar Sharif, Christos-Savvas Bouganis |
DATE | 1 |
| 2024 | Characterizing Information Seeking Events in Health-Related Social DiscourseabstractSocial media sites have become a popular platform for individuals to seek and share health information. Despite the progress in natural language processing for social media mining, a gap remains in analyzing health-related texts on social discourse in the context of events. Event-driven analysis can offer insights into different facets of healthcare at an individual and collective level, including treatment options, misconceptions, knowledge gaps, etc. This paper presents a paradigm to characterize health-related information-seeking in social discourse through the lens of events. Events here are board categories defined with domain experts that capture the trajectory of the treatment/medication. To illustrate the value of this approach, we analyze Reddit posts regarding medications for Opioid Use Disorder (OUD), a critical global health concern. To the best of our knowledge, this is the first attempt to define event categories for characterizing information-seeking in OUD social discourse. Guided by domain experts, we develop TREAT-ISE, a novel multilabel treatment information-seeking event dataset to analyze online discourse on an event-based framework. This dataset contains Reddit posts on information-seeking events related to recovery from OUD, where each post is annotated based on the type of events. We also establish a strong performance benchmark (77.4% F1 score) for the task by employing several machine learning and deep learning classifiers. Finally, we thoroughly investigate the performance and errors of ChatGPT on this task, providing valuable insights into the LLM's capabilities and ongoing characterization efforts. Omar Sharif, Madhusudan Basak, Tanzia Parvin, Ava Scharfstein, Alphonso Bradham, Jacob T. Borodovsky, Sarah E. Lord, Sarah Masud Preum |
AAAI | 1 |
| 2024 | Deciphering Hate: Identifying Hateful Memes and Their TargetsabstractInternet memes have become a powerful means for individuals to express emotions, thoughts, and perspectives on social media.While often considered a source of humor and entertainment, memes can also disseminate hateful content targeting individuals or communities.Most existing research focuses on the negative aspects of memes in high-resource languages, overlooking the distinctive challenges associated with low-resource languages like Bengali (also known as Bangla).Furthermore, while previous work on Bengali memes has focused on detecting hateful memes, there has been no work on detecting their targeted entities.To bridge this gap and facilitate research in this arena, we introduce a novel multimodal dataset for Bengali, BHM (Bengali Hateful Memes).The dataset consists of 7,148 memes with Bengali as well as code-mixed captions, tailored for two tasks: (i) detecting hateful memes, and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).To solve these tasks, we propose DORA (Dual cO-attention fRAmework), a multimodal deep neural network that systematically extracts the significant modality features from the memes and jointly evaluates them with the modality-specific features to understand the context better.Our experiments show that DORA is generalizable on other low-resource hateful meme datasets and outperforms several state-of-the-art rivaling baselines. Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah Masud Preum |
ACL (1) | 2 |
| 2024 | A Framework for Designing Scalable Gaussian Belief Propagation Accelerators for use in SLAMabstractGaussian Belief Propagation (GBP) is an iterative method for factor graph inference that provides an approximate solution to the probability distribution of a system. It has been shown to be a powerful tool in numerous applications including SLAM, where the estimation of the robot's position and the map of the environment is required. State-of-the-art implementations suffer from scalability issues, or exhibit performance degradation when off-chip memory access is required. This paper addresses these challenges using a streaming architecture via a chain of parameterizable Processing Elements (PE) that can be tuned to the problem's characteristics through the use of an optimizer. This work overcomes the limitations of existing GBP implementations achieving 142x-168x performance improvements over an embed-ded CPU for large graphs. Omar Sharif, Christos-Savvas Bouganis |
DATE | 1 |
| 2024 | A Multimodal Framework to Detect Target Aware Aggression in MemesabstractShawly Ahsan, Eftekhar Hossain, Omar Sharif, Avishek Das, Mohammed Moshiul Hoque, M. Dewan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shawly Ahsan, Eftekhar Hossain, Omar Sharif, Avishek Das, Mohammed Moshiul Hoque, M. Ali Akber Dewan |
EACL (1) | 3 |
| 2024 | Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex ArgumentsabstractPrior works formulate the extraction of eventspecific arguments as a span extraction problem, where event arguments are explicit -i.e.assumed to be contiguous spans of text in a document.In this study, we revisit this definition of Event Extraction (EE) by introducing two key argument types that cannot be modeled by existing EE frameworks.First, implicit arguments are event arguments which are not explicitly mentioned in the text, but can be inferred through context.Second, scattered arguments are event arguments that are composed of information scattered throughout the text.These two argument types are crucial to elicit the full breadth of information required for proper event modeling.To support the extraction of explicit, implicit, and scattered arguments, we develop a novel dataset, DiscourseEE, which includes 7,464 argument annotations from online health discourse.Notably, 51.2% of the arguments are implicit, and 17.4% are scattered, making Dis-courseEE a unique corpus for complex event extraction.Additionally, we formulate argument extraction as a text generation problem to facilitate the extraction of complex argument types.We provide a comprehensive evaluation of state-of-the-art models and highlight critical open challenges in generative event extraction.Our data and codebase are available at https://omar-sharif03.github.io/DiscourseEE. Omar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud Preum |
EMNLP | 1 |
| 2024 | Theme-Driven Keyphrase Extraction to Analyze Social Media DiscourseabstractSocial media platforms are vital resources for sharing self-reported health experiences, offering rich data on various health topics. Despite advancements in Natural Language Processing (NLP) enabling large-scale social media data analysis, a gap remains in applying keyphrase extraction to health-related content. Keyphrase extraction is used to identify salient concepts in social media discourse without being constrained by predefined entity classes. This paper introduces a theme-driven keyphrase extraction framework tailored for social media, a pioneering approach designed to capture clinically relevant keyphrases from user-generated health texts. Themes are defined as broad categories determined by the objectives of the extraction task. We formulate this novel task of theme-driven keyphrase extraction and demonstrate its potential for efficiently mining social media text for the use case of treatment for opioid use disorder. This paper leverages qualitative and quantitative analysis to demonstrate the feasibility of extracting actionable insights from social media data and efficiently extracting keyphrases using minimally supervised NLP models. Our contributions include the development of a novel data collection and curation framework for theme-driven keyphrase extraction and the creation of SuboxoPhrase, the first dataset of its kind comprising human-annotated keyphrases from a Reddit community. We also identify the scope of minimally supervised NLP models to extract keyphrases from social media data efficiently. Lastly, we found that a large language model (ChatGPT) outperforms unsupervised keyphrase extraction models, showcasing its efficacy in this task. William Romano, Omar Sharif, Madhusudan Basak, Joseph Gatto, Sarah Masud Preum |
ICWSM | 2 |
| 2024 | SAWTab: Smoothed Adaptive Weighting for Tabular Data in Semi-supervised Learning
Morteza Mohammady Gharasuie, Fengjiao Wang, Omar Sharif, Ravi Mukkamala |
PAKDD (3) | 3 |
| 2022 | MemoSen: A Multimodal Dataset for Sentiment Analysis of MemesabstractPosting and sharing memes have become a powerful expedient of expressing opinions on social media in recent days. Analysis of sentiment from memes has gained much attention to researchers due to its substantial implications in various domains like finance and politics. Past studies on sentiment analysis of memes have primarily been conducted in English, where low-resource languages gain little or no attention. However, due to the proliferation of social media usage in recent years, sentiment analysis of memes is also a crucial research issue in low resource languages. The scarcity of benchmark datasets is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali. This paper presents a novel multimodal dataset (named MemoSen) for Bengali containing 4417 memes with three annotated labels positive, negative, and neutral. A detailed annotation guideline is provided to facilitate further resource development in this domain. Additionally, a set of experiments are carried out on MemoSen by constructing twelve unimodal (i.e., visual, textual) and ten multimodal (image+text) models. The evaluation exhibits that the integration of multimodal information significantly improves (about 1.2%) the meme sentiment classification compared to the unimodal counterparts and thus elucidate the novel aspects of multimodality. Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque |
LREC | 2 |
| 2022 | Tackling cyber-aggression: Identification and fine-grained categorization of aggressive texts on social media using weighted ensemble of transformers
Omar Sharif, Mohammed Moshiul Hoque |
Neurocomputing | 1 |
| 2020 | SentiLSTM: A Deep Learning Approach for Sentiment Analysis of Restaurant Reviews
Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Iqbal H. Sarker |
HIS | 2 |