EDBT 2026 Demo / reviewers in the wild / expert
Khaled Ammar
dblp:29/9884 · also Khalid Ammar
· DBLP profile ↗
14ranked-venue papers
7as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Building an Enterprise World ModelabstractEnterprises face a dual challenge when deploying large language models (LLMs) in customer-facing environments. The first challenge is ensuring knowledge retrieval accuracy and alignment with legal compliance. The second is ensuring alignment with their brand identity while optimizing engagement outcomes. Generic LLMs are trained on public data and world knowledge but do not represent enterprise-specific behaviour. When LLMs are supported with context-aware retrieval methods, they can address the first challenge. The second challenge, however, requires careful fine-tuning. For an enterprise, the fine-tuning process is similar to hiring someone with public knowledge about business, but is not yet trained to act as an experienced employee. This talk presents a hybrid fine-tuning and context-orchestration framework, an important step towards constructing an enterprise world model that unifies structured, unstructured, and conversational data for compliant, brand-aligned LLMs. Khaled Ammar |
WSDM | 1 |
| 2024 | KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceabstractIn recent years, we have witnessed the growing interest from academia and industry in applying data science technologies to analyze large amounts of data. In this process, a myriad of artifacts (datasets, pipeline scripts, etc.) are created. However, there has been no systematic attempt to holistically collect and exploit all the knowledge and experiences that are implicitly contained in those artifacts. Instead, data scientists recover information and expertise from colleagues or learn via trial and error. Hence, this paper presents a scalable platform, KGLiDS, that employs machine learning and knowledge graph technologies to abstract and capture the semantics of data science artifacts and their connections. Based on this information, KGLiDS enables various downstream applications, such as data discovery and pipeline automation. Our comprehensive evaluation covers use cases in data discovery, data cleaning, transformation, and AutoML. It shows that KGLiDS is significantly faster with a lower memory footprint than the state-of-the-art systems while achieving comparable or better accuracy. Mossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier, Ahmed Helal, Antonio Cavalcante, Khaled Ammar, Katja Hose, Essam Mansour 0001 |
ICDE | 7 |
| 2022 | Optimizing Differentially-Maintained Recursive Queries on Dynamic GraphsabstractDifferential computation (DC) is a highly general incremental computation/view maintenance technique that can maintain the output of an arbitrary and possibly recursive dataflow computation upon changes to its base inputs. As such, it is a promising technique for graph database management systems (GDBMS) that support continuous recursive queries over dynamic graphs. Although differential computation can be highly efficient for maintaining these queries, it can require prohibitively large amount of memory. This paper studies how to reduce the memory overhead of DC with the goal of increasing the scalability of systems that adopt it. We propose a suite of optimizations that are based on dropping the differences of operators, both completely or partially, and recomputing these differences when necessary. We propose deterministic and probabilistic data structures to keep track of the dropped differences. Extensive experiments demonstrate that the optimizations can improve the scalability of a DC-based continuous query processor. Khaled Ammar, Siddhartha Sahu, Semih Salihoglu, M. Tamer Özsu |
Proc. VLDB Endow. | 1 |
| 2021 | A Demonstration of KGLac: A Data Discovery and Enrichment Platform for Data ScienceabstractData science growing success relies on knowing where a relevant dataset exists, understanding its impact on a specific task, finding ways to enrich a dataset, and leveraging insights derived from it. With the growth of open data initiatives, data scientists need an extensible set of effective discovery operations to find relevant data from their enterprise datasets accessible via data discovery systems or open datasets accessible via data portals. Existing portals and systems suffer from limited discovery support and do not track the use of a dataset and insights derived from it. We will demonstrate KGLac, a system that captures metadata and semantics of datasets to construct a knowledge graph (GLac) interconnecting data items, e.g., tables and columns. KGLac supports various data discovery operations via SPARQL queries for table discovery, unionable and joinable tables, plus annotation with related derived insights. We harness a broad range of Machine Learning (ML) approaches with GLac to enable automatic graph learning for advanced and semantic data discovery. The demo will showcase how KGLac facilitates data discovery and enrichment while developing an ML pipeline to evaluate potential gender salary bias in IT jobs. Ahmed Helal, Mossad Helali, Khaled Ammar, Essam Mansour 0001 |
Proc. VLDB Endow. | 3 |
| 2018 | Distributed Evaluation of Subgraph Queries Using Worst-case Optimal and Low-Memory DataflowsabstractWe study the problem of finding and monitoring fixed-size subgraphs in a continually changing large-scale graph. We present the first approach that (i) performs worst-case optimal computation and communication, (ii) maintains a total memory footprint linear in the number of input edges, and (iii) scales down per-worker computation, communication, and memory requirements linearly as the number of workers increases, even on adversarially skewed inputs. Our approach is based on worst-case optimal join algorithms, recast as a data-parallel dataflow computation. We describe the general algorithm and modifications that make it robust to skewed data, prove theoretical bounds on its resource requirements in the massively parallel computing model, and implement and evaluate it on graphs containing as many as 64 billion edges. The underlying algorithm and ideas generalize from finding and monitoring subgraphs to the more general problem of computing and maintaining relational equi-joins over dynamic relations. Khaled Ammar, Frank McSherry, Semih Salihoglu, Manas Joglekar |
Proc. VLDB Endow. | 1 |
| 2018 | Experimental Analysis of Distributed Graph SystemsabstractThis paper evaluates eight parallel graph processing systems: Hadoop, HaLoop, Vertica, Giraph, GraphLab (PowerGraph), Blogel, Flink Gelly, and GraphX (SPARK) over four very large datasets (Twitter, World Road Network, UK 200705, and ClueWeb) using four workloads (PageRank, WCC, SSSP and K-hop). The main objective is to perform an independent scale-out study by experimentally analyzing the performance, usability, and scalability (using up to 128 machines) of these systems. In addition to performance results, we discuss our experiences in using these systems and suggest some system tuning heuristics that lead to better performance. Khaled Ammar, M. Tamer Özsu |
Proc. VLDB Endow. | 1 |
| 2016 | Sapphire: Querying RDF Data Made SimpleabstractThere is currently a large amount of publicly accessible structured data available as RDF data sets. For example, the Linked Open Data (LOD) cloud now consists of thousands of RDF data sets with over 30 billion triples, and the number and size of the data sets is continuously growing. Many of the data sets in the LOD cloud provide public SPARQL endpoints to allow issuing queries over them. These end-points enable users to retrieve data using precise and highly expressive SPARQL queries. However, in order to do so, the user must have sufficient knowledge about the data sets that she wishes to query, that is, the structure of data, the vocabulary used within the data set, the exact values of literals, their data types, etc. Thus, while SPARQL is powerful, it is not easy to use. An alternative to SPARQL that does not require as much prior knowledge of the data is some form of keyword search over the structured data. Keyword search queries are easy to use, but inherently ambiguous in describing structured queries. This demonstration introduces Sapphire, a system for querying RDF data that strikes a middle ground between ambiguous keyword search and difficult-to-use SPARQL. Our system does not replace either, but utilizes both where they are most effective. Sapphire helps the user construct expressive SPARQL queries that represent her information needs without requiring detailed knowledge about the queried data sets. These queries are then executed over public SPARQL endpoints from the LOD cloud. Sapphire guides the user in the query writing process by showing suggestions of query terms based on the queried data, and by recommending changes to the query based on a predictive user model. Ahmed El-Roby, Khaled Ammar, Ashraf Aboulnaga, Jimmy Lin |
Proc. VLDB Endow. | 2 |
| 2015 | DISTINGER: A distributed graph data structure for massive dynamic graph processingabstractLarge and dynamic graphs with streaming updates have been gaining traction recently, along with the need for enabling graph analytics in a commodity cluster instead of a high-performance computing facility. Surprisingly, there is a lack of study on scaling out graph data structures to represent sparse dynamic graphs in a commodity cluster, and even the latest work [1] based upon the most common in-memory graph representation CSR [2] is a single-machine case. In this paper we present DISTINGER, a distributed graph representation that handles massive graph analytics with streaming updates. DISTINGER successfully extends a scale-up design to a scale-out graph data structure while maintains its efficiency and scalability. We implement our design and algorithms as a prototype, and compare it to single-site STINGER and state-of-art graph systems. Our experimental evaluation in a real cluster shows that DISTINGER can handle larger graphs than STINGER, and perform graph tasks (PageRank and edge updates) more efficiently than GraphLab and Giraph. Guoyao Feng, Khaled Ammar |
IEEE BigData | 3 |
| 2015 | BusMate: Understanding Mobility Behavior for Trajectory-Based AdvertisingabstractMobile advertising is a rapidly growing form of advertising, carrying with it the promise of more finely tuned, targeted advertisements. In this paper, we study mobile advertising in the use of public transportation (specifically, buses). Through a qualitative study, we identify a number of factors that can influence the effectiveness of mobile advertising for bus passengers. We then use these findings to inform a prototype advertising system for bus passengers, called Bus Mate. Collectively, our research suggests that the bus passenger's trip timeline (i.e., How long they will be on the bus, and where they are on the trip) is one of the most important factors in defining the users' attention level. We propose an attention graph and validate it using our experimental mobile application. Finally, we propose some design recommendations, based on our findings, for platforms aiming at targeting mobile ads to bus passengers. Khaled Ammar, Abdullah Elsayed, Mohamed M. Sabri, Michael Terry |
MDM (2) | 1 |
| 2015 | Continuous Median Queries in Wireless Sensor NetworksabstractA Wireless Sensor Network (WSN) consists of a set of small and autonomous sensing nodes, which possess limited energy and computational capabilities, and are typically used to monitor events. In many applications, one is interested in continuous and robust statistical summaries of the observed values, and in this paper we focus on reporting the median of the observed values gathered by the WSN. Keeping in mind that the nodes' energy consumption is paramount in a WSN, we propose a distributed and energy-efficient approach based on the use of multiple, suitably designed, histogram queries which efficiently explore existing cached results minimizing the number of bytes transmitted. Our experimental results, using synthetic and real datasets, show that our proposed solution is indeed able to substantially extend the lifespan of the WSN when compared to the current state-of-the-art. Khaled Ammar, Mario A. Nascimento |
MDM (1) | 1 |
| 2014 | An Experimental Comparison of Pregel-like Graph Processing SystemsabstractThe introduction of Google's Pregel generated much interest in the field of large-scale graph data processing, inspiring the development of Pregel-like systems such as Apache Giraph, GPS, Mizan, and GraphLab, all of which have appeared in the past two years. To gain an understanding of how Pregel-like systems perform, we conduct a study to experimentally compare Giraph, GPS, Mizan, and GraphLab on equal ground by considering graph and algorithm agnostic optimizations and by using several metrics. The systems are compared with four different algorithms (PageRank, single source shortest path, weakly connected components, and distributed minimum spanning tree) on up to 128 Amazon EC2 machines. We find that the system optimizations present in Giraph and GraphLab allow them to perform well. Our evaluation also shows Giraph 1.0.0's considerable improvement since Giraph 0.1 and identifies areas of improvement for all systems. Minyang Han, Khuzaima Daudjee, Khaled Ammar, M. Tamer Özsu, Xingfang Wang, Tianqi Jin |
Proc. VLDB Endow. | 3 |
| 2013 | Cost-Based Quantile Query Processing in Wireless Sensor NetworksabstractIn this paper we investigate how to efficiently and effectively use histogram queries for processing quantile queries in wireless sensor networks. A major concern when processing queries within such an environment is to minimize the energy consumption by the network nodes, thus extending the networks lifetime, e.g., the time when the first node runs out of energy. Towards that goal, we define a cost model for a refinement-based algorithm that performs a series of refining histogram queries in order to determine the exact quantile value. Given that the histogram size, i.e., its number of bins, is an important factor in the query processing cost, we use the defined cost model to estimate the histogram size that minimizes the maximum energy cost per-node when processing the quantile query. This is equivalent to maximizing the time until the first node dies and therefore to extending the network's lifetime. In our experiments, using synthetic and real datasets, we evaluate the performance of the proposed solutions in a variety of different settings. Johannes Niedermayer, Mario A. Nascimento, Matthias Renz, Peer Kröger, Khaled Ammar, Hans-Peter Kriegel |
MDM (1) | 5 |
| 2011 | Histogram and Other Aggregate Queries in Wireless Sensor Networks
Khaled Ammar, Mario A. Nascimento |
SSDBM | 1 |
| 1988 | An algorithm for polygon conversion to boxes for VLSI layouts
Asim J. Al-Khalili, Dhamin Al-Khalili, Khaled Ammar |
Integr. | 3 |