VLDB 2026 Research / reviewers in the wild / expert
Ahmed Elgohary
dblp:03/8978
· DBLP profile ↗
19ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorComputer networks · 4Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Question answering and dialogue systems · 29% Trustworthy machine learning · 28% Information extraction and text analysis · 20% | |
| Computer networks
4 papers |
Wireless sensing and localization · 100% | |
| Databases, data mining, and information retrieval
3 papers |
Machine learning and data management · 58% Data models and query languages · 25% Data mining · 17% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025 |
Wireless sensing and localization
indoor localization |
0.7 | 4 | 2016 | SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016 Video: Unsupervised indoor localization (UnLoc): beyond the prototype · MobiSys 2014 Demo: unsupervised indoor localization · MobiSys 2012 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.7 | 2 | 2019 | Can You Unpack That? Learning to Rewrite Questions-in-Context · EMNLP/IJCNLP (1) 2019 A dataset and baselines for sequential open-domain question answering · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.4 | 1 | 2020 | It Takes Two to Lie: One to Lie, and One to Listen · ACL 2020 |
Natural language and speech › Question answering and dialogue systems › natural language interface
interactive semantic parsing |
0.4 | 1 | 2020 | Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.4 | 1 | 2020 | Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.4 | 1 | 2020 | Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020 |
Wireless sensing and localization › indoor localization
landmark-based localization |
0.4 | 2 | 2016 | SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016 No need to war-drive: unsupervised indoor localization · MobiSys 2012 |
Natural language and speech › Question answering and dialogue systems
question rewriting |
0.4 | 1 | 2019 | Can You Unpack That? Learning to Rewrite Questions-in-Context · EMNLP/IJCNLP (1) 2019 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.3 | 1 | 2018 | Generating Natural Language Adversarial Examples · EMNLP 2018 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.3 | 1 | 2018 | A dataset and baselines for sequential open-domain question answering · EMNLP 2018 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2018 | Generating Natural Language Adversarial Examples · EMNLP 2018 |
Machine learning and data management
scalable machine learning |
0.3 | 1 | 2018 | Compressed linear algebra for large-scale machine learning · VLDB J. 2018 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 1 | 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
pluralistic alignment |
0.3 | 1 | 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025 |
Wireless sensing and localization
simultaneous localization and mapping |
0.2 | 1 | 2016 | SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016 |
Data mining › sampling
representative selection |
0.2 | 1 | 2013 | Distributed Column Subset Selection on MapReduce · ICDM 2013 |
Algorithms and data structures › matrix approximation
column subset selection |
0.2 | 1 | 2013 | Distributed Column Subset Selection on MapReduce · ICDM 2013 |
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing |
0.1 | 3 | 2014 | Video: Unsupervised indoor localization (UnLoc): beyond the prototype · MobiSys 2014 Demo: unsupervised indoor localization · MobiSys 2012 No need to war-drive: unsupervised indoor localization · MobiSys 2012 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2018 | Compressed linear algebra for large-scale machine learning · VLDB J. 2018 |
Natural language and speech › Language models and text generation
text generation |
0.1 | 1 | 2018 | Generating Natural Language Adversarial Examples · EMNLP 2018 |
Robotics › Robot navigation and mapping
SLAM |
0.1 | 1 | 2016 | SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.0 | 1 | 2013 | Distributed Column Subset Selection on MapReduce · ICDM 2013 |
Methods — techniques the papers use, named apart from their topics
dead reckoning · 1.2safety configs · 0.9in-context alignment · 0.9data-centric alignment · 0.9landmark sensing · 0.7random projection · 0.5low-rank approximation · 0.5sequence-to-structure models · 0.4question rewriting · 0.4population-based optimization · 0.3linear algebra compression · 0.3dataset construction · 0.3black-box optimization · 0.3baseline · 0.3adversarial training · 0.3wifi fingerprinting · 0.3magnetometer sensing · 0.3accelerometer sensing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety RequirementsabstractThe current paradigm for safety alignment of large language models (LLMs) follows a _one-size-fits-all_ approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In addition, users may have diverse safety needs, making a model with _static_ safety standards too restrictive to be useful, as well as too costly to be re-aligned.
We propose _Controllable Safety Alignment_ (CoSA), a framework designed to adapt models to diverse safety requirements without re-training. Instead of aligning a fixed model, we align models to follow _safety configs_—free-form natural language descriptions of the desired safety behaviors—that are provided as part of the system prompt. To adjust model safety behavior, authorized users only need to modify such safety configs at inference time. To enable that, we propose CoSAlign, a data-centric method for aligning LLMs to easily adapt to diverse safety configs. Furthermore, we devise a novel controllability evaluation protocol that considers both helpfulness and configured safety, summarizing them into CoSA-Score, and construct CoSApien, a _human-authored_ benchmark that consists of real-world LLM use cases with diverse safety requirements and corresponding evaluation prompts. We show that CoSAlign leads to substantial gains of controllability over strong baselines including in-context alignment. Our framework encourages better representation and adaptation to pluralistic human values in LLMs, and thereby increasing their practicality. Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme |
ICLR | 2 |
| 2021 | NL-EDIT: Correcting Semantic Parse Errors through Natural Language InteractionabstractAhmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo Ramos, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo A. Ramos, Ahmed Awadallah 0001 |
NAACL-HLT | 1 |
| 2020 | Speak to your Parser: Interactive Text-to-SQL with Natural Language FeedbackabstractWe study the task of semantic parse correction with natural language feedback.Given a natural language utterance, most semantic parsing systems pose the problem as one-shot translation where the utterance is mapped to a corresponding logical form.In this paper, we investigate a more interactive scenario where humans can further interact with the system by providing free-form natural language feedback to correct the system when it generates an inaccurate interpretation of an initial utterance.We focus on natural language to SQL systems and construct, SPLASH, a dataset of utterances, incorrect SQL interpretations and the corresponding natural language feedback.We compare various reference models for the correction task and show that incorporating such a rich form of feedback can significantly improve the overall semantic parsing accuracy while retaining the flexibility of natural language interaction.While we estimated human correction accuracy is 81.5%, our best model achieves only 25.1%, which leaves a large gap for improvement in future research.SPLASH is publicly available at https:// aka.ms/Splash_dataset. Exact Match Accuracy (%) BaselineCorrection End-to-End Without Feedback ⇒ Seq2Struct N/A 41.30 ⇒ Re- Ahmed Elgohary, Saghar Hosseini, Ahmed Awadallah 0001 |
ACL | 1 |
| 2020 | It Takes Two to Lie: One to Lie, and One to ListenabstractDenis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan Boyd-Graber. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan L. Boyd-Graber |
ACL | 3 |
| 2019 | Can You Unpack That? Learning to Rewrite Questions-in-ContextabstractAhmed Elgohary, Denis Peskov, Jordan Boyd-Graber. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ahmed Elgohary, Denis Peskov, Jordan L. Boyd-Graber |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Assessing Composition in Sentence Vector RepresentationsabstractAn important component of achieving language understanding is mastering the composition of sentence meaning, but an immediate challenge to solving this problem is the opacity of sentence vector representations produced by current neural sentence composition models. We present a method to address this challenge, developing tasks that directly target compositional meaning information in sentence vector representations with a high degree of precision and control. To enable the creation of these controlled tasks, we introduce a specialized sentence generation system that produces large, annotated sentence sets meeting specified syntactic, semantic and lexical constraints. We describe the details of the method and generation system, and then present results of experiments applying our method to probe for compositional information in embeddings from a number of existing sentence composition models. We find that the method is able to extract useful information about the differing capacities of these models, and we discuss the implications of our results with respect to these systems’ capturing of sentence information. We make available for public use the datasets used for these experiments, as well as the generation system. Allyson Ettinger, Ahmed Elgohary, Colin Phillips, Philip Resnik |
COLING | 2 |
| 2018 | Generating Natural Language Adversarial ExamplesabstractDeep neural networks (DNNs) are vulnerable to adversarial examples, perturbations to correctly classified examples which can cause the model to misclassify.In the image domain, these perturbations are often virtually indistinguishable to human perception, causing humans and state-of-the-art models to disagree.However, in the natural language domain, small perturbations are clearly perceptible, and the replacement of a single word can drastically alter the semantics of the document.Given these challenges, we use a black-box population-based optimization algorithm to generate semantically and syntactically similar adversarial examples that fool well-trained sentiment analysis and textual entailment models with success rates of 97% and 70%, respectively.We additionally demonstrate that 92.3% of the successful sentiment analysis adversarial examples are classified to their original label by 20 human annotators, and that the examples are perceptibly quite similar.Finally, we discuss an attempt to use adversarial training as a defense, but fail to yield improvement, demonstrating the strength and diversity of our adversarial examples.We hope our findings encourage researchers to pursue improving the robustness of DNNs in the natural language domain. Moustafa Farid Alzantot, Yash Sharma 0001, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava 0001, Kai-Wei Chang 0001 |
EMNLP | 3 |
| 2018 | A dataset and baselines for sequential open-domain question answeringabstractPrevious work on question-answering systems mainly focuses on answering individual questions, assuming they are independent and devoid of context.Instead, we investigate sequential question answering, asking multiple related questions.We present QBLink, a new dataset of fully human-authored questions.We extend existing strong question answering frameworks to include previous questions to improve the overall question-answering accuracy in open-domain question answering.The dataset is publicly available at http:// sequential.qanta.org. Ahmed Elgohary, Chen Zhao 0013, Jordan L. Boyd-Graber |
EMNLP | 1 |
| 2018 | Compressed linear algebra for large-scale machine learning
Ahmed Elgohary, Matthias Boehm 0001, Peter J. Haas, Frederick Reiss 0001, Berthold Reinwald |
VLDB J. | 1 |
| 2016 | Compressed Linear Algebra for Large-Scale Machine LearningabstractLarge-scale machine learning (ML) algorithms are often iterative, using repeated read-only data access and I/O-bound matrix-vector multiplications to converge to an optimal model. It is crucial for performance to fit the data into single-node or distributed main memory. General-purpose, heavy- and lightweight compression techniques struggle to achieve both good compression ratios and fast decompression speed to enable block-wise uncompressed operations. Hence, we initiate work on compressed linear algebra (CLA), in which lightweight database compression techniques are applied to matrices and then linear algebra operations such as matrix-vector multiplication are executed directly on the compressed representations. We contribute effective column compression schemes, cache-conscious operations, and an efficient sampling-based compression algorithm. Our experiments show that CLA achieves in-memory operations performance close to the uncompressed case and good compression ratios that allow us to fit larger datasets into available memory. We thereby obtain significant end-to-end performance improvements up to 26x or reduced memory requirements. Ahmed Elgohary, Matthias Boehm 0001, Peter J. Haas, Frederick Reiss 0001, Berthold Reinwald |
Proc. VLDB Endow. | 1 |
| 2016 | SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor LocalizationabstractIndoor localization using mobile sensors has gained momentum lately. Most of the current systems rely on an extensive calibration step to achieve high accuracy. We propose SemanticSLAM, a novel unsupervised indoor localization scheme that bypasses the need for war-driving. SemanticSLAM leverages the idea that certain locations in an indoor environment have a unique signature on one or more phone sensors. Climbing stairs, for example, has a distinct pattern on the phone's accelerometer; a specific spot may experience an unusual magnetic interference while another may have a unique set of Wi-Fi access points covering it. SemanticSLAM uses these unique points in the environment as landmarks and combines them with dead-reckoning in a new Simultaneous Localization And Mapping (SLAM) framework to reduce both the localization error and convergence time. In particular, the phone inertial sensors are used to keep track of the user's path, while the observed landmarks are used to compensate for the accumulation of error in a unified probabilistic framework. Evaluation in two testbeds on Android phones shows that the system can achieve 0.53 meters human median localization errors. In addition, the system can detect the location of landmarks with 0.83 meters median error. This is 62 percent better than a system that does not use SLAM. Moreover, SemanticSLAM has a 33 percent lower convergence time compared to the same systems. This highlights the promise of SemanticSLAM as an unconventional approach for indoor localization. Heba Abdelnasser, Reham Mohamed Aburas, Ahmed Elgohary, Moustafa Farid Alzantot, He Wang 0008, Souvik Sen, Romit Roy Choudhury, Moustafa Youssef 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Greedy column subset selection for large-scale data sets
Ahmed K. Farahat, Ahmed Elgohary, Ali Ghodsi 0001, Mohamed S. Kamel |
Knowl. Inf. Syst. | 2 |
| 2014 | Video: Unsupervised indoor localization (UnLoc): beyond the prototypeabstractThis video presents a demo of indoor localization in multiple settings. In the demo, a user walks with a smartphone and the user's location is shown on the phone's screen in real time. Our system, called Unsupervised Indoor Localization (UnLoc) utilizes the sensor data from smartphones to learn "invisible landmarks" in the environment. Example landmarks could be a unique magnetic fluctuation experienced when the phone is near a water-cooler, or a distinct gyroscope rotation when the user turns a corner. We use these indoor "landmarks" to periodically reset the user's location. To track the user between these landmarks, we use an optimized variant of dead reckoning, ultimately leading to a robust location tracking system. We call our system UnLoc, since the landmarks are generated in an unsupervised manner, requiring no manual effort or floorplan of the building. The demo describes the high level intuitions, shows UnLoc in operation, and shares experiences from running UnLoc in various real-world environments. He Wang 0008, Souvik Sen, Alexander Mariakakis, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001, Romit Roy Choudhury |
MobiSys | 4 |
| 2014 | Embed and Conquer: Scalable Embeddings for Kernel k-Means on MapReduceabstractThe kernel k-means is an effective method for data clustering which extends the commonly-used k-means algorithm to work on a similarity matrix over complex data structures. It is, however, computationally very complex as it requires the complete kernel matrix to be calculated and stored. Further, its kernelized nature hinders the parallelization of its computations on modern scalable infrastructures for distributed computing. In this paper, we are defining a family of kernelbased low-dimensional embeddings that allows for scaling kernel k-means on MapReduce via an efficient and unified parallelization strategy. Afterwards, we propose two practical methods for low-dimensional embedding that adhere to our definition of the embeddings family. Exploiting the proposed parallelization strategy, we present two scalable MapReduce algorithms for kernel k-means. We demonstrate the effectiveness and efficiency of the proposed algorithms through an empirical evaluation on benchmark datasets. Ahmed Elgohary, Ahmed K. Farahat, Mohamed S. Kamel, Fakhri Karray |
SDM | 1 |
| 2013 | Distributed Column Subset Selection on MapReduceabstractGiven a very large data set distributed over a cluster of several nodes, this paper addresses the problem of selecting a few data instances that best represent the entire data set. The solution to this problem is of a crucial importance in the big data era as it enables data analysts to understand the insights of the data and explore its hidden structure. The selected instances can also be used for data preprocessing tasks such as learning a low-dimensional embedding of the data points or computing a low-rank approximation of the corresponding matrix. The paper first formulates the problem as the selection of a few representative columns from a matrix whose columns are massively distributed, and it then proposes a MapReduce algorithm for selecting those representatives. The algorithm first learns a concise representation of all columns using random projection, and it then solves a generalized column subset selection problem at each machine in which a subset of columns are selected from the sub-matrix on that machine such that the reconstruction error of the concise representation is minimized. The paper then demonstrates the effectiveness and efficiency of the proposed algorithm through an empirical evaluation on benchmark data sets. Ahmed K. Farahat, Ahmed Elgohary, Ali Ghodsi 0001, Mohamed S. Kamel |
ICDM | 2 |
| 2012 | No need to war-drive: unsupervised indoor localizationabstractWe propose UnLoc, an unsupervised indoor localization scheme that bypasses the need for war-driving. Our key observation is that certain locations in an indoor environment present identifiable signatures on one or more sensing dimensions. An elevator, for instance, imposes a distinct pattern on a smartphone's accelerometer; a corridor-corner may overhear a unique set of WiFi access points; a specific spot may experience an unusual magnetic fluctuation. We hypothesize that these kind of signatures naturally exist in the environment, and can be envisioned as internal landmarks of a building. Mobile devices that "sense" these landmarks can recalibrate their locations, while dead-reckoning schemes can track them between landmarks. Results from 3 different indoor settings, including a shopping mall, demonstrate median location errors of 1:69m. War-driving is not necessary, neither are floorplans the system simultaneously computes the locations of users and landmarks, in a manner that they converge reasonably quickly. We believe this is an unconventional approach to indoor localization, holding promise for real-world deployment. He Wang 0008, Souvik Sen, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001, Romit Roy Choudhury |
MobiSys | 3 |
| 2012 | Demo: unsupervised indoor localizationabstractWe propose UnLoc [1], an unsupervised indoor localization scheme that bypasses the need for war-driving. Our key observation is that certain locations in an indoor environment present an identifiable signature on one or more sensing dimensions. An elevator, for instance, imposes a distinct pattern on a smartphone's accelerometer; a specific spot may experience an unusual magnetic fluctuation. This form of urban sensing and activity recognition has already been demonstrated in literature [2, 3], but not yet applied in pure localization applications. We hypothesize that these kind of signatures naturally exist in the environment and can be envisioned as internal landmarks of a building. Mobile devices that "sense" these landmarks can recalibrate their locations, while dead-reckoning schemes can track them between landmarks. Neither war-driving nor floorplans are necessary - the system simultaneously computes the locations of users and landmarks, in a manner so that they converge reasonably quickly. We believe this is an unconventional approach to indoor localization, holding promise for real-world deployment. He Wang 0008, Souvik Sen, Alexander Mariakakis, Romit Roy Choudhury, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001 |
MobiSys | 5 |
| 2011 | Efficient data clustering over peer-to-peer networksabstractDue to the dramatic increase of data volumes in different applications, it is becoming infeasible to keep these data in one centralized machine. It is becoming more and more natural to deal with distributed databases and networks. That is why distributed data mining techniques have been introduced. One of the most important data mining problems is data clustering. While many clustering algorithms exist for centralized databases, there is a lack of efficient algorithms for distributed databases. In this paper, an efficient algorithm is proposed for clustering distributed databases. The proposed methodology employs an iterative optimization technique to achieve better clustering objective. The experimental results reported in this paper show the superiority of the proposed technique over a recently proposed algorithm based on a distributed version of the well known K-Means algorithm (Datta et al. 2009) [1]. Ahmed Elgohary, Mohamed A. Ismail |
ISDA | 1 |
| 2010 | Wiki-rec: A semantic-based recommendation system using Wikipedia as an ontologyabstractNowadays, satisfying user needs has become the main challenge in a variety of web applications. Recommender systems play a major role in that direction. However, as most of the information is present in a textual form, recommender systems face the challenge of efficiently analyzing huge amounts of text. The usage of semantic-based analysis has gained much interest in recent years. The emergence of ontologies has yet facilitated semantic interpretation of text. However, relying on an ontology for performing the semantic analysis requires too much effort to construct and maintain the used ontologies. Besides, the currently known ontologies cover a small number of the world's concepts especially when a non-domain-specific concepts are needed. This paper proposes the use of Wikipedia as ontology to solve the problems of using traditional ontologies for the text analysis in text-based recommendation systems. A full system model that unifies semantic-based analysis with a collaborative via content recommendation system is presented. Ahmed Elgohary, Hussein Nomir, Ibrahim Sabek, Mohamed Samir, Moustafa Badawy, Noha A. Yousri |
ISDA | 1 |