Ahmed Elgohary

dblp:03/8978 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorComputer networks · 4Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Question answering and dialogue systems · 29% Trustworthy machine learning · 28% Information extraction and text analysis · 20%
Computer networks
4 papers
Wireless sensing and localization · 100%
Databases, data mining, and information retrieval
3 papers
Machine learning and data management · 58% Data models and query languages · 25% Data mining · 17%

Topics — the 24 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025
Wireless sensing and localization
indoor localization
0.742016
SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016
Video: Unsupervised indoor localization (UnLoc): beyond the prototype · MobiSys 2014
Demo: unsupervised indoor localization · MobiSys 2012
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.722019
Can You Unpack That? Learning to Rewrite Questions-in-Context · EMNLP/IJCNLP (1) 2019
A dataset and baselines for sequential open-domain question answering · EMNLP 2018
Natural language and speech › Information extraction and text analysis › text classification
deception detection
0.412020
It Takes Two to Lie: One to Lie, and One to Listen · ACL 2020
Natural language and speech › Question answering and dialogue systems › natural language interface
interactive semantic parsing
0.412020
Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020
Natural language and speech › Information extraction and text analysis
semantic parsing
0.412020
Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.412020
Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback · ACL 2020
Wireless sensing and localization › indoor localization
landmark-based localization
0.422016
SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016
No need to war-drive: unsupervised indoor localization · MobiSys 2012
Natural language and speech › Question answering and dialogue systems
question rewriting
0.412019
Can You Unpack That? Learning to Rewrite Questions-in-Context · EMNLP/IJCNLP (1) 2019
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.312018
Generating Natural Language Adversarial Examples · EMNLP 2018
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.312018
A dataset and baselines for sequential open-domain question answering · EMNLP 2018
Machine learning › Trustworthy machine learning
robustness
0.312018
Generating Natural Language Adversarial Examples · EMNLP 2018
Machine learning and data management
scalable machine learning
0.312018
Compressed linear algebra for large-scale machine learning · VLDB J. 2018
Machine learning › Trustworthy machine learning
fairness
0.312025
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.312025
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements · ICLR 2025
Wireless sensing and localization
simultaneous localization and mapping
0.212016
SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016
Data mining › sampling
representative selection
0.212013
Distributed Column Subset Selection on MapReduce · ICDM 2013
Algorithms and data structures › matrix approximation
column subset selection
0.212013
Distributed Column Subset Selection on MapReduce · ICDM 2013
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing
0.132014
Video: Unsupervised indoor localization (UnLoc): beyond the prototype · MobiSys 2014
Demo: unsupervised indoor localization · MobiSys 2012
No need to war-drive: unsupervised indoor localization · MobiSys 2012
Machine learning › Efficient and distributed learning
model compression
0.112018
Compressed linear algebra for large-scale machine learning · VLDB J. 2018
Natural language and speech › Language models and text generation
text generation
0.112018
Generating Natural Language Adversarial Examples · EMNLP 2018
Robotics › Robot navigation and mapping
SLAM
0.112016
SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization · IEEE Trans. Mob. Comput. 2016
Parallel and multicore computing › data-parallel programming
mapreduce
0.012013
Distributed Column Subset Selection on MapReduce · ICDM 2013

Methods — techniques the papers use, named apart from their topics

dead reckoning · 1.2safety configs · 0.9in-context alignment · 0.9data-centric alignment · 0.9landmark sensing · 0.7random projection · 0.5low-rank approximation · 0.5sequence-to-structure models · 0.4question rewriting · 0.4population-based optimization · 0.3linear algebra compression · 0.3dataset construction · 0.3black-box optimization · 0.3baseline · 0.3adversarial training · 0.3wifi fingerprinting · 0.3magnetometer sensing · 0.3accelerometer sensing · 0.3
YearPublicationVenuePosition
2025 Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
abstract
The current paradigm for safety alignment of large language models (LLMs) follows a _one-size-fits-all_ approach: the model refuses to interact with any content deemed unsafe by the model provider. This approach lacks flexibility in the face of varying social norms across cultures and regions. In addition, users may have diverse safety needs, making a model with _static_ safety standards too restrictive to be useful, as well as too costly to be re-aligned. We propose _Controllable Safety Alignment_ (CoSA), a framework designed to adapt models to diverse safety requirements without re-training. Instead of aligning a fixed model, we align models to follow _safety configs_—free-form natural language descriptions of the desired safety behaviors—that are provided as part of the system prompt. To adjust model safety behavior, authorized users only need to modify such safety configs at inference time. To enable that, we propose CoSAlign, a data-centric method for aligning LLMs to easily adapt to diverse safety configs. Furthermore, we devise a novel controllability evaluation protocol that considers both helpfulness and configured safety, summarizing them into CoSA-Score, and construct CoSApien, a _human-authored_ benchmark that consists of real-world LLM use cases with diverse safety requirements and corresponding evaluation prompts. We show that CoSAlign leads to substantial gains of controllability over strong baselines including in-context alignment. Our framework encourages better representation and adaptation to pluralistic human values in LLMs, and thereby increasing their practicality.
Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, Benjamin Van Durme
ICLR2
2021 NL-EDIT: Correcting Semantic Parse Errors through Natural Language Interaction
abstract
Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo Ramos, Ahmed Hassan Awadallah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo A. Ramos, Ahmed Awadallah 0001
NAACL-HLT1
2020 Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback
abstract
We study the task of semantic parse correction with natural language feedback.Given a natural language utterance, most semantic parsing systems pose the problem as one-shot translation where the utterance is mapped to a corresponding logical form.In this paper, we investigate a more interactive scenario where humans can further interact with the system by providing free-form natural language feedback to correct the system when it generates an inaccurate interpretation of an initial utterance.We focus on natural language to SQL systems and construct, SPLASH, a dataset of utterances, incorrect SQL interpretations and the corresponding natural language feedback.We compare various reference models for the correction task and show that incorporating such a rich form of feedback can significantly improve the overall semantic parsing accuracy while retaining the flexibility of natural language interaction.While we estimated human correction accuracy is 81.5%, our best model achieves only 25.1%, which leaves a large gap for improvement in future research.SPLASH is publicly available at https:// aka.ms/Splash_dataset. Exact Match Accuracy (%) BaselineCorrection End-to-End Without Feedback ⇒ Seq2Struct N/A 41.30 ⇒ Re-
Ahmed Elgohary, Saghar Hosseini, Ahmed Awadallah 0001
ACL1
2020 It Takes Two to Lie: One to Lie, and One to Listen
abstract
Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan Boyd-Graber. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, Jordan L. Boyd-Graber
ACL3
2019 Can You Unpack That? Learning to Rewrite Questions-in-Context
abstract
Ahmed Elgohary, Denis Peskov, Jordan Boyd-Graber. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ahmed Elgohary, Denis Peskov, Jordan L. Boyd-Graber
EMNLP/IJCNLP (1)1
2018 Assessing Composition in Sentence Vector Representations
abstract
An important component of achieving language understanding is mastering the composition of sentence meaning, but an immediate challenge to solving this problem is the opacity of sentence vector representations produced by current neural sentence composition models. We present a method to address this challenge, developing tasks that directly target compositional meaning information in sentence vector representations with a high degree of precision and control. To enable the creation of these controlled tasks, we introduce a specialized sentence generation system that produces large, annotated sentence sets meeting specified syntactic, semantic and lexical constraints. We describe the details of the method and generation system, and then present results of experiments applying our method to probe for compositional information in embeddings from a number of existing sentence composition models. We find that the method is able to extract useful information about the differing capacities of these models, and we discuss the implications of our results with respect to these systems’ capturing of sentence information. We make available for public use the datasets used for these experiments, as well as the generation system.
Allyson Ettinger, Ahmed Elgohary, Colin Phillips, Philip Resnik
COLING2
2018 Generating Natural Language Adversarial Examples
abstract
Deep neural networks (DNNs) are vulnerable to adversarial examples, perturbations to correctly classified examples which can cause the model to misclassify.In the image domain, these perturbations are often virtually indistinguishable to human perception, causing humans and state-of-the-art models to disagree.However, in the natural language domain, small perturbations are clearly perceptible, and the replacement of a single word can drastically alter the semantics of the document.Given these challenges, we use a black-box population-based optimization algorithm to generate semantically and syntactically similar adversarial examples that fool well-trained sentiment analysis and textual entailment models with success rates of 97% and 70%, respectively.We additionally demonstrate that 92.3% of the successful sentiment analysis adversarial examples are classified to their original label by 20 human annotators, and that the examples are perceptibly quite similar.Finally, we discuss an attempt to use adversarial training as a defense, but fail to yield improvement, demonstrating the strength and diversity of our adversarial examples.We hope our findings encourage researchers to pursue improving the robustness of DNNs in the natural language domain.
Moustafa Farid Alzantot, Yash Sharma 0001, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava 0001, Kai-Wei Chang 0001
EMNLP3
2018 A dataset and baselines for sequential open-domain question answering
abstract
Previous work on question-answering systems mainly focuses on answering individual questions, assuming they are independent and devoid of context.Instead, we investigate sequential question answering, asking multiple related questions.We present QBLink, a new dataset of fully human-authored questions.We extend existing strong question answering frameworks to include previous questions to improve the overall question-answering accuracy in open-domain question answering.The dataset is publicly available at http:// sequential.qanta.org.
Ahmed Elgohary, Chen Zhao 0013, Jordan L. Boyd-Graber
EMNLP1
2018 Compressed linear algebra for large-scale machine learning
Ahmed Elgohary, Matthias Boehm 0001, Peter J. Haas, Frederick Reiss 0001, Berthold Reinwald
VLDB J.1
2016 Compressed Linear Algebra for Large-Scale Machine Learning
abstract
Large-scale machine learning (ML) algorithms are often iterative, using repeated read-only data access and I/O-bound matrix-vector multiplications to converge to an optimal model. It is crucial for performance to fit the data into single-node or distributed main memory. General-purpose, heavy- and lightweight compression techniques struggle to achieve both good compression ratios and fast decompression speed to enable block-wise uncompressed operations. Hence, we initiate work on compressed linear algebra (CLA), in which lightweight database compression techniques are applied to matrices and then linear algebra operations such as matrix-vector multiplication are executed directly on the compressed representations. We contribute effective column compression schemes, cache-conscious operations, and an efficient sampling-based compression algorithm. Our experiments show that CLA achieves in-memory operations performance close to the uncompressed case and good compression ratios that allow us to fit larger datasets into available memory. We thereby obtain significant end-to-end performance improvements up to 26x or reduced memory requirements.
Ahmed Elgohary, Matthias Boehm 0001, Peter J. Haas, Frederick Reiss 0001, Berthold Reinwald
Proc. VLDB Endow.1
2016 SemanticSLAM: Using Environment Landmarks for Unsupervised Indoor Localization
abstract
Indoor localization using mobile sensors has gained momentum lately. Most of the current systems rely on an extensive calibration step to achieve high accuracy. We propose SemanticSLAM, a novel unsupervised indoor localization scheme that bypasses the need for war-driving. SemanticSLAM leverages the idea that certain locations in an indoor environment have a unique signature on one or more phone sensors. Climbing stairs, for example, has a distinct pattern on the phone's accelerometer; a specific spot may experience an unusual magnetic interference while another may have a unique set of Wi-Fi access points covering it. SemanticSLAM uses these unique points in the environment as landmarks and combines them with dead-reckoning in a new Simultaneous Localization And Mapping (SLAM) framework to reduce both the localization error and convergence time. In particular, the phone inertial sensors are used to keep track of the user's path, while the observed landmarks are used to compensate for the accumulation of error in a unified probabilistic framework. Evaluation in two testbeds on Android phones shows that the system can achieve 0.53 meters human median localization errors. In addition, the system can detect the location of landmarks with 0.83 meters median error. This is 62 percent better than a system that does not use SLAM. Moreover, SemanticSLAM has a 33 percent lower convergence time compared to the same systems. This highlights the promise of SemanticSLAM as an unconventional approach for indoor localization.
Heba Abdelnasser, Reham Mohamed Aburas, Ahmed Elgohary, Moustafa Farid Alzantot, He Wang 0008, Souvik Sen, Romit Roy Choudhury, Moustafa Youssef 0001
IEEE Trans. Mob. Comput.3
2015 Greedy column subset selection for large-scale data sets
Ahmed K. Farahat, Ahmed Elgohary, Ali Ghodsi 0001, Mohamed S. Kamel
Knowl. Inf. Syst.2
2014 Video: Unsupervised indoor localization (UnLoc): beyond the prototype
abstract
This video presents a demo of indoor localization in multiple settings. In the demo, a user walks with a smartphone and the user's location is shown on the phone's screen in real time. Our system, called Unsupervised Indoor Localization (UnLoc) utilizes the sensor data from smartphones to learn "invisible landmarks" in the environment. Example landmarks could be a unique magnetic fluctuation experienced when the phone is near a water-cooler, or a distinct gyroscope rotation when the user turns a corner. We use these indoor "landmarks" to periodically reset the user's location. To track the user between these landmarks, we use an optimized variant of dead reckoning, ultimately leading to a robust location tracking system. We call our system UnLoc, since the landmarks are generated in an unsupervised manner, requiring no manual effort or floorplan of the building. The demo describes the high level intuitions, shows UnLoc in operation, and shares experiences from running UnLoc in various real-world environments.
He Wang 0008, Souvik Sen, Alexander Mariakakis, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001, Romit Roy Choudhury
MobiSys4
2014 Embed and Conquer: Scalable Embeddings for Kernel k-Means on MapReduce
abstract
The kernel k-means is an effective method for data clustering which extends the commonly-used k-means algorithm to work on a similarity matrix over complex data structures. It is, however, computationally very complex as it requires the complete kernel matrix to be calculated and stored. Further, its kernelized nature hinders the parallelization of its computations on modern scalable infrastructures for distributed computing. In this paper, we are defining a family of kernelbased low-dimensional embeddings that allows for scaling kernel k-means on MapReduce via an efficient and unified parallelization strategy. Afterwards, we propose two practical methods for low-dimensional embedding that adhere to our definition of the embeddings family. Exploiting the proposed parallelization strategy, we present two scalable MapReduce algorithms for kernel k-means. We demonstrate the effectiveness and efficiency of the proposed algorithms through an empirical evaluation on benchmark datasets.
Ahmed Elgohary, Ahmed K. Farahat, Mohamed S. Kamel, Fakhri Karray
SDM1
2013 Distributed Column Subset Selection on MapReduce
abstract
Given a very large data set distributed over a cluster of several nodes, this paper addresses the problem of selecting a few data instances that best represent the entire data set. The solution to this problem is of a crucial importance in the big data era as it enables data analysts to understand the insights of the data and explore its hidden structure. The selected instances can also be used for data preprocessing tasks such as learning a low-dimensional embedding of the data points or computing a low-rank approximation of the corresponding matrix. The paper first formulates the problem as the selection of a few representative columns from a matrix whose columns are massively distributed, and it then proposes a MapReduce algorithm for selecting those representatives. The algorithm first learns a concise representation of all columns using random projection, and it then solves a generalized column subset selection problem at each machine in which a subset of columns are selected from the sub-matrix on that machine such that the reconstruction error of the concise representation is minimized. The paper then demonstrates the effectiveness and efficiency of the proposed algorithm through an empirical evaluation on benchmark data sets.
Ahmed K. Farahat, Ahmed Elgohary, Ali Ghodsi 0001, Mohamed S. Kamel
ICDM2
2012 No need to war-drive: unsupervised indoor localization
abstract
We propose UnLoc, an unsupervised indoor localization scheme that bypasses the need for war-driving. Our key observation is that certain locations in an indoor environment present identifiable signatures on one or more sensing dimensions. An elevator, for instance, imposes a distinct pattern on a smartphone's accelerometer; a corridor-corner may overhear a unique set of WiFi access points; a specific spot may experience an unusual magnetic fluctuation. We hypothesize that these kind of signatures naturally exist in the environment, and can be envisioned as internal landmarks of a building. Mobile devices that "sense" these landmarks can recalibrate their locations, while dead-reckoning schemes can track them between landmarks. Results from 3 different indoor settings, including a shopping mall, demonstrate median location errors of 1:69m. War-driving is not necessary, neither are floorplans the system simultaneously computes the locations of users and landmarks, in a manner that they converge reasonably quickly. We believe this is an unconventional approach to indoor localization, holding promise for real-world deployment.
He Wang 0008, Souvik Sen, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001, Romit Roy Choudhury
MobiSys3
2012 Demo: unsupervised indoor localization
abstract
We propose UnLoc [1], an unsupervised indoor localization scheme that bypasses the need for war-driving. Our key observation is that certain locations in an indoor environment present an identifiable signature on one or more sensing dimensions. An elevator, for instance, imposes a distinct pattern on a smartphone's accelerometer; a specific spot may experience an unusual magnetic fluctuation. This form of urban sensing and activity recognition has already been demonstrated in literature [2, 3], but not yet applied in pure localization applications. We hypothesize that these kind of signatures naturally exist in the environment and can be envisioned as internal landmarks of a building. Mobile devices that "sense" these landmarks can recalibrate their locations, while dead-reckoning schemes can track them between landmarks. Neither war-driving nor floorplans are necessary - the system simultaneously computes the locations of users and landmarks, in a manner so that they converge reasonably quickly. We believe this is an unconventional approach to indoor localization, holding promise for real-world deployment.
He Wang 0008, Souvik Sen, Alexander Mariakakis, Romit Roy Choudhury, Ahmed Elgohary, Moustafa Farid Alzantot, Moustafa Youssef 0001
MobiSys5
2011 Efficient data clustering over peer-to-peer networks
abstract
Due to the dramatic increase of data volumes in different applications, it is becoming infeasible to keep these data in one centralized machine. It is becoming more and more natural to deal with distributed databases and networks. That is why distributed data mining techniques have been introduced. One of the most important data mining problems is data clustering. While many clustering algorithms exist for centralized databases, there is a lack of efficient algorithms for distributed databases. In this paper, an efficient algorithm is proposed for clustering distributed databases. The proposed methodology employs an iterative optimization technique to achieve better clustering objective. The experimental results reported in this paper show the superiority of the proposed technique over a recently proposed algorithm based on a distributed version of the well known K-Means algorithm (Datta et al. 2009) [1].
Ahmed Elgohary, Mohamed A. Ismail
ISDA1
2010 Wiki-rec: A semantic-based recommendation system using Wikipedia as an ontology
abstract
Nowadays, satisfying user needs has become the main challenge in a variety of web applications. Recommender systems play a major role in that direction. However, as most of the information is present in a textual form, recommender systems face the challenge of efficiently analyzing huge amounts of text. The usage of semantic-based analysis has gained much interest in recent years. The emergence of ontologies has yet facilitated semantic interpretation of text. However, relying on an ontology for performing the semantic analysis requires too much effort to construct and maintain the used ontologies. Besides, the currently known ontologies cover a small number of the world's concepts especially when a non-domain-specific concepts are needed. This paper proposes the use of Wikipedia as ontology to solve the problems of using traditional ontologies for the text analysis in text-based recommendation systems. A full system model that unifies semantic-based analysis with a collaborative via content recommendation system is presented.
Ahmed Elgohary, Hussein Nomir, Ibrahim Sabek, Mohamed Samir, Moustafa Badawy, Noha A. Yousri
ISDA1