EDBT 2026 Demo / reviewers in the wild / expert
Mohammad Taha Khan
dblp:154/3627
· DBLP profile ↗
9ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0003-4743-5610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 2 since 2021Computer networks · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the Impact of Code Changes on the Fault Localizability of Large Language ModelsabstractGenerative Large Language Models (LLMs) are increasingly used in non-generative software maintenance tasks, such as fault localization (FL). Success in FL depends on a models ability to reason about program semantics beyond surface-level syntactic and lexical features. However, widely used LLM benchmarks primarily evaluate code generation, which differs fundamentally from semantic program reasoning. Meanwhile, traditional FL benchmarks such as Defect4J and BugsInPy are either not scalable or obsolete, as their datasets have become part of LLM training data, leading to biased results. This paper presents the first large-scale empirical investigation into the robustness of LLMs fault localizability. Inspired by mutation testing, we develop an end-to-end evaluation framework that addresses key limitations in existing LLM evaluation, including data contamination, scalability, automation, and extensibility. Using real-world programs with specifications, we inject unseen faults and ask LLMs to localize them, filtering out underspecified programs where localization is ambiguous. For each successfully localized program, we apply semantic-preserving mutations (SPMs) and rerun localization to assess robustness and determine whether LLM reasoning relies on syntactic cues rather than semantics. We evaluate 10 state-of-the-art LLMs on 750,013 fault localization tasks from over 1,300 Java and Python programs. We find that SPMs cause LLMs to fail on previously localized faults in 78% of cases, and that reasoning is stronger when relevant code appears earlier in context. These results indicate that LLM code reasoning is often tied to features irrelevant to semantics. We also identify code patterns that are challenging for LLMs to reason about. Overall, our findings motivate fundamental advances in how LLMs represent, interpret, and prioritize code semantics to reason more deeply about program logic Sabaat Haroon, Ahmad Khan 0001, Ahmad Humayun, Waris Gill, Abdul Haddi Amjad, Ali Raza Butt, Mohammad Taha Khan, Muhammad Ali Gulzar |
ICST | 7 |
| 2021 | Helping Users Automatically Find and Manage Sensitive, Expendable Files in Cloud Storage
Mohammad Taha Khan, Christopher Tran 0001, Dimitri Vasilkov, Chris Kanich, Blase Ur, Elena Zheleva |
USENIX Security Symposium | 1 |
| 2021 | Blind In/On-Path Attacks and Applications to VPNs
William J. Tolley, Beau Kujath, Mohammad Taha Khan, Narseo Vallina-Rodriguez, Jedidiah R. Crandall |
USENIX Security Symposium | 3 |
| 2019 | Moving Beyond Set-It-And-Forget-It Privacy Settings on Social MediaabstractWhen users post on social media, they protect their privacy by choosing an access control setting that is rarely revisited. Changes in users' lives and relationships, as well as social media platforms themselves, can cause mismatches between a post's active privacy setting and the desired setting. The importance of managing this setting combined with the high volume of potential friend-post pairs needing evaluation necessitate a semi-automated approach. We attack this problem through a combination of a user study and the development of automated inference of potentially mismatched privacy settings. A total of 78 Facebook users reevaluated the privacy settings for five of their Facebook posts, also indicating whether a selection of friends should be able to access each post. They also explained their decision. With this user data, we designed a classifier to identify posts with currently incorrect sharing settings. This classifier shows a 317% improvement over a baseline classifier based on friend interaction. We also find that many of the most useful features can be collected without user intervention, and we identify directions for improving the classifier's accuracy. Mainack Mondal, Günce Su Yilmaz, Noah Hirsch, Mohammad Taha Khan, Michael Tang, Christopher Tran 0001, Chris Kanich, Blase Ur, Elena Zheleva |
CCS | 4 |
| 2018 | Forgotten But Not Gone: Identifying the Need for Longitudinal Data Management in Cloud StorageabstractUsers have accumulated years of personal data in cloud storage, creating potential privacy and security risks. This agglomeration includes files retained or shared with others simply out of momentum, rather than intention. We presented 100 online-survey participants with a stratified sample of 10 files currently stored in their own Dropbox or Google Drive accounts. We asked about the origin of each file, whether the participant remembered that file was stored there, and, when applicable, about that file's sharing status. We also recorded participants' preferences moving forward for keeping, deleting, or encrypting those files, as well as adjusting sharing settings. Participants had forgotten that half of the files they saw were in the cloud. Overall, 83% of participants wanted to delete at least one file they saw, while 13% wanted to unshare at least one file. Our combined results suggest directions for retrospective cloud data management. Mohammad Taha Khan, Maria Hyun, Chris Kanich, Blase Ur |
CHI | 1 |
| 2018 | An Empirical Analysis of the Commercial VPN Ecosystem
Mohammad Taha Khan, Joe DeBlasio, Geoffrey M. Voelker, Alex C. Snoeren, Chris Kanich, Narseo Vallina-Rodriguez |
Internet Measurement Conference | 1 |
| 2016 | Sneak-Peek: High speed covert channels in data center networksabstractWith the advent of big data, modern businesses face an increasing need to store and process large volumes of sensitive customer information on the cloud. In these environments, resources are shared across a multitude of mutually untrusting tenants increasing propensity for data leakage. This problem stands to grow further in severity with increasing use of clouds in all aspects of our daily lives and the recent spate of high-profile data exfiltration attacks are evidence. To highlight this serious issue, we present a novel and highspeed network-based covert channel that is robust and circumvents a broad set of security mechanisms currently deployed by cloud vendors. We successfully test our channel on numerous network environments, including commercial clouds such as EC2 and Azure. Using an information theoretic model of the channel, we derive an upper bound on the maximum information rate and propose an optimal coding scheme. Our adaptive decoding algorithm caters to the cross traffic in the channel and maintains high bit rates and extremely low error rates. Finally, we discuss several effective avenues for mitigation of the aforementioned channel and provide insights into how data exfiltration can be prevented in such shared environments. Rashid Tahir, Mohammad Taha Khan, Xun Gong 0001, AmirEmad Ghassami, Hasanat Kazmi, Matthew Caesar 0001, Fareed Zaffar, Negar Kiyavash |
INFOCOM | 2 |
| 2015 | Every Second Counts: Quantifying the Negative Externalities of Cybercrime via TyposquattingabstractWhile we have a good understanding of how cyber crime is perpetrated and the profits of the attackers, the harm experienced by humans is less well understood, and reducing this harm should be the ultimate goal of any security intervention. This paper presents a strategy for quantifying the harm caused by the cyber crime of typo squatting via the novel technique of intent inference. Intent inference allows us to define a new metric for quantifying harm to users, develop a new methodology for identifying typo squatting domain names, and quantify the harm caused by various typo squatting perpetrators. We find that typo squatting costs the typical user 1.3 seconds per typo squatting event over the alternative of receiving a browser error page, and legitimate sites lose approximately 5% of their mistyped traffic over the alternative of an unregistered typo. Although on average perpetrators increase the time it takes a user to find their intended site, many typo squatters actually improve the latency between a typo and its correction, calling into question the necessity of harsh penalties or legal intervention against this flavor of cyber crime. Mohammad Taha Khan, Xiang Huo, Zhou Li 0001, Chris Kanich |
IEEE Symposium on Security and Privacy | 1 |
| 2014 | Efficient relaying strategy selection and signal combining using error estimation codesabstractIn this paper, we use recently developed error estimation coding (EEC) to devise an efficient relaying framework for a multi-relay network. EEC utilizes an added redundancy to form an estimate of the bit-error rate (BER) of data received over a noisy channel. Utilizing the BER estimate, we propose a relaying selection strategy that uses the BER estimate at any given relay to switch between Amplify-Forward (AF) and Detect-Forward (DF) cooperation. In addition, we propose a signal combining rule at the destination that weighs different copies of the same data using the corresponding BER estimates. For performance evaluations, we implement the proposed scheme on a USRP-based platform, and test its performance by conducting experiments in an indoor office environment. Our results validate the efficacy of the proposed strategy; the proposed scheme significantly outperforms standard equal gain combining-based AF and DF cooperation. Mohammad Taha Khan, Talha A. Anwar, Muhammad Kumail Haider, Momin Uppal |
WCNC | 1 |