Ha Dao

dblp:232/9577 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-1950-8384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 "Nobody should control the end user": Exploring Privacy Perspectives of Indian Internet Users in Light of DPDPA
abstract
With the rapid increase in online interactions, concerns over data privacy and transparency of data processing practices have become more pronounced. While regulations like the GDPR have driven the widespread adoption of cookie banners in the EU, India's Digital Personal Data Protection Act (DPDPA) promises similar changes domestically, aiming to introduce a framework for data protection. However, certain clauses within the DPDPA raise concerns about potential infringements on user privacy, given the exemptions for government accountability and user consent requirements. In this study, for the first time, we explore Indian Internet users' awareness and perceptions of cookie banners, online privacy, and privacy regulations, especially in light of the newly passed DPDPA. We conducted an online anonymous survey with 428 Indian participants, which addressed: (1) users' perspectives on cookie banners, (2) their attitudes towards online privacy and privacy regulations, and (3) their acceptance of 10 contentious DPDPA clauses that favor state authorities and may enable surveillance. Our findings reveal that privacy-conscious users often lack consistent awareness of privacy mechanisms, and their concerns do not always lead to protective actions. Our thematic analysis of 143 open-ended responses shows that users' privacy and data protection concerns are rooted in skepticism towards the government, shaping their perceptions of the DPDPA and fueling demands for policy revisions. Our study highlights the need for clearer communication regarding the DPDPA, user-centric consent mechanisms, and policy refinements to enhance data privacy practices in India.
Sana Athar, Devashish Gosain, Anja Feldmann, Mannat Kaur 0001, Ha Dao
AsiaCCS5
2026 Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Nicolas Steinacker-Olsztyn, Devashish Gosain, Ha Dao
WWW3
2026 Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads
abstract
YouTube is one of the largest video platforms on the web, with Google Ads deeply integrated into the viewing experience. While users may expect some level of tracking during ad delivery, the extent and mechanics of Google Ads tracking on YouTube, particularly the tracking behaviors triggered by user interactions with ads, remain underexplored. To address this gap, for the first time, we develop YT-AdTrack, a fully automated framework that measures tracking initiated by YouTube ads across 430 top-trending videos, three widely used browsers, and six geographic locations. In our baseline measurement campaign, YT-AdTrack strategically accepts cookie banners on YouTube, interacts with displayed ads, and subsequently accepts cookie banners on advertiser landing pages to capture downstream tracking behavior. Our findings show that every ad click consistently carries a unique Google Click Identifier, gclid, which is propagated through redirection chains and ultimately embedded in the advertiser’s landing page. We further observe that 64 (out of 76) advertisers persist this identifier as a first-party cookie, thereby transforming a short-lived click token into a durable user identifier. In addition, gclid values frequently leak across parties from advertisers, exposing them to both Google-controlled services and external third-party ad networks, which exacerbates the tracking nexus by extending well beyond standard conversion measurement. Strikingly, even when cookie banners are rejected, ad interactions remain consistently tagged: 55.4% of advertisers store gclidas a cookie, and 18.9% enable auto-tagging, which allows Googleto directly persist the identifier. This demonstrates that banner rejection does not safeguard users from gclid-based tracking. We find these behaviors to be consistent across browsers, with advertisers persisting the identifier in 76.4–84.2% of cases and Google directly storing it in over two-thirds of interactions. Similarly, across six geographic locations, gclid-based tracking persists, with advertisers storing it in 72.5–88.5% of cases and Google’s auto-tagging active in all locations. Overall, our analysis reveals that a single ad click can initiate durable cross-site tracking that persists across various banner choices, browser environments, and regional contexts.
Ha Dao, Abhishek Shinde, Sana Athar, Devashish Gosain
Proc. Priv. Enhancing Technol.1
2025 Harnessing the Power of LLMs for Code Smell Detection in Terraform Infrastructure as Code
abstract
Terraform is a widely used Infrastructure as Code (IaC) tool that simplifies cloud resource management through declarative configuration. However, Terraform configurations often exhibit code smells, which can introduce security vulnerabilities, maintainability challenges, and operational inefficiencies. While code smells have been extensively studied in other IaC platforms, research on Terraform remains limited, and existing static analysis tools struggle to detect a broad range of code smells. In this paper, we harness large language models (LLMs) for automated Terraform code smell detection, leveraging their ability to generalize beyond predefined rule-based heuristics. We construct a synthetic benchmark dataset of 42 Terraform configurations covering 14 distinct code smells and evaluate detection performance across traditional linters and LLM-based approaches. Our results show that LLMs significantly outperform static analysis tools, detecting a broader range of code smells, especially logic-related code smells. We extend our analysis to real-world repositories with 71 high-quality Terraform projects from GitHub. Our findings reveal that 84.5% of repositories contain at least one code smell. Notably, some repositories accumulate a high number of unique code smells, with one exhibiting nine distinct issues, underscoring severe quality concerns. Moreover, we find that code smells persist across repositories regardless of popularity, affecting even widely used projects with thousands of stars and forks.
Quoc-Huy Vo, Ha Dao, Kensuke Fukuda
COMPSAC2
2025 A First Look at Cookies Having Independent Partitioned State
Maximilian Zöllner, Anja Feldmann, Ha Dao
PAM3
2025 Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies
abstract
In response to the ePrivacy Directive and the consent requirements introduced by the GDPR, websites began deploying consent banners to obtain user permission for data collection and processing. However, due to shared third-party services and technical loopholes, non-consensual cross-site tracking can still occur. In fact, contrary to user expectations of seemingly isolated consent, a user's decision on one website may affect tracking behavior on others. In this study, we investigate the technical and behavioral mechanisms behind these discrepancies. Specifically, we disclose a persistent tracking mechanism exploiting web cookies. These cookies, which we refer to as intractable, are initially set on websites with accepted banners, persist in the browser, and are subsequently sent to trackers before the user provides explicit consent on other websites. To meticulously analyze this covert tracking behavior, we conduct an extensive measurement study performing stateful crawls on over 20k domains from the Tranco top list, strategically accepting banners in the first half of domains and measuring intractable cookies in the second half. Our findings reveal that around 50% of websites send at least one intractable cookie, with the majority set to expire after more than 10 days. In addition, enabling the Global Privacy Control (GPC) signal initially reduces the number of intractable cookies by 30% on average, with a further 32% reduction possible on subsequent visits by rejecting the banners. Moreover, websites with Consent Management Platform (CMP) banners, on average, send 6.9 times more intractable cookies compared to those with native banners. Our research further reveals that even if users reject all other banners, they still receive a large number of intractable cookies set by websites with cookie paywalls. Additionally, our measurement on the partitioned cookies---cookies that are restricted to the top-level site and thus mitigate cross-site tracking---shows that only 1.3% of tracking cookies are marked as such, indicating their minimal impact on cross-site tracking via intractable cookies.
Ali Rasaii, Ha Dao, Anja Feldmann, Mohammadmahdi Javid, Oliver Gasser, Devashish Gosain
Proc. Priv. Enhancing Technol.2
2025 Unmasking the Shadows: A Cross-Country Study of Online Tracking in Illegal Movie Streaming Services
abstract
The proliferation of Illegal Movie Streaming Services (IMSS) has posed significant challenges to legitimate streaming services and law enforcement alike, causing financial losses and complicating efforts to combat copyright infringement. Motivated by the absence of a comprehensive list of IMSS, and recognizing that IMSS websites often have short-lived domains, we first introduce a methodology to detect IMSS sites. Our evaluation demonstrates that our method achieves a recall of 84.31% in identifying IMSS. Applying this method on the Tranco Top 1M domains, we find 283 new websites hosting IMSS. When characterizing the IMSS ecosystem, our findings reveal that four specific IMSS sites attract considerable attention, appearing in the Tranco Top 10K domains. Additionally, these sites employ complex redirection patterns, with one site using up to 11 hops to evade detection. Using Google Identifiers, we then uncover 11 cases of co-ownership, where multiple sites share the same identifiers, indicating common operation. Finally, by crawling IMSS sites from seven vantage points (VPs), we investigate online tracking practices on these services — an area that has previously lacked thorough investigation. We find that more than 95% of IMSS include at least one third-party tracker on their websites. Interestingly, tracker presence is lower in the European Union (EU) countries than in other VPs. Furthermore, third-party tracking cookies are not the primary mechanism on IMSS sites; instead, the more invasive and unavoidable fingerprinting techniques are predominantly used for tracking across different VPs.
Hussein Sheaib, Anja Feldmann, Ha Dao
Proc. Priv. Enhancing Technol.3
2021 Alternative to third-party cookies: investigating persistent PII leakage-based web tracking
abstract
Many popular websites give users the ability to sign up for their services, which requires personally identifiable information (PII). However, these websites embed third-party tracking and advertising resources, and as a consequence, the authentication flow can intentionally or unintentionally leak PII to these services. Since a user can be identified with PII, trackers can use it for tracking purposes, leading to further privacy leaks when cross-site, cross-browser, and cross-device tracking occur.
Ha Dao, Kensuke Fukuda
CoNEXT1
2021 CNAME Cloaking-Based Tracking on the Web: Characterization, Detection, and Protection
abstract
Third-party tracking on the Web has been used for collecting and correlating user's browsing behavior. Due to the increasing use of ad-blocking and third-party tracking protections, tracking providers introduced a new technique called CNAME cloaking. It misleads Web browsers into believing that a request for a subdomain of the visited website originates from this particular website, while this subdomain uses a CNAME to resolve to a tracking-related third-party domain. This technique thus circumvents the third-party targeting privacy protections. The goals of this paper are to characterize, detect, and protect the end-user against CNAME cloaking based tracking. Firstly, we characterize CNAME cloaking-based tracking by crawling top pages of the Alexa Top 300,000 sites and analyzing the usage of CNAME cloaking with CNAME blocklist, including websites and tracking providers using this technique to track users' activities. We also point out that browsers and privacy protection extensions are largely ineffective to deal with CNAME cloaking-based tracking except for Firefox with a developer's version of the uBlock Origin extension. Secondly, we propose a supervised machine learning-based approach to detect CNAME cloaking-based tracking without the on-demand DNS lookup. We show that the proposed approach outperforms well-known tracking filter lists. Finally, to circumvent the lack of DNS API in Chrome-based browsers, we design and implement a prototype of the supervised machine learning-based browser extension to detect and filter out CNAME cloaking tracking, called CNAMETracking Uncloaker. Our evaluation shows that CNAMETracking Uncloaker is able to filter out CNAME cloaking-based tracking requests without performance degradation when compared with the vanilla setting on the Chrome browser.
Ha Dao, Johan Mazel, Kensuke Fukuda
IEEE Trans. Netw. Serv. Manag.1
2020 A machine learning approach for detecting CNAME cloaking-based tracking on the Web
abstract
Various in-browser privacy protection techniques have been designed to protect end-users from third-party tracking. In an arms race against these counter-measures, the tracking providers developed a new technique called CNAME cloaking based tracking to avoid issues with browsers that block third-party cookies and requests. To detect this tracking technique, browser extensions require on-demand DNS lookup APIs. This feature is however only supported by the Firefox browser.In this paper, we propose a supervised machine learning-based method to detect CNAME cloaking-based tracking without the on-demand DNS lookup. Our goal is to detect both sites and requests linked to CNAME cloaking-related tracking. We crawl a list of target sites and store all HTTP/HTTPS requests with their attributes. Then we label all instances automatically by looking up CNAME record of subdomain, and applying wildcard matching based on well-known tracking filter lists. After extracting features, we build a supervised classification model to distinguish site and request related to CNAME cloaking-based tracking. Our evaluation shows that the proposed approach outperforms well-known tracking filter lists: F1 scores of 0.790 for sites and 0.885 for requests. By analyzing the feature permutation importance, we demonstrate that the number of scripts and the proportion of XMLHttpRequests are discriminative for detecting sites, and the length of URL request is helpful in detecting requests. Finally, we analyze concept drift by using the 2018 dataset to train a model and obtain a reasonable performance on the 2020 dataset for detecting both sites and requests using CNAME cloaking-based tracking.
Ha Dao, Kensuke Fukuda
GLOBECOM1