VLDB 2026 Research / reviewers in the wild / expert
Hao Cui 0004
dblp:98/75-4
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-7574-2004ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding Privacy Norms through Web FormsabstractWeb forms are one of the primary ways to collect personal information online, yet they are relatively under-studied. Unlike web tracking, data collection through web forms is explicit and contextualized. Users (i) are asked to input specific personal information types, and (ii) know the specific context (i.e., on which website and for what purpose). For web forms to be trusted by users, they must meet the common sense standards of appropriate data collection practices within a particular context (i.e., privacy norms). In this paper, we extract the privacy norms embedded within web forms through a measurement study. First, we build a specialized crawler to discover web forms on websites. We run it on 11,500 popular websites, and we create a dataset of 293K web forms. Second, to process data of this scale, we develop a cost-efficient way to annotate web forms with form types and personal information types, using text classifiers trained with assistance of large language models (LLMs). Third, by analyzing the annotated dataset, we reveal common patterns of data collection practices. We find that (i) these patterns are explained by functional necessities and legal obligations, thus reflecting privacy norms, and that (ii) deviations from the observed norms often signal unnecessary data collection. In addition, we analyze the privacy policies that accompany web forms. We show that, despite their wide adoption and use, there is a disconnect between privacy policy disclosures and the observed privacy norms. Hao Cui 0004, Rahmadi Trimananda, Athina Markopoulou |
Proc. Priv. Enhancing Technol. | 1 |
| 2025 | BehaVR: User Identification Based on VR Sensor DataabstractVirtual reality (VR) platforms enable a wide range of applications, however, pose unique privacy risks. In particular, VR devices are equipped with a rich set of sensors that collect personal and sensitive information (e.g., body motion, eye gaze, hand joints, and facial expression). The data from these newly available sensors can be used to uniquely identify a user, even in the absence of explicit identifiers. In this paper, we seek to understand the extent to which a user can be identified based solely on VR sensor data, within and across real-world apps from diverse genres. We consider adversaries with capabilities that range from observing APIs available within a single app (app adversary) to observing all or selected sensor measurements across multiple apps on the VR device (device adversary). To that end, we introduce BehaVR, a framework for collecting and analyzing data from all sensor groups collected by multiple apps running on a VR device. We use BehaVR to collect data from real users that interact with 20 popular real-world apps. We use that data to build machine learning models for user identification within and across apps, with features extracted from available sensor data. We show that these models can identify users with an accuracy of up to 100%, and we reveal the most important features and sensor groups, depending on the functionality of the app and the adversary. To the best of our knowledge, BehaVR is the first to analyze user identification in VR comprehensively, i.e., considering all sensor measurements available on consumer VR devices, collected by multiple real-world, as opposed to custom-made, apps. Ismat Jarin, Rahmadi Trimananda, Hao Cui 0004, Salma Hosni Emam Mohamed Elmalaki, Athina Markopoulou |
Proc. Priv. Enhancing Technol. | 4 |
| 2023 | Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker EcosystemabstractSmart speakers collect voice commands, which can be used to infer sensitive information about users. Given the potential for privacy harms, there is a need for greater transparency and control over the data collected, used, and shared by smart speaker platforms as well as third party skills supported on them. To bridge this gap, we build a framework to measure data collection, usage, and sharing by the smart speaker platforms. We apply our framework to the Amazon smart speaker ecosystem. Our results show that Amazon and third parties, including advertising and tracking services that are unique to the smart speaker ecosystem, collect smart speaker interaction data. We also find that Amazon processes smart speaker interaction data to infer user interests and uses those inferences to serve targeted ads to users. Smart speaker interaction also leads to ad targeting and as much as 30X higher bids in ad auctions, from third party advertisers. Finally, we find that Amazon's and third party skills' data practices are often not clearly disclosed in their policy documents. Umar Iqbal 0002, Pouneh Nikkhah Bahrami, Rahmadi Trimananda, Hao Cui 0004, Alexander Gamero-Garrido, Daniel J. Dubois, David R. Choffnes, Athina Markopoulou, Franziska Roesner, Zubair Shafiq |
IMC | 4 |
| 2021 | CDAE: Color decomposition-based adversarial examples for screen devices
Huanyu Bian, Hao Cui 0004, Kunlin Liu, Hang Zhou 0007, Dongdong Chen 0001, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
Inf. Sci. | 2 |
| 2020 | Model Watermarking for Image Processing NetworksabstractDeep learning has achieved tremendous success in numerous industrial applications. As training a good model often needs massive high-quality data and computation resources, the learned models often have significant business values. However, these valuable deep models are exposed to a huge risk of infringements. For example, if the attacker has the full information of one target model including the network structure and weights, the model can be easily finetuned on new datasets. Even if the attacker can only access the output of the target model, he/she can still train another similar surrogate model by generating a large scale of input-output training pairs. How to protect the intellectual property of deep models is a very important but seriously under-researched problem. There are a few recent attempts at classification network protection only.In this paper, we propose the first model watermarking framework for protecting image processing models. To achieve this goal, we leverage the spatial invisible watermarking mechanism. Specifically, given a black-box target model, a unified and invisible watermark is hidden into its outputs, which can be regarded as a special task-agnostic barrier. In this way, when the attacker trains one surrogate model by using the input-output pairs of the target model, the hidden watermark will be learned and extracted afterward. To enable watermarks from binary bits to high-resolution images, both traditional and deep spatial invisible watermarking mechanism are considered. Experiments demonstrate the robustness of the proposed watermarking mechanism, which can resist surrogate models learned with different network structures and objective functions. Besides deep models, the proposed method is also easy to be extended to protect data and traditional image processing algorithms. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Han Fang 0004, Weiming Zhang 0001, Wenbo Zhou 0004, Hao Cui 0004, Nenghai Yu |
AAAI | 7 |
| 2020 | A Camera Shooting Resilient Watermarking Scheme for Underpainting DocumentsabstractThis paper designs a novel underpainting based camera shooting resilient (CSR) document watermarking algorithm for dealing with the leak source tracking problem. By applying such algorithm, we can extract the authentication watermark information from the candid photographs. The watermarked underpainting contains three significant properties. 1) Inconspicuousness. The watermarked underpainting is inconspicuous and it will not easily be maliciously attacked. 2) Robustness. We propose DCT-based watermark embedding algorithm and distortion compensation based extracting algorithm, which make the watermark robust to camera shooting process. 3) Autocorrelation. We design the flip-based method to arrange the watermarked underpainting. So that a complete watermark region can be accurately located even if part of the document is recorded. Compared with previous watermarking algorithms, the proposed scheme guaranteed content independent embedding as well as the robustness to the camera shooting process. Besides, the proposed scheme satisfies the accuracy of extraction even when the captured document is incomplete. Han Fang 0004, Weiming Zhang 0001, Zehua Ma, Hang Zhou 0007, Hao Cui 0004, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | UnseenCode: Invisible On-screen Barcode with Image-based ExtractionabstractScreen-camera communication techniques achieve one-way communication with widely-used screen and camera devices. Popular barcode methods use visible spatial patterns to represent data, which has been considered obtrusive to human observers. Recent works borrow ideas from visual light communication (VLC), and use inter-frame pixel change to modulate data. To recover pixel change, the receiver has to record and process video frames. Such video-based extraction has high hardware requirements and lacks reliability. Inspired by VLC-based methods, we propose UnseenCode, an invisible on-screen barcode scheme. It leverages inter-frame embedment from VLC-based methods to embed invisible barcodes into arbitrary on-screen contents. Unlike VLC-based methods, UnseenCode does not require video-based extraction. We propose an image-based extraction method based on cross-component correlation of color images. Any off-the-shelf smartphones with camera capability can be used to read UnseenCode by capturing on-screen contents. We propose the full implementation of UnseenCode for evaluation. Experimental results show that UnseenCode decoding algorithm is reliable and applicable under various screen and camera settings. UnseenCode provides up to 2.5 kbit capacity with less than 5% error rate. Hao Cui 0004, Huanyu Bian, Weiming Zhang 0001, Nenghai Yu |
INFOCOM | 1 |
| 2019 | Screen-Shooting Resilient WatermarkingabstractThis paper proposes a novel screen-shooting resilient watermarking scheme, which means that if the watermarked image is displayed on the screen and the screen information is captured by the camera, we can still extract the watermark message from the captured photo. To realize such demands, we analyzed the special distortions caused by the screen-shooting process, including lens distortion, light source distortion, and moiré distortion. To resist the geometric deformation caused by lens distortion, we proposed an intensity-based scale-invariant feature transform (I-SIFT) algorithm which can accurately locate the embedding regions. As for the loss of image details caused by light source distortion and moiré distortion, we put forward a small-size template algorithm to repeatedly embed the watermark into different regions, so that at least one complete information region can survive from distortions. At the extraction side, we designed a cross-validation-based extraction algorithm to cope with repeated embedding. The validity and correctness of the extraction method are verified by hypothesis testing. Furthermore, to boost the extraction speed, we proposed a SIFT feature editing algorithm to enhance the intensity of the keypoints, based on which, the extraction accuracy and extraction speed can be greatly improved. The experimental results show that the proposed watermarking scheme achieves high robustness for screen-shooting process. Compared with the previous schemes, our algorithm provides significant improvement in robustness for screen-shooting process and extraction efficiency. Han Fang 0004, Weiming Zhang 0001, Hang Zhou 0007, Hao Cui 0004, Nenghai Yu |
IEEE Trans. Inf. Forensics Secur. | 4 |