VLDB 2026 Research / reviewers in the wild / expert
Xiaoyuan Wu
dblp:01/3193
· DBLP profile ↗
25ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorSecurity and privacy · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive ScenariosabstractLarge language models (LLMs) are rapidly being adopted for tasks like drafting emails, summarizing meetings, and answering health questions.In these settings, users may need to share private information (e.g., contact details, health records).To evaluate LLMs' ability to identify and redact such information, prior work introduced real-life, scenario-based benchmarks (e.g., ConfAIde, PrivacyLens) and found that LLMs can leak private information in complex scenarios.However, these evaluations relied on proxy LLMs to judge the helpfulness and privacy-preservation quality of LLM responses, rather than directly measuring users' perceptions.To understand how users perceive the helpfulness and privacy-preservation quality of LLM responses to privacy-sensitive scenarios, we conducted a user study (n = 94) using 90 PrivacyLens scenarios.We found that users had low agreement with each other when evaluating identical LLM responses.In contrast, five proxy LLMs reached high agreement, yet each proxy LLM had low correlation with users' evaluations.These results indicate that proxy LLMs cannot accurately estimate users' wide range of perceptions of utility and privacy in privacy-sensitive scenarios.We discuss the need for more user-centered studies to measure LLMs' ability to help users while preserving privacy, and for improving alignment between LLMs and users in estimating perceived privacy and utility. Xiaoyuan Wu, Roshni Kaushik, Lujo Bauer, Koichi Onoue |
ACL (1) | 1 |
| 2026 | Passing Down Passwords: How Older Adults Approach Postmortem Account Access and Digital Estate PlanningabstractTraditional estate planning practices enable people to provide their heirs access to the assets left behind but are often insufficient for the transfer and management of online accounts. To understand how estate planning practices could be improved, we conducted 21 semi-structured interviews with older adults in the United States that explored their practices, concerns, and needs regarding postmortem online account access and management. We encountered few formalized digital estate planning practices; many participants use their credential management practices—primarily pen-and-paper—to provide postmortem account access. How participants envision account transfer is motivated by trust in their current practices and in their heirs, while concerns regarding technology hinder adoption of new methods. Participants consistently prioritize accounts with financial assets, and expectations surrounding postmortem account management vary based on individual circumstances, with the common goal of reducing burdens on executors and heirs. Our results suggest the need for developing technical standardization and expert guidance for digital estate planning. Jenny Tang, Xiaoyuan Wu, Lujo Bauer, Nicolas Christin, Lorrie Faith Cranor |
CHI | 2 |
| 2026 | An Ultra-Low-Power, High-Output-Power and High-Sensitivity Bioelectronic Transceiver
Xiaoyuan Wu, Kangjie Zhao, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang |
ISCAS | 2 |
| 2025 | Measuring Risks to Users' Health Privacy Posed by Third-Party Web Tracking and Targeted Advertising
Eric Zeng 0001, Xiaoyuan Wu, Emily N. Ertmann, Lily Huang, Danielle F. Johnson, Anusha T. Mehendale, Brandon T. Tang, Karolina Zhukoff, Michael Adjei-Poku, Lujo Bauer, Ari B. Friedman, Matthew S. McCoy |
CHI | 2 |
| 2025 | Estimating LLM Consistency: A User Baseline vs Surrogate MetricsabstractLarge language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text.Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses-the model's confidence in the response or likelihood of generating a similar response when resampled.In previous work, measuring LLM response consistency often relied on calculating the probability of a response appearing within a pool of resampled responses, analyzing internal states, or evaluating logits of resopnses.However, it was not clear how well these approaches approximated users' perceptions of consistency of LLM responses.To find out, we performed a user study (n = 2, 976) demonstrating that current methods for measuring LLM response consistency typically do not align well with humans' perceptions of LLM consistency.We propose a logit-based ensemble method for estimating LLM consistency and show that our method matches the performance of the bestperforming existing metric in estimating human ratings of LLM consistency.Our results suggest that methods for estimating LLM consistency without human evaluation are sufficiently imperfect to warrant broader use of evaluation with human input; this would avoid misjudging the adequacy of models because of the imperfections of automated consistency metrics. Xiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo Bauer |
EMNLP | 1 |
| 2025 | A Reinforcement Learning-Based Approach for Storage Assignment in Coded BlockchainabstractCoded blockchain leverages error correction codes to create coded fragments that are then stored in a distributed manner by nodes. It significantly reduces the storage requirement of conventional blockchain. A key limitation, however, is that as nodes store arbitrary coded fragments, which leads to high transmission overhead or decoding failure. To this end, this paper introduces a novel reinforcement learning approach to assign coded fragments to nodes. Specifically, nodes learn to store coded fragments based on feedback from other nodes, which considers their block decoding process. To study its efficacy, this paper compares the proposed approach with existing centralized and distributed coded fragments assignment solutions. The simulation results show that our approach outperforms existing solutions by 6 % in terms of loss function and by 20 % in transmission overhead. Changlin Yang, Weijian Xia, Xiaoyuan Wu, Xiangping Chen |
ICPADS | 3 |
| 2025 | Transparency or Information Overload? Evaluating Users' Comprehension and Perceptions of the iOS App Privacy Report
Xiaoyuan Wu, Lydia Hu, Eric Zeng 0001, Hana Habib, Lujo Bauer |
NDSS | 1 |
| 2024 | The Sword of Damocles: Upgradeable Smart Contract in EthereumabstractAlthough smart contracts are immutable once they are deployed, the reality is that they need upgrades to fix bugs or add new features. Nowadays, there are a few upgrade methods in Ethereum, some of which can change the contract without changing the contract address that users interact with. This upgrade way increases potential danger and results in users' distrust, because it may secretly change the function of the contract and cause users financial loss. We examine two of these upgrade methods, i.e., proxy pattern and metamorphic contract. For the proxy pattern, we propose a bytecode-based method for detecting these upgradeable contracts, which achieves a 99.37% F1-score. We use the bytecode-based method to detect the contracts in the first 12 million blocks of Ethereum and find 126,500 upgradeable contracts. For the metamorphic contracts, we employ an Ethereum replay tool to replay the transactions and find the metamorphic contracts according to the SELFDESTRUCT and CREATE2 instructions. We find that 64.3% of the contracts upgraded using this way are malicious MEV bots. Finally, we summarize the reasons for smart contract upgrades and make development recommendations. Yuan Huang 0002, Xiaoyuan Wu, Quanqi Wang, Ziang Qian, Xiangping Chen, Mingdong Tang, Zibin Zheng |
ICPC | 2 |
| 2023 | Poster: Longitudinal Measurement of the Adoption Dynamics in Apple's Privacy Label EcosystemabstractThis work reports on a large scale, longitudinal analysis of the adoption dynamics of privacy labels in the iOS App Store, measuring this first-of-its kind ecosystem as it reaches maturity over two and a half years after launching in December 2020. The motivation is to shed light on the factors affecting the shifts in privacy labels and provide insights into how and when an app's label changes. By collecting nearly weekly snapshots of over 1.6 million apps for over a year, we analyze the dynamics of privacy label adoption and the accuracy of reported labels. Our analysis of 74.5% of apps having labels after two years provides important context into this mature ecosystem where labels are becoming the standard. However, we find compelling evidence that labels may not fully capture behavior, as 28.9% of apps indicate no data collection and distributions differ between voluntary versus mandatory adoptions. Once set, labels rarely change but additions reflect more data collection. In addition to our measurement, we also plan to release a new (and growing) data set that can be used by future researchers. David G. Balash, Mir Masood Ali, Monica Kodwani, Xiaoyuan Wu, Chris Kanich, Adam J. Aviv |
CCS | 4 |
| 2023 | A 88%-Peak-Efficiency 10-mV-Voltage-Ripple Dual-Mode Switched-Capacitor DC-DC Converter for Ultra-Low-Power Battery ManagementabstractThis paper proposes a high-efficiency low-ripple dual-mode switched-capacitor (SC) DC-DC converter for low-power IoT and wearable device applications. A hybrid self-biased current scheme (HSBC) is developed to achieve low output voltage ripple and fast response. Two supply voltage domains of HSBC and clock drive controller circuit are introduced to reduce the power loss of the control circuit. An on-chip ultra-low-power bias circuit is also designed to minimize the power loss of the bias generator. The proposed DC-DC converter is implemented in a 40 nm CMOS process. Post-layout simulation results show that the converter realizes 1.6-1.8 V to 0.4 V conversion. The peak efficiency is up to 88% at$5\ \mu\mathrm{A}$, and the voltage ripple is less than 10 mV over a load range of 10 nA-$10\ \mu\mathrm{A}$. The response time of the converter is less than$15\ \mu\mathrm{s}$. Xiaoyuan Wu, Leilei Huang, Boxiao Liu, Chunqi Shi, Jinghong Chen, Runxi Zhang |
ISCAS | 1 |
| 2023 | An effective electrocardiogram segments denoising method combined with ensemble empirical mode decomposition, empirical mode decomposition, and wavelet packetabstractAbstract Electrocardiogram (ECG) is the most extensively applied diagnostic approach for heart diseases. However, an ECG signal is a weak bioelectrical signal and is easily disturbed by baseline wander, powerline interference, and muscle artefacts, which make detection of heart diseases more difficult. Therefore, it is very important to denoise the contaminated ECG signal in practical application. In this article, an effective ECG segments denoising method combining the ensemble empirical mode decomposition (EEMD), empirical mode decomposition (EMD), and wavelet packet (WP) is designed. The ECG signal is decomposed using the EEMD for the first time, and then the highest frequency component is decomposed by the EMD for the second time, and the high frequency components obtained from the second time are decomposed and reconstructed by the WP for the third time. Finally, the processed signal components are fused to obtain the denoised ECG signal. Furthermore, the signal‐to‐noise ratio (SNR), mean square error (MSE), root mean square error (RMSE), and normalised cross correlation coefficient (R) are used to evaluate the noise reduction algorithm. The mean SNR, MSE, RMSE, and R are 5.7427, 0.0071, 0.0551, and 0.9050 in the China Physiological Signal Challenge 2018 dataset. Compared with others denoising methods, the experimental results not only exhibit that the SNR of the ECG signal is effectively improved, but also show that the details of the ECG signal are fully retained, laying a solid foundation for the automatic detection of ECG segments. Yaru Yue, Chengdong Chen, Xiaoyuan Wu, Xiaoguang Zhou |
IET Signal Process. | 3 |
| 2022 | User Perceptions of Five-Word PasswordsabstractHuman-chosen passwords are often short, selected non-uniformly, and thus, susceptible to automated guessing attacks. To help users to select more secure but memorable passwords, experts have recommended the use of passphrases of multiple words or phrases. In this paper, we explore a strategy for passphrase selection, so-called five-word passwords, where users are assigned five random words for a passphrase. Such a password composition policy was recently adopted at Georgetown University in December 2020. Through a two-part online survey (n = 150 and n = 116), participants selected a five-word password under different conditions. We find that computer-generated five-word passwords are more diverse and likely more secure than five-word passwords users select themselves. While all cases of five-word passwords are likely more secure than a human-generated, traditional password, participants expressed misconceptions regarding the security of five-word passwords (and passwords generally). Five-word passwords also appear to negatively impact usability, only 39.7 % of participants successfully recalled their password after two weeks. While five-word passwords offer improvements for security, more outreach is needed to explain their security benefits and reduce usability burdens. Xiaoyuan Wu, Collins W. Munyendo, Eddie Cosic, Genevieve A. Flynn, Olivia Legault, Adam J. Aviv |
ACSAC | 1 |
| 2022 | Security and Privacy Perceptions of Third-Party Application Access for Google Accounts
David G. Balash, Xiaoyuan Wu, Miles Grant, Irwin Reyes, Adam J. Aviv |
USENIX Security Symposium | 2 |
| 2022 | Assessing integrated coal production and land reconstruction systems under extreme temperatures
Xiaoyuan Wu, Yung-Ho Chiu, Qinghua Pang |
Expert Syst. Appl. | 2 |
| 2019 | A knowledge discovery and reuse method for time estimation in ship block manufacturing planning using DEA
Miaomiao Sun, Duanfeng Han, Xuezhang Mao, Xiaoyuan Wu |
Adv. Eng. Informatics | 6 |
| 2009 | Predicting the conversion probability for items on C2C ecommerce sitesabstractOnline ecommerce has been booming for a decade. For instance, as the largest online C2C marketplace (eBay), millions of new items are listed daily. Due to the overwhelming number of items, the process of finding the right items to buy is sometimes daunting. In order to address this problem, this paper describes the idea of predicting the probability that a newly listed item will be sold successfully. And adjust the item exposure chances proportional according to their conversion possibility. Hence, by ranking higher items that users are likely to buy, the chance that users make the purchases could be increased as well as their user satisfaction. For catalog products that have been listed repeatedly, this probability can be measured empirically. However, on C2C sites like eBay, lots of items are not product-based. They are unique, and from different sellers. Therefore, in order to predict whether a new listing will be sold, we collect a large scale item set as the training data, and a set of features were used to model the average buyer shopping decision on C2C sites. Experimental results verified our system's feasibility and effectiveness. Xiaoyuan Wu, Alvaro Bolivar |
CIKM | 1 |
| 2009 | Substitutes or complements: another step forward in recommendationsabstractIn this paper, we introduce the method tagging substitute-complement attributes on miscellaneous recommending relations, and elaborate how this step contributes to electronic merchandising. Jiaqian Zheng, Xiaoyuan Wu, Junyu Niu, Alvaro Bolivar |
EC | 2 |
| 2009 | Rare item detection in e-commerce siteabstractAs the largest online marketplace in the world, eBay has a huge inventory where there are plenty of great rare items with potentially large, even rapturous buyers. These items are obscured in long tail of eBay item listing and hard to find through existing searching or browsing methods. It is observed that there are great rarity demands from users according to eBay query log. To keep up with the demands, the paper proposes a method to automatically detect rare items in eBay online listing. A large set of features relevant to the task are investigated to filter items and further measure item rareness. The experiments on the most rarity-demand-intensitive domains show that the method may effectively detect rare items (>90% precision). Xiaoyuan Wu, Alvaro Bolivar |
WWW | 2 |
| 2008 | Keyword extraction for contextual advertisementabstractAs the largest online marketplace, eBay strives to promote its inventory throughout the Web via different types of online advertisement. Contextually relevant links to eBay assets on third party sites is one example of such advertisement avenues. Keyword extraction is the task at the core of any contextual advertisement system. In this paper, we explore a machine learning approach to this problem. The proposed solution uses linear and logistic regression models learnt from human labeled data, combined with document, text and eBay specific features. In addition, we propose a solution to identify the prevalent category of eBay items in order to solve the problem of keyword ambiguity. Xiaoyuan Wu, Alvaro Bolivar |
WWW | 1 |
| 2007 | Optimizing web search using social annotationsabstractThis paper explores the use of social annotations to improve web search. Nowadays, many services, e.g. del.icio.us, have been developed for web users to organize and share their favorite web pages on line by using social annotations. We observe that the social annotations can benefit web search in two aspects: 1) the annotations are usually good summaries of corresponding web pages; 2) the count of annotations indicates the popularity of web pages. Two novel algorithms are proposed to incorporate the above information into page ranking: 1) SocialSimRank (SSR) calculates the similarity between social annotations and web queries; 2) SocialPageRank (SPR) captures the popularity of web pages. Preliminary experimental results show that SSR can find the latent semantic association between queries and annotations, while SPR successfully measures the quality (popularity) of a web page from the web users ’ perspective. We further evaluate the proposed methods empirically with 50 manually constructed queries and 3000 auto-generated queries on a dataset crawled from del.icio.us. Experiments show that both SSR and SPR benefit web search significantly. Shenghua Bao, Gui-Rong Xue, Xiaoyuan Wu, Yong Yu 0001, Ben Fei, Zhong Su |
WWW | 3 |
| 2006 | Salient Phrases-based Clustering and Ranking in Chinese Bulletin Board System
Xiaoyuan Wu, Shen Huang, Yong Yu 0001 |
SEKE | 1 |
| 2006 | Automatic Identification of Chinese Weblogger's Interests Based on Text ClassificationabstractChinese Weblogs have been expanded in an incredible speed in recent years. There is plentiful personal information in Weblogs. In this paper, we propose a text classification based approach to automatically identify the interests of a Weblogger. To solve the problems arising out of class Weblog documents, the technique of heterogeneous classifiers combination is used. We also use hierarchical classification technique to identify much specific interests. Experiments show that our interest identification approach has a high accuracy and, for most Webloggers in our experiments, their interests implied in the contents of blogs could be well identified by using this approach Xiaochuan Ni 0001, Xiaoyuan Wu, Yong Yu 0001 |
Web Intelligence | 2 |
| 2005 | Ontology based semantic conflicts resolution in collaborative editing of design documents
Xiaoyuan Wu, Jiang-Ming Yang |
Adv. Eng. Informatics | 3 |
| 2002 | An Improved Concurrency Control Algorithm of Cooperative Design Document Based on MarkingabstractHigh reactivity cannot be achieved without each object being replicated on every site, when assuming a nonnegligible network latency in a cooperative design document system. It is an austerity challenge for concurrency control, and causes some new problems, such as intention-violation. The author designed a concurrency control method based on marking, which accords with integrity consistency model. However, it is inefficient at judging causality and concurrent operation. This paper brings forward an improved algorithm to solve this problem. Xiaoyuan Wu, Jiaomao Liu |
CSCWD | 1 |
| 2001 | Extracting WEB Table Information in Cooperative Learning Activities Based on Abstract Semantic ModelabstractA great deal of Web table information exists in cooperative learning activities. The paper presents a new method that extracts information from tables of Web documents. Using a tabled abstract semantic model to describe complicated tables and understand tables from the point of view of semantics, the method reduces the dependence for the design difference of table constructions in the extraction process. At the same time, it utilizes the characteristics of HTML and the techniques of natural language processing to design some heuristic rules, and thus aids the identification of table items. On the above basis, we design a prototype, "EXTable", and then gain a better result according to experimentation. Guowen Wu, Xiaoyuan Wu, Baile Shi |
CSCWD | 3 |