VLDB 2026 Research / reviewers in the wild / expert
Lu Xian
dblp:276/5983
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-8120-1012ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Collaborative Podcast Creation: Practices, Challenges, and Future Technology Roles in Creative WorkflowsabstractCreative workflows increasingly rely on technological support, yet less is known about how creators navigate the complex collaborative ecosystems and what roles technology might play in supporting these ecosystems. We investigate podcast production as a case to understand collaborative creative practices and identify opportunities for future technology design. Through diary study and interviews with 14 podcast creators, we examine how teams navigate complex stakeholder ecosystems involving distributed teams, episodic guests, multi-tiered audiences, and networks. Our findings reveal three underlying tensions that lead to frictions in current collaborative creation practices: fragmentation vs. coherence in creative coordination, structure vs. spontaneity in live production, and intimacy vs. scale in listener engagement. We outline design opportunities for technology roles to address these tensions: a Story Architect synthesizing fragmented inputs while preserving accessibility, a Live Collaborator providing real-time support during work, and a Feedback Proxy simulating audience feedback at scale. This work contributes a rich empirical foundation for designing technology that reduces social coordination burden while preserving authentic creative relationships in collaborative creation. Ye Yuan 0010, Ruotong Wang 0002, Jan Cao, Lu Xian, Svetlana Yarosh |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | Layered, Overlapping, and Inconsistent: A Large-Scale Analysis of the Multiple Privacy Policies and Controls of U.S. BanksabstractPrivacy policies are often complex. An exception is the two-page standardized notice that U.S. financial institutions must provide under the Gramm-Leach-Bliley Act (GLBA). However, banks now operate websites, mobile apps, and other services that involve complex data sharing practices that require additional privacy notices and do-not-sell opt-outs. We conducted a large-scale analysis of how U.S. banks implement privacy policies and controls in response to GLBA; other federal privacy policy requirements; and the California Consumer Privacy Act (CCPA), a key example for U.S. state privacy laws. We focused on the disclosure and control of a set of especially privacy-invasive practices: third-party data sharing for marketing-related purposes. We collected privacy policies for the 2,067 largest U.S. banks, 45.2% of which provided multiple policies. Across disclosures and controls for the same bank, we identified frequent, concerning inconsistencies---53.8% of banks with multiple privacy policies indicated in GLBA notices that they do not share with third parties but disclosed sharing in other policies. This multiplicity of policies, with the inconsistencies it causes, may create consumer confusion and undermine the transparency goals of the very laws that require them. Our findings call into question whether current policy requirements, such as the GLBA notice, are achieving their intended goals in today's online banking landscape. We discuss potential avenues for reforming and harmonizing privacy policies and control requirements across federal and state laws. Lu Xian, Van Hong Tran, Lauren Lee, Meera Kumar, Florian Schaub |
CCS | 1 |
| 2025 | SoAC and SoACer: A Sector-Based Corpus and LLM-Based Framework for Sectoral Website ClassificationabstractOne approach to understanding the vastness and complexity of the web is to categorize websites into sectors that reflect the specific industries or domains in which they operate. However, existing website classification approaches often struggle to handle the noisy, unstructured, and lengthy nature of web content, and current datasets lack a universal sector classification labeling system specifically designed for the web. To address these issues, we introduce SoAC (Sector of Activity Corpus), a large-scale corpus comprising 195, 495 websites categorized into 10 broad sectors tailored for web content, which serves as the benchmark for evaluating our proposed classification framework, SoACer (Sector of Activity Classifier). Building on this resource, SoACer is a novel end-to-end classification framework that first fetches website information, then incorporates extractive summarization to condense noisy and lengthy content into a concise representation, and finally employs large language model (LLM) embeddings (Llama3-8B) combined with a classification head to achieve accurate sectoral prediction. Through extensive experiments, including ablation studies and detailed error analysis, we demonstrate that SoACer achieves an overall accuracy of 72.6% on our proposed SoAC dataset. Our ablation study confirms that extractive summarization not only reduces computational overhead but also enhances classification performance, while our error analysis reveals meaningful sector overlaps that underscore the need for multi-label and hierarchical classification frameworks. These findings provide a robust foundation for future exploration of advanced classification techniques that better capture the complex nature of modern website content.1 Shahriar Shayesteh, Mukund Srinath, Lee Matheson, Lu Xian, Sinjoy Saha, C. Lee Giles, Shomir Wilson |
DocEng | 4 |
| 2025 | The Informal Labor in Creator Economy: The Making of Key Opinion Consumers From Ordinary Users on XiaohongshuabstractThis paper critically examines flexible content creation conducted by Key Opinion Consumers (KOCs) on a prominent social media and e-commerce platform in China, Xiaohongshu (RED). Drawing on nine-month ethnographic work conducted online, we find that the production of the KOC role on RED is predicated on the interactions and negotiations among multiple stakeholders---content creators, marketers, consumer brands (corporations), and the platform. KOCs are instrumental in RED influencer marketing tactics and amplify the mundane and daily life content popular on the platform. They navigate the dynamics in the triangulated relations with other stakeholders in order to secure economic opportunities for producing advertorial content, and yet, the labor involved in producing such content is deliberately obscured to make it appear as spontaneous, ordinary user posts for the sake of marketing campaigns. Meanwhile, the commercial value of their work is often underestimated and overshadowed in corporate paperwork, platform technological mechanisms, and business models, resulting in and reinforcing inadequate recognition and compensation of KOCs. We illustrate the precarious nature of a form of creativity-driven digital employment through the case of KOCs on the RED platform by demonstrating how this work is made informal. This perspective offers a new lens to understand content creation labor that is indispensable yet unrecognized by the social media industry. We advocate for a contextualized and nuanced examination of how labor is valued and compensated and urge for better protections and working conditions for informal laborers like KOCs. Huiran Yi, Lu Xian |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2025 | User-Centric Textual Descriptions of Privacy-Enhancing Technologies for Ad Tracking and AnalyticsabstractDescribing Privacy Enhancing Technologies (PETs) to the general public is challenging but essential to convey the privacy protections they provide. Existing research has explored the explanation of differential privacy in health contexts. Our study adapts well-performing textual descriptions of local differential privacy from prior work to a new context and broadens the investigation to the descriptions of additional PETs. Specifically, we develop user-centric textual descriptions for popular PETs in ad tracking and analytics, including local differential privacy, federated learning with and without local differential privacy, and Google's Topics. We examine the applicability of previous findings to these expanded contexts, and evaluate the PET descriptions with quantitative and qualitative survey data (n=306). We find that adapting a process- and implications-focused approach to the ad tracking and analytics context achieved similar effects in facilitating user understanding compared to health contexts, and that our descriptions developed with this process+implications approach for the additional, understudied PETs help users understand PETs' processes. We also find that incorporating an implications statement into PET descriptions did not hurt user comprehension but also did not achieve a significant positive effect, which contrasts prior findings in health contexts. We note that the use of technical terms as well as the machine learning aspect of PETs, even without delving into specifics, led to confusion for some respondents. Based on our findings, we offer recommendations and insights for crafting effective user-centric descriptions of privacy-enhancing technologies. Lu Xian, Song Mi Lee-Kan, Jane Im, Florian Schaub |
Proc. Priv. Enhancing Technol. | 1 |
| 2024 | Landscape of Large Language Models in Global English News: Topics, Sentiments, and Spatiotemporal AnalysisabstractGenerative AI has exhibited considerable potential to transform various industries and public life. The role of news media coverage of generative AI is pivotal in shaping public perceptions and judgments about this significant technological innovation. This paper provides in-depth analysis and rich insights into the temporal and spatial distribution of topics, sentiment, and substantive themes within global news coverage focusing on the latest emerging technology—generative AI. We collected a comprehensive dataset of English news articles (January 2018 to November 2023, N = 24,827) through ProQuest databases. For topic modeling, we employed the BERTopic technique and combined it with qualitative coding to identify semantic themes. Subsequently, sentiment analysis was conducted using the RoBERTa-base model. Analysis of temporal patterns in the data reveals notable variability in coverage across key topics—business, corporate technological development, regulation and security, and education—with spikes in articles coinciding with major AI developments and policy discussions. Sentiment analysis shows a predominantly neutral to positive media stance, with the business-related articles exhibiting more positive sentiment, while regulation and security articles receive a reserved, neutral to negative sentiment. Our study offers a valuable framework to investigate global news discourse and evaluate news attitudes and themes related to emerging technologies. Lu Xian, Lingyao Li, Ben Zefeng Zhang, Libby Hemphill |
ICWSM | 1 |
| 2023 | Conceptualizing Algorithmic StigmatizationabstractAlgorithmic systems have infiltrated many aspects of our society, mundane to high-stakes, and can lead to algorithmic harms known as representational and allocative. In this paper, we consider what stigma theory illuminates about mechanisms leading to algorithmic harms in algorithmic assemblages. We apply the four stigma elements (i.e., labeling, stereotyping, separation, status loss/discrimination) outlined in sociological stigma theories to algorithmic assemblages in two contexts : 1) "risk prediction" algorithms in higher education, and 2) suicidal expression and ideation detection on social media. We contribute the novel theoretical conceptualization of algorithmic stigmatization as a sociotechnical mechanism that leads to a unique kind of algorithmic harm: algorithmic stigma. Theorizing algorithmic stigmatization aids in identifying theoretically-driven points of intervention to mitigate and/or repair algorithmic stigma. While prior theorizations reveal how stigma governs socially and spatially, this work illustrates how stigma governs sociotechnically. Nazanin Andalibi, Cassidy Pyle, Kristen Barta, Lu Xian, Abigail Z. Jacobs, Mark S. Ackerman |
CHI | 4 |