VLDB 2026 Research / reviewers in the wild / expert
Chenyan Jia
dblp:278/8322
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-8407-9224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | My Body, Their Business: User Perspectives on Commercial Data Practices in FemTech mHealth AppsabstractFemTech, including apps for fertility, menstruation, and menopause, increasingly shapes how users manage intimate aspects of their health. Yet these apps are often built on opaque commercial models, raising ethical concerns about consent, privacy, and misuse of sensitive health data. While prior work has documented these risks, less is known about how users perceive and negotiate commercial data practices in FemTech apps. We conducted an online survey with 187 participants, combining factorial vignettes with provotypes— interface prototypes designed to provoke reflection— to examine user boundaries and discomforts around FemTech data collection and commercial use. Participants drew sharp distinctions across data types, resisting peripheral data collection and pervasive tracking. Commercial practices were often judged conditionally: tolerated only when functionally relevant. Notably, our provotypes, even under exaggerated transparency, elicited more forgiving responses to commercial practices compared to brief text descriptions in the vignettes. We discuss implications for designing transparent, accountable, and user-aligned FemTech. Ghada Alsebayel, Ximena Lainfiesta, Ayesha Fatima, Giovanni Maria Troiano, Chenyan Jia, Casper Harteveld |
CHI | 5 |
| 2026 | Beyond Accuracy: Experts See AI Fact-Checks as Accurate but Less UsefulabstractAs misinformation proliferates online, large language models (LLMs) have been proposed as a promising tool to accelerate fact-checking workflows. While LLMs demonstrate strong performance in tasks such as text annotation, their capabilities in generating fact-checking reports remain uncertain. To investigate how media experts evaluate LLM-generated fact-checking reports, we conducted a 2 (Source: human vs. LLM) X 2 (Disclosure of Source: yes or no) between-subjects online experiment with media professionals (N=274). Our analyses reveal that experts perceive LLM-generated reports as significantly less useful than human-written reports; and such differences become larger when participants are not aware of the source. However, LLM-generated fact-checking reports were rated as accurate and logical as human-authored ones. Party affiliation plays a role in predicting perceived logicalness. Our findings advance the understanding of experts’ evaluation of LLM-generated content within the context of misinformation, which provides important theoretical contributions to HCI and communication theories as well as practical implications for the field. Chenyan Jia, Apoorva Gondimalla, Angie Zhang, David Joseph Mullings, Alexander Boltz, Min Kyung Lee |
CHI | 1 |
| 2024 | MPRNet: Multi-scale Pointwise Regression Network for Crowd Counting and Localization
Chenyan Jia, Zhitao Cheng, Yanlin Leng, Yong Tang 0002 |
ICIC (12) | 1 |
| 2024 | Training Socially Aligned Language Models on Simulated Social InteractionsabstractThe goal of social alignment for AI systems is to make sure these models can conduct themselves appropriately following social values. Unlike humans who establish a consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly recite the corpus in social isolation, which causes poor generalization in unfamiliar cases and the lack of robustness under adversarial attacks. In this work, we introduce a new training paradigm that enables LMs to learn from simulated social interactions. Compared with existing methods, our method is much more scalable and efficient, and shows superior performance in alignment benchmarks and human evaluation. Ruibo Liu, Chenyan Jia, Diyi Yang, Soroush Vosoughi |
ICLR | 3 |
| 2024 | Embedding Democratic Values into Social Media AIs via Societal Objective FunctionsabstractMounting evidence indicates that the artificial intelligence (AI) systems that rank our social media feeds bear nontrivial responsibility for amplifying partisan animosity: negative thoughts, feelings, and behaviors toward political out-groups. Can we design these AIs to consider democratic values such as mitigating partisan animosity as part of their objective functions? We introduce a method for translating established, vetted social scientific constructs into AI objective functions, which we term societal objective functions, and demonstrate the method with application to the political science construct of anti-democratic attitudes. Traditionally, we have lacked observable outcomes to use to train such models-however, the social sciences have developed survey instruments and qualitative codebooks for these constructs, and their precision facilitates translation into detailed prompts for large language models. We apply this method to create a democratic attitude model that estimates the extent to which a social media post promotes anti-democratic attitudes, and test this democratic attitude model across three studies. In Study 1, we first test the attitudinal and behavioral effectiveness of the intervention among US partisans (N=1,380) by manually annotating (alpha=.895) social media posts with anti-democratic attitude scores and testing several feed ranking conditions based on these scores. Removal (d=.20) and downranking feeds (d=.25) reduced participants' partisan animosity without compromising their experience and engagement. In Study 2, we scale up the manual labels by creating the democratic attitude model, finding strong agreement with manual labels (rho=.75). Finally, in Study 3, we replicate Study 1 using the democratic attitude model instead of manual labels to test its attitudinal and behavioral impact (N=558), and again find that the feed downranking using the societal objective function reduced partisan animosity (d=.25). This method presents a novel strategy to draw on social science theory and methods to mitigate societal harms in social media AIs. Chenyan Jia, Michelle S. Lam, Minh Chau Mai, Jeffrey T. Hancock, Michael S. Bernstein |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | Non-Parallel Text Style Transfer with Self-Parallel Supervision
Ruibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu, Soroush Vosoughi |
ICLR | 3 |
| 2022 | Second Thoughts are Best: Learning to Re-Align With Human Values from Text EditsabstractWe present Second Thoughts, a new learning paradigm that enables language models (LMs) to re-align with human values. By modeling the chain-of-edits between value-unaligned and value-aligned text, with LM fine-tuning and additional refinement through reinforcement learning, Second Thoughts not only achieves superior performance in three value alignment benchmark datasets but also shows strong human-value transfer learning ability in few-shot scenarios. The generated editing steps also offer better interpretability and ease for interactive error correction. Extensive human evaluations further confirm its effectiveness. Ruibo Liu, Chenyan Jia, Ziyu Zhuang, Tony X. Liu, Soroush Vosoughi |
NeurIPS | 2 |
| 2022 | Quantifying and alleviating political bias in language models
Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Soroush Vosoughi |
Artif. Intell. | 2 |
| 2022 | Trust in COVID-19 public health informationabstractAbstract Understanding the factors that influence trust in public health information is critical for designing successful public health campaigns during pandemics such as COVID‐19. We present findings from a cross‐sectional survey of 454 US adults—243 older (65+) and 211 younger (18–64) adults—who responded to questionnaires on human values, trust in COVID‐19 information sources, attention to information quality, self‐efficacy, and factual knowledge about COVID‐19. Path analysis showed that trust in direct personal contacts (B = 0.071, p = .04) and attention to information quality (B = 0.251, p < .001) were positively related to self‐efficacy for coping with COVID‐19. The human value of self‐transcendence, which emphasizes valuing others as equals and being concerned with their welfare, had significant positive indirect effects on self‐efficacy in coping with COVID‐19 (mediated by attention to information quality; effect = 0.049, 95% CI 0.001–0.104) and factual knowledge about COVID‐19 (also mediated by attention to information quality; effect = 0.037, 95% CI 0.003–0.089). Our path model offers guidance for fine‐tuning strategies for effective public health messaging and serves as a basis for further research to better understand the societal impact of COVID‐19 and other public health crises. Nitin Verma, Kenneth R. Fleischmann, Bo Xie 0001, Min Kyung Lee, Katherine Rich, Kristina Shiroma, Chenyan Jia, Tara Zimmerman |
J. Assoc. Inf. Sci. Technol. | 8 |
| 2022 | Understanding Effects of Algorithmic vs. Community Label on Perceived Accuracy of Hyper-partisan MisinformationabstractHyper-partisan misinformation has become a major public concern. In order to examine what type of misinformation label can mitigate hyper-partisan misinformation sharing on social media, we conducted a 4 (label type: algorithm, community, third-party fact-checker, and no label) X 2 (post ideology: liberal vs. conservative) between-subjects online experiment (N = 1,677) in the context of COVID-19 health information. The results suggest that for liberal users, all labels reduced the perceived accuracy and believability of fake posts regardless of the posts' ideology. In contrast, for conservative users, the efficacy of the labels depended on whether the posts were ideologically consistent: algorithmic labels were more effective in reducing the perceived accuracy and believability of fake conservative posts compared to community labels, whereas all labels were effective in reducing their belief in liberal posts. Our results shed light on the differing effects of various misinformation labels dependent on people's political ideology. Chenyan Jia, Alexander Boltz, Angie Zhang, Anqing Chen, Min Kyung Lee |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Mitigating Political Bias in Language Models through Reinforced CalibrationabstractCurrent large-scale language models can be politically biased as a result of the data they are trained on, potentially causing serious problems when they are deployed in real-world settings. In this paper, we describe metrics for measuring political bias in GPT-2 generation and propose a reinforcement learning (RL) framework for mitigating political biases in generated text. By using rewards from word embeddings or a classifier, our RL framework guides debiased generation without having access to the training data or requiring the model to be retrained. In empirical experiments on three attributes sensitive to political bias (gender, location, and topic), our methods reduced bias according to both our metrics and human evaluation, while maintaining readability and semantic coherence. Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Soroush Vosoughi |
AAAI | 2 |
| 2021 | Political Depolarization of News Articles Using Attribute-Aware Word Embeddings
Ruibo Liu, Chenyan Jia, Soroush Vosoughi |
ICWSM | 3 |
| 2021 | A Transformer-based Framework for Neutralizing and Reversing the Political Polarity of News ArticlesabstractPeople often prefer to consume news with similar political predispositions and access like-minded news articles, which aggravates polarized clusters known as "echo chamber". To mitigate this phenomenon, we propose a computer-aided solution to help combat extreme political polarization. Specifically, we present a framework for reversing or neutralizing the political polarity of news headlines and articles. The framework leverages the attention mechanism of a Transformer-based language model to first identify polar sentences, and then either flip the polarity to the neutral or to the opposite through a GAN network. Tested on the same benchmark dataset, our framework achieves a 3%-10% improvement on the flipping/neutralizing success rate of headlines compared with the current state-of-the-art model. Adding to prior literature, our framework not only flips the polarity of headlines but also extends the task of polarity flipping to full-length articles. Human evaluation results show that our model successfully neutralizes or reverses the polarity of news without reducing readability. We release a large annotated dataset that includes both news headlines and full-length articles with polarity labels and meta-data to be used for future research. Our framework has a potential to be used by social scientists, content creators and content consumers in the real world. Ruibo Liu, Chenyan Jia, Soroush Vosoughi |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2020 | Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationabstractData augmentation is proven to be effective in many NLU tasks, especially for those suffering from data scarcity.In this paper, we present a powerful and easy to deploy text augmentation framework, Data Boost, which augments data through reinforcement learning guided conditional generation.We evaluate Data Boost on three diverse text classification tasks under five different classifier architectures.The result shows that Data Boost can boost the performance of classifiers especially in low-resource data scenarios.For instance, Data Boost improves F1 for the three tasks by 8.7% on average when given only 10% of the whole data for training.We also compare Data Boost with six prior text augmentation methods.Through human evaluations (N =178), we confirm that Data Boost augmentation has comparable quality as the original data with respect to readability and class consistency. Ruibo Liu, Guangxuan Xu, Chenyan Jia, Soroush Vosoughi |
EMNLP (1) | 3 |