VLDB 2026 Research / reviewers in the wild / expert
Diyi Yang
dblp:70/11145
· DBLP profile ↗
21ranked-venue papers in the field
7as first author
9since 2021 · last 2024
0000-0003-1220-3983ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (7 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HILITE: Human-in-the-loop Interactive Tool for Image EditingabstractImage editing tools have a plethora of commercial and creative applications — content-creation, digital photography, advertisements, graphic design, and development of educational media. The shortcomings of image editing software include difficulty of use and, for AI-based software, reliance on single image editing models, which often poses the dilemma of a tradeoff between image editing quality and user-friendliness. While the performances of individual image editing models have improved with their evolution over time, these singular models are often specialized on specific image editing tasks. In this work, we introduce HILITE, an open-source interactive image editing platform with a human-in-the-loop design that combines six diffusion-based image editing models. For one, HILITE’s accessible and easily-understandable user interface provides a straightforward user workflow from image input and prompt entry to selection of desired output. Secondly, the combination of several models with diverse specializations in turn allows HILITE to generalize on a wide variety of image editing tasks, essentially creating a "one-stop shop" for image editing. Third, HILITE iteratively takes user feedback, which both enhances the user experience and enables collection of crowd-sourced data for image editing. HILITE outperforms two major image editing softwares, OpenAI’s DALL•E 3 and Google’s Imagen 3, across two widely-user quantitative metrics for image editing evaluation. Considering the growing demand for readily-available and high-performing image editing tools, HILITE provides a novel platform design with multifaceted use cases in both business and academia. The platform can be found at https://platform.opennlplabs.org/ or https://platform-deployment.vercel.app/. Arya Pasumarthi, Armaan Sharma, Jainish H. Patel, Ayush Bheemaiah, Subhadra Vadlamannati, Seth Chang, Sophia Li, Eshaan Barkataki, Yutong Zhang 0011, Diyi Yang, Graham Neubig, Simran Khanuja |
IEEE Big Data | 10 |
| 2024 | A Crisis of Civility? Modeling Incivility and Its Effects in Political Discourse OnlineabstractGrowing concerns have been raised about the detrimental effects of uncivil comments on the web towards democracy. However, there is still a lack of understanding about online incivility's nuanced and complicated nature and its impact on conversation development and user behaviors. This work aims to fill that research gap by modeling incivility and its relationship to political discussions. We develop a comprehensive and fine-grained taxonomy that characterizes incivility with vulgarity, name-calling (inter-personal and third-party attacks), aspersion, and stereotypes, and then apply the framework to quantify the level of each incivility category in over 40 million comments from Reddit. Using large-scale quantitative analysis, we investigate the types of interactions and contexts in which incivility is more likely to occur, model how incivility shapes subsequent conversations, and examine user engagement patterns and behavioral changes after exposure to incivility. Our findings show that conversations that start out uncivil tend to become more uncivil in responses, and exposure to different incivility categories has differing effects on community members' engagement. We conclude with the implications of our research in assisting the design and moderation of online political communities. Wenna Qin, Aniruddha Murali, Christopher Eckart, Jacob Beel, Yi-Chia Wang, Diyi Yang |
ICWSM | 8 |
| 2023 | Just Another Day on Twitter: A Complete 24 Hours of Twitter DataabstractAt the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change. Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter |
ICWSM | 9 |
| 2023 | Graph Vulnerability and Robustness: A SurveyabstractThe study of network robustness is a critical tool in the characterization and sense making of complex interconnected systems such as infrastructure, communication and social networks. While significant research has been conducted in these areas, gaps in the surveying literature still exist. Answers to key questions are currently scattered across multiple scientific fields and numerous papers. In this survey, we distill key findings across numerous domains and provide researchers crucial access to important information by(1) summarizing and comparing recent and classical graph robustness measures; (2) exploring which robustness measures are most applicable to different categories of networks (e.g., social, infrastructure); (3) reviewing common network attack strategies, and summarizing which attacks are most effective across different network topologies; and (4) extensive discussion on selecting defense techniques to mitigate attacks across a variety of networks. This survey guides researchers and practitioners in navigating the expansive field of network robustness, while summarizing answers to key questions. We conclude by highlighting current research directions and open problems. Scott Freitas, Diyi Yang, Srijan Kumar, Hanghang Tong, Polo Chau |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Linguistic Characterization of Divisive Topics Online: Case Studies on Contentiousness in Abortion, Climate Change, and Gun Control
Jacob Beel, Tong Xiang, Sandeep Soni, Diyi Yang |
ICWSM | 4 |
| 2022 | Know It to Defeat It: Exploring Health Rumor Characteristics and Debunking Efforts on Chinese Social Media during COVID-19 Crisis
Wenjie Yang 0004, Sitong Wang 0001, Zhenhui Peng, Chuhan Shi, Xiaojuan Ma, Diyi Yang |
ICWSM | 6 |
| 2021 | Racism is a virus: anti-asian hate and counterspeech in social media during the COVID-19 crisisabstractThe spread of COVID-19 has sparked racism and hate on social media targeted towards Asian communities. However, little is known about how racial hate spreads during a pandemic and the role of counterspeech in mitigating this spread. In this work, we study the evolution and spread of anti-Asian hate speech through the lens of Twitter. We create COVID-HATE, the largest dataset of anti-Asian hate and counterspeech spanning 14 months, containing over 206 million tweets, and a social network with over 127 million nodes. By creating a novel hand-labeled dataset of 3,355 tweets, we train a text classifier to identify hateful and counterspeech tweets that achieves an average macro-F1 score of 0.832. Using this dataset, we conduct longitudinal analysis of tweets and users. Analysis of the social network reveals that hateful and counterspeech users interact and engage extensively with one another, instead of living in isolated polarized communities. We find that nodes were highly likely to become hateful after being exposed to hateful content in the year 2020. Notably, counterspeech messages discourage users from turning hateful, potentially suggesting a solution to curb hate on web and social media platforms. Data and code is available at http://claws.cc.gatech.edu/covid. Bing He 0002, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, Srijan Kumar |
ASONAM | 5 |
| 2021 | Evaluating Graph Vulnerability and Robustness using TIGERabstractNetwork robustness plays a crucial role in our understanding of complex interconnected systems such as transportation, communication, and computer networks. While significant research has been conducted in the area of network robustness, no comprehensive open-source toolbox currently exists to assist researchers and practitioners in this important topic. This lack of available tools hinders reproducibility and examination of existing work, development of new research, and dissemination of new ideas. We contribute TIGER, an open-sourced Python toolbox to address these challenges. TIGER contains 22 graph robustness measures with both original and fast approximate versions; 17 failure and attack strategies; 15 heuristic and optimization-based defense techniques; and 4 simulation tools. By democratizing the tools required to study network robustness, our goal is to assist researchers and practitioners in analyzing their own networks; and facilitate the development of new research in the field. TIGER has been integrated into the Nvidia Data Science Teaching Kit available to educators across the world; and Georgia Tech's Data and Visual Analytics class with over 1,000 students. TIGER is open sourced at: https://github.com/safreita1/TIGER Scott Freitas, Diyi Yang, Srijan Kumar, Hanghang Tong, Polo Chau |
CIKM | 2 |
| 2021 | Understanding the Invitation Acceptance in Agent-initiated Social E-commerce
Fengli Xu, Guozhen Zhang 0001, Yuan Yuan 0032, Hongjia Huang, Diyi Yang, Depeng Jin, Yong Li 0008 |
ICWSM | 5 |
| 2020 | Characterizing Collective Attention via Descriptor Context: A Case Study of Public Discussions of Crisis Events
Diyi Yang, Jacob Eisenstein |
ICWSM | 2 |
| 2018 | Understanding Self-Narration of Personally Experienced Racism on Reddit
Diyi Yang, Scott Counts |
ICWSM | 1 |
| 2017 | Self-Disclosure and Channel Difference in Online Health Support Groups
Diyi Yang, Zheng Yao 0006, Robert E. Kraut |
ICWSM | 1 |
| 2016 | Who Did What: Editor Role Identification in Wikipedia
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy |
ICWSM | 1 |
| 2014 | Constrained Question Recommendation in MOOCs via SubmodularityabstractA recent area in which recommender systems have shown their value is in online discussion forums and question-answer sites. Earlier work in this space has focused on the problem of matching participants to opportunities but has not adequately addressed the problem that in these social contexts, multiple dimensions of constraints must be satisfied, including limitations on capacity and minimal requirements for expertise. In this work, we propose such a constrained question recommendation problem with load balance constraints in discussion forums and use flow based model to generate the optimal solution. In particular, to address the introduced computation complexity, we investigate the concept of submodularity of the objective function and propose a specific submodular method to give an approximated solution. We present experiments conducted on two Massive Open Online Course (MOOC) discussion forum datasets, and demonstrate the effectiveness and efficiency of our submodular method in solving constrained question recommendation tasks. Diyi Yang, Jingbo Shang, Carolyn P. Rosé |
CIKM | 1 |
| 2014 | Linguistic Reflections of Student Engagement in Massive Open Online Courses
Miaomiao Wen, Diyi Yang, Carolyn P. Rosé |
ICWSM | 2 |
| 2014 | Question recommendation with constraints for massive open online coursesabstractMassive Open Online Courses (MOOCs) have experienced a recent boom in interest. Problems students struggle with in the discussion forums, such as difficultly in finding interesting discussion opportunities or attracting helpers to address posted problems, provide new opportunities for recommender systems. In contrast to traditional product recommendation, question recommendation in discussion forums should simultaneously consider constraints on both students and questions. These considerations include (1) Load Balancing - students should not be over-burdened with too many requests; and (2) Expertise Matching - students should not be requested to address problems they are not capable of addressing. In this work, we formulate a novel constrained question recommendation problem to address the above considerations. We design a context-aware matrix factorization model to predict students' preferences over questions, then build a max cost flow model to manage the constraints. Experimental results conducted on three MOOC datasets demonstrate that our method significantly outperforms baseline methods in optimizing overall forum welfare, and in predicting which specific questions students might be interested in. Diyi Yang, David Adamson, Carolyn P. Rosé |
RecSys | 1 |
| 2013 | Predicting advertiser bidding behaviors in sponsored search by rationality modelingabstractWe study how an advertiser changes his/her bid prices in sponsored search, by modeling his/her rationality. Predicting the bid changes of advertisers with respect to their campaign performances is a key capability of search engines, since it can be used to improve the offline evaluation of new advertising technologies and the forecast of future revenue of the search engine. Previous work on advertiser behavior modeling heavily relies on the assumption of perfect advertiser rationality; however, in most cases, this assumption does not hold in practice. Advertisers may be unwilling, incapable, and/or constrained to achieve their best response. In this paper, we explicitly model these limitations in the rationality of advertisers, and build a probabilistic advertiser behavior model from the perspective of a search engine. We then use the expected payoff to define the objective function for an advertiser to optimize given his/her limited rationality. By solving the optimization problem with Monte Carlo, we get a prediction of mixed bid strategy for each advertiser in the next period of time. We examine the effectiveness of our model both directly using real historical bids and indirectly using revenue prediction and click number prediction. Our experimental results based on the sponsored search logs from a commercial search engine show that the proposed model can provide a more accurate prediction of advertiser bid behaviors than several baseline methods. Bin Gao 0001, Diyi Yang, Tie-Yan Liu |
WWW | 3 |
| 2012 | Feature Based Informative Model for Discriminating Favorite Items from Unrated Ones
Tianqi Chen 0001, Diyi Yang, Weinan Zhang 0001, Yong Yu 0001 |
APWeb | 3 |
| 2012 | Local implicit feedback mining for music recommendationabstractDigital music has experienced a quite fascinating transformation during the past decades. Thousands of people share or distribute their music collections on the Internet, resulting in an explosive increase of information and more user dependence on automatic recommender systems. Though there are many techniques such as collaborative filtering, most approaches focus mainly on users' global behaviors, neglecting local actions and the specific properties of music. In this paper, we propose a simple and effective local implicit feedback model mining users' local preferences to get better recommendation performance in both rating and ranking prediction. Moreover, we design an efficient training algorithm to speed up the updating procedure, and give a method to find the most appropriate time granularity to assist the performance. We conduct various experiments to evaluate the performance of this model, which show that it outperforms baseline model significantly. Integration with existing temporal models achieves a great improvement compared to the reported best single model for Yahoo! Music. Diyi Yang, Tianqi Chen 0001, Weinan Zhang 0001, Qiuxia Lu, Yong Yu 0001 |
RecSys | 1 |
| 2012 | Collaborative filtering with short term preferences miningabstractRecently, recommender systems have fascinated researchers and benefited a variety of people's online activities, enabling users to survive the explosive web information. Traditional collaborative filtering techniques handle the general recommendation well. However, most such approaches usually focus on long term preferences. To discover more short term factors influencing people's decisions, we propose a short term preferences model, implemented with implicit user feedback. We conduct experiments comparing the performances of different short term models, which show that our model outperforms significantly compared to those long term models. Diyi Yang, Tianqi Chen 0001, Weinan Zhang 0001, Yong Yu 0001 |
SIGIR | 1 |
| 2012 | Serendipitous Personalized Ranking for Top-N RecommendationabstractSerendipitous recommendation has benefitted both e-retailers and users. It tends to suggest items which are both unexpected and useful to users. These items are not only profitable to the retailers but also surprisingly suitable to consumers' tastes. However, due to the imbalance in observed data for popular and tail items, existing collaborative filtering methods fail to give satisfactory serendipitous recommendations. To solve this problem, we propose a simple and effective method, called serendipitous personalized ranking. The experimental results demonstrate that our method significantly improves both accuracy and serendipity for top-N recommendation compared to traditional personalized ranking methods in various settings. Qiuxia Lu, Tianqi Chen 0001, Weinan Zhang 0001, Diyi Yang, Yong Yu 0001 |
Web Intelligence | 4 |