Jiaqing Yuan

dblp:276/1194 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-0539-3806ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Analyzing Reddit Stories of Sexual Violence: Incidents, Effects, and Requests for Advice
abstract
Warning: This paper may contain triggering language for some readers, especially survivors of sexual violence. Survivors of sexual violence sometimes share their experiences on social media, revealing their feelings and emotions and seeking advice. On platforms such as Reddit, some stories can be long---up to 40,000 characters. We posit that such long stories are demanding for helpers to read and respond to. Prior research has indicated that parts of these stories describing the incident, the effects on the poster, and advice requested by the poster are important. Highlighting those parts can draw helpers' attention toward key information and assist them in reading and responding to long stories. We first examine the stories posted on Reddit for the prevalence of these parts. Second, we develop a computational model to highlight these parts of a story. On ten-fold cross-validation of a dataset, our model achieves a macro F1 score of 0.82. In addition, we contribute METHREE, a dataset comprising 8,947 labeled sentences for these parts from Reddit stories. A survey of users who are helpers on some relevant subreddits shows that the parts highlighted by our tool represent important information and assist them while reading and responding to long stories. We find that these tool-generated highlights statistically significantly reduce the demandingness of long stories. Moreover, almost all helpers felt that highlighted stories are helpful and easier to read, understand, and respond to than nonhighlighted ones. In particular, on a 4-point Likert scale, there is about 0.7 point reduction in demandingess when stories were presented with highlights.
Hannah Javidi, Jiaqing Yuan, Ruijie Xi, Munindar P. Singh
ICWSM3
2025 A Benchmark for Cross-Domain Argumentative Stance Classification on Social Media
abstract
Argumentative stance classification plays a key role in identifying authors' viewpoints on specific topics. However, generating diverse pairs of argumentative sentences across various domains is challenging. Existing benchmarks often come from a single domain or focus on a limited set of topics. Additionally, manual annotation for accurate labeling is time-consuming and labor-intensive. To address these challenges, we propose leveraging platform rules, readily available expert-curated content, and large language models to bypass the need for human annotation. Our approach produces a multidomain benchmark comprising 4,498 topical claims and 30,961 arguments from three sources, spanning 21 domains. We benchmark the dataset in fully supervised, zero-shot, and few-shot settings, shedding light on the strengths and limitations of different methodologies.
Jiaqing Yuan, Ruijie Xi, Munindar P. Singh
ICWSM1
2023 Conversation Modeling to Predict Derailment
abstract
Conversations among online users sometimes derail, i.e., break down into personal attacks. Derailment interferes with the healthy growth of communities in cyberspace. The ability to predict whether an ongoing conversation will derail could provide valuable advance, even real-time, insight to both interlocutors and moderators. Prior approaches predict conversation derailment retrospectively without the ability to forestall the derailment proactively. Some existing works attempt to make dynamic predictions as the conversation develops, but fail to incorporate multisource information, such as conversational structure and distance to derailment. We propose a hierarchical transformer-based framework that combines utterance-level and conversation-level information to capture fine-grained contextual semantics. We propose a domain-adaptive pretraining objective to unite conversational structure information and a multitask learning scheme to leverage the distance from each utterance to derailment. An evaluation of our framework on two conversation derailment datasets shows an improvement in F1 score for the prediction of derailment. These results demonstrate the effectiveness of incorporating multisource information for predicting the derailment of a conversation.
Jiaqing Yuan, Munindar P. Singh
ICWSM1
2020 DNA4mC-LIP: a linear integration method to identify N4-methylcytosine site in multiple species
abstract
MOTIVATION: DNA N4-methylcytosine (4mC) is a crucial epigenetic modification. However, the knowledge about its biological functions is limited. Effective and accurate identification of 4mC sites will be helpful to reveal its biological functions and mechanisms. Since experimental methods are cost and ineffective, a number of machine learning-based approaches have been proposed to detect 4mC sites. Although these methods yielded acceptable accuracy, there is still room for the improvement of the prediction performance and the stability of existing methods in practical applications. RESULTS: In this work, we first systematically assessed the existing methods based on an independent dataset. And then, we proposed DNA4mC-LIP, a linear integration method by combining existing predictors to identify 4mC sites in multiple species. The results obtained from independent dataset demonstrated that DNA4mC-LIP outperformed existing methods for identifying 4mC sites. To facilitate the scientific community, a web server for DNA4mC-LIP was developed. We anticipated that DNA4mC-LIP could serve as a powerful computational technique for identifying 4mC sites and facilitate the interpretation of 4mC mechanism. AVAILABILITY AND IMPLEMENTATION: http://i.uestc.edu.cn/DNA4mC-LIP/. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qiang Tang 0013, Juanjuan Kang, Jiaqing Yuan, Hua Tang, Xianhai Li, Hao Lin 0001, Jian Huang 0004, Wei Chen 0064
Bioinform.3