VLDB 2026 Research / reviewers in the wild / expert
Baoxi Liu
dblp:216/7481
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An improved retrieval-augmented long-term grouting power prediction method: Rejecting low-similarity retrievals
Baoxi Liu, Liangsi Xu, Bingyu Ren, Chengyu Yu, Hongling Yu |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | UAD-ICL: Uncertainty-aware semantic control for trustworthy latent context in-context learning
Yilu Hu, Baoxi Liu |
Pattern Recognit. | 4 |
| 2024 | JargonFM: A Framework With Multiple Interpretation Modes for Jargon Understanding in Online CommunitiesabstractJargon words are commonly used in the communication of online communities. These words are characterized by special and implicit meanings that can only be comprehended by a small group of users, which brings challenges to community regulation and user engagement. For this problem, we present JargonFM, a framework with multiple interpretation modes for jargon understanding in online communities. JargonFM is designed based on the scientific explanation framework and supports three interpretation modes: jargon category prediction based on a jargon classifier, similar word identification based on a jargon synonyms selector, and representative text selection based on an example sentence selector. A jargon interpreter was also implemented to demonstrate the usage and usefulness of the interpretation framework. Automatic and human evaluations suggest that JargonFM can explain jargon words more accurately and more efficiently than the existing interpretation methods, leading to its wide acceptance among the evaluation participants. Zhengqing Guan, Peng Zhang 0060, Hansu Gu, Tun Lu, Baoxi Liu, Ning Gu 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2022 | A Personalized Cross-Platform Post Style Transfer Method Based on Transformer and Bi-Attention MechanismabstractTo meet different social purposes, users usually share content related to the same topic or event to multiple social media platforms (cross-platform content sharing). As the differences of social norms and audiences among these social ecosystems, there are differences in the use of words and expressions in different platforms, resulting in different language styles among different platforms. In reality, it is usually difficult for users to grasp the consistency between the language style of posts to be published and that of a platform as the problem of context collapse. To address this problem, firstly, we conduct an study to investigate users' content sharing practices across two Chinese popular social media platforms (Douban and Weibo). The results indicate that: 1) there are significant linguistic differences between different platforms; 2) users' content sharing practices are personalized, and the style of their newly shared content is correlated with their historical posts. Secondly, based on the above findings, we propose a personalized cross-platform post style transfer model. The model can automatically transfer users' posts from one platform's language style to the target platform's language style, while preserving the content and reflecting users' personalized characteristics as much as possible. Experiments on the datasets collected from Douban and Weibo show that our model generally outperforms other comparison models on both style transfer and personalization metrics. Baoxi Liu, Peng Zhang 0060, Tun Lu, Hansu Gu, Ning Gu 0001 |
WSDM | 2 |
| 2022 | A Semantic Embedding Enhanced Topic Model For User-Generated Textual Content Modeling In Social EcosystemsabstractAbstract The development of Information and Communication Technologies (ICT) and Web 2.0 promotes the emergence of diverse social ecosystems like social Internet of Things (IoT), social media and online communities. User-generated textual content (UGTC), which consists of unstructured texts, is the most important and common type of user-generated content in social ecosystems. UGTC in social ecosystems is generated according to two types of context information—global context (topics) and local context (semantic regularities). For UGTC modeling, topic models just consider global context but ignore semantic regularities, while semantic embedding models are on the opposite. So only utilizing topic models or semantic embedding models to model UGTC suffers from some drawbacks. For this problem, we propose a semantic embedding enhanced topic model named SEE-Twitter-LDA for accurately modeling UGTC in social ecosystems. The core of SEE-Twitter-LDA is that words are generated according to mutual semantic information of topics and semantic regularities. So global context and local context are jointly considered for UGTC modeling. By utilizing 553 098 tweets sampled from Twitter and 211 233 posts sampled from Weibo, we validate SEE-Twitter-LDA’s better performance on perplexity, topic divergence and topic coherence versus existing related models. Peng Zhang 0060, Baoxi Liu, Tun Lu, Hansu Gu, Xianghua Ding, Ning Gu 0001 |
Comput. J. | 2 |
| 2022 | Building a Personalized Model for Social Media Textual Content CensorshipabstractSocial media users often suffer from the problem of content over-disclosure. Most existing studies attempt to solve this problem by recommending proper audiences for users when sharing content. However, the audience management strategy cannot filter out sensitive information from the post and narrow the scope of content permeation. On the contrary, this paper conducts research from the content perspective and aims to design a content censorship model to help users evaluate the publicity of a post and find the sensitive information from it. The user can revise the content accordingly to achieve goals of sensitive information protection and broader content permeation. For this intention, we first built a dataset to explore the factors related to the public level of a post and the sensitive information. Based on the findings, a novel personalized multi-task content censorship model was built using several state-of-the-art deep learning techniques such as Seq2Seq and Co-training. We also implemented a prototype, i.e. a Browser plugin-based content censorship tool, by utilizing Weibo as a research site. Our model and its prototype were evaluated through automatic and human evaluations. The automatic evaluation suggests that our model outperforms the baseline methods on several metrics including precision, recall, and F1-score. The human evaluation also reveals that our model and prototype play an important role in helping users identify sensitive information. Based on these results, we proposed several insights for the future design of the social media content censorship system. Baoxi Liu, Peng Zhang 0060, Yubo Shu, Zhengqing Guan, Tun Lu, Hansu Gu, Ning Gu 0001 |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | Building User-oriented Personalized Machine Translator based on User-Generated Textual ContentabstractMachine Translation (MT) has been a very useful tool to assist multilingual communication and collaboration. In recent years, by taking advantage of the exciting developments of neural networks and deep learning, the accuracy and speed of machine translation have been continuously improved. However, most machine translation methods and systems are data-driven. They tend to select a consensus response represented in training data, while a user's preferred linguistic style, which is important for translation comprehension and user experience, is ignored. For this problem, we aim to build a user-oriented personalized machine translation model in this paper. The model aims to learn each user's linguistic style from the textual content that is generated by her/him (User-Generated Textual Content, UGTC) in social media context and generate personalized translation results utilizing several state-of-the-art deep learning techniques like Transformer and pre-training. We also implemented a user-oriented personalized machine translator using Weibo as a case of the source of UGTC to provide a systematical implementation scheme of a user-oriented personalized machine translation system based on our model. The translator was evaluated by automatic evaluation in combination with human evaluation. The results suggest that our model can generate more personalized, natural and lively translation results and enhance the comprehensibility of translation results, which makes its generations more preferred by users versus general translation results. Peng Zhang 0060, Zhengqing Guan, Baoxi Liu, Xianghua Ding, Tun Lu, Hansu Gu, Ning Gu 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Jointly Predicting Future Content in Multiple Social Media Sites Based on Multi-task LearningabstractUser-generated contents (UGC) in social media are the direct expression of users’ interests, preferences, and opinions. User behavior prediction based on UGC has increasingly been investigated in recent years. Compared to learning a person’s behavioral patterns in each social media site separately, jointly predicting user behavior in multiple social media sites and complementing each other (cross-site user behavior prediction) can be more accurate. However, cross-site user behavior prediction based on UGC is a challenging task due to the difficulty of cross-site data sampling, the complexity of UGC modeling, and uncertainty of knowledge sharing among different sites. For these problems, we propose a Cross-Site Multi-Task (CSMT) learning method to jointly predict user behavior in multiple social media sites. CSMT mainly derives from the hierarchical attention network and multi-task learning. Using this method, the UGC in each social media site can obtain fine-grained representations in terms of words, topics, posts, hashtags, and time slices as well as the relevances among them, and prediction tasks in different social media sites can be jointly implemented and complement each other. By utilizing two cross-site datasets sampled from Weibo, Douban, Facebook, and Twitter, we validate our method’s superiority on several classification metrics compared with existing related methods. Peng Zhang 0060, Baoxi Liu, Tun Lu, Xianghua Ding, Hansu Gu, Ning Gu 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Studying and Understanding Characteristics of Post-Syncing Practice and Goal in Social Network SitesabstractMany popular social network sites (SNSs) provide the post-syncing functionality, which allows users to synchronize posts automatically among different SNSs. Nowadays there exists divergence on this functionality from the view of sink SNS. The key to solving this problem is to understand the characteristics of users’ post-syncing practice and goals and evaluate whether they are consistent with an SNS’s norms, cultures, and goals. However, studying and understanding the characteristics of post-syncing practice and goal are challenging tasks as a result of the difficulty of data sampling and the complexity of post-syncing behavior. In this article, we focus on investigating this question by quantitative analysis in combination with qualitative analysis. In the quantitative study, by utilizing 211,233 synced-posts sampled from Weibo, we aim to investigate characteristics of post-syncing from three perspectives: user, content, and goal. The results suggest that post-syncing plays an important role in exhibiting one’s current activities, creations, and skills as well as advertisements but involves a risk of exhibiting personal sensitive profiles. To understand the results, we present an interview-based qualitative study based on thematic analysis. It indicates that the publicity, urgency, and remarkableness of contents and differences of social affordances and social circles between sink SNS and source SNS as well as the one-time consent of post-syncing authentication jointly account for the major role of post-syncing. Based on these results, we propose insights for post-syncing functionality’s adoption, design, and promotion. Peng Zhang 0060, Baoxi Liu, Xianghua Ding, Tun Lu, Hansu Gu, Ning Gu 0001 |
ACM Trans. Web | 2 |
| 2020 | Understanding Social Interaction across Social Network SitesabstractPeople tend to utilize multiple social network sites (SNSs) simultaneously to maintain some social relationships, which results that there are many overlapping relationships and interactions among SNSs. Although many studies have focused on social interaction and cross-SNS user footprint analysis and understanding, little research investigates social interaction from a perspective of two or more SNSs, and the interplay of interaction among SNSs has been unknown. In this paper, we aim to explore whether interaction building in a new SNS hinders the interaction frequency in an existing site, and if so, what kinds of users and relationships’ interactions are more or less likely to be affected. For these questions, we sampled 7,015 pairs of overlapping identities, 23,590 pairs of overlapping relationships and 6,771 pairs of overlapping interactions from Weibo and Douban and made analysis by combining multiple methods like Regression Discontinuity Design and Random-effects Negative Binomial Regression model. Our results suggest that no matter from the perspective of individuals or from the perspective of relationships, interaction construction in a new SNS is detrimental to interaction frequency in an existing site. Based on our findings, we also propose several valuable insights about how to enhance social interaction and promote its retention when users are involved into interacting practice in multiple platforms. Peng Zhang 0060, Tun Lu, Baoxi Liu, Hansu Gu, Ning Gu 0001 |
Int. J. Hum. Comput. Interact. | 3 |
| 2020 | A reliable cross-site user generated content modeling method based on topic model
Baoxi Liu, Peng Zhang 0060, Tun Lu, Ning Gu 0001 |
Knowl. Based Syst. | 1 |
| 2018 | Crowd evacuation simulation approach based on navigation knowledge and two-layer control mechanism
Hong Liu 0013, Baoxi Liu, Hao Zhang 0009, Guijuan Zhang |
Inf. Sci. | 2 |