EDBT 2026 Demo / reviewers in the wild / expert
Guangyuan Piao
dblp:177/5856
· DBLP profile ↗
22ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0003-0516-2802ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Text-Based Hierarchical Multilabel Classification for Mobile Applications via Contrastive LearningabstractA hierarchical labeling system for mobile applications (apps) benefits a wide range of downstream businesses that integrate the labeling with their proprietary user data, to improve user modeling. Such a label hierarchy can define more granular labels that capture detailed app features beyond the limitations of traditional broad app categories. In this paper, we address the problem of hierarchical multilabel classification for apps by using their textual information such as names and descriptions. We present: 1) HMCN (Hierarchical Multilabel Classification Network) for handling the classification from two perspectives: the first focuses on a multilabel classification without hierarchical constraints, while the second predicts labels sequentially at each hierarchical level considering such constraints; 2) HMCL (Hierarchical Multilabel Contrastive Learning), a scheme that is capable of learning more distinguishable app representations to enhance the performance of HMCN. Empirical results on our Tencent App Store dataset and two public datasets demonstrate that our approach performs well compared with state-of-the-art methods. The approach has been deployed at Tencent and the multilabel classification outputs for apps have helped a downstream task--credit risk management of users--improve its performance by 10.70% with regard to the Kolmogorov-Smirnov metric, for over one year. Yang Xiao 0014, Weipeng Huang, Guangyuan Piao |
KDD (2) | 4 |
| 2025 | Demystifying Chains, Trees, and Graphs of ThoughtsabstractThe field of natural language processing (NLP) has witnessed significant progress in recent years, with a notable focus on improving large language models' (LLM) performance through innovative prompting techniques. Among these, prompt engineering coupled with structures has emerged as a promising paradigm, with designs such as Chain-of-Thought, Tree of Thoughts, or Graph of Thoughts, in which the overall LLM reasoning is guided by a structure such as a graph. As illustrated with numerous examples, this paradigm significantly enhances the LLM's capability to solve numerous tasks, ranging from logical or mathematical reasoning to planning or creative writing. To facilitate the understanding of this growing field and pave the way for future developments, we devise a general blueprint for effective and efficient LLM reasoning schemes. For this, we conduct an in-depth analysis of the prompt execution pipeline, clarifying and clearly defining different concepts. We then build the first taxonomy of structure-enhanced LLM reasoning schemes. We focus on identifying fundamental classes of harnessed structures, and we analyze the representations of these structures, algorithms executed with these structures, and many others. We refer to these structures as reasoning topologies, because their representation becomes to a degree spatial, as they are contained within the LLM context. Our study compares existing prompting schemes using the proposed taxonomy, discussing how certain design choices lead to different patterns in performance and cost. We also outline theoretical underpinnings, relationships between prompting and other parts of the LLM ecosystem such as knowledge bases, and the associated research challenges. Our work will help to advance future prompt engineering techniques. Maciej Besta, Florim Memedi, Robert Gerstenberger, Guangyuan Piao, Nils Blach, Piotr Nyczyk, Marcin Copik, Grzegorz Kwasniewski, Lukas Gianinazzi, Ales Kubicek, Hubert Niewiadomski, Aidan O'Mahony, Onur Mutlu, Torsten Hoefler |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Rényi Divergence Deep Mutual Learning
Weipeng Fuzzy Huang, Junjie Tao, Changbo Deng, Ming Fan 0002, Wenqiang Wan, Guangyuan Piao |
ECML/PKDD (2) | 7 |
| 2021 | Geometric Heuristics for Transfer Learning in Decision TreesabstractMotivated by a network fault detection problem, we study how recall can be boosted in a decision tree classifier, without sacrificing too much precision. This problem is relevant and novel in the context of transfer learning(TL), in which few target domain training samples are available. We define a geometric optimization problem for boosting the recall of a decision tree classifier, and show it is NP-hard. To solve it efficiently, we propose several near-linear time heuristics, and experimentally validate these heuristics in the context of TL. Our evaluation includes 7 public datasets, as well as 6 network fault datasets, and we compare our heuristics with several existing TL algorithms, as well as exact mixed integer linear programming(MILP) solutions to our optimization problem. We find that our heuristics boost recall in a manner similar to optimal MILP solutions, yet require several orders of magnitude less compute time. In many cases the F1 score of our approach is competitive, and often better, than other TL algorithms. Moreover, our approach can be used as a building block to apply transfer learning to more powerful ensemble methods, such as random forests. Siddhesh Chaubal, Mateusz Rzepecki, Patrick K. Nicholson, Guangyuan Piao, Alessandra Sala |
CIKM | 4 |
| 2021 | Recommending Knowledge Concepts on MOOC Platforms with Meta-path-based Representation Learning
Guangyuan Piao |
EDM | 1 |
| 2021 | Deadline-Aware TDMA Scheduling for Multihop Networks Using Reinforcement LearningabstractTime division multiple access (TDMA) is the medium access control strategy of choice for multihop networks with deterministic delay guarantee requirements. As such, many Internet of Things applications use protocols based on time division multiple access. Optimal slot assignment in such networks is NP-hard when there are strict deadline requirements and is generally done using heuristics that give suboptimal transmission schedules in linear time. However, existing heuristics make a scheduling decision at each time slot based on the same criterion without considering its effect on subsequent network states or scheduling actions. Here, we first identify a set of node features that capture the information necessary for network state representation to aid building schedules using Reinforcement Learning (RL). We then propose three different centralized approaches to RL-based TDMA scheduling that vary in training and network representation methods. Using RL allows applying diverse criteria at different time slots while considering the effect of a scheduling action on meeting the scheduling objective for the entire TDMA frame, resulting in better schedules. We compare the three proposed schemes in terms of how well they meet the scheduling objectives and their applicability to networks with memory and time constraints. One of the schemes proposed is RLSchedule, which is particularly suited to constrained networks. Simulation results for a variety of network scenarios show that RLSchedule reduces the percentage of packets missing deadlines by up to 60% compared to the best available baseline heuristic. Shanti Chilukuri, Guangyuan Piao, Diego Lugones, Dirk Pesch |
Networking | 2 |
| 2021 | Inferring Hierarchical Mixture Structures: A Bayesian Nonparametric Approach
Weipeng Huang, Nishma Laitonjam, Guangyuan Piao, Neil J. Hurley |
PAKDD (3) | 3 |
| 2021 | Learning to Predict the Departure Dynamics of Wikidata Editors
Guangyuan Piao, Weipeng Huang |
ISWC | 1 |
| 2021 | Data-Driven Energy Conservation in Cellular Networks: A Systems ApproachabstractThe energy consumption of mobile networks is already substantial nowadays, and only expected to further increase with the roll-out of 5G. Base stations are the key elements in this context: reducing their energy consumption is of paramount importance for network operators, not only to lower operating costs, but also to meet sustainable development goals. Today's base stations are typically over-provisioned, i.e., they comprise multiple cells to meet the peak load in a region. Therefore, substantial energy savings are possible by switching off cells that are under-utilized. This article proposes a data-driven approach to determine the time periods when a cell can be switched off. Forecasting is used to accurately predict network utilization and automatically find the time intervals to reliably switch off a cell. We carefully analyze the requirements of the system as a whole, from data collection to forecasting methods, to enable effective energy savings in practice. Considering several real-world traces from LTE networks, we show that an average of 10.24% energy savings is possible. We explore the trade-offs between energy savings and overhead in switching off cells, and provide insights into the choice of methods accordingly. In particular, we show that the accuracy of forecasting is not the most important factor in achieving energy savings; instead, the prediction (uncertainty) interval plays a key role in being able to achieve energy savings with less impact on end-users. Finally, we propose a model to generate utilization traces that match the distribution of real-world traces obtained from cellular networks. Gopika Premsankar, Guangyuan Piao, Patrick K. Nicholson, Mario Di Francesco, Diego Lugones |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | Partially Observable Markov Decision Process Modelling for Assessing HierarchiesabstractHierarchical clustering has been shown to be valuable in many scenarios. Despite its usefulness to many situations, there is no agreed methodology on how to properly evaluate the hierarchies produced from different techniques, particularly in the case where ground-truth labels are unavailable. This motivates us to propose a framework for assessing the quality of hierarchical clustering allocations which covers the case of no ground-truth information. This measurement is useful, e.g., to assess the hierarchical structures used by online retailer websites to display their product catalogues. Our framework is one of the few attempts for the hierarchy evaluation from a decision theoretic perspective. We model the process as a bot searching stochastically for items in the hierarchy and establish a measure representing the degree to which the hierarchy supports this search. We employ Partially Observable Markov Decision Processes (POMDP) to model the uncertainty, the decision making, and the cognitive return for searchers in such a scenario. Weipeng Huang, Guangyuan Piao, Neil J. Hurley |
ACML | 2 |
| 2020 | Mining User Interests from Social MediaabstractSocial media users readily share their preferences, life events, sentiment and opinions, and implicitly signal their thoughts, feelings, and psychological behavior. This makes social media a viable source of information to accurately and effectively mine users' interests with the hopes of enabling more effective user engagement, better quality delivery of appropriate services and higher user satisfaction. In this tutorial, we cover five important aspects related to the effective mining of user interests: (1) the foundations of social user interest modeling, such as information sources, various types of representation models and temporal features, (2) techniques that have been adopted or proposed for mining user interests, (3) different evaluation methodologies and benchmark datasets, (4) different applications that have been taking advantage of user interest mining from social media platforms, and (5) existing challenges, open research questions and exciting opportunities for further work. Fattane Zarrinkalam, Guangyuan Piao, Stefano Faralli 0001, Ebrahim Bagheri |
CIKM | 2 |
| 2020 | Env2Vec: accelerating VNF testing with deep learningabstractThe adoption of fast-paced practices for developing virtual network functions (VNFs) allows for continuous software delivery and creates a market advantage for network operators. This adoption, however, is problematic for testing engineers that need to assure, in shorter development cycles, certain quality of highly-configurable product releases running on heterogeneous clouds. Machine learning (ML) can accelerate testing workflows by detecting performance issues in new software builds. However, the overhead of maintaining several models for all combinations of build types, network configurations, and other stack parameters, can quickly become prohibitive and make the application of ML infeasible. Guangyuan Piao, Patrick K. Nicholson, Diego Lugones |
EuroSys | 1 |
| 2018 | Transfer Learning for Item Recommendations and Knowledge Graph Completion in Item Related Domains via a Co-Factorization Model
Guangyuan Piao, John G. Breslin |
ESWC | 1 |
| 2018 | Learning to Rank Tweets with Author-Based Long Short-Term Memory Networks
Guangyuan Piao, John G. Breslin |
ICWE | 1 |
| 2018 | Inferring user interests in microblogging social networks: a survey
Guangyuan Piao, John G. Breslin |
User Model. User Adapt. Interact. | 1 |
| 2017 | Inferring User Interests for Passive Users on Twitter by Leveraging Followee Biographies
Guangyuan Piao, John G. Breslin |
ECIR | 1 |
| 2017 | Factorization Machines Leveraging Lightweight Linked Open Data-Enabled Features for Top-N Recommendations
Guangyuan Piao, John G. Breslin |
WISE (2) | 1 |
| 2016 | User Modeling on Twitter with WordNet Synsets and DBpedia Concepts for Personalized RecommendationsabstractUser modeling of individual users on the Social Web platforms such as Twitter plays a significant role in providing personalized recommendations and filtering interesting information from social streams. Recently, researchers proposed the use of concepts (e.g., DBpedia entities) for representing user interests instead of word-based approaches, since Knowledge Bases such as DBpedia provide cross-domain background knowledge about concepts, and thus can be used for extending user interest profiles. Even so, not all concepts can be covered by a Knowledge Base, especially in the case of microblogging platforms such as Twitter where new concepts/topics emerge everyday. In this short paper, instead of using concepts alone, we propose using synsets from WordNet and concepts from DBpedia for representing user interests. We evaluate our proposed user modeling strategies by comparing them with other bag-of-concepts approaches. The results show that using synsets and concepts together for representing user interests improves the quality of user modeling significantly in the context of link recommendations on Twitter. Guangyuan Piao, John G. Breslin |
CIKM | 1 |
| 2016 | Interest Representation, Enrichment, Dynamics, and Propagation: A Study of the Synergetic Effect of Different User Modeling Dimensions for Personalized Recommendations on Twitter
Guangyuan Piao, John G. Breslin |
EKAW | 1 |
| 2016 | Towards Comprehensive User Modeling on the Social Web for Personalized Link RecommendationsabstractUser modeling for individual users on the Social Web plays a significant role and is a fundamental step for personalization as well as recommendations. Previous studies have proposed various user modeling strategies in different dimensions such as (1) interest representation, (2) interest propagation, (3) content enrichment and (4) temporal dynamics of user interests. This research mainly focuses on the first two dimensions interest representation and propagation. In addition, we also investigate the combination of these four dimensions and their synergistic effect on the quality of user modeling. Different user modeling strategies will then be evaluated in the context of personalized link recommender systems using standard evaluation methodologies such as Mean Reciprocal Rank (MRR), recall ([email protected]) and success ([email protected]) at rank N. Guangyuan Piao |
UMAP | 1 |
| 2016 | Analyzing Aggregated Semantics-enabled User Modeling on Google+ and Twitter for Personalized Link RecommendationsabstractIn this paper, we study if reusing Google+ profiles can provide reliable recommendations on Twitter to resolve the cold start problem. Next, we investigate the impact of giving different weights for aggregating user profiles from two OSNs and present that giving a higher weight to the targeted OSN profile for aggregation allows the best performance in the context of a personalized link recommender system. Finally, we propose a user modeling strategy which combines entity-and category-based user profiles using with a discounting strategy. Results show that our proposed strategy improves the quality of user modeling significantly compared to the baseline method. Guangyuan Piao, John G. Breslin |
UMAP | 1 |
| 2016 | Analyzing MOOC Entries of Professionals on LinkedIn for User Modeling and Personalized MOOC RecommendationsabstractThe main contribution of this work is the comparison of three user modeling strategies based on job titles, educational fields and skills in LinkedIn profiles, for personalized MOOC recommendations in a cold start situation. Results show that the skill-based user modeling strategy performs best, followed by the job- and edu-based strategies. Guangyuan Piao, John G. Breslin |
UMAP | 1 |