Guangyuan Piao

dblp:177/5856 · DBLP profile ↗
← Back
22ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0003-0516-2802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorComputer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Text-Based Hierarchical Multilabel Classification for Mobile Applications via Contrastive Learning
abstract
A hierarchical labeling system for mobile applications (apps) benefits a wide range of downstream businesses that integrate the labeling with their proprietary user data, to improve user modeling. Such a label hierarchy can define more granular labels that capture detailed app features beyond the limitations of traditional broad app categories. In this paper, we address the problem of hierarchical multilabel classification for apps by using their textual information such as names and descriptions. We present: 1) HMCN (Hierarchical Multilabel Classification Network) for handling the classification from two perspectives: the first focuses on a multilabel classification without hierarchical constraints, while the second predicts labels sequentially at each hierarchical level considering such constraints; 2) HMCL (Hierarchical Multilabel Contrastive Learning), a scheme that is capable of learning more distinguishable app representations to enhance the performance of HMCN. Empirical results on our Tencent App Store dataset and two public datasets demonstrate that our approach performs well compared with state-of-the-art methods. The approach has been deployed at Tencent and the multilabel classification outputs for apps have helped a downstream task--credit risk management of users--improve its performance by 10.70% with regard to the Kolmogorov-Smirnov metric, for over one year.
Yang Xiao 0014, Weipeng Huang, Guangyuan Piao
KDD (2)4
2025 Demystifying Chains, Trees, and Graphs of Thoughts
abstract
The field of natural language processing (NLP) has witnessed significant progress in recent years, with a notable focus on improving large language models' (LLM) performance through innovative prompting techniques. Among these, prompt engineering coupled with structures has emerged as a promising paradigm, with designs such as Chain-of-Thought, Tree of Thoughts, or Graph of Thoughts, in which the overall LLM reasoning is guided by a structure such as a graph. As illustrated with numerous examples, this paradigm significantly enhances the LLM's capability to solve numerous tasks, ranging from logical or mathematical reasoning to planning or creative writing. To facilitate the understanding of this growing field and pave the way for future developments, we devise a general blueprint for effective and efficient LLM reasoning schemes. For this, we conduct an in-depth analysis of the prompt execution pipeline, clarifying and clearly defining different concepts. We then build the first taxonomy of structure-enhanced LLM reasoning schemes. We focus on identifying fundamental classes of harnessed structures, and we analyze the representations of these structures, algorithms executed with these structures, and many others. We refer to these structures as reasoning topologies, because their representation becomes to a degree spatial, as they are contained within the LLM context. Our study compares existing prompting schemes using the proposed taxonomy, discussing how certain design choices lead to different patterns in performance and cost. We also outline theoretical underpinnings, relationships between prompting and other parts of the LLM ecosystem such as knowledge bases, and the associated research challenges. Our work will help to advance future prompt engineering techniques.
Maciej Besta, Florim Memedi, Robert Gerstenberger, Guangyuan Piao, Nils Blach, Piotr Nyczyk, Marcin Copik, Grzegorz Kwasniewski, Lukas Gianinazzi, Ales Kubicek, Hubert Niewiadomski, Aidan O'Mahony, Onur Mutlu, Torsten Hoefler
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Rényi Divergence Deep Mutual Learning
Weipeng Fuzzy Huang, Junjie Tao, Changbo Deng, Ming Fan 0002, Wenqiang Wan, Guangyuan Piao
ECML/PKDD (2)7
2021 Geometric Heuristics for Transfer Learning in Decision Trees
abstract
Motivated by a network fault detection problem, we study how recall can be boosted in a decision tree classifier, without sacrificing too much precision. This problem is relevant and novel in the context of transfer learning(TL), in which few target domain training samples are available. We define a geometric optimization problem for boosting the recall of a decision tree classifier, and show it is NP-hard. To solve it efficiently, we propose several near-linear time heuristics, and experimentally validate these heuristics in the context of TL. Our evaluation includes 7 public datasets, as well as 6 network fault datasets, and we compare our heuristics with several existing TL algorithms, as well as exact mixed integer linear programming(MILP) solutions to our optimization problem. We find that our heuristics boost recall in a manner similar to optimal MILP solutions, yet require several orders of magnitude less compute time. In many cases the F1 score of our approach is competitive, and often better, than other TL algorithms. Moreover, our approach can be used as a building block to apply transfer learning to more powerful ensemble methods, such as random forests.
Siddhesh Chaubal, Mateusz Rzepecki, Patrick K. Nicholson, Guangyuan Piao, Alessandra Sala
CIKM4
2021 Recommending Knowledge Concepts on MOOC Platforms with Meta-path-based Representation Learning
Guangyuan Piao
EDM1
2021 Deadline-Aware TDMA Scheduling for Multihop Networks Using Reinforcement Learning
abstract
Time division multiple access (TDMA) is the medium access control strategy of choice for multihop networks with deterministic delay guarantee requirements. As such, many Internet of Things applications use protocols based on time division multiple access. Optimal slot assignment in such networks is NP-hard when there are strict deadline requirements and is generally done using heuristics that give suboptimal transmission schedules in linear time. However, existing heuristics make a scheduling decision at each time slot based on the same criterion without considering its effect on subsequent network states or scheduling actions. Here, we first identify a set of node features that capture the information necessary for network state representation to aid building schedules using Reinforcement Learning (RL). We then propose three different centralized approaches to RL-based TDMA scheduling that vary in training and network representation methods. Using RL allows applying diverse criteria at different time slots while considering the effect of a scheduling action on meeting the scheduling objective for the entire TDMA frame, resulting in better schedules. We compare the three proposed schemes in terms of how well they meet the scheduling objectives and their applicability to networks with memory and time constraints. One of the schemes proposed is RLSchedule, which is particularly suited to constrained networks. Simulation results for a variety of network scenarios show that RLSchedule reduces the percentage of packets missing deadlines by up to 60% compared to the best available baseline heuristic.
Shanti Chilukuri, Guangyuan Piao, Diego Lugones, Dirk Pesch
Networking2
2021 Inferring Hierarchical Mixture Structures: A Bayesian Nonparametric Approach
Weipeng Huang, Nishma Laitonjam, Guangyuan Piao, Neil J. Hurley
PAKDD (3)3
2021 Learning to Predict the Departure Dynamics of Wikidata Editors
Guangyuan Piao, Weipeng Huang
ISWC1
2021 Data-Driven Energy Conservation in Cellular Networks: A Systems Approach
abstract
The energy consumption of mobile networks is already substantial nowadays, and only expected to further increase with the roll-out of 5G. Base stations are the key elements in this context: reducing their energy consumption is of paramount importance for network operators, not only to lower operating costs, but also to meet sustainable development goals. Today's base stations are typically over-provisioned, i.e., they comprise multiple cells to meet the peak load in a region. Therefore, substantial energy savings are possible by switching off cells that are under-utilized. This article proposes a data-driven approach to determine the time periods when a cell can be switched off. Forecasting is used to accurately predict network utilization and automatically find the time intervals to reliably switch off a cell. We carefully analyze the requirements of the system as a whole, from data collection to forecasting methods, to enable effective energy savings in practice. Considering several real-world traces from LTE networks, we show that an average of 10.24% energy savings is possible. We explore the trade-offs between energy savings and overhead in switching off cells, and provide insights into the choice of methods accordingly. In particular, we show that the accuracy of forecasting is not the most important factor in achieving energy savings; instead, the prediction (uncertainty) interval plays a key role in being able to achieve energy savings with less impact on end-users. Finally, we propose a model to generate utilization traces that match the distribution of real-world traces obtained from cellular networks.
Gopika Premsankar, Guangyuan Piao, Patrick K. Nicholson, Mario Di Francesco, Diego Lugones
IEEE Trans. Netw. Serv. Manag.2
2020 Partially Observable Markov Decision Process Modelling for Assessing Hierarchies
abstract
Hierarchical clustering has been shown to be valuable in many scenarios. Despite its usefulness to many situations, there is no agreed methodology on how to properly evaluate the hierarchies produced from different techniques, particularly in the case where ground-truth labels are unavailable. This motivates us to propose a framework for assessing the quality of hierarchical clustering allocations which covers the case of no ground-truth information. This measurement is useful, e.g., to assess the hierarchical structures used by online retailer websites to display their product catalogues. Our framework is one of the few attempts for the hierarchy evaluation from a decision theoretic perspective. We model the process as a bot searching stochastically for items in the hierarchy and establish a measure representing the degree to which the hierarchy supports this search. We employ Partially Observable Markov Decision Processes (POMDP) to model the uncertainty, the decision making, and the cognitive return for searchers in such a scenario.
Weipeng Huang, Guangyuan Piao, Neil J. Hurley
ACML2
2020 Mining User Interests from Social Media
abstract
Social media users readily share their preferences, life events, sentiment and opinions, and implicitly signal their thoughts, feelings, and psychological behavior. This makes social media a viable source of information to accurately and effectively mine users' interests with the hopes of enabling more effective user engagement, better quality delivery of appropriate services and higher user satisfaction. In this tutorial, we cover five important aspects related to the effective mining of user interests: (1) the foundations of social user interest modeling, such as information sources, various types of representation models and temporal features, (2) techniques that have been adopted or proposed for mining user interests, (3) different evaluation methodologies and benchmark datasets, (4) different applications that have been taking advantage of user interest mining from social media platforms, and (5) existing challenges, open research questions and exciting opportunities for further work.
Fattane Zarrinkalam, Guangyuan Piao, Stefano Faralli 0001, Ebrahim Bagheri
CIKM2
2020 Env2Vec: accelerating VNF testing with deep learning
abstract
The adoption of fast-paced practices for developing virtual network functions (VNFs) allows for continuous software delivery and creates a market advantage for network operators. This adoption, however, is problematic for testing engineers that need to assure, in shorter development cycles, certain quality of highly-configurable product releases running on heterogeneous clouds. Machine learning (ML) can accelerate testing workflows by detecting performance issues in new software builds. However, the overhead of maintaining several models for all combinations of build types, network configurations, and other stack parameters, can quickly become prohibitive and make the application of ML infeasible.
Guangyuan Piao, Patrick K. Nicholson, Diego Lugones
EuroSys1
2018 Transfer Learning for Item Recommendations and Knowledge Graph Completion in Item Related Domains via a Co-Factorization Model
Guangyuan Piao, John G. Breslin
ESWC1
2018 Learning to Rank Tweets with Author-Based Long Short-Term Memory Networks
Guangyuan Piao, John G. Breslin
ICWE1
2018 Inferring user interests in microblogging social networks: a survey
Guangyuan Piao, John G. Breslin
User Model. User Adapt. Interact.1
2017 Inferring User Interests for Passive Users on Twitter by Leveraging Followee Biographies
Guangyuan Piao, John G. Breslin
ECIR1
2017 Factorization Machines Leveraging Lightweight Linked Open Data-Enabled Features for Top-N Recommendations
Guangyuan Piao, John G. Breslin
WISE (2)1
2016 User Modeling on Twitter with WordNet Synsets and DBpedia Concepts for Personalized Recommendations
abstract
User modeling of individual users on the Social Web platforms such as Twitter plays a significant role in providing personalized recommendations and filtering interesting information from social streams. Recently, researchers proposed the use of concepts (e.g., DBpedia entities) for representing user interests instead of word-based approaches, since Knowledge Bases such as DBpedia provide cross-domain background knowledge about concepts, and thus can be used for extending user interest profiles. Even so, not all concepts can be covered by a Knowledge Base, especially in the case of microblogging platforms such as Twitter where new concepts/topics emerge everyday. In this short paper, instead of using concepts alone, we propose using synsets from WordNet and concepts from DBpedia for representing user interests. We evaluate our proposed user modeling strategies by comparing them with other bag-of-concepts approaches. The results show that using synsets and concepts together for representing user interests improves the quality of user modeling significantly in the context of link recommendations on Twitter.
Guangyuan Piao, John G. Breslin
CIKM1
2016 Interest Representation, Enrichment, Dynamics, and Propagation: A Study of the Synergetic Effect of Different User Modeling Dimensions for Personalized Recommendations on Twitter
Guangyuan Piao, John G. Breslin
EKAW1
2016 Towards Comprehensive User Modeling on the Social Web for Personalized Link Recommendations
abstract
User modeling for individual users on the Social Web plays a significant role and is a fundamental step for personalization as well as recommendations. Previous studies have proposed various user modeling strategies in different dimensions such as (1) interest representation, (2) interest propagation, (3) content enrichment and (4) temporal dynamics of user interests. This research mainly focuses on the first two dimensions interest representation and propagation. In addition, we also investigate the combination of these four dimensions and their synergistic effect on the quality of user modeling. Different user modeling strategies will then be evaluated in the context of personalized link recommender systems using standard evaluation methodologies such as Mean Reciprocal Rank (MRR), recall ([email protected]) and success ([email protected]) at rank N.
Guangyuan Piao
UMAP1
2016 Analyzing Aggregated Semantics-enabled User Modeling on Google+ and Twitter for Personalized Link Recommendations
abstract
In this paper, we study if reusing Google+ profiles can provide reliable recommendations on Twitter to resolve the cold start problem. Next, we investigate the impact of giving different weights for aggregating user profiles from two OSNs and present that giving a higher weight to the targeted OSN profile for aggregation allows the best performance in the context of a personalized link recommender system. Finally, we propose a user modeling strategy which combines entity-and category-based user profiles using with a discounting strategy. Results show that our proposed strategy improves the quality of user modeling significantly compared to the baseline method.
Guangyuan Piao, John G. Breslin
UMAP1
2016 Analyzing MOOC Entries of Professionals on LinkedIn for User Modeling and Personalized MOOC Recommendations
abstract
The main contribution of this work is the comparison of three user modeling strategies based on job titles, educational fields and skills in LinkedIn profiles, for personalized MOOC recommendations in a cold start situation. Results show that the skill-based user modeling strategy performs best, followed by the job- and edu-based strategies.
Guangyuan Piao, John G. Breslin
UMAP1