Guangxia Li

dblp:23/8127 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 11 · 6 first-author · 3 since 2021Computer networks · 10 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Sharpness-aware Federated Graph Learning
abstract
One of many impediments to applying graph neural networks (GNNs) in processing large-volume real-world graph-structured data is that it disapproves of a centralized training scheme which involves gathering data belonging to different organizations due to privacy concerns. As a distributed data processing scheme, federated graph learning (FGL) enables learning GNN models collaboratively without sharing participants' private data. Though theoretically feasible, a core challenge in FGL systems is the variation of local training data distributions among clients, also known as the data heterogeneity problem. Most existing solutions suffer from two problems: (1) The typical optimizer based on empirical risk minimization tends to cause local models to fall into sharp valleys and weakens their generalization to out-of-distribution graph data. (2) The prevalent dimensional collapse in the learned representations of local graph data has an adverse impact on the classification capacity of the GNN model. To this end, we formulate a novel optimization objective that is aware of the sharpness (i.e., the curvature of the loss surface) of local GNN models. By minimizing the loss function and its sharpness simultaneously, we seek out model parameters in a flat region with uniformly low loss values, thus improving the generalization over heterogeneous data. By introducing a regularizer based on the correlation matrix of local representations, we relax the correlations of representations generated by individual local graph samples, so as to alleviate the dimensional collapse of the learned model. The proposed Sharpness-aware fEderated grAph Learning (SEAL) algorithm can enhance the classification accuracy and generalization ability of local GNN models in federated graph learning. Experimental studies on several graph classification benchmarks show that SEAL consistently outperforms SOTA FGL baselines and provides gains for more participants.
Ruiyu Li, Peige Zhao, Guangxia Li, Xingyu Gao 0001, Zhiqiang Xu 0003
WSDM3
2025 Learning to Explain: Towards Human-Aligned Explainability in Deep Reinforcement Learning via Attention Guidance
abstract
Recent advances in explainable deep reinforcement learning (DRL) have provided insights into the reasoning behind decisions made by DRL agents. However, existing methods often overlook the subjective nature of explanations and fail to consider human cognitive styles and preferences. Such ignorance tends to reduce the interpretability and relevance of the generated explanations from a human evaluator's perspective. To address this issue, we introduce human cognition into the explaining procedure by integrating DRL with attention guidance in a novel manner. The proposed concept proximal policy optimization (Concept-PPO) learns to generate human-aligned explanations by jointly optimizing the DRL performance and the discrepancy between generated explanations and human annotations. Its key component is a specially designed spatial concept transformer that can enhance explaining efficiency by premasking decision-irrelevant information. Experiments on the ATARI benchmark demonstrate that Concept-PPO achieves better policies than its black-box counterparts, and user studies confirm its superiority in generating human-aligned explanations compared to existing explainable DRL methods.
Bokai Ji, Guangxia Li
IJCAI2
2025 Matching Game Based Robust Service Recovery in Space-Air-Ground Integrated Network
abstract
As an important issue in the sixth generation communication technologies, the space-air-ground integrated network (SAG IN), mainly composed of satellites, unmanned aerial vehicles (UAVs), and ground stations, can provide global information services. However, it is challenging to provide robust services due to the dynamic characteristics of UAV s and satellites, as well as the resource incompatibility among different nodes. By introducing the network function virtualization technique to SAGIN, tasks can be converted into service function chains (SFCs) composed of multiple virtual network functions in series, and the resource allocation of SAGIN is deemed as the SFC deployment and scheduling. However, the node failure or link disconnections may occur in SAG IN, resulting in failures of SFC implementation. Hence, how to guarantee the robust service recovery of SFCs is challenging. In this paper, we propose the SFC deployment and recovery model to cope with the resource failure. The problem is formulated to minimize the total time consumption to complete the SFC deployment and recovery. Since the problem is an integer linear programming and intractable to solve, we propose an algorithm based on two-sided matching game to implement robust recovery of affected SFCs. Finally, simulation results verify the effectiveness and advantages of the proposed algorithm over other benchmark algorithms.
Yilu Cao, Ziye Jia, Lijun He 0005, Kun Guo 0002, Guangxia Li, Qihui Wu 0001
VTC2025-Spring5
2025 QoS-Guarantee Resource Allocation of Slicing Services in Integrated Satellite-Terrestrial Networks Based on Deep Reinforcement Learning
abstract
Integrated satellite-terrestrial networks (ISTNs) enable global connectivity but face challenges in efficient resource allocation due to increasing service demands. To address Quality of Service (QoS) degradation caused by inefficient resource allocation in ISTN's heterogeneous network, we propose a network slicing (NS) resource allocation algorithm based on deep reinforcement learning (DRL). First, an ISTN system model is constructed using NS, along with an evaluation approach for slicing services. Next, a satisfaction utility function is defined to quantify the QoS of slicing services, and an optimization problem is formulated. Then, based on the Markov decision process (MDP) and dueling double deep Q-learning (D3QN) theory, an NS resource allocation algorithm is designed, comprising both training and execution phases. Simulation results demonstrate that the proposed algorithm outperforms baseline approaches in system satisfaction, bandwidth allocation, and satellite network utilization.
Siying Hu, Xiaoqin Song, Ruizheng Ye, Guangxia Li
VTC2025-Spring4
2025 Training-free Graph Anomaly Detection: A Simple Approach via Singular Value Decomposition
abstract
Graph anomaly detection (GAD) is essential for identifying irregular behavior within graphs. Recent advances in GAD rely on deep learning techniques and have shown promise. However, prior deep learning-based GAD methods suffer from various limitations such as low accuracy, long training time, and limited scalability. To tackle these limitations, we propose TFGAD, a training-free graph anomaly detection approach. Our main idea is to process node attributes and local structures separately using distinct matrices, which are optimally determined via singular value decomposition, thus eliminating the need for additional training. For anomaly detection, we propose a lightweight scoring function that combines the reconstruction errors of node attributes with the projection lengths of local structures to quantify node abnormalities. Extensive experiments demonstrate that TFGAD significantly outperforms state-of-the-art deep learning-based baselines while reducing runtime and memory overhead. The results highlight TFGAD's potential as an effective and efficient solution for GAD, particularly in scenarios where computational resources are constrained.
Guangxia Li, Hao Weng, Yiyu Xiang
WWW2
2025 Online parallel multi-task relationship learning via alternating direction method of multipliers
Ruiyu Li, Peilin Zhao, Guangxia Li, Zhiqiang Xu 0003
Neurocomputing3
2024 A Simple and Effective Method for Anomaly Detection on Attributed Graphs via Feature Consistency
abstract
Anomaly detection on attributed graphs aims to identify rare nodes that deviate significantly from the majority of nodes. Although recent graph self-supervised learning methods have demonstrated great potential, their complex training and detection schemes may lead to suboptimal efficiency and effectiveness. In this study, we propose a simple and effective anomaly detection method for attributed graphs, based on the intuition that normal nodes can retain strong consistency between their attributes and links when compared to abnormal nodes. Our method measures the underlying consistency between the attributes and links of each node by transforming them into a common subspace through subspace projection and alignment. The resulting consistency measure serves as an indicator for quantifying the abnormality of nodes. As a natural combination with minimal additional overhead, our method further exploits the reconstruction errors of node attributes resulting from subspace projection to formulate a more complete and powerful anomaly indicator. Despite its simplicity, the experimental results demonstrate that the proposed method can achieve competitive or superior performance in a fully unsupervised manner. Moreover, we extend our method using deep learning techniques, leading to significant improvements over state-of-the-art methods on benchmark attributed graph datasets.
Guangxia Li
ICASSP2
2024 A Context Augmented Multi-Play Multi-Armed Bandit Algorithm for Fast Channel Allocation in Opportunistic Spectrum Access
abstract
We study the restless contextual multi-play multi-armed bandit (MP-MAB) problem for channel allocation in the opportunity spectrum access (OSA) scenario. Most existing MP-MAB methods are impractical for real-world OSA systems as they assume many ideal conditions, incur a heavy computational cost, and most importantly, ignore the impact of channel noise which is directly related to the quality of service. In this study, we embody this impact by modeling channel noise as a perturbation of the arm’s reward function in MP-MAB. As there is an implicit correlation between channel state information and channel noise, we take the former as a context for MP-MAB to present the perturbation caused by the latter. We investigate two types of correlation between the context and the perturbation—linear and nonlinear, and derive two index policies, respectively. These policies learn the correlations through a linear model and a neural network, and use estimated noise value to adjust the upper confidence bound. Numerical experiments demonstrate that the proposed policies can achieve lower regret and select sub-optimal arms in a more reasonable way.
Ruiyu Li, Guangxia Li, Xiao Lu 0001, Jichao Liu
ISCC2
2024 DNN Tasks Offloading and Bandwidth Optimization for Satellite-Terrestrial Collaborative Intelligence
abstract
Deep Neural Networks (DNNs) are now widely used in Low Earth Orbit (LEO) satellites, such as in remote sensing and environmental monitoring. DNN tasks are generally resource-intensive, while the resources of LEO satellites including computation and storage resources are usually limited, which implies directly running high-precision and complex DNNs on them is extremely challenging. A promising way is leveraging the layered structure of DNNs and executing DNN tasks collaboratively between satellites and ground, i.e., satellite-terrestrial collaborative inference. However, most existing works about satellite- terrestrial collaborative inference mainly focus on the optimization of DNN offloading strategy in terms of latency and energy minimization, without considering how to minimize the highly precious satellite communication resources in the collaboration. In this paper, we study how to jointly optimize the offloading decision and satellites' communication bandwidth, to achieve the minimization of weighted sum of latency, energy consumption, and communication bandwidth consumption. The aforementioned problem is a Mixed Integer Nonlinear Programming (MINLP) problem and hard to resolve. We design an alternating optimization algorithm combining branch-and-bound and gradient descent methods (AO-SA) to obtain an efficient solution. Extensive simulations validate the efficiency of the proposed algorithm: compared to existing satellite-terrestrial offloading algorithms, it improves the performance in terms of latency and energy consumption by up to 31 %, while saving the bandwidth resource of satellites by 28 % on average.
Haochun Lei, Yuben Qu, Lei Zhang 0038, Lingyuan Zhao, Guangxia Li, Qihui Wu 0001
MSN7
2023 Enhancing the Interpretability of Deep Multi-agent Reinforcement Learning via Neural Logic Reasoning
Bokai Ji, Guangxia Li
ICANN (10)2
2023 Transmit Power Optimization and Precoding Design in Multiuser Satellite MIMO Downlink With SINR Constraints
abstract
In this article, a multiuser (MU) satellite multiple-input–multiple-output (MIMO) downlink including a multibeam GEO satellite and multiple terrestrial users is considered. In order to meet the signal-to-interference-plus-noise ratio (SINR) requirements of different users while minimizing the satellite transmit power, the onboard precoding problem is investigated. Zero-forcing (ZF) and regularized ZF (RZF) precoding are first adopted and it is found that they can satisfy the SINR constraints by multiplying by a power allocation matrix. The required power allocation matrix is divided into two types: 1) equal power factor schemes and 2) unequal power factor schemes, and they are derived separately, where Karush–Kuhn–Tucker (KKT) conditions are used in RZF unequal power factor scheme and a heuristic algorithm is proposed to optimize the regularization factor in RZF schemes to further reduce the transmit power. Then, a semidefinite programming beamforming (SDPBF) is proposed to jointly optimize precoding and power allocation matrices. The closed-form expressions of the transmit power of the above five precoding schemes are derived and their performance is compared in the simulation part. The research results show that under the same conditions, the performance of RZF is better than that of ZF, the unequal power factor scheme is better than the equal power factor scheme, and SDPBF has the best performance. Besides, it is also found that in addition to changing the precoding method, activating more beams will further reduce the transmit power.
Hongpeng Zhu, Guangxia Li
IEEE Internet Things J.5
2023 User Selection in ZF Precoding Multi-User Satellite MIMO Downlink With QoS Constraints
abstract
In this paper, the user selection problem is investigated in the multi-user satellite multiple-input multiple-output (MIMO) downlink, which includes a multi-beam geostationary earth orbit (GEO) satellite in space and multiple terrestrial users on the ground. Considering that the number of users is much larger than that of beams or feeds, user selection should be performed so that multiple users can be served successively in different time slots. In order to meet the quality of service (QoS) requirements of users and maximize the system sum rate, three user selection algorithms are proposed: greedy user selection (GUS) algorithm, add swap user selection (AS) algorithm, and random add swap user selection (RAS) algorithm. Among the three algorithms, only the “add” operation is adopted in the GUS algorithm, both the “add”, and “swap” operations are considered in the AS algorithm. On this basis, heuristic search and “random swap” operation are introduced and the RAS algorithm is obtained. The complexity of the above three algorithms is analyzed analytically and the RAS algorithm is found to have the highest complexity. Besides, the performance of the three algorithms is analyzed and compared through simulations, where the random user selection (RUS) algorithm is taken as a benchmark. Numerical results show that the system sum rate and energy efficiency are greatly improved by the proposed algorithms and the RAS algorithm has the best performance.
Hongpeng Zhu, Yinxia Zhu, Dongming Bian, Guangxia Li
IEEE Trans. Commun.7
2023 Load Balancing of Double Queues and Utility-Workload Tradeoff in Heterogeneous Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a popular service paradigm by which mobile devices can offload their latency-sensitive and computation-intensive workloads to edge servers. The MEC service scheduling problem has been investigated in recent years. However, most MEC service scheduling mechanisms only consider workloads on homogeneous edge servers, causing servers’ queue backlogs to be too large when innumerable user requests arrive concurrently. In this paper, we are the first to propose a double-queue workloads scheduling model innovatively, and formulate a system (including user ends and edge server ends) utility into a scheduling optimization problem. To tackle such an NP scheduling problem, we present a Lyapunov-based decomposition strategy to convert the original problem into three equivalent subproblems. By aggregating three subproblem solving strategies, we propose the Lyapunov-based online matching algorithm for edge service scheduling, named LOMES, to obtain an optimal system utility while guaranteeing the load balancing of mobile devices and heterogeneous edge servers. Simulations further validate that LOMES realizes the load balancing of two queue lengths and a$[O(1/V); O(V)]$tradeoff between the system’s utility and workloads with a utility-workload tradeoff parameter${V}$.
Xuewen Dong, Zijie Di, Liangmin Wang 0001, Qingsong Yao, Guangxia Li, Yulong Shen 0001
IEEE Trans. Wirel. Commun.5
2022 Towards Relational Multi-Agent Reinforcement Learning via Inductive Logic Programming
Guangxia Li
ICANN (2)1
2022 Dual Adversarial Federated Learning on Non-IID Data
Tao Zhang 0029, Shaojing Yang, Anxiao Song, Guangxia Li, Xuewen Dong
KSEM (3)4
2022 Cooperative Relative Localization for UAV Swarm in GNSS-Denied Environment: A Coalition Formation Game Approach
abstract
Unmanned aerial vehicle (UAV) swarms require accurate relative localization to safeguard flight missions in the global navigation satellite system-denied environment due to a lack of absolute position information. The existing work of relative localization faces challenges, such as the ranging information loss and the low localization frequency caused by long-distance ranging and large-scale characteristic of the UAV swarm. This article proposes a clustering-based cooperative relative localization scheme for UAV swarm, which contains a two-level framework: inter/intra-cluster localization. In order to investigate the tradeoff between intracluster cooperation and intercluster packet loss, the clustering-based problem is constructed as a coalition formation game (CFG) model. Given the designed coalition value, preference relation, and the coalition formation principles, it is proved that the proposed CFG model has a Nash stable partition. The designed coalition formation algorithm includes coalition heads and beacon drones selection mechanism. Simulation results show that the proposed CFG algorithms shorten the ranging time compared with global localization and achieve better localization performance (localization error and success rate) than contrast algorithms.
Lang Ruan, Guangxia Li, Weiheng Dai, Shiwei Tian, Guangteng Fan, Jian Wang 0007, Xiaoqi Dai
IEEE Internet Things J.2
2020 Collaborative online ranking algorithms for multitask learning
Guangxia Li, Peilin Zhao, Tao Mei 0001, Peng Yang 0010, Yulong Shen 0001, Kuiyu Chang, Steven C. H. Hoi
Knowl. Inf. Syst.1
2020 Robust Hybrid Cooperative Positioning Via a Modified Distributed Projection-Based Method
abstract
Cooperative positioning is attracting an increasing amount of attention due to its ability to enhance the accuracy and availability of positioning performance. Current algorithms for cooperative positioning are sensitive to the initial guess as a result of their nonconvex objective functions, which is especially true in hybrid wireless networks. Perfect a priori information about the locations is needed, which is rather problematic in many scenarios. With strong convergence, the iterative parallel projection method (IPPM) is extended to hybrid wireless networks (H-IPPM) in this paper. Motivated by the fact that normal weighted methods cannot achieve the optimal solution, the position uncertainty is modeled, and two distributed weighted parallel projection algorithms, namely, an inexact weighted algorithm called the HBFW-IPPM and an exact weighted algorithm called the HCPW-IPPM, are developed when considering both the range measurement errors and position uncertainty. Experiments in a realistic outdoor scenario are conducted. The results indicate that the exact weighted algorithm HCPW-IPPM shows superior and robust performance in both warm-start and cold-start conditions, and this is true even when non-line of sight (NLOS) measurements and weight estimation errors are taken into account.
Tianwei Liu, Guangxia Li, Siming Li, Shiwei Tian
IEEE Trans. Wirel. Commun.2
2020 An incentive mechanism with bid privacy protection on multi-bid crowdsourced spectrum sensing
Xuewen Dong, Guangxia Li, Tao Zhang 0029, Di Lu 0001, Yulong Shen 0001, Jianfeng Ma 0001
World Wide Web2
2019 Data Analytics for Fog Computing by Distributed Online Learning with Asynchronous Update
abstract
Fog computing extends the cloud computing paradigm by allocating substantial portions of computations and services towards the edge of a network, and is, therefore, particularly suitable for large-scale, geo-distributed, and data-intensive applications. As the popularity of fog applications increases, there is a demand for the development of smart data analytic tools, which can process massive data streams in an efficient manner. To satisfy such requirements, we propose a system in which data streams generated from distributed sources are digested almost locally, whereas a relatively small amount of distilled information is converged to a center. The center extracts knowledge from the collected information, and shares it across all subordinates to boost their performances. Upon the proposed system, we devise a distributed machine learning algorithm using the online learning approach, which is well known for its high efficiency and innate ability to cope with streaming data. An asynchronous update strategy with rigorous theoretical support is applied to enhance the system robustness. Experimental results demonstrate that the proposed method is comparable with a model trained over a centralized platform in terms of the classification accuracy, whereas the efficiency and scalability of the overall system are improved.
Guangxia Li, Peilin Zhao, Xiao Lu 0001, Jia Liu 0009, Yulong Shen 0001
ICC1
2019 Detecting cyberattacks in industrial control systems using online learning algorithms
Guangxia Li, Yulong Shen 0001, Peilin Zhao, Xiao Lu 0001, Jia Liu 0009, Steven C. H. Hoi
Neurocomputing1
2018 Performance Analysis of Wireless-Powered Relaying with Ambient Backscattering
abstract
With the increasing use of smart objects, such as wearable health gadgets, household automation devices, and personal electronics, there is a growing demand for a globally interconnected information network, known as the Internet of Things (IoT). IoT is featured with low-power communications among a massive number of ubiquitously-deployed and energy-constrained electronics, like sensors and actuators. In this context, wireless-powered cooperative relaying emerges as a promising solution to extend coverage and solve energy scarcity problems for IoT devices. In this paper, we propose a novel hybrid relay by combining wireless-powered communications and ambient backscattering functions for improved applicability and performance. To well adapt the hybrid relay to the network environments, we design a mode selection protocol to coordinate between the two functions. Moreover, we analyze the successful transmission probability of a dual-hop relaying system with the hybrid relay. Through numerical results, we demonstrate the performance gain of the hybrid relay and the impact of the system parameters.
Xiao Lu 0001, Guangxia Li, Hai Jiang 0001, Dusit Niyato, Ping Wang 0001
ICC2
2017 A Cloud-Based Stream Processing Platform for Traffic Monitoring Using Large-Scale Probe Vehicle Data
abstract
Probe vehicle data, also known as floating car data or connected vehicle data, is the data collected from GPS-enabled sensors on vehicles. With the advancement in wireless communications and localization technologies, more and more vehicles are expected to be equipped with such sensors. Existing studies only focus on using small-scale probe vehicle data. In this paper, we are interested in developing a real-time parallel stream processing framework to extract traffic flow KPIs from large-scale probe vehicle data. The developed framework is implemented using Apache Storm on Amazon AWS, and can process one million probe vehicle messages per second. Various design considerations, such as data partition and delay processing are discussed. To evaluate the performance of stream processing framework, simulated probe vehicle data based on the actual traffic flows in Jurong Lake District (JLD) of Singapore, is generated using the microscopic simulation software VISSIM. The JLD data is replicated multiple times to represent the one million population of vehicles in Singapore. GPS errors and communication delays are added to represent the real situations before the data is fed to stream processing module. The estimated KPIs from our stream processing model are validated against the ground truth values under different penetration levels.
Yiyang Pei, Guangxia Li, Hai-Heng Ng, Kah Eng Hoe, Chee-Wei Ang, Wee Siong Ng, Kenji Takao, Hirokazu Shibata, Koichiro Okada
WCNC4
2016 Deceptive Review Spam Detection via Exploiting Task Relatedness and Unlabeled Data
Zhen Hai, Peilin Zhao, Peng Cheng 0008, Peng Yang 0010, Xiaoli Li 0001, Guangxia Li
EMNLP6
2016 Learning Correlative and Personalized Structure for Online Multi-Task Classification
abstract
Multi-Task Learning (MTL) can enhance the classifier's generalization performance by learning multiple related tasks simultaneously. Conventional MTL works under the offline or batch learning setting and suffers from the expensive training cost together with the poor scalability. To address such inefficiency issues, online learning technique has been applied to solve MTL problems. However, most existing algorithms for online MTL constrain task relatedness into a presumed structure via a single weight matrix, a strict restriction that does not always hold in practice. In this paper, we propose a general online MTL framework that overcomes this restriction by decomposing the weight matrix into two components: the first component captures the correlative structure among tasks in a low-rank subspace, and the second component identifies the personalized patterns for the outlier tasks. A projected gradient scheme is devised to learn such components adaptively. Theoretical analysis shows that the proposed algorithm can achieve a sub-linear regret with respect to the best linear model in hindsight. Experimental investigation on a number of real-world datasets also verifies the efficacy of our approach.
Peng Yang 0010, Guangxia Li, Peilin Zhao, Xiaoli Li 0001, Sujatha Das Gollapalli
SDM2
2015 Fast PageRank approximation by adaptive sampling
Guangxia Li, James Cheng
Knowl. Inf. Syst.2
2015 Performance of ML Range Estimator in Radio Interferometric Positioning Systems
abstract
The radio interferometric positioning system (RIPS) is a novel positioning solution used in wireless sensor networks. This letter explores the ranging accuracy of RIPS in two configurations. In the linear step-frequency configuration, we derive the mean squared error (MSE) of the maximum likelihood (ML) estimator. In the random step-frequency configuration, we introduce average MSE to characterize the performance of the ML estimator. The simulation results fit well with theoretical analysis.
Wangdong Qi, Guangxia Li
IEEE Signal Process. Lett.3
2014 Extracting rate changes in transcriptional regulation from MEDLINE abstracts
abstract
BACKGROUND: Time delays are important factors that are often neglected in gene regulatory network (GRN) inference models. Validating time delays from knowledge bases is a challenge since the vast majority of biological databases do not record temporal information of gene regulations. Biological knowledge and facts on gene regulations are typically extracted from bio-literature with specialized methods that depend on the regulation task. In this paper, we mine evidences for time delays related to the transcriptional regulation of yeast from the PubMed abstracts. RESULTS: Since the vast majority of abstracts lack quantitative time information, we can only collect qualitative evidences of time delays. Specifically, the speed-up or delay in transcriptional regulation rate can provide evidences for time delays (shorter or longer) in GRN. Thus, we focus on deriving events related to rate changes in transcriptional regulation. A corpus of yeast regulation related abstracts was manually labeled with such events. In order to capture these events automatically, we create an ontology of sub-processes that are likely to result in transcription rate changes by combining textual patterns and biological knowledge. We also propose effective feature extraction methods based on the created ontology to identify the direct evidences with specific details of these events. Our ontologies outperform existing state-of-the-art gene regulation ontologies in the automatic rule learning method applied to our corpus. The proposed deterministic ontology rule-based method can achieve comparable performance to the automatic rule learning method based on decision trees. This demonstrates the effectiveness of our ontology in identifying rate-changing events. We also tested the effectiveness of the proposed feature mining methods on detecting direct evidence of events. Experimental results show that the machine learning method on these features achieves an F1-score of 71.43%. CONCLUSIONS: The manually labeled corpus of events relating to rate changes in transcriptional regulation for yeast is available in https://sites.google.com/site/wentingntu/data. The created ontologies summarized both biological causes of rate changes in transcriptional regulation and corresponding positive and negative textual patterns from the corpus. They are demonstrated to be effective in identifying rate-changing events, which shows the benefits of combining textual patterns and biological knowledge on extracting complex biological events.
Kui Miao, Guangxia Li, Kuiyu Chang, Jie Zheng 0002, Jagath C. Rajapakse
BMC Bioinform.3
2014 Collaborative Online Multitask Learning
abstract
We study the problem of online multitask learning for solving multiple related classification tasks in parallel, aiming at classifying every sequence of data received by each task accurately and efficiently. One practical example of online multitask learning is the micro-blog sentiment detection on a group of users, which classifies micro-blog posts generated by each user into emotional or non-emotional categories. This particular online learning task is challenging for a number of reasons. First of all, to meet the critical requirements of online applications, a highly efficient and scalable classification solution that can make immediate predictions with low learning cost is needed. This requirement leaves conventional batch learning algorithms out of consideration. Second, classical classification methods, be it batch or online, often encounter a dilemma when applied to a group of tasks, i.e., on one hand, a single classification model trained on the entire collection of data from all tasks may fail to capture characteristics of individual task; on the other hand, a model trained independently on individual tasks may suffer from insufficient training data. To overcome these challenges, in this paper, we propose a collaborative online multitask learning method, which learns a global model over the entire data of all tasks. At the same time, individual models for multiple related tasks are jointly inferred by leveraging the global model through a collaborative online learning approach. We illustrate the efficacy of the proposed technique on a synthetic dataset. We also evaluate it on three real-life problems-spam email filtering, bioinformatics data classification, and micro-blog sentiment detection. Experimental results show that our method is effective and scalable at the online classification of multiple related tasks.
Guangxia Li, Steven C. H. Hoi, Kuiyu Chang, Ramesh Jain 0001
IEEE Trans. Knowl. Data Eng.1
2012 Real-time stereo matching based on fast belief propagation
Xueqin Xiang, Guangxia Li, Yuyong He
Mach. Vis. Appl.3
2012 Multiview Semi-Supervised Learning with Consensus
abstract
Obtaining high-quality and up-to-date labeled data can be difficult in many real-world machine learning applications. Semi-supervised learning aims to improve the performance of a classifier trained with limited number of labeled data by utilizing the unlabeled ones. This paper demonstrates a way to improve the transductive SVM, which is an existing semi-supervised learning algorithm, by employing a multiview learning paradigm. Multiview learning is based on the fact that for some problems, there may exist multiple perspectives, so called views, of each data sample. For example, in text classification, the typical view contains a large number of raw content features such as term frequency, while a second view may contain a small but highly informative number of domain specific features. We propose a novel two-view transductive SVM that takes advantage of both the abundant amount of unlabeled data and their multiple representations to improve classification result. The idea is straightforward: train a classifier on each of the two views of both labeled and unlabeled data, and impose a global constraint requiring each classifier to assign the same class label to each labeled and unlabeled sample. We also incorporate manifold regularization, a kind of graph-based semi-supervised learning method into our framework. The proposed two-view transductive SVM was evaluated on both synthetic and real-life data sets. Experimental results show that our algorithm performs up to 10 percent better than a single-view learning approach, especially when the amount of labeled data is small. The other advantage of our two-view semi-supervised learning approach is its significantly improved stability, which is especially useful when dealing with noisy data in real-world applications.
Guangxia Li, Kuiyu Chang, Steven C. H. Hoi
IEEE Trans. Knowl. Data Eng.1
2011 Collaborative online learning of user generated content
abstract
We study the problem of online classification of user generated content, with the goal of efficiently learning to categorize content generated by individual user. This problem is challenging due to several reasons. First, the huge amount of user generated content demands a highly efficient and scalable classification solution. Second, the categories are typically highly imbalanced, i.e., the number of samples from a particular useful class could be far and few between compared to some others (majority class). In some applications like spam detection, identification of the minority class often has significantly greater value than that of the majority class. Last but not least, when learning a classification model from a group of users, there is a dilemma: A single classification model trained on the entire corpus may fail to capture personalized characteristics such as language and writing styles unique to each user. On the other hand, a personalized model dedicated to each user may be inaccurate due to the scarcity of training data, especially at the very beginning; when users have written just a few articles. To overcome these challenges, we propose learning a global model over all users' data, which is then leveraged to continuously refine the individual models through a collaborative online learning approach. The class imbalance problem is addressed via a cost-sensitive learning approach. Experimental results show that our method is effective and scalable for timely classification of user generated content.
Guangxia Li, Kuiyu Chang, Steven C. H. Hoi, Ramesh Jain 0001
CIKM1
2010 Fast and Simple Super Resolution for Range Data
abstract
Current active 3D range sensors, such as time-of-flight cameras, enable acquiring of range maps at video frame rate. Unfortunately, the resolution of the range maps is quite limited and the captured data are typically contaminated by noise. We therefore present a simple pipeline to enhance the quality as well as improve the spatial and depth resolution of range data in real time by up sampling the depth information with the data from high resolution video camera and utilizing a new strategy to increase the sub-pixel accuracy. Our algorithm can greatly improve the reconstruction quality, boost the resolution of the range data to that of video sensor while achieving high computational efficiency for a real-time application.
Xueqin Xiang, Guangxia Li, Jing Tong
CW2
2010 Micro-blogging Sentiment Detection by Collaborative Online Learning
abstract
We study the online micro-blog sentiment detection problem, which aims to determine whether a micro-blog post expresses emotions. This problem is challenging because a micro-blog post is very short and individuals have distinct ways of expressing emotions. A single classification model trained on the entire corpus may fail to capture characteristics unique to each user. On the other hand, a personalized model for each user may be inaccurate due to the scarcity of training data, especially at the very beginning where users have just posted a few entries. To overcome these challenges, we propose learning a global model over all micro-bloggers, which is then leveraged to continuously refine the individual models through a collaborative online learning way. We evaluate our algorithm on a real-life micro-blog dataset collected from the popular micro-blog site - Twitter. Results show that our algorithm is effective and efficient for timely sentiment detection in real micro-blogging applications.
Guangxia Li, Steven C. H. Hoi, Kuiyu Chang, Ramesh Jain 0001
ICDM1
2010 Two-View Transductive Support Vector Machines
abstract
Obtaining high-quality and up-to-date labeled data can be difficult in many real-world machine learning applications, especially for Internet classification tasks like review spam detection, which changes at a very brisk pace. For some problems, there may exist multiple perspectives, so called views, of each data sample. For example, in text classification, the typical view contains a large number of raw content features such as term frequency, while a second view may contain a small but highly-informative number of domain specific features. We thus propose a novel two-view transductive SVM that takes advantage of both the abundant amount of unlabeled data and their multiple representations to improve the performance of classifiers. The idea is fairly simple: train a classifier on each of the two views of both labeled and unlabeled data, and impose a global constraint that each classifier assigns the same class label to each labeled and unlabeled data. We applied our two-view transductive SVM to the WebKB course dataset, and a real-life review spam classification dataset. Experimental results show that our proposed approach performs up to 5% better than a single view learning algorithm, especially when the amount of labeled data is small. The other advantage of our two-view approach is its significantly improved stability, which is especially useful for noisy real world data.
Guangxia Li, Steven C. H. Hoi, Kuiyu Chang
SDM1