Minghua Xu 0003

dblp:33/2798-3 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
6since 2021 · last 2024
0009-0002-7746-2468ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Enhancing Distribution and Label Consistency for Graph Out-of-Distribution Generalization
abstract
To deal with distribution shifts in graph data, various graph out-of-distribution (OOD) generalization techniques have been recently proposed. These methods often employ a two-step strategy that first creates augmented environments and subsequently identifies invariant subgraphs to improve generalizability. Nevertheless, this approach could be suboptimal from the perspective of consistency. First, the process of augmenting environments by altering the graphs while preserving labels may lead to graphs that are not realistic or meaningfully related to the origin distribution, thus lacking distribution consistency. Second, the extracted subgraphs are obtained from directly modifying graphs, and may not necessarily maintain a consistent predictive relationship with their labels, thereby impacting label consistency. In response to these challenges, we introduce an innovative approach that aims to enhance these two types of consistency for graph OOD generalization. We propose a modifier to obtain both augmented and invariant graphs in a unified manner. With the augmented graphs, we enrich the training data without compromising the integrity of label-graph relationships. The label consistency enhancement in our framework further preserves the supervision information in the invariant graph. We conduct extensive experiments on real-world datasets to demonstrate the superiority of our framework over other state-of-the-art baselines.
Song Wang 0013, Rashidul Islam, Huiyuan Chen, Minghua Xu 0003, Jundong Li, Yiwei Cai
ICDM5
2024 Masked Graph Transformer for Large-Scale Recommendation
abstract
Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the scalability of Graph Transformers, particularly for large-scale recommendation. Here we propose an efficient Masked Graph Transformer, named MGFormer, capable of capturing all-pair interactions among nodes with a linear complexity. To achieve this, we treat all user/item nodes as independent tokens, enhance them with positional embeddings, and feed them into a kernelized attention module. Additionally, we incorporate learnable relative degree information to appropriately reweigh the attentions. Experimental results show the superior performance of our MGFormer, even with a single attention layer.
Huiyuan Chen, Zhe Xu 0007, Chin-Chia Michael Yeh, Vivian Lai, Yan Zheng 0001, Minghua Xu 0003, Hanghang Tong
SIGIR6
2024 Can One Embedding Fit All? A Multi-Interest Learning Paradigm Towards Improving User Interest Diversity Fairness
abstract
Recommender systems (RSs) have gained widespread applications across various domains owing to the superior ability to capture users' interests. However, the complexity and nuanced nature of users' interests, which span a wide range of diversity, pose a significant challenge in delivering fair recommendations. In practice, user preferences vary significantly; some users show a clear preference toward certain item categories, while others have a broad interest in diverse ones. Even though it is expected that all users should receive high-quality recommendations, the effectiveness of RSs in catering to this disparate interest diversity remains under-explored.
Yuying Zhao, Minghua Xu 0003, Huiyuan Chen, Yuzhong Chen 0004, Yiwei Cai, Rashidul Islam, Yu Wang 0160, Tyler Derr
WWW2
2023 A Plug-n-Play Framework for Scaling Private Set Intersection to Billion-Sized Sets
Saikrishna Badrinarayanan, Ranjit Kumaresan, Mihai Christodorescu, Vinjith Nagaraja, Karan Patel, Srinivasan Raghuraman, Peter Rindal, Minghua Xu 0003
CANS9
2023 From Trainable Negative Depth to Edge Heterophily in Graphs
abstract
Finding the proper depth $d$ of a graph convolutional network (GCN) that provides strong representation ability has drawn significant attention, yet nonetheless largely remains an open problem for the graph learning community. Although noteworthy progress has been made, the depth or the number of layers of a corresponding GCN is realized by a series of graph convolution operations, which naturally makes $d$ a positive integer ($d \in \mathbb{N}+$). An interesting question is whether breaking the constraint of $\mathbb{N}+$ by making $d$ a real number ($d \in \mathbb{R}$) can bring new insights into graph learning mechanisms. In this work, by redefining GCN's depth $d$ as a trainable parameter continuously adjustable within $(-\infty,+\infty)$, we open a new door of controlling its signal processing capability to model graph homophily/heterophily (nodes with similar/dissimilar labels/attributes tend to be inter-connected). A simple and powerful GCN model TEDGCN, is proposed to retain the simplicity of GCN and meanwhile automatically search for the optimal $d$ without the prior knowledge regarding whether the input graph is homophilic or heterophilic. Negative-valued $d$ intrinsically enables high-pass frequency filtering functionality via augmented topology for graph heterophily. Extensive experiments demonstrate the superiority of TEDGCN on node classification tasks for a variety of homophilic and heterophilic graphs.
Yuzhong Chen 0004, Huiyuan Chen, Minghua Xu 0003, Mahashweta Das, Hao Yang 0007, Hanghang Tong
NeurIPS4
2023 Enhancing Transformers without Self-supervised Learning: A Loss Landscape Perspective in Sequential Recommendation
abstract
Transformer and its variants are a powerful class of architectures for sequential recommendation, owing to their ability of capturing a user’s dynamic interests from their past interactions. Despite their success, Transformer-based models often require the optimization of a large number of parameters, making them difficult to train from sparse data in sequential recommendation. To address the problem of data sparsity, previous studies have utilized self-supervised learning to enhance Transformers, such as pre-training embeddings from item attributes or contrastive data augmentations. However, these approaches encounter several training issues, including initialization sensitivity, manual data augmentations, and large batch-size memory bottlenecks.
Vivian Lai, Huiyuan Chen, Chin-Chia Michael Yeh, Minghua Xu 0003, Yiwei Cai, Hao Yang 0007
RecSys4
2010 Filter Design and Analysis in Frequency Domain for Server Scheduling and Optimization
abstract
Internet traffic often exhibits a structure with rich high-order statistical properties like self-similarity and long-range dependency (LRD). This greatly complicates the problem of server performance modeling and optimization. Existing tools like queuing models in most cases only hold in mean value analysis under the assumption of simplified traffic structures. In this paper, we present a filter model to characterize the relationship among the factors of server capacity, request scheduling, and service quality for general input traffic. By the model, a server scheduler operates as an finite-duration impulse response (FIR) filter that transforms request processes into workload processes with the objective of minimizing load variation or overload probability, and meanwhile, without violating request response deadlines as defined in service-level agreements. We present a design and analysis of the filter for traffic with strong LRD in the frequency domain. Most Internet traffic has monotonically decreasing strength of variation functions over frequency. For this type of input traffic, we prove that optimal schedulers must have a convex structure. Uniform resource allocation is an extreme case of the convexity and is proved to be optimal for Poisson traffic. We integrate the convex structural principle with the Generalized Processor Sharing (GPS) discipline and show that the enhanced GPS policy improves the service quality significantly. Furthermore, we show that the presence of LRD in the input traffic results in shift of variation strength from high frequency to lower frequency bands and consequently leads to a degradation of the service quality.
Cheng-Zhong Xu 0001, Minghua Xu 0003, Le Yi Wang, Gang George Yin
IEEE Trans. Parallel Distributed Syst.2
2008 Frequency Domain Filter Design and Analysis of Request Scheduling in Internet Servers
abstract
Internet traffic has a characteristic of strong correlation. This traffic characteristic greatly complicates the problem of server performance modeling and optimization. Conventional time domain analysis has limitations in the study of the impact of complex traffic on server performance, because self-similarity of Internet traffic is often characterized in frequency domain. In this paper, we present a frequency domain filter model to characterize the relationship between server capacity, resource allocation, and service quality for general input traffic. Power spectral density (PSD) shows the strength of variations (power) as a function of frequency. By the model, server scheduler operates as a filter of input traffic that transforms its PSD function into another PSD function of server utilization process. The optimality of the scheduler in second-order statistics is to minimize the power leakage in the transformation. Most Internet traffic has monotonically decreasing PSD functions. For this type of input traffic, we prove that the optimal schedulers have a convex structure. Uniform allocation is an extreme case of the convexity and is proven to be optimal for traffic of independent arrivals. We integrate the convex-structured scheduling principle with GPS discipline and show that the enhanced GPS policy improves the service quality significantly.
Minghua Xu 0003, Cheng-Zhong Xu 0001
ICDCS1
2005 Optimal Time-Variant Resource Allocation for Internet Servers with Delay
abstract
The increasing popularity of high-volume performance-critical Internet applications is a challenge for servers to provide individual response-time guarantees. Considering the fact that most Internet applications can tolerate a small percentage of deadline misses, we define delay constraint as a statistical guarantee to relax server resource requirements. A recent decay function model characterizes the relationship between the request delay constraint, deadline misses, and server capacity in a transfer function based filter system. A time-invariant scheduler was proposed to minimize system load variances in support of requests with the same delay constraints. This paper extends the model to support requests with different deadlines and describes an optimal time-variant scheduling policy that minimizes load variances and capacity requirement. The resultant capacity bound is further tightened by utilizing the information of request arrival distribution. Simulation results validate the extended decay function model and show the superiority of the scheduler in comparison with other scheduling algorithms.
Xiliang Zhong, Cheng-Zhong Xu 0001, Minghua Xu 0003, Jianbin Wei
IEEE Real-Time and Embedded Technology and Applications Symposium3
2004 Decay function model for resource configuration and adaptive allocation on Internet servers
abstract
Server-side resource configuration and allocation for QoS guarantee is a challenge in performance-critical Internet applications. To overcome the difficulties caused by the high-variability and burstiness of Internet traffic, this paper presents a decay function model of request scheduling algorithms for the resource configuration and allocation problem. Under the decay function model, request scheduling is modelled as a transfer-function based filter system that has an input process of requests and an output process of server load. Unlike conventional queueing network models that rely on mean-value analysis for input renewal or Markovian processes, this decay function model works for general time-series based or measurement based processes and hence facilitates the study of statistical correlations between the request traffic, server load, and QoS of requests. Based on the model, we apply filter design theories in signal processing in the optimality analysis of various scheduling algorithms. We reveal a relationship between the server capacity, scheduling policy, service deadline, and other request properties in a formalism and present the optimality condition with respect to the second moments of request properties for an important class of fixed-time scheduling policies. Simulation results verify the relationship and show that optimal fixed-time scheduling can effectively reduce the server workload variance and guarantee service deadlines with high robustness on the Internet.
Minghua Xu 0003, Cheng-Zhong Xu 0001
IWQoS1