Mingqing Hu

dblp:66/2666 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
4since 2021 · last 2024
0009-0007-5528-6869ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A Multi-Node Multi-GPU Distributed GNN Training Framework for Large-Scale Online Advertising
abstract
Graph Neural Networks (GNNs) have become critical in various domains such as online advertising but face scalability challenges due to the growing size of graph data, leading to the needs for advanced distributed GPU computation strategies across multiple nodes. This paper presents PGLBox-Cluster, a robust distributed graph learning framework constructed atop the PaddlePaddle platform, implemented to efficiently process graphs comprising billions of nodes and edges. Through strategic partitioning of the model, node attributes, and graph data and leveraging industrial-grade RPC and NCCL for communication, PGLBox-Cluster facilitates effective distributed computation. The extensive experimental results confirm that PGLBox-Cluster achieves a 1.94x to 2.93x speedup over the single-node configuration, significantly elevating graph neural network scalability and efficiency by handling datasets exceeding 3 billion nodes and 120 billion edges with its novel asynchronous communication and graph partitioning techniques. The repository is released at This Link.
Xuewu Jiao, Xinsheng Luo, Jiang Bian 0003, Junchao Yang 0001, Mingqing Hu, Weipeng Lu, Shikun Feng, Danlei Feng, Haoyi Xiong, Shuanglong Li
CIKM7
2023 PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation
abstract
While having been used widely for large-scale recommendation and online advertising, the Graph Neural Network (GNN) has demonstrated its representation learning capacity to extract embeddings of nodes and edges through passing, transforming, and aggregating information over the graph. In this work, we propose PGLBox1 - a multi-GPU graph learning framework based on PaddlePaddle [24], incorporating with optimized storage, computation, and communication strategies, to train deep GNNs based on web-scale graphs for the recommendation. Specifically, PGLBox adopts a hierarchical storage system with three layers to facilitate I/O, where graphs and embeddings are stored in the HBMs and SSDs, respectively, with MEMs as the cache. To fully utilize multi-GPUs and I/O bandwidth, PGLBox proposes an asynchronous pipeline with three stages - it first samples the subgraphs from the input graph, then pulls & updates embeddings and trains GNNs on the subgraph with parameters updating queued at the end of the pipeline. Thanks to the capacity of PGLBox in handling web-scale graphs, it becomes feasible to unify the view of GNN-based recommendation tasks for multiple advertising verticals and fuse all these graphs into a unified yet huge one. We evaluate PGLBox using a bucket of realistic GNN training tasks for the recommendation, and compare the performance of PGLBox on top of a multi-GPU server (Tesla A100×8) and the legacy training system based on a 40-node MPI cluster at Baidu. The overall comparisons show that PGLBox could save up to 55% monetary cost for training GNN models, and achieve up to 14× training speedup with the same accuracy as the legacy trainer. The open-source implementation of PGLBox is available at https://github.com/PaddlePaddle/PGL/tree/main/apps/PGLBox.
Xuewu Jiao, Weibin Li 0004, Xinxuan Wu, Jiang Bian 0003, Siming Dai, Xinsheng Luo, Mingqing Hu, Zhengjie Huang, Danlei Feng, Junchao Yang 0001, Shikun Feng, Haoyi Xiong, Dianhai Yu, Shuanglong Li, Jingzhou He, Yanjun Ma
KDD9
2022 PaddleBox: Communication-Efficient TeraByte-Scale Model Training Framework for Online Advertising
abstract
Click-through rate (CTR) prediction is one of the most crucial components in the online advertising industry. In order to produce a personalized CTR prediction, an industry-level CTR prediction model commonly takes a high-dimensional (∼ 1012) sparse vector (that is encoded from query keywords, user portraits, etc.) as input. As a result, the model requires Terabyte scale parameters to embed the high-dimensional input. Hierarchical distributed GPU parameter server has been developed at Baidu to enable GPU with limited memory to train the massive network by leveraging CPU main memory and SSDs as secondary storage. In this work, we identify two major challenges in the existing GPU training framework for massive-scale ad models and propose a collection of optimizations to tackle these challenges: (a) the GPU, CPU, SSD rapidly communicate with each other during the training. The connections between GPUs and CPUs are non-uniform due to the hardware topology. The data communication route should be optimized according to the hardware topology; (b) GPUs in different computing nodes frequently communicates to synchronize parameters. It is thus required to optimize the communications so that the distributed system can become scalable. In this paper, we propose a hardware-aware training workflow that couples the hardware topology into the algorithm design. To reduce the extensive communication between computing nodes, we introduce a k-step model merging algorithm for Adam and provide its convergence rate in non-convex optimization. To the best of our knowledge, this is the first application of k-step adaptive optimization method in industrial CTR model training. Experiments on commercial search ads data confirm the effectiveness of our proposed training framework.
Weijie Zhao 0001, Xuewu Jiao, Mingqing Hu, Ping Li 0001
IEEE Big Data3
2022 Prevention of GAN-Based Privacy Inferring Attacks Towards Federated Learning
Hongbo Cao, Yongsheng Zhu, Yuange Ren, Bin Wang 0062, Mingqing Hu, Wanqi Wang, Wei Wang 0012
CollaborateCom (2)5
2015 Semantics-preserving hashing for cross-view retrieval
abstract
With benefits of low storage costs and high query speeds, hashing methods are widely researched for efficiently retrieving large-scale data, which commonly contains multiple views, e.g. a news report with images, videos and texts. In this paper, we study the problem of cross-view retrieval and propose an effective Semantics-Preserving Hashing method, termed SePH. Given semantic affinities of training data as supervised information, SePH transforms them into a probability distribution and approximates it with to-be-learnt hash codes in Hamming space via minimizing the Kullback-Leibler divergence. Then kernel logistic regression with a sampling strategy is utilized to learn the nonlinear projections from features in each view to the learnt hash codes. And for any unseen instance, predicted hash codes and their corresponding output probabilities from observed views are utilized to determine its unified hash code, using a novel probabilistic approach. Extensive experiments conducted on three benchmark datasets well demonstrate the effectiveness and reasonableness of SePH.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001
CVPR3
2015 Image auto-annotation via tag-dependent random search over range-constrained visual neighbours
Zijia Lin, Guiguang Ding, Mingqing Hu
Multim. Tools Appl.3
2014 Multi-label Classification via Feature-aware Implicit Label Space Encoding
abstract
To tackle a multi-label classification problem with many classes, recently label space dimension reduction (LSDR) is proposed. It encodes the original label space to a low-dimensional latent space and uses a decoding process for recovery. In this paper, we propose a novel method termed FaIE to perform LSDR via Feature-aware Implicit label space Encoding. Unlike most previous work, the proposed FaIE makes no assumptions about the encoding process and directly learns a code matrix, i.e. the encoding result of some implicit encoding function, and a linear decoding matrix. To learn both matrices, FaIE jointly maximizes the recoverability of the original label space from the latent space, and the predictability of the latent space from the feature space, thus making itself feature-aware. FaIE can also be specified to learn an explicit encoding function, and extended with kernel tricks to handle non-linear correlations between the feature space and the latent space. Extensive experiments conducted on benchmark datasets well demonstrate its effectiveness.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001
ICML3
2014 Image tag completion via dual-view linear sparse reconstructions
Zijia Lin, Guiguang Ding, Mingqing Hu, Yunzhen Lin, Shuzhi Sam Ge
Comput. Vis. Image Underst.3
2013 Image Tag Completion via Image-Specific and Tag-Specific Linear Sparse Reconstructions
abstract
Though widely utilized for facilitating image management, user-provided image tags are usually incomplete and insufficient to describe the whole semantic content of corresponding images, resulting in performance degradations in tag-dependent applications and thus necessitating effective tag completion methods. In this paper, we propose a novel scheme denoted as LSR for automatic image tag completion via image-specific and tag-specific Linear Sparse Reconstructions. Given an incomplete initial tagging matrix with each row representing an image and each column representing a tag, LSR optimally reconstructs each image (i.e. row) and each tag (i.e. column) with remaining ones under constraints of sparsity, considering image-image similarity, image-tag association and tag-tag concurrence. Then both image-specific and tag-specific reconstruction values are normalized and merged for selecting missing related tags. Extensive experiments conducted on both benchmark dataset and web images well demonstrate the effectiveness of the proposed LSR.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001, Xiaojun Ye 0001
CVPR3
2013 Multi-source image auto-annotation
abstract
Though the field of image auto-annotation has been extensively researched, most previous work concentrated on the single-source problem, assuming that both labelled and unseen to-be-annotated images are from a single source (e.g. an identical website), while in practice they are generally collected from multiple sources (e.g. different websites). In that case, treating each source independently may suffer from the insufficiency of labelled data for model training, while merging with labelled images from other sources can bring risky biases to the source-specific model. In this paper, we propose a multi-task learning model to alleviate the multi-source image auto-annotation problem, with each task defined as performing auto-annotation for the corresponding source. Specifically, the proposed model trains annotation models for all sources in parallel with the introduction of inter-source structure regularizers and parameter constraints for sharing information and enhancing the overall performance. Experiments conducted on three different-source benchmark datasets and their combinations yield inspiring results and demonstrate that the proposed model can well utilize the shared information and relieve the risky biases.
Zijia Lin, Guiguang Ding, Mingqing Hu
ICIP3
2012 Automatic image annotation using tag-related random search over visual neighbors
abstract
In this paper, we propose a novel image auto-annotation model using tag-related random search over range-constrained visual neighbors of the to-be-annotated image. The proposed model, termed as TagSearcher, observes that the annotating performances of many previous visual-neighbor-based models are generally sensitive to the quantity setting of visual neighbors, and the probabilities for visual neighbors to be selected is better to be tag-dependent, meaning that each candidate tag can have its own trustworthy part of visual neighbors for score prediction. And thus TagSearcher uses a constrained range rather than an identical and fixed number of visual neighbors for auto-annotation. By performing a novel tag-related random search process over the graphical model made up of range-constrained visual neighbors, TagSearcher can find the trustworthy part for each candidate tag, and further utilize both visual similarities and tag correlations for score prediction. With the range constraint for visual neighbors and the tag-related random search process, TagSearcher can not only achieve satisfactory annotating performances, but also reduce the performance sensitivity. Experiments conducted on benchmark Corel5k well demonstrate its rationality and effectiveness.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001, Jia-Guang Sun 0001
CIKM3
2009 Building Sparse Multiple-Kernel SVM Classifiers
abstract
The support vector machines (SVMs) have been very successful in many machine learning problems. However, they can be slow during testing because of the possibly large number of support vectors obtained. Recently, Wu (2005) proposed a sparse formulation that restricts the SVM to use a small number of expansion vectors. In this paper, we further extend this idea by integrating with techniques from multiple-kernel learning (MKL). The kernel function in this sparse SVM formulation no longer needs to be fixed but can be automatically learned as a linear combination of kernels. Two formulations of such sparse multiple-kernel classifiers are proposed. The first one is based on a convex combination of the given base kernels, while the second one uses a convex combination of the so-called "equivalent" kernels. Empirically, the second formulation is particularly competitive. Experiments on a large number of toy and real-world data sets show that the resultant classifier is compact and accurate, and can also be easily trained by simply alternating linear program and standard SVM solver.
Mingqing Hu, Yiqiang Chen 0001, James T. Kwok
IEEE Trans. Neural Networks1
2008 WiFi-Based Power Aware Pervasive Device
abstract
In this paper, we propose a new kind of WiFi-based embedded device and two power efficient killer applications on it. First, the definition of our pervasive device and the key challenge research issues on it is introduced. Then we specially present our work on it, including hardware and software part. For hardware part, we will show our new Loongson SOC (system on a chip) chip based hardware architecture, which is power efficient and flexible connective interface one. For software part, we focus on two killer applications for this kind of pervasive device: location estimation and video codec. The power aware method we used in above two applications will be introduced in detail.
Yiqiang Chen 0001, Mingqing Hu, Qingsheng Yuan, Junfa Liu
PerCom2
2006 An Improved Reduced Set Method to Control the Run-time Complexity of SVM in Wireless Sensor Networks
abstract
One prominent disadvantage of SVM when implemented in wireless sensor networks (WSNs) is the run-time complexity of classifier, which linearly increases with the number of support vectors (SVs). This disadvantage prevents applying SVM in some applications. In this paper, we propose an improved reduced set method to find solutions characterized by few number of vectors and having good generalization properties. The idea behind our improved method is to combine finding patterns with maximum absolute margin and performing gradient-descent to find new patterns in new decision function. Our method can partially overcome the non-convexity difficulty. The application context is that of WSNs, where a general sensor node is equipped with fixed point CPU. The performance of fixed point implementation of our algorithm is also provided.
Mingqing Hu, Andrea Boni
ETFA1