Kaiqiang Xu

dblp:54/2203 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TurboBus: Pooling PCIe Bandwidth for LLM Workloads via Scale-Up Fabrics
Kaiqiang Xu, Kai Chen 0005
SIGCOMM2
2026 Reconstructing shared visual experiences from human brain activity across individuals
Yanyan Huang, Kaiqiang Xu, Yannan Chen, Lequan Yu, Zhijun Yao, Yu Fu 0008
Medical Image Anal.4
2025 Design and Operation of Shared Machine Learning Clusters on Campus
abstract
The rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems.
Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005
ASPLOS (1)1
2025 GREEN: Carbon-efficient Resource Scheduling for Machine Learning Clusters
Kaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 0001, Kai Chen 0005
NSDI1
2025 Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005
OSDI6
2025 Coflow Scheduling for LLM Training
abstract
Training large language models (LLMs) generates diverse coflows within a cluster, requiring optimized scheduling to enhance communication-computation overlap and minimize training time. Existing schedulers inadequately handle contention both across and within coflows, resulting in suboptimal performance.
Xinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin, Yijun Sun, Zhenghang Ren, Han Tian, Kai Chen 0005
SIGCOMM3
2025 Adversarial perturbation and defense for generalizable person re-identification
Hongchen Tan, Kaiqiang Xu, Pingping Tao, Xiuping Liu
Neural Networks2
2025 Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed Data
abstract
Privacy-preserving machine learning (PPML) algorithms use secure computation protocols to allow multiple data parties to collaboratively train machine learning (ML) models while maintaining their data confidentiality. However, current PPML frameworks couple secure protocols with ML models in PPML algorithm implementations, making it challenging for non-experts to develop and optimize PPML applications, limiting their accessibility and performance. We propose Sequoia, a novel PPML framework that decouples ML models and secure protocols to optimize the development and execution of PPML applications across data parties. Sequoia offers JAX-compatible APIs for users to program their ML models, while using a compiler-executor architecture to automatically apply PPML algorithms and system optimizations for model execution over distributed data. The compiler in Sequoia incorporates cross-party PPML processes into user-defined ML models by transparently adding computation, encryption, and communication steps with extensible policies, and the executor efficiently schedules code execution across multiple data parties, considering data dependencies and device heterogeneity. Compared to existing PPML frameworks, Sequoia requires 64%-92% fewer lines of code for users to implement the same PPML algorithms, and achieves 88% speedup of training throughput in horizontal PPML.
Kaiqiang Xu, Di Chai, Junxue Zhang 0001, Fan Lai 0001, Kai Chen 0005
Proc. ACM Manag. Data1
2024 Multi-view deep subspace clustering via level-by-level guided multi-level features learning
Kaiqiang Xu, Kewei Tang, Zhixun Su
Appl. Intell.1
2024 Clean and robust multi-level subspace representations learning for deep multi-view subspace clustering
Kaiqiang Xu, Kewei Tang, Zhixun Su, Hongchen Tan
Expert Syst. Appl.1
2024 Attention-Bridged Modal Interaction for Text-to-Image Generation
abstract
We propose a novel Text-to-Image Generation Network, Attention-bridged Modal Interaction Generative Adversarial Network (AMI-GAN), to better explore modal interaction and perception for high-quality image synthesis. The AMI-GAN contains two novel designs: an Attention-bridged Modal Interaction (AMI) module and a Residual Perception Discriminator (RPD). In AMI, we mainly design a multi-scale attention mechanism to exploit semantics alignment, fusion, and enhancement between text and image, to better refine details and context semantics of the synthesized image. In RPD, we design a multi-scale information perception mechanism with our proposed novel information adjustment function, to encourage the discriminator to better perceive visual differences between the real and synthesized image. Consequently, the discriminator will drive the generator to improve the visual quality of the synthesized image. Besides, based on these novel designs, we can design two versions, a single-stage generation framework (AMI-GAN-S), and a multi-stage generation framework (AMI-GAN-M), respectively. The former can synthesize high-resolution images because of its low computational cost; the latter can synthesize images with realistic detail. Experimental results on two widely used T2I datasets showed that our AMI-GANs achieve competitive performance in T2I task.
Hongchen Tan, Kaiqiang Xu, Huasheng Wang, Xiuping Liu, Xin Li 0003
IEEE Trans. Circuits Syst. Video Technol.3
2023 Multi-view subspace clustering via consistent and diverse deep latent representations
Kewei Tang, Kaiqiang Xu, Zhixun Su, Nan Zhang 0014
Inf. Sci.2
2023 Deep multi-view subspace clustering via structure-preserved multi-scale features fusion
Kaiqiang Xu, Kewei Tang, Zhixun Su
Neural Comput. Appl.1
2023 Scalable and Efficient Full-Graph GNN Training for Large Graphs
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools to capture structural information from graph-structured data, achieving state-of-the-art performance on applications such as recommendation, knowledge graph, and search. Graphs in these domains typically contain hundreds of millions of nodes and billions of edges. However, previous GNN systems demonstrate poor scalability because large and interleaved computation dependencies in GNN training cause significant overhead in current parallelization methods. We present G3, a distributed system that can efficiently train GNNs over billion-edge graphs at scale. G3 introduces GNN hybrid parallelism which synthesizes three dimensions of parallelism to scale out GNN training by sharing intermediate results peer-to-peer in fine granularity, eliminating layer-wise barriers for global collective communication or neighbor replications as seen in prior works. G3 leverages locality-aware iterative partitioning and multi-level pipeline scheduling to exploit acceleration opportunities by distributing balanced workload among workers and overlapping computation with communication in both inter-layer and intra-layer training processes. We show via a prototype implementation and comprehensive experiments that G3 can achieve as much as 2.24x speedup in a 16-node cluster, and better final accuracy over prior works.
Xinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin, Kai Chen 0005, Xin Jin 0008
Proc. ACM Manag. Data2
2023 Selecting the Best Part From Multiple Laplacian Autoencoders for Multi-View Subspace Clustering
abstract
The multi-view subspace clustering attracts much attention in recent years. Most methods follow the framework of fusing the affinity graph learned in each view. In this framework, both the fusion strategy and built graph of each view are very important. In this paper, we propose novel methods for multi-view subspace clustering to address these two aspects. On the one hand, we adopt the autoencoders with Laplacian regularization to construct the affinity graph in each view. Compared with previous work employing the autoencoders, the Laplacian term in our method can guide the learned latent representation favoring affinity extraction. Besides, we also discuss the reasons for adding Laplacian regularization. On the other hand, we propose a novel fusion strategy distinguished from the related literature. If the affinity graph of some view is not extracted well, the performance of previous fusion strategies will be seriously affected. Since our strategy can choose the best part from each affinity graph, it can overcome this limitation to some extent. Extensive experimental results on multiple benchmark data sets confirm the effectiveness of our method.
Kewei Tang, Kaiqiang Xu, Wei Jiang 0007, Zhixun Su, Xiyan Sun
IEEE Trans. Knowl. Data Eng.2
2020 Gated Convolutional Networks with Hybrid Connectivity for Image Classification
abstract
We propose a simple yet effective method to reduce the redundancy of DenseNet by substantially decreasing the number of stacked modules by replacing the original bottleneck by our SMG module, which is augmented by local residual. Furthermore, SMG module is equipped with an efficient two-stage pipeline, which aims to DenseNet-like architectures that need to integrate all previous outputs, i.e., squeezing the incoming informative but redundant features gradually by hierarchical convolutions as a hourglass shape and then exciting it by multi-kernel depthwise convolutions, the output of which would be compact and hold more informative multi-scale features. We further develop a forget and an update gate by introducing the popular attention modules to implement the effective fusion instead of a simple addition between reused and new features. Due to the Hybrid Connectivity (nested combination of global dense and local residual) and Gated mechanisms, we called our network as the HCGNet. Experimental results on CIFAR and ImageNet datasets show that HCGNet is more prominently efficient than DenseNet, and can also significantly outperform state-of-the-art networks with less complexity. Moreover, HCGNet also shows the remarkable interpretability and robustness by network dissection and adversarial defense, respectively. On MS-COCO, HCGNet can consistently learn better features than popular backbones.
Chuanguang Yang, Zhulin An, Hui Zhu 0002, Kun Zhang 0045, Kaiqiang Xu, Chao Li 0028, Yongjun Xu 0001
AAAI6
2020 DRNet: Dissect and Reconstruct the Convolutional Neural Network via Interpretable Manners
abstract
Convolutional neural networks (ConvNets) are widely used in real life. People usually use ConvNets which pre-trained on a fixed number of classes. However, for different application scenarios, we usually do not need all of the classes, which means ConvNets are redundant when dealing with these tasks. This paper focuses on the redundancy of ConvNet channels. We proposed a novel idea: using an interpretable manner to find the most important channels for every single class (dissect), and dynamically run channels according to classes in need (reconstruct). For VGG16 pre-trained on CIFAR-10, we only run 11\% parameters for two-classes sub-tasks on average with negligible accuracy loss. For VGG16 pre-trained on ImageNet, our method averagely gains 14.29\% accuracy promotion for two-classes sub-tasks. In addition, analysis show that our method captures some semantic meanings of channels, and uses the context information more targeted for sub-tasks of ConvNets.
Zhulin An, Chuanguang Yang, Hui Zhu 0002, Kaiqiang Xu, Yongjun Xu 0001
ECAI5
2020 Towards More Efficient And Effective Inference: The Joint Decision Of Multi-Participants
abstract
Existing approaches to improve the performances of convolutional neural networks by optimizing the local architectures or deepening the networks tend to increase the size of models significantly. In order to deploy and apply the neural networks to edge devices which are in great demand, reducing the scale of networks is quite crucial. However, It is easy to degrade the performance of image processing by compressing the networks. In this paper, we propose a method which is suitable for edge devices while improving the efficiency and effectiveness of inference. The joint decision of multiparticipants, mainly contain multi-layers and multi-networks, can achieve higher classification accuracy (0.26% on CFAR-10 and 4.49% on CFAR-100 at most) with similar total number of parameters for classical convolutional neural networks.
Hui Zhu 0002, Zhulin An, Kaiqiang Xu, Yongjun Xu 0001
ICIP3
2020 HLNet: Modeling High and Low Frequencies for Scene Parsing
abstract
In this paper we propose to model high and low frequencies of segmentation map, based on the observation that the map can be seen as a mixture of different frequencies. Based on the sparsity of high frequencies and local similarity of low frequencies, we design special building blocks and further a novel High and Low frequency Network (HLNet) with two branches based on FCN to predict high and low frequencies of the segmentation map, respectively. Specifically, we design a high frequency branch with a small kernel size and high-resolution features to predict a sparse high frequency component. Mean-while, a low frequency branch with similarity computing and low-resolution features is employed to predict a low frequency component. On top of two branches, we combine two different frequency components to generate final result for scene parsing. We empirically demonstrate that the designed model achieves superior performance 44.07% on ADE20K, and 80.14% mIoU on Cityscapes datasets.
Kaiqiang Xu, Zhulin An, Hui Zhu 0002, Yongjun Xu 0001
IJCNN1
2020 Efficient Search for the Number of Channels for Convolutional Neural Networks
abstract
Latest algorithms for automatic neural architecture search perform remarkably but few of them can effectively design the number of channels for convolutional neural networks and consume less computational efforts. In this paper, we propose a method for efficient automatic search which is special to the widths of networks instead of the connections within neural architectures. Our method, functionally incremental search based on function-preserving, will explore the number of channels for almost any convolutional neural network rapidly while controlling the number of parameters and even the amount of computations (FLOPs). On CIFAR-10 and CIFAR-100 classification, our method using minimal computational resources (0.41 ~ 1.29 GPU-days) can discover more effective rules of the widths of networks to improve the accuracy (a ~ 1.08 on CIFAR-10 and b ~ 2.33 on CIFAR-100) with fewer number of parameters.
Hui Zhu 0002, Zhulin An, Chuanguang Yang, Kaiqiang Xu, Yongjun Xu 0001
IJCNN5
2017 Retinex-based perceptual contrast enhancement in images using luminance adaptation
abstract
In this paper, we propose retinex-based perceptual contrast enhancement in images using luminance adaptation. We use the retinex theory to decompose an image into illumination and reflectance layers, and adopt luminance adaptation to handle the illumination layer which causes detail loss. First, we obtain the illumination layer using adaptive Gaussian filtering to remove halo artifacts. Then, we adaptively remove illumination of the illumination layer in the multi-scale retinex (MSR) process based on luminance adaptation to preserve details. Finally, we perform contrast enhancement on the MSR result. Experimental results demonstrate that the proposed method successfully enhances contrast in images while keeping textures in highlight regions.
Kaiqiang Xu, Cheolkon Jung
ICASSP1
2017 Naturalness-preserved tone mapping in images based on perceptual quantization
abstract
In this paper, we propose naturalness-preserved tone mapping in images based on perceptual quantization (PQ). PQ is a transfer function based on Barten's contrast sensitive function (CSF) which represents human visual perception on luminance, and we adopt it to generate a limit curve for perceptual contrast enhancement. First, we obtain a limit curve in an image based on PQ transfer function to adjust the degree of contrast enhancement. Second, we redistribute the histogram using the limit curve and achieve perceptual contrast enhancement. Experimental results demonstrate that the proposed method effectively enhances contrast in images while successfully preserving naturalness.
Cheolkon Jung, Kaiqiang Xu
ICIP2