Ziji Shi

dblp:220/3155 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-9398-6507ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
Chaoyi Ruan, Chao Bi, Ziji Shi, Jialin Li 0001
NSDI4
2025 TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
abstract
Tensor parallelism is an essential technique for distributed training of large neural networks. However, automatically determining an optimal tensor parallel strategy is challenging due to the gigantic search space, which grows exponentially with model size and tensor dimension. This prohibits the adoption of auto-parallel systems on larger models.
Ziji Shi, Ang Wang, Jie Zhang 0135, Chencan Wu, Yong Li 0045, Xiaokui Xiao, Wei Lin 0016, Jialin Li 0001
ICPP1
2024 ParaGAN: A Scalable Distributed Training Framework for Generative Adversarial Networks
abstract
Recent advances in Generative Artificial Intelligence have fueled numerous applications, particularly those involving Generative Adversarial Networks (GANs), which are essential for synthesizing realistic photos and videos. However, efficiently training GANs remains a critical challenge due to their computationally intensive and numerically unstable nature. Existing methods often require days or even weeks for training, posing significant resource and time constraints.
Ziji Shi, Jialin Li 0001, Yang You 0001
SoCC1
2022 Go Wider Instead of Deeper
abstract
More transformer blocks with residual connections have recently achieved impressive results on various tasks. To achieve better performance with fewer trainable parameters, recent methods are proposed to go shallower by parameter sharing or model compressing along with the depth. However, weak modeling capacity limits their performance. Contrastively, going wider by inducing more trainable matrixes and parameters would produce a huge model requiring advanced parallelism to train and inference. In this paper, we propose a parameter-efficient framework, going wider instead of deeper. Specially, following existing works, we adapt parameter sharing to compress along depth. But, such deployment would limit the performance. To maximize modeling capacity, we scale along model width by replacing feed-forward network (FFN) with mixture-of-experts (MoE). Across transformer blocks, instead of sharing normalization layers, we propose to use individual layernorms to transform various semantic representations in a more parameter-efficient way. To evaluate our plug-and-run framework, we design WideNet and conduct comprehensive experiments on popular computer vision and natural language processing benchmarks. On ImageNet-1K, our best model outperforms Vision Transformer (ViT) by 1.5% with 0.72 times trainable parameters. Using 0.46 times and 0.13 times parameters, our WideNet can still surpass ViT and ViT-MoE by 0.8% and 2.1%, respectively. On four natural language processing datasets, WideNet outperforms ALBERT by 1.8% on average and surpass BERT using factorized embedding parameterization by 0.8% with fewer parameters.
Fuzhao Xue, Ziji Shi, Futao Wei, Yuxuan Lou, Yong Liu 0020, Yang You 0001
AAAI2
2022 Whale: Efficient Giant Model Training over Heterogeneous GPUs
Xianyan Jia, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang 0135, Langshi Chen, Yong Li 0045, Zhen Zheng, Wei Lin 0016
USENIX ATC5
2020 Domain Adaptation for Degraded Remote Scene Classification
abstract
Remote scene classification serves a vital role in many applications. However, satellite images are often blurred and degraded due to aerosol scattering under fog, haze, and other weather conditions, reducing the image contrast and color fidelity. State-of-the-art remote sensing classification models building upon convolutional neural networks (CNNs) are mostly trained on annotated datasets of clear satellite images. When applied to blurred images, they will suffer a great degradation in performance. To address this problem, we adopt the domain adaptation algorithm TADA and propose Transferable Attention enhanced Adversarial Adaptation Network (TA3N), which utilizes annotated data in clear images by applying knowledge transferring from clear image domain to blurred image domain. Our TA3N first integrates spatial attention to focus on salient areas which are discriminative and transferable. In addition, domain discriminator and adversarial training via gradient reversal layer are used to minimize the discrepancies in extracted features from clear and degraded domains. We synthesize degraded remote scene classification dataset SSI based on FoHIS model. Experiments on degraded SSI showed that TA3N significantly outperforms baseline and other state-of-the-art domain adaptation methods.
Jianfei Yang 0001, Hailin Chen, Yuecong Xu, Ziji Shi, Ruikang Luo, Lihua Xie 0001, Rong Su 0001
ICARCV4