EDBT 2026 Demo / reviewers in the wild / expert
Shixun Zhang
dblp:78/9671
· DBLP profile ↗
6ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-9487-4697ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFTabstractParameter-Efficient Fine-Tuning (PEFT) methods are crucial for adapting large pre-trained models. Among these, LoRA is considered a foundational approach. Building on this, the influential DoRA method enhances performance by decomposing weight updates into magnitude and direction. However, its underlying mechanism remains unclear, and it introduces significant computational overhead. In this work, we first identify that DoRA's success stems from its capacity to increase the singular value entropy of the weight update matrix, which promotes a more uniform update distribution akin to full fine-tuning. We then reformulate DoRA into a mathematically equivalent and more efficient matrix form, revealing it as a learnable weight conditioning method. Based on this insight, we propose a unified framework for designing advanced PEFT methods by exploring two orthogonal dimensions: the architectural placement and the transformation type of the conditioning matrix. Within this framework, we introduce two novel methods: (1) Pre-Diag, which applies a diagonal conditioning matrix before the LoRA update to efficiently calibrate the pre-trained weights, thereby enhancing performance while reducing training time; and (2) Skewed Orthogonal Rotation Adaptation (SORA), which employs a parameter-efficient orthogonal rotation to perform a more powerful, norm-preserving transformation of the feature space. Extensive experiments on natural language understanding and generation tasks demonstrate that our proposed methods achieve superior performance and efficiency compared to both LoRA and DoRA. Da Chang, Yu Li 0036, Yongxiang Liu, Pengxiang Xu, Shixun Zhang |
AAAI | 6 |
| 2026 | Synchronized Dual-Ring: A Synergistic Algorithm for Bandwidth-Efficient Collective Communication Leveraging Die-to-Die Direct Interconnects
Dongxiang Zhang, Erkun Zhang, Shixun Zhang, Bingqiang Wang, Fangjiong Chen |
IPDPS | 4 |
| 2025 | Improving the Energy Efficiency of AI Clusters Through Variability-Aware Frequency Scaling and Task Allocation
Dongxiang Zhang, Bingqiang Wang, Qiang Wang 0060, Shixun Zhang |
ICA3PP (6) | 5 |
| 2025 | AUE: A Normalized Energy Efficiency Metric for AI Servers Under LLM WorkloadsabstractUnder the rapid advancement of large model-driven artificial intelligence, the surging energy consumption of AI training and inference tasks has created an urgent need for precise and comparable energy efficiency metrics to guide the design and deployment of green computing systems. While existing metrics such as PUE and Green500 metrics focus on infrastructure or traditional numerical computations, they cannot reflect the characteristics of AI workloads. Although applicationoriented metrics like J/response and J/token are designed for LLMs, they remain susceptible to biases induced by model scale and output strategies, lacking cross-model comparability. This paper proposes a novel AI energy efficiency metric, AUE, defined as the energy consumed per thousand tokens per billion activated parameters. By normalizing model size effects, AUE accurately reflects the energy efficiency of underlying computational resources. We theoretically justify the validity of AUE and conduct experiments on a server equipped with$4 \times$Ascend NPU 910C accelerators, evaluating dense Transformer and MoE architectures across both training and inference workloads. Experimental results demonstrate that traditional J/token metrics disproportionately favor smaller models, whereas AUE reveals true energy utilization efficiency. For instance, while Qwen3 0.6B shows superior J/token values compared to Qwen3 14B, the 14B model achieves a significantly better AUE of 12.88 J/(KToken GParam) versus 24.35 J/(KToken GParam) for the 0.6 B model, consistent with measured FLOPs where the 14B model outperforms its smaller counterpart. With advantages including simple measurement procedures and compatibility across platforms and model architectures, AUE provides a viable pathway toward establishing a normalized and standardized AI energy efficiency evaluation framework. Dongxiang Zhang, Qiang Wang 0060, Bingqiang Wang, Shixun Zhang, Yonghong Tian 0001 |
ICPADS | 6 |
| 2020 | FMore: An Incentive Scheme of Multi-dimensional Auction for Federated Learning in MECabstractPromising federated learning coupled with Mobile Edge Computing (MEC) is considered as one of the most promising solutions to the AI-driven service provision. Plenty of studies focus on federated learning from the performance and security aspects, but they neglect the incentive mechanism. In MEC, edge nodes would not like to voluntarily participate in learning, and they differ in the provision of multi-dimensional resources, both of which might deteriorate the performance of federated learning. Also, lightweight schemes appeal to edge nodes in MEC. These features require the incentive mechanism to be well designed for MEC. In this paper, we present an incentive mechanism FMore with multi-dimensional procurement auction of K winners. Our proposal FMore not only is lightweight and incentive compatible, but also encourages more high-quality edge nodes with low cost to participate in learning and eventually improve the performance of federated learning. We also present theoretical results of Nash equilibrium strategy to edge nodes and employ the expected utility theory to provide guidance to the aggregator. Both extensive simulations and real-world experiments demonstrate that the proposed scheme can effectively reduce the training rounds and drastically improve the model accuracy for challenging AI tasks. Rongfei Zeng, Shixun Zhang, Xiaowen Chu 0001 |
ICDCS | 2 |
| 2013 | Exploiting Execution Order and Parallelism from Processing Flow Applying Pipeline-Based Programming Method on Manycore AcceleratorsabstractMany core architecture promotes a massively parallel computing on the accelerators. Especially GPU is one of the main series of the high performance computing, which is also employed by top supercomputers in the world. However, the programming method on such accelerators needs the double programming, in which the programmer needs to develop a control program executed on the CPU side to schedule the invocation of the accelerator's kernel program. Moreover, the programmer needs to consider the stream computing paradigm. To overcome the difficulty, the author of this paper has proposed and implemented a command line-based programming tool called CarSh that eliminates to develop the CPU program from the programmer. Using the CarSh, it is available to implement a GUI-based programming tool for the accelerators that visualizes a pipeline-based processing flow by connecting the kernel programs via the I/O data streams. In the case of applying the GUI-based programming, it is very hard to find the starting point of a complex processing flow. Moreover, although the processing pipeline should include the potential parallelism, it is hard for the programmer to exploit it intuitively. This paper proposes an algorithm that addresses those difficulties, and also evaluates the performance aspect using CarSh environment. Shinichi Yamagiwa, Ryo Jozaki, Shixun Zhang, Ryo Zaizen, Dewen Xu |
ICPP | 3 |