Jianing Zheng

dblp:145/0054 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RheoSparse: Exploring Finer Grained Structured Sparsity for Small Language Models
abstract
Small Language Models (SLMs) are designed for efficient on-device deployment, but compressing them without significant accuracy loss remains challenging. Structured sparsity methods like N:M pruning, which removes N parameters out of every M, are widely used to improve hardware efficiency. However, their rigid patterns, such as the commonly adopted 4:8 format, often degrade SLM performance, since these models are more sensitive to parameter removal than larger ones. We observe that finer-grained patterns like N:64 can better preserve accuracy under the same sparsity budget, yet current inference systems do not efficiently support them, especially during token generation. Furthermore, applying such fine-grained sparsity uniformly across all layers is suboptimal, as different layers respond differently to pruning. To address this, we propose RheoSparse. First, we use coarse-to-fine evolutionary search to assign sparsity levels across layers under a global budget. Second, we design a highly-optimized Sparse Matrix-Vector Multiplication (SpMV) kernel that efficiently supports arbitrary structured sparsity patterns during token generation. For example, on Qwen2.5-1.5B, RheoSparse reduces perplexity (PPL) by 33.09% and improves downstream task performance by 9.3% compared to 4:8 sparsity, while maintaining the same parameter count. Furthermore, our SpMV kernel outperforms the state-of-the-art sparse kernel SpInfer by up to 49.6%.
Jianing Zheng, Gang Chen 0023
DATE1
2026 Maximizing Personalized Energy-efficiency for Swarm Learning in 6G Networks
Jianing Zheng, Jiadong Yu
ICC1
2026 Recursive rolling-enhanced intrinsic feature-adaptive framework: avoiding information leakage for dynamic multi-step interval prediction
Yinan Peng, Peizhi Li, Jianing Zheng
Expert Syst. Appl.3
2026 Terafly: A Multinode FPGA-Based Accelerator Design for Efficient Cooperative Inference in LLMs
abstract
In this paper, we propose Terafly, a multi-node accelerator design tailored for efficient Large Language Model (LLM) deployment and inference. Conventional accelerator architectures struggle to effectively handle both the prefill and decode stages during inference. To address this limitation, we introduce a hybrid spatial-temporal architecture that combines the high-throughput advantages of spatial architectures with the flexibility of temporal architectures, enabling it to accommodate the diverse inference patterns of LLMs. In addition, we propose a generation framework to streamline the customization of our LLM-friendly design for various deployment scenarios. Within this framework, users can specify their requirements such as model type, target platform, and performance goals. The framework then generates multiple accelerator nodes and maps them to distinct Super Logic Regions (SLRs) within a single FPGA, enabling cooperative inference under a model parallelism scheme. Through experiments, our generated accelerator can be easily deployed on both Alveo U250 and U50lv cards, serving models ranging from OPT-350M to OPT-1.3B under various performance settings. Notably, when running OPT-1.3B using the generated dual-node accelerator on a single Alveo U50lv card, we achieve an average 1.1x speed-up and a 3.4x improvement in energy efficiency compared to the Nvidia A100 GPU.
Jianing Zheng, Gang Chen 0023, Libo Huang 0002, Xin Lou 0001, Wei-Shi Zheng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Scheduling and Fusion for Multimodal Federated Learning in Energy-Constrained Wireless Networks
abstract
The rise of privacy-preserving applications, such as medical diagnostics and the Metaverse, highlights the importance of federated learning (FL) for distributed model training at the wireless edge. These applications often rely on multimodal data (e.g., text, images, audio), necessitating advances in multimodal federated learning (MMFL). However, MMFL faces challenges like energy efficiency, multimodal fusion, and heterogeneity. To address these, a scheduling and fusion-based MMFL framework (SFMMFL) is proposed that focuses on improving both the scheduling mechanism and aggregation strategy. To improve the training performance under energy constraint, a Lyapunov-based scheduling algorithm is proposed, in which long-term optimization is transformed into immediate optimization. After that, to tackle the issue of model separation caused by multimodal datasets, a multimodal model aggregation strategy based on Knowledge Distillation (KD) is introduced for multimodal fusion. Convergence analysis proves its feasibility, and simulation results demonstrate that it can achieve faster and more stable convergence performance while improving model training accuracy. Specifically, our proposed SFMMFL can lower the energy consumption of the system by about$20\%$for computing and$16.67\%$for transmission.
Jianing Zheng, Jiadong Yu, Xiaolan Liu 0001
IEEE Trans. Mob. Comput.1
2025 LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
abstract
In this paper, we propose LoopLynx, a scalable dataflow architecture for efficient LLM inference that optimizes FPGA usage through a hybrid spatial-temporal design. The design of LoopLynx incorporates a hybrid temporal-spatial architecture, where computationally intensive operators are implemented as large dataflow kernels. This achieves high throughput similar to spatial architecture, and organizing and reusing these kernels in a temporal way together enhances FPGA peak performance. Furthermore, to overcome the resource limitations of a single device, we provide a multi-FPGA distributed architecture that overlaps and hides all data transfers so that the distributed accelerators are fully utilized. By doing so, LoopLynx can be effectively scaled to multiple devices to further explore model parallelism for large-scale LLM inference. Evaluation of GPT-2 model demonstrates that LoopLynx can achieve comparable performance to state-of-the-art single FPGA-based accelerations. In addition, compared to Nvidia A100, our accelerator with a dual-FPGA configuration delivers a 2.52x speed-up in inference latency while consuming only 48.1% of the energy.
Jianing Zheng, Gang Chen 0023
DATE1
2025 Knowledge Distillation-based Aggregation and Energy-constrained User Scheduling for Multimodal Federated Learning
abstract
The emergence of intelligent, privacy-preserving applications—such as medical diagnostics and the Metaverse—at the wireless network edge has made federated learning (FL) a promising approach for supporting distributed model training. Given that these applications rely on multimodal data (e.g., text, images, audio) collected from users, there is a growing need for research on multimodal federated learning (MMFL) to enable more robust and comprehensive learning. However, MMFL in wireless edge networks introduces unique challenges beyond traditional FL, particularly in managing the fusion of multimodal data and coordinating user scheduling across various data modalities. To address these issues, we propose a knowledge distillation-based MMFL (KD-MMFL) which considers unimodal model aggregation and adaptive multimodal aggregation via multi-teacher knowledge distillation (MTKD), to build a robust multimodal global model that effectively integrates insights from diverse data sources. Furthermore, to tackle the challenges of device heterogeneity in terms of data distribution and energy availability, we introduce a Lyapunov-based scheduling algorithm to map the long-term optimal problem that maximizes the training performance while considering users’ energy constraints into an immediate optimization problem. Experiment results demonstrate that the KD-MMFL framework successfully learns a robust multimodal global model from distributed multimodal datasets, while the Lyapunov-based scheduling method achieves an improvement of approximately 25% or more in energy efficiency compared to other scheduling methods.
Jianing Zheng, Jiadong Yu
GLOBECOM1
2024 AoU-Based Local Update and User Scheduling for Semi-Asynchronous Online Federated Learning in Wireless Networks
abstract
With the advent of the 5G and 6G eras and the explosive growth of mobile users, machine learning (ML) is increasingly used for extracting important information from a large amount of generated data and making intelligent decisions for complex environments. Especially, distributed ML techniques are getting more attention to enable training ML models in a distributed manner by exploiting distributed computational resources at the network edge. Federated learning (FL) as a classical distributed learning approach can not only protect data privacy but also reduce communication overhead. However, it requires synchrony among users, which is hard to satisfy due to the heterogeneity of the wireless networks. Hence, we first propose a clustering-based semi-asynchronous Online FL with AoU-based local update (CSAOFL-ALU) with importance-based user clustering and AsynFL-ALU-based local update. After that, the BS aggregates the cluster model of each cluster with synchronous FL. We also provide mathematical convergence analysis of the CSAOFL-ALU algorithm. The results show that the global model convergence rate is inversely proportional to the users’ AoU, at the same time, the convergence bound of the global loss function is inversely proportional to the size and the importance of the user dataset. The experiments are conducted on the non-IID MINST dataset. Numerical results demonstrate that the proposed AsynFL-ALU with priority-based user scheduling achieves better learning performance than fully AsynFL, and converges faster than the baseline user scheduling schemes. The CSAOFL-ALU converges faster with less communication time than the baseline algorithms and increases the fairness of user participation.
Jianing Zheng, Xiaolan Liu 0001, Zhuang Ling, Fengye Hu
IEEE Internet Things J.1
2023 When Materials Meet Sound: Discovering the Meaning of Deformable Materials in Musical Interaction
abstract
Research on Digital Musical Instruments (DMIs) design highlights that materiality plays an important role in DMI design and musical interaction. However, DMI design research often focuses on technology-oriented factors, with less exploration of the meaning of materials in design practice. In this paper, we explore how DMI designers understand deformable sensor materials and how they use these as a resource for creative aesthetic design. Eleven DMI designers were invited to use a selection of deformable sensor materials to create prototype DMIs with them in a design activity. Three design approaches emerged, determined by how designers perceived and explored sensor materials. We discuss the potential of the methodology for exploring strongly entangled elements, such as material, gesture, and sound, in DMI design. The results contribute to the design practice for DMI designers and to further exploration of material-based design research in Human-Computer Interaction.
Jianing Zheng, Andrew P. McPherson, Nick Bryan-Kinns
Conference on Designing Interactive Systems1
2022 Material Matters: Exploring Materiality in Digital Musical Instruments Design
abstract
Research on the design of Digital Musical Instruments (DMIs) has highlighted the importance of musical gestures and embodied interaction in DMI design. However, this research often focuses on technical and sonic factors of design, with less attention on how materials influence the design process and DMI design idea generation. Thus, this paper explores materiality in DMIs design through a material probe approach with deformable materials. This paper reports on a study with fifteen DMI designers investigating the evoked meaning of material properties in a musical context beyond their digital interactivity. Results suggest that material properties inspired participants’ design thinking, and there was a strong connection between tactility and imagined sound production. We also reported the patterns of gestural interaction of deformable materials in DMIs. We reflect on these results to report lessons learned that could inform interactive systems’ material design within and beyond the musical domain.
Jianing Zheng, Nick Bryan-Kinns, Andrew P. McPherson
Conference on Designing Interactive Systems1
2022 Risk evaluation for industrial smart product-service systems: An integrated method considering failure mode correlations
Wenyan Song, Jianing Zheng, Zixuan Niu, Pai Zheng
Adv. Eng. Informatics2