EDBT 2026 Demo / reviewers in the wild / expert
Qihua Zhou
dblp:213/0984
· DBLP profile ↗
33ranked-venue papers
16as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Systems, architecture and hardware · 9 · 8 first-author · 5 since 2021Computer networks · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeNC++: Efficient Diffusion-Enhanced Neural Codec for End-to-end Semantic Streaming at the EdgeabstractThe neural-enhanced video streaming (NeVS) has been an emerging technique to integrate neural models into video codecs for higher streaming efficiency. The state-of-the-art methods, e.g., DeNC and Gemino, typically compress videos in RGB space and restore video quality via a neural enhancement model hosted on the external media server. However, these methods are not always accessible in resource-constrained edge environments due to their heavy reliance on the media server's computation, which undermines end-to-end performance and restricts NeVS's usage boundary. This limitation raises an interesting question: is it possible to make NeVS lightweight so that all neural codec operations can be handled directly by clients' edge devices? In this paper, we present the answer yes and develop a new plug-and-play module called DeNC++, which significantly improves the compression-restoration-overhead trade-off over existing methods. Our core design philosophy is to wrap all the codec operations within a latent semantic space, in which the original high-dimensional visual signals are efficiently embedded into low-dimensional semantic representations. With this fundamental transformation, DeNC++'s neural encoder introduces the triple semantic-bitwidth-resolution compression to effectively lower the streaming traffic. Meanwhile, we make DeNC++'s neural decoder aware of the perceptual loss caused by its encoder and design tiny generative models to guarantee high restoration quality. We also strictly restrict the runtime computational overhead and accelerate the neural enhancement process, making DeNC++ compatible with commodity edge devices. Real-world evaluations reveal that DeNC++ consistently provides higher restoration quality while achieving 24-55 times higher compression ratio and 5-7 times end-to-end speedup over the latest NeVS solutions. Qihua Zhou, Wangjiang Gong, Zili Meng, Yaxiong Xie, Yaodong Huang, Junchen Jiang, Laizhong Cui |
AAAI | 1 |
| 2026 | Physical Embedding for Radio Map Construction
Zheng Xing 0001, Liang Xie 0011, Tao Guo 0004, Qi Tan 0003, Qihua Zhou, Weibing Zhao, Ruikang Zhong, Laizhong Cui |
ICC | 5 |
| 2026 | SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB VideosabstractBrain-inspired Spiking Neural Networks (SNNs) leverage a sparse, event-driven computational paradigm and have shown great potential for low-power object tracking. However, most existing SNN-based object tracking rely on event camera data, whereas traditional RGB video remains the dominant input modality in real-world applications. Research on RGB-based SNN multi-object tracking, particularly directly trained deep SNN models, is still in its infancy. To address this, we propose SMTrack, the first directly trained deep SNN framework for end-to-end multi-object tracking on standard RGB data. To handle the challenges caused by scale and density variations among objects, we introduce an Adaptive Scale-aware Normalized Wasserstein Distance Loss (Asa-NWDLoss), which dynamically adjusts the normalization factor based on the average object size within each training batch. For the identity association stage, we integrate the TrackTrack to maintain robust and consistent trajectory tracking. Extensive experiments on BEE24, MOT17, MOT20, and DanceTrack demonstrate that SMTrack achieves comparable performance to mainstream ANN-based approaches with only a few time steps. Codes is available at https://github.com/OpenCodeGithub/SMTrack. Pengzhi Zhong, Dan Zeng 0002, Qihua Zhou, Feixiang He, Shuiwang Li |
IEEE Internet Things J. | 4 |
| 2025 | Mjölnir: Breaking the Shield of Perturbation-Protected Gradients via Adaptive DiffusionabstractPerturbation-based mechanisms, such as differential privacy, mitigate gradient leakage attacks by introducing noise into the gradients, thereby preventing attackers from reconstructing clients' private data from the leaked gradients. However, can gradient perturbation protection mechanisms truly defend against all gradient leakage attacks? In this paper, we present the first attempt to break the shield of gradient perturbation protection in Federated Learning for the extraction of private information. We focus on common noise distributions, specifically Gaussian and Laplace, and apply our approach to DNN and CNN models. We introduce Mjölnir, a perturbation-resilient gradient leakage attack that is capable of removing perturbations from gradients without requiring additional access to the original model structure or external data. Specifically, we leverage the inherent diffusion properties of gradient perturbation protection to develop a novel diffusion-based gradient denoising model for Mjölnir. By constructing a surrogate client model that captures the structure of perturbed gradients, we obtain crucial gradient data for training the diffusion model. We further utilize the insight that monitoring disturbance levels during the reverse diffusion process can enhance gradient denoising capabilities, allowing Mjölnir to generate gradients that closely approximate the original, unperturbed versions through adaptive sampling steps. Extensive experiments demonstrate that Mjölnir effectively recovers the protected gradients and exposes the Federated Learning process to the threat of gradient leakage, achieving superior performance in gradient denoising and private data recovery. Xuan Liu 0001, Siqi Cai 0001, Qihua Zhou, Song Guo 0001, Ruibin Li, Kaiwei Lin |
AAAI | 3 |
| 2025 | DeNC: Unleash Neural Codecs in Video Streaming with Diffusion EnhancementabstractRecent years have witnessed the rise of Neural-enhanced Video Streaming (NeVS), which integrates neural restoration models into video codecs for higher compression-restoration performance. Despite its benefit, existing work has not well explored the full potential of NeVS paradigm, due to: (1) post-streaming restoration by decoder while lacking the proactive collaboration of encoder, (2) end-to-end optimization based on conventional rate-distortion theory, which has been verified that low distortion is not always a synonym for high perceptual quality, and (3) coupled design for domain-specific tasks that cannot generalize to various video codecs. Observing these limitations, our objective is not to incrementally present an improved restoration model. Instead, we focus on the encoder-decoder synergy, i.e., the codec, which is non-trivial since it inherently strikes the rate-distortion-perception trade-off of NeVS. Aiming at this target, we propose the Diffusion-enhanced Neural Codec (DeNC), a plug-and-play module for current NeVS paradigm, to significantly reduce the required bitrates while preserving high perceptual quality of restored videos. Our key design is twofold. First, DeNC improves the encoder's compression efficiency by simultaneously reducing the resolution and color bit-depth of frame referencing. Second, DeNC empowers the decoder with perception-oriented restoration capability by making its diffusion-based restoration process aware of the encoder's compression conditions. Real-world evaluations show that DeNC improves compression ratios with nearly an order of magnitude and achieves much higher restoration quality (e.g., 93+ VMAF and 23% higher MOS) over the latest baselines. Qihua Zhou, Ruibin Li, Jingcai Guo, Yaodong Huang, Zhenda Xu, Laizhong Cui, Song Guo 0001 |
AAAI | 1 |
| 2025 | D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingabstractThe mixture of experts (MoE) model is a sparse variant of large language models (LLMs), designed to hold a better balance between intelligent capability and computational overhead. Despite its benefits, MoE is still too expensive to deploy on resource-constrained edge devices, especially with the demands of on-device inference services. Recent research efforts often apply model compression techniques, such as quantization, pruning and merging, to restrict MoE complexity. Unfortunately, due to their predefined static model optimization strategies, they cannot always achieve the desired quality-overhead trade-off when handling multiple requests, finally degrading the on-device quality of service. These limitations motivate us to propose the D2MoE, an algorithm-system co-design framework that matches diverse task requirements by dynamically allocating the most proper bit-width to each expert. Specifically, inspired by the nested structure of matryoshka dolls, we propose the matryoshka weight quantization (MWQ) to progressively compress expert weights in a bit-nested manner and reduce the required runtime memory. On top of it, we further optimize the I/O-computation pipeline and design a heuristic scheduling algorithm following our hottest-expert-bit-first (HEBF) principle, which maximizes the expert parallelism between I/O and computation queue under constrained memory budgets, thus significantly reducing the idle temporal bubbles waiting for the experts to load. Evaluations on real edge devices show that D2MoE improves the overall inference throughput by up to 1.39× and reduces the peak memory footprint by up to 53% over the latest on-device inference frameworks, while still preserving comparable serving accuracy as its INT8 counterparts. Qihua Zhou, Zicong Hong, Song Guo 0001 |
MobiCom | 2 |
| 2025 | Prompt-Ladder: Memory-efficient prompt tuning for vision-language models on edge devices
Siqi Cai 0001, Xuan Liu 0001, Jingling Yuan, Qihua Zhou |
Pattern Recognit. | 4 |
| 2025 | Collaborative Neural Architecture Search for Personalized Federated LearningabstractPersonalized federated learning (pFL) is a promising approach to train customized models for multiple clients over heterogeneous data distributions. However, existing works on pFL often rely on the optimization of model parameters and ignore the personalization demand on neural network architecture, which can greatly affect the model performance in practice. Therefore, generating personalized models with different neural architectures for different clients is a key issue in implementing pFL in a heterogeneous environment. Motivated by Neural Architecture Search (NAS), a model architecture searching methodology, this paper aims to automate the model design in a collaborative manner while achieving good training performance for each client. Specifically, we reconstruct the centralized searching of NAS into the distributed scheme called Personalized Architecture Search (PAS), where differentiable architecture fine-tuning is achieved via gradient-descent optimization, thus making each client obtain the most appropriate model. Furthermore, to aggregate knowledge from heterogeneous neural architectures, a knowledge distillation-based training framework is proposed to achieve a good trade-off between generalization and personalization in federated learning. Extensive experiments demonstrate that our architecture-level personalization method achieves higher accuracy under the non-iid settings, while not aggravating model complexity over state-of-the-art benchmarks. Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Zicong Hong, Yufeng Zhan, Qihua Zhou |
IEEE Trans. Computers | 6 |
| 2025 | Model Decomposition and Reassembly for Purified Knowledge Transfer in Personalized Federated LearningabstractPersonalized federated learning (pFL) is to collaboratively train non-identical machine learning models for different clients to adapt to their heterogeneously distributed datasets. State-of-the-art pFL approaches pay much attention on exploiting clients’ inter-similarities to facilitate the collaborative learning process, meanwhile, can barely escape from the irrelevant knowledge pooling that is inevitable during the aggregation phase, and thus hindering the optimization convergence and degrading the personalization performance. To tackle such conflicts between facilitating collaboration and promoting personalization, we propose a novel pFL framework, dubbed pFedC, to first decompose the global aggregated knowledge into several compositional branches, and then selectively reassemble the relevant branches for supporting conflicts-aware collaboration among contradictory clients. Specifically, by reconstructing each local model into a shared feature extractor and multiple decomposed task-specific classifiers, the training on each client transforms into a mutually reinforced and relatively independent multi-task learning process, which provides a new perspective for pFL. Besides, we conduct a purified knowledge aggregation mechanism via quantifying the combination weights for each client to capture clients’ common prior, as well as mitigate potential conflicts from the divergent knowledge caused by the heterogeneous data. Extensive experiments over various models and datasets demonstrate the effectiveness and superior performance of the proposed algorithm. Jie Zhang 0076, Song Guo 0001, Xiaosong Ma, Wenchao Xu 0001, Qihua Zhou, Jingcai Guo, Zicong Hong, Jun Shan |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Feature Correlation-Guided Knowledge Transfer for Federated Self-Supervised LearningabstractExtensive attention has been paid to the application of self-supervised learning (SSL) approaches on federated learning (FL) to tackle the label scarcity problem. Previous works on federated SSL (FedSSL) generally fall into two categories: parameter-based model aggregation or data-based feature sharing to achieve knowledge transfer among multiple unlabeled clients. Despite the progress, they inevitably rely on some assumptions, such as homogeneous models or the existence of an additional public dataset, which hinder the universality of the training frameworks for more general scenarios (e.g., unlabeled clients with heterogeneous models). Therefore, in this article, we propose a novel and general method named federated self-supervised learning with feature-correlation-based aggregation (FedFoA) to tackle the above limitations. By exchanging feature correlation instead of model parameters or feature mappings, our approach reduces the discrepancies of local representations learning processes, thus promoting collaboration between heterogeneous clients. A factorization-based method is designed to extract the cross-feature relation matrix from local representations, which serves as a knowledge medium for the aggregation phase. We demonstrate that FedFoA is a heterogeneity-supportive and privacy-preserving training framework and can be easily compatible with state-of-the-art FedSSL methods. Extensive empirical experiments demonstrate our proposed approach outperforms the state-of-the-art methods by a significant margin. Yi Liu 0057, Song Guo 0001, Jie Zhang 0076, Yufeng Zhan, Qihua Zhou, Yingchun Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | On the Robustness of Neural-Enhanced Video Streaming against Adversarial AttacksabstractThe explosive growth of video traffic on today's Internet promotes the rise of Neural-enhanced Video Streaming (NeVS), which effectively improves the rate-distortion trade-off by employing a cheap neural super-resolution model for quality enhancement on the receiver side. Missing by existing work, we reveal that the NeVS pipeline may suffer from a practical threat, where the crucial codec component (i.e., encoder for compression and decoder for restoration) can trigger adversarial attacks in a man-in-the-middle manner to significantly destroy video recovery performance and finally incurs the malfunction of downstream video perception tasks. In this paper, we are the first attempt to inspect the vulnerability of NeVS and discover a novel adversarial attack, called codec hijacking, where the injected invisible perturbation conspires with the malicious encoding matrix by reorganizing the spatial-temporal bit allocation within the bitstream size budget. Such a zero-day vulnerability makes our attack hard to defend because there is no visual distortion on the recovered videos until the attack happens. More seriously, this attack can be extended to diverse enhancement models, thus exposing a wide range of video perception tasks under threat. Evaluation based on state-of-the-art video codec benchmark illustrates that our attack significantly degrades the recovery performance of NeVS over previous attack methods. The damaged video quality finally leads to obvious malfunction of downstream tasks with over 75% success rate. We hope to arouse public attention on codec hijacking and its defence. Qihua Zhou, Jingcai Guo, Song Guo 0001, Ruibin Li, Jie Zhang 0076, Zhenda Xu |
AAAI | 1 |
| 2024 | ParsNets: A Parsimonious Composition of Orthogonal and Low-Rank Linear Networks for Zero-Shot Learning
Jingcai Guo, Qihua Zhou, Xiaocheng Lu, Ruibin Li, Jie Zhang 0076, Junyang Chen 0001, Xin Xie 0001, Song Guo 0001 |
IJCAI | 2 |
| 2024 | FreePIH: Training-Free Painterly Image Harmonization with Diffusion ModelabstractThis paper provides an efficient training-free painterly image harmonization (PIH) method, dubbed FreePIH, that leverages only a pre-trained diffusion model to achieve state-of-the-art harmonization results. Unlike existing methods that require either training auxiliary networks or fine-tuning a large pre-trained backbone, or both, to harmonize a foreground object with a painterly-style background image, our FreePIH tames the denoising process as a plug-in module for foreground image style transfer. Specifically, we find that the very last few steps of the denoising (i.e., generation) process strongly correspond to the stylistic information of images, and based on this, we propose to augment the latent features of both the foreground and background images with Gaussians for a direct denoising-based harmonization. To guarantee the fidelity of the harmonized image, we make use of latent features to enforce the consistency of the content and stability of the foreground objects in the latent space, and meanwhile, aligning both fore-/back-grounds with the same style. Moreover, to accommodate the generation with more structural and textural details, we further integrate text prompts to attend to the latent features, hence improving the generation quality. Quantitative and qualitative evaluations on COCO and LAION 5B datasets demonstrate that our method can surpass representative baselines by large margins. Ruibin Li, Jingcai Guo, Qihua Zhou, Song Guo 0001 |
ACM Multimedia | 3 |
| 2024 | PASS: Patch Automatic Skip Scheme for Efficient On-Device Video PerceptionabstractReal-time video perception tasks are often challenging on resource-constrained edge devices due to the issues of accuracy drop and hardware overhead, where saving computations is the key to performance improvement. Existing methods either rely on domain-specific neural chips or priorly searched models, which require specialized optimization according to different task properties. These limitations motivate us to design a general and task-independent methodology, called Patch Automatic Skip Scheme (PASS), which supports diverse video perception settings by decoupling acceleration and tasks. The gist is to capture inter-frame correlations and skip redundant computations at patch level, where the patch is a non-overlapping square block in visual. PASS equips each convolution layer with a learnable gate to selectively determine which patches could be safely skipped without degrading model accuracy. Specifically, we are the first to construct a self-supervisory procedure for gate optimization, which learns to extract contrastive representations from frame sequences. The pre-trained gates can serve as plug-and-play modules to implement patch-skippable neural backbones, and automatically generate proper skip strategy to accelerate different video-based downstream tasks, e.g., outperforming state-of-the-art MobileHumanPose in 3D pose estimation and FairMOT in multiple object tracking, by up to 9.43 × and 12.19 × speedups, respectively, on NVIDIA Jetson Nano devices. Qihua Zhou, Song Guo 0001, Jiacheng Liang, Jingcai Guo, Zhenda Xu, Jingren Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Tree Learning: Towards Promoting Coordination in Scalable Multi-Client Training AccelerationabstractIteration based collaborative learning (CL) paradigms, such as federated learning (FL) and split learning (SL), faces challenges in training neural models over the rapidly growing yet resource-constrained edge devices. Such devices have difficulty in accommodating a full-size large model for FL or affording an excessive waiting time for the mandatory synchronization step in SL. To deal with such challenge, we propose a novel CL framework which adopts an tree-aggregation structure with an adaptive partition and ensemble strategy to achieve optimal synchronization and fast convergence at scale. To find the optimal split point for heterogeneous clients, we also design a novel partitioning algorithm by minimizing the idleness during communication and achieving the optimal synchronization between clients. In addition, a parallelism paradigm is proposed to unleash the potential of optimum synchronization between the clients and server to boost the distributed training process without losing model accuracy for edge devices. Furthermore, we theoretically prove that our framework can achieve better convergence rate than state-of-the-art CL paradigms. We conduct extensive experiments and show that our framework is 4.6× in training speed as compared with the traditional methods, without compromising training accuracy. Tao Guo 0004, Song Guo 0001, Feijie Wu, Wenchao Xu 0001, Jiewei Zhang, Qihua Zhou, Quan Chen 0003, Weihua Zhuang |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Chiron: A Robustness-Aware Incentive Scheme for Edge Learning via Hierarchical Reinforcement LearningabstractOver the past few years, edge learning has achieved significant success in mobile edge networks. Few works have designed incentive mechanism that motivates edge nodes to participate in edge learning. However, most existing works only consider myopic optimization and assume that all edge nodes are honest, which lacks long-term sustainability and the final performance assurance. In this paper, we propose Chiron, an incentive-driven Byzantine-resistant long-term mechanism based on hierarchical reinforcement learning (HRL). First, our optimization goal includes both learning-algorithm performance criteria (i.e., global accuracy) and systematical criteria (i.e., resource consumption), which aim to improve the edge learning performance under a given resource budget. Second, we propose a three-layer HRL architecture to handle long-term optimization, short-term optimization, and byzantine resistance, respectively. Finally, we conduct experiments on various edge learning tasks to demonstrate the superiority of the proposed approach. Specifically, our system can successfully exclude malicious nodes and lazy nodes out of the edge learning participation and achieves 14.96% higher accuracy and 12.66% higher total utility than the state-of-the-art methods under the same budget limit. Yi Liu 0057, Song Guo 0001, Yufeng Zhan, Leijie Wu, Zicong Hong, Qihua Zhou |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Graph Knows Unknowns: Reformulate Zero-Shot Learning as Sample-Level Graph RecognitionabstractZero-shot learning (ZSL) is an extreme case of transfer learning that aims to recognize samples (e.g., images) of unseen classes relying on a train-set covering only seen classes and a set of auxiliary knowledge (e.g., semantic descriptors). Existing methods usually resort to constructing a visual-to-semantics mapping based on features extracted from each whole sample. However, since the visual and semantic spaces are inherently independent and may exist in different manifolds, these methods may easily suffer from the domain bias problem due to the knowledge transfer from seen to unseen classes. Unlike existing works, this paper investigates the fine-grained ZSL from a novel perspective of sample-level graph. Specifically, we decompose an input into several fine-grained elements and construct a graph structure per sample to measure and utilize element-granularity relations within each sample. Taking advantage of recently developed graph neural networks (GNNs), we formulate the ZSL problem to a graph-to-semantics mapping task, which can better exploit element-semantics correlation and local sub-structural information in samples. Experimental results on the widely used benchmark datasets demonstrate that the proposed method can mitigate the domain bias problem and achieve competitive performance against other representative methods. Jingcai Guo, Song Guo 0001, Qihua Zhou, Xiaocheng Lu, Fushuo Huo |
AAAI | 3 |
| 2023 | PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge DevicesabstractReal-time video perception tasks are often challenging over the resource-constrained edge devices due to the concerns of accuracy drop and hardware overhead, where saving computations is the key to performance improvement. Existing methods either rely on domain-specific neural chips or priorly searched models, which require specialized optimization according to different task properties. In this work, we propose a general and task-independent Patch Automatic Skip Scheme (PASS), a novel end-to-end learning pipeline to support diverse video perception settings by decoupling acceleration and tasks. The gist is to capture the temporal similarity across video frames and skip the redundant computations at patch level, where the patch is a non-overlapping square block in visual. PASS equips each convolution layer with a learnable gate to selectively determine which patches could be safely skipped without degrading model accuracy. As to each layer, a desired gate needs to make flexible skip decisions based on intermediate features without any annotations, which cannot be achieved by conventional supervised learning paradigm. To address this challenge, we are the first to construct a tough self-supervisory procedure for optimizing these gates, which learns to extract contrastive representation, i.e., distinguishing similarity and difference, from frame sequence. These high-capacity gates can serve as a plug-and-play module for convolutional neural network (CNN) backbones to implement patch-skippable architectures, and automatically generate proper skip strategy to accelerate different video-based downstream tasks, e.g., outperforming the state-of-the-art MobileHumanPose (MHP) in 3D pose estimation and FairMOT in multiple object tracking, by up to 9.43 times and 12.19 times speedups, respectively. By directly processing the raw data of frames, PASS can generalize to real-time video streams on commodity edge devices, e.g., NVIDIA Jetson Nano, with efficient performance in realistic deployment. Qihua Zhou, Song Guo 0001, Jiacheng Liang, Zhenda Xu, Jingren Zhou 0001 |
AAAI | 1 |
| 2023 | Development of Deep Learning Algorithms for Automated Scoliosis and Abnormal Posture Screening Using 2D Back ImageabstractAdolescent idiopathic scoliosis is becoming a common spinal disorder among adolescents. The traditional methods of scoliosis screening are labor-intensive and can result in unnecessary referrals and radiological exposure for adolescents due to their low positive predictive value. For early screening of scoliosis and abnormal posture, a mobile-based cost-free, accurate and radiation-free scoliosis screening system is proposed in this paper. We establish a database with labeled 2D unclothed back images and corresponding whole-spine standing posterior-anterior X-ray images, and innovatively propose a new network topology of the 2D back image to localize the back landmarks. With only an unclothed back image, this system can automatically classify normal, abnormal posture and scoliosis with an overall classification accuracy of 88.1%. This system has the potential to overcome the time and space limitations of conventional screening for scoliosis and abnormal posture. Zhenda Xu, Donghua Hang, Qihua Zhou, Song Guo 0001, Aiqian Gan |
ICME | 5 |
| 2022 | 2D Photogrammetry Image of Adolescent Idiopathic Scoliosis Screening Using Deep Learning
Zhenda Xu, Jiazi Ouyang, Aiqian Gan, Qihua Zhou, Song Guo 0001 |
ISBRA | 5 |
| 2022 | Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative LearningabstractIt witnesses that the collaborative learning (CL) systems often face the performance bottleneck of limited bandwidth, where multiple low-end devices continuously generate data and transmit intermediate features to the cloud for incremental training. To this end, improving the communication efficiency by reducing traffic size is one of the most crucial issues for realistic deployment. Existing systems mostly compress features at pixel level and ignore the characteristics of feature structure, which could be further exploited for more efficient compression. In this paper, we take new insights into implementing scalable CL systems through a hierarchical compression on features, termed Stripe-wise Group Quantization (SGQ). Different from previous unstructured quantization methods, SGQ captures both channel and spatial similarity in pixels, and simultaneously encodes features in these two levels to gain a much higher compression ratio. In particular, we refactor feature structure based on inter-channel similarity and bound the gradient deviation caused by quantization, in forward and backward passes, respectively. Such a double-stage pipeline makes SGQ hold a sublinear convergence order as the vanilla SGD-based optimization. Extensive experiments show that SGQ achieves a higher traffic reduction ratio by up to 15.97 times and provides 9.22 times image processing speedup over the uniform quantized training, while preserving adequate model accuracy as FP32 does, even using 4-bit quantization. This verifies that SGQ can be applied to a wide spectrum of edge intelligence applications. Qihua Zhou, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Jiewei Zhang, Tao Guo 0004, Zhenda Xu, Zhihao Qu |
NeurIPS | 1 |
| 2022 | A Comprehensive Survey on Training Acceleration for Large Machine Learning Models in IoTabstractThe ever-growing artificial intelligence (AI) applications have greatly reshaped our world in many areas, e.g., smart home, computer vision, natural language processing, etc. Behind these applications are usually machine learning (ML) models with extremely large size, which require huge data sets for accurate training to mine the value contained in the big data. Large ML models, however, can consume tremendous computing resources to achieve decent performance and thus, it is difficult to train them in resource-constrained Internet of Things (IoT) environments, which would prevent further development and application of AI techniques in the future. To deal with such challenges, there are many efforts on accelerating the training process for large ML models in IoT. In this article, we provide a comprehensive review on the recent advances toward reducing the computing cost during the training stage while maintaining comparable model accuracy. Specifically, the optimization algorithms that aim to improve the convergence rate are emphasized over various distributed learning architectures that exploit ubiquitous computing resources. Then, the article elaborates the computation hardware acceleration and communication optimization for collaborative training among multiple learning entities. Finally, the remaining challenges, future opportunities, and possible directions are discussed. Haozhao Wang, Zhihao Qu, Qihua Zhou, Haobo Zhang 0002, Boyuan Luo, Wenchao Xu 0001, Song Guo 0001, Ruixuan Li 0001 |
IEEE Internet Things J. | 3 |
| 2021 | Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning
Qihua Zhou, Song Guo 0001, Zhihao Qu, Jingcai Guo, Zhenda Xu, Jiewei Zhang, Tao Guo 0004, Boyuan Luo, Jingren Zhou 0001 |
USENIX ATC | 1 |
| 2021 | On-Device Learning Systems for Edge Intelligence: A Software and Hardware Synergy PerspectiveabstractModern machine learning (ML) applications are often deployed in the cloud environment to exploit the computational power of clusters. However, this in-cloud computing scheme cannot satisfy the demands of emerging edge intelligence scenarios, including providing personalized models, protecting user privacy, adapting to real-time tasks, and saving resource cost. In order to conquer the limitations of conventional in-cloud computing, there comes the rise of on-device learning, which makes the end-to-end ML procedure totally on user devices, without unnecessary involvement of the cloud. In spite of the promising advantages of on-device learning, implementing a high-performance on-device learning system still faces with many severe challenges, such as insufficient user training data, backward propagation (BP) blocking, and limited peak processing speed. Observing the substantial improvement space in the implementation and acceleration of on-device learning systems, we intend to present a comprehensive analysis of the latest research progress and point out potential optimization directions from the system perspective. This survey presents a software and hardware synergy of on-device learning techniques, covering the scope of model-level neural network design, algorithm-level training optimization, and hardware-level instruction acceleration. We hope this survey could bring fruitful discussions and inspire the researchers to further promote the field of edge intelligence. Qihua Zhou, Zhihao Qu, Song Guo 0001, Boyuan Luo, Jingcai Guo, Zhenda Xu, Rajendra Akerkar |
IEEE Internet Things J. | 1 |
| 2021 | Falcon: Addressing Stragglers in Heterogeneous Parameter Server Via Multiple ParallelismabstractThe parameter server architecture has shown promising performance advantages when handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving stragglers may not fully exploit the computation resource of the cluster as evidenced by our experiments, especially in the heterogeneous environment. This motivates us to design a heterogeneity-aware parameter server paradigm that addresses stragglers and accelerates DL training from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines to solve this problem in two aspects: (1) controlling each worker's training speed via elastic training parallelism control and (2) transferring blocked tasks from stragglers to pioneers to fully utilize the computation resource. Following these guidelines, we propose the abstraction of parallelism as an infrastructure and design the Elastic-Parallelism Synchronous Parallel (EPSP) algorithm to handle distributed training and parameter synchronization, supporting both enforcedand slack-synchronization schemes. The whole idea has been implemented into a prototype called Falcon which effectively accelerates the DL training speed with the presence of stragglers. Evaluation under various benchmarks with baseline comparison demonstrates the superiority of our system. Specifically, Falcon reduces the training convergence time, by up to 61.83, 55.19, 38.92, and 23.68 percent shorter than FlexRR, Sync-opt, ConSGD, and DynSGD, respectively. Qihua Zhou, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun, Kun Wang 0005 |
IEEE Trans. Computers | 1 |
| 2021 | Petrel: Heterogeneity-Aware Distributed Deep Learning Via Hybrid SynchronizationabstractThe parameter server (PS) paradigm has achieved great success in deploying large-scale distributed Deep Learning (DL) systems. However, these systems implicitly assume that the cluster is homogeneous and this assumption does not hold in many realworld cases. Although the previous efforts are paid to address heterogeneity, they mainly prioritize the contribution of fast workers and reduce the involvement of slow workers, resulting in the limitations of workload imbalance and computation inefficiency. We reveal that grouping workers into communities, an abstraction proposed by us, and handling parameter synchronization at the community level can conquer these limitations and accelerate the training convergence progress. The inspiration of community comes from our exploration of prior knowledge about the similarity between workers, which is often neglected by previous work. These observations motivate us to propose a new synchronization mechanism named Community-aware Synchronous Parallel (CASP), which uses the Asynchronous Advantage Actor-Critic (A3C)-based algorithm to intelligently determine community configuration and fully improve the synchronization performance. The whole idea has been implemented in a prototype system called Petrel that achieves a good balance between convergence efficiency and communication overhead. The evaluation under various benchmarks with multiple metrics and baseline comparison demonstrates the effectiveness of Petrel. Specifically, Petrel accelerates the training convergence speed by up to 1.87 x faster and reduces communication traffic by up to 26.85 percent, on average, over the non-community synchronization mechanisms. Qihua Zhou, Song Guo 0001, Zhihao Qu, Peng Li 0017, Li Li 0012, Minyi Guo, Kun Wang 0005 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Canary: Decentralized Distributed Deep Learning Via Gradient Sketch and Partition in Multi-Interface NetworksabstractThe multi-interface networks are efficient infrastructures to deploy distributed Deep Learning (DL) tasks as the model gradients generated by each worker can be exchanged to others via different links in parallel. Although this decentralized parameter synchronization mechanism can reduce the time of gradient exchange, building a high-performance distributed DL architecture still requires the balance of communication efficiency and computational utilization, i.e., addressing the issues of traffic burst, data consistency, and programming convenience. To achieve this goal, we intend to asynchronously exchange gradient pieces without the central control in multi-interface networks. We propose the Piece-level Gradient Exchange and Multi-interface Collective Communication to handle parameter synchronization and traffic transmission, respectively. Specifically, we design the gradient sketch approach based on 8-bit uniform quantization to compress gradient tensors and introduce the colayerabstraction to better handle gradient partition, exchange and pipelining. Also, we provide general programming interfaces to capture the synchronization semantics and build the Gradient Exchange Index (GEI) data structures to make our approach online applicable. We implement our algorithms into a prototype system called Canary by using PyTorch-1.4.0. Experiments conducted in Alibaba Cloud demonstrate that Canary reduces 56.28 percent traffic on average and completes the training by up to 1.61x, 2.28x, and 2.84x faster than BML, Ako on PyTorch, and PS on TensorFlow, respectively. Qihua Zhou, Kun Wang 0005, Haodong Lu 0001, Wenyao Xu, Yanfei Sun, Song Guo 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Petrel: Community-aware Synchronous Parallel for Heterogeneous Parameter ServerabstractAs to address the impact of heterogeneity in distributed Deep Learning (DL) systems, most previous approaches focus on prioritizing the contribution of fast workers and reducing the involvement of slow workers, incurring the limitations of workload imbalance and computation inefficiency. We reveal that grouping workers into communities, an abstraction proposed by us, and handling parameter synchronization in community level can conquer these limitations and accelerate the training convergence progress. The inspiration of community comes from our exploration of prior knowledge about the similarity between workers, which is often neglected by previous work. These observations motivate us to propose a new synchronization mechanism named Community-aware Synchronous Parallel (CSP), which uses the Asynchronous Advantage Actor-Critic (A3C), a Reinforcement Learning (RL) based algorithm, to intelligently determine community configuration and fully improve the synchronization performance. The whole idea has been implemented in a system called Petrel that achieves a good balance between convergence efficiency and communication overhead. The evaluation under different benchmarks demonstrates our approach can effectively accelerate the training convergence speed and reduce synchro-nization traffic. Qihua Zhou, Song Guo 0001, Peng Li 0017, Yanfei Sun, Li Li 0012, Minyi Guo, Kun Wang 0005 |
ICDCS | 1 |
| 2020 | Dual-view Attention Networks for Single Image Super-ResolutionabstractOne non-negligible flaw of the convolutional neural networks (CNNs) based single image super-resolution (SISR) models is that most of them are not able to restore high-resolution (HR) images containing sufficient high-frequency information. Worse still, as the depth of CNNs increases, the training easily suffers from the vanishing gradients. These problems hinder the effectiveness of CNNs in SISR. In this paper, we propose the Dual-view Attention Networks to alleviate these problems for SISR. Specifically, we propose the local aware (LA) and global aware (GA) attentions to deal with LR features in unequal manners, which can highlight the high-frequency components and discriminate each feature from LR images in the local and global views, respectively. Furthermore, the local attentive residual-dense (LARD) block that combines the LA attention with multiple residual and dense connections is proposed to fit a deeper yet easy to train architecture. The experimental results verified the effectiveness of our model compared with other state-of-the-art methods. Jingcai Guo, Shiheng Ma, Jie Zhang 0076, Qihua Zhou, Song Guo 0001 |
ACM Multimedia | 4 |
| 2019 | Falcon: Towards Computation-Parallel Deep Learning in Heterogeneous Parameter ServerabstractParameter server paradigm has shown great performance superiority for handling deep learning (DL) applications. One crucial issue in this regard is the presence of stragglers, which significantly retards DL training progress. Previous solutions for solving straggler may not fully exploit the computation capacity of a cluster as evidenced by our experiments. This motivates us to make an attempt at building a new parameter server architecture that mitigates and addresses stragglers in heterogeneous DL from the perspective of computation parallelism. We introduce a novel methodology named straggler projection to give a comprehensive inspection of stragglers and reveal practical guidelines for resolving this problem: (1) reducing straggler emergence frequency via elastic parallelism control and (2) transferring blocked tasks to pioneer workers for fully exploiting cluster computation capacity. Following the guidelines, we propose the abstraction of parallelism as an infrastructure and elaborate the Elastic-Parallelism Synchronous Parallel (EPSP) that supports both enforced-and slack-synchronization schemes. The whole idea has been implemented in a prototype called Falcon which efficiently accelerates the DL training progress with the presence of stragglers. Evaluation under various benchmarks with baseline comparison evidences the superiority of our system. Specifically, Falcon yields shorter convergence time, by up to 61.83%, 55.19%, 38.92% and 23.68% reduction over FlexRR, Sync-opt, ConSGD and DynSGD, respectively. Qihua Zhou, Kun Wang 0005, Song Guo 0001, Haodong Lu 0001, Li Li 0012, Minyi Guo, Yanfei Sun |
ICDCS | 1 |
| 2019 | Fast Coflow Scheduling via Traffic Compression and Stage Pipelining in Datacenter NetworksabstractBig data analytics in datacenters often involve scheduling of data-parallel jobs. Traditional scheduling techniques based on improving network resource utilization are subject to limited bandwidth in datacenter networks. To alleviate the shortage of bandwidth, some cluster frameworks employ techniques of traffic compression to reduce transmission consumption. However, they tackle scheduling in a coarse-grained manner at task level and do not perform well in terms of flow-level metrics due to high complexity. Fortunately, the abstraction of coflow pioneers a new perspective to facilitate scheduling efficiency. In this paper, we introduce a coflow compression mechanism to minimize the completion time in data-intensive applications. Due to the NP-hardness, we propose a heuristic algorithm called Fastest-Volume-Disposal-First (FVDF) to solve this problem. For online applicability, FVDF supports stage pipelining to accelerate scheduling and exploits recurrent neural networks (RNNs) to predict compression speed. Meanwhile, we build Swallow, an efficient scheduling system that implements our proposed algorithms. It minimizes coflow completion time (CCT) while guaranteeing resource conservation and starvation freedom. The results of both trace-driven simulations and real experiments show the superiority of our algorithm, over existing one. Specifically, Swallow speeds up CCT and job completion time (JCT) by up to 1.47χ and 1.66χ on average, respectively, over the SEBF in Varys, one of the most efficient coflow scheduling algorithms so far. Moreover, with coflow compression, Swallow reduces data traffic by up to 48.41 percent on average. Qihua Zhou, Kun Wang 0005, Peng Li 0017, Deze Zeng, Song Guo 0001, Minyi Guo |
IEEE Trans. Computers | 1 |
| 2018 | Swallow: Joint Online Scheduling and Coflow Compression in Datacenter NetworksabstractBig data analytics in datacenters often involves scheduling of data-parallel job, which are bottlenecked by limited bandwidth of datacenter networks. To alleviate the shortage of bandwidth, some existing work has proposed traffic compression to reduce the amount of data transmitted over the network. However, their proposed traffic compression works in a coarse-grained manner at job level, leaving a large optimization space unexplored for further performance improvement. In this paper, we propose a flow-level traffic compression and scheduling system, called Swallow, to accelerate data-intensive applications. Specifically, we target on coflows, which is an elegant abstraction of parallel flows generated by big data jobs. With the objective of minimizing coflow completion time (CCT), we propose a heuristic algorithm called Fastest-Volume-Disposal-First (FVDV) and implement Swallow based on Spark. The results of both trace-driven simulations and real experiments show the superiority of our system, over existing algorithms. Swallow can reduce CCT and job completion time (JCT) by up to 1.47 × and 1.66 × on average, respectively, over the SEBF in Varys, one of the most efficient coflow scheduling algorithms so far. Moreover, with coflow compression, Swallow reduces data traffic by up to 48.41% on average. Qihua Zhou, Peng Li 0017, Kun Wang 0005, Deze Zeng, Song Guo 0001, Minyi Guo |
IPDPS | 1 |
| 2017 | Promoting Security and Efficiency in D2D Underlay Communication: A Bargaining Game ApproachabstractDevice-to-device (D2D) communication is a promising technology for expanding the next generation wireless cellular network. To deal with the security challenges and optimize the system communication quality, this paper investigates the security and efficiency problem in D2D underlay communication with the presence of malicious eavesdroppers. Fairness and strategy space of both D2D user equipment (DUE) and cellular user equipment (CUE) are taken into consideration under the control of proposed efficiency functions. Problems are formulated as a series of utility functions built on the unit price of jamming power and the amount of jamming service. Extracting system model into a price negotiation under Bargaining Game (PNBG) that a buyer and a seller both desiring maximum its profits, we solve the problems by reaching an agreement of the two sides. The step number of bargain process is also a restriction under consideration. For the Non-Step scheme, an Evaluation Function (EF) and a Comprehensive Utility Function (CUF) are demonstrated to analyze the negotiation process. For Step-Contained scheme, the step number of iteration is involved and an Attenuation Function (AF) is introduced to modify the Bargaining Game. Algorithms of two schemes are designed to derive the equilibrium point for reaching an agreement. Finally, simulations are illustrated for verifying proposed approach. Qihua Zhou, Weifeng Lu, Siguang Chen, Kun Wang 0005 |
GLOBECOM | 1 |