Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhenda Xu

dblp:266/6962 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0003-0655-5405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 83% Generative modeling · 8% Video understanding and tracking · 6%
Computer networks
3 papers
Content delivery and video streaming · 64% Wireless sensing and localization · 30% Edge and fog computing · 7%
Computer graphics and multimedia
1 paper
Image and video coding · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Embedded and real-time systems · 54% Distributed systems · 46%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Content delivery and video streaming › video delivery
neural-enhanced video streaming
1.622025
DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement · AAAI 2025
On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks · AAAI 2024
Machine learning › Efficient and distributed learning
model acceleration
1.422024
PASS: Patch Automatic Skip Scheme for Efficient On-Device Video Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2024
PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge Devices · AAAI 2023
Image and video coding › neural compression
neural image and video compression
0.912025
DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement · AAAI 2025
Image and video coding › rate-distortion analysis
rate-distortion-perception tradeoff
0.912025
DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement · AAAI 2025
Wireless sensing and localization
adversarial robustness
0.812024
On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks · AAAI 2024
Security and privacy of machine learning
adversarial attack
0.812024
On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks · AAAI 2024
Machine learning › Efficient and distributed learning › token reduction
patch pruning
0.712023
PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge Devices · AAAI 2023
Distributed systems
edge computing
0.712023
PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge Devices · AAAI 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.612022
Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
feature compression
0.612022
Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning · NeurIPS 2022
Machine learning › Efficient and distributed learning
federated learning
0.612022
Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning · NeurIPS 2022
Machine learning › Efficient and distributed learning › low-precision training
INT8 training
0.512021
Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning · USENIX ATC 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning · USENIX ATC 2021
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
0.512021
Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning · USENIX ATC 2021
Machine learning › Efficient and distributed learning › model compression › quantization
quantized training
0.512021
Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning · USENIX ATC 2021
Machine learning › Generative modeling › diffusion model › image restoration
diffusion-based image restoration
0.312025
DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement · AAAI 2025
Machine learning › Generative modeling
diffusion model
0.312025
DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement · AAAI 2025
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.212024
PASS: Patch Automatic Skip Scheme for Efficient On-Device Video Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › Video understanding and tracking
multi-object tracking
0.212024
PASS: Patch Automatic Skip Scheme for Efficient On-Device Video Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Edge and fog computing › distributed learning
collaborative training
0.212022
Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

contrastive representation learning · 2.8neural codec · 2.6diffusion model · 2.6spatial-temporal bit allocation · 1.5self-supervisory gate optimization · 1.5adversarial perturbation · 1.5self-supervised learning · 1.3learnable gating · 1.3stripe-wise group quantization · 1.1quantization · 1.1loss-aware compensation · 0.5backward quantization · 0.5
YearPublicationVenuePosition
2025 DeNC: Unleash Neural Codecs in Video Streaming with Diffusion Enhancement
abstract
Recent years have witnessed the rise of Neural-enhanced Video Streaming (NeVS), which integrates neural restoration models into video codecs for higher compression-restoration performance. Despite its benefit, existing work has not well explored the full potential of NeVS paradigm, due to: (1) post-streaming restoration by decoder while lacking the proactive collaboration of encoder, (2) end-to-end optimization based on conventional rate-distortion theory, which has been verified that low distortion is not always a synonym for high perceptual quality, and (3) coupled design for domain-specific tasks that cannot generalize to various video codecs. Observing these limitations, our objective is not to incrementally present an improved restoration model. Instead, we focus on the encoder-decoder synergy, i.e., the codec, which is non-trivial since it inherently strikes the rate-distortion-perception trade-off of NeVS. Aiming at this target, we propose the Diffusion-enhanced Neural Codec (DeNC), a plug-and-play module for current NeVS paradigm, to significantly reduce the required bitrates while preserving high perceptual quality of restored videos. Our key design is twofold. First, DeNC improves the encoder's compression efficiency by simultaneously reducing the resolution and color bit-depth of frame referencing. Second, DeNC empowers the decoder with perception-oriented restoration capability by making its diffusion-based restoration process aware of the encoder's compression conditions. Real-world evaluations show that DeNC improves compression ratios with nearly an order of magnitude and achieves much higher restoration quality (e.g., 93+ VMAF and 23% higher MOS) over the latest baselines.
Qihua Zhou, Ruibin Li, Jingcai Guo, Yaodong Huang, Zhenda Xu, Laizhong Cui, Song Guo 0001
AAAI5
2024 On the Robustness of Neural-Enhanced Video Streaming against Adversarial Attacks
abstract
The explosive growth of video traffic on today's Internet promotes the rise of Neural-enhanced Video Streaming (NeVS), which effectively improves the rate-distortion trade-off by employing a cheap neural super-resolution model for quality enhancement on the receiver side. Missing by existing work, we reveal that the NeVS pipeline may suffer from a practical threat, where the crucial codec component (i.e., encoder for compression and decoder for restoration) can trigger adversarial attacks in a man-in-the-middle manner to significantly destroy video recovery performance and finally incurs the malfunction of downstream video perception tasks. In this paper, we are the first attempt to inspect the vulnerability of NeVS and discover a novel adversarial attack, called codec hijacking, where the injected invisible perturbation conspires with the malicious encoding matrix by reorganizing the spatial-temporal bit allocation within the bitstream size budget. Such a zero-day vulnerability makes our attack hard to defend because there is no visual distortion on the recovered videos until the attack happens. More seriously, this attack can be extended to diverse enhancement models, thus exposing a wide range of video perception tasks under threat. Evaluation based on state-of-the-art video codec benchmark illustrates that our attack significantly degrades the recovery performance of NeVS over previous attack methods. The damaged video quality finally leads to obvious malfunction of downstream tasks with over 75% success rate. We hope to arouse public attention on codec hijacking and its defence.
Qihua Zhou, Jingcai Guo, Song Guo 0001, Ruibin Li, Jie Zhang 0076, Zhenda Xu
AAAI7
2024 PASS: Patch Automatic Skip Scheme for Efficient On-Device Video Perception
abstract
Real-time video perception tasks are often challenging on resource-constrained edge devices due to the issues of accuracy drop and hardware overhead, where saving computations is the key to performance improvement. Existing methods either rely on domain-specific neural chips or priorly searched models, which require specialized optimization according to different task properties. These limitations motivate us to design a general and task-independent methodology, called Patch Automatic Skip Scheme (PASS), which supports diverse video perception settings by decoupling acceleration and tasks. The gist is to capture inter-frame correlations and skip redundant computations at patch level, where the patch is a non-overlapping square block in visual. PASS equips each convolution layer with a learnable gate to selectively determine which patches could be safely skipped without degrading model accuracy. Specifically, we are the first to construct a self-supervisory procedure for gate optimization, which learns to extract contrastive representations from frame sequences. The pre-trained gates can serve as plug-and-play modules to implement patch-skippable neural backbones, and automatically generate proper skip strategy to accelerate different video-based downstream tasks, e.g., outperforming state-of-the-art MobileHumanPose in 3D pose estimation and FairMOT in multiple object tracking, by up to 9.43 × and 12.19 × speedups, respectively, on NVIDIA Jetson Nano devices.
Qihua Zhou, Song Guo 0001, Jiacheng Liang, Jingcai Guo, Zhenda Xu, Jingren Zhou 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge Devices
abstract
Real-time video perception tasks are often challenging over the resource-constrained edge devices due to the concerns of accuracy drop and hardware overhead, where saving computations is the key to performance improvement. Existing methods either rely on domain-specific neural chips or priorly searched models, which require specialized optimization according to different task properties. In this work, we propose a general and task-independent Patch Automatic Skip Scheme (PASS), a novel end-to-end learning pipeline to support diverse video perception settings by decoupling acceleration and tasks. The gist is to capture the temporal similarity across video frames and skip the redundant computations at patch level, where the patch is a non-overlapping square block in visual. PASS equips each convolution layer with a learnable gate to selectively determine which patches could be safely skipped without degrading model accuracy. As to each layer, a desired gate needs to make flexible skip decisions based on intermediate features without any annotations, which cannot be achieved by conventional supervised learning paradigm. To address this challenge, we are the first to construct a tough self-supervisory procedure for optimizing these gates, which learns to extract contrastive representation, i.e., distinguishing similarity and difference, from frame sequence. These high-capacity gates can serve as a plug-and-play module for convolutional neural network (CNN) backbones to implement patch-skippable architectures, and automatically generate proper skip strategy to accelerate different video-based downstream tasks, e.g., outperforming the state-of-the-art MobileHumanPose (MHP) in 3D pose estimation and FairMOT in multiple object tracking, by up to 9.43 times and 12.19 times speedups, respectively. By directly processing the raw data of frames, PASS can generalize to real-time video streams on commodity edge devices, e.g., NVIDIA Jetson Nano, with efficient performance in realistic deployment.
Qihua Zhou, Song Guo 0001, Jiacheng Liang, Zhenda Xu, Jingren Zhou 0001
AAAI5
2023 Development of Deep Learning Algorithms for Automated Scoliosis and Abnormal Posture Screening Using 2D Back Image
abstract
Adolescent idiopathic scoliosis is becoming a common spinal disorder among adolescents. The traditional methods of scoliosis screening are labor-intensive and can result in unnecessary referrals and radiological exposure for adolescents due to their low positive predictive value. For early screening of scoliosis and abnormal posture, a mobile-based cost-free, accurate and radiation-free scoliosis screening system is proposed in this paper. We establish a database with labeled 2D unclothed back images and corresponding whole-spine standing posterior-anterior X-ray images, and innovatively propose a new network topology of the 2D back image to localize the back landmarks. With only an unclothed back image, this system can automatically classify normal, abnormal posture and scoliosis with an overall classification accuracy of 88.1%. This system has the potential to overcome the time and space limitations of conventional screening for scoliosis and abnormal posture.
Zhenda Xu, Donghua Hang, Qihua Zhou, Song Guo 0001, Aiqian Gan
ICME1
2022 2D Photogrammetry Image of Adolescent Idiopathic Scoliosis Screening Using Deep Learning
Zhenda Xu, Jiazi Ouyang, Aiqian Gan, Qihua Zhou, Song Guo 0001
ISBRA1
2022 Hierarchical Channel-spatial Encoding for Communication-efficient Collaborative Learning
abstract
It witnesses that the collaborative learning (CL) systems often face the performance bottleneck of limited bandwidth, where multiple low-end devices continuously generate data and transmit intermediate features to the cloud for incremental training. To this end, improving the communication efficiency by reducing traffic size is one of the most crucial issues for realistic deployment. Existing systems mostly compress features at pixel level and ignore the characteristics of feature structure, which could be further exploited for more efficient compression. In this paper, we take new insights into implementing scalable CL systems through a hierarchical compression on features, termed Stripe-wise Group Quantization (SGQ). Different from previous unstructured quantization methods, SGQ captures both channel and spatial similarity in pixels, and simultaneously encodes features in these two levels to gain a much higher compression ratio. In particular, we refactor feature structure based on inter-channel similarity and bound the gradient deviation caused by quantization, in forward and backward passes, respectively. Such a double-stage pipeline makes SGQ hold a sublinear convergence order as the vanilla SGD-based optimization. Extensive experiments show that SGQ achieves a higher traffic reduction ratio by up to 15.97 times and provides 9.22 times image processing speedup over the uniform quantized training, while preserving adequate model accuracy as FP32 does, even using 4-bit quantization. This verifies that SGQ can be applied to a wide spectrum of edge intelligence applications.
Qihua Zhou, Song Guo 0001, Yi Liu 0057, Jie Zhang 0076, Jiewei Zhang, Tao Guo 0004, Zhenda Xu, Zhihao Qu
NeurIPS7
2021 Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device Learning
Qihua Zhou, Song Guo 0001, Zhihao Qu, Jingcai Guo, Zhenda Xu, Jiewei Zhang, Tao Guo 0004, Boyuan Luo, Jingren Zhou 0001
USENIX ATC5
2021 On-Device Learning Systems for Edge Intelligence: A Software and Hardware Synergy Perspective
abstract
Modern machine learning (ML) applications are often deployed in the cloud environment to exploit the computational power of clusters. However, this in-cloud computing scheme cannot satisfy the demands of emerging edge intelligence scenarios, including providing personalized models, protecting user privacy, adapting to real-time tasks, and saving resource cost. In order to conquer the limitations of conventional in-cloud computing, there comes the rise of on-device learning, which makes the end-to-end ML procedure totally on user devices, without unnecessary involvement of the cloud. In spite of the promising advantages of on-device learning, implementing a high-performance on-device learning system still faces with many severe challenges, such as insufficient user training data, backward propagation (BP) blocking, and limited peak processing speed. Observing the substantial improvement space in the implementation and acceleration of on-device learning systems, we intend to present a comprehensive analysis of the latest research progress and point out potential optimization directions from the system perspective. This survey presents a software and hardware synergy of on-device learning techniques, covering the scope of model-level neural network design, algorithm-level training optimization, and hardware-level instruction acceleration. We hope this survey could bring fruitful discussions and inspire the researchers to further promote the field of edge intelligence.
Qihua Zhou, Zhihao Qu, Song Guo 0001, Boyuan Luo, Jingcai Guo, Zhenda Xu, Rajendra Akerkar
IEEE Internet Things J.6