EDBT 2026 Demo / reviewers in the wild / expert
Youbing Hu
dblp:193/3081
· DBLP profile ↗
17ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-2181-9659ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TeamTTA: Efficient Multi-Device Collaboration for Open-Set Test-Time Adaptation via Cloud IntegrationabstractDeep neural networks (DNNs) deployed on edge devices often suffer from severe performance degradation when exposed to dynamic and continually shifting environments. Test-time adaptation (TTA) has emerged as a promising solution by updating models online with incoming test data. However, edge deployment poses unique challenges: limited computational resources, latency caused by adaptation delays, and knowledge isolation across devices. The situation becomes even more complex in open-world scenarios, where the presence of unknown categories further disrupts adaptation. To overcome these limitations, we propose TeamTTA, a cloud-integrated framework designed for efficient multi-device collaboration open-set test-time adaptation. Specifically, TeamTTA aggregates reliable samples from multiple edge devices through crowdsourcing, uploads them to the cloud, and maintains a memory buffer for continual adaptation. A large vision model (LVM) in the cloud leverages its zero-shot generalization ability to filter out open-set samples and acts as a teacher model, distilling its knowledge into a replicated student edge model stored in the cloud. The adapted model parameters, or alternatively global statistics under poor network conditions, are then transmitted back to the edge devices for efficient inference. Extensive experiments on standard public TTA benchmarks, including corrupted and open-set datasets, show that TeamTTA achieves superior adaptation accuracy, robustness to distribution shifts, and communication efficiency, outperforming state-of-the-art TTA baselines. These results validate the effectiveness of integrating cloud-edge collaboration and LVM-driven knowledge distillation for real-world edge intelligence. Anqi Lu, Youbing Hu, Dawei Wei, Zhiqiang Cao 0001, Jie Liu 0001, Zhijun Li 0002 |
J. Artif. Intell. Res. | 2 |
| 2025 | FoCTTA: Low-Memory Continual Test-Time Adaptation with FocusabstractContinual adaptation to domain shifts at test time (CTTA) is crucial for enhancing the intelligence of deep learning enabled IoT applications. However, prevailing CTTA methods, which typically update all batch normalization (BN) layers, exhibit two memory inefficiencies. First, the reliance on BN layers for adaptation necessitates large batch sizes, leading to high memory usage. Second, updating all BN layers requires storing the activations of all BN layers for backpropagation, exacerbating the memory demand. Both factors lead to substantial memory costs, making existing solutions impractical for IoT devices. In this paper, we present FoCTTA, a low-memory CTTA strategy. The key is to automatically identify and adapt a few drift-sensitive representation layers, rather than blindly update all BN layers. The shift from BN to representation layers eliminates the need for large batch sizes. Also, by updating adaptation-critical layers only, FoCTTA avoids storing excessive activations. This focused adaptation approach ensures that FoCTTA is not only memory-efficient but also maintains effective adaptation. Evaluations show that FoCTTA improves the adaptation accuracy over the state-of-the-arts by 4.5%, 4.9%, and 14.8% on CIFAR10-C, CIFAR100-C, and ImageNet-C under the same memory constraints. Across various batch sizes, FoCTTA reduces the memory usage by 3-fold on average, while improving the accuracy by 8.1%, 3.6%, and 0.2%, respectively, on the three datasets. Youbing Hu, Zimu Zhou, Anqi Lu, Zhiqiang Cao 0001, Zhijun Li 0002 |
ICME | 1 |
| 2025 | Edge-Cloud Collaborated Object Detection via Bandwidth Adaptive Difficult-Case DiscriminatorabstractObject detection, a fundamental task in computer vision, is crucial for various intelligent edge computing applications. However, object detection algorithms are usually heavy in computation, hindering their deployments on resource-constrained edge devices. Traditional edge-cloud collaboration schemes, like deep neural network (DNN) partitioning across edge and cloud, are unfit for object detection due to the significant communication costs incurred by the large size of intermediate results. To this end, we propose a Difficult-Case based Small-Big model (DCSB) framework. It employs a difficult-case discriminator on the edge device to control data transfer between the small model on the edge and the large model in the cloud. We also adopt regional sampling to further reduce the bandwidth consumption and create a discriminator zoo to accommodate the varying networking conditions. Additionally, we extend DCSB to video tasks by developing an adaptive sampling rate update algorithm, aiming to minimize computational demands without sacrificing detection accuracy. Extensive experiments show that DCSB can detect 97.26%-97.96% objects while saving 74.37%-82.23% network bandwidth, compared to cloud-only methods. Furthermore, DCSB significantly outperforms the latest DNN partitioning methods, reducing inference time by 92.60%-95.10% given an 8Mbps transmission bandwidth. In video tasks, DCSB matches the detection accuracy of leading video analysis methods while cutting the computational overhead by 40%. Zhiqiang Cao 0001, Zimu Zhou, Yongrui Chen 0001, Youbing Hu, Anqi Lu, Jie Liu 0001, Zhijun Li 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Enhancing Remote Sensing Image Scene Classification With Satellite-Terrestrial Collaboration and Attention-Aware Transmission PolicyabstractAdvancements in Earth observation sensors on low Earth orbit (LEO) satellites have significantly increased the volume of remote sensing images. This growth has led to challenges such as higher storage demands, downlink bandwidth stress, and transmission delays, particularly for real-time remote sensing image scene classification (RSISC). To address this, we propose a novel Satellite-Terrestrial Collaborative Scene Classification (STCSC) framework that integrates transmission and computation. The framework employs an attention-aware policy on the satellite, which adaptively determines the sequence of images and selection of image blocks for transmission, as well as these blocks' sampling rates. This policy is based on image complexity and the real-time data transmission rate, prioritizing blocks crucial for downstream tasks. On the ground, a classification model processes the received image blocks, balancing classification accuracy and transmission delay. Moreover, we have developed a comprehensive simulation system to validate the performance of our framework, including simulations of the satellite, transmission, and ground modules. Simulation results demonstrate that our STCSC framework can reduce transmission delay by 76.6% while enhancing classification accuracy on the ground by 0.6%. Additionally, our attention-aware policy is compatible with any ground classification model. Anqi Lu, Youbing Hu, Zhiqiang Cao 0001, Jie Liu 0001, Lingzhi Li 0001, Zhijun Li 0002 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image RecognitionabstractThe Vision Transformer (ViT) excels in accuracy when handling high-resolution images, yet it confronts the challenge of significant spatial redundancy, leading to increased computational and memory requirements. To address this, we present the Localization and Focus Vision Transformer (LF-ViT). This model operates by strategically curtailing computational demands without impinging on performance. In the Localization phase, a reduced-resolution image is processed; if a definitive prediction remains elusive, our pioneering Neighborhood Global Class Attention (NGCA) mechanism is triggered, effectively identifying and spotlighting class-discriminative regions based on initial findings. Subsequently, in the Focus phase, this designated region is used from the original image to enhance recognition. Uniquely, LF-ViT employs consistent parameters across both phases, ensuring seamless end-to-end optimization. Our empirical tests affirm LF-ViT's prowess: it remarkably decreases Deit-S's FLOPs by 63% and concurrently amplifies throughput twofold. Code of this project is at https://github.com/edgeai1/LF-ViT.git. Youbing Hu, Anqi Lu, Zhiqiang Cao 0001, Dawei Wei, Jie Liu 0001, Zhijun Li 0002 |
AAAI | 1 |
| 2024 | Optimized Click Prediction on Mobile Devices via Device-Cloud SynergyabstractThe rapid growth of deep learning-based services and applications underscores the need for efficient neural network model deployment. Traditional cloud-centric solutions, despite their computational power, face significant challenges such as high energy consumption, network transmission delays, and user privacy concerns. Conversely, performing high-performance inference on resource-constrained mobile devices, especially for tasks like advertising click prediction, presents its own set of difficulties. To address these challenges, we propose a device-cloud collaboration system utilizing a difficult-case discriminator. This system classifies input samples based on semantic information into difficult and simple cases. Difficult cases are processed in the cloud using a large model, while simple cases are handled on the device by a smaller model. This approach maximizes system resources and protects user privacy. Evaluations on public datasets show that our system significantly outperforms other advertisement methods in click prediction accuracy and uploading efficiency. Compared to the device-only approach, our system improves the area under the curve (AUC) by 8.9%, and compared to the cloud-centric approach, it reduces the upload ratio by 34%. Moreover, deploying our system on a specific smartphone demonstrates substantial improvements in private real datasets. Shuyuan Pan, Anqi Lu, Youbing Hu, Lingzhi Li 0001, Zhijun Li 0002 |
ICPADS | 3 |
| 2024 | GlareShell: Graph learning-based PHP webshell detection for web server of industrial internet
Pengbin Feng, Dawei Wei, Qiaoyang Li, Youbing Hu, Ning Xi 0002 |
Comput. Networks | 5 |
| 2024 | Using Physical Dynamics: Accurate and Real-Time Object Detection for High-Resolution Video Streaming on Internet of Things DevicesabstractObject detection is crucial in video analytics pipelines, but there is a need to optimize deep neural networks (DNNs)-based object detection for resource-constrained Internet of Things (IoT) devices devices. The computational constraints inherent to the IoT device inevitably curtail its precision and real-time efficacy in the domain of object detection, with pronounced challenges arising, particularly when confronted with high-resolution video streams. To overcome these limitations, we propose UPD (Using Physical Dynamics), a novel on-device system that enables real-time and accurate object detection for high-resolution video streams. UPD employs a lightweight tracking algorithm for the detection of the majority of video frames, concurrently executing the object detector in a parallel fashion only in select instances. UPD addresses tracking errors by eliminating inaccurate feature points and correcting tracking results using physical information about the object. Unlike previous approaches that depend solely on the high-latency object detector to offset errors, our method is unaffected by the video resolution level. Extensive experiments demonstrate that UPD facilitates real-time analysis of high-resolution videos on IoT devices and significantly improves the overall accuracy (mIoU) compared to state-of-the-art DBT (Detection-Based-Tracking) frameworks, achieving a 100% accuracy improvement on three commonly used datasets. A video demo can be found at https://youtu.be/gKRQPHJ6gmY. Zhiqiang Cao 0001, Youbing Hu, Anqi Lu, Jie Liu 0001, Zhijun Li 0002 |
IEEE Internet Things J. | 3 |
| 2024 | RAPNet: Resolution-Adaptive and Predictive Early Exit Network for Efficient Image RecognitionabstractDeploying compute-intensive deep neural networks (DNNs) on resource-constrained end devices has become a prominent trend, enabling localized intelligence. However, efficiently deploying these DNNs at scale poses challenges. To address this, extensive research has focused on the early exit architecture based on convolutional neural networks (CNNs), which dynamically adapt network depth to reduce inference computation. Nevertheless, the sequential execution of all internal classifiers (ICs) and subsequent termination based on an exit criterion is inefficient. Motivated by these insights, we introduce a resolution-adaptive prediction network (RAPNet) architecture. RAPNet comprises a lightweight prediction network that captures global image features and an inference network integrated with an early exit architecture. The prediction network accurately determines the optimal IC position conditioned on the input images for efficient image classification. Additionally, we incorporate resolution-adaptive inference and feature fusion mechanisms by computational reuse, to effectively mitigate image spatial redundancy and improve the accuracy of ICs. We conduct extensive experiments across various data sets and architectures to demonstrate that RAPNet achieves a significantly better accuracy versus computational tradeoff than other recently proposed early exit methods. For instance, when using MobileNet as the base network, RAPNet achieves significant accuracy improvements of 12% and 5.7% on the Tiny Imagenet and CIFAR-100 data sets, respectively, surpassing other early exit methods with similar computational constraints. Youbing Hu, Zimu Zhou, Zhiqiang Cao 0001, Anqi Lu, Jie Liu 0001, Min Zhang 0005, Zhijun Li 0002 |
IEEE Internet Things J. | 1 |
| 2024 | Patching in Order: Efficient On-Device Model Fine-Tuning for Multi-DNN Vision ApplicationsabstractThe increasing deployment of multiple deep neural networks (DNNs) on edge devices is revolutionizing mobile vision applications, spanning autonomous vehicles, augmented reality, and video surveillance. These applications demand adaptation to contextual and environmental drifts, typically through fine-tuning on edge devices without cloud access, due to increasing data privacy concerns and the urgency for timely responses. However, fine-tuning multiple DNNs on edge devices faces significant challenges due to the substantial computational workload. In this paper, we present PatchLine, a novel framework tailored for efficient on-device training in the form of fine-tuning for multi-DNN vision applications. At the core of PatchLine is an innovative lightweight adapter design called patches coupled with a strategic patch updating approach across models. Specifically, PatchLine adopts drift-adaptive incremental patching, correlation-aware warm patching, and entropy-based sample selection, to holistically reduce the number of trainable parameters, training epochs, and training samples. Experiments on four datasets, three vision tasks, four backbones, and two platforms demonstrate that PatchLine reduces the total computational cost by an average of 55% without sacrificing accuracy compared to the state-of-the-art. Zhiqiang Cao 0001, Zimu Zhou, Anqi Lu, Youbing Hu, Jie Liu 0001, Min Zhang 0005, Zhijun Li 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Edge-Cloud Collaborated Object Detection via Difficult-Case DiscriminatorabstractAs one of the basic tasks of computer vision, object detection has been widely used in many intelligent applications. However, object detection algorithms are usually heavyweight in computation, hindering their implementations on resource-constrained edge devices. Current edge-cloud collaboration methods, such as CNN partition over edge-cloud devices, are not suitable for object detection since the large data size of the intermediate results will introduce extravagant communication costs. To address this challenge, we propose a difficult-case based small-big model (DCSB) framework that deploys a difficult-case discriminator on the edge device to control the data transfer between the small model (edge) and the big model (cloud). Upon receiving data, the edge device operates a difficult-case discriminator to classify images into easy cases and difficult cases according to the specific semantics of the images. The difficult cases will be uploaded to the cloud. To reduce bandwidth consumption, we propose a regional sampling method that adaptively down-samples some regions of the difficult case to reduce the amount of transferred data based on the primary results of the lightweight model. Experimental results on VOC, COCO, and HELMET datasets using two object detection algorithms demonstrate that DCSB can detect 93.77%-97.05% objects but save 77.19% -80.55% of network bandwidth compared with the cloud-only method, while the edge-only method can only detect 54.90%-68.28% objects in the same condition. In addition, compared with the state-of-the-art model partition method - CAS, DCSB saves 95.19%-95.80% of the inference time when the transmission bandwidth is 8Mbps. Zhiqiang Cao 0001, Zhijun Li 0002, Yongrui Chen 0001, Youbing Hu, Jie Liu 0001 |
ICDCS | 5 |
| 2023 | Content-Aware Adaptive Device-Cloud Collaborative Inference for Object DetectionabstractMany intelligent applications based on deep neural networks (DNNs) are increasingly running on Internet of Things (IoT) devices. Unfortunately, the computing resources of these IoT devices are limited, which will seriously hinder the widespread deployment of various smart applications. A popular solution is to offload part of computation tasks from the IoT device to cloud by way of device–cloud collaboration. However, existing collaboration approaches may suffer from long network transmission delay or degraded accuracy due to the large amount of intermediate results, bring enormous challenges to the tasks, such as object detection, that require massive computing resources. In this article, we propose an efficient device–cloud collaborative inference (DCCI) object detection framework, which dynamically adjusts the amount of transferred data according to the content of input images. Specifically, a content-aware hard-case discriminator is proposed to automatically classify the input images as hard-cases or simple-cases, the hard-cases are uploaded to the cloud to be processed by a deployed heavyweight model, and the simple cases are processed by a lightweight model deployed to the IoT device, where the lightweight model is automatically compressed based on reinforcement learning according to the resource constraints of the IoT device. Furthermore, a collaborative scheduler based on the runtime load and network transmission capability of IoT devices is proposed to optimize the collaborative computation between IoT devices and the cloud. Extensive experimental evaluations show that compared to the Device-only approach, DCCI can reduce the memory footprint and compute resources of IoT devices by more than 90.0% and 30.87%, respectively. Compared to Cloud-centric, DCCI can save$2.0\times $of network bandwidth. In addition, compared with the state-of-the-art DNN partitioning method, DCCI can save$1.2\times $of inference latency, and$1.3\times $of IoT device energy consumption with the same accuracy constraint. Youbing Hu, Zhijun Li 0002, Yongrui Chen 0001, Zhiqiang Cao 0001, Jie Liu 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Satellite-Terrestrial Collaborative Object Detection via Task-Inspired FrameworkabstractRecently, buoyed by advances in the space industry, low Earth orbit (LEO) satellites have become an important part of the Internet of Things (IoT). LEO satellites have entered the era of a big data link with IoT, how to deal with the data from the satellite IoT is a problem worthy of consideration. Conventional object detection method in optical remote sensing simply transmits the raw data to the ground. However, it ignores the properties of the images and the connection with the downstream task. To obtain efficient data transmission and accurate object detection, we propose a task-inspired satellite–terrestrial collaborative object detection framework called STCOD. It detects regions of interest (ROIs) and adopts a block-based adaptive sampling method to compress the background (BG) in optical remote sensing images by introducing satellite edge computing (SEC) on satellites. The STCOD framework also sets the transmission priority of image blocks according to their contributions to the task and uses fountain code to ensure the reliable transmission of important image blocks. We build a whole software simulation framework to validate our method, including the satellite module, the transmission module, and the terrestrial module. Extensive experimental results show that the STCOD framework can reduce the amount of downlink data decreased by 50.04% while losing the detection accuracy by 0.54%. In our simulated satellite–terrestrial link, the STCOD framework can reduce the number of satellite-to-terrestrial transmissions by half. When the packet loss rate is between 5% and 20%, the detection accuracy is lost only 0.05% to 0.5%. Anqi Lu, Youbing Hu, Zhiqiang Cao 0001, Yongrui Chen 0001, Zhijun Li 0002 |
IEEE Internet Things J. | 3 |
| 2018 | Multi-Constrained Routing Based on Particle Swarm Optimization and Fireworks AlgorithmabstractThis paper sets up a mathematical model that satisfies the multiconstrained routing optimization problem. By adding a penalty, multiple constraints are mapped to a fitness that satisfies multiple constraints. Then, it uses a heeristic routing algorithm based on particle swarm optimization (PSO) to perform heuristic routing search. Introducing the fireworks algorithm (FWA) based on the PSO search algorithm, our algorithm searches the optimal solution more quickly. Besides, it reduces the defect of PSO falling into the local optimum. Simulation shows the algorithm can effectively solve the multiconstrained routing problem in large-scale networks. While searching for optimal solutions, the success rate of the algorithm is about 5.21% higher than that of the standard PSO algorithm. That is improved by using the ant colony algorithm. The PSO-ACO algorithm is about 2.57% higher than the problem. The average cost of the final search is about 4.36% higher than that of the standard PSO algorithm. It is about 1.34% higher than the PSO-ACO algorithm improved by the ant colony algorithm. Youbing Hu, Jinjiang Wan, Kaidong Wang |
IECON | 1 |
| 2018 | Multi-Constrained Routing Optimization Algorithm Based on DAGabstractA new multiconstrained routing algorithm based on quality of Service (QoS), DAG_DMCOP, is proposed. The algorithm is divided into two parts: (1) Pruning strategy. Find all paths that meet the needs of multiple constraints and convert the network topology into a Directed Acyclic Graph (DAG). (2) Search strategy. The link synthesis cost function is introduced to adjust the multiconstrained conditions (bandwidth, delay, delay jitter and other QoS parameters) adaptively. Then an improved Dijkstra algorithm is used to find out the optimal path that meets the needs of multiple constraints. Simulation show the algorithm can quickly find the path with the least cost and is suitable for large-scale multiconstrained routing networks. It is an efficient new algorithm for solving multiconstrained routing problems. Kaidong Wang, Jinjiang Wang, Youbing Hu, Shuangqin Wang |
IECON | 5 |
| 2017 | A new hybrid data-driven model for event-based rainfall-runoff simulation
Guangyuan Kan, Jiren Li, Xingnan Zhang, Liuqian Ding, Xiaoyan He, Ke Liang 0004, Xiaoming Jiang, Minglei Ren, Zhongbo Zhang, Youbing Hu |
Neural Comput. Appl. | 12 |
| 2017 | A multi-core CPU and many-core GPU based fast parallel shuffled complex evolution global optimization approachabstractIn the field of hydrological modelling, the global and automatic parameter calibration has been a hot issue for many years. Among automatic parameter optimization algorithms, the shuffled complex evolution developed at the University of Arizona (SCE-UA) is the most successful method for stably and robustly locating the global “best” parameter values. Ever since the invention of the SCE-UA, the profession suddenly has a consistent way to calibrate watershed models. However, the computational efficiency of the SCE-UA significantly deteriorates when coping with big data and complex models. For the purpose of solving the efficiency problem, the recently emerging heterogeneous parallel computing (parallel computing by using the multi-core CPU and many-core GPU) was applied in the parallelization and acceleration of the SCE-UA. The original serial and proposed parallel SCE-UA were compared to test the performance based on the Griewank benchmark function. The comparison results indicated that the parallel SCE-UA converged much faster than the serial version and its optimization accuracy was the same as the serial version. It has a promising application prospect in the field of fast hydrological model parameter optimization. Guangyuan Kan, Tianjie Lei, Ke Liang 0004, Jiren Li, Liuqian Ding, Xiaoyan He, Depeng Zuo, Zhenxin Bao, Mark Amo-Boateng, Youbing Hu, Mengjie Zhang 0003 |
IEEE Trans. Parallel Distributed Syst. | 12 |