EDBT 2026 Demo / reviewers in the wild / expert
Boyu Diao
dblp:161/2139
· DBLP profile ↗
29ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0002-8360-7718ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Group-TopK: Optimizing Distributed Training on Edge Devices via Communication Compression
Yifan Wang 0005, Xiaohui Peng 0002, Haohao Ma, Hui Sun 0002, Deke Guo, Boyu Diao |
HPDC | 6 |
| 2025 | MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsabstractDiffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse. Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Renshuai Tao, Yongjun Xu 0001, Michele Magno |
AAAI | 6 |
| 2025 | A Sensitivity-Driven Expert Allocation Method in LoRA-MoE for Efficient Fine-TuningabstractAs deep learning models expand, the pre-training-fine-tuning paradigm has become the standard approach for handling various downstream tasks. However, shared parameters can lead to diminished performance when dealing with complex datasets involving multiple tasks. While introducing Mixture-of-Experts (MoE) methods has alleviated this issue to some extent, it also significantly increases the number of parameters required for fine-tuning and training time, introducing greater parameter redundancy. To address these challenges, we propose a method for allocating expert numbers based on parameter sensitivity-LoRA-SMoE (A Sensitivity-Driven Expert Allocation Method in LoRA-MoE for Efficient Fine-Tuning). This method rapidly assesses the sensitivity of different tasks to parameters by sampling a small amount of data and using gradient information. It then adaptively allocates expert numbers within a given budget. The process maintains comparable memory consumption to LoRA (Low-Rank Adaptation) while ensuring an efficient and resource-friendly fine-tuning procedure. Experimental results demonstrate that compared to SOTA fine-tuning methods, our LoRA-SMoE approach can enhance model performance while reducing the number of trainable parameters. This significantly improves model performance in resource-constrained environments. Additionally, due to its efficient parameter sensitivity evaluation mechanism, LoRA-SMoE requires minimal computational overhead to optimize expert allocation, making it particularly suitable for scenarios with limited computational resources. All the code in this study will be made publicly available following the acceptance of the paper for publication. Source code is at https://github.com/EMLS-ICTCAS/LoRA-SMoE Junzhou Xu, Boyu Diao, Chengqiang Qi, Shaobo Zhao, Yongjun Xu 0001 |
CCGrid | 2 |
| 2025 | A Nonlinear Hash-Based Optimization Method for SpMV on GPUsabstractSparse matrix-vector multiplication (SpMV) is a fundamental operation with a wide range of applications in scientific computing and artificial intelligence. However, the large scale and sparsity of sparse matrix often make it a performance bottleneck. In this paper, we highlight the effectiveness of hash-based techniques in optimizing sparse matrix reordering, introducing the Hash-based Partition (HBP) format, a lightweight SpMV approach. HBP retains the performance benefits of the 2Dpartitioning method while leveraging the hash transformation's ability to group similar elements, thereby accelerating the preprocessing phase of sparse matrix reordering. Additionally, we achieve parallel load balancing across matrix blocks through a competitive method. Our experiments, conducted on both Nvidia Jetson AGX Orin and Nvidia RTX 4090, show that in the preprocessing step, our method offers an average speedup of 3.53 times compared to the sorting approach and 3.67 times compared to the dynamic programming method employed in Regu2D. Furthermore, in SpMV, our method achieves a maximum speedup of 3.32 times on Orin and 3.01 times on RTX4090 against the CSR format in sparse matrices from the University of Florida Sparse Matrix Collection. Boyu Diao, Hangda Liu, Zhulin An, Yongjun Xu 0001 |
CCGrid | 2 |
| 2025 | IOR: Inversed Objects Replay for Incremental Object DetectionabstractExisting Incremental Object Detection (IOD) methods partially alleviate catastrophic forgetting when incrementally detecting new objects in real-world scenarios. However, many of these methods rely on the assumption that unlabeled old-class objects may co-occur with labeled new-class objects in the incremental data. When unlabeled old-class objects are absent, the performance of existing methods tends to degrade. The absence can be mitigated by generating old-class samples, but it incurs high costs. This paper argues that previous generation-based IOD suffers from redundancy, both in the use of generative models, which require additional training and storage, and in the overproduction of generated samples, many of which do not contribute significantly to performance improvements. To eliminate the redundancy, we propose Inversed Objects Replay (IOR). Specifically, we generate old-class samples by inversing the original detectors, thus eliminating the necessity of training and storing additional generative models. We propose augmented replay to reuse the objects in generated samples, reducing redundant generations. Moreover, we propose high-value knowledge distillation focusing on the positions of old-class objects overwhelmed by the background, which transfers the knowledge to the incremental detector. Extensive experiments conducted on MS COCO 2017 demonstrate that our method can efficiently improve detection performance in IOD scenarios with the absence of old-class objects. Zijia An, Boyu Diao, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
ICASSP | 2 |
| 2025 | Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersabstractDiffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters.
Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present **Q-VDiT**, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the *Token aware Quantization Estimator* (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce *Temporal Maintenance Distillation* (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency score of 23.40, setting a new benchmark and outperforming the current state-of-the-art quantization methods by **1.9$\times$**. Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Zhulin An, Libo Huang 0001, Boyu Diao, Zixiang Zhao, Yongjun Xu 0001, Michele Magno |
ICML | 8 |
| 2025 | Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learningabstract3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D FSCIL approaches primarily utilize multimodal pre-trained models to extract the semantic features, heavily dependent on meticulously designed high-quality prompts and fine-tuning strategies. To reduce this dependence, this paper proposes a novel method for **3D** **F**SCI**L** with **E**mbedded **G**eometric features (**3D-FLEG**). Specifically, 3D-FLEG develops a point cloud *geometric feature extraction module* to capture category-related geometric characteristics. To address the modality heterogeneity issues that arise from integrating geometric and text features, 3D-FLEG introduces a *geometric feature embedding module*. By augmenting text prompts with spatial geometric features through these modules, 3D-FLEG can learn robust representations of new categories even with limited samples, while mitigating forgetting of the previously learned categories. Experiments conducted on several publicly available 3D point cloud datasets, including ModelNet, ShapeNet, ScanObjectNN, and CO3D, demonstrate 3D-FLEG's superiority over existing state-of-the-art 3D FSCIL methods. Code is available at https://github.com/lixiangqi707/3D-FLEG. Xiangqi Li, Libo Huang 0001, Zhulin An, Weilun Feng, Chuanguang Yang, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ICML | 6 |
| 2025 | Gensor: A Graph-Based Construction Tensor Compilation Method for Deep LearningabstractHigh-performance deep learning depends on efficient tensor programs. In recent years, automatic tensor program optimization, also known as tensor compilation, has emerged as the primary approach to generating efficient tensor programs. However, how to generate kernels with higher performance in a shorter time is still the key challenge. In this paper, we present Gensor, a graph-based construction tensor compilation method for deep learning, to further improve the performance of construction tensor compilation. Unlike existing tree-based methods, Gensor abstracts construction space into a graph structure. Gensor then explores the construction space with Markov analysis. Gensor takes tensor programs as states and models scheduling primitives as transition actions between these states. Therefore, the process of tensor program construction optimization is abstracted as a graph traversal process. This approach expands the optimization space, improving operator performance while ensuring rapid optimization. Extensive experiments with typical operators demonstrate that Gensor significantly outperforms the state-of-the-art methods on GPUs for both cloud servers and edge devices. As a result, Gensor can generate operator kernels in seconds, with performance increasing by 18 % on average, reaching a maximum of 30 %. It also achieves high speedup for end-to-end models like ResNet50 and GPT-2, with an average acceleration of 20 %. Hangda Liu, Boyu Diao, Xiaohui Peng 0002, Yongjun Xu 0001 |
IPDPS | 2 |
| 2025 | BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting ModelsabstractThe advent of universal time series forecasting models has revolutionized zero-shot forecasting across diverse domains, yet the critical role of data diversity in training these models remains underexplored. Existing large-scale time series datasets often suffer from inherent biases and imbalanced distributions, leading to suboptimal model performance and generalization. To address this gap, we introduce BLAST, a novel pre-training corpus designed to enhance data diversity through a balanced sampling strategy. First, BLAST incorporates 321 billion observations from publicly available datasets and employs a comprehensive suite of statistical metrics to characterize time series patterns. Then, to facilitate pattern-oriented sampling, the data is implicitly clustered using grid-based partitioning. Furthermore, by integrating grid sampling and grid mixup techniques, BLAST ensures a balanced and representative coverage of diverse patterns. Experimental results demonstrate that models pre-trained on BLAST achieve state-of-the-art performance with a fraction of the computational resources and training tokens required by existing methods. Our findings highlight the pivotal role of data diversity in improving both training efficiency and model performance for the universal forecasting task. Zezhi Shao, Yujie Li 0008, Fei Wang 0014, Chengqing Yu, Yisong Fu, Tangwen Qian, Bin Xu 0019, Boyu Diao, Yongjun Xu 0001, Xueqi Cheng 0001 |
KDD (2) | 8 |
| 2025 | On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series ForecastingabstractTransformers have gained attention in atmospheric time series forecasting (ATSF) for their ability to capture global spatial-temporal correlations. However, their complex architectures lead to excessive parameter counts and extended training times, limiting their scalability to large-scale forecasting. In this paper, we revisit ATSF from a theoretical perspective of atmospheric dynamics and uncover a key insight: spatial-temporal position embedding (STPE) can inherently model spatial-temporal correlations even without attention mechanisms. Its effectiveness arises from integrating geographical coordinates and temporal features, which are intrinsically linked to atmospheric dynamics. Based on this, we propose **STELLA**, a **S**patial-**T**emporal knowledge **E**mbedded **L**ightweight mode**L** for ASTF, utilizing only STPE and an MLP architecture in place of Transformer layers. With 10k parameters and one hour of training, STELLA achieves superior performance on five datasets compared to other advanced methods. The paper emphasizes the effectiveness of spatial-temporal knowledge integration over complex architectures, providing novel insights for ATSF. Yisong Fu, Fei Wang 0014, Zezhi Shao, Boyu Diao, Lin Wu 0006, Zhulin An, Chengqing Yu, Yujie Li 0008, Yongjun Xu 0001 |
NeurIPS | 4 |
| 2025 | A resource-aware workload scheduling method for unbalanced GEMMs on GPUsabstractAbstract GEMM (General Matrix Multiplication) serves as a fundamental operator for deep learning computations. Especially in attention-based deep learning models, such as Bert, GPT, and SAM, the sizes of matrices involved in GEMMs exhibit an unbalanced distribution due to the variable input, resulting in the low utilization of hardware resources. To address the issue, this paper proposes inserting a novel GEMM processing layer into the deep learning inference stack and using an adaptive load balancing method to partition and schedule GEMM computation tasks. The method is implemented with hardware runtime resource information, such as the occupancy of computing units, etc. Experiment results show the remarkable performance of our method in unbalanced input GEMM scenarios, achieving an average performance improvement of 2.3x. The method also performs well in attention-based models (GPT-2 and SAM), achieving an average inference speed improvement of 1.1x. These findings highlight the effectiveness of resource-aware algorithm optimization, especially for computation task scheduling. Hangda Liu, Boyu Diao, Yongjun Xu 0001 |
Comput. J. | 2 |
| 2025 | Low-redundancy distillation for continual learning
Boyu Diao, Libo Huang 0001, Zijia An, Hangda Liu, Zhulin An, Yongjun Xu 0001 |
Pattern Recognit. | 2 |
| 2024 | eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented GenerationabstractClass incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) to train a feature extractor incrementally, or store prototypes (features of learned tasks) to estimate the incremental feature distribution. However, the stored exemplars would violate the data privacy concerns, while the fixed prototypes might not reasonably be consistent with the incremental feature distribution, hindering the exploration of real-world CIL applications. In this paper, we propose a data-free CIL method with embedding distillation and Task-oriented generation (eTag), which requires neither exemplar nor prototype. Embedding distillation prevents the feature extractor from forgetting by distilling the outputs from the networks' intermediate blocks. Task-oriented generation enables a lightweight generator to produce dynamic features, fitting the needs of the top incremental classifier. Experimental results confirm that the proposed eTag considerably outperforms state-of-the-art methods on several benchmark datasets. Libo Huang 0001, Yan Zeng 0002, Chuanguang Yang, Zhulin An, Boyu Diao, Yongjun Xu 0001 |
AAAI | 5 |
| 2024 | CLIP-KD: An Empirical Study of CLIP Model DistillationabstractContrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by a large teacher CLIP model. We propose several distillation strategies, including relation, feature, gradient and contrastive paradigms, to examine the effectiveness of CLIP-Knowledge Distillation (KD). We show that a simple feature mimicry with Mean Squared Error loss works surprisingly well. Moreover, interactive contrastive learning across teacher and student encoders is also effective in performance improvement. We explain that the success of CLIP-KD can be attributed to maximizing the feature similarity between teacher and student. The unified method is applied to distill several student models trained on CC3M+12M. CLIP-KD improves student CLIP models consistently over zero-shot ImageNet classification and cross-modal retrieval bench-marks. When using ViT-U14 pretrained on Laion-400M as the teacher, CLIP-KD achieves 57.5% and 55.4% zero-shot top-1 ImageNet accuracy over ViT-B/16 and ResNet-50, surpassing the original CLIP without KD by 20.5% and 20.1% margins, respectively. Our code is released on https://github.com/winycg/CLIP-KD. Chuanguang Yang, Zhulin An, Libo Huang 0001, Junyu Bi, Xinqiang Yu, Boyu Diao, Yongjun Xu 0001 |
CVPR | 7 |
| 2024 | SCP: A Structure Combination Pruning Method via Structured Sparse for Deep Convolutional Neural Networks
Qiyun Chen, Boyu Diao, Yongjun Xu 0001 |
ICPR (8) | 2 |
| 2024 | An Evolutionary Search-Based Operator Fusion Method with Binary Representation for Deep Learning Inference Acceleration
Boyu Diao, Hangda Liu, Qiyun Chen, Qi Wang 0025, Yongjun Xu 0001 |
ICPR (6) | 2 |
| 2024 | Relational Diffusion Distillation for Efficient Image Generation
Weilun Feng, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ACM Multimedia | 5 |
| 2024 | Continual Learning in the Frequency DomainabstractContinual learning (CL) is designed to learn new tasks while preserving existing knowledge. Replaying samples from earlier tasks has proven to be an effective method to mitigate the forgetting of previously acquired knowledge. However, the current research on the training efficiency of rehearsal-based methods is insufficient, which limits the practical application of CL systems in resource-limited scenarios. The human visual system (HVS) exhibits varying sensitivities to different frequency components, enabling the efficient elimination of visually redundant information. Inspired by HVS, we propose a novel framework called Continual Learning in the Frequency Domain (CLFD). To our knowledge, this is the first study to utilize frequency domain features to enhance the performance and efficiency of CL training on edge devices. For the input features of the feature extractor, CLFD employs wavelet transform to map the original input image into the frequency domain, thereby effectively reducing the size of input feature maps. Regarding the output features of the feature extractor, CLFD selectively utilizes output features for distinct classes for classification, thereby balancing the reusability and interference of output features based on the frequency domain similarity of the classes across various tasks. Optimizing only the input and output features of the feature extractor allows for seamless integration of CLFD with various rehearsal-based methods. Extensive experiments conducted in both cloud and edge environments demonstrate that CLFD consistently improves the performance of state-of-the-art (SOTA) methods in both precision and training efficiency. Specifically, CLFD can increase the accuracy of the SOTA CL method by up to 6.83% and reduce the training time by 2.6×. Boyu Diao, Libo Huang 0001, Zijia An, Zhulin An, Yongjun Xu 0001 |
NeurIPS | 2 |
| 2024 | DTuner: A Construction-Based Optimization Method for Dynamic Tensor Operators Accelerating
Boyu Diao, Hangda Liu, Yongjun Xu 0001 |
NPC (1) | 2 |
| 2024 | AKGF: Automatic Kernel Generation for DNN on CPU-FPGAabstractAbstract While tensor accelerated compilers have proven effective in deploying deep neural networks (DNN) on general-purpose hardware, optimizing for FPGA remains challenging due to the complex DNN architectures and the heterogeneous, semi-open compute units. This paper introduces the Automatic Kernel Generation for DNN on CPU-FPGA (AKGF) framework for efficient deployment of DNN on heterogeneous CPU-FPGA platforms. AKGF generates an intermediate representation (IR) of the DNN using TVM’s Halide IR, annotates the operators of model layers in the IR to compute them on the corresponding hardware cores, and further optimizes the operator code for CPU and FPGA using ARM’s function library and the polyhedral model to enhance model inference speed and power consumption. The experimental tests conducted on a CPU-FPGA board validate the effectiveness of AKGF, demonstrating significant acceleration ratios (up to 6.7x) compared to state-of-the-art accelerators while achieving a 2x power optimization. AKGF effectively leverages the computational capabilities of both CPU and FPGA for high-performance deployment of DNN on CPU-FPGA platforms. Boyu Diao |
Comput. J. | 3 |
| 2024 | Sketch-fusion: A gradient compression method with multi-layer fusion for communication-efficient distributed training
Li-Rong Dai 0001, Luqi Gong, Zhulin An, Yongjun Xu 0001, Boyu Diao |
J. Parallel Distributed Comput. | 5 |
| 2023 | Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical ViewabstractThis paper aims to explain the generalization of deep-fake detectors from the novel perspective of multi-order interactions among visual concepts. Specifically, we propose three hypotheses: 1. Deepfake detectors encode multi-order interactions among visual concepts, in which the low-order interactions usually have substantially negative contributions to deepfake detection. 2. Deepfake detectors with better generalization abilities tend to encode low-order interactions with fewer negative contributions. 3. Generalized deepfake detectors usually weaken the negative contributions of low-order interactions by suppressing their strength. Accordingly, we design several mathematical metrics to evaluate the effect of low-order interaction for deepfake detectors. Extensive comparative experiments are conducted, which verify the soundness of our hypotheses. Based on the analyses, we further propose a generic method, which directly reduces the toxic effects of low-order interactions to improve the generalization of deepfake detectors to some extent. Kelu Yao, Jin Wang 0039, Boyu Diao, Chao Li 0028 |
ICCV | 3 |
| 2022 | Interpretable Generative Adversarial NetworksabstractLearning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator encode disentangled localized visual concepts. Each filter in the layer is supposed to consistently generate image regions corresponding to the same visual concept when generating different images. The interpretable GAN learns to automatically discover meaningful visual concepts without any annotations of visual concepts. The interpretable GAN enables people to modify a specific visual concept on generated images by manipulating feature maps of the corresponding filters in the layer. Our method can be broadly applied to different types of GANs. Experiments have demonstrated the effectiveness of our method. Chao Li 0028, Kelu Yao, Jin Wang 0039, Boyu Diao, Yongjun Xu 0001, Quanshi Zhang |
AAAI | 4 |
| 2022 | Towards Understanding the Effect of Node Features on the Predictions of Graph Neural Networks
Zhao Zhang 0011, Boyu Diao, Yongjun Xu 0001, Chao Li 0028 |
ICANN (2) | 3 |
| 2022 | Pruning Filter via Gaussian Distribution Feature for Deep Neural Networks AccelerationabstractDeep learning has achieved impressive results in many areas, but the deployment of edge intelligent devices is still very slow. To solve this problem, we propose a novel compression and acceleration method based on data distribution characteristics for deep neural networks, namely Pruning Filter via Gaussian Distribution Feature (PFGDF). Compared with previous advanced pruning methods, PFGDF compresses the model by filters with insignificance in distribution, regardless of the contribution and sensitivity information of the convolution filter. PFGDF is significantly different from weight sparsification pruning because it does not require the special accelerated library to process the sparse weight matrix and introduces no more extra parameters. The pruning process of PFGDF is automated. Furthermore, the model compressed by PFGDF can restore the same performance as the uncompressed model. We evaluate PFGDF through extensive experiments, on CIFAR-10, PFGDF compresses the convolution filter on VGG-16 by 66.62% with more than 90% parameter reduced, while the inference time is accelerated by 83.73% on Huawei MATE 10. Jianrong Xu, Boyu Diao, Bifeng Cui, Chao Li 0028, Hailong Hong |
IJCNN | 2 |
| 2019 | PDMAC-SIC: Priority-based Distributed Low Delay MAC with Successive Interference Cancellation for Industrial Wireless NetworksabstractCommunications in industrial applications like wireless factory automation demands different timing requirements. Providing timely medium access of the critical traffic and its prioritization over regular traffic is a significant challenge in industrial wireless networks. Successive Interference Cancellation (SIC) technique is an effective way to decrease access delay by allowing multiple transmissions concurrently. A series of novel Medium Access Control (MAC) protocols are proposed to differentiate access delay for various traffic types or only exploit SIC for unique traffic type. However, to the best of our knowledge, this work is the first priority-based distributed MAC protocol that employs SIC (PDMAC-SIC) to provide low delay and accommodate different types of traffic for industrial wireless networks. There are two major contributions of our work: first, an extra power contention procedure other than traditional RTS/CTS contention in CSMA/CA is introduced in our PDMAC-SIC. This power contention procedure allows multiple transmitters to access the same channel simultaneously and thus access delay is decreased. Second, PDMAC-SIC is modeled by Markov chain and then the access delay is minimized by optimizing the size of power contention window. Our analytical model is verified through simulation. Results reveal that PDMAC-SIC performs better on access delay and packet loss rate than the existing good performing priority based CSMA/CA. Qi Wang 0025, Jianmin Liu, Chentao He, Boyu Diao, Yongjun Xu 0001 |
APNOMS | 5 |
| 2019 | Multi-objective Pruning for CNNs Using Genetic Algorithm
Chuanguang Yang, Zhulin An, Chao Li 0028, Boyu Diao, Yongjun Xu 0001 |
ICANN (2) | 4 |
| 2016 | A Reliable Depth-Based Routing Protocol with Network Coding for Underwater Sensor NetworksabstractWith the rapid development of marine technology, underwater sensor networks (UWSNs) are gradually evolving from research to practice in recent years. Practicability and reliability are two major concerns for routing protocols in UWSNs. As localization is not necessary in depth-based routing protocol (DBR), it has an outstanding practicability than other geographic routing protocols. However, the reliability is not well ensured. In this paper, we propose an innovative depth-based routing with network coding improving routing reliability while preserving the intrinsic distributed manner of DBR and introducing little time delay and energy cost. Moreover, a simple analytical performance model where ideal MAC is assumed is proposed to derive the analytical delivery ratio for our DBR-NC and DBR protocols. This analytical model is validated by simulation results. The extensive simulation results show that the proposed DBR-NC protocol outperforms (over 15%) the state of art DBR protocols in terms of packet delivery ratio. We also show that our DBR-NC will not introduce much extra delay and energy consumptions. Boyu Diao, Yongjun Xu 0001, Qi Wang 0025, Zhao Chen 0007, Chao Li 0028, Zhulin An, Guangjie Han |
ICPADS | 1 |
| 2015 | Vessel trajectory partitioning based on hierarchical fusion of position data
Xianbin Wu, Lin Wu 0006, Yongjun Xu 0001, Zhulin An, Boyu Diao |
FUSION | 5 |