EDBT 2026 Demo / reviewers in the wild / expert
Weiguang Pang
dblp:301/3561
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-0208-4677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Abusing DDS Discovery: Denial-of-Service Attacks Against ROS 2abstractThe Data Distribution Service (DDS) provides data-centric publish-subscribe messaging with a mandatory discovery protocol, enabling distributed applications to automatically locate and communicate with each other. ROS 2, the de facto middleware for robotic systems, adopts DDS as its communication backbone. In this paper, we demonstrate that the DDS discovery mechanism can be exploited to mount Denial-of-Service attacks against ROS 2 applications. By repeatedly triggering discovery traffic, an adversary can significantly inflate pipeline latency during runtime. We validate the attack on ROS 2 Humble with two widely used DDS implementations and a real UAV case study, confirming its effectiveness across different configurations. Jiafu Xu, Songran Liu, Zilong Wang 0018, Minghe Yu 0001, Yue Tang 0001, Yang Wang 0082, Weiguang Pang, Wang Yi 0001 |
DATE | 7 |
| 2025 | Dual Focus-Attention Transformer for Robust Point Cloud RegistrationabstractRecently, coarse-to-fine methods for point cloud registration have achieved great success, but few works deeply explore the impact of feature interaction at both coarse and fine scales. By visualizing attention scores and correspondences, we find that existing methods fail to achieve effective feature aggregation at the two scales during the feature interaction. To tackle this issue, we propose a Dual Focus-Attention Transformer framework, which only focuses on points relevant to the current point for feature interaction, avoiding interactions with irrelevant points. For the coarse scale, we design a superpoint focus-attention transformer guided by sparse keypoints, which are selected from the neighborhood of superpoints. For the fine scale, we only perform feature interaction between the point sets that belong to the same superpoint. Experiments show that our method achieve the state-of-the-art performance on three standard benchmarks. The code and pre-trained models are available at https://github.com/fukexue/DFAT.git. Kexue Fu 0001, Mingzhi Yuan, Changwei Wang 0001, Weiguang Pang, Jing Chi, Manning Wang, Longxiang Gao |
CVPR | 4 |
| 2025 | AASD: Accelerate Inference by Aligning Speculative Decoding in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) have achieved notable success in visual instruction tuning, yet their inference is time-consuming due to the auto-regressive decoding of Large Language Model (LLM) backbone. Traditional methods for accelerating inference, including model compression and migration from language model acceleration, often compromise output quality or face challenges in effectively integrating multimodal features. To address these issues, we propose AASD, a novel framework for Accelerating inference with refined KV Cache and Aligning speculative decoding in MLLMs. Our approach leverages the target model’s cached KeyValue (KV) pairs to extract vital information for generating draft tokens, enabling efficient speculative decoding. To reduce the computational burden associated with long multimodal token sequences, we introduce a KV Projector to compress the KV Cache while maintaining representational fidelity. Additionally, we design a Target-Draft Attention mechanism that optimizes the alignment between the draft model and the target model, achieving the benefits of real inference scenarios with minimal computational overhead. Extensive experiments on mainstream MLLMs demonstrate that our method achieves up to a $2 \times$ inference speedup without sacrificing accuracy. This study not only provides an effective and lightweight solution for accelerating MLLM inference but also introduces a novel alignment strategy for speculative decoding in multimodal contexts, laying a strong foundation for future research in efficient MLLMs. Code is availiable at https://github.com/transcend-0/ASD Muyang Zhang, Weiguang Pang, Yuzhi Chen, Rongtao Xu, Kexue Fu 0001, Changwei Wang 0001, Longxiang Gao |
DAC | 4 |
| 2025 | A Mamba-KAN Joint UNet Framework for Medical Image Segmentation
Haoyu Zhou, Changwei Wang 0001, Weiguang Pang, Lei Cui 0006, Shujun Gu, Longxiang Gao, Kexue Fu 0001, Youyang Qu |
PRCV (3) | 3 |
| 2024 | Control Flow Divergence Optimization by Exploiting Tensor CoresabstractKernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem. Weiguang Pang, Xu Jiang 0004, Songran Liu, Lei Qiao 0002, Kexue Fu 0001, Longxiang Gao, Wang Yi 0001 |
DAC | 1 |
| 2023 | Efficient CUDA stream management for multi-DNN real-time inference on embedded GPUs
Weiguang Pang, Xiantong Luo, Kailun Chen, Dong Ji, Lei Qiao 0002, Wang Yi 0001 |
J. Syst. Archit. | 1 |
| 2022 | Toward the Predictability of Dynamic Real-Time DNN InferenceabstractDeep neural networks (DNNs) have been widely used in many cyber–physical systems (CPSs). However, it is still a challenging work to deploy DNNs in real-time systems. In particular, the execution time of DNN inference must be predictable, s.t. it could be known whether the runtime inference can complete within a required timing constraint. Moreover, the timing constraints may change dynamically with the runtime environment in many embedded applications, such as autonomous cars. A possible way to meet such dynamic real-time requirements is to execute different subnetworks of a DNN at runtime. However, improper construction of subnetworks may not only introduce unpredictable inference time, s.t. the real-timing constraints could be violated unexpectedly, but also has poor compatibility with the well-optimized machine learning framework (e.g., TensorFlow). In this article, we study the predictability when executing different subnetworks of a DNN. In particular, we present a featurewise runtime adaptation framework for DNN inference, which is implemented and validated on NVIDIA Jetson TX2 and Nano with TensorFlow. The experimental results show that our method can achieve predictable inference time in comparison with the state-of-the-art methods. Weiguang Pang, Xu Jiang 0004, Mingsong Lv, Teng Gao, Di Liu 0002, Wang Yi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |