VLDB 2026 Research / reviewers in the wild / expert
Yongjiang Xue
dblp:271/4906
· DBLP profile ↗
19ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-3487-0309ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HESP-Stream: A Sensor-Direct, Instruction-Driven FPGA System for Real-Time Sparse Event ProcessingabstractTo address the security threats posed by micro-UAVs, this work presents HESP-Stream, a real-time event-based detection system enabled by a sensor-direct, instruction-driven hardware–software co-design. Unlike conventional CPU-assisted pipelines, HESP-Stream eliminates the host-side preprocessing bottleneck by integrating on-chip event voxelization, instruction-driven sparse convolution, and fully streaming FPGA execution into a unified end-to-end datapath. Implemented on a Xilinx Kintex-7 FPGA, HESP-Stream achieves 59.51% mIoU and 85.34% accuracy after quantization. Under typical high-sparsity micro-UAV scenes, the system reaches an end-to-end latency as low as 5.622ms, delivering 2.930× and 2.989× speedups over GPU end-to-end and CPU+GPU pipelines, respectively. Jihui Qi, Yongjiang Xue, Qingzeng Song |
FCCM | 2 |
| 2026 | ReCoVLM: A Reconfigurable FPGA-GPU Co-Design for Edge Vision-Language InferenceabstractVision-Language Models (VLMs) demonstrate remarkable capabilities in open-world understanding and interaction. However, edge deployment remains challenging due to strict constraints on power, latency, reliability, and privacy requirements. The VLM inference pipeline typically consists of three stages: visual encoding, cross-modal prefill, and autoregressive decoding. These stages exhibit distinct characteristics regarding operator types, parallel granularity, and memory access patterns. Consequently, a single homogeneous edge device (e.g., CPU, GPU, NPU, or FPGA) often fails to achieve simultaneous optimality in performance and energy efficiency across the entire pipeline.To address this, we propose a configurable FPGA–GPU heterogeneous collaborative system tailored for edge VLM inference. Leveraging specific hardware strengths, we map the compute-intensive visual encoding and cross-modal Prefill stages to the GPU, while offloading the bandwidth-sensitive autoregressive decode stage to the FPGA. This workload-aware mapping results in improved energy efficiency. To mitigate edge resource constraints, we designed a model compression method for edge GPUs and FPGAs. By compressing the model parameters to a range acceptable by the hardware and using only 5% of visual tokens, we achieved 88.9% of the baseline performance. On the FPGA, we implement a pipelined decoder that efficiently utilizes DDR bandwidth and propose an efficient deployment method for MoE (Mixture of Experts). Including aligned block layout of expert weights, operator fusion, and on-chip caching strategies to reduce bandwidth fluctuations caused by weight switching and improve throughput stability, enables the effective bandwidth utilization rate to reach 87.2%. Experimental results on a heterogeneous platform comprising a Jetson Orin Nano (8GB) and a Xilinx VU9P FPGA demonstrate an end-to-end inference throughput of 18.1 tokens/s. Compared to an NVIDIA RTX 4090 baseline, our system achieves a 2.67× improvement in energy efficiency. Yongjiang Xue, Kailai Zhuang, Qingzeng Song |
FCCM | 2 |
| 2026 | A Unified Low-Latency ML-KEM Accelerator with Deterministic Packetization and Conflict-Free Memory MappingabstractNational Institute of Standards and Technology (NIST) Federal Information Processing Standards (FIPS) 203 standardizes Module-Lattice-Based Key-Encapsulation Mechanism (ML-KEM), driving the demand for low-latency accelerators that handle heterogeneous pipeline rates and complex on-chip memory provisioning. We present a speed-prioritized, unified ML-KEM-512 accelerator with a 16-lane packet interface across Key Generation (KeyGen), Encapsulation (Encaps), and Decapsulation (Decaps). To bridge irregular data streams, an output shaper aggregates accepted coefficients from rejection sampling into deterministic 16-lane blocks, shortening control critical paths. To sustain parallel Number Theoretic Transform (NTT) and Inverse Number Theoretic Transform (INTT) utilization, memory is partitioned by access semantics: sequential boundary traffic uses hierarchical buffers, while the strided transform domain uses a 16-bank store via distributed Look-Up Table Random Access Memory (LUTRAM). XOR bank mapping and stage-aware scheduling eliminate access conflicts inherent in standard mappings, while a Latest Value Table (LVT)-based scheme provides logical Two-Write Two-Read (2W2R) semantics. Implemented on a Xilinx Artix-7 FPGA, the design operates at 225 MHz, completing KeyGen, Encaps, and Decaps in 1131, 1228, and 1641 cycles. Achieving an 17.78 µs lifecycle latency, it demonstrates superior execution efficiency for the complete ML-KEM-512 flow. Xiangrui Jia, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | An Agile Design Framework for Resource-Efficient and Parameterizable Edge ISPsabstractFacing the real-time and resource constraints of edge imaging systems, this paper presents a highly parameterizable ISP hardware pipeline in SpinalHDL. To overcome the poor reusability of manually maintained Verilog designs, we adopt a configuration-driven generation paradigm to construct an end-to-end RAW-to-YUV streaming architecture. Under a unified ISPConfig, the pipeline supports flexible module chaining, compile-time structural specialization, and timing-consistent integration. At the microarchitectural level, we introduce a generic Window Generator that combines parameterized line buffers with a deterministic Skid Buffer-based flow-control shell to preserve pixel-level synchronization under backpressure. Experiments on the AMD Xilinx Kria KV260 show that the generated pipeline sustains real-time 1080p60 processing and reduces LUT and DSP utilization by 19.0% and 28.8%, respectively, relative to a functionally matched hand-coded Verilog baseline. The resource gains are mainly attributed to the shared Window Generator, unified inter-stage interfaces, and generator-time elimination of duplicated glue logic. The generated design also agrees well with software references, achieving 39.71 dB PSNR, 0.9636 SSIM, and 1.255 MAE, while representative requirement changes can be completed within a 20-minute RTL-to-simulation regression loop on average across representative modification tasks. Xitong Jiang, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | FPGA-based Adaptive Texture-Aware Stereo Matching Accelerator for Co-axial RGB-NIR SensorsabstractEdge robots require reliable all-day depth perception, but deploying RGB-NIR heterogeneous sensor fusion is severely constrained by strict size, weight, and power limits. Furthermore, traditional cross-modal registration and dynamic vertical residual compensation rely on external memory buffering, which stalls the continuous hardware dataflow. This paper presents a hardware-algorithm co-designed, DDR-less stereo matching accelerator tailored for co-axial RGB-NIR sensors to enable robust all-day edge vision. By employing a fully pipelined datapath with on-chip line buffers, the design bypasses all external memory dependencies. A cost-domain adaptive texture fusion (CATF) module reuses the Census transform register array to extract texture confidence, adaptively modulating semi-global matching penalty parameters and cross-modal fusion weights with near-zero overhead. To ensure robustness against mechanical vibrations and thermal drifts, an epipolar redundancy mechanism dynamically searches adjacent scanlines, tolerating up to 1-pixel vertical residuals. Implemented on a Zynq-7020 SoC, the system processes 720×540 video streams at 128 FPS. Consuming 35,868 LUTs and 2.43 Mb BRAM, it reduces the mean absolute error by up to 19.9% in all-day scenarios compared to OpenCV SGBM baselines, achieving an effective balance between robust multimodal perception and hardware efficiency. Chenghe Zhang, Qingzeng Song, Yongjiang Xue, Weigang Kong, Fei Qiao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | FPGA-Based High-Parallelism ORB Feature Extraction Accelerator for Visual SLAM
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ISCAS | 3 |
| 2026 | X2Fashion: temporally consistent fashion video generation guided by image, pose and text
Yongjiang Xue, Congwei Guo, Aizhe Wu, Yongzhen Ke, Peirong Tang |
Multim. Syst. | 1 |
| 2025 | ELTC: An End-to-End Large Language Model-Based Tensor Compilation Optimization Framework
WenBo Ma, Qingzeng Song, Fei Qiao, Yongjiang Xue |
APLAS | 4 |
| 2025 | Multi-task Low-Level Vision Network for FPGA Deployment: SL-SYENet Joint Optimization Framework
Shenao Li, Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICA3PP (7) | 2 |
| 2025 | FPGA-Based Hybrid GCN-Transformer Accelerator for 3D Pose Estimation
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICA3PP (6) | 2 |
| 2025 | DB-MFNet: A Dual-Branch Cross-Modal Fusion Network for High-Resolution Remote Sensing Semantic Segmentation
Huiren Hao, Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICIC (5) | 2 |
| 2025 | RFW-YOLO: A Multi-scale Feature Fusion Method for Infrared Anti-UAV Detection Based on WTConv
Zijuan Li, Yongjiang Xue, Qingzeng Song |
ICIC (2) | 2 |
| 2025 | EFF-ViT: A Vision Transformer with Feature Enhancement and Fusion for Fine-Grained Visual Classification
Yongjiang Xue, Qingzeng Song |
ICIC (21) | 2 |
| 2025 | RCMUNet: An End-to-End Hybrid U-Net Architecture of CNN-Mamba for Low-Light RAW Image Enhancement
Yongjiang Xue, Fei Qiao, Qingzeng Song |
ICIC (14) | 2 |
| 2025 | FDSI-RTDETR : A Lightweight Unmanned Aerial Vehicle (UAV) Aerial Image Small Object Detection NetworkabstractUnmanned aerial vehicle (UAV) aerial image detection tasks frequently encounter challenges, including small object detection, significant occlusion, dense target presence, and inconsistent illumination conditions. Traditional object detection algorithms sometimes produce erroneous detections. This study proposes a lightweight enhanced method based on the design concept of RT-DETR. Initially, a lightweight FRESE-Block module is introduced to enhance the feature extraction capabilities of the backbone network. FasterNet is improved through RepConv, while an efficient channel attention mechanism is implemented to represent the interrelations across feature mapping channels, enhancing the capture of small target features. Subsequently, the Dynamic-range Histogram Self-Attention is incorporated into the intra-scale feature interaction module. This approach enables thorough feature extraction at both local and global levels, hence minimizing the false detection rate. Furthermore, the ESlimneck-DASF architecture is suggested to enhance crossscale feature fusion. This framework fully utilizes the benefits of dynamic upsampling and feature fusion, enriching semantic information across various scales. Ultimately, the Inner-MPDIoU loss function with an optimal ratio was chosen to enhance convergence speed and detection accuracy. Experimental findings on the VisDrone2019 and NWPU VHR-10 datasets indicate that the mAP values are 44.1% and 90.7%, representing increases of 1.7% and 2.4% over RT-DETR-r18. Meanwhile, parameters and GFLOPs are diminished by 12.17% and 9.95%. Tianyu Guo 0012, Qingzeng Song, Yongjiang Xue, Fei Qiao |
IJCNN | 3 |
| 2025 | FPGA-Accelerated CNN-Transformer Hybrid Model for Real-Time Semantic Segmentation in Autonomous Driving
Yongjiang Xue, Fei Qiao, Qingzeng Song |
NPC (1) | 2 |
| 2025 | High-Performance FPGA-Based System for Diabetic Retinopathy ClassificationabstractDiabetic retinopathy (DR), as a major complication of diabetes mellitus, presents significant diagnostic challenges including limited medical resources, insufficient manpower, and time-consuming examination procedures. To address these issues, this study proposes a high-performance FPGA-based intelligent DR classification system. By integrating an innovative MIL-ViT network architecture with FQ- ViT quantization techniques, we have achieved efficient real-time diagnosis on the Xilinx ZCU102 plat- form. The system demonstrates classification accuracies of 85.5% and 90.8% on the APTOS2019 and RFMiD2020 datasets respectively. Through hardware optimization, we have reduced the single-image inference time to under 20 milliseconds, meeting clinical real-time requirements. Notably, our specialized acceleration architecture achieves an energy efficiency of 36.91 GOPs/W, representing significant improvement over existing solutions. The complete algorithm-to- hardware co-design maintains quantization accuracy loss below 4% while accomplishing full-stack optimization. This work provides valuable insights for developing embedded medical diagnostic devices. Fengtao Shi, Yongjiang Xue, Jinzhu Zhang, Qingzeng Song |
SMC | 2 |
| 2021 | Innovative Design and Simulation of a Transformable Robot with Flexibility and Versatility, RHex-T3abstractThis paper presents a transformable RHex-inspired robot, RHex-T3, with high energy efficiency, excellent flexibility and versatility. By using the innovative 2-DoF transformable structure, RHex-T3 inherits most of RHex’s mobility, and can also switch to other 4 modes for handling various missions. The wheel-mode improves the efficiency of RHex-T3, and the leg-mode helps to generate a smooth locomotion when RHex-T3 is overcoming obstacles. In addition, RHex-T3 can switch to the claw-mode for transportation missions, and even climb ladders by using the hook-mode. The simulation model is conducted based on the mechanical structure, and thus the properties in different modes are verified and analyzed through numerical simulations. Yue Lin 0006, Yujia Tian, Yongjiang Xue, Shujun Han, Huaiyu Zhang, Wenxin Lai |
ICRA | 3 |
| 2021 | Lywal: a Leg-Wheel Transformable Quadruped Robot with Picking up and Transport FunctionsabstractThis paper introduces a leg-wheel transformable quadruped robot named Lywal which can switch to the leg-mode and the wheel-mode for locomotion, and the claw-mode for picking up and transport functions. First, the mechanical structure of Lywal is designed by using an innovative 2-DoF transformable mechanism. Second, the calculation of kinematics is analyzed in detail. Then, the switching-mode strategy and the mobile control strategies in different modes are designed. Finally, the prototype of Lywal is built. The properties of the mobile modes are analyzed, and the picking-up and transport functions of the claw-mode are verified through physical experiments. Yongjiang Xue, Xichen Yuan, Yuhai Wang, Juezhu Lai |
ICRA | 1 |