Jiale Dong

dblp:295/8350 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Video object segmentation based on feature compression and attention correction
Jiale Dong, Chenxu Wang 0012, Sugang Ma, Wangsheng Yu
Signal Process. Image Commun.2
2026 MoE-Sched: Enabling Efficient FPGA Deployment of Mixture-of-Experts Vision Transformers via Coordinated Scheduling
abstract
Compared to traditional Vision Transformers (ViT), Mixture-of-Experts ViTs (MoE-ViTs) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, field-programmable gate array (FPGA)-based accelerators for MoE-ViTs emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of efficiently managing the highly dynamic and sparse computation patterns, resulting in suboptimal tradeoffs between resource utilization and performance. To address the inefficiencies in deploying MoE-ViTs on FPGAs, we present MoE-Sched, a novel end-to-end accelerator that embraces a scheduling-centric design philosophy. Rather than optimizing isolated kernels, MoE-Sched coordinates multilevel scheduling, from fine-grained intrakernel streaming to module reuse and multidie mapping, to holistically balance latency, bandwidth (BW), and resource usage. We further integrate a hardware-aware quantization scheme tailored for streaming attention and sparse expert execution, preserving accuracy while minimizing overhead. Experimental results demonstrate that our accelerator achieves nearly 100 frames/s on M3ViT-tiny, a$3.13\times $improvement in throughput, and over 75% energy reduction compared to state-of-the-art (SOTA) FPGA MoE accelerators, while maintaining less than 1% accuracy loss across vision benchmarks. Our implementation will be open-sourced.
Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou
IEEE Trans. Very Large Scale Integr. Syst.1
2025 CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
Jiale Dong, Wenqi Lou, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou
Euro-Par (2)1
2025 UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
abstract
Compared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, FPGA-based accelerators for MoE-ViT emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of fully exploring the design space, leading to suboptimal trade-offs between resource utilization and performance. To overcome this problem, we introduce UbiMoE, a novel end-to-end FPGA accelerator tailored for MoE-ViT. Leveraging the unique computational and memory access patterns of MoE-ViTs, we develop a latency-optimized streaming attention kernel and a resource-efficient reusable linear kernel, effectively balancing performance and resource consumption. To further enhance design efficiency, we propose a two-stage heuristic search algorithm that optimally tunes hardware parameters for various FPGA resource constraints. Compared to state-of-the-art (SOTA) FPGA designs, UbiMoE achieves 1.34× and 3.35× throughput improvements for MoE-ViT on Xilinx ZCU102 and Alveo U280 platforms, respectively, while enhancing energy efficiency by 1.75× and 1.54×. Our implementation is available at https://github.com/DJ000011/UbiMoE.
Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou
ISCAS1
2025 Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation
Chenxu Wang 0012, Sugang Ma, Jiale Dong, Yunchen Wang, Wangsheng Yu
Knowl. Based Syst.4
2024 EVRACE: Enhanced Visual Retrieval and Analysis for Large-Scale E-Bike Charging Curve Exploration
abstract
This Paper introduces a novel Enhanced Visual Retrieval and Analysis for Charging Efficiency (EVRACE) system, which presents a novel two-stage framework for visual retrieval and advanced analytics for massive charging curves of electric vehicles. The EVRACE is becoming an important system for enhancing battery performance management, detecting and mitigating charging anomalies, and informing strategic decision-making for stakeholders in the EV ecosystem. EVRACE utilizes deep time-series clustering to categorize charging curves with distinct characteristics and retrieves the most relevant cluster and the top k curves for comprehensive data analysis and risk assessment. This innovative framework substantially enhances retrieval efficiency, mitigates computational complexity, and offers a robust solution for real-time processing of large-scale charging data. Validation using real-world datasets demonstrates that the EVRACE system significantly improves retrieval efficiency and achieves high accuracy in anomaly detection compared to traditional methods.
Saisai Hu, Donghui Ding, Zhijun Pan, Jiale Dong
IEEE Big Data7
2024 Video object segmentation based on dynamic perception update and feature fusion
Fucheng Li, Jiale Dong, Nan Dai, Sugang Ma, JiuLun Fan 0001
Image Vis. Comput.3
2021 A DMP-based Online Adaptive Stiffness Adjustment Method
abstract
Learning from demonstration (LfD) is a promising method for robots to learn and generalize human-like skills. It has the advantages of high programming efficiency, easy optimization, and non-professionals can also operate. There is a lot of research work that learn motion trajectories and stiffness curves from human demonstrations simutaneously to make the robot compliant, but previous work rarely consider the changes of environment. In this article, we propose an adaptive stiffness method that enables the robot to learn motion and stiffness trajectories from a single demonstration. When the environment changes, it can spontaneously tune the stiffness according to environmental feedback to ensure the smoothness of the task. Thus the robot has the ability to adapt to environmental changes. We first proved the theoretical feasibility of the method, and then we conducted physical experiments on the Baxter robot to verify the effectiveness of the proposed method.
Jiale Dong, Weiyong Si, Chenguang Yang 0001
IECON1