EDBT 2026 Demo / reviewers in the wild / expert
Weirong Dong
dblp:397/6268
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-6328-2816ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gundam: A Generalized Unified Design and Analysis Model for Matrix Multiplication on Edge
Weirong Dong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ASP-DAC | 3 |
| 2026 | LMESN: A Leakage-Driven MOSFET Reservoir for Scalable and Ultra-Low-Power Temporal InferenceabstractEdge-based temporal inference demands energy-efficient and scalable computing architectures, but existing analog reservoir computing models often face high energy costs and limited reconfigurability. We present LMESN, a leakage-current-driven, pulse-based reservoir computing architecture that exploits intrinsic threshold-voltage variation in standard CMOS to realize ultra-low-power stochastic dynamics. To overcome physical array size constraints, we propose a Shift-Multi-Mask (SMM) technique that emulates large virtual reservoirs through cyclic mask shifts, reducing update energy by over $100 \times$ and enabling single-cycle reconfiguration. To further boost task-level performance, we develop a hardware-software co-optimization framework that jointly tunes the ADC quantization range and reservoir mask structure via a discrete genetic algorithm. Post-layout simulations in 22 nm CMOS and evaluations on eight time-series datasets demonstrate up to 13.7% accuracy improvement and $5 \times$ variance reduction over unoptimized LMESN baselines. Compared to prior analog and neural network models, LMESN achieves 3–7 orders of magnitude lower energy consumption while delivering competitive or superior accuracy. Together, these innovations make LMESN a scalable, energy-efficient, and task-adaptive platform for edge temporal processing, setting a new direction in physical reservoir computing. Haoyuan, Masami Utsunomiya, Ryuko Seki, Weirong Dong, Feng Liang 0001, Takashi Sato 0001 |
ASP-DAC | 5 |
| 2026 | Mitigating Conductance Drift via In-Situ Calibration for Reliable RRAM-Based CIM Edge Inference
Zhen Kong, Weirong Dong, Zhengke Yang, Yida Liang, Jiamin Li 0008, Yida Li 0004, Longyang Lin |
ISCAS | 2 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 3 |
| 2026 | A Low-Power Speech-Based Depression Recognition Processor With Hierarchical Local-Global NetworkabstractDepression is a critical public health concern characterized by underdiagnosis, often due to stigma, lack of awareness, and reluctance to seek help. Cases of delayed intervention could be alleviated by wearable solutions which enable continuous and unobtrusive monitoring of depression indicators. Compared to electroencephalogram (EEG)-based and video-based depression recognition, speech-based approaches can be performed without deliberate user attention. However, due to the limited accuracy of existing algorithms and constrained resources at edge, performing accurate speech-based depression recognition on wearable platforms remains a challenge. Therefore, to achieve unobtrusive, accurate, and efficient depression recognition at edge, this work presents a hierarchical local–global network (HLG-Net) and optimized processor design for speech-based depression recognition. The proposed HLG-Net integrates convolutional neural networks (CNNs) with multihead attention (MHA) mechanism to simultaneously capture local acoustic features and global utterance-level coherence, enhancing depression stage recognition. For efficient processor design, a cross-layer buffered dataflow is proposed for efficient data handling, reducing data storage by 98.77%. The computing unit (CU) employs layer fusion, operator optimization, and quantization techniques to further improve resource utilization and reduce power consumption while preserving recognition accuracy. System-level low-power techniques such as clock/input gating and near-threshold design for application specific integrated circuit (ASIC) further reduce power consumption. The proposed processor implemented on field-programmable gate array (FPGA) (XC7Z100-2FFG900) achieves the lowest reported mean absolute error (MAE) of 5.13 on AVEC 2014 database. The 180-nm ASIC implementation shows a simulated power consumption of$17.4~\mu $W at 0.4 V. The results demonstrate the feasibility of accurate and efficient speech-based depression recognition on wearables. Yuxing Zhi, Weirong Dong, Huaijun Wang, Longyang Lin, Jiamin Li 0008 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |