EDBT 2026 Demo / reviewers in the wild / expert
Mingyu Wang 0001
dblp:65/8491-1
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0001-5722-9752ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EE-Extractor: a near-sensor real-time effective event extractor for dynamic vision sensor
Feiqiang Li, Mingyu Wang 0001, Wenhong Li, Ming-e Jing, Xiaoyang Zeng |
Sci. China Inf. Sci. | 3 |
| 2026 | EDMOT: A 25-65 ns Latency Event-Driven Multiobject Tracker for Dynamic Vision SensorsabstractDynamic vision sensors (DVSs), renowned for their low latency and sparse event-driven output, have garnered significant attention in machine vision applications, particularly in latency-sensitive applications like object tracking. However, current DVS-based object tracking systems have not fully utilized the advantages of DVS due to their frame-based processing or the high computing intensity of their algorithms. This brief presents EDMOT, a fully event-driven multiobject tracking system for DVS, achieving 25–65 ns latency at 200 MHz. EDMOT introduces three key innovations: 1) a novel event-driven update mechanism that processes only the latest event and the expired oldest event, minimizing computational overhead; 2) a dual-threshold tracking strategy that decouples object formation and motion phases, significantly improving tracking accuracy; and 3) a row–column feature memory with flag registers, enabling object separation within eight clock cycles. The proposed EDMOT is evaluated on public datasets, demonstrating superior tracking accuracy compared to prior methods. Finally, EDMOT was implemented at the HLMC 55 nm, supporting 20–100 Me/s throughput. To the best of our knowledge, this is a multiobject tracking system with minimal delay and the highest event processing throughput. Feiqiang Li, Yaoyi Chen, Mingyu Wang 0001, Minge Jing, Wenhong Li, Xiaoyang Zeng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | An FPGA-Based Event-Driven SNN Accelerator for DVS Applications With Structured Sparsity and Early-StopabstractThis paper proposed an algorithm-hardware co-design of an event-driven spiking neural network (SNN) accelerator for classification tasks of event-based data from dynamic vision sensors (DVS), which can implement a feed-forward SNN with a maximum network size of 1 million synapses. Configurable structured sparsity is introduced between the first layer and the second layer to improve energy efficiency and balance the workload between different processing elements (PEs). The number of available neurons in the accelerator is sparsity-dependent and ranges from 1024 to 4096. The modified leaky-integrate-fire (LIF) neuron model and an event-driven neuron update scheme are employed in both algorithm and hardware to fully utilize the natural sparsity of DVS event stream. Early-stop inference strategy on the hardware enables a trade-off between inference accuracy and efficiency. A three-layer fully connected SNN is trained through backpropagation through time (BPTT) and is implemented and evaluated on Xilinx ZCU104 FPGA. Our design can achieve 96.0% accuracy on the N-MNIST dataset and 79.0% accuracy on the DVS128-Gesture dataset both at 50% sparsity. The top performance of the accelerator on ZCU104 is 3.22 GSOP/S and 3.99 GSOP/W at 250 MHz. Shu Cao, Shangmei Wang, Mingyu Wang 0001, Wenhong Li, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Defective Pixel Corrector for Line Scan and Area Scan Image SensorsabstractThe defective pixel corrector is an essential component of the image processor, which detects and corrects defective pixels in the image. Processing discrete defective pixels that exist in the output of area-scan image sensors is the focus of current research. However, the case of defective columns produced by line scan image sensors is a massive challenge for existing algorithms. In this paper, we develop novel algorithms for the detection and correction of columnar defective pixels by modeling the properties of defective columns in the output images of line scan image sensors. In addition, we design the non-extremum verdict and texture adaptive correction for clustered defective pixels in the area scan sensors and apply them to current algorithms to get enhancement, and new algorithms are obtained. Moreover, we propose a generalized hardware architecture for detection and correction. In the experimental stage, the proposed methods are compared with the state-of-the-art methods, both at the algorithmic level and in hardware implementation. The experimental results show that, compared to the current widely used algorithms, our methods for line-scan image sensors improve the detection rate by 28%, the detection precision by 90%, and the image restoration quality by 35% with a 10% increase in hardware consumption. For area-scan image sensors, the proposed non-extremum verdict improves the detection rate by 25%, improves the precision from 0.056 to 0.983, and the texture adaptive correction mitigates the image blurring caused by the correction process, with a 37% improvement in image quality, while incurring only a 1% increase in hardware overhead. In addition, our proposed algorithms achieve comparable detection and correction results with far fewer computational and storage resource requirements than machine learning-based algorithms. Liyuan Peng, Mingyu Wang 0001, Wenhong Li, Minge Jing, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | High Signal-to-Noise Ratio and High-Sensitivity 4-D LiDAR Imaging ReceiverabstractThis brief designs and implements a 4-D imaging light detection and ranging (LiDAR) receiver. It employs a reconfigurable transimpedance amplifier (TIA) that alternates between two modes to separately achieve ranging and light intensity quantification functions. A new mode-switching method based on a monostable multivibrator is proposed, allowing the TIA to automatically switch modes during measurement. The reconfigurable TIA and mode-switching method enable the application of charge sampling in 4-D LiDAR imaging receiver, resulting in a higher signal-to-noise ratio (SNR) compared with traditional designs. In addition, the TIA mode used for distance measurement achieves a bandwidth of 140 MHz, a gain of$99.8~\text {dB}\Omega $, and an input-referred noise of ~20-nA rms, indicating high detection sensitivity. A prototype implemented in 0.18-um CMOS verifies the feasibility of the proposed receiver, consuming 24.5 mW. Measurement results demonstrate that the prototype can effectively acquire distance and light intensity information of target objects within a range of 5 m. Jianping Guo 0002, Zhengping Gao, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A 1024-Neuron 1M-Synapse Event-Driven SNN Accelerator for DVS ApplicationsabstractThis paper proposed a hardware-algorithm co-design of an event-driven Spiking Neural Network (SNN) accelerator with structured sparsity for Dynamic Vision Sensors (DVS) applications. The accelerator can accommodate up to 1024 neurons and 1 million synapses for a feed-forward fully connected SNN implementation. Configurable structured sparsity is introduced by modular arithmetic both in the algorithm and hardware to improve the energy efficiency, reduce the memory requirement, and balance the workload between different processing elements. With an event-driven neuron update scheme, the accelerator can fully utilize the benefits of structured sparsity and can directly process DVS output data for classification tasks without encoding. A three-layer SNN is trained through backpropagation through time (BPTT) and is implemented on Xilinx ZCU104 FPGA, which achieves 96% accuracy on the N-MNIST dataset and 79% accuracy on the DVS-Gesture dataset both at 50% sparsity. The top performance of the accelerator on ZCU104 is 3.82 GSOP/S and 5.31 GSOP/W at 250 MHz. Shu Cao, Shangmei Wang, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
ISCAS | 6 |
| 2024 | A Lossless Compression Algorithm with Hardware Implementation for Dynamic Vision SensorabstractNowadays, with the increase resolution of Dynamic Vision Sensor (DVS), efficient compression algorithm for event stream is needed urgently. Conventional DVS system encodes event data in address event representation (AER) for output while ignores the data redundancy imposed by the correlation of events. To address this challenge, this paper first analyzes the spatiotemporal characteristics of event stream and the impact of readout circuits. Based on the analysis, the context-based encoding strategies for spatial address, timestamp and polarity of events are proposed respectively with the consideration of data flow in DVS hardware. Besides, the hardware architecture with high parallelism is presented to implement the compression algorithm, which achieves high throughput at an affordable cost. The hardware is implemented in the 55nm process as part of a 512x512 resolution DVS. The experimental results demonstrate that our methods achieves higher average compression ratio compared to conventional and DVS-specific coding algorithms. Zewei Ding, Shangmei Wang, Yujie Cai, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
ISCAS | 6 |
| 2024 | A Multimode Neuromorphic Vision Sensor With Improved Brightness Measurement Performance by Pulse Coding MethodabstractThis article proposes a multimode neuromorphic event-frame integrated vision sensor that enables event detection (ED) with simultaneous brightness measurement based on the pulse width modulation mechanism. The logarithmic voltage is directly taken as intensity information. Brightness measurement involves in-pixel voltage-to-pulse conversion and out-pixel pulse coding. The maximum event bandwidth is improved to 366 Meps by pipelining the time-prior arbiter along with the address-events grouping circuit. A wide intensity dynamic range of 105 dB can theoretically be achieved through logarithmic photoelectric conversion and pulse coding. Our sensor supports a$128\times 64$frame-like image with an improved signal-to-noise ratio of 49 dB. The experimental results indicate that the log sensitivity of the optimized logarithmic photoreceptor was measured as 164 mV/dec. The equivalent frame rate for both event and intensity reaches kilo fps, making it a promising candidate in high-speed wireless sensing applications. Zewei Ding, Qijuan Wu, Mingyu Wang 0001, Jingjing Liu 0004, Xiaoyang Zeng, Wenhong Li, Zhi Liu 0004, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 3 |
| 2024 | KBStyle: Fast Style Transfer Using a 200 KB Network With Symmetric Knowledge DistillationabstractConvolutional Neural Networks (CNNs) have achieved remarkable progress in arbitrary artistic style transfer. However, the model size of existing state-of-the-art (SOTA) style transfer algorithms is immense, leading to enormous computational costs and memory demand. It makes real-time and high resolution hard for GPUs with limited memory and limits the application on mobile devices. This paper proposes a novel arbitrary artistic style transfer algorithm, KBStyle, whose model size is only 200 KB. Firstly, we design a style transfer network where the style encoder, content encoder, and corresponding decoder are custom designed to guarantee low computational cost and high shape retention. Besides, the weighted style loss function is presented to improve the performance of style migration. Then, we propose a novel knowledge distillation method (Symmetric Knowledge Distillation, SKD) for encoder-decoder-based style transfer models, which redefines the knowledge and symmetrically compresses the encoder and decoder. With the SKD, the proposed style transfer network is further compressed by 14 times to achieve the KBStyle. Experimental results demonstrate that the proposed SKD method achieves comparable results with other SOTA knowledge distillation algorithms for style transfer. Besides, the proposed KBStyle achieves high-quality stylized images. And the inference time of the KBStyle on an Nvidia TITAN RTX GPU is only 20 ms when the resolutions of the content image and style image are both 2k-resolution ( 2048×1080 ). Moreover, the 200 KB model size of KBStyle is much smaller than the SOTA models and facilitates style transfer on mobile devices. Wenshu Chen, Mingyu Wang 0001, Xiaolin Wu 0001, Xiaoyang Zeng |
IEEE Trans. Image Process. | 3 |
| 2023 | Denoising Method for Dynamic Vision Sensor Based on Two-Dimensional Event DensityabstractThe Dynamic Vision Sensor (DVS) is a new type of bionic vision image sensor that offers the advantages of low latency, low power consumption, and high dynamics range compared to conventional sensors. However, background activity (BA) noise will degrade the quality of the DVS output data and lead to unnecessary bandwidth overhead. In dark environments, pixel arrays generate abundant noise, and conventional spatiotemporal filters can hardly achieve satisfactory results. To solve this problem, we exploit the difference in event density distribution between the actual event and noise and propose a denoising method that utilizes the event densities with two neighbors of different radii. Compared to spatiotemporal filters, our approach reduces the error rate on synthetic datasets by at least 35%. Meanwhile, our approach is subjectively more visually appealing. With our denoising method, the performance of DVS can be better in dark conditions. Yaoyi Chen, Feiqiang Li, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
ISCAS | 6 |
| 2023 | Queue-based Spatiotemporal Filter and Clustering for Dynamic Vision SensorabstractDynamic vision sensors (DVS) have significant potential in scenes involving high-speed motion and extreme light. However, DVS is sensitive to background active noise, which will degrade the quality of the output. The ordinary$O(N^{2})$-Space spatiotemporal filter's memory complexity is high. It needs$N\times N$memory cells ($N\times N$is the resolution on the sensor). Some works reduce memory complexity by sacrificing the performance of the filter. To ensure the filtering effect and reduce the filter's memory complexity, this paper proposes a novel filter: Queue-based spatiotemporal filter. Moreover, based on the Queue-based spatiotemporal filter, this paper proposes a clustering algorithm that can cluster while filtering. Experiments show that the proposed filter's performance is similar to the$O(N^{2})$-Space spatiotemporal filter while having a lower memory complexity. Besides, using the proposed clustering algorithm, the objects in motion can be clustered with low calculation complexity. Feiqiang Li, Yaoyi Chen, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
ISCAS | 6 |
| 2023 | TSDN: Two-Stage Raw Denoising in the DarkabstractDenoising is one of the most significant procedures in the image processing pipeline. Nowadays, deep-learning-based algorithms have achieved superior denoising quality than traditional algorithms. However, the noise becomes severe in the dark environment, where even the SOTA algorithms fail to achieve satisfactory performance. Besides, the high computational complexity of deep-learning-based denoising algorithms makes them hardware unfriendly and difficult to process high-resolution images in real-time. To address these issues, a novel low-light RAW denoising algorithm Two-Stage-Denoising (TSDN), is proposed in this paper. In TSDN, denoising consists of two procedures: noise removal and image restoration. Firstly, in the noise-removal stage, most noise is removed from the image, and an intermediate image that is easier for the network to recover the clean image is obtained. Then, in the restoration stage, the clean image is restored from the intermediate image. The TSDN is designed to be light-weight for real-time and hardware friendly. However, the tiny network will be insufficient for satisfactory performance if directly trained from scratch. Therefore, we present an Expand-Shrink-Learning (ESL) method to train the TSDN. In the ESL method, firstly, the tiny network is expanded to a larger one with similar architecture but more channels and layers, which enhances the learning ability of the network because of more parameters. Secondly, the larger network is shrunk and restored to the original small network in fine-grained learning procedures, including Channel-Shrink-Learning (CSL) and Layer-Shrink-Learning (LSL). Experimental results demonstrate that the proposed TSDN achieves better performance (PSNR and SSIM) than other SOTA algorithms in the dark environment. Besides, the model size of TSDN is one-eighth of that of the U-Net for denoising (a classical denoising network). Wenshu Chen, Mingyu Wang 0001, Xiaolin Wu 0001, Xiaoyang Zeng |
IEEE Trans. Image Process. | 3 |
| 2023 | LineDL: Processing Images Line-by-Line With Deep LearningabstractAlthough deep learning-based (DL-based) image processing algorithms have achieved superior performance, they are still difficult to apply on mobile devices (e.g., smartphones and cameras) due to the following reasons: 1) the high memory demand and 2) large model size. To adapt DL-based methods to mobile devices, motivated by the characteristics of image signal processors (ISPs), we propose a novel algorithm named LineDL. In LineDL, the default mode of the whole-image processing is reformulated as a line-by-line mode, eliminating the need to store large amounts of intermediate data for the whole image. An information transmission module (ITM) is designed to extract and convey the interline correlation and integrate the interline features. Furthermore, we develop a model compression method to reduce the model size while maintaining competitive performance; that is, knowledge is redefined, and compression is performed in two directions. We evaluate LineDL on general image processing tasks, including denoising and superresolution. The extensive experimental results demonstrate that LineDL achieves image quality comparable to that of state-of-the-art (SOTA) DL-based algorithms with a much smaller memory demand and competitive model size. Wenshu Chen, Liyuan Peng, Yuhao Liu 0001, Mingyu Wang 0001, Xiao-Ping Zhang 0002, Xiaoyang Zeng |
IEEE Trans. Image Process. | 5 |
| 2021 | Computing Utilization Enhancement for Chiplet-based Homogeneous Processing-in-Memory Deep Learning ProcessorsabstractThis paper presents a design strategy of chiplet-based processing-in-memory systems for deep neural network applications. Monolithic silicon chips are area and power limited, failing to catch the recent rapid growth of deep learning algorithms. The paper first demonstrates a straightforward layer-wise method that partitions the workload of a monolithic accelerator to a multi-chiplet pipeline. A quantitative analysis shows that the straightforward separation degrades the overall utilization of computing resources due to the reduced on-chiplet memory size, thus introducing a higher memory wall. A tile interleaving strategy is proposed to overcome such degradation. This strategy can segment one layer to different chiplets which maximizes the computing utilization. To facilitate the strategy, the modification of the chiplet system hardware is also discussed. To validate the proposed strategy, a nine-chiplet processing-in-memory system is evaluated with a custom-designed object detection network. Each chiplet can achieve a peak performance of 204.8GOPS at a 100-MHz rate. The peak performance of the overall system is 1.711TOPS, where no off-chip memory access is needed. By the tile interleaving strategy, the utilization is improved from 53.9 to 92.8 Bo Jiao 0003, Haozhe Zhu, Jinshan Zhang 0006, Shunli Wang 0001, Xiaoyang Kang 0001, Lihua Zhang 0002, Mingyu Wang 0001, Chixiao Chen |
ACM Great Lakes Symposium on VLSI | 7 |
| 2021 | Manifold constrained joint sparse learning via non-convex regularization
Jingjing Liu 0004, Xianchao Xiu, Wanquan Liu, Xiaoyang Zeng, Mingyu Wang 0001, Hui Chen 0007 |
Neurocomputing | 6 |