Taoyi Wang

dblp:189/9848 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0003-1878-5451ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 4/8b High-Precision Fully-Parallel In-Sensor Computing Chip with Subthreshold Digital Pixel and Hybrid Pulse Modulation
Junda Zhao, Yimo Du, Taoyi Wang, Junzhan Liu, He Zhang 0011, Wang Kang 0001
ISCAS3
2026 Espresso: Exploiting the Sparsity Property in Brain-Inspired Vision Sensors With Spatiotemporal Ordering
abstract
Brain-inspired vision sensors (BVSs), drawing inspiration from the human visual system, produce sparse, high-temporal-resolution data stream capable of capturing rapid object motion. However, real-time processing of such data while preserving its inherent sparsity presents a critical challenge for practical deployment. A key challenge is to leverage the spatiotemporal correlations among event stream. Overlooking these correlations leads to prohibitively high query costs that even negate the benefits of sparsity, while exploiting them introduces complications such as input-output order conflicts and trade-offs between memory usage and latency. To address this, we present Espresso, an efficient hardware architecture that leverages the spatiotemporal order of events while explicitly preserving sparsity, achieving low-latency stream processing of event data. We first formalize a spatiotemporal order representation that identifies key features for stream processing on sparse events. Building on this, Espresso decouples output window address from input event address, resolving order conflicts via a dedicated queue mechanism and minimizing memory overhead through an optimized hash table. This enables immediate window-wise processing with minimal latency. To coordinate the pipeline, we design the Event-Scheduler, a streamlined finite state machine that prunes computations on zero values and aligns input-output stream order discrepancies. Integrated together, these modules deliver scalable, high-throughput processing for event-driven vision tasks. Espresso achieves up to 5000 fps, offering a 5.1× performance improvement over embedded GPUs. With parallel instantiations, it exceeds over 7000 fps in structured scenes and maintains over 2000 fps under complex-environment scenarios with minimal hardware overhead. These results establish Espresso as an efficient and scalable solution for real-time event-based vision processing, demonstrating the importance of spatiotemporal ordering in unlocking the full potential of BVSs.
Leshan Li, Taoyi Wang, Mingtao Ou, Xinglong Ji
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Diffusion-Based Extreme High-Speed Scenes Reconstruction with the Complementary Vision Sensor
Yapeng Meng, Taoyi Wang, Lijian Wang
ICCV3
2025 CSVO: Complementary-Pathway Spatial-Enhanced Visual Odometry for Extreme Environments with Brain-Inspired Vision Sensors
abstract
Visual Odometry (VO) estimates the pose and motion trajectory of the camera based on visual input, serving as a fundamental technique for robotic positioning and navigation. However, existing VO methods face challenges in visual degradation in extreme environments, e.g., high dynamic range or fast-motion conditions. Although event-based sensing schemes offer partial solutions to this problem, they are limited by unstable features and noise. Recently, a novel brain-inspired vision sensor, Tianmouc, has been reported, incorporating two complementary pathways: a cognition-oriented pathway (COP) for precise color intensity and an action-oriented pathway (AOP) for fast spatiotemporal sensing, considered a promising visual input for VO tasks. Here, we develop Complementary Pathway Spatial Enhanced Visual Odometry (CSVO) to cope with extreme scenarios by fusing the COP and AOP information of Tianmouc. To leverage the dynamic range expansion brought about by dual-pathway fusion, as well as the low-latency spatial difference data in AOP to address high-speed motion, we design an asynchronous dual-pathway feature encoder considering synchronous multimodal fusion and asynchronous cross-modal feature matching. To train and evaluate CSVO, we transform two conventional VO datasets, TartanAir and Apollo, to Tianmouc modality through simulation and collect a real- world Tianmouc-VO dataset in challenging scenes. Our results demonstrate state-of-the-art performance over existing methods on these datasets. Our work sheds light on the generalizability of agents working in extreme scenarios. The codes and data sets are available at https://github.com/Tianmouc/CSVO.
Taoyi Wang
IROS4
2024 HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks
abstract
The rapid development of deep learning has propelled many real-world artificial intelligence applications. Many of these applications integrate multiple neural networks (multi-NN) to cater to various functionalities. There are two challenges of multi-NN acceleration: (1) competition for shared resources becomes a bottleneck, and (2) heterogeneous workloads exhibit remarkably different computing-memory characteristics and various synchronization requirements. Therefore, resource isolation and fine-grained resource allocation for each task are two fundamental requirements for multi-NN computing systems. Although a number of multi-NN acceleration technologies have been explored, few can completely fulfill both of these requirements, especially for mobile scenarios. This paper reports a Hierarchical Asynchronous Parallel Model (HASP) to enhance multi-NN performance to meet both requirements. HASP can be implemented on a multicore processor that adopts Multiple Instruction Multiple Data (MIMD) or Single Instruction Multiple Thread (SIMT) architectures, with minor adaptive modification needed. Further, a prototype chip is developed to validate the hardware effectiveness of this design. A corresponding mapping strategy is also developed, allowing the proposed architecture to simultaneously promote resource utilization and throughput. With the same workload, the prototype chip demonstrates 3.62$\boldsymbol{\times}$, and 3.51$\boldsymbol{\times}$higher throughput over Planaria and 8.68$\boldsymbol{\times}$, 2.61$\boldsymbol{\times}$over Jetson AGX Orin for MobileNet-V1 and ResNet50, respectively.
Songchen Ma, Taoyi Wang, Guanrui Wang, Chenhang Song, Huanyu Qu, Junfeng Lin, Jing Pei
IEEE Trans. Computers3
2023 CBKI: A confidence-based knowledge integration framework for multi-choice machine reading comprehension
Xianghui Meng, Yang Song 0010, Qingchun Bai, Taoyi Wang
Knowl. Based Syst.4