EDBT 2026 Demo / reviewers in the wild / expert
Tianyang Yu
dblp:247/6252
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neuromorphic Hyperdimensional Computing for Efficiently Processing Event-Based DataabstractThe neuromorphic sensor’s event-based data output offers significant benefits, including minimal data redundancy and exceptional time resolution, which guarantee low power consumption and heightened sensitivity during the data acquisition process. Spiking neural network (SNN), with its inherent event-driven characteristic, is well-suited for processing event-based data, and its spike-based computing mechanism enhances the efficiency of data processing. Recent studies are exploring the integration of brain-inspired hyperdimensional computing (HDC) with SNN to leverage HDC’s advantages, including the low inference and training complexity, aiming to further reduce hardware overhead associated with SNN deployment. However, existing works have not effectively harnessed the information output by SNN during hyperdimensional encoding, leading to considerable area and energy overhead. In this article, an efficient neuromorphic HDC method is proposed, featuring a simplified hyperdimensional encoding approach that considers the temporal dynamics of SNN. In addition, a lightweight accelerator design matching the proposed method is also given. Experimental results show that the proposed accelerator achieves over 50% area reduction and reduces energy consumption by 20%–90%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | MIRACLE: Multimodal Information Retrieval via a Combined In-Memory Processing and Content Addressable Memory ApproachabstractThe rapid advancement of information technology has brought multimodal information retrieval into the research spotlight. Neural networks, particularly Transformers, have emerged as the dominant solution for extracting multimodal feature vectors. While neural network acceleration has been extensively explored, the subsequent retrieval stage in multimodal scenarios remains under-optimized. Conventional retrieval approaches, such as cosine similarity sorting on von Neumann architectures, suffer from significant data migration and computational inefficiencies. Hashing methods enhance storage and computation efficiency but encounter challenges in energy-efficient implementation and mitigating accuracy losses due to modal heterogeneity. This paper presents a hybrid architecture that integrates in-memory processing (PIM) and content-addressable memory (CAM) to address these challenges. Transformer-extracted features are processed via in-memory random hashing leveraging device-intrinsic properties, with CAM facilitating parallel search space reduction. A final cosine similarity reranking stage refines the results while balancing accuracy with energy efficiency. Experimental evaluations validate that the proposed method, when compared to the baseline traditional CPU-based cosine similarity retrieval, 1) achieves almost identical level of accuracy, dramatically outperforming other pure CAMbased Hamming distance retrieval approaches; and 2) reduces latency by $9.45 \times$ and energy consumption by $30.20 \times$. Xuehui Liu, Tianyang Yu, Shuo Ran, Bi Wu 0002, Xiaotao Jia, Weiqiang Liu 0001, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 3 |
| 2025 | Learning-Based Realtime Synthetic Aperture Radar Imaging for Embedded System on SatelliteabstractSpaceborne Synthetic Aperture Radar (SAR), due to its ability of all-weather and all-day sensing, is extensively utilized across various fields. Realtime imaging on satellite is highly advantageous as it significantly reduces the communication cost and delay between satellite and ground station, making it exceptionally suitable for time-sensitive applications like maritime search and rescue. However, the limited computing power of embedded system on satellite, together with the substantial computational demands of traditional imaging algorithms, pose challenges for realtime imaging. To address these issues, a lightweight learning-based model with adjustable complexity is proposed for realtime imaging. Furthermore, we collect a large amount of real-world echo data from satellite and construct the first large-scale dataset, for training and evaluating the learning-based SAR imaging task. Experimental results and evaluation in the embedded system show that, the proposed learning-based model achieves up to 19× speedup than traditional algorithm. Tianyang Yu, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 1 |
| 2025 | LAHDC: Logic-Aggregation-Based Query for Embedded Hyperdimensional Computing AcceleratorabstractWith low complexity and robustness, hyperdimensional computing (HDC) has become a promising paradigm for edge-side applications. HDC employs hypervectors (generally with 2–10 K dimensions) to represent input samples, and performs logical operations in hyperdimensional space to complete perceptual tasks. Compared to deep neural network (DNN), HDC is more suitable for lightweight edge-side applications (i.e., speech, activity recognition), due to its low complexity and less computational scheduling. However, existing HDC’s querying process relies on trained class hypervectors, resulting in on-chip storage and transmission overhead which limits the application of ASIC-based or FPGA-based HDC accelerators in embedded systems. In this article, a logic-aggregation-based query method called LAHDC is proposed to eliminate such overhead. In addition, an ultratiny HDC accelerator design matching LAHDC is also proposed, as well as an automated tool to search for optimal structure and generate hardware design code. Experimental results show that, compared to existing ASIC-based HDC accelerators, the proposed design reduce the area/energy by more than 95%/80%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Utilizing Large Language Models (LLMs) in Data Analysis Pipeline for Digital Phenotyping: Description, Prediction, and VisualizationabstractDigital phenotyping is the "moment-by-moment quantification of the individual-level human phenotype in situ using data from personal digital devices," according to Onnela and Rauch. Digital phenotyping research has historically contained many entry barriers due to high costs and complexity. However, the growing popularity of personal devices such as mobile phones has enabled researchers to collect participant data with more convenience and lower costs than ever before. This paper presents the Intelligent Phenotype Analysis Suite (IPAS), a new AI-powered tool for streamlined phenotype investigation. This tool provides researchers the ability to generate descriptions, predictions, and visualizations of their digital phenotyping data through a simple, novel chat-bot interface. IPAS combines an array of data science techniques with natural language processing capabilities of large language models to accelerate the data analysis process for researchers. IPAS extracts raw Beiwe data, an intuitive data collection platform which only requires participants to install an application on their mobile phone and permit data collection. Furthermore, we evaluate the accuracy of IPAS by using key LLM performance metrics: Precision, Recall, and F1 Score. While testing, IPAS sometimes struggled to generate code. However, in all test prompts, IPAS correctly identified the pre-written function needed to perform the requested task. Altogether, IPAS improves prior methods by enabling researchers of all levels of experience to analyze digital phenotyping data using natural language queries. Derek Nissen, Tianyang Yu, Reyva Babtista, Yi Shang |
IEEE Big Data | 2 |
| 2024 | A Combined Content Addressable Memory and In-Memory Processing Approach for k-Clique Counting Accelerationabstractk-Clique counting problem plays an important role in graph mining which has seen a growing number of applications. However, current k-Clique counting accelerators cannot meet the performance requirement mainly because they struggle with high data transfer issue incurred by the intensive set intersection operations and the inability of load balancing. In this paper, we propose to solve this problem with a hybrid framework of content addressable memory (CAM) and in-memory processing (PIM). Specifically, we first utilize CAM for binary induced subgraph generation in order to reduce the search space, then we use PIM to implement in-place parallel k-Clique counting through iterative Boolean logic "AND" like operation. To take full advantage of this combined CAM and PIM framework, we develop dynamic task scheduling strategies that can achieve near optimal load balancing among the PIM arrays. Experimental results demonstrate that, compared with state-of-the-art CPU and GPU platforms, our approach achieves speedups of 167.5× and 28.8×, respectively. Meanwhile, the energy efficiency is improved by 788.3× over the GPU baseline. Xidi Ma, Tianyang Yu, Bi Wu 0002, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 4 |
| 2024 | Fully Learnable Hyperdimensional Computing Framework With Ultratiny Accelerator for Edge-Side ApplicationsabstractBrain-inspired hyperdimensional computing (HDC) is a new computational paradigm that encodes input sample into a hypervector (generally with dimensions of$2K-10K$), and performs simple arithmetic and logic operations in the hyperdimensional space to complete perceptual tasks like human brain. Due to its simplicity, interpretability, and robustness, HDC has gradually become a competitor and substitute for deep neural network (DNN) in many tasks. However, there exists an accuracy gap between existing heuristic HDC algorithms and DNN in computer vision tasks, as existing encoding methods have difficulty in filtering out large amount of background and noise in the images, and effectively extracting the spatial structure features of images. In addition, the existing hardware for HDC deployment mainly focuses on in-memory computing (IMC), application specific integrated circuit (ASIC), or high-capacity field programmable gate array (high-capacity FPGA), which cannot meet the flexibility, small area, and low power requirements of edge-side applications. In this paper, a fully learnable HDC framework with learnable preprocessing, encoding and querying, is proposed to boost the accuracy in computer vision tasks, as well as an ultra-tiny accelerator based on edge-side FPGA which matches the proposed framework. Experiments show that on multiple commonly-used image datasets, the proposed HDC framework has an average computation reduction of 80% compared to other most advanced strategies, while achieves a 1.2% accuracy increase. Evaluation on edge-side FPGA shows that compared to other FPGA based state-of-the-art designs, the proposed accelerator saves more than$10\boldsymbol{\times}$hardware resource and power consumption. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Edge-Side Fine-Grained Sparse CNN Accelerator With Efficient Dynamic Pruning SchemeabstractWith the rapid development of the Internet of Things (IoT), it has become a common concern of academia and industry to provide real-time high performance services for edge-side applications and to bestow intelligence on massive edge-side devices. Due to the limitations of storage space, volume and power consumption of edge side devices, it is difficult for existing convolutional neural networks with large number of parameters and large amount of computation to match them. Network pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, fine-grained pruning is not hardware friendly, while other structured pruning schemes will result in a much higher loss of accuracy under the same compression ratio. In this paper, an model compression strategy is given including the proposed efficient fine-grained pruning scheme, a dynamic pruning & training method, and a weight importance judgment method. Depending on this strategy, sparse VGG16 (ResNet50) model can be obtained by training from scratch, and achieves a total of$16\times $compression ratio with 1/32 indexing overhead. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed. Implementing VGG16 and ResNet50 on the proposed accelerator, the experimental results show that compared with the most advanced design, the proposed accelerator can achieve 8.13 Frames Per Second (FPS) with$2.17\times $better power efficiency and at most$4.14\times $better calculation density. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | A Potential Enabler for High-Performance In-Memory Multi-Bit Arithmetic Schemes With Unipolar Switching SOT-MRAMabstractDue to the physical separation of data processing and storage, the conventional Von Neumann architecture exists excessive data migration overhead to curtail the progress of data-intensive applications. In this way, the Computing-in-Memory (CiM) architecture is proposed. Due to the boolean property of the memory cell, the current CiM mainly focuses on single-bit logic design. For the multi-bit arithmetic design, a prevalent patchwork approach is employed using single-bit logic, leaving the design with insufficient parallelism. This paper proposes a high-performance in-memory multi-bit addition (M-Add) and multiplication (M-Mul) scheme based on unipolar switching SOT-MRAM. For the M-Add scheme, transmission logic-based circuit design is proposed to realize single-step inter-column XOR operations, which is logically fits perfectly the g operator of parallel prefix algorithm. Further, the oBK algorithm is presented to maximize the g operator occupancy. For the M-Mul scheme, mapping the Booth decoder to the control signal of the proposed modified flip-flop queue, only two steps are required to realize the decoding of three encoded signals in parallel. The simulation results indicate the proposed design reduces the latency of N-bit Add (N-bit Mul) by an average of 82.6% (31.5%) compared to state-of-the-art CiM designs. Further, a CNN application based on proposed operations achieves 1.23 TOPS/w on the CIFAR-10 dataset, with an average of 47.57% increase over other CiM designs. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Toward Efficient Retraining: A Large-Scale Approximate Neural Network Framework With Cross-Layer OptimizationabstractLeveraging approximate multipliers in approximate neural networks (ApproxNNs) can effectively reduce hardware area and power consumption, making them suitable for edge-side applications. However, the propagation of layer-by-layer errors limits the application of approximate multipliers to large-scale ApproxNNs and complex tasks. Currently, retraining techniques that consider approximate multiplication errors are commonly used to compensate for the accuracy loss. However, due to the irregularity of the errors introduced by approximate multiplier, it is difficult for the existing generic acceleration hardware (e.g., GPU) to efficiently simulate its function and accelerate retraining, which thereby leads to a huge retraining overhead in ApproxNNs’ application. In this article, we propose an ApproxNN framework that introduces errors with regular and controlled positions for high-efficiency retraining of large-scale ApproxNNs. An approximate multiplier design that matches this framework is also presented to verify the effectiveness of the proposed ApproxNN framework. Experiment results demonstrate that the proposed ApproxNN framework is able to achieve up to 46$\times$speedup in retraining, and the proposed approximate multiplier reduces area/power-delay product (PDP) by 31%/63% compared to the exact multiplier. Compared with the floating-point neural network (NN) model, an accuracy decrease of only 1.13% is achieved when applied to ResNet50 on ImageNet dataset with only 15-epochs retraining, which surpasses other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Data Stream Oriented Fine-grained Sparse CNN Accelerator with Efficient Unstructured Pruning StrategyabstractNetwork pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, unstructured pruning is not hardware friendly, while structured pruning will result in a significant loss of accuracy. In this paper, an unstructured fine-grained pruning strategy is proposed and achieves a 16X compression ratio with a top-1 accuracy loss of 1.4% for VGG-16. Combined with the proposed hardware-oriented hyperparameter selection method, compression rates of up to 64X can be obtained while fully meeting the edge-side accuracy requirements. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed for pruned VGG-16. The experimental results show that compared with the most advanced design, the proposed accelerator can achieve 21 Frames Per Second (FPS) with 3X better power efficiency and 2.19X better calculation density. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Energy-efficient Oriented Approximate Quantization Scheme for Fine-Grained Sparse Neural Network AccelerationabstractFor edge-side applications with severe power constraints, using pruning and quantization to compress models while maintaining network accuracy has become a widely deployed form of Convolutional Neural Networks (CNNs). For the same model accuracy, fine-grained non-regular pruning can bring higher model compression rate than coarse-grained regular pruning, but also introduces a larger indexing overhead. Besides, this overhead increases dramatically as the pruning granularity decreases. In this work, an approximate quantization scheme for fine-grained pruning is proposed. By reusing part of the quantized data bits, the proposed scheme can merge quantized data and indexes approximately, reducing the indexing overhead as well. Meanwhile, since the approximate compensation of index bits, the proposed scheme achieves an effective model accuracy improvement compared to the case of direct quantization to low bit-width. Experimental results show that, for 2:4 fine-grained pruning and 8-bit quantization scenario, the proposed method can save 20% of memory space and transmission cost. Compared with the direct quantization to 6-bit approach, the proposed scheme improves the accuracy by nearly 0.5% in the simulation of ImageNet dataset on ResNet50, despite occupying the same storage space. When deploying Yolov2-tiny at 16 × compression ratio, the energy efficiency of the CNN accelerator with the proposed approximate quantization is 1.33-3.82× that of other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ICCD | 1 |
| 2021 | Efficient Reinforcement Learning Development with RLzooabstractMany multimedia developers are exploring for adopting Deep Reinforcement Learning (DRL) techniques in their applications. They however often find such an adoption challenging. Existing DRL libraries provide poor support for prototyping DRL agents (i.e., models), customising the agents, and comparing the performance of DRL agents. As a result, the developers often report low efficiency in developing DRL agents. In this paper, we introduce RLzoo, a new DRL library that aims to make the development of DRL agents efficient. RLzoo provides developers with (i) high-level yet flexible APIs for prototyping DRL agents, and further customising the agents for best performance, (ii) a model zoo where users can import a wide range of DRL agents and easily compare their performance, and (iii) an algorithm that can automatically construct DRL agents with custom components (which are critical to improve agent's performance in custom applications). Evaluation results show that RLzoo can effectively reduce the development cost of DRL agents, while achieving comparable performance with existing DRL libraries. Tianyang Yu, Hongming Zhang 0003, Yanhua Huang, Quancheng Guo, Luo Mai, Hao Dong 0003 |
ACM Multimedia | 2 |