Songchen Ma

dblp:322/8380 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-7913-046XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Optimizing Spatial Data Structure with Near-Cache Acceleration by Exploiting Physical Locality
Haoran Pei, Zijian Pan, Songchen Ma, Leshan Li, Xinglong Ji
ISCA6
2026 Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit Computation
abstract
Large Language Models (LLMs) have transformed society, but their computational and energy needs hinder efficient inference. The memory wall, the growing processor-memory speed disparity, remains a critical bottleneck for LLM. While Process-in-Memory (PIM) architectures address this challenge by co-locating computation with memory, achieving 5-20 × higher bandwidth than GPUs, existing scalable PIM solutions face critical trade-offs in flexibility, capacity, and efficiency when handling LLMs' dynamic memory-compute patterns and operator diversity. DRAM-PIM suffers from inter-bank communication overhead despite its vector parallelism. SRAM-PIM offers sub10ns latency for matrix operation but is constrained by limited capacity. This work introduces CompAir, a scalable PIM architecture that integrates DRAM-PIM and SRAM-PIM with hybrid bonding, enabling efficient linear computations while unlocking multi-granularity data pathways. We further develop CompAirNoC, an advanced network-on-chip (NoC) with an embedded arithmetic logic unit that performs non-linear operations during data movement. Such a design offloads the centralized communication bottleneck in the channel level to distributed banks, simultaneously reducing communication overhead and area cost for scalability. Finally, we develop a hierarchical Instruction Set Architecture that ensures both flexibility and programmability of the hybrid PIM. Experiments show CompAir delivers 1.83-7.98× faster prefill and 1.95 - 6.28 × faster decoding versus state-ofthe-art PIM designs, with 3.52 × lower energy than GPU-PIM hybrids. This work presents the first systematic exploration of hybrid DRAM-PIM and SRAM-PIM architectures with innetwork computation, paving the way towards a scalable PIM system for LLM inference.
Songchen Ma, Huanyu Qu, Jia Chen 0032, Junfeng Lin, Fengbin Tu
ISCA2
2025 CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator Design
abstract
The rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture.
Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu
ICCAD1
2025 RoboSpike: Fully Utilizing the Heterogeneous System With Subcallback Scheduling in ROS 2
abstract
The advancement in artificial intelligence (AI) has greatly propelled the development of robotics, requiring the adoption of heterogeneous computing architectures with multicore CPUs, GPUs, and accelerators to meet the growing computational needs of edge computing. Such heterogeneity, coupled with the inherently IO-intensive nature of robotic applications, poses substantial challenges for task scheduling and resource management. These challenges are particularly acute for systems striving to maximize computational resource utilization, which cannot be effectively addressed through callback-level scheduling. To overcome these obstacles, we developed RoboSpike, a systematic solution built on the Robot Operating System 2 (ROS 2). We first implemented a subcallback scheduling mechanism utilizing coroutines to utilize the blocked CPUs which wait for I/O operations. Building on this mechanism, we extended the design to incorporate the coprocessor and introduced an auto-tuning algorithm to adapt to system performance variations. Finally, we performed the response time analysis to ensure that the RoboSpike is predictable in time. The evaluation results demonstrate that RoboSpike achieves substantial improvements, increasing throughput by 1.65–2.25 times in real-world scenarios. RoboSpike enhances the scheduling capabilities of ROS 2 by refining the granularity from the callback level, thus opening up new opportunities for performance improvement in robotic systems, especially in resource-limited scenarios with complex workloads.
Songchen Ma, Xinglong Ji
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 HASP: Hierarchical Asynchronous Parallelism for Multi-NN Tasks
abstract
The rapid development of deep learning has propelled many real-world artificial intelligence applications. Many of these applications integrate multiple neural networks (multi-NN) to cater to various functionalities. There are two challenges of multi-NN acceleration: (1) competition for shared resources becomes a bottleneck, and (2) heterogeneous workloads exhibit remarkably different computing-memory characteristics and various synchronization requirements. Therefore, resource isolation and fine-grained resource allocation for each task are two fundamental requirements for multi-NN computing systems. Although a number of multi-NN acceleration technologies have been explored, few can completely fulfill both of these requirements, especially for mobile scenarios. This paper reports a Hierarchical Asynchronous Parallel Model (HASP) to enhance multi-NN performance to meet both requirements. HASP can be implemented on a multicore processor that adopts Multiple Instruction Multiple Data (MIMD) or Single Instruction Multiple Thread (SIMT) architectures, with minor adaptive modification needed. Further, a prototype chip is developed to validate the hardware effectiveness of this design. A corresponding mapping strategy is also developed, allowing the proposed architecture to simultaneously promote resource utilization and throughput. With the same workload, the prototype chip demonstrates 3.62$\boldsymbol{\times}$, and 3.51$\boldsymbol{\times}$higher throughput over Planaria and 8.68$\boldsymbol{\times}$, 2.61$\boldsymbol{\times}$over Jetson AGX Orin for MobileNet-V1 and ResNet50, respectively.
Songchen Ma, Taoyi Wang, Guanrui Wang, Chenhang Song, Huanyu Qu, Junfeng Lin, Jing Pei
IEEE Trans. Computers2
2024 SongC: A Compiler for Hybrid Near-Memory and In-Memory Many-Core Architecture
abstract
Building hybrid systems that incorporate various processing-in-memory (PIM) devices and processing-near-memory (PNM) technologies can offer complementary advantages in both efficiency and flexibility, while many-core architectures show great potential in deploying data-centric parallel applications with high performance. Compilers for the hybrid PN/IM architecture are critical for enabling such computing systems to be put into practical use. However, most of the existing neural network compilers for PIM or PNM are optimized from the perspective of an operator, and cannot effectively take advantage of a decentralized core-level dataflow with large on-chip memory access bandwidth. Here, we propose a full-stack System-on-graph Compiler (SongC) framework for many-core architecture, which optimizes the efficiency of the PIM devices and leverages the flexibility of the PNM architectures. SongC establishes multi-level graph abstractions to clarify the critical deployment challenges at different levels and generalizes the standard optimizations, decoupling versatile algorithms and diverse types of hardware. To handle the complexity of many-core resource utilization, we also establish a simulation-compilation interaction flow, including a just-in-time evaluator to boost the scheduling search and an extended Roofline model, referred to as the Palace model, to guide the search. Experiments demonstrate the various optimizations and overall performance of SongC and reveal the capability of strategy exploration.
Junfeng Lin, Huanyu Qu, Songchen Ma, Xinglong Ji, Xiaochuan Li 0003, Chenhang Song
IEEE Trans. Computers3