Jiayi Qian

dblp:392/9357 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0006-8238-8687ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Compositional AI Beyond LLMs: System Implications of Neuro-Symbolic-Probabilistic Architectures
abstract
Large Language Models (LLMs) have driven remarkable progress in artificial intelligence (AI), but their rapid growth faces challenges of unsustainable computation, limited robustness, and poor explainability. Compositional AI, which integrates LLMs with symbolic reasoning and probabilistic inference, has emerged as a promising paradigm to enable interpretability, robustness, trustworthiness, and data-efficient learning. Recent neuro-symbolic-probabilistic systems demonstrate strong potential in agentic applications, advancing reasoning and cognitive capabilities toward human-like intelligence.
Zishen Wan, Hanchen Yang 0001, Jiayi Qian, Ritik Raj, Joongun Park, Arijit Raychowdhury, Tushar Krishna
ASPLOS (1)3
2026 REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic Intelligence
abstract
Neuro-symbolic AI systems integrate neural perception with symbolic and probabilistic reasoning to enable dataefficient, interpretable, and robust intelligence beyond purely neural models. Although this compositional paradigm has shown superior performance in domains such as mathematical reasoning, planning, and verification, its deployment remains challenging due to severe inefficiencies in symbolic and probabilistic inference. Through systematic analysis of representative neurosymbolic workloads, we identify probabilistic logical reasoning as the inefficiency bottleneck, characterized by irregular control flow, low arithmetic intensity, uncoalesced memory accesses, and poor hardware utilization on CPUs and GPUs. This paper presents REASON, an integrated acceleration framework for probabilistic logical reasoning in neuro-symbolic AI. At the algorithm level, REASON introduces a unified directed acyclic graph representation that captures common structure across symbolic and probabilistic models, coupled with adaptive pruning and regularization. At the architecture level, REASON features a reconfigurable, tree-based processing fabric optimized for irregular traversal, symbolic deduction, and probabilistic aggregation. At the system level, REASON is tightly integrated with GPU streaming multiprocessors through a programmable interface and multi-level pipeline that efficiently orchestrates neural, symbolic, and probabilistic execution. Evaluated across six neurosymbolic workloads, REASON achieves 1 2 − 5 0 × speedup and 310-681 × energy efficiency over desktop and edge GPUs under TSMC 28 nm node. REASON enables real-time probabilistic logical reasoning, completing end-to-end tasks in 0.8 s with$6 \text{mm}^{2}$area and 2.12 W power, demonstrating that targeted acceleration of probabilistic logical reasoning is critical for practical and scalable neuro-symbolic AI and positioning REASON as a foundational system architecture for next-generation cognitive intelligence.
Zishen Wan, Che-Kai Liu, Jiayi Qian, Hanchen Yang 0001, Arijit Raychowdhury, Tushar Krishna
HPCA3
2026 Towards System-2 AI: Workloads and Characterizations of Energy-Based Models
abstract
The rapid progress of artificial intelligence (AI) has been largely driven by deep neural networks, yet their insufficient compositional reasoning abilities, limited robustness, and lack of explainability expose fundamental limitations for next-generation cognitive systems. Energy-Based Models (EBMs), adhering to System-2-style thinking, have recently shown remarkable performance across generative, compositional, and reasoning tasks, offering a pathway toward more robust and cognitively grounded AI. However, the system behavior of EBMs remains poorly understood. Their long sampling loops, strict execution data dependencies, model compositionality and heterogeneity make them highly inefficient on off-the-shelf hardware, creating significant barriers for scalable and real-time deployment. In this paper, we present the first comprehensive workload characterization and system study of EBMs. We first introduce a taxonomy that unifies all the diverse EBM algorithms (around twenty), then formalize their end-to-end runtime to expose key optimization opportunities, and experimentally profile representative models across CPU, GPU, and TPU platforms to understand their runtime breakdowns, memory behavior, computational operators, roofline model, algorithmic performance, and MCMC (Markov Chain Monte Carlo) sampling impacts. Our analysis reveals fundamental system bottlenecks unique to EBMs, including highly repetitive forward and backward neural operations, long and inefficient sampling trajectories, and severe hardware under-utilization caused by heterogeneous model components and operations. Furthermore, based on the experiments and analysis, we suggest optimization solutions across the algorithm, system and architecture levels, to improve EBM computing performance, efficiency, and scalability, aiming to pinpoint the challenges and future directions of EBM research.
Hanchen Yang 0001, Jiayi Qian, Zishen Wan, Jingtian Dang, Yilun Du, Tushar Krishna
ISPASS2
2026 2MCD-YOLO11: a lightweight algorithm based on cross-scale feature fusion for gesture recognition
Yueqiang Feng, Jiayi Qian, Wenxing Zuo, Yutong Tan, Zhengwei Shui
J. Supercomput.2
2025 ReCA: Integrated Acceleration for Real-Time and Efficient Cooperative Embodied Autonomous Agents
Zishen Wan, Yuhang Du, Mohamed Ibrahim 0002, Jiayi Qian, Jason Jabbour, Yang Zhao 0013, Tushar Krishna, Arijit Raychowdhury, Vijay Janapa Reddi
ASPLOS (2)4
2025 Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
abstract
Diffusion models have shown remarkable success across generative tasks, yet their high computational demands challenge deployment on resource-limited platforms. This paper investigates a critical question for compute-optimal diffusion model deployment: Under a post-training setting without fine-tuning, is it more effective to reduce the number of denoising steps or to use a cheaper per-step inference? Intuitively, reducing the number of denoising steps increases the variability of the distributions across steps, making the model more sensitive to compression. In contrast, keeping more denoising steps makes the differences smaller, preserving redundancy, and making post-training compression more feasible. To systematically examine this, we propose PostDiff, a training-free framework for accelerating pre-trained diffusion models by reducing redundancy at both the input level and module level in a post-training manner. At the input level, we propose a mixed-resolution denoising scheme based on the insight that reducing generation resolution in early denoising steps can enhance low-frequency components and improve final generation fidelity. At the module level, we employ a hybrid module caching strategy to reuse computations across denoising steps. Extensive experiments and ablation studies demonstrate that (1) PostDiff can significantly improve the fidelity-efficiency trade-off of state-of-the-art diffusion models, and (2) to boost efficiency while maintaining decent generation fidelity, reducing per-step inference cost is often more effective than reducing the number of denoising steps. Our code is available at https://github.com/GATECH-EIC/PostDiff.
Zhenbang Du, Yonggan Fu, Jiayi Qian, Yingyan (Celine) Lin
ICCV4
2025 Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability
abstract
Embodied systems, where generative autonomous agents engage with the physical world through integrated perception, cognition, action, and advanced reasoning powered by large language models (LLMs), hold immense potential for addressing complex, long-horizon, multi-objective tasks in realworld environments. However, deploying these systems remains challenging due to prolonged runtime latency, limited scalability, and heightened sensitivity, leading to significant system inefficiencies. In this paper, we aim to understand the workload characteristics of embodied agent systems and explore optimization solutions. We systematically categorize these systems into four paradigms and conduct benchmarking studies to evaluate their task performance and system efficiency across various modules, agent scales, and embodied tasks. Our benchmarking studies uncover critical challenges, such as prolonged planning and communication latency, redundant agent interactions, complex low-level control mechanisms, memory inconsistencies, exploding prompt lengths, sensitivity to self-correction and execution, sharp declines in success rates, and reduced collaboration efficiency as agent numbers increase. Leveraging these profiling insights, we suggest system optimization strategies to improve the performance, efficiency, and scalability of embodied agents across different paradigms. This paper presents the first system-level analysis of embodied AI agents, and explores opportunities for advancing future embodied system design.
Zishen Wan, Jiayi Qian, Yuhang Du, Jason Jabbour, Yilun Du, Yang Zhao 0013, Arijit Raychowdhury, Tushar Krishna, Vijay Janapa Reddi
ISPASS2
2024 AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment
abstract
Motivated by the transformative capabilities of large language models (LLMs) across various natural language tasks, there has been a growing demand to deploy these models effectively across diverse real-world applications and platforms. However, the challenge of efficiently deploying LLMs has become increasingly pronounced due to the varying application-specific performance requirements and the rapid evolution of computational platforms, which feature diverse resource constraints and deployment flows. These varying requirements necessitate LLMs that can adapt their structures (depth and width) for optimal efficiency across different platforms and application specifications. To address this critical gap, we propose AmoebaLLM, a novel framework designed to enable the instant derivation of LLM subnets of arbitrary shapes, which achieve the accuracy-efficiency frontier and can be extracted immediately after a one-time fine-tuning. In this way, AmoebaLLM significantly facilitates rapid deployment tailored to various platforms and applications. Specifically, AmoebaLLM integrates three innovative components: (1) a knowledge-preserving subnet selection strategy that features a dynamic-programming approach for depth shrinking and an importance-driven method for width shrinking; (2) a shape-aware mixture of LoRAs to mitigate gradient conflicts among subnets during fine-tuning; and (3) an in-place distillation scheme with loss-magnitude balancing as the fine-tuning objective. Extensive experiments validate that AmoebaLLM not only sets new standards in LLM adaptability but also successfully delivers subnets that achieve state-of-the-art trade-offs between accuracy and efficiency.
Yonggan Fu, Zhongzhi Yu, Jiayi Qian, Yongan Zhang, Xiangchi Yuan, Dachuan Shi, Roman Yakunin, Yingyan (Celine) Lin
NeurIPS4