Zhiwen Mo

dblp:99/3235 · also Zhi-Wen Mo · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
YearPublicationVenuePosition
2026 FastTTS: Accelerating Test-Time Scaling for Edge LLM Reasoning
Hao Mark Chen, Zhiwen Mo, Guanxi Lu, Shuang Liang 0012, Lingxiao Ma, Wayne Luk, Hongxiang Fan
ASPLOS (2)2
2026 Coset Ensemble Decoder for Quantum Error Correction with Algorithm-Hardware Co-Design
Shuang Liang 0012, Jubo Xu, Giulio Bassanino, Qianzhou Wang, Yuncheng Lu, Zhiwen Mo, Paul H. J. Kelly, Wayne Luk, Hongxiang Fan
ISCA7
2026 Combating the Memory Walls: Optimization Pathways for Long-Context Agentic Llm Inference
abstract
Large Language Models (LLMs) serve as the core components of AI agents used across a wide range of applications, including enterprise workflow automation, software engineering, web automation, computer use, and research. These agentic LLM inference tasks are fundamentally different from traditional chatbot-focused inference — they often have much larger context lengths to capture complex, prolonged inputs, such as an entire webpage DOM or complicated tool call trajectories. This, in turn, generates significant off-chip memory traffic for hardware at the inference stage and causes the workload to be constrained by the two memory walls, namely the bandwidth and capacity walls, preventing the compute units from achieving high utilization. In this paper, we introduce PLENA, a hardware–software co-designed system that applies three core optimization pathways. PLENA features a novel flattened systolic-array architecture (Pathway 1) and efficient compute and memory units that support an asymmetric quantization scheme (Pathway 2). It also provides native support for FlashAttention (Pathway 3). In addition, PLENA is developed with a complete software–hardware stack, including a custom ISA, a compiler, a transaction-level simulator, and an automated design-space exploration flow. Experimental results show that PLENA delivers up to 2.23× and 4.70× higher throughput than the A100 GPU and TPU v6e, respectively, under identical multiplier counts and memory configurations during LLaMA agentic inference. PLENA also achieves up to 4.04× higher energy efficiency than the A100 GPU.
Can Xiao, Jiayi Nie, Binglei Lou, Jeffrey T. H. Wong, Zhiwen Mo, Przemyslaw Forys, Chengyang Ai, Timi Adeniran, Wayne Luk, Hongxiang Fan, Jianyi Cheng, Timothy M. Jones 0001, Rika Antonova, Robert Mullins 0001, Aaron Zhao
ISCA7
2026 Three-dimensional dual hesitant fuzzy constructions of underlying operations, aggregation operators, decision frameworks and their applications of multiple attribute group decision-making
Huarong Feng, Xianyong Zhang, Zhiwen Mo
Eng. Appl. Artif. Intell.4
2025 LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
abstract
Large Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research.
Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004
ISCA1
2025 Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
abstract
Test-time scaling (TTS) has proven effective in enhancing the reasoning capabilities of large language models (LLMs). Verification plays a key role in TTS, simultaneously influencing (1) reasoning performance and (2) compute efficiency, due to the quality and computational cost of verification. In this work, we challenge the conventional paradigms of verification, and make the first attempt toward systematically investigating the impact of verification granularity—that is, how frequently the verifier is invoked during generation, beyond verifying only the final output or individual generation steps. To this end, we introduce Variable Granularity Search (VG-Search), a unified algorithm that generalizes beam search and Best-of-N sampling via a tunable granularity parameter $g$. Extensive experiments with VG-Search under varying compute budgets, generator-verifier configurations, and task attributes reveal that dynamically selecting $g$ can improve the compute efficiency and scaling behavior. Building on these findings, we propose adaptive VG-Search strategies that achieve accuracy gains of up to 3.1\% over Beam Search and 3.6\% over Best-of-N, while reducing FLOPs by over 52\%. We will open-source the code to support future research.
Hao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo, Masato Motomura, Hongxiang Fan
NeurIPS4
2025 PipeThreader: Software-Defined Pipelining for Efficient DNN Execution
Yu Cheng 0030, Lei Wang 0222, Yining Shi 0001, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang 0053, Zhiwen Mo, Fan Yang 0024, Mao Yang 0004, Zhi Yang 0001
OSDI8
2025 Feature selections based on uncertainty measurements from dual-quantitative improvement and double-hierarchical fusion
Xianyong Zhang, Zhiying Lv, Zhiwen Mo
Appl. Intell.4
2025 Improved attribute reductions based on combined measures from dependency-roughness-entropy fusion and classification-class-level integration
Yixiao Yuan, Xianyong Zhang, Zhiwen Mo
Appl. Intell.4
2024 Enabling Multiple Tensor-wise Operator Fusion for Transformer Models on Spatial Accelerators
abstract
In transformer models, data reuse within an operator is insufficient, which prompts more aggressive multiple tensor-wise operator fusion (multi-tensor fusion). Due to the complexity in tensor-wise operator dataflow, conventional fusion techniques often fall short by limited dataflow options and short fusion length. In this study, we first identify three challenges on multi-tensor fusion that result in inferior fusions. Then we propose dataflow adaptive tiling (DAT), a novel inter-operator dataflow to enable an efficient fusion of multiple operators connected in any form and chained in any length. Then, we broaden the dataflow exploration from intraoperator to inter-operator and develop an exploration framework to quickly find the best dataflow on spatial accelerators with given on-chip buffer size. Experiment results show that DAT delivers 2.24× and 1.74× speedup and 35.5% and 15.5% energy savings on average for edge and cloud accelerators, respectively, comparing to the state-of-the-art dataflow explorer FLAT. DAT is open-sourced at https://github.com/lxu28973/DAT.git.
Zhiwen Mo, Qin Wang 0009, Jianfei Jiang 0001, Naifeng Jing
DAC2
2023 Novel Aczel-Alsina operations-based linguistic Z-number aggregation operators and their applications in multi-attribute group decision-making process
Bo Chen 0040, Qiang Cai 0002, Guiwu Wei 0001, Zhiwen Mo
Eng. Appl. Artif. Intell.4
2023 An extended Exp-TODIM method for multiple attribute decision making based on the Z-Wasserstein distance
Qiang Cai 0002, Guiwu Wei 0001, Zhiwen Mo
Expert Syst. Appl.5
2012 Comparative study of variable precision rough set model and graded rough set model
Xianyong Zhang, Zhiwen Mo, Fang Xiong
Int. J. Approx. Reason.2
2009 Products of Mealy-type fuzzy finite state machines
Zhiwen Mo, Dong Qiu, Yang Wang 0053
Fuzzy Sets Syst.2
2009 On starshaped fuzzy sets
Dong Qiu, Lan Shu, Zhiwen Mo
Fuzzy Sets Syst.3
2009 Notes on fuzzy complex analysis
Dong Qiu, Lan Shu, Zhiwen Mo
Fuzzy Sets Syst.3
2007 Notes on "the lower and upper approximations in a fuzzy group" and "rough ideals in semigroups"
Zhiwen Mo
Inf. Sci.2
2004 Minimization algorithm of fuzzy finite automata
Zhiwen Mo
Fuzzy Sets Syst.2
2002 On convexity of fuzzy syntopogenous spaces
Zhiwen Mo
Fuzzy Sets Syst.1
1995 Fuzzy Ideals Generated by Fuzzy Sets in Semigroups
Zhiwen Mo, Xue-ping Wang 0001
Inf. Sci.1
1993 On pointwise depiction of fuzzy regularity of semigroups
Zhiwen Mo, Xue-ping Wang 0001
Inf. Sci.1