Yasuyuki Okoshi

dblp:316/3509 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0005-8472-7841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 55% Deep learning architectures and training · 33% Language models and text generation · 12%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 67% Hardware accelerators and domain-specific architectures · 33%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.422025
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression · NeurIPS 2025
Multicoated Supermasks Enhance Hidden Networks · ICML 2022
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention
1.012026
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms · AAAI 2026
Machine learning › Deep learning architectures and training › overparameterized neural network
strong lottery ticket hypothesis
1.012026
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms · AAAI 2026
Machine learning › Deep learning architectures and training
transformer
1.012026
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms · AAAI 2026
Hardware accelerators and domain-specific architectures › quantization
activation quantization
1.012026
AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization · HPCA 2026
Memory systems › cache
key-value cache
1.012026
AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization · HPCA 2026
Memory systems
processing-in-memory
1.012026
AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization · HPCA 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.912025
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression · NeurIPS 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.612022
Multicoated Supermasks Enhance Hidden Networks · ICML 2022
Machine learning › Efficient and distributed learning › model compression › weight masking
supermask
0.612022
Multicoated Supermasks Enhance Hidden Networks · ICML 2022
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.312026
AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization · HPCA 2026
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
0.312026
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms · AAAI 2026
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.312025
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

product quantization · 2.0clustering-based vector quantization · 2.0theoretical approximation analysis · 1.0random subnetwork pruning · 1.0matrix quantization · 0.9binary coding · 0.9best-of-n sampling · 0.9beam search · 0.9adaptive search · 0.9backpropagation · 0.6
YearPublicationVenuePosition
2026 The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
abstract
The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks. Although recent theoretical studies have established the SLTH across various neural architectures, the SLTH for transformer architectures still lacks theoretical understanding. In particular, the current theory of the SLTH does not yet account for the multi-head attention (MHA) mechanism, a core component of transformers. To address this gap, we introduce a theoretical analysis of the existence of SLTs within MHAs. We prove that, if a randomly initialized MHA of H heads and input dimension d has the hidden dimension O(d log(Hd^(3/2))) for the key and value, it contains an SLT that approximates an arbitrary MHA with the same input dimension with high probability. Furthermore, by leveraging this theory for MHAs, we extend the SLTH to transformers without normalization layers. We empirically validate our theoretical findings, demonstrating that the approximation error between the SLT within a source model (MHA and transformer) and an approximate target counterpart decreases exponentially by increasing the hidden dimension of the source model.
Hikari Otsuka, Daiki Chijiwa, Yasuyuki Okoshi, Daichi Fujiki, Susumu Takeuchi, Masato Motomura
AAAI3
2026 AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization
abstract
Processing-in-Memory (PIM) architectures offer a promising solution to the memory bottlenecks in data-intensive machine learning, yet often overlook the growing challenge of activation memory footprint. Conventional PIM approaches struggle with massive KV cache sizes generated in long-context scenarios by Transformer-based models, frequently exceeding PIM's limited memory capacity, while techniques like sparse attention can conflict with PIM's need for data locality. Existing PIM approaches and quantization methods are often insufficient or poorly suited for leveraging the unique characteristics of activations. This work identifies an opportunity for PIMspecialized activation quantization to enhance bandwidth and compute efficiency. We explore clustering-based vector quantization approaches, which align well with activation characteristics and PIM's internal bandwidth capabilities. Building on this, we introduce AQPIM, a novel PIM-aware activation quantization framework based on Product Quantization (PQ), optimizing it for modern Large Language Models (LLMs). By performing quantization directly within memory, AQPIM leverages PIM's high internal bandwidth and enables direct computation on compressed data, significantly reducing both memory footprint and computational overhead for attention computation. AQPIM addresses PQ's accuracy challenges by introducing several algorithmic optimizations. Evaluations demonstrate that AQPIM achieves significant performance improvements, drastically reducing of GPU-CPU communication that can account for$90 \sim 98.5 \%$of decoding latency, together with$3.4 \times$speedup over a SOTA PIM approach.
Kosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi Fujiki
HPCA2
2025 Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
abstract
Test-time scaling (TTS) has proven effective in enhancing the reasoning capabilities of large language models (LLMs). Verification plays a key role in TTS, simultaneously influencing (1) reasoning performance and (2) compute efficiency, due to the quality and computational cost of verification. In this work, we challenge the conventional paradigms of verification, and make the first attempt toward systematically investigating the impact of verification granularity—that is, how frequently the verifier is invoked during generation, beyond verifying only the final output or individual generation steps. To this end, we introduce Variable Granularity Search (VG-Search), a unified algorithm that generalizes beam search and Best-of-N sampling via a tunable granularity parameter $g$. Extensive experiments with VG-Search under varying compute budgets, generator-verifier configurations, and task attributes reveal that dynamically selecting $g$ can improve the compute efficiency and scaling behavior. Building on these findings, we propose adaptive VG-Search strategies that achieve accuracy gains of up to 3.1\% over Beam Search and 3.6\% over Best-of-N, while reducing FLOPs by over 52\%. We will open-source the code to support future research.
Hao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo, Masato Motomura, Hongxiang Fan
NeurIPS3
2025 Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
abstract
This paper proposes a novel matrix quantization method, Binary Quadratic Quan- tization (BQQ). In contrast to conventional first-order quantization approaches— such as uniform quantization and binary coding quantization—that approximate real-valued matrices via linear combinations of binary bases, BQQ leverages the expressive power of binary quadratic expressions while maintaining an extremely compact data format. We validate our approach with two experiments: a matrix compression benchmark and post-training quantization (PTQ) on pretrained Vision Transformer-based models. Experimental results demonstrate that BQQ consistently achieves a superior trade-off between memory efficiency and reconstruction error than conventional methods for compressing diverse matrix data. It also delivers strong PTQ performance, even though we neither target state-of-the-art PTQ accuracy under tight memory constraints nor rely on PTQ-specific binary matrix optimization. For example, our proposed method outperforms the state-of- the-art PTQ method by up to 2.2% and 59.1% on the ImageNet dataset under the calibration-based and data-free scenarios, respectively, with quantization equivalent to 2 bits. These findings highlight the surprising effectiveness of binary quadratic expressions for efficient matrix approximation and neural network compression.
Kyo Kuroki, Yasuyuki Okoshi, Thiem Van Chu, Kazushi Kawamura, Masato Motomura
NeurIPS2
2022 Multicoated Supermasks Enhance Hidden Networks
abstract
Hidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing backpropagated information. To address this we propose a method that extends Hidden Networks by training an overlay of multiple hierarchical supermasks{—}a multicoated supermask. This method shows that using multiple supermasks for a single task achieves higher accuracy without additional training cost. Experiments on CIFAR-10 and ImageNet show that Multicoated Supermasks enhance the tradeoff between accuracy and model size. A ResNet-101 using a 7-coated supermask outperforms its Hidden Networks counterpart by 4%, matching the accuracy of a dense ResNet-50 while being an order of magnitude smaller.
Yasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando, Kazushi Kawamura, Thiem Van Chu, Masato Motomura, Jaehoon Yu
ICML1