EDBT 2026 Demo / reviewers in the wild / expert
Yuma Ichikawa
dblp:334/5344
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0004-4216-7017ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 63% Optimization for machine learning · 26% Language models and text generation · 11% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Quantum computing and quantum information · 50% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference efficiency |
1.0 | 1 | 2026 | PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
1.0 | 1 | 2026 | PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language Generation · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression
large language model compression |
0.9 | 1 | 2025 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization · NeurIPS 2025 |
Mathematical optimization
combinatorial optimization |
0.9 | 1 | 2025 | Optimization by Parallel Quasi-Quantum Annealing with Gradient-Based Sampling · ICLR 2025 |
Quantum computing and quantum information › quantum computational models
quantum annealing |
0.9 | 1 | 2025 | Optimization by Parallel Quasi-Quantum Annealing with Gradient-Based Sampling · ICLR 2025 |
Machine learning › Optimization for machine learning
combinatorial optimization |
0.8 | 1 | 2024 | Controlling Continuous Relaxation for Combinatorial Optimization · NeurIPS 2024 |
Machine learning › Optimization for machine learning › convex relaxation
continuous relaxation |
0.8 | 1 | 2024 | Controlling Continuous Relaxation for Combinatorial Optimization · NeurIPS 2024 |
GPUs and heterogeneous computing
GPU computing |
0.3 | 1 | 2025 | Optimization by Parallel Quasi-Quantum Annealing with Gradient-Based Sampling · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
parallel tempering · 1.7gradient-based sampling · 1.7continuous relaxation · 1.7latent variable model · 1.0autoregressive modeling · 1.0quantization error propagation · 0.9layer-wise post-training quantization · 0.9penalty term · 0.8continuous relaxation annealing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PHOTON: Hierarchical Autoregressive Modeling for Lightspeed and Memory-Efficient Language GenerationabstractTransformers operate as horizontal token-bytoken scanners; at each generation step, attending to an ever-growing sequence of tokenlevel states.This access pattern increases prefill latency and makes long-context decoding more memory-bound, as KV-cache reads and writes dominate inference time over arithmetic operations.We propose Parallel Hierarchical Operation for TOp-down Networks (PHOTON), a hierarchical autoregressive model that replaces horizontal scanning with vertical, multi-resolution context scanning.PHOTON maintains a hierarchy of latent streams: a bottom-up encoder compresses tokens into low-rate contextual states, while lightweight top-down decoders reconstruct fine-grained token representations in parallel.We further introduce recursive generation that updates only the coarsest latent stream and eliminates bottom-up re-encoding.Experimental results show that PHOTON is superior to competitive Transformer-based language models regarding the throughput-quality tradeoff, providing advantages in long-context and multi-query tasks.In particular, this reduces decode-time KV-cache traffic, yielding up to 10 3 × higher throughput per unit memory. Yuma Ichikawa, Naoya Takagi, Takumi Nakagawa, Yuzi Kanazawa, Akira Sakai |
ACL (1) | 1 |
| 2025 | Optimization by Parallel Quasi-Quantum Annealing with Gradient-Based SamplingabstractLearning-based methods have gained attention as general-purpose solvers due to their ability to automatically learn problem-specific heuristics, reducing the need for manually crafted heuristics. However, these methods often face scalability challenges. To address these issues, the improved Sampling algorithm for Combinatorial Optimization (iSCO), using discrete Langevin dynamics, has been proposed, demonstrating better performance than several learning-based solvers. This study proposes a different approach that integrates gradient-based update through continuous relaxation, combined with Quasi-Quantum Annealing (QQA). QQA smoothly transitions the objective function, starting from a simple convex function, minimized at half-integral values, to the original objective function, where the relaxed variables are minimized only in the discrete space. Furthermore, we incorporate parallel run communication leveraging GPUs to enhance exploration capabilities and accelerate convergence. Numerical experiments demonstrate that our method is a competitive general-purpose solver, achieving performance comparable to iSCO and learning-based solvers across various benchmark problems. Notably, our method exhibits superior speed-quality trade-offs for large-scale instances compared to iSCO, learning-based solvers, commercial solvers, and specialized algorithms. Yuma Ichikawa, Yamato Arai |
ICLR | 1 |
| 2025 | Quantization Error Propagation: Revisiting Layer-Wise Post-Training QuantizationabstractLayer-wise PTQ is a promising technique for compressing large language models (LLMs), due to its simplicity and effectiveness without requiring retraining. However, recent progress in this area is saturating, underscoring the need to revisit its core limitations and explore further improvements. We address this challenge by identifying a key limitation of existing layer-wise PTQ methods: the growth of quantization errors across layers significantly degrades performance, particularly in low-bit regimes. To address this fundamental issue, we propose Quantization Error Propagation (QEP), a general, lightweight, and scalable framework that enhances layer-wise PTQ by explicitly propagating quantization errors and compensating for accumulated errors. QEP also offers a tunable propagation mechanism that prevents overfitting and controls computational overhead, enabling the framework to adapt to various architectures and resource budgets. Extensive experiments on several LLMs demonstrate that QEP-enhanced layer-wise PTQ achieves substantially higher accuracy than existing methods. Notably, the gains are most pronounced in the extremely low-bit quantization regime. Yamato Arai, Yuma Ichikawa |
NeurIPS | 2 |
| 2024 | Learning Dynamics in Linear VAE: Posterior Collapse Threshold, Superfluous Latent Space Pitfalls, and Speedup with KL AnnealingabstractVariational autoencoders (VAEs) face a notorious problem wherein the variational posterior often aligns closely with the prior, a phenomenon known as posterior collapse, which hinders the quality of representation learning. To mitigate this problem, an adjustable hyperparameter $\beta$ and a strategy for annealing this parameter, called KL annealing, are proposed. This study presents a theoretical analysis of the learning dynamics in a minimal VAE. It is rigorously proved that the dynamics converge to a deterministic process within the limit of large input dimensions, thereby enabling a detailed dynamical analysis of the generalization error. Furthermore, the analysis shows that the VAE initially learns entangled representations and gradually acquires disentangled representations. A fixed-point analysis of the deterministic process reveals that when $\beta$ exceeds a certain threshold, posterior collapse becomes inevitable regardless of the learning period. Additionally, the superfluous latent variables for the data-generative factors lead to overfitting of the background noise; this adversely affects both generalization and learning convergence. The analysis further unveiled that appropriately tuned KL annealing can accelerate convergence. Yuma Ichikawa, Koji Hukushima |
AISTATS | 1 |
| 2024 | Adaptive Flip Graph Algorithm for Matrix MultiplicationabstractThis study proposes the “adaptive flip graph algorithm”, which combines adaptive searches with the flip graph algorithm for finding fast and efficient methods for matrix multiplication. The adaptive flip graph algorithm addresses the inherent limitations of exploration and inefficient search encountered in the original flip graph algorithm, particularly when dealing with large matrix multiplication. For the limitation of exploration, the proposed algorithm adaptively transitions over the flip graph, introducing a flexibility that does not strictly reduce the number of multiplications. Concerning the issue of inefficient search in large instances, the proposed algorithm adaptively constraints the search range instead of relying on a completely random search, facilitating more effective exploration. In particular, a formal proof is provided that the introduction of plus transitions in the proposed algorithm ensures the connectivity of any node in the flip graph, which represents a method of matrix multiplication. Numerical experimental results demonstrate the effectiveness of the adaptive flip graph algorithm, which involves applying matrices calculated in characteristic 2. This algorithm reduces the number of multiplications for a 4 × 5 matrix multiplied by a 5 × 5 matrix from 76 to 73 and that for a 5 × 5 matrix multiplied by another 5 × 5 matrix from 95 to 94. Yamato Arai, Yuma Ichikawa, Koji Hukushima |
ISSAC | 2 |
| 2024 | Controlling Continuous Relaxation for Combinatorial OptimizationabstractUnsupervised learning (UL)-based solvers for combinatorial optimization (CO) train a neural network that generates a soft solution by directly optimizing the CO objective using a continuous relaxation strategy. These solvers offer several advantages over traditional methods and other learning-based methods, particularly for large-scale CO problems. However, UL-based solvers face two practical issues: (I) an optimization issue, where UL-based solvers are easily trapped at local optima, and (II) a rounding issue, where UL-based solvers require artificial post-learning rounding from the continuous space back to the original discrete space, undermining the robustness of the results. This study proposes a Continuous Relaxation Annealing (CRA) strategy, an effective rounding-free learning method for UL-based solvers. CRA introduces a penalty term that dynamically shifts from prioritizing continuous solutions, effectively smoothing the non-convexity of the objective function, to enforcing discreteness, eliminating artificial rounding. Experimental results demonstrate that CRA significantly enhances the performance of UL-based solvers, outperforming existing UL-based solvers and greedy algorithms in complex CO problems. Additionally, CRA effectively eliminates artificial rounding and accelerates the learning process. Yuma Ichikawa |
NeurIPS | 1 |