VLDB 2026 Research / reviewers in the wild / expert
Ranajoy Sadhukhan
dblp:277/5965
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 50% Language models and text generation · 32% Deep learning architectures and training · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › memory mechanism
associative memory networks |
0.9 | 1 | 2025 | Memory Mosaics · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | MagicPIG: LSH Sampling for Efficient LLM Generation · ICLR 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization |
0.9 | 1 | 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding · ICLR 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.9 | 1 | 2025 | MagicPIG: LSH Sampling for Efficient LLM Generation · ICLR 2025 |
Natural language and speech › Language models and text generation
language modeling |
0.9 | 1 | 2025 | Memory Mosaics · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.9 | 1 | 2025 | Kinetics: Rethinking Test-Time Scaling Law · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.9 | 1 | 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding · ICLR 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Kinetics: Rethinking Test-Time Scaling Law · NeurIPS 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2025 | Memory Mosaics · ICLR 2025 |
Machine learning › Efficient and distributed learning
KV cache |
0.3 | 1 | 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding · ICLR 2025 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
long chain-of-thought reasoning |
0.3 | 1 | 2025 | Kinetics: Rethinking Test-Time Scaling Law · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding · ICLR 2025 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
heterogeneous CPU-GPU inference |
0.3 | 1 | 2025 | MagicPIG: LSH Sampling for Efficient LLM Generation · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
topk attention approximation · 1.7sampling · 1.7locality-sensitive hashing · 1.7speculative decoding · 0.9sparse KV cache · 0.9scaling law analysis · 0.9predictive disentanglement · 0.9drafting strategy · 0.9best-of-n sampling · 0.9associative memory · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MagicPIG: LSH Sampling for Efficient LLM GenerationabstractLarge language models (LLMs) with long context windows have gained significant attention. However, the KV cache, stored to avoid re-computation, becomes a bottleneck. Various dynamic sparse or TopK-based attention approximation methods have been proposed to leverage the common insight that attention is sparse. In this paper, we first show that TopK attention itself suffers from quality degradation in certain downstream tasks because attention is not always as sparse as expected. Rather than selecting the keys and values with the highest attention scores, sampling with theoretical guarantees can provide a better estimation for attention output. To make the sampling-based approximation practical in LLM generation, we propose MagicPIG, a heterogeneous system based on Locality Sensitive Hashing (LSH). MagicPIG significantly reduces the workload of attention computation while preserving high accuracy for diverse tasks. MagicPIG stores the LSH hash tables and runs the attention computation on the CPU, which allows it to serve longer contexts and larger batch sizes with high approximation accuracy. MagicPIG can improve decoding throughput by up to $5\times$ across various GPU hardware and achieve 54ms decoding latency on a single RTX 4090 for Llama-3.1-8B-Instruct model with a context of 96k tokens. Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye 0001, Niklas Nolte, Yuandong Tian, Matthijs Douze, Léon Bottou, Beidi Chen |
ICLR | 2 |
| 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingabstractLarge Language Models (LLMs) have become more prevalent in long-context applications such as interactive chatbots, document analysis, and agent workflows, but it is challenging to serve long-context requests with low latency and high throughput. Speculative decoding (SD) is a widely used technique to reduce latency losslessly, but the conventional wisdom suggests that its efficacy is limited to small batch sizes. In MagicDec, we show that surprisingly SD can achieve speedup even for a high throughput inference regime for moderate to long sequences. More interestingly, an intelligent drafting strategy can achieve better speedup with increasing batch size based on our rigorous analysis. MagicDec first identifies the bottleneck shifts with increasing batch size and sequence length, and uses these insights to deploy SD more effectively for high throughput inference. We leverage draft model with sparse KV cache to address the KV bottleneck, which scales with both sequence length and batch size. Additionally, we propose a theoretical model to select the optimal drafting strategy for maximum speedup. Our work highlights the broad applicability of speculative decoding in long-context serving, as it can enhance throughput and reduce latency without compromising accuracy. For moderate to long sequences, we demonstrate up to 2.51x speedup for LLaMA-3.1-8B when serving batch sizes ranging from 32 to 256 on various types of hardware and tasks. Ranajoy Sadhukhan, Zhuoming Chen, Vashisth Tiwari, Ruihang Lai, Jinyuan Shi, Ian En-Hsu Yen, Avner May, Tianqi Chen 0001, Beidi Chen |
ICLR | 1 |
| 2025 | Memory MosaicsabstractMemory Mosaics are networks of associative memories working in concert to achieve a prediction task of interest. Like transformers, memory mosaics possess compositional capabilities and in-context learning capabilities. Unlike transformers, memory mosaics achieve these capabilities in comparatively transparent way (“predictive disentanglement”). We illustrate these capabilities on a toy example and also show that memory mosaics perform as well or better than transformers on medium-scale language modeling tasks. Niklas Nolte, Ranajoy Sadhukhan, Beidi Chen, Léon Bottou |
ICLR | 3 |
| 2025 | Kinetics: Rethinking Test-Time Scaling LawabstractWe rethink test-time scaling laws from a practical efficiency perspective, revealing that the effectiveness of smaller models is significantly overestimated. Prior work, grounded in compute-optimality, overlooks critical memory access bottlenecks introduced by inference-time strategies (e.g., Best-of-N, long CoTs). Our holistic analysis, spanning models from 0.6B to 32B parameters, reveals a new Kinetics Scaling Law that better guides resource allocation by incorporating both computation and memory access costs. The Kinetics Scaling Law suggests that test-time compute is more effective when used on models above a threshold (14B) than on smaller ones. A key reason is that in test-time scaling, attention—rather than parameter count—emerges as the dominant cost factor. Motivated by this, we propose a new scaling paradigm centered on sparse attention, which lowers per-token cost and enables longer generations and more parallel samples within the same resource budget. Empirically, we show that sparse attention models consistently outperform dense counterparts, achieving over 60-point gains in low-cost regimes and over 5-point gains in high-cost regimes for problem-solving accuracy on AIME and LiveCodeBench. These results suggest that sparse attention is essential for realizing the full potential of test-time scaling because, unlike training where parameter scaling saturates, test-time accuracy continues to improve through increased generation. Ranajoy Sadhukhan, Zhuoming Chen, Haizhong Zheng, Beidi Chen |
NeurIPS | 1 |
| 2022 | Taxonomy Driven Learning Of Semantic Hierarchy Of ClassesabstractStandard pre-trained convolutional neural networks are deployed on different task-specific limited class applications. These applications require classifying images of a much smaller subset of classes than that of the original large domain dataset on which the network is pre-trained. Therefore, a computationally inefficient and over-represented network is obtained. Hierarchically Self Decomposing CNN (HSD-CNN) addresses this issue by dissecting the network into sub-networks in an automated hierarchical fashion such that each sub-network is useful for classifying images of closely related classes. However, visual similarities are not always well-aligned with the semantic understanding of humans. In this paper, we propose a method that aids the pre-trained network to learn the hierarchy of classes derived from standard taxonomy, WordNet and, produce sub-networks corresponding to semantically meaningful classes upon decomposition. Experimental results show that the cluster of classes obtained for each sub-network is semantically closer according to WordNet hierarchy without degradation in overall accuracy. Ranajoy Sadhukhan, Ankita Chatterjee, Jayanta Mukhopadhyay, Amit Patra |
ICIP | 1 |
| 2020 | Knowledge Distillation Inspired Fine-Tuning Of Tucker Decomposed CNNS and Adversarial Robustness AnalysisabstractThe recent works in tensor decomposition of convolutional neural networks have paid little attention to finetuning the decomposed models more effectively. We propose to improve the accuracy as well as the adversarial robustness of decomposed networks over existing noniterative methods by distilling knowledge from the computationally intensive undecomposed (teacher) model to the decomposed (student) model. Through a series of experiments, we demonstrate the effectiveness of knowledge distillation with different loss functions and compare it to the existing fine-tuning strategy of minimizing crossentropy loss with ground truth labels. Finally, we conclude that the student networks obtained by the proposed approach are superior not only in terms of accuracy but also adversarial robustness, which is often compromised in the existing methods. Ranajoy Sadhukhan, Avinab Saha, Jayanta Mukhopadhyay, Amit Patra |
ICIP | 1 |