EDBT 2026 Demo / reviewers in the wild / expert
Tianyi Zhang 0011
dblp:17/322-11
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2025
0009-0007-1753-9933ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 82% Learning paradigms · 8% Generative modeling · 6% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 50% Knowledge graphs · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 74% GPUs and heterogeneous computing · 26% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
3.4 | 4 | 2025 | 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11) · NeurIPS 2025 Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation · ICML 2025 LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
2.4 | 3 | 2025 | 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11) · NeurIPS 2025 NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization › transformer quantization
large language model quantization |
0.9 | 1 | 2025 | LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid · ICLR 2025 |
Machine learning › Generative modeling
lossless compression |
0.9 | 1 | 2025 | 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11) · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.9 | 1 | 2025 | LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation · ICML 2025 |
Machine learning › Efficient and distributed learning
attention computation |
0.8 | 1 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.8 | 1 | 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization |
0.8 | 1 | 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Data mining › structured data mining
graph mining |
0.8 | 1 | 2024 | Learning Scalable Structural Representations for Link Prediction with Bloom Signatures · WWW 2024 |
Knowledge graphs
link prediction |
0.8 | 1 | 2024 | Learning Scalable Structural Representations for Link Prediction with Bloom Signatures · WWW 2024 |
Processor architecture and microarchitecture
SIMD |
0.8 | 1 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.6 | 1 | 2022 | Retaining Knowledge for Learning with Dynamic Definition · NeurIPS 2022 |
Machine learning › Learning paradigms
continual learning |
0.6 | 1 | 2022 | Retaining Knowledge for Learning with Dynamic Definition · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › model inference
transformer inference |
0.5 | 2 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
GPUs and heterogeneous computing
GPU kernel |
0.3 | 1 | 2025 | 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11) · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
lookup table · 1.7entropy coding · 1.7custom GPU kernel · 1.7sketching · 0.9quantization · 0.9low-rank adaptation · 0.9loss-error-aware grid · 0.9locality-sensitive hashing · 0.9hessian-based quantization · 0.9message passing · 0.8hashing · 0.8graph neural network · 0.8entropy analysis · 0.8coupled quantization · 0.8bloom signatures · 0.8SIMD lookup · 0.84-bit quantization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware GridabstractLarge language models (LLMs) have shown immense potential across various domains, but their high memory requirements and inference costs remain critical challenges for deployment. Post-training quantization (PTQ) has emerged as a promising technique to reduce memory requirements and decoding latency. However, recent accurate quantization methods often depend on specialized computations or custom data formats to achieve better model quality, which limits their compatibility with popular frameworks, as they require dedicated inference kernels tailored to specific hardware and software platforms, hindering wider adoption. Furthermore, many competitive methods have high resource requirements and computational overhead for quantizing models, making it challenging to scale them to hundreds of billions of parameters. In response to these challenges, we propose LeanQuant (Loss-error-aware network Quantization), a novel quantization method that is accurate, versatile, and scalable. In the existing popular iterative loss-error-based quantization framework, we identify a critical limitation in prior methods: the min-max affine quantization grid fails to preserve model quality due to outliers in inverse Hessian diagonals. To overcome this fundamental issue, we propose learning loss-error-aware grids, instead of using non-adaptive min-max affine grids. Our approach not only produces quantized models that are more accurate but also generalizes to a wider range of quantization types, including affine and non-uniform quantization, enhancing compatibility with more frameworks. Extensive experiments with recent LLMs demonstrate that LeanQuant is highly accurate, comparing favorably against competitive baselines in model quality, and scalable, achieving very accurate quantization of Llama-3.1 405B, one of the largest open-source LLMs to date, using two Quadro RTX 8000-48GB GPUs in 21 hours. Our code is available at https://github.com/LeanModels/LeanQuant. Tianyi Zhang 0011, Anshumali Shrivastava |
ICLR | 1 |
| 2025 | Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM AdaptationabstractAdapting pre-trained large language models (LLMs) is crucial but challenging due to their enormous size. Parameter-efficient fine-tuning (PEFT) techniques typically employ additive adapters applied to frozen model weights. To further reduce memory usage, model weights are often compressed through quantization. However, existing PEFT methods often yield suboptimal model quality because they rely on restrictive assumptions, such as low-rank constraints on adapters to limit the number of trainable parameters. We find that sketching, a popular data compression technique, can serve as an efficient LLM adaptation strategy while avoiding the low-rank assumption. We introduce SketchTune, a compressive adaptation strategy that compresses LLM weights into compact fine-tunable sketches, integrating compression and adaptation into a unified framework. This integration eliminates the need for complex two-path computation in existing PEFT techniques, enabling faster and more memory-efficient training and inference. SketchTune is supported by mathematical insights into matrix classes that are better approximated using sketching rather than low-rank methods. Our extensive evaluations with Llama and Mistral models demonstrate that SketchTune outperforms leading PEFT methods across diverse tasks while using substantially smaller base models and comparable trainable parameters. As a highlight, SketchTune outperforms LoRA, DoRA, and S2FT on commonsense and math benchmarks using 2.6-3.5$\times$ smaller base models and exceeds LoftQ in accuracy by 14.48\% on GSM8K with 7.3$\times$ fewer trainable parameters. Tianyi Zhang 0011, Junda Su, Aditya Desai, Oscar Wu, Zhaozhuo Xu, Anshumali Shrivastava |
ICML | 1 |
| 2025 | IDentity with Locality: An Ideal Hash for Gene Sequence Search
Tianyi Zhang 0011, Aditya Desai, Anshumali Shrivastava |
KDD (1) | 1 |
| 2025 | 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)abstractLarge-scale AI models, such as Large Language Models (LLMs) and Diffusion Models (DMs), have grown rapidly in size, creating significant challenges for efficient deployment on resource-constrained hardware. In this paper, we introduce Dynamic-Length Float (DFloat11), a lossless compression framework that reduces LLM and DM size by 30\% while preserving outputs that are bit-for-bit identical to the original model. DFloat11 is motivated by the low entropy in the BFloat16 weight representation of LLMs, which reveals significant inefficiency in the existing storage format. By applying entropy coding, DFloat11 assigns dynamic-length encodings to weights based on frequency, achieving near information-optimal compression without any loss of precision. To facilitate efficient inference with dynamic-length encodings, we develop a custom GPU kernel for fast online decompression. Our design incorporates the following: (i) compact, hierarchical lookup tables (LUTs) that fit within GPU SRAM for efficient decoding, (ii) a two-phase GPU kernel for coordinating thread read/write positions using lightweight auxiliary variables, and (iii) transformer-block-level decompression to minimize latency. Experiments on Llama 3.3, Qwen 3, Mistral 3, FLUX.1, and others validate our hypothesis that DFloat11 achieves around 30\% model size reduction while preserving bit-for-bit identical outputs. Compared to a potential alternative of offloading parts of an uncompressed model to the CPU to meet memory constraints, DFloat11 achieves 2.3--46.2$\times$ higher throughput in token generation. With a fixed GPU memory budget, DFloat11 enables 5.7--14.9$\times$ longer generation lengths than uncompressed models. Notably, our method enables lossless inference of Llama 3.1 405B, an 810GB model, on a single node equipped with 8$\times$80GB GPUs. Tianyi Zhang 0011, Mohsen Hariri, Shaochen Zhong, Vipin Chaudhary, Yang Sui 0001, Xia Ben Hu, Anshumali Shrivastava |
NeurIPS | 1 |
| 2025 | Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM PretrainingabstractLow-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically project gradients onto a low-rank subspace, reducing the memory cost of storing optimizer states. A key challenge in these methods is selecting suitable subspaces to ensure an effective optimization trajectory. Most existing approaches select the dominant subspace to preserve gradient information, as this intuitively provides the best approximation. However, we find that in practice, the dominant subspace stops changing during pretraining, thereby constraining weight updates to similar subspaces. In this paper, we propose importance sampling for low-rank optimization in LLM pretraining with a provable convergence guarantee, which the dominant subspace approach does not have. Empirically, we demonstrate that our method significantly outperforms previous methods in LLM pretraining tasks. Junze Yin, Guanchu Wang, Zirui Liu 0001, Lin Yang 0011, Tianyi Zhang 0011, Anshumali Shrivastava, Vladimir Braverman |
NeurIPS | 6 |
| 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled QuantizationabstractEfficient deployment of Large Language Models (LLMs) requires batching multiple requests together to improve throughput. As batch size, context length, or model size increases, the size of key and value (KV) cache quickly becomes the main contributor to GPU memory usage and the bottleneck of inference latency and throughput. Quantization has emerged as an effective technique for KV cache compression, but existing methods still fail at very low bit widths. Currently, KV cache quantization is performed per-channel or per-token independently. Our analysis shows that distinct channels of a key/value activation embedding are highly interdependent, and the joint entropy of multiple channels grows at a slower rate than the sum of their marginal entropy, which implies that per-channel independent quantization is sub-optimal. To mitigate this sub-optimality, we propose Coupled Quantization (CQ), which couples multiple key/value channels together for quantization to exploit their interdependence and encode the activations in a more information-efficient manner. Extensive experiments reveal that CQ compares favorably with existing baselines in preserving model quality, and improves inference throughput by 1.4–3.5$\times$ relative to the uncompressed baseline. Furthermore, we demonstrate that CQ can preserve model quality reasonably with KV cache quantized down to 1 bit. Tianyi Zhang 0011, Jonah Yi, Zhaozhuo Xu, Anshumali Shrivastava |
NeurIPS | 1 |
| 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free AttentionabstractLarge Language Model (LLM) inference on Central Processing Units (CPU) is challenging due to the vast quantities of Multiply-Add (MAD) matrix operations in the attention computations. This paper highlights a rare gem in modern CPUs, Single-Instruction-Multiple-Data (SIMD) registers, which allows for ultra-low-latency lookups in a batch. We leverage this unique capability to propose NoMAD-Attention, an efficient attention algorithm that replaces MAD operations with in-register lookups. Through hardware-aware algorithmic designs, NoMAD-Attention achieves the computation of attention scores using repeated fast accesses to SIMD registers. NoMAD-Attention works with pre-trained attention-based LLMs without model finetuning. Extensive empirical evaluations demonstrate that NoMAD-Attention maintains the quality of the original LLMs well and speeds up the 4-bit quantized LLaMA-7B-based model by up to $2 \times$ at 16k context length. Tianyi Zhang 0011, Jonah Yi, Zhaozhuo Xu, Anshumali Shrivastava |
NeurIPS | 1 |
| 2024 | Learning Scalable Structural Representations for Link Prediction with Bloom SignaturesabstractGraph neural networks (GNNs) have shown great potential in learning on graphs, but they are known to perform sub-optimally on link prediction tasks. Existing GNNs are primarily designed to learn node-wise representations and usually fail to capture pairwise relations between target nodes, which proves to be crucial for link prediction. Recent works resort to learning more expressive edge-wise representations by enhancing vanilla GNNs with structural features such as labeling tricks and link prediction heuristics, but they suffer from high computational overhead and limited scalability. To tackle this issue, we propose to learn structural link representations by augmenting the message-passing framework of GNNs with Bloom signatures. Bloom signatures are hashing-based compact encodings of node neighborhoods, which can be efficiently merged to recover various types of edge-wise structural features. We further show that any type of neighborhood overlap-based heuristic can be estimated by a neural network that takes Bloom signatures as input. GNNs with Bloom signatures are provably more expressive than vanilla GNNs and also more scalable than existing edge-wise models. Experimental results on five standard link prediction benchmarks show that our proposed model achieves comparable or better performance than existing edge-wise GNN models while being 3-200x faster and more memory-efficient for online inference. Source code is available at https://github.com/tonyzhang617/BloomSigLP. Tianyi Zhang 0011, Haoteng Yin, Rongzhe Wei, Pan Li 0005, Anshumali Shrivastava |
WWW | 1 |
| 2023 | Graph Self-supervised Learning via Proximity Distribution MinimizationabstractSelf-supervised learning (SSL) for graphs is an essential problem since graph data are ubiquitous and labeling can be costly. We argue that existing SSL approaches for graphs have two limitations. First, they rely on corruption techniques such as node attribute perturbation and edge dropping to generate graph views for contrastive learning. These unnatural corruption techniques require extensive tuning efforts and provide marginal improvements. Second, the current approaches require the computation of multiple graph views, which is memory and computationally inefficient. These shortcomings of graph SSL call for a corruption-free single-view learning approach, but the strawman approach of using neighboring nodes as positive examples suffers two problems: it ignores the strength of connections between nodes implied by the graph structure on a macro level, and cannot deal with the high noise in real-world graphs. We propose Proximity Divergence Minimization (PDM), a corruption-free single-view graph SSL approach that overcomes these problems by leveraging node proximity to measure connection strength and denoise the graph structure. Through extensive experiments, we show that PDM achieves up to 4.55% absolute improvement in ROC-AUC on graph SSL tasks over state-of-the-art approaches while being more memory efficient. Moreover, PDM even outperforms supervised training on node classification tasks of ogbn-proteins dataset. Our code is publicly available. Tianyi Zhang 0011, Zhenwei Dai, Zhaozhuo Xu, Anshumali Shrivastava |
UAI | 1 |
| 2022 | Retaining Knowledge for Learning with Dynamic DefinitionabstractMachine learning models are often deployed in settings where they must be constantly updated in response to the changes in class definitions while retaining high accuracy on previously learned definitions. A classical use case is fraud detection, where new fraud schemes come one after another. While such an update can be accomplished by re-training on the complete data, the process is inefficient and prevents real-time and on-device learning. On the other hand, efficient methods that incrementally learn from new data often result in the forgetting of previously-learned knowledge. We define this problem as Learning with Dynamic Definition (LDD) and demonstrate that popular models, such as the Vision Transformer and Roberta, exhibit substantial forgetting of past definitions. We present the first practical and provable solution to LDD. Our proposal is a hash-based sparsity model \textit{RIDDLE} that solves evolving definitions by associating samples only to relevant parameters. We prove that our model is a universal function approximator and theoretically bounds the knowledge lost during the update process. On practical tasks with evolving class definition in vision and natural language processing, \textit{RIDDLE} outperforms baselines by up to 30\% on the original dataset while providing competitive accuracy on the update dataset. Zichang Liu, Benjamin Coleman, Tianyi Zhang 0011, Anshumali Shrivastava |
NeurIPS | 3 |