VLDB 2026 Research / reviewers in the wild / expert
Marta Andronic
dblp:356/2624
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0004-6153-0491ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bit-Serial Acceleration of LLM Inference With Mixture-of-Datatype QuantizationabstractLarge language models (LLMs) have achieved significant breakthroughs on machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their wide deployment. In this paper, we propose BitMoD, an algorithm-hardware co-design solution for efficient LLM deployment. On the algorithm side, BitMoD introduces “fine-grained data type adaptation”, which uses a different data type to quantize a group (e.g., 128) of weights and key-value-cache (KV-cache). Through the careful design of these data types, BitMoD is able to quantize LLM weights and KV-cache to sub-4-bit precision while maintaining high accuracy. On the hardware side, BitMoD employs the bit-serial computing paradigm to easily support multiple numerical precisions and data types, thus providing a flexible trade-off between model accuracy and hardware efficiency. Furthermore, we design low-cost hardware components to effectively handle online KV-cache quantization and per-group partial sum dequantization. Our evaluation on a diverse set of LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization methods on both discriminative and generative tasks. Combining the superior model performance with an efficient accelerator design, BitMoD surpasses the state-of-the-art LLM accelerator in terms of both hardware performance and energy efficiency. Yuzong Chen 0001, Chi-Chih Chang, Xilai Dai, Ahmed F. AbouElhamayed, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah |
IEEE Trans. Computers | 5 |
| 2025 | NeuraLUT-Assemble: Hardware-Aware Assembling of Sub-Neural Networks for Efficient LUT InferenceabstractEfficient neural networks (NNs) leveraging lookup tables (LUTs) have demonstrated significant potential for emerging AI applications, particularly when deployed on field-programmable gate arrays (FPGAs) for edge computing. These architectures promise ultra-low latency and reduced resource utilization, broadening neural network adoption in fields such as particle physics. However, existing LUT-based designs suffer from accuracy degradation due to the large fan-in required by neurons being limited by the exponential scaling of LUT resources with input width. In practice, in prior work this tension has resulted in the reliance on extremely sparse models. We present NeuraLUT-Assemble, a novel framework that addresses these limitations by combining mixed-precision techniques with the assembly of larger neurons from smaller units, thereby increasing connectivity while keeping the number of inputs of any given LUT manageable. Additionally, we intro-duce skip-connections across entire LUT structures to improve gradient flow. NeuraLUT-Assemble closes the accuracy gap between LUT-based methods and (fully-connected) MLP-based models, achieving competitive accuracy on tasks such as network intrusion detection, digit classification, and jet classification, demonstrating up to 8.42x reduction in the area-delay product compared to the state-of-the-art at the time of the publication. Marta Andronic, George A. Constantinides |
FCCM | 1 |
| 2025 | Ph.D. Project Hardware-Aware Neural NetworksabstractThe challenges of deploying overparameterized neural networks (NNs) on resource-constrained hardware necessitate innovative approaches that transcend traditional precision and overparameterization paradigms. Our research explores NN design with a deep awareness of hardware limitations, rethinking traditional approaches to radically reduce inference cost on field-programmable gate arrays (FPGAs). We focus on developing hardware-aware NNs that integrate multiple levels of precision and expressivity to enhance performance. Our PolyLUT, NeuraLUT and NeuraLUT-Assemble methodologies leverage the flexibility of Boolean lookup tables (LUTs) to enhance network expressivity and efficiency. By strategically integrating diverse data representations and arithmetic operations, these models redefine NN design, offering a pathway to highly efficient, hardware-optimized networks that do not compromise on accuracy or generalization. Future research will aim to investigate how these models scale up to more complex tasks, even broadening the scope to language models, where LUT-based approaches could offer significant improvements in memory efficiency and inference speed. Moreover, along the practical advancements in efficient AI design, we plan to advance the theoretical understanding of these unconventional topologies. Marta Andronic, George A. Constantinides |
FCCM | 1 |
| 2025 | ReducedLUT: Table Decomposition with "Don't Care" ConditionsabstractLookup tables (LUTs) are frequently used to efficiently store arrays of precomputed values for complex mathematical computations. When used in the context of neural networks, these functions exhibit a lack of recognizable patterns which presents an unusual challenge for conventional logic synthesis techniques. Several approaches are known to break down a single large lookup table into multiple smaller ones that can be recombined. Traditional methods, such as plain tabulation, piecewise linear approximation, and multipartite table methods, often yield inefficient hardware solutions when applied to LUT-based NNs. Oliver Cassidy, Marta Andronic, Samuel Coward, George A. Constantinides |
FPGA | 2 |
| 2025 | Greater than the Sum of its LUTs: Scaling Up LUT-based Neural Networks with AmigoLUTabstractApplications like high-energy physics and cybersecurity require extremely high throughput and low latency neural network (NN) inference. Lookup-table-based NNs address these constraints by implementing NNs as lookup tables (LUTs), achieving inference latency on the order of nanoseconds. Since LUTs are a fundamental FPGA building block, LUT-based NNs efficiently map to FPGAs. LogicNets (and its successors) form one class of LUT-based NNs that target FPGAs, mapping neurons directly to LUTs to meet low latency constraints with minimal resources. However, it is difficult to build larger, more performant LUT-based NNs like LogicNets because LUT usage increases exponentially with respect to neuron fan-in (i.e., number of synapses X synapse bitwidth). A large LUT-based NN quickly runs out of LUTs on an FPGA. Our work AmigoLUT addresses this issue by creating ensembles of smaller LUT-based NNs that scale linearly with respect to the number of models. AmigoLUT improves the scalability of LUT-based NNs, reaching higher throughput with up to an order of magnitude fewer LUTs than the largest LUT-based NNs. Olivia Weng, Marta Andronic, Danial Zuberi, Caleb Geniesse, George A. Constantinides, Nicholas J. Fraser, Javier M. Duarte, Ryan Kastner |
FPGA | 2 |
| 2025 | BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationabstractLarge language models (LLMs) have demonstrated remarkable performance across various machine learning tasks. Yet the substantial memory footprint of LLMs significantly hinders their deployment. In this paper, we improve the accessibility of LLMs through BitMoD1, an algorithm-hardware co-design solution that enables efficient LLM acceleration at low weight precision. On the algorithm side, BitMoD introduces fine-grained data type adaptation that uses a different numerical data type to quantize a group of (e.g., 128) weights. Through the careful design of these new data types, BitMoD is able to quantize LLM weights to very low precision (e.g., 4 bits and 3 bits) while maintaining high accuracy. On the hardware side, BitMoD employs a bitserial processing element to easily support multiple numerical precisions and data types; our hardware design includes two key innovations: First, it employs a unified representation to process different weight data types, thus reducing the hardware cost. Second, it adopts a bit-serial dequantization unit to rescale the per-group partial sum with minimal hardware overhead. Our evaluation on six representative LLMs demonstrates that BitMoD significantly outperforms state-of-the-art LLM quantization and acceleration methods. For discriminative tasks, BitMoD can quantize LLM weights to 4 -bit with1Code is available at: https://github.com/yc2367/BitMoD-HPCA-25 Yuzong Chen 0001, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang 0053, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah |
HPCA | 5 |
| 2025 | PolyLUT: Ultra-Low Latency Polynomial Inference With Hardware-Aware Structured Pruning
Marta Andronic, George A. Constantinides |
IEEE Trans. Computers | 1 |
| 2024 | NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable FunctionsabstractField-Programmable Gate Array (FPGA) accelerators have proven successful in handling latency- and resource-critical deep neural network (DNN) inference tasks. Among the most computationally intensive operations in a neural network (NN) is the dot product between the feature and weight vectors. Thus, some previous FPGA acceleration works have proposed mapping neurons with quantized inputs and outputs directly to lookup tables (LUTs) for hardware implementation. In these works, the boundaries of the neurons coincide with the boundaries of the LUTs. We propose relaxing these boundaries and mapping entire sub-networks to a single LUT. As the sub-networks are absorbed within the LUT, the NN topology and precision within a partition do not affect the size of the lookup tables generated. Therefore, we utilize fully connected layers with floating-point precision inside each partition, which benefit from being universal function approximators, but with rigid sparsity and quantization enforced between partitions, where the NN topology becomes exposed to the circuit topology. Although cheap to implement, this approach can lead to very deep NNs, and so to tackle challenges like vanishing gradients, we also introduce skip connections inside the partitions. The resulting methodology can be seen as training DNNs with a specific FPGA hardware-inspired sparsity pattern that allows them to be mapped to much shallower circuit-level networks, thereby significantly improving latency. We validate our proposed method on a known latency-critical task, jet substructure tagging, and on the classical computer vision task, digit classification using MNIST. Our approach allows for greater function expressivity within the LUTs compared to existing work, leading to up to $4.3 \times$ lower latency NNs for the same accuracy. Marta Andronic, George A. Constantinides |
FPL | 1 |