VLDB 2026 Research / reviewers in the wild / expert
Mohammad Nouri
dblp:163/8333
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Retrieval-Augmented GenerationabstractAn evolving solution to address hallucination and enhance accuracy in large language models (LLMs) is Retrieval-Augmented Generation (RAG), which involves augmenting LLMs with information retrieved from an external knowledge source, such as the web. This paper profiles several RAG execution pipelines and demystifies the complex interplay between their retrieval and generation phases. We demonstrate that while exact retrieval schemes are expensive, they can reduce inference time compared to approximate retrieval variants because an exact retrieval model can send a smaller but more accurate list of documents to the generative model while maintaining the same end-to-end accuracy. This observation motivates the acceleration of the exact nearest neighbor search for RAG. Derrick Quinn, Mohammad Nouri, John Salihu, Alireza Salemi, Sukhan Lee 0002, Hamed Zamani, Mohammad Alian |
ASPLOS (1) | 2 |
| 2025 | A Reconfigurable and Accurate Circuit-Level Substrate for DRAM Design and Analysis
S. M. Mojahidul Ahsan, Mohammad Nouri, Ramesh Reddy Ganapam, Mohammad Alian, Tamzidul Hoque |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | A Flexible and Accurate Circuit-Level Substrate for Future DRAM Design and AnalysisabstractWe present a reconfigurable, circuit-level substrate for DRAM design and analysis, implemented in SPICE. Existing DRAM substrates exhibit critical inaccuracies, particularly in access transistor modeling and sense amplifier sizing. This leads to inaccurate cell retention time affecting reliable power and performance projections. Our framework incorporates an end-to-end datapath substrate from DRAM cells to chip I/O and a custom Verilog-A model for DRAM access transistor with validated$I_{ON} / I_{OFF}$characteristics. The model achieves less than 1% error in cell retention time and aligns with JEDEC timing standards. In addition, we developed a Python-based design space exploration (DSE) tool to enable rapid customization across technology nodes and access device models. S. M. Mojahidul Ahsan, Mohammad Nouri, Ramesh Reddy Ganapam, Mohammad Alian, Tamzidul Hoque |
ISPASS | 2 |
| 2024 | SmartDIMM: In-Memory Acceleration of Upper Layer ProtocolsabstractThere has been significant focus on offloading upperlayer network protocols (ULPs) to accelerators located on CPUs and SmartNICs. However, restricting accelerator placement to these locations limits both the variety of ULPs that can be accelerated and the overall performance. In particular, it overlooks the opportunity to accelerate ULPs running atop a stateful transport protocol in the face of high cache contention. That is, at high network rates, the frequent DRAM accesses and SmartNIC-CPU synchronizations outweigh the benefits of hardware acceleration. This work introduces SmartDIMM, which unlocks the opportunity for accelerating ULPs running atop stateful transport protocols that primarily operate on data stored in DRAM. We prototyped SmartDIMM using Samsung's AxDIMM and implemented endto-end offloading of (de/en)cryption and (de)compression– two ULPs widely employed in datacenters. We then compared the performance of SmartDIMM with accelerator placements on the CPU, SmartNIC, and PCIe cards. Our results demonstrate that ULP offloading on SmartDIMM outperforms CPU, SmartNIC and PCIe-based offload configurations. In comparison to a server executing (de/en)cryption and (de)compression on the CPU, SmartDIMM achieves 21.0% to 10.28 × higher requests per second and 36.3% to 88.9% lower memory bandwidth utilization. Amin Mamandipoor, Mohammad Nouri, Mohammad Alian |
HPCA | 3 |