EDBT 2026 Demo / reviewers in the wild / expert
Preyas Shah
dblp:77/10335
· DBLP profile ↗
4ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Electronic design automation · 50% Hardware accelerators and domain-specific architectures · 44% GPUs and heterogeneous computing · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › spatial architecture
dataflow accelerator |
0.6 | 1 | 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022 |
Electronic design automation › high-level synthesis
accelerator generation |
0.4 | 1 | 2020 | DSAGEN: Synthesizing Programmable Spatial Accelerators · ISCA 2020 |
Electronic design automation › design space exploration
automated design space exploration |
0.4 | 1 | 2020 | DSAGEN: Synthesizing Programmable Spatial Accelerators · ISCA 2020 |
Electronic design automation
hardware/software co-design |
0.4 | 1 | 2020 | DSAGEN: Synthesizing Programmable Spatial Accelerators · ISCA 2020 |
Compilers and program optimization › code generation
compiler mapping |
0.1 | 1 | 2020 | DSAGEN: Synthesizing Programmable Spatial Accelerators · ISCA 2020 |
Methods — techniques the papers use, named apart from their topics
modular transformations · 0.9hardware primitive composition · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial productabstractIn this paper we introduce the Mozart Processor, which implements a new processing paradigm called Reuse Exposed Dataflow (RED). RED is a counterpart to existing execution models of Von-Neumann, SIMT, Dataflow, and FPGA. Dataflow and data reuse are the fundamental architecture primitives in RED, implemented with mechanisms for inter-worker communication and synchronization. The paper defines the processor architecture, the details of the microarchitecture, chip implementation, software stack development, and performance results. The architecture's goal is to achieve near-CPU like flexibility while having ASIC-like efficiency for a large-class of data-intensive workloads. An additional goal was software maturity --- have large coverage of applications immediately, avoiding the need for a long-drawn hand-tuning software development phase. The architecture was defined with this software-maturity/compiler friendliness in mind. In short, the goal was to do to GPUs, what GPUs did to CPUs --- i.e. be a better solution for a large range of workloads, while preserving flexibility and programmability. The chip was implemented with HBM and PCIe interfaces and taken to production on a 16nm TSMC FFC process. For ML inference tasks with batch-size=4, Mozart is integer factors better than state-of-the-art GPUs even while being nearly 2 technology nodes behind. We conclude with a set of lessons learned, the unique challenges of a clean-slate architecture in a commercial setting, and pointers for uncovered research problems. Karthikeyan Sankaralingam, Tony Nowatzki, Vinay Gangadhar, Preyas Shah, William Galliher, Ziliang Guo, Jitu Khare, Deepak Vijay, Poly Palamuttam, Maghawan Punde, Alex Tan, Vijayraghavan Thiruvengadam, Rongyi Wang, Shunmiao Xu |
ISCA | 4 |
| 2021 | Mozart: Designing for Software Maturity and the Next Paradigm for Chip ArchitecturesabstractWhere does AI hardware/software stand today? 1. The computational diversity needed to support AI is increasing2. The software user experience expectations is increasing3. GPU software maturity* is unrivalled in completeness and hence allows near complete dominance among AI industry deployment and researchers.4. This support for model diversity is fuelling these trends and increasing GPU adoption!* NVIDIA DL stack - cuDNN, TensorRT, etc. Karthikeyan Sankaralingam, Tony Nowatzki, Greg Wright, Poly Palamuttam, Jitu Khare, Vinay Gangadhar, Preyas Shah |
HCS | 7 |
| 2020 | DSAGEN: Synthesizing Programmable Spatial AcceleratorsabstractDomain-specific hardware accelerators can provide orders of magnitude speedup and energy efficiency over general purpose processors. However, they require extensive manual effort in hardware design and software stack development. Automated ASIC generation (eg. HLS) can be insufficient, because the hardware becomes inflexible. An ideal accelerator generation framework would be automatable, enable deep specialization to the domain, and maintain a uniform programming interface. Our insight is that many prior accelerator architectures can be approximated by composing a small number of hardware primitives, specifically those from spatial architectures. With careful design, a compiler can understand how to use available primitives, with modular and composable transformations, to take advantage of the features of a given program. This suggests a paradigm where accelerators can be generated by searching within such a rich accelerator design space, guided by the affinity of input programs for hardware primitives and their interactions. We use this approach to develop the DSAGEN framework, which automates the hardware/software co-design process for reconfigurable accelerators. For several existing accelerators, our evaluation demonstrates that the compiler can achieve 89% of the performance of manually tuned versions. For automated design space exploration, we target multiple sets of workloads which prior accelerators are design for; the generated hardware has mean 1.3× perf2/mm2over prior programmable accelerators. Jian Weng 0002, Sihao Liu, Vidushi Dadu, Zhengrong Wang, Preyas Shah, Tony Nowatzki |
ISCA | 5 |
| 2011 | Capacitive skin sensors for robot impact monitoringabstractA new generation of robots is being designed for human occupied workspaces where safety is of great concern. This research demonstrates the use of a capacitive skin sensor for collision detection. Tests demonstrate that the sensor reduces impact forces and can detect and characterize collision events, providing information that may be used in the future for force reduction behaviors. Various parameters that affect collision severity, including interface friction, interface stiffness, end tip velocity and joint stiffness irrespective of controller bandwidth are also explored using the sensor to provide information about the contact force at the site of impact. Joint stiffness is made independent of controller bandwidth limitations using passive torsional springs of various stiffnesses. Results indicate a positive correlation between peak impact force and joint stiffness, skin friction and interface stiffness, with implications for future skin and robot link designs and post-collision behaviors. Samson Phan, Zhan Fan Quek, Preyas Shah, Dongjun Shin, Oussama Khatib, Mark R. Cutkosky |
IROS | 3 |