VLDB 2026 Research / reviewers in the wild / expert
Nupur Sumeet
dblp:290/4072
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2022
0000-0002-8023-6990ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | High-Performance Deployment of Text Detection Model: Compression and Hardware Platform considerationsabstractNetwork compression is often adopted for high throughput implementation on commercial accelerators. We propose a heuristic based approach to obtain compressed networks with a hardware-friendly architecture as an alternative to conventional NAS algorithms that are computationally expensive. The proposed compressed network introduces 142 $\times$ memory-footprint reduction and provide throughput improvement of 5-8 $\times$ on target hardware platforms, while retaining accuracy within 5% of the baseline trained model. We report performance acceleration on CPU, GPU, and FPGAs for a text detection task. Nupur Sumeet, Karan Rawat, Manoj Nambiar 0001 |
ISPASS | 1 |
| 2022 | Performance Model and Profile Guided Design of a High-Performance Session Based Recommendation EngineabstractSession-based recommendation (SBR) systems are widely used in transactional systems to make personalized recommendations to the end-user. In online retail systems, recommendations-based decisions need to be made at a very high rate especially during peak hours. The required computational workload is very high especially when there is a larger number of products involved. Session Based Recommendation (SBR) models incorporate the learning-based product buying pattern from various user interaction sessions and try to recommend the top-K products, the user is likely to purchase. These models comprise several functional layers that widely vary in their compute and data access patterns. To support high recommendation rates, all these layers need a performance optimal implementation, which can be a challenge given the diverse nature of the computations involved. For this reason, one compute platform - whether it is CPU, GPU, or a Field Programmable Gate Array (FPGA) may not be able to provide an optimal implementation for all the layers. In this paper, we describe performance modeling and profile-based design approach to arrive at an optimal implementation, comprising of the hybrid CPU, GPU, and FPGA platforms for NISER - a session-based recommendation model that avoids popularity bias in recommendations. In addition, the design for the CPU-FPGA hybrid platform is implemented for NISER and we observed that experimental results closely follow the results predicted by the performance model for the implemented deployment option. Ashwin Krishnan, Manoj Nambiar 0001, Nupur Sumeet, Sana Iqbal |
ICPE | 3 |
| 2022 | HLS_Profiler: Non-Intrusive Profiling Tool for HLS based ApplicationsabstractThe High-Level Synthesis (HLS) tools aid in simplified and faster design development without familiarity with Hardware Description Language (HDL) and Register Transfer Logic (RTL) design flow. However, it is not straight forward to associate every line of source code to a clock-cycle of synthesized hardware design. On the other hand, the traditional RTL-based design development flow provides the fine-grained performance profile through waveforms. With the same level of visibility in HLS designs, the designers can identify the performance-bottlenecks and obtain the target performance by iteratively fine-tuning the source code. Although, the HLS development tools provide the low-level waveforms, interpreting them in terms of source code variables is a challenging and tedious task. Addressing this gap, we propose an automated profiler tool, HLS\_Profiler, that provides performance profile of source code in a cycle-accurate manner. The HLS\_Profiler tool is non-intrusive and collectively uses the $łangle$static analysis, dynamic trace$\rangle$ of the source code to present the performance profile report to attribute latent clock cycles to each line of source code. Additionally, we developed a set of associative rules to maintain correctness in performance profile of the HLS design. To verify correctness, we demonstrate the HLS\_Profiler tool on MachSuite Benchmarks and an industry-grade recommendation application. The proposed HLS\_Profiler framework provides visibility into the cycle-by-cycle hardware execution of source-code and aids the designer in making performance-centric decisions. Nupur Sumeet, Deeksha Deeksha, Manoj Nambiar 0001 |
ICPE | 1 |
| 2021 | HLS_PRINT: High Performance Logging Framework on FPGAabstractRecent availability of High-Level Synthesis (HLS) development flow from FPGA vendors like Xilinx [1] and Intel [2] have simplified hardware design development to a great extent. A HLS design development flow includes a compiler which can compile a high-level language, such as C/C++, into the corresponding HDL (Hardware Description Language) representation. Developing applications using HLS would entice enterprise users, given the simplicity of coding in a high-level language. Almost all data center applications running in the data center would log important data. This included the run time contextual data, intermediate steps and final results. This information could serve many purposes like auditing, data for machine learning or just troubleshooting application execution issues. The last requirement is very essential to ensure reduced downtime as availability issues could result in significant loss in business. Logging was not possible with FPGA applications developed even in HLS. Collecting log data necessitates the use of hardware vendor specific modules called integrated logic analyzers (ILA) [3] which require manual integration and support small buffers which limits the amount of data that can be captured.Addressing these issues, we built a logging framework which would enable logging for FPGA implemented applications just as they would for software-based applications. The data would be available to operations staff in similar form as software applications. The logging framework is very similar to the use of printf function available in the C stdio library that comes with standard C-based software compilers.Our logging framework, HLS_PRINT [4], is a hardware-software solution to enable print functionality in the HLS design platforms to bridge the hardware logging gap. We use source-to-source transformations to create HLS synthesizable print constructs. In addition to this, the HLS_PRINT offers a push-button integration into an existing HDL project. The software part of the framework presents logged data in human readable format. Nupur Sumeet, Manoj Nambiar 0001 |
FPL | 1 |
| 2021 | HLS_PRINT: High Performance Logging Framework on FPGAabstractFPGAs have been tipped to be useful for implementing low latency transaction processing systems. Getting computationally powerful over time, they are making their way into enterprise data centers. Another factor is the availability of C compilers for FPGAs as opposed to hardware description languages (HDLs) that requires special skills. However, data center operations staff were concerned about real time troubleshooting in production. Tracing FPGA implemented application execution require special skills and vendor specific tools that can capture limited by the amount of data. To address this, we designed and implemented a logging frame-work on the FPGA. This paper presents the design and implementation of the framework. We present an algorithm that checks and generates alerts for performance overheads introduced due to the use of logging. Finally, experimental results are presented which demonstrate zero or low overhead of the logging framework. Nupur Sumeet, Manoj Nambiar 0001 |
ICPE | 1 |