EDBT 2026 Demo / reviewers in the wild / expert
Ajay Dholakia
dblp:27/5574
· DBLP profile ↗
5ranked-venue papers in the field
1as first author
1since 2021 · last 2021
0009-0007-8973-6063ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4 (1 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Recent Efficiency Gains in Deep Learning: Performance, Power, and SustainabilityabstractDeep learning (DL) continues to develop at a rapid pace with improvements coming from both hardware and software sides. In this work we evaluate the strides made over the last 2 years, during which a new generation of GPU accelerators has been introduced and significant algorithmic progress has been made. We find a dramatic improvement in runtime and power usage for a standard AI training workload. Specifically, a 4x improvement in runtime and energy consumption is demonstrated. The improvements are about equally split between hardware and algorithms. Additionally, we examine further ways to improve AI training power consumption on data center servers and identity 3 system level tunings that make most difference. These yield up to 20% more energy savings without any changes to the user code. Implications for the field and ways to make DL more energy-efficient going forward are also discussed. (Abstract) Miro Hodak, Ajay Dholakia |
IEEE BigData | 2 |
| 2019 | Towards Power Efficiency in Deep Learning on Data Center HardwareabstractDeep learning (DL) is a computationally intensive workload that is expected to grow rapidly in data centers in the near future. Its high energy demand necessitates finding ways to improve computational efficiency. In this work, we directly measure power used by the whole system as well as that used by GPU, CPU, and RAM during DL training to determine their contributions to the overall energy consumption. We find that while GPUs use most of the power - about 70 % - the consumption of other components is also significant and their optimizations can bring important power savings. Evaluating a multitude of options, we identify the parameters that bring in the most power savings. Overall, an energy savings of over 20% of can be obtained by adjusting system settings alone without changing the workload, at the cost of a minor increase in runtime. Alternatively, if runtime needs to stay constant, an 18% energy savings is identified. In distributed multi-server DL, we find that scale-out overhead has only a small energy cost, making distributed training more energy-efficient than expected. Implications for the field and ways to make DL more energy-efficient going forward are also discussed. Miro Hodak, Masha Gorkovenko, Ajay Dholakia |
IEEE BigData | 3 |
| 2018 | Performance Implications of Big Data in Scalable Deep Learning: On the Importance of Bandwidth and CachingabstractDeep learning techniques have revolutionized many areas including computer vision and speech recognition. While such networks require tremendous amounts of data, the requirement for and connection to Big Data storage systems is often undervalued and not well understood. In this paper, we explore the relationship between Big Data storage, networking, and Deep Learning workloads to understand key factors for designing Big Data/Deep Learning integrated solutions. We find that storage and networking bandwidths are the main parameters determining Deep Learning training performance. Local data caching can provide a performance boost and eliminate repeated network transfers, but it is mainly limited to smaller datasets that fit into memory. On the other hand, local disk caching is an intriguing option that is overlooked in current state-of-the-art systems. Finally, we distill our work into guidelines for designing Big Data/Deep Learning solutions. Miro Hodak, David Ellison, Peter Seidel, Ajay Dholakia |
IEEE BigData | 4 |
| 2017 | Designing a high performance cluster for large-scale SQL-on-hadoop analyticsabstractExecuting and optimizing SQL analytics on Data Lakes and Enterprise Data Warehouses (EDW) are areas of significant and growing interest. Achieving high performance for SQL analytics on large-scale data repositories remains a key challenge for data practitioners. The SQL-on-Hadoop Analytics solution described in this paper is very well suited for implementing the infrastructure to support these modern analytics initiatives while meeting requirements such as higher performance, lower cost, more efficient data center footprint, lower power consumption, appropriate storage needs and increased reliability. By using a TPC-DS derived workload applied to 100 TB of data, the work demonstrates for the first time the feasibility of designing such an extremely high-performance cluster. Furthermore, it enables investigation of large-scale SQL-on-Hadoop systems as the Spark SQL framework matures and enables similar investigations into machine learning and related Spark capabilities. Ajay Dholakia, Prasad Venkatachar, Kshitij A. Doshi, Ravikanth Durgavajhala, Stewart Tate, Berni Schiefer, Matthew Sheard, Ramnath Sai Sagar |
IEEE BigData | 1 |
| 2003 | A Nanotechnology-based Approach to Data Storage
Evangelos Eleftheriou, Peter Bächtold, Giovanni Cherubini, Ajay Dholakia, Christoph Hagleitner, Teddy Loeliger, Angeliki Pantazi, Haralampos Pozidis, T. R. Albrecht, Gerd Karl Binnig, Michel Despont, Ute Drechsler, Urs Dürig, Bernd Gotsmann, Daniel Jubin, Walter Häberle, Mark A. Lantz, Hugo E. Rothuizen, Richard Stutz, Peter Vettiger, Dorothea Wiesmann |
VLDB | 4 |