EDBT 2026 Demo / reviewers in the wild / expert
Pankaj Mehra
dblp:91/1566
· DBLP profile ↗
24ranked-venue papers
8as first author
3since 2021 · last 2023
0009-0003-3188-243XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 74% Storage systems · 18% Distributed systems · 4% | |
| Software engineering, system software, and programming languages
5 papers |
Operating systems · 81% Services computing and microservices · 18% Program synthesis and code generation · 1% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 99% Machine learning and data management · 1% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 96% Representation and self-supervised learning · 4% |
Topics — the 21 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
0.9 | 2 | 2021 | Twizzler: A Data-centric OS for Non-volatile Memory · ACM Trans. Storage 2021 Twizzler: a Data-Centric OS for Non-Volatile Memory · USENIX ATC 2020 |
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory |
0.5 | 1 | 2021 | Twizzler: A Data-centric OS for Non-volatile Memory · ACM Trans. Storage 2021 |
Storage systems › computational storage
in-storage computing |
0.4 | 1 | 2019 | Practical Near-Data Processing to Evolve Memory and Storage Devices into Mainstream Heterogeneous Computing Systems · DAC 2019 |
Memory systems › processing-in-memory
near-data processing |
0.4 | 1 | 2019 | Practical Near-Data Processing to Evolve Memory and Storage Devices into Mainstream Heterogeneous Computing Systems · DAC 2019 |
Memory systems
processing-in-memory |
0.4 | 1 | 2019 | Practical Near-Data Processing to Evolve Memory and Storage Devices into Mainstream Heterogeneous Computing Systems · DAC 2019 |
Distributed systems › distributed system architecture
heterogeneous distributed systems |
0.1 | 1 | 2019 | Practical Near-Data Processing to Evolve Memory and Storage Devices into Mainstream Heterogeneous Computing Systems · DAC 2019 |
Information retrieval › information filtering
document routing |
0.1 | 1 | 2007 | Content-based document routing and index partitioning for scalable similarity-based searches in a large corpus · KDD 2007 |
Information retrieval › indexing
index partitioning |
0.1 | 1 | 2007 | Content-based document routing and index partitioning for scalable similarity-based searches in a large corpus · KDD 2007 |
Information retrieval
similarity search |
0.1 | 1 | 2007 | Content-based document routing and index partitioning for scalable similarity-based searches in a large corpus · KDD 2007 |
Performance modeling and evaluation
performance prediction |
0.0 | 2 | 1995 | Automated Performance Prediction of Message-Passing Parallel Programs · SC 1995 A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994 |
Performance modeling and evaluation › performance model construction
automated performance modeling |
0.0 | 1 | 1995 | Automated Performance Prediction of Message-Passing Parallel Programs · SC 1995 |
Parallel and multicore computing › parallel computing
parallel program analysis |
0.0 | 1 | 1995 | Automated Performance Prediction of Message-Passing Parallel Programs · SC 1995 |
Performance modeling and evaluation › parallel system performance
parallel performance modeling |
0.0 | 1 | 1994 | A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994 |
Parallel and multicore computing › load balancing
dynamic load balancing |
0.0 | 1 | 1992 | Physical-Level Synthetic Workload Generation for Load-Balancing Experiments · HPDC 1992 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 1992 | Physical-Level Synthetic Workload Generation for Load-Balancing Experiments · HPDC 1992 |
Performance modeling and evaluation › workload characterization › workload modeling
synthetic workload generation |
0.0 | 1 | 1992 | Physical-Level Synthetic Workload Generation for Load-Balancing Experiments · HPDC 1992 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 1992 | Physical-Level Synthetic Workload Generation for Load-Balancing Experiments · HPDC 1992 |
Machine learning › Representation and self-supervised learning › automated feature generation
constructive induction |
0.0 | 1 | 1989 | Principled Constructive Induction · IJCAI 1989 |
Program synthesis and code generation › inductive logic programming
constructive induction |
0.0 | 1 | 1989 | Constructive Induction Framework · ML 1989 |
Parallel and multicore computing › parallel programming models › message passing
message-passing parallel programs |
0.0 | 1 | 1995 | Automated Performance Prediction of Message-Passing Parallel Programs · SC 1995 |
Parallel and multicore computing
message-passing programs |
0.0 | 1 | 1994 | A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel Programs · SIGMETRICS 1994 |
Methods — techniques the papers use, named apart from their topics
cross-object pointers · 1.0NVM DIMM evaluation · 1.0unsupervised information extraction · 0.3parsing · 0.3linguistic patterns · 0.3feature-based routing · 0.1distributed search · 0.1synthetic workload generation · 0.0resource-utilization replay · 0.0analytic execution time modeling · 0.0statistical regression · 0.0simulation · 0.0profiling · 0.0constructive induction · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TMC: Near-Optimal Resource Allocation for Tiered-Memory SystemsabstractMain memory dominates data center server cost, and hence data center operators are exploring alternative technologies such as CXL-attached and persistent memory to improve cost without jeopardizing performance. Introducing multiple tiers of memory introduces new challenges, such as selecting the appropriate memory configuration for a given workload mix. In particular, we observe that inefficient configurations increase cost by up to 2.6× for clients, and resource stranding increases cost by 2.2× for cloud operators. To address this challenge, we introduce TMC, a system for recommending cloud configurations according to workload characteristics and the dynamic resource utilization of a cluster. Whereas prior work utilized extensive simulation or costly machine learning techniques, incurring significant search costs, our approach profiles applications to reveal internal properties that lead to fast and accurate performance estimations. Our novel configuration-selection algorithm incorporates a new heuristic, packing penalty, to ensure that recommended configurations will also achieve good resource efficiency. Our experiments demonstrate that TMC reduces the search cost by up to 4× over the state-of-the-art, while improving resource utilization by up to 17% as compared to a naive policy that requests optimal tiered memory allocations in isolation. Yuanjiang Ni, Pankaj Mehra, Ethan L. Miller, Heiner Litz |
SoCC | 2 |
| 2021 | Don't Let RPCs Constrain Your APIabstractAs data becomes increasingly distributed, traditional RPC and data serialization limits performance, result in rigidity, and hamper expressivity. We believe that technology trends including high-density persistent memory, high-speed networks, and programmable switches make this the right time to revisit prior research on distributed shared memory, global addressing, and content-based networking. Our vision combines the code mobility of RPC with first-class data references in a global address space by co-designing the OS and the network around pervasive data identity. We have initial results showing the promise of the proposed co-design. Daniel Bittman, Robert Soulé, Ethan L. Miller, Vishal Shrivastav, Pankaj Mehra, Matthew Boisvert, Avi Silberschatz, Peter Alvaro |
HotNets | 5 |
| 2021 | Twizzler: A Data-centric OS for Non-volatile MemoryabstractByte-addressable, non-volatile memory (NVM) presents an opportunity to rethink the entire system stack. We present Twizzler, an operating system redesign for this near-future. Twizzler removes the kernel from the I/O path, provides programs with memory-style access to persistent data using small (64 bit), object-relative cross-object pointers, and enables simple and efficient long-term sharing of data both between applications and between runs of an application. Twizzler provides a clean-slate programming model for persistent data, realizing the vision of Unix in a world of persistent RAM. We show that Twizzler is simpler, more extensible, and more secure than existing I/O models and implementations by building software for Twizzler and evaluating it on NVM DIMMs. Most persistent pointer operations in Twizzler impose less than 0.5 ns added latency. Twizzler operations are up to faster than Unix , and SQLite queries are up to faster than on PMDK. YCSB workloads ran 1.1– faster on Twizzler than on native and NVM-optimized SQLite backends. Daniel Bittman, Peter Alvaro, Pankaj Mehra, Darrell D. E. Long, Ethan L. Miller |
ACM Trans. Storage | 3 |
| 2020 | Twizzler: a Data-Centric OS for Non-Volatile Memory
Daniel Bittman, Peter Alvaro, Pankaj Mehra, Darrell D. E. Long, Ethan L. Miller |
USENIX ATC | 3 |
| 2019 | Practical Near-Data Processing to Evolve Memory and Storage Devices into Mainstream Heterogeneous Computing SystemsabstractThe capacity of memory and storage devices is expected to increase drastically with adoption of the forthcoming memory and integration technologies. This is a welcome improvement especially for datacenter servers running modern data-intensive applications. Nonetheless, for such servers to fully benefit from the increasing capacity, the bandwidth of interconnects between processors and these devices must also increase proportionally, which becomes ever costlier under unabating physical constraints. As a promising alternative to tackle this challenge cost-effectively, a heterogeneous computing paradigm referred to as near-data processing (NDP) has emerged. However, NDP has not yet been widely adopted by the industry because of significant gaps between existing software stacks and demanded ones for NDP-capable memory and storage devices. Aiming to overcome the gaps, we propose to turn memory and storage devices into familiar heterogeneous distributed computing systems. Then, we demonstrate potentials of such computing systems for existing data-intensive applications with two recently implemented NDP-capable devices. Finally, we conclude with a practical blueprint to exploit the NDP-based computing systems for speeding up solving future computer-aided design and optimization problems. Nam Sung Kim, Pankaj Mehra |
DAC | 2 |
| 2012 | Mining Business Contracts for Service ExceptionsabstractA contract is a legally binding agreement between real-world business entities whom we treat as providing services to one another. We focus on business rather than technical services. We think of a business contract as specifying the functional and nonfunctional behaviors of and interactions among the services. In current practice, contracts are produced as text documents. Thus the relevant service capabilities, requirements, qualities, and risks are hidden and difficult to access and reason about. We describe a simple but effective unsupervised information extraction approach and tool, Contract Miner, for discovering service exceptions at the phrase level from a large contract repository. Our approach involves preprocessing followed by an application of linguistic patterns and parsing to extract the service exception phrases. Identifying such (noun) phrases can help build service exception vocabularies that support the development of a taxonomy of business terms, and also facilitate modeling and analyzing service engagements. A lightweight online tool that comes with Contract Miner highlights the relevant text in service contracts and thereby assists users in reviewing contracts. Contract Miner produces promising results in terms of precision and recall when evaluated over a corpus of manually annotated contracts. Xibin Gao, Munindar P. Singh, Pankaj Mehra |
IEEE Trans. Serv. Comput. | 3 |
| 2008 | Growing Fields of Interest - Using an Expand and Reduce Strategy for Domain Model ExtractionabstractDomain hierarchies are widely used as models underlying information retrieval tasks. Formal ontologies and taxonomies enrich such hierarchies further with properties and relationships but require manual effort; therefore they are costly to maintain, and often stale. Folksonomies and vocabularies lack rich category structure. Classification and extraction require the coverage of vocabularies and the alterability of folksonomies and can largely benefit from category relationships and other properties. With Doozer, a program for building conceptual models of information domains, we want to bridge the gap between the vocabularies and Folksonomies on the one side and the rich, expert-designed ontologies and taxonomies on the other. Doozer mines Wikipedia to produce tight domain hierarchies, starting with simple domain descriptions. It also adds relevancy scores for use in automated classification of information. The output model is described as a hierarchy of domain terms that can be used immediately for classifiers and IR systems or as a basis for manual or semi-automatic creation of formal ontologies. Christopher Thomas 0001, Pankaj Mehra, Roger Brooks, Amit P. Sheth |
Web Intelligence | 2 |
| 2007 | Content-based document routing and index partitioning for scalable similarity-based searches in a large corpusabstractWe present a document routing and index partitioning scheme for scalable similarity-based search of documents in a large corpus. We consider the case when similarity-based search is performed by finding documents that have features in common with the query document. While it is possible to store all the features of all the documents in one index, this suffers from obvious scalability problems. Our approach is to partition the feature index into multiple smaller partitions that can be hosted on separate servers, enabling scalable and parallel search execution. When a document is ingested into the repository, a small number of partitions are chosen to store the features of the document. To perform similarity-based search, also, only a small number of partitions are queried. Our approach is stateless and incremental. The decision as to which partitions the features of the document should be routed to (for storing at ingestion time, and for similarity based search at query time) is solely based on the features of the document. Deepavali Bhagwat, Kave Eshghi, Pankaj Mehra |
KDD | 3 |
| 2004 | Fast and Flexible Persistence: The Magic Potion for Fault-Tolerance, Scalability and Performance in Online Data StoresabstractSummary form only given. We examine the architecture of computer systems designed to update, integrate and serve enterprise information. Our analysis of their scale, performance, availability and data integrity draws attention to the so-called 'storage gap'. We propose 'persistent memory' as the technology to bridge that gap. Its impact is demonstrated using practical business information processing scenarios. We conclude with a discussion of our prototype, the results achieved, and the challenges that lie ahead. Pankaj Mehra, Samuel A. Fineberg |
IPDPS | 1 |
| 2003 | The Quest for the Perfect Server for Network Computing ApplicationsabstractThis talk begins with an examination of server-side architecture and workload trends. Severe performance and availability gaps in server architecture, compounded by the growing cost and complexity of development and deployment, are shown hampering the successful adoption of Web services. Two architectural remedies are proposed: (I) unification of I/O and memory; and (II) whole system programming. The proposed Apsara architecture for servers combines these new techniques with the best of existing server architecture concepts. The talk concludes with identification of key technical challenges and opportunities, and insights from early experiences in implementing Apsara. Pankaj Mehra |
NCA | 1 |
| 2002 | Flow Control in ServerNet® Clusters
Vladimir Shurbanov, Dimiter R. Avresky, Pankaj Mehra, William J. Watson |
J. Supercomput. | 3 |
| 2001 | A Queueing Model for Space-Division Packets Switches and Its Application to the Performance Evaluation of Computer NetworksabstractA closed queueing model of a generalized input-queueing space-division packet switch is presented. The model explicitly reflects the architecture of the router in terms of the number of outputs and the number of packets available for transmission, as well as the distribution of traffic over the ports. The model provides the upper bound on the utilization (throughput) of the router given the number of input and output ports of the router and the distribution of traffic among them. The results obtained by the model are verified by comparison with data generated by a simulator. It is demonstrated that the results are also useful far accurately estimating the throughput of networks. Vladimir Shurbanov, Dimiter R. Avresky, Pankaj Mehra |
IPDPS | 3 |
| 2001 | Optimal Utilization of Equivalent Paths in Computer Networks with Static RoutingabstractFocuses on the utilization of alternative communication paths in local and system area networks with static routing. A lot of research work has been devoted to employing such paths for fault tolerance, but the issue of utilizing them for performance enhancement has been largely neglected, especially for static routing networks. This work formally proves that the throughput of multiple paths is maximal if the traffic is uniformly distributed over them. Based on this, a procedure for destination partitioning in static routing networks is introduced. It is applicable to arbitrary multi-path topologies and traffic patterns that lend themselves to partitioning. The procedure is applied to several topologies with different degree of equivalent paths coverage and their performance is evaluated through simulations. The results demonstrate that the network performance is significantly improved when the proposed partitioning procedure is applied. Dimiter R. Avresky, Vladimir Shurbanov, Natcho H. Natchev, F. Zuccarino, Pankaj Mehra |
NCA | 5 |
| 2001 | Trends in System Area NetworkingabstractAs bus architectures approach the end of their lifetimes, and as networking extends its reach deeper into systems, the future of systems and networking appears rife with possibilities. This paper attempts to look beyond the rivalry between computer companies over buses, to a world in which computation and networking are both pervasive. In order to delve deeper into some of these issues, the discussion focuses on the data center microcosm, where the impact of these trends is both visible and relevant. What we find there is components that defy description in current terminology, new locales for existing functions, and perhaps most important, a whole new class of workloads. Pankaj Mehra |
NCA | 1 |
| 2000 | Flow Control in ServerNetR Clusters
Vladimir Shurbanov, Dimiter R. Avresky, Pankaj Mehra, William J. Watson |
Euro-Par | 3 |
| 1996 | The Effect of Interrupts on Software Pipeline Execution on Message-Passing ArchitecturesabstractObservationsshow that fine-grain software pipelines on Rob F. Van der Wijngaart, Sekhar R. Sarukkai, Pankaj Mehra |
International Conference on Supercomputing | 3 |
| 1996 | Analysis and Optimization of Software Pipeline Performance on MIMD Parallel Computers
Rob F. Van der Wijngaart, Sekhar R. Sarukkai, Pankaj Mehra |
J. Parallel Distributed Comput. | 3 |
| 1995 | Automated Performance Prediction of Message-Passing Parallel ProgramsabstractThe increasing use of massively parallel supercomputers to solve large-scale scientific problems has generated a need for tools that can predict scalability trends of applications written for these machines. Much work has been done to create simple models that represent important characteristics of parallel programs, such as latency, network contention, and communication volume. But many of these methods still require substantial manual effort to represent an application in the model's format. The MK toolkit described in this paper is the result of an on-going effort to automate the formation of analytic expressions of program execution time, with a minimum of programmer assistance. In this paper we demonstrate the feasibility of our approach, by extending previous work to detect and model communication patterns automatically, with and without overlapped computations. The predictions derived from these models agree, within reasonable limits, with execution times of programs measured on the Intel iPSC/860 and Paragon. Further, we demonstrate the use of MK in selecting optimal computational grain size and studying various scalability metrics. Robert J. Block, Sekhar R. Sarukkai, Pankaj Mehra |
SC | 3 |
| 1995 | Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs using the AIMS ToolkitabstractAbstract Writing large‐scale parallel and distributed scientific applications that make optimum use of the multiprocessor is a challenging problem. Typically, computational resources are underused due to performance failures in the application being executed. Performance‐tuning tools are essential for exposing these performance failures and for suggesting ways to improve program performance. In this paper, we first address fundamental issues in building useful performance‐tuning tools and then describe our experience with the AIMS toolkit for tuning parallel and distributed programs on a variety of platforms. AIMS supports source‐code instrumentation, run‐time monitoring, graphical execution profiles, performance indices and automated modeling techniques as ways to expose performance problems of programs. Using several examples representing a broad range of scientific applications, we illustrate AIMS' effectiveness in exposing performance problems in parallel and distributed programs. Jerry C. Yan, Sekhar R. Sarukkai, Pankaj Mehra |
Softw. Pract. Exp. | 3 |
| 1994 | A Comparison of Two Model-Based Performance-Prediction Techniques for Message-Passing Parallel ProgramsabstractThis paper describes our experience in modeling two significant parallel applications: ARC2D, a 2-dimensional Euler solver; and, Xtrid, a tridiagonal linear solver. Both of these models were expressed in BDL (Behavior Description language) and simulated on an iPSC/860 Hypercube modeled using Axe (Abstract eXecution Environment). BDL models consist of abstract communicating objects: blocks of sequential code are modeled by single RUN statements; all communication operations in the original code are mirrored by corresponding BDL operations in the model. Our ARC2D model was built by first profiling the program to locate the significant loops and then timing the basic blocks within those loops. Simulated completion times were (except in one case) within 8% of measured execution times. Lengthy simulations were necessary for predicting the performance of large-scale runs. For Xtrid, only the loops surrounding communications were modeled; other loops were absorbed into large sequential blocks whose complexity was estimated using statistical regression. This approach yielded a much smaller model whose computation and communication complexities were clearly manifest. Analysis of complexity allowed rapid prediction of large-scale performance without lengthy simulations! Analytically predicted speed-ups were within 7% of those predicted by simulation. Simulated completion times were within 5% of measured execution times. The second approach provides a more effective methodology for simulation-based performance-tuning. Pankaj Mehra, Catherine H. Schulbach, Jerry C. Yan |
SIGMETRICS | 1 |
| 1993 | Automated Learning of Workload Measures for Load Balancing on a Distributed SystemabstractLoad-balancing systems use workload indices to dynamically schedule jobs. We present a novel method of automatically learning such indices. Our approach uses comparator neural networks, one per site, which learn to predict the relative speedup of an incoming job using only the resource-utilization patterns observed prior to the job's arrival. Our load indices combine information from the key resources of contention: CPU, disk, network, and memory. Pankaj Mehra, Benjamin W. Wah |
ICPP (3) | 1 |
| 1992 | Physical-Level Synthetic Workload Generation for Load-Balancing ExperimentsabstractSynthetic workload generation uses artificial programs to mimic the resource-utilization patterns of real workloads. It is important for systematic evaluation of dynamic load-balancing strategies because load-balancing experiments require the measurement of task-completion times under realistic and reproducible workloads. The authors describe a generator that permits accurate replay of measured system-wide loads, thus providing an ideal setting for conducting load-balancing experiments. The generator is implemented inside the operating-system kernel and, therefore, has complete control over the local resources. It controls the useage levels of four key resources: CPU, memory, disk, and network. In order to reproduce accurately the behavior of the process population generating the measured load, the generator gives up a fraction of its resources in response to the arrival of new jobs, and reclaims these resources when the jobs terminate. The authors results show near-perfect reproduction of background load even in the presence of interfering foreground load.> Pankaj Mehra, Benjamin W. Wah |
HPDC | 1 |
| 1989 | Constructive Induction Framework
Pankaj Mehra |
ML | 1 |
| 1989 | Principled Constructive Induction
Pankaj Mehra, Larry A. Rendell, Benjamin W. Wah |
IJCAI | 1 |