Herman Lam

dblp:l/HermanLam · DBLP profile ↗
← Back
61ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-2388-0819ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14Software engineering, systems software and programming languages · 13 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorArtificial intelligence and machine learning · 5Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 RISCBench: Benchmarking RISC-V Orchestration Efficiency in FPGA and FPGA-Like Computing Engines: An Industry Evaluation of Control-Plane Bottlenecks and Sustained Throughput Metrics
abstract
Heterogeneous systems increasingly rely on RISC-V cores as orchestration engines to manage data movement, synchronization, and scheduling across accelerators and reconfigurable fabrics. Conventional performance metrics, such as FLOPs, TOPS/W, or energy per operation, do not capture orchestration efficiency, even though it often dictates sustained system behavior. This gap is increasingly relevant as systems evolve toward tightly coupled heterogeneous fabrics and co-packaged accelerators, where control-plane behavior determines whether these platforms achieve their promised performance. We present RISCBench, a kernel benchmark suite and open methodology for quantifying orchestration efficiency. RISCBench introduces the Sustained Instantaneous Throughput (SIT) metric, which accumulates instantaneous throughput over near-aggregate execution intervals, capturing sustained efficiency beyond peak rates. The methodology is evaluated across representative platforms spanning soft and hard RISC-V orchestration engines, including FPGA-based prototyping and accelerator-class implementations. Results highlight synchronization and data residency driven tradeoffs that limit realized throughput beyond peak performance, motivating SIT as a practical, platform-independent descriptor for evaluating orchestration efficiency in heterogeneous systems and AI inference applications.
David Ojika, Projjal Gupta, Preethi Budi, Herman Lam, Shreya Mehrotra
FPGA4
2026 L-PCN: A Point Cloud Accelerator Exploiting Spatial Locality through Octree-Based Islandization
Jieming Yin, ZhiLei Chai, Jiliang Zhang 0011, Herman Lam
ISCA8
2025 APR-OIS: A Near-Sensor Point Cloud Pre-Processing Accelerator on FPGA
abstract
Raw point clouds generated by 3D sensors (such as Li-DARs, and RGB-D cameras) are typically large in scale, containing an enormous number of points. For this reason, the raw point clouds usually require an expensive down-sampling phase to reduce the number of points while preserving their spatial structure, before being used as input to most point cloud applications. In recent years, the Octree spatial indexing method [2] for the point clouds has been used to accelerate the down-sampling phase. Based on the existing Octree-based down-sampling method and spatial proximity function of Octree nodes, we propose an approximation method, Approximate Octree-Indexed Sampling (APR-OIS), to further accelerate Octree-based down-sampling with a minimal loss of accuracy. It is implemented as an FPGA-based accelerator, functioning as a near-sensor processor
Herman Lam
FCCM2
2024 HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
abstract
Point cloud is an important type of geometric data structure for many embedded applications such as autonomous driving and augmented reality. Current Point Cloud Networks (PCNs) have proven to achieve great success in using inference to perform point cloud analysis, including object part segmentation, shape classification, and so on. However, point cloud applications on the computing edge require more than just the inference step. They require an end-to-end (E2E) processing of the point cloud workloads: pre-processing of raw data, input preparation, and inference to perform point cloud analysis. Current PCN approaches to support end-to-end processing of point cloud workload cannot meet the real-time latency requirement on the edge, i.e., the ability of the AI service to keep up with the speed of raw data generation by 3D sensors. Latency for end-to-end processing of the point cloud workloads stems from two reasons: memory-intensive down-sampling in the pre-processing phase and the data structuring step for input preparation in the inference phase. In this paper, we present HgPCN, an end-to-end heterogeneous architecture for real-time embedded point cloud applications. In HgPCN, we introduce two novel methodologies based on spatial indexing to address the two identified bottlenecks. In the Pre-processing Engine of HgPCN, an Octree-Indexed-Sampling method is used to optimize the memory-intensive down-sampling bottleneck of the pre-processing phase. In the Inference Engine, HgPCN extends a commercial DLA with a customized Data Structuring Unit which is based on a Voxel-Expanded Gathering method to fundamentally reduce the workload of the data structuring step in the inference phase. The initial prototype of HgPCN has been implemented on an Intel PAC (Xeon+FPGA) platform. Four commonly available point cloud datasets were used for comparison, running on three baseline devices: Intel Xeon W-2255, Nvidia Xavier NX Jetson GPU, and Nvidia 4060ti GPU. These point cloud datasets were also run on two existing PCN accelerators for comparison: PointACC and Mesorasi. Our results show that for the inference phase, depending on the dataset size, HgPCN achieves speedup from 1.3× To 10.2× vs. PointACC, 2.2× To 16.5× vs. Mesorasi, and 6.4× To 21× vs. Jetson NX GPU. Along with optimization of the memory-intensive down-sampling bottleneck in pre-processing phase, the overall latency shows that HgPCN can reach the real-time requirement by providing end-to-end service with keeping up with the raw data generation rate.
Wesley Piard, Bhavesh Patel, Herman Lam
MICRO6
2021 Incorporating Fault-Tolerance Awareness into System-Level Modeling and Simulation
abstract
Designing High Performance Computing (HPC) systems involves exploring large design spaces composed of many hardware and software options to meet system objectives. Various modeling and simulation (MODSIM) techniques help design space exploration (DSE) by providing a low-cost way to predict system metrics before building. Larger and more heterogeneous systems created to meet higher system objectives lead to an increase in system fault rates, which negatively affects performance and other system objectives. Therefore, systems employ various fault-tolerance (FT) techniques, such as checkpoint restart (C/R), to mitigate these negative effects. However, because FT techniques incur some degree of overhead, it is important to include the effects of faults and fault-tolerance in modeling and simulation.
Trokon Johnson, Herman Lam
CLUSTER2
2021 Optimized FPGA-based Deep Learning Accelerator for Sparse CNN using High Bandwidth Memory
abstract
Large Convolutional Neural Networks (CNNs) are often pruned and compressed to reduce the amount of parameters and memory requirement. However, the resulting irregularity in the sparse data makes it difficult for FPGA accelerators that contains systolic arrays of Multiply-and-Accumulate (MAC) units, such as Intel's FPGA-based Deep Learning Accelerator (DLA), to achieve their maximum potential. Moreover, FPGAs with low-bandwidth off-chip memory could not satisfy the memory bandwidth requirement for sparse matrix computation. In this paper, we present 1) a sparse matrix packing technique that condenses sparse inputs and filters before feeding them into the systolic array of MAC units in the Intel DLA, and 2) a customization of the Intel DLA which allows the FPGA to efficiently utilize a high bandwidth memory (HBM2) integrated in the same package. For end-to-end inference with randomly pruned ResNet-50/MobileNet CNN models, our experiments demonstrate 2.7x/3x performance improvement compared to an FPGA with DDR4, 2.2x/2.1x speedup against a server-class Intel SkyLake CPU, and comparable performance with 1.7x/2x power efficiency gain as compared to an NVidia V100 GPU.
David Ojika, Bhavesh Patel, Herman Lam
FCCM4
2019 Accelerating Scientific Discovery with SCAIGATE Science Gateway
abstract
The demand for computational accelerators (GPUs, FPGAs, ASICs, etc.) is growing due to the widening variety of datacenter applications fueled by recent scientific breakthroughs that leverage artificial intelligence (AI). As much as these applications (e.g., cosmology, physics, etc.) have continued to witness record-breaking accuracy in predictive capabilities due to AI widespread influence, the infrastructure and workflow to take these applications out of research labs into production and business use-cases continues to lag. To address these important infrastructural challenges, we present SCAIGATE, a prototype science gateway with a simplified workflow aimed at facilitating model building/validation workflows in large-scale scientific applications.
David Ojika, Bhavesh Patel, Ann Gordon-Ross, Herman Lam
eScience5
2018 Scalable Behavioral Emulation of Extreme-Scale Systems Using Structural Simulation Toolkit
abstract
With extremely large design spaces for algorithm and architecture to be explored, there is a need for fast and scalable performance modeling tools for preparing HPC application codes. Behavioral Emulation (BE) is a recent coarse-grained modeling and simulation methodology that has been proposed to solve this co-design problem. In this paper, we introduce a distributed parallel simulation library for Behavioral Emulation called BE-SST, integrated into the Structural Simulation Toolkit (SST). BE-SST provides simple interfaces and framework for development of coarse-grained BE models which can be extended to model new notional architectures. BE-SST also supports Monte Carlo simulations to generate meaningful distributions and summary statistics rather than a single datum for performance. In this paper, we present BE-SST simulations of two existing large DOE machines (Vulcan and Titan), which have been validated against actual testbed measurements and showed 5-10% error. These validated system models (up to 128k cores) are used to make blind predictions of application performance on systems larger than the current machines (up to 512k cores) - a crucial simulator feature for design-space exploration of notional systems. We further studied BE-SST in terms of scalability and performance, simulating up to a million cores, with BE-SST running on more than 2k parallel processes. BE-SST shows good scalability with a linear increase in memory usage and simulation time with increase in simulated system size, and a peak speedup of 7x over single process simulation. With ease of use and good scaling, we assert that BE-SST can significantly speed up design-space exploration.
Ajay Ramaswamy, Nalini Kumar, Aravind Neelakantan, Herman Lam, Greg Stitt
ICPP4
2016 Novo-G#: a multidimensional torus-based reconfigurable cluster for molecular dynamics
abstract
Summary Molecular dynamics (MD) is a large‐scale, communication‐intensive problem that has been the subject of high‐performance computing research and acceleration for years. Not surprisingly, the most success in accelerating MD comes from specialized systems such as the Anton machine. In this paper, we describe Novo‐G# (novo‐jee‐sharp), a multi‐node reconfigurable system designed for the acceleration of communication‐intensive scientific problems in general, and MD in particular. This system provides a high‐bandwidth, low‐latency 3D torus network to allow direct communication between kernels running on multiple field‐programmable gate arrays. We also present a performance model for Novo‐G# running the 3D Fast Fourier Transform (FFT) kernel that forms the core of MD simulations. We validate the model against published Anton performance data and through initial hardware experiments on Novo‐G#. Finally, through simulation studies, we show that this system at scale performs better than specialized systems like Anton and outperforms established CPU‐based clusters like Blue Gene/Q by an order of magnitude for the 3D FFT kernel, with greater flexibility and lower costs. Copyright © 2015 John Wiley & Sons, Ltd.
Abhijeet Lawande, Alan D. George, Herman Lam
Concurr. Comput. Pract. Exp.3
2016 Analysis of Fixed, Reconfigurable, and Hybrid Devices with Computational, Memory, I/O, & Realizable-Utilization Metrics
abstract
The modern processor landscape is a varied and diverse community. As such, developers need a way to quickly and fairly compare various devices for use with particular applications. This article expands the authors’ previously published computational-density metrics and presents an analysis of a new generation of various device architectures, including CPU, DSP, FPGA, GPU, and hybrid architectures. Also, new memory metrics are added to expand the existing suite of metrics to characterize the memory resources on various processing devices. Finally, a new relational metric, realizable utilization (RU) , is introduced, which quantifies the fraction of the computational density metric that an application achieves within an individual implementation. The RU metric can be used to provide valuable feedback to application developers and architecture designers by highlighting the upper bound on specific application optimization and providing a quantifiable measure of theoretical and realizable performance. Overall, the analysis in this article quantifies the performance tradeoffs among the architectures studied, the memory characteristics of different device types, and the efficiency of device architectures.
Justin Richardson, Alan D. George, Herman Lam
ACM Trans. Reconfigurable Technol. Syst.4
2015 Comparative analysis of OpenCL vs. HDL with image-processing kernels on Stratix-V FPGA
abstract
Application development with hardware description languages (HDLs) such as VHDL or Verilog involves numerous productivity challenges, limiting the potential impact of reconfigurable computing (RC) with FPGAs in high-performance computing. Major challenges with HDL design include steep learning curves, large and complex codes, long compilation times, and lack of development standards across platforms. A relative newcomer to RC, the Open Computing Language (OpenCL) reduces productivity hurdles by providing a platform-independent, C-based programming language. In this study, we conduct a performance and productivity comparison between three image-processing kernels (Canny edge detector, Sobel filter, and SURF feature-extractor) developed using Altera's SDK for OpenCL and traditional VHDL. Our results show that VHDL designs achieved a more efficient use of resources (59% to 70% less logic), however, both OpenCL and VHDL designs resulted in similar timing constraints (255MHzmax<; 325MHz). Furthermore, we observed a 6× increase in productivity when using OpenCL development tools, as well as the ability to efficiently port the same OpenCL designs without change to three different RC platforms, with similar performance in terms of frequency and resource utilization.
Kenneth Hill, Stefan Craciun, Alan D. George, Herman Lam
ASAP4
2015 CMT-bone: A Mini-App for Compressible Multiphase Turbulence Simulation Software
abstract
Designed with the goal of mimicking key features of real HPC workloads, mini-apps have become an important tool for co-design. An investigation of mini-app behavior can provide system designers with insight into the impact of architectures, programming models, and tools on application performance. Mini-apps can also serve as a platform for fast algorithm design space exploration, allowing the application developers to evaluate their design choices before significantly redesigning the application codes. Consequently, it is prudent to develop a mini-app alongside the full blown application it is intended to represent. In this paper, we present CMT-bone a mini-app for the compressible multiphase turbulence (CMT) application, CMT-nek, being developed to extend the physics of the CESAR Nek5000 application code. CMT-bone consists of the most computationally intensive kernels of CMT-nek and the communication operations involved in nearest-neighbor updates and vector reductions. The mini-app represents CMT-nek in its most mature state and going forward it will be developed in parallel with the CMT-nek application to keep pace with key new performance impacting changes. We describe these kernels and discuss the role that CMT-bone has played in enabling interdisciplinary collaboration by allowing application developers to work with computer scientists on performance optimization on current architectures and performance analysis on notional future systems.
Nalini Kumar, Mrugesh Sringarpure, Tania Banerjee, Jason Hackl, S. Balachandar 0001, Herman Lam, Alan D. George, Sanjay Ranka
CLUSTER6
2015 CE2016: Updated computer engineering curriculum guidelines
abstract
Joint ACM/IEEE Computer Society undergraduate computer engineering curriculum guidelines are slated for release in 2016. These update the 2004 guidelines commonly known as CE2004. The presenters are part of the task group leading the revisions and will give an overview of the latest draft. Participants will engage in discussions on potential improvements to the guidelines to ensure that they are useful to programs as they work to ensure their curricula reflect the state-of-the-art in computer engineering education and practice and are relevant for the coming decade.
Eric Durant, John Impagliazzo, Susan Conry, Robert B. Reese, Herman Lam, Victor P. Nelson, Joseph L. A. Hughes, Junlin Lu, Andrew D. McGettrick
FIE5
2015 Low-level PGAS computing on many-core processors with TSHMEM
abstract
Summary Diminishing returns from increased clock frequencies and instruction‐level parallelism have forced computer architects to adopt architectures that exploit wider parallelism through multiple processor cores. While emerging many‐core architectures have progressed at a remarkable rate, concerns arise regarding the performance and productivity of numerous parallel‐programming tools for application development. Development of parallel applications on many‐core processors often requires developers to familiarize themselves with unique characteristics of a target platform while attempting to maximize performance and maintain correctness of their applications. The family of partitioned global address space (PGAS) programming models comprises the current state of the art in balancing performance and programmability. One such PGAS approach is SHMEM, a lightweight, shared‐memory programming library that has demonstrated high performance and productivity potential for parallel‐computing systems with distributed‐memory architectures. In the paper, we present research, design, and analysis of a new SHMEM infrastructure specifically crafted for low‐level PGAS on modern and emerging many‐core processors featuring dozens of cores and more. Our approach (with a new library known as TSHMEM) is investigated and evaluated atop two generations of Tilera architectures, which are among the most sophisticated and scalable many‐core processors to date, and is intended to enable similar libraries atop other architectures now emerging. In developing TSHMEM, we explore design decisions and their impact on parallel performance for the Tilera TILE‐Gx and TILEPro many‐core architectures, and then evaluate the designs and algorithms within TSHMEM through microbenchmarking and applications studies with other communication libraries. Our results with barrier primitives provided by the Tilera libraries show dissimilar performance between the TILE‐Gx and TILEPro; therefore, TSHMEM's barrier design takes an alternative approach and leverages the on‐chip mesh network to provide consistent low‐latency performance. In addition, our experiments with TSHMEM show that naive collective algorithms consistently outperformed linear distributed collective algorithms when executed in an SMP‐centric environment. In leveraging these insights for the design of TSHMEM, our approach outperforms the OpenSHMEM reference implementation, achieves similar to positive performance over OpenMP and OSHMPI atop MPICH, and supports similar libraries in delivering high‐performance parallel computing to emerging many‐core systems. Copyright © 2015 John Wiley & Sons, Ltd.
Bryant C. Lam, Alan D. George, Herman Lam, Vikas Aggarwal
Concurr. Comput. Pract. Exp.3
2014 Setting the stage for CE2016: A revised body of knowledge
abstract
The audience will discuss the current state of the effort to update the 2004 document titled "Curriculum Guidelines for Undergraduate Degree Programs in Computer Engineering," also known as CE2004. The presenters represent the ACM and the IEEE Computer Society (IEEE-CS), which are leading the effort. They will engage participants on ways of improving the body of knowledge so that the document reflects the state-of-the-art of computer engineering education and practice that is relevant for the coming decade.
Eric Durant, John Impagliazzo, Susan Conry, Robert B. Reese, Mitchell A. Thornton, Herman Lam, Victor P. Nelson
FIE6
2013 A scalable RC architecture for mean-shift clustering
abstract
The mean-shift algorithm provides a unique non-parametric and unsupervised clustering solution to image segmentation and has a proven record of very good performance for a wide variety of input images. It is essential to image processing because it provides the initial and vital steps to numerous object recognition and tracking applications. However, image segmentation using mean-shift clustering is widely recognized as one of the most compute-intensive tasks in image processing, and suffers from poor scalability with respect to the image size (N pixels) and number of iterations (k): O(kN2). Our novel approach focuses on creating a scalable hardware architecture fine-tuned to the computational requirements of the mean-shift clustering algorithm. By efficiently parallelizing and mapping the algorithm to reconfigurable hardware, we can effectively cluster hundreds of pixels independently. Each pixel can benefit from its own dedicated pipeline and can move independently of all other pixels towards its respective cluster. By using our mean-shift FPGA architecture, we achieve a speedup of three orders of magnitude with respect to our software baseline.
Stefan Craciun, Gongyu Wang, Alan D. George, Herman Lam, José C. Príncipe
ASAP4
2013 Reconfigurable computing middleware for application portability and productivity
abstract
Reconfigurable computing (RC) devices such as field-programmable gate arrays (FPGAs) offer significant advantages over fixed-logic, many-core CPU and GPU architectures, including increased performance for many computationally challenging applications, superior power efficiency, and reconfigurability. Difficulties of using FPGAs, however, has limited their acceptance in high-performance computing (HPC) and high-performance embedded computing (HPEC) applications. These difficulties stem from a lack of standards between FPGA platforms and the complexities of hardware design, and lead to higher costs and time to market over competing technologies. Differences in FPGA platform resources such as the type and number of FPGAs, memories and interconnects, as well as vendor-specific procedural APIs and hardware interfaces, inhibits application portability and code reusability. Despite efforts to reduce FPGA application design complexity through technologies such as high-level synthesis (HLS) tools, platform support and portability remains limited, and is typically left as a challenge for application developers. In this paper, we present a novel RC Middleware (RCMW), an extensible framework which enables FPGA application portability and enhances developer productivity by providing an application-centric development environment. Developers focus specifically on the optimal resources and interfaces required by their application, and RCMW handles the mapping and translation of those resources onto a target platform. We demonstrate that RCMW enables application portability over three heterogeneous platforms from two vendors, using both Xilinx and Altera FPGAs, with less than 10% performance and area overhead for several application kernels, and microbenchmarks for the common case. We present the productivity benefits of RCMW, showing that RCMW reduces required number of hardware and software driver lines of code and total development time with respect to native platform deployment methods for several application kernels.
Robert Kirchgessner, Alan D. George, Herman Lam
ASAP3
2012 VirtualRC: a virtual FPGA platform for applications and tools portability
abstract
Numerous studies have shown significant performance and power benefits of field-programmable gate arrays (FPGAs). Despite these benefits, FPGA usage has been limited by application design complexity caused largely by the lack of code and tool portability across different FPGA platforms, which prevents design reuse. This paper addresses the portability challenge by introducing a framework of architecture and middleware for virtualization of FPGA platforms, collectively named VirtualRC. Experiments show modest overhead of 5-6% in performance and 1% in area, while enabling portability of 11 applications and two high-level synthesis tools across three physical platforms.
Robert Kirchgessner, Greg Stitt, Alan D. George, Herman Lam
FPGA4
2012 RCML: An Environment for Estimation Modeling of Reconfigurable Computing Systems
abstract
Reconfigurable computing (RC) is emerging as a promising area for embedded computing, in which complex systems must balance performance, flexibility, cost, and power. The difficulty associated with RC development suggests improved strategic planning and analysis techniques can save significant development time and effort. This article presents a new abstract modeling language and environment, the RC Modeling Language (RCML), to facilitate efficient design space exploration of RC systems at the estimation modeling level, that is, before building a functional implementation. Two integrated analysis tools and case studies, one analytical and one simulative, are presented illustrating relatively accurate automated analysis of systems modeled in RCML.
Casey Reardon, Brian Holland, Alan D. George, Greg Stitt, Herman Lam
ACM Trans. Embed. Comput. Syst.5
2012 Reconfigurable Fault Tolerance: A Comprehensive Framework for Reliable and Adaptive FPGA-Based Space Computing
abstract
Commercial SRAM-based, field-programmable gate arrays (FPGAs) have the potential to provide space applications with the necessary performance to meet next-generation mission requirements. However, mitigating an FPGA’s susceptibility to single-event upset (SEU) radiation is challenging. Triple-modular redundancy (TMR) techniques are traditionally used to mitigate radiation effects, but TMR incurs substantial overheads such as increased area and power requirements. In order to reduce these overheads while still providing sufficient radiation mitigation, we propose a reconfigurable fault tolerance (RFT) framework that enables system designers to dynamically adjust a system’s level of redundancy and fault mitigation based on the varying radiation incurred at different orbital positions. This framework includes an adaptive hardware architecture that leverages FPGA reconfigurable techniques to enable significant processing to be performed efficiently and reliably when environmental factors permit. To accurately estimate upset rates, we propose an upset rate modeling tool that captures time-varying radiation effects for arbitrary satellite orbits using a collection of existing, publically available tools and models. We perform fault-injection testing on a prototype RFT platform to validate the RFT architecture and RFT performability models. We combine our RFT hardware architecture and the modeled upset rates using phased-mission Markov modeling to estimate performability gains achievable using our framework for two case-study orbits.
Adam Jacobs, Grzegorz Cieslewski, Alan D. George, Ann Gordon-Ross, Herman Lam
ACM Trans. Reconfigurable Technol. Syst.5
2011 SHMEM+: A multilevel-PGAS programming model for reconfigurable supercomputing
abstract
Reconfigurable Computing (RC) systems based on FPGAs are becoming an increasingly attractive solution to building parallel systems of the future. Applications targeting such systems have demonstrated superior performance and reduced energy consumption versus their traditional counterparts based on microprocessors. However, most of such work has been limited to small system sizes. Unlike traditional HPC systems, lack of integrated, system-wide, parallel-programming models and languages presents a significant design challenge for creating applications targeting scalable, reconfigurable HPC systems. In this article, we extend the traditional Partitioned Global Address Space (PGAS) model to provide a multilevel integration of memory, which simplifies development of parallel applications for such systems and improves developer productivity. The new multilevel-PGAS programming model captures the unique characteristics of reconfigurable HPC systems, such as the existence of multiple levels of memory hierarchy and heterogeneous computation resources. Based on this model, we extend and adapt the SHMEM communication library to become what we call SHMEM+, the first known SHMEM library enabling coordination between FPGAs and CPUs in a reconfigurable, heterogeneous HPC system. Applications designed with SHMEM+ yield improved developer productivity compared to current methods of multidevice RC design and exhibit a high degree of portability. In addition, our design of SHMEM+ library itself is portable and provides peak communication bandwidth comparable to vendor-proprietary versions of SHMEM. Application case studies are presented to illustrate the advantages of SHMEM+.
Vikas Aggarwal, Alan D. George, Changil Yoon, Kishore Yalamanchili, Herman Lam
ACM Trans. Reconfigurable Technol. Syst.5
2011 An analytical model for multilevel performance prediction of Multi-FPGA systems
abstract
Power limitations in semiconductors have made explicitly parallel device architectures such as Field-Programmable Gate Arrays (FPGAs) increasingly attractive for use in scalable systems. However, mitigating the significant cost of FPGA development requires efficient design-space exploration to plan and evaluate a range of potential algorithm and platform choices prior to implementation. The authors propose the RC Amenability Test for Scalable Systems (RATSS), an analytical model which enables straightforward, fast, and reasonably accurate performance prediction prior to implementation by extending current modeling concepts to multi-FPGA designs. RATSS provides a comprehensive strategic model to evaluate applications based on the computation and communication requirements of the algorithm and capabilities of the FPGA platform. The RATSS model targets data-parallel applications on current scalable FPGA systems. Three case studies with RATSS demonstrate nearly 90% prediction accuracy as compared to corresponding implementations.
Brian Holland, Alan D. George, Herman Lam, Melissa C. Smith
ACM Trans. Reconfigurable Technol. Syst.3
2010 Characterization of Fixed and Reconfigurable Multi-Core Devices for Application Acceleration
abstract
As on-chip transistor counts increase, the computing landscape has shifted to multi- and many-core devices. Computational accelerators have adopted this trend by incorporating both fixed and reconfigurable many-core and multi-core devices. As more, disparate devices enter the market, there is an increasing need for concepts, terminology, and classification techniques to understand the device tradeoffs. Additionally, computational performance, memory performance, and power metrics are needed to objectively compare devices. These metrics will assist application scientists in selecting the appropriate device early in the development cycle. This article presents a hierarchical taxonomy of computing devices, concepts and terminology describing reconfigurability, and computational density and internal memory bandwidth metrics to compare devices.
Chris Massie, Alan D. George, Justin Richardson, Kunal Gosrani, Herman Lam
ACM Trans. Reconfigurable Technol. Syst.6
2008 A Rule-Based Approach for Availability of Web Service
abstract
Sustainable success of service oriented applications relies on capabilities to manage possible service failures. To substitute a failed service with some other equivalent service is unavoidable in recovering a suspended application due to failure of a constituent service. In this paper, we report a rule based approach to Web service substitution in order to secure availability of services. Availability provides delivery assurance for each Web service so that Simple Object Access Protocol (SOAP) messages cannot be lost undetectably, especially in a Web service composition. The rules are written in Semantic Web Rule Language. The rules are a formal representation of a categorization-based scheme to identify exchangeable Web services. This scheme not only tackles the issue of heterogeneity of domain ontology in describing the Web services, it also adapts itself by learning newly discovered ontology instances. A technical framework of Web service substitution using rule based deduction is demonstrated. Experiments on service substitution based on the proposed framework achieve a best precision of 85%.
Qianhui Althea Liang, Herman Lam, Lalita Narupiyakul, Patrick C. K. Hung
ICWS2
2007 Science gateways made easy: the In-VIGO approach
abstract
Abstract Science gateways require the easy enabling of legacy scientific applications on computing Grids and the generation of user‐friendly interfaces that hide the complexity of the Grid from the user. This paper presents the In‐VIGO approach to the creation and management of science gateways. First, we discuss the virtualization of machines, networks and data to facilitate the dynamic creation of secure execution environments that meet application requirements. Then we discuss the virtualization of applications, i.e. the execution on shared resources of multiple isolated application instances with customized behavior, in the context of In‐VIGO. A Virtual Application Service (VAS) architecture for automatically generating, customizing, deploying, and using virtual applications as Grid services is then described. Starting with a grammar‐based description of the command‐line syntax, the automated process generates the VAS description and the VAS implementation (code for application encapsulation and data binding) that is deployed and made available through a Web interface. A VAS can be customized on a per‐user basis by restricting the capabilities of the original application or by adding to it features such as parameter sweeping. This is a scalable approach to the integration of scientific applications as services into Grids and can be applied to any tool with an arbitrarily complex command‐line syntax. Copyright © 2006 John Wiley & Sons, Ltd.
Andréa M. Matsunaga, Maurício O. Tsugawa, Sumalatha Adabala, Renato J. O. Figueiredo, Herman Lam, José A. B. Fortes
Concurr. Comput. Pract. Exp.5
2005 Application Modeling and Representation for Automatic Grid-Enabling of Legacy Applications
abstract
To best exploit the potential of the grid, it is necessary to "grid-enable" legacy applications that were not originally developed to run on a grid. In this paper, a generic framework to grid-enable legacy applications is presented. In particular, this paper focuses on a general approach to model and represent an application in a way that is supportive of the key properties of the grid-enabling framework: generality, automatic generation and integration of grid applications, plug-and-play deployment, and interoperability. Based on the model, a configuration language is developed to describe a wide range of command-line applications and their execution requirements. It is easy to use and does not require the application enabler to know the details of the grid middleware. A case study is presented to illustrate the automated grid-enabling of a legacy application in In-VIGO (in-virtual information grid organizations), a grid computing infrastructure that makes extensive use of visualization technology
Andréa M. Matsunaga, Vivekananthan Sanjeepan, Herman Lam, José A. B. Fortes
e-Science4
2005 On the Use of Virtualization and Service Technologies to Enable Grid-Computing
Andréa M. Matsunaga, Maurício O. Tsugawa, Ming Zhao 0002, Vivekananthan Sanjeepan, Sumalatha Adabala, Renato J. O. Figueiredo, Herman Lam, José A. B. Fortes
Euro-Par8
2005 In-VIGO virtual networks and virtual application services: automated grid-enabling and deployment of applications
abstract
This poster briefly introduces two resource-virtualization techniques needed for the creation of virtual(ized) grids: virtual networks and virtual application services. The former provides bidirectional network connectivity even in the presence of firewalls, network address translation gateways and proxies by creating virtual routers and virtual IP space. The later allows automated creation and deployment of legacy applications into grids by generating a virtual application service that allows the execution in shared resources of multiple isolated application instances with customized behavior.
Maurício O. Tsugawa, Andréa M. Matsunaga, Vivekananthan Sanjeepan, Herman Lam, Renato J. O. Figueiredo, José A. B. Fortes
HPDC5
2005 A Service-Oriented, Scalable Approach to Grid-Enabling of Legacy Scientific Applications
abstract
This paper describes a scalable approach to the enabling of legacy scientific applications on computing grids using a service-oriented architecture. In the context of this paper grid-enabling means turning an existing application, installed on a grid resource, into a service and generating the application-specific user interfaces to use that application through a Web portal. Scalability is achieved by providing a common abstraction for a category of applications and providing a "generic" application service to wrap those applications as services. The focus of this paper's approach is on grid-enabling "command-oriented" scientific applications. The novel aspect of the approach is that the entire process -from turning an application into a service to the user-interface generation for that application - is done automatically, without requiring coding or grid-system downtime. Portlet technology is used to dynamically generate application-specific interfaces. Further, the approach makes it possible to customize the applications for different user groups by way of simplifying, restricting or composing the functionalities of applications. The approach is useful for building grid portals on which a large number of applications need to be dynamically enabled.
Vivekananthan Sanjeepan, Andréa M. Matsunaga, Herman Lam, José A. B. Fortes
ICWS4
2004 Constraint Specification and Processing in Web Services Publication and Discovery
abstract
Much effort is being made by the IT industry towards the establishment of a Web services infrastructure and the refinement of its component technologies to enable the sharing of heterogeneous application resources. Traditional roles of the service provider, service requestor and service broker and their interactions are now being improved upon to enable more effective services. The implementation of the Web service broker is currently limited to being an interface to the service repository for service registration, browsing and/or programmatic access. In this work, we have extended the functionality of the Web services broker to include constraint specification and processing, which enables the broker to find a good match between a service provider's capabilities and a service requestor's requirements. This paper presents the extension made to the Web Services Description Language to include constraint specifications in service descriptions and requests, the architecture of a constraint-based broker, the constraint matching technique, some implementation details, and preliminary evaluation results.
Seema Degwekar, Stanley Y. W. Su, Herman Lam
ICWS3
2004 Adaptive Grid Service Flow Management: Framework and Model
abstract
Grid computing provides the basic software infrastructure for integrating geographically distributed resources and services through standardized grid services. One of the key challenges to enable the broader use of grid services beyond the domain of scientific computing is the ability to perform complex tasks that require the modeling and coordination of the enactment of a number of distributed grid services. Workflow technology is a good candidate for supporting grid service flow. However, traditional workflow is static, thus unable to exploit the dynamic information available in the grid and respond to the dynamic nature of the grid. In this paper, we present an adaptive framework that provides adaptive management of grid service flows. The framework is based on an adaptive grid service flow model and is supported by an event-trigger-rule (ETR) technology that will be used to trigger rules in a distributed fashion to adapt a grid service flow to the dynamic grid environment and the changing requirements of a grid application.
Yu Long 0002, Herman Lam, Stanley Y. W. Su
ICWS2
2004 Event and rule services for achieving a Web-based knowledge network
Minsoo Lee, Stanley Y. W. Su, Herman Lam
Knowl. Based Syst.3
2003 Integration of Business Event and Rule Management with the Web Services Model
Karthik Nagarajan, Herman Lam, Stanley Y. W. Su
ICWS2
2003 A Cost-Benefit Evaluation Server for decision support in e-business
Youzhong Liu, Fahong Yu, Stanley Y. W. Su, Herman Lam
Decis. Support Syst.4
2001 An Information Infrastructure and E-Services for Supporting Internet-Based Scalable E-Business Enterprises
abstract
The paper presents an information infrastructure for supporting Internet-based scalable e-business enterprises (ISEE). The information infrastructure is formed by a network of ISEE hubs, each of which has a number of replicable e-business servers providing various e-services to individuals and businesses. The servers are the implementations of a number of core technologies developed to facilitate collaborative e-business, including business event and rule management, active distributed objects, constraint satisfaction processing, and cost benefit analysis and selection. Using the e-services provided by these servers, other e-services such as constraint-based brokering, supplier selection, active business process management, and automated negotiation can be developed. Supply chain management is used as an example of collaborative e-business in the descriptions of these technologies and their implementations.
Stanley Y. W. Su, Herman Lam, Minsoo Lee, Sherman X. Bai, Zuo-Jun Max Shen
EDOC2
2001 Event and Rule Services for Achieving a Web-Based Knowledge Network
Minsoo Lee, Stanley Y. W. Su, Herman Lam
Web Intelligence3
2001 An Internet-based negotiation server for e-commerce
Stanley Y. W. Su, Chunbo Huang, Joachim Hammer, Haifei Li 0002, Liu Wang 0003, Youzhong Liu, Charnyote Pluempitiwiriyawej, Minsoo Lee, Herman Lam
VLDB J.10
2001 A Web-Based Knowledge Network for Supporting Emerging Internet Applications
Minsoo Lee, Stanley Y. W. Su, Herman Lam
World Wide Web3
2000 Distributed and Concurrent Processing of Business Object Documents in Support of e-Enterprise Integration
abstract
The Internet and distributed object technologies have made it possible for different business enterprises to draw upon the best of their resources for conducting joint business as a virtual e-enterprise (VEE). To enable virtual e-enterprises, the integration of legacy applications and the modeling and enactment of concurrent business processes are necessary. The authors combine the features of the messaging approach and the distributed object approach to system integration. Business Object Documents (BOD) are used for transmitting business operations and data among application systems. Message transmission is supported by two underlying communication infrastructures: CORBA and Java RMI. The separation of messaging from communication infrastructure allows the underlying infrastructure to be changed without impacting application systems. Also, business processes are modeled as sequences or network structures of BOD transmissions. The process models are replicated at all sites and used by an extended information infrastructure to enable distributed, concurrent enactment of processes.
Stanley Y. W. Su, Youzhong Liu, Minsoo Lee, Herman Lam
EDOC5
1996 NCL: A Common Language for Achieving Rule-Based Interoperability Among Heterogeneous Systems
Stanley Y. W. Su, Herman Lam, Tsae-Feng Yu, Javier A. Arroyo-Figueroa, Zhidong Yang, Sooha Lee
J. Intell. Inf. Syst.2
1996 The Design and Implementation of K: A High-Level Knowledge-Base Programming Language of OSAM*.KBMS
Yuh-Ming Shyy, Javier Arroyo, Stanley Y. W. Su, Herman Lam
VLDB J.4
1995 An Extensible Knowledge Base Management System for Supporting Rule-based Interoperability among Heterogeneous Systems
abstract
The main objective of a virtual enterprise (VE) is to allow a number of organizations to rapidly develop a working environment to manage a collection of resources contributed by the organizations toward the attainment of some common goals. One of the key requirements of a virtual enterprise is to develop an information infrastructure to support the interoperability of distributed and heterogeneous systems for controlling and conducting the business of the virtual enterprise. In order to achieve the objective and to meet this requirement, it is necessary to model all things of interest to a virtual enterprise such as data, human and hardware resources, organizational structures, business constraints, production processes, and activities in work management. Additionally, a system is needed to manage the meta-information and the shared data and to provide both build-time and run-time services to the heterogeneous systems to achieve their interoperability. In this paper, we describe the mo...
Stanley Y. W. Su, Herman Lam, Javier A. Arroyo-Figueroa, Tsae-Feng Yu, Zhidong Yang
CIKM2
1995 Algorithms for Asynchronous Parallel Processing of Object-Oriented Databases
abstract
Management of large quantities of complex data is essential in many advanced application areas. Object-oriented (OO) database management system have been developed to effectively model and process the complex domain knowledge. They have been shown to outperform some existing relational systems. The existing implementations of OO database management systems attempt to improve the efficiency of OO queries by explicitly capturing the relationships among objects. However, the execution of complex queries involving the retrieval of objects from many classes and relationships among them causes the existing system to operate inefficiently. In this paper, we present parallel algorithms for the processing of queries against a large OO database. The algorithms are based on a closed model of query processing pattern-based access instead of the conventional value-based access. During processing, the algorithms avoid the execution of time-consuming join operations by making use of the explicitly stored object associations. Generation of large quantities of temporary data is avoided by marking objects using their identifiers and by employing a two-phase query processing strategy. A query is processed by concurrent multiple waves, thereby improving parallelism avoiding the complexities introduced in their sequential implementation. The correctness and the performance of the parallel algorithms have been tested and analyzed by running parallel programs on a 32-node transputer based parallel machine designed and developed at the IBM Research Center at Yorktown Heights, New York. Benchmark queries of different semantic complexities are generated, and their performance is analyzed for various data and query parameters.>
Arun K. Thakore, Stanley Y. W. Su, Herman Lam
IEEE Trans. Knowl. Data Eng.3
1993 OSAM*KBMS: An Object-Oriented Knowledge Base Management System for Supporting Advanced Applications
Stanley Y. W. Su, Herman Lam, Srinivasa Eddula, Javier Arroyo, Neeta Prasad, Ronghao Zhuang
SIGMOD Conference2
1993 An Object Flow Computer for Database Applications: Design and Performance Evaluation
Chiang Lee, Herman Lam, Stanley Y. W. Su
J. Parallel Distributed Comput.2
1993 Association Algebra: A Mathematical Foundation for Object-Oriented Databases
abstract
The application of the object-oriented (O-O) paradigm in the database management field has gained much attention in recent years. Several experimental and commercial O-O database management systems have become available. However, the existing O-O DBMSs still lack a solid mathematical foundation for the manipulation of O-O databases, the optimization of queries, and the design and selection of storage structures for supporting O-O database manipulations. This paper presents an association algebra (A-algebra) to serve as a mathematical foundation for processing O-O databases, which is analogous to the relational algebra used for processing relational databases. In this algebra, objects and their associations in an O-O database are uniformly represented by association patterns which are manipulated by a number of operators to produce other association patterns. Different from the relational algebra, in which set operations operate on relations with union-compatible structures, the A-algebra operators can operate on association patterns of homogeneous and heterogeneous structures. Different from the traditional record-based relational processing, the A-algebra allows very complex patterns of object associations to be directly manipulated. The pattern-based query formulation and the A-algebra operators are described. Some mathematical properties of the algebraic operators are presented together with their application in query decomposition and optimization. The completeness of the A-algebra is also defined and proven. The A-algebra has been used as the basis for the design and implementation of an object-oriented query language, OQL, which is the query language used in a prototype Knowledge Base Management System OSAM*.KBMS.>
Stanley Y. W. Su, Mingsen Guo, Herman Lam
IEEE Trans. Knowl. Data Eng.3
1991 An Association Algebra For Processing Object-Oriented Databases
abstract
An association algebra (A-algebra) is presented for manipulating object-oriented (O-O) databases which is analogous to the relational algebra for relational databases. In this algebra, objects and their associations in an O-O database are uniformly represented by association patterns and are manipulated by a number of operators. These operators are defined to operate on association patterns of both heterogeneous and homogeneous structures. Very complex structures (e.g. network structures of object associations across several classes) can be directly manipulated by these operators. Therefore, the association algebra has greater expressive powers than the relational algebra which manipulates on relations of compatible structures. Some mathematical properties of these operators are described together with their application in query decomposition and optimization. The algebra has been used as the basis for the design and implementation of an O-O query language called OQL and a knowledge rule specification language.>
Mingsen Guo, Stanley Y. W. Su, Herman Lam
ICDE3
1991 An Extensible Kernel Object Management System
Rahim Yaseen, Stanley Y. W. Su, Herman Lam
OOPSLA3
1991 A Special Function Unit for Database Operations (SFU-DB): Design and Performance Evaluation
abstract
The design and analysis of a special function unit for database operations (SFU-DB) that uses a novel hardware sorting module, the automatic retrieval memory (ARM), are described. The SFU-DB is a functionally independent unit that efficiently performs certain nonnumeric operations. It can function as a coprocessor for a host CPU or as a special processing unit in a highly parallel processing system. The ARM implements in hardware a true distribution-based sort algorithm that requires no comparison operations. Without performing any comparison, the SFU-DB avoids the lower bound constraint on comparison-based sorting algorithms and achieves, for the worst case, a complexity of O(n) for both execution time and main memory size. Using the fundamental sort algorithm with slight modifications. the SFU-DB also uses the ARM as an engine for other primitive database operations such as relational join, elimination of duplicates, set union, set intersection, and set difference, also with complexity of O(n). The SFU-DB/ARM architecture is rather simple and requires only a modest amount of specialized hardware. The specialized hardware has been designed and simulated for fabrication using CMOS gate arrays, and the remainder of the SFU-DB has been simulated in software using Turbo Pascal running on an IBM-PC.>
Herman Lam, Chiang Lee, Stanley Y. W. Su
IEEE Trans. Computers1
1990 A graphical interface for an object-oriented query language
abstract
A graphical user interface for an object-oriented query language, GOQL, is presented, GOQL is a part of a prototype knowledge base management system which is based on an object-oriented semantic association model, OSAM. GOQL consists of a graphical browser and a graphical querying module. The browser allows a user to browse through a complex knowledge-base schema graphically and prune it into a desired level of abstraction and details before the querying process. In the querying module, there are two modes, OQL and graphical OQL. The OQL mode is provided for knowledgeable users to directly type in the OQL command. In the graphical OQL mode, the user is guided through the formation of the query. It is noted that the object-oriented nature and the increased semantics of the underlying model and query language pose new challenges in user interface design due to their added complexity. On the other hand, these features also provide more information to the system in order to make the user interface more intelligent.>
Herman Lam, H. More Chen, Frederick S. Ty, Jiwen Qiu, Stanley Y. W. Su
COMPSAC1
1990 Heuristic algorithms for path determination in a semantic network
abstract
The authors present two heuristic algorithms for determining traversal paths in a semantic network which models an object-oriented database. The first algorithm is an extension of Dijkstra's shortest path algorithm, and it identifies the most likely interpretation of an incomplete specified query. The second algorithm finds all possible interpretations of the query and ranks them in order of likely interpretations. Cost assignment for different paths is based on a set of heuristic rules which allows costs to be dynamically determined during a path traversal. The algorithms have been implemented and are in use in a graphics interface developed for an object-oriented knowledge base management system.>
Stanley Y. W. Su, Shirish Puranik, Herman Lam
COMPSAC3
1990 Conceptual Design for Non-Database Experts with an Interactive Schema Tailoring Tool
Shamkant B. Navathe, Seong Geum, Dinesh K. Desai, Herman Lam
ER4
1990 A Rule-based Language for Deductive Object-Oriented Databases
abstract
A deductive rule-based language for object-oriented databases is presented. A deductive rule in this language derives new patterns of associations among objects of some selected classes if these objects fall in certain 'base' on other derived patterns. The patterns of object associations derived by a rule are held in a subdatabase whose intention consists of some selected classes and their associations. In other words, the structure of a derived subdatabase is represented using the structural constructs provided by the object-oriented data model and hence can be uniformly operated on by other rules to further derive new subdatabases. Therefore, the world of subdatabases is closed under this rule-based language.>
Abdallah M. Alashqur, Stanley Y. W. Su, Herman Lam
ICDE3
1990 Asynchronous Parallel Processing of Object Bases Using Multiple Wavefronts
Arun K. Thakore, Stanley Y. W. Su, Herman Lam, Dennis G. Shea
ICPP (1)3
1989 Integrating the concepts and techniques of semantic modeling and the object-oriented paradigm
abstract
The object orientation of a semantic association model (OSAM) is presented. It integrates the concepts and techniques of semantic modeling and those introduced by the object-oriented paradigm. Unlike conventional data models such as the relational model, the object orientation of OSAM allows the user to model an application in terms of complex objects, classes and their associations, instead of tuples (or records) and relations (or record types). The primitives (objects, class, instance and link) and the perspectives (class and object) of an OSAM database are described. Key differences between OSAM and a conventional object-oriented model are discussed. The features of object orientation and explicit definition of semantic associations among objects allow the database of an application domain to be modeled, accessed and manipulated at a higher conceptual level and thus simplify the tasks of the users in the development of their applications.>
Herman Lam, Stanley Y. W. Su, Abdallah M. Alashqur
COMPSAC1
1989 OQL: A Query Language for Manipulating Object-oriented Databases
Abdallah M. Alashqur, Stanley Y. W. Su, Herman Lam
VLDB3
1988 IMDAS - An Integrated Manufacturing Data Administration System
Vishu Krishnamurthy, Stanley Y. W. Su, Herman Lam, Mary Mitchell, Edward Barkmeyer
Data Knowl. Eng.3
1988 A Physical Database Design Evaluation System for CODASYL Databases
abstract
An interactive design tool for designing CODASYL databases is described. The system is composed of three main modules: a user interface, a transaction analyzer, and a core module. The user interface allows a designer to enter interactively information concerning a database design which is to be evaluated. The transaction analyzer allows the designer to specify the processing requirements in terms of typical logical transactions to be executed against the database and translates these logical transaction into physical transaction which access and manipulate the physical databases. The core module is the implementation of a set of analytical models and cost formulas developed for the manipulation of indexed sequential and hash-based files and CODASYL sets. These models and formulas account for the situation in which occurrences of multiple record types are stored in the same area. Also presented are the results of a series of experiments in which key design parameters are varied. The system is implemented in UCSD Pascal running on IBM PCs.>
Herman Lam, Stanley Y. W. Su, Nageshwar R. Koganti
IEEE Trans. Software Eng.1
1987 A Special Function Unit for Database Operations Within a Data-Control Flow System
Herman Lam, Stanley Y. W. Su, F. L. C. Seeger, William R. Eisenstadt
ICPP1
1986 A Special-Function Unit for Sorting and Sort-Based Database Operations
abstract
Achieving efficiency in database management functions is a fundamental problem underlying many computer applications. Efficiency is difficult to achieve using the traditional general-purpose von Neumann processors. Recent advances in microelectronic technologies have prompted many new research activities in the design, implementation, and application of database machines which are tailored for processing database management functions. To build an efficient system, the software algorithms designed for this type of system need to be tailored to take advantage of the hardware characteristics of these machines. Furthermore, special hardware units should be used, if they are cost- effective, to execute or to assist the execution of these software algorithms.
Louiqa Raschid, Tinghe Fei, Herman Lam, Stanley Y. W. Su
IEEE Trans. Computers3
1981 Transformation of Data Traversals and Operations in Application Programs to Account for Semantic Changes of Databases
abstract
This paper addresses the problem of application program conversion to account for changes in database semantics that result in changes in the schema and database contents. With the observation that the existing data models can be viewed as alternative ways of modeling the same database semantics, a methodology of application program analysis and conversion based on an existing-DBMS-model-and schema-independent representation of both the database and programs is presented. In this methodology, the source and target databases are described in terms of the association types of a semantic association model. The structural properties, the integrity constraints, and the operational characteristics (storage operation behaviors) of the association types are more explicitly defined to reveal the semantics that is generally hidden in application programs. The explicit descriptions of the source and target databases are used as the basis for program analysis and conversion. Application programs are described in terms of a small number of “access patterns” which define the data traversals and operations of the programs. In addition to the methodology, this paper (1) describes a model of a generalized application program conversion system that serves as a framework for research, (2) presents an analysis of access patterns that serve as the primitives for program description, (3) delineates some meaningful semantic changes to databases and their corresponding transformation rules for program conversion, (4) illustrates the application of these rules to two different approaches to program conversion problems, and (5) reports on the development effort undertaken at the University of Florida.
Stanley Y. W. Su, Herman Lam, Der Her Lo
ACM Trans. Database Syst.2