Charles C. Weems

dblp:w/CharlesCWeems · also Charles C. Weems Jr. · DBLP profile ↗
← Back
63ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-8535-0258ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 2 first-authorArtificial intelligence and machine learning · 11 · 4 first-authorHuman-computer interaction and ubiquitous computing · 11 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 4Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Modernizing the Introductory Computing Sequence: Integrating Parallel and Distributed Computing in CS1 and CS2
abstract
The rapid evolution of computing demands curricula that reflect modern practices, yet many CS1 and CS2 courses continue to emphasize only sequential programming. This NSF-funded project addresses that gap by designing and disseminating exemplar CS1 and CS2 courses that integrate parallel, distributed, and event-driven computing as core concepts. The materials include unplugged activities and programming labs for both C++ and Java. To ensure broad applicability and adoption, development occurred in collaboration with instructors from six diverse institutions who are now implementing the materials. Evaluation includes surveys, assignment-specific instruments, and cross-team analysis. This poster presents the project’s vision, methods, and resources, highlighting how others can adopt and adapt them to teach modern computing.
April Renee Crockett, David P. Bunde, Gerald C. Gannod, Sushil K. Prasad, Jaime Spacco, Alan Sussman, Neena Thota, Charles C. Weems, Ramachandran Vaidyanathan
SIGCSE (2)8
2026 Envisioning CS1 and CS2: The Future of Introductory Problem Solving and Programming
abstract
Computer Science education, and all education for that matter, is being disrupted by Generative AI. While there have been few truly transformational technologies similar to AI, other incremental but impactful advances have helped shape the computing ecosystem. Other recent examples include the transition to multicore systems (requiring the promotion of parallel computing from an elective topic), the shift to graphical interfaces (raising expectations for assignments and motivating the creation of Media Computation), and the emergence of object-oriented programming. In this Birds of a Feather Session, we ask the question ''How might we redesign our CS1 and CS2 courses to better prepare students for emerging and future computing paradigms while maintaining strong foundations in problem solving, programming, and computational thinking?'' Using collaborative brainstorming techniques, participants will create a list of potential future paradigms (either disruptive or incremental) that are relevant to CS1/CS2, and develop proposed roadmaps that identify how those paradigms can be leveraged as contexts for teaching the existing CS1 and CS2 courses within the CS2023 curriculum.
Gerald C. Gannod, David P. Bunde, April Renee Crockett, Alan Sussman, Sushil K. Prasad, Charles C. Weems, Ramachandran Vaidyanathan, Suzanne Matthews, Jaime Spacco
SIGCSE (2)6
2026 Modernizing the CS Introductory Sequence with Parallel and Distributed Computing (and some AI)
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching a 20th century model of algorithmic problem solving, where sequence, branch, and loop are the only organizing principles needed for algorithms. We invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as a GPU in many cases. Most of their favorite applications use multiple cores and distributed resources. Often concurrency offers simpler solutions than sequential approaches. In this tutorial we overview key PDC concepts and provide examples of how they may naturally be incorporated in early computing classes. We lead participants through plugged and unplugged curriculum modules that have been successfully integrated and tested in existing computing classes at multiple institutions. We also discuss recent efforts at integrating AI methods, including LLMs, into early classes. In addition, we highlight other CDER activities for integration of PDC and AI into undergraduate computing curricula. Additional Information: No equipment or prior PDC experience is required, although a laptop that can run C++, Java and Python is recommended for following along with some code examples if desired.
Charles C. Weems, April Renee Crockett, David P. Bunde, Alan Sussman, Ramachandran Vaidyanathan, Sushil K. Prasad, Gerald C. Gannod, Jaime Spacco
SIGCSE (2)1
2024 WIP: Updating CS1 to a 21st-Century Model of Computing
abstract
This work in progress innovative practice paper documents ways in which current introductory computing courses are designed for an earlier generation of computers. We describe our plans for updating these courses for modern systems and programming practices and share details of the development of exemplar courses that will be adoptable by diverse institutions and programs teaching introductory programming courses.
David P. Bunde, April Renee Crockett, Gerald C. Gannod, Jaime Spacco, Neena Thota, Charles C. Weems
FIE6
2024 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, so it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning of their computing education. With all computing devices that students use currently having multiple cores as well as a GPU in many cases, many students' favorite applications use multiple cores and/or distributed processors. However, we are still teaching them to solve problems using only sequential thinking. Why?
Alan Sussman, Sushil K. Prasad, Charles C. Weems, Sheikh K. Ghafoor, Ramachandran Vaidyanathan
SIGCSE (2)3
2023 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching to a 20th century model of algorithmic problem solving. Sequence, branch, and loop are taught in our early courses as the only organizing principles needed for algorithms, and we invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as a GPU in many cases. Most of their favorite applications use multiple cores and numbers of distributed processors. Often concurrency offers simpler solutions than sequential approaches. Industry is desperate for software engineers who think naturally in terms of exploiting these capabilities, rather than seeing them as an exotic upper-level topic that gets layered over a sequential solution. However, we are still teaching students to solve problems using sequential thinking. In this workshop we overview key PDC concepts and provide examples of how they may naturally be incorporated in early computing classes. We will introduce plugged and unplugged curriculum modules that have been successfully integrated in existing computing classes at multiple institutions. We will highlight the upcoming summer training workshop, for which we have funding to support attendance, as well as other CDER (Center for Parallel and Distributed Computing Curriculum Development and Educational Resources) activities.
Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman, Ramachandran Vaidyanathan, Sushil K. Prasad
SIGCSE (2)2
2023 NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing for Undergraduates - Version II - Big Data, Energy, and Distributed Computing
abstract
This special session will report on the updated NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing released in Nov 2020 by the Center for Parallel and Distributed Computing Curriculum Development and Educational Resources (CDER). The purpose of the special session is to obtain SIGCSE community feedback on this curriculum in a highly interactive manner employing the hybrid modality and supported by a full-time CDER booth for the duration of SIGCSE. In this era of big data, cloud, and multi- and many-core systems, it is essential that the computer science (CS) and computer engineering (CE) graduates have basic skills in parallel and distributed computing (PDC). The topics are primarily organized into the areas of architecture, programming, and algorithms topics. A set of pervasive concepts that percolate across area boundaries are also identified. Version 1 of this curriculum was released in December 2012. That curriculum guideline has over 140 early adopter institutions worldwide and has been incorporated into the 2013 ACM/IEEE Computer Science curricula. This Version-II represents a major revision. The updates have focused on enhancing coverage related to the topical aspects of Big Data, Energy, and Distributed Computing.
Sushil K. Prasad, Charles C. Weems, Alan Sussman, Trilce Estrada, Ramachandran Vaidyanathan, Sheikh K. Ghafoor, Krishna Kant 0001, Craig B. Stunkel
SIGCSE (2)2
2022 Integrating Parallel and Distributed Computing in Early CS Courses
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching to a 20th century model of algorithmic problem solving. Sequence, branch, and loop are taught in our early courses as the only organizing principles needed for algorithms, and we invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as GPU in many cases. Most of their favorite applications use multiple cores and numbers of distributed processors. Often concurrency offers simpler solutions than sequential approaches. ACM and ABET have recommended including PDC in the undergraduate CS curriculum. However, we are still teaching them to solve problems using sequential thinking. In this workshop we overview the key PDC concepts and provide examples of how they may naturally be incorporated in early CS classes. We will introduce plugged and unplugged curriculum modules that have been successfully integrated in existing CS classes at multiple institutions. We will highlight the upcoming summer training that we are organizing, for which we have funding to support attendance.
Sheikh K. Ghafoor, Sushil K. Prasad, Charles C. Weems
SIGCSE (2)3
2021 DPF-ECC: A Framework for Efficient ECC With Double Precision Floating-Point Computing Power
abstract
Used ubiquitously in a huge amount of security protocols or applications, elliptic curve cryptography (ECC) is one of the most important cryptographic primitives, featuring efficiency and short key size compared with other public-key cryptosystems such as DSA and RSA. However, as a computation-intensive public-key cryptographic primitive, ECC arithmetic is still the bottleneck that restrains the overall performance of the end applications. In this paper, instead of the conventional and straightforward integer-based methods, we present a general framework to accelerate ECC schemes over prime field, called DPF-ECC, that deeply exploits double precision floating-point (DPF) computing power. The DPF-ECC framework finely manages each bit of the DPF numbers and minimizes the overhead brought by additional data format conversion, by making use of the DPF representation, the rounding operations, and fused multiply-add instruction supported by the IEEE 754 floating point standard. We also conduct two comprehensive case studies on Crandall primes and Solinas primes to demonstrate how the DPF-ECC framework is applied to the prevailing ECC schemes. To evaluate the proposed DPF-ECC framework in the real world, leveraging the floating-point computing power of GPUs, we implement Curve25519/448 and Edwards25519/448, the popular ECC schemes widely used in TLS 1.3, SSH, etc. The experimental result in Tesla P100 achieves a record-setting performance that outperforms the existing fastest integer work with 2x to 3x throughput. With dependency only on the very commonly supported IEEE 754 floating point standard, DPF-ECC framework can be a very competent and promising candidate for ECC implementation in most of general-purpose platforms.
Fangyu Zheng, Rong Wei, Jiankuo Dong, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IEEE Trans. Inf. Forensics Secur.8
2020 DPF-ECC: Accelerating Elliptic Curve Cryptography with Floating-Point Computing Power of GPUs
abstract
Driven by artificial intelligence (AI) and computer vision industries, Graphics Processing Units (GPUs) are now rapidly achieving extraordinary computing power. In particular, the floating-point computing power, which is heavily relied on by graphics rendering and AI computation workload, is developing much faster in GPUs. Meanwhile, in many fields such as ecommerce and online finance, the demand for cryptographic operations for secure communications and authentication is also expanding.In this contribution, targeting the important cryptographic primitives widely used in TLS 1.3, etc., we implement Curve25519 and Edwards25519 with GPUs' floating-point computing power, where various performance optimization methods are customized for the target platform, including novel big-number representations combined with a new floating-point-based computing algorithm, efficient merged reduction strategies, and curve-level acceleration. This paper reports record-setting performance for the elliptic-curve method: on TITAN V, we respectively achieve 7.21 and 77.30 million operations per second of unknown and known point multiplication of Edwards25519, and 13.55 million operations per second of point multiplication of Curve25519. To the best of our knowledge, this contribution is the first to show that floating-point-based ECC implementations can outperform the integer-based ones by a huge margin. The experimental result in Tesla P100 achieves over double performance of the existing fastest integer work on the same platform, and the result in TITAN V sets a record for the throughput which is 4.43 times better than the second.
Fangyu Zheng, Niall Emmart, Jiankuo Dong, Jingqiang Lin 0001, Charles C. Weems
IPDPS6
2019 Modernizing Early CS Courses with Parallel and Distributed Computing
abstract
Parallel and distributed computing (PDC) is now a pervasive aspect of deployed systems, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Our students all have multicore laptops. Most of their favorite applications use vast numbers of distributed processors. Why are we still teaching them to solve problems using only sequential thinking? Come to this workshop to see how easy it is to open their eyes to exploiting concurrency in problem solving, starting in their earliest courses. You'll hear about and experience some unplugged activities, learn how to help students recognize examples of concurrency in the world around them, see how event driven user interfaces can easily exemplify issues related to multithreading, and how freely available libraries can be used to naturally exploit parallelism in working with large data structures. We will also highlight the two summer training programs that we are organizing, for which we have funding to support attendance by instructors. Having a laptop that can run Java and C++ will allow you to follow along with some code examples, but isn't necessary.
Sushil K. Prasad, Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman
SIGCSE3
2018 Faster Modular Exponentiation Using Double Precision Floating Point Arithmetic on the GPU
abstract
This paper presents a new approach to integer multiple precision (MP) modular exponentiation, using double-precision floating point (DPF) operations, that is suitable for GPU implementation. We show speedups ranging from 20 % to 34 % over the best prior GPU times for sizes corresponding to common RSA cryptographic operations (2048 to 4096 bits). Three techniques are described. First, by adding 2104to the high half of the product, and 252to the low half, we set the implicit leading 1 in the DPF mantissa so that the full 52 explicit bits are available for each half of the 104-bit products of samples. Second, the DPF values are cast bitwise to 64-bit integers for adding the column sums to get the MP result. Normally the cast would require masking off the exponents, but because they are constant, we can include them in the column sums and correct just once for their total. Third, by initializing the column sums with the appropriate negative value to compensate for the exponent sums, no corrective subtraction is needed. Our implementation on an NVIDIA GTX Titan Black GPU achieves between 132.5K and 161.9K modular exponentiations per second of size 1024 bits, with latencies ranging from 21.7 ms to 17.8 ms, making it practical for online RSA applications. Proportional results are shown for 1536 and 2048 bits. The implementation is so efficient that its maximum sustained performance is actually bounded by the thermal limit of the GPU.
Niall Emmart, Fangyu Zheng, Charles C. Weems
ARITH3
2018 A New Variant of the Barrett Algorithm Applied to Quotient Selection
abstract
Quotient Selection (QS) is a key step in the classic O(n2) multiple precision division algorithm. On processors with fast hardware division, it is a trivial problem, but on GPUs, division is quite slow. In this paper we investigate the effectiveness of Brent and Zimmermann's variant as well as our own novel variant of Barrett's algorithm. Our new approach is shown to be suitable for low radix (single precision) QS. Three highly optimized implementations, two of the Brent and Zimmerman variant and one based on our new approach, have been developed and we show that each is many times faster than using the division operation built in to the compiler. In addition, our variant is on average 22 % faster than the other two implementations. We also sketch proofs of correctness for all of the implementations and our new algorithm.
Niall Emmart, Fangyu Zheng, Charles C. Weems
ARITH3
2018 sDPF-RSA: Utilizing Floating-point Computing Power of GPUs for Massive Digital Signature Computations
abstract
In financial, electronic and other security-sensitive industries, data centers require various protocols and algorithms to secure massive volumes of transactions. It is well known that digital signature is a computationally expensive task and a potential bottleneck that can restrict overall performance. In this paper, we make the following contributions. First, we propose a novel method called sDPF-RSA to accelerate the core algorithm of RSA, Montgomery multiplication, for Graphics Processing Units (GPUs). The sDPF approach takes advantage of the sign bit to increase the amount of information processed with each double precision floating point value and considerably improves performance. Second, we have comprehensively reviewed and tested the algorithms to ensure they all run in constant time. In particular we improve the standard carry resolution algorithm, introducing two constant time parallel techniques. We thus minimize the potential for timing attacks against GPU based RSA crypto-systems. Finally, we propose a full implementation of RSA, optimized for our GPU-accelerated computing platform to maximize its computing power. With protection against timing attacks, the throughputs of RSA-2048/3072/4096 on an NVIDIA GeForce GTX TITAN Black set a record of 52,747/15,179/6,435 (for signature generation) and 1,237,694/584,083/354,139 (for signature verification with public key 65,537) operations per second with modest latency, outperforming the contemporaneous CPU and many-core processor Xeon Phi by 3.9-11 times.
Jiankuo Dong, Fangyu Zheng, Niall Emmart, Jingqiang Lin 0001, Charles C. Weems
IPDPS5
2018 NSF/IEEE-TCPP Curriculum Initiative on Parallel and Distributed Computing: Status Report
abstract
No abstract available.
Sushil K. Prasad, Charles C. Weems, John P. Dougherty, Debzani Deb
SIGCSE2
2016 Optimizing Modular Multiplication for NVIDIA's Maxwell GPUs
abstract
In this paper we show how we were able to achieve record rates of multiple precision (MP) modular multiplication (mulmod) operations in the new NVIDIA MP math library (XMP) on Maxwell, NVIDIA's most recent generation of graphics processing units (GPUs). Mulmod is a key operation that is used in multiple places within the MP library, and has many real world applications, especially in cryptography, which makes it important to achieve a highly optimized implementation. Here we reveal how multiple techniques were combined to make the best use of the GPU'sinstructions, registers, memory, and threads. A particularly interesting algorithmic aspect, designed to work with the 16-bit hardware multipliers found in Maxwell, is the use of a two-pass process to first compute unaligned partial products, then shift the result 16 bits to the left, then compute the aligned partial products. The new algorithms are much faster than the prior, state of the art, row-oriented multiply and reduce approach, achieving speedups of 61% at 256 bits, and 117% at 512 bits, with peaks rates of 4027 million mulmod operations at 256 bits and 1081 million at 512 bits on a GTX 980Ti.
Niall Emmart, Justin Luitjens, Charles C. Weems, Cliff Woolley
ARITH3
2016 Asymptotic Optimality of Parallel Short Division
abstract
In 2011 we published a practical algorithm for short division (division of a multiple precision dividend by a single precision divisor) on a parallel processor (HiPC 2011) with a run time of O(n/p+log p). Our algorithm, based on parallel computation of remainder sequences, is an improvement of Takahashi's earlier work (LSSC 2007) which has a run time of O((n/p) log p). Here we prove that Omega(n/p+log p) is a tight lower bound for short division (using a conventional fixed radix number system) on EREW and CREW PRAMs when the divisor d is not simply a power of two. The proof is based on an application of Cook, Dwork, and Reischuk's work on Boolean function complexity. The result itself is especially significant because it establishes a novel tight lower bound for two fundamental arithmetic operations, short division and division by a fixed constant, on an important class of parallel machines.
Niall Emmart, Charles C. Weems
IPDPS2
2015 Pushing the Performance Envelope of Modular Exponentiation Across Multiple Generations of GPUs
abstract
Multiprecision modular exponentiation is a key operation in popular encryption schemes such as RSA, but is computationally expensive. Contexts such as handling many secure web connections in a server can demand higher rates of exponent operations than a traditional multicore can support. Graphics processors offer an opportunity to accelerate batches of exponent calculations both by executing them in parallel as well as through parallelizing the operations within the multiprecision arithmetic itself. However, obtaining performance close to the theoretical peak can be extremely challenging. Furthermore, each new generation of GPU architecture can require a substantially different approach to achieve maximum performance. In this paper we show how we improve modular exponentiation performance over prior results by at factors ranging from 2.6 to 24, across generations of NVIDIA GPU, from compute capability 1.1 onward. Of particular interest is the parameter space that must be searched to find the optimal configuration of memory layout, launch geometry, and algorithm for each architecture at different problem sizes. Our efforts have resulted in a set of tools for generating library functions in the PTX assembly language and searching to find these optima. From our experience it can be argued that a new programming paradigm is needed to achieve full performance potential on core library components as GPUs evolve through multiple generations.
Niall Emmart, Charles C. Weems
IPDPS2
2015 A New Memory-Disk Integrated System with HW Optimizer
abstract
Current high-performance computer systems utilize a memory hierarchy of on-chip cache, main memory, and secondary storage due to differences in device characteristics. Limiting the amount of main memory causes page swap operations and duplicates data between the main memory and the storage device. The characteristics of next-generation memory, such as nonvolatility, byte addressability, and scaling to greater capacity, can be used to solve these problems. Simple replacement of secondary storage with new forms of nonvolatile memory in a traditional memory hierarchy still causes typical problems, such as memory bottleneck, page swaps, and write overhead. Thus, we suggest a single architecture that merges the main memory and secondary storage into a system called a Memory-Disk Integrated System (MDIS). The MDIS architecture is composed of a virtually decoupled NVRAM and a nonvolatile memory performance optimizer combining hardware and software to support this system. The virtually decoupled NVRAM module can support conventional main memory and disk storage operations logically without data duplication and can reduce write operations to the NVRAM. To increase the lifetime and optimize the performance of this NVRAM, another hardware module called a Nonvolatile Performance Optimizer (NVPO) is used that is composed of four small buffers. The NVPO exploits spatial and temporal characteristics of static/dynamic data based on program execution characteristics. Enhanced virtual memory management and address translation modules in the operating system can support these hardware components to achieve a seamless memory-storage environment. Our experimental results show that the proposed architecture can improve execution time by about 89% over a conventional DRAM main memory/HDD storage system, and 77% over a state-of-the-art PRAM main memory/HDD disk system with DRAM buffer. Also, the lifetime of the virtually decoupled NVRAM is estimated to be 40% longer than that of a traditional hierarchy based on the same device technology.
Do-Heon Lee, Su-Kyung Yoon, Jung-Geun Kim, Charles C. Weems, Shin-Dug Kim
ACM Trans. Archit. Code Optim.4
2014 VLSI Design of a Large-Number Multiplier for Fully Homomorphic Encryption
abstract
This paper presents the design of a power- and area-efficient high-speed 768000-bit multiplier, based on fast Fourier transform multiplication for fully homomorphic encryption operations. A memory-based in-place architecture is presented for the FFT processor that performs 64000-point finite-field FFT operations using a radix-16 computing unit and 16 dual-port SRAMs. By adopting a special prime as the base of the finite field, the radix-16 calculations are simplified to requiring only additions and shift operations. A two-stage carry-look-ahead scheme is employed to resolve carries and obtain the multiplication result. The multiplier design is validated by comparing its results with the GNU Multiple Precision (GMP) arithmetic library. The proposed design has been synthesized using 90-nm process technology with an estimated die area of 45.3 mm2. At 200 MHz, the large-number multiplier offers roughly twice the performance of a previous implementation on an NVIDIA C2050 graphics processor unit and is 29 times faster than the Xeon X5650 CPU, while at the same time consuming a modest 0.97 W.
Wei Wang 0053, Xinming Huang 0001, Niall Emmart, Charles C. Weems
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Characterizing the microarchitectural side effects of operating system calls
abstract
We measure the collateral effect on microarchitectural state of system calls using a validated, cycle-accurate simulator. Our results demonstrate that, in some cases, the disruption on user-mode performance is significant. This disruption varies by the operating system and even the kernel version in use.
Addison Mayberry, Matthew Laquidara, Charles C. Weems
ISPASS3
2012 A Pattern Adaptive NAND Flash Memory Storage Structure
abstract
To enhance performance of flash memory-based solid state disk (SSD), large logically chained blocks can be assembled by binding adjacent flash blocks across several flash memory chips. However, flash memory does not allow in-place overwriting and thus the operations that merge writes on these blocks suffer a visible decrease in performance. Furthermore, when small random writes are spread over the disk address space, performance tends to be degraded significantly. We thus present a technique to manage random writes efficiently to achieve stable SSD performance. In this paper, we propose a pattern adaptive SSD structure, which classifies access patterns as either random or sequential. The structure primarily consists of a write cache and a flash translation layer that separates groups of writes by access pattern (S-FTL). Separately managing the two types of write patterns enables greater parallelism and reduces the cost of large block management, thus enhancing the performance of the proposed SSD. Simulation experiments show that the proposed pattern adaptive structure can provide 39 percent decrease in extra flash block erase overhead on the average, and write performance can be improved by around 60 percent, compared with a basic FTL applied to existing parallel SSD structures.
Seung-Ho Park, Jung-Wook Park, Shin-Dug Kim, Charles C. Weems
IEEE Trans. Computers4
2011 Parallel multiple precision division by a single precision divisor
abstract
We report an algorithm for division of a multi- precision integer by a single-precision value using a graphics processing unit (GPU). Our algorithm combines a parallel version of Jebelean's exact division algorithm with a left-to- right algorithm for computing the borrow chain, to relax the requirement of exactness. We also employ Takahashi's recently reported cyclic reduction technique [10] for GPU division to further enhance performance. The result is that our algorithm is asymptotically faster, at O(n/p + log p), than Takahashi's algorithm at 0(n/p log p). We report results for dividends with precisions of 1024, 2048, and 4096 bits running on an NVIDIA GTX 480, and show that, for non-constant divisors, our algorithm is 20% slower at 1024 bits (due to startup overhead), by 2048 we are 40% faster, and at 4096 bits we are able to run 2.5 times faster. For division by constants, with precomputed tables, our algorithm is faster at all sizes with a speedup ranging from 2.3 to 6 times faster.
Niall Emmart, Charles C. Weems
HiPC2
2011 NSF/IEEE-TCPP curriculum initiative on parallel and distributed computing: core topics for undergraduates
abstract
No abstract available.
Sushil K. Prasad, Almadena Yu. Chtchelkanova, Sajal K. Das 0001, Frank Dehne, Mohamed G. Gouda, Joseph F. JáJá, Krishna Kant 0001, Anita La Salle, Richard LeBlanc, Manish Lumsdaine, David A. Padua, Manish Parashar, Viktor Prasanna 0001, Yves Robert, Arnold L. Rosenberg, Sartaj Sahni, Behrooz A. Shirazi, Alan Sussman, Charles C. Weems, Jie Wu 0001
SIGCSE20
2010 An instruction-systolic programmable shader architecture for multi-threaded 3D graphics processing
Jung-Wook Park, Hoon-Mo Yang, Gi-Ho Park, Shin-Dug Kim, Charles C. Weems
J. Parallel Distributed Comput.5
2009 A Profile-Based Multimedia Sharing Scheme With Virtual Community, Based on Personal Space in a Ubiquitous Computing Environment
abstract
For ubiquitous computing environments, an important parameter is whether all the components in the specific environment can connect with one another. Given this capability, we can share various kinds of content across mobile terminals. This paper introduces an effective scheme to manage multimedia sharing based on specially designed profiles and a virtual community. A virtual community is defined as any specific group of users connected for a common interest. Specifically the proposed scheme consists of two layers, i.e., a community construction layer and a multimedia sharing layer, based on personal spaces, which are responsible for constructing and managing the multimedia sharing community. The community construction layer, which is designed to be run on the mobile terminals, provides an effective way to find community members simultaneously, based on specially designed profiles, such as a user profile and an abstract profile. The multimedia sharing layer is responsible for sharing multimedia content, and is constructed as a specially designed scheme based on locality. The proposed scheme provides an effective multimedia sharing mechanism within a community. Simulation results show that the number of messages and the time required for community member discovery is reduced by 32% and 73%, respectively, in comparison with the conventional DHT-based scheme. The approach also reduces the time to exchange content by 50% with respect to the same baseline.
Chung-Pyo Hong, Eo-Hyung Lee, Charles C. Weems, Shin-Dug Kim
IEEE Trans. Multim.3
2008 Towards universal code generator generation
abstract
One of the most difficult tasks a compiler writer faces is the construction of the code generator. The code generator is that part of the compiler that translates compiler intermediate representation (IR) into instructions for a target machine. Unfortunately, implementing a code generator "by hand" is a difficult, time consuming, and error prone task. The details of both the IR and target instruction set must be carefully considered in order to generate correct and efficient code. This, in turn, requires an expert in both the compiler internals as well as the target machine. Even an expert, however, can produce a code generator that is difficult to verify and debug. In this paper we present a universal approach for automating the construction of correct code generators. In particular, we show that both the compiler IR and target instruction set semantics can be described by a machine description language and leveraged by a heuristic search procedure to derive code generator patterns. We then utilize formal methods to determine if the IR and target sequence pairs that make up these patterns are semantically equivalent.
Timothy Richards, Edward K. Walters II, J. Eliot B. Moss, Trek S. Palmer, Charles C. Weems
IPDPS5
2008 CASL: A rapid-prototyping language for modern micro-architectures
Edward K. Walters II, J. Eliot B. Moss, Trek S. Palmer, Timothy Richards, Charles C. Weems
Comput. Lang. Syst. Struct.5
2008 An effective vertical handoff scheme based on service management for ubiquitous computing
Chung-Pyo Hong, Charles C. Weems, Shin-Dug Kim
Comput. Commun.2
2007 Modeling Modern Micro-architectures using CASL
abstract
We overview CASL, the CoGenT architecture specification language, a mixed behavioral-structure architecture description language designed to facilitate fast prototyping and tool generation for computer architectures with deep pipelines and complicated timing. We show how CASL can describe pipelines, dynamic information contexts, and contention using the DLX/MIPS architecture as an example.
Edward K. Walters II, J. Eliot B. Moss, Trek S. Palmer, Timothy Richards, Charles C. Weems
IPDPS5
2003 Guided Region Prefetching: A Cooperative Hardware/Software Approach
abstract
Despite large caches, main-memory access latencies still cause significant performance losses in many applications. Numerous hardware and software prefetching schemes have been proposed to tolerate these latencies. Software prefetching typically provides better prefetch accuracy than hardware, but is limited by prefetch instruction overheads and the compiler's limited ability to schedule prefetches sufficiently far in advance to cover level-two cache miss latencies. Hardware prefetching can be effective at hiding these large latencies, but generates many useless prefetches and consumes considerable memory bandwidth. We propose a cooperative hardware-software prefetching scheme called guided region prefetching (GRP), which uses compiler-generated hints encoded in load instructions to regulate an aggressive hardware prefetching engine. We compare GRP against a sophisticated pure hardware stride prefetcher and a scheduled region prefetching (SRP) engine. SRP and GRP show the best performance, with respective 22% and 21% gains over no prefetching, but SRP incurs 180% extra memory traffic-nearly tripling bandwidth requirements. GRP achieves performance close to SRP, but with a mere eighth of the extra prefetching traffic, a 23% increase over no prefetching. The GRP hardware-software collaboration thus combines the accuracy of compiler-based program analysis with the performance potential of aggressive hardware prefetching, bringing the performance gap versus a perfect L2 cache under 20%.
Zhenlin Wang 0003, Doug Burger, Steven K. Reinhardt, Kathryn S. McKinley, Charles C. Weems
ISCA5
2003 An Intelligent Cache System with Hardware Prefetching for High Performance
abstract
We present a high performance cache structure with a hardware prefetching mechanism that enhances exploitation of spatial and temporal locality. The proposed cache, which we call a selective-mode intelligent (SMI) cache, consists of three parts: a direct-mapped cache with a small block size, a fully associative spatial buffer with a large block size, and a hardware prefetching unit. Temporal locality is exploited by selectively moving small blocks into the direct-mapped cache after monitoring their activity in the spatial buffer for a time period. Spatial locality is enhanced by intelligently prefetching a neighboring block when a spatial buffer hit occurs. The overhead of this prefetching operation is shown to be negligible. We also show that the prefetch operation is highly accurate: Over 90 percent of all prefetches generated are for blocks that are subsequently accessed. Our results show that the system enables the cache size to be reduced by a factor of four to eight relative to a conventional direct-mapped cache while maintaining similar performance. Also, the SMI cache can reduce the miss ratio by around 20 percent and the average memory access time by 10 percent, compared with a victim-buffer cache configuration.
Seh-Woong Jeong, Shin-Dug Kim, Charles C. Weems
IEEE Trans. Computers4
2003 Exploration of the Performance of a Data Mining Application via Hardware Based Monitoring
Mathew S. Thoennes, Charles C. Weems
J. Supercomput.2
2002 A banked-promotion translation lookaside buffer system
Seh-Woong Jeong, Shin-Dug Kim, Charles C. Weems
J. Syst. Archit.4
2002 Application-adaptive intelligent cache memory system
abstract
This article presents the design of a simple hardware-controlled, high performance cache system. The design supports fast access time, optimal utilization of temporal and spatial localities adaptive to given applications, and a simple dynamic fetching mechanism with different fetch sizes. Support for dynamically varying the fetch size makes the cache equally effective for general-purpose as well as multimedia applications. Our cache organization and operational mechanism are especially designed to maximize temporal locality and spatial locality, selectively and adaptively. Simulation shows that the average memory access time of the proposed cache is equal to that of a conventional direct-mapped cache with eight times as much space. In addition, the simulations show that our cache achieves better performance than a 2-way or 4-way set associative cache with twice as much space. The average miss ratio, compared with the victim cache with 32-byte block size, is improved by about 41% or 60% for general applications and multimedia applications, respectively. It is also shown that power consumption of the proposed cache is around 10% to 60% lower than other cache systems that we examine. Our cache system thus offers high performance with low power consumption and low hardware cost.
Shin-Dug Kim, Charles C. Weems
ACM Trans. Embed. Comput. Syst.3
2001 Guest Editor's Introduction, Special Issue: International Parallel and Distributed Processing Symposium 2000
Charles C. Weems
J. Parallel Distributed Comput.1
1999 Using Emulations to Enhance the Performance of Parallel Architectures
abstract
We illustrate the potential of techniques and results from the theory of network emulations to enhance the performance of a parallel architecture. The vehicle for this demonstration is a suite of algorithms that endow an N-processor bit-serial processor array A with a "meta-instruction" GAUGE k, which (logically) reconfigures A into an N/k-processor virtual machine B/sub k/ that has: 1) a datapath and memory bus whose emulated width is k bits, as opposed to A's 1-bit width and 2) an instruction set that operates on k-bit words, in contrast to A's instruction set, which operates on 1-bit words. In order to stress the strength of the approach, we show (via pseudocode) how our emulation techniques can be implemented efficiently even if A operates in strict SIMD mode, with only single-bit masking capabilities and with no indexed memory accesses. We describe at an algorithmic level how to implement our technique-including datapath conversion ("corner-turning") and the creation of the word-parallel instruction sets-on arrays of any regular network topology. We instantiate our technique in detail for arrays based on topologies with quite disparate characteristics: the hypercube, the de Bruijn network, and a genre of mesh with reconfigurable buses. Importantly, the emulations that underlie our technique do not alter the native machine's instruction set, hence allowing an invariant programming model across gauges.
Bojana Obrenic, Martin C. Herbordt, Arnold L. Rosenberg, Charles C. Weems
IEEE Trans. Parallel Distributed Syst.4
1999 The spring scheduling coprocessor: a scheduling accelerator
abstract
The spring scheduling coprocessor is a novel very large scale integration (VLSI) accelerator for multiprocessor real-time systems. The coprocessor can be used for static as well as online scheduling. Many different policies and their combinations can be used (e.g., earliest deadline first, highest value first, or resource-oriented policies such as earliest available time first). In this paper, we describe a coprocessor architecture, a CMOS implementation, an implementation of the host/coprocessor interface and a study of the overall performance improvement. We show that the current VLSI chip speeds up the main portion of the scheduling operation by over three orders of magnitude. We also present an overall system improvement analysis by accounting for the operating system overheads and identify the next set of bottlenecks to improve. The scheduling coprocessor includes several novel VLSI features. It is implemented as a parallel architecture for scheduling that is parameterized for different numbers of tasks, numbers of resources, and internal wordlengths. The architecture was implemented using a single-phase clocking style in several novel ways. The 328 000 transistor custom 2-/spl mu/m VLSI accelerator running with a 100-MHz clock, combined with careful hardware/software co-design results in a considerable performance improvement, thus removing a major bottleneck in real-time systems.
Wayne P. Burleson, Jason Ko, Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems
IEEE Trans. Very Large Scale Integr. Syst.7
1997 Preprototyping SIMD Coprocessors Using Virtual Machine Emulation and Trace Compilation
abstract
The use of massively parallel SIMD array architectures is proliferating in the area of domain specific coprocessors. Even so, they have undergone few systematic empirical studies. The underlying problems include the size of the architecture space, the lack of portability of the test programs, and the inherent complexity of simulating up to hundreds of thousands of processing elements. We address the computational cost problem with a novel approach to trace-based simulation. Code is run on an abstract virtual machine to generate a coarse-grained trace, which is then refined through a series of transformations (a process we call trace compilation) wherein greater resolution is obtained with respect to the details of the target machine. We have found this technique to be one to two orders of magnitude faster than instruction-level simulation while still retaining much of the accuracy of the model. Furthermore, abstract machine traces must be regenerated for only a small fraction of the possible parameter combinations. Using virtual machine emulation and trace compilation also addresses program portability by allowing the user to code in a single data parallel language with a single compiler, regardless of the target architecture. This technique has already been used to generate significant results with respect to SIMD array architectures, a sample of which are presented here.
Martin C. Herbordt, Owais Kidwai, Charles C. Weems
SIGMETRICS3
1995 An empirical study of datapath, memory hierarchy, and network in SIMD array architectures
abstract
Although SIMD arrays have been built for 30 years, they have as a class been the subject of few empirical design studies. Using ENPASSANT, a simulation environment developed for that purpose, we analyze several aspects of SIMD array architecture with respect to a test suite of spatially mapped applications. Several surprising results are obtained. With respect to memory hierarchy, we find that adding a level of cache to current PE designs is likely to be advantageous, but that such a cache will look quite different than expected. In particular, we find that associativity has unusual significance and that performance varies inversely with block size. Router network results indicate the importance of support for local transfers, broadcast, and reduction even at the expense of arbitrary permutations. Other communication results point to the appropriate dimensionality of k-ary n-cube networks (2 or 3), and the criticality of supporting bidirectional transfers, even if the overall bandwidth remains unchanged.
Martin C. Herbordt, Charles C. Weems
ICCD2
1995 Experimental Analysis of Some SIMD Array Memory Hierarchies
Martin C. Herbordt, Charles C. Weems
ICPP (1)2
1995 Enpassant: An Environment for Evaluating Massively Parallel Array Architectures for Spatially Mapped Applications
abstract
Although massively parallel arrays for spatially mapped applications have been proposed since the 1950s42 and built since the 1960s,12 there have been very few systematic empirical studies that cover more than a small fraction of the design space. The problems have included the lack of a test suite of non-trivial application codes; inadequate language support; the difficulties of balancing evaluation performance with flexibility; and balancing test suite portability with accuracy of evaluation. We describe an environment that addresses these problems. A realistic workload including a series of applications currently being used as building blocks in vision research has been constructed. Both flexibility in architectural parameter selection and simulation efficiency are maintained with a novel new technique that combines virtual machine emulation with trace-driven simulation. The trade-off between fairness to diverse target architectures and programmability of the test suite is addressed through the use of operator and application libraries for a small set of critical functions. We also present examples of the type of results we are obtaining, including the effects of changing ALU designs and datapath widths, finding critical points in register set and cache sizes, the benefits of various types of router networks, and the performance cost of processor virtualization.
Martin C. Herbordt, Charles C. Weems
Int. J. Pattern Recognit. Artif. Intell.2
1994 Practical Algorithms for Online Routing on Fixed and Reconfigurable Meshes
Martin C. Herbordt, James C. Corbett, Charles C. Weems, John Spalding
J. Parallel Distributed Comput.3
1993 Parallel dense depth-from-motion on the image understanding architecture
abstract
The design and implementation of a single instruction multiple data depth-from-motion algorithm on the image understanding architecture simulator are described. Correspondences are established in parallel for two temporarily separated images through correlation. The correspondences are used to determine the translational and rotational motion parameters of the camera through a parallel motion algorithm. This is done by first determining the appropriate translational parameters and then constraining the search for the exact translational and rotational parameters. The dense depth map is computed from the image correspondences and the computed motion parameters. Results are analyzed for three image sequences acquired from mobile vehicles. Depths are obtained at an average accuracy of about 8% in outdoor image sequences. The depth maps are processed to locate relatively small obstacles, like cans and cones, to a distance of about 60 ft. Large obstacles, like hills, are located even when they are much further away.>
Rabindranath Dutta, Charles C. Weems
CVPR2
1993 The Spring Scheduling Co-Processor: A Scheduling Accelerator
abstract
We present a novel co-processor for multiprocessor scheduling in the Spring real-time operating system. Since most dynamic scheduling problems are NP-complete, we use a heuristic algorithm which uses a smart searching scheme to find a feasible schedule for a set of specified tasks and hard deadlines. A parallel VLSI architecture for scheduling is developed that can be scaled for different numbers of tasks, numbers of resources, internal wordlengths, and future IC technologies. The scheduling architecture is implemented in a 0.8/spl mu/ CMOS technology and uses an advanced clocking scheme to allow further scaling to future technologies. With an internal clock rate of 100 MHz, a speed increase of two orders of magnitude is expected for scheduling tasks, thus removing a major bottleneck in real-time systems.>
Wayne P. Burleson, Jason Ko, Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems
ICCD7
1993 The Spring Scheduling Co-Processor: Design, Use, and Performance
abstract
We present a novel VLSI co-processor for real-time multiprocessor scheduling. The co-processor can be used for sophisticated static scheduling as well as for online scheduling using many different algorithms such as earliest deadline first, highest value first, or the Spring scheduling algorithm. When such an algorithm is used online it is important to assess the performance impact of the interface of the co-processor to the host system, in this case, the Spring kernel. We focus on the interface and its implications for overall scheduling performance. We show that the current VLSI chip speeds up the main portion of the scheduling operation by over three orders of magnitude and speeds up the overall scheduling operation 30 fold. The parallel VLSI architecture for scheduling is briefly presented. This architecture can be scaled for different numbers of tasks, resources, and internal word lengths. The implementation uses an advanced clocking scheme to allow further scaling using future IC technologies.>
Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems, Wayne P. Burleson, Jason Ko
RTSS5
1993 Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial Intelligence
abstract
The findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.>
Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems
IEEE Trans. Knowl. Data Eng.24
1992 Nonuniform region processing on SIMD arrays using the coterie network
Martin C. Herbordt, Charles C. Weems, Michael J. Scudder
Mach. Vis. Appl.2
1992 Next generation architectures integrating sensory and symbolic processing
Charles C. Weems
Mach. Vis. Appl.1
1991 A computational framework and SIMD algorithms for low-level support of intermediate level vision processing
abstract
The authors propose an additional level of parallelism, called multi-associativity, as a framework for simultaneously performing associative computation on data sets mapped to irregular, non-uniform, aggregates of processing elements (PEs). They introduce algorithms developed for the CAAPP to simulate efficiently within aggregates of PEs simultaneously the associative algorithms typically supported in hardware at the array level. Some of the results are: the efficient application of existing associative algorithms to arbitrary aggregates of PEs in parallel and the development of multi-associative algorithms, among them parallel prefix and convex hull. The multi-associative framework also extends the associative paradigm by allowing operation on and among aggregates themselves, operations not defined when the entity in question is always an entire array.>
Martin C. Herbordt, Charles C. Weems, Michael J. Scudder
CVPR2
1991 Multi-associativity: A Framework for Solving Multiple Non-uniform Problem Instances Simultaneously on SIMD Arrays
Martin C. Herbordt, Charles C. Weems
ICPP (3)2
1991 The DARPA Image Understanding Benchmark for Parallel Computers
Charles C. Weems, Edward M. Riseman, Allen R. Hanson, Azriel Rosenfeld
J. Parallel Distributed Comput.1
1991 Architectural requirements of image understanding with respect to parallel processing
abstract
An overview of the architectural requirements for parallel processing in support of real-time, knowledge-based computer vision is given. One of the goals of this work is to provide an appreciation for the diversity, complexity, and computational intensity of vision processing. It begins with a description of common vision algorithms, analyzes their requirements in terms of the inherent structures that are present, and relates them to computation, communication, and control in parallel processing. After concluding that traditional architectural approaches to parallel processors are suboptimal, discussion focuses on heterogeneous processors. Different forms of parallelism can be applied at three levels of computational granularity, each having unique requirements and corresponding to levels of abstraction in the image interpretation process. In addition, interaction between levels must take place via parallel data and control paths. The paper concludes with a brief discussion of image understanding architecture, a multilevel parallel processor designed for image understanding.>
Charles C. Weems
Proc. IEEE1
1990 A feedback concentrator for the Image Understanding Architecture
abstract
The Image Understanding Architecture (IUA) is a massively parallel, multilevel system. The hardware implementation of two important summary feedback mechanisms-some/none response and count responders-for the lower two processing levels of the first generation IUA are described. Both mechanisms are implemented using multiple copies of a single custom VLSI chip. A brief overview of the custom chip is provided. The performance of the IUA's low-level processor with the feedback concentrator is compared to a similar mesh-connected parallel processor without the feedback concentrator mechanism and is shown to be significantly faster. An overview of the plants for the feedback concentrator for the second generation IUA is provided.>
Deepak Rana, Charles C. Weems
ASAP2
1990 A multiple-level heterogeneous architecture for image understanding
abstract
The image understanding architecture (IUA) system is designed specifically for computer vision processing that relies heavily on artificial intelligence techniques to classify objects. To provide for low-, intermediate-, and high-level computer vision processing, the IUA system combines three heterogeneous levels of parallelism with associative processing mechanisms. The lowest level of the IUA contains a SIMD array of per pixel bit-serial processing elements with networking capabilities that exceed those of present bit-serial machines. The SIMD level is closely mated with a middle-level MIMD array of high-performance digital signal processing chips that communicate via a flexible, modular message-passing network. The dual-capability architecture is much more effective than a solely SIMD or MIMD organization. The highest level comprises an array of general-purpose microprocessors for symbol manipulation. The first- and second-generation implementations of the IUA architecture are described.>
David B. Shu, J. Greg Nash, Charles C. Weems
ASAP3
1990 Routing on the CAAPP
abstract
A collection of routing algorithms for the content-addressable array parallel processor (CAAPP) which allow many more classes of interprocessor communication to be executed efficiently than otherwise on machines with conventional mesh-connected topologies is presented. It is shown that routing on the CAAPP can be executed with simplicity and performance similar to that of a dedicated routing network. Experimental results are presented from random permutations as well as from several common machine vision applications.>
Martin C. Herbordt, Charles C. Weems, David B. Shu
ICPR (2)2
1990 The IUA feedback concentrator
abstract
A feedback concentrator for a massively parallel, multilevel image understanding architecture (IUA) is presented. A brief overview of the IUA is given. The details of the feedback concentrator mechanism, which was implemented using a custom VLSI chip, are presented. The custom chip uses a combination of circuit techniques to achieve high speed. A description of the custom VLSI concentrator chip is provided. The performance of the architecture is compared with a mesh-connected processor without the feedback mechanism for three common, low-level vision tasks and is shown to be significantly faster.>
Deepak Rana, Charles C. Weems
ICPR (2)2
1990 A multiple-level heterogeneous architecture for image understanding
abstract
First- and second-generation implementations of an image understanding architecture, IUA, are described. The IUA system is designed specifically for computer vision processing that relies heavily on artificial intelligence techniques to classify objects. To provide for low-, intermediate-, and high-level computer vision processing-required for model-/knowledge-based interpretation of sensor data-the IUA system combines three heterogeneous levels of parallelism with associative processing mechanisms. The tightly coupled symbolic and numeric processing capabilities constitute a unique computing paradigm ideally suited to computer vision applications, which require both control- and data-parallel processing. The IUA architecture, hardware, and software are described.>
David B. Shu, J. Greg Nash, Charles C. Weems
ICPR (2)3
1990 An overview of architecture research for image understanding at the University of Massachusetts
abstract
Architectural research in support of knowledge-based computer vision is described. Efforts are focused on two major areas: the development of the image understanding architecture (IUA) and the benchmarking of parallel processors for vision. Two aspects of the IUA research are discussed. One is the development of the low-level processor chip, which integrates 64 one-bit (serial) processors, data caches, a communication network, and parallel interfaces to image input/output (I/O) hardware and the intermediate-level processors. The other is the ICAP, which is designed to manipulate tokens (symbolic descriptions of extracted image events and their associated attributes) at the intermediate level and to support database functions that allow access to these tokens. Work on an image understanding work bench is briefly described.>
Charles C. Weems, Deepak Rana, Allen R. Hanson, Edward M. Riseman, David B. Shu, J. Greg Nash
ICPR (2)1
1990 Message-Passing Algorithms for a SIMD Torus with Coteries
abstract
This paper describes the results of an investigation into routing algorithms to be used when programming the CAAPP (Content Addressable Array Parallel Processor) [19], a SIMD mesh-connected array processor enhanced with the coterie network, a mechanism similar to reconfigurable buses. We will show that the coterie network gives the CAAPP a capability far beyond solely meshconnected processors; in fact, the performance of routing on many classes of permutations is more comparable to the Connection Machine which has a dedicated hypercube routing network. Most of the current routing algorithms for meshconnected array processors (with N PEs in an n \\Theta n
Martin C. Herbordt, Charles C. Weems, James C. Corbett
SPAA2
1989 The image understanding architecture
Charles C. Weems, Steven P. Levitan, Allen R. Hanson, Edward M. Riseman, David B. Shu, J. Greg Nash
Int. J. Comput. Vis.1
1988 IU parallel processing benchmark
abstract
A benchmark is presented that was designed to evaluate the merits of various parallel architectures as applied to image understanding (IU). This benchmark exercise addresses the issue of system performance on an integrated set of tasks, where the task interactions that are typical of complex vision application are present. The goal of this exercise is to gain a better understanding of vision architecture requirements, which can be used to guide the development of the next generation of vision architectures.>
Charles C. Weems, Edward M. Riseman, Allen R. Hanson, Azriel Rosenfeld
CVPR1
1983 Determination of the Rotational and Translational Components of a Flow Field Using a Content Addressable Parallel Processor
Martha Steenstrup, Daryl T. Lawton, Charles C. Weems
ICPP3