EDBT 2026 Demo / reviewers in the wild / expert
Richard W. Linderman
dblp:67/6524
· DBLP profile ↗
18ranked-venue papers
4as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9Systems, architecture and hardware · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 30% Emerging computing paradigms · 29% Reconfigurable computing and FPGAs · 22% | |
| Artificial intelligence
1 paper |
Image recognition and object detection · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2013 | A Parallel Neuromorphic Text Recognition System and Its Implementation on a Heterogeneous High-Performance Computing Cluster · IEEE Trans. Computers 2013 |
Parallel and multicore computing
parallel architecture |
0.2 | 1 | 2013 | A Parallel Neuromorphic Text Recognition System and Its Implementation on a Heterogeneous High-Performance Computing Cluster · IEEE Trans. Computers 2013 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.1 | 1 | 2006 | Poster reception - Improving the performance of parallel backprojection on a reconfigurable supercomputer · SC 2006 |
Reconfigurable computing and FPGAs › high-performance reconfigurable computing
reconfigurable supercomputing |
0.1 | 1 | 2006 | Poster reception - Improving the performance of parallel backprojection on a reconfigurable supercomputer · SC 2006 |
Computer vision › Image recognition and object detection
text recognition |
0.0 | 1 | 2013 | A Parallel Neuromorphic Text Recognition System and Its Implementation on a Heterogeneous High-Performance Computing Cluster · IEEE Trans. Computers 2013 |
Embedded and real-time systems
embedded signal processing |
0.0 | 1 | 1998 | A Dependable High Performance Wafer Scale Architecture for Embedded Signal Processing · IEEE Trans. Computers 1998 |
Hardware accelerators and domain-specific architectures
wafer-scale architecture |
0.0 | 1 | 1998 | A Dependable High Performance Wafer Scale Architecture for Embedded Signal Processing · IEEE Trans. Computers 1998 |
High-performance computing › cluster computing
heterogeneous clusters |
0.0 | 1 | 2006 | Poster reception - Improving the performance of parallel backprojection on a reconfigurable supercomputer · SC 2006 |
Parallel and multicore computing
multiprocessor system |
0.0 | 1 | 1998 | A Dependable High Performance Wafer Scale Architecture for Embedded Signal Processing · IEEE Trans. Computers 1998 |
Electronic design automation
logic synthesis |
0.0 | 1 | 1989 | Design and application of an optimizing XROM silicon compiler · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989 |
Integrated circuit design
memory circuit design |
0.0 | 1 | 1989 | Design and application of an optimizing XROM silicon compiler · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989 |
Memory systems › non-volatile memory
read-only memory |
0.0 | 1 | 1989 | Design and application of an optimizing XROM silicon compiler · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989 |
Electronic design automation › system-level design › system synthesis
silicon compiler |
0.0 | 1 | 1989 | Design and application of an optimizing XROM silicon compiler · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989 |
Methods — techniques the papers use, named apart from their topics
parallel computing · 0.3neuromorphic computing models · 0.3watchdog checking · 0.0coprocessing · 0.0traveling salesman problem · 0.0graph partitioning heuristic · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | AnRAD: A Neuromorphic Anomaly Detection Framework for Massive Concurrent Data StreamsabstractThe evolution of high performance computing technologies has enabled the large-scale implementation of neuromorphic models and pushed the research in computational intelligence into a new era. Among the machine learning applications, unsupervised detection of anomalous streams is especially challenging due to the requirements of detection accuracy and real-time performance. Designing a computing framework that harnesses the growing computing power of the multicore systems while maintaining high sensitivity and specificity to the anomalies is an urgent research topic. In this paper, we propose anomaly recognition and detection (AnRAD), a bioinspired detection framework that performs probabilistic inferences. We analyze the feature dependency and develop a self-structuring method that learns an efficient confabulation network using unlabeled data. This network is capable of fast incremental learning, which continuously refines the knowledge base using streaming data. Compared with several existing anomaly detection approaches, our method provides competitive detection quality. Furthermore, we exploit the massive parallel structure of the AnRAD framework. Our implementations of the detection algorithm on the graphic processing unit and the Xeon Phi coprocessor both obtain substantial speedups over the sequential implementation on general-purpose microprocessor. The framework provides real-time service to concurrent data streams within diversified knowledge contexts, and can be applied to large problems with multiple local patterns. Experimental results demonstrate high computing performance and memory efficiency. For vehicle behavior detection, the framework is able to monitor up to 16000 vehicles (data streams) and their interactions in real time with a single commodity coprocessor, and uses less than 0.2 ms for one testing subject. Finally, the detection network is ported to our spiking neural network simulator to show the potential of adapting to the emerging neuromorphic architectures. Qiuwen Chen, Ryan S. Luley, Qing Wu 0002, Morgan Bishop, Richard W. Linderman, Qinru Qiu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Self-structured confabulation network for fast anomaly detection and reasoningabstractInference models such as the confabulation network are particularly useful in anomaly detection applications because they allow introspection to the decision process. However, building such network model always requires expert knowledge. In this paper, we present a self-structuring technique that learns the structure of a confabulation network from unlabeled data. Without any assumption of the distribution of data, we leverage the mutual information between features to learn a succinct network configuration, and enable fast incremental learning to refine the knowledge bases from continuous data streams. Compared to several existing anomaly detection methods, the proposed approach provides higher detection performance and excellent reasoning capability. We also exploit the massive parallelism that is inherent to the inference model and accelerate the detection process using GPUs. Experimental results show significant speedups and the potential to be applied to real-time applications with high-volume data streams. Qiuwen Chen, Qing Wu 0002, Morgan Bishop, Richard W. Linderman, Qinru Qiu |
IJCNN | 4 |
| 2014 | Memristor Crossbar-Based Neuromorphic Computing System: A Case StudyabstractBy mimicking the highly parallel biological systems, neuromorphic hardware provides the capability of information processing within a compact and energy-efficient platform. However, traditional Von Neumann architecture and the limited signal connections have severely constrained the scalability and performance of such hardware implementations. Recently, many research efforts have been investigated in utilizing the latest discovered memristors in neuromorphic systems due to the similarity of memristors to biological synapses. In this paper, we explore the potential of a memristor crossbar array that functions as an autoassociative memory and apply it to brain-state-in-a-box (BSB) neural networks. Especially, the recall and training functions of a multianswer character recognition process based on the BSB model are studied. The robustness of the BSB circuit is analyzed and evaluated based on extensive Monte Carlo simulations, considering input defects, process variations, and electrical fluctuations. The results show that the hardware-based training scheme proposed in the paper can alleviate and even cancel out the majority of the noise issue. Miao Hu 0002, Hai Li 0001, Yiran Chen 0001, Qing Wu 0002, Garrett S. Rose, Richard W. Linderman |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2013 | A Parallel Neuromorphic Text Recognition System and Its Implementation on a Heterogeneous High-Performance Computing ClusterabstractGiven the recent progress in the evolution of high-performance computing (HPC) technologies, the research in computational intelligence has entered a new era. In this paper, we present an HPC-based context-aware intelligent text recognition system (ITRS) that serves as the physical layer of machine reading. A parallel computing architecture is adopted that incorporates the HPC technologies with advances in neuromorphic computing models. The algorithm learns from what has been read and, based on the obtained knowledge, it forms anticipations of the word and sentence level context. The information processing flow of the ITRS imitates the function of the neocortex system. It incorporates large number of simple pattern detection modules with advanced information association layer to achieve perception and recognition. Such architecture provides robust performance to images with large noise. The implemented ITRS software is able to process about 16 to 20 scanned pages per second on the 500 trillion floating point operations per second (TFLOPS) Air Force Research Laboratory (AFRL)/Information Directorate (RI) Condor HPC after performance optimization. Qinru Qiu, Qing Wu 0002, Morgan Bishop, Robinson E. Pino, Richard W. Linderman |
IEEE Trans. Computers | 5 |
| 2011 | Parallel flux tensor analysis for efficient moving object detection
Kannappan Palaniappan, Ilker Ersoy, Guna Seetharaman, Shelby R. Davis, Praveen Kumar 0005, Raghuveer M. Rao, Richard W. Linderman |
FUSION | 7 |
| 2011 | Unified perception-prediction model for context aware text recognition on a heterogeneous many-core platformabstractExisting optical character recognition (OCR) software tools can perform text image detection and pattern recognition with fairly high accuracy, however their performance will be significantly impaired when the image of the character is partially blocked or smudged. Such missing information does not hinder the human perception because we predict the missing part based on the word level and sentence level context of the character. In order to mimic the human cognitive behavior, we developed a hybrid cognitive architecture combining two neuromorphic computing models, i.e. brain-state-in-a-box (BSB) and cogent confabulation, to achieve context-aware text recognition. The BSB model performs the character recognition from input image while the confabulation models perform the context-aware prediction based on the word and sentence knowledge bases. The software tool is implemented on an 1824-core computing cluster. Its accuracy and performance are analyzed in the paper. Qinru Qiu, Qing Wu 0002, Richard W. Linderman |
IJCNN | 3 |
| 2010 | Affordable emerging computer hardware for neuromorphic computing applicationsabstractWe are pursuing an investigation of neuromorphic computational models and architectures in order to leverage present understanding of how the estimated 1011neurons and 1015neuron connections in the mammalian brain are able to do some of the things a human does, and as quickly as it does it, using slow base components, while consuming very little power on affordable synthetic non-biological computing hardware. Understanding and harvesting neurologically based methods is a promising approach with great potential that may help us achieve massively parallel computation far beyond the scope of traditional computing. Morgan Bishop, Michael J. Moore, Daniel J. Burns, Robinson E. Pino, Richard W. Linderman |
IJCNN | 5 |
| 2010 | A columnar primary visual cortex (V1) model emulation using a PS3 Cell-BE arrayabstractA model of portions of the cerebral cortex is being developed to explore neuromorphic computing strategies in the context of highly parallel platforms. The interest is driven by the value of applications which can make use of highly parallel architectures we expect to see surpassing one thousand cores per die in the next few years. A central question we seek to answer is what the architecture of hyper-parallel machines should be. We also seek to understand computational methods akin to how a brain deals with sensing, perception, memory, and cognition. The model is being developed incrementally, starting with the primary visual cortex (V1) field. It is based upon structures roughly corresponding to neocortical minicolumn and functional column structures. Gaps in neuroscience, such as inter-cell connectivity, are filled using estimates of functionality that are plausible given current understanding of the micro-anatomy. The success we encountered with achieving real-time performance is evidence validating the use of Cell-Be architecture in some classes of neuromorphic emulation. In this study we identified a particular gap-fill algorithm for lateral connections within V1 that is suggestive of a learning strategy whereby the lateral network subsumes expectation affect, reducing perception time and improving perception affect. Michael J. Moore, Richard W. Linderman, Morgan Bishop, Robinson E. Pino |
IJCNN | 2 |
| 2010 | Neuromorphic algorithms on clusters of PlayStation 3sabstractThere is a significant interest in the research community to develop large scale, high performance implementations of neuromorphic models. These have the potential to provide significantly stronger information processing capabilities than current computing algorithms. In this paper we present the implementation of five neuromorphic models on a 50 TeraFLOPS 336 node Playstation 3 cluster at the Air Force Research Laboratory. The five models examined span two classes of neuromorphic algorithms: hierarchical Bayesian and spiking neural networks. Our results indicate that the models scale well on this cluster and can emulate between 108 to 1010 neurons. In particular, our study indicates that a cluster of Playstation 3s can provide an economical, yet powerful, platform for simulating large scale neuromorphic models. Tarek M. Taha, Pavan Yalamanchili, Mohammad Ashraf Bhuiyan, Rommel Jalasutram, Richard W. Linderman |
IJCNN | 6 |
| 2008 | Performance optimization for pattern recognition using associative neural memoryabstractIn this paper, we present our work in the implementation and performance optimization of the recall operation of the Brain-State-in-a-Box (BSB) model on the Cell Broadband Engine processor. We have applied optimization techniques on different parts of the algorithm to improve the overall computing and communication performance of the BSB recall algorithm. Runtime measurements show that, we have been able to achieve about 70% of the theoretical peak performance of the processor. Qing Wu 0002, Prakash Mukre, Richard W. Linderman, Thomas Renz, Daniel J. Burns, Michael J. Moore, Qinru Qiu |
ICME | 3 |
| 2008 | Accelerating cogent confabulation: An exploration in the architecture design spaceabstractCogent confabulation is a computation model that mimics the Hebbian learning, information storage, inter-relation of symbolic concepts, and the recall operations of the brain. The model has been applied to cognitive processing of language, audio and visual signals. In this project, we focus on how to accelerate the computation which underlie confabulation based sentence completion through software and hardware optimization. On the software implementation side, appropriate data structures can improve the performance of the software by more than 5,000X. On the hardware implementation side, the cogent confabulation algorithm is an ideal candidate for parallel processing and its performance can be significantly improved with the help of application specific, massively parallel computing platforms. However, as the complexity and parallelism of the hardware increases, cost also increases. Architectures with different performance-cost tradeoffs are analyzed and compared. Our analysis shows that although increasing the number of processors or the size of memories per processor can increase performance, the hardware cost and performance improvements do not always exhibit a linear relation. Hardware configuration options must be carefully evaluated in order to achieve good cost performance tradeoffs. Qinru Qiu, Daniel J. Burns, Michael J. Moore, Richard W. Linderman, Thomas Renz, Qing Wu 0002 |
IJCNN | 4 |
| 2007 | Architectural Design and Complexity Analysis of Large-Scale Cortical Simulation on a Hybrid Computing PlatformabstractResearch and development in modeling and simulation of human cognizance functions requires a high-performance computing platform for manipulating large-scale mathematical models. Traditional computing architectures cannot fulfill the attendant needs in terms of arithmetic computation and communication bandwidth. In this work, we propose a novel hybrid computing architecture for the simulation and evaluation of large-scale associative neural memory models. The proposed architecture achieves very high computing and communication performances by combining the technologies of hardware-accelerated computing, parallel distributed data operation and the publish/subscribe protocol. Analysis has been done on the computation and data bandwidth demands for implementing a large-scale brain-state-in-a-box (BSB) model. Compared to the traditional computing architecture, the proposed architecture can achieve at least 100X speedup. Qing Wu 0002, Qinru Qiu, Richard W. Linderman, Daniel J. Burns, Michael J. Moore, Dennis Fitzgerald |
CISDA | 3 |
| 2006 | Limiting Optimism: Time or Event Count?abstractIn optimistically synchronized parallel discrete event simulators, unlimited optimism can lead to excessive rollbacks and simulation thrashing. Artificially throttling the simulation is a well-known technique for improving the performance by avoiding the effects of uncontrolled optimism. Simulation throttling has been attempted based on time or event count beyond global virtual time (GVT). In this paper, we carry out a simulation study of both approaches within time warp as well as breathing time warp (BTW) synchronization algorithms in the context of the SPEEDES simulation framework. We discover that anomalies arise when limiting optimism based on event count (as is the case in the default SPEEDES BTW algorithm) and that this gives rise to forced rollbacks that generally do not arise without simulation throttling. We implement a version of BTW that limits optimism based on time beyond GVT and show that it outperforms BTW for two large scale applications. We discuss initial experiences with adaptively estimating the optimism limit Nael B. Abu-Ghazaleh, Richard W. Linderman |
DS-RT | 2 |
| 2006 | Poster reception - Improving the performance of parallel backprojection on a reconfigurable supercomputerabstractReconfigurable supercomputing is a new direction of research in the quest for novel high-performance architectures. A handful of products are available in this field, but no architecture has emerged as the clear winner in performance. Using the US Air Force Research Labs' Heterogeneous High Performance Cluster (HHPC), we have developed a successful reconfigurable supercomputing implementation of synthetic aperture radar (SAR) image formation using backprojection, a highly parallel algorithm. By studying the architectural features of the HHPC and hand-tuning the application to those features, we estimate that better than 1000x speedup over a single-node software-only solution is feasable. In particular, we consider the movement of data between nodes, the movement of data between host PC and FPGA board, and the use of on-board memories available to the FPGA for data storage. These results indicate that reconfigurable supercomputers are an excellent choice to tackle highly-parallel large-dataset problems. Ben Cordes, Miriam Leeser, Eric L. Miller 0001, Richard W. Linderman |
SC | 4 |
| 1998 | A Dependable High Performance Wafer Scale Architecture for Embedded Signal ProcessingabstractA high performance, programmable, floating point multiprocessor architecture has been specifically designed to exploit advanced two- and three-dimensional hybrid wafer scale packaging to achieve low size, weight, and power, and improve reliability for embedded systems applications. Processing elements comprised of a 0.8 micron CMOS dual processor chip and commercial synchronous SRAMs achieve more than 100 MFLOPS/Watt. This power efficiency allows up to 32 processing elements to be incorporated into a single 3D multichip module, eliminating multiple discrete packages and thousands of wirebonds. The dual processor chip can dynamically switch between independent processing, watchdog checking, and coprocessing modes. A flat, SRAM memory provides predictable instruction set timing and independent and accurate performance prediction. Richard W. Linderman, Ralph L. R. Kohler, Mark H. Linderman |
IEEE Trans. Computers | 1 |
| 1989 | Design and application of an optimizing XROM silicon compilerabstractIt is demonstrated that optimization techniques incorporated within a silicon compiler for read-only memories (ROMs) can achieve significant yield, power, and speed improvements by minimizing the number of transistors, drains, and metal interconnections in the ROM. Transistor minimization adopts a heuristic solution to the NP-complete graph partitioning problem with a powerful technique applicable to various ROM design styles and technologies. If diffusion mask personalization is permitted, the design can be further improved by solving the traveling salesman problem to minimize transistor source/drain regions. In table look-up ROMs compiled for 3- and 1.2- mu m CMOS with diffusion mask programming, the compiler eliminated over 45% of the transistors and drains. Test results show that 3- mu m CMOS ROMs have access times between 50 and 70 ns. ROMs with 1.2- mu m features achieve simulated access times below 20 ns. A simple interface allows the optimizing compiler to work easily with other CAD tools such as microcode assemblers.> Richard W. Linderman, Paul C. Rossbach, David M. Gallagher |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1984 | A three dimensional systolic array architecture for fast matrix multiplicationabstractIn recent years one and two dimensional systolic arrays have been designed to implement a wide range of matrix operations and signal processing algorithms. This paper proposes a three dimensional systolic array architecture of N3processors to further extend the performance advantages which can be achieved through regular local data transfer. The three dimensional array is discussed in the context of fixed point matrix-matrix multiplication which requires O(N3) multiplications and additions. The array pipelines N of these problems to attain a throughput rate which is practically independent of N. The performance, size, and fault tolerance of the array are discussed for the case N=32. Richard W. Linderman, Walter H. Ku |
ICASSP | 1 |
| 1984 | Digital signal processing capabilities of CUSP, a high performance bit-serial VLSI processorabstractThe Cornell University Signal Processor (CUSP) is a high performance CMOS processor which has been custom designed in VLSI to efficiently compute digital signal processing algorithms based on the Cooley-Tukey Radix-4 Fast Fourier Transform. One CUSP chip can be used as a stand-alone peripheral in a microprocessor system, or CUSP units can be combined into arrays in order to process signals with sampling rates of several MHz. It can attain a high numerical accuracy while maintaining a throughput superior to most available systems. In this paper we discuss the signal processing capabilities of CUSP, its simple programmability, and I/O interface. We also present simulation results which quantify the signal-to-noise ratio advantages of the bit-serial architecture. Richard W. Linderman, Peter P. Reusens, Paul M. Chau, Walter H. Ku |
ICASSP | 1 |