VLDB 2026 Research / reviewers in the wild / expert
Jean-Luc Gaudiot
dblp:g/JeanLucGaudiot
· DBLP profile ↗
138ranked-venue papers
28as first author
12since 2021 · last 2026
0000-0001-9164-8731ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 103 · 23 first-author · 4 since 2021Software engineering, systems software and programming languages · 15 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IEEE Computer Society in the Age of AI: A Strategic Position Paper for the Next Decade
Jean-Luc Gaudiot, Shaoshan Liu |
COMPSAC | 1 |
| 2025 | A Data-Driven Approach Using Hardware Performance Counters to Detect Micro-Architectural AttacksabstractModern processors remain vulnerable to microarchitectural attacks like Spectre, which exploit speculative execution to leak sensitive data. Traditional detection methods struggle with evolving threats, while hardware redesigns are impractical for legacy systems. This paper presents a hardware performance counter (HPC)-based method that achieves 99% detection accuracy in identifying Spectre attack patterns. Compared to existing approaches, our method remains robust against noise and background variations, ensuring high detection accuracy across system loads without necessitating hardware modifications. Jaya Keshava Chandra Kotha, Jean-Luc Gaudiot |
COMPSAC | 2 |
| 2024 | Dataflow Accelerator Architecture for Autonomous Machine ComputingabstractCommercial autonomous machines is a thriving sector, one that is likely the next ubiquitous computing platform, after Personal Computers (PC), cloud computing, and mobile computing. Nevertheless, a suitable computing substrate for autonomous machines is missing, and many companies are forced to develop ad hoc computing solutions that are neither principled nor extensible. By analyzing the demands of autonomous machine computing, this article proposes Dataflow Accelerator Architecture (DAA), a modern instantiation of the classic dataflow principle, that matches the characteristics of autonomous machine software. Shaoshan Liu, Yuhao Zhu 0001, Bo Yu 0014, Jean-Luc Gaudiot, Guangrong Gao |
ICCAD | 4 |
| 2024 | Hierarchical Heterogeneous Cluster Systems for Scalable Distributed Deep LearningabstractDistributed deep learning framework tools should aim at high efficiency of training and inference of distributed exascale deep learning algorithms. There are three major challenges in this endeavor: scalability, adaptivity and efficiency. Any future framework will need to be adaptively utilized for a variety of heterogeneous hardware and network environments and will thus be required to be capable of scaling from single compute node up to large clusters. Further, it should be efficiently integrated into popular frameworks such as TensorFlow, PyTorch, etc. This paper proposes a dynamically hybrid (hierarchy) distribution structure for distributed deep learning, taking advantage of flexible synchronization on both centralized and decentralized architectures, implementing multi-level fine-grain parallelism on distributed platforms. It is scalable as the number of compute nodes increases, and can also adapt to various compute abilities, memory structures and communication costs. Tongsheng Geng, Ericson Silva, Jean-Luc Gaudiot |
ISORC | 4 |
| 2023 | A Novel Spatial-Temporal Multi-Scale Alignment Graph Neural Network Security Model for Vehicles PredictionabstractTraffic flow forecasting is indispensable in today’s society and regarded as a key problem for Intelligent Transportation Systems (ITS), as emergency delays in vehicles can cause serious traffic security accidents. However, the complex dynamic spatial-temporal dependency and correlation between different locations on the road make it a challenging task for security in transportation. To date, most existing forecasting frames make use of graph convolution to model the dynamic spatial-temporal correlation of vehicle transportation data, ignoring semantic similarity between nodes and thus, resulting in accuracy degradation. In addition, traffic data does not strictly follow periodicity and hard to be captured. To solve the aforementioned challenging issues, we propose in this article CRFAST-GCN, a multi-branch spatial-temporal attention graph convolution network. First, we capture the multi-scale (e.g., hour, day, and week) long- short-term dependencies through three identical branches, then introduce conditional random field (CRF) enhanced graph convolution network to capture the semantic similarity globally, so then we exploit the attention mechanism to captures the periodicity. For model evaluation using two real-world datasets, performance analysis shows that the proposed CRFAST-GCN successfully handles the complex spatial-temporal dynamics effectively and achieves improvement over the baselines at 50% (maximum), outperforming other advanced existing methods. Chunyan Diao, Da-Fang Zhang 0001, Wei Liang 0005, Kuanching Li, Yujie Hong, Jean-Luc Gaudiot |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Concept drift detection for distributed multi-model machine learning systemsabstractMany works focus on optimizing machine learning models during their training phase, but fail to account how these models adapt into their model-serving phase once they are deployed into real world applications. In this phase models must process through streams of data that can evolve over time and distort the relationship between incoming data, causing concept drift. This paper proposes leveraging the advantages of emerging features stores in order to improve concept drift detection on unlabeled, dynamic data streams across multiple models. Firstly, we introduce Drift Detection on Distributed Datasets (QuaD), which combines classical drift detectors to make use of labeled and unlabeled data, and create local context (i.e. per live model) and global context (i.e. across multiple models). Secondly, we propose using feature store entities, SHAP values, and Collaborative Filtering (CF) to augment unlabeled data across multiple models. To the best of our knowledge, QuaD is the first work that examines the collective behavior of concept drift across multiple models and discerns associations between models that may share a susceptibility in a dynamic setting. QuaD uses a combination of performance-based and data distribution-based drift detectors and CF to capture varying types of concept drifts for labeled and unlabeled data streams and is modeled around the data abstraction provided by emerging feature stores. Beverly Abadines Quon, Jean-Luc Gaudiot |
COMPSAC | 2 |
| 2022 | Programming Autonomous Machines : Special Session PaperabstractOne key technical challenge in the age of autonomous machines is the programming of autonomous machines, which demands the synergy across multiple domains, including fundamental computer science, computer architecture, and robotics, and requires expertise from both academia and industry. This paper discusses the programming theory and practices tied to producing real-life autonomous machines, and covers aspects from high-level concepts down to low-level code generation in the context of specific functional requirements, performance expectation, and implementation constraints of autonomous machines. Shaoshan Liu, Xiaoming Li 0010, Tongsheng Geng, Stéphane Zuckerman, Jean-Luc Gaudiot |
EMSOFT | 5 |
| 2022 | TransMigrator: A Transformer-Based Predictive Page Migration Mechanism for Heterogeneous Memory
Songwen Pei, Yihuan Qian, Jie Tang 0003, Jean-Luc Gaudiot |
NPC | 5 |
| 2022 | Detecting Spectre Attacks Using Hardware Performance CountersabstractSpectre attacks can be catastrophic and widespread because they exploit common design flaws caused by the speculative capabilities in modern processors to leak sensitive data through side channels. Completely fixing the problem would require a redesign of the architecture for transient execution or the implementation of a new design on re-configurable hardware. However, such fixes cannot be backported to old machines with fixed hardware design. Completely replacing those machines will take a long time. Moreover, existing software patches may cause significant performance overhead. This paper proposes to detect Spectre by monitoring deviations in microarchitectural events using hardware performance counters with promising accuracy above 90% under a variety of workload conditions. However, the attacker may attempt to evade detection by slowing down the attack or mimicking benign programs. This paper thus compares different evasion strategies quantitatively and demonstrates that it is possible for the attacker to avoid detection when operating the attacks at a lower speed while maintaining a reasonable attack success rate. Then, we show that, in order to resist evasion, the original detector must be enhanced by randomly switching between a set of detectors using different features and sampling periods so we can keep the detection accuracy above 80%. Congmiao Li, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2021 | Streaming Data Priority Scheduling Framework for Autonomous Driving by EdgeabstractIn recent years, intelligent vehicles like autonomous vehicles generate a huge amount of sensing data continuously. The computations on those data streams are far beyond the processing capacity of on-board computing. To deal with the streaming data process in real-time, the deployment of streaming data processing system by edge turns to the first choice in terms of performance. However, the existing frameworks cannot satisfy the complicated demands from autonomous driving tasks and lack the ability in supporting the task priority scheduling. In this paper, we propose a streaming data priority scheduling framework for autonomous driving by edge on Spark Streaming and make an implementation on Spark 2.3.0. The proposed framework can identify the priorities among different data processing tasks and implement the task scheduling based on non-preemptive priority queuing theory. To meet differentiated service level requirements, the proposed non-preemptive priority queuing scheduling mechanism considers the priority category of tasks, the distance between vehicles and edge nodes, and the priority weight of vehicles. Experiments show that this mechanism can effectively identify the priority information of different tasks from different vehicles and reduce the end-to-end latency of high-priority tasks by up to 46% than low-priority tasks. Lingbing Yao, Hang Zhao 0016, Jie Tang 0003, Shaoshan Liu, Jean-Luc Gaudiot |
COMPSAC | 5 |
| 2021 | Genetic scheduling policy on codelet modelabstractSummary The Codelet Model is a fine‐grained event‐driven hybrid parallel model inspired by dataflow, whose computing performance depends on the scheduling policy. An approximate optimal codelet scheduling policy based on the features of the task graphs is important to accelerate the performance of dataflow computer system. Therefore, we have proposed an adaptive genetic scheduling policy (GSP) for codelet by improving a “pure” genetic algorithm (IPGA) for given tasks with complex dependencies. It is verified that the genetic scheduling policy is effective according to bunches of experimental results. Songwen Pei, Linhua Jiang, Naixue Xiong, Jean-Luc Gaudiot |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | A novel data representation framework based on nonnegative manifold regularisationabstractRepresentation learning techniques have been frequently applied in multimedia content analysis and retrieval. In this study, an efficient multimedia data clustering method is presented, which consists of two independent algorithms. First, we propose a new representation framework by incorporating sparse coding and manifold regularisation in an optimisation objective function, the cluster indicator matrix is estimated by introducing ℓ1 sparsity norm coarsely. Second, we refine the estimated cluster indicator matrix by performing spectral rotation such that an optimal assignment for clustering can be learned. Compared with existing methods, we have the following merits: our method takes into account the global matrix reconstruction information and locality manifold information simultaneously. Therefore, global and locality information both are respected. Additionally, theoretical justification about the novel representation method is presented in this study. Comprehensive experiments demonstrate the effectiveness and efficiency of our method in comparison with the state-of-the-art clustering methods on six real-world image datasets. Wei Liang 0005, Jintian Tang, Hongbo Zhou 0012, Kuanching Li, Jean-Luc Gaudiot |
Connect. Sci. | 6 |
| 2020 | PDAWL: Profile-Based Iterative Dynamic Adaptive WorkLoad Balance on Heterogeneous Architectures
Tongsheng Geng, Marcos Amaris, Stéphane Zuckerman, Alfredo Goldman, Guang R. Gao, Jean-Luc Gaudiot |
JSSPP | 6 |
| 2020 | π-Hub: Large-scale video learning, storage, and retrieval on heterogeneous hardware platforms
Jie Tang 0003, Shaoshan Liu, Jie Cao 0003, Bolin Ding, Jean-Luc Gaudiot, Weisong Shi |
Future Gener. Comput. Syst. | 6 |
| 2020 | Secure Data Storage and Recovery in Industrial Blockchain Network EnvironmentsabstractThe massive redundant data storage and communication in network 4.0 environments have issues of low integrity, high cost, and easy tampering. To address these issues, in this article, a secure data storage and recovery scheme in the blockchain-based network is proposed by improving the decentration, tampering-proof, real-time monitoring, and management of storage systems, as such design supports the dynamic storage, fast repair, and update of distributed data in the data storage system of industrial nodes. A local regenerative code technology is used to repair and store data between failed nodes while ensuring the privacy of user data. That is, as the data stored are found to be damaged, multiple local repair groups constructed by vector code can simultaneously yet efficiently repair multiple distributed data storage nodes. Based on the unique chain storage structure, such as data consensus mechanism and smart contract, the storage structure of blockchain distributed coding not only quickly repair the nearby local regenerative codes in the blockchain but also reduce the resource overhead in the data storage process of industrial nodes. Experimental results show that the proposed scheme improves the repair rate of multinode data by 9% and data storage rate increased by 8.6%, indicating to be promising with good security and real-time performance. Wei Liang 0005, Yongkai Fan, Kuanching Li, Da-Fang Zhang 0001, Jean-Luc Gaudiot |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Detecting Malicious Attacks Exploiting Hardware Vulnerabilities Using Performance CountersabstractOver the past decades, the major objectives of computer design have been to improve performance and to reduce cost, energy consumption, and size, while security has remained a secondary concern. Meanwhile, malicious attacks have rapidly grown as the number of Internet-connected devices, ranging from personal smart embedded systems to large cloud servers, have been increasing. Traditional antivirus software cannot keep up with the increasing incidence of these attacks, especially for exploits targeting hardware design vulnerabilities. For example, as DRAM process technology scales down, it becomes easier for DRAM cells to electrically interact with each other. For instance, in Rowhammer attacks, it is possible to corrupt data in nearby rows by reading the same row in DRAM. As Rowhammer exploits a computer hardware weakness, no software patch can completely fix the problem. Similarly, there is no efficient software mitigation to the recently reported attack Spectre. The attack exploits microarchitectural design vulnerabilities to leak protected data through side channels. In general, completely fixing hardware-level vulnerabilities would require a redesign of the hardware which cannot be backported. In this paper, we demonstrate that by monitoring deviations in microarchitectural events such as cache misses, branch mispredictions from existing CPU performance counters, hardware-level attacks such as Rowhammer and Spectre can be efficiently detected during runtime with promising accuracy and reasonable performance overhead using various machine learning classifiers. Congmiao Li, Jean-Luc Gaudiot |
COMPSAC (1) | 2 |
| 2018 | Teaching Autonomous Driving Using a Modular and Integrated ApproachabstractIntroduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving. Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot |
COMPSAC (1) | 7 |
| 2018 | Online Detection of Spectre Attacks Using Microarchitectural Traces from Performance CountersabstractTo improve processor performance, computer architects have adopted such acceleration techniques as speculative execution and caching. However, researchers have recently discovered that this approach implies inherent security flaws, as exploited by Meltdown and Spectre. Attacks targeting these vulnerabilities can leak protected data through side channels such as data cache timing by exploiting mis-speculated executions. The flaws can be catastrophic because they are fundamental and widespread and they affect many modern processors. Mitigating the effect of Meltdown is relatively straightforward in that it entails a software-based fix which has already been deployed by major OS vendors. However, to this day, there is no effective mitigation to Spectre. Fixing the problem may require a redesign of the architecture for conditional execution in future processors. In addition, a Spectre attack is hard to detect using traditional software-based antivirus techniques because it does not leave traces in traditional log files. In this paper, we proposed to monitor microarchitectural events such as cache misses, branch mispredictions from existing CPU performance counters to detect Spectre during attack runtime. Our detector was able to achieve 0% false negatives with less than 1 % false positives using various machine learning classifiers with a reasonable performance overhead. Congmiao Li, Jean-Luc Gaudiot |
SBAC-PAD | 2 |
| 2017 | Return of experience on the mean-shift clustering for heterogeneous architecture use caseabstractThe exponential increment in data size poses new challenges for computer scientists, giving rise to a new set of methodologies under the term Big Data. Many efficient algorithms for machine learning have been proposed, facing up time and memory requirements. Nevertheless, with hardware acceleration, multiple software instructions can be integrated and executed into a single hardware die. Current researches aim at eliminating the burden for the user in using multiple processor types. In this paper we propose our return of experience on a new way of implementing machine learning algorithms on heterogeneous hardware. To explore our vision, we use a parallel Mean-shift algorithm, developed at LIPN as our case study to investigate issues in building efficient Machine Learning libraries for heterogeneous systems. The ultimate goal is to provide a core set of building blocks for Machine Learning programming that could serve either to build new applications on heterogeneous architectures or to control the evolution of the underlying platform. We thus examine the difficulties encountered during the implementation of the algorithm with the aim to discover methodologies for building systems based on heterogeneous hardware. We also discover issues and building blocks for solving concrete machine learning (ML) problems on the Chisel software stack we use for this purpose. Christophe Cérin, Jean-Luc Gaudiot, Mustapha Lebbah, Foutse Yuehgoh |
IEEE BigData | 2 |
| 2016 | PETS: Performance, energy and thermal aware scheduler for job mapping with resource allocation in heterogeneous systemsabstractMany computing systems today are heterogeneous in that they consist of a mix of different types of processing units (e.g., CPUs, GPUs). Each of these processing units has different performance capabilities and energy consumption characteristics. Job mapping and scheduling play a crucial role in such systems as they strongly affect the overall system performance, energy consumption, peak power and peak temperature. Allocating resources (e.g., core scaling, threads allocation) is another challenge since different sets of resources exhibit different behavior in terms of performance and energy consumption. Many studies have been conducted on job scheduling with an eye on performance improvement. However, few of them consider both performance and energy. We thus propose our approach we call Performance, Energy and Thermal aware Scheduler (PETS) to combine job mapping, core scaling, and threads allocation into one scheduler. We apply an evolutionary algorithm (a Genetic Algorithm - GA) to find an efficient job schedule in terms of both execution time and energy consumption, under peak power and peak temperature constraints. Compared to a performance-based GA and other schedulers, on average, our PETS scheduler can achieve up to a 4.7x of speedup and an energy saving of up to 195%. Shouq Alsubaihi, Jean-Luc Gaudiot |
IPCCC | 2 |
| 2015 | How can Garbage Collection be energy efficient by dynamic offloading?abstractGarbage Collection (GC) is still a major issue in JVM for both mobile and cluster computing. GC offloading is proposed to improve the performance of GC by delivering part or all of the operations into another dedicated GC hardware. However, the traditional offloading just offloads directly not considering the phase change of GC behavior, which can be classified into two different groups: minor GC and major GC. The minor GC is fast and frequently invoked, while major GC is expensive in terms of time but seldom takes place. The direct offloading made GC workload frequently hopping between main processor and GC hardware, introduced a noticeable overhead and offset any possible benefits of workload loading. To solve this issue, we propose to offload GC dynamically by a careful selection of profitable and harmful GC operations. We also made a case study on Apache Spark, a lightning-fast cluster computing platform. It shows dynamic offloading can yield nearly 42.6% performance improvement with a concurrent 32.1% in energy cost reduction. Jie Tang 0003, Chen Liu 0001, Jean-Luc Gaudiot |
ASAP | 3 |
| 2015 | Design of configurable I/O pin control block for improving reusability in multimedia SoC platforms
Myoung-Seo Kim, Cheong-Ghil Kim, Shin-Dug Kim, Jean-Luc Gaudiot |
Multim. Tools Appl. | 4 |
| 2015 | Network Variation and Fault Tolerant Performance Acceleration in Mobile Devices with Simultaneous Remote ExecutionabstractAs mobile applications provide increasingly richer features to end users, it has become imperative to overcome the constraints of a resource-limited mobile hardware. Remote execution is one promising technique to resolve this important problem. Using this technique, the computation intensive part of the workload is migrated to resource-rich servers, and then once the computation is completed, the results can be returned to the client devices. To enable this operation, strong wireless connectivity is required. However, unstable wireless connections are the staple of real-life. This makes performance unpredictable, sometimes offsetting the benefits brought by this technique and leading to performance degradation. To address this problem, in this paper, we present a Simultaneous Remote Execution (SRE) model for mobile devices. Our SRE model performs concurrent executions both locally and remotely. Therefore, the worst-case execution time on fluctuating network condition is significantly reduced. In addition, SRE provides inherent tolerance for abrupt network failure. We designed and implemented an SRE-based offloading system consisting of a real smartphone and a remote server connected via 3G and Wifi networks. The experimental results under various real-life network variation scenarios show that SRE outperforms the alternative schemes in highly fluctuating network environments. Keunsoo Kim, Benjamin Y. Cho, Won Woo Ro, Jean-Luc Gaudiot |
IEEE Trans. Computers | 4 |
| 2014 | How many cores do we need to run a parallel workload: A test drive of the Intel SCC platform?
Chen Liu 0001, Pollawat Thanarungroj, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 3 |
| 2014 | Complexity-Effective Contention Management with Dynamic Backoff for Transactional Memory SystemsabstractReducing memory access conflicts is a crucial part of the design of Transactional Memory (TM) systems since the number of running threads increases and long latency transactions gradually appear: without an efficient contention management, there will be repeated aborts and wasteful rollback operations. In this paper, we present a dynamic backoff control algorithm developed for complexity-effective and distributed contention management in Hardware Transactional Memory (HTM) systems. Our approach aims at controlling the restarting intervals of aborted transactions, and can be easily applied to the various TM systems. To this end, we have profiled the applications of the STAMP benchmark suite and have identified those “problem” transactions which repeatedly cause aborts in the applications with the attendant high contention rate. The proposed algorithm alleviates the impact of these repeated aborts by dynamically adjusting the initial exponent value of the traditional backoff approach. In addition, the proposed scheme decreases the number of wasted cycles down to 82% on average compared to the baseline TM system. Our design has been integrated in LogTM-SE where we observed an average performance improvement of 18%. Dongmin Choi, Won Woo Ro, Jean-Luc Gaudiot |
IEEE Trans. Computers | 4 |
| 2014 | $C\!\!-\!\!Lock$ : Energy Efficient Synchronization for Embedded Multicore SystemsabstractData synchronization among multiple cores has been one of the critical issues which must be resolved in order to optimize the parallelism of multicore architectures. Data synchronization schemes can be classified as lock-based methods (“pessimistic”) and lock-free methods (“optimistic”). However, none of these methods consider the nature of embedded systems which have demanding and sometimes conflicting requirements not only for high performance, but also for low power consumption. As an answer to these problems, we propose$C\!\!- \!\! Lock$, an energy- and performance-efficient data synchronization method for multicore embedded systems.$C\!\!- \!\! Lock$achieves balanced energy- and performance-efficiency by combining the advantages of lock-based methods and transactional memory (TM) approaches; in$C\!\!- \!\! Lock$, the core is blocked only when true conflicts exist (advantage of TM), while avoiding roll-back operations which can cause huge overhead with regard to both performance and energy (this is an advantage of locks). Also, in order to save more energy,$C\!\!- \!\! Lock$disables the clocks of the cores which are blocked for the access to the shared data until the shared data become available. We compared our$C\!\!- \!\! Lock$approach against traditional locks and transactional memory systems and found that$C\!\!- \!\! Lock$can reduce the energy-delay product by up to 1.94 times and 13.78 times compared to the baseline and TM, respectively. Sang Hyong Lee, Minje Jun, Byunghoon Lee, Won Woo Ro, Eui-Young Chung, Jean-Luc Gaudiot |
IEEE Trans. Computers | 7 |
| 2013 | Mark-Sharing: A Parallel Garbage Collection Algorithm for Low Synchronization OverheadabstractTwo main problems prevent a parallel garbage collection (GC) scheme with lock-based synchronization from providing a high level of scalability: the load imbalance and the runtime overhead of thread synchronization operations. These problems become even more serious as the number of available threads increases. We propose the Mark-Sharing algorithm to improve the performance of parallel GC using transactional memory (TM) systems. The Mark-Sharing algorithm guarantees that all threads access the shared resource by using both the task-stealing and task-releasing mechanisms appropriately. In addition, we introduce a selection manager that minimizes the contention and idle time of garbage collectors by maintaining task information. The proposed algorithm outperforms the prior pool-sharing algorithm of GC in the HTM, providing more than 90% performance improvement on average. Hyunkyu Park 0004, Changmin Lee 0002, Won Woo Ro, Jean-Luc Gaudiot |
ICPADS | 5 |
| 2013 | Pinned OS/Services: A Case Study of XML Parsing on Intel SCC
Jie Tang 0003, Pollawat Thanarungroj, Chen Liu 0001, Shaoshan Liu, Zhimin Gu, Jean-Luc Gaudiot |
J. Comput. Sci. Technol. | 6 |
| 2013 | Acceleration of XML Parsing through PrefetchingabstractExtensible Markup Language (XML) has become a widely adopted standard for data representation and exchange. However, its features also introduce significant overhead threatening the performance of modern applications. In this paper, we present a study of XML parsing and determine that memory-side data loading in the parsing stage incurs a significant performance overhead, as much as the computation does. Hence, we propose memory-side acceleration which incorporates of data prefetching techniques, and can be applied on top of computation-side acceleration to speed up the XML data parsing. To this end, we study here the impact of our proposed scheme on the performance and energy consumption and demonstrated how it is capable of improving performance by up to 20 percent as well as produce up to 12.77 percent of energy saving when implemented in 32-nm technology. In addition, we implement a prefetcher on an platform in an effort to evaluate its implementation feasibility in terms of area and energy overhead. Jie Tang 0003, Shaoshan Liu, Chen Liu 0001, Zhimin Gu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 5 |
| 2013 | Importance of Coherence Protocols with Network Applications on Multicore ProcessorsabstractAs Internet and information technology have continued developing, the necessity for fast packet processing in computer networks has also grown in importance. All emerging network applications require deep packet classification as well as security-related processing and they should be run at line rates. Hence, network speed and the complexity of network applications will continue increasing and future network processors should simultaneously meet two requirements: high performance and high programmability. We will show that the performance of single processors will not be sufficient to support future demands. Instead, we will have to turn to multicore processors, which can exploit the parallelism in network workloads. In this paper, we focus on the cache coherence protocols which are central to the design of multicore-based network processors. We investigate the effects of two main categories of various cache coherence protocols with several network workloads on multicore processors. Our simulation results show that token protocols have a significantly higher performance than directory protocols. With an 8-core configuration, token protocols improves the performance compared to directory protocols by a factor of nearly 4 on average. Kyueun Yi, Won Woo Ro, Jean-Luc Gaudiot |
IEEE Trans. Computers | 3 |
| 2013 | Design and evaluation of random linear network coding Accelerators on FPGAs
Won Seob Jeong, Won Woo Ro, Jean-Luc Gaudiot |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | Achieving energy efficiency through runtime partial reconfiguration on reconfigurable systemsabstractOne major advantage of reconfigurable computing systems is their ability to reconfigure hardware at runtime. In this paper, we study the feasibility of achieving energy efficiency in reconfigurable computing systems (e.g., FPGAs) through runtime partial reconfiguration (PR) techniques. In the ideal scenario, we use a hardware accelerator to accelerate certain parts of the program execution; when the accelerator is not active, we use partial reconfiguration to unload it to reduce power consumption. Since the reconfiguration process may introduce a high energy overhead, it is unclear whether this approach is efficient. To approach this problem, we first analytically identify the conditions under which partial reconfiguration can reduce energy consumption. Our results indicate that the key to reduce partial reconfiguration energy overhead is to minimize the time overhead of the reconfiguration process. Based on this analysis, we design and implement a fast reconfiguration engine that achieves close-to-ideal throughput on Xilinx Virtex-4 FPGAs. Our fast reconfiguration engine utilizes a master-slave DMA pair to stream data between the SRAM and the Internal Configuration Access Port (ICAP). We experimentally verify our proposed solutions and compare our design to existing energy reduction techniques, such as clock gating. The results of our study show that by using partial reconfiguration to eliminate the power consumption of the accelerator when it is inactive, we can accelerate program execution and at the same time reduce the overall energy consumption by half. Shaoshan Liu, Richard Neil Pittman, Alessandro Forin, Jean-Luc Gaudiot |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | Parallel Sparse Approximate Inverse Preconditioning on Graphic Processing UnitsabstractAccelerating numerical algorithms for solving sparse linear systems on parallel architectures has attracted the attention of many researchers due to their applicability to many engineering and scientific problems. The solution of sparse systems often dominates the overall execution time of such problems and is mainly solved by iterative methods. Preconditioners are used to accelerate the convergence rate of these solvers and reduce the total execution time. Sparse approximate inverse (SAI) preconditioners are a popular class of preconditioners designed to improve the condition number of large sparse matrices. We propose a GPU accelerated SAI preconditioning technique called GSAI, which parallelizes the computation of this preconditioner on NVIDIA graphic cards. The preconditioner is then used to enhance the convergence rate of the BiConjugate Gradient Stabilized (BiCGStab) iterative solver on the GPU. The SAI preconditioner is generated on average 28 and 23 times faster on the NVIDIA GTX480 and TESLA M2070 graphic cards, respectively, compared to ParaSails (a popular implementation of SAI preconditioners on CPU) single processor/core results. The proposed GSAI technique computes the SAI preconditioner in approximately the same time as ParaSails generates the same preconditioner on 16 AMD Opteron 252 processors. Maryam Mehri Dehnavi, David M. Fernandez, Jean-Luc Gaudiot, Dennis Giannacopoulos |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Synchronization-Aware Energy Management for VFI-Based Multicore Real-Time SystemsabstractVoltage and frequency island (VFI) was recently adopted as an effective energy management technique for multicore processors. For a set of periodic real-time tasks that access shared resources running on a VFI-based multicore system with dynamic voltage and frequency scaling (DVFS) capability, we study both static and dynamic synchronization-aware energy management schemes. First, based on the enhanced MSRP resource access protocol with a suspension mechanism, we devise a synchronization-aware task mapping heuristic for partitioned-EDF scheduling, which assigns tasks that access similar set of resources to the same core to reduce the synchronization overhead and thus improve schedulability. Then, static schemes that assign uniform and different scaled frequencies for tasks on different VFIs are studied. To further exploit dynamic slack, we propose an integrated synchronization-aware slack management framework to appropriately reclaim, preserve, release and steal slack at runtime to slow down the execution of tasks subject to the common voltage/frequency limitation of VFIs and timing/synchronization constraints of tasks. Taking the additional delay due to task synchronization into consideration, the new scheme allocates slack in a fair manner and scales down the execution of both noncritical and critical sections of tasks for more energy savings. Simulation results show that, the synchronization-aware mapping can significantly improve the schedulability of tasks. The energy savings obtained by the static scheme with different frequencies for tasks on different VFIs is close to that of an optimal Integer Nonlinear Programming (INLP) solution. Moreover, compared to the simple extension of existing solutions for uniprocessor systems, our schemes can obtain much better energy savings (up to 40 percent) with comparable DVFS overhead. Jian-Jun Han, Dakai Zhu 0001, Hai Jin 0001, Laurence T. Yang, Jean-Luc Gaudiot |
IEEE Trans. Computers | 6 |
| 2012 | Packer: Parallel Garbage Collection Based on Virtual SpacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and normal object space (non-LOS). When either space is full, garbage collection is triggered even though the other space may still have plenty of room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection inefficient on multiple spaces. To address these problems, we propose Packer, a parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physical space. With multiple virtual spaces, Packer offers efficient memory management. With one physical space, Packer avoids the problem of an inefficient space utilization. To reduce the garbage collection pause time, we also propose a novel parallelization method that is applicable to multiple virtual spaces. Specifically, we reduce the compacting GC parallelization problem into a discreted acyclic graph (DAG) traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Jie Tang 0003, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IEEE Trans. Computers | 5 |
| 2012 | Minimizing the runtime partial reconfiguration overheads in reconfigurable systems
Shaoshan Liu, Richard Neil Pittman, Alessandro Forin, Jean-Luc Gaudiot |
J. Supercomput. | 4 |
| 2012 | Achieving middleware execution efficiency: hardware-assisted garbage collection operationsabstractAlthough virtualization technologies bring many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we achieved an order of magnitude speedup and significant improvement on energy efficiency. In addition, the results of our performance estimation study indicate that the hardware-assisted GC instructions can reduce the GC execution time by half and lead to a 7% improvement on the overall execution time. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
J. Supercomput. | 5 |
| 2011 | Keynote talk: Fighting Amdahl's law in many-core and GPU parallel architectures with value predictionabstractSummary form only given. The next grail sought by HPC community is the exascale, 100 times the current scale. This target will not be reached easily as many challenges are uprising. The first challenge, the Energy consumption, has become a strict constraint now with a limit set to 20MW (twice as the current top supercomputers). Multiplying the computing elements will imply to drastically reduce the power consumption of each of them. The second challenge will be to keep it cool as: first the overall power envelope, 20MW, include the energy for cooling and second, because 20MW will be turned into heat by joule effect. And the operating temperature of electronic must be bounded otherwise, the leakage (and thus the power consumption) increases and the reliability decreases. This brings us to a third challenge regarding the reliability of the machine, the number of components will be tremendous, thus, the probability of having failing ones will increase. It has to be managed in such a way that applications will not be impacted by the failures. Finally, The last challenge is related to the software stack of these supercomputers, how will we manage billions of threads, how will we debug it, … New paradigms are currently being studied, for instance Bag of tasks, that try to tackle these aspects. These are the challenges we have to solve!! In this presentation, brightened up with insight into Bull roadmap, we present a possible future. Jean-Luc Gaudiot |
AICCSA | 1 |
| 2011 | Memory-Side Acceleration for XML Parsing
Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Chen Liu 0001, Jean-Luc Gaudiot |
NPC | 5 |
| 2011 | Workload Characterization of Cryptography Algorithms for Hardware AccelerationabstractData encryption/decryption has become an essential component for modern information exchange. However, executing these cryptographic algorithms is often associated with huge overhead and the need to reduce this overhead arises correspondingly. In this paper, we select nine widely adopted cryptography algorithms and study their workload characteristics. Different from many previous works, we consider the overhead not only from the perspective of computation but also focusing on the memory access pattern. We break down the function execution time to identify the software bottleneck suitable for hardware acceleration. Then we categorize the operations needed by these algorithms. In particular, we introduce a concept called 'Load-Store Block' (LSB) and perform LSB identification of various algorithms. Our results illustrate that for cryptographic algorithms, the execution rate of most hotspot functions is more than 60%; memory access instruction ratio is mostly more than 60%; and LSB instructions account for more than 30% for selected benchmarks. Based on our findings, we suggest future directions in designing either the hardware accelerator associated with microprocessor or specific microprocessor for cryptography applications. Jed Kao-Tung Chang, Chen Liu 0001, Shaoshan Liu, Jean-Luc Gaudiot |
ICPE | 4 |
| 2011 | Reducing Power in All Major CAM and SRAM-Based Processor Units via Centralized, Dynamic Resource Size ManagementabstractPower minimization has become a primary concern in microprocessor design. In recent years, many circuit and micro-architectural innovations have been proposed to reduce power in many individual processor units. However, many of these prior efforts have concentrated on the approaches which require considerable redesign and verification efforts. Also it has not been investigated whether these techniques can be combined. Therefore a challenge is to find a centralized and simple algorithm which can address power issues for more than one unit, and ultimately the entire chip and comes with the least amount of redesign and verification efforts, the lowest possible design risk and the least hardware overhead. This paper proposes such a centralized approach that attempts to simultaneously reduce power in processor units with highest dissipation: reorder buffer, instruction queue, load/store queue, and register files. It is based on an observation that utilization for the aforementioned units varies significantly, during cache miss period. Therefore we propose to dynamically adjust the size and thus power dissipation of these resources during such periods. Circuit level modifications required for such resource adaptation are presented. Simulation results show a substantial power reduction at the cost of a negligible performance impact and a small hardware overhead. Houman Homayoun, Avesta Sasan, Jean-Luc Gaudiot, Alexander V. Veidenbaum |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | On energy efficiency of reconfigurable systems with run-time partial reconfigurationabstractIn this paper we study whether partial reconfiguration can be used to reduce FPGA energy consumption. In an ideal scenario, we will have a hardware accelerator to assist with certain parts of program execution. When the accelerator is not active, we use partial reconfiguration to unload it to reduce both static and dynamic power. However, the reconfiguration process may introduce a high energy overhead, thus it is unclear whether this approach is feasible. To approach this problem, we identify the conditions under which partial reconfiguration can be used to reduce energy consumption, and we propose solutions to minimize the configuration energy overhead. The results of our study show that by using partial reconfiguration to reduce the power consumption of the accelerator when it is inactive, we can accelerate program execution and at the same time halve the overall energy consumption. Shaoshan Liu, Richard Neil Pittman, Alessandro Form, Jean-Luc Gaudiot |
ASAP | 4 |
| 2010 | Hardware-assisted middleware: Acceleration of garbage collection operationsabstractAlthough the virtualization technology brings many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we managed to achieve an order of magnitude speedup. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
ASAP | 5 |
| 2010 | A Theoretical Framework for Value Prediction in Parallel SystemsabstractWe present here a theoretical framework towards a fundamental understanding of the effects of value prediction. Our framework consists of two parts: first, an identification of the theoretical limit of value prediction and an indication of the potential to improve parallelism through the exploitation of value predictability; second, a demonstration of the feasibility of data prediction and a theoretical support to verify this feasibility. The experiment results demonstrate the immense potential of value prediction in enhancing the performance of many-core architectures. Shaoshan Liu, Christine Eisenbeis, Jean-Luc Gaudiot |
ICPP | 3 |
| 2010 | Speculative Execution on GPU: An Exploratory StudyabstractWe explore the possibility of using GPUs for speculative execution: we implement software value prediction techniques to accelerate programs with limited parallelism, and software speculation techniques to accelerate programs that contain runtime parallelism, which are hard to parallelize statically. Our experiment results show that due to the relatively high overhead, mapping software value prediction techniques on existing GPUs may not bring any immediate performance gain. On the other hand, although software speculation techniques introduce some overhead as well, mapping these techniques to existing GPUs can already bring some performance gain over CPU. Shaoshan Liu, Christine Eisenbeis, Jean-Luc Gaudiot |
ICPP | 3 |
| 2010 | Hardware-assisted security mechanism: The acceleration of cryptographic operations with low hardware costabstractThis paper presents generic cryptographic accelerator. Certain "hotspot function" are found in the cryptographic algorithm which consume a substantial amount of execution time of the specific algorithm. INTEL performance analyzer VTune was used which analyzes the software performance on IA-32 and Intel64-based machines to examne the hotsport function. By moving the operations to hardware, we can reduce the overheads introduced by the crypto-computation so that the computing resource can focus on the useful work. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
IPCCC | 3 |
| 2010 | Power Efficient Scheduling for Hard Real-Time Systems on a Multiprocessor Platform
Peter J. Nistler, Jean-Luc Gaudiot |
NPC | 2 |
| 2010 | Energy-Efficient Scheduling of Real-Time Periodic Tasks in Multicore Systems
Jian-Jun Han, Jean-Luc Gaudiot |
NPC | 4 |
| 2010 | The Performance Analysis and Hardware Acceleration of Crypto-computations for Enhanced SecurityabstractSecurity is very important in modern life due to most information is now stored in digital format. A good security mechanism will keep information secrecy and integrity, hence, plays an important role in modern information exchange. However, cryptography algorithms are extremely expensive in terms of execution time. To make data not easily being cracked, many arithmetic and logical operations will be executed in the encryption/decryption process with many data movement. This means the cryptographic applications are both computation and memory intensive. Using a general-purpose processor for this scenario would not be very cost-effective. This study addresses this problem. Compared to the previous designs, we used a performance analyzer to identify “hotspot” functions across a set of benchmarks. The hotspot function consumes a substantial amount of the execution time of the specific algorithm. Then we translate these hotspot functions into hardware accelerators to improve the performance. Overall we achieve 34 - 83 folds of speedup. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
PRDC | 3 |
| 2010 | Network Applications on Simultaneous Multithreading ProcessorsabstractAs network applications become increasingly sophisticated and Internet traffic is getting heavier, future network processors must continue processing computation-intensive network applications at line rates. Most programmable network processors on the market today, such as the Intel IXP2800, target relatively low performance (from 100 Mbps to 10 Gbps). However, low cost edge routers will find it hard to cope with the forthcoming sophistication of network applications to be processed at those speeds. Hence, new architectures should be designed for the programmable network processors of the future. The goal of this paper is to evaluate the applicability and efficiency of Simultaneous MultiThreaded (SMT) as the base architecture of a network processor. Indeed, the SMT model inherently allows the multiple parallel threads which must be dealt with in network processor applications. In this paper, we investigate the architectural implications of network applications on the SMT architecture. We demonstrate that, when executed as independent threads, applications chosen from different network layers show an improved Instructions Per Cycle (IPC) and cache behavior when compared with the situation where the program executed comes from a single network application. Finally, a new architectural solution to cope with packet dependency is proposed and evaluated. Kyueun Yi, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2009 | The Impact of Resource Sharing Control on the Design of Multicore Processors
Chen Liu 0001, Jean-Luc Gaudiot |
ICA3PP | 2 |
| 2009 | Packer: An innovative space-time-efficient parallel garbage collection algorithm based on virtual spacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and non-large object space (non-LOS). When one of the spaces is full, garbage collection is triggered even though the other space may still have a lot of free room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection not efficient on multiple spaces. To address these problems, we propose Packer, a space-and-time-efficient parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physically shared space. With multiple virtual spaces, Packer offers the advantage of efficient memory management. At the same time, with one physically shared space, Packer avoids the problem of inefficient space utilization. To reduce the garbage collection pause time of Packer, we also propose a novel parallelization method that is applicable to multiple virtual spaces. We reduce the compacting GC parallelization problem into a tree traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IPDPS | 4 |
| 2009 | A complexity-effective microprocessor design with decoupled dispatch queues and prefetching
Won Woo Ro, Jean-Luc Gaudiot |
Parallel Comput. | 2 |
| 2009 | Potential Impact of Value Prediction on Communication in Many-Core ArchitecturesabstractThe newly emerging many-core-on-a-chip designs have renewed an intense interest in parallel processing. By applying Amdahl's formulation to the programs in the PARSEC and SPLASH-2 benchmark suites, we find that most applications may not have sufficient parallelism to efficiently utilize modern parallel machines. The long sequential portions in these application programs are caused by computation as well as communication latency. However, value prediction techniques may allow the ldquoparallelizationrdquo of the sequential portion by predicting values before they are produced. In conventional superscalar architectures, the computation latency dominates the sequential sections. Thus, value prediction techniques may be used to predict the computation result before it is produced. In many-core architectures, since the communication latency increases with the number of cores, value prediction techniques may be used to reduce both the communication and computation latency. In this paper, we extend Amdahl's formulation to model the data redundancy inherent to each benchmark, thereby identifying the potential of value prediction techniques. Our analysis shows that the performance of PARSEC benchmarks may improve by a factor of 180 and 230 percent for the SPLASH-2 suite, compared to when only the intrinsic parallelism is considered. This demonstrates the immense potential of fine-grained value prediction in reducing the communication latency in many-core architectures. Shaoshan Liu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2009 | Special issue of Supercomputing Journal on secure, manageable and controllable grid services
Christophe Cérin, Jean-Luc Gaudiot, Kuanching Li |
J. Supercomput. | 2 |
| 2008 | Adaptive techniques for leakage power management in L2 cache peripheral circuitsabstractRecent studies indicate that a considerable amount of an L2 cache leakage power is dissipated in its peripheral circuits, e.g., decoders, word-lines and I/O drivers. In addition, L2 cache is becoming larger, thus increasing the leakage power. This paper proposes two adaptive architectural techniques (ADM and ASM) to reduce leakage in the L2 cache peripheral circuits. The adaptive techniques use the product of cache hierarchy miss rates to guide the leakage control in accordance with program behavior. The result for SPEC2K benchmarks show that the first technique (ASM) achieves a 34% average leakage power reduction with a 1.8% average IPC reduction. The second technique (ADM) achieves a 52% average savings with a 1.9% average IPC reduction. This corresponds to a 2 to 3 X improvement over recently proposed static techniques. Houman Homayoun, Alexander V. Veidenbaum, Jean-Luc Gaudiot |
ICCD | 3 |
| 2008 | Automatic object and image alignment using Fourier Descriptors
Weisheng Duan, Falko Kuester, Jean-Luc Gaudiot, Omar Hammami |
Image Vis. Comput. | 3 |
| 2008 | A low-complexity microprocessor design with speculative pre-execution
Won Woo Ro, Jean-Luc Gaudiot |
J. Syst. Archit. | 2 |
| 2008 | An Efficient Data-Distribution Mechanism in a Processor-In-Memory (PIM) Architecture Applied to Motion EstimationabstractIn general, the main purpose of using processor-in-memory (PIM) modules is to dramatically increase the data-level parallelism (DLP) and avoid the limited issue rate of current systems (even when they include SIMD extensions) caused by the limited data bandwidth and functional units. Our approach is to divide the PIM module into hundreds of smaller pieces so that each of these smaller PIMs can execute motion estimation for a group of macro blocks in a parallel fashion. We also design the logic in each PIM to execute in a highly pipelined fashion so that even more parallelism can be exploited. The main contribution of this paper is the presentation of architectural techniques that can be used in the PIM module to overcome the addressing and data sharing overhead when these smaller PIMs are used. Our architectural techniques have been applied to motion estimation. Indeed, it has been reported that motion estimation takes the majority of the execution time of MPEG encoding and it has been researched by many because of its importance in MPEG encoding. With our paradigm and techniques, the host processor can be relieved from the most computationally demanding and data-intensive portions of the workload, which should therefore yield a significant performance gain. Indeed, we observed (when 512 of these smaller PIMs were used) a reduction in the number of memory accesses by a factor of up to 2,034 times. At the same time, the performance improved by a multiplicative factor as high as 439 times. Jung-Yup Kang, Sandeep Gupta 0001, Jean-Luc Gaudiot |
IEEE Trans. Computers | 3 |
| 2007 | Architectural Support for Network Applications on Simultaneous MultiThreading ProcessorsabstractAs network applications become increasingly sophisticated and Internet traffic is getting heavier, future network processors must continue processing computation-intensive network applications at line rates. Most programmable network processors on the market today, such as the Intel IXP2800, target low performance (from 100 Mbps to 10 Gbps). However, low cost edge routers find it hard to cope with the forthcoming sophistication of network applications to be processed at those speeds. Hence, new architectures should be designed for the programmable network processors of the future. The goal of this paper is to evaluate the applicability and efficiency of simultaneous multi-threaded (SMT) as a network processor. Indeed, the SMT model inherently allows the multiple parallel threads which must be dealt with in network processor applications. In this paper, we investigate the architectural implications of network applications on the SMT architecture. We demonstrate that, when executed as independent threads, applications chosen from different network layers show an improved IPC and cache behavior when compared with the situation where the program executed comes from a single network application. Finally, a new architectural solution to cope with packet dependency is proposed and evaluated. Kyueun Yi, Jean-Luc Gaudiot |
IPDPS | 2 |
| 2007 | Architectural Implications of Cache Coherence Protocols with Network Applications on Chip MultiProcessorsabstractNetwork processors are specialized integrated circuits used to process packets in such network equipment as core routers, edge routers, and access routers. As predicted by Gilder’s law, Internet traffic has doubled each year since 1997 and this trend is showing no signs of abating. Since all emerging network applications which require deep packet classification and security-related processing should be run at line rates and since network speed and network applications complexity continue increasing, future network processors should simultaneously meet two requirements: high performance and high programmability. Single processor performance will not be sufficient to support the requirements which will be imposed on future network processors. In this paper, we consider the CMP model as the baseline architecture of future network processors. We investigate the architectural implications of cache coherence protocols with network workloads on CMPs. Our results show that the token protocol which uses the tokens to control read/write permission of shared data blocks shows better performance than the directory protocol by a factor of 13.4%. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Kyueun Yi, Jean-Luc Gaudiot |
NPC | 2 |
| 2006 | Design and Effectiveness of Small-Sized Decoupled Dispatch Queues
Won Woo Ro, Jean-Luc Gaudiot |
Euro-Par | 2 |
| 2006 | Low Power Microprocessor Design for Embedded Systems
Seong-Won Lee, Neungsoo Park, Jean-Luc Gaudiot |
ICCSA (4) | 3 |
| 2006 | The Walls of Computer Design
Jean-Luc Gaudiot |
ISPA | 1 |
| 2006 | Design Trade-Offs and Deadlock Prevention in Transient Fault-Tolerant SMT ProcessorsabstractSince the very concept of simultaneous multi-threading (SMT) entails inherent redundancy, some proposals have been made to run two copies of the same thread on top of SMT platforms in order to detect and correct soft errors. This allows, upon detection of an error, for the rolling back of the processor state to a known safe point, and then a retry of the instructions, thereby resulting in a completely error-free execution. This paper focuses on two crucial implementation issues introduced by this concept: (i) the design trade-off between the fault detection coverage versus the design costs; (ii) the possible occurrence of deadlock situations. To achieve the largest possible fault detection coverage, we replicate the instructions fetched in order to generate the redundant thread copies. Further, we apply the SMT thread scheduling at the instruction dispatch stage so as to lower the performance overhead. As a result, when compared to the baseline processor, our simulation results show that by using our two new schemes, the performance overhead can be reduced down to as little as 34% on the average, down from 42%. Finally, in the fault-tolerant execution mode, since the two copied threads are cooperating with one another, deadlock situations could be quite common. We thus present a detailed deadlock analysis and then conclude that allocating some entries of ROB, LQ, and SQ for the trailing thread is sufficient to prevent such deadlocks Jean-Luc Gaudiot |
PRDC | 2 |
| 2006 | Introduction
Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 1 |
| 2006 | Speculative pre-execution assisted by compiler (SPEAR)
Won Woo Ro, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 2 |
| 2006 | Adaptive dynamic thread scheduling for simultaneous multithreaded architectures with a detector thread
Chulho Shin, Seong-Won Lee, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 3 |
| 2006 | A Simple High-Speed Multiplier DesignabstractThe performance of multiplication is crucial for multimedia applications such as 3D graphics and signal processing systems, which depend on the execution of large numbers of multiplications. Previously reported algorithms mainly focused on rapidly reducing the partial products rows down to final sums and carries used for the final accumulation. These techniques mostly rely on circuit optimization and minimization of the critical paths. In this paper, an algorithm to achieve fast multiplication in two's complement representation is presented. Rather than focusing on reducing the partial products rows down to final sums and carries, our approach strives to generate fewer partial products rows. In turn, this influences the speed of the multiplication, even before applying partial products reduction techniques. Fewer partial products rows are produced, thereby lowering the overall operation time. In addition to the speed improvement, our algorithm results in a true diamond-shape for the partial product tree, which is more efficient in terms of implementation. The synthesis results of our multiplication algorithm using the Artisan TSMC 0.13um 1.2-Volt standard-cell library show 13 percent improvement in speed and 14 percent improvement in power savings for 8-bit \times 8-bit multiplications (10 percent and 3 percent, respectively, for 16-bit \times 16-bit multiplications) when compared to conventional multiplication algorithms. Jung-Yup Kang, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2006 | Throttling-Based Resource Management in High Performance Multithreaded ArchitecturesabstractUp to now, the power problems which could be caused by the huge amount of hardware resources present in modern systems have not been a primary concern. More recently, however, power consumption has begun limiting the number of resources which can be safely integrated into a single package, lest the heat dissipation exceed physical limits (before actual package meltdown). At the same time, new architectural techniques such as simultaneous multithreading (SMT), whose goal it is to efficiently use the resources of a superscalar machine without introducing excessive additional control overhead, have appeared on the scene. In this paper, we present a new resource management scheme which enables an efficient low power mode in SMT architectures. The proposed scheme is based on a modified pipeline throttling technique which introduces a throttling point at the last stage of the processor pipeline in order to reduce power consumption. We demonstrate that resource utilization plays an important role in efficient power management and that our strategy can significantly improve performance in the power-saving mode. Since the proposed resource management scheme tests the processor condition cycle by cycle, we evaluate its performance by setting a target IPC as one sort of immediate power measure. Our analysis shows that an SMT processor with our dynamic resource management scheme can yield significantly higher overall performance Seong-Won Lee, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2006 | Design and evaluation of a hierarchical decoupled architecture
Won Woo Ro, Stephen P. Crago, Alvin M. Despain, Jean-Luc Gaudiot |
J. Supercomput. | 4 |
| 2005 | A Low-Complexity Issue Queue Design with Speculative Pre-execution
Won Woo Ro, Jean-Luc Gaudiot |
HiPC | 2 |
| 2005 | Area and System Clock Effects on SMT/CMP ThroughputabstractTwo approaches to high throughput processors are chip multiprocessing (CMP) and simultaneous multithreading (SMT). CMP increases layout efficiency, which allows more functional units and a faster clock rate. However, CMP suffers from hardware partitioning of functional resources. SMT increases functional unit utilization by issuing instructions simultaneously from multiple threads. However, a wide-issue SMT suffers from layout and technology implementation problems. We use silicon resources as our basis for comparison and find that area and system clock have a large effect on the optimal SMT/CMP design trade. We show the area overhead of SMT on each processor and how it scales with the width of the processor pipeline and the number of SMT threads. The wide issue SMT delivers the highest single-thread performance with improved multithread throughput. However, multiple smaller cores deliver the highest throughput. Also, alternate processor configurations are explored that trade off SMT threads for other microarchitecture features. The result is a small increase to single-thread performance, but a fairly large reduction in throughput. James Burns, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 2004 | A Fast and Well-Structured MultiplierabstractThe performance of multiplication is crucial for multimedia applications such as 3D graphics and signal processing systems which depend on extensive numbers of multiplications. Previously reported multiplication algorithms mainly focus on rapidly reducing the partial products rows down to final sums and carries used for the final accumulation. These techniques mostly rely on circuit optimization and minimization of the critical paths. In this paper, an algorithm to achieve fast multiplication in two's complement representation is presented. Indeed, our approach focuses on reducing the number of partial product rows. In turn, this directly influences the speed of the multiplication, even before applying partial products reduction techniques. Fewer partial products rows are produced, thereby lowering the overall operation time. This results in a true diamond-shape for the partial product tree which is more efficient in terms of implementation. Jung-Yup Kang, Jean-Luc Gaudiot |
DSD | 2 |
| 2004 | Speculation Control for Simultaneous MultithreadingabstractSummary form only given. Speculative executions help modern processors to expose independent instructions on the fly and accordingly exploit more instruction-level parallelism. However, when incorrect speculations occur, useless work is performed for those incorrectly speculated instructions. This lowers a sustained performance and leads to a significant waste of power. Unlike superscalar processors, simultaneous multithreading (SMT) processors can concurrently execute multiple threads. Thus, they have a chance to control speculative executions by deliberately choosing threads from which instructions will be fetched at each cycle, considering the dynamic characteristics of running threads. We present an efficient front-end mechanism, called SAFE-T (speculation-aware front-end throttling), for scheduling threads in SMT processors. It involves thread prioritizing and throttling; priority given to a thread can be overridden when that thread seems to suffer from an excessive amount of incorrect speculations, therefore preventing instructions from being fetched. Simulation results show that our policy provides an average reduction of 41.6% in the number of wrong-path instructions and improves the instruction throughput by up to 14.5%. A cost-effective implementation for the proposed policy is shown as well. Dongsoo Kang, Jean-Luc Gaudiot |
IPDPS | 2 |
| 2004 | SPEAR: A Hybrid Model for Speculative Pre-ExecutionabstractSummary form only given. Speculative preexecution achieves efficient data prefetching by running additional prefetching threads on spare hardware contexts. Various implementations for speculative preexecution have been proposed, including compiler-based static approaches and hardware-based dynamic approaches. A static approach defines the p-thread at compile time and executes it as a stand-alone running thread. Therefore, it cannot efficiently take the dynamic events into account and requires a higher fetch bandwidth. Conversely, a hardware approach is, by essence, able to dynamically use the runtime information. However, it requires more complex hardware and also lacks global program information on data and control flow. We propose SPEAR (Speculative Preexecution Assisted by CompileR), a preexecution model which is a hybrid of the two approaches. It relies on a post-compiler to extract the p-thread code from program binaries and uses specially designed hardware to trigger the execution of the p-thread. For this purpose, an automated software tool for p-thread identification has been developed and a modified SMT model with the specially designed front-end is proposed. Won Woo Ro, Jean-Luc Gaudiot |
IPDPS | 2 |
| 2003 | An Efficient PIM (Processor-In-Memory) Architecture for Motion EstimationabstractMotion estimation is the most time consuming stage of MPEG family encodings and it reportedly absorbs up to 90% of the total execution time of MPEG processing. Therefore, we propose a hardware/software co-design paradigm that uses a PIM module to efficiently execute motion estimation operations. We use a PIM module to reduce the memory access penalty caused by a large number of memory accesses. We segment the PIM module into small pieces so that each smaller PIM module can execute the operations in parallel fashion. However, in order to execute the operations in parallel, there are critical overheads that involve replicating a huge amount of data to many of these smaller PIM modules. Not only do these replications require a huge amount of additional memory accesses but also calculations when generating addresses. Therefore, we also present an efficient data distribution mechanism to effectively support parallel executions among these smaller PIM modules. With our paradigm, the host processor can be relieved from computationally-intensive and data-intensive workloads of motion estimation. We observed up to 2034/spl times/ improvement in reduction of the number of memory accesses and up to 439/spl times/ performance improvement for the execution of motion estimation operations when using our computing paradigm. Jung-Yup Kang, Sandeep Gupta 0001, Saurabh Shah, Jean-Luc Gaudiot |
ASAP | 4 |
| 2003 | Clustered Microarchitecture Simultaneous Multithreading
Seong-Won Lee, Jean-Luc Gaudiot |
Euro-Par | 2 |
| 2003 | Incomplete Information Processing for Optimization of Distributed Applications
Alfredo Cristóbal-Salas, Andrei Tchernykh, Jean-Luc Gaudiot |
SNPD | 3 |
| 2003 | Introducing the New Editor-in-Chief of the IEEE Transactions on Computers
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2002 | Parallel Computer Architecture and Instruction-Level Parallelism
Jean-Luc Gaudiot |
Euro-Par | 1 |
| 2002 | On a scheme for parallel sorting on heterogeneous clusters
Christophe Cérin, Jean-Luc Gaudiot |
Future Gener. Comput. Syst. | 2 |
| 2002 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2002 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2002 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2002 | SMT Layout Overhead and ScalabilityabstractSimultaneous Multi-Threading (SMT) is a hardware technique that increases processor throughput by issuing instructions simultaneously from multiple threads. However, while SMT can be added to an existing microarchitecture with relatively low overhead, this additional chip area could be used for other resources such as more functional units, larger caches, or better branch predictors. How large is the SMT overhead and at what point does SMT no longer pay off for maximum throughput compared to adding other architecture features? This paper evaluates the silicon overhead of SMT by performing a transistor/interconnect-level analysis of the layout. We discuss microarchitecture issues that impact SMT implementations and show how the Instruction Set Architecture (ISA) and microarchitecture can have a large effect on the SMT overhead and performance. Results show that SMT yields large performance gains with small to moderate area overhead. James Burns, Jean-Luc Gaudiot |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | MCOMA: A Multithreaded COMA ArchitectureabstractThe authors present a new Cache Only Memory Architecture, MCOMA, whose performance supersedes that of existing COMA architectures. This performance gain is obtained by basing the execution on multithreading principles and the provision of a separate search interconnection network between group directories to optimize the data search. As a result, our proposed architecture benefits from the data locality of COMA as well as the latency tolerance of multithreaded execution. To demonstrate the performance of our proposed model, we developed an execution-driven simulator, on top of the MINT, a MIPS interpreter (J.E. Veenstra and R.J. Fowler, 1994). Halima El Naga, Jean-Luc Gaudiot |
ICCD | 2 |
| 2001 | Alias Analysis for Java with Reference-Set RepresentationabstractProposes a flow-sensitive, context-insensitive alias analysis in Java that is more efficient and precise than previous analyses in C++. For this, we propose a reference-set alias representation and we present the propagation rules for this representation. For the type determination, the type table is built with reference variables and with all possible types of those variables. We propose an algorithm in a popular iterative loop method with a structural traversal of a context-free grammar. Finally, we show that our reference-set representation has better performance for the alias analysis algorithm than the existing object-pair representation does. Jongwook Woo, Jehak Woo, Isabelle Attali, Denis Caromel, Jean-Luc Gaudiot, Andrew L. Wendelborn |
ICPADS | 5 |
| 2001 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2000 | An Over-partitioning Scheme for Parallel Sorting on Clusters with Processors Running at different Speeds
Christophe Cérin, Jean-Luc Gaudiot |
CLUSTER | 2 |
| 2000 | Parallel Sorting Algorithms with Sampling Techniques on Clusters with Processors Running at Different Speeds
Christophe Cérin, Jean-Luc Gaudiot |
HiPC | 2 |
| 2000 | Quantifying the SMT Layout Overhead-Does SMT Pull Its Weight?abstractSimultaneous Multi-Threading (SMT) is a hardware technique that increases processor throughput by issuing instructions simultaneously from multiple threads. However, while SMT can be added to an existing microarchitecture with relatively low overhead, this additional chip area could be used for other resources such as more functional units, larger caches or better branch predictors. How large is the SMT overhead, and at what point does SMT no longer pay off compared to adding other architecture features? This paper evaluates the silicon overhead of SMT by performing a transistor/interconnect level analysis of the layout. We discuss micro-architecture issues that impact SMT implementations, and show how the Instruction Set Architecture (ISA) and microarchitecture can have a large effect on the SMT overhead and performance. Results show that SMT yields large performance gains with small to moderate area overhead. James Burns, Jean-Luc Gaudiot |
HPCA | 2 |
| 2000 | Caching Single-Assignment Structures to Build a Robust Fine-Grain Multi-Threading SystemabstractWe present the design, implementation, and evaluation of single assignment data structures and of a software controlled cache in an existing multi-threaded architecture platform-the Efficient Architecture for Running Threads (EARTH). The software-controlled cache (ISSC) exploits temporal and spatial locality of EARTH split-phased memory transactions for single-assignment memory references. Our experimental evaluation indicates that the caching mechanism for single-assignment storage makes the EARTH memory system more robust to variations in the latency of memory operations. As a consequence the system can be ported to a wider range of machine platforms and deliver speedup for both regular and irregular application. Wen-Yen Lin, Jean-Luc Gaudiot, José Nelson Amaral, Guang R. Gao |
IPDPS | 2 |
| 2000 | An efficient heuristic for code partitioning
Moez Ayed, Jean-Luc Gaudiot |
Parallel Comput. | 2 |
| 2000 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2000 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 2000 | Editor's Note
Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 1999 | Communication Generation for Aligned and Cyclic(K) Distributions Using Integer LatticeabstractOptimizing communication is a key issue in generating efficient SPMD codes in compiling distributed arrays on data parallel languages, such as High Performance Fortran. In HPF, the array distribution may involve alignment and cyclic(k)-distribution such that the enumeration of the local set and the enumeration of the communication set exhibit regular patterns which can be modeled as integer lattices. In the special case of unit-strided alignment, many techniques of the communication set enumeration have been proposed, while in the general case of the non-unit-strided alignment, inspector-like run-time codes are needed to build repeating pattern table or to scan over local elements such that the communication set can be constructed. Unlike other works on this problem of the general alignment and cyclic(k) distribution, our approach derives an algebraic solution for such an integer lattice that models the communication set by using the Smith-Normal-Form analysis, therefore, efficient enumeration of the communication set can be generated. Based on the integer lattice, we also present our algorithm for the SPMD code generation. In our approach, when the parameters are known, the SPMD program can be efficiently constructed without any inspector-like run-time codes. Eric Hung-Yu Tseng, Jean-Luc Gaudiot |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1998 | Flat Indexing: A Compilation Technique to Enhance Parallelism of Logic ProgramsabstractThe paper presents a systematic approach to the compilation of logic programs for efficient clause indexing. As the kernel of the approach, we propose the indexing tree which provides a simple, but precise representation of average parallelism per node (i.e., choice point) as well as the amount of clause trials. It also provides the way to evaluate the number of the cases that the control is passed to the failure code by the indexing instruction such as switch on term, switch on constant, or switch on structure. By analyzing the indexing tree created when using the indexing scheme implemented in the WAM, we show the drawback of the WAM indexing scheme in terms of parallelism exposition and scheduling. Subsequently we propose a new indexing scheme, which we call Flat indexing. Experimental results show that over one half of the benchmarks benefit from the Flat indexing, such that compared with the WAM indexing scheme, the number of choice points is reduced by 15%. Moreover, the amount of failures which occur during the execution of indexing instructions is reduced by 35%. Hiecheol Kim, Jean-Luc Gaudiot |
ICPADS | 2 |
| 1998 | Analysis of a Heuristic for Code Partitioning
Moez Ayed, Jean-Luc Gaudiot |
J. Supercomput. | 2 |
| 1997 | Data and Workload Distribution in a Multithreaded Architecture
Andrew Sohn, Mitsuhisa Sato, Namhoon Yoo, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 4 |
| 1996 | Parallel Computing with the Sisal Applicative Language: Programmability and Performance IssuesabstractThe traditional argument for applicative languages has been programmability. Indeed, due to high-level abstractions and the implicit parallelism provided by applicative languages, programmers are free to concentrate on the implementation of the algorithm at hand without being burdened with low-level machine execution details. However, it has long been believed that the implementation and raw performance of applicative languages would be their downfall. We report here that it is easy to deliver both programmability and performance through applicative programming. To demonstrate the viability of applicative programming in the context of parallel computing, quantitative results from an experiment which consists of developing a multigrid elliptic Partial Differential Equation (PDE) solver are presented. Chinhyun Kim, Jean-Luc Gaudiot, Wlodzimierz Proskurowski |
Softw. Pract. Exp. | 2 |
| 1995 | Incorporating Input/Output Operations Into Dynamic Data-Flow Graphs
Paraskevas Evripidou, Jean-Luc Gaudiot |
Parallel Comput. | 2 |
| 1995 | Implementing regularly structured neural networks on the DREAM machineabstractHigh-throughput implementations of neural network models are required to transfer the technology from small prototype research problems into large-scale "real-world" applications. The flexibility of these implementations in accommodating for modifications to the neural network computation and structure is of paramount importance. The performance of many implementation methods today is greatly dependent on the density and the interconnection structure of the neural network model being implemented. A principal contribution of this paper is to demonstrate an implementation method which exploits maximum amount of parallelism from neural computation, without enforcing stringent conditions on the neural network interconnection structure, to achieve this high implementation efficiency. We propose a new reconfigurable parallel processing architecture, the Dynamically Reconfigurable Extended Array Multiprocessor (DREAM) machine, and an associated mapping method for implementing neural networks with regular interconnection structures. Details of the system execution rate calculation as a function of the neural network structure are presented. Several example neural network structures are used to demonstrate the efficiency of our mapping method and the DREAM machine architecture on implementing diverse interconnection structures. We show that due to the reconfigurable nature of the DREAM machine, most of the available parallelism of neural networks can be efficiently exploited. Soheil Shams, Jean-Luc Gaudiot |
IEEE Trans. Neural Networks | 2 |
| 1994 | Authors' Reply
Walid A. Najjar, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 1993 | Special Issue on DataFlow and Multithreaded Architectures - Guest Editors' Introduction
Guang R. Gao, Jean-Luc Gaudiot, Lubomir F. Bic |
J. Parallel Distributed Comput. | 2 |
| 1993 | Block Scheduling of Iterative Algorithms and Graph-Level Priority Scheduling in a Simulated Data-Flow MultiprocessorabstractIterative methods for solving linear systems are discussed. Although these methods are inherently highly sequential, it is shown that much parallelism could be exploited in a data-flow system by scheduling the iterative part of the algorithms in blocks and by looking ahead across several iterations. This approach is general and will apply to other iterative and loop-based problems. It is also demonstrated by simulation that relying solely on data-driven scheduling of parallel and unrolled loops results in low resource utilization and poor performance. A graph-level priority scheduling mechanism has been developed that greatly improves resource utilization and yields higher performance.> Paraskevas Evripidou, Jean-Luc Gaudiot |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1992 | Efficient implementation of neural networks on the DREAM machineabstractHigh-throughput implementations of neural networks are needed in order to expand the use of this technology from small research problems into practical 'real-world' applications. Due to the wide range of possible neural network paradigms and the rapid evolution of these models, high degree of implementation flexibility if essential. The Dynamically Reconfigurable Extended Array Multiprocessor (DREAM) Machine has been specifically designed for implementation of neural networks. The architecture offers sufficient flexibility for use on a wide range of neural network applications. The authors describe the basic computational and communicational requirements of neural network models. A mapping method is proposed that can take advantage of the DREAM Machine's reconfigurable interconnection network to achieve efficient implementations for a diverse range of neural network structures. The effectiveness of the architecture and the mapping method is demonstrated through the use of several examples and shown to be superior to previous systolic implementation methods.> Soheil Shams, Jean-Luc Gaudiot |
ICPR (4) | 2 |
| 1992 | Parallelism Profiling of an Ops5 Production System InterpreterabstractThe issue of processing nonnumeric algorithms has been one of the major research foci of parallel processing. However, nonnumeric algorithms are known to be difficult to parallelize due to their inherent non-regularity. The authors investigate the potential parallelism in a typical nonnumeric problem. In particular, they selected an OPS5 production system interpreter. The parallelism they investigate is a very fine-grain parallelism at the instruction level. For the implementation of a complete OPS5, they selected SISAL, a purely functional language. A complete OPS5 is implemented in SISAL. Benchmark production system programs are executed to investigate potential parallelism in the production system interpreter. Experimental results indicated that most of the parallelism can be extracted from the matching step of production systems. This particular implementation shows that the production system interpreter can give up to 8-fold speed-up in the absence of communication overhead.> Andrea Sohn, Jean-Luc Gaudiot, Michel Guyot |
SEKE | 2 |
| 1991 | A scheme to extract run-time parallelism form sequential loopsabstractArticle Free Access Share on A scheme to extract run-time parallelism form sequential loops Authors: Chinhyan Kim Electrical Engineering-Systems, University of Southern California, Los Angeles, California Electrical Engineering-Systems, University of Southern California, Los Angeles, CaliforniaView Profile , Jean-Luc Gaudiot Electrical Engineering-Systems, University of Southern California, Los Angeles, California Electrical Engineering-Systems, University of Southern California, Los Angeles, CaliforniaView Profile Authors Info & Claims ICS '91: Proceedings of the 5th international conference on SupercomputingJune 1991 Pages 75–84https://doi.org/10.1145/109025.109049Published:01 June 1991Publication History 1citation1,024DownloadsMetricsTotal Citations1Total Downloads1,024Last 12 Months16Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Chinhyun Kim, Jean-Luc Gaudiot |
ICS | 2 |
| 1991 | A Macro Actor/Token Implementation of Production Systems on a Data-Mow Multiprocessor
Andrea Sohn, Jean-Luc Gaudiot |
IJCAI | 2 |
| 1990 | A Decoupled Graph/Computation Data-Driven Architecture with Variable-Resolution Actors
Paraskevas Evripidou, Jean-Luc Gaudiot |
ICPP (1) | 2 |
| 1990 | Compiling Programs to Direct Access Data-Flow Graphs
Yi-Hsiu Wei, Jean-Luc Gaudiot |
ICPP (2) | 2 |
| 1990 | Representing and Processing Production Systems in Connectionist ArchitecturesabstractMuch effort has been expended on developing special architectures dedicated to the efficient execution of problems in artificial intelligence (AI), especially production systems. While artificial neural networks (ANNs) offer the promise of solving various problems in pattern recognition and classification, we demonstrate here that the ANN approach can be applied to the AI production system paradigm. Among various types of neural networks, the three-layers of ring-structured feedback network is considered in this paper to suit the problem domain under investigation. Characteristics of the production system paradigm are identified. Various aspects of the use of feedback neural networks in mapping production systems are discussed. Two types of representation techniques are studied: local and hierarchical representations. A hierarchical representation derives features from patterns in production systems and constructs a 3-dimensional space called feature space, where a pattern can be uniquely defined by a vector. To demonstrate the efficient use of the neural network approach, a mapping of the generic production system is detailed throughout the paper. The results of a deterministic simulation demonstrate that the three layers of ring-structured feedback neural network architecture can be an efficient processing mechanism for the AI production system paradigm. Andrew Sohn, Jean-Luc Gaudiot |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1990 | Special Issue on Data-Flow Computing: Guest Editors' Introduction
Lubomir F. Bic, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 2 |
| 1990 | Network design and allocation considerations in the Hughes data-flow machine
M. R. Exum, Jean-Luc Gaudiot |
Parallel Comput. | 2 |
| 1990 | Network Resilience: A Measure of Network Fault ToleranceabstractA probabilistic measure of network fault tolerance expressed as the probability of a disconnection is proposed. Qualitative evaluation of this measure is presented. As expected, the single-node disconnection probability is the dominant factor irrespective of the topology under consideration. The authors derive an analytical approximation to the disconnection probability and verify it with a Monte Carlo simulation. On the basis of this model, the measures of network resilience and relative network resilience are proposed as probabilistic measures of network fault tolerance. These are used to evaluate the effects of the disconnection probability on the reliability of the system.> Walid A. Najjar, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 1990 | Data-Driven Parallel Production SystemsabstractMuch effort has been expended on developing special architectures dedicated to the efficient execution of production systems. While data-flow principles of execution offer the promise of high programmability for numerical computations, it is shown that the data-driven principles can also be applied to symbolic computations. In particular, a mapping of the RETE match algorithm along the line of production systems is considered. Bottlenecks of the RETE match algorithm in a multiprocessor environment are identified and possible solutions are suggested. The modifications to the actor set as well as the program graph design are shown for execution on the tagged data-flow computer. The results of a deterministic simulation of this multiprocessor architecture demonstrate that artificial intelligence production systems can be efficiently mapped on data-driven architectures.> Jean-Luc Gaudiot, Andrew Sohn |
IEEE Trans. Software Eng. | 1 |
| 1989 | Parallel Computing: One Opportunity, Four ChallengesabstractIntegration technologies for parallel computing are discussed. Requirements imposed by the need to properly interconnect large numbers of processing elements are emphasized. Four topics discussed as cornerstones in the field of parallel computing are programmability, communication network design, reliable operation, and performance evaluation and benchmarking.> Jean-Luc Gaudiot |
ICDE | 1 |
| 1989 | Limits on Scalability in Gracefully Degradable Large-Scale SystemsabstractThe authors present an analysis of the scalability of large-scale degradable homogeneous multiprocessors by assessing the limitations imposed by reliability considerations on the number of processors. They demonstrate that graceful degradation in large-scale systems is not scalable. An increase in the number of processors must be matched by a significant increase in the coverage factor in order to maintain the same performance and reliability levels.> Walid A. Najjar, Jean-Luc Gaudiot |
SRDS | 2 |
| 1989 | Occamflow: A Methodology for Programming Multiprocessor Systems
Jean-Luc Gaudiot, Liang-Teh Lee |
J. Parallel Distributed Comput. | 1 |
| 1989 | Token Relabeling in a Tagged Token Data-Flow ArchitecureabstractA direct structure access approach called token relabeling scheme is presented in which all array operations can be performed without the use of any intermediary structure memory. The graph constructs for both approaches are described. Four numerical algorithms including fast Fourier transform, bitonic sort, LU decomposition, and matrix multiplication are implemented in both approaches. Token relabeling is directly applicable to arrays which are entirely consumed. Graph constructs which allow partial consumption are also designed. The effectiveness and efficiency of this parallel random access construct verifies the feasibility and applicability of the token relabeling scheme. A deterministic simulation of Arvind and K.P. Gostelow's tagged-token data-flow architecture (1982) is undertaken to validate the graph designs and evaluate the performance of the graphs. It is shown that the direct access graphs present better performance with regard to execution time, speedup, and resource utilization.> Jean-Luc Gaudiot, Yi-Hsiu Wei |
IEEE Trans. Computers | 1 |
| 1988 | Network Disconnection in Distributed SystemsabstractA probabilistic measure of network fault-tolerance expressed as the probability of occurrence of a disconnection is proposed. A simple analytical approximation to the disconnection probability is derived and verified using a Monte Carlo simulation. The effects of the disconnection probability on the overall system's reliability are analyzed using a proposed measure of network resilience.> Walid A. Najjar, Jean-Luc Gaudiot |
ICDCS | 2 |
| 1988 | Iterative Algorithms in a Data-Driven Environment
Paraskevas Evripidou, Jean-Luc Gaudiot |
ICPP (1) | 2 |
| 1988 | Data-Driven Multiprocessor Implementation of the Rete Match Algorithm
Jean-Luc Gaudiot, Andrew Sohn |
ICPP (1) | 1 |
| 1988 | Solving Partial Differential Equations in a Data-Driven Multiprocessor EnvironmentabstractThe implementation of some partial differential equation (PDE) solvers (such as the Jacobi method) on a tagged token data-flow graph is demonstrated. Asynchronous methods (e.g. chaotic relaxation) are studied and other scheduling approaches (such as the token no-labeling scheme) are introduced to support the implementation of the asynchronous methods in a data-driven environment. High-level data-flow-language program constructs are introduced in order to handle chaotic operations. The performance of the program graphs is demonstrated by a deterministic simulation of a message-passing data-flow multiprocessor. An analysis of the overhead in the data-flow graphs is undertaken to demonstrate the limits of parallel operations in data-flow PDE program graphs.> Jean-Luc Gaudiot, C. M. Lin, M. Hosseiniyar |
ISCA | 1 |
| 1988 | Program graph allocation in distributed multicomputers
Jean-Luc Gaudiot, Jen-I Pi, Michael L. Campbell |
Parallel Comput. | 1 |
| 1988 | Demand-Driven Interpretation of FP Programs on a Data-Flow MultiprocessorabstractPresents a demand-driven evaluation system for list-structure language systems, using the functional language FP. It enables execution in a data-driven environment. A formal approach for transforming FP programs into lazy programs, which contain the notion of demands, is used. The superset language of FP is called DFP (demand-driven FP). A demand reduction scheme is used to remove unnecessary demand propagations on DFP programs to reduce run-time overhead. The DFP programs are translated into data-flow graphs according to the graph schemata developed from the FP-DFP transformation rules. The execution characteristics of the DFP graphs are identified and the architecture supports for efficient execution are suggested. The system allows programming in FP by infinite data structures and the application of partial-function-value evaluation. Examples of these applications are used to demonstrate the transformation process, the principles of run-time interpretation, the effectiveness of the transformation, and the power of the evaluation system.> Yi-Hsiu Wei, Jean-Luc Gaudiot |
IEEE Trans. Computers | 2 |
| 1987 | Multi-Level Execution In Data-Flow Architectures
Walid A. Najjar, Jean-Luc Gaudiot |
ICPP | 2 |
| 1987 | Data-driven multicomputers in digital signal processingabstractNew technologies of integration allow the design of powerful systems which may include several thousands of elementary processors. These multiprocessors may be used for a range of applications in signal and data processing. However, assuring the proper interaction of a large number of processors and the ultimate safe execution of the user programs presents a crucial scheduling problem. The scheduling of operations upon the availability of their operands has been termed the data-driven mode of execution and offers an elegant solution to the issue. This approach is described in this paper and several architectures which have been proposed or implemented (systolic arrays, data-flow machines, etc.) are examined in detail. The problems associated with data-driven execution are also studied. A multi-level approach to high-speed digital signal processing is then evaluated. Jean-Luc Gaudiot |
Proc. IEEE | 1 |
| 1986 | : Token Relabeling in a Tagged Data-Flow Architecture
Yi-Hsiu Wei, Jean-Luc Gaudiot |
ICPP | 2 |
| 1986 | Structure Handling in Data-Flow SystemsabstractData-flow languages have been hailed as the solution to the programmability of general-purpose multiprocessors. However, data-flow semantics introduce constructs that lead to much overhead at compilation, allocation, and execution time. Indeed, due to its functionality, the data-flow model of computation does not handle repetitive program constructs very efficiently. This is due to the fact that the cornerstone of data flow, namely the concept of single assignment, is opposed to the idea of reexecution of a portion of program as in a loop. A corollary of this problem is the effective representation, storage, and processing of data structures, as these will most often be used in loops. In this paper, various aspects of this issue are explailned in detail. Several solutions that have been put forward in the current literature are then surveyed and analyzed. In order to offset some of the disadvantages presented by these, we introduce new methods for handling arrays. In the first one, we raise the level of computation to that of arrays for more efficient operation. In the two others, the opposite approach is taken, and the notion of array is done away with entirely at the execution level in order to take advantage of the data-flow semantics at their best logical level of performance. Jean-Luc Gaudiot |
IEEE Trans. Computers | 1 |
| 1985 | Fault-Tolerance and Data-Flow Systems
Jean-Luc Gaudiot, Cauligi S. Raghavendra |
ICDCS | 1 |
| 1985 | Methods for Handling Structures in Data-Flow SystemsabstractData-flow languages are an elegant solution to the programmability of general-purpose multiprocessors.However, data-flow semantics introduce large overhead at execution time.Indeed, the effective representation, storage, and processing of data is diltScult due to the functionality of the dataflow model of computation.In this paper, various aspects of ~,his issue are explained in detail and several solutions that have been put forward in the current literature are surveyed.We present also alternatives: in the first one, we raise the level of computation to that of arrays for more efficient operation.In the two others, the opposite approach is taken, and the notion of array is entirely done away with at the execution level in order to take advantage of the data-flow semantics at their best logical level of performance. I. INTRODUCTION Elementary data-flow principles Jean-Luc Gaudiot |
ISCA | 1 |
| 1985 | Performance evaluation of a simulated data-flow computer with low-resolution actors
Jean-Luc Gaudiot, Milos D. Ercegovac |
J. Parallel Distributed Comput. | 1 |
| 1985 | A Distributed VLSI Architecture for Efficient Signal and Data ProcessingabstractThe machine described, the Hughes Data-Flow Multiprocessor (HDFM), is a high-performance, scalable, fault-tolerant, highly programmable multicomputer designed for embedded signal and data processing applications. The architecture of the machine is described in detail, and the influences on the final design of various requirements such as weight, size, power consumption, performance level and reliability are shown. The processing elements have been designed to reduce the number of VLSI component types required and for modularity of the physical system. The modular nature of the architecture allows a range of throughput and reliability requirements to be met. The model of execution, derived from original data-flow principles, is presented, as well as the various software tools which give the system its high-level language programmability. Complex constructs (such as large structure handling) are demonstrated. The results of a deterministic simulation of the machine show that a 64-processing-element machine may provide real throughput of 64 million instructions per second (MIPS). Jean-Luc Gaudiot, Rex W. Vedder, George K. Tucker, Dennis Finn, Michael L. Campbell |
IEEE Trans. Computers | 1 |
| 1984 | Performance Analysis of a Data-Flow Computer with Variable Resolution Actors
Jean-Luc Gaudiot, Milos D. Ercegovac |
ICDCS | 1 |
| 1982 | A scheme for handling arrays in data-flow systems
Jean-Luc Gaudiot, Milos D. Ercegovac |
ICDCS | 1 |