Gabriela Nicolescu

dblp:39/1586 · DBLP profile ↗
← Back
61ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-5205-9931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 31 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Integrating formal methods and automated tools for DO-178C compliance in UAV software
abstract
The development of software for Unmanned Aerial Vehicles (UAVs) is governed by stringent safety-critical regulations, with DO-178C serving as the primary standard for airborne systems. Ensuring compliance requires extensive verification, validation, and traceability across the software lifecycle, which becomes increasingly complex for autonomous and adaptive UAV functions. This paper proposes an integrated methodology for regulatory compliance checking that combines formal methods with automated verification tools to generate certification-ready evidence under DO-178C. Formal methods are applied at multiple levels: Alloy is used for requirements consistency checking, the SPIN model checker for architectural interaction properties, and bounded model checking for code-level analysis. These techniques are integrated with automated toolchains that provide continuous bidirectional traceability, structural coverage analysis, and automated test execution across the software lifecycle. The approach is evaluated on a Design Assurance Level (DAL) B UAV Collision Avoidance System. The case study demonstrates the production of certification-ready evidence bundles, including closed bidirectional traceability from system requirements through high and low-level software requirements to source code and tests, formal proof summaries linked to requirements, and decision coverage reports on mission-critical logic. Results indicate that tightly integrating formal analysis with automated verification improves early defect detection, reduces manual evidence assembly, and strengthens the auditability of DO-178C compliance. The combined use of formal methods and automation offers a scalable pathway for UAVs and other autonomous systems to achieve compliance with evolving safety regulations. The findings highlight that integrating regulatory compliance checking into development processes can simultaneously enhance rigour and efficiency, providing a model for certifiable autonomy software in civil airspace. • Engineered a workflow that combines formal methods with automation for DO-178C UAV compliance. • End-to-end methodology validated on a UAV Collision Avoidance System case study. • Produced certification-ready evidence: traceability, proofs, and coverage. • Improved early defect detection and reduced manual certification effort. • Illustrates a scalable path for certifiable autonomy in safety-critical UAVs.
Rim Zrelli, Henrique Amaral Misson, Sorelle Audrey K. Kamkuimo, Maroua Ben Attia, Abdo Shabah, Felipe G. Magalhaes, Gabriela Nicolescu
Inf. Softw. Technol.7
2026 Automatic translation of natural language requirements into CTL specifications using Large Language Models: A multi-approach evaluation
abstract
Translating natural language (NL) requirements into formal specifications such as Computation Tree Logic (CTL) is essential for improving the efficiency and scalability of formal verification, especially in safety-critical systems. This study evaluates the ability of Large Language Models (LLMs) to automate this process. We compare three approaches: fine-tuning the Mistral model, using GPT-4 in a few-shot learning setup, and a hybrid that feeds a BERT pattern classifier’s prediction to GPT-4. Using the Natural2CTL dataset, we assess strict logical accuracy and an ambiguity-tolerant accuracy, complemented by auxiliary semantic and structural operator similarity measures. Fine-tuning yields the strongest strict correctness and operator-structure fidelity, while the hybrid narrows the gap to fine-tuning and substantially improves over few-shot prompting alone. Residual errors across methods concentrate in path-quantifier selection, temporal granularity, and scoping in multi-clause requirements. Overall, LLMs can draft CTL candidates that are usable after lightweight normalisation, but they should be integrated into human-in-the-loop workflows with basic automated checks before use in high-assurance settings. • Benchmarks three LLM-based methods for NL-to-CTL translation. • Fine-tuned Mistral achieves 47.6% strict logical accuracy and 71.4% ambiguity-tolerant accuracy for CTL specification generation. • GPT-4 few-shot learning offers rapid prototyping but lower syntactic precision. • BERT-GPT hybrid balances pattern recognition and generative translation. • LLM automation reduces expert effort, but expert review remains critical for safety.
Rim Zrelli, Henrique Amaral Misson, Marwa Ben Attia, Felipe G. Magalhaes, Abdo Shabah, Gabriela Nicolescu
J. Syst. Softw.6
2026 NOProbe: A NOP-Based Dynamic Binary Instrumentation Framework Using Binary Rewriting on x86
abstract
Dynamic Binary Instrumentation (DBI) in user space often suffers from low probe insertion success rates and high execution overhead, due to challenges in handling the compact instruction layouts ($\lt $5 bytes) and complex trampoline placement constraints. Existing techniques are either limited in scope, incur high runtime overhead, or rely on heavyweight code relocation. This paper introduces NOProbe, a lightweight, user-space DBI framework that enables safe and efficient probe insertion using two novel strategies. The first strategy locates trampoline sites by leveraging compiler-generated NOP paddings; the second employs pseudo-NOP instructions to support trampoline placement even when instructions overlap. Additionally, we propose a thread-safe patching algorithm,lock-redirect-load-arm, for safe runtime code modification. Experimental results show that NOProbe achieves 97%-99% probe effectiveness, reduces probe insertion latency, and maintains very low per-probe execution overhead, even under high probe density and multithreaded workloads.
Ahmad Shahnejat Bushehri, Anas Balboul, Adel Belkhiri, Samira Keivanpour, Gabriela Nicolescu, Michel R. Dagenais
IEEE Trans. Dependable Secur. Comput.5
2025 Focus Bug: Learning Environmental Awareness for Efficient Mapless Navigation
abstract
Tiny robots such as nano quadcopters or micro rovers are highly beneficial for various applications as they are inexpensive, agile, and safe for humans. However, their extreme size, weight, and power (SWAP) constraints lead to extremely limited compute, making autonomous navigation challenging. Existing approaches have enabled navigation within these tight constraints, but struggle in dynamic and cluttered scenes.To this end, we present Focus Bug, a novel and robust mapless navigation algorithm that can run on extremely limited hardware. Focus Bug reduces the amount of processed sensory data using a tiny reinforcement learning policy, only processing the inputs necessary for navigation. We use deep reinforcement learning (DRL) to identify critical parts of the robot’s range data and combine it with classical mapless navigation methods to benefit from their robustness and established performance. We implement and evaluate Focus Bug both on a drone in simulation and a micro-rover in the real world to show it can be applied across embodiments. Our hybrid approach outperforms the state-of-the-art in DRL navigation (57% less collisions in dynamic environments) while reducing the amount of range data processed by 87%, and achieving a 2.6X improvement in processing time compared to classical methods. Focus bug is the first method to achieve the high success rate of robust methods (97%) within such a tight compute budget.
Charles Dansereau, Bardienus Pieter Duisterhof, Gabriela Nicolescu
IROS3
2024 Natural2CTL: A Dataset for Natural Language Requirements and Their CTL Formal Equivalents
Rim Zrelli, Henrique Amaral Misson, Maroua Ben Attia, Felipe G. Magalhaes, Abdo Shabah, Gabriela Nicolescu
REFSQ6
2024 Advancing Formal Verification: Fine-Tuning LLMs for Translating Natural Language Requirements to CTL Specifications
abstract
In the domain of formal verification, translating natural language (NL) requirements into Computation Tree Logic (CTL) specifications presents a notable challenge due to the disparity between human-readable documents and formal specifications. This paper introduces a novel approach that leverages Large Language Models (LLMs) to automate this translation process, thereby enhancing the accuracy and efficiency of formal verification practices. We fine-tune three state-of-the-art LLMs—LLAMA3, Mistral, and Qwen2—with a particular focus on optimizing the Mistral model due to its superior performance. Our methodology is supported by the Natural2CTL dataset, consisting of 2,095 NL requirements and their corresponding CTL specifications. We employ evaluation metrics such as validation loss, accuracy, semantic similarity, and Structural Operator Jaccard Similarity (SOJS) for a comprehensive assessment of model performance. Additionally, a comparative analysis with human translators, trained in CTL logic, underscores the LLMs’ potential to match or even surpass human accuracy in translating NL requirements into formal specifications. Our findings reveal that the fine-tuned Mistral model significantly outperforms the other LLMs and human participants, demonstrating superior accuracy in generating CTL specifications. This study advances the field of formal verification by proposing a scalable solution to the NL-to-CTL translation challenge, setting a new benchmark for the integration of AI tools in complex specification tasks.
Rim Zrelli, Henrique Amaral Misson, Maroua Ben Attia, Felipe G. Magalhaes, Abdo Shabah, Gabriela Nicolescu
RSP6
2023 Efficient Defense Against Model Stealing Attacks on Convolutional Neural Networks
abstract
Model stealing attacks have become a serious concern for deep learning models, where an attacker can steal a trained model by querying its black-box API. This can lead to intellectual property theft and other security and privacy risks. The current state-of-the-art defenses against model stealing attacks suggest adding perturbations to the prediction probabilities. However, they suffer from heavy computations and make impracticable assumptions about the adversary. They often require the training of auxiliary models. This can be time-consuming and resource-intensive which hinders the deployment of these defenses in real-world applications. In this paper, we propose a simple yet effective and efficient defense alternative. We introduce a heuristic approach to perturb the output probabilities. The proposed defense can be easily integrated into models without additional training. We show that our defense is effective in defending against three state-of-the-art stealing attacks. We evaluate our approach on large and quantized (i.e., compressed) Convolutional Neural Networks (CNNs) trained on several vision datasets. Our technique outperforms the state-of-the-art defenses with a ×37 faster inference latency without requiring any additional model and with a low impact on the model's performance. We validate that our defense is also effective for quantized CNNs targeting edge devices.
Kacem Khaled, Mouna Dhaouadi, Felipe G. Magalhaes, Gabriela Nicolescu
ICMLA4
2023 SerIOS: Enhancing Hardware Security in Integrated Optoelectronic Systems
abstract
Silicon photonics (SiPh) has different applications, from enabling fast and high-bandwidth communication for high-performance computing systems to realizing energy-efficient optical computation for AI hardware accelerators. However, integrating SiPh with electronic sub-systems can introduce new security vulnerabilities that cannot be adequately addressed using existing hardware security solutions for electronic systems. This paper introduces SerIOS, the first framework aimed at enhancing hardware security in optoelectronic systems by leveraging the unique properties of optical lithography. SerIOS employs cryptographic keys generated based on imperfections in the optical lithography process and an online detection mechanism to detect attacks. Simulation and synthesis results demonstrate SerIOS's effectiveness in detecting and preventing attacks, with a small area footprint of less than 15% and a 100% detection rate across various attack scenarios and optoelectronic architectures, including photonic AI accelerators.
Felipe G. Magalhaes, Mahdi Nikdast, Gabriela Nicolescu
RSP3
2023 ReDaML: A Modeling Language for DO-178C High-Level Requirements in Airspace Systems
abstract
Software development in critical airspace cyber-physical systems is challenging, mainly because of its safety-critical nature. Safety standards and regulations, such as DO-178C, provide guidelines for the development of software to ensure they adhere to the essential safety requirements in the certification processes. The requirements process proposed in the standard, which is responsible for developing the high-level requirements, is one of the most crucial steps in the life cycle since it serves as the basis for the subsequent processes. Having safety as a major concern, specifying safety requirements is of fundamental importance, allowing engineers to evaluate them and propose measures to mitigate the impact of a system failure, which can be catastrophic. In this paper, we present ReDaML, a domain-specific modelling language designed to support the development of safety-critical software systems, focused on the specification of high-level requirements in accordance with the DO-178C guidelines. Finally, a scenario of applying the approach to an UAS collision avoidance system is demonstrated.
Henrique Amaral Misson, Rim Zrelli, Maroua Ben Attia, Felipe G. Magalhaes, Gabriela Nicolescu
RSP5
2023 Security assessment of a commercial router using physical access: a case study
abstract
Physical access to a device can greatly help in vulnerability research as it opens up new vectors for exploitation. This is especially true for embedded devices, which often come with open serial ports and various types of debugging features. Therefore, security assessments of these devices should take into consideration the hardware components and their means of communication. In this paper, we explore a testing methodology that transforms a black box test into a white box test by using physical access to retrieve the code of the applications running on the device. To demonstrate its advantages, we apply this methodology to assess the security risks on a commercial router. We use it to uncover multiple code execution vulnerabilities in the router's firmware. We discuss secure coding guidelines to remediate those vulnerabilities and the importance of IoT security.
Colin Stephenne, Felipe G. Magalhaes, Frédéric Cuppens, Jean-Yves Ouattara, Militza Jean, Gabriela Nicolescu
RSP7
2022 Careful What You Wish For: on the Extraction of Adversarially Trained Models
abstract
Recent attacks on Machine Learning (ML) models such as evasion attacks with adversarial examples and models stealing through extraction attacks pose several security and privacy threats. Prior work proposes to use adversarial training to secure models from adversarial examples that can evade the classification of a model and deteriorate its performance. However, this protection technique affects the model’s decision boundary and its prediction probabilities, hence it might raise model privacy risks. In fact, a malicious user using only a query access to the prediction output of a model can extract it and obtain a high-accuracy and high-fidelity surrogate model. To have a greater extraction, these attacks leverage the prediction probabilities of the victim model. Indeed, all previous work on extraction attacks do not take into consideration the changes in the training process for security purposes. In this paper, we propose a framework to assess extraction attacks on adversarially trained models with vision datasets. To the best of our knowledge, our work is the first to perform such evaluation. Through an extensive empirical study, we demonstrate that adversarially trained models are more vulnerable to extraction attacks than models obtained under natural training circumstances. They can achieve up to ×1.2 higher accuracy and agreement with a fraction lower than ×0.75 of the queries. We additionally find that the adversarial robustness capability is transferable through extraction attacks, i.e., extracted Deep Neural Networks (DNNs) from robust models show an enhanced accuracy to adversarial examples compared to extracted DNNs from naturally trained (i.e. standard) models.
Kacem Khaled, Gabriela Nicolescu, Felipe G. Magalhaes
PST2
2021 HyCo: A Low-Latency Hybrid Control Plane for Optical Interconnection Networks
abstract
Next-generation multiprocessor systems point to the integration of a large number of cores (e.g., processing and memory) where electrical networks-on-chip (eNoCs) can improve the communication performance. As the number of integrated cores increases, metallic interconnect in eNoCs becomes a bottleneck, leading to communication performance degradation and increased power consumption. Optical interconnection networks (OINs) have emerged to outperform the communication infrastructure in multiprocessor systems. Nevertheless, OINs’ full capability is curbed by high latency electrical controllers required to orchestrate and (re)configure the underlying photonic components, realizing a path between sending and receiving cores. Control techniques impose a high latency to perform the network routing, limiting the full utilization of OINs. In this paper, we design a novel low-latency Hybrid Controller (HyCo) that employs acceleration techniques to reduce its execution time. HyCo is developed based on integrating centralized and distributed control techniques as well as by using pre-calculated network routes and a Bloom filter, all of which result in a considerable reduction in HyCo’s latency. Simulation and prototyping results for networks up to 64×64 indicate a latency smaller than 50 ns, in the worst-case scenario.
Felipe G. Magalhaes, Mahdi Nikdast, Fabiano Hessel, Odile Liboiron-Ladouceur, Gabriela Nicolescu
RSP5
2019 Cache Locking Content Selection Algorithms for ARINC-653 Compliant RTOS
abstract
Avionic software is the subject of stringent real time, determinism and safety constraints. Software designers face several challenges, one of them being the interferences that appear in common situations, such as resource sharing. The interferences introduce non-determinism and delays in execution time. One of the main interference prone resources are cache memories. In single-core processors, caches comprise multiple private levels. This breaks the isolation principle imposed by avionic standards, such as the ARINC-653. This standard defines partitioned architectures where one partition should never directly interfere with another one. In cache-based architectures, one partition can modify the cache content of another partition. In this paper, we propose a method based on cache locking to reduce the non-determinism and the contention on lower level memories while improving the time performances.
Alexy Torres Aurora Dugo, Jean-Baptiste Lefoul, Felipe G. Magalhaes, Dahman Assal, Gabriela Nicolescu
ACM Trans. Embed. Comput. Syst.5
2019 Thermal-Aware Design Method for Laser Group Control in Nanophotonic Interconnects
abstract
On-chip integrated lasers are key devices to deliver the high bandwidth expected from nanophotonic interconnects. However, lasers are highly sensitive to temperature variation, which influences the lasing efficiency and the wavelengths of emitted optical signals, both of which are key factors in interconnect power efficiency. It is, thus, necessary to develop techniques for efficient thermal-aware control of lasers. In this brief, we propose the grouping of lasers for efficient power control of their temperature. Laser grouping is carried out taking into account the layout symmetries, and a design method allows the definition of control laws.
Amira Aouina, Hui Li 0034, Ian O'Connor, Gabriela Nicolescu, Sébastien Le Beux
IEEE Trans. Very Large Scale Integr. Syst.5
2018 Silicon Photonic Interconnects: Minimizing the Controller Latency
abstract
Silicon photonic interconnects (SPIs) have emerged as a promising solution to outperform the communication infrastructure in multiprocessor systems-on-chip (MPSoCs). Routing a message from one node to another in an MPSoC integrating SPIs, several photonic components (e.g., switching elements) need to be configured to realize an optical path between sending and receiving nodes. Such configurations are performed in an electronic controller, which, if not fast, imposes high latency in SPIs, constraining the application of SPIs in MPSoCs. Realizing a full exploitation of SPIs, this paper presents a look-up-table-based centralized controller (LUCC). We indicate that LUCC has the lowest latency among the state-of-the-art controllers for SPIs while it can be applied to different SPI architectures. Employing acceleration techniques based on off-line routings, we report (simulation and prototyping) a worst-case control latency smaller than 5 ns. Moreover, LUCC is experimentally integrated with a photonic switch in the lab, where we show contention resolution in one clock cycle.
Felipe G. Magalhaes, Mahdi Nikdast, Yule Xiong, Fabiano Hessel, Odile Liboiron-Ladouceur, Gabriela Nicolescu
ACM Great Lakes Symposium on VLSI6
2018 DeEPeR: Enhancing Performance and Reliability in Chip-Scale Optical Interconnection Networks
abstract
This paper presents an efficient device-level design method to enhance the performance and reliability (DeEPeR) in optical interconnection networks (OINs) under fabrication process variations (PV). Considering different range of variations, DeEPeR explores the design space of fundamental optical components in OINs (e.g., microresonators (MRs)) to improve the overall system performance and reliability. Our study also includes the design and fabrication of several MRs to experimentally validate our proposed method. Moreover, as a system-level case study, we apply DeEPeR to a general passive OIN under PV. Results indicate that DeEPeR considerably improves the optical signal-to-noise ratio (OSNR) in optical interconnection networks.
Mahdi Nikdast, Gabriela Nicolescu, Jelena Trajkovic, Odile Liboiron-Ladouceur
ACM Great Lakes Symposium on VLSI2
2018 From Swarms to Stars: Task Coverage in Robot Swarms with Connectivity Constraints
abstract
Swarm robotics carries the potential of solving complex tasks using simple devices. To do so, however, one must define distributed control algorithms capable of producing globally coordinated behaviours. We propose a methodology to address the problem of the spatial coverage of multiple tasks with a swarm of robots that must not lose global connectivity. Our methodology comprises two layers: (i) a distributed Robot Navigation Controller (RNC) is responsible for simultaneously guaranteeing connectivity and pursuit of multiple tasks; and (ii) a global Task Scheduling Controller approximates the optimal strategy for the RNC with minimal computational load. Our contributions include: (i) a qualitative analysis of the literature on connectivity assessment, (ii) our proposed methodology, (iii) simulations in a multi-physics environment, (iv) real-life robot experiments, and (v) the experimental validation of connectivity, coverage optimality, and fault-tolerance.
Jacopo Panerati, Luca Gianoli, Carlo Pinciroli, Abdo Shabah, Gabriela Nicolescu, Giovanni Beltrame
ICRA5
2017 An analysis of random cache effects on real-time multi-core scheduling algorithms
abstract
The effect of sharing the last-level cache (LLC) among cores in a multi-core system has not been thoroughly investigated especially in the design of efficient scheduling algorithms. And with the growing interest in random caches, which allow for an easier estimation of the worst-case execution time of tasks in critical real-time embedded systems, tools that analyse the sensitivity of workloads to sharing the LLC become necessary. In this paper, we extend a realtime multiprocessor scheduling simulator, SimSo, with a framework that incorporates a random cache model for multi-level caches to evaluate emerging scheduling algorithms under the influence of shared caches. A set of experiments were performed to study the behavior of workloads with respect to worst-case response time, average slack time, and maximum utilization, with varying cache designs under different scheduling algorithms.
Imane Hafnaoui, Chao Chen 0031, Rabeh Ayari, Gabriela Nicolescu, Giovanni Beltrame
RSP4
2016 Modeling fabrication non-uniformity in chip-scale silicon photonic interconnects
Mahdi Nikdast, Gabriela Nicolescu, Jelena Trajkovic, Odile Liboiron-Ladouceur
DATE2
2016 Schedulability-guided exploration of multi-core systems
abstract
Efficient mapping of tasks onto heterogeneous multi-core systems is very challenging especially under hard timing constraints. Assigning tasks to processors is an NP-hard problem and solving it requires the use of meta-heuristics. Relevantly, genetic algorithms have already proven to be one of the most powerful and widely used stochastic tools to solve this problem. Conventional genetic algorithms were initially defined as a general evolutionary algorithm based on blind operators. It is commonly admitted that the use of these operators is quite poor for an efficient exploration. Like-wise, since exhaustive exploration of the solution space is unrealistic, a potent option is often to guide the exploration process by hints, derived by problem structure. This guided exploration prioritizes fitter solutions to be part of next generations and avoids exploring unpromising configurations by transmitting a set of predefined criteria from parents to children. Consequently, genetic operators, such as crossover, must incorporate specific domain knowledge to intelligently guide the exploration of the solution space. In this paper, we illustrate and evaluate the impact of crossover operators and we propose a hybrid genetic algorithm based on a novel schedulability-guided operator that easily outperforms the classical operators by offering at least 21% improvement in terms of ratio of certainly schedulable tasks.
Rabeh Ayari, Imane Hafnaoui, Giovanni Beltrame, Gabriela Nicolescu
RSP4
2016 Tuning framework for stencil computation in heterogeneous parallel platforms
Taieb Lamine Ben Cheikh, Alexandra Aguiar, Sofiène Tahar, Gabriela Nicolescu
J. Supercomput.4
2015 Energy-efficient optical crossbars on chip with multi-layer deposited silicon
abstract
The many cores design research community have shown high interest in optical crossbars on chip for more than a decade. Key properties of optical crossbars, namely a) contention-free data routing b) low-latency communication and c) potential for high bandwidth through the use of WDM, motivate several implementations. These implementations demonstrate very different scalability and power efficiency ability depending on three key design factors: a) the network topology, b) the considered layout and c) the insertion losses induced by the fabrication process. The emerging design technique relying on multi-layer deposited silicon allows reducing optical losses, which may lead to significant reduction of the power consumption. In this paper, multi-layer deposited silicon based crossbars are proposed and compared. The results indicate that the proposed ring-based network exhibits, on average, 22% and 51.4% improvement for worst-case and average losses respectively compared to the most power-efficient related crossbars.
Hui Li 0034, Sébastien Le Beux, Gabriela Nicolescu, Ian O'Connor
ASP-DAC3
2015 Thermal aware design method for VCSEL-based on-chip optical interconnect
Hui Li 0034, Alain Fourmigue, Sébastien Le Beux, Xavier Letartre, Ian O'Connor, Gabriela Nicolescu
DATE6
2014 Chameleon: Channel efficient Optical Network-on-Chip
abstract
The next generation of MPSoC points to the integration of thousands of IP cores, requiring high performance interconnect for high throughput communications. Optical on-chip interconnect enables significantly increased bandwidth and decreased latency in MPSoC. However, the interface between electrical and photonic devices implies strong layout constraints that may impact the system performance and scalability. In this paper, we propose a novel optical interconnect named Chameleon. The interface simplifies the layout and allows the bandwidth between IP cores to be adapted according to the communication requirements. Compared to related networks, Chameleon demonstrates improved scalability and flexibility at the cost of minor increase in power consumption.
Sébastien Le Beux, Hui Li 0034, Ian O'Connor, Kazem Cheshmi, Xuchen Liu 0003, Jelena Trajkovic, Gabriela Nicolescu
DATE7
2014 Efficient transient thermal simulation of 3D ICs with liquid-cooling and through silicon vias
abstract
Three-dimensional integrated circuits (3D ICs) with advanced cooling systems are emerging as a viable solution for many-core platforms. These architectures generate a high and rapidly changing thermal flux. Their design requires accurate transient thermal models. Several models have been proposed, either with limited capabilities, or poor simulation performance. This work introduces an efficient algorithm based on the Finite Difference Method to compute the transient temperature in liquid-cooled 3D ICs. Our experiments show a 5x speedup versus state-of-the-art models, while maintaining the same level of accuracy, and demonstrate the effect of large through silicon vias arrays on thermal dissipation.
Alain Fourmigue, Giovanni Beltrame, Gabriela Nicolescu
DATE3
2014 Optical crossbars on chip, a comparative study based on worst-case losses
abstract
SUMMARY The many‐core design research community has shown high interest in optical crossbar on chip for more than a decade. Key properties of optical crossbars, namely (1) contention‐free data routing, (2) low latency communication, and (3) potential for high bandwidth through the use of wavelength division multiplexing, motivate several implementations of this type of interconnect. These implementations demonstrate very different scalability and power efficiency abilities depending on three key design factors: (1) network topology, (2) considered layout, and (3) insertion losses induced by the fabrication process. In this paper, the worst‐case optical losses of crossbar implementations are compared according to the factors mentioned earlier. The comparison results have the potential to help many‐core system designer to select the most appropriate crossbar implementation according to, for instance, the number of IP cores and the die size. Copyright © 2014 John Wiley & Sons, Ltd.
Sébastien Le Beux, Hui Li 0034, Gabriela Nicolescu, Jelena Trajkovic, Ian O'Connor
Concurr. Comput. Pract. Exp.3
2013 Explicit transient thermal simulation of liquid-cooled 3D ICs
abstract
The high heat flux and compact structure of three-dimensional circuits (3D ICs) make conventional air-cooled devices more subsceptible to overheating. Liquid cooling is an alternative that can improve heat dissipation, and reduce thermal issues. Fast and accurate thermal models are needed to appropriately dimension the cooling system at design time. Several models have been proposed to study different designs, but generally with low simulation performance. In this paper, we present an efficient model of the transient thermal behaviour of liquid-cooled 3D ICs. In our experiments, our approach is 60 times faster and uses 600 times less memory than state-of-the-art models, while maintaining the same level of accuracy.
Alain Fourmigue, Giovanni Beltrame, Gabriela Nicolescu
DATE3
2013 Potential and pitfalls of silicon photonics computing and interconnect
abstract
Trends in SoC design are leading to 3D integration of thousands of high-performance computing resources and high-throughput interconnects, opening up new research directions for hybrid electronic/photonic architectures. In this paper, we introduce how state of the art silicon-photonic devices can realize elementary operations that are traditionally performed by electronic devices, e.g. circuit switching and Boolean function computation. We then highlight how these devices need to be assembled in order to realize more complex functions, taking into account the constraints specific to silicon-photonic technology. In the last part, we summarize the main research directions for the near future.
Sébastien Le Beux, Ian O'Connor, Zhen Li 0046, Xavier Letartre, Christelle Monat, Jelena Trajkovic, Gabriela Nicolescu
ISCAS7
2013 Embedded system verification through constraint-based scheduling
abstract
Verification has become one of the main bottlenecks in the design process of embedded systems, particularly for Multiprocessor Systems-on-Chip (MPSoCs). Efficiently proving the correctness of a design is of extreme importance to reduce cost and time-to-market. Simulation is a common verification method, but complex systems usually require long simulation times. This work advocates Constraint Programming (CP) as a powerful tool for the verification of performance metrics of MPSoCs. Our methodology was evaluated using streaming applications mapped onto a target MPSoC. The resulting constraint-based scheduling problem allowed us to identify performance constraint violations in a fraction of the time required by simulation-based verification.
Olfat El-Mahi, Gilles Pesant, Gabriela Nicolescu, Giovanni Beltrame
RSP3
2012 MpAssign: a framework for solving the many-core platform mapping problem
abstract
SUMMARY Many‐core platforms, providing large numbers of parallel execution resources, emerge as a response to the increasing computation needs of embedded applications. A major challenge raised by this trend is the efficient mapping of applications on parallel resources. This is a nontrivial problem because of the number of parameters to be considered for characterizing both the applications and the underlying platform architectures. Recently, several authors have proposed to use multi‐objective evolutionary algorithm to solve this problem within the context of mapping applications on network‐on‐chips. However, these proposals have several limitations: (1) only few metaheuristics are explored (mainly Nondominated Sorting Genetic Algorithm II and Strength Pareto Evolutionary Algorithm 2), (2) only few objective functions are provided, and (3) they only deal with a small number of the application and architecture constraints. In this paper, we propose a new framework that avoids all of the problems cited previously. Our framework is implemented on top of the jMetal framework, which offers an extensible environment. Our framework allows designers to (1) explore several new metaheuristics, (2) easily add a new objective function (or to use an existing one), and (3) take into account any number of architecture and application constraints. The paper also presents experiments illustrating how our framework is applied to the problem of mapping streaming applications on an NoC‐based many‐core platform. Our results show that several new metaheuristics outperform the classical multi‐objective metaheuristics such as Nondominated Sorting Genetic Algorithm II and Strength Pareto Evolutionary Algorithm 2. Moreover, a parallel multi‐objective evolutionary algorithm is implemented in our framework in order to increase the explored space of solutions by simultaneously running several metaheuristics. Copyright © 2011 John Wiley & Sons, Ltd.
Youcef Bouchebaba, Ali Erdem Özcan, Pierre G. Paulin, Gabriela Nicolescu
Softw. Pract. Exp.4
2012 Integrating Memory Optimization with Mapping Algorithms for Multi-Processors System-on-Chip
abstract
Due to their great ability to parallelize at a very high integration level, Multi-Processors Systems-on-Chip (MPSoCs) are good candidates for systems and applications such as multimedia. Memory is becoming a key player for significant improvements in these applications (power, performance and area). The large amount of data manipulated by these applications requires high-capacity computing and memory. Lately, new programming models have been introduced. This leads to the need of new optimization and mapping techniques suitable for embedded systems and their programming models. This article presents novel approaches for combining memory optimization with mapping of data-driven applications while considering anti-dependence conflicts. Two different approaches are studied and integrated with existing mapping algorithms. The first approach (based on heuristic algorithms) keeps the graph transformation for memory optimization stage from the mapping stage and enables their combination in a design flow. The second approach (based on evolutionary algorithms) combines these two stages and integrates them in a unique stage. Some significant improvements are obtained for memory gain, communication load and physical links.
Bruno Girodias, Luiza Gheorghe Iugan, Youcef Bouchebaba, Gabriela Nicolescu, El Mostapha Aboulhamid, Michel Langevin, Pierre G. Paulin
ACM Trans. Embed. Comput. Syst.4
2011 A multi-objective decision-theoretic exploration algorithm for platform-based design
abstract
This paper presents an efficient technique to perform multi-objective design space exploration of a multiprocessor platform. Instead of using semi-random search algorithms (like simulated annealing, tabu search, genetic algorithms, etc.), we use the domain knowledge derived from the platform architecture to set-up the exploration as a discrete-space multi-objective Markov Decision Process (MDP). The system walks the design space changing its parameters, performing simulations only when probabilistic information becomes insufficient for a decision. The algorithm employs a novel multi-objective value function and exploration strategy, which guarantees high accuracy and minimizes the number of necessary simulations. The proposed technique has been tested with a small benchmark (to compare the results against exhaustive exploration) and two large applications (to prove effectiveness in a real case), namely the ffmpeg transcoder and pigz parallel compressor. Results show that the exploration can be performed with 10% of the simulations necessary for state-of-the-art exploration algorithms and with unrivaled accuracy (0.6 ± 0.05% error).
Giovanni Beltrame, Gabriela Nicolescu
DATE2
2011 Optical Ring Network-on-Chip (ORNoC): Architecture and design methodology
abstract
State-of-the-art System-on-Chip (SoC) consists of hundreds of processing elements, while trends in design of the next generation of SoC point to integration of thousand of processing elements, requiring high performance interconnect for high throughput communications. Optical on-chip interconnects are currently considered as one of the most promising paradigms for the design of such next generation Multi-Processors System on Chip (MPSoC). They enable significantly increased bandwidth, increased immunity to electromagnetic noise, decreased latency, and decreased power. Therefore, defining new architectures taking advantage of optical interconnects represents today a key issue for MPSoC designers. Moreover, new design methodologies, considering the design constraints specific to these architectures are mandatory. In this paper, we present a contention-free new architecture based on optical network on chip, called Optical Ring Network-on-Chip (ORNoC). We also show that our network scales well with both large 2D and 3D architectures. For the efficient design, we propose automatic wavelength-/waveguide assignment and demonstrate that the proposed architecture is capable of connecting 1296 nodes with only 102 waveguides and 64 wavelengths per waveguide.
Sébastien Le Beux, Jelena Trajkovic, Ian O'Connor, Gabriela Nicolescu, Guy Bois, Pierre G. Paulin
DATE4
2011 Multi-granularity thermal evaluation of 3D MPSoC architectures
abstract
Three-dimensional (3D) integrated circuits (IC) are emerging as a viable solution to enhance the performance of Multi-processor System-On-Chip (MPSoC). The use of highspeed hardware and the increased density of 3D architectures present novel challenges concerning thermal dissipation and power management. Most approaches at power and thermal modeling use either static analytical models or slow low-level analog simulations. In this paper, we propose a novel thermal modeling methodology for evaluation of 3D MPSoCs. The integration of this methodology in a virtual platform enables effcient dynamic thermal evaluation of a chip. We present initial results for an architecture based on a 3D Network-On-Chip (NoC) interconnecting 2D processing elements (PE). Our methodology is based on the finite difference method: we perform an initial static characterization, after which high-speed dynamic simulation is possible.
Alain Fourmigue, Giovanni Beltrame, Gabriela Nicolescu, El Mostapha Aboulhamid, Ian O'Connor
DATE3
2011 Layout guidelines for 3D architectures including Optical Ring Network-on-Chip (ORNoC)
abstract
Trends in design of the next generation of Multi-Processors System on Chip (MPSoC) point to 3D integration of thousand of processing elements, requiring high performance interconnect for high throughput and low latency communications. Optical on-chip interconnects enable significantly increased bandwidth and decreased latency. They are thus considered as one of the most promising paradigms for the design of such system. However, existence of interfaces between electronic and photonic signals implies strong constraints on the layout of the 3D architecture and may impact the architecture scalability. In this paper, we propose and evaluate a possible layout for an optical Network-on-Chip used to interconnect processing elements located on different electrical layers.
Sébastien Le Beux, Jelena Trajkovic, Ian O'Connor, Gabriela Nicolescu
VLSI-SoC4
2011 Matrix Nanodevice-Based Logic Architectures and Associated Functional Mapping Method
abstract
This article describes a novel computing architecture organization based on nanoscale logic cells. We propose the use of a cluster of matrix arrangements of cells. In order to interconnect such fine-grained logic cells within a matrix, conventional techniques are not suitable due to a large interconnect overhead. Therefore, we propose the use of static and incomplete interconnect topologies to create matrices of cells. We also propose a method to map functions onto such architectures. We then explore the main parameters of the structure (size of matrices and interconnect topologies) and their impact on the main performance metrics (packing efficiency, speed, and fault tolerance). A cluster packing method also allows the evaluation of the number of matrices used by complex functions and the fill factor for various matrix sizes. The analyses show that this approach is particularly suited for matrices of 16 cells interconnected by modified omega networks. We can conclude that this architecture could improve the scalability of traditional FPGAs by a factor of 8.5.
Pierre-Emmanuel Gaillardon, Fabien Clermidy, Ian O'Connor, Maimouna Amadou, Gabriela Nicolescu
ACM J. Emerg. Technol. Comput. Syst.6
2010 A system-level exploration flow for optica network on chip (ONoC) in 3D MPSoC
abstract
Optical on-chip interconnects and 3D die stacking are currently considered to be two promising paradigms for the design of next generation Multi-Processors System on Chip architectures (MPSoC). New architectures based on these paradigms are currently emerging and new system-level approaches are required for their efficient design and prototype. The paper investigates a system-level flow for evaluating design feasibility, interconnect architecture performance and application execution efficiency as early as possible in the MPSoC design cycle.
Sébastien Le Beux, Gabriela Nicolescu, Guy Bois, Pierre G. Paulin
ISCAS2
2010 Combining mapping and partitioning exploration for NoC-based embedded systems
Sébastien Le Beux, Guy Bois, Gabriela Nicolescu, Youcef Bouchebaba, Michel Langevin, Pierre G. Paulin
J. Syst. Archit.3
2009 Co-simulation based platform for wireless protocols design explorations
abstract
Longer range, faster speed and stronger link are today's wireless mandatory characteristics. Tremendous efforts are being deployed to create new and improved wireless protocols. However, these new protocols are being tested in harsh and uncontrolled environments. Simulation tools help to capture the expected behavior, but the proposed designs might not work in real life situations due to lack of accurate simulation models. Testbed platforms are able to test designs in real life settings, but the flexibility of the design is reduced and design exploration becomes a complex task. This paper presents a hybrid platform composed of a simulation tool and a testbed environment, which makes it possible easily design and accurately test new wireless protocols.
Alain Fourmigue, Bruno Girodias, Gabriela Nicolescu, El Mostapha Aboulhamid
DATE3
2009 Embedded tutorial - Understanding multicore technologies
abstract
Summary form only given. Multicore SoCs integrate an increasing number of heterogeneous programmable units and sophisticated communication interconnects. Unlike classic computers, the design of SoC includes the building of application specific architecture and specific interconnect and other hardware components required to execute the software for a well defined class of applications. In this case, the programming model hides both hardware and software interfaces that may include sophisticated communication and synchronisation concepts to handle parallel programs running on the processors. This embedded tutorial introduces the key technologies for the design of such complex devices.
Ahmed Amine Jerraya, Gabriela Nicolescu
DATE2
2009 Emerging Technologies and Nanoscale Computing Fabrics
abstract
6-8 July 2014
Ian O'Connor, Kotb Jabeur, Nataliya Yakymets, Renaud Daviot, David Navarro, Pierre-Emmanuel Gaillardon, Fabien Clermidy, Maimouna Amadou, Gabriela Nicolescu
VLSI-SoC10
2008 Semantics for Model-Based Validation of Continuous/Discrete Systems
abstract
Continuous and discrete components can be integrated in diverse systems including defense, medical, electronic, communication, and automotive applications. Given the heterogeneity of concepts that have to be taken into consideration, their design involves overcoming specific global modeling and validation challenges. This paper presents semantics for model-based validation of continuous/discrete systems. It focuses on the simulation interfaces semantics, representation and verification. The proposed approach is applied for the validation of a continuous/discrete medical system, an automatic glycemia level regulator.
Luiza Gheorghe Iugan, Faouzi Bouchhima, Gabriela Nicolescu, Hanifa Boucheneb
DATE3
2007 Two-level tiling for MPSoC architecture
abstract
Multiprocessor systems-on-a-chip (MPSoCs architectures) have received a lot of attention in the past years, but few advances in compilation techniques target these architectures. This is particularly true for the exploitation of several level of memory hierarchy. Usually tiling is applied to one loop nest; in this paper we apply simultaneously loop fusion with two-level tiling to several loop nests in the context of a MPSoC architecture. The two level-tiling allows the simultaneous optimization of caches and registers. To optimize the memory space used by temporary arrays, buffers and registers are used as a replacement. The experiments show that these techniques yield a significant reduction in the number of data cache misses and in processing time.
Youcef Bouchebaba, Essaid Bensoudane, Bruno Lavigueur, Pierre G. Paulin, Gabriela Nicolescu
ASAP5
2007 System level assessment of an optical NoC in an MPSoC platform
Matthieu Briere, Bruno Girodias, Youcef Bouchebaba, Gabriela Nicolescu, Fabien Mieyeville, Frédéric Gaffiot, Ian O'Connor
DATE4
2007 MPSoC memory optimization for digital camera applications
abstract
Multiprocessor system-on-a-chip architectures have received a lot of attention in the past years, but few advances in compilation techniques are targeting these architectures. This is particularly true for the exploitation of data locality. Most of the compilation techniques discussed in the literature for parallel architectures are based on single loop nest. However, most multimedia and image processing applications are composed of several loop nests. In this paper, new techniques based on program transformations are proposed to optimize these types of applications. In a monoprocessor architecture, the loop fusion technique is well known. In this paper, the loop fusion is generalized and adapted to a MPSoC architecture. Another technique called "computation propagation " is proposed. It completely removes the temporary arrays and significantly reduces the memory accesses, the memory space and the processing time. Experimental results show that this new technique yields a significant reduction in the number of data cache misses (35%), in processing time (30%) and in channel transactions (85%).
Youcef Bouchebaba, Bruno Lavigueur, Bruno Girodias, Gabriela Nicolescu, Pierre G. Paulin
DSD4
2007 MPSoC memory optimization using program transformation
abstract
Multiprocessor system-on-a-chip (MPSoC) architectures have received a lot of attention in the past years, but few advances in compilation techniques target these architectures. This is particularly true for the exploitation of data locality. Most of the compilation techniques for parallel architectures discussed in the literature are based on a single loop nest. This article presents new techniques that consist in applying loop fusion and tiling to several loop nests and to parallelize the resulting code across different processors. These two techniques reduce the number of memory accesses. However, they increase dependencies and thereby reduce the exploitable parallelism in the code. This article tries to address this contradiction. To optimize the memory space used by temporary arrays, smaller buffers are used as a replacement. Different strategies are studied to optimize the processing time spent accessing these buffers. The experiments show that these techniques yield a significant reduction in the number of data cache misses (30%) and in processing time (50%).
Youcef Bouchebaba, Bruno Girodias, Gabriela Nicolescu, El Mostapha Aboulhamid, Bruno Lavigueur, Pierre G. Paulin
ACM Trans. Design Autom. Electr. Syst.3
2006 Buffer and register allocation for memory space optimization
abstract
In today's embedded systems, memory hierarchy is rapidly becoming a major factor in terms of power, performance and area. This is especially true for embedded multimedia applications using temporary multi-dimensional arrays that are typically used to store intermediate results during multimedia processing. In this paper, the authors propose a new technique that optimizes the use of caches and registers. It consists in combining buffer and register allocation to reduce the size of the temporary arrays. The authors use the concept of live data to replace each array by a smaller buffer. The authors then replace references to this buffer by registers. The experiments are made on a Unix environment and on the StepNP simulator. The results show that the technique yields significant reduction of the number of data cache misses
Youcef Bouchebaba, Gabriela Nicolescu, El Mostapha Aboulhamid, Fabien Coelho
ASAP2
2006 Soft-error classification and impact analysis on real-time operating systems
abstract
This paper investigates the sensitivity of real-time systems running applications under operating systems that are subject to soft-errors. We consider applications using different real-time operating system services: scheduling, time and memory management, intertask communication and synchronization. We report results of a detailed analysis regarding the impact of soft-errors on real-time operating systems cores, taking into account the application timing constraints. Our results show the extent to which soft-errors occurring in a real-time operating system's kernel impact its reliability
N. Ignat, Bogdan Nicolescu, Yvon Savaria, Gabriela Nicolescu
DATE4
2006 A new efficient EDA tool design methodology
abstract
New sophisticated EDA tools and methodologies will be needed to make products viable in the future marketplace by simplifying the various design stages. These tools will permit system design at a high abstraction level and enable automatic refinement through several abstraction levels to obtain a final prototype. They will have to be based on representations that are clean, complete, and easy to manipulate. In order to develop these new EDA tools, key features such as standardization, metadata programming, reflectivity, and introspection are needed. This work proposes a .Net Framework-based methodology, which possesses all these required key features. This methodology simplifies specification, synthesis, and validation of systems and enables the efficient creation/customization of EDA tools at low cost and development time. We show the effectiveness of this methodology by presenting its application for the design of a new EDA tool called ESys .Net (Embedded System design with .Net). We emphasize the specification and simulation aspects of this tool.
James Lapalme, El Mostapha Aboulhamid, Gabriela Nicolescu
ACM Trans. Embed. Comput. Syst.3
2006 Parallel programming models for a multiprocessor SoC platform applied to networking and multimedia
abstract
The MultiFlex system is an application-to-platform mapping tool that integrates heterogeneous parallel components-H/W or S/W- into a homogeneous platform programming environment. This leads to higher quality designs through encapsulation and abstraction. Two high-level parallel programming models are supported by the following MultiFlex platform mapping tools: a distributed system object component (DSOC) object-oriented message passing model and a symmetrical multiprocessing (SMP) model using shared memory. We demonstrate the combined use of the MultiFlex multiprocessor mapping tools, supported by high-speed hardware-assisted messaging, context-switching, and dynamic scheduling using the StepNP demonstrator multiprocessor system-on-chip platform, for two representative applications: 1) an Internet traffic management application running at 2.5 Gb/s and 2) an MPEG4 video encoder (VGA resolution, at 30 frames/s). For these applications, a combination of the DSOC and SMP programming models were used in interoperable fashion. After optimization and mapping, processor utilization rates of 85%-91% were demonstrated for the traffic manager. For the MPEG4 decoder, the average processor utilization was 88%
Pierre G. Paulin, Chuck Pilkington, Michel Langevin, Essaid Bensoudane, Damien Lyonnard, Olivier Benny, Bruno Lavigueur, David Lo 0002, Giovanni Beltrame, Vincent Gagné, Gabriela Nicolescu
IEEE Trans. Very Large Scale Integr. Syst.11
2004 .NET Framework - A Solution for the Next Generation Tools for System-Level Modeling and Simulation
abstract
Nowadays, the use of system level description languages is mandatory for the efficient design of complex systems. These description languages are exemplified by SystemC and SystemVerilog. In this paper, we propose a new .NET framework based system level modeling and simulation environment called Esys.NET (embedded systems design with .NET). It allows (1) cooperation - by enabling Web-based design and multi-language features, (2) easy systems specification task - by enabling integration of software components running application and operating systems and by alleviating memory management, (3) link to automatic refinement tools - by enabling translation of specification models into a standard intermediate format and annotation of specification models, and (4) comparative performances with existing environments.
James Lapalme, El Mostapha Aboulhamid, Gabriela Nicolescu, Luc Charest, François R. Boyer, J. P. David, Guy Bois
DATE3
2004 ESys.Net: a new solution for embedded systems modeling and simulation
abstract
The next generation of tools for embedded systems design will represent a common arena for several cooperating groups. These tools will permit system design at a high abstraction level and enable automatic refinement through several abstraction levels to obtain the final prototype. To facilitate this evolution, we propose a new .Net Framework based system level modeling and simulation environment. This environment allows (1) cooperation -- by enabling web-based design and multi-language features, (2) easy systems specification task -- by enabling the integration of software components and by alleviating memory management and (3) the linking to automatic refinement tools -- by enabling the translation of model specifications into a standard intermediate format and by permitting the annotation of model specifications.
James Lapalme, El Mostapha Aboulhamid, Gabriela Nicolescu, Luc Charest, François R. Boyer, J. P. David, Guy Bois
LCTES3
2004 Object-based hardware/software component interconnection model for interface design in system-on-a-chip circuits
Wander O. Cesário, Lovic Gauthier, Damien Lyonnard, Gabriela Nicolescu, Ahmed Amine Jerraya
J. Syst. Softw.4
2002 Component-based design approach for multicore SoCs
abstract
This paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. Component-based design provides primitives to build complex architectures from basic components. This bottom-up approach allows design-architects to explore efficient custom solutions with best performances. This paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. The system specifications are represented as a virtual architecture described in a SystemC-like model and annotated with a set of configuration parameters. Our component-based design environment provides automatic wrapper-generation tools able to synthesize hardware interfaces, device drivers, and operating systems that implement a high-level interconnect API. This approach, experimented over a VDSL system, shows a drastic design time reduction without any significant efficiency loss in the final circuit.
Wander O. Cesário, Amer Baghdadi, Lovic Gauthier, Damien Lyonnard, Gabriela Nicolescu, Yanick Paviot, Sungjoo Yoo, Ahmed Amine Jerraya, Mario Diaz-Nava
DAC5
2002 Automatic Generation of Fast Timed Simulation Models for Operating Systems in SoC Design
abstract
To enable fast and accurate evaluation of HW/SW implementation choices of on-chip communication, we present a method to automatically generate timed OS simulation models. The method generates the OS simulation models with the simulation environment as a virtual processor Since the generated OS simulation models use final OS code, the presented method can mitigate the OS code equivalence problem. The generated model also simulates different types of processor exceptions. This approach provides two orders of magnitude higher simulation speedup compared to the simulation using instruction set simulators for SW simulation.
Sungjoo Yoo, Gabriela Nicolescu, Lovic Gauthier, Ahmed Amine Jerraya
DATE2
2001 Scalable and flexible cosimulation of SoC designs with heterogeneous multi-processor target architectures
abstract
In this paper, we present a cosimulation environment that provides modularity, scalability, and flexibility in cosimulation of SoC designs with heterogeneous multi-processor target architectures. Our cosimulation environment is based on an object-oriented simulation environment, SystemC. Exploiting the object orientation in SystemC representation, we achieve modularity and scalability of cosimulation by developing modular cosimulation interfaces. The object orientation also enables mixed-level cosimulation to be easily implemented thereby the designer can have flexibility in trade off between simulation performance and accuracy. Experiments with an IS-95 CDMA cellular phone system design show the effectiveness of the cosimulation environment.
Patrice Gerin, Sungjoo Yoo, Gabriela Nicolescu, Ahmed Amine Jerraya
ASP-DAC3
2001 A higher level system communication model for object-oriented specification and design of embedded systems
abstract
The design starting point for current embedded systems design is getting higher and higher on the abstraction level scale in order to meet the challenge of the increasing design gap. Up till now the state-of-the-art tools and methods have used as a highest abstraction of communication the send-receive over a channel, e.g. as in SDL and COSSAP. We introduce a novel higher level communication mechanism for system-level specification which has features supporting object-oriented descriptions and client-server type communication modelling as in CORBA. The communication primitives have been implemented as extensions to System-C, and simulation experiments have been performed.
Kjetil Svarstad, Nezih Ben-Fredj, Gabriela Nicolescu, Ahmed Amine Jerraya
ASP-DAC3
2001 Mixed-level cosimulation for fine gradual refinement of communication in SoC design
abstract
In this paper we propose a method of mixed-level cosimulation that enables gradual refinement of SoC communication from protocol-neutral communication to protocol-fixed communication. For fine granularity in refinement, the method enables the designer to perform channel refinement and module refinement. Thus, the designer can perform more extensive design space exploration in communication refinement. We show the effectiveness of the proposed method in a case study of communication refinement is an IS-95 CDMA cellular phone system design.
Gabriela Nicolescu, Sungjoo Yoo, Ahmed Amine Jerraya
DATE1
2001 A model for describing communication between aggregate objects in the specification and design of embedded systems
abstract
The elevation of design description abstractions is a well accepted technique for handling the complexity and shortening the design time of modern embedded systems. It is shown that abstractions for communication are as important as for behaviour for specification and system level abstractions, and an extension on a novel higher level communication mechanism which has features for supporting the description of complex aggregate associations between objects in specifications such as UML is investigated. The communication primitives have been implemented as extensions to SystemC, and a comprehensive example from a UML specification through functional specification down to an executable SystemC description is included.
Kjetil Svarstad, Gabriela Nicolescu, Ahmed Amine Jerraya
DATE2
2000 Towards design and validation of mixed-technology SOCs
abstract
This paper illustrates an approach to design and validation of heterogeneous systems. The emphasis is placed on devices which incorporate MEMS parts in either a single mixed-technology (CMOS + micromachining) SOC device, or alternatively as a hybrid system with the MEMS part in a separate chip. The design flow is general, and it is illustrated for the case of applications embedding CMOS sensors. In particular, applications based on finger-print recognition are considered since a rich variety of sensors and data processing algorithms can be considered. A high level multi-language/multi-engine approach is used for system specification and co-simulation. This also allows for an initial high-level architecture exploration, according to performance and cost requirements imposed by the target application. Thermal simulation of the overall device, including packaging, is also considered since this can have a significant impact in sensor performance. From the selected system specification, the actual architecture is finally generated via a multi-language co-design approach which can result in both hardware and software parts. The hardware parts are composed of available IP cores. For the case of a single chip implementation, the most important issue of embedded-core-based testing is briefly considered, and current techniques are adapted for testing the embedded cores in the SOC devices discussed.
Salvador Mir, Benoît Charlot, Gabriela Nicolescu, Philippe Coste, Fabien Parrain, Nacer-Eddine Zergainoh, Bernard Courtois, Ahmed Amine Jerraya, Márta Rencz
ACM Great Lakes Symposium on VLSI3
2000 Multi-Level Communication Synthesis of Heterogeneous Multilanguage Specification
abstract
The complexity of modern embedded systems requires the cooperation of several teams belonging to different cultures and using different languages as well as the reuse of software, hardware and communication IP modules at the early design steps. The key issue for the design of such systems is the overall system validation and the synthesis of the communication between the different subsystems. In this paper we focus on the problem of multi-level communication synthesis and show the results of the application of this methodology on an example. Designers get feedback at all design steps via the cosimulation engine that permits fast evaluation.
Fabiano Hessel, Philippe Coste, Gabriela Nicolescu, P. LeMarrec, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya
ICCD3