Sébastien Le Beux

dblp:44/1238 · DBLP profile ↗
← Back
40ranked-venue papers
10as first author
8since 2021 · last 2026
0000-0003-1778-0253ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 10 first-author · 8 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Leveraging Recurrent Patterns in Graph Accelerators
abstract
Graph accelerators have emerged as a promising solution for processing large-scale sparse graphs, leveraging the in-situ computation of ReRAM-based crossbars to maximize computational efficiency. However, existing designs suffer from memristor access overhead due to the large number of graph partitions. This leads to increased execution time, higher energy consumption, and reduced circuit lifetime. This paper proposes a graph processing method that minimizes memristor write operations by identifying frequent subgraph patterns and assigning them to graph engines, referred to as static, allowing most subgraphs to be processed without a need for crossbar reconfiguration. Experimental results show speed up to 2.38× speedup and 7.23× energy savings compared to state-of-the-art accelerators. Furthermore, our method extends the circuit lifetime by 2× compared to state-of-the-art ReRAM graph accelerators.
Sébastien Le Beux
DATE2
2024 Signed Convolution in Photonics with Phase-Change Materials using Mixed-Polarity Bitstreams
abstract
As AI continues to grow in importance, in order to reduce its carbon footprint and utilization of computer resources, numerous alternatives are under investigation to improve its hardware building blocks. In particular, in convolutional neural networks (CNNs), the convolution function represents the most important operation and one of the best targets for optimization. A new approach to convolution had recently emerged using optics, phase-change materials (PCMs) and stochastic computing, but is thus far limited to unsigned operands. In this paper, we propose an extension in which the convolutional kernels are signed, using mixed-polarity bitstreams. We present a proof of validity for our method, while also showing that, in simulation and under similar operating conditions, our approach is less affected by noise than the common approach in the literature.
Raphael Cardoso, Clément Zrounba, Mohab Abdalla, Paul Jiménez, Mauricio Gomes de Queiroz, Benoît Charbonnier, Fabio Pavanello, Ian O'Connor, Sébastien Le Beux
ASPDAC9
2023 Towards a Robust Multiply-Accumulate Cell in Photonics using Phase-Change Materials
abstract
In this paper we propose a novel approach to multiply-accumulate (MAC) operations in photonics. This approach is based on stochastic computing and on the dynamic behavior of phase-change materials (PCMs), leading to the unique characteristic of automatically storing the result in non-volatile memory. We demonstrate that, even with perfect look-up tables, the standard approach to PCM scalar multiplication is highly susceptible to perturbations as small as 0.1% of the input power, causing repetitive peaks of 600% relative error. In the same operating conditions, the proposed method achieves an average of 7× improvement in precision.
Raphael Cardoso, Clément Zrounba, Mohab Abdalla, Paul Jiménez, Mauricio Gomes de Queiroz, Benoît Charbonnier, Fabio Pavanello, Ian O'Connor, Sébastien Le Beux
DATE9
2022 Non-Volatile Phase Change Material based Nanophotonic Interconnect
abstract
Integrated optics is a promising technology to take advantage of light propagation for high throughput chip-scale interconnects in many core architectures. A key challenge for the deployment of nanophotonic interconnects is their high static power, which is induced by signal losses and devices calibration. To tackle this challenge, we propose to use Phase Change Material (PCM) to configure optical paths between writers and readers. The non-volatility of PCM elements and the high contrast between crystalline and amorphous phase states allow to bypass unused readers, thus reducing losses and calibration requirements. We evaluate the efficiency of the proposed PCM-based interconnects using system level simulations carried out with SNIPER manycore simulator. For this purpose, we have modified the simulator to partition clusters according to executed applications. Simulation results show that bypassing readers using PCM leads up to 52% communication power saving.
Parya Zolfaghari, Joel Ortiz, Cédric Killian, Sébastien Le Beux
DATE4
2022 Toward Large Scale All-Optical Spiking Neural Networks
abstract
Silicon Photonics is a promising technology to develop neuromorphic hardware accelerators. Most optical neural networks rely on wavelength division multiplexing (WDM), which calls for power hungry calibration to compensate for non-uniformity fabrication process and thermal variations of microring resonators (MRR). This imposes practical limits on neuromorphic photonic hardware since only a small number of synaptic connections per neuron can be implemented. As a result, the mapping of neural networks (NN) on a hardware platform require pruning of synaptic connections, which drastically affects the accuracy. In this work, we propose a method to efficiently map pre-trained NN on an all-optical spiking neural network (SNN), with the aim to optimize hardware utilization while minimizing accuracy loss. The method relies on weight partitioning and unrolling to reduce synaptic connections. The resulting neural networks are mapped on an architecture we propose, allowing to estimate accuracy and power consumption. Results show the capability of weight partitioning to implement a realistic NN while attaining 58% reduction in energy consumption compared with unrolling.
Milad Eslaminia, Sébastien Le Beux
VLSI-SoC2
2022 Towards All-optical Stochastic Computing Using Photonic Crystal Nanocavities
abstract
Stochastic computing allows a drastic reduction in hardware complexity using serial processing of bit streams. While the induced high computing latency can be overcome using integrated optics technology, the design of realistic optical stochastic computing architectures calls for energy efficient switching devices. Photonics Crystal (PhC) nanocavities are μm 2 scale devices offering 100fJ switching operation under picoseconds-scale switching speed. Fabrication process allows controlling the Quality factor of each nanocavity resonance, leading to opportunities to implement architectures involving cascaded gates and multi-wavelength signaling. In this paper, we investigate the design of cascaded gates architecture using nanocavities in the context of stochastic computing. We propose a transmission model considering key nanocavity device parameters, such as Quality factors, resonance wavelength, and switching efficiency. The model is calibrated with experimental measurements. We propose the design of XOR gate and multiplexer. We illustrate the use of the gates to design an edge detection filter. System-level exploration of laser power, bit-stream length and bit-error rate is carried out for the processing of gray-scale images. The results show that the proposed architecture leads to 8.5nJ/pixel energy consumption and 512ns/pixel processing time.
Hassnaa El-Derhalli, Léa Constans, Sébastien Le Beux, Alfredo De Rossi, Fabrice Raineri, Sofiène Tahar
ACM J. Emerg. Technol. Comput. Syst.3
2022 Distance-aware Approximate Nanophotonic Interconnect
abstract
The energy consumption of manycore architectures is dominated by data movement, which calls for energy-efficient and high-bandwidth interconnects. To overcome the bandwidth limitation of electrical interconnects, integrated optics appear as a promising technology. However, it suffers from high power overhead related to low laser efficiency, which calls for the use of techniques and methods to improve its energy costs. Besides, approximate computing is emerging as an efficient method to reduce energy consumption and improve execution speed of embedded computing systems. It relies on allowing accuracy reduction on data at the cost of tolerable application output error. In this context, the work presented in this article exploits both features by defining approximate communications for error-tolerant applications. We propose a method to design realistic and scalable nanophotonic interconnect supporting approximate data transmission and power adaption according to the communication distance to improve the energy efficiency. For this purpose, the data can be sent by mixing low optical power signal and truncation for the Least Significant Bits (LSB) of the floating-point numbers, while the overall power is adapted according to the communication distance. We define two ranges of communications, short and long, which require only four power levels. This reduces area and power overhead to control the laser output power. A transmission model allows estimating the laser power according to the targeted BER and the number of truncated bits, while the optical network interface allows configuring, at runtime, the number of approximated and truncated bits and the laser output powers. We explore the energy efficiency provided by each communication scheme, and we investigate the error resilience of the benchmarks over several approximation and truncation schemes. The simulation results of ApproxBench applications show that, compared to an interconnect involving only robust communications, approximations in the optical transmission led to up to 53% laser power reduction with a limited degradation at the application level with less than 9% of output error. Finally, we show that our solution is scalable and leads to 10% reduction in the total energy consumption, 35× reduction in the laser driver size, and 10× reduction in the laser controller compared to state-of-the-art solution.
Jaechul Lee, Cédric Killian, Sébastien Le Beux, Daniel Chillet
ACM Trans. Design Autom. Electr. Syst.3
2021 A Reconfigurable Nanophotonic Architecture based on Phase Change Material
abstract
Silicon photonics is an emerging technology allowing to take the advantage of high-speed light propagation to accelerate computing kernels in integrated systems. Micrometer-scale optical devices call for reconfigurable architectures to maximize resources utilization. Typical reconfigurable optical computing architectures involve micro-ring resonators for electro-optic modulation. However, such devices require voltage and thermal tuning to compensate for fabrication process variability and thermal sensitivity. This power-hungry calibration leads to significant static power overhead, thus limiting the scalability of optical architectures. In this paper, we propose to use non-volatile Phase Change Materials (PCM) elements to route optical signals only through the required resonators, hence saving calibration energy of bypassed resonators. The non-volatility of PCM elements allows maintaining the optical path. We investigate the efficiency of the PCM elements on the Reconfigurable Directed Logic (RDL) architecture. Results show that the static power is reduced by 32.8% on average and that 30% power saving is obtained from 158kHz reconfiguration frequency.
Parya Zolfaghari, Sébastien Le Beux
VLSI-SoC2
2020 OSCAR: An Optical Stochastic Computing AcceleRator for Polynomial Functions
abstract
Approximate computing allows improving design energy efficiency at the cost of computing accuracy. Stochastic computing is an approximate computing technique, where numbers are represented as probabilities using stochastic bit streams. The serial processing of the bit streams leads to reduced hardware complexity but induces high processing latency. Silicon photonics has the potential to overcome this limitation thanks to high propagation speed of signals and high bandwidth. However, the technology remains costly, which calls for optical accelerators capable to adapt to application-specific requirements. In this paper, we propose a reconfigurable optical accelerator capable to adapt to computing accuracy, energy efficiency, and throughput objectives. The architecture can be configured to execute i) 4thorder function for high accuracy processing or ii) 2ndorder function for high-energy efficiency or high throughput purposes. Evaluations are carried out using image processing Gamma correction application. Compared to a static architecture for which accuracy is defined at design time, the proposed architecture leads to 36.8% energy overhead but increases the range of reachable accuracy by 65%.
Hassnaa El-Derhalli, Sébastien Le Beux, Sofiène Tahar
DATE2
2020 3D Logic Cells Design and Results Based on Vertical NWFET Technology Including Tied Compact Model
abstract
Gate-all-around Vertical Nanowire Field Effect Transistors (VNWFET) are emerging devices., which are well suited to pursue scaling beyond lateral scaling limitations around 7nm. This work explores the relative merits and drawbacks of the technology in the context of logic cell design. We describe a junctionless nanowire technology and associated compact model., which accurately describes fabricated device behavior in all regions of operations for transistors based on between 16 and 625 parallel nanowires of diameters between 22 and 50nm. We used this model to simulate the projected performance of inverter logic gates based on passive load., active load and complementary topologies and carry out an performance exploration for the number of nanowires in transistors. In terms of compactness., through a dedicated full 3D layout design., we also demonstrate a 48% reduction in lateral dimensions for the complementary structure with respect to 7nm FinFET-based inverters.
C. Mukherjee 0001, Marina Deng, François Marc, Cristell Maneux, Arnaud Poittevin, Ian O'Connor, Sébastien Le Beux, Cédric Marchand 0002, Aurélie Lecestre, Guilhem Larrieu
VLSI-SOC7
2019 Stochastic Computing with Integrated Optics
abstract
Stochastic computing (SC) allows reducing hardware complexity and improving energy efficiency of error resilient applications. However, a main limitation of the computing paradigm is the low throughput induced by the intrinsic serial computing of bit-streams. In this paper, we address the implementation of SC in the optical domain, with the aim to improve the computation speed. We implement a generic optical architecture allowing the execution of polynomial functions. We propose design methods to explore the design space in order to optimize key metrics such as circuit robustness and power consumption. We show that a circuit implementing a 2ndorder polynomial degree function and operating at 1Ghz leads to 20.1pJ laser consumption per computed bit.
Hassnaa El-Derhalli, Sébastien Le Beux, Sofiène Tahar
DATE2
2019 Approximate nanophotonic interconnects
abstract
The energy consumption of manycore is dominated by data movement, which calls for energy-efficient and high-bandwidth interconnects. Integrated optics is promising technology to overcome the bandwidth limitations of electrical interconnects. However, it suffers from high power overhead related to low efficiency lasers, which calls for the use of approximate communications for error tolerant applications. In this context, this paper investigates the design of an Optical NoC supporting the transmission of approximate data. For this purpose, the least significant bits of floating point numbers are transmitted with low power optical signals. A transmission model allows estimating the laser power according to the targeted BER and a micro-architecture allows configuring, at run-time, the number of approximated bits and the laser output powers. Simulations results show that, compared to an interconnect involving only robust communications, approximations in the optical transmission lead to up to 42% laser power reduction for image processing application with a limited degradation at the application level.
Jaechul Lee, Cédric Killian, Sébastien Le Beux, Daniel Chillet
NOCS3
2019 Thermal-Aware Design Method for Laser Group Control in Nanophotonic Interconnects
abstract
On-chip integrated lasers are key devices to deliver the high bandwidth expected from nanophotonic interconnects. However, lasers are highly sensitive to temperature variation, which influences the lasing efficiency and the wavelengths of emitted optical signals, both of which are key factors in interconnect power efficiency. It is, thus, necessary to develop techniques for efficient thermal-aware control of lasers. In this brief, we propose the grouping of lasers for efficient power control of their temperature. Laser grouping is carried out taking into account the layout symmetries, and a design method allows the definition of control laws.
Amira Aouina, Hui Li 0034, Ian O'Connor, Gabriela Nicolescu, Sébastien Le Beux
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Large scale, high density integration of all spin logic
abstract
Spintronics brings new features that make it a viable candidate technology to implement non-conventional processing for new computing paradigms in an efficient way. The first milestone of the spintronics roadmap was the fabrication of hybrid systems where the data processing relies mostly on charge-based electronics devices (CMOS), while the memory hierarchy is partially or totally replaced by MRAM. In the next step, spintronics can also be used for data processing, still in conjunction with CMOS. Nevertheless, replacing all the processing by pure spintronic circuits, without any charge current, remains the ultimate objective of spintronics. All spin logic (ASL) paves the way towards that goal, even if some CMOS control circuits are still necessary. However, as ASL does not rely on the same computing principle as CMOS, it is necessary to address some specific issues. Pure spin current propagates in every direction, including backwards in the presence of multiple inputs; and is divided when crossings are encountered. It combines mainly linearly, while logic operations require non-linear binary decisions. Interconnect between logic gates requires directionality from inputs to outputs, and fanout with negligible signal attenuation. In this context, we develop new strategies for ASL modeling and logic design. We propose an architecture and a design strategy based on a high-density array to address the specific issues of directionality, attenuation and linearity. Moreover, the feasibility is supported through the modeling and the simulation of its basic block. This implies modularity to simulate complex circuits, even when they are ahead of today's experimental demonstrations.
Sébastien Le Beux, Ian O'Connor, Jacques-Olivier Klein
DATE2
2018 Offline Optimization of Wavelength Allocation and Laser Power in Nanophotonic Interconnects
abstract
Optical Network-on-Chip (ONoC) is a promising communication medium for large-scale multiprocessor systems-on-chips. Indeed, ONoC can outperform classical electrical NoCs in terms of energy efficiency and bandwidth density, in particular, because this medium can support multiple transactions at the same time on different wavelengths by using Wavelength Division Multiplexing (WDM). However, multiple signals sharing simultaneously the same part of a waveguide can lead to inter-channel crosstalk noise. This problem impacts the signal-to-noise ratio of the optical signals, which leads to an increase in the Bit Error Rate (BER) at the receiver side. If a specific BER is targeted, an increase of laser power should be necessary to satisfy the SNR. In this context, an important issue is to evaluate the laser power needed to satisfy the various desired communication bandwidths based on the BER performance requirements. In this article, we propose an off-line approach that concurrently optimizes the laser power scaling and execution time of a global application. A set of different levels of power is introduced for each laser, to ensure that optical signals can be emitted with just-enough power to ensure targeted BER. As a result, most promising solutions are highlighted for mapping a defined application onto a 16-core ring-based WDM ONoC.
Jiating Luo, Cédric Killian, Sébastien Le Beux, Daniel Chillet, Olivier Sentieys, Ian O'Connor
ACM J. Emerg. Technol. Comput. Syst.3
2017 Energy and Performance Trade-off in Nanophotonic Interconnects using Coding Techniques
abstract
Nanophotonic is an emerging technology considered as one of the key solutions for future generation on-chip interconnects. Indeed, this technology provides high bandwidth for data transfers and can be a very interesting alternative to bypass the bottleneck induced by classical NoC. However, their implementation in fully integrated 3D circuits remains uncertain due to the high power consumption of on-chip lasers. However, if a specific bit error rate is targeted, digital processing can be added in the electrical domain to reduce the laser power and keep the same communication reliability. This paper addresses this problem and proposesto transmit encoded data on the optical interconnect, which allows for a reduction of the laser power consumption, thus increasing nanophotonics interconnects energy efficiency. The results presented in this paper show that using simple Hamming coder and decoder permits to reduce the laser power by nearly 50% without loss in communication data rate and with a negligible hardware overhead.
Cédric Killian, Daniel Chillet, Sébastien Le Beux, Van-Dung Pham, Olivier Sentieys, Ian O'Connor
DAC3
2017 Performance and energy aware wavelength allocation on ring-based WDM 3D optical NoC
abstract
Optical Network-on-Chip (ONoC) is a promising communication medium for large-scale Multiprocessor System on Chip (MPSoC). ONoC outperforms classical electrical NoC in terms of throughput and latency. The medium can support multiple transactions at the same time on different wavelengths by using Wavelength Division Multiplexing (WDM). Moreover multiple wavelengths can be used as high-bandwidth channel to reduce transmission time. However, multiple signals sharing simultaneously a waveguide can lead to inter-channel crosstalk noise. This problem impacts the Signal to Noise Ratio (SNR) of the optical signal, which leads to an increase in the Bit Error Rate (BER) at the receiver side. In this paper we first formulate the crosstalk noise and execution time models and then propose a Wavelength Allocation (WA) method in a ring-based WDM ONoC allowing to search for performance and energy trade-offs, based on the application constraints. As result, most promising WA solutions are highlighted for a defined application mapping onto 16-core WDM ONoC.
Jiating Luo, A. Elantably, Van-Dung Pham, Cédric Killian, Daniel Chillet, Sébastien Le Beux, Olivier Sentieys, Ian O'Connor
DATE6
2017 Energy-Efficiency Comparison of Multi-Layer Deposited Nanophotonic Crossbar Interconnects
abstract
Single-layer optical crossbar interconnections based on Wavelength Division Multiplexing stand among other nanophotonic interconnects because of their low latency and low power. However, such architectures suffer from a poor scalability due to losses induced by long propagation distances on waveguides and waveguide crossings. Multi-layer deposited silicon technology allows the stacking of optical layers that are connected by means of Optical Vertical Couplers. This allows significant reduction in the optical losses, which contributes to improve the interconnect scalability but also leads to new challenges related to network designs and layouts. In this article, we investigate the design of optical crossbars using multi-layer silicon deposited technology. We propose implementations for Ring-, Matrix-, λ-router-, and Snake-based topologies. Layouts avoiding waveguide crossings are compared to those minimizing the waveguide length according to worst-case and average losses. The laser output power is estimated from the losses, which allows us to evaluate the energy efficiency improvement induced by multi-layer technology over traditional planar implementations (33% on average). Finally, networks comparison has been carried out and the results show that the ring topology leads to a 43% reduction in the laser output power.
Hui Li 0034, Sébastien Le Beux, Martha Johanna Sepúlveda, Ian O'Connor
ACM J. Emerg. Technol. Comput. Syst.2
2016 Coherent and Incoherent Crosstalk Noise Analyses in Interchip/Intrachip Optical Interconnection Networks
abstract
Recently, interchip/intrachip optical interconnection networks have been proposed for ultrahigh-bandwidth and low-latency communications. These networks employ the microresonators (MRs) to modulate, direct, or detect the optical signal. However, utilized MRs suffer from intrinsic crosstalk noise and signal power loss, degrading the network efficiency via the signal-to-noise ratio (SNR). The amount of crosstalk noise and signal power loss may differ from network to network. Hence, there exists a need to systematically analyze the effect of the crosstalk noise and the power loss issues. In this paper, we have developed the analytical models considering both coherent and incoherent crosstalk for both the interchip and intrachip optical networks. The interchip/intrachip optical interconnection networks—the$\text{I}^{2}$CON—are analyzed as a case study. The quantitative results on the individual networks have demonstrated that the architectural design determines the impact of crosstalk on the SNR. We have also demonstrated that the optical interconnection networks with interchip/intrachip interconnects result in better bit error rate (BER) compared with that of only intrachip interconnect. Our analyses of the worst case can be utilized as a platform to compare the realistic performance among different optical interconnection networks via the degradation of SNR/BER and data bandwidth.
Luan H. K. Duong, Zhehui Wang, Mahdi Nikdast, Jiang Xu 0001, Peng Yang 0003, Zhe Wang 0003, Rafael Kioji Vivas Maeda, Haoran Li 0002, Xuan Wang 0001, Sébastien Le Beux, Yvain Thonnart
IEEE Trans. Very Large Scale Integr. Syst.11
2015 Energy-efficient optical crossbars on chip with multi-layer deposited silicon
abstract
The many cores design research community have shown high interest in optical crossbars on chip for more than a decade. Key properties of optical crossbars, namely a) contention-free data routing b) low-latency communication and c) potential for high bandwidth through the use of WDM, motivate several implementations. These implementations demonstrate very different scalability and power efficiency ability depending on three key design factors: a) the network topology, b) the considered layout and c) the insertion losses induced by the fabrication process. The emerging design technique relying on multi-layer deposited silicon allows reducing optical losses, which may lead to significant reduction of the power consumption. In this paper, multi-layer deposited silicon based crossbars are proposed and compared. The results indicate that the proposed ring-based network exhibits, on average, 22% and 51.4% improvement for worst-case and average losses respectively compared to the most power-efficient related crossbars.
Hui Li 0034, Sébastien Le Beux, Gabriela Nicolescu, Ian O'Connor
ASP-DAC2
2015 Complementary communication path for energy efficient on-chip optical interconnects
abstract
Optical interconnects are considered to be one of the key solutions for future generation on-chip interconnects. However, energy efficiency is mainly limited by the losses incurred by the optical signals, which considerably reduces the optical power received by the photodetectors. In this paper we propose a differential transmission of the modulated signals, which contributes to improve the transmission of the optical signal power on the receiver side. With this approach, it is possible to reduce the input laser power and increase the energy efficiency of the optical communication. The approach is generic and can be applied to SWSR-, MWSR-, SWMR- and MWMR-like architectures.
Hui Li 0034, Sébastien Le Beux, Yvain Thonnart, Ian O'Connor
DAC2
2015 Coherent crosstalk noise analyses in ring-based optical interconnects
Luan H. K. Duong, Mahdi Nikdast, Jiang Xu 0001, Zhehui Wang, Yvain Thonnart, Sébastien Le Beux, Peng Yang 0003, Xiaowen Wu
DATE6
2015 Thermal aware design method for VCSEL-based on-chip optical interconnect
Hui Li 0034, Alain Fourmigue, Sébastien Le Beux, Xavier Letartre, Ian O'Connor, Gabriela Nicolescu
DATE3
2014 Chameleon: Channel efficient Optical Network-on-Chip
abstract
The next generation of MPSoC points to the integration of thousands of IP cores, requiring high performance interconnect for high throughput communications. Optical on-chip interconnect enables significantly increased bandwidth and decreased latency in MPSoC. However, the interface between electrical and photonic devices implies strong layout constraints that may impact the system performance and scalability. In this paper, we propose a novel optical interconnect named Chameleon. The interface simplifies the layout and allows the bandwidth between IP cores to be adapted according to the communication requirements. Compared to related networks, Chameleon demonstrates improved scalability and flexibility at the cost of minor increase in power consumption.
Sébastien Le Beux, Hui Li 0034, Ian O'Connor, Kazem Cheshmi, Xuchen Liu 0003, Jelena Trajkovic, Gabriela Nicolescu
DATE1
2014 CLAP: a crosstalk and loss analysis platform for optical interconnects
abstract
Basic photonic devices in inter- and intra-chip optical networks suffer from inevitable power loss and crosstalk noise. Incoherent crosstalk introduces quick power fluctuations, while coherent crosstalk varies the optical power of the optical signal in optical interconnection networks (OINs). As a result, the accumulative crosstalk in large scale OINs considerably hurts the signal-to-noise ratio (SNR) and imposes high power penalties. In this work, we aim at studying the worst-case incoherent and coherent crosstalk in OINs at the system level. The proposed analytical models are integrated into a newly developed crosstalk and loss analysis platform, called CLAP, to facilitate the SNR analyses in arbitrary OINs.
Mahdi Nikdast, Luan H. K. Duong, Jiang Xu 0001, Sébastien Le Beux, Xiaowen Wu, Zhehui Wang, Peng Yang 0003, Yaoyao Ye
NOCS4
2014 Introduction to the special session on "Silicon photonic interconnects: an illusion or a realistic solution?"
abstract
The performance of a multiprocessor system-on-chip (MPSoC) is determined not only by the performance of its processing cores and memories, but also by how efficiently they collaborate with one another. It is the MPSoCs communication architecture which determines the collaboration efficiency. The migration towards MPSoCs is propelled by the shrinking feature sizes in each generation of process technology. On the one hand, smaller transistors allow for more processor cores and memories on a single chip and result in more on-chip computations as well as communications. On the other hand, reducing feature sizes makes on-chip communication more difficult. The International Roadmap for Semiconductors (ITRS) shows that the latency of metallic interconnects increases exponentially as feature sizes decrease. On-chip communication using metallic interconnects will need more than one clock cycle to send information from sources to destinations. Moreover, metallic interconnects consume a significant amount of power. Studies shows that global metallic interconnects could consume kilowatts of power to achieve required communication bandwidth by 2020 [1].
Jiang Xu 0001, Sébastien Le Beux, Yvain Thonnart
NOCS2
2014 Complementary logic interface for high performan optical computing with OLUT
abstract
The Optical LUT (OLUT) has been proposed as a parallel and energy-efficient logic architecture for building prospective on-chip optical FPGAs in order to replace traditional power-hungry electronic computing circuits. In this paper, we present a new OLUT implementation that computes a pair of complementary Boolean logic functions through wavelength multiplexing, allowing the computational capacity to be doubled for a reasonable optical power and area overhead depending on the OLUT size. Worst-case evaluation of the optical laser power needed to perform reliable logic operations demonstrates the potential of the proposed OLUT for energy- and hardware-efficient photonic reconfigurable computing.
Zhen Li 0046, Sébastien Le Beux, Ian O'Connor, Christelle Monat, Xavier Letartre
VLSI-SoC2
2014 Optical crossbars on chip, a comparative study based on worst-case losses
abstract
SUMMARY The many‐core design research community has shown high interest in optical crossbar on chip for more than a decade. Key properties of optical crossbars, namely (1) contention‐free data routing, (2) low latency communication, and (3) potential for high bandwidth through the use of wavelength division multiplexing, motivate several implementations of this type of interconnect. These implementations demonstrate very different scalability and power efficiency abilities depending on three key design factors: (1) network topology, (2) considered layout, and (3) insertion losses induced by the fabrication process. In this paper, the worst‐case optical losses of crossbar implementations are compared according to the factors mentioned earlier. The comparison results have the potential to help many‐core system designer to select the most appropriate crossbar implementation according to, for instance, the number of IP cores and the die size. Copyright © 2014 John Wiley & Sons, Ltd.
Sébastien Le Beux, Hui Li 0034, Gabriela Nicolescu, Jelena Trajkovic, Ian O'Connor
Concurr. Comput. Pract. Exp.1
2013 Optical look up table
abstract
The computation capacity of conventional FPGAs is directly proportional to the size and expressive power of Look Up Table (LUT) resources. Individual LUT performance is limited by transistor switching time and power dissipation, defined by the CMOS fabrication process. In this paper we propose OLUT, an optical core implementation of LUT, which has the potential for low latency and low power computation. In addition, the use of Wavelength Division Multiplexing (WDM) allows parallel computation, which can further increase computation capacity. Preliminary experimental results demonstrate the potential for optically assisted on-chip computation.
Zhen Li 0046, Sébastien Le Beux, Christelle Monat, Xavier Letartre, Ian O'Connor
DATE2
2013 Potential and pitfalls of silicon photonics computing and interconnect
abstract
Trends in SoC design are leading to 3D integration of thousands of high-performance computing resources and high-throughput interconnects, opening up new research directions for hybrid electronic/photonic architectures. In this paper, we introduce how state of the art silicon-photonic devices can realize elementary operations that are traditionally performed by electronic devices, e.g. circuit switching and Boolean function computation. We then highlight how these devices need to be assembled in order to realize more complex functions, taking into account the constraints specific to silicon-photonic technology. In the last part, we summarize the main research directions for the near future.
Sébastien Le Beux, Ian O'Connor, Zhen Li 0046, Xavier Letartre, Christelle Monat, Jelena Trajkovic, Gabriela Nicolescu
ISCAS1
2013 Reconfigurable photonic switching: Towards all-optical FPGAs
abstract
FPGA performance is limited by transistor switching time and power dissipation defined by the CMOS technology. Similarly to the case of interconnects, silicon photonics can also be leveraged to replace traditional, slow and power consuming, electrical computing circuits. In this paper we propose an all Optical LUT, an optical core implementation of LUT, which has the potential for low latency and relatively low power computation. The OLUT unique feature is its compliant input and output optical interfaces, which allows assembly to realize complex functions. The use of Wavelength Division Multiplexing (WDM) allows both parallel communication and computation, which can further increase the system performance. Preliminary results illustrate how the proposed OLUT can be assembled, giving the trends to the design of an all optical FPGA.
Sébastien Le Beux, Zhen Li 0046, Christelle Monat, Xavier Letartre, Ian O'Connor
VLSI-SoC1
2012 Ambipolar double-gate FETs for the design of compact logic structures
abstract
We present in this paper a circuit design approach to achieve compact logic circuits with ambipolar double-gate devices, using the in-field controllability of such devices. The approach is demonstrated for complementary static logic design style. We apply this approach in a case study focused on Double Gate Carbon Nanotube FET (DG-CNTFET) technology and show that, with respect to conventional CMOS-like static logic structures and for comparable power consumption, time delay and integration density can both be improved by a factor of 1.5x and 2x, respectively. Compared with a predictive model for 16nm CMOS technology, the gates built according to the design approach described in this work and based on DG-CNTFET offer a gain of 30% concerning Power-Delay-Product (PDP).
Kotb Jabeur, Ian O'Connor, Nataliya Yakymets, Sébastien Le Beux
ACM Great Lakes Symposium on VLSI4
2011 Optical Ring Network-on-Chip (ORNoC): Architecture and design methodology
abstract
State-of-the-art System-on-Chip (SoC) consists of hundreds of processing elements, while trends in design of the next generation of SoC point to integration of thousand of processing elements, requiring high performance interconnect for high throughput communications. Optical on-chip interconnects are currently considered as one of the most promising paradigms for the design of such next generation Multi-Processors System on Chip (MPSoC). They enable significantly increased bandwidth, increased immunity to electromagnetic noise, decreased latency, and decreased power. Therefore, defining new architectures taking advantage of optical interconnects represents today a key issue for MPSoC designers. Moreover, new design methodologies, considering the design constraints specific to these architectures are mandatory. In this paper, we present a contention-free new architecture based on optical network on chip, called Optical Ring Network-on-Chip (ORNoC). We also show that our network scales well with both large 2D and 3D architectures. For the efficient design, we propose automatic wavelength-/waveguide assignment and demonstrate that the proposed architecture is capable of connecting 1296 nodes with only 102 waveguides and 64 wavelengths per waveguide.
Sébastien Le Beux, Jelena Trajkovic, Ian O'Connor, Gabriela Nicolescu, Guy Bois, Pierre G. Paulin
DATE1
2011 Fine-grain reconfigurable logic cells based on double-gate CNTFETs
abstract
This paper presents 2-inputs cells designed to perform reconfigurable operations in nanometric systems exploiting the ambipolar property of double-gate (DG) carbon nanotube (CNT) FETs. Previous work [1] described a dynamic logic cell generating only 14 functions instead of 16 normally performed by the multiplexer-based logic part of a CLB (Configurable Logic Block) of an FPGA for 2-inputs. In this work, a reconfigurable 2-input dynamic logic cell designed using DG-CNTFET devices is able to achieve the whole set of 16 functions exploiting a specific correlation between input and configuration signals to offer full functionality over previous version. We also built a reconfigurable 2-input static logic cell which performs 16 functions. Both cells demonstrate a significant reduction in circuit complexity with respect to conventional CMOS-based reconfigurable cells for equivalent functionality. Compared with a 2-LUT, the dynamic cell improve the time delay by a factor of 2X to the detriment of 2X increase in power consumption, while the static logic cell shows an improvement of 2X in term of power consumption and time delay.
Kotb Jabeur, Nataliya Yakymets, Ian O'Connor, Sébastien Le Beux
ACM Great Lakes Symposium on VLSI4
2011 Layout guidelines for 3D architectures including Optical Ring Network-on-Chip (ORNoC)
abstract
Trends in design of the next generation of Multi-Processors System on Chip (MPSoC) point to 3D integration of thousand of processing elements, requiring high performance interconnect for high throughput and low latency communications. Optical on-chip interconnects enable significantly increased bandwidth and decreased latency. They are thus considered as one of the most promising paradigms for the design of such system. However, existence of interfaces between electronic and photonic signals implies strong constraints on the layout of the 3D architecture and may impact the architecture scalability. In this paper, we propose and evaluate a possible layout for an optical Network-on-Chip used to interconnect processing elements located on different electrical layers.
Sébastien Le Beux, Jelena Trajkovic, Ian O'Connor, Gabriela Nicolescu
VLSI-SoC1
2011 A Model-Driven Design Framework for Massively Parallel Embedded Systems
abstract
Modern embedded systems integrate more and more complex functionalities. At the same time, the semiconductor technology advances enable to increase the amount of hardware resources on a chip for the execution. Massively parallel embedded systems specifically deal with the optimized usage of such hardware resources to efficiently execute their functionalities. The design of these systems mainly relies on the following challenging issues: first, how to deal with the parallelism in order to increase the performance; second, how to abstract their implementation details in order to manage their complexity; third, how to refine these abstract representations in order to produce efficient implementations. This article presents the Gaspard design framework for massively parallel embedded systems as a solution to the preceding issues. Gaspard uses the repetitive Model of Computation (MoC), which offers a powerful expression of the regular parallelism available in both system functionality and architecture. Embedded systems are designed at a high abstraction level with the MARTE (Modeling and Analysis of Real-time and Embedded systems) standard profile, in which our repetitive MoC is described by the so-called Repetitive Structure Modeling (RSM) package. Based on the Model-Driven Engineering (MDE) paradigm, MARTE models are refined towards lower abstraction levels, which make possible the design space exploration. By combining all these capabilities, Gaspard allows the designers to automatically generate code for formal verification, simulation and hardware synthesis from high-level specifications of high-performance embedded systems. Its effectiveness is demonstrated with the design of an embedded system for a multimedia application.
Abdoulaye Gamatié, Sébastien Le Beux, Éric Piel, Rabie Ben Atitallah, Anne Etien, Philippe Marquet, Jean-Luc Dekeyser
ACM Trans. Embed. Comput. Syst.2
2010 A system-level exploration flow for optica network on chip (ONoC) in 3D MPSoC
abstract
Optical on-chip interconnects and 3D die stacking are currently considered to be two promising paradigms for the design of next generation Multi-Processors System on Chip architectures (MPSoC). New architectures based on these paradigms are currently emerging and new system-level approaches are required for their efficient design and prototype. The paper investigates a system-level flow for evaluating design feasibility, interconnect architecture performance and application execution efficiency as early as possible in the MPSoC design cycle.
Sébastien Le Beux, Gabriela Nicolescu, Guy Bois, Pierre G. Paulin
ISCAS1
2010 Combining mapping and partitioning exploration for NoC-based embedded systems
Sébastien Le Beux, Guy Bois, Gabriela Nicolescu, Youcef Bouchebaba, Michel Langevin, Pierre G. Paulin
J. Syst. Archit.1
2007 A Design Flow to Map Parallel Applications onto FPGAs
abstract
This paper introduces a new flow able to fit a parallel application onto an FPGA according to the FPGA characteristics such as computing power and IOs. The flow is based on iterative refactoring and transformations of the application. From the resulting application, a VHDL code is generated. This code is finally used to simulate or synthesize the application. Significant experiments have validated the approach.
Sébastien Le Beux, Philippe Marquet, Jean-Luc Dekeyser
FPL1
2006 FPGA Implementation of Embedded Cruise Control and Anti-Collision Radar
abstract
The ModEasy project seeks to develop techniques and software tools to aid in the development of reliable microprocessor based electronic (embedded) systems using advanced development and verification systems. The tools are to be evaluated in practical domains such as the automotive sector for reactive cruise control and anti-collision radar. We choose to define specific IPs using FPGA techniques to cover this application domain. This paper presents the implementation of such a complex and safety application on a single FPGA. The target system is composed of a reactive cruise control, a detection radar and the associated treatments
Sébastien Le Beux, Philippe Marquet, Ouassila Labbani-Narsis, Jean-Luc Dekeyser
DSD1