Amara Amara

dblp:43/3768 · DBLP profile ↗
← Back
23ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-9511-0899ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 4 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Fully-Parallel Digital MRAM Computing-in-Memory Macro Featuring a High-Efficient Dynamic Adder Tree and Bit-Splitting MAC
Zhongzhen Tong, Jiye Yao, Shaohui Ma, Yulong Qiu, Zhaohao Wang, Amara Amara, Xiaoyang Lin
ISCAS6
2025 ASP: A daptive Sparse LUT-based Bit-Slice Accelerator for Efficient CNN Inference
abstract
Recent advances in convolutional neural network (CNN) accelerators have leveraged data sparsity, significantly enhancing energy efficiency in resource- and energy-constrained hardware platforms. However, fully exploiting data sparsity and further enhancing energy efficiency require the deployment of specialized hardware accelerators. In this work, we present a bit-slice accelerator that integrates an LUT-based multiplier and a bit manipulation unit to achieve efficient resource utilization and low-latency multi-bit-width operations. Furthermore, the bit-slice processing element (PE) and dynamic PE array utilize adaptive subdivision for irregular matrices, optimizing data handling across varying levels of sparsity. Our synthesized RTL implementation demonstrates a significant improvement in hardware utilization, achieving a throughput of 66.23 GOPS/W on AlexNet and 61.13 GOPS/W on VGG16. This design achieves 2.05x higher energy efficiency and 2.75x lower latency compared to state-of-the-art designs, outperforming them on both the AlexNet and VGG16 benchmarks.
Yulong Qiu, Amara Amara
ISCAS3
2025 A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFET
abstract
Spintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT.
Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.10
2022 Real-Time and Cost-Effective Smart Mat System Based on Frequency Channel Selection for Sleep Posture Recognition in IoMT
abstract
Sleep posture, which affects the quality of sleep and could lead to medical conditions, such as pressure ulcers, is a key metric for sleep analysis in Internet of Medical Things (IoMT). In this article, a real-time and low-cost smart mat system for sleep posture recognition based on frequency channel selection is proposed. The system can recognize postures unobtrusively with a dense flexible sensor array. In addition, to enable real-time recognition with a relatively low-cost STM32 processor system, a lightweight algorithm that includes frequency channel selection, model pretraining, and real-time classification is proposed. Through a series of short-term and overnight experiments with 21 subjects, the feasibility and reliability of the proposed system were evaluated. Experimental results show that the accuracy of the short-term experiment is up to 95.43% and of the overnight experiment is up to 86.80% for four posture categories (supine, prone, right, and left) classification. The model size is just 56 kB which is much smaller than other methods. The runtime of the complete algorithm is about 6 ms with a low-power STM32 embedded system, which shows the system’s ability to provide real-time posture recognition. As an edge device, the proposed system could lead to the development of fast, convenient, and low-cost sleep posture recognition products for IoMT.
Haikang Diao, Chen Chen 0039, Wei Yuan 0005, Amara Amara, Toshiyo Tamura, Benny P. L. Lo, Long Meng, Sio-Hang Pun, Yuan-Ting Zhang, Wei Chen 0015
IEEE Internet Things J.5
2021 Unobtrusive Smart Mat System for Sleep Posture Recognition
abstract
Sleep posture, as a crucial index for sleep quality assessment and pressure ulcer prevention, has been widely studied for medical diagnoses and sleep disease treatment. In this paper, an unobtrusive smart mat system for sleep posture recognition is proposed, which is based on a dense flexible sensor array and printed electrodes and along with an algorithmic framework. With the dense flexible sensor array, the system offers a comfortable and high-resolution solution for long-term pressure sensing. Meanwhile, compared with other large-area and low-density mat systems, it reduces the area to minimize manufacturing cost and computational complexity, while also increases the density of the sensor to improve accuracy. To distinguish the sleep postures, the algorithmic framework that includes pre-processing, feature extraction, and posture classification is developed. Pilot studies in two scenarios including subject-dependent and subject- independent classification are performed with 7 persons for 4 different postures recognition. The experimental results show that the accuracy of the smart mat system can achieve over 78% using Support Vector Machines (SVMs) and k-Nearest Neighbor (kNN) for the subject-independent scenario. For the subject-dependent scenario, the accuracy can reach over 95%. It proves that the proposed method can recognize different sleep postures effectively.
Haikang Diao, Chen Chen 0039, Wei Chen 0015, Wei Yuan 0005, Amara Amara
ISCAS5
2019 ICTs as catalysts in child protection programmes: current landscape in South Asia & a concept to inform future use
abstract
Information and Communications Technologies (ICT) are increasingly, and increasingly effectively, being used in development and humanitarian work. Whereas health and education lead this use, application to child protection remains sparse and ill-understood. This paper helps address these two gaps. On the one hand, it enhances understanding of the use of ICT in child protection, by presenting the global and south Asia-specific landscapes, focusing on notable initiatives, partnerships and tools; on the other, it hopes to guide future use by putting forth a concept for an ICT-strengthened, child-centred system for case identification and management. The concept can be implemented using existing solutions of proven utility and efficacy. Therefore, we show that ICTs have a very high potential for use in child protection and can build on an increasingly solid evidence base and important successes.
Balwant Godara, Nihaalini Kumar, Frederique Boursin, Gatienne Jobit, Amara Amara, Thierry Agagliate
ICTD5
2017 Tunnel FET based refresh-free-DRAM
abstract
A refresh free and scalable ultimate DRAM (uDRAM) bitcell and architecture is proposed for embedded application. uDRAM 1T1C bitcell is designed using access Tunnel FETs. Proposed design is able to store the data statically during retention eliminating the need for refresh. This is achieved using negative differential resistance property of TFETs and storage capacitor leakage. uDRAM allows scaling of storage capacitor by 87% and 80% in comparison to DDR and eDRAMs, respectively. Bitcell area of 0.0275μm2is achieved in 28nm FDSOI-CMOS and is scalable further with technology shrink. Estimated throughput gain is 3.8% to 18% in comparison to CMOS DRAMs by refresh removal.
Navneet Gupta, Adam Makosiej, Andrei Vladimirescu, Amara Amara, Costin Anghel
DATE4
2016 3T-TFET bitcell based TFET-CMOS hybrid SRAM design for Ultra-Low Power applications
Navneet Gupta, Adam Makosiej, Andrei Vladimirescu, Amara Amara, Costin Anghel
DATE4
2016 Ultra-compact SRAM design using TFETs for low power low voltage applications
abstract
This paper presents a hybrid TFET/CMOS SRAM architecture designed to address the requirements for ULP (Ultra-Low Power) applications, like the IoT (Internet of Things). A novel 3-Transistor TFET SRAM cell is used for array while periphery is maintained in standard CMOS. The simulation extractions for power and speed are done including wiring and device parasitics extracted from 4Kb SRAM designed in 28nm FDSOI CMOS process. The proposed 3T-TFET SRAM cell supports aggressive voltage scaling without impacting data stability and allows application of performance boosting techniques without impacting cell leakage. The memory array leakage current is less than 1 fA/bit at sub-0.5V supply voltages, showing up-to 50x and 104x improvement compared with state-of-the-art TFET and CMOS SRAM bitcells, respectively. Bitcell area is reduced by 3x in comparison to existing TFET designs. Evaluated static noise margin (SNM) is 100mV for supply voltages range from 0.2V to 0.6V. Minimum read and write access pulse is evaluated at 15ns at 0.45V supply voltage.
Navneet Gupta, Adam Makosiej, Andrei Vladimirescu, Amara Amara, Costin Anghel
ISCAS4
2016 A cost-effective approach for ubiquitous broadband access based on hybrid PLC-VLC system
abstract
Visible light communication (VLC) using the light emitting diode (LED) will become an appealing alternative to the radio frequency communication technology for indoor wireless broadband access. However, VLC needs a ubiquitous network as its backbone to avoid becoming an information isolated island. Power line communication (PLC) systems could easily solve the informative problem of VLC while powering the LED lamps at the same time, which is considered as a good partner of VLC for the cost-effective implementation. In this paper, a novel and cost-effective framework of ubiquitous indoor broadband access based on deeply integrated VLC and PLC technology with only low-cost modification to the current infrastructure is therefore proposed. The broadband access network supports duplex transmission through each LED using the decode-and-forward (DF) working mode. This paper will present our recent research progress in this area, including a prototyping of duplex voice communications network based on hybrid PLC and VLC in our lab. Our research and development plan in this area for the near future will also be covered.
Jian Song 0004, Sicong Liu 0002, Guangxin Zhou, Bingyan Yu, Wenbo Ding 0001, Fang Yang 0001, Hongming Zhang 0010, Xun Zhang 0002, Amara Amara
ISCAS9
2015 Ultra-low leakage sub-32nm TFET/CMOS hybrid 32kb pseudo DualPort scratchpad with GHz speed for embedded applications
abstract
In this paper, an ultra-low-leakage TFET/CMOS hybrid Dual-Port SRAM (DPSRAM) based scratchpad memory is proposed. DPSRAM cells are designed using TFETs to reduce the leakage power in the memory array as compared to CMOS. Peripheral circuits are designed using 28nm FDSOI technology to increase speed and to reduce area as compared to full TFET based memories. Performance and stability of the memory is analyzed for different supply voltages to support dynamic voltage frequency scaling (DVFS). Imbalanced single-ended sensing is proposed in the paper and different write-assist techniques are analyzed for the proposed TFET memory cell. In the analysis of TFET DPSRAM bitcell at 1V supply voltage the evaluated noise margins are 114mV and 185 mV for read and write, respectively, with a 5 orders of magnitude reduction in leakage as compared to a similar CMOS bitcell. Results of performance evaluation of the designed 32Kb TFET/CMOS DPSRAM show a gain of up to 79.2% in write speed using write assist at sub-1V supply voltages and less than 1 ns read/write cycle time for more than 1V supply voltages.
Navneet Gupta, Adam Makosiej, Amara Amara, Andrei Vladimirescu, Costin Anghel
ISCAS4
2015 0.5-V sub-ns open-BL SRAM array with mid-point-sensing multi-power 5T cell
abstract
To achieve 0.5-V high-speed SRAMs, two proposals are demonstrated. One is a multi-power-supply five-transistor cell (5T cell), combined with a boosted word-line voltage and a mid-point sensing enabled by precharging bit-lines to VDD/2. The other is a partial activation of a multi-divided open-bit-line array without significant area penalty. Layout and post-layout simulation with a 28-nm fully-depleted planar-logic SOI MOSFET reveal that a 5T-cell 4-kb array in a 128-kb SRAM core is able to achieve x6 faster and x14 lower power than the counterpart 6T-cell array, suggesting a possibility of a 540-ps cycle time at 0.5 V.
Kiyoo Itoh 0002, Khaja Ahmad Shaik, Amara Amara
ISCAS3
2013 CMOS SRAM scaling limits under optimum stability constraints
abstract
This paper presents a predictive analysis of the high-density SRAM cell scaling from the stability and low power perspective. Based on a subthreshold SRAM analytical model [5] and a SRAM area-scaling model the Data Retention Voltage (DRV) defined as the lowest VDDthat can be applied during standby without losing data, as well as the minimum supply voltage for reliable read and write (VMIN), are investigated. The analysis is performed for several future technology nodes down to the 18 nm node. It takes into account the impact of MOS key parameters: threshold voltage (VT), subthreshold slope, DIBL, body factor and Pelgrom's Coefficient AVT. It is demonstrated, that due to process variations, the use of bulk CMOS for sub-28 nm becomes very challenging and severely limits area and supply scaling. Thin-film technology such as Ultra-Thin Body and BOX (UTBB) FDSOI however, should allow stable and power- and area-efficient SRAM design scaling below the 22 nm node with DRV lower than 0.4 V.
Adam Makosiej, Olivier Thomas, Amara Amara, Andrei Vladimirescu
ISCAS3
2012 Stability and yield-oriented ultra-low-power embedded 6T SRAM cell design optimization
abstract
This paper presents a methodology for the optimal design of CMOS 6T SRAM ultra-low-power (ULP) bitcells minimizing power consumption under strict stability constraints in all operating modes. An accurate analytical SRAM subthreshold model is developed for characterizing the cell behavior and optimizing its performance. The proposed design approach is demonstrated for an SRAM implemented in a 32nm CMOS UTBB-FDSOI technology. Stable operation in both read and write is obtained for the optimized cell at VDD=0.4V. Moreover, in the optimization process the standby and active power were reduced up to 10x and 3x, respectively.
Adam Makosiej, Olivier Thomas, Andrei Vladimirescu, Amara Amara
DATE4
2012 A 32nm tunnel FET SRAM for ultra low leakage
abstract
This paper describes the applicability of Tunnel FETs to commercial embedded Static Random-Access Memories (SRAM). Numerical device simulations were used first to optimize the performance of the TFET. The optimized TFETs show a steeper subthreshold slope than CMOS leading to a 5 orders of magnitude reduction in standby current. A look-up table model for circuit simulation of the TFET was developed based on characteristics obtained from TCAD simulations. A TFET SRAM cell is proposed and its stability is analyzed. Our novel 8T TFET SRAM cell operates at VDD=1V. The Read and Write Static Noise Margins are evaluated at 120mV and 200mV, with the operation speed of 300MHz and 1GHz in read and write respectively. The cell leakage is less than 10fA at VDD=1V. Our results show that TFETs are excellent candidates for embedded SRAMs due to their Ultra-Low Standby Power (LSTP).
Adam Makosiej, Rutwick Kumar Kashyap, Andrei Vladimirescu, Amara Amara, Costin Anghel
ISCAS4
2010 32nm and beyond Multi-VT Ultra-Thin Body and BOX FDSOI: From device to circuit
abstract
A low-cost and high-manufacturability Multi-VTUltra-Thin BOX and Body (UT2B) FDSOI technology is proposed for high-performance and low-leakage digital circuits. This concept allows setting up low, standard and high threshold voltage (VT) devices without degrading the good channel electrostatic control and the low VTdispersion of the FDSOI technology. Device electrical characteristics, process flow and physical design are described and the performance of digital circuits is evaluated.
Olivier Thomas, Jean-Philippe Noël, Claire Fenouillet-Béranger, Marie-Anne Jaud, J. Dura, P. Perreau, Frédéric Boeuf, François Andrieu, D. Delprat, F. Boedt, Konstantin Bourdelle, Bich-Yen Nguyen, Andrei Vladimirescu, Amara Amara
ISCAS14
2009 SRAM Voltage and Current Sense Amplifiers in sub-32nm Double-gate CMOS Insensitive to Process Variations and Transistor Mismatch
abstract
This paper presents a comparative study of two novel sub-32 nm current (CSA) and voltage (VSA) sense amplifiers in fully depleted (FD) double-gate (DG) silicon-on-insulator (SOI) technology with planar independent self-aligned gates. The proposed sense amplifiers (SA) need 40% to 4 times less power, achieve a 10-15% increase in speed and have a 2.5 to 5 times larger tolerance to Vthand L mismatch compared to published DG SAs. Both architectures take advantage of the back gate in order to improve circuit properties. The new CSA is 12% faster and reduces power consumption 3.3 times compared to the new VSA, with the latter having a significant advantage in size.
Piotr Nasalski, Adam Makosiej, Bastien Giraud, Andrei Vladimirescu, Amara Amara
ISCAS5
2008 A novel 4T asymmetric single-ended SRAM cell in sub-32 nm double gate technology
abstract
This paper presents a 4T asymmetric single-ended (ASE) SRAM cell in sub-32 nm CMOS fully depleted (FD) double-gate (DG) silicon-on-insulator (SOI) technology with planar self-aligned gates. Both independent- and connected- gates operation is analyzed either with symmetrical or asymmetrical transistors which have been adjusted according to the current and future process possibilities. The proposed cell is compared with the conventional 6T and an efficient 4T cell. A second version of the new cell is also proposed to improve the write operation. Both novel cells take advantage of the additional gate, offered by the DG technology, to improve stability and write criteria. The results of read-, retention- and write margins, power consumption, access time, write disturb and area are displayed for all cells.
Bastien Giraud, Amara Amara
ISCAS2
2007 A Comparative Study of 6T and 4T SRAM Cells in Double-Gate CMOS with Statistical Variation
abstract
This paper presents a comparative study of sub-32 nm CMOS 6T and 4T SRAM cells in fully depleted (FD) double-gate (DG) silicon-on-insulator (SOI) technology with planar independent self-aligned gates. Both independent- and connected-gate operation is analyzed by modulating the drain current with both front and back gate voltages. An improved 4T driver-less (DL) SRAM cell is proposed which takes advantage of the back gate to improve stability in read and retention mode by applying feedback between access transistor and storage node. The results of statistical characterization of read-, retention- and write margins, power and access time are presented for all cells in the presence of process variability.
Bastien Giraud, Amara Amara, Andrei Vladimirescu
ISCAS2
2005 Iris identification and robustness evaluation of a wavelet packets based algorithm
abstract
This paper presents an iris recognition system, based on a wavelet packet analysis using orthogonal wavelets. The identification of the different packets that carry discriminating information about the iris texture is carried out through an energy measure. Tests, conducted on a database of 149 high quality iris images show good robustness in relation to changes in illumination, blurring, optical axis deviation or local defects in the images.
Florence Rossant, Frédéric Amiel, Thomas Ea, Amara Amara, Manuel Torres Eslava
ICIP (3)4
2004 IRIS features extraction using wavelet packets
abstract
This article presents an application of wavelet packet analysis to the features extraction part of an iris recognition system. An energy measure is used to identify the particular packets that carries discriminating information about the iris texture. Several different orthogonal wavelets arc tested and a comparison to nonorthogonal analysis using Gabor wavelets is done. The experimental results show 100% correct classifications when applying the algorithm on an iris image database and the new algorithm is therefore an interesting alternative to Gabor based methods.
Erik Rydgren, Thomas Ea, Frédéric Amiel, Florence Rossant, Amara Amara
ICIP5
2004 Modeling subthreshold SOI logic for static timing analysis
abstract
A simple, yet realistic physics-based model is introduced to describe the subthreshold drain current of a MOSFET taking into account the body- and drain-voltage dependencies, including the short channel effects. This model, verified by SPICE simulations, describes adequately the pseudotriode and pseudosaturation regions of MOS transistors operated below V/sub T/. It can be applied for predicting bulk- or partially depleted (PD) SOI CMOS circuit operation. Analytical expressions derived for the logic switching threshold and delay are applied to predict the performance of CMOS-SOI inverters.
Alexandre Valentian, Olivier Thomas, Andrei Vladimirescu, Amara Amara
IEEE Trans. Very Large Scale Integr. Syst.4
1996 A 1.0ns 64-bits GaAs Adder using Quad tree algorithm
abstract
This paper describes a full custom 64-bits adder targeting the VITESSE E/D MESFET process HGaAsIII. This adder which respects a bit slice topology is part of the project of GaAs data-path compiler for ALLIANCE CAD TOOLs. GaAs's specific properties have been exploited in a full custom approach. Original architecture have been used to increase the parallelism of carries' computation. The layout is portable using a symbolic approach and could also be used with other E/D MESFET process.
Philippe Royannez, Amara Amara
Great Lakes Symposium on VLSI2