VLDB 2026 Research / reviewers in the wild / expert
Ibrahim M. Elfadel
dblp:29/2561 · also Ibrahim Abe M. Elfadel
· DBLP profile ↗
77ranked-venue papers
16as first author
22since 2021 · last 2025
0000-0003-3220-9987ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 10 first-author · 16 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorComputer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Topological and Temporal Feature Extraction for EEG-Based Seizure Detection
Andreas Henschel, Ibrahim M. Elfadel |
HealthCom | 3 |
| 2025 | A BCI Framework Using Multifractal Detrended Fluctuation Analysis for EEG Feature Extraction
Waleed Bin Owais, Anant Sharma, Herbert F. Jelinek, Ibrahim M. Elfadel |
HealthCom | 4 |
| 2025 | PMIL: A Topology Module to Improve MIL-based WSI ClassificationabstractDeep learning models have achieved remarkable success in pathology image analysis. However, they still face challenges in effectively modeling fine-grained, object-level features. Topological Data Analysis (TDA) has shown promise for addressing these issues but remains underexplored, particularly for whole-slide pathology applications. Additionally, the effectiveness of TDA has yet to be firmly established, as current studies largely use small-scale datasets. In this work, we address these gaps by introducing Persistent Homology in Multiple Instance Learning (PMIL), the first adaptable TDA-based module within the MIL framework. We validate our approach on a large-scale classification dataset, benchmarking against multiple state-of-the-art methods. Ahmad Obeid 0001, Anabia Sohail, Said Boumaraf, Xiabi Liu, Sajid Javed, Hasan Almarzouqi, Jorge Dias 0001, Mohammed Bennamoun, Naoufel Werghi, Ibrahim M. Elfadel |
ISCAS | 10 |
| 2024 | Order-Preserving Cryptography for the Confidential Inference in Random Forests: FPGA Design and ImplementationabstractPrior work has addressed the problem of confidential inference in decision trees. Both traditional order-preserving cryptography (OPE) and order-preserving NTRU cryptography have been used to ensure data and model privacy in decision trees. Furthermore, FPGA architectures and implementations have been proposed for implementing such confidential inference algorithms on resource-limited, edge-based platforms such as low-cost FPGA boards. In this paper, we address the challenging problem of scalability of order-preserving confidential inference to random forests, which are ensembles of decision trees that are meant to improve their classification accuracy and reduce their overfitting. The paper develops a methodology and an FPGA implementation strategy for scaling up OPE to random forests. In particular, a framework is used to study the multifaceted tradeoffs that exist between the number of trees in the random forest, the strength of the encryption, the accuracy of the inferences, and the resources of the edge platform. Extensive experiments are conducted using the MNIST dataset and the Intel DE10 Standard FPGA board. Rupesh Raj Karn, Kashif Nawaz, Ibrahim M. Elfadel |
DAC | 3 |
| 2024 | Live Demonstration: W3M Wearable Weight and Walk Monitoring SystemabstractThe "Wearable Weight and Walk Monitoring System (W3M)" is a patented wearable technology, designed to monitor sudden changes in body weight. Such sudden changes occur in a variety of medical conditions such as congestive heart failure and Crohne's disease, and may indicate an imminent need for urgent medical care. Asim Arif, Adedayo Adegbile, Qiraat Khan, Hamda Memon, Ibrahim M. Elfadel |
ISCAS | 5 |
| 2024 | On Various Extensions of the Shannon-Hagelbarger Concavity TheoremabstractThe Shannon-Hagelbarger concavity theorem asserts that the one-port driving-point resistance of a network of linear two-terminal resistors is a concave function of its branch resistances. This theorem has been applied in many domains, including circuit optimization, graph algorithms, and network analysis. In this paper, we show that this concavity theorem can be extended to any linear networks consisting solely of two-terminal inductors or two-terminal capacitors. The main contribution of this paper is a unified proof methodology based on various forms of Tellegen’s theorem. The methodology provides circuit insight into the concavity results and avoids the use of additional circuit elements as was the case in the original proof of Shannon and Hagelbarger. Ibrahim M. Elfadel |
ISCAS | 1 |
| 2024 | OSHDA: A Containerized CAD Tool for the Design and Analysis of Behavioral FSM Logic LockingabstractThis paper introduces the Open-source Secure Hardware Design and Analysis (OSHDA) toolchain for the logic locking of finite-state machines (FSMs) at the behavioral level. OSHDA's FSM obfuscation method is based on the recently developed State Permutation Logic Locking (SPeLL) algorithm which obfuscates the behavioral transition graph of the FSM, thus avoiding the use of dummy states and reducing exposure to reverse engineering attacks. In addition to implementing the SPeLL algorithm, the toolchain implements a full logic synthesis flow, including the evaluation of the gate-level SPeLL hardware overhead for both FPGA and ASIC designs. In particular, OSHDA enables the automation of trade-off analysis between the strength of SPeLL security and its hardware overhead. The paper further describes attempted attacks on SPeLL using state-of-the-art de-obfuscation tools and identifies research gaps in behavioral de-obfuscation that must be addressed before one can successfully de-obfuscate SPeLL. OSHDA comes with its own scripting subsystem for augmenting its analysis, adding de-obfuscation methods, and integrating physical design tools. Finally, OSHDA is deployed as a hardware security microservice using the Docker framework. Esrat Khan, Shahzad Muzaffar, Lamees M. Al Qassem, Ibrahim M. Elfadel |
VLSI-SoC | 4 |
| 2024 | Containerized Microservices: A Survey of Resource Management FrameworksabstractThe growing adoption of microservice architectures (MSAs) has led to major research and development efforts to address their challenges and improve their performance, reliability, and robustness. Important aspects of MSA that are not sufficiently covered in the open literature include efficient cloud resource allocation and optimal power management. Other aspects of MSA remain widely scattered in the literature, including cost analysis, service level agreements (SLAs), and demand-driven scaling. In this article, we examine recent cloud frameworks for containerized microservices with a focus on efficient resource utilization using auto-scaling. We classify these frameworks on the basis of their resource allocation models and underlying hardware resources. We highlight current MSA trends and identify workload-driven resource sharing within microservice meshes and SLA streamlining as two key areas for future microservice research. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | Secure Edge-Coded Signaling IoT Transceiver With Reduced Encryption OverheadabstractThe edge-coded signaling (ECS) protocol enables single-wire signaling in IoT devices and sensors using two important neuromorphic attributes. The first is the coding of bits as a stream of pulses (spikes), and the second is the circumvention of clock and data recovery (CDR) at the receiver. In addition, ECS can be endowed with strong, yet lightweight, security features using an ultralow-latency version of the A5/1 stream cipher. Such strong security comes at the expense of decreased data rates and significant area overhead. In this article, we introduce a new generation of secure ECS protocols that incorporates two notable improvements. The first is a more compact pulse stream definition that results in improved data rates for the plain ECS protocol. The second is a coding-aware version of the low-latency A5/1 stream cipher that results in minimal impact on the effective data rate of the transmission. Consequently, a new all-digital and secure ECS transceiver design is proposed, prototyped, and functionally verified in 65-nm technology. Compared with previous generations of secure ECS transceivers, this new design achieves an increase of approximately 138%, 199%, and 640% in minimum, average, and maximum data rates, respectively, and results in increased resiliency against brute-force attacks by a factor of 16. Furthermore, the ASIC implementation shows that it maintains the compact and energy-efficient features of the ECS architecture, using only$28~\mu $W with an average energy efficiency of 2.745 pJ/bit and a gate count of approximately 2880 gates. This is more than 40% decrease in the equivalent gate count relative to the previous secure ECS generation. Mizan Abraha Gebremichael, Ibrahim M. Elfadel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | OSμS: An Open-Source Microservice Prototyping PlatformabstractOne major advantage of microservice cloud architectures is the agility with which microservices can be replicated to help improve the overall quality of service and meet service-level contracts. Their challenge is to carefully balance the horizontal microservice replicas with the vertical resources of CPU, memory, and IO that are allocated to each microservice. The objective of such balancing act is, of course, to avoid both service bottlenecks and resource wastage. In this paper, we present OSμS, a new open-source microservice prototyping platform that has been developed and instrumented from the ground up with the objective of collecting fine-grained, non-proprietary metrology on microservice mesh performance. We will illustrate the use of OSμS for developing and evaluating machine-learning algorithms for the horizontal and vertical autoscaling of microservice architectures. A hybrid algorithm based on decision-tree learning will be implemented on OSμS and compared with the academic state of the art and existing cloud-provider solutions. The advantages of such algorithm in improving horizontal and vertical resource utilization will be highlighted. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
CloudCom | 4 |
| 2023 | Network Theorems for Fractional-Order CircuitsabstractIn this paper, we provide a circuit-theoretic analysis of linear fractional-order circuits based on the use of Tellegen's theorem. The advantages of the circuit-theoretic analysis over the more common control-theoretic one are illustrated with general theorems on the driving-point impedances and admittances of multi-type, fractional-order circuits made of an arbitrary, finite, connected mesh of fractional-order elements. The explicit use of Kirchhoff's circuit laws leads to sharp results on the stability and resonant behavior of such networks. In particular, it results in more intuitive conditions on the range of fractional orders needed to guarantee network stability or resonant behavior. Ibrahim M. Elfadel |
ISCAS | 1 |
| 2023 | Post-Quantum, Order-Preserving Encryption for the Confidential Inference in Decision Trees: FPGA Design and ImplementationabstractOne main objective of this paper is to show how to adapt the well-known, lattice-based NTRU post-quantum encryption to the confidential inference in decision trees. Another objective is to describe a resource-efficient FPGA implementation of the adapted NTRU. The typical use case of such encryption is that of two parties where one party has proprietary ownership of the decision tree model while the other party has proprietary ownership of the data. Confidential inference in decision trees can be insured using order-preserving cryptography, which has much weaker requirements and is therefore easier to implement than fully homomorphic cryptography. Post-quantum NTRU is not order-preserving, but interestingly, it can be modified to obey the order-preserving property. We call the resulting cipher OP-NTRU. Lossless compression can be applied to the ciphertext produced by OP-NTRU to facilitate its hardware acceleration. OP-NTRU has been implemented on an FPGA with the HDL code automatically compiled from the machine learning framework. Confidential inference experiments report more than 96% compression without degrading inference accuracy in FPGA for the MNIST dataset. Rupesh Raj Karn, Kashif Nawaz, Ibrahim M. Elfadel |
VLSI-SoC | 3 |
| 2022 | On a Generalization of Tellegen's Theorem to Quantum CircuitsabstractTellegen’s theorem is one of the foundational topological theorems of circuit theory. In its most elementary form, it expresses an energy balance between the power supplied at the ports of the electrical circuit network and the power consumed in its internal branches. But the topological roots of Tellegen’s theorem allow it to take other forms that make it into a powerful tool for deriving several fundamental results in electrical circuit networks, linear or non-linear, reciprocal or nonreciprocal. Tellegen’s theorem admits a generalization to the field case of Maxwell’s equations that makes it useful in the analysis and modeling of microwave devices. In this paper, the field case of Shrödinger’s equation is examined, and a generalization of Tellegen’s theorem to the case of quantum wave functions is proposed. This generalization has the potential of extending the application domain of Tellegen’s theorem to the circuits used in quantum computing. Ibrahim M. Elfadel |
ISCAS | 1 |
| 2022 | Hyper-parameter Tuning for Progressive Learning and its Application to Network Cyber SecurityabstractThe long-term deployment of data-driven AI technology using artificial neural networks (ANNs) should be scalable and maintainable when new data becomes available. To insure smooth adaptation, the learning must be cumulative so that the network consumes new data without compromising its inference performance based on past data. Such incremental accumulation of learning experience is known as progressive learning. In this paper, we address the open problem of tuning the hyperparameters of neural networks during progressive learning. A hyper-parameter optimization framework is proposed that selects the best hyper-parameter values on a task-by-task basis. The neural network model adapts to each progressive learning task by adjusting the hyper-parameters under which the neural architecture is incrementally grown. Several hyper-parameter search strategies are explored and compared in support of progressive learning. In contrast to the predominant practice of using imaging datasets in machine learning, we have used cybersecurity datasets to illustrate the advantages of the proposed hyper-parameter tuning algorithms. Rupesh Raj Karn, Matthew M. Ziegler, Jinwook Jung, Ibrahim M. Elfadel |
ISCAS | 4 |
| 2022 | Confidential Inference in Decision Trees: FPGA Design and ImplementationabstractIn confidential computing, algorithms operate on encrypted inputs to produce encrypted outputs. Specifically, in confidential inference, Alice has the parameters of the machine-learning model but does not want to reveal them to Bob who has the data. Bob wants to use Alice’s model for inference but does not want to reveal his data. Alice and Bob agree to use homomorphic encryption for running the inference engine in full confidence without revealing either model or data. They find that full homomorphic encryption is very time consuming and very challenging to accelerate on hardware. In this particular case, homomorphic encryption can be made computationally efficient and can even be readily accelerated on hardware. In this paper, we reveal how Alice and Bob run the inference engine in full confidence and show an FPGA implementation of the specialized homomorphic computing algorithm they used. We further evaluate the resources needed to implement the encrypted decision tree and compare them with those of a plain decision tree. Confidential inference tests are run on the encrypted FPGA design using the MNIST dataset. Rupesh Raj Karn, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2022 | Logic Locking of Finite-State Machines Using Transition ObfuscationabstractIn this paper, we introduce a novel algorithm for securing sequential circuits at the Register-Transfer Level (RTL) that does not require any state augmentation. The algorithm is based on the encryption of the state encodings with a key that is known only to the IP provider. When the correct key is input at runtime, the sequential circuit will operate as designed, otherwise it will operate according to a state transition map that is defined by the wrong key. We call this mode of operation: transition obfuscation. One important advantage of the proposed method is that using the wrong key does not necessarily result in the sequential circuit getting stuck at any one state or getting trapped within any black hole. As a result, the secured sequential circuit is more immune to reverse engineering attacks, and because of the large number of wrong full-state transition maps, more immune to side-channel attacks. A full low-complexity, RTL design methodology based on the new algorithm is presented along with extensive experiments quantifying its design overhead and illustrating its advantages in terms of immunity to reverse engineering attacks. Shahzad Muzaffar, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2022 | Unsupervised Land-Cover Segmentation Using Accelerated Balanced Deep Embedded ClusteringabstractIn this letter, we address the issue of the automatic labeling of remote sensing datasets using a novel deep learning clustering algorithm. The proposed algorithm addresses the inherent susceptibility of the deep embedded clustering (DEC) algorithm to data imbalance using additional search and extraction steps. Furthermore, the proposed algorithm is highly parallelizable. A graphics processing unit (GPU) implementation is shown to achieve 40X to 2600X of performance speedup and improved clustering accuracy with respect to DEC and other clustering approaches. Ahmad Obeid 0001, Ibrahim M. Elfadel, Naoufel Werghi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | FPGAaaS: A Survey of Infrastructures and SystemsabstractThe popularity of cloud computing services for delivering and accessing infrastructure on demand has significantly increased over the last few years. Concurrently, the usage of FPGAs to accelerate compute-intensive applications has become more widespread in different computational domains due to their ability to achieve high throughput and predictable latency while providing programmability and improved energy efficiency. Computationally intensive applications such as big data analytics, machine learning, and video processing have been accelerated by FPGAs. With the exponential workload increase in data centers, major cloud service providers have made FPGAs and their capabilities available as cloud services. However, enabling FPGAs in the cloud is not a trivial task due to incompatibilities with existing cloud infrastructure and operational challenges related to abstraction, virtualization, partitioning, and security. In this article, we survey recent frameworks for offering FPGA hardware acceleration as a cloud service, classify them based on their virtualization mode, tenancy model, communication interface, software stack, and hardware infrastructure. We further highlight current FPGAaaS trends and identify FPGA resource sharing, security, and microservicing as important areas for future research. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Convergent Time-Stepping Schemes for Analog ReLU NetworksabstractWith the phenomenal growth of deep learning paradigms based on the use of the rectified linear unit (ReLU) as activation function and the importance attached to the hardware acceleration of such learning approaches, there is a pressing need for the development of numerical simulation algorithms that are tailored for the specific context of analog ReLU networks. In this paper, we propose two time-stepping schemes for the transient analysis of analog ReLU networks and provide rigorous proofs of their convergence under mild conditions on the ReLU network connectivity matrix. Simulation examples are provided that illustrate the numerical stability of these schemes and contrast their convergence rates. Ibrahim M. Elfadel |
ISCAS | 1 |
| 2021 | Beyond Arduino: A Guide for the PerplexedabstractArduino IDE and boards have been used in training, hobby projects, and lab experiments because of their ease of use, flexibility, and readily available libraries. However, they cannot be used for commercial products and customized systems that target low-cost and low-power operation, are resource-constrained and need to comply with industrial and regulatory standards. This paper provides a methodology and guidelines on how to migrate a lab prototype developed using the Arduino environment and to a near-product prototype using an industry- strength IDE and MDK development environment. The guidelines involve a sequence of steps whose main goal is to minimize the hardware and software debugging effort. Each step is focused on bringing one aspect under active development while keeping everything else fixed and bug-free. Moreover, at each step, a working reference is set up to enable cross-checks in case of any unexpected problem. The guidelines are illustrated with a case study from our own work in which an Arduino prototype has been successfully transformed into a wearable with an extremely small footprint. In addition to their value for embedded system design, these guidelines have the educational value of contrasting the learning outcomes of embedded system courses based on the Arduino framework vs. those based on an industry-driven MDK and IDE. Shahzad Muzaffar, Ibrahim M. Elfadel |
ISCAS | 2 |
| 2021 | On the Stability of Analog ReLU NetworksabstractRectified linear unit (ReLU) networks have become widely used in machine learning and automated inference using neural networks. Various forms of hardware accelerators based on ReLU networks have also been under development. In this brief, the stability problem in analog ReLU networks is addressed. Using the Lyapunov stability theory, it is shown that the origin of an unforced, analog ReLU dynamical system is globally asymptotically stable if the induced Euclidean norm of its connectivity matrix is less than one. An example is given to demonstrate that this upper bound is the best that can be achieved. In particular, the stability result holds for the case of a nonsymmetric connectivity matrix as may occur in some mathematical models of neurobiology. Ibrahim M. Elfadel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Cryptomining Detection in Container Clouds Using System Calls and Explainable Machine LearningabstractThe use of containers in cloud computing has been steadily increasing. With the emergence of Kubernetes, the management of applications inside containers (or pods) is simplified. Kubernetes allows automated actions like self-healing, scaling, rolling back, and updates for the application management. At the same time, security threats have also evolved with attacks on pods to perform malicious actions. Out of several recent malware types, cryptomining has emerged as one of the most serious threats with its hijacking of server resources for cryptocurrency mining. During application deployment and execution in the pod, a cryptomining process, started by a hidden malware executable can be run in the background, and a method to detect malicious cryptomining software running inside Kubernetes pods is needed. One feasible strategy is to use machine learning (ML) to identify and classify pods based on whether or not they contain a running process of cryptomining. In addition to such detection, the system administrator will need an explanation as to the reason(s) of the ML's classification outcome. The explanation will justify and support disruptive administrative decisions such as pod removal or its restart with a new image. In this article, we describe the design and implementation of an ML-based detection system of anomalous pods in a Kubernetes cluster by monitoring Linux-kernel system calls (syscalls). Several types of cryptominers images are used as containers within an anomalous pod, and several ML models are built to detect such pods in the presence of numerous healthy cloud workloads. Explainability is provided using SHAP, LIME, and a novel auto-encoding-based scheme for LSTM models. Seven evaluation metrics are used to compare and contrast the explainable models of the proposed ML cryptomining detection engine. Rupesh Raj Karn, Prabhakar Kudva, Hai Huang 0002, Sahil Suneja, Ibrahim M. Elfadel |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Lessons Learned the Hard Wayabstract“Fail often to succeed sooner” is a common mantra that we are told is the secret to success. When reporting research results, however, scholars rarely write about their failed attempts and only focus on the successful ones. Perhaps the source of this disconnect between what we preach and what we do can be found in the underlying assumption that published work is meant to move the field forward and failed attempts supposedly do not. The goal of the confessions presented in this paper is to show that even failed attempts are genuine and valuable contributions to our field provided that we learn from our mistakes and correct them. The 27 confessions span from planning oversights, digital and analog design errors, misunderstanding of devices, overlooked parasitics, LVS errors, and troubles in testing. Tobi Delbruck, Ibrahim M. Elfadel, Shahzad Muzaffar, Germain Haessig, Bo Wang 0012, Amine Bermak, Rui Graca, Luis A. Camuñas-Mesa, Bathiya Senevirathna, Pamela Abshire, Bernabé Linares-Barranco, Saeed Afshar, Shih-Chii Liu, Runchun Wang, Piotr Dudek, Stephen J. Carey, José M. de la Rosa 0001, Marc Dandin, Sheung Lu, Vincent Frick, Teresa Serrano-Gotarredona, Paula López Martinez 0001, Melika Payvand, Advait Madhavan, Eric R. Fossum, Juan Camilo Vasquez Tieck, Yan Liu 0016, Timothy G. Constandinou, Alexander Serb, Ricardo Carmona-Galán, Robert Nawrocki, Walter D. Leon-Salas |
ISCAS | 2 |
| 2020 | An Inference Hardware Accelerator for EEG-Based Emotion DetectionabstractThe wearability of emotion classifiers is a must if they are to significantly improve the social integration of patients suffering from neurological disorders. Such wearability requires the use of low-power hardware accelerators that would enable near real-time classification and extended periods of operations. In this paper, we architect, design, implement, and test a handcrafted, hardware Convolutional Neural Network, named BioCNN, optimized for EEG-based emotion detection and other similar bio-medical applications. The architecture of BioCNN is based on aggressive pipelining and hardware parallelism that maximizes resource re-use and minimizes memory footprint. The FEXD and DEAP datasets are used to test the BioCNN prototype that is implemented using the Digilent Atlys Board with a low-cost Spartan-6 FPGA. The experimental results show that BioCNN has a competitive energy efficiency of 11GOps/W, a throughput of 1.65GOps that is in line with the real-time specification of a wearable device, and a latency of less than 1ms, which is much smaller than the 150ms required for human interaction times. Its emotion inference accuracy is competitive with the top software-based emotion detectors. Hector A. Gonzalez, Shahzad Muzaffar, Jerald Yoo, Ibrahim M. Elfadel |
ISCAS | 4 |
| 2020 | Dynamically Generated Compact Neural Networks for Task Progressive LearningabstractTask progressive learning is often required where the training data become available in batches over the time. Such learning has the characteristic of using an existing model trained over a set of tasks to learn a new task while maintaining the accuracy of older tasks. Artificial Neural Networks (ANNs) have a higher capacity for progressive learning than other traditional machine learning models due to the availability of a large number of ANN parameters. A progressive model that uses a fully connected ANN suffers from long training time, overfitting, and excessive resource usage. It is therefore necessary to generate the ANN incrementally as new tasks arrive and new training is needed. In this paper, an incremental algorithm is presented to dynamically generate a compact neural network by pruning and expanding the synaptic weights based on the learning requirements of the new tasks. The algorithm is implemented, analyzed, and validated using the cloud network security datasets, UNSW and AWID, as well as the image dataset, MNIST. Rupesh Raj Karn, Prabhakar Kudva, Ibrahim M. Elfadel |
ISCAS | 3 |
| 2020 | A Remote FPGA Laboratory as a Cloud MicroserviceabstractIn this paper, we propose a web-based microservice for programming FPGAs remotely in the cloud without requiring additional hardware or third party software. Docker and Docker-compose are used to implement a scalable, secure environment for FPGA design and operation. The paper further describes in detail the architectural design, user interface, tool flow, and hardware back-end of the FPGA cloud service. The Docker files are made publicly available at GitHub. Lamees M. Al Qassem, Thanos Stouraitis, Ernesto Damiani, Ibrahim M. Elfadel |
ISCAS | 4 |
| 2020 | Dynamic Edge-coded Protocols for Low-power, Device-to-device CommunicationabstractClock and Data Recovery (CDR) has been a foundational receiver component in serial communications. Yet this component is known to add significant design complexity to the receiver and to consume significant resources in area and power. In the resource-limited world of constrained IoT nodes, the need of including CDR in the communication link is being re-assessed and new techniques for achieving reliable serial transmission without CDR have been emerging. These new techniques are distinguished by their use of transition edges rather than bit times for coding and detection. This article presents the design, implementation, and testing of a novel CDR-less transmission protocol that achieves significant improvements in data rate, reliability, packet security, and power efficiency with respect to state-of-the-art CDR-less techniques. The new protocol further tolerates significant jitters and clock discrepancies between transmitter and receiver. An FPGA and an ASIC (65 nm technology) implementation of the protocol have shown it to consume around 19μ W of power at a clock rate of 25 MHz, and to have a small footprint with a gate count of approximately 2,098 gates. In particular, the new protocol reduces area by more than 87% and power by more than 78% in comparison with CDR-based serial bit transfer protocols. Furthermore, the new protocol is shown to be versatile in its applications to available communication media, including wired, wireless, infrared, and human-body channels, under a variety of digital modulation schemes. Shahzad Muzaffar, Ibrahim M. Elfadel |
ACM Trans. Sens. Networks | 2 |
| 2019 | Double Data Rate Dynamic Edge-Coded Signaling for Low-Power IoT CommunicationabstractDynamic Edge-Coded Signaling (ECS) is a recently introduced protocol for single-channel signaling between constrained IoT nodes. One of the most distinguishing features of ECS is that its receiver does not require any circuitry for clockand-data recovery. Other important ECS features include its tolerance with respect to clock variations between transmitter and receiver and its amenability to seamlessly integrate lightweight cryptographic algorithms to ensure secure communication. ECS encodes information using pulse counts with the counting based on one of the pulse edges. In this paper, we address the problem of improving the ECS data rate for a given clock frequency and under a given power envelop by using both pulse edges of the ECS pulse stream. We call the novel protocol double-data-rate ECS (DDR-ECS) in analogy with DDR memory systems. While the concept is intuitive and attractive, its hardware implementation is not. This paper, therefore, presents an efficient hardware design of the DDR-ECS transceiver that preserves the ECS built-in features while essentially doubling the data rate at the same clock frequency and within the same power budget. A 65nm ASIC synthesis of the transceiver shows that DDR-ECS consumes the ECS equivalent power of 19μW, uses a small form factor of only 1934 gates, and doubles the dynamic data rate to the 7.8-44.4 Mb/s range with an average of 12 Mb/s at a clock rate of 25MHz. Shahzad Muzaffar, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2019 | Domain-Specific Architecture for IMU Array Data FusionabstractTo achieve high accuracy at low cost in a navigational system, an array of several low-cost MEMS Inertial Measurement Units (IMU's) may be used rather than one single high-performance but high-cost and power hungry mechanical IMU. To combine and predict the outputs and internal states of the IMU array, signal processing algorithms, such as the Kalman Filter (KF), are used with their prediction accuracy increasing with the number of array elements. While large IMU arrays are beneficial for accurate and precise estimation of linear and angular accelerations, they are detrimental to the KF computations since the underlying matrix dimensions of each KF variable increase drastically with array size. This paper discusses a domain-specific processor architecture implemented on an Artix-7 FPGA that can efficiently support the KF matrix operations and improve the throughput of the KF data fusion component. The processor instruction set, design and firmware are discussed in detail. The processor performance and hardware resource utilization are also fully quantified. To constrain the resource and power requirements of the processor-to-array interface, the kinematic model of the IMU accelerometer is used to devise a model-based approximation technique that reduces the number of sensor interface units to just one. This is then implemented using the time multiplexing of the data from various array sensors. Experimental results show that the RMSE of the estimated linear acceleration remains below 0.3m/s2for a sensor noise standard deviation of less than 0.04m/s2. The proposed combination of the model-based approximation with the domain specific processor results in a compact data fusion processing system with minimal footprint that vastly outperforms a general purpose processor. Owais Talaat Waheed, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2019 | Dynamic Autoselection and Autotuning of Machine Learning Models for Cloud Network AnalyticsabstractCloud network monitoring data is dynamic and distributed. Signals to monitor the cloud can appear, disappear or change their importance and clarity over time. Machine learning (ML) models tuned to a given data set can therefore quickly become inadequate. A model might be highly accurate at one point in time but may lose its accuracy at a later time due to changes in input data and their features. Distributed learning with dynamic model selection is therefore often required. Under such selection, poorly performing models (although aggressively tuned for the prior data) are retired or put on standby while new or standby models are brought in. The well-known method of Ensemble ML (EML) may potentially be applied to improve the overall accuracy of a family of ML models. Unfortunately, EML has several disadvantages, including the need for continuous training, excessive computational resources, requirement for large training datasets, high risks of overfitting, and a time-consuming model-building process. In this paper, we propose a novel cloud methodology for automatic ML model selection and tuning that automates model building and selection and is competitive with existing methods. We use unsupervised learning to better explore the data space before the generation of targeted supervised learning models in an automated fashion. In particular, we create a Cloud DevOps architecture for autotuning and selection based on container orchestration and messaging between containers, and take advantage of a new autoscaling method to dynamically create and evaluate instantiations of ML algorithms. The proposed methodology and tool are demonstrated on cloud network security datasets. Rupesh Raj Karn, Prabhakar Kudva, Ibrahim M. Elfadel |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 14 |
| 2019 | A Domain-Specific Processor Microarchitecture for Energy-Efficient, Dynamic IoT CommunicationabstractIn this paper, we present a domain-specific processor architecture, named pulsed-index communication interface architecture (PICIA), for single-channel IoT communication based on the recently introduced pulsed-signaling protocols, according to which information is encoded as series of pulses representing ON bits. In addition to the traditional aspects of instruction set architecture (ISA) design such as addressing modes, instruction types, instruction formats, registers, interrupts, and external I/O, the ISA includes domain-specific instructions that facilitate bit stream encoding and decoding based on the pulsed-signaling techniques. The domain-specific PICIA microarchitecture employs a set of optimized processing blocks that can be used programmatically to encode and decode the transmitted data in the most economical way. The PICIA allows customizations that support both standard pulsed-signaling techniques and specialized protocols that belong to the same family. The PICIA design further allows an amalgamation of software and hardware that significantly reduces the number of instructions required to implement a given communication interface without impacting the data rates and reliability of the pulsed-signaling protocols. The PICIA processor has been implemented in Verilog HDL and tested using a Xilinx Spartan-6 field-programmable gate array (FPGA). Furthermore, a 65-nm application-specific integrated circuits (ASIC) synthesis of the design confirms the small-footprint and low-power features of PICIA. The consumed power has been evaluated at 31.14μW with an energy efficiency of less than 10 pJ/bit. Shahzad Muzaffar, Ibrahim M. Elfadel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | An Instruction Set Architecture for Low-power, Dynamic IoT CommunicationabstractThis paper presents an instruction set architecture (ISA) dedicated to the rapid and efficient implementation of single-channel IoT communication interfaces. The architecture is meant to provide a programming interface for the implementation of signaling protocols based on the recently introduced pulsed-index schemes. In addition to the traditional aspects of ISA design such as addressing modes, instruction types, instruction formats, registers, interrupts, and external I/O, the ISA includes special-purpose instructions that facilitate bit stream encoding and decoding based on the pulsed-index techniques. Verilog HDL is used to synthesize a fully functional processor based on this ISA and provide both an FPGA implementation and a synthesised ASIC design in GLOBALFOUNDRIES 65nm. The ASIC design confirms the low-power features of this ISA with consumed power around 31µW and energy efficiency of less than 10pJ/bit. Shahzad Muzaffar, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2018 | Multi-Objective 3D Floorplanning with Integrated Voltage AssignmentabstractVoltage assignment is a well-known technique for circuit design, which has been applied successfully to reduce power consumption in classical 2D integrated circuits (ICs). Its usage in the context of 3D ICs has not been fully explored yet although reducing power in 3D designs is of crucial importance, for example, to tackle the ever-present challenge of thermal management. In this article, we investigate the effective and efficient partitioning of 3D designs into multiple voltage domains during the floorplanning step of physical design. In particular, we introduce, implement, and evaluate novel algorithms for effective integration of voltage assignment into the inner floorplanning loops. Our algorithms are compatible not only with the traditional objectives of 2D floorplanning but also with the additional objectives and constraints of 3D designs, including the planning of through-silicon vias (TSVs) and the thermal management of stacked dies. We test our 3D floorplanner extensively on the GSRC benchmarks as well as on an augmented version of the IBM-HB+ benchmarks. The 3D floorplans are shown to achieve effective trade-offs for power and delays throughout different configurations—our results surpass naïve low-power and high-performance voltage assignment by 17% and 10%, on average. Finally, we release our 3D floorplanning framework as open-source code. Johann Knechtel, Jens Lienig, Ibrahim M. Elfadel |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | A pulsed decimal technique for single-channel, dynamic signaling for IoT applicationsabstractPulsed-Index Communication (PIC) is a recent technique for single-channel communication which is based on the principle of transferring the indices of only the ON bits in the form of a series of pulse streams. In this paper, we present a modified version of PIC which is based on the same underlying idea but with key improvements in data rate and reliability. The proposed technique is called Pulsed Decimal Communication (PDC). Like PIC, PDC is a protocol for single-channel, high-data rate, low-power dynamic signaling that does not require any clock and data recovery. It however achieves higher data rates by introducing a three-step algorithm, comprising a segmentation, an encoding, and a sub-segmentation step. The segmentation step is used to split the data word into smaller segments and therefore smaller decimal numbers to represent them. The encoding step reduces the number of ON bits in the data and relocates them to lower indices. The sub-segmentation step is used to split further the segments into smaller sub-segments. The complete process reduces the number of pulses required to transmit binary data, thus improving the data rate. Compared with PIC, PDC achieves a 78% improvement in data rate and is more reliable as it eliminates the variations in the number of symbols to be transmitted. An FPGA and an ASIC (65nm technology) implementation of the protocol show that the low-power operation and small footprint of PIC are maintained in PDC, which consumes around 25of power at a clock frequency of 25MHz with a gate count of approximately 2150. Shahzad Muzaffar, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 13 |
| 2016 | Invited - Ultra low power integrated transceivers for near-field IoTabstractIn this paper, we propose mm-Waves for Near-Field IoT, ultralow power transceivers. With small footprint and no external components, the transceivers could be integrated with the sensors, with the wireless sensor nodes organized in a Master-Slave, asymmetrical network. With low complexity and high energy efficiency, the slave nodes benefit from a minimalist design approach with integrated antennas and integrated resonators for absolute frequency accuracy. Two designs are presented. The first is a K-band, super-regenerative, logarithmic-mode, OOK receiver achieving a peak energy efficiency of 200pJ/bit at 4Mb/s and a BER of 10−3. With 800μW peak and 8μW average power, the sensitivity of the receiver is −60dBm for the same data and bit-error rates. Realized in a 65nm CMOS process from GF, the active area of the receiver is 740×670μm2. The second design is a 100Kb/s, V-band transceiver with integrated antenna. It achieves 20pJ/bit energy efficiency (Rx mode) and it provides means for 1/f noise mitigation. Mihai Sanduleanu, Ibrahim M. Elfadel |
DAC | 2 |
| 2016 | Automatic protocol configuration in single-channel low-power dynamic signaling for IoT devicesabstractPulsed-Index Communication (PIC) is a novel technique for single-channel, high-data rate, low-power dynamic signaling that does not require any clock and data recovery. It is fully adapted to the simple yet robust communication needs of Internet of Things (IoT) devices and sensors. However, its error-free operation with maximum data rate requires a careful and judicious setting of PIC data packet and pulse timing parameters. In this paper, we present a new algorithm for automatically detecting and setting the PIC protocol parameters at the power-on phase while removing the restriction on the IoT devices in the PIC network to communicate at a baud rate. The hardware realization of the algorithm is power-efficient and uses closed-form formulas that assign suitable protocol parameters to both ends of the transmission link based on clock rate differences. This difference is determined by a preliminary exchange of clock pulse streams between the transmitter and the receiver. The automatic parameter setting remains operational even in the presence of variations between the local clock frequencies of the IoT devices communicating via PIC. The algorithm is illustrated in the case of several IoT devices with different local clock frequencies that are in need to synchronize their communication parameters with respect to the clock frequency of a master gateway node. A power-on PIC parameter configuration process is rigorously specified, and both an FPGA and an ASIC implementations are presented. In particular, we show that for an ASIC implementation in 65nm technology, the low-power operation of PIC is maintained, consuming only 4.35µW of power at a clock frequency of 25MHz. This architecture is experimentally verified and tested on a point-to-point communication link between two IoT devices connected via a single PIC channel in a master-slave mode. Shahzad Muzaffar, Numan Saeed, Ibrahim M. Elfadel |
VLSI-SoC | 3 |
| 2016 | Compact Model Parameter Extraction Using Bayesian Inference, Incomplete New Measurements, and Optimal Bias SelectionabstractIn this paper, we propose a novel MOSFET parameter extraction method to enable early technology evaluation. The distinguishing feature of the proposed method is that it enables the extraction of MOSFET model parameters using limited and incomplete current–voltage measurements from on-chip monitor circuits. An important step in this method is the use of maximuma posterioriestimation where past measurements of transistors from various technologies are used to learn a prior distribution and its uncertainty matrix for the parameters of the target technology. The framework then utilizes Bayesian inference to facilitate extraction using a very small set of additional measurements. The proposed method is validated using various past technologies and post-silicon measurements for a commercial 28-nm process. The proposed extraction can be used to characterize the statistical variations of MOSFETs with the significant benefit that the restrictions imposed by the backward propagation of variance algorithm are relaxed. We also study the lower bound requirement for the number of transistor measurements needed to extract a full set of parameters for a compact model. Finally, we propose an efficient algorithm for selecting the optimal transistor biases by minimizing a cost function derived from information-theoretic concept of average marginal information gain. Li Yu 0009, Sharad Saxena, Christopher Hess, Ibrahim M. Elfadel, Dimitri A. Antoniadis, Duane S. Boning |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Compact modeling of microbatteries using behavioral linearization and model-order reductionabstractThin-film, solid-state microbatteries represent now a viable alternative for powering small form-factor microsystems or storing the power harvested by energy microsensors. One major obstacle to their widespread use in integrated systems has been the absence of a high-fidelity, physics-based, compact model describing their operation and enabling their design and verification in the same CAD environment as integrated systems or energy harvesters. In this work, we develop and validate such a model using a thorough analysis of the electrochemistry of a thin-film, solid-state lithium-ion microbattery. Our compact model is based on a behavioral linearisation step where the nonlinear partial differential equations (PDEs) describing the microbattery electrochemistry are replaced with linear ones without virtually any loss in accuracy. Unlike Taylor series and other local techniques, our behavioral linearisation is global and is based on the careful examinations and validation of global electroneutrality in the thin-film, solid-state electrolyte. We then apply the well-established methodology of Arnoldi-based model-order reduction (MOR) techniques to develop a compact microbattery model capable of reproducing its input(current)-output(voltage) electrical behavior with less than 1% error with respect to the full discretised PDEs. The use of the reduced-order model results in more than 30X speedup in transient simulation. Mohammed Shemsu Nesro, Lizhong Sun, Ibrahim M. Elfadel |
ASP-DAC | 3 |
| 2015 | A pulsed-index technique for single-channel, low-power, dynamic signaling
Shahzad Muzaffar, Jerald Yoo, Ayman Shabra, Ibrahim M. Elfadel |
DATE | 4 |
| 2015 | Statistical library characterization using belief propagation across multiple technology nodes
Li Yu 0009, Sharad Saxena, Christopher Hess, Ibrahim M. Elfadel, Dimitri A. Antoniadis, Duane S. Boning |
DATE | 4 |
| 2015 | Power management of pulsed-index communication protocolsabstractPulsed-Index Communication (PIC) is a novel technique for single-channel, high-data-rate, low-power dynamic signaling that does not require any clock and data recovery (CDR). It is fully adapted to the simple yet robust communication needs of IoT devices and sensors. Prior work has focused on the power savings that this protocol can achieve as a result of the elimination of circuitry devoted to clock and data recovery. In this paper, we show that further power saving can be achieved using the duty cycle of the pulse as a power control parameter. This power control policy is applied to a single-wire link with significant power saving achieved above and beyond the savings due to the CDR elimination. These power savings are obtained without any impact on data rate. The pulse control policy is implemented using 45nm CMOS technology and verified on various, single-channel communication links. Shahzad Muzaffar, Ibrahim M. Elfadel |
ICCD | 2 |
| 2015 | Multicore power proxies using least-angle regressionabstractThe use of performance counters (PCs) to develop per-core power proxies for multicore processors is now well established. These proxies are typically obtained using traditional linear regression techniques. These techniques have the disadvantage of requiring the full PC set regardless of the workload run by the multicore processor. Typically a computationally expensive principal component analysis is conducted to find the PCs most correlated with each workload. In this paper, we use the more recent algorithm of least-angle regression to efficiently develop power proxies that include only PCs most relevant to the workload. Such PCs can be considered workload signatures and used to categorize the workload and to trigger specific power management action. Our new power proxies are trained and tested on workloads from the PARSEC and SPEC CPU 2006 benchmarks with an average error of less than 3%. Rupesh Raj Karn, Ibrahim M. Elfadel |
ISCAS | 2 |
| 2015 | Timing and robustness analysis of Pulsed-Index protocols for single-channel IoT communicationsabstractPulsed-Index Communication (PIC) is a novel technique for single-channel, high-data-rate, low-power dynamic signaling that does not require any clock and data recovery. It is fully adapted to the simple yet robust communication needs of IoT devices and sensors. In this paper, we present a full quantitative analysis of the timing and robustness properties of PIC protocols, including the impact of important protocol parameters such as pulse width and inter-symbol delays on average data rate and protocol robustness with respect to clock variations. The main result of this paper is a theoretical upper bound on clock variability between transmitter and receiver below which the protocol operates with zero decoding error over an ideal channel. This bound is verified experimentally using a full FPGA implementation that includes point-to-point transmission between two TI MSP430 microcontrollers, acting as two IoT sensor nodes over a single-wire connection. Shahzad Muzaffar, Ibrahim M. Elfadel |
VLSI-SoC | 2 |
| 2014 | Remembrance of Transistors Past: Compact Model Parameter Extraction Using Bayesian Inference and Incomplete New MeasurementsabstractIn this paper, we propose a novel MOSFET parameter extraction method to enable early technology evaluation. The distinguishing feature of the proposed method is that it enables the extraction of an entire set of MOSFET model parameters using limited and incomplete IV measurements from on-chip monitor circuits. An important step in this method is the use of maximum-a-posteriori estimation where past measurements of transistors from various technologies are used to learn a prior distribution and its uncertainty matrix for the parameters of the target technology. The framework then utilizes Bayesian inference to facilitate extraction using a very small set of additional measurements. The proposed method is validated using various past technologies and post-silicon measurements for a commercial 28-nm process. The proposed extraction could also be used to characterize the statistical variations of MOSFETs with the significant benefit that some constraints required by the backward propagation of variance (BPV) method are relaxed. Li Yu 0009, Sharad Saxena, Christopher Hess, Ibrahim M. Elfadel, Dimitri A. Antoniadis, Duane S. Boning |
DAC | 4 |
| 2014 | Unified, ultra compact, quadratic power proxies for multi-core processorsabstractPer-core power proxies for multi-core processors are known to use several dozens of hardware activity monitors to achieve a 2% accuracy on core power estimation. These activity monitors are typically not accessible to the user, and even if they were accessible, there would be a significant overhead in using them at the kernel or OS level for power monitoring or control. Furthermore, when scaled up to hundreds of cores per chip, such power proxies become a computational bottleneck for power management operations such as chip power capping. In this paper, we show that a 4% accuracy or better for per-core power estimation can be achieved using an ultra compact power proxy based on a hybrid set of only four user-accessible parameters, namely core frequency, core temperature, instruction-per-cycle and active-state residency. Our proxy is nonlinear, valid across all P and C states, and is based on a randomized power data collection strategy that aims at exercising all the P and C levels of each core. We illustrate the accuracy of the model using the full suite of the SPEC CPU 2006 benchmarks on a 12-core processor. Muhammad Yasin, Anas Shahrour, Ibrahim M. Elfadel |
DATE | 3 |
| 2014 | Efficient performance estimation with very small sample size via physical subspace projection and maximum a posteriori estimationabstractIn this paper, we propose a novel integrated circuits performance estimation algorithm through a physical subspace projection and maximum-a-posteriori (MAP) estimation. Our goal is to estimate the distribution of a target circuit performance with very small measurement sample size from on-chip monitor circuits. The key idea in this work is to exploit the fact that simulation and measurement data are physically correlated under different circuit configurations and topologies. First, different groups of measurements are projected to a subspace spanned by a set of physical variables. The projection is achieved by performing a sensitivity analysis of measurement parameters with respect to the subspace variables using a virtual source MOSFET compact model. Then a Bayesian treatment is developed by introducing prior distributions over these subspace variables. Maximum a posteriori estimation is then applied using the prior, and an expectation-maximization (EM) algorithm is used to estimate the circuit performance. The proposed method is validated by postsilicon measurement for a commercial 28-nm process. An average error reduction of 2x is achieved which can be translated to 32x reduction on data size needed for samples on the same die. A 150x and 70x sample size reduction on training dies is also achieved compared to traditional least-square fitting method and least-angle regression method, respectively, without reducing accuracy. Li Yu 0009, Sharad Saxena, Christopher Hess, Ibrahim M. Elfadel, Dimitri A. Antoniadis, Duane S. Boning |
DATE | 4 |
| 2014 | Calculation of Generalized Polynomial-Chaos Basis Functions and Gauss Quadrature Rules in Hierarchical Uncertainty QuantificationabstractStochastic spectral methods are efficient techniques for uncertainty quantification. Recently they have shown excellent performance in the statistical analysis of integrated circuits. In stochastic spectral methods, one needs to determine a set of orthonormal polynomials and a proper numerical quadrature rule. The former are used as the basis functions in a generalized polynomial chaos expansion. The latter is used to compute the integrals involved in stochastic spectral methods. Obtaining such information requires knowing the density function of the random input a-priori. However, individual system components are often described by surrogate models rather than density functions. In order to apply stochastic spectral methods in hierarchical uncertainty quantification, we first propose to construct physically consistent closed-form density functions by two monotone interpolation schemes. Then, by exploiting the special forms of the obtained density functions, we determine the generalized polynomial-chaos basis functions and the Gauss quadrature rules that are required by a stochastic spectral simulator. The effectiveness of our proposed algorithm is verified by both synthetic and practical circuit examples. Zheng Zhang 0005, Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | An ultra-compact virtual source FET model for deeply-scaled devices: Parameter extraction and validation for standard cell libraries and digital circuitsabstractIn this paper, we present the first validation of the virtual source (VS) charge-based compact model for standard cell libraries and large-scale digital circuits. With only a modest number of physically meaningful parameters, the VS model accounts for the main short-channel effects in nanometer technologies. Using a novel DC and transient parameter extraction methodology, the model is verified with simulated data from a well-characterized, industrial 40-nm bulk silicon model. The VS model is used to fully characterize a standard cell library with timing comparisons showing less than 2.7% error with respect to the industrial design kit. Furthermore, a 1001-stage inverter chain and a 32-bit ripple-carry adder are employed as test cases in a vendor CAD environment to validate the use of the VS model for large-scale digital circuit applications. Parametric Vdd sweeps show that the VS model is also ready for usage in low-power design methodologies. Finally, runtime comparisons have shown that the use of the VS model results in a speedup of about 7.6×. Li Yu 0009, Omar Mysore, Luca Daniel, Dimitri A. Antoniadis, Ibrahim M. Elfadel, Duane S. Boning |
ASP-DAC | 6 |
| 2013 | Closed-loop control for power and thermal management in multi-core processors: formal methods and industrial practiceabstractThe need to use feedback to come up with context-dependent and workload-aware strategies for runtime power and thermal management (PTM) in high-end and mobile processors has been advocated since the early 2000. Two seminal papers that appeared in 2002 [1], [2] defined a framework for the use of feedback mechanisms for power and temperature control. In [1], the focus was on power management with the goal being to extend battery life on the AMD Mobile Athlon. This was one of the earliest papers to use DVFS settings as actuators to guarantee a given energy level in the battery at the end of a given time interval. The controller was implemented using a combination of OS files and Linux kernel modules. Almost simultaneously, [2] posed the dynamic thermal management task as a formal control-theoretic problem requiring the thermal modeling of the processor and the use of the established control structures of classical feedback theory. Some of the defining features of [2] include the development of layout-based thermal RC models for the processor; the use of an architecturally-driven control mechanism, namely, the instruction fetching rate; and the use of the SPEC2000 benchmarks to illustrate temperature control action under various workloads. The controller used in [2] is a Proportional-Integral-Differential (PID) structure whose input is the deviation of the sensed temperature from the target temperature and whose output is the toggle rate of the instruction fetching mechanism. Ibrahim M. Elfadel, Radu Marculescu, David Atienza 0001 |
DATE | 1 |
| 2013 | Statistical modeling with the virtual source MOSFET modelabstractA statistical extension of the ultra-compact Virtual Source (VS) MOSFET model is developed here for the first time. The characterization uses a statistical extraction technique based on the backward propagation of variance (BPV) with variability parameters derived directly from the nominal VS model. The resulting statistical VS model is extensively validated using Monte Carlo simulations, and the statistical distributions of several figures of merit for logic and memory cells are compared with those of a BSIM model from a 40-nm CMOS industrial design kit. The comparisons show almost identical distributions with distinct run time advantages for the statistical VS model. Additional simulations show that the statistical VS model accurately captures non-Gaussian features that are important for low-power designs. Li Yu 0009, Dimitri A. Antoniadis, Ibrahim M. Elfadel, Duane S. Boning |
DATE | 4 |
| 2013 | Uncertainty quantification for integrated circuits: stochastic spectral methodsabstractDue to significant manufacturing process variations, the performance of integrated circuits (ICs) has become increasingly uncertain. Such uncertainties must be carefully quantified with efficient stochastic circuit simulators. This paper discusses the recent advances of stochastic spectral circuit simulators based on generalized polynomial chaos (gPC). Such techniques can handle both Gaussian and non-Gaussian random parameters, showing remarkable speedup over Monte Carlo for circuits with a small or medium number of parameters. We focus on the recently developed stochastic testing and the application of conventional stochastic Galerkin and stochastic collocation schemes to nonlinear circuit problems. The uncertainty quantification algorithms for static, transient and periodic steady-state simulations are presented along with some practical simulation results. Some open problems in this field are discussed. Zheng Zhang 0005, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 2 |
| 2013 | Stochastic Testing Method for Transistor-Level Uncertainty Quantification Based on Generalized Polynomial ChaosabstractUncertainties have become a major concern in integrated circuit design. In order to avoid the huge number of repeated simulations in conventional Monte Carlo flows, this paper presents an intrusive spectral simulator for statistical circuit analysis. Our simulator employs the recently developed generalized polynomial chaos expansion to perform uncertainty quantification of nonlinear transistor circuits with both Gaussian and non-Gaussian random parameters. We modify the nonintrusive stochastic collocation (SC) method and develop an intrusive variant called stochastic testing (ST) method. Compared with the popular intrusive stochastic Galerkin (SG) method, the coupled deterministic equations resulting from our proposed ST method can be solved in a decoupled manner at each time point. At the same time, ST requires fewer samples and allows more flexible time step size controls than directly using a nonintrusive SC solver. These two properties make ST more efficient than SG and than existing SC methods, and more suitable for time-domain circuit simulation. Simulation results of several digital, analog and RF circuits are reported. Since our algorithm is based on generic mathematical models, the proposed ST algorithm can be applied to many other engineering problems. Zheng Zhang 0005, Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Model order reduction of fully parameterized systems by recursive least square optimizationabstractThis paper presents an approach for the model order reduction of fully parameterized linear dynamic systems. In a fully parameterized system, not only the state matrices, but also can the input/output matrices be parameterized. The algorithm presented in this paper is based on neither conventional moment-matching nor balanced-truncation ideas. Instead, it uses “optimal (block) vectors” to construct the projection matrix, such that the system errors in the whole parameter space are minimized. This minimization problem is formulated as a recursive least square (RLS) optimization and then solved at a low cost. Our algorithm is tested by a set of multi-port multi-parameter cases with both intermediate and large parameter variations. The numerical results show that high accuracy is guaranteed, and that very compact models can be obtained for multi-parameter models due to the fact that the ROM size is independent of the number of parameters in our approach. Zheng Zhang 0005, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 2 |
| 2009 | An efficient resistance sensitivity extraction algorithm for conductors of arbitrary shapesabstractDue to technology scaling, integrated circuit manufacturing techniques are producing structures with large variabilities in their dimensions. To guarantee high yield, the manufactured structures must have the proper electrical characteristics despite such geometrical variations. For a designer, this means extracting the electrical characteristics of a whole family of structure realizations in order to guarantee that they all satisfy the required electrical characteristics. Sensitivity extraction provides an efficient algorithm to extract all realizations concurrently. This paper presents a complete framework for efficient resistance sensitivity extraction. The framework is based on both the Finite Element Method (FEM) for resistance extraction and the adjoint method for sensitivity analysis. FEM enables the calculation of resistances of interconnects of arbitrary shapes, while the adjoint method enables sensitivity calculation in a computational complexity that is independent of the number of varying parameters. The accuracy and efficiency of the algorithm are demonstrated on a variety of complex examples. Tarek A. El-Moselhy, Ibrahim M. Elfadel, Bill Dewey |
DAC | 2 |
| 2009 | Rewritable Channels With Data-Dependent NoiseabstractWe present some recent results on rewritable channels, that is, storage channels that admit optional reading and rewriting of the content at a given cost. This is a general class of channels that models many nonvolatile memories. We focus on the storage capacity of rewritable channels affected by data-dependent noise. We prove tight upper and lower bounds on the storage capacity of a simple yet significant channel model and suggest some simple capacity-achieving coding techniques. Lower bounds on the storage capacity of Gaussian rewritable channels with data-dependent noise are also shown. Thomas Mittelholzer, Michele Franceschini, Luis A. Lastras, Ibrahim M. Elfadel |
ICC | 4 |
| 2009 | A hierarchical floating random walk algorithm for fabric-aware 3D capacitance extractionabstractWith the adoption of ultra regular fabric paradigms for controlling design printability at the 22nm node and beyond, there is an emerging need for a layout-driven, pattern-based parasitic extraction of alternative fabric layouts. In this paper, we propose a hierarchical floating random walk (HFRW) algorithm for computing the 3D capacitances of a large number of topologically different layout configurations that are all composed of the same layout motifs. Our algorithm is not a standard hierarchical domain decomposition extension of the well established floating random walk technique, but rather a novel algorithm that employs Markov Transition Matrices. Specifically, unlike the fast-multipole boundary element method and hierarchical domain decomposition (which use a far-field approximation to gain computational efficiency), our proposed algorithm is exact and does not rely on any tradeoff between accuracy and computational efficiency. Instead, it relies on a tradeoff between memory and computational efficiency. Since floating random walk type of algorithms have generally minimal memory requirements, such a tradeoff does not result in any practical limitations. The main practical advantage of the proposed algorithm is its ability to handle a set of layout configurations in a complexity that is basically independent of the set size. For instance, in a large 3D layout example, the capacitance calculation of 120 different configurations made of similar motifs is accomplished in the time required to solve independently just 2 configurations, i.e. a 60x speedup. Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 2 |
| 2009 | Convergence of Transverse Waveform Relaxation for the Electrical Analysis of Very Wide Transmission Line BusesabstractIn this paper, we study the convergence and approximation error of the transverse waveform relaxation (TWR) method for the analysis of very wide on-chip multiconductor transmission line systems. Significant notational simplicity is achieved in the analysis using a splitting framework for the per-unit-length matrix parameters of the transmission lines. This splitting enables us to show that the state-transition matrix of the coupled lines satisfies a linear Volterra integral equation of the second kind, whose solution is generated by the TWR method as a summable series of iterated kernels with decreasing norms. The upper bounds on these norms are proved to be$O(k^{r}/r!)$, where$r$is the number of iterations and$k$is a measure of the electromagnetic couplings between the lines. Very fast convergence is guaranteed in the case of weak coupling$(k \ll 1)$. These favorable convergence properties are illustrated using a test suite of industrial very large scale integration global buses in a modern 65-nm CMOS process, where it is shown that few$(\approx 3)$Gauss–Jacobi iterations are sufficient for convergence to the exact solution. Ibrahim M. Elfadel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Efficient algorithm for the computation of on-chip capacitance sensitivities with respect to a large set of parametersabstractRecent CAD methodologies of Design-for-Manufacturability (DFM) have naturally led to a significant increase in the number of process and layout parameters that have to be taken into account in design-rule checking. Methodological consistency requires that a similar number of parameters be taken into account during layout parasitic extraction. Because of the inherent variability of these parameters, the issue of efficiently extracting deterministic parasitic sensitivities with respect to such a large number of parameters must be addressed. In this paper, we tackle this very issue in the context of capacitance sensitivity extraction. In particular, we show how the adjoint sensitivity method can be efficiently integrated within a finite-difference (FD) scheme to compute the sensitivity of the capacitance with respect to a large set of BEOL parameters. If np is the number of parameters, the speedup of the adjoint method is shown to be a factor of np/2 with respect to direct FD sensitivity techniques. The proposed method has been implemented and verified on a 65nm BEOL cross section having 10 metal layers and a total number of 59 parameters. Because of its speed, the method can be advantageously used to prune out of the CAD flow those BEOL parameters that yield a capacitance sensitivity less than a given threshold. Tarek A. El-Moselhy, Ibrahim M. Elfadel, David Widiger |
DAC | 2 |
| 2008 | A capacitance solver for incremental variation-aware extractionabstractLithographic limitations and manufacturing uncertainties are resulting in fabricated shapes on wafer that are topologically equivalent, but geometrically different from the corresponding drawn shapes. While first-order sensitivity information can measure the change in pattern parasitics when the shape variations are small, there is still a need for a high-order algorithm that can extract parasitic variations incrementally in the presence of a large number of simultaneous shape variations. This paper proposes such an algorithm based on the wellknown method of floating random walk (FRW). Specifically, we formalize the notion of random path sharing between several conductors undergoing shape perturbations and use it as a basis of a fast capacitance sensitivity extraction algorithm and a fast incremental variational capacitance extraction algorithm. The efficiency of these algorithms is further improved with a novel FRW method for dealing with layered media. Our numerical examples show a 10X speed up with respect to the boundary-element method adjoint or finite-difference sensitivity extraction, and more than 560X speed up with respect to a non-incremental FRW method for a high-order variational extraction. Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 2 |
| 2004 | A CAD Methodology and Tool for the Characterization of Wide On-Chip BusesabstractIn this paper, we describe a CAD methodology for the full electrical characterization of high-performance, on-chip data buses. The goal of this methodology is to allow the accurate modeling and analysis of wide, on-chip data buses as early as possible in the design cycle. The modeling is based on a manufacturing (rather than design-manual) description of the back-end-of-the-line (BEOL) cross section of a given technology and on a full yet contained description of the power-ground mesh in which the data bus is embedded. One major aspect of the resulting electrical models is that they allow the designer to evaluate the wide bus from the three viewpoints of signal timing, crosstalk (both inductive and capacitive), and common-mode signal integrity. Another major aspect is that they take into account such important high-frequency phenomena as the dependence of the current return-path resistance on frequencies. The CAD methodology described in this paper has been extensively correlated with on-chip hardware measurements. Ibrahim M. Elfadel, Alina Deutsch, Gerard V. Kopcsay, Bradley Rubin, Howard H. Smith |
DATE | 1 |
| 1999 | Gradient-Based Optimization of Custom Circuits Using a Static-Timing FormulationabstractArticle Free Access Share on Gradient-based optimization of custom circuits using a static-timing formulation Authors: A. R. Conn IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile , I. M. Elfadel IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile , W. W. Molzen IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile , P. R. O'Brien IBM Electronic Design Automation, 11400 Burnet Road, M. S. 9460, Austin, TX IBM Electronic Design Automation, 11400 Burnet Road, M. S. 9460, Austin, TXView Profile , P. N. Strenski IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile , C. Visweswariah IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile , C. B. Whan IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NY IBM Thomas J. Watson Research Center, Route 134 and Taconic, Yorktown Heights, NYView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 452–459https://doi.org/10.1145/309847.309979Published:01 June 1999Publication History 29citation367DownloadsMetricsTotal Citations29Total Downloads367Last 12 Months21Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Andrew Conn 0001, Ibrahim M. Elfadel, W. W. Molzen, P. R. O'Brien, Philip N. Strenski, Chandramouli Visweswariah, C. B. Whan |
DAC | 2 |
| 1999 | Advances in transistor timing, simulation, and optimization (tutorial abstract)
Jacob K. White 0001, Jacob Avidan, Ibrahim M. Elfadel, Martin D. F. Wong |
ICCAD | 3 |
| 1998 | Interconnect in high speed designs: problems, methodologies and toolsabstractNo abstract available. Phillip J. Restle, Joel R. Phillips, Ibrahim M. Elfadel |
ICCAD | 3 |
| 1997 | Zeros and Passivity of Arnoldi-Reduced-Order Models for Interconnect NetworksabstractCAD tools and research in the area of reduced-ordermodeling of large linear interconnect networks have evolvedfrom merely finding a Pad' e approximation for the givennetwork transfer function to finding an approximate transferfunction that preserves such circuit-theoretic propertiesof the network as stability, passivity, and RLC synthesizability.In particular, preserving passivity guarantees thatthe reduced-order models will be well-behaved when embeddedback in the circuit where the interconnect networkoriginated. While stability can be ascertained by studyingthe poles of the reduced-order transfer function, passivitydepends on both the poles and zeros of the networkdriving-point impedance. In this paper, we present a novelmethod for studying the zeros of reduced-order transferfunctions and show how it yields conclusions about passivityand synthesizability. Moreover, in order to obtain aguaranteed-passive reduced-order model for multiport RCnetworks, a new algorithm based on the Arnoldi iteration ispresented. This algorithm is as computationallyefficient asthe one used to generate guaranteed-stable reduced-ordermodels [Coordinate-transformed Arnoldi for generating guranteed stable reduced-order models for RLC circuits]. Ibrahim M. Elfadel, David D. Ling |
DAC | 1 |
| 1997 | A block rational Arnoldi algorithm for multipoint passive model-order reduction of multiport RLC networksabstractWork in the area of model-order reduction for RLC interconnect networks has focused on building reduced-order models that preserve the circuit-theoretic properties of the network, such as stability, passivity, and synthesizability (Silveira et al., 1996). Passivity is the one circuit-theoretic property that is vital for the successful simulation of a large circuit netlist containing reduced-order models of its interconnect networks. Non-passive reduced-order models may lead to instabilities even if they are themselves stable. We address the problem of guaranteeing the accuracy and passivity of reduced-order models of multiport RLC networks at any finite number of expansion points. The novel passivity-preserving model-order reduction scheme is a block version of the rational Arnoldi algorithm (Ruhe, 1994). The scheme reduces to that of (Odabasioglu et al., 1997) when applied to a single expansion point at zero frequency. Although the treatment of this paper is restricted to expansion points that are on the negative real axis, it is shown that the resulting passive reduced-order model is superior in accuracy to the one that would result from expanding the original model around a single point. Nyquist plots are used to illustrate both the passivity and the accuracy of the reduced order models. Ibrahim M. Elfadel, David D. Ling |
ICCAD | 1 |
| 1996 | Stability criteria for Arnoldi-based model-order reductionabstractPade approximation is an often-used method for reducing the order of a finite-dimensional, linear, time invariant, signal model. It is known to suffer from two problems: numerical instability during the computation of the Pade coefficients and lack of guaranteed stability for the resulting reduced model even when the original system is stable. We show how the numerical instability problem can be avoided using the Arnoldi algorithm applied to an appropriately chosen Krylov subspace. Moreover, we give an easily computable sufficient condition on the system matrix that guarantees the stability of the reduced model at any approximation order. Ibrahim M. Elfadel, Luís Miguel Silveira, Jacob K. White 0001 |
ICASSP | 1 |
| 1996 | A coordinate-transformed Arnoldi algorithm for generating guaranteed stable reduced-order models of RLC circuitsabstractSince the first papers on asymptotic waveform evaluation (AWE), Pade-based reduced order models have become standard for improving coupled circuit-interconnect simulation efficiency. Such models can be accurately computed using bi-orthogonalization algorithms like Pade via Lanczos (PVL), but the resulting Pade approximates can still be unstable even when generated from stable RLC circuits. For certain classes of RC circuits it has been shown that congruence transforms, like the Arnoldi algorithm, can generate guaranteed stable and passive reduced-order models. In this paper we present a computationally efficient model-order reduction technique, the coordinate-transformed Arnoldi algorithm, and show that this method generates arbitrarily accurate and guaranteed stable reduced-order models for RLC circuits. Examples are presented which demonstrates the enhanced stability and efficiency of the new method. Luís Miguel Silveira, Mattan Kamon, Ibrahim M. Elfadel, Jacob K. White 0001 |
ICCAD | 3 |
| 1995 | Global dynamics in principal singular subspace networksabstractA left (resp. right) principal singular subspace of dimension p is the subspace spanned by the p left (resp. right) singular vectors corresponding to the p largest singular values of the cross-correlation matrix of two stochastic processes. We study the global dynamics of a system of nonlinear ordinary differential equations (ODEs) that govern the unsupervised Hebbian learning of left and right principal singular subspaces from samples of the two stochastic processes. In particular, we show that these equations admit a simple Lyapunov function when they are restricted to a well defined smooth, compact manifold, and that they are related to a matrix Riccati differential equation. Moreover, we show that in the case p=1, the solutions of these ODEs can be given in closed form. Ibrahim M. Elfadel |
ICASSP | 1 |
| 1995 | Convex Potentials and their Conjugates in Analog Mean-Field OptimizationabstractThis paper deals with the problem of mapping hybrid (i.e., both discrete and continuous) constrained optimization problems onto analog networks. The saddle-point paradigm of mean-field methods in statistical physics provides a systematic procedure for finding such a mapping via the notion of effective energy. Specifically, it is shown that within this paradigm, to each closed bounded constraint set is associated a smooth convex potential function. Using the conjugate (or the Legendre-Fenchel transform) of the convex potential, the effective energy can be transformed to yield a cost function that is a natural generalization of the analog Hopfield energy. Descent dynamics and deterministic annealing can then be used to find the global minimum of the original minimization problem. When the conjugate is hard to compute explicitly, it is shown that a minimax dynamics, similar to that of Arrow and Hurwicz in Lagrangian optimization, can be used to find the saddle points of the effective energy. As an illustration of its wide applicability, the effective energy framework is used to derive Hopfield-like energy functions and descent dynamics for two classes of networks previously considered in the literature, winner-take-all networks and rotor networks, even when the cost function of the original optimization problem is not quadratic. Ibrahim M. Elfadel |
Neural Comput. | 1 |
| 1995 | Time-Domain Solutions of Oja's EquationsabstractOja's equations describe a well-studied system for unsupervised Hebbian learning of principal components. This paper derives the explicit time-domain solution of Oja's equations for the single-neuron case. It also shows that, under a linear change of coordinates, these equations are a gradient system in the general multi-neuron case. This latter result leads to a new Lyapunov-like function for Oja's equations. John L. Wyatt Jr., Ibrahim M. Elfadel |
Neural Comput. | 2 |
| 1994 | An Efficient Approach to Transmission Line Simulation Using Measured or Tabulated S-parameter DataabstractIn this paper we describe an algorithm for ecient circuit-level simulation of transmission lines which can be speci ed by tables of frequency-dependent scattering parameters.The approach uses a forced stable section-by-section `2 minimization approach to construct a high order rational function approximation to the frequency domain data, and then applies guaranteed stable balanced realization techniques to reduce the order of the rational function.The rational function is then incorporated in a circuit simulator using fast recursive convolution.An example of a transmission line with skin-eect is examined to both demonstrate the eectiveness of the approach and to show its generality. 31 Luís Miguel Silveira, Ibrahim M. Elfadel, Jacob K. White 0001, Moni Chilukuri, Kenneth S. Kundert |
DAC | 2 |
| 1994 | Approximate covariance functions for gray-level Gibbs random fieldsabstractIn this paper, we depart from the usual engineering approach for estimating the means and correlations of gray-level Gibbs random field (GRF) image models using Monte Carlo simulation and show that it is possible to obtain analytical estimates based on the mean-field approximation of statistical physics. In particular, we give a closed-form formula for the covariance matrix of a gray-level GRF model in the case where the coupling between gray levels is quadratic.> Ibrahim M. Elfadel |
ICASSP (5) | 1 |
| 1994 | Gibbs Random Fields, Cooccurrences, and Texture ModelingabstractGibbs random field (GRF) models and features from cooccurrence matrices are typically considered as separate but useful tools for texture discrimination. The authors show an explicit relationship between cooccurrences and a large class of GRF's. This result comes from a new framework based on a set-theoretic concept called the "aura set" and on measures of this set, "aura measures." This framework is also shown to be useful for relating different texture analysis tools. The authors show how the aura set can be constructed with morphological dilation, how its measure yields cooccurrences, and how it can be applied to characterizing the behavior of the Gibbs model for texture. In particular, they show how the aura measure generalizes, to any number of gray levels and neighborhood order, some properties previously known for just the binary, nearest-neighbor GRF. Finally, the authors illustrate how these properties can guide one's intuition about the types of GRF patterns which are most likely to form.> Ibrahim M. Elfadel, Rosalind W. Picard |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | The Softmax Nonlinearity: Derivation Using Statistical Mechanics and Useful Properties as a Multiterminal Analog Circuit Element
Ibrahim M. Elfadel, John L. Wyatt Jr. |
NIPS | 1 |
| 1991 | Markov/Gibbs texture modeling: aura matrices and temperature effectsabstractAn 'aura' framework is used to rewrite the nonlinear energy function of a homogeneous anisotropic Markov/Gibbs random field (MRF) as a linear sum of aura measures. The formulation relates MRFs to co-occurrence matrices. It also provides a physical interpretation of MRF textures in terms of the mixing and separation of gray-level sets, and in terms of boundary maximization and minimization. Within this framework, the authors introduce the use of temperature for texture modeling and show how the parameters of the MRF can be interpreted as temperature annealing rates. In particular, they show evidence for a transition temperature, above which all patterns generated will be visually similar, and below which a pattern evolves down to its ground state. Results which characterize the ground state patterns are described.> Rosalind W. Picard, Ibrahim M. Elfadel, Alex Pentland |
CVPR | 2 |