EDBT 2026 Demo / reviewers in the wild / expert
Tooraj Nikoubin
dblp:19/8271
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-1724-3503ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLS-LCR: Classification Subspace Learning with Learnable Categorical Regularization in Forward Forward NetworksabstractThe Forward–Forward (FF) algorithm provides a biologically motivated alternative to backpropagation by relying on layer-wise local updates computed through forward passes only. Despite its conceptual appeal, FF exhibits limited scalability in deeper networks, where rigid goodness aggregation over the full activation space couples representation learning with discrimination and leads to unstable behavior as depth increases. In this work, we introduce Classification Subspace Learning with Learnable Categorical Regularization in Forward–Forward Networks (CLS-LCR), a structural modification that explicitly decouples feature propagation from goodness computation within each layer. The proposed method partitions activations into a feature subspace and a dedicated CLS subspace for discrimination, and employs a learnable, depth-aware routing mechanism to regulate neuron contributions across layers. By improving the separation between positive and negative goodness signals and mitigating early saturation of discrimination neurons, CLS-LCR enables deeper forward-only architectures to maintain stable learning dynamics. Experiments on MNIST and Fashion-MNIST show that CLS-LCR improves with depth, achieving 96.28% accuracy at depth 8 compared to 89.66% for vanilla FF, while preserving the strictly forward, locally trained nature of the algorithm. Ali Karkehabadi, Zuxiong Tan, Tooraj Nikoubin, Houman Homayoun, Avesta Sasan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | SuperGate-Net: CDM-Based MAC for Scalable Neural Inference via Cross-Layer AnalysisabstractLogic-based neural inference is often assumed to be hardware-efficient, yet its gate-level cost has rarely been validated beyond FPGA-resource abstractions. We present a cross-layer framework connecting quantized neural network training to logic- and circuit-level hardware evaluation. Low-bit sparse MLPs are trained under signed and unsigned quantization; gate-level Boolean synthesis of extracted truth tables yields hundreds of gates per first-layer neuron in the evaluated configurations, indicating limited gate-level scalability for direct truth-table realization. More importantly, low-bit signed quantization transforms each weight-activation product from a general multiplication into a conditional add/subtract/shift primitive, a structural regularization that induces arithmetic regularity across the network. This regularity is precisely what Cell Design Methodology (CDM) supergates are designed to exploit: by merging transistor stacks of repeated arithmetic patterns, CDM achieves up to 47% lower transistor count and 77% lower power at the representative arithmetic-block level in GF22nm; when projected to the network level using those measured primitives, FOM gains of up to 26 × are observed across four benchmarks. Signed quantization consistently outperforms unsigned in accuracy. These results establish a coherent cross-layer principle: quantization induces structure, structured arithmetic enables compact MAC realization, and CDM provides an efficient transistor-level implementation of that structure. Harshith Navin Lachappa, Yogeswar Reddy Thota, Mahathi Ellanti, Srija Vuppala, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | PRISM: Pruning via Rectified-gradient Importance and Saliency Mapping - making models sparse for execution on edgeabstractDeep vision models routinely exceed the memory and latency budgets of edge devices, making pruning a practical necessity. However, existing approaches face a three-way trade-off: methods tailored to specific architectures lack generality, hardware-friendly structured sparsity can hurt accuracy, and accurate importance estimates are often computationally expensive. We present PRISM, a saliency-based pruning framework that resolves this tension by using gated (rectified) gradients to denoise per-sample signals and produce reliable weight-level importance in a single backward pass. These scores accumulate over data and can be aggregated along structural axes—channels, neurons, attention heads, or fixed N: M blocks—so the same criterion supports both unstructured and structured sparsity with linear-time scoring. On ImageNet-1K, PRISM prunes ResNet-50, reducing parameters by 41.9% and MACs by 51.2% while improving Top-1 by +0.66 percentage points; under 2:4 sparsity it reaches 78.2% Top-1. On Transformers, PRISM matches or surpasses strong baselines, e.g., 74.1% Top-1 on DeiT-Tiny with 2: 4 sparsity, and outperforms prior structured methods. By coupling rectified-gradient saliency with lightweight aggregation, PRISM delivers an architecture-agnostic, hardware-aligned, and interpretable route to efficient deep learning across CNNs and ViTs. Zuxiong Tan, Ali Karkehabadi, Houman Homayoun, Tooraj Nikoubin, Avesta Sasan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | Agentic Hardware Synthesis with CDM Supergatesfor Efficient Design GenerationabstractWe present an end-to-end hardware agent that translates multimodal specifications including free-form text, PDF, DOCX, audio, and video into verified hardware across four levels of abstraction: behavioral RTL, gate-level netlist with GLS, transistor-level CMOS SPICE verified by ngspice, and layout-preview SVG with ITRS-based physical estimates (area, power, latency, energy) across configurable technology nodes (130 nm–5 nm). A novel Cell Design Methodology (CDM) supergate knowledge-injection mechanism embeds a structured manifest of multi-output, multi-functional CDM cell patterns into the generation and repair stages, grounding transistor-level synthesis in provably correct pmos/nmos topologies rather than unconstrained behavioral inference. When processing multi-source inputs. A structured intermediate representation canonicalizes all inputs, and Gemini Embeddings rank evidence chunks for retrieval-augmented specification extraction. A two-oracle strategy uses a deterministic spec-derived testbench as the primary oracle and a smoke testbench as fallback. Verification is performed entirely by external tools, including Verilator, Icarus Verilog, Yosys, SymbiYosys, OpenSTA, and ngspice. The agent reports RTL complexity metrics, transistor counts, layout dimensions, physical estimates, formal assertions, and per-iteration audit data for reproducibility. Srija Vuppala, Yogeswar Reddy Thota, Mahathi Ellanti, Harshith Navin Lachappa, Avesta Sasan, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | TinyML Based Stress Detection utilizing PPG Signals: A Lightweight Approach for Smart Wearable Devices
Priyanka Ganesan, Yogeswar Reddy Thota, Hashem Shehata, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | TinyML Enabled Real-Time Bearing Fault Classification in Motors Using Vibration Signals
Yogeswar Reddy Thota, Mojtaba Afshar, Samantha Boden, Brendan Dunlap, Bilal Akin, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | TinyML Based Biometric Authentication Using PPG Signals for Edge Devices
Yogeswar Reddy Thota, Jeffrey Scott Nixon, Bhavya Chandran, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | Area-power and Energy Efficient Substitution box (S-box) in Advanced Encryption Standard (AES)abstractAdvanced Encryption Standard (AES) is a widely used Symmetric-key algorithm. A small characteristics improvement of AES circuits can significantly affect the system's characteristics. The Substitution box (S-box) is the only nonlinear transformation of the AES algorithm, which dominates the complexity of any hardware implementation of the AES. This paper proposes an area-power and energy-efficient design of the S-box based on multiple optimizations including simplifying the input-output true table rather than the direct implementation of the Galois Field inversions GF (28) equations. The gate-level and architecture optimization of the proposed S-box reduces the area of squaring, multiplication with λ, and multiplicative inverse over the GF (24) blocks, using minimum number of XOR gates. The power, area, and power-delay product (PDP) improvement in the selected technologies, 90nm, 45nm, and 32 nm are, respectively; 18-75%, 18-31%, and 30-60%, accounting for a respective improvement of AES, when compared with conventional S-box substituted AES block. Omid Bazgir, Satwik Gali, Tooraj Nikoubin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Hybrid Logical Effort for Hybrid Logic Style Full Adders in Multistage StructuresabstractOne of the critical issues in the advancement of very large scale of integration circuit design is the estimation of timing behavior of the arithmetic circuits. The concept of logical effort provides a proficient approach to comprehend and assess the timing behavior of circuits with conventional CMOS (C-CMOS) structure. However, this technique is not working for circuits with a hybrid structure. On the other hand, numerous circuits with the hybrid structure which are faster and consume less power than C-CMOS one have been proposed for different applications such as portable and IoT devices. In this regard, the necessity of having and use of a simple and efficient timing behavior method like conventional logical effort for analysis of the hybrid adder circuits is inevitable. This paper proposes an efficient analysis and modeling technique that enables designers to assess the timing behavior of hybrid full adder circuits at the block level and anticipate their performance in multistage circuits. The gain and selection factor are introduced as a criterion for accurate selection and optimization of the hybrid adder cells measurable on the single test bench for management of energy efficiency and performance tradeoff. The proposed method is investigated using 32-nm CMOS and FinFET technologies. Hareesh-Reddy Basireddy, Karthikeya Challa, Tooraj Nikoubin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Energy and Area Efficient Three-Input XOR/XNORs With Systematic Cell Design MethodologyabstractIn this brief, we propose three efficient three-input XOR/XNOR circuits as the most significant blocks of digital systems with a new systematic cell design methodology (SCDM) in hybrid-CMOS logic style. SCDM, which is an extension of CDM, plays the essential role in designing efficient circuits. At first, it is deliberately given priority to general design goals in a base structure of circuits. This structure is generated systematically by employing binary decision diagram. After that, concerning high flexibility in design targets, SCDM aims to specific ones in the remaining three steps, which are wise selections of basic cells and amend mechanisms, as well as transistor sizing. In the end, the resultant three-input XOR/XNORs enjoy full-swing and fairly balanced outputs. They perform well with supply voltage scaling, and their critical path contains only two transistors. They also outperform their counterparts exhibiting 27%-77% reduction in average energy-delay product in HSPICE simulation based on TSMC 0.13-μm technology. The symmetric schematic topologies significantly simplify and minimize the layout, as 26%-32% improvement in area is demonstrated. Tooraj Nikoubin, Mahdieh Grailoo, Changzhi Li |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |