VLDB 2026 Research / reviewers in the wild / expert
Sk. Noor Mahammad
dblp:152/3821 · also S. K. Noor Mahammad
· DBLP profile ↗
20ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0003-4708-4769ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 15 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Energy efficient and high throughput prefix-based pattern matching technique on TCAMs for NIDS
Sameera Shaik, S. M. Srinivasavarma Vegesna, Sk. Noor Mahammad |
Integr. | 3 |
| 2024 | TeRa: Ternary and Range based packet classification engine
Dhayalakumar M, Sk. Noor Mahammad |
Integr. | 2 |
| 2024 | A twofold bio-inspired system for mitigating SEUs in the controllers of digital system deployed on FPGA
S. Deepanjali, Sk. Noor Mahammad |
J. Supercomput. | 2 |
| 2024 | Scalable and Accelerated Self-healing Control Circuit Using Evolvable HardwareabstractControllers are mission-critical components of any electronic design. By sending control signals, they decide which and when other data path elements must operate. Faults, especially Single Event Upset (SEU) occurrence in these components, can lead to functional/mission failure of the system when deployed in harsh environments. Hence, competence to self-heal from SEU is highly required in the control path of the digital system. Reconfiguration is critical for recovering from a faulty state to a non-faulty state. Compared to native reconfiguration, the Virtual Reconfigurable Circuit (VRC) is an FPGA-generic reconfiguration mechanism. The non-partial reconfiguration in VRC and extensive architecture are considered hindrances in extending the VRC-based Evolvable Hardware (EHW) to real-time fault mitigation. To confront this challenge, we have proposed an intrinsic constrained evolution to improve the scalability and accelerate the evolution process for VRC-based fault mitigation in mission-critical applications. Experimentation is conducted on complex ACM/SIGDA benchmark circuits and real-time circuits used in space missions, which are not included in related works. In addition, a comparative study is made between existing and proposed methodologies for brushless DC motor control circuits. The hardware utilization in the multiplexer has been significantly reduced, resulting in up to a 77% reduction in the existing VRC architecture. The proposed methodology employs a fault localization approach to narrow the search space effectively. This approach has yielded an 87% improvement on average in convergence speed, as measured by the evolution time, compared to the existing work. S. Deepanjali, Sk. Noor Mahammad |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | Intrinsic Based Self-healing Adder Design Using Chromosome Reconstruction Algorithm
Raghavendra Kumar Sakali, Sk. Noor Mahammad |
J. Electron. Test. | 2 |
| 2023 | Energy efficient multiply-accumulate unit using novel recursive multiplication for error-tolerant applications
S. Skandha Deepsita, Sk. Noor Mahammad |
Integr. | 3 |
| 2023 | Preferential fault-tolerance multiplier design to mitigate soft errors in FPGAs
Raghavendra Kumar Sakali, Sreehari Veeramachaneni, Sk. Noor Mahammad |
Integr. | 3 |
| 2023 | Deterministic Approach for Range-enhanced Reconfigurable Packet Classification EngineabstractReconfigurable hardware is a promising technology for implementing firewalls, routing mechanisms, and new protocols for evolving high-performance network systems. This work presents a novel deterministic approach for a Range-enhanced Reconfigurable Packet Classification Engine based on the number of rules on FPGAs. The proposed framework uses a RAM-established Ternary Match to represent the prefix and the range prefix and efficient rule-reordering for priority selection to get both best-match and multi-match in the same architecture. The recommended framework exhibits 3.2 Mbits of LUT-RAM-based ternary content addressable memory (TCAM) to hold a maximum of 31.3 K of 104- bit rules with 520 MPPS . LUT-RAM, along with BRAM, shows 4 Mbits of TCAM space to implement 38.5 K of 104- bit rules to sustain a throughput of 400 MPPS on Virtex-7 FPGA. The complete architecture offers scalability, better resource utilization (minimum of 50% ), representation of inverse prefix with single entry, range expansion with a single rule, getting best- and multi-match, and determination of the required number of FPGA resources for a particular dataset. Dhayalakumar M, Sk. Noor Mahammad |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Self Healing Controllers to Mitigate SEU in the Control Path of FPGA Based System: A Complete Intrinsic Evolutionary Approach
S. Deepanjali, Sk. Noor Mahammad |
J. Electron. Test. | 2 |
| 2022 | A New Approximate 4-2 Compressor using Merged Sum and Carry
Chinthalgiri Jyothi, Saranya Karunamurthi, Bhaskara Rao Jammu 0001, Sreehari Veeramachaneni, Sk. Noor Mahammad |
J. Electron. Test. | 5 |
| 2022 | Low power, high speed approximate multiplier for error resilient applications
S. Skandha Deepsita, Sk. Noor Mahammad |
Integr. | 2 |
| 2022 | Hardware-based multi-match packet classification in NIDS: an overview and novel extensions for improving the energy efficiency of TCAM-based classifiers
S. M. Srinivasavarma Vegesna, Shanmukha Rao Pydi, Sk. Noor Mahammad |
J. Supercomput. | 3 |
| 2022 | Energy Efficient Error Resilient Multiplier Using Low-power CompressorsabstractThe approximate hardware design can save huge energy at the cost of errors incurred in the design. This article proposes the approximate algorithm for low-power compressors, utilized to build approximate multiplier with low energy and acceptable error profiles. This article presents two design approaches (DA1 and DA2) for higher bit size approximate multipliers. The proposed multiplier of DA1 have no propagation of carry signal from LSB to MSB, resulted in a very high-speed design. The increment in delay, power, and energy are not exponential with increment of multiplier size ( n ) for DA1 multiplier. It can be observed that the maximum combinations lie in the threshold Error Distance of 5% of the maximum value possible for any particular multiplier of size n . The proposed 4-bit DA1 multiplier consumes only 1.3 fJ of energy, which is 87.9%, 78%, 94%, 67.5%, and 58.9% less when compared to M1, M2, LxA, MxA, accurate designs respectively. The DA2 approach is recursive method, i.e., n -bit multiplier built with n/2-bit sub-multipliers. The proposed 8-bit multiplication has 92% energy savings with Mean Relative Error Distance (MRED) of 0.3 for the DA1 approach and at least 11% to 40% of energy savings with MRED of 0.08 for the DA2 approach. The proposed multipliers are employed in the image processing algorithm of DCT, and the quality is evaluated. The standard PSNR metric is 55 dB for less approximation and 35 dB for maximum approximation. S. Skandha Deepsita, Dhayalakumar M, Sk. Noor Mahammad |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2021 | Energy Efficient and Multiplierless Approximate Integer DCT Implementation for HEVCabstractInteger discrete cosine transform (DCT) is one of the popularly used mathematical operation on image/video applications. These applications has a unique property that the error presence in the output will not degrade the output visual quality. Hardware complexity integer DCT for image/video processing is very high, approximation in the hardware results reduction in area, power and delay without compromising the perceptual quality of the output image. This paper proposes an efficient technique to implement the approximate integer DCT for high efficiency video coding (HEVC). The proposed technique approximates the coefficients of the integer DCT using their structural property and hence multiplier hardware is not required for DCT. The proposed design is approximately 77% efficient in terms of power delay product compared with accurate integer DCT. The Quality of Approximate DCT on Images is estimated using Average PSNR and are found to be 32.17 and 30.13 for 512, 256 size images respectively. S. Skandha Deepsita, Kuchipudi Divya, Sk. Noor Mahammad |
VLSI-SoC | 3 |
| 2021 | A TCAM-based Caching Architecture Framework for Packet ClassificationabstractPacket Classification is the enabling function for performing many networking applications like Integrated Services, Differentiated Services, Access Control/Firewalls, and Intrusion Detection. To cope with high-speed links and ever-increasing bandwidth requirements, time-efficient solutions are needed for which Ternary Content Addressable Memories (TCAMs) are popularly used. However, high cost, heavy power consumption, and poor scalability limit their use in many commercial switches. In this work, an efficient framework for caching the packet classification rules on TCAMs in accordance with traffic characteristics is proposed. The proposed design will have a two-level classification engine in which level-1 is a TCAM classifier with a smaller rule capacity and level-2 is a software classifier. The classifiers are assisted by a rule update engine that monitors the rule temporal behavior and performs timely updates of the rules onto level-1. Crucial challenges with respect to the proposed framework design are defined and addressed effectively in this work. Simulation results shows that the architecture can achieve a throughput of 250 Gbps on average by caching only 10% of the total rules for rule databases of sizes 10,000. The proposed architecture, to the best of our knowledge, is the only traffic-aware architecture using TCAMs that provides a completely deployable framework and also can scale for speeds beyond 250 Gbps (OC-1920 and beyond). S. M. Srinivasavarma Vegesna, Shiv Vidhyut, Sk. Noor Mahammad |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | A Novel Rule Mapping on TCAM for Power Efficient Packet ClassificationabstractPacket Classification is the enabling function performed in commodity switches for providing various services such as access control, intrusion detection, load balancing, and so on. Ternary Content Addressable Memories (TCAMs) are the de facto standard for performing packet classification at high speeds. However, TCAMs are highly costlier both in terms of cost and power consumption, forcing the switch vendors towards placing lots of effort for power management. Hence, power-efficient solutions for TCAM-based packet classification are highly relevant even today. In this article, we propose a novel rule placement algorithm based on the unique field values’ presence within the rule databases. We evaluate the total search that is needed to be inspected with respect to the traditional placement approach and the proposed placement approach based on the information content within the fields. Simulation results showed an average reduction of 30.55% in the search space by the proposed placement approach, thereby resulting in an average reduction of 18.85% per search energy over TCAM. With typical TCAM clock speeds ranging between 200--400MHz, this reduction in the per-search energy maps to a huge reduction in the total energy consumed by the TCAM-based network switches. The proposed solution is plug-and-play type requiring only minimal pre-processing within the Network Processing Unit (NPU) of the switches and edge routers. S. M. Srinivasavarma Vegesna, Ashok Chakravarthy Nara, Sk. Noor Mahammad |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | A Novel Range Matching Architecture for Packet Classification Without Rule ExpansionabstractThe speed requirement for the routing table lookup and the packet classification is rapidly increasing due to the increase in the number of packets needed to be processed per second. The hardware-based packet classification relies on ternary content addressable memory (TCAM) to meet this speed requirement. However, TCAM consumes huge power and also supports only for longest prefix match and exact match, where the classification rule also has a range match (RM) field. Hence, it is mandatory to encode the RM into prefix match to accommodate the rule in TCAM. In the worst case, one rule is encoded into (2 W -2) 2 rules (where W is a number of bits to represent range). This work proposes a novel RM architecture, and a detailed analysis about the range field on the standard dataset and the real-life classifier rules are presented. In the literature, the existing RM architecture is used to avoid the range to prefix conversion, but due to the serial operation, it lacks in performance. For constant time lookup, TCAM is the best option, but it does not support RM. The proposed architecture takes one clock cycle for RM and does not require any encoding/ conversion. Hence, there will be a single entry for every rule. It is observed that just 4% of the two-dimensional range rules are present in this dataset, and it will increase the rule set size by 4 times in the best case and nearly 30 times in the worst case. The proposed RM circuit is operated in parallel with TCAM without compromising the speed, and this circuit saves huge power around 70% and area around 61%, where the range to prefix conversion/encoding is completely avoided. The proposed architecture is well suited for current IPv4- and IPv6-based networks, as well as in software-defined networks in the near future. Shanmugakumar Murugesan, Sk. Noor Mahammad |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Multi-mode parallel and folded VLSI architectures for 1D-fast Fourier transform
M. Mohamed Asan Basiri, Sk. Noor Mahammad |
Integr. | 2 |
| 2015 | Genetic annealing with efficient strategies to improve the performance for the NP-hard and routing problemsabstractproblem which cannot be solved in polynomial time for asymptotically large values of and travelling salesman problem (TSP) is important in operations research and theoretical computer science. In this paper a balanced combination of genetic algorithm and simulated annealing has been applied. To improve the performance of finding an optimal solution from huge search space, we have incorporated the use of tournament and rank as selection operators, and inver-over operator mechanism for crossover and mutation. This proposed technique is applied for some routing resource problems in a chip design process and a best optimal solution was obtained, and the TSP appears as a sub-problem in many areas and is used as a benchmark for many optimisation methods. Rajesh Eswarawaka, Sk. Noor Mahammad, Eswara Reddy B. |
J. Exp. Theor. Artif. Intell. | 2 |
| 2014 | An Efficient Hardware-Based Higher Radix Floating Point MAC DesignabstractThis article proposes an effective way of implementing a multiply accumulate circuit (MAC) for high-speed floating point arithmetic operations. The real-world applications related to digital signal processing and the like demand high-performance computation with greater accuracy. In general, digital signals are represented as a sequence of signed/unsigned fixed/floating point numbers. The final result of a MAC operation can be computed by feeding the mantissa of the previous MAC result as one of the partial products to a Wallace tree multiplier or Braun multiplier. Thus, the separate accumulation circuit can be avoided by keeping the circuit depth still within the bounds of the Wallace tree multiplier, namely O ( log 2 n ), or Braun multiplier, namely O ( n ). In this article, three kinds of floating point MACs are proposed. The experimental results show 48.54% of improvement in worst path delay achieved by the proposed floating point MAC using a radix-2 Wallace structure compared with a conventional floating point MAC without a pipeline using a 45nm technology library. The same proposed design gives 39.92% of improvement in worst path delay without a pipeline using a radix-4 Braun structure as compared with a conventional design. In this article, a radix-32 Q 32.32 -format-based floating point MAC is proposed using a Wallace tree/Braun multiplier. Also this article discusses the msb prediction problem and its solution in floating point arithmetic that is not available in modern fused multiply-add designs. The performance results show comparisons between the proposed floating point MAC with various floating point MAC designs for radix-2,-4,-8, and -16. The proposed design has lesser depth than a conventional floating point MAC as well as a lower area requirement than other ways of floating point MAC implementation, both with/without a pipeline. M. Mohamed Asan Basiri, Sk. Noor Mahammad |
ACM Trans. Design Autom. Electr. Syst. | 2 |