Shashi Kumar

dblp:68/5519 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
abstract
Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance.
Shashi Kumar, Srikanth R. Madikeri, Esaú Villatoro-Tello, Sergio Burdisso, Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Petr Motlícek, D. S. Karthik Pandia, Shankar Venkatesan, Kadri Hacioglu, Andreas Stolcke
ASRU1
2025 XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
abstract
Self-supervised pretrained models exhibit competitive performance in automatic speech recognition (ASR) on finetuning, even with limited in-domain supervised data. However, popular pretrained models are not suitable for streaming ASR because they are trained with full attention context. In this paper, we introduce XLSR-Transducer, where the XLSR-53 model is used as encoder in transducer setup. Our experiments on the AMI dataset reveal that the XLSR-Transducer achieves 4% absolute WER improvement over Whisper large-v2 and 8% over a Zipformer transducer model trained from scratch. To enable streaming capabilities, we investigate different attention masking patterns in the self-attention computation of transformer layers within the XLSR-53 model. We validate XLSR-Transducer on AMI and 5 languages from CommonVoice under low-resource scenarios. Finally, with the introduction of attention sinks, we reduce the left context by half while achieving a relative 12% improvement in WER.
Shashi Kumar, Srikanth R. Madikeri, Juan Zuluaga-Gomez, Esaú Villatoro-Tello, Iuliia Thorbecke, Petr Motlícek, Manjunath K. E, Aravind Ganapathiraju
ICASSP1
2025 Speech Data Selection for Efficient ASR Fine-Tuning using Domain Classifier and Pseudo-Label Filtering
abstract
In real-world speech data processing, the scarcity of annotated data and the abundance of unlabelled speech data present a significant challenge. To address this, we propose an efficient data selection pipeline for fine-tuning ASR models by generating pseudo-labels using WhisperX pipeline and selecting efficient labels for fine-tuning. In our work, we propose a domain classifier system developed with a computationally inexpensive TFIDF and classical machine learning algorithm. Later, we filter data from the classifier output using a novel metric that assesses word ratio and perplexity distribution. The filtered pseudo labels are then used for fine-tuning standard encoder-decoder Whisper models and Zipformer. Our proposed data selection pipeline reduces the dataset size by approximately 1/100thwhile maintaining performance comparable to the full dataset, outperforming random domain-independent selection strategies.
Pradeep Rangappa, Juan Zuluaga-Gomez, Srikanth R. Madikeri, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Sergio Burdisso, Shashi Kumar, Esaú Villatoro-Tello, Iuliia Nigmatulina, Petr Motlícek, D. S. Karthik Pandia, Aravind Ganapathiraju
ICASSP7
2025 Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering
abstract
Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost.
Roberto Andrés Vasco Carofilis, Pradeep Rangappa, Srikanth R. Madikeri, Shashi Kumar, Sergio Burdisso, Jeena J. Prakash, Esaú Villatoro-Tello, Petr Motlícek, Bidisha Sharma, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke
INTERSPEECH4
2025 Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering
abstract
Fine-tuning pretrained ASR models for specific domains is challenging for small organizations with limited labeled data and computational resources. Here, we explore different data selection pipelines and propose a robust approach that improves ASR adaptation by filtering pseudo-labels generated using Whisper (encoder-decoder) and Zipformer (transducer) models. Our approach integrates multiple selection strategies -- including word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis -- to extract high-quality training segments. We evaluate our method on Whisper and Zipformer using a 7500-hour baseline, comparing it to a CER-based approach relying on hypotheses from three ASR systems. Fine-tuning on 7500 hours of pseudo-labeled call center data achieves 12.3% WER, while our filtering reduces the dataset to 100 hours (1.4%) with similar performance; a similar trend is observed on Fisher English.
Pradeep Rangappa, Roberto Andrés Vasco Carofilis, Jeena J. Prakash, Shashi Kumar, Sergio Burdisso, Srikanth R. Madikeri, Esaú Villatoro-Tello, Bidisha Sharma, Petr Motlícek, Kadri Hacioglu, Shankar Venkatesan, Saurabh Vyas, Andreas Stolcke
INTERSPEECH4
2025 Latent Space Factorization in LoRA
abstract
Low-rank adaptation (LoRA) is a widely used method for parameter-efficient finetuning. However, existing LoRA variants lack mechanisms to explicitly disambiguate task-relevant information within the learned low-rank subspace, potentially limiting downstream performance. We propose Factorized Variational Autoencoder LoRA (FVAE-LoRA), which leverages a VAE to learn two distinct latent spaces. Our novel Evidence Lower Bound formulation explicitly promotes factorization between the latent spaces, dedicating one latent space to task-salient features and the other to residual information. Extensive experiments on text, audio, and image tasks demonstrate that FVAE-LoRA consistently outperforms standard LoRA. Moreover, spurious correlation evaluations confirm that FVAE-LoRA better isolates task-relevant signals, leading to improved robustness under distribution shifts. Our code is publicly available at: https://github.com/idiap/FVAE-LoRA
Shashi Kumar, Yacouba Kaloga, John Mitros, Petr Motlícek, Ina Kodrasi
NeurIPS1
2024 TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
abstract
Shashi Kumar, Srikanth Madikeri, Juan Pablo Zuluaga Gomez, Iuliia Thorbecke, Esaú Villatoro-tello, Sergio Burdisso, Petr Motlicek, Karthik Pandia D S, Aravind Ganapathiraju. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Shashi Kumar, Srikanth R. Madikeri, Juan Zuluaga-Gomez, Iuliia Thorbecke, Esaú Villatoro-Tello, Sergio Burdisso, Petr Motlícek, Karthik S, Aravind Ganapathiraju
EMNLP1
2024 Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers
abstract
Traditionally, automatic speech recognition (ASR) and speaker change detection (SCD) systems have been independently trained to generate comprehensive transcripts accompanied by speaker turns. Recently, joint training of ASR and SCD systems, by inserting speaker turn tokens in the ASR training text, has been shown to be successful. In this work, we present a multitask alternative to the joint training approach. Results obtained on the mix-headset audios of AMI corpus show that the proposed multitask training yields an absolute improvement of 1.8% in coverage and purity based F1 score on SCD task without ASR degradation. We also examine the trade-offs between the ASR and SCD performance when trained using multitask criteria. Additionally, we validate the speaker change information in the embedding spaces obtained after different transformer layers of a self-supervised pre-trained model, such as XLSR-53, by integrating an SCD classifier at the output of specific transformer layers. Results reveal that the use of different embedding spaces from XLSR-53 model for multitask ASR and SCD is advantageous.1
Shashi Kumar, Srikanth R. Madikeri, Iuliia Nigmatulina, Esaú Villatoro-Tello, Petr Motlícek, D. S. Karthik Pandia, S. Pavankumar Dubagunta, Aravind Ganapathiraju
ICASSP1
2024 Probability-Aware Word-Confusion-Network-To-Text Alignment Approach for Intent Classification
abstract
Spoken Language Understanding (SLU) technologies have greatly improved due to the effective pretraining of speech representations. A common requirement of industry-based solutions is the portability to deploy SLU models in voice-assistant devices. Thus, distilling knowledge from large text-based language models has become an attractive solution for achieving good performance and guaranteeing portability. In this paper, we introduce a novel architecture that uses a cross-modal attention mechanism to extract bin-level contextual embeddings from a word-confusion network (WNC) encoding such that these can be directly compared and aligned with traditional text-based contextual embeddings. This alignment is achieved using a recently proposed tokenwise constrastive loss function. We validate our architecture’s effectiveness by fine-tuning our WCN-based pretrained model to do intent classification (IC) on the well-known SLURP dataset. Obtained accuracy on the IC task (81%), depicts a 9.4% relative improvement compared to a recent/equivalent E2E method.
Esaú Villatoro-Tello, Srikanth R. Madikeri, Bidisha Sharma, Driss Khalil, Shashi Kumar, Iuliia Nigmatulina, Petr Motlícek, Aravind Ganapathiraju
ICASSP5
2021 Whisper Speech Enhancement Using Joint Variational Autoencoder for Improved Speech Recognition
Vikas Agrawal, Shashi Kumar, Shakti P. Rath
Interspeech2
2021 Speaker Normalization Using Joint Variational Autoencoder
Shashi Kumar, Shakti P. Rath
Interspeech1
2019 Joint Distribution Learning in the Framework of Variational Autoencoders for Far-Field Speech Enhancement
abstract
Far-field speech recognition is a challenging task as speech recognizers trained on close-talk speech do not generalize well to far-field speech. In order to handle such issues, neural network based speech enhancement is typically applied using denoising autoencoder (DA). Recently generative models have become more popular particularly in the field of image generation and translation. One of the popular techniques in this generative framework is variational autoencoder (VAE). In this paper we consider VAE for speech enhancement task in the context of automatic speech recognition (ASR). We propose a novel modification in the conventional VAE to model joint distribution of the far-field and close-talk features for a common latent space representation, which we refer to as joint-VAE. Unlike conventional VAE, joint-VAE involves one encoder network that projects the far-field features onto a latent space and two decoder networks that generate close-talk and far-field features separately. Experiments conducted on the AMI corpus show that it gives a relative WER improvement of 9% compared to conventional DA and a relative improvement of 19.2% compared to mismatched train and test scenario.
Mahesh K. Chelimilla, Shashi Kumar, Shakti P. Rath
ASRU2
2019 X-Band Polarimetric Sar Copolar Phase Difference for Fresh Snow Depth Estimation in the Northwestern Himalayan Watershed
abstract
The estimation of fresh snow depth (FSD) using X-band synthetic aperture radar (SAR) is feasible but challenging depending on the hydrometeorological conditions and data availability. In this study, the FSD is computed for the Beas river watershed in the northwestern Himalayas near Manali, India. It incorporates the recent copolar phase difference (CPD) based FSD inversion model. Moreover, the TerraSAR-X and TANDEM-X bistatic data acquired in January 2016 are used as inputs to the model along with the snow density measurements at the Dhundi ground station. Additionally, apart from applying layover and forest masks, the potential uncertainty sources in the complex mountainous terrains are identified using the H-A-α decomposition and unsupervised Wishart classification techniques. Furthermore, due to the limited number of weather stations, the results are validated using a 3×3 neighbourhood window surrounding the Dhundi site. Also, the effects of different FSD ensemble window sizes are tested for performing sensitivity analysis.
Sayantan Majumdar, Praveen K. Thakur, Ling Chang 0002, Shashi Kumar
IGARSS4
2019 Far-Field Speech Enhancement Using Heteroscedastic Autoencoder for Improved Speech Recognition
Shashi Kumar, Shakti P. Rath
INTERSPEECH1
2018 Remote Compositional Pyroxene Estimates in the Reiner Gamma Formation Using Feature-Oriented Pca: New Insights Into Lunar Swirls
abstract
Moon, being the geological `Rosetta' of the Earth, has its surface mainly composed of basalts and ferroan anorthosite rock suits whose intriguing elemental compositional variations outstands lunar geology. Pyroxene attributes to lunar sub-crustal evolution. The present work focuses on quantifying pyroxene-rich lithologies in the Reiner Gamma Formation (RGF) using an improved technique based on Feature-oriented Principal Component Analysis (FPCA). High resolution data from Moon Mineralogy Mapper (M3) facilitates spectral analysis of lunar soil. Optical Maturity (OMAT) has been operated revealing a more immature soil along the albedo trail of the RGF as compared to the surrounding. Further, FPCA-based band selection exhibited pyroxene-rich occurrences in the central RGF along with traces of olivine orthopyroxene lithology near wrinkled ridges. In support of this, spectra are compositionally analyzed using Modified Gaussian Model (MGM) portraying Mg-rich and moderate to high Ca bearing orthopyroxenes towards the eastern RGF attributing sub-crustal petrography.
Shashwat Shukla, Shashi Kumar
IGARSS2
2013 An Efficient Router Architecture and Its FPGA Prototyping to Support Junction Based Routing in NoC Platforms
abstract
As mesh topology NoC is becoming a standard for implementing multi-core and multi-processor SoCs, there is a focus on developing routing algorithms for efficient on-chip communication. Junction Based Routing (JBR) is one such routing algorithm suitable for large NoC platforms. In this paper, we describe a router architecture as well as its FPGA prototyping for supporting the new routing algorithm. The router architecture required is much more complex because of the need of a routing table in each router and requires more complicated control to manage flow of packets through the router. Router design is described in detail and has a flit latency of only two clock cycles at zero load. The router design was prototyped using ALTERA DE2 board. We present FPGA utilization results for the router design and show that it is feasible to prototype large NoC platforms on available FPGA chips using our router design.
Muhammad Awais Aslam, Shashi Kumar, Rickard Holsmark
DSD2
2010 Designing Efficient Source Routing for Mesh Topology Network on Chip Platforms
abstract
Efficient on-chip communication is very important for exploiting enormous computing power available on a multi-core chip. Network on Chip (NoC) has emerged as a competitive candidate for implementing on-chip communication. Routing algorithms significantly affect the performance of a NoC. Most of the existing NoC architectural proposals advocate distributed routing algorithms for building NoC platforms. Although source routing offers many advantages, researchers avoided it due to its apparent disadvantage of larger header size requirement that results in lower bandwidth utilization. In this paper we make a strong case for the use of source routing for NoCs, especially for platforms with small sizes and regular topologies. We present a methodology to compute application specific efficient paths for communication among cores with a high degree of load balancing. The methodology first selects the most appropriate deadlock free routing algorithm, from a set of routing algorithms, based on the application's traffic patterns. Then the selected (possibly adaptive) routing algorithm is used to compute efficient static paths with the goal of link load balancing. We demonstrate through simulation based evaluation that source routing has a potential of achieving higher performance, for example up to 28% lower latency even at medium load, as compared to distributed routing. A simple scheme is proposed for encoding of router ports to reduce the header overhead. A generic simulator was developed for evaluation and performance comparison between source routing and distributed routing. We also designed a router to support source routing for mesh topology NoC platforms.
Saad Mubeen, Shashi Kumar
DSD2
2010 An Efficient Technique for In-order Packet Delivery with Adaptive Routing Algorithms in Networks on Chip
abstract
Although adaptive routing algorithms promise higher communication performance, as compared to deterministic routing algorithms, they suffer from the out-of-order packet delivery problem. In the context of Network on Chip, the area and computational overhead of ordering packets at the destination is high and may reverse any gain achieved through the use of adaptivity of the routing algorithm. In this paper, we describe a novel scheme for ensuring in-order packet delivery while retaining the performance advantages of adaptive routing. The hardware architecture of a router that supports the proposed scheme is described. Although the basic idea in our proposal is topology independent we evaluate and compare the performance of our scheme with both deterministic as well as adaptive routing algorithms for 2D mesh NoC. As compared to the XY routing algorithm, our technique significantly reduces the packet delay and improves the saturation point. The impact on router area and power dissipation is also discussed. Although the power consumption of routers increase, the energy consumption per flit increases less than 2% on average, since the higher performance allows for draining more traffic during a certain time window.
Maurizio Palesi, Rickard Holsmark, Xiaohang Wang 0001, Shashi Kumar, Mei Yang 0001, Yingtao Jiang, Vincenzo Catania
DSD4
2010 Leveraging Partially Faulty Links Usage for Enhancing Yield and Performance in Networks-on-Chip
abstract
The communication infrastructure of a complex multicore system-on-a-chip is getting an increasing fraction of the overall chip area. According to the International Technology Roadmap for Semiconductors, killer defect density does not decrease over successive technology generations. For this reason, the probability that a manufacturing defect affects the communication system is predicted to increase. In this paper, we deal with manufacturing defects which affect the links in a network-on-chip-based interconnection system. The goal of this paper is to show that by using effective routing functions, supported by appropriate selection policies and with a limited amount of extra logic in the router, it is easy to exploit partially faulty links to improve the performance of the system. We show that, instead of discarding partially faulty links, they can be used at reduced capacity to improve the distribution of the traffic over the network, yielding performance and power improvements. We couple an application-specific routing function with a set of selection policies which are aware of link fault distribution and evaluate them on both synthetic traffic and a real complex multimedia application. We also present an implementation of the router, augmented with the extra logic, to support both the proposed selection functions and the transmission of messages over partially faulty links. We analyze the router in terms of silicon area, timing, and power dissipation.
Maurizio Palesi, Shashi Kumar, Vincenzo Catania
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 HiRA: A methodology for deadlock free routing in hierarchical networks on chip
abstract
Complexity of designing large and complex NoCs can be reduced/managed by using the concept of hierarchical networks. In this paper, we propose a methodology for design of deadlock free routing algorithms for hierarchical networks, by combining routing algorithms of component subnets. Specifically, our methodology ensures reachability and deadlock freedom for the complete network if routing algorithms for subnets are deadlock free. We evaluate and compare the performance of hierarchical routing algorithms designed using our methodology with routing algorithms for corresponding flat networks. We show that hierarchical routing, combining best routing algorithm for each subnet, has a potential for providing better performance than using any single routing algorithm. This is observed for both synthetic as well as traffic from real applications. We also demonstrate, by measuring jitter in throughput, that hierarchical routing algorithms leads to smoother flow of network traffic. A router architecture that supports scalable table-based routing is briefly outlined.
Rickard Holsmark, Shashi Kumar, Maurizio Palesi, Andres Mejia
NOCS2
2009 Application Specific Routing Algorithms for Networks on Chip
abstract
In this paper we present a methodology to develop efficient and deadlock free routing algorithms for Network-on-Chip (NoC) platforms which are specialized for an application or a set of concurrent applications. The proposed methodology, called application specific routing algorithm (APSRA), exploits the application specific information regarding pairs of cores which communicate and other pairs which never communicate in the NoC platform to maximize communication adaptivity and performance. The methodology also exploits the known information regarding concurrency/non-concurrency of communication transactions among cores for the same purpose. We demonstrate, through analysis of adaptivity as well as simulation based evaluation of latency and throughput, that algorithms produced by the proposed methodology give significantly higher performance as compared to other deadlock free algorithms for both homogeneous as well as heterogeneous 2D mesh topology NoC systems. For example, for homogeneous mesh NoC, APSRA results in approximately 30% less average delay as compared to odd-even algorithm just below saturation load. Similarly the saturation load point for APSRA is significantly higher as compared to other adaptive routing algorithms for both homogeneous and non-homogeneous mesh networks.
Maurizio Palesi, Rickard Holsmark, Shashi Kumar, Vincenzo Catania
IEEE Trans. Parallel Distributed Syst.3
2009 Region-Based Routing: A Mechanism to Support Efficient Routing Algorithms in NoCs
abstract
An efficient routing algorithm is important for large on-chip networks [network-on-chip (NoC)] to provide the required communication performance to applications. Implementing NoC using table-based switches provide many advantages, including possibility of changing routing algorithms and fault tolerance, due to the option of table reconfigurations. However, table-based switches have been considered unsuitable for NoCs due to their perceived high area and power consumption. In this paper, we describe the region-based routing (RBR) mechanism which groups destinations into network regions allowing an efficient implementation with logic blocks. RBR can also be viewed as a mechanism to reduce the number of entries in routing tables. RBR is general and can be used in conjunction with any adaptive routing algorithm. In particular, we have evaluated the proposed scheme in conjunction with a general routing algorithm, namely segment-based routing (SR) and an application specific routing algorithm (APSRA) using regular and irregular mesh topologies. Our study shows that the number of entries in the table is significantly reduced, especially for large networks. Evaluation results show that RBR requires only four regions to support several routing algorithms in a 2-D mesh with no performance degradation. Considering link failures, our results indicate that RBR combined with SR is able to tolerate up to 7 link failures in an 8times8 mesh. RBR also reduces area and power dissipation of an equivalent table-based implementation by factors of 8 and 10, respectively. Moreover, the degradation in performance of the network is insignificant when using APSRA combined with RBR.
Andres Mejia, Maurizio Palesi, José Flich, Shashi Kumar, Pedro López 0001, Rickard Holsmark, José Duato
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Efficient Application Specific Routing Algorithms for NoC Systems utilizing Partially Faulty Links
abstract
In this paper we propose a series of efficient routing strategies to effectively utilize NoC systems with partially faulty links. These strategies try to use partially faulty links when the load is high and distribute traffic uniformly on links. Evaluation of our strategies for 8times8 mesh with 7% partially faulty links shows that, using our best strategy, it is possible to achieve an average reduction of up to 50% on packet delay when the load is high. We have also worked out complete designs of routers which can tolerate partial link faults and implement our routing strategies. Approximately 25% extra area and 5% extra power consumption is required for the design of the upgraded router incorporating link fault tolerance and the best routing strategy. However, the overall performance improvement counter balances such overhead resulting in an overall saving in energy consumption of up to 20%. The proposed strategies offer a way to increase the effective yield of large and complex NoC systems.
Dario Frazzetta, Giuseppe Dimartino, Maurizio Palesi, Shashi Kumar, Vincenzo Catania
DSD4
2008 Design of Bandwidth Aware and Congestion Avoiding Efficient Routing Algorithms for Networks-on-Chip Platforms
Maurizio Palesi, Giuseppe Longo, Salvatore Signorino, Rickard Holsmark, Shashi Kumar, Vincenzo Catania
NOCS5
2008 Deadlock free routing algorithms for irregular mesh topology NoC systems with rectangular regions
Rickard Holsmark, Maurizio Palesi, Shashi Kumar
J. Syst. Archit.3
2007 Exploiting Communication Concurrency for Efficient Deadlock Free Routing in Reconfigurable NoC Platforms
abstract
In this paper we make a case for the use of NoC paradigm to develop future FPGAs in which large computational blocks (cores) are connected to each other through a packet switched communication network. We propose a methodology to develop efficient and deadlock free routing algorithms for such NoC platforms which can be specialized for an application or a set of concurrent applications. Application specific topology of communicating cores as well as information about their communication concurrency over time is exploited to maximize communication adaptivity and performance. We demonstrate, both through analysis of adaptivity as well as simulation based evaluation of latency and throughput, that our algorithm gives significantly higher performance as compared to general purpose deadlock free algorithms like XY and odd-even.
Maurizio Palesi, Shashi Kumar, Rickard Holsmark, Vincenzo Catania
IPDPS2
2007 Prediction of flow stress for carbon steels using recurrent self-organizing neuro fuzzy networks
Shashi Kumar, PKS Prakash, Ravi Shankar 0001, Manoj Kumar Tiwari, Shashi Bhushan Kumar
Expert Syst. Appl.1
2006 Off-Line Testing of Delay Faults in NoC Interconnects
abstract
Testing of high density SoCs operating at high clock speeds is an important but difficult problem. Many faults, like delay faults, in such sub-micron chips may only appear when the chip works at normal operating speed. In this paper, we propose a methodology for at-speed testing of delay faults in links connecting two distinct clock domains in a SoC. We give an analytical analysis about the efficiency of this method. We also propose a simple digital hardware structure for the receiver end of the link under test to detect delay faults. It is possible to extend our method to combine it with functional testing of the link and adapt it for online testing
Tomas Bengtsson, Artur Jutman, Shashi Kumar, Raimund Ubar, Zebo Peng
DSD3
2006 Deadlock Free Routing Algorithms for Mesh Topology NoC Systems with Regions
abstract
Region concept helps to accommodate cores larger than the tile size in mesh topology NoC architectures. In addition, it offers many new opportunities for NoC design, as well as provides new design issues and challenges. The most important among these is the design of a deadlock free routing algorithm. In this paper, we present and compare two routing algorithms for mesh topology NoC with regions. The first algorithm is borrowed from the area of fault tolerant networks and is adapted for the NoC context. We compare this with an algorithm designed using a methodology for design of application specific routing algorithms for communication networks. Our study shows that the application specific routing algorithm not only provides much higher adaptivity, but also superior performance as compared to the other algorithm in all traffic cases
Rickard Holsmark, Maurizio Palesi, Shashi Kumar
DSD3
2006 Solving Part-Type Selection and Operation Allocation Problems in an FMS: An Approach Using Constraints-Based Fast Simulated Annealing Algorithm
abstract
Production planning of a flexible manufacturing system (FMS) is plagued by two interrelated problems, namely 1) part-type selection and 2) operation allocation on machines. The combination of these problems is termed a machine loading problem, which is treated as a strongly NP-hard problem. In this paper, the machine loading problem has been modeled by taking into account objective functions and several constraints related to the flexibility of machines, availability of machining time, tool slots, etc. Minimization of system unbalance (SU), maximization of system throughput (TH), and the combination of SU and TH are the three objectives of this paper, whereas two main constraints to be satisfied are related to time and tool slots available on machines. Solutions for such problems even for a moderate number of part types and machines are marked by excessive computational complexities and thus entail the application of some random search optimization techniques to resolve the same. In this paper, a new algorithm termed as constraints-based fast simulated annealing (SA) is proposed to address a well-known machine loading problem available in the literature. The proposed algorithm enjoys the merits of simple SA and simple genetic algorithm and is designed to be free from some of their drawbacks. The enticing feature of the algorithm is that it provides more opportunity to escape from the local minimum. The application of the algorithm is tested on standard data sets, and superiority of the same is witnessed. Intensive experimentations were carried out to evaluate the effectiveness of the proposed algorithm, and the efficacy of the same is authenticated by efficiently testing the performance of algorithm over well-known functions
Manoj Kumar Tiwari, Shashi Kumar, PKS Prakash, Ravi Shankar 0001
IEEE Trans. Syst. Man Cybern. Part A3
2003 A Two-step Genetic Algorithm for Mapping Task Graphs to a Network on Chip Architecture
abstract
Network on Chip (NoC) is a new paradigm for designing core based System on Chip which supports high degree of reusability and is scalable. In this paper we describe an efficient two-step genetic algorithm that has been used to build a tool for mapping an application, described by a parameterized task graph, on to a NoC architecture with a two dimensional mesh of switches as a communication backbone. The computational resources in NoC consist of a set of heterogeneous IP cores. Our algorithm finds a mapping of the vertices of the task graph to available cores so that the overall execution time of the task graph is minimized. We have developed a NoC architecture specific communication delay model to estimate the execution time. Our algorithm is able to handle large task graphs and provide near optimal mapping in a few minutes on a PC platform. Our tool also provides facilities for specifying NoC architecture, generation and viewing synthetic task graphs and viewing the progress of the genetic algorithm as it converges to a solution.
Tang Lei, Shashi Kumar
DSD2
2002 Multi-hop routing of multi-terminal nets for evaluation of hybrid multi-FPGA boards
abstract
In rapid prototyping system application, any large digital circuit can be implemented onto Multi-FPGA Board(MFB). Key MFB architectural feature is its inter-FPGA connections consisting of fixed connections(FC) i.e. FPGA-FPGA connections and programmable connections(PC) i.e. FPGA-programmable switch like FPID-FPGA. MFBs consisting of both the types of connections are known as hybrid MFBs. Since, PC requires two wires as against one wire in FC, MFB must have minimum number of PCs to keep fabrication easy. In partitioned circuit, multi-terminal nets (MTNs) are distributed over one or more circuit parts. When each circuit part is implemented over one FPGA, the MTNs between circuit parts will be routed over PCs and FCs between corresponding FPGAs. Multi-hop routers are used to minimize the use of PCs, but they increase source to sink delay with increasing number of hops. A generic multi-hop router to route two-terminal nets, which obeys the given limit on hops, was presented in our previous work [2002]. In this paper, we extend the same to route multi-terminal nets.
Sushil Chandra Jain, Shashi Kumar
FPT3
1999 Lowering Power Consumption in Clock by Using Globally Asynchronous Locally Synchronous Design Style
abstract
Power consumption in clock of large high performance VLSIs can be reduced by adopting Globally Asynchronous, Locally Synchronous design style (GALS).GALS has small overheads for the global asynchronous communication and local clock generation.We propose methods to a) evaluate the benefits of GALS and account for its overheads, which can be used as the basis for partitioning the system into optimal number/size of synchronous blocks, and b) automate the synthesis of the global asynchronous communication.Three realistic ASICs, ranging in complexity from 1 to 3 million gates, were used to evaluate GALS benefits and overheads.The results show an average power saving of about 70% in clock with negligible overheads. Lowering power consumption in clock by using Globally AsynchronousLocally Synchronous design style.
Ahmed Hemani, Thomas Meincke, Shashi Kumar, Adam Postula, Thomas Olsson 0001, Peter Nilsson 0001, Johnny Öberg, Peeter Ellervee, Dan Lundqvist
DAC3
1999 The Rugby Model: A Conceptual Frame for the Study of Modelling, Analysis and Synthesis Concepts of Electronic Systems
abstract
We propose a conceptual framework, called the Rugby Model, in which designs, design processes and design tools can be studied. It is an extension of the Y chart and adds two dimensions for design representation, namely Data and Tune. The behavioural domain of Y chart is replaced by a more restricted domain called Computation. The structural and physical domains of Y chart are merged into a more general domain called Communication. A fifth dimension deals with design manipulations and transformations at three abstraction levels. The model shall establish a common understanding of modelling and design process concepts for communication and education in the community. In a case study we illustrate how a design can be characterized with the concepts the Rugby model.
Axel Jantsch, Shashi Kumar, Ahmed Hemani
DATE2
1996 Iterative Deepening Multiobjective A
S. Harikumar, Shashi Kumar
Inf. Process. Lett.2
1995 Circuit partitioning with partial order for mixed simulation emulation environment
abstract
A low-cost hybrid simulator for VLSI circuits has been under development at IIT Delhi. The simulator uses a Reconfigurable System (RS) consisting of a limited number of FPGAs for hardware emulation and blends the ideas of hardware emulation with conventional software simulation. A crucial preparatory step is to partition a given circuit into as few parts as possible. The parts are then downloaded onto the RS one by one and emulated in stand alone mode or in conjunction with software simulator. The hybrid simulation environment poses some unique requirements on the partitioner. This paper presents can efficient partitioning algorithm for this purpose. A study of performance of the algorithm on 92 benchmark circuits for various I/O and size constraints of FPGAs has been carried out and good results have been obtained.
Gurmeet Singh Manku, Shashi Kumar
RSP3
1991 A heuristic search strategy for optimization of trade-off cost measures
abstract
The problem of optimization in a multiple-cost search space by combining admissible heuristic estimates for the different cost parameters is investigated. The authors propose an algorithm, MULT* for solving trade-off optimization problems for which there are good admissible heuristics available for each of the associated cost parameters. It is shown that MULT* is an admissible algorithm and it has many of the important properties of A*. Conditions are also given under which it is possible to prune paths in the search graph. A method is also given to relax admissibility of the heuristics to have a more efficient version of MULT* with a bounded decrease in solution quality.>
Shashi Kumar, Vikraman Arvind
ICTAI2
1989 Automatic Synthesis of Microprogrammed Control Units from Behavioral Descriptions
abstract
This paper presents an approach for automatic synthesis of a microprogrammed control unit from a behavioral description, incorporating two new features:
Shashi Kumar, P. Kulshreshtha, Sudipto Ghose
DAC2