Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kaushik De

dblp:14/2886 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
0since 2021 · last 2018
0000-0002-5647-4489ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 8 first-authorSoftware engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Electronic design automation · 34% Embedded and real-time systems · 26% Distributed systems · 26%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
grid computing
0.012004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Electronic design automation
physical design
0.012004
Accurate pre-layout estimation of standard cell characteristics · DAC 2004
Embedded and real-time systems
timing behavior
0.012004
Accurate pre-layout estimation of standard cell characteristics · DAC 2004
High-performance computing
distributed computing infrastructure
0.012004
The Grid2003 Production Grid: Principles and Practice · HPDC 2004
Electronic design automation
logic synthesis
0.011994
A portable parallel algorithm for logic synthesis using transduction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Parallel and multicore computing
parallel programming models and runtimes
0.011994
A portable parallel algorithm for logic synthesis using transduction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Electronic design automation › logic synthesis
combinational logic synthesis
0.011994
A portable parallel algorithm for logic synthesis using transduction · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994

Methods — techniques the papers use, named apart from their topics

layout effect estimation · 0.0runtime system portability · 0.0asynchronous message-driven dataflow · 0.0
YearPublicationVenuePosition
2018 Towards Exascale Computing for High Energy Physics: The ATLAS Experience at ORNL
abstract
Traditionally, the ATLAS experiment at Large Hadron Collider (LHC) has utilized distributed resources as provided by the Worldwide LHC Computing Grid (WLCG) to support data distribution, data analysis and simulations. For example, the ATLAS experiment uses a geographically distributed grid of approximately 200,000 cores continuously (250 000 cores at peak), (over 1,000 million core-hours per year) to process, simulate, and analyze its data (todays total data volume of ATLAS is more than 300 PB). After the early success in discovering a new particle consistent with the long-awaited Higgs boson, ATLAS is continuing the precision measurements necessary for further discoveries. Planned high-luminosity LHC upgrade and related ATLAS detector upgrades, that are necessary for physics searches beyond Standard Model, pose serious challenge for ATLAS computing. Data volumes are expected to increase at higher energy and luminosity, causing the storage and computing needs to grow at a much higher pace than the flat budget technology evolution (see Fig. 1). The need for simulation and analysis will overwhelm the expected capacity of WLCG computing facilities unless the range and precision of physics studies will be curtailed.
V. Ananthraj, Kaushik De, Shantenu Jha, Alexei Klimentov, Danila Oleynik, Sarp Oral, André Merzky, Ruslan Mashinistov, Sergey Panitkin, P. Svirin, Matteo Turilli, Jack C. Wells, Sean R. Wilkinson
eScience2
2018 Fine-Grained Processing Towards HL-LHC Computing in ATLAS
abstract
During LHC's Run-2 ATLAS has been developing and evaluating new fine-grained approaches to workflows and dataflows able to better utilize computing resources in terms of storage, processing and networks. The compute-limited physics of ATLAS has driven the collaboration to aggressively harvest opportunistic cycles from what are often transiently available resources, including HPCs, clouds, volunteer computing, and grid resources in transitional states. Fine-grained processing (with typically a few minutes' granularity, corresponding to one event for the present ATLAS full simulation) enables agile workflows with a light footprint on the resource such that cycles can be more fully and efficiently utilized than with conventional workflows processing O(GB) files per job.
Doug Benjamin, Paolo Calafiura, Taylor Childers, Kaushik De, Alessandro Di Girolamo, Esteban Fullana, Wen Guan, Tadashi Maeno, Nicolò Magini, Paul Nilsson, Danila Oleynik, Shaojun Sun, Vakho Tsulaia, Peter van Gemmeren, Torre J. Wenaus
eScience4
2018 Modeling Impact of Execution Strategies on Resource Utilization
abstract
The analysis of the hundreds of petabytes of raw and derived HEP (High Energy Physics) data will necessitate exascale computing. In addition to unprecedented volume, these data are distributed over hundreds of computing centers. In response to these application requirement, as well as performance requirement by using parallel processing (i.e., parallelism), and as a consequence of technology trends, there has been an increase in the uptake of supercomputers by HEP projects.
Alexey A. Poyda, Mikhail Titov, Alexei Klimentov, Jack C. Wells, Sarp Oral, Kaushik De, Danila Oleynik, Shantenu Jha
eScience6
2017 High-Throughput Computing on High-Performance Platforms: A Case Study
abstract
The computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size re-source. In spite of the impressive scale of the existing distributed computing solutions, the federation of small to mid-size resources will be insufficient to meet projected future demands. This paper is a case study of how the ATLAS experiment has embraced Titan - a DOE leadership facility in conjunction with traditional distributed high-throughput computing to reach sustained production scales of approximately 52M core-hours a years. The three main contributions of this paper are: (i) a critical evaluation of design and operational considerations to support the sustained, scalable and production usage of Titan; (ii) a preliminary characterization of a next generation executor for PanDA to support new workloads and advanced execution modes; and (iii) early lessons for how current and future experimental and observational systems can be integrated with production supercomputers and other platforms in a general and extensible manner.
Danila Oleynik, Sergey Panitkin, Matteo Turilli, Alessio Angius, Sarp Oral, Kaushik De, Alexei Klimentov, Jack C. Wells, Shantenu Jha
eScience6
2004 Accurate pre-layout estimation of standard cell characteristics
abstract
With the advent of deep-submicron technologies, it has become essential to model the impact of physical/layout effects up front in all design flows \citeITRS02. The effect of layout parasitics is considerable even at the intra-cell level in standard cells. Hence, it has become critically important for any transistor-level optimization to consider the effect of these layout parasitics as an integral part of the optimization process. However, since it is not computationally feasible for the actual layout to be a part of any such optimization procedures, we propose a technique that estimates cell layout characteristics without actually performing the layout and subsequent extraction steps. We demonstrate in this work that it is indeed feasible to estimate the layout effects to get timing characteristics that are on average within about 1.5\% of post-layout timing and that the technique is thousands of times faster than the actual creation of layout.
Hiroaki Yoshida, Kaushik De, Vamsi Boppana
DAC2
2004 The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon
HPDC90
2000 Fine-Grained Parallel VLSI Synthesis for Commercial CAD on a Network of Workstations
abstract
We present a fine-grained parallel processing scheme for speeding up an industrial VLSI synthesis tool on a network of workstations without sacrificing the quality of results. The synthesis tool is Ambit BuildGates, a high-capacity ASIC logic synthesis software from Cadence Design Systems. We examine some necessary operating conditions for a practical parallel implementation of such a software, and propose a parallel approach which accommodates for the highly-irregular computation requirements in synthesis and the high-latency, low-bandwidth conditions of the target environment. For pragmatic as well as performance concerns, we designed a parallel algorithm which produces results (synthesized logic) that are identical to those of the original uniprocessor algorithm. We employ heuristic load assessment and adaptive cyclic distribution in order to actively balance the unpredictable load throughout execution, which enables a considerable reduction in runtime (i.e. 51.3 hours down to 23.4 hours) on actual customer design benchmarks.
Victor Kim, Prithviraj Banerjee, Kaushik De
ICPP3
1997 Test methodology for embedded cores which protects intellectual property
abstract
Testing of embedded cores poses a great challenge. These cores cannot be tested in isolation because core I/Os are not directly accessible from ASIC I/Os. A novel test methodology is developed which generates a partial netlist for protection of intellectual property (IP) by performing structural analysis. This partial netlist is used in ASIC level test generation. For the remaining gates of the core, patterns are supplied to test those gates, which can be applied through only core scan chain. Another scheme is developed to select a few I/Os optimally to add boundary scan circuits to improve IP protection.
Kaushik De
VTS1
1995 Failure Analysis for Full-Scan Circuits
abstract
We present a complete system for failure analysis of full-scan circuits. A novel scheme has been proposed to handle multiple faults up to a certain extent by ranking the faults according to the likelihood of being present in the defective part. The user can interactively recompute the suspect fault list by changing some parameters. If the suspect fault list is large, we generate new test patterns to distinguish the faults in the suspect list. User can iterate over a defective part several times until the suspect fault list is reasonably small. Then each suspect site is probed using E-beam. This tool is integrated into design environment of LSI Logic Corporation and produced good results when applied on a few industry circuits.
Kaushik De, Arun Gunda
ITC1
1994 Parallel Logic Synthesis Using Partitioning
abstract
In this paper, we present a partitioning approach of parallel logic synthesis, which is different from the previous approaches which involved parallelization of individual operations within the synthesis algorithm. We partition the given logic circuits and distribute the partitions to different processors for synthesis. For good load balancing, partitioning algorithm is tuned so that the estimated synthesis times of individual partitions are equal. To improve the quality of synthesized circuits, we propose a novel iterative repartitioning and resynthesis approach to parallel logic synthesis. Experimental evaluation in several large circuits are shown on a network of workstations, and results are compared with MIS.
Kaushik De, Prithviraj Banerjee
ICPP (3)1
1994 A portable parallel algorithm for logic synthesis using transduction
abstract
Combinational logic synthesis is a very important phase of VLSI system design. But the logic synthesis process requires large computing times if near optimal quality of the logic network is desired. Parallel processing is fast becoming an attractive solution to reduce the computational time. Recently, researchers have started to investigate parallel algorithms for problems in logic synthesis and verification. Much of the work in parallel algorithms for CAD reported to date, however, suffers from a major limitation. The parallel algorithms proposed for the CAD applications are designed with a specific underlying parallel architecture in mind. Moreover, incompatibilities in programming environments also make it difficult to port these programs across different parallel machines. As a result, a parallel algorithm needs to be developed afresh for every target parallel architecture. The ongoing project of ProperCAD offers an attractive solution to that problem. It allows the development and implementation of a parallel algorithm on the CHARM runtime system such that it can be executed in all the parallel machines without any change in the program. In this paper, we describe a portable parallel algorithm for logic synthesis based on the Transduction method, called ProperSYN. This algorithm uses an asynchronous message driven data-flow model of computation, with no explicit synchronizing barriers separating different phases of parallel computation as used in many previously developed parallel algorithms. Our algorithm is therefore more scalable to large numbers of processors. The algorithm has been implemented and it runs on a variety of parallel machines. We present results on several benchmark circuits for shared memory MIMD machines like Sequent Symmetry and Encore Multimax, distributed memory MIMD machine like the Intel/860 hypercube and distributed processing systems like networks of SUN workstations.>
Kaushik De, Balkrishna Ramkumar, Prithviraj Banerjee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1994 RSYN: a system for automated synthesis of reliable multilevel circuits
abstract
Conventional logic synthesis systems are targeted towards reducing the area required by a logic block, as measured by the literal count or gate count; or, improving the performance in terms of gate delays; or, improving the testability of the synthesized circuit, as measured by the irredundancy of the resultant circuit. In this paper, we address the problem of developing reliability driven logic synthesis algorithms for multilevel logic circuits, which are integrated within the MIS synthesis system. Our procedures are based on concurrent error detection techniques that have been proposed in the past for two level circuits, and adapting those techniques to multilevel logic synthesis algorithms. Three schemes for concurrent error detection in a multilevel circuit are proposed in this paper, using which all the single stuck at faults in the circuit can be detected concurrently. The first scheme uses duplication of a given multilevel circuit with the addition of a totally self-checking comparator. The second scheme proposes a procedure to generate the multilevel circuit from a two level representation under some constraint such that, the Berger code of the output vector can be used to detect any single fault inside the circuit, except at the inputs. A constrained technology mapping procedure is also presented in this paper. The third scheme is based on parity codes on the outputs. The outputs are partitioned using a novel partitioning algorithm, and each partition is implemented using a multilevel circuit. Some additional parity coded outputs are generated. In all three schemes, all the necessary checkers are generated automatically and the whole circuit is placed and routed using the Timberwolf layout package. The area overheads for several benchmark examples are reported in this paper. The entire procedure is integrated into a new system called RSYN.>
Kaushik De, Chitra Natarajan, Devi Nair, Prithviraj Banerjee
IEEE Trans. Very Large Scale Integr. Syst.1
1993 PREST: a system for logic partitioning and resynthesis for testability
abstract
The authors propose a heuristic procedure for partitioning a circuit into several blocks so that after the resynthesis of each block and subsequent reconnection there is a near-minimal number of redundant faults in the circuit. A probabilistic technique is used to estimate the size of a don't care set, and the partitioning approach tries to reduce the don't care size across the partitions. The approach, called PREST (for Partitioning and RESynthesis for Testability), has been applied on various MCNC and ISCAS benchmark circuits, and excellent results in terms of the size and testability of the synthesized circuit have been obtained.>
Kaushik De, Prithviraj Banerjee
IEEE Trans. Very Large Scale Integr. Syst.1
1992 ProperSYN: a portable parallel algorithm for logic synthesis
abstract
An algorithm based on the transduction method and implemented in the ProperCAD environment is described. The parallel ProperSYN algorithm attempts to make the execution time manageably small. The algorithm uses an asynchronous message driven computing model with no synchronizing barriers, and hence it is scalable to a larger number of processors. Also, the algorithm is portable across a wide variety parallel machines. Experimental results on various parallel machines are presented. The algorithm is built around a well-defined sequential algorithm interface such that there can be benefits from future expansion of the sequential algorithm.>
Kaushik De, Balkrishna Ramkumar, Prithviraj Banerjee
ICCAD1
1991 Logic Partitioning and Resynthesis for Testability
Kaushik De, Prithviraj Banerjee
ITC1
1989 Synthesis of delay fault testable combinational logic
abstract
The synthesis of combinational logic which is robust delay fault testable is developed. In a circuit, any reconvergent fanout may result in the presence of blocked paths and/or paths which can be sensitized only if some other path is also sensitized. Implicit don't care terms are used to detect these problems and a local transformation at the reconvergence point is used to upgrade the delay fault testability of the circuit. The sharing of terms in a multilevel circuit is preserved to the greatest extent possible. Good results have been obtained based on an implementation of the algorithm in the LISP programming language on a TI Explorer machine.>
Kaushik Roy 0001, Jacob A. Abraham, Kaushik De, Stephen L. Lusky
ICCAD3