EDBT 2026 Demo / reviewers in the wild / expert
Rajeev Jain
dblp:48/6486
· DBLP profile ↗
27ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-authorSystems, architecture and hardware · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 3Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking community drug response prediction models: datasets, models, tools, and metrics for cross-dataset generalization analysisabstractDeep learning and machine learning models have shown promise in drug response prediction (DRP), yet their ability to generalize across datasets remains an open question, raising concerns about their real-world applicability. Due to the lack of standardized benchmarking approaches, model evaluations and comparisons often rely on inconsistent datasets and evaluation criteria, making it difficult to assess true predictive capabilities. In this work, we introduce a benchmarking framework for evaluating cross-dataset prediction generalization in DRP models. Our framework incorporates five publicly available drug screening datasets, seven standardized DRP models, and a scalable workflow for systematic evaluation. To assess model generalization, we introduce a set of evaluation metrics that quantify both absolute performance (e.g. predictive accuracy across datasets) and relative performance (e.g. performance drop compared to within-dataset results), enabling a more comprehensive assessment of model transferability. Our results reveal substantial performance drops when models are tested on unseen datasets, underscoring the importance of rigorous generalization assessments. While several models demonstrate relatively strong cross-dataset generalization, no single model consistently outperforms across all datasets. Furthermore, we identify CTRPv2 as the most effective source dataset for training, yielding higher generalization scores across target datasets. By sharing this standardized evaluation framework with the community, our study aims to establish a rigorous foundation for model comparison, and accelerate the development of robust DRP models for real-world applications. Alexander Partin, Priyanka Vasanthakumari, Oleksandr Narykov, Andreas Wilke, Natasha Koussa, Sara E. Jones, Yitan Zhu, Jamie C. Overbeek, Rajeev Jain, Gayara Demini Fernando, Cesar Sanchez-Villalobos, Cristina Garcia-Cardona, Jamaludin Mohd-Yusof, Nicholas Chia, Justin M. Wozniak, Souparno Ghosh, Ranadip Pal, Thomas S. Brettin, M. Ryan Weil, Rick L. Stevens |
Briefings Bioinform. | 9 |
| 2024 | DE-HNN: An effective neural model for Circuit Netlist representationabstractThe run-time for optimization tools used in chip design has grown with the complexity of designs to the point where it can take several days to go through one design cycle which has become a bottleneck. Designers want fast tools that can quickly give feedback on a design. Using the input and output data of the tools from past designs, one can attempt to build a machine learning model that predicts the outcome of a design in significantly shorter time than running the tool. The accuracy of such models is affected by the representation of the design data, which is usually a netlist that describes the elements of the digital circuit and how they are connected. Graph representations for the netlist together with graph neural networks have been investigated for such models. However, the characteristics of netlists pose several challenges for existing graph learning frameworks, due to the large number of nodes and the importance of long-range interactions between nodes. To address these challenges, we represent the netlist as a directed hypergraph and propose a Directional Equivariant Hypergraph Neural Network (DE-HNN) for the effective learning of (directed) hypergraphs. Theoretically, we show that our DE-HNN can universally approximate any node or hyperedge based function that satisfies certain permutation equivariant and invariant properties natural for directed hypergraphs. We compare the proposed DE-HNN with several State-of-the-art (SOTA) machine learning models for (hyper)graphs and netlists, and show that the DE-HNN significantly outperforms them in predicting the outcome of optimized place-and-route tools directly from the input netlists. Zhishang Luo, Truong Son Hy, Puoya Tabaghi, Michaël Defferrard, Elahe Rezaei, Ryan Carey, William Rhett Davis, Rajeev Jain, Yusu Wang 0001 |
AISTATS | 8 |
| 2024 | Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on CircuitsabstractWhile graph neural networks (GNNs) have gained popularity for learning circuit representations in various electronic design automation (EDA) tasks, they face challenges in scalability when applied to large graphs and exhibit limited generalizability to new designs. These limitations make them less practical for addressing large-scale, complex circuit problems. In this work we propose HOGA, a novel attention-based model for learning circuit representations in a scalable and generalizable manner. HOGA first computes hop-wise features per node prior to model training. Subsequently, the hop-wise features are solely used to produce node representations through a gated self-attention module, which adaptively learns important features among different hops without involving the graph topology. As a result, HOGA is adaptive to various structures across different circuits and can be efficiently trained in a distributed manner. To demonstrate the efficacy of HOGA, we consider two representative EDA tasks: quality of results (QoR) prediction and functional reasoning. Our experimental results indicate that (1) HOGA reduces estimation error over conventional GNNs by 46.76% for predicting QoR after logic synthesis; (2) HOGA improves 10.0% reasoning accuracy over GNNs for identifying functional blocks on unseen gate-level netlists after complex technology mapping; (3) The training time for HOGA almost linearly decreases with an increase in computing resources. Source code of HOGA is freely available at: github.com/cornell-zhang/HOGA. Chenhui Deng, Zichao Yue, Cunxi Yu, Gokce Sarar, Ryan Carey, Rajeev Jain, Zhiru Zhang |
DAC | 6 |
| 2023 | An Automation Framework for Comparison of Cancer Response Models Across ConfigurationsabstractMachine learning has made significant advancements in precision medicine, resulting in the development of various deep learning applications. For instance, in cancer drug response prediction, numerous deep learning models have been created. However, comparing these models across vast configurations of hyperparameters and data sets can be challenging. In this paper, we introduce a new scalable workflow suite that aims to answer questions that arise when comparing different models developed by different teams on similar or the same problems. We explain the problem in more detail and discuss our approach using near-exascale or exascale computers. Justin M. Wozniak, Rajeev Jain, Andreas Wilke, Rylie Weaver, Alexander Partin, Thomas S. Brettin, Rick L. Stevens |
e-Science | 2 |
| 2022 | Automated Design of Analog Circuits Using Reinforcement LearningabstractAnalog and mixed-signal (AMS) blocks are often a crucial and time-consuming part of System-on-Chip (SoC) design, primarily due to a manual circuit and layout iterations. Existing automated solutions for selecting circuit parameters for a given target specification are often not efficient, accurate, or reliable. In order for an automated sizing tool to be practical, we posit that it must: 1) return valid results for a large range of target specifications; 2) understand where and why it is unable to meet certain specifications; 3) consider true layout parasitic simulations for complete end-to-end design; and 4) be automated, allowing most of the design effort to fall on the tool. In this article, we address these critical points by establishing an automated reinforcement learning framework, AutoCkt, by 1) successfully deploying it on a complex two-stage transimpedance amplifier and two-stage folded cascode with biasing in the 16-nm FinFet technology; 2) implementing a new combined distribution deployment algorithm to improve efficiency; 3) analyzing in-depth the efficacy of the trained agent; and 4) demonstrating the functionality of this tool when considering a topology that is highly sensitive to layout parasitics. Our algorithm not only successfully reaches unique, valid, and practical performances, but also does so in state-of-the-art run time, up to 38X more efficient than prior work. In addition, our tool averages just four parasitic simulations obtained by using the Berkeley Analog Generator, to achieve a target specification post-layout for the folded cascode. AutoCkt successfully generates LVS-passed designs with validation in process corner variation results. Keertana Settaluri, Zhaokai Liu, Rishubh Khurana, S. Arash Mirhaj, Rajeev Jain, Borivoje Nikolic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Fast and Accurate PPA Modeling with Transfer LearningabstractThe power, performance and area (PPA) of digital blocks can vary 10:1 based on their synthesis, place, and route tool recipes. With rapid increase in number of PVT corners and complexity of logic functions approaching 10M gates, industry has an acute need to minimize the human resources, compute servers, and EDA licenses needed to achieve a Pareto optimal recipe. We first present models for fast accurate PPA prediction that can reduce the manual optimization iterations with EDA tools. Secondly we investigate techniques to automate the PPA optimization using evolutionary algorithms. For PPA prediction, a baseline model is trained on a known design using Latin hypercube sample runs of the EDA tool, and transfer learning is then used to train the model for an unseen design. For a known design the baseline needed 150 training runs to achieve a 95% accuracy. With transfer learning the same accuracy was achieved on a different (unseen) design in only 15 runs indicating the viability of transfer learning to generalize PPA models. The PPA optimization technique, based on evolutionary algorithms, effectively combines the PPA modeling and optimization. Our approach reached the same PPA solution as human designers in the same or fewer runs for a CORTEX-M0 system design. This shows potential for automating the recipe optimization without needing more runs than a human designer would need. William Rhett Davis, Paul D. Franzon, Luis Francisco, Billy Huggins, Rajeev Jain |
ICCAD | 5 |
| 2018 | Efficient reinforcement learning for automating human decision-making in SoC designabstractThe exponential growth in PVT corners due to Moore's law scaling, and the increasing demand for consumer applications and longer battery life in mobile devices, has ushered in significant cost and power-related challenges for designing and productizing mobile chips within a predictable schedule. Two main reasons for this are the reliance on human decision-making to achieve the desired performance within the target area and power budget, and significant increases in complexity of the human decision-making space. The problem is that to-date human design experience has not been replaced by design automation tools, and tasks requiring experience of past designs are still being performed manually. Shankar Sadasivam, Rajeev Jain |
DAC | 4 |
| 2018 | CANDLE/Supervisor: a workflow framework for machine learning applied to cancer researchabstractBACKGROUND: Current multi-petaflop supercomputers are powerful systems, but present challenges when faced with problems requiring large machine learning workflows. Complex algorithms running at system scale, often with different patterns that require disparate software packages and complex data flows cause difficulties in assembling and managing large experiments on these machines. RESULTS: This paper presents a workflow system that makes progress on scaling machine learning ensembles, specifically in this first release, ensembles of deep neural networks that address problems in cancer research across the atomistic, molecular and population scales. The initial release of the application framework that we call CANDLE/Supervisor addresses the problem of hyper-parameter exploration of deep neural networks. CONCLUSIONS: Initial results demonstrating CANDLE on DOE systems at ORNL, ANL and NERSC (Titan, Theta and Cori, respectively) demonstrate both scaling and multi-platform execution. Justin M. Wozniak, Rajeev Jain, Prasanna Balaprakash, Jonathan Ozik, Nicholson T. Collier, John Bauer, Fangfang Xia, Thomas S. Brettin, Rick L. Stevens, Jamaludin Mohd-Yusof, Cristina Garcia-Cardona, Brian Van Essen, Matt Baughman |
BMC Bioinform. | 2 |
| 2009 | Spectrum Sensing Design Framework Based on Cross-Layer Optimization of Detection EfficiencyabstractCross-layer optimization of the spectrum sensing (SS) sub-system in a cognitive radio (CR) is essential for achieving the best spectrum utilization because the three SS layers - sensing radio, sensing PHY and sensing MAC - jointly determine the overall performance of the SS sub-system. In this paper we introduce a function called detection efficiency that measures how much of the idle primary user spectra is actually utilized by the CR users as a function of regulatory constraints, network specifications and the parameters of the sensing radio, MAC and PHY. By optimizing these parameters to maximize detection efficiency of the SS system, the designer can compare and select the best SS design for a given wireless network requirement. By applying this design method to evaluate several design options for the SS layers, we show that cross-layer optimization can in some cases improve the detection efficiency (and therefore spectrum utilization) from 35% to as much as 99%. Rajeev Jain, Danijela Cabric |
ICC | 2 |
| 1999 | Adaptive radio for multimedia wireless linksabstractThe quality of wireless links suffers from time-varying channel degradations such as interference, flat-fading, and frequency-selective fading. Current radios are limited in their ability to adapt to these channel variations because they are designed with fixed values for most system parameters such as frame length, error control, and processing gain. The values for these parameters are usually a compromise between the requirements for worst-case channel conditions and the need for low implementation cost. Therefore, in benign channel conditions these commercial radios can consume more battery energy than needed to maintain a desired link quality, while in a severely degraded channel they can consume energy without providing any quality-of-service (QoS). While techniques for adapting radio parameters to channel variations have been studied to improve link performance, in this paper they are applied to minimize battery energy. Specifically, an adaptive radio is being designed that adapts the frame length, error control, processing gain, and equalization to different channel conditions, while minimizing battery energy consumption. Experimental measurements and simulation results are presented to illustrate the adaptive radio's energy savings. Charles Chien, Mani Srivastava 0001, Rajeev Jain, Paul Lettieri, Vipin Aggarwal, Robert Sternowski |
IEEE J. Sel. Areas Commun. | 3 |
| 1996 | A low power architecture for wireless multimedia systems: lessons learned from building a power hogabstractUCLA has constructed a network testbed which serves as an environment for developing wireless multimedia systems. The first generation testbed consists of a set of battery powered terminals with low-bit rate video codecs connected to a peer-to-peer multi-hop network over spread spectrum radios. Not surprisingly, this set of features draws considerable current during operation, and achieves a battery life of approximately one hour using a 24 Wh NiMH battery. This report discusses the power consumption of the existing testbed, lessons learned during its development, and a course of action directed to improving battery life for a portable multimedia terminal. The key advancement is the development of a complete system architecture focused on turning off power to components and subsystems. Starting with communications protocols, and progressing up through processors to APIs, this new design extends useful battery life through a significant reduction in current drain. William H. Mangione-Smith, Phil Seong Ghang, Sean Nazareth, Paul Lettieri, Walt Boring, Rajeev Jain |
ISLPED | 6 |
| 1995 | Techniques for FPGA Implementation of Video Compression SystemsabstractReal-time video compression is a challenging subject for FPGA implementation because it typically has a large computational complexity and requires high data throughput. Previous implementations have used parallel banks of FPGAs or DSPs to meet these requirements. Using design techniques that maximize FPGA utilization, we have implemented two video compression systems, each of which uses a single FPGA. In this first system, algorithmic optimizations are made to create a low-complexity implementation that exploits the in-system programmability of the FPGA. This low-complexity implementation performs well, but is limited to a single compression algorithm. In the second system, the FPGA is augmented with an external, low-complexity, video signal processor (VSP) This combination of ASIC and FPGA is flexible enough to implement four common compression algorithms, and powerful enough to execute them in real time. Brian Schoner, John D. Villasenor, Steve Molloy, Rajeev Jain |
FPGA | 4 |
| 1994 | An Algorithm-Driven Processor Design for Video CompressionabstractA design approach for video signal processors is presented that is driven by characteristics of the algorithms, rather than by technological enhancements to conventional microprocessor-style DSP architectures. The goal of this algorithm-driven approach is the design of processors that possess not only the flexibility to execute several algorithms, but an order of magnitude lower complexity than conventional processors. This design approach is applied to the design of a high-speed signal processor targeted at video compression algorithms including the 8/spl times/8 DCT, wavelet/subband coding, and vector quantization. The fabricated processor chip can execute each algorithm at up 25 MPixels/sec and has been implemented with only 80,000 transistors in a 1.2-/spl mu/m CMOS process.> Steve Molloy, Brian Schoner, Avanindra Madisetti, Rajeev Jain, Roy Matic |
ICIP (3) | 4 |
| 1993 | Performance Analysis of an All-Digital BPSK Direct-Sequence Spread-Spectrum IF Receiver ArchitectureabstractA VLSI architecture for an all-digital binary phase shift keying (BPSK) direct-sequence (DS) spread spectrum (SS) intermediate frequency (IF) receiver is presented, and an in-depth performance analysis is given. The all-digital architecture incorporates a Costas loop for carrier recovery and a delay-locked loop for clock recovery. For the pseudorandom noise (PN) acquisition block, a robust energy detection scheme is proposed to reduce false PN locks over a broad range of signal-to-noise ratios. The proposed architecture is intended for use in the 902-928 MHz unlicensed spread spectrum radio band. A 100 kbs information rate and a 12.7 Mchips/second PN code rate are assumed. The IF center frequency is 12.7 MHz and the IF sampling rate is 50.8 Msamples/second, which is the Nyquist rate for the 25.4 MHz bandwidth signal. Finite wordlength effects have been simulated to optimize the architecture, thereby minimizing the chip area, and results of the finite wordlength simulations demonstrate that the chip architecture achieves a bit error rate performance within 1 dB of theory in an additive white Gaussian noise channel.> Bong-Young Chung, Charles Chien, Henry Samueli, Rajeev Jain |
IEEE J. Sel. Areas Commun. | 4 |
| 1992 | Real Time Implementation of Pruned Tree Search Vector QuantizationabstractDiscusses the design of a CMOS integrated circuit for real time vector quantization (VQ) of images at MPEG rates. It has been designed as a slave processor which can implement binary, non-binary, and pruned tree search VQ algorithms. Inputs include the image source vectors, the VQ codevectors and external control signals that direct the search. The chip outputs the index of the codevector that best approximates the input in a mean square error sense. The layout has been generated using a 1.2 mu CMOS library and measures 5.76*6.6 mm/sup 2/. Critical path simulation with SPICE indicates a maximum clock rate of 40 MHz.> Avanindra Madisetti, Rajeev Jain, Richard L. Baker |
Data Compression Conference | 2 |
| 1992 | Hi-PASS: a computer-aided synthesis system for maximally parallel digital signal processing ASICsabstractHi-PASS, a CAD system for digital signal processor (DSP) architecture synthesis, has been developed to automatically produce maximally parallel VLSI designs for real-time applications. The target DSP applications are the class for which desired sample rates are so high that time sharing of hardware is not feasible. Hi-PASS accepts a C code description of the design to be synthesized and produces a structural description of the final design which can be fabricated using standard cells provided by the Lager IV silicon assembly system. Hi-PASS is a collection of modular tools, several of which contain new techniques within the field of synthesis. C code is converted into a flowgraph using a technique known as symbolic interpretation, including new methods for handling certain types of input-dependent control flow. The HOPS program is a heuristic search flowgraph optimizer tailored for the arithmetic intensive structures necessary for real-time DSP applications. The high throughput rates necessary for these applications also require that retiming be used to provide very short inter-register delay times. A new retiming tool has been developed which allows for computationally efficient bit-level retiming.> Philip Duncan, Shobana Swamy, Steve Sprouse, Daniel Potasz, Rajeev Jain, Neal M. Gafter, William Cammack, Yiwan Wong, Wanda Gass |
ICASSP | 5 |
| 1992 | Decimation filter compiler for oversampling A/D applicationsabstractA silicon compiler (DECGEN) which generates a decimation filter for oversampling A/D converters is described. It generates a multistage filter consisting of two stages of sinc filter and an equiripple bandshaping FIR filter. A heuristic procedure is derived for dividing the decimation ratio between the two stages of sinc filter in an optimum fashion. Digit serialization and hardware multiplexing are used for the target architecture in order to reduce silicon area.> Soei-Shin Hang, Rajeev Jain |
ICASSP | 2 |
| 1992 | Architectures and integrated circuits for real time vector quantization of imagesabstractAn architecture for implementing pruned tree search vector quantization (VQ) which has a performance close to the optimal rate distortion bound is presented. A layout design is also presented to verify the feasibility of such an implementation. The chip has been designed as a slave processor which can implement binary, non-binary and pruned tree search VQ algorithms which allow near optimal performance in a mean square error sense while still permitting real time VQ codevectors and external control signals that direct the search. The chip outputs the index of the codevector that best approximates the input in a mean square error sense. The input source and codevectors have a wordlength of 8 bits each. An internal wordlength of 24 bits allows a maximum of 256 pixels per vector. The layout has been generated using a 1.2- mu library and is 5.76 mm*6.6 mm. Critical path simulation with SPICE indicates a maximum clock rate of 40 MHz.> Avanindra Madisetti, Rajeev Jain, Richard L. Baker, Raffi Dianysian |
ICASSP | 2 |
| 1992 | An integrated circuit design for pruned tree-search vector quantization encoding with an off-chip controllerabstractThe design of an encoder for pruned tree-search vector quantization (VQ) is discussed. This allows near-optimal performance in a mean square error sense while keeping the hardware complexity low. The encoder is partitioned into a slave processor chip that computes the distance and performs minimizations and an off-chip controller that directs the search. Pointer addressing is exploited in the codebook memory to keep the controller hardware simple. Inputs to the slave processor include the source vectors, the code vectors; and external control signals. The slave processor outputs the index of the code vector that best approximates the input in a mean square error sense. The layout for the slave processor has been generated using a 1.2- mu m CMOS library and measures 5.76*6.6 mm/sup 2/. Critical path simulation with SPICE indicates a throughput of 89 million multiply-accumulates per second. This implies that real-time processing at MPEG rates can be achieved if the number of levels (N7) and the number of children at any node (M) obey the constraint M*N> Rajeev Jain, Avanindra Madisetti, Richard L. Baker |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1991 | Computer-aided design of high-speed lattice wave digital filter integrated circuitsabstractA silicon compiler for generating IIR filter integrated circuits based on the lattice wave digital filter structure is presented. A novel architecture has been developed to minimize the chip area while achieving a high throughput. A fifth-order filter synthesized with the compiler measures 4.17 mm/sup 2/ and has an estimated sample rate of at least 60 MHz in 0.8 mu m BiCMOS gate-array technology. This represents an order of magnitude improvement over the lattice wave digital filter compiler.> Lynette Liu, Toshiaki Yoshino, Steven Sprouse, Rajeev Jain |
ICASSP | 4 |
| 1991 | An integrated CAD system for algorithm-specific IC designabstractLAGER is an integrated computer-aided design system for algorithm-specific integrated circuit design, targeted at applications such as speech processing, image processing, telecommunications, and robot control. LAGER provides user interfaces at behavioral, structural, and physical levels and allows easy integration of novel CAD tools. LAGER consists of a behavioral mapper and a silicon assembler. The behavioral mapper maps the behavior onto a parameterized structure to produce microcode and parameter values. The silicon assembler then translates the filled-out structural description into a physical layout, and, with the aid of simulation tools, the user can fine tune the data path by iterating this process. The silicon assembler can also be used without the behavioral mapper for high-sample-rate applications. A number of algorithm-specific ICs designed with LAGER have been fabricated and tested, and as examples, a robot arm controller chip and a real-time image segmentation chip are described.> C. Bernard Shung, Rajeev Jain, Ken Rimey, Mani Srivastava 0001, Brian C. Richards, Erik Lettang, Syed Khalid Azim, Lars E. Thon, Paul N. Hilfinger, Jan M. Rabaey, Robert W. Brodersen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1990 | A functional silicon compiler for high speed FIR digital filtersabstractA functional compiler system for the implementation of high-speed finite impulse-response (FIR) digital filters on gate-array ICs is presented. The system is capable of implementing complex digital filters directly from frequency-domain specifications. Fast turnaround and sample rates in excess of 100 MHz are achieved by using a combination of architectural optimization and advanced 0.8- mu m BiCMOS gate-array technology. A 64-tap FIR digital filter synthesized using this new functional compiler system is presented. It has been fabricated and tested fully functional at a sample frequency of 100 MHz.> Paul T. Yang, Rajeev Jain, Toshiaki Yoshino, Wanda Gass, Ashwin Shah |
ICASSP | 2 |
| 1989 | FDT-a design tool for switched capacitor filtersabstractThe authors present FDT (filter design tool), a complete design tool for switched-capacitor filter synthesis from specifications to circuit schematic generation. The tool supports both ladder and biquad structure synthesis. It provides analysis and optimization capabilities at both pole-zero level and circuit level. It also extends these capabilities for noise calculation and noise minimization of the synthesized filter. These facilities together with a completely interactive and flexible user interface provide an efficient design tool for filter synthesis and design evaluation in a short time. FDT was successfully used in the design of a high-pass filter and two delay equalizers in an analog interface chip.> Sattam Dasgupta, Mahesh Mehendale, V. R. Sudershan, Rajeev Jain, Nagaraj Subramanyam, James Hochschild |
ICCAD | 4 |
| 1988 | An ASIC architecture for contour line filteringabstractProposes an application-specific IC architecture for implementing a contour line filter for a real-time fingerprint recognition system. The architecture demonstrates the fact that a library-based set of data-path modules which are steered by powerful, synchronized controllers can result in real-time performance on a single chip, even for a complex image processing application. Moreover, the external frame memories required are minimized in the design by working in-place and by appropriate scheduling at the cost of little extra hardware and few additional evaluation cycles.> Francky Catthoor, Lawrence O'Gorman, Rajeev Jain |
ICASSP | 3 |
| 1986 | A fast adder-based multiplication unit for customised digital signal processorsabstractThis paper presents the design of a parameterisable multiplication-accumulation unit for integrated digital signal processing devices. It is especially suited for macrocell based customised signal processors as proposed by [1]. The complete architecture including the data-path and associated control are described. Estimates of the chip area and speed are presented based on a 3 µm CMOS cell library. The essential advantage of the proposed design is that the time taken for ann-bit byn-bit multiplication of two signals is(n/4)+1processor cycles where each cycle only requires the shifted values of two numbers to be added to a third number using a carry-save and a carry-propagate adder. Compared to conventional shifter-adder based multiplication units this leads to an improvement in the throughput by approximately a factor 4 with one-third the area of a fully hardwired array multiplier for a 16×16 bit multiplication. Luc de Vos, Rajeev Jain, Hugo De Man, Walter Ulbrich |
ICASSP | 2 |
| 1985 | CAD Tools for the optimized design of custom VLSI wave digital filtersabstractCAD tools to support a top-down custom design methodology for integrated digital filters are presented. The methodology is based on a tool-box concept, which makes use of specialised analysis, synthesis and optimization programs at each design level: the network, the architecture and the circuit layout. In this paper, CAD tools for filter synthesis, network optimization and architecture optimization are developed. These tools complement design aids for architecture synthesis, and automatic layout generation (silicon compilers) to create a complete design environment. By combining both synthesis as well as optimization aids at each design level, it is possible to achieve complete automation while retaining efficient use of silicon area, speed and power consumption. Application of these tools to the custom integration of wave digital filters with bit-serial architectures is demonstrated. Rajeev Jain, Gert Goossens, Luc Claesen, Joos Vandewalle, Hugo De Man, L. Gazsi, Alfred Fettweis |
ICASSP | 1 |
| 1984 | Efficient CAD tools for the coefficient optimisation of arbitrary integrated digital filtersabstractAn interactive CAD methodology is presented for designing digital filters with optimised discrete coefficients. The objective is to minimise the cost of integrationC(\undertilde{a})subject to the constraint that the desired specifications on the complex frequency transfer functionH(w,\undertilde{a})are satisfied. The vector a represents the fixed coefficients in the filter network. For VLSI DSP architectures, where multiplication is implemented using software controlled or hardwired shift-and-add operations,C(\undertilde{a})is defined as the total number of non-zero bits in the discrete representation of\undertilde{a}. An accurate multiparameter analysis technique, based on the bilinear property ofH(w,\undertilde{a}), is exploited to develop numerically reliable and computationally efficient tools for function evaluation and feasible region determination. These are built into a CAD framework together with tools for computingC(\undertilde{a})and predicting the location of the minima ofC(\undertilde{a})Extremely fast optimisation strategies using this framework are demonstrated, for different filter topologies with up to 10 coefficients. Rajeev Jain, Joos Vandewalle, Hugo De Man |
ICASSP | 1 |