EDBT 2026 Demo / reviewers in the wild / expert
Satyajit Das
dblp:134/3948
· DBLP profile ↗
21ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0002-7550-2641ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology and Reliability Aware Qubit Mapping for Quantum Core SystemsabstractScaling quantum processors motivates multicore (modular) quantum architectures, where inter-core teleportation and state transfer are substantially slower and less reliable than intra-core operations. A key compilation/runtime challenge in such systems is inter-core qubit mapping: assigning logical qubits to cores across circuit timeslices to minimize non-local transfers under core-capacity and network constraints. Existing approaches such as FGP-rOEE, QUBO-based mapping, and Hungarian Qubit Assignment (HQA) reduce inter-core interactions, but are often evaluated under simplified interconnect assumptions and do not explicitly optimize for topology and reliability-dependent transfer costs. In this work, we present TR-HQA, a topology and reliability aware extension of HQA that minimizes expected inter-core transfer cost over a weighted interconnect graph, where edge weights capture hop distance and link success probability. We also propose GTR-QA, a greedy approximation that reduces placement time while maintaining competitive transfer cost. Using a multicore mapping simulator built on a pytket front-end, we evaluate TR-HQA and GTR-QA across circuit families (structured benchmarks and random circuits) and multiple interconnect topologies. Our results show that TR-HQA consistently reduces expected inter-core communication cost compared to topology-agnostic baselines, while GTR-QA provides a favorable quality–runtime trade-off. Across standard quantum benchmarks (QFT, Cuccaro Adder, and Draper Adder), TR-HQA reduces inter-core communication cost by 2.15 × on average (1.70 × –2.67 ×) and improves runtime by 11.8 × on average (1.02 × –30 ×). Rajeswari Suance P. S, Akansh Khandelwal, Surankan De, Satyajit Das, John Jose, Maurizio Palesi |
CF | 4 |
| 2026 | ALISTA: Accelerator using LSH-based maximum inner-product search in transformer attention
Shine Parekkadan Sunny, Satyajit Das |
Integr. | 2 |
| 2025 | SimEx-ViT: Explainable Vision Transformer with Similarity-Based Attention Modulation
R. Selventhiran, Vish Rajalingam, Satyajit Das |
CAIP (2) | 3 |
| 2024 | SplitMS: Split Modulo-Scheduling for Accelerating Loops Onto CGRAsabstractCoarse-Grained Reconfigurable Array (CGRA) ar-chitectures are popular for accelerating loop kernels due to a good balance between energy efficiency and flexibility. Modulo scheduling (MS) is the preferred solution for efficiently mapping loops onto CGRAs. Existing CGRA MS algorithms suffer from low resource utilization if the number of operation nodes in the Data Flow Graph (DFG) is less than the number of Processing Elements (PEs) in the CGRA. To improve instruction level parallelism (ILP), the common approaches unroll the loop before applying MS. However, finding valid MS solutions for larger DFGs becomes difficult for CGRAs with resource constraints. This paper proposes a novel Split Modulo-Scheduling (SplitMS) technique to improve the ILP by segmenting the target CGRA into clusters and mapping loop chunks. We also present a lightweight hardware approach to support the cluster execution. Experiments show that SplitMS for a$4\times[2\times 2]$CGRA cluster achieves an average speedup of$2.8\times$over MS for a$4\times 4$target CGRA with 8 Load-Store Units (LSUs). SplitMS increases an average of$2.9\times$the PE utilization and$3\times$the energy efficiency over the conventional MS approach. Christie Sajitha Sajan, Kevin J. M. Martin, Satyajit Das, Philippe Coussy |
DSD | 3 |
| 2024 | ByteZip: Efficient Lossless Compression for Structured Byte Streams Using DNNsabstractData compression plays a key role in efficient data storage, transmission, and processing. With the fast development of deep learning techniques, deep neural networks have been used in this field to achieve a higher compression rate. Deep learning-based general-purpose lossless compression techniques are formulated as an autoregressive sequential prediction problem. These methods are state-of-the-art in terms of compression ratio but not practical due to runtime and resource constraints. Recent advances in lossless image compression using non-autoregressive methods for probability modeling prove to be a faster and more practical approach. In this paper, we propose ByteZip, a lossless compression method based on the non-autoregressive approach for known or defined structured byte streams. ByteZip involves hierarchical probabilistic modeling using autoencoders and density mixture models. This approach reduces the overhead of sequential processing. The goal is to design a practical lossless compressor with faster compression and decompression along with a competitive compression ratio. Experiments show that the proposed approach achieves a 64× higher compression speed than the state-of-the-art transformer-based model TRACE with an overhead of only 5% less size reduction on average. Our approach outperforms general-purpose compressors such as Gzip (23% more size reduction on average) and 7z (16% more size reduction on average). Parvathy Ramakrishnan P, Satyajit Das |
IJCNN | 2 |
| 2024 | Efficient FFT-Based CNN Acceleration with Intra-Patch Parallelization and Flex-Stationary DataflowabstractThis paper presents a novel Hadamard Product Generator (HPG) for FFT-based CNN acceleration that effectively addresses computation and energy bottlenecks. The proposed block uses Intra-Patch parallelization to optimize Complex Multiply and Accumulate (CMAC) unit utilization and maintains identical reuse behavior across patch elements. This scheme also offers multiple spatial unrolling schemes to increase resource reuse. Additionally, the proposed HPG leverages the Flex-Stationary dataflow to adaptively store tensors with high reuse opportunities in the on-chip memory. The prototype is implemented on the Zynq MPSoC (XCZU7CG). It showcases throughput gains of 8.16× for VGG-16 and 9.30× for AlexNet compared to the state-of-the-art frequency domain accelerator with an area overhead of only 28.51%. It achieves an 8.82× average improvement over Eyeriss and a 7.87× improvement over Flexflow in EDP with a similar hardware configuration. Shine Parekkadan Sunny, Satyajit Das |
ISCAS | 2 |
| 2024 | CREPE: Concurrent Reverse-Modulo-Scheduling and Placement for CGRAsabstractCoarse-Grained Reconfigurable Array (CGRA) architectures are popular as high-performance and energy-efficient computing devices. Compute-intensive loop constructs of complex applications are mapped onto CGRAs by modulo-scheduling the innermost loop dataflow graph (DFG). In the state-of-the-art approaches, mapping quality is typically determined by initiation interval (II), whileschedule lengthfor one iteration is neglected. However, for nested loops,schedule lengthbecomes important. In this article, we propose CREPE, aConcurrentReverse-modulo-scheduling andPlacement technique for CGRAs that minimizes bothIIandschedule length. CREPE performs simultaneous modulo-scheduling and placement coupled with dynamic graph transformations, generating good-quality mappings with high success rates. Furthermore, we introduce a compilation flow that maps nested loops onto the CGRA and modulo-schedules the innermost loop using CREPE. Experiments show that the proposed solution outperforms the conventional approaches in mapping success rate and total execution time with no impact on the compilation time. CREPE maps all kernels considered while state-of-the-art techniques Crimson and Epimap failed to find a mapping or mapped at very highIIs. On a 2×4 CGRA, CREPE reports a 100% success rate and a speed-up up to 5.9× and 1.4× over Crimson with 78.5% and Epimap with 46.4% success rates respectively. Chilankamol Sunny, Satyajit Das, Kevin J. M. Martin, Philippe Coussy |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | An Efficient and Flexible Stochastic CGRA Mapping ApproachabstractCoarse-Grained Reconfigurable Array (CGRA) architectures are promising high-performance and power-efficient platforms. However, mapping applications efficiently on CGRA is a challenging task. This is known to be an NP complete problem. Hence, finding good mapping solutions for a given CGRA architecture within a reasonable time is complex. Additionally, finding scalability in compilation time and memory footprint for large heterogeneous CGRAs is also a well known problem. In this article, we present a stochastic mapping approach that can efficiently explore the architecture space and allows finding best of solutions while having limited and steady use of memory footprint. Experimental results show that our compilation flow allows to reach performances with low-complexity CGRA architectures that are as good as those obtained with more complex ones thanks to the better exploration of the mapping solution space. Parameters considered in our experiments are number of tiles, Register File (RF) size, number of load/store (LS) units, network topologies, and so on. Our results demonstrate that high-quality compilation for a wide range of applications is possible within reasonable run-times. Experiments with several DSP benchmarks show that the best CGRA configuration from the architectural exploration surpasses an ultra low-power DSP optimized RISC-V CPU to achieve up to 15.28× (with an average of 6× and minimum of 3.4×) performance gain and 29.7× (with an average of 13.5× and minimum of 6.3×) energy gain with an area overhead of 1.5× only. Satyajit Das, Kevin J. M. Martin, Thomas Peyret, Philippe Coussy |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2020 | TRANSPIRE: An energy-efficient TRANSprecision floating-point Programmable archItectuREabstractIn recent years, Coarse Grain Reconfigurable Architecture (CGRA) accelerators have been increasingly deployed in Internet-of-Things (IoT) end nodes. A modern CGRA has to support and efficiently accelerate both integer and floating-point (FP) operations. In this paper, we propose an ultra-low-power tunable-precision CGRA architectural template, called TRANSprecision floating-point Programmable archItectuRE (TRANSPIRE), and its associated compilation flow supporting both integer and FP operations. TRANSPIRE employs transprecision computing and multiple Single Instruction Multiple Data (SIMD) to accelerate FP operations while boosting energy efficiency as well. Experimental results show that TRANSPIRE achieves a maximum of 10.06× performance gain and consumes 12.91× less energy w.r.t. a RISC-V based CPU with an enhanced ISA supporting SIMD-style vectorization and FP data-types, while executing applications for near-sensor computing and embedded machine learning, with an area overhead of 1.25× only. Rohit Prasad, Satyajit Das, Kevin J. M. Martin, Giuseppe Tagliavini, Philippe Coussy, Luca Benini, Davide Rossi 0001 |
DATE | 2 |
| 2020 | Energy Efficient Acceleration Of Floating Point Applications Onto CGRAabstractIn this paper, we propose a novel CGRA architecture and associated compilation flow supporting both integer and floating-point computations for energy efficient acceleration of DSP applications. Experimental results show that the proposed accelerator achieves a maximum of 4.61 × speedup compared to a DSP optimized, ultra low power RISC-V based CPU while executing seizure detection, a representative of wide range of EEG signal processing applications with an area overhead of 1.9×. The proposed CGRA achieves a maximum of 6.5× energy efficiency compared to the CPU. Satyajit Das, Rohit Prasad, Kevin J. M. Martin, Philippe Coussy |
ICASSP | 1 |
| 2019 | Context-memory Aware Mapping for Energy Efficient Acceleration with CGRAsabstractCoarse Grained Reconfigurable Arrays (CGRAs) are emerging as low power computing alternative providing a high grade of acceleration. However, the area and energy efficiency of these devices are bottlenecked by the configuration/context memory when they are made autonomous and loosely coupled with CPUs. The size of these context memories is of prime importance due to their high area and impact on the power consumption. For instance, a 64-word context memory typically represents 40% of a processing element area. In this context, since traditional mapping approaches do not take the size of the context memory into account, CGRAs often become oversized which strongly degrade their performance and interest. In this paper, we propose a context memory aware mapping for CGRAs to achieve better area and energy efficiency. This paper motivates the need of constraining the size of the context memory inside the processing element (PE) for ultra low power acceleration. It also describes the mapping approach which tries to find at least one mapping solution for a given set of constraints defined by the context memories of the PEs. Experiments show that our proposed solution achieves an average of 2.3× energy gain (with a maximum of 3.1× and a minimum of 1.4×) compared to the mapping approach without the memory constraints, while using 2× less context memory. When compared to the CPU, the proposed mapping achieves an average of 14× (with a maximum of 23× and minimum of 5×) energy gain. Satyajit Das, Kevin J. M. Martin, Philippe Coussy |
DATE | 1 |
| 2019 | On the orness of Bonferroni mean and its variantsabstractThis article addresses orness measures to reflect the or-like degree of the Bonferroni mean (BM) and its variants. Some properties of these operators associated with their orness measures are portrayed analytically. However, the general orness measure involves the multiple integrals with the integral fold number being the number of the aggregated elements and as a result, the computation becomes complicated when the number of the aggregated elements is large. Furthermore, the analytical formula of the orness measure often cannot be obtained. For this reason, this study concentrates on Monte Carlo simulation to validate the result. We estimate the two parameters of the BM for a predefined orness value and a fixed length of the input vector. Besides the theoretical study of orness measure related to BM and its variants, the article also explores the simulation-based results. To support this, we provide four numerical examples. Bapi Dutta, José Rui Figueira, Satyajit Das |
Int. J. Intell. Syst. | 3 |
| 2019 | Attribute weight computation in a decision making problem by particle swarm optimization
Satyajit Das, Debashree Guha |
Neural Comput. Appl. | 1 |
| 2019 | An Energy-Efficient Integrated Programmable Array Accelerator and Compilation Flow for Near-Sensor Ultralow Power ProcessingabstractIn this paper, we give a fresh look to coarse grained reconfigurable arrays (CGRAs) as ultralow power accelerators for near-sensor processing. We present a general-purpose integrated programmable-array accelerator (IPA) exploiting a novel architecture, execution model, and compilation flow for application mapping that can handle kernels containing complex control flow, without the significant energy overhead incurred by state of the art predication approaches. To optimize the performance and energy efficiency, we explore the IPA architecture with special focus on shared memory access, with the help of the flexible compilation flow presented in this paper. We achieve a maximum energy gain of 2×, and performance gain of 1.33× and 1.8× compared with state of the art partial and full predication techniques, respectively. The proposed accelerator achieves an average energy efficiency of 1617 MOPS/mW operating at 100 MHz, 0.6 V in 28 nm UTBB FD-SOI technology, over a wide range of near-sensor processing kernels, leading to an improvement up to 18×, with an average of 9.23× (as well as a speed-up up to 20.3×, with an average of 9.7×) compared to a core specialized for ultralow power near-sensor processing. Satyajit Das, Kevin J. M. Martin, Davide Rossi 0001, Philippe Coussy, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | A Heterogeneous Cluster with Reconfigurable Accelerator for Energy Efficient Near-Sensor Data AnalyticsabstractIoT end-nodes require high performance and extreme energy efficiency to cope with complex near-sensor data analytics algorithms. Processing on multiple programmable processors operating in near-threshold is emerging as a promising solution to exploit the energy boost given by low-voltage operation, while recovering the related frequency degradation with parallelism. In this work, we present a heterogeneous cluster architecture extending a traditional parallel processor cluster with a reconfigurable Integrated Programmable Array (IPA) accelerator. While programmable processors guarantee programming legacy to easily manage peripherals, radio software stacks as well as the global program flow, offloading data-intensive and control-intensive kernels to the IPA leads to much higher system level performance and energy-efficiency. Experimental results show that the proposed heterogeneous cluster outperforms an 8-core homogeneous architecture by up to 4.8× in performance and 4.5× in energy efficiency when executing a mix of control-intensive and data-intensive kernels typical of near-sensor data analytics applications. Satyajit Das, Kevin J. M. Martin, Philippe Coussy, Davide Rossi 0001 |
ISCAS | 1 |
| 2018 | Information Measures in the Intuitionistic Fuzzy Framework and Their RelationshipsabstractThis contribution first addresses axiomatic definitions of some important information measures of Atanassov's intuitionistic fuzzy sets (AIFSs), such as, distance measure, similarity measure, entropy measure, and recently developed knowledge measure. After that, we study the relationship between similarity measures and knowledge measures, and that between distance measures and knowledge measures for AIFSs. In particular, we investigate the transformations of similarity measures into knowledge measures for AIFSs and vice versa. These transformations can induce more formulae to define the distance measures, the similarity measures, and the knowledge measures for AIFSs. In addition, to support these processes, suitable examples are provided for each of the introduced transformations. The contribution ends by introducing numerical examples with a discussion of results by using several types of knowledge measures and similarity measures. Satyajit Das, Debashree Guha, Radko Mesiar |
IEEE Trans. Fuzzy Syst. | 1 |
| 2017 | Efficient mapping of CDFG onto coarse-grained reconfigurable array architecturesabstractIn the approaching era of IoT, flexible and low power accelerators have become essential to meet aggressive energy efficiency targets. During the last few decades, Coarse Grain Reconfigurable Arrays (CGRA) have demonstrated high energy efficiency as accelerators, especially for high-performance streaming applications. While existing CGRAs mostly rely on partial and full predication techniques to support conditional branches, inefficient architecture and mapping support for handling control flow limits the use of CGRAs in accelerating either only inner loop bodies, or transformed loops specifically adapted to the target CGRA. This paper proposes a novel CGRA architecture with support for jump and conditional jump instructions and a lightweight global synchronization mechanism to enable complete Control Data Flow Graph (CDFG) mapping in an ultra-low-power environment. The architecture is coupled with a complete design flow that efficiently maps applications with heavy control flow starting from a generic C language description. The proposed mapping approach reduces the impact of wasteful instruction issues in the conventional approaches of predication providing an average energy improvement of 1.44× and 1.6× when compared to the state of the art partial and full predication techniques. Moreover, the proposed method achieves an average speed-up up to 21× and an energy improvement up to 50.42× while executing applications with heavy control flow with respect to sequential execution on a low-power embedded CPU, demonstrating its suitability for next generation IoT applications. Satyajit Das, Kevin J. M. Martin, Philippe Coussy, Davide Rossi 0001, Luca Benini |
ASP-DAC | 1 |
| 2017 | A 142MOPS/mW integrated programmable array accelerator for smart visual processingabstractDue to increasing demand of low power computing, and diminishing returns from technology scaling, industry and academia are turning with renewed interest toward energy-efficient programmable accelerators. This paper proposes an Integrated Programmable-Array accelerator (IPA) architecture based on an innovative execution model, targeted to accelerate both data and control-flow parts of deeply embedded vision applications typical of edge-nodes of the Internet of Things (IoT). In this paper we demonstrate the performance and energy efficiency of IPA implementing a smart visual trigger application. Experimental results show that the proposed accelerator delivers 507 MOPS and 142 MOPS/mW on the target application, surpassing a low-power processor optimized for DSP applications by 6x in performance and by 10x in energy efficiency. Moreover, it surpasses performance of state of the art CGRAs only capable of implementing data-flow portion of applications by 1.6x, demonstrating the effectiveness of the proposed architecture and computational model. Satyajit Das, Davide Rossi 0001, Kevin J. M. Martin, Philippe Coussy, Luca Benini |
ISCAS | 1 |
| 2017 | Extended Bonferroni Mean Under Intuitionistic Fuzzy Environment Based on a Strict t-ConormabstractThe Bonferroni mean (BM) operator, originally introduced by Bonferroni, assumes homogeneous relationship among the input arguments. Recently, extended BM (EBM) operator is developed to capture heterogeneous relationship among real arguments and linguistic 2-tuple data. In this paper, we investigate the EBM operator with Atanassov's intuitionistic fuzzy sets (AIFSs) by using additive generators of strict t-conorms in the operations of AIFSs. This new operator is referred to as AIF-EBM operator. We also define weighted AIF-EBM (WAIF-EBM) operator by using the said operations. Moreover, we investigate several desirable properties of the proposed operators and we prove that some known specific intuitionistic fuzzy aggregation operators are special cases of the proposed AIF-EBM and WAIF-EBM operators. Subsequently, by utilizing the proposed operators, we develop a new approach for determining criteria weights, where criteria are heterogeneously interrelated. The contribution ends by introducing a numerical example with a comparative analysis of the proposed approach and the existing processes. The results of the numerical example are also analyzed with parameterized strict t-conorms. Satyajit Das, Debashree Guha, Radko Mesiar |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Medical diagnosis with the aid of using fuzzy logic and intuitionistic fuzzy logic
Satyajit Das, Debashree Guha, Bapi Dutta |
Appl. Intell. | 1 |
| 2016 | Weight computation of criteria in a decision-making problem by knowledge measure with intuitionistic fuzzy set and interval-valued intuitionistic fuzzy set
Satyajit Das, Bapi Dutta, Debashree Guha |
Soft Comput. | 1 |