EDBT 2026 Demo / reviewers in the wild / expert
Shuvra S. Bhattacharyya
dblp:b/ShuvraSBhattacharyya · also Shuvra Shikhar Bhattacharyya
· DBLP profile ↗
124ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0001-7719-1106ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorComputer networks · 4Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Theory of computation · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynPlay: Large-Scale Synthetic Human Data with Real-World Diversity for Aerial-View PerceptionabstractWe introduce SynPlay, a large-scale synthetic human dataset purpose-built for advancing multi-perspective human localization, with a predominant focus on aerial-view perception. SynPlay departs from traditional synthetic datasets by addressing a critical but underexplored challenge: localizing humans in aerial scenes where subjects often occupy only tens of pixels in the image. In such scenarios, fine-grained details like facial features or textures become irrelevant, shifting the burden of recognition to human motion, behavior, and interactions. To meet this need, SynPlay implements a novel rule-guided motion generation framework that combines real-world motion capture with motion evolution graphs. This design enables human actions to evolve dynamically through high-level game rules rather than predefined scripts, resulting in effectively uncountable motion variations. Unlike existing synthetic datasets—which either focus on static visual traits or reuse a limited set of mocap-driven actions—SynPlay captures a wide spectrum of spontaneous behaviors, including complex interactions that naturally emerge from unscripted gameplay scenarios. SynPlay also introduces an extensive multi-camera setup that spans UAVs at random altitudes, CCTVs, and a freely roaming UGV, achieving true near-to-far perspective coverage in a single dataset. The majority of instances are captured from aerial viewpoints at varying scales, directly supporting the development of models for long-range human analysis—a setting where existing datasets fall short. Our data contains over 73k images and 6.5M human instances, with detailed annotations for detection, segmentation, and keypoint tasks. Extensive experiments demonstrate that training with SynPlay significantly improves human localization performance, especially in few-shot and data-scarce scenarios. Jinsub Yim, Hyungtae Lee, Sungmin Eum, Yi-Ting Shen, Heesung Kwon, Shuvra S. Bhattacharyya |
WACV | 7 |
| 2025 | Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsabstractComposed pose retrieval (CPR) enables users to search for human poses by specifying a reference pose and a transition description, but progress in this field is hindered by the scarcity and inconsistency of annotated pose transitions. Existing CPR datasets rely on costly human annotations or heuristic-based rule generation, both of which limit scalability and diversity. In this work, we introduce AutoComPose, the first framework that leverages multimodal large language models (MLLMs) to automatically generate rich and structured pose transition descriptions. Our method enhances annotation quality by structuring transitions into fine-grained body part movements and introducing mirrored/swapped variations, while a cyclic consistency constraint ensures logical coherence between forward and reverse transitions. To advance CPR research, we construct and release two dedicated benchmarks, AIST-CPR and PoseFixCPR, supplementing prior datasets with enhanced attributes. Extensive experiments demonstrate that training retrieval models with AutoComPose yields superior performance over human-annotated and heuristic-based methods, significantly reducing annotation costs while improving retrieval quality. Our work pioneers the automatic annotation of pose transitions, establishing a scalable foundation for future CPR research. Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya |
ICCV | 7 |
| 2025 | Diversifying Human Pose In Synthetic Data For Aerial-View Human DetectionabstractSynthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances—particularly diverse poses—remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra’s algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size. Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
ICIP | 4 |
| 2025 | Towards Incorporating Social and Spatial Dependencies in Machine Learning Models for Crime PredictionabstractIn order to better allocate police resources, it is crucial to accurately predict crime risk across time and space. Considering crime prediction as a typical time series forecasting task (i.e., using crime history in a geographic space to predict future occurrences) has been shown to be effective. However, contextual features (e.g., socio-economic factors, characteristics of the built environment, etc.) can also be introduced to improve the accuracy of the predictions as they relate to the formation of crime. Since such contextual features are usually considered to be static relative to the time scales of crime history in a geographic space, we first propose a parallel branch model with dedicated branches for each type of data so that temporal crime history and time-invariant contextual features can be processed coherently. Then, we incorporate both spatial and social dependencies into the model, based on the observation that crime prediction can be enhanced by considering geographic spaces that are spatially adjacent as well as socially similar. The experimental results confirm the effectiveness of our proposed model. In particular, our model achieves state-of-the-art performance with an accuracy of 75.3% and an AUC-ROC (area under the receiver operating characteristic curve) of 0.79. Xiaowen Qi, Kiminori Nakamura, Shuvra S. Bhattacharyya |
SMC | 3 |
| 2025 | ShaderNN: A lightweight and efficient inference engine for real-time applications on mobile GPUsabstractInference using deep neural networks on mobile devices has been an active area of research in recent years. The design of a deep learning inference framework targeted for mobile devices needs to consider various factors, such as the limited computational capacity of the devices, low power budget, varied memory access methods, and I/O bus bandwidth governed by the underlying processor's architecture. Furthermore, integrating an inference framework with time-sensitive applications - such as games and video-based software to perform tasks like ray tracing denoising and video processing - introduces the need to minimize data movement between processors and increase data locality in the target processor. In this paper, we propose Shader Neural Network (ShaderNN), an OpenGL-based, fast, and power-efficient inference framework designed for mobile devices to address these challenges. Our contributions include the following: (1) the texture-based input/output provides an efficient, zero-copy integration with real-time graphics pipelines or image processing applications, thereby saving expensive data transfers between CPU and GPU, which are unavoidable in most existing inference engines; (2) we are the first to leverage fragment shaders based on the OpenGL backend in neural network inference operators, which has an advantage in deploying parametrically small neural network models; (3) a hybrid implementation of the compute shader and fragment shader is proposed that enables layer-level shader selection to boost performance; and (4) we utilize OpenGL features - such as normalization, interpolation and texture padding - to improve performance. Experiments illustrate the favorable performance of ShaderNN over other popular on-device deep learning frameworks such as TensorFlow-Lite on the latest mobile devices powered by Qualcomm and MediaTek chips. A case study further demonstrates the usability and integration of the ShaderNN framework with a media processing Android application seamlessly. ShaderNN is available open source at Github (https://github.com/inferenceengine/shadernn). Yuzhong Yan, Abhishek Saxena, Jiangong Chen, Rong Chen 0005, Shuvra S. Bhattacharyya |
Neurocomputing | 8 |
| 2025 | Real-time Fixed Priority Scheduling Synthesis Using Affine DataFlow Graphs: from Theory to PracticeabstractThe major drawback of using static schedules to execute dataflow applications is their high inflexibility. In real-time systems, periodic schedules make it easier to assert safety guarantees and to decrease the schedule size, but their characteristics remain hard to compute. This article presents an approach to automatically generate fixed priority schedules from a dataflow specification. To do so, precedence dependencies between actors in the dataflow graphs are abstracted, as well as the task periods, by using affine relations . This abstraction allows us to synthesize schedules efficiently considering two main objectives: the maximization of throughput and the minimization of buffer sizes. Given a dataflow graph to execute in a real-time environment, we transform it into an Affine Dataflow Graph (ADFG) and compute the task priorities, their mapping, the number of delays in the buffers, and the buffer sizes. This article is the first to present an overview of both theoretical and practical aspects of ADFG. On the theoretical side, it presents corrections and improvements on the fixed priority case. On the practical side, benchmark evaluations demonstrate the robustness and maturity of the approach that our scheduling synthesizer implements. Synthesized schedules are evaluated by using scheduling simulation and real-time implementation. Last but not least, the synthesized periods reach the optimal throughput if enough processors are available, and most of the time the periods reach the maximal processor utilization factor in the uni-processor case. Moreover, execution time of the synthesis is about only 1 second for the main proposed algorithms. Alexandre Honorat, Hai Nam Tran, Loïc Besnard, Shuvra S. Bhattacharyya, Jean-Pierre Talpin |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2024 | View DiffGait: View Pyramid Diffusion for Gait RecognitionabstractView transformation is crucial for gait recognition. Most existing methods use a view transformation models (VTM) or generative models (VAE or GAN) to achieve transformation. These approaches commonly adopt a paradigm of transforming a gait feature from one view to another. However, most existing methods attempt to use a single or multiple, large-view transformation model to directly transform a source view image to the target view image. Such transformations usually suffer from precision problems under large viewpoint variations due to the lack of fine view prediction. To overcome this challenge, we introduce a novel framework, ViewDiffGait, employing a view pyramid structure and diffusion models. ViewDiffGait is formulated as an iterative refinement generation task in a biologically interpretable way, capable of generating more accurate lateral view images from coarse to fine. Unlike the typical diffusion model that directly adds and removes Gaussian noise in the original image, the ViewDiffGait diffusion process involves a view pyramid structure to capture fine view transformations. The diffusion process adds view noise from the pyramid top to the bottom, while the denoising process removes view noise from the pyramid bottom to the top. We conducted extensive experiments on the CASIA-B and OUMVLP datasets, demonstrating that ViewDiffGait can generate more realistic images, remove variations effectively, and lead to high performance in real applications. Rijun Liao, Zhu Li 0001, Shuvra S. Bhattacharyya, George York |
FG | 3 |
| 2024 | Exploring the Potential of Synthetic Data to Replace Real DataabstractThe potential of synthetic data to replace real data creates a huge demand for synthetic data in data-hungry AI. This potential is even greater when synthetic data is used for training along with a small number of real images from domains other than the test domain. We find that this potential varies depending on (i) the number of cross-domain real images and (ii) the test set on which the trained model is evaluated. We introduce two new metrics, the train2test distance and $\mathrm{AP}_{\mathrm{t} 2 \mathrm{t}}$, to evaluate the ability of a cross-domain training set using synthetic data to represent the characteristics of test instances in relation to training performance. Using these metrics, we delve deeper into the factors that influence the potential of synthetic data and uncover some interesting dynamics about how synthetic data impacts training performance. We hope these discoveries will encourage more widespread use of synthetic data. Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
ICIP | 4 |
| 2024 | Balancing Fairness and Accuracy for Predictive Models in Criminal Justice Applications Using Multi-Objective Optimization MethodsabstractIn the field of predictive modeling for criminal justice applications, the dual challenges of ensuring fairness and maintaining interpretability are crucial. This paper addresses these challenges by introducing a new approach to optimizing decision trees using evolutionary algorithms (EAs). Our approach focuses on refining decision trees to achieve a balance between accuracy and algorithmic fairness, a task complicated by potential bias present in historical data. By leveraging the principles of multi-objective optimization, our model systematically trades off prediction accuracy and fair-ness. The evolutionary process characterized by selection, crossover, and mutation is tailored to fit the decision tree structure, ensuring that model development is not only accurate but also promoting measurable fairness. Experimental results demonstrate the effectiveness of our approach in providing interpretable and fair predictive models that can be considered for high-stakes applications in criminal justice. More broadly, this research makes a significant contribution to the field of explainable machine learning, providing a powerful framework for engineering systems that are transparent, fair, and adaptable to different data environments. Xiaowen Qi, Yujunrong Ma, Kiminori Nakamura, Shuvra S. Bhattacharyya |
SMC | 4 |
| 2024 | HashReID: Dynamic Network with Binary Codes for Efficient Person Re-identificationabstractBiometric applications, such as person re-identification (ReID), are often deployed on energy constrained devices. While recent ReID methods prioritize high retrieval performance, they often come with large computational costs and high search time, rendering them less practical in real-world settings. In this work, we propose an input-adaptive network with multiple exit blocks, that can terminate computation early if the retrieval is straightforward or noisy, saving a lot of computation. To assess the complexity of the input, we introduce a temporal-based classifier driven by a new training strategy. Furthermore, we adopt a binary hash code generation approach instead of relying on continuous-valued features, which significantly improves the search process by a factor of 20. To ensure similarity preservation, we utilize a new ranking regularizer that bridges the gap between continuous and binary features. Extensive analysis of our proposed method is conducted on three datasets: Market1501, MSMT17 (Multi-Scene Multi-Time), and the BGC1 (BRIAR Government Collection). Using our approach, more than 70% of the samples with compact hash codes exit early on the Market1501 dataset, saving 80% of the networks computational cost and improving over other hash-based methods by 60%. These results demonstrate a significant improvement over dynamic networks and showcase comparable accuracy performance to conventional ReID methods. Kshitij Nikhal, Yujunrong Ma, Shuvra S. Bhattacharyya, Benjamin S. Riggan |
WACV | 3 |
| 2024 | HoloCamera: Advanced Volumetric Capture for Cinematic-Quality VR ApplicationsabstractHigh-precision virtual environments are increasingly important for various education, simulation, training, performance, and entertainment applications. We present HoloCamera, an innovative volumetric capture instrument to rapidly acquire, process, and create cinematic-quality virtual avatars and scenarios. The HoloCamera consists of a custom-designed free-standing structure with 300 high-resolution RGB cameras mounted with uniform spacing spanning the four sides and the ceiling of a room-sized studio. The light field acquired from these cameras is streamed through a distributed array of GPUs that interleave the processing and transmission of 4K resolution images. The distributed compute infrastructure that powers these RGB cameras consists of 50 Jetson AGX Xavier boards, with each processing unit dedicated to driving and processing imagery from six cameras. A high-speed Gigabit Ethernet network fabric seamlessly interconnects all computing boards. In this systems paper, we provide an in-depth description of the steps involved and lessons learned in constructing such a cutting-edge volumetric capture facility that can be generalized to other such facilities. We delve into the techniques employed to achieve precise frame synchronization and spatial calibration of cameras, careful determination of angled camera mounts, image processing from the camera sensors, and the need for a resilient and robust network infrastructure. To advance the field of volumetric capture, we are releasing a high-fidelity static light-field dataset, which will serve as a benchmark for further research and applications of cinematic-quality volumetric light fields. Jonathan Heagerty, Shuvra S. Bhattacharyya, Sujal Bista, Barbara Brawn, Brandon Yushan Feng, Susmija Jabbireddy, Joseph F. JáJá, Hernisa Kacorri, David Li 0001, Derek Yarnell, Matthias Zwicker, Amitabh Varshney |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Progressive Transformation Learning for Leveraging Virtual Images in TrainingabstractTo effectively interrogate UAV-based images for detecting objects of interest, such as humans, it is essential to acquire large-scale UAV-based datasets that include human instances with various poses captured from widely varying viewing angles. As a viable alternative to laborious and costly data curation, we introduce Progressive Transformation Learning (PTL), which gradually augments a training dataset by adding transformed virtual images with enhanced realism. Generally, a virtual2real transformation generator in the conditional GAN framework suffers from quality degradation when a large domain gap exists between real and virtual images. To deal with the domain gap, PTL takes a novel approach that progressively iterates the following three steps: 1) select a subset from a pool of virtual images according to the domain gap, 2) transform the selected virtual images to enhance realism, and 3) add the transformed virtual images to the training set while removing them from the pool. In PTL, accurately quantifying the domain gap is critical. To do that, we theoretically demonstrate that the feature representation space of a given object detector can be modeled as a multivariate Gaussian distribution from which the Mahalanobis distance between a virtual object and the Gaussian distribution of each object category in the representation space can be readily computed. Experiments show that PTL results in a substantial performance increase over the baseline, especially in the small data and the cross-domain regime. Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya |
CVPR | 4 |
| 2023 | Towards Interpretable, Attention-Based Crime ForecastingabstractWhile the use of machine learning techniques in high stake fields, such as medical diagnosis and criminal justice, has been increasing in recent years, concerns have been raised regarding the lack of transparency and interpretability of the algorithms used. In this paper, we propose the use of interpretable attention-based ConvLSTM models for crime forecasting application. This approach combines the power of ConvLSTM models in capturing spatio-temporal patterns with the interpretability of attention mechanisms, allowing for the identification of key geographic areas in the input data that contribute to the prediction. We demonstrate the effectiveness of this approach through experiments on real-world crime data, showing that our model demonstrates high accuracy in crime predictions while providing insightful visualization that enhances the interpretability of prediction results. Yujunrong Ma, Xiaowen Qi, Kiminori Nakamura, Shuvra S. Bhattacharyya |
SMC | 4 |
| 2022 | EADTC: An Approach to Interpretable and Accurate Crime PredictionabstractMachine learning applications related to high-stakes decisions are often surrounded by significant amounts of controversy. This has led to increasing interest in interpretable machine learning models. A well-known class of interpretable models is that of decision trees (DTs), which mirror a common strategy used by humans to arrive at solutions through a series of well-defined decisions. However, much of previous research on DTs for criminal justice predictions has focused primarily on collections (ensembles) of DTs whose results are aggregated together. Such DT ensembles are used to help improve accuracy; however, their increased complexity and deviation from human decision-making processes makes them much less interpretable compared to single-DT approaches. In this paper, we present a new DT model for criminal recidivism prediction that is designed with high interpretability, accuracy, and fairness as core objectives. The interpretability of the model stems from its formulation in terms of a single DT structure, while accuracy is achieved through an intensive optimization process of DT parameters that is carried out using a novel evolutionary algorithm. Through extensive experiments, we analyze the performance of our proposed EADTC (Evolutionary Algorithm Decision Tree for Crime prediction) method on relevant datasets. Our experiments show that the EADTC approach achieves competitive accuracy and fairness with respect to state-of-the-art ensemble DT models, while achieving higher interpretability due to the simpler, single-DT structure. Yujunrong Ma, Kiminori Nakamura, Eungjoo Lee 0001, Shuvra S. Bhattacharyya |
SMC | 4 |
| 2022 | PoseMapGait: A model-based gait recognition method with pose estimation maps and graph convolutional networks
Rijun Liao, Zhu Li 0001, Shuvra S. Bhattacharyya, George York |
Neurocomputing | 3 |
| 2021 | Aerial Image Classification with Label Splitting and Optimized Triplet Loss LearningabstractWith the development of airplane platforms, aerial image classification plays an important role in a wide range of remote sensing applications. The number of most of aerial image dataset is very limited compared with other computer vision datasets. Unlike many works that use data augmentation to solve this problem, we adopt a novel strategy, called, label splitting, to deal with limited samples. Specifically, each sample has its original semantic label, we assign a new appearance label via unsupervised clustering for each sample by label splitting. Then an optimized triplet loss learning is applied to distill domain specific knowledge. This is achieved through a binary tree forest partitioning and triplets selection and optimization scheme that controls the triplet quality. Simulation results on NWPU, UCM and AID datasets demonstrate that proposed solution achieves the state-of-the-art performance in the aerial image classification. Rijun Liao, Zhu Li 0001, Shuvra S. Bhattacharyya, George York |
VCIP | 3 |
| 2021 | Plug-and-Play Deblurring for Robust Object DetectionabstractObject detection is a classic computer vision task, which learns the mapping between an image and object bounding boxes + class labels. Many applications of object detection involve images which are prone to degradation at capture time, notably motion blur from a moving camera like UAVs or object itself. One approach to handling this blur involves using common deblurring methods to recover the clean pixel images and then the apply vision task. This task is typically ill-posed. On top of this, application of these methods also add onto the inference time of the vision network, which can hinder performance of video inputs. To address the issues, we propose a novel plug-and-play (PnP) solution that insert deblurring features into the target vision task network without the need to retrain the task network. The deblur features are learned from a classification loss network on blur strength and directions, and the PnP scheme works well with the object detection network with minimum inference time complexity, compared with the state of the art deblur and then detection solution. Gerald Xie, Zhu Li 0001, Shuvra S. Bhattacharyya, Asif Mehmood |
VCIP | 3 |
| 2021 | Feature Extraction and Classification for Communication Channels in Wireless Mechatronic SystemsabstractFor accurate characterization and evaluation of wireless mechatronic systems, effective modeling of wireless communication channels is of paramount importance, especially to simulation-oriented methods. Conventional simulation methods employ mathematical models to abstract details of prototype channels. Although such mathematical models often have rigorous theoretical underpinnings, they can be weak in capturing complex environmental characteristics and complex forms of diversity that are exhibited in industrial communication environments. To address this problem, we develop, in this paper, a new approach to deriving effective simulation models for industrial communication channels. Our approach involves field measurements from actual wireless mechatronic environments together with feature extraction from the measurements, and data-driven classification based on the extracted features. Our approach leads to a general framework for simulating wireless mechatronic systems in a way that realistically incorporates the complex channel characteristics of these systems. Mohamed Kashef, Richard Candell, Yongkang Liu 0001, Karl Montgomery, Shuvra S. Bhattacharyya |
WFCS | 6 |
| 2021 | A novel view synthesis approach based on view space covering for gait recognition
Rijun Liao, Weizhi An, Zhu Li 0001, Shuvra S. Bhattacharyya |
Neurocomputing | 4 |
| 2021 | Hyperspectral Image Classification With Attention-Aided CNNsabstractConvolutional neural networks (CNNs) have been widely used for hyperspectral image classification. As a common process, small cubes are first cropped from the hyperspectral image and then fed into CNNs to extract spectral and spatial features. It is well known that different spectral bands and spatial positions in the cubes have different discriminative abilities. If fully explored, this prior information will help improve the learning capacity of CNNs. Along this direction, we propose an attention-aided CNN model for spectral-spatial classification of hyperspectral images. Specifically, a spectral attention subnetwork and a spatial attention subnetwork are proposed for spectral and spatial classifications, respectively. Both of them are based on the traditional CNN model and incorporate attention modules to aid networks that focus on more discriminative channels or positions. In the final classification phase, the spectral classification result and the spatial classification result are combined together via an adaptively weighted summation method. To evaluate the effectiveness of the proposed model, we conduct experiments on three standard hyperspectral data sets. The experimental results show that the proposed model can achieve superior performance compared with several state-of-the-art CNN-related models. Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Pedram Ghamisi, Shuvra S. Bhattacharyya |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Decidable Variable-Rate Dataflow for Heterogeneous Signal Processing SystemsabstractDynamic dataflow models of computation have become widely used through their adoption to popular programming frameworks such as TensorFlow and GNU Radio. Although dynamic dataflow models offer more programming freedom, they lack analyzability compared to their static counterparts (such as synchronous dataflow). In this paper we advocate the use of a boundedly dynamic dataflow model of computation, VR-PRUNE, that remains analyzable but still offers more programming freedom than a fully static dataflow model. The paper presents the VR-PRUNE model of computation and runtime, and illustrates its applicability to practical signal processing applications by two use cases: an adaptive convolutional neural network, and a predistortion filter for wireless communications. By runtime experiments on two heterogeneous computing platforms we show that VR-PRUNE is both flexible and efficient. Yujunrong Ma, Jiahao Wu 0001, Shuvra S. Bhattacharyya, Jani Boutellier |
ICASSP | 3 |
| 2020 | Prinet: A Prior Driven Spectral Super-Resolution NetworkabstractSpectral super-resolution aims to reconstruct hyperspectral images from RGB images directly. In recent years, convolutional networks have been successfully employed to this task. However, few of them take into account the specific properties of hyperspectral images. In this paper, we attempt to design a super-resolution network, named PriNET, based on two prior knowledge about hyperspectral images. The first one is spectral correlation. According to this property, we design a decomposition network to reconstruct hyperspectral images. In this network, the whole spectral bands of hyperspectral images are divided into several groups, and multiple residual networks are proposed to reconstruct them separately. The second knowledge is that the hyperspectral image should be able to generate its corresponding RGB image. Inspired from it, we design a self-supervised network to fine-tune the reconstruction results of the decomposition network. Finally, these two networks are combined together to constitute PriNET. Experimental results on two hyperspectral datasets demonstrate that the proposed PriNET can achieve better performance than several state-of-the-art networks. Renlong Hang, Zhu Li 0001, Qingshan Liu 0001, Shuvra S. Bhattacharyya |
ICME | 4 |
| 2020 | Efficient Model Solving for Markov Decision ProcessesabstractMarkov decision processes provide powerful tools for adaptive management of computing and communication in cyber-physical systems. However, efficient solvers are required to provide these capabilities on embedded computing platforms. This paper describes two new MDP solvers for embedded applications: Sparse Value Iteration (SVI) uses sparse matrix methods and runs on small, single-threaded CPU platforms; Sparse Parallel Value Iteration (SPVI) extends this approach to leverage the parallelism of embedded graphics processing units (GPUs) to further improve performance on more sophisticated embedded platforms. Both solvers improve running time and reduce power consumption. Adrian E. Sapio, Shuvra S. Bhattacharyya, Marilyn Wolf |
ISCC | 2 |
| 2020 | Integrating Field Measurements into a Model-Based Simulator for Industrial Communication NetworksabstractEfficient and accurate simulation methods are of increasing importance in the design and evaluation of factory communication systems. Model-based simulation methods are based on formal models that govern the interactions between components and subsystems in the systems that are being simulated. The formal models facilitate systematic integration across the system, and enable powerful methods for analysis and optimization of system performance. However conventional simulation approaches utilize communication channel models that do not fully reflect the characteristics and diversity of industrial communication channels. To help bridge this gap, we develop in this paper new methods for channel model construction for link-layer simulation that systematically incorporate field measurements of wireless communication channels from industrial networks, and derive corresponding channel modeling library components. The generated library components capture channel characteristics in the form of lookup tables, which can be flexibly integrated into system-level simulators or co-simulation tools. We integrate our new table-generation methods into a model-based co-simulator that jointly simulates the interactions among process flows, physical layouts of workcells, and communication channels in factory systems that are integrated with wireless networks. Experimental results using our lookup-table-augmented co-simulator demonstrate the utility of the proposed methods for flexibly and accurately integrating realistic industrial network channel conditions into simulation processes. Honglei Li 0005, Mohamed Kashef, Yongkang Liu 0001, Richard Candell, Shuvra S. Bhattacharyya |
WFCS | 6 |
| 2020 | Runtime Adaptation in Wireless Sensor Nodes Using Structured LearningabstractMarkov Decision Processes (MDPs) provide important capabilities for facilitating the dynamic adaptation and self-optimization of cyber physical systems at runtime. In recent years, this has primarily taken the form of Reinforcement Learning (RL) techniques that eliminate some MDP components for the purpose of reducing computational requirements. In this work, we show that recent advancements in Compact MDP Models (CMMs) provide sufficient cause to question this trend when designing wireless sensor network nodes. In this work, a novel CMM-based approach to designing self-aware wireless sensor nodes is presented and compared to Q-Learning, a popular RL technique. We show that a certain class of CPS nodes is not well served by RL methods and contrast RL versus CMM methods in this context. Through both simulation and a prototype implementation, we demonstrate that CMM methods can provide significantly better runtime adaptation performance relative to Q-Learning, with comparable resource requirements. Adrian E. Sapio, Shuvra S. Bhattacharyya, Marilyn Wolf |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2019 | Incremental Deep Neural Network Pruning Based on Hessian ApproximationabstractIn this paper, based on the Hessian approximation, an incremental pruning method is proposed to compress the deep neural network. The proposed method starts from the idea of using the Hessian to measure the "importance" of each weight in a deep neural network, and it mainly has the following key contributions. First, we propose to use the second moment in Adam optimizer as a measure of the "importance" of each weight to avoid calculating the Hessian matrix. Second, an incremental method is proposed to prune the neural network step by step. The incremental method can adjust the remaining non-zero weights of the whole network after each pruning to help boost the performance of the pruned network. Last but not least, the proposed method applies an automatically-generated global threshold for all the weights among all the layers, which achieves the inter-layer bit allocation automatically. Such a method can improve performance and save the complexity of adjusting the pruning threshold layer by layer. We perform a number of experiments on MNIST and ImageNet using commonly used neural networks such as AlexNet and VGG16 to show the benefits of the proposed algorithm. The experimental results show that the proposed algorithm is able to compress the network significantly with almost no loss of accuracy, which demonstrates the effectiveness of the proposed algorithm. Li Li 0040, Zhu Li 0001, Yue Li 0015, Birendra Kathariya, Shuvra S. Bhattacharyya |
DCC | 5 |
| 2019 | Gradient Image Super-resolution for Low-resolution Image RecognitionabstractIn visual object recognition problems essential to surveillance and navigation problems in a variety of military and civilian use cases, low-resolution and low-quality images present great challenges to this problem. Recent advancements in deep learning based methods like EDSR/VDSR have boosted pixel domain image super-resolution (SR) performances significantly in terms of signal to noise ratio(SNR)/ mean square error(MSE) metrics of the super-resolved image. However, these pixel domain signal quality metrics may not directly correlate to the machine vision tasks like key points detection and object recognition. In this work, we develop a machine vision tasks-friendly super-resolution technique which enhances the gradient images and associated features from the low-resolution images that benefit the high level machine vision tasks. Here, a residual learning deep neural network based gradient image super-resolution solution is developed with scale space adaptive network depth, and simulation results demonstrate the performance gains in both gradient image quality as well as key points repeatability. Dewan Fahim Noor, Yue Li 0015, Zhu Li 0001, Shuvra S. Bhattacharyya, George York |
ICASSP | 4 |
| 2019 | Multi-Frame Super Resolution with Deep Residual Learning on Flow Registered Non-Integer Pixel ImagesabstractSuper-Resolution (SR) of low-quality images is an important topic of research in image processing and computer vision field. Using multi-frame, super-resolution algorithm can reconstruct high-resolution images by incorporating the information of the subsequent images. Most of the super-resolution techniques for multi-frames either use a more traditional or mathematical approach or deep learning based approach with optical flow in consideration. In this paper, we develop a way to combine the optical flow enabled sub-pixel registration method for mapping into the high-resolution grid and a deep residual learning approach for restoring features with noise removal. The results exhibit a significant gain over the state of art methods and the bi-cubic interpolation method. Dewan Fahim Noor, Li Li 0040, Zhu Li 0001, Shuvra S. Bhattacharyya |
ICIP | 4 |
| 2019 | Low Resolution Recognition of Aerial ImagesabstractRemote sensing for classification has been widely studied and is useful for a lot of applications like precision agriculture, surveillance, and military applications. Recently, due to tremendous results achieved by deep learning using Convolutional Neural Networks (CNN) for Imagenet dataset, there have been a large number of works which use deep learning for aerial image classification. Most of the works concentrate on original resolution and there are no works on low-resolution recognition of aerial images. This work is critical because aerial images are taken from a very high distance from the ground and the cost of installing high definition cameras is high, so it is hard to get a high resolution of the image. In this paper, we explore how we can do the better classification of aerial images for original spatial resolution and low spatial resolution in deep learning by using texture information. In our framework, we use YUV color space which is generally used for video coding and we also use Laplacian of Gaussian (LOG) information to exploit the texture information. We decouple RGB information into luminance information (Y channel), color information (UV) and texture information (LOG) and we train a separate CNN for each feature and combine them using autoencoder and with our results, we show that we do better than RGB images in original resolution and low resolution. Raghunath Sai Puttagunta, Renlong Hang, Zhu Li 0001, Shuvra S. Bhattacharyya |
VCIP | 4 |
| 2019 | An integrated hardware/software design methodology for signal processing systemsabstractThis paper presents a new methodology for design and implementation of signal processing systems on system-on-chip (SoC) platforms. The methodology is centered on the use of lightweight application programming interfaces for applying principles of dataflow design at different layers of abstraction. The development processes integrated in our approach are software implementation, hardware implementation, hardware-software co-design, and optimized application mapping. The proposed methodology facilitates development and integration of signal processing hardware and software modules that involve heterogeneous programming languages and platforms. As a demonstration of the proposed design framework, we present a dataflow-based deep neural network (DNN) implementation for vehicle classification that is streamlined for real-time operation on embedded SoC devices. Using the proposed methodology, we apply and integrate a variety of dataflow graph optimizations that are important for efficient mapping of the DNN system into a resource constrained implementation that involves cooperating multicore CPUs and field-programmable gate array subsystems. Through experiments, we demonstrate the flexibility and effectiveness with which different design transformations can be applied and integrated across multiple scales of the targeted computing system. Lin Li 0029, Carlo Sau, Tiziana Fanni, Jingui Li, Timo Viitanen, François Christophe, Francesca Palumbo, Luigi Raffo, Heikki Huttunen, Jarmo Takala, Shuvra S. Bhattacharyya |
J. Syst. Archit. | 11 |
| 2019 | Multi-Scale Gradient Image Super-Resolution for Preserving SIFT Key Points in Low-Resolution Images
Dewan Fahim Noor, Yue Li 0015, Zhu Li 0001, Shuvra S. Bhattacharyya, George York |
Signal Process. Image Commun. | 4 |
| 2018 | A design tool for high performance image processing on multicore platformsabstractDesign and implementation of smart vision systems often involve the mapping of complex image processing algorithms into efficient, real-time implementations on multicore platforms. In this paper, we describe a novel design tool that is developed to address this important challenge. A key component of the tool is a new approach to hierarchical dataflow scheduling that integrates a global scheduler and multiple local schedulers. The local schedulers are lightweight modules that work independently. The global scheduler interacts with the local schedulers to optimize overall memory usage and execution time. The proposed design tool is demonstrated through a case study involving an image stitching application for large scale microscopy images. Jiahao Wu 0001, Timothy Blattner, Walid Keyrouz, Shuvra S. Bhattacharyya |
DATE | 4 |
| 2018 | A Joint Target Localization and Classification Framework for Sensor NetworksabstractIn this paper, we propose a joint framework for target localization and classification using a single generalized model for non-imaging based multi-modal sensor data. For target localization, we exploit both sensor data and estimated dynamics within a local neighborhood. We validate the capabilities of our framework by using a multi-modal dataset, which includes ground truth GPS information (e.g., time and position) and data from co-located seismic and acoustic sensors. Experimental results show that our framework achieves better classification accuracy compared to recent fusion algorithms using temporal accumulation and achieves more accurate target localizations than multilateration. Kyunghun Lee, Benjamin S. Riggan, Shuvra S. Bhattacharyya |
ICASSP | 3 |
| 2018 | Toward Efficient Many-core Scheduling of Partial Expansion GraphsabstractTransformation of synchronous data flow graphs (SDF) into equivalent homogeneous SDF representations has been extensively applied as a pre-processing stage when mapping signal processing algorithms onto parallel platforms. While this transformation helps fully expose task and data parallelism, it also presents several limitations such as an exponential increase in the number of actors and excessive communication overhead. Partial expansion graphs were introduced to address these limitations for multi-core platforms. However, existing solutions are not well-suited to achieve efficient scheduling on many-core architectures. In this article, we develop a new approach that employs cyclo-static data flow techniques to provide a simple but efficient method of coordinating the data production and consumption in the expanded graphs. We demonstrate the advantage of our approach through experiments on real application models. Hai Nam Tran, Shuvra S. Bhattacharyya, Jean-Pierre Talpin |
SCOPES | 2 |
| 2018 | Reproducible Evaluation of System Efficiency With a Model of Architecture: From Theory to PracticeabstractCurrent trends in high performance and embedded computing include design of increasingly complex hardware architectures with high parallelism, heterogeneous processing elements, and nonuniform communication resources. In order to take hardware and software design decisions, early evaluations of the system nonfunctional properties are needed. These evaluations of system efficiency require electronic system-level information on both algorithms and architecture. Contrary to algorithm models for which a major body of work has been conducted on defining formal models of computation (MoCs), architecture models from the literature are mostly empirical models from which reproducible experimentation requires the accompanying software. In this paper, a precise definition of a model of architecture (MoA) is proposed that focuses on reproducibility and abstraction and removes the overlap previously existing between the notions of MoA and MoC. A first MoA, called the linear system-level architecture model (LSLA), is presented. To demonstrate the generic nature of the proposed new architecture modeling concepts, we show that the LSLA model can be integrated flexibly with different MoCs. LSLA is then used to model the energy consumption of a state-of-the-art multiprocessor system-on-chip (MPSoC) when running an application described using the synchronous dataflow MoC. A method to automatically learn LSLA model parameters from platform measurements is introduced. Despite the high complexity of the underlying hardware and software, a simple LSLA model is demonstrated to estimate the energy consumption of the MPSoC with a fidelity of 86%. Maxime Pelcat, Alexandre Mercat, Karol Desnos, Luca Maggiani, Yanzhou Liu 0001, Julien Heulot, Jean-François Nezan, Wassim Hamidouche, Daniel Ménard, Shuvra S. Bhattacharyya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2018 | Memory-Constrained Vectorization and Scheduling of Dataflow Graphs for Hybrid CPU-GPU PlatformsabstractThe increasing use of heterogeneous embedded systems with multi-core CPUs and Graphics Processing Units (GPUs) presents important challenges in effectively exploiting pipeline, task, and data-level parallelism to meet throughput requirements of digital signal processing applications. Moreover, in the presence of system-level memory constraints, hand optimization of code to satisfy these requirements is inefficient and error prone and can therefore, greatly slow down development time or result in highly underutilized processing resources. In this article, we present vectorization and scheduling methods to effectively exploit multiple forms of parallelism for throughput optimization on hybrid CPU-GPU platforms, while conforming to system-level memory constraints. The methods operate on synchronous dataflow representations, which are widely used in the design of embedded systems for signal and information processing. We show that our novel methods can significantly improve system throughput compared to previous vectorization and scheduling approaches under the same memory constraints. In addition, we present a practical case-study of applying our methods to significantly improve the throughput of an orthogonal frequency division multiplexing receiver system for wireless communications. Shuoxin Lin, Jiahao Wu 0001, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Design and implementation of adaptive signal processing systems using Markov decision processesabstractIn this paper, we propose a novel framework, called Hierarchical MDP framework for Compact System-level Modeling (HMCSM), for design and implementation of adaptive embedded signal processing systems. The HMCSM framework applies Markov decision processes (MDPs) to enable autonomous adaptation of embedded signal processing under multidimensional constraints and optimization objectives. The framework integrates automated, MDP-based generation of optimal reconfiguration policies, dataflow-based application modeling, and implementation of embedded control software that carries out the generated reconfiguration policies. HMCSM systematically decomposes a complex, monolithic MDP into a set of separate MDPs that are connected hierarchically, and that operate more efficiently through such a modularized structure. We demonstrate the effectiveness of our new MDP-based system design framework through experiments with an adaptive wireless communications receiver. Lin Li 0029, Adrian E. Sapio, Jiahao Wu 0001, Yanzhou Liu 0001, Kyunghun Lee, Marilyn Wolf, Shuvra S. Bhattacharyya |
ASAP | 7 |
| 2017 | An accumulative fusion architecture for discriminating people and vehicles using acoustic and seismic signalsabstractIn this paper, we develop new multiclass classification algorithms for detecting people and vehicles by fusing data from a multimodal, unattended ground sensor node. The specific types of sensors that we apply in this work are acoustic and seismic sensors. We investigate two alternative approaches to multiclass classification in this context - the first is based on applying Dempster-Shafer Theory to perform score-level fusion, and the second involves the accumulation of local similarity evidences derived from a feature-level fusion model that combines both modalities. We experiment with the proposed algorithms using different datasets obtained from acoustic and seismic sensors in various outdoor environments, and evaluate the performance of the two algorithms in terms of receiver operating characteristic and classification accuracy. Our results demonstrate overall superiority of the proposed new feature-level fusion approach for multiclass discrimination among people, vehicles and noise. Kyunghun Lee, Benjamin S. Riggan, Shuvra S. Bhattacharyya |
ICASSP | 3 |
| 2017 | Hardware design methodology using lightweight dataflow and its integration with low power techniques
Tiziana Fanni, Lin Li 0029, Timo Viitanen, Carlo Sau, Renjie Xie, Francesca Palumbo, Luigi Raffo, Heikki Huttunen, Jarmo Takala, Shuvra S. Bhattacharyya |
J. Syst. Archit. | 10 |
| 2016 | Design space exploration and constrained multiobjective optimization for digital predistortion systemsabstractIn this paper, we develop new models and methods for exploring multidimensional design spaces associated with digital predistortion (DPD) systems. DPD systems are important components for power amplifier linearization in wireless communication transceivers. In contrast to conventional DPD implementation methods, which are focused on optimizing a single objective — most commonly, the adjacent channel power ratio (ACPR) — without systematically taking into account other relevant metrics, we consider DPD system implementation in a multiobjective optimization context. In our targeted multiobjective context, trade-offs among power consumption and multiple DPD performance metrics are jointly optimized subject to performance constraints imposed by the given modulation scheme. Through synthesis and simulation results, we demonstrate that DPD systems derived through our design space exploration techniques exhibit significantly improved trade-offs among multidimensional implementation criteria, including energy consumption, ACPR, and symbol error-rate. Additionally, we perform experiments using three different LTE modulation schemes, and we demonstrate that our multiobjective optimization approach significantly enhances system adaptivity in response to changes in the employed modulation scheme. Lin Li 0029, Amanullah Ghazi, Jani Boutellier, Lauri Anttila, Mikko Valkama, Shuvra S. Bhattacharyya |
ASAP | 6 |
| 2016 | A Design Framework for Mapping Vectorized Synchronous Dataflow Graphs onto CPU-GPU PlatformsabstractHeterogeneous computing platforms with multicore central processing units (CPUs) and graphics processing units (GPUs) are of increasing interest to designers of embedded signal processing systems since they offer the potential for significant performance boost while maintaining the flexibility of software-based design flows. Developing optimized implementations for CPU-GPU platforms is challenging due to complex, inter-related design issues, including task scheduling, interprocessor communication, memory management, and modeling and exploitation of different forms of parallelism. In this paper, we present an automated, dataflow based, design framework called DIF-GPU for application mapping and software synthesis on heterogeneous CPU-GPU platforms. DIF-GPU is based on novel extensions to the dataflow interchange format (DIF) package, which is a software environment for developing and experimenting with dataflow-based design methods and synthesis techniques for embedded signal processing systems. DIF-GPU exploits multiple forms of parallelism by deeply incorporating efficient vectorization and scheduling techniques for synchronous dataflow specifications, and incorporating techniques for streamlining interprocessor communication. DIF-GPU also provides software synthesis capabilities to help accelerate the process of moving from high-level application models to optimized implementations. Shuoxin Lin, Yanzhou Liu 0001, William Plishker, Shuvra S. Bhattacharyya |
SCOPES | 4 |
| 2014 | Low power implementation of digital predistortion filter on a heterogeneous application specific multiprocessorabstractPower-constrained mobile radio communication transmitters drive their transmit power amplifiers close to their saturation regions, which results in nonlinear intermodulation distortion that is especially harmful in multi-cluster and carrier aggregation transmission scenarios. Digital predistortion is a method for linearizing the transmitter and suppressing the most harmful spurious emissions at the transmitter power amplifier output. This paper describes a programmable implementation of a digital predistortion filter on a heterogeneous Transport Trigger Architecture (TTA) multiprocessor. The predistortion algorithm is based on a parallel Hammerstein polynomial model and the experimental results show that the proposed programmable architecture is capable of linearizing a 20 MHz LTE carrier in realtime with a power consumption that is suitable for mobile devices. Amanullah Ghazi, Jani Boutellier, Mahmoud Abdelaziz, Xiaojia Lu, Lauri Anttila, Joseph R. Cavallaro, Shuvra S. Bhattacharyya, Mikko Valkama, Markku Juntti |
ICASSP | 7 |
| 2014 | Efficient architecture mapping of FFT/IFFT for cognitive radio networksabstractCognitive radio networks require flexibility to support a variety of wireless communication system standards. Many modern systems utilize some form of orthogonal frequency division multiplexing (OFDM) and single-carrier frequency-division multiple access (SC-FDMA) often augmented with multiple input multiple output (MIMO) antenna schemes. A common module in these standards is the fast Fourier transform (FFT) and its inverse. Although many architectures exist for traditional power-of-two FFT lengths, the recent 3GPP LTE standards define non-power-of-two transform lengths. The various FFT and IFFT lengths for both the uplink and downlink processing require support for radix-2, radix-3, and radix-5 modules. In this paper, we propose a highly flexible FFT/IFFT architecture that can support a broad variety of transform sizes and efficient mapping to programmable testbed platforms for cognitive radio networks. This novel architecture will provide a range of transform sizes of the general form (2n3k5l), and for use in emerging algorithms for massive MIMO detectors. Bei Yin, Inkeun Cho, Joseph R. Cavallaro, Shuvra S. Bhattacharyya, Jarmo Takala |
ICASSP | 5 |
| 2013 | Configurable, resource-optimized FFT architecture for OFDM communicationabstractIn this paper, we present a designer-configurable, resource efficient FPGA architecture for OFDM system implementation. Our design achieves a significant improvement in resource efficiency for a given data rate. This efficiency improvement is achieved through careful analysis of how FFT computation is performed within the context of OFDM systems, and streamlining memory management and control logic based on this analysis. In particular, our OFDM-targeted FFT design eliminates redundant buffer memory, and simplifies control logic to save FPGA resources. We have synthesized and tested our design using the Xilinx ISE 13.4 synthesis tool, and compared the results with the Xilinx FFT v7.1, which is a widely used commercial FPGA IP core. We have demonstrated that our design provides at least 8.8% enhancement in terms of resource efficiency compared to Xilinx FFT v7.1 when it is embedded within the same OFDM configuration. Inkeun Cho, Chung-Ching Shen, Yahia Tachwali, Chia-Jui Hsu, Shuvra S. Bhattacharyya |
ICASSP | 5 |
| 2013 | A novel framework for design and implementation of adaptive stream mining systemsabstractWith the increasing need for accurate mining and classification from multimedia data content, and the growth of such multimedia applications in mobile and distributed architectures, stream mining systems require increasing amounts of flexibility, extensibility, and adaptivity for effective deployment. To address this challenge, we propose a novel approach that rigorously integrates foundations of dataflow modeling for high level signal processing system design, and adaptive stream mining based on dynamic topologies of classifiers. In particular, we introduce a new design environment, called the lightweight dataflow for dynamic data driven application systems (LiD4E) environment. LiD4E provides formal semantics, rooted in dataflow principles, for design and implementation of a broad class of multimedia stream mining topologies. We demonstrate the capabilities of LiD4E using a face detection application that systematically adapts the type of classifier used based on dynamically changing application constraints. Kishan Sudusinghe, Stephen Won, Mihaela van der Schaar, Shuvra S. Bhattacharyya |
ICME | 4 |
| 2013 | High-performance and low-energy buffer mapping method for multiprocessor DSP systemsabstractWhen implementing digital signal processing (DSP) applications onto multiprocessor systems, one significant problem in the viewpoints of performance is the memory wall. In this paper, to help alleviate the memory wall problem, we propose a novel, high-performance buffer mapping policy for SDF-represented DSP applications on bus-based multiprocessor systems that support the shared-memory programming model. The proposed policy exploits the bank concurrency of the DRAM main memory system according to the analysis of hierarchical parallelism. Energy consumption is also a critical parameter, especially in battery-based embedded computing systems. In this paper, we apply a synchronization back-off scheme on the top of the proposed high-performance buffer mapping policy to reduce energy consumption. The energy saving is attained by minimizing the number of non-essential synchronization transactions. We measure throughput and energy consumption on both synthetic and real benchmarks. The simulation results show that the proposed buffer mapping policy is very useful in terms of performance, especially in memory-intensive applications where the total execution time of computational tasks is relatively small compared to that of memory operations. In addition, the proposed synchronization back-off scheme provides a reduction in the number of synchronization transactions without degrading performance, which results in system energy saving. Dongwon Lee 0003, Marilyn Wolf, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | Partial Expansion Graphs: Exposing Parallelism and Dynamic Scheduling Opportunities for DSP ApplicationsabstractEmerging Digital Signal Processing (DSP) algorithms and wireless communications protocols require dynamic adaptation and online reconfiguration for the implemented systems at runtime. In this paper, we introduce the concept of Partial Expansion Graphs (PEGs) as an implementation model and associated class of scheduling strategies. PEGs are designed to help realize DSP systems in terms of forms and granularities of parallelism that are well matched to the given applications and targeted platforms. PEGs also facilitate derivation of both static and dynamic scheduling techniques,depending on the amount of variability in task execution times and other operating conditions. We show how to implement efficient PEG-based scheduling methods using real time operating systems, and to re-use pre-optimized libraries of DSP components within such implementations. Empirical results show that the PEG strategy can 1) achieve significant speedups on a state of the art multicore signal processor platform for static dataflow applications with predictable execution times,and 2) exceed classical scheduling speedups for application shaving execution times that can vary dynamically. This ability to handle variable execution times is especially useful as DSP applications and platforms increase in complexity and adaptive behavior, thereby reducing execution time predictability. George F. Zaki, William Plishker, Shuvra S. Bhattacharyya, Frank Fruth |
ASAP | 3 |
| 2012 | Parameterized scheduling for signal processing systems using topological patternsabstractIn recent work, a graphical modeling construct called “topological patterns” has been shown to enable concise representation and direct analysis of repetitive dataflow graph sub-structures in the context of design methods and tools for digital signal processing systems. In this paper, we present a formal design method for specifying topological patterns and deriving parameterized schedules from such patterns based on a novel schedule model called the scalable schedule tree. The approach represents an important class of parameterized schedule structures in a form that is intuitive for representation and efficient for code generation. We demonstrate our methods for topological pattern representation, scalable schedule tree derivation, and associated dataflow graph code generation using a case study for image processing. Shenpei Wu, Chung-Ching Shen, Nimish Sane, Kelly Davis, Shuvra S. Bhattacharyya |
ICASSP | 5 |
| 2012 | Design and Synthesis for Multimedia Systems Using the Targeted Dataflow Interchange FormatabstractDevelopment of multimedia systems that can be targeted to different platforms is challenging due to the need for rigorous integration between high-level abstract modeling, and low-level synthesis and optimization. In this paper, a new dataflow-based design tool called the targeted dataflow interchange format is introduced for retargetable design, analysis, and implementation of embedded software for multimedia systems. Our approach provides novel capabilities, based on principles of task-level dataflow analysis, for exploring and optimizing interactions across design components; object-oriented data structures for encapsulating contextual information for components; a novel model for representing parameterized schedules that are derived from repetitive graph structures; and automated code generation for programming interfaces and low-level customizations that are geared toward high-performance embedded-processing architectures. We demonstrate our design tool for cross-platform application design, parameterized schedule representation, and associated dataflow graph-code generation using a case study centered around an image registration application. Chung-Ching Shen, Shenpei Wu, Nimish Sane, Hsiang-Huang Wu, William Plishker, Shuvra S. Bhattacharyya |
IEEE Trans. Multim. | 6 |
| 2011 | Modeling and optimization of dynamic signal processing in resource-aware sensor networksabstractSensor node processing in resource-aware sensor networks is often critically dependent on dynamic signal processing functionality - i.e., signal processing functionality in which computational structure must be dynamically assessed and adapted based on time-varying environmental conditions, operating constraints or application requirements. In dynamic signal processing systems, it is important to provide flexibility for run-time adaptation of application behavior and execution characteristics, but in the domain of resource-aware sensor networks, such flexibility cannot come with significant costs in terms of power consumption overhead or reduced predictability. In this paper, we review a variety of complementary models of computation that are being developed as part of the dataflow interchange format (DIF) project to facilitate efficient and reliable implementation of dynamic signal processing systems. We demonstrate these methods in the context of resource-aware sensor networks. Shuvra S. Bhattacharyya, William Plishker, Nimish Sane, Chung-Ching Shen, Hsiang-Huang Wu |
AVSS | 1 |
| 2011 | A design tool for efficient mapping of multimedia applications onto heterogeneous platformsabstractDevelopment of multimedia systems on heterogeneous platforms is a challenging task with existing design tools due to a lack of rigorous integration between high level abstract modeling, and low level synthesis and analysis. In this paper, we present a new dataflow-based design tool, called the targeted dataflow interchange format (TDIF), for design, analysis, and implementation of embedded software for multimedia systems. Our approach provides novel capabilities, based on the principles of task-level dataflow analysis, for exploring and optimizing interactions across application behavior; operational context; heterogeneous platforms, including high performance embedded processing architectures; and implementation constraints. Chung-Ching Shen, Hsiang-Huang Wu, Nimish Sane, William Plishker, Shuvra S. Bhattacharyya |
ICME | 5 |
| 2011 | Design methods for Wireless Sensor Network Building Energy Monitoring SystemsabstractIn this paper, we present a new energy analysis method for evaluating energy consumption of embedded sensor nodes at the application level and the network level. Then we apply the proposed energy analysis method to develop new energy management schemes in order to maximize lifetime for Wireless Sensor Network Building Energy Monitoring Systems (WSNBEMS). At the application level, we develop a new design approach that uses dataflow techniques to model the application-level interfacing behavior between the processor and sensors on an embedded sensor node. At the network level, we analyze the energy consumption of the IEEE 802.15.4 MAC functionality. Based on our techniques for modeling and energy analysis, we have implemented an optimized WSNBEMS for a real building, and validated our energy analysis techniques through measurements on this implementation. The performance of our implementation is also evaluated in terms of monitoring accuracy and energy consumption savings. We have demonstrated that by applying the proposed scheme, system lifetime can be improved significantly without affecting monitoring accuracy. Inkeun Cho, Chung-Ching Shen, Siddharth Potbhare, Shuvra S. Bhattacharyya, Neil Goldsman |
LCN | 4 |
| 2011 | Multithreaded Simulation for Synchronous Dataflow GraphsabstractFor system simulation, Synchronous DataFlow (SDF) has been widely used as a core model of computation in design tools for digital communication and signal processing systems. The traditional approach for simulating SDF graphs is to compute and execute static schedules in single-processor desktop environments. Nowadays, however, multicore processors are increasingly popular desktop platforms for their potential performance improvements through thread-level parallelism. Without novel scheduling and simulation techniques that explicitly explore thread-level parallelism for executing SDF graphs, current design tools gain only minimal performance improvements on multicore platforms. In this article, we present a new multithreaded simulation scheduler, called MSS, to provide simulation runtime speedup for executing SDF graphs on multicore processors. MSS strategically integrates graph clustering, intracluster scheduling, actor vectorization, and intercluster buffering techniques to construct InterThread Communication (ITC) graphs at compile-time. MSS then applies efficient synchronization and dynamic scheduling techniques at runtime for executing ITC graphs in multithreaded environments. We have implemented MSS in the Advanced Design System (ADS) from Agilent Technologies. On an Intel dual-core, hyper-threading (4 processing units) processor, our results from this implementation demonstrate up to 3.5 times speedup in simulating modern wireless communication systems (e.g., WCDMA3G, CDMA 2000, WiMax, EDGE, and Digital TV). Chia-Jui Hsu, José Luis Pino, Shuvra S. Bhattacharyya |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2010 | Loop transformations for interface-based hierarchies IN SDF graphsabstractData-flow has proven to be an attractive computation model for programming digital signal processing (DSP) applications. A restricted version of data-flow, termed synchronous data-flow (SDF), offers strong compile-time predictability properties, but has limited expressive power. A new type of hierarchy (Interface-based SDF) has been proposed allowing more expressivity while maintaining its predictability. One of the main problems with this hierarchical SDF model is the lack of trade-off between parallelism and network clustering. This paper presents a systematic method for applying an important class of loop transformation techniques in the context of interface-based SDF semantics. The resulting approach provides novel capabilities for integrating parallelism extraction properties of the targeted loop transformations with the useful modeling, analysis, and code reuse properties provided by SDF. Jonathan Piat, Shuvra S. Bhattacharyya, Mickaël Raulet |
ASAP | 2 |
| 2010 | FPGA-based design and implementation of the 3GPP-LTE physical layer using parameterized synchronous dataflow techniquesabstractSynchronous dataflow (SDF) is an ubiquitous dataflow model of computation that has been studied extensively for efficient simulation and software synthesis of DSP applications. In recent years, parameterized SDF (PSDF) has evolved as a useful framework for modeling SDF graphs in which arbitrary parameters can be changed dynamically. However, the potential to enable efficient hardware synthesis has been treated relatively sparsely in the literature for SDF and even more so for the newer, more general PSDF model. This paper investigates efficient FPGA-based design and implementation of the physical layer for 3GPP-Long Term Evolution (LTE), a next generation cellular standard. To capture the SDF behavior of the functional core of LTE along with higher level dynamics in the standard, we use a novel PSDF-based FPGA architecture framework. We implement our PSDF-based, LTE design framework using National Instrument's LabVIEW FPGA, a recently-introduced commercial platform for reconfigurable hardware implementation. We show that our framework can effectively model the dynamics of the LTE protocol, while also providing a synthesis framework for efficient FPGA implementation. Hojin Kee, Shuvra S. Bhattacharyya, Ian C. Wong, Yong Rao |
ICASSP | 2 |
| 2010 | Buffer management for multi-application image processing on multi-core platforms: Analysis and case studyabstractDue to the limited amounts of on-chip memory, large volumes of data, and performance and power consumption overhead associated with interprocessor communication, efficient management of buffer memory is critical to multi-core image processing. To address this problem, this paper develops new modeling and analysis techniques based on dataflow representations, and demonstrates these techniques on a multi-core implementation case study involving multiple, concurrently-executing image processing applications. Our techniques are based on careful representation and exploitation of frame- or block-based operations, which involve repeated invocations of the same computations across regularly- arranged subsets of data. Using these new approaches to manage block-based image data, this paper demonstrates methods to analyze synchronization overhead and FIFO buffer sizes when mapping image processing applications onto heterogeneous, multi core architectures. Dong-Ik Ko, Nara Won, Shuvra S. Bhattacharyya |
ICASSP | 3 |
| 2010 | Simulating dynamic communication systems using the core functional dataflow modelabstractThe latest communication technologies invariably consist of modules with dynamic behavior. There exists a number of design tools for communication system design with their foundation in dataflow modeling semantics. These tools must not only support the functional specification of dynamic communication modules and subsystems but also provide accurate estimation of resource requirements for efficient simulation and implementation. We explore this trade-off - between flexible specification of dynamic behavior and accurate estimation of resource requirements - using a representative application employing an adaptive modulation scheme. We propose an approach for precise modeling of such applications based on a recently-introduced form of dynamic dataflow called core functional dataflow. From our proposed modeling approach, we show how parameterized looped schedules can be generated and analyzed to simulate applications with low run-time overhead as well as guaranteed bounded memory execution. We demonstrate our approach using the Advanced Design System from Agilent Technologies, Inc., which is a commercial tool for design and simulation of communication systems. Nimish Sane, Chia-Jui Hsu, José Luis Pino, Shuvra S. Bhattacharyya |
ICASSP | 4 |
| 2010 | Design and implementation of embedded computer vision systems based on particle filters
Sankalita Saha, Neal K. Bambha, Shuvra S. Bhattacharyya |
Comput. Vis. Image Underst. | 3 |
| 2010 | Analysis of SystemC actor networks for efficient synthesisabstractApplications in the signal processing domain are often modeled by dataflow graphs. Due to heterogeneous complexity requirements, these graphs contain both dynamic and static dataflow actors. In previous work, we presented a generalized clustering approach for these heterogeneous dataflow graphs in the presence of unbounded buffers. This clustering approach allows the application of static scheduling methodologies for static parts of an application during embedded software generation for multiprocessor systems. It systematically exploits the predictability and efficiency of the static dataflow model to obtain latency and throughput improvements. In this article, we present a generalization of this clustering technique to dataflow graphs with bounded buffers, therefore enabling synthesis for embedded systems without dynamic memory allocation. Furthermore, a case study is given to demonstrate the performance benefits of the approach. Joachim Falk, Christian Zebelein, Joachim Keinert, Christian Haubelt, Jürgen Teich, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2010 | Energy-driven distribution of signal processing applications across wireless sensor networksabstractWireless sensor network (WSN) applications have been studied extensively in recent years. Such applications involve resource-limited embedded sensor nodes that have small size and low power requirements. Based on the need for extended network lifetimes in WSNs in terms of energy use, the energy efficiency of computation and communication operations in the sensor nodes becomes critical. Digital Signal Processing (DSP) applications typically require intensive data processing operations and as a result are difficult to implement directly in resource-limited WSNs. In this article, we present a novel design methodology for modeling and implementing computationally intensive DSP applications applied to wireless sensor networks. This methodology explores efficient modeling techniques for DSP applications, including data sensing and processing; derives formulations of Energy-Driven Partitioning (EDP) for distributing such applications across wireless sensor networks; and develops efficient heuristic algorithms for finding partitioning results that maximize the network lifetime. To address such an energy-driven partitioning problem, this article provides a new way of aggregating data and reducing communication traffic among nodes based on application analysis. By considering low data token delivery points and the distribution of computation in the application, our approach finds energy-efficient trade-offs between data communication and computation. Chung-Ching Shen, William Plishker, Dong-Ik Ko, Shuvra S. Bhattacharyya, Neil Goldsman |
ACM Trans. Sens. Networks | 4 |
| 2009 | Mode grouping for more effective generalized scheduling of dynamic dataflow applicationsabstractFor a number of years, dataflow concepts have provided designers of digital signal processing systems with environments capable of expressing high-level software architectures as well as low-level, performance-oriented kernels. To apply these proven techniques to new complex, dynamic applications, we identify repetitive sequences of atomic, repeatable actions (“modes”) inside dynamic actors to expose more of the static nature of the application. In this work, we propose a mode grouping strategy that aids in the decomposition of a dynamic dataflow graph into a set of static dataflow graphs that interact dynamically. Mode grouping enables the discovery of larger static subgraphs improving scheduling results. We show that grouping modes results in improved schedules with lower memory requirements for implementations by up to 37 % including a common imaging benchmark with dynamic behavior: 3D B-spline interpolation. William Plishker, Nimish Sane, Shuvra S. Bhattacharyya |
DAC | 3 |
| 2009 | A generalized scheduling approach for dynamic dataflow applicationsabstractFor a number of years, dataflow concepts have provided designers of digital signal processing systems with environments capable of expressing high-level software architectures as well as low-level, performance-oriented kernels. But analysis of system-level trade-offs has been inhibited by the diversity of models and the dynamic nature of modern dataflow applications. To facilitate design space exploration for software implementations of heterogeneous dataflow applications, developers need tools capable of deeply analyzing and optimizing the application. To this end, we present a new scheduling approach that leverages a recently proposed general model of dynamic dataflow called core functional dataflow (CFDF). CFDF supports high-level application descriptions with multiple models of dataflow by structuring actors with sets of modes that represent fixed behaviors. In this work we show that by decomposing a dynamic dataflow graph as directed by its modes, we can derive a set of static dataflow graphs that interact dynamically. This enables designers to readily experiment with existing dataflow model specific scheduling techniques to all or some parts of the application while applying custom schedulers to others. We demonstrate this generalized dataflow scheduling method on dynamic mixed-model applications and show that run-time and buffer sizes significantly improve compared to a baseline dynamic dataflow scheduler and simulator. William Plishker, Nimish Sane, Shuvra S. Bhattacharyya |
DATE | 3 |
| 2009 | Exploiting statically schedulable regions in dataflow programsabstractDataflow descriptions have been used in a wide range of Digital Signal Processing (DSP) applications, such as multi-media processing, and wireless communications. Among various forms of dataflow modeling, Synchronous Dataflow (SDF) is geared towards static scheduling of computational modules, which improves system performance and predictability. However, many DSP applications do not fully conform to the restrictions of SDF modeling. More general dataflow models, such as CAL, have been developed to describe dynamically-structured DSP applications. Such generalized models can express dynamically changing functionality, but lose the powerful static scheduling capabilities provided by SDF. This paper focuses on detection of SDF-like regions in dynamic dataflow descriptions - in particular, in the generalized specification framework of CAL. This is an important step for applying static scheduling techniques within a dynamic dataflow framework. Our techniques combine the advantages of different dataflow languages and tools, including CAL, DIF and CAL2C. The techniques are demonstrated on the IDCT module of MPEG Reconfigurable Video Coding (RVC). Ruirui Gu, Jörn W. Janneck, Mickaël Raulet, Shuvra S. Bhattacharyya |
ICASSP | 4 |
| 2009 | Exploring the Concurrency of an MPEG RVC Decoder Based on Dataflow Program AnalysisabstractThis paper presents an in-depth case study on dataflow-based analysis and exploitation of parallelism in the design and implementation of a MPEG reconfigurable video coding decoder. Dataflow descriptions have been used in a wide range of digital signal processing (DSP) applications, such as applications for multimedia processing and wireless communications. Because dataflow models are effective in exposing concurrency and other important forms of high level application structure, dataflow techniques are promising for implementing complex DSP applications on multicore systems, and other kinds of parallel processing platforms. In this paper, we use the client access license (CAL) language as a concrete framework for representing and demonstrating dataflow design techniques. Furthermore, we also describe our application of the differential item functioning dataflow interchange format package (TDP), a software tool for analyzing dataflow networks, to the systematic exploitation of concurrency in CAL networks that are targeted to multicore platforms. Using TDP, one is able to automatically process regions that are extracted from the original network, and exhibit properties similar to synchronous dataflow (SDF) models. This is important in our context because powerful techniques, based on static scheduling, are available for exploiting concurrency in SDF descriptions. Detection of SDF-like regions is an important step for applying static scheduling techniques within a dynamic dataflow framework. Furthermore, segmenting a system into SDF-like regions also allows us to explore cross-actor concurrency that results from dynamic dependences among different regions. Using SDF-like region detection as a preprocessing step to software synthesis generally provides an efficient way for mapping tasks to multicore systems, and improves the system performance of video processing applications on multicore platforms. Ruirui Gu, Jörn W. Janneck, Shuvra S. Bhattacharyya, Mickaël Raulet, Matthieu Wipliez, William Plishker |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Multithreaded simulation for synchronous dataflow graphsabstractSynchronous dataflow (SDF) has been successfully used in design tools for system-level simulation of wireless communication systems. Modern wireless communication standards involve large complexity and highly-multirate behavior, and typically result in long simulation time. The traditional approach for simulating SDF graphs is to compute and execute static single-processor schedules. Nowadays, multi-core processors are increasingly popular for their potential performance improvements through on-chip, thread-level parallelism. However, without novel scheduling and simulation techniques that explicitly explore multithreading capability, current design tools gain only minimal performance improvements. In this paper, we present a new multithreaded simulation scheduler, called MSS, to provide simulation runtime speed-up for executing SDF graphs on multi-core processors. We have implemented MSS in the Advanced Design System (ADS) from Agilent Technologies. On an Intel dualcore, hyper-threading (4 processing units) processor, our results from this implementation demonstrate up to 3.5 times speed-up in simulating modern wireless communication systems (e.g., WCDMA3G, CDMA 2000, WiMax, EDGE, and Digital TV). Chia-Jui Hsu, José Luis Pino, Shuvra S. Bhattacharyya |
DAC | 3 |
| 2008 | An Optimized Message Passing Framework for Parallel Implementation of Signal Processing ApplicationsabstractNovel reconfigurable computing platforms enable efficient realizations of complex signal processing applications by allowing exploitation of parallelization resulting in high throughput in a cost-efficient way. However, the design of such systems poses various challenges due to the complexities posed by the applications themselves as well as the heterogeneous nature of the targeted platforms. One of the most significant challenges is communication between the various computing elements for parallel implementation. In this paper, we present a communication interface, called the signal passing interface (SPI), that attempts to overcome this challenge by integrating relevant properties of two different yet important paradigms in this context - dataflow and the message passing interface (MPI). SPI is targeted towards signal processing applications and, due to its careful specialization, more performance-efficient for their embedded implementation. It is also more easier and intuitive to use. Earlier, a preliminary version of SPI was presented [12] which was restricted to static dataflow behavior. Here, we present a more complete version of SPI with new features to address both static and dynamic dataflow behavior, and to provide new optimization techniques. We develop a hardware description language (HDL) realization of the SPI library, and demonstrate its functionality on the Xilinx Virtex-4 FPGA. Details of the HDL-based SPI library along with experiments with two signal processing applications on the FPGA are also presented. Sankalita Saha, Jason Schlessman, Sebastian Puthenpurayil, Shuvra S. Bhattacharyya, Marilyn Wolf |
DATE | 4 |
| 2008 | A generalized static data flow clustering algorithm for mpsoc scheduling of multimedia applicationsabstractAbstract—In this paper, an efficient embedded software synthesis approach based on a generalized clustering algorithm for static dataflow subgraphs embedded in general dataflow graphs is proposed. The clustered subgraph is quasi-statically scheduled, thus improving performance of the synthesized software in terms of latency and throughput compared to a dynamically scheduled execution. The proposed clustering algorithm outperforms previous approaches by a faster computation and a more compact representation of the derived quasi-static schedules. This is achieved by a rule-based approach, which avoids an explicit enumeration of the state space. Experimental results show significant improvements in both performance and code size when compared to a state-of-the-art clustering algorithm. Joachim Falk, Joachim Keinert, Christian Haubelt, Jürgen Teich, Shuvra S. Bhattacharyya |
EMSOFT | 5 |
| 2008 | Multiobjective Optimization of FPGA-Based Medical Image RegistrationabstractWith a multitude of technological innovations, one emerging trend in image processing, and medical image processing, in particular, is custom hardware implementation of computationally intensive algorithms in the quest to achieve real-time performance. For reasons of area-efficiency and performance, these implementations often employ limited-precision datapaths. Identifying effective wordlengths for these datapaths while accounting for tradeoffs between design complexity and accuracy is a critical and time consuming aspect of this design process. Having access to optimized tradeoff curves can equip designers to adapt their designs to different performance requirements and target specific devices while reducing design time. This paper presents a multiobjective optimization strategy developed in the context of field-programmable gate array-based implementation of medical image registration. Within this framework, we compare several search methods and demonstrate the applicability of an evolutionary algorithm-based search for efficiently identifying superior multiobjective tradeoff curves. This strategy can easily be adapted to a wide range of signal processing applications, including areas of image and video processing beyond the medical domain. Omkar Dandekar, William Plishker, Shuvra S. Bhattacharyya, Raj Shekhar |
FCCM | 3 |
| 2008 | Systematic generation of FPGA-based FFT implementationsabstractIn this paper, we propose a systemic approach for synthesizing field-programmable gate array (FPGA) implementations of fast Fourier transform (FFT) computations. Our approach considers both cost (in terms of FPGA resource requirements), and performance (in terms of throughput), and optimizes for both of these dimensions based on user-specified requirements. Our approach involves two orthogonal techniques-FFT inner loop unrolling and outer loop unrolling - to perform design space exploration in terms of cost and performance. By appropriately combining these two forms unrolling, we can achieve cost-optimized FFT implementations in terms of FPGA slices or block RAMs in FPGA, subject to the required throughput. We compared the results of our synthesis approach with a recently-introduced commercial FPGA intellectual property (IP) core - the FFT IP module in the Xilinx LogiCore Library, which provides different FFT implementations that are optimized for a limited set of performance levels. Our results demonstrate efficiency levels that are in some cases better than these commercial IP blocks. At the same time, our approach provides the advantages of being able to optimize implementations based on arbitrary, user-specified performance levels, and of being based on general formulations of FFT loop unrolling trade-offs, which can be retargeted to different kinds of FPGA devices. Hojin Kee, Newton Petersen, Jacob Kornerup, Shuvra S. Bhattacharyya |
ICASSP | 4 |
| 2008 | Parameterized design framework for hardware implementation of particle filtersabstractParticle filtering methods provide powerful techniques for solving non-linear state-estimation problems, and are applied to a variety of application areas in signal processing. Because of their vast computational complexity, real-time hardware implementation of particle-filter-based systems is a challenging task. However, many particle filter applications share common characteristics, and the same system design can be reused with appropriate streamlining. To achieve this, a parameterized design framework for particle filters is proposed in this paper. In this framework, parameterization of system features that vary over specific implementations enables reuse of a generic design for a wide range of applications with minimal re-design effort. Using this framework, we explore different design options for implementing two different particle filtering applications on field-programmable gate arrays (FPGAs), and we present associated results on trade-offs between area (FPGA resource requirements) and execution speed. Sankalita Saha, Neal K. Bambha, Shuvra S. Bhattacharyya |
ICASSP | 3 |
| 2008 | Design and optimization of a distributed, embedded speech recognition systemabstractIn this paper, we present the design and implementation of a distributed sensor network application for embedded, isolated-word, real-time speech recognition. In our system design, we adopt a parameterized-dataflow-based modeling approach to model the functionalities associated with sensing and processing of acoustic data, and we implement the associated embedded software on an off-the-shelf sensor node platform that is equipped with an acoustic sensor. The topology of the sensor network deployed in this work involves a clustered network hierarchy. A customized time division multiple access protocol is developed to manage the wireless channel. We analyze the distribution of the overall computation workload across the network to improve energy efficiency. In our experiments, we demonstrate the recognition accuracy for our speech recognition system to verify its functionality and utility. We also evaluate improvements in network lifetime to demonstrate the effectiveness of our energy-aware optimization techniques. Chung-Ching Shen, William Plishker, Shuvra S. Bhattacharyya |
IPDPS | 3 |
| 2008 | The Signal Passing Interface and Its Application to Embedded Implementation of Smart Camera ApplicationsabstractEmbedded smart camera systems comprise computation- and resource-hungry applications implemented on small, complex but resource-hardy platforms. Efficient implementation of such applications can benefit significantly from parallelization. However, communication between different processing units is a nontrivial task. In addition, new and emerging distributed smart cameras require efficient methods of communication for optimized distributed implementations. In this paper, a novel communication interface, called the signal passing interface (SPI), is presented that attempts to overcome this challenge by integrating relevant properties of two different, yet important, paradigms in this context-dataflow and message passing interface (MPI). Dataflow is a widely used modeling paradigm for signal processing applications, while MPI is an established communication interface in the general-purpose processor community. SPI is targeted toward computation-intensive signal processing applications, and due to its careful specialization, more performance-efficient for embedded implementation in this domain. SPI is also much easier and more intuitive to use. In this paper, successful application of this communication interface to two smart camera applications has been presented in detail to validate a new methodology for efficient distributed implementation for this domain. Sankalita Saha, Sebastian Puthenpurayil, Jason Schlessman, Shuvra S. Bhattacharyya, Wayne Wolf |
Proc. IEEE | 4 |
| 2007 | Energy-Aware Data Compression for Wireless Sensor NetworksabstractData compression techniques have extensive applications in power-constrained digital communication systems, such as in the rapidly-developing domain of wireless sensor network applications. This paper explores energy consumption tradeoffs associated with data compression, particularly in the context of lossless compression for acoustic signals. Such signal processing is relevant in a variety of sensor network applications, including surveillance and monitoring. Applying data compression in a sensor node generally reduces the energy consumption of the transceiver at the expense of additional energy expended in the embedded processor due to the computational cost of compression. This paper introduces a methodology for comparing data compression algorithms in sensor networks based on the figure of merit D/ E, where D is the amount of data (before compression) that can be transmitted under a given energy budget E for computation and communication. We develop experiments to evaluate, using this figure of merit, different variants of linear predictive coding. We also demonstrate how different models of computation applied to the embedded software design lead to different degrees of processing efficiency, and thereby have significant effect on the targeted figure of merit. Sebastian Puthenpurayil, Ruirui Gu, Shuvra S. Bhattacharyya |
ICASSP (2) | 3 |
| 2007 | Design Techniques for Streamlined Integration and Fault Tolerance in a Distributed Sensor System for Line-crossing RecognitionabstractDistributed sensor system applications (e.g., wireless sensor networks) have been studied extensively in recent years. Such applications involve resource-limited embedded sensor nodes that communicate with each other through self-organizing protocols. Depending on application requirements, distributed sensor system design may include protocol and prototype implementation. Prototype implementation is especially useful in establishing and maintaining system functionality as the design is customized to satisfy size, energy, and cost constraints. In this paper, we present a streamlined, application-specific approach to incorporating fault tolerance into a TDMA-based distributed sensor system for line-crossing recognition. The objective of this approach is to prevent node failures from translating into failures in the overall system. Our approach is specialized and light-weight so that fault tolerance is achieved without significant degradation in energy efficiency. We also present an asynchronous handshaking approach for providing synchronization between the transceiver and digital processing subsystem in sensor node. This provides a general method for achieving such synchronization with reduced hardware requirements and reduced energy consumption compared to conventional approaches, which rely on generic interface protocols. We demonstrate the capabilities of our approaches to fault tolerance and transceiver-processor integration through experiments involving a complete prototype wireless sensor network test-bed, and a distributed line-crossing recognition application that runs on this test-bed. Chung-Ching Shen, Roni Kupershtok, Shuvra S. Bhattacharyya, Neil Goldsman |
ICCCN | 3 |
| 2007 | Compact, Low Power Wireless Sensor Network System for Line Crossing RecognitionabstractMany application-specific wireless sensor network (WSN) systems require small size and low power features due to their limited resources, and their use in distributed, wireless environments. In this paper, we present a light-weight distributed algorithm for line-crossing recognition, together with its analysis, implementation, and experimental evaluation within a prototype wireless sensor network platform. The algorithm is developed in conjunction with a TDMA-based communication protocol such that the proposed system provides for low duty cycle and energy efficient operation. An accurate lifetime model is proposed with consideration of detailed energy usage to analyze and estimate the system lifetime. Our experimental results demonstrate the accuracy of this lifetime model, and its utility in optimizing network implementation. The design and experimental evaluation of our prototype network demonstrates the compactness and functionality of the proposed distributed WSN system for line-crossing recognition. Chung-Ching Shen, Roni Kupershtok, Felice Maria Vanin, Xi Shao, Datta Sheth, Neil Goldsman, Quirino Balzano, Shuvra S. Bhattacharyya |
ISCAS | 9 |
| 2007 | An Energy-Driven Design Methodology for Distributing DSP Applications across Wireless Sensor NetworksabstractWireless sensor network (WSN) applications have been studied extensively in recent years. Such applications involve resource-limited embedded sensor nodes that have small size and low power requirements. Based on the need for extended network lifetimes in WSNs in terms of energy use, the energy efficiency of computation and communication operations in the embedded sensor nodes becomes critical. Digital signal processing (DSP) applications typically require intensive data processing operations. They are difficult to apply directly in resource-limited WSNs because their operational complexity can strongly influence the network lifetime. In this paper, we present a design methodology for modeling and implementing DSP applications applied to wireless sensor networks. This methodology explores efficient modeling techniques for DSP applications, including acoustic sensing and data processing; derives formulations of energy-driven partitioning for distributing such applications across wireless sensor networks; and develops efficient heuristic algorithms for finding partitioning results that maximize the network lifetime. A case study involving a speech recognition system demonstrates the capabilities of our proposed methodology. Chung-Ching Shen, William Plishker, Shuvra S. Bhattacharyya, Neil Goldsman |
RTSS | 3 |
| 2007 | Probabilistic design of multimedia embedded systemsabstractIn this paper, we propose the novel concept of probabilistic design for multimedia embedded systems, which is motivated by the challenge of how to design, but not overdesign, such systems while systematically incorporating performance requirements of multimedia application, uncertainties in execution time, and tolerance for reasonable execution failures. Unlike most present techniques that are based on either worst- or average-case execution times of application tasks, where the former guarantees the completion of each execution, but often leads to overdesigned systems, and the latter fails to provide any completion guarantees, the proposed probabilistic design method takes advantage of unique features mentioned above of multimedia systems to relax the rigid hardware requirements for software implementation and avoid overdesigning the system. In essence, this relaxation expands the design space and we further develop an off-line on-line minimum effort algorithm for quick exploration of the enlarged design space at early design stages. This is the first step toward our goal of bridging the gap between real-time analysis and embedded software implementation for rapid and economic multimedia system design. It is our belief that the proposed method has great potential in reducing system resource while meeting performance requirements. The experimental results confirm this as we achieve significant saving in system's energy consumption to provide a statistical completion ratio guarantee (i.e., the expected number of completions over a large number of iterations is greater than a given value). Shaoxiong Hua, Gang Qu 0001, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2007 | Beyond single-appearance schedules: Efficient DSP software synthesis using nested procedure callsabstractSynthesis of digital signal-processing (DSP) software from dataflow-based formal models is an effective approach for tackling the complexity of modern DSP applications. In this paper, an efficient method is proposed for applying subroutine call instantiation of module functionality when synthesizing embedded software from a dataflow specification. The technique is based on a novel recursive decomposition of subgraphs in a cluster hierarchy that is optimized for low buffer size. Applying this technique, one can achieve significantly lower buffer sizes than what is available for minimum code size inlined schedules, which have been the emphasis of prior work on software synthesis. Furthermore, it is guaranteed that the number of procedure calls in the synthesized program is polynomially bounded in the size of the input dataflow graph, even though the number of module invocations may increase exponentially. This recursive decomposition approach provides an efficient means for integrating subroutine-based module instantiation into the design space of DSP software synthesis. The experimental results demonstrate a significant improvement in buffer cost, especially for more irregular multirate DSP applications, with moderate code and execution time overhead. Ming-Yung Ko, Praveen K. Murthy, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2007 | Efficient simulation of critical synchronous dataflow graphsabstractSystem-level modeling, simulation, and synthesis using electronic design automation (EDA) tools are key steps in the design process for communication and signal processing systems, and the synchronous dataflow (SDF) model of computation is widely used in EDA tools for these purposes. Behavioral representations of modern wireless communication systems typically result in critical SDF graphs : These consist of hundreds of components (or more) and involve complex intercomponent connections with highly multirate relationships (i.e., with large variations in average rates of data transfer or component execution across different subsystems). Simulating such systems using conventional SDF scheduling techniques generally leads to unacceptable simulation time and memory requirements on modern workstations and high-end PCs. In this article, we present a novel simulation-oriented scheduler (SOS) that strategically integrates several techniques for graph decomposition and SDF scheduling to provide effective, joint minimization of time and memory requirements for simulating critical SDF graphs. We have implemented SOS in the advanced design system (ADS) from Agilent Technologies. Our results from this implementation demonstrate large improvements in simulating real-world, large-scale, and highly multirate wireless communication systems (e.g., 3GPP, Bluetooth, 802.16e, CDMA 2000, XM radio, EDGE, and Digital TV). Chia-Jui Hsu, Ming-Yung Ko, Shuvra S. Bhattacharyya, Suren Ramasubbu, José Luis Pino |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Affine Nested Loop Programs and their Binary Parameterized Dataflow Graph CounterpartsabstractParameterized static affine nested loop programs can be automatically converted to input-output equivalent Kahn Process Network specifications. These networks turn out to be close relatives of parameterized cyclo-static dataflow graphs. Token production and consumption can be cyclic with a finite number of cycles or finite non-cyclic. Moreover the token production and consumption sequences are binary. Ed F. Deprettere, Todor P. Stefanov, Shuvra S. Bhattacharyya, Mainak Sen |
ASAP | 3 |
| 2006 | Efficient simulation of critical synchronous dataflow graphsabstractSimulation and verification using electronic design automation (EDA) tools are key steps in the design process for communication and signal processing systems. The synchronous dataflow (SDF) model of computation is widely used in EDA tools for system modeling and simulation in the communication and signal processing domains. Behavioral representations of practical wireless communication systems typically result in critical SDF graphs - they consist of hundreds of components (or more) and involve complex inter-component connections with highly multirate relationships (i.e., with large variations in average rates of data transfer or component execution across different subsystems). Simulating such systems using conventional SDF scheduling techniques generally leads to unacceptable simulation time and memory requirements on modern workstations and high-end PCs. In this paper, we present a novel simulation-oriented SDF scheduler (SOS) that strategically integrates several techniques for graph decomposition and SDF scheduling to provide effective, joint minimization of time and memory requirements for simulating large-scale and heavily multirate SDF graphs. We have implemented the SOS scheduler in the Advanced Design System (ADS) from Agilent Technologies. Our results from this implementation demonstrate large improvements in simulating real-world wireless communication systems (e.g. 3GPP, Bluetooth, 802.16e, CDMA 2000, and XM radio). Chia-Jui Hsu, Suren Ramasubbu, Ming-Yung Ko, José Luis Pino, Shuvra S. Bhattacharyya |
DAC | 5 |
| 2006 | Mapping Multimedia Applications Onto Configurable Hardware With Parameterized Cyclo-Static Dataflow GraphsabstractThis paper develops methods for model-based design and implementation of image processing applications. We apply our previously developed meta-modeling technique of homogeneous parameterized dataflow (HPDF)[9] to the framework of cyclostatic dataflow (CSDF) [1], and demonstrate this integrated modeling methodology through hardware mapping of a gesture recognition application. We also provide a comparative study between HPDF/CSDF-based representation of the gesture recognition application, and a previously developed version based on applying HPDF in conjunction with conventional synchronous dataflow (SDF) semantics [9]. Fiorella Haim, Mainak Sen, Dong-Ik Ko, Shuvra S. Bhattacharyya, Marilyn Wolf |
ICASSP (3) | 4 |
| 2006 | Energy-efficient embedded software implementation on multiprocessor system-on-chip with multiple voltagesabstractThis paper develops energy-driven completion ratio guaranteed scheduling techniques for the implementation of embedded software on multiprocessor systems with multiple supply voltages. We leverage application's performance requirements, uncertainties in execution time, and tolerance for reasonable execution failures to scale each processor's supply voltage at run-time to reduce the multiprocessor system's total energy consumption. Specifically, we study how to trade the difference between the system's highest achievable completion ratio Q max and the required completion ratio Q 0 for energy saving. First, we propose a best-effort energy minimization algorithm (BEEM1) that achieves Q max with the provably minimum energy consumption. We then relax its unrealistic assumption on the application's real execution time and develop algorithm BEEM2 that only requires the application's best- and worst-case execution times. Finally, we propose a hybrid offline on-line completion ratio guaranteed energy minimization algorithm (QGEM) that provides the required Q 0 with further energy reduction based on the probabilistic distribution of the application's execution time. We implement the proposed algorithms and verify their energy efficiency on real-life DSP applications and the TGFF random benchmark suite. BEEM1, BEEM2, and QGEM all provide the required completion ratio with average energy reduction of 28.7, 26.4, and 35.8%, respectively. Shaoxiong Hua, Gang Qu 0001, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2006 | Efficient Techniques for Clustering and Scheduling onto Embedded MultiprocessorsabstractMultiprocessor mapping and scheduling algorithms have been extensively studied over the past few decades and have been tackled from different perspectives. In the late 1980's, the two-step decomposition of schedulingnto clustering and cluster-scheduling - was introduced. Ever since, several clustering and merging algorithms have been proposed and individually reported to be efficient. However, it is not clear how effective they are and how well they compare against single-step scheduling algorithms or other multistep algorithms. In this paper, we explore the effectiveness of the two-phase decomposition of scheduling and describe efficient and novel techniques that aggressively streamline interprocessor communications and can be tuned to exploit the significantly longer compilation time that is available to embedded system designers. We evaluate a number of leading clustering and merging algorithms using a set of benchmarks with diverse structures. We present an experimental setup for comparing the single-step against the two-step scheduling approach. We determine the importance of different steps in scheduling and the effect of different steps on overall schedule performance and show that the decomposition of the scheduling process indeed improves the overall performance. We also show that the quality of the solutions depends on the quality of the clusters generated in the clustering step. Based on the results, we also discuss why the parallel time metric in the clustering step may not provide an accurate measure for the final performance of cluster-scheduling Vida Kianzad, Shuvra S. Bhattacharyya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2005 | CASPER: An Integrated Energy-Driven Approach for Task Graph Scheduling on Distributed Embedded SystemsabstractFor multiprocessor embedded systems, the dynamic voltage scaling (DVS) technique can be applied to scheduled applications for energy reduction. DVS utilizes slack in the schedule to slow down processes and save energy. Therefore, it is generally believed that the maximal energy saving is achieved on a schedule with the minimum makespan (maximal slack). Most current approaches treat task assignment, scheduling, and DVS separately. In this paper, we present a framework called CASPER (combined assignment, scheduling, and power-management) that challenges this common belief by integrating task scheduling and DVS under a single iterative optimization loop via genetic algorithm. We have conducted extensive experiments to validate the energy efficiency of CASPER. For homogeneous multiprocessor systems (in which all processors are of the same type), we consider a recently proposed slack distribution algorithm (PDP-SPM) by S. Hua and G. Qu (2005): applying PDP-SPM on the schedule with the minimal makespan gives an average of 53.8% energy saving; CASPER finds schedules with slightly larger makespan but a 57.3% energy saving, a 7.8% improvement. For heterogeneous systems, we consider the power variation DVS (PV-DVS) algorithm by Schmitz et al. (2004), CASPER improves its energy efficiency by 8.2%. Finally, our results also show that the proposed single loop CASPER framework saves 23.3% more energy over GMA+EE-GLSA by Schmitz et al. (2002), the only other known integrated approach with a nested loop that combines scheduling and power management in the inner loop but leaves assignment in the outer loop. Vida Kianzad, Shuvra S. Bhattacharyya, Gang Qu 0001 |
ASAP | 2 |
| 2005 | Communication strategies for shared-bus embedded multiprocessorsabstractThis paper explores the problem of efficiently ordering interprocessor communication operations in both statically and dynamically-scheduled multiprocessors for iterative dataflow graphs with probabilistic execution times. In most digital signal processing applications, the throughput of the system is significantly affected by communication costs. We explicitly model these costs within an effective graph-theoretic analysis framework. We show that ordered transaction schedules can significantly outperform both self-timed schedules and dynamic schedules for moderate task execution time variability. As the task execution time variability increases, we show that first self-timed and then dynamic scheduling policies are preferred. We perform an extensive experimental comparison on both real and simulated benchmarks to gauge the effect of synchronization and communication overhead costs on these crossover points. Neal K. Bambha, Shuvra S. Bhattacharyya |
EMSOFT | 2 |
| 2005 | Dynamic configuration of dataflow graph topology for DSP system design [video encoder example]abstractDataflow is widely used for designing DSP applications. Despite its intrinsic advantages, one weak point is its difficulty in flexible expression of applications with data dependent change in execution structure. This paper suggests an approach to providing dynamically configured dataflow graph topologies using a new modeling and synthesis technique called DGT (dynamic graph topology). DGT builds on PSDF semantics (B. Bhattacharya et al, IEEE Tran. on Sig. Proc., vol.49(10), p.2408-2421, 2001). All possible graph topologies for a given graph are obtained at compile time and the corresponding graph based on parameters and data is dynamically set up in an efficient manner at runtime before the invocation of the associated graph. Systematic methods for reducing code and buffer size are applied based on characteristics of each configured graph. We have compared DGT against conventional modeling approaches through a detailed case study of an MPEG 2 video encoder system, and our experiments demonstrate the efficiency of the DGT approach. Dong-Ik Ko, Shuvra S. Bhattacharyya |
ICASSP (5) | 2 |
| 2005 | Modeling image processing systems with homogeneous parameterized dataflow graphsabstractWe describe a new dataflow model called homogeneous parameterized dataflow (HPDF). This form of dynamic dataflow graph takes advantage of the fact that in a large number of image processing applications, data production and consumption rates, though dynamic, are equal across graph edges for any particular iteration, which leads to a homogeneous rate of actor execution, even though data production and consumption values are dynamic and vary across graph edges. We discuss existing dataflow models and formulate in detail the HPDF model. We develop examples of applications that are described naturally in terms of HPDF semantics and present experimental results that demonstrate the efficacy of the HPDF approach. Mainak Sen, Shuvra S. Bhattacharyya, Tiehan Lv, Marilyn Wolf |
ICASSP (5) | 2 |
| 2005 | An Extended Motion-Estimation Architecture Applied to Shape RecognitionabstractAn architecture for shape recognition is presented, with emphasis on low-latency and power efficiency. This architecture is an extension of an existing architecture used for motion estimation. A number of algorithms were mapped to this architecture. Bounds related to power are given per frame for memory access rates. Face detection within CIPR CIF sequences was used as a target application, with feasible frame rates of 30 fps attained. Power results for this extended architecture correlate with power consumption of the existing architecture Jason Schlessman, Sankalita Saha, Marilyn Wolf, Shuvra S. Bhattacharyya |
ICME | 4 |
| 2005 | Software Synthesis from the Dataflow Interchange FormatabstractSpecification, validation, and synthesis are important aspects of embedded systems design. The use of dataflow-based design environments for these purposes is becoming increasingly popular in the domain of digital signal processing (DSP). The dataflow inter-change format (DIF) [11] and the associated DIF package have been developed for specifying, working with, and transferring dataflow-based DSP designs across tools. In this paper, we present the newly developed DIF-to-C software synthesis framework for automatically generating monolithic C-code implementations from DSP system specifications that are programmed in DIF. This framework allows designers to efficiently explore the complex range of implementation tradeoffs that are available through various dataflow-based techniques for scheduling and memory management. Furthermore, the DIF-to-C framework provides a standard, vendor-neutral mechanism for linking coarse grain data-flow optimizations with fine grain hand-optimized libraries and the large body of optimization techniques in the area of C compilers for DSP. Through experiments involving several DSP applications, we demonstrate the novel and useful capabilities of our DIF-to-C software synthesis framework. Chia-Jui Hsu, Shuvra S. Bhattacharyya |
SCOPES | 2 |
| 2005 | DSP Address Optimization Using Evolutionary AlgorithmsabstractOffset assignment has been studied as a highly effective approach to code optimization in modern digital signal processors (DSPs). In this paper, we propose two evolutionary algorithms to solve the general offset assignment problem with k address registers and an arbitrary auto-modify range. These algorithms differ from previous algorithms by having the capability of visiting the entire search space. We implement and analyze a variety of existing general offset assignment algorithms and test them on a set of standard benchmarks. The algorithms we propose can achieve a performance improvement of up to 31% over the best existing algorithm. We also achieve an average of 14% improvement over the union of recently proposed algorithms. Sean Leventhal, Neal K. Bambha, Shuvra S. Bhattacharyya, Gang Qu 0001 |
SCOPES | 4 |
| 2005 | Joint Application Mapping/Interconnect Synthesis Techniques for Embedded Chip-Scale MultiprocessorsabstractAs transistor sizes shrink, interconnects represent an increasing bottleneck for chip designers. Several groups are developing new interconnection methods and system architectures to cope with this trend. New architectures require new methods for high-level application mapping and hardware/software codesign. We present high-level scheduling and interconnect topology synthesis techniques for embedded multiprocessor systems-on-chip that are streamlined for one or more digital signal processing applications. That is, we seek to synthesize an application-specific interconnect topology. We show that flexible interconnect topologies utilizing low-hop communication between processors offer advantages for reduced power and latency. We show that existing multiprocessor scheduling algorithms can deadlock if the topology graph is not strongly connected, or if a constraint is imposed on the maximum number of hops allowed for communication. We detail an efficient algorithm that can be used in conjunction with existing scheduling algorithms for avoiding this deadlock. We show that it is advantageous to perform application scheduling and interconnect synthesis jointly, and present a probabilistic scheduling/interconnect algorithm that utilizes graph isomorphism to pare the design space. Neal K. Bambha, Shuvra S. Bhattacharyya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | CHARMED: A Multi-Objective Co-Synthesis Framework for Multi-Mode Embedded Systems
Vida Kianzad, Shuvra S. Bhattacharyya |
ASAP | 2 |
| 2004 | Java-through-C Compilation: An Enabling Technology for Java in Embedded SystemsabstractThe Java programming language is acheiving greater acceptance in high-end embedded systems such as cellphones and PDAs. However, current embedded implementations of Java impose tight constraints on functionality, while requiring significant storage space. In addition, they require that a JVM be ported to each such platform. We demonstrate the first Java-to-C compilation strategy that is suitable for a wide range of embedded systems, thereby enabling broad use of Java on embedded platforms. This strategy removes many of the constraints on functionality and reduces code size without sacrificing performance. The compilation framework described is easily retargetable, and is also applicable to bare-bones embedded systems with no operating system or JVM. On an average, we found the size of the generated executables to be over 25 times smaller than those generated by a cutting-edge Java-to-native-code compiler, while providing performance comparable to the best of various Java implementation strategies. Ankush Varma, Shuvra S. Bhattacharyya |
DATE | 2 |
| 2004 | Systematic Integration of Parameterized Local Search Techniques in Evolutionary Algorithms
Neal K. Bambha, Shuvra S. Bhattacharyya, Jürgen Teich, Eckart Zitzler |
GECCO (2) | 2 |
| 2004 | Systematic exploitation of data parallelism in hardware synthesis of DSP applicationsabstractWe describe an approach that we have explored for low-power synthesis and optimization of image, video, and digital signal processing (DSP) applications. In particular, we consider the systematic exploitation of data parallelism across the operations of an application dataflow graph when synthesizing a dedicated hardware implementation. Data parallelism occurs commonly in DSP applications, and provides flexible opportunities to increase throughput or lower power consumption. Exploiting this parallelism in a dedicated hardware implementation comes at the expense of increased resource requirements, which must be balanced carefully when applying the technique in a design tool. We propose a high level synthesis algorithm to determine the data parallelism factor for each computation, and, based on the area and performance trade-off curve, design an efficient hardware representation of the dataflow graph. For performance estimation, our approach uses a cyclostatic dataflow intermediate representation of the hardware structure under synthesis. We then apply an automatic hardware generation framework to build the actual circuit. Mainak Sen, Shuvra S. Bhattacharyya |
ICASSP (5) | 2 |
| 2004 | Compact Procedural Implementation in DSP Software Synthesis Through Recursive Graph Decomposition
Ming-Yung Ko, Praveen K. Murthy, Shuvra S. Bhattacharyya |
SCOPES | 3 |
| 2004 | Systematic integration of parameterized local search into evolutionary algorithmsabstractApplication-specific, parameterized local search algorithms (PLSAs), in which optimization accuracy can be traded off with run time, arise naturally in many optimization contexts. We introduce a novel approach, called simulated heating, for systematically integrating parameterized local search into evolutionary algorithms (EAs). Using the framework of simulated heating, we investigate both static and dynamic strategies for systematically managing the tradeoff between PLSA accuracy and optimization effort. Our goal is to achieve maximum solution quality within a fixed optimization time budget. We show that the simulated heating technique better utilizes the given optimization time resources than standard hybrid methods that employ fixed parameters, and that the technique is less sensitive to these parameter settings. We apply this framework to three different optimization problems, compare our results to the standard hybrid methods, and show quantitatively that careful management of this tradeoff is necessary to achieve the full potential of an EA/PLSA combination. Neal K. Bambha, Shuvra S. Bhattacharyya, Jürgen Teich, Eckart Zitzler |
IEEE Trans. Evol. Comput. | 2 |
| 2004 | Buffer merging - a powerful technique for reducing memory requirements of synchronous dataflow specificationsabstractWe develop a new technique called buffer merging for reducing memory requirements of synchronous dataflow (SDF) specifications. SDF has proven to be an attractive model for specifying DSP systems, and is used in many commercial tools like System Canvas, SPW, and Cocentric. Good synthesis from an SDF specification depends crucially on scheduling, and memory is an important metric for generating efficient schedules. Previous techniques on memory minimization have either not considered buffer sharing at all, or have done so at a fairly coarse level (the meaning of this will be made more precise in the article). In this article, we develop a buffer overlaying strategy that works at the level of an input/output edge pair of an actor. It works by algebraically encapsulating the lifetimes of the tokens on the input/output edge pair, and determines the maximum amount of the input buffer space that can be reused by the output. We develop the mathematical basis for performing merging operations, and develop several algorithms and heuristics for using the merging technique for generating efficient implementations. We show improvements of up to 48% over previous techniques. Praveen K. Murthy, Shuvra S. Bhattacharyya |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2003 | Logic foundry: rapid prototyping of FPGA-based DSP systemsabstractThe Logic Foundry is a system for the creation and integration of FPGA-based DSP systems. Recognizing that some of the greatest challenges in creating FPGA-based systems occur in the integration of the various components, we have developed a system that addresses the following four areas of integration: design flow integration, component integration, platform integration, and software integration. Using the Logic Foundry, a system can easily be specified, and then automatically constructed and integrated with system level software. Gary Spivey, Shuvra S. Bhattacharyya, Kazuo Nakajima |
ASP-DAC | 2 |
| 2003 | Energy reduction techniques for multimedia applications with tolerance to deadline missesabstractMany embedded systems such as PDAs require processing of the given applications with rigid power budget. However, they are able to tolerate occasional failures due to the imperfect human visual/auditory systems. The problem we address in this paper is how to utilize such tolerance to reduce multimedia system's energy consumption for providing guaranteed quality of service at the user level in terms of completion ratio. We explore a range of offline and on-line strategies that take this tolerance into account in conjunction with the modest non-determinism in application's execution time. First, we give a simple best-effort approach that achieves the maximum completion ratio; then we propose an enhanced on-line best-e.ort energy minimization (BEEM) approach and a hybrid offline/on-line minimum-effort (O2ME) approach. We prove that BEEM maintains the maximum completion ratio while consuming the provably least amount of energy and O2ME guarantees the required completion ratio statistically. We apply both approaches to a variety of benchmark task graphs, most from popular DSP applications. Simulation results show that significant energy savings (38% for BEEM and 54% for O2ME, both over the simple best-e.ort approach) can be achieved while meeting the required completion ratio requirements. Shaoxiong Hua, Gang Qu 0001, Shuvra S. Bhattacharyya |
DAC | 3 |
| 2003 | Energy-Efficient Multi-processor Implementation of Embedded Software
Shaoxiong Hua, Gang Qu 0001, Shuvra S. Bhattacharyya |
EMSOFT | 3 |
| 2003 | Partitioning for DSP Software Synthesis
Ming-Yung Ko, Shuvra S. Bhattacharyya |
SCOPES | 2 |
| 2003 | Introduction to the two special issues on memoryabstractintroduction Introduction to the two special issues on memory Share on Authors: Bruce Jacob View Profile , Shuvra Bhattacharyya View Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 2Issue 1February 2003 pp 1–4https://doi.org/10.1145/605459.605460Online:01 February 2003Publication History 1citation2,104DownloadsMetricsTotal Citations1Total Downloads2,104Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Bruce L. Jacob, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2002 | A Component Architecture for FPGA-Based, DSP System DesignabstractIntroducing FPGA components into DSP system implementations creates an assortment of challenges across system architecture and logic design. Recognizing that some of the greatest challenges occur in the integration of the various components, we have developed a component architecture and an associated set of software tools, collectively called the Logic Foundry. Using the Logic Foundry, an FPGA-based DSP system can be easily constructed from pre-built components and implemented on a variety of back-end FPGA platforms. The resulting implementation can then be encapsulated and integrated into a variety of front-end software application environments. This paper develops the component architecture and integration capabilities of the Logic Foundry, and examines a number of application case studies that we have experimented with using the Logic Foundry. Gary Spivey, Shuvra S. Bhattacharyya, Kazuo Nakajima |
ASAP | 2 |
| 2002 | Introduction to the two special issues on memoryabstractintroduction Introduction to the two special issues on memory Share on Authors: Bruce Jacob View Profile , Shuvra Bhattacharyya View Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 1Issue 1November 2002 pp 2–5https://doi.org/10.1145/581888.581890Online:01 November 2002Publication History 0citation861DownloadsMetricsTotal Citations0Total Downloads861Last 12 Months7Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Bruce L. Jacob, Shuvra S. Bhattacharyya |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2001 | Multiprocessor Clustering for Embedded Systems
Vida Kianzad, Shuvra S. Bhattacharyya |
Euro-Par | 2 |
| 2001 | An efficient timing model for hardware implementation of multirate dataflow graphsabstractWe consider the problem of representing timing information associated with functions in a dataflow graph used to represent a signal processing system in the context of high-level hardware (architectural) synthesis. This information is used for synthesis of appropriate architectures for implementing the graph. Conventional models for timing suffer from shortcomings that make it difficult to represent timing information in a hierarchical manner, especially for multirate signal processing systems. We identify some of these shortcomings, and provide an alternate model that does not have these problems. We show that with some reasonable assumptions on the way hardware implementations of multirate systems operate, we can derive general hierarchical descriptions of multirate systems similarly to single rate systems. Several analytical results such as the computation of the iteration period bound, that previously applied only to single rate systems can also easily be extended to multirate systems under the new assumptions. We have applied our model to several multirate signal processing applications, and obtained favorable results. We present results of the timing information computed for several multirate DSP applications that show how the new treatment can streamline the problem of performance analysis and synthesis of such systems. Nitin Chandrachoodan, Shuvra S. Bhattacharyya, K. J. Ray Liu |
ICASSP | 2 |
| 2001 | Shared buffer implementations of signal processing systems usinglifetime analysis techniquesabstractThere has been a proliferation of block-diagram environments for specifying and prototyping digital signal processing (DSP) systems. These include tools from academia such as Ptolemy and commercial tools such as DSPCanvas from Angeles Design Systems, signal processing work system (SPW) from Cadence, and COSSAP from Synopsys. The block diagram languages used in these environments are usually based on dataflow semantics because various subsets of dataflow have proven to be good matches for expressing and modeling signal processing systems. In particular, synchronous dataflow (SDF) has been found to be a particularly good match for expressing multirate signal processing systems. One of the key problems that arises during synthesis from an SDF specification is scheduling. Past work on scheduling from SDF has focused on optimization of program memory and buffer memory under a model that did not exploit sharing opportunities. In this paper, we build on our previously developed analysis and optimization framework for looped schedules to formally tackle the problem of generating optimally compact schedules for SDF graphs. We develop techniques for computing these optimally compact schedules in a manner that also attempts to minimize buffering memory under the assumption that buffers will be shared. This results in schedules whose data memory usage is drastically lower than methods in the past have achieved. The method we use is that of lifetime analysis; we develop a model for buffer lifetimes in SDF graphs and develop scheduling algorithms that attempt to generate schedules that minimize the maximum number of live tokens under the particular buffer lifetime model. We develop several efficient algorithms for extracting the relevant lifetimes from the SDF schedule. We then use the well-known first-fit heuristic for packing arrays efficiently into memory. We report extensive experimental results on applying these techniques to several practical SDF systems and show improvements that average 50% over previous techniques, with some systems exhibiting up to an 83% improvement over previous techniques. Praveen K. Murthy, Shuvra S. Bhattacharyya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Contention-Conscious Transaction Ordering in Embedded MultiprocessorsabstractThis paper explores the problem of efficiently ordering interprocessor communication operations in statically-scheduled multiprocessors for iterative dataflow graphs. In most digital signal processing applications, the throughput of the system is significantly affected by communication costs. By explicitly modeling these costs within an effective graph-theoretic analysis framework, we show that ordered transaction schedules can significantly outperform self-timed schedules even when synchronization costs are low. However, we also show that when communication latencies are non-negligible, finding an optimal transaction order given a static schedule is an NP-complete problem, and that this intractability holds both under iterative and non-iterative execution. We develop new heuristics for finding efficient transaction orders, and perform an experimental comparison to gauge the performance of these heuristics. Mukul Khandelia, Shuvra S. Bhattacharyya |
ASAP | 2 |
| 2000 | Optimizing the efficiency of parameterized local search within global search: a preliminary studyabstractApplication-specific, parameterized local search algorithms (PLSAs), in which optimization accuracy can be traded-off with run-time, arise naturally in many optimization contexts. We introduce a novel approach, called simulated heating, for systematically integrating parameterized local search into global search algorithms (GSAs) in general and evolutionary algorithms in particular. Using the framework of simulated heating, we investigate both static and dynamic strategies for systematically managing the trade-off between PLSA accuracy and optimization effort. We show quantitatively that careful management of this trade-off is necessary to achieve the full potential of a GSA/PLSA combination. Furthermore, we provide preliminary results which demonstrate the effectiveness of our simulated heating techniques in the context of code optimization for embedded software implementation, a practical problem that involves vast and complex search spaces. Eckart Zitzler, Jürgen Teich, Shuvra S. Bhattacharyya |
CEC | 3 |
| 2000 | Shared Memory Implementations of Synchronous Dataflow SpecificationsabstractThere has been a proliferation of block-diagram environments for specifying and prototyping DSP systems. These include tools from academia like Ptolemy and GRAPE, and commercial tools like SPW from Cadence Design Systems, Cossap from Synopsys, and the HP ADS tool from HP. The block diagram languages used in these environments are usually based on dataflow semantics because various subsets of dataflow have proven to be good matches for expressing and modeling signal processing systems. In particular synchronous dataflow (SDF) has been found to be a particularly good match for expressing multirate signal processing systems. One of the key problems that arises during synthesis from an SDF specification is scheduling. Past work on scheduling from SDF has focused on optimization of program memory and buffer memory. However, no attempt was made for overlaying or sharing buffers. In this paper we formally tackle the problem of generating optimally compact schedules for SDF graphs, that also attempt to minimize buffering memory under the assumption that buffers will be shared. This will result in schedules whose data memory usage is drastically lower (up to 83%) than methods in the past have achieved. Praveen K. Murthy, Shuvra S. Bhattacharyya |
DATE | 2 |
| 2000 | Parameterized dataflow modeling of DSP systemsabstractDataflow has proven to be an attractive computation model for programming DSP applications. A restricted version of dataflow, termed synchronous dataflow (SDF), that offers strong compile-time predictability properties, but has limited expressive power, has been studied extensively in the DSP context. Many extensions to synchronous dataflow have been proposed to increase its expressivity, while maintaining its compile-time predictability properties as much as possible. We propose a parameterized data-flow framework that can be applied as a meta-modeling technique to significantly improve the expressive power of an arbitrary data-flow model that possesses a well-defined concept of a graph iteration. Indeed, the parameterized dataflow framework is compatible with many of the existing dataflow models for DSP including SDF, CSDF, and SSDF. We develop a precise, formal semantics for parameterized synchronous dataflow that allows data-dependent dynamic DSP systems to be modeled in a natural and intuitive fashion. Desirable properties of a modeling environment like dynamic re-configurability and design re-use emerge as inherent characteristics of the parameterized framework. An example of a speech compression application is used to illustrate the efficacy of the parameterized modeling techniques in real-life data-dependent DSP systems. Bishnupriya Bhattacharya, Shuvra S. Bhattacharyya |
ICASSP | 2 |
| 2000 | The CBP parameter - a useful annotation to aid block-diagram compilers for DSPabstractMemory consumption is an important metric during software synthesis from block-diagram specifications of DSP applications. Conventionally, no assumption is made about when, during the execution of a functional block (actor), the associated data values (tokens) are actually consumed and produced. However, we show in this paper that it is possible to concisely and precisely capture key properties pertaining to the relative times at which tokens are produced and consumed by an actor. We show this by introducing the consumed-before-produced (CBP) parameter, which provides a general method for characterizing the token transfer of an actor. Good bounds on the CBP parameter can aid a block-diagram compiler in performing more aggressive optimizations for reducing buffer sizes on the edges between actors. We formally define the CBP parameter; derive some useful properties of this parameter; illustrate how the value of the parameter can be derived by examining in derail the multi-rate FIR filtering operation; and examine CBP parameterizations for several other practical DSP actors. Shuvra S. Bhattacharyya, Praveen K. Murthy |
ISCAS | 1 |
| 2000 | Evolutionary algorithms for the synthesis of embedded softwareabstractThis paper addresses the problem of trading off between the minimization of program and data memory requirements of single-processer Implementations of dataflow programs. Based on the formal model of synchronous dataflow (SDF) graphs, so called single appearance schedules are known to be program-memory optimal. Among these schedules, buffer memory schedules are investigated and explored based on a two-step approach: 1) an evolutionary algorithm (EA) is applied to efficiently explore the (in general) exponential search space of actor firing orders; 2) for each order, the buffer costs are evaluated by applying a dynamic programming post-optimization step (GDPPO). This iterative approach is compared to existing heuristics for buffer memory optimization. Eckart Zitzler, Jürgen Teich, Shuvra S. Bhattacharyya |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1998 | Buffer Memory Optimization in DSP Applications - An Evolutionary Approach
Jürgen Teich, Eckart Zitzler, Shuvra S. Bhattacharyya |
PPSN | 3 |
| 1997 | Optimized software synthesis for synchronous dataflowabstractThis paper reviews a set of techniques for compiling dataflow-based, graphical programs for digital signal processing (DSP) applications into efficient implementations on programmable digital signal processors. This is a critical problem because programmable digital signal processors have very limited amounts of on-chip memory and the speed power, and financial cost penalties for using off-chip memory are often prohibitively high for the types of applications, typically embedded systems, in which these processors are used. The compilation techniques described in this paper are developed for the synchronous dataflow model of computation, a model that has found widespread use for specifying and prototyping DSP systems. Shuvra S. Bhattacharyya, Praveen K. Murthy, Edward A. Lee |
ASAP | 1 |
| 1997 | Joint Minimization of Code and Data for Synchronous Dataflow Programs
Praveen K. Murthy, Shuvra S. Bhattacharyya, Edward A. Lee |
Formal Methods Syst. Des. | 2 |
| 1996 | Latency-constrained Resynchronization for Multiprocessor DSP ImplementationabstractResynchronization is a post-optimization for static multiprocessor schedules in which extraneous synchronization operations are introduced in such a way that the number of original synchronizations that consequently become redundant significantly exceeds the number of additional synchronizations. Redundant synchronizations are synchronization operations whose corresponding sequencing requirements are enforced completely by other synchronizations in the system. The amount of run-time overhead required for synchronization can be reduced significantly by eliminating redundant synchronizations. However, since additional serialization is imposed by the new synchronizations resynchronization can produce significant increase in latency. This paper addresses the problem of computing an optimal resynchronization (one that results in the lowest average rate at which synchronization operations have to be performed) among all resynchronizations that do not increase the latency beyond a prespecified upper bound L/sub max/. Our study is based on the context of self-timed execution of iterative data flow programs, which is an implementation model that has been applied extensively for digital signal processing systems. Shuvra S. Bhattacharyya, Sundararajan Sriram, Edward A. Lee |
ASAP | 1 |
| 1995 | Minimizing Synchronization Overhead in Statically Scheduled Multiprocessor SystemsabstractSynchronization overhead can significantly degrade performance in embedded multiprocessor systems. This paper develops techniques to determine a minimal set of processor synchronizations that are essential for correct execution in an embedded multiprocessor implementation. Our study is based in the context of self-timed execution of iterative dataflow programs; dataflow programming in this form has been applied extensively, particularly in the context of signal processing software. Self-timed execution refers to a combined compile-time/run-time scheduling strategy in which processors synchronize with one another only based on inter-processor communication requirements, and thus, synchronization of processors at the end of each loop iteration does not generally occur. We introduce a new graph-theoretic framework, based on a data structure called the synchronization graph, for analyzing and optimizing synchronization overhead in self-timed, iterative dataflow programs. We also present an optimization that involves converting a synchronization graph that is not strongly connected into a strongly connected graph. Shuvra S. Bhattacharyya, Sundararajan Sriram, Edward A. Lee |
ASAP | 1 |
| 1995 | Converting graphical DSP programs into memory constrained software prototypesabstractSince software prototypes of DSP applications are most efficient when their code and data space requirements can be accommodated entirely within the on-chip memory of the target processor it is crucial to employ efficient memory-minimizing compilation techniques in a DSP software prototyping system. In this paper, we introduce two techniques for the combined minimization of code and data when compiling graphical programs that are based on the synchronous dataflow (SDF) model. The first method is a customization to acyclic graphs of a bottom-up technique, called Pairwise Grouping of Adjacent Nodes (PGAN), that was proposed earlier for general SDF graphs. We show that our customization significantly reduces the complexity of the general PGAN algorithm and performs optimally for a certain class of applications. The second approach is a top-down technique, called Recursive Partitioning by Minimum Cuts (RPMC), that is based on a generalized minimum cut operation. From an extensive experimental study, we conclude that RPMC and our customization of PGAN are complementary, and both should be incorporated into SDF-based prototyping environments in which the minimization of memory requirements is important. Shuvra S. Bhattacharyya, Praveen K. Murthy, Edward A. Lee |
RSP | 1 |
| 1994 | Minimizing memory requirements for chain-structured synchronous dataflow programsabstractThis paper addresses trade-offs between the minimization of program memory and data memory requirements in the compilation of dataflow programs for multirate signal processing. Our techniques are specific to the synchronous dataflow (SDF) model of Lee and Messerschmitt (1987), which has been used extensively in software synthesis environments for DSP. We focus on programs that are represented as chain-structured SDF graphs. We show that there is an O(n/sup 3/) dynamic programming algorithm for determining a schedule that minimizes data memory usage among the set of schedules that minimize program memory usage. A practical example to illustrate the efficacy of this approach is given. Some extensions of this algorithm are also given; for example, we show that the algorithm applies to the more general class of well-ordered graphs.> Praveen K. Murthy, Shuvra S. Bhattacharyya, Edward A. Lee |
ICASSP (2) | 2 |
| 1994 | Looped Schedules for Dataflow Descriptions of Multirate Signal Processing Algorithms
Shuvra S. Bhattacharyya, Edward A. Lee |
Formal Methods Syst. Des. | 1 |
| 1989 | GABRIEL: A Design Environment for Programmable DSPsabstractGabriel is a retargetable software system for the development of assembly code and microcode for single or multiple programmable DSPs. It is intended to ease code development even for processors that are not easy targets for conventional compilers. Code generation for the Motorola DSP56001 is emphasized. A Thor-based simulator supplies a variety of target multi-DSP architectures based on the DSP56001. The top-level algorithm description is a large grain data flow graph, and a graphical interface using OCT and VEM provides a natural representation of the high level structure of the algorithm. Edward A. Lee, E. Goei, H. Heine, W.-H. Ho, Shuvra S. Bhattacharyya, Jeffery C. Bier, E. Guntvedt |
DAC | 5 |