Sourav Bhattacharya

dblp:69/3637 · DBLP profile ↗
← Back
72ranked-venue papers
19as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 14 since 2021Computer networks · 18 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Systems, architecture and hardware · 13 · 7 first-authorSoftware engineering, systems software and programming languages · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 7 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author
YearPublicationVenuePosition
2026 HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
abstract
State-of-the-art text-to-image diffusion models (DMs) achieve remarkable quality, yet their massive parameter scale (8-11B) poses significant challenges for inferences on resource-constrained devices. In this paper, we present HierarchicalPrune, a novel compression framework grounded in a key observation: DM blocks exhibit distinct functional hierarchies, where early blocks establish semantic structures while later blocks handle texture refinements. HierarchicalPrune synergistically combines three techniques: (1) Hierarchical Position Pruning, which identifies and removes less essential later blocks based on position hierarchy; (2) Positional Weight Preservation, which systematically protects early model portions that are essential for semantic structural integrity; and (3) Sensitivity-Guided Distillation, which adjusts knowledge-transfer intensity based on our discovery of block-wise sensitivity variations. As a result, our framework brings billion-scale diffusion models into a range more suitable for on-device inference, while preserving the quality of the output images. Specifically, combined with INT4 weight quantisation, HierarchicalPrune achieves 77.5-80.4% memory footprint reduction (e.g., from 15.8 GB to 3.2 GB) and 27.9-38.0% latency reduction, measured on server and consumer grade GPUs, with the minimum drop of 2.6% in GenEval score and 7% in HPSv2 score compared to the original model. Finally, our comprehensive user study with 85 participants demonstrates that HierarchicalPrune maintains perceptual quality comparable to the original model while significantly outperforming prior works.
Young D. Kwon, Rui Li 0052, Da Li 0001, Sourav Bhattacharya, Stylianos I. Venieris
AAAI5
2025 Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
abstract
Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure contamination impact, LLMs trained with/without contamination are compared. A contaminated LLM is more likely to generate test sentences it has seen during training. Then, speech recognisers based on LLMs are compared. They show only subtle error rate differences if the LLM is contaminated, but assign significantly higher probabilities to transcriptions seen during LLM training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.
Yuan Tseng, Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
ASRU5
2025 Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
abstract
Self-attention relies on positional embeddings to encode input order. Relative Position (RelPos) embeddings are widely used in Automatic Speech Recognition (ASR). However, RelPos has quadratic time complexity to input length and is often incompatible with fast GPU implementations of attention. In contrast, Rotary Positional Embedding (RoPE) rotates each input vector based on its absolute position, taking linear time to sequence length, implicitly encoding relative distances through self-attention dot products. Thus, it is usually compatible with efficient attention. However, its use in ASR remains underexplored. This work evaluates RoPE across diverse ASR tasks with training data ranging from 100 to 50,000 hours, covering various speech types (read, spontaneous, clean, noisy) and different accents in both streaming and non-streaming settings. ASR error rates are similar or better than RelPos, while training time is reduced by up to $21 \%$. Code is available via the SpeechBrain toolkit.
Shucong Zhang, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
ASRU4
2025 Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
abstract
Automatic speech recognition (ASR) with an encoder equipped with self-attention, whether streaming or non-streaming, takes quadratic time in the length of the speech utterance. This slows down training and decoding, increase the cost, and limits the deployment of the ASR in constrained devices. SummaryMixing is a promising linear-time complexity alternative to self-attention for non-streaming speech recognition that, for the first time, preserves or outperforms the accuracy of self-attention models. Unfortunately, the original definition of SummaryMixing is not suited to streaming speech recognition. Hence, this work extends SummaryMixing to a Conformer Transducer that works in both a streaming and an offline mode. It shows that this new linear-time complexity speech encoder outperforms self-attention in both scenarios while requiring less compute and memory during training and decoding.
Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
ICASSP4
2025 Edit: Efficient Diffusion Transformers with Linear Compressed Attention
abstract
Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation with higher resolution or on devices with limited resources. This work introduces an efficient diffusion transformer (EDiT) to alleviate these efficiency bottlenecks in conventional DiTs and Multimodal DiTs (MM-DiTs). First, we present a novel linear compressed attention method that uses a multi-layer convolutional network to modulate queries with local information while keys and values are aggregated spatially. Second, we formulate a hybrid attention scheme for multimodal inputs that combines linear attention for image-to-image interactions and standard scaled dot-product attention for interactions involving prompts. Merging these two approaches leads to an expressive, linear-time Multimodal Efficient Diffusion Transformer (MM-EDiT). We demonstrate the effectiveness of the EDiT and MM-EDiT architectures by integrating them into PixArt-Sigma (conventional DiT) and Stable Diffusion 3.5-Medium (MM-DiT), achieving up to 2.2x speedup with comparable image quality after distillation.
Philipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick, Luca Morreale, Mehdi Noroozi, Alberto Gil C. P. Ramos, Sourav Bhattacharya
ICCV8
2025 Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities
abstract
Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existing approaches typically require resource-intensive re-training or additional parameters to accommodate for the new tasks, which makes the model inefficient for on-device deployment. We propose *Multi-Task Upcycling* (MTU), a simple yet effective recipe that extends the capabilities of a pre-trained text-to-image diffusion model to support a variety of image-to-image generation tasks. MTU replaces Feed-Forward Network (FFN) layers in the diffusion model with smaller FFNs, referred to as *experts*, and combines them with a dynamic routing mechanism. To the best of our knowledge, MTU is the first multi-task diffusion modeling approach that seamlessly blends multi-tasking with on-device compatibility, by mitigating the issue of parameter inflation. We show that the performance of MTU is on par with the single-task fine-tuned diffusion models across several tasks including *image editing, super-resolution*, and *inpainting*, while maintaining similar latency and computational load (GFLOPs) as the single-task fine-tuned models.
Ruchika Chavhan, Abhinav Mehrotra, Malcolm Chadwick, Alberto Gil C. P. Ramos, Luca Morreale, Mehdi Noroozi, Sourav Bhattacharya
ICML7
2025 Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
Rogier C. van Dalen, Shucong Zhang, Titouan Parcollet, Sourav Bhattacharya
INTERSPEECH4
2025 Fast Sampling Through The Reuse Of Attention Maps In Diffusion Models
abstract
Abstract Text-to-image diffusion models have demonstrated unprecedented capabilities for flexible and realistic image synthesis. Nevertheless, these models rely on a time-consuming sampling procedure, which has motivated attempts to reduce their latency. When improving efficiency, researchers often use the original diffusion model to train an additional network designed specifically for fast image generation. In contrast, our approach seeks to reduce latency directly, without any retraining, fine-tuning, or knowledge distillation. In particular, we find the repeated calculation of attention maps to be costly yet redundant, and instead suggest reusing them during sampling. Our specific reuse strategies are based on ODE theory, which implies that the later a map is reused, the smaller the distortion in the final image. We empirically compare our reuse strategies with few-step sampling procedures of comparable latency, finding that reuse generates images that are closer to those produced by the original high-latency diffusion model.
Rosco Hunter, Lukasz Dudziak, Mohamed S. Abdelfattah, Abhinav Mehrotra, Sourav Bhattacharya, Hongkai Wen 0001
Int. J. Comput. Vis.5
2024 SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
INTERSPEECH4
2024 Linear-Complexity Self-Supervised Learning for Speech Processing
Shucong Zhang, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
INTERSPEECH4
2023 On the (In)Efficiency of Acoustic Feature Extractors for Self-Supervised Speech Representation Learning
abstract
International audience
Titouan Parcollet, Shucong Zhang, Rogier C. van Dalen, Alberto Gil C. P. Ramos, Sourav Bhattacharya
INTERSPEECH5
2023 Real-Time Personalised Speech Enhancement Transformers with Dynamic Cross-attended Speaker Representations
Shucong Zhang, Malcolm Chadwick, Alberto Gil C. P. Ramos, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
INTERSPEECH6
2022 Conditioning Sequence-to-sequence Networks with Learned Activations
Alberto Gil C. P. Ramos, Abhinav Mehrotra, Nicholas D. Lane, Sourav Bhattacharya
ICLR4
2021 Defensive Tensorization
Adrian Bulat, Jean Kossaifi, Sourav Bhattacharya, Yannis Panagakis, Timothy M. Hospedales, Georgios Tzimiropoulos, Nicholas D. Lane, Maja Pantic
BMVC3
2021 NAS-Bench-ASR: Reproducible Neural Architecture Search for Speech Recognition
Abhinav Mehrotra, Alberto Gil C. P. Ramos, Sourav Bhattacharya, Lukasz Dudziak, Ravichander Vipperla, Thomas C. P. Chau, Mohamed S. Abdelfattah, Samin Ishtiaq, Nicholas D. Lane
ICLR3
2020 3D-Stacked Memory For Shared-Memory Multithreaded Workloads
abstract
This paper aims to address the issue of CPU-memory intercommunication latency with the help of 3D stacked memory. We propose a 3D-stacked memory configuration, where a DRAM module is mounted on top of the CPU to reduce latency. We have used a comprehensive simulation environment to assure both fabrication feasibility and energy efficiency of the proposed 3D stacked memory modules. We have evaluated our proposed architecture by running PARSEC 2.1, a benchmark suite for shared-memory multithreaded workloads. The results demonstrate an average of 40% improvement over conventional DDR3/4 memory architectures.
Sourav Bhattacharya, Horacio González-Vélez
ECMS1
2020 Iterative Compression of End-to-End ASR Model Using AutoML
abstract
Increasing demand for on-device Automatic Speech Recognition (ASR) systems has resulted in renewed interests in developing automatic model compression techniques. Past research have shown that AutoML-based Low Rank Factorization (LRF) technique, when applied to an end-to-end Encoder-Attention-Decoder style ASR model, can achieve a speedup of up to 3.7x, outperforming laborious manual rank-selection approaches. However, we show that current AutoML-based search techniques only work up to a certain compression level, beyond which they fail to produce compressed models with acceptable word error rates (WER). In this work, we propose an iterative AutoML-based LRF approach that achieves over 5x compression without degrading the WER, thereby advancing the state-of-the-art in ASR compression.
Abhinav Mehrotra, Lukasz Dudziak, Jinsu Yeo, Young-Yoon Lee, Ravichander Vipperla, Mohamed S. Abdelfattah, Sourav Bhattacharya, Samin Ishtiaq, Alberto Gil C. P. Ramos, Nicholas D. Lane
INTERSPEECH7
2020 Bunched LPCNet: Vocoder for Low-Cost Neural Text-To-Speech Systems
abstract
LPCNet is an efficient vocoder that combines linear prediction and deep neural network modules to keep the computational complexity low. In this work, we present two techniques to further reduce it's complexity, aiming for a low-cost LPCNet vocoder-based neural Text-to-Speech (TTS) System. These techniques are: 1) Sample-bunching, which allows LPCNet to generate more than one audio sample per inference; and 2) Bit-bunching, which reduces the computations in the final layer of LPCNet. With the proposed bunching techniques, LPCNet, in conjunction with a Deep Convolutional TTS (DCTTS) acoustic model, shows a 2.19x improvement over the baseline run-time when running on a mobile device, with a less than 0.1 decrease in TTS mean opinion score (MOS).
Ravichander Vipperla, Kihyun Choo, Samin Ishtiaq, Kyoungbo Min, Sourav Bhattacharya, Abhinav Mehrotra, Alberto Gil C. P. Ramos, Nicholas D. Lane
INTERSPEECH6
2020 Augmenting Conversational Agents with Ambient Acoustic Contexts
abstract
Conversational agents are rich in content today. However, they are entirely oblivious to users’ situational context, limiting their ability to adapt their response and interaction style. To this end, we explore the design space for a context augmented conversational agent, including analysis of input segment dynamics and computational alternatives. Building on these, we propose a solution that redesigns the input segment intelligently for ambient context recognition, achieved in a two-step inference pipeline. We first separate the non-speech segment from acoustic signals and then use a neural network to infer diverse ambient contexts. To build the network, we curated a public audio dataset through crowdsourcing. Our experimental results demonstrate that the proposed network can distinguish between 9 ambient contexts with an average F1 score of 0.80 with a computational latency of 3 milliseconds. We also build a compressed neural network for on-device processing, optimised for both accuracy and latency. Finally, we present a concrete manifestation of our solution in designing a context-aware conversational agent and demonstrate use cases.
Chunjong Park, Chulhong Min, Sourav Bhattacharya, Fahim Kawsar
MobileHCI3
2020 Dynamic Distributed Edge Resource Provisioning via Online Learning across Timescales
abstract
The strategic management of distributed resources of mobile edge computing networks often requires managing different system components over different timescales. In this paper, we formulate a nonlinear mixed-integer program to capture the online optimization of the edge network’s long-term cost, where we distribute workload more frequently on the fast timescale and provision resources less frequently on the slow timescale. We design a novel online learning framework consisting of three algorithms to make fast-timescale and slow-timescale fractional decisions, respectively, and round such decisions into integers. Our algorithms run in polynomial time in an online manner, jointly solving the original NP-hard problem that can contain arbitrary and unpredictable inputs. Via rigorous formal analysis, we prove a parameterized-constant competitive ratio as the performance guarantee for our approach. We conduct extensive evaluations with real-world data and confirm our approach’s superiority over existing practices and state-of-the-arts.
Wencong You, Lei Jiao 0002, Sourav Bhattacharya, Yuan Zhang 0013
SECON3
2019 AudiDoS: Real-Time Denial-of-Service Adversarial Attacks on Deep Audio Models
abstract
Deep learning has enabled personal and IoT devices to rethink microphones as a multi-purpose sensor for understanding conversation and the surrounding environment. This resulted in a proliferation of Voice Controllable Systems (VCS) around us. The increasing popularity of such systems is also prone to attracting miscreants, who often want to take advantage of the VCS without the knowledge of the user. Consequently, understanding the robustness of VCS, especially under adversarial attacks, has become an important research topic. Although there exists some previous work on audio adversarial attacks, their scopes are limited to embedding the attacks onto pre-recorded music clips, which when played through speakers cause VCS to misbehave. As an attack-audio needs to be played, the occurrence of this type of attacks can be suspected by a human listener. In this paper, we focus on audio-based Denial-of-Service (DoS) attack, which is unexplored in the literature. Contrary to previous work, we show that adversarial audio attacks in real-time and overthe-air are possible, while a user interacts with VCS. We show that the attacks are effective regardless of the user's command and interaction timings. In this paper, we present a first-of-itskind imperceptible and always-on universal audio perturbation technique that enables such DoS attack to be successful. We thoroughly evaluate the performance of the attacking scheme across (i) two learning tasks, (ii) two model architectures and (iii) three datasets. We demonstrate that the attack can introduce as high as 78% error rate in audio recognition tasks.
Taesik Gong, Alberto Gil C. P. Ramos, Sourav Bhattacharya, Akhil Mathur, Fahim Kawsar
ICMLA3
2019 MobiSR: Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors
abstract
In recent years, convolutional networks have demonstrated unprecedented performance in the image restoration task of super-resolution (SR). SR entails the upscaling of a single low-resolution image in order to meet application-specific image quality demands and plays a key role in mobile devices. To comply with privacy regulations and reduce the overhead of cloud computing, executing SR models locally on-device constitutes a key alternative approach. Nevertheless, the excessive compute and memory requirements of SR workloads pose a challenge in mapping SR networks on resource-constrained mobile platforms. This work presents MobiSR, a novel framework for performing efficient super-resolution on-device. Given a target mobile platform, the proposed framework considers popular model compression techniques and traverses the design space to reach the highest performing trade-off between image quality and processing speed. At run time, a novel scheduler dispatches incoming image patches to the appropriate model-engine pair based on the patch's estimated upscaling difficulty in order to meet the required image quality with minimum processing latency. Quantitative evaluation shows that the proposed framework yields on-device SR designs that achieve an average speedup of 2.13x over highly-optimized parallel difficulty-unaware mappings and 4.79x over highly-optimized single compute engine implementations.
Royson Lee, Stylianos I. Venieris, Lukasz Dudziak, Sourav Bhattacharya, Nicholas D. Lane
MobiCom4
2019 Poster: MobiSR - Efficient On-Device Super-Resolution through Heterogeneous Mobile Processors
abstract
In recent years, convolutional networks have demonstrated unprecedented performance in the image restoration task of super-resolution (SR). SR entails the upscaling of a single low-resolution image in order to meet application-specific image quality demands and plays a key role in mobile devices. To comply with privacy regulations and reduce the overhead of cloud computing, executing SR models locally on-device constitutes a key alternative approach. Nevertheless, the excessive compute and memory requirements of SR workloads pose a challenge in mapping SR networks on resource-constrained mobile platforms. This work presents MobiSR, a novel framework for performing efficient super-resolution on-device. Given a target mobile platform, the proposed framework considers popular model compression techniques and traverses the design space to reach the highest performing trade-off between image quality and processing speed. At run time, a novel scheduler dispatches incoming image patches to the appropriate model-engine pair based on the patch's estimated upscaling difficulty in order to meet the required image quality with minimum processing latency. Quantitative evaluation shows that the proposed framework yields on-device SR designs that achieve an average speedup of 2.13x over highly-optimized parallel difficulty-unaware mappings and 4.79x over highly-optimized single compute engine implementations.
Royson Lee, Stylianos I. Venieris, Lukasz Dudziak, Sourav Bhattacharya, Nicholas D. Lane
MobiCom4
2018 Deterministic Binary Filters for Convolutional Neural Networks
abstract
We propose Deterministic Binary Filters, an approach to Convolutional Neural Networks that learns weighting coefficients of predefined orthogonal binary basis instead of the conventional approach of learning directly the convolutional filters. This approach results in model architectures with significantly fewer parameters (4x to 16x) and smaller model sizes (32x due to the use of binary rather than floating point precision). We show our deterministic filter design can be integrated into well-known network architectures (such as ResNet and SqueezeNet) with as little as 2% loss of accuracy (under datasets like CIFAR-10). Under ImageNet, they result in 3x model size reduction compared to sub-megabyte binary networks while reaching comparable accuracy levels.
Vincent W. S. Tseng, Sourav Bhattacharya, Javier Fernández-Marqués, Milad Alizadeh, Catherine Tong, Nicholas D. Lane
IJCAI2
2018 Using deep data augmentation training to address software and hardware heterogeneities in wearable and smartphone sensing devices
abstract
A small variation in mobile hardware and software can potentially cause a significant heterogeneity or variation in the sensor data each device collects. For example, the microphone and accelerometer sensors on different devices can respond very differently to the same audio or motion phenomena. Other factors, like the instantaneous computational load on a smartphone, can cause key behavior like sensor sampling rates to fluctuate, further polluting the data. When sensing devices are deployed in unconstrained and real-world conditions, examples of sharply lower classification accuracy are observed due to what is collectively known as the sensing system heterogeneity. In this work, we take an unconventional approach and argue against solving individual forms of heterogeneity, e.g., improving OS behavior, or the quality/uniformity of components. Instead, we propose and build classifiers that themselves are more tolerant of these variations by leveraging deep learning and a data-augmented training process. Neither augmentation nor deep learning has previously been attempted to cope with sensor heterogeneity. We systematically investigate how these two machine learning methodologies can be adapted to solve such problems, and identify when and where they are able to be successful. We find that our proposed approach is able to reduce classifier errors on an average by 9% and 17% for a range of inertial-and audio-based mobile classification tasks.
Akhil Mathur, Sourav Bhattacharya, Petar Velickovic, Leonid Joffe, Nicholas D. Lane, Fahim Kawsar, Pietro Liò
IPSN3
2018 Deterministic binary filters for keyword spotting applications
abstract
We present a binary architecture with 60% fewer parameters and 50% fewer operations during inference compared to the current state of the art for keyword spotting (KWS) applications at the cost of 3.4% accuracy. We construct convolutional filters on-the-fly using orthogonal binary codes and results in a compact architecture that would fit in devices with less than 30kB of memory.
Javier Fernández-Marqués, Vincent W. S. Tseng, Sourav Bhattacharya, Nicholas D. Lane
MobiSys3
2018 Guest Editorial Special Issue on Emerging Social Internet of Things: Recent Advances and Applications
abstract
The concept of Social Internet of Things (SIoT) has emerged from the integration of social networking into the core of the Internet of Things (IoT). It envisions IoT objects and devices to have social interactions with each other autonomously, cooperate with other agents, and exchange information with human users and surrounding computing devices. These objects are able to sense/actuate, store, and interpret information in an opportunistic and loosely coupled fashion. The objects in the SIoT paradigm can exhibit multiple forms of social relationships derived from their collaborative activities or functional, temporal and spatial dependencies to meet a particular need of human users, which signify the difference between the SIoT domain to that of social-based mobile networks or sensor networks. The social interaction among the SIoT objects contribute a huge volume of data to be processed and used by various applications such as social VANET, social connected health, SIoT-based recommendation service, traffic service, policing, energy management etc, in the area of Smart Cities, Smart Homes, Smart Grid, and Smart Factories to satisfy human needs, interests, and objectives. Such a dynamic landscape with billions of social communities of objects and devices requires new models, theories, and approaches of interaction and collaboration, which could be established by referring to the experience that people have already gained in social networking domain over the past few years.
Giancarlo Fortino, Mohammad Mehedi Hassan, MengChu Zhou, Andrzej M. Goscinski, Md. Zakirul Alam Bhuiyan, Jianqiang Li 0002, Sourav Bhattacharya
IEEE Internet Things J.7
2017 DeepEye: Resource Efficient Local Execution of Multiple Deep Vision Models using Wearable Commodity Hardware
abstract
Wearable devices with built-in cameras present interesting opportunities for users to capture various aspects of their daily life and are potentially also useful in supporting users with low vision in their everyday tasks. However, state-of-the-art image wearables available in the market are limited to capturing images periodically and do not provide any real-time analysis of the data that might be useful for the wearers. In this paper, we present DeepEye - a match-box sized wearable camera that is capable of running multiple cloud-scale deep learn- ing models locally on the device, thereby enabling rich analysis of the captured images in near real-time without offloading them to the cloud. DeepEye is powered by a commodity wearable processor (Snapdragon 410) which ensures its wearable form factor. The software architecture for DeepEye addresses a key limitation with executing multiple deep learning models on constrained hardware, that is their limited runtime memory. We propose a novel inference software pipeline that targets the local execution of multiple deep vision models (specifically, CNNs) by interleaving the execution of computation-heavy convolutional layers with the loading of memory-heavy fully-connected layers. Beyond this core idea, the execution framework incorporates: a memory caching scheme and a selective use of model compression techniques that further minimizes memory bottlenecks. Through a series of experiments, we show that our execution framework outperforms the baseline approaches significantly in terms of inference latency, memory requirements and energy consumption.
Akhil Mathur, Nicholas D. Lane, Sourav Bhattacharya, Aidan Boran, Claudio Forlivesi, Fahim Kawsar
MobiSys3
2016 DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices
abstract
Breakthroughs from the field of deep learning are radically changing how sensor data are interpreted to extract the high-level information needed by mobile apps. It is critical that the gains in inference accuracy that deep models afford become embedded in future generations of mobile apps. In this work, we present the design and implementation of DeepX, a software accelerator for deep learning execution. DeepX signif- icantly lowers the device resources (viz. memory, computation, energy) required by deep learning that currently act as a severe bottleneck to mobile adoption. The foundation of DeepX is a pair of resource control algorithms, designed for the inference stage of deep learning, that: (1) decompose monolithic deep model network architectures into unit- blocks of various types, that are then more efficiently executed by heterogeneous local device processors (e.g., GPUs, CPUs); and (2), perform principled resource scaling that adjusts the architecture of deep models to shape the overhead each unit-blocks introduces. Experiments show, DeepX can allow even large-scale deep learning models to execute efficently on modern mobile processors and significantly outperform existing solutions, such as cloud-based offloading.
Nicholas D. Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Lei Jiao 0002, Lorena Qendro, Fahim Kawsar
IPSN2
2016 Demonstration Abstract: Accelerating Embedded Deep Learning Using DeepX
abstract
Deep learning has revolutionized the way sensor measurements are interpreted and application of deep learning has seen a great leap in inference accuracies in a number of fields. However, the significant requirement for memory and computational power has hindered the wide scale adoption of these novel computational techniques on resource constrained wearable and mobile platforms. In this demonstration we present DeepX, a software accelerator for efficiently running deep neural networks and convolutional neural networks on resource constrained embedded platforms, e.g., Nvidia Tegra K1 and Qualcomm Snapdragon 400.
Nicholas D. Lane, Sourav Bhattacharya, Petko Georgiev, Claudio Forlivesi, Fahim Kawsar
IPSN2
2016 Sparsification and Separation of Deep Learning Layers for Constrained Resource Inference on Wearables
abstract
Deep learning has revolutionized the way sensor data are analyzed and interpreted. The accuracy gains these approaches offer make them attractive for the next generation of mobile, wearable and embedded sensory applications. However, state-of-the-art deep learning algorithms typically require a significant amount of device and processor resources, even just for the inference stages that are used to discriminate high-level classes from low-level data. The limited availability of memory, computation, and energy on mobile and embedded platforms thus pose a significant challenge to the adoption of these powerful learning techniques. In this paper, we propose SparseSep, a new approach that leverages the sparsification of fully connected layers and separation of convolutional kernels to reduce the resource requirements of popular deep learning algorithms. As a result, SparseSep allows large-scale DNNs and CNNs to run efficiently on mobile and embedded hardware with only minimal impact on inference accuracy. We experiment using SparseSep across a variety of common processors such as the Qualcomm Snapdragon 400, ARM Cortex M0 and M3, and Nvidia Tegra K1, and show that it allows inference for various deep models to execute more efficiently; for example, on average requiring 11.3 times less memory and running 13.3 times faster on these representative platforms.
Sourav Bhattacharya, Nicholas D. Lane
SenSys1
2015 Checksum gestures: continuous gestures as an out-of-band channel for secure pairing
abstract
We propose the use of a single continuous gesture as a novel, intuitive, and efficient mechanism to authenticate a secure communication channel. Our approach builds on a novel algorithm for encoding (at least 20-bits) authentication information as a single continuous gesture, referred to as a checksum gesture. By asking the user to perform the generated gesture, a secure channel can be authenticated. Results from a controlled user experiment (N = 13 participants, 1022 trials) demonstrate the feasibility of our technique, showing over 90% success rate in establishing a secure communication channel despite relying on complex gesture patterns. The authentication times of our method are over three-folds faster than with previous gesture-based solutions. The average execution time of a gesture is 5:7 seconds in our study, which is comparable to the input time of conventional text input based PIN authentication. Our approach is particularly well-suited for scenarios involving wearable devices that lack conventional input capabilities, e.g., pairing a smartwatch with an interactive display.
Imtiaj Ahmed, Yina Ye, Sourav Bhattacharya, N. Asokan, Giulio Jacucci, Petteri Nurmi, Sasu Tarkoma
UbiComp3
2015 Smart Devices are Different: Assessing and MitigatingMobile Sensing Heterogeneities for Activity Recognition
abstract
The widespread presence of motion sensors on users' personal mobile devices has spawned a growing research interest in human activity recognition (HAR). However, when deployed at a large-scale, e.g., on multiple devices, the performance of a HAR system is often significantly lower than in reported research results. This is due to variations in training and test device hardware and their operating system characteristics among others. In this paper, we systematically investigate sensor-, device- and workload-specific heterogeneities using 36 smartphones and smartwatches, consisting of 13 different device models from four manufacturers. Furthermore, we conduct experiments with nine users and investigate popular feature representation and classification techniques in HAR research. Our results indicate that on-device sensor and sensor handling heterogeneities impair HAR performances significantly. Moreover, the impairments vary significantly across devices and depends on the type of recognition technique used. We systematically evaluate the effect of mobile sensing heterogeneities on HAR and propose a novel clustering-based mitigation technique suitable for large-scale deployment of HAR, where heterogeneity of devices and their usage scenarios are intrinsic.
Allan Stisen, Henrik Blunck, Sourav Bhattacharya, Thor S. Prentow, Mikkel Baun Kjærgaard, Anind K. Dey, Tobias Sonne, Mads Møller Jensen
SenSys3
2015 Robust and Energy-Efficient Trajectory Tracking for Mobile Devices
abstract
Many mobile location-aware applications require the sampling of trajectory data accurately over an extended period of time. However, continuous trajectory tracking poses new challenges to the overall battery life of the device, and thus novel energy-efficient sensor management strategies are necessary for improving the lifetime of such applications. Additionally, such sensor management strategies are required to provide a high and application-adjustable level of robustness regardless of the user’s transportation mode. In this article, we extend and further analyze the sensor management strategies of the EnTracked$_{T}$system that intelligently determines when to sample different on-device sensors (e.g., accelerometer, compass and GPS) for trajectory tracking. Specifically, we propose the concept of situational bounding to improve and parameterize the robustness of sensor management strategies for trajectory tracking. We demonstrate the effectiveness of our proposed approach by performing a series of emulation experiments on real world data sets collected from different modes of transportation (including walking, running, biking and commuting by car) on mobile devices from two different platforms. Thorough experimental analyses indicate that our system can save significant amounts of battery power compared to the state-of-the-art position tracking systems, while simultaneously maintaining robustness and accuracy bounds as required by diverse location-aware applications.
Sourav Bhattacharya, Henrik Blunck, Mikkel Baun Kjærgaard, Petteri Nurmi
IEEE Trans. Mob. Comput.1
2014 The company you keep: mobile malware infection rates and inexpensive risk indicators
abstract
There is little information from independent sources in the public domain about mobile malware infection rates. The only previous independent estimate (0.0009%) [11], was based on indirect measurements obtained from domain-name resolution traces. In this paper, we present the first independent study of malware infection rates and associated risk factors using data collected directly from over 55,000 Android devices. We find that the malware infection rates in Android devices estimated using two malware datasets (0.28% and 0.26%), though small, are significantly higher than the previous independent estimate. Based on the hypothesis that some application stores have a greater density of malicious applications and that advertising within applications and cross-promotional deals may act as infection vectors, we investigate whether the set of applications used on a device can serve as an indicator for infection of that device. Our analysis indicates that, while not an accurate indicator of infection by itself, the application set does serve as an inexpensive method for identifying the pool of devices on which more expensive monitoring and analysis mechanisms should be deployed. Using our two malware datasets we show that this indicator performs up to about five times better at identifying infected devices than the baseline of random checks. Such indicators can be used, for example, in the search for new or previously undetected malware. It is therefore a technique that can complement standard malware scanning. Our analysis also demonstrates a marginally significant difference in battery use between infected and clean devices.
Hien Thi Thu Truong, Eemil Lagerspetz, Petteri Nurmi, Adam J. Oliner, Sasu Tarkoma, N. Asokan, Sourav Bhattacharya
WWW7
2014 Using unlabeled data in a sparse-coding framework for human activity recognition
Sourav Bhattacharya, Petteri Nurmi, Nils Y. Hammerla, Thomas Plötz
Pervasive Mob. Comput.1
2013 An Un-tethered Mobile Shopping Experience
Venkatraman Ramakrishna, Jerome White, Nitendra Rajput, Kundan Srivastava, Sourav Bhattacharya, Yetesh Chaudhary
MobiQuitous6
2012 Ma$$iv€ - An Intelligent Mobile Grocery Assistant
abstract
We present Ma$$iv€, an intelligent mobile grocery assistant that provides support for the customer during the entire shopping process. To guide the design of Ma$$iv€, we conducted a user study that explored customer preferences regarding features in a mobile grocery aid. We first describe the study and its results, after which we introduce the design principles and design of Ma$$iv€. We also describe the features that Ma$$iv€ supports and discuss functionalities that we are integrating into Ma$$iv€. As part of the discussion, we describe technical challenges that we have encountered during our development efforts.
Sourav Bhattacharya, Patrik Floréen, Andreas Forsblom, Samuli Hemminki, Petri Myllymäki, Petteri Nurmi, Teemu Pulkkinen, Antti Salovaara
Intelligent Environments1
2011 Enriching location information: an energy-efficient approach
abstract
Off-the-shelf modern mobile devices come with a number of inbuilt sensors, e.g., GPS, WiFi, GSM, accelerometer, compass, gyroscope, NFC and Bluetooth. Equipped with all these sensors and internet connectivity, modern mobile phones are enabling continuous sensing and increasingly many emergent mobile applications are using sensed context on the phone to understand users' needs and improve usability. However, limited battery power is a big hindrance to the deployment of continuous sensing on mobile devices and without any intelligent sensor management, the battery lasts only few hours. In this research, we emphasize on location-awareness and address the challenges in developing ubiquitous positioning solutions, cross-device indoor localization, position and trajectory tracking and inferring high-level contexts using machine-learning techniques on sensor data in an energy-efficient way.
Sourav Bhattacharya
UbiComp1
2011 Influence of landmark-based navigation instructions on user attention in indoor smart spaces
abstract
Using landmark-based navigation instructions is widely considered to be the most effective strategy for presenting navigation instructions. Among other things, landmark-based instructions can reduce the user's cognitive load, increase confidence in navigation decisions and reduce the number of navigational errors. Their main disadvantage is that the user typically focuses considerable amount of attention on searching for landmark points, which easily results in poor awareness of the user's surroundings. In indoor spaces, this implies that landmark-based instructions can reduce the attention the user pays on advertisements and commercial displays, thus rendering the assistance commercially inviable. To better understand how landmark-based instructions influence the user's awareness of her surroundings, we conducted a user study with $20$ participants in a large national supermarket that investigated how the attention the user pays on her surroundings varies across two types of landmark-based instructions that vary in terms of their visual demand. The results indicate that an increase in the visual demand of landmark-based instructions does not necessarily improve the participant's recall of their surrounding environment and that this increase can cause a decrease in navigation efficiency. The results also indicate that participants generally pay little attention to their surroundings and are more likely to rationalize than to actually remember much from their surroundings. Implications of the findings on navigation assistants are discussed.
Petteri Nurmi, Antti Salovaara, Sourav Bhattacharya, Teemu Pulkkinen, Gerrit Kahl
IUI3
2011 Energy-efficient trajectory tracking for mobile devices
abstract
Emergent location-aware applications often require tracking trajectories of mobile devices over a long period of time. To be useful, the tracking has to be energy-efficient to avoid having a major impact on the battery life of the mobile device. Furthermore, when trajectory information needs to be sent to a remote server, on-device simplification of the trajectories is needed to reduce the amount of data transmission. While there has recently been a lot of work on energy-efficient position tracking, the energy-efficient tracking of trajectories has not been addressed in previous work. In this paper we propose a novel on-device sensor management strategy and a set of trajectory updating protocols which intelligently determine when to sample different sensors (accelerometer, compass and GPS) and when data should be simplified and sent to a remote server. The system is configurable with regards to accuracy requirements and provides a unified framework for both position and trajectory tracking. We demonstrate the effectiveness of our approach by emulation experiments on real world data sets collected from different modes of transportation (walking, running, biking and commuting by car) as well as by validating with a real-world deployment. The results demonstrate that our approach is able to provide considerable savings in the battery consumption compared to a state-of-the-art position tracking system while at the same time maintaining the accuracy of the resulting trajectory, i.e., support of specific accuracy requirements and different types of applications can be ensured.
Mikkel Baun Kjærgaard, Sourav Bhattacharya, Henrik Blunck, Petteri Nurmi
MobiSys2
2010 A grid-based algorithm for on-device GSM positioning
abstract
We propose a grid-based GSM positioning algorithm that can be deployed entirely on mobile devices. The algorithm uses Gaussian distributions to model signal intensity variations within each grid cell. Position estimates are calculated by combining a probabilistic centroid algorithm with particle filtering. In addition to presenting the positioning algorithm, we describe methods that can be used to create, update and maintain radio maps on a mobile device. We have implemented the positioning algorithm on Nokia S60 and Nokia N900 devices and we evaluate the algorithm using a combination of offline and real world tests. The results indicate that the accuracy of our method is comparable to state-of-the-art methods, while at the same time having significantly smaller storage requirements.
Petteri Nurmi, Sourav Bhattacharya, Joonas Kukkonen
UbiComp2
2003 Cascade of Distributed and Cooperating Firewalls in a Secure Data Network
abstract
Security issues are critical in networked information systems, e.g., with financial information, corporate proprietary information, contractual and legal information, human resource data, medical records, etc. The paper addresses such diversity of security needs among the different information and resources connected over a secure data network. Installation of firewalls across the data network is a popular approach to providing a secure data network. However, single, individual firewalls may not provide adequate security protection to meet the users needs. The cost of super firewalls, design flaws, as well as implementation inappropriateness with such firewalls may retain security loopholes. The idea proposed is to introduce a cascade of (potentially simpler and less expensive) firewalls in the secure data network, where, between the attacker node and the attacked node, multiple firewalls are expected to provide an added degree of protection. This approach, broadly following the theme of redundancy in engineering systems' design, will increase the confidence and provide more completeness in the level of security protection by the firewalls. The cascade of (i.e., multiple) firewalls can be placed across the secure data network in many ways, not all of which are equally attractive from cost and end-to-end delay perspectives. Toward this, we present heuristics for placement of these firewalls across the different nodes and links of the network in a way that different users can have the level of security they individually need, without having to pay added hardware costs or excess network delay. Three metrics are proposed to evaluate these heuristics: cost, delay, and reduction of attacker's traffic. Performance of these heuristics is presented using simulation, along with some early analytical results. Our research also extends the firewall technology into the well-known advantages of distributed firewalls. Furthermore, the distributed firewalls can be designed to cooperate and stop an attacker's traffic closest to the attack point, thereby reducing the amount of hacker's traffic into the network.
Robert N. Smith, Sourav Bhattacharya
IEEE Trans. Knowl. Data Eng.3
1999 Multimedia Tools and Applications in the Development of a Large Complex Network System
abstract
The paper demonstrates the capability and use of integrating multimedia tools and events in the life-cycle development of large, complex networks. Network engineering (NE), detailed by the authors' previous research, defines process steps for the requirement specification, design, evaluation, testing, layout and maintenance of computer networks (S. Palangala et al., 1998). The paper shows that the integration of multimedia in the process steps for NE results in better networks being designed. The paper details a multimedia based simulation executed on the network prototype, prior to submitting the designed network to a more rigorous simulation software, like OPNET (OPtimized Network Engineering Tools). The tool detailed in the paper allows the user to specify events during the network simulation. Events are user defined conditions that can identify and alert the user regarding abnormal behavior of the network system. The key ideas presented in the paper include the usage of multimedia in the NE process, and the role of user-defined events in identifying design aspects of the network.
Sourav Bhattacharya, Srihari Palangala, Tolety Siva Perraju
COMPSAC1
1999 A Protocol and Simulation for Distributed Communicating Firewalls
abstract
The concept of distributing firewalls into the Internet was previously presented for the purpose of pushing LAN attacks away from a single firewall (R.N. Smith and S. Bhattacharya, 1997; 1999). The paper presents a protocol for firewalls to communicate information to enable distributed firewalls to isolate LAN attacks. Currently firewalls are used to protect a single LAN or extranet of collaborating units. However, each firewall in these configurations are individually managed. Our approach is to place firewalls out into the Internet that will cooperate and push the attack to a firewall that is nearer to the source of the attack. These distributed firewalls can be considered as gateway firewalls. We present a protocol of command and information packets used to take the offensive in the Internet war against hackers and crackers. The communicating firewalls would be placed in routers or switches acting as gateways throughout the Internet. The proposed protocol can be encapsulated as a security agent into any one of the popular router protocols (e.g., BGP and PNNI). We have currently chosen to place our protocol over BGP-4. In order to evaluate our new protocol, we have developed a distributed network protocol simulator which we also describe.
Robert N. Smith, Sourav Bhattacharya
COMPSAC2
1999 Operating firewalls outside the LAN perimeter
abstract
Firewalls are well known for their task of securing the enterprise intranet from untrusted users attempting to gain access. The concept of firewalls got its start when routers began to be used to balance network load. The effort to balance network traffic load at the transport level was extended to the server operating system where application proxy service and application level filtering is provided. Firewalls allow selected communications data to pass from one side of the corporate network perimeter to the other side. Since the firewall is the primary entry point to a corporate LAN from the Internet, the firewall frequently comes under attack by hackers and crackers. One form of attack is "denial-of-service". "Denial-of-service" attacks are easier to detect than are attacks that allow the attacker through the firewall on a valid password that they obtained by performing social engineering. Spamming the corporate email system is one form of "denial-of-service" attack, while many other forms simply flood the firewall with useless packets to prevent other authorized users from gaining access through the firewall. The paper presents a plan to place firewalls outside the corporate network boundaries, into the Internet. By having firewalls out in the Internet acting as agents for the corporations we expect to see attackers stopped closer to their source gateway. This changes the firewall task from a defensive mode to an offensive one. By having firewalls working together to seek out and locate or block the attacker at the source gateway, we gain several benefits. The paper proposes that the gateway protocol be modified to include this filtering function.
Robert N. Smith, Sourav Bhattacharya
IPCCC2
1999 Working Set in Channel Management for Cellular Networks
abstract
Channel resource management (e.g., allocation and borrowing) is a critical aspect of wireless cellular networks. Since channel borrowing is an expensive task, incurring long delay, a key design goal of efficient channel resource management is to minimize the total number of channel borrowing operations. This paper focuses on the 'Quantum' issue (number of channels) in the channel borrowing and releasing steps. The proposed concept employs the Working Set approach, popular in Operating Systems, into channel resource management, where the quantum of channels transacted during each borrowing and release operation reaches a "middle point" and would not be too small or too large. An analytical model and steady-state simulation have been used to demonstrate the benefits.
Haihong Zheng, Sourav Bhattacharya
ISADS2
1998 Business rule extraction techniques for COBOL programs
abstract
Business rules are operational rules, often coded into software, that business organizations follow to perform various activities, such as transaction processing, quality control, business planning and database management. Over time, business rules evolve and the software that implemented them are also changed and maintained. As the encompassing software becomes large and aged, the business rules embedded are difficult to extract and understand, evolving with the encompassing software over time. Furthermore, the encompassing software is changed without changing the corresponding text documents, and thus, often, the business organization trusts the code more than any other documents. This paper proposes several techniques to extract business rules from legacy code. It is possible to use a generic software maintenance tool to extract business rules; however this can be an expensive exercise. We propose a tailored solution approach for the business rule extraction (BRE) problem, which combines variable classifications, program slicing, heuristics for identifying slicing criteria, multiple representations of business rules, bottom-up data flow analysis and hierarchical abstraction, among other maintenance techniques. The proposed solution approach has been implemented as a system and successfully tried with a number of industrial programs. © 1998 John Wiley & Sons, Ltd.
Hai Huang 0011, Wei-Tek Tsai, Sourav Bhattacharya
J. Softw. Maintenance Res. Pract.3
1997 Fault propagation analysis based variable length checkpoint placement for fault-tolerant parallel and distributed systems
abstract
The paper proposes optimal checkpoint placement strategies using failure propagation analysis in a distributed rollback recovery system. The authors' previously proposed idea of failure propagation analysis (FPA) based checkpoint placement strategy is enhanced by incorporating link failures, task grouping/allocation, and loop stabilization aspects. Owing to the empirical observation that a large number of faults occur around message communication instructions, the checkpoint placement strategy places more checkpoints around message send/receive regions of the code. Allocation of tasks (or, threads) onto different processors can lead to varied communication patterns, which in turn can affect the FPA process and the checkpoint placement strategies. Thus, another key contribution of our research is to show the cyclic relationship between checkpointing and task allocation, as well as recursion in parallel or distributed programs. The proposed ideas and FPA approaches are illustrated using a typical parallel algorithm-the fast Fourier transform (FFT).
Viral Shah, Sourav Bhattacharya
COMPSAC2
1997 Dynamic scheduling of real-time messages over an optical network
abstract
This paper proposes real-time traffic management algorithms over a multihop optical network. Our algorithms consider dynamic scheduling of (time, wavelength) slots in a multihop TWDM transmission schedule, so that a set of time-critical messages can meet their hard (and, soft) deadlines. Hard real-time messages are assigned to the highest priority level, while soft real-time messages are assigned to user-defined, progressively lower priority levels. The objective is to schedule these messages, all or as many as possible following the priority ordering. Previous research in this direction are either static scheduling policies (which cannot adapt to varying traffic conditions), or dynamic scheduling but for non-realtime traffic. Performance is measured in simulation. Where input factors include load, proximity of the deadline values, relative mix of hard and soft real-time messages, and priority levels of the messages.
Cheng-Chi Yu, Sourav Bhattacharya
ICCCN2
1997 Design and Analysis of Fault-Tolerant Star Networks
abstract
The star graph has been proposed as an attractive alternative to the hypercube offering a lower degree, a smaller diameter, and a smaller distance for a similar number of nodes. In this paper, we investigate the fault-tolerant design for the Star interconnection network using modules called fault tolerant building blocks (FTBBs). Each FTBB module contains several primary and few spare nodes. The spare nodes within each FTBB can replace the primary nodes when a failure occurs. If each spare node within an FTBB can replace any primary node we ascribe the situation to fall spare utilization. We propose fault tolerant Star networks constructed from smaller FTBBs with full spare utilization and a fault-tolerant routing scheme to reconfigure the system when a failure occurs.
Chungti Liang, Sourav Bhattacharya, Jack Tan
ICPP2
1997 Real-time multicast in wireless communication
abstract
This paper presents the reconfiguration of multicast tree at the instance of node migration in cellular wireless networks. We consider a novel, and highly practical, formulation of the real-time multicast problem. Unlike the traditional notion of source to leaf node message transmission time being the measure for real-time, we consider the multicast tree re-construction time (in the event of a node migration) as the measure for real-time constraints. The overall goal is to minimize both the delays-i): delay of re-constructing the multicast tree and ii) source to destination transmission delay. In this paper we introduce and provide solutions for the first objective, while the current research in real-time multicast tree considers the latter objective only. We propose three heuristics to solve the real-time multicast tree re-construction problem and present analysis and simulation results for different factors that contribute differently to the reconstruction delay.
Sandeepan Sanyal, Laila Nahar, Sourav Bhattacharya
ICPP3
1996 Business Rule Extraction from Legacy Code
abstract
Business rules are operational rules that business organizations follow to perform various activities. Over time, business rules evolve and the software that implemented them are also changed. As the encompassing software becomes large and aged the business rules embedded are difficult to extract and understand. Furthermore, the encompassing software is changed without changing the corresponding documents, so the business organization often trusts the code more than any other documents. It is possible to use a generic tool to extract business rules, but this can be an expensive exercise. The paper proposes a tailored solution approach to the business rule extraction problem, which combines variable classifications, program slicing, and hierarchical abstraction among other maintenance techniques. The proposed approach has been implemented as a system and successfully experimented with a number of industrial programs. The prototype has been demonstrated at several industrial software maintenance sites since June 1995.
Hai Huang 0011, Wei-Tek Tsai, Sourav Bhattacharya
COMPSAC3
1996 Dynamic Network Management for Firmware Controlled Network Topology
abstract
High assurance is a collective term implying real time security, reliability and safety. We consider a load balancing utopia for dynamic media access control (MAC) in firmware controlled network media (e.g., wireless, optical network) which provides high assurance. In a firmware controlled network, frequency and time assignment can be embedded in the logical channel on the fly-that is, logical channels with various frequency, time and code assignments can be created and updated dynamically. Consequences are that, when high assurance like real time traffic is required, we can create additional logical channels to support it without leading to congestion on the existing logical channels.
L. K. Nahar, Sourav Bhattacharya
COMPSAC2
1996 Marriage of Wired and Wireless Networks to Build Tomorrows Internet
abstract
Summary form only given. Tomorrow's Internet will neither be completely wired nor be completely wireless. Our proposed solution is the interim of the wired and wireless network, where roaming nodes are connected back to the wired media when they reach a wired site. The idea is to utilize the large bandwidth infrastructure available over the wired backbone network to connect all static nodes and at the same time use the wireless media for only those nodes which are in transition from one wired site to another.
Sandeepan Sanyal, Sourav Bhattacharya
COMPSAC2
1996 Software Engineering Practices and Tools for Real-Time Systems (Part I): Guest Editor's Introduction
Sourav Bhattacharya, Ramin Mojdehbakhsh, Wei-Tek Tsai
Int. J. Softw. Eng. Knowl. Eng.1
1996 Software Engineering Practices and Tools for Real-Time Systems (Part II): Guest Editor's Introduction
Sourav Bhattacharya, Ramin Mojdehbakhsh, Wei-Tek Tsai
Int. J. Softw. Eng. Knowl. Eng.1
1995 Real-Time Communication Benchmark
Sourav Bhattacharya
ICCCN1
1995 Optimal multihop routing. An iterative approach to TWDM embedding
abstract
This paper introduces the concept of time slot synchronization in time-wave division multiplexed (TWDM) multihop lightwave networks. It is shown that the time slot assignments of the intermediate nodes in a multihop path have significant effect on the end-to-end message delay. This is different from traditional notions in routing, where the hop count and congestion control are the primary concerns. Assuming a TWDM embedding of a given logical topology already exists, we formulate the optimal routing problem for arbitrary propagation delays. We propose a graph unfolding technique which converts this problem into the shortest path routing problem for weighted graphs with well known solutions. We show a method to estimate the buffering cost at the intermediate nodes (in multihop routing) to accommodate high-bandwidth traffic (e.g., video). Next, we address the question: given a logical topology what should the TWDM embedding be so that the routing delays are optimized P This leads us to an iterative approach for TWDM embedding. Given an initial embedding an optimal route is estimated for each pair of nodes. A weighted average of these optimal route distances is used as a metric to evaluate the TWDM embedding. Using a heuristic the TWDM embedding is modified and the process iterated to improve along this metric. Preliminary performance results illustrating the improvement in routing delay are provided. For example, for a 4-cube network on average delay minimization of 10% to 20% is observed.
Sourav Bhattacharya, Aloke Guha, Allalaghatta Pavan, David Hung-Chang Du
LCN1
1995 Covert Channel Secure Hypercube Message Communication
Sourav Bhattacharya, Thomas F. Keefe, Wei-Tek Tsai
J. Parallel Distributed Comput.1
1994 Recursive Binary Tree Layout Mixing
Sourav Bhattacharya, Wei-Tek Tsai
Inf. Sci.1
1994 Multicasting in Generalized Multistage Interconnection Networks
Sourav Bhattacharya, Gary Elsesser, Wei-Tek Tsai, Ding-Zhu Du
J. Parallel Distributed Comput.1
1994 Fault-Tolerant Multicasting on Hypercubes
Albert C. Liang, Sourav Bhattacharya, Wei-Tek Tsai
J. Parallel Distributed Comput.2
1994 Processor preallocation and load balancing of DOALL loops
Gary Elsesser, Viet N. Ngo, Sourav Bhattacharya, Wei-Tek Tsai
J. Supercomput.3
1993 Reverse Channel Augmented Multihop Lightwave Networks
abstract
The idea of providing reverse channels with minimal hardware costs in Shuffle-nets to reduce the network diameter improve mean delay, and provide other advantages is discussed. The reverse channel idea can be applied to any multistage network with wrapped around connections. The advantages of reverse channels are demonstrated by providing a simple reverse connection between adjacent stages of the Shuffle-net. It is shown how a time- and wavelength-division-multiplexed media access protocol for the Shuffle-net can be easily adapted for the reverse channel augmented Shuffle-net (RC-Shuffle-net). The performance of the RC-Shuffle-net and its advantages over the Shuffle-net are shown.>
Allalaghatta Pavan, Sourav Bhattacharya, David Hung-Chang Du
INFOCOM2
1992 Quadtree interconnection network layout
abstract
Quadtree data structure has been used in a number of applications. However, VLSI embedding of quadtree based parallel architecture using grid model has not been studied. This paper studies VLSI embedding of quadtree using grid model. H-tree layout for binary tree is extended for trivial quadtree layout, followed by two layout strategies for rectangular grids. Two generic layout styles (standard layout and X-layout) are proposed for higher order grids (e.g., hexagonal and octagonal grids). Base tile layout patterns are proposed for area compaction with recursive X-layout. In each case, layout dimensions and I/O bandwidth are computed. The authors demonstrate how the two generic layouts can be mixed to obtain higher I/O bandwidth and estimate the area sacrifice. An improved recursive layout mixing strategy is proposed.>
Sourav Bhattacharya, Shekhar H. Kirani, Wei-Tek Tsai
Great Lakes Symposium on VLSI1
1992 Array Covering: A Technique4 for Enabling Lloop Parallelization
Viet N. Ngo, Gary Elsesser, Sourav Bhattacharya, Wei-Tek Tsai
ICPP (2)3
1991 I/O bound binary tree layout
abstract
The authors propose a VLSI layout strategy for a full binary tree. This layout can support more border leaf processing elements (PEs) and thus can give a higher I-O bandwidth. It is superior to the H-tree layout in terms of the number of boundary leaves. The approach uses H-tree pattern for constructing subtree layouts and then combines a number of such subtrees following standard tree style to get a larger sized tree layout. Finally at the top level H-tree layout style is used to get the overall tree layout. Different I/O bandwidths can be obtained varying the subtree height. The authors derive expression for layout area, longest link and aspect ratio of the chip. It is observed that I/O bandwidth can be significantly increased without much area overhead using this approach.>
Sourav Bhattacharya, Yoon-Hwa Choi, Wei-Tek Tsai
Great Lakes Symposium on VLSI1
1991 Uni-directional cube-connected cycles
abstract
Cube connected cycles (CCC), a popular and layout/efficient alternative to hypercube, can emulate the performance of hypercube for many parallel algorithms. Recently, interconnection networks based on simplex links rather than duplex have been proposed. The uni-directional architectures have layout advantages and reduces complexity of each processing element (PE). The authors propose directed cube connected cycles (DCCC) as a uni-directional alternative to CCC. They have developed PE-to-PE routing algorithm for DCCC. A method for porting algorithms (designed to run on bi-directional CCC) to DCCC is provided. The extent of slowdown due to using simplex links is evaluated. They also provide loop embedding on DCCC. DCCC is found competitive to CCC in algorithmic performance though DCCC is much layout inexpensive.>
Sourav Bhattacharya, Yoon-Hwa Choi, Wei-Tek Tsai
Great Lakes Symposium on VLSI1
1991 Area efficient binary tree layout
abstract
H-Tree layout for binary trees can utilize only 50% of the available nodes. Improved binary tree layout techniques have been developed only after relaxing the rectangular grid model assumptions. The authors propose an area-efficient VLSI layout strategy for full binary trees without relaxing the rectangular grid model assumptions. For a height-5 full binary tree they developed a (5*8) layout pattern on an ad hoc basis. This tile is more area efficient than an equivalent H-Tree layout of a height-5 full binary tree. Using this tile, higher level trees are built in a way identical to H-Tree. The area efficiency remains for any level of tree construction. The proposed layout has an improved aspect ratio compared with H-Tree and features a reduced length of the longest link.>
Sourav Bhattacharya, Wei-Tek Tsai
Great Lakes Symposium on VLSI1
1991 Inverted Memory
Sourav Bhattacharya, Chungti Liang, Wei-Tek Tsai
ICPP (1)1
1990 Fault-Tolerant Cube-Connected Cycles Structures Through Dimensional Substitution
Nian-Feng Tzeng, Sourav Bhattacharya, Po-Jen Chuang
ICPP (1)2