EDBT 2026 Demo / reviewers in the wild / expert
Shirish S. Karande
dblp:136/8377 · also Shirish Karande 0001, Shirish Subhash Karande
· DBLP profile ↗
49ranked-venue papers
20as first author
8since 2021 · last 2026
0009-0000-2693-845XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 7 since 2021Computer networks · 15 · 8 first-authorArtificial intelligence and machine learning · 5 · 2 since 2021Theory of computation · 3 · 2 first-authorSecurity and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLARIS: Clear and Intelligible Speech from Whispered and Dysarthric VoicesabstractWhispered and dysarthric speech hinder effective communication and undermine the reliability of voice-enabled systems. We present CLARIS, a compact speech-to-speech restoration system that turns such atypical input into clear, expressive speech. CLARIS requires no disorder-specific architectural tuning, generalizes across languages, and adapts quickly to new accents and speakers, enabling practical personalization. On whispered English, Hindi, and clinically challenging dysarthric speech, CLARIS delivers state-of-the-art intelligibility and naturalness, with listener studies confirming gains in quality, intelligibility, naturalness, and prosody. The system runs in real time, converting one second of input in about 30ms and enables inclusive, private, and personalized voice interaction. Audio samples are available at https://claris-w2s.github.io/CLARIS/ Neil Shah, Yash Sonkar, Shirish S. Karande, Vineet Gandhi |
CHI | 3 |
| 2026 | DoTA: Latent Distribution Conditioned Data Attribution for Diffusion ModelsabstractDiffusion models have emerged as the backbone of several modern generative AI models for effective visual content generation. However, their opaque nature raises fundamental questions about which training samples are responsible for specific generations, especially in applications involving bias detection, model auditing, and dataset curation. Data attribution seeks to identify the training samples that highly influence the output of generative models, a task that becomes especially challenging when targeting fine-scale attributes for attribution. Prior work has focused on broad concepts such as global features or entire images, often overlooking the nuances of fine-grained attributes and relying on group-based strategies that dilute individual influence. We propose a novel latent distribution conditioned method DoTA for data attribution. DoTA presents an effective search space pruning technique based on the latent distribution matching between the generated and training data for effective and controlled attribution. We demonstrate the effectiveness of attribution through extensive quantitative and qualitative evaluations across challenging settings such as counterfactual evaluation and robustness to adversarial attack. Ninad Joshi, Shirish S. Karande |
WACV | 3 |
| 2025 | Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM DatasetabstractCurrent Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well across different speakers. To address this issue, we focus on learning phoneme-level alignments from paired whispers and text and employ a Text-to-Speech (TTS) system to simulate the ground-truth. To reduce dependence on whispers, we learn phoneme alignments directly from NAMs, though the quality is constrained by the available training data. To further mitigate reliance on NAM/whisper data for ground-truth simulation, we propose incorporating the lip modality to infer speech and introduce a novel diffusion-based method that leverages recent advancements in lip-to-speech technology. Additionally, we release the MultiNAM dataset with over 7.96 hours of paired NAM, whisper, video, and text data from two speakers and benchmark all methods on this dataset. Speech samples and the dataset are available at https://diff-nam.github.io/DiffNAM/ Neil Kumar Shah, Shirish S. Karande, Vineet Gandhi |
ICASSP | 2 |
| 2025 | MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRIabstractPrevious real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in poor intelligibility. We introduce a novel approach that adapts the multi-modal self-supervised AV-HuBERT model for text prediction from rtMRI and incorporates a new flow-based duration predictor for speaker-specific alignment. The predicted text and durations are then used by a speech decoder to synthesize aligned speech in any novel voice. We conduct thorough experiments on two datasets and demonstrate our method’s generalization ability to unseen speakers. We assess our framework’s performance by masking parts of the rtMRI video to evaluate the impact of different articulators on text prediction. Our method achieves a 15.18% Word Error Rate (WER) on the USC-TIMIT MRI corpus, marking a huge improvement over the current state-of-the-art. Speech samples are available at https://mri2speech.github.io/MRI2Speech/ Neil Kumar Shah, Ayan Kashyap, Shirish S. Karande, Vineet Gandhi |
ICASSP | 3 |
| 2025 | HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian LanguagesabstractWe present a multilingual, multi-speaker, low-latency speech synthesis system developed by the HamaraAwaz team for Track 1 of the LIMMITS’25 challenge. To improve speaker similarity and naturalness in Indic languages, we build on ParrotTTS. We utilize disentangled self-supervised speech representations and incorporate enhancements such as Byte-Pair Encoding for text representation to reduce latency and relative positional representations to enhance speech quality. The proposed model achieved a naturalness Mean Opinion Score (MOS) of 3.51 and a speaker similarity score of 3.53 in the LIMMITS’25 grand challenge held as part of ICASSP-25. Speech samples are available at https://parrot-tts.github.io/HamaraAwaz/ Neil Kumar Shah, Parth Khadse, Shirish S. Karande, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2025 | NAM-to-Speech Conversion with Multitask-Enhanced Autoregressive Models
Neil Shah, Shirish S. Karande, Vineet Gandhi |
INTERSPEECH | 2 |
| 2025 | VisualFusion: Enhancing Blog Content with Advanced Infographic PipelineabstractInfographics represent a key component of any blog or article, facilitating effective communication of ideas while fos-tering reader engagement. However, many content creators possess limited expertise in crafting visually striking info-graphics. This gap is effectively addressed by our proposed pipeline, designed to aid writers in generating compelling infographics tailored to their written content. Our pipeline uses textual content and tabular data from the blog to gen-erate anchor plots. Leveraging LLM for prompt generation, the pipeline integrates the generated prompts with these anchor plots through a Image to Image (121) generation Model. We observe that majority of resulting images generated using this approach align with the article's narrative and effectively represent the underlying tabular data. Additionally, we introduce our proposed AADaT (Aesthetical Adherence to Data and Text) Score, adept at comprehen-sively assessing aesthetics, textual alignment, data fidelity, and overall image quality concurrently. In comparative evaluations, our pipeline has demonstrated around 15% superior performance relative to state-of-the-art models such as DALL-e and Stable Diffusion Large by showcasing much better data adherence and aesthetics. While state-of-the-art models excel in some metrics but falter in others, our pipeline demonstrates a balanced performance across all metrics. The source code and data corpus may be available on request. Anurag Deo, Savita Bhat, Shirish S. Karande |
WACV | 3 |
| 2024 | Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
Neil Kumar Shah, Shirish S. Karande, Vineet Gandhi |
INTERSPEECH | 2 |
| 2020 | Understanding Advertisements with BERTabstractWe consider a task based on CVPR 2018 challenge dataset on advertisement (Ad) understanding.The task involves detecting the viewer's interpretation of an Ad image captured as text.Recent results have shown that the embedded scene-text in the image holds a vital cue for this task.Motivated by this, we fine-tune the base BERT model for a sentencepair classification task.Despite utilizing the scene-text as the only source of visual information, we could achieve a hit-or-miss accuracy of 84.95% on the challenge test data.To enable BERT to process other visual information, we append image captions to the scene-text.This achieves an accuracy of 89.69%, which is an improvement of 4.7%.This is the best reported result for this task. Kanika Kalra, Bhargav Kurma, Silpa Vadakkeeveetil Sreelatha, Manasi Patwardhan 0001, Shirish S. Karande |
ACL | 5 |
| 2018 | Towards automating disambiguation of regulations: using the wisdom of crowdsabstractCompliant software is a critical need of all modern businesses. Disambiguating regulations to derive requirements is therefore an important software engineering activity. Regulations however are ridden with ambiguities that make their comprehension a challenge, seemingly surmountable only by legal experts. Since legal experts' involvement in every project is expensive, approaches to automate the disambiguation need to be explored. These approaches however require a large amount of annotated data. Collecting data exclusively from experts is not a scalable and affordable solution. In this paper, we present the results of a crowd sourcing experiment to collect annotations on ambiguities in regulations from professional software engineers. We discuss an approach to automate the arduous and critical step of identifying ground truth labels by employing crowd consensus using Expectation Maximization (EM). We demonstrate that the annotations reaching a consensus match those of experts with an accuracy of 87%. Manasi Patwardhan 0001, Abhishek Sainani, Shirish S. Karande, Smita Ghaisas |
ASE | 4 |
| 2017 | Deep Learning Based Car Damage ClassificationabstractImage based vehicle insurance processing is an important area with large scope for automation. In this paper we consider the problem of car damage classification, where some of the categories can be fine-granular. We explore deep learning based techniques for this purpose. Initially, we try directly training a CNN. However, due to small set of labeled data, it does not work well. Then, we explore the effect of domain-specific pre-training followed by fine-tuning. Finally, we experiment with transfer learning and ensemble learning. Experimental results show that transfer learning works better than domain specific fine-tuning. We achieve accuracy of 89.5% with combination of transfer and ensemble learning. Kalpesh Patil, Mandar Kulkarni, Anand Sriraman, Shirish S. Karande |
ICMLA | 4 |
| 2016 | Unsupervised Word Clustering Using Deep FeaturesabstractDigitization is crucial especially in the Indian context. OCR engines fail on Indian scripts mainly because character segmentation is non-trivial. Even word based recognition approaches suffer from the issues such as time degradations, word segmentation errors, font style/size variations. In this paper, we propose a deep learning architecture based approach for unsupervised word clustering. An edge responsive untrained Convolutional Neural Network (CNN) is used as a feature extractor. Graph connected component analysis is applied on the similarity graph computed from the word features. Our approach inherently detects similar shape patterns at word level and hence, it is language agnostic. We validated our approach against multiple state of art word matching techniques. Experimental results show that our approach significantly outperforms all of them on variety of data sets. In addition, the approach is observed to be robust to word segmentation errors, font style/size variations. Mandar Kulkarni, Shirish S. Karande, Sachin Lodha |
DAS | 2 |
| 2016 | On Employing a Highly Mismatched Crowd for Speech Transcription
Purushotam G. Radadia, Kanika Kalra, Shirish S. Karande, Sachin Lodha |
INTERSPEECH | 4 |
| 2015 | Multi-view image inpainting with sparse representationsabstractThis paper proposes a patch based image inpainting algorithm for multi-view images. In our framework we fill the holes which are created by removing objects from an image pair. We assume that a user provides two masks to remove objects from an image pair. Our algorithm consists of two stages. In the first stage we align the images and construct an exemplar dictionary with the patches sampled from the reference as well as the warped image. In the second stage, the reference image is iteratively filled by choosing patches along the boundary of the hole. We use l1-minimization framework to estimate the unknown pixels. This proposed method is observed to outperform existing techniques that are built upon exemplar based sparse reconstruction. Sandhya Thaskani, Shirish S. Karande, Sachin Lodha |
ICIP | 2 |
| 2014 | Breaching IM session privacy using causalityabstractThe breach of privacy in encrypted instant messenger (IM) service is a serious threat to user anonymity. Performance of previous de-anonymization strategies was limited to 65%. We perform network de-anonymization by taking advantage of the cause-effect relationship between sent and received packet streams and demonstrate this approach on a data set of Yahoo! IM service traffic traces. An investigation of various measures of causality shows that IM networks can be breached with a hit rate of 99%. A KCI Causality based approach alone can provide a true positive rate of about 97%. Individual performances of Granger, Zhang and IGCI causality are limited owing to the very low SNR of packet traces and variable network delays. Saad Saleh, Mamoon Raja, Muhammad Shahnawaz, Muhammad Usman Ilyas, Khawar Khurshid, Zubair Shafiq, Alex X. Liu, Hayder Radha, Shirish S. Karande |
GLOBECOM | 9 |
| 2013 | On the real-time masking of the sound of credit cards using hot patchingabstractPhone based card payments utilize inband DTMF signaling to convey data. Since the DTMF signals are audible to a human ear, a call operator is in position to carry out a privacy attack. We investigate real-time techniques that can obfuscate the 'digit' values without deteriorating the voice quality. Furthermore, we consider a setting where the privacy solution is being provided by a third party which does not have the benefit of open interfaces to the communication application. Our experiments reveal the efficacy of binary interception to 'inject' the signal filtering. Meanwhile, we observe that several DTMF suppression techniques that have been proposed in literature can leave a residue that is sufficient for de-anonymizing the digit value. In light of these observations, we argue in favor of more modest privacy guarantees, which can be achieved by suppressing only the higher frequency. We show that margin crossings and peak variances can be used for fast pre-filtering of audio to detect the presence of a tone, thus reducing the computational needs. Manish Shukla 0001, Purushotam G. Radadia, Shirish S. Karande, Sachin Lodha |
CCS | 3 |
| 2012 | Retargeting LT codes using XORs at the relayabstractWe consider a network where multiple sources communicate via a single relay to sinks with non-uniform and unequal demands. The LT distributions employed at the sources may be ill-matched to the demands of the sinks connected to the relay. We consider a probabilistic “morphing” of two or more fountain-encoded streams into a single stream better suited for the demand patterns downstream of the relay. The relay observes symbols generated from two distinct fountain codes, and can decide to forward a symbol from one source, or the other source, or the X-OR of the two symbols. We propose a linear programming based design of Generalized LT codes, which with appropriate substitutions is utilized to design the probabilities for the relay. Simulation results show that the designs obtained on the basis of the proposed optimization problem, when compared with multiplexing or simple mixing, reduce the number of symbols that have to be downloaded to guarantee a desired probability of successful decoding. Shirish S. Karande |
ITW | 1 |
| 2011 | Multicast Throughput Order of Network Coding in Wireless Ad-hoc NetworksabstractWe consider a network with n nodes distributed uniformly in a unit square. We show that, under the protocol model, when ns= Ω (log(n)1+α) out of the n nodes, each act as source of independent information for a multicast group consisting of m randomly chosen destinations, the per-session capacity in the presence of network coding (NC) has a tight bound of Θ(√n/ns√mlog(n)) when m = O(n/log(n)) and Θ(1/ns) when m = Ω(n/log(n)). In the case of the physical model, we consider ns= n and show that the per-session capacity under the physical model has a tight bound of Θ(1/√mn) when m = O(n/(log(n))3), and Θ(1/n) when m = Ω(n/log(n)). Prior work has shown that these same order bounds are achievable utilizing only traditional store-and-forward methods. Consequently, our work implies that the network coding gain is bounded by a constant for all values of m. For the physical model we have an exception to the above conclusion when m is bounded by O(n/(log(n))3) and Ω(n/log(n)). In this range, the network coding gain is bounded by O((log(n))1/2). Shirish S. Karande, Zheng Wang 0006, Hamid R. Sadjadpour, J. J. Garcia-Luna-Aceves |
IEEE Trans. Commun. | 1 |
| 2010 | Maximal Recovery Network Coding Under Topology ConstraintabstractNetwork coding (NC) within wireless sensor networks (WSNs) can be viewed as the mapping of efficient channel codes to the data generated within the network. In particular, this perspective of code-on-network-graphs (CNG) can be exploited to map source data generated within WSN (of size K) to a variable nodes subset in low-density parity check (LDPC) codes. The resulting fixed size symbol stream when transmitted through the network suffers erasures. At sink, an average ofzsource symbols can be recovered by employing belief propagation decoding. In this paper, we determine CNG code ensembles that achieve maximal recovery (z/K) for different erasure rates and network topological constraints corresponding to node transmission range. An analytic framework to predict code performance under transmission range constraints is developed. Additionally, necessary condition for code stability was derived using fixed-point stability analysis. Optimal solutions for a WSN with 1000 nodes are determined using differential evolution algorithm. We outline a distributed algorithm for generating a sequence of encoded symbols adhering to the designed code ensemble. The performance of the designed CNG code is demonstrated to be superior to random NC and growth code based ensembles, as well as resilient to network size and inter-connectivity variations. Kiran Misra, Shirish S. Karande, Hayder Radha |
IEEE Trans. Inf. Theory | 2 |
| 2009 | Network Coding Does Not Change the Multicast throughput Order of Wireless Ad Hoc NetworksabstractWe demonstrate that the gain attained by network coding (NC) on the multicast capacity of random wireless ad hoc networks is bounded by a constant factor. We consider a network with n nodes distributed uniformly in a unit square, with each node acting as a source for independent information to be sent to a multicast group consisting of m randomly chosen destinations. We show that, under the protocol model, the per- session capacity in the presence of arbitrary NC has a tight bound of Theta (1/radic(mnlog(n))) when m = O(n/(log(n))) and Theta(1/n) when m = Omega(n/(log(n))). Our result follows from the fact that prior work has shown that the same order bounds are achievable with pure routing based only on traditional store-and-forward methods. Shirish S. Karande, Zheng Wang 0006, Hamid R. Sadjadpour, J. J. Garcia-Luna-Aceves |
ICC | 1 |
| 2009 | Multicast Throughput Order of Network Coding in Wireless Ad-hoc NetworksabstractWe show that network coding (NC) does not provide any order gain in the multicast capacity of random wireless ad hoc networks. We consider a network with n nodes distributed uniformly in a unit square, with each node acting as a source for independent information to be sent to a multicast group consisting of m randomly chosen destinations. We show that, in the presence of NC, the per-session capacity under the protocol model has a tight bound of Theta (1/(mnlog(n))) when m = O (n/log(n)) Theta (1/n) when m = Omega (n/log/n). Furthermore, we show that the per-session capacity under the physical model has a tight bound of Theta (1/(mn)) when m = O (n/(log(n))3), and Theta (1/n) when m = Omega (n/log(n)). Prior work has shown that these same order bounds are achievable utilizing only traditional store-and- forward methods. Shirish S. Karande, Zheng Wang 0006, Hamid R. Sadjadpour, J. J. Garcia-Luna-Aceves |
SECON | 1 |
| 2009 | Optimal Unicast Capacity of Random Geometric Graphs: Impact of Multipacket Transmission and ReceptionabstractWe establish a tight max-flow min-cut theorem for multi-commodity routing in random geometric graphs. We show that, as the number of nodes in the network n tends to infinity, the maximum concurrent flow (MCF) and the minimum cut-sparsity scale as ¿(n2r3(n)/k), for a random choice of k = ¿(n) source-destination pairs, where n and r(n) are the number of nodes and the communication range in the network respectively. The MCF equals the interference-free capacity of an ad-hoc network. We exploit this fact to develop novel graph theoretic techniques that can be used to deduce tight order bounds on the capacity of ad-hoc networks. We generalize all existing capacity results reported to date by showing that the per-commodity capacity of the network scales as ¿(1/r(n)k) for the single-packet reception model suggested by Gupta and Kumar, and as ¿(nr(n)/k) for the multiple-packet reception model suggested by others. More importantly, we show that, if the nodes in the network are capable of (perfect) multiple-packet transmission (MPT) and reception (MPR), then it is feasible to achieve the optimal scaling of ¿(n2r3(n)/k), despite the presence of interference. In comparison to the Gupta-Kumar model, the realization of MPT and MPR may require the deployment of a large number of antennas at each node or bandwidth expansion. Nevertheless, in stark contrast to the existing literature, our analysis presents the possibility of actually increasing the capacity of ad-hoc networks with n even while the communication range tends to zero! J. J. Garcia-Luna-Aceves, Zheng Wang 0006, Hamid R. Sadjadpour, Shirish S. Karande |
IEEE J. Sel. Areas Commun. | 4 |
| 2009 | Fundamental limits of information dissemination in wireless ad hoc networks-part I: single-packet receptionabstractWe present the first unified modeling framework for the computation of the capacity-delay tradeoff of random wireless ad hoc networks. This framework considers information dissemination by means of unicast routing, multicast routing, broadcasting, or different forms of anycasting. We introduce (n, m, k) -casting as a generalization of all forms of one-toone, one-to-many, and many-to-many information dissemination in wireless networks. In this context, n, m, and k denote the total number of nodes in the network, the number of destinations for each communication group, and the actual number of communication-group members that receive information (k ¿ m), respectively. We describe the capacity-delay tradeoff for (n, m, k) -casting in wireless ad hoc networks in which receivers perform single-packet reception (SPR). Our results are consistent with prior results in wireless networks and extend them to the general (n,m,k) -cast case. Zheng Wang 0006, Hamid R. Sadjadpour, J. J. Garcia-Luna-Aceves, Shirish S. Karande |
IEEE Trans. Wirel. Commun. | 4 |
| 2008 | Complexity reduction using power-law based scheduling for exploiting spatial correlation in distributed video codingabstractIn pixel-domain distributed video coding (DVC), due to the largely translational nature of motion, residue errors in the side-information frame are often clustered together. These clusterings can be exploited to reduce the number of syndrome bits required to successfully perform low density parity check (LDPC) decoding, and therefore improve the overall rate-distortion performance. We shall see that using alternate iterations of LDPC syndrome decoding and Baum-Welch channel estimation proves to be an efficient scheme for exploiting the spatial clustering of errors in pixel-domain DVC. In this paper we demonstrate that a sparser power-law based scheduling of the channel estimation iteration leads to significant reduction in estimation complexity (around 83% reduction) for a small loss in rate-distortion performance (less than 0.75 dB). This sparser scheduling of channel estimation iterations can potentially improve decoding delays. Kiran Misra, Shirish S. Karande, Keyur Desai, Hayder Radha |
ICIP | 2 |
| 2008 | Maximal Recovery Network Coding under Topology ConstraintabstractRecent advances have shown that channel codes can be mapped onto networks to realize efficient Network Coding (NC); this has led to the emergence of Code-on-Network-Graphs (CNG). Traditional CNG approaches (e.g Decentralized Erasure Codes) focus on a generating a sequence of encoded symbols from a given input source (of size K), such that the original symbols can be recovered from any subset of the encoded symbols of size equal to or slightly larger than K. However in all cases the number of source symbols recovered falls rapidly if the number of encoded symbols received falls below K. In this paper we determine the CNG code-ensembles (under statistical toplogy constraint) which result in maximal recovery of WSN source data (for different erasure-rates), thereby minimizing the deterioration in data recovery. We also perform fixed point stability analysis on the underlying LDPC code ensemble. We then propose a distributed algorithm for generating a sequence of encoded symbols adhering to the designed code ensemble. Optimal solutions for a sensor network with 1000 nodes is determined using the Differential Evolution algorithm, and the solution sensitivity to variance in number of sensor nodes and node-interconnectivity is evaluated. Kiran Misra, Shirish S. Karande, Hayder Radha |
INFOCOM | 2 |
| 2008 | Design and analysis of Generalized LT-codes using colored ripplesabstractResearch has shown that fluid limits of Markov processes can be used to obtain closed form expressions for the evolution of the ripple-size. In this work we extend the above analysis to generalized LT (GLT) codes, which can be used to represent LT encoding (with priorities) over multiple data segments. In our analysis, we segregate the ripple into multiple colored ripples, where each color corresponds to a segment. We derive closed form expressions for the size of each ripple. We utilize these expressions to design GLT distributions, optimized for a desired intermediate and unequal recovery. Shirish S. Karande, Kiran Misra, Sohraab Soltani, Hayder Radha |
ISIT | 1 |
| 2008 | Optimal scaling of multicommodity flows in wireless ad hoc networks: Beyond the Gupta-Kumar barrierabstractWe establish a tight max-flow min-cut theorem for multicommodity routing in random geometric graphs. We show that, as the number of nodes in the network n tends to infinity, the maximum concurrent flow (MCF) and the minimum cut-capacity scale as Theta(n2r3(n)/k) for a random choice of k ges Theta(n) source-destination pairs, where r(n) is the communication range in the network. We exploit the fact, that the MCF in a random geometric graph equals the interference-free capacity of an ad-hoc network under the protocol model, to derive scaling laws for interference-constrained network capacity. We generalize all existing results reported to date by showing that the per-commodity capacity of the network scales as Theta(1/r(n)k) for the single-packet reception model suggested by Gupta and Kumar, and as Theta(nr(n)/k) for the multiple-packet reception model suggested by others. More importantly, we show that, if the nodes in the network are capable of multiple-packet transmission and reception, then it is feasible to achieve the optimal scaling of Theta(n2r3(n)/k), despite the presence of interference. This result provides an improvement of Theta(nr2(n)) over the highest achieved capacity reported to date. In stark contrast to the conventional wisdom that has evolved from the Gupta-Kumar results, our results show that the capacity of ad-hoc networks can actually increase with n while the communication range tends to zero! Shirish S. Karande, Zheng Wang 0006, Hamid R. Sadjadpour, J. J. Garcia-Luna-Aceves |
MASS | 1 |
| 2008 | On Channel Capacity Estimation and Prediction for Rate-Adaptive Wireless VideoabstractPacket drops caused byresidueerrors(MAC-layer errors) can severely deteriorate the wireless video quality. Prior studies have shown that this loss of quality can be circumvented by using forward error correction (FEC) to recover information from the corrupted packets. The performance of FEC encoded video streaming is critically dependent upon the choice of source and channel coding rates. In practice, the wireless channel conditions can vary significantly, thus altering the optimal rate choices. Thus, it is essential to develop an architecture which can estimate the channel capacity and utilize this estimate for rate allocation. In this paper we develop such a framework. Our contributions consist of two parts. In the first part we develop a prediction framework that leverages the received packets' signal to silence ratio (SSR) indications and MAC-layer checksum as side information to predict the operational channel capacity. In the second part, we use this prediction framework for rate allocation. The optimal rate allocation is dependent upon thechannelcapacity, the distribution of the (capacity)predictionerrorand the rate-distortion (RD) characteristics of the video source. Consequently, we propose a framework that utilizes the aforementioned statistics for RD optimal rate adaptation. We exhibit the efficacy of the proposed scheme by simulations using actual 802.11b wireless traces, an RD model for the video source and an ideal FEC model. Simulations using source RD models derived from five different popular video codecs (including H.264), show that the proposed framework provides up-to 5-dB improvements in peak signal-to-noise ratio (PSNR) when compared with conventional rate-adaptive schemes. Yongju Cho, Shirish S. Karande, Kiran Misra, Hayder Radha, Jeongju Yoo, Jinwoo Hong |
IEEE Trans. Multim. | 2 |
| 2007 | On Channel State Inference and Prediction Using Observable Variables in 802.11b NetworkabstractPerformance of cross-layer protocols that recommend the relay of corrupted packets to higher layers can be improved significantly by accurately inferring/predicting the bit error rate (BER) in the packets. In practice, higher layers observe the bits only after some hard decision. Hence physical layer link-quality indications, such as the signal strength of each individual bit, are not observable at higher layers. Therefore, it is essential to identify practically observable variables, which can be used for reasonably robust channel state inference/prediction (CSI/CSP). Here, inference specifically refers to estimating the BER in an already received packet, while prediction refers to anticipating the BER in a future packet. In this paper, we note that, in practical 802.11b devices, it is possible to acquire a Signal to Silence Ratio (SSR) indication and measure the background traffic intensity (p) on a per packet basis. This paper, thus presents a measurement-based study that analyzes the utility of SSR andpas side-information for CSI/CSP. In this work, we exploit the method of types to measure the robustness of the observable side-information. Our analysis and simulations based on an extensive set of actual 802.11b traces exhibit the practical utility of the considered observable variables. Shirish S. Karande, Syed Ali Khayam, Yongju Cho, Kiran Misra, Hayder Radha, Jae-Gon Kim, Jinwoo Hong |
ICC | 1 |
| 2007 | Optimally Mapping an Iterative Channel Decoding Algorithm to a Wireless Sensor NetworkabstractRetransmission based schemes are not suitable for energy constrained wireless sensor networks. Hence, there is an interest in including parity bits in each packet for error control. From an information-theoretic perspective the most efficient usage of network capacity can be achieved by performing full encoding/decoding at each node and using a variable rate in accordance with the link-quality. However, such an approach represents a major burden on power-constrained sensors. In this paper, we propose a more practical approach that is based on optimally distributing iterative channel decoding over sensor networks. In such a paradigm, the guarantee with which the base station, orcollector, gets the data from a sensor is a function of the processing within the intermediate nodes between source and destination (in-network processing). There are two extreme cases: a) Complete channel decoding at each hop and b) decoding only at the final destination. In this paper, we present a novel scheme in which intermediate nodes conduct partial decoding of LDPC coded packets. In this scheme each node is assigned some number of decoding iterations. The relay node conducts LPDC decoding for that number of iterations and forwards the packet, without ensuring a complete error correction. We show that such partial processing is sufficient to improve the end-to-end reliability significantly. Additionally, we show that it is feasible to tradeoff complexity/energy usage with distortion/reliability by varying the assignment of number of iterations. Finally, we present a low-complexity dynamic programming algorithm that optimally assigns iterations within the network to facilitate operation along an optimalenergy-distortioncurve. Saad B. Qaisar, Shirish S. Karande, Kiran Misra, Hayder Radha |
ICC | 2 |
| 2007 | Transmission-Distortion Tradeoffs in Network Channel CodingabstractNetwork channel coding (NCC) is a framework under which intermediate router/nodes employ encoding/decoding operations to facilitate an efficient multicast delivery of video. The network usage and video distortion is a function of the channel coding rates assigned to nodes in the network. In this paper, we investigate the tradeoff between the total bandwidth usage and a global distortion measure. We propose a dynamic programming based framework to identify the optimal transmission-distortion operating points. The proposed optimal allocation of coding rates is compared with an earlier version of NCC, network embedded FEC (NEF), which optimally places codecs of a fixed channel coding rate within the network. The proposed NCC scheme is shown to achieve significantly improved transmission distortion tradeoffs. Shirish S. Karande, Hayder Radha |
ICIP (5) | 1 |
| 2007 | Hybrid Erasure-Error Protocols for Wireless VideoabstractMany recently proposed cross-layer protocols for wireless video, have advocated the relay of corrupted packet to higher layers. Such protocols lead to both errors and erasures at the compressed video application layer. We generically refer to such schemes as hybrid erasure-error protocols (HEEPs). In this paper, we analyze the utility of HEEPs for efficient transmission of video over wireless channels. In order to maintain the generic nature of the deductions in this paper, we base our analysis on two (rather abstract) communication schemes for wireless video: hybrid error-erasure cross-layer design (CLD) and hybrid error-erasure cross-layer design with side-information (CLDS). We make a comparative analysis of the channel capacities of these schemes over single and multi-hop wireless channels to identify the conditions under which the HEEPs can provide improved performance over conventional (CON) protocols. In addition, we employ Reed Solomon (RS) and low-density parity check (LDPC)-code-based forward-error correction (FEC) schemes to illustrate that the improvement in capacity can easily enable an FEC scheme employed in conjunction with a HEEP to provide improved throughput. Finally we compare the performance of CON, CLD, and CLDS in terms of video quality using the H.264 video standard. The simulation results show a significant advantage for the HEEPs Shirish S. Karande, Hayder Radha |
IEEE Trans. Multim. | 1 |
| 2007 | Header Detection to Improve Multimedia Quality Over Wireless NetworksabstractWireless multimedia studies have revealed that forward error correction (FEC) on corrupted packets yields better bandwidth utilization and lower delay than retransmissions. To facilitate FEC-based recovery, corrupted packets should not be dropped so that maximum number of packets is relayed to a wireless receiver's FEC decoder. Previous studies proposed to mitigate wireless packet drops by a partial checksum that ignored payload errors. Such schemes require modifications to both transmitters and receivers, and incur packet-losses due to header errors. In this paper, we introduce a receiver-based scheme which uses the history of active multimedia sessions to detect transmitted values of corrupted packet headers, thereby improving wireless multimedia throughput. Header detection is posed as the decision-theoretic problem of multihypothesis detection of known parameters in noise. Performance of the proposed scheme is evaluated using trace-driven video simulations on an 802.11b local area network. We show that header detection with application layer FEC provides significant throughput and video quality improvements over the conventional UDP/IP/802.11 protocol stack Syed Ali Khayam, Shirish S. Karande, Muhammad Usman Ilyas, Hayder Radha |
IEEE Trans. Multim. | 2 |
| 2006 | Improving Wireless Multimedia Quality using Header Detection with PriorsabstractRecent wireless multimedia studies have revealed that forward error correction (FEC) on corrupted packets yields better bandwidth utilization and lesser delay than retransmissions. To facilitate FEC decoding at a wireless receiver, it is desirable to relay maximum number of (error-free and corrupted) packets to the receiver's application layer. To that end, most cross-layer multimedia schemes perform partial checksum only on packet headers. However, even with a partial checksum, bursty wireless errors introduce frequent header corruptions, thereby causing considerable packet drops. In this paper, we extend our work in [1], which proposed receiver-based schemes to correct a packet's critical (and corrupted) header fields. This paper: (a) poses header detection as the well-known decision theoretic problem of detecting known parameters in noise, (b) evaluates two detectors using bit-error traces collected over an 802.11b network, (c) provides throughput results for the detectors, and (d) uses video in conjunction with FEC as an example to highlight the efficacy of the proposed schemes. We show that header detection provides significant improvements in throughput and video quality over the conventional UDP/IP protocol stack. Syed Ali Khayam, Shirish S. Karande, Muhammad Usman Ilyas, Hayder Radha |
ICC | 2 |
| 2006 | CLIX: Network Coding and Cross Layer Information Exchange of Wireless VideoabstractNetwork coding (NC) can be efficiently combined with the "physical layer broadcast" property of wireless mediums to facilitate mutual exchange of independent information. At the same time, experimental/theoretical analysis of wireless networks has shown the efficacy of cross-layer protocols that relay corrupted packets in bandwidth hungry video applications. The integration of NC-based information exchange and cross-layer (CL) protocols for wireless video is the primary theme of this work. A particular issue addressed in this paper is the impact of errors in a packet, on the performance of a network code. Thus, we identify the operating conditions under which NC, despite the presence of residue errors, is beneficial. Based on theoretical analysis and experiments using 802.11b wireless traces it is established that the combination of NC with relay of corrupted packets can perform better than (i) conventional schemes that drop all corrupted packets (ii) a scheme that deploys network coding only (iii) a scheme that only deploys a CL strategy that recovers information from corrupted packets. The proposed cross-layer exchange (CLIX) scheme significantly improves the performance of an H.264 based video codec for wireless networks. Shirish S. Karande, Kiran Misra, Hayder Radha |
ICIP | 1 |
| 2006 | Utilizing SSR Indications for Improved Video Communication in Presence of 802.11B Residue ErrorsabstractRadio hardware used for the reception of 802.11b frames is capable of associating a signal to silence ratio (SSR) with each received frame. If a received frame is corrupted, then these SSR indications can be used to provide robust apriori estimate of the bit error rate in the packet. In many recently proposed cross-layer protocols, for transmission of video over wireless networks, recovery of information from partially corrupted packets has shown significant utility. In this paper, based on experiments with actual 802.11b error traces, we show that the channel state information (CSI) provided by the SSR indications can be used to improve the error recovery performance of an FEC scheme employed in conjunction with a cross-layer protocol. H.264 based simulation are used to establish the efficacy of the proposed work for video applications; specifically for video over 802.11b WLAN Shirish S. Karande, Utpal Parrikar, Kiran Misra, Hayder Radha |
ICME | 1 |
| 2006 | Influence of Graph Properties of Peer-to-Peer Topologies on Video Streaming with Network Channel CodingabstractNetwork channel coding (NCC) distributes channel coding functions over network nodes participating in common or diverse communication sessions. A particular case of NCC is network embedded FEC (NEF), which has been shown to exhibit significant improvements in the performance of video streaming applications over multicast peer-to-peer (p2p) trees. The placement of NCC/NEF codecs and its utility in improving the throughput performance is in general a function of the underlying p2p graph topology. In this paper we consider two major forms of p2p topologies: (1) perfectly structured k-ary tree topologies that can be built from (virtually) ideal p2p graphs and (2) unstructured random tree topologies where new nodes randomly join as children to any of the existing peers. The two topologies represent an optimal low diameter structured p2p topology and a trivial randomly evolving sub-optimal topology, respectively. In this paper, we show the impact of key graph parameters, such as the maximum node-degree k and minimum tree-height D, on the performance of NCC in terms of NEF throughput as well as video quality for both structured and unstructured topologies. The utility of NCC/NEF for low-degree and/or less structured p2p topologies is especially highlighted by demonstrating that, embedding of additional codecs can render the performance of less structured topologies or higher diameter topologies to be almost as good as that of the very well structured low diameter topologies. We also investigate the impact of the graph properties on the placement of the NEF codecs Shirish S. Karande, Hayder Radha |
ICME | 1 |
| 2005 | The utility of hybrid error-erasure LDPC (HEEL) codes for wireless multimediaabstractTraditional wireless communication protocols do not relay corrupted packets towards the application layer and neither do they forward such packets over multiple hops. Such an approach can lead to a significant number of packet drops and thus a severe deterioration in performance of high bandwidth applications. Cross-layer protocols which do relay and forward corrupted packets have exhibited substantial promise to mitigate the above problem and thus their utility for wireless multimedia needs to be explored further. Moreover, there is a need to identify efficient channel coding methods for the cross-layer channel. Unlike the traditional schemes, where the channel observed at the application layer is a pure erasure channel, in the cross-layer schemes the application layer channel exhibits hybrid erasure-error impairments. Thus in this paper, we use a rather abstract link-layer model on the basis of which we compare the performance of cross-layer and conventional schemes. We identify the modifications required to be made to RS and LDPC based FEC schemes in order to use them over hybrid erasure-error channels. Finally we compare the considered schemes in terms of video quality using the emerging H.264 video standard. Our video analysis is based on employing a hybrid error-erasure channel coding FEC for the cross-layer schemes versus employing erasure recovery FEC for the traditional protocols. We show that cross-layer schemes can lead to a significant improvement in video quality. Shirish S. Karande, Hayder Radha |
ICC | 1 |
| 2005 | A statistical receiver-based approach for improved throughput of multimedia communications over wireless LANsabstractDelay-sensitivity of real-time applications stipulates resilience against errors and losses in multimedia content. Such resilience is particularly important for bandwidth-constrained and error-prone wireless networks. Recent wireless studies have highlighted that significant improvements in bandwidth utilization and multimedia quality can be achieved if the decision to retain or drop "corrupted" packets is made at the application layer instead of medium access, network or transport layers. Such a strategy, however, necessitates that the maximum number of (good and bad) packets are relayed to the application layer while minimizing any modifications (if any) to the widely-deployed UDP/IP protocol stack. Previous studies have proposed that, while ignoring corruptions in packet payload, only packets with errors in the headers should be dropped. We introduce a receiver-based scheme which uses the history of a multimedia session to correct packet header errors, thereby improving the throughput of real-time applications over wireless local area networks (LANs). The proposed scheme is truly receiver-based and, therefore, does not require any modifications to the source or any intermediate network node. Only minor modifications are required to the protocol stack of the multimedia receiver. Specifically, the proposed scheme generates statistics based on the history of "critical" protocol header fields for active multimedia sessions. These statistics are in turn employed to (a) correct errors in the critical fields and (b) determine if the relevant packets belong to the receiver's session of interest. Syed Ali Khayam, Muhammad Usman Ilyas, Klaus Porsch, Shirish S. Karande, Hayder Radha |
ICC | 4 |
| 2005 | Network embedded FEC (NEF) for video multicast in presence of packet loss correlationabstractNetwork embedded forward error correction (NEF) framework is a paradigm shift from a conventional approach of providing forward error correction (FEC) only on an end-to-end basis. Previous work on NEF has shown promise in greatly improving the decodable probability, message throughput and video quality available to the end receiver. However the utility of NEF for video applications in presence of packet loss correlation has not been evaluated as yet. In practice, packet losses are often correlated and occur in bursts. The distortion in video can be sensitive to the bursty nature of the losses. Thus in this paper we analyze the performance of NEF within random multicast distribution trees that exhibit losses-with-memory over their branches. We utilized the Gilbert model for packet losses over each link. We quantify the performance improvement of NEF over conventional end-to-end FEC in terms of improvement in video quality in presence of packet loss correlation. Embedding NEF codecs can impact the sensitivity of video quality to packet loss correlation. We explicitly evaluate as to how NEF can alter the dependence of video quality on packet loss burstiness. Finally we show that in addition to an average improvement in video quality, NEF can improve the performance in terms of quality guarantees also. Shirish S. Karande, Mingquan Wu, Hayder Radha |
ICIP (1) | 1 |
| 2005 | Does relay of corrupted packets increase capacity?abstractCross-layer protocols that are designed for wireless exhibit both errors and erasures. The nature of these error/erasure impairments and their impact on capacity is a function of the particular cross-layer scheme used. In this paper, we consider two (rather abstract) communication schemes, cross layer design (CLD) and cross layer design with side-information (CLDS). We make a comparative analysis of the channel capacities of these schemes over single and multi-hop wireless channels to identify the conditions under which the cross-layer protocols/channels can provide improved performance over traditional (pure) erasure channels. We show that cross-layer schemes like CLD lead to a capacity increase in most realistic channel scenarios (especially over a single hop) and schemes such as CLDS always lead to a capacity increase. Shirish S. Karande, Hayder Radha |
WCNC | 1 |
| 2005 | Network-embedded FEC for optimum throughput of multicast packet video
Mingquan Wu, Shirish S. Karande, Hayder Radha |
Signal Process. Image Commun. | 2 |
| 2004 | Density and irregularity dependence of partial recovery codesabstractClassical linear block codes, such as Reed-Solomon (RS)-based codes, fail to recover any lost message symbols when the total losses exceed the redundant symbols. Under such adverse channel conditions, source and/or channel rate adaptation to incorporate additional redundancy might not be a viable option. In this paper, we explore a novel method of code adaptation which alters the degree distribution in order to achieve partial recovery of information when complete recovery is not possible. In particular, we change the degree distribution by adjusting the density and irregularity of the code. First, we illustrate that, while maintaining a constant rate, a partial recovery code can be optimized by density modification. Then, we focus on the Partial Reed-Solomon (PRS) codes, which are a family of RS-based codes that are capable of achieving different levels of partial recovery by adjustment to their order. We analyze the dependence of erasure recovery of these codes on density and regularity for a given number of losses. Finally, we present results and analysis which demonstrate that, for a given number of erasures, the PRS codes of order-1 render optimal (erasure-recovery) performance. Shirish S. Karande, Hayder Radha |
ICC | 1 |
| 2004 | Rate-constrained adaptive fec for video over erasure channels with memoryabstractCurrent adaptive FEC schemes used for video streaming applications alter the redundancy in a block of message packets to adapt to varying channel conditions. However, for many popular streaming applications, both the source-rate and the available bandwidth are constrained. In this paper, we present FEC codes that can adapt in real-time to provide higher source-packets recovery without changing the FEC block (N, K) pair constraint. The FEC code profile is changed as function of the number of losses to facilitate an improved data recovery even under severe channel conditions (e.g., number of losses within an N-packet FEC block is larger than N-K). We present a feedback based adaptive FEC scheme, which can adapt in a rate-constrained manner. We also illustrate the utility of this scheme for video streaming applications by analyzing the results of extensive video simulations and comparing our performance to adaptive Reed Solomon FEC schemes. We consider a variety of video sequences and use actual packet traces from WLAN (802.11b) and wired Internet environments. Comparison between the two schemes is conducted on the basis of message packet recovery, PSNR, model based perceptual evaluation and visual subjective evaluation. It is shown that the proposed scheme can significantly improve the video quality and in particular reduce the jerkiness in the received video. Shirish S. Karande, Hayder Radha |
ICIP | 1 |
| 2003 | Analysis and modeling of errors at the 802.11b link layerabstractIn this paper, we analyze the errors observed at the link layer of an 802.11b network. Our analysis at all supported bitrates (i.e., 2, 5.5. and 11 Mbps) establishes that the error patterns are not memoryless, and therefore, they exhibit a certain level of temporal dependencies. Thus, we evaluate the suitability of a two-state Markov model to capture the channel behavior. Non-stationarity of the error patterns renders such a simplistic model inadequate, and hence, we consider higher order models. This formulates a key contribution of this paper, and that is, a hierarchical Markov model, which captures the non-stationarity of the channel while employing real-time application-specific considerations to determine state-transition probabilities. Shirish S. Karande, Syed Ali Khayam, Michael Krappel, Hayder Radha |
ICME | 1 |
| 2003 | A new family of channel coding schemes for real-time visual communicationsabstractIn this paper, we extend our work in [Karande, S. S. and Radha, H., 2003] which employed partial Reed-Solomon (PRS) codes at coding rates near channel capacity on a binary erasure channel (BEC). We demonstrated that an appropriately designed PRS code outperforms the classical Reed-Solomon (RS) code for a performance criterion tailored for realtime applications. This paper extends this analysis for a general BEC with varying channel conditions (under/above channel capacity). We illustrate that PRS codes exhibit a graceful degradation in erasure recovery performance and, hence, are suitable for multimedia communication. Our video simulation results outline that the enhanced erasure recovery yields a profound improvement in the perceived media quality. Finally we investigate the performance of the dividend rendered by PRS codes operating above channel capacity. In particular we define a paradigm for a unique "fixed rate" adaptive FEC scheme based on PRS codes. Shirish S. Karande, Hayder Radha |
ICME | 1 |
| 2003 | Cross-layer protocol design for real-time multimedia applications over 802.11 b networksabstractInherent vulnerability of the wireless medium renders it more susceptible to errors and losses than classical wired media. In this paper, we evaluate the suitability of protocols and strategies across different layers of the stack to provide real-time services over 802.11b wireless LANs. More specifically, within the context of cross-layer design, we compare the performance of UDP with UDP lite - a proposed framework, which improves bandwidth utilization by delivering partially damaged packets to the realtime application. First, we study the high-level end-to-end throughput improvement achieved by making cross-layer modifications to support a UDP lite framework. We compare the quality of perceived media rendered by UDP (dropped packets only) and UDP lite (dropped and corrupted packets). This formulates one of the key findings of this study, that is, although UDP lite improves the overall high-level throughput by relaying corrupted packets to the real-time application, it fails to provide significant enhancement in perceived media quality. This can, in part, be attributed to the bursty nature of errors and losses that we observed at the application layer regardless of the selected transport protocol. Finally, we compare the error-recovery/concealment overhead required by UDP and UDP lite in order to deliver lossless multimedia. We conclude that the overhead required by UDP lite is considerably lower than UDP, since the received corrupted packets that are delivered by UDP lite (but not by UDP) facilitate error-recovery. Syed Ali Khayam, Shirish S. Karande, Michael Krappel, Hayder Radha |
ICME | 2 |
| 2003 | Partial Reed Solomon codes for erasure channelsabstractWe introduce a new family of linear block codes, which we refer to as partial Reed Solomon (PRS) codes. These codes are specifically designed and optimized for real-time multimedia communication over packet-based erasure channels. Based on the constraints and flexibilities of real-time applications, we define a performance measure, message throughput (/spl tau//sub m/), which is suitable for these applications. This measure differentiates the notion of optimum codes for the target multimedia applications as compared to performance measures that are used for nonreal-time data. Based on the proposed measure, we combine the advantages of lowering the density of a code for near capacity performance with the high decoding efficiency of Reed Solomon (RS) codes, in order to design optimum PRS codes. Then, we demonstrate, through an example of a binary erasure channel (BEC), that at near-capacity coding rates, the appropriate design of a PRS code can outperform an RS-code. We extend this analysis and optimization for a general BEC over a wide range of channel conditions. Moreover, as compared with RS codes, the proposed PRS codes provide a significantly improved graceful degradation when the number of losses exceeds the number of parity symbols within the code block. This is a highly desirable feature for real-time multimedia applications. Shirish S. Karande, Hayder Radha |
ITW | 1 |
| 2003 | Performance analysis and modeling of errors and losses over 802.11b LANs for high-bit-rate real-time multimedia
Syed Ali Khayam, Shirish S. Karande, Hayder Radha, Dmitri Loguinov |
Signal Process. Image Commun. | 2 |