Mallesham Dasari

dblp:154/8616 · DBLP profile ↗
← Back
30ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-8855-036XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Computer networks · 11 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ISM: Intelligent Multi-Path Scheduler for Multi-Camera Networked Systems
Alireza Mohammadhosseini, Jacob Chakareski, Mallesham Dasari
MMSys3
2026 LMG: Efficient Streaming of Layered Mesh-Gaussian 3D Scenes
Yuan-Chun Sun, Guodong Chen 0004, Sam Ziaie Kondori, Mallesham Dasari, Cheng-Hsin Hsu
MMSys4
2026 SPARC: Proximity-aware Scheduling of AR Mapping and Cloud-based GenAI Upsampling for Efficient Multi-User SLAM
abstract
The scalability of multi-user SLAM is fundamentally limited by the constrained network and computational resources. Existing approaches either focus on SLAM for single-user scenarios or overload networks and servers by streaming dense, uniform camera data and treating all users equally. This results in poor pose estimation accuracy or slow updates to multiple users. Our key insight is that the sparsity and heterogeneity of user activity reveal that not all users or frames contribute equally to the shared map. Building on this, we propose SPARC - Proximity-aware Scheduling of AR Mapping and a blur-aware adaptive Cloud-based GenAI sampling method, which together form a cloud-native framework for efficient multi-user SLAM. On the client side, adaptive, context-aware frame transmission selectively forwards high-value frames. On the server side, generative AI (GenAI)-based upsampling reconstructs dense scene features from sparse inputs, while a proximity-aware scheduler prioritizes updates for users with higher drift or critical interactions. Together, these components reduce redundant transmission, improve resource allocation, and enable fairness without sacrificing accuracy. We show through extensive experimentation that our method reduces the latency by 2× to 4× compared to state-of-the-art while maintaining similar or better tracking accuracy. More broadly, this work reimagines SLAM as a cloud-native service, paving the way for scalable, real-time AR/VR applications where many users seamlessly interact in shared environments.
Shneka Muthu Kumara Swamy, Mallesham Dasari, Nicholas Mastronarde, Jacob Chakareski
MMSys2
2026 4DGStream: Variable Bitrate Dynamic Gaussian Splatting Streaming
abstract
While 3D Gaussian Splatting (3DGS) has revolutionized static scene representation, the extension to dynamic scene, i.e., 3DGS video (GSV), faces challenges related to reconstruction quality, rendering speed, and storage requirements. The substantial data volume of current GSV poses significant hurdles for streaming applications, particularly in the realm of AR, VR and MR. To tackle these challenges, we introduce 4DGStream, a novel framework that integrates an efficient GSV compression method, Light4D, and a bitrate adaptation streaming strategy, QoSmooth, to ensure smooth playback while maintaining high visual quality. Light4D employs a binarizationassisted spatiotemporal deformation network to model the deformation of Gaussian primitive attributes over time, while a spatiotemporal-aware masking module prunes trivial Gaussians, further enhancing long-term reconstruction quality. To reduce storage, Light4D uses a binary hash grid to model the entropy of attributes for arithmetic coding, with its binary nature allowing efficient entropy modeling via a Bernoulli distribution. These components enable Light4D to improve the FPS/Storage metric by up to 12.4× over SpacetimeGS and 26.4× over 4DGS on the Neu3D dataset, with performance gains exceeding 3× orders of magnitude compared to other NeRF-based state-of-the-art (SOTA) methods. Here, FPS/Storage reflects the balance between rendering speed and data storage. Despite significant model size reductions, Light4D maintains or surpasses the reconstruction quality of 4DGS. Furthermore, QoSmooth provides effective rate control to enhance playback smoothness, reducing bitrate level switches by 61.6% and increasing time-average utility by 26.2%. All these improvements make 4DGStream highly suited for GSV streaming, improving QoE by 36.7% compared to SOTA methods.
Zhicheng Liang, Dayou Zhang, Linfeng Shen, Miao Zhang 0003, Jian Zhang 0054, Bin Ju, Mallesham Dasari, Fangxin Wang 0001, Jiangchuan Liu
IEEE Trans. Multim.7
2026 Implicit Surface Compression - with Good Old Discrete Cosine Transform and Motion Compensation
abstract
The rapid adoption of volumetric capture technologies has created a pressing need for efficient storage and streaming of dynamic 3D content. Unfortunately, current compression standards often treat dynamic sequences as independent frames or rely on computationally expensive non-rigid registration, making them unsuitable for real-time applications or large-scale environments. In this paper, we present a novel end-to-end compression framework for dynamic Truncated Signed Distance Field volumes derived from captured 3D content, leveraging a representation that is temporally stable, easily parallelizable, and already widely used in scene reconstruction and volumetric fusion pipelines. We then adapt classic 2D video coding paradigms such as spatial coding via Discrete Cosine Transform and temporal coding using a real-time motion compensation pipeline to provide robust, real-time, and training-free encoding and decoding for 3D content. Extensive evaluations on human performance captures demonstrate that our codec achieves ~35% bitrate savings at equal distortion while operating in real time at 30 FPS, while stronger temporal coherence in large-scale synthetic environments yields up to 12× bitrate reduction at equal distortion.
Shengxi Wu, Tianshu Huang, Mallesham Dasari, Srinivasan Seshan, Anthony Rowe 0001
ACM Trans. Graph.4
2026 SceneHub4D: A Dataset and Evaluation Framework for 6-DoF 4D VR Scenes
abstract
Volumetric video and 6-DoF scene capture are becoming central to immersive applications such as telepresence and mixed reality content delivery. However, existing volumetric datasets are often short in duration, restricted to studio-captured human subjects, and provide only limited geometric representations. Consequently, evaluating real-world immersive applications in full-scene contexts often necessitates custom capture and 3D reconstruction setups, creating high practical barriers and ultimately hindering reproducibility. To this end, we present SceneHub4D, a new dataset and evaluation framework. Our dataset captures long, dynamic sequences across diverse real-world indoor environments with synchronized multi-view RGB-D streams, calibrated camera poses, and high-resolution background geometry reconstructed via photogrammetry and LiDAR. We provide multiple 3D representations, including point clouds, textured meshes, and Gaussian splats, along with a software toolkit for format conversion, rendering, and metric evaluation. To support structured comparison and perceptual analysis, we provide supplementary metrics including Geometry Complexity Score and Volumetric Temporal Information, and evaluate rendering performance across desktop GPUs and VR headsets. By lowering the practical barriers to capture, reconstruction, and evaluation, SceneHub4D enables researchers to study immersive 3D streaming and rendering systems without requiring custom hardware setups or complex data collection pipelines. We expect it will serve as a useful foundation for advancing volumetric media research.
Jaehong Kim 0002, Mallesham Dasari, Srinivasan Seshan, Anthony Rowe 0001
IEEE Trans. Vis. Comput. Graph.3
2025 SVD: Spatial Video Dataset
abstract
Stereoscopic video has long been the subject of research due to its ability to deliver immersive three-dimensional content to a wide range of applications. The dual-view format inherently provides binocular disparity cues that enhance depth perception and realism, making it indispensable for fields such as telepresence, 3D mapping, and robotic vision. Until recently, however, end-to-end pipelines for capturing, encoding, and viewing high-quality stereoscopic video were neither widely accessible nor optimized for consumer-grade devices. Today's smartphones, such as the iPhone Pro, and modern Head-Mounted Displays (HMDs) like the Apple Vision Pro, offer built-in support for stereoscopic video capture, hardware-accelerated encoding, and seamless playback on devices like the Apple Vision Pro and Meta Quest 3, which require minimal user intervention. Apple refers to this streamlined workflow as spatial Video. Making the full stereoscopic video process available to everyone has made new applications possible. Despite these advances, there remains a notable absence of publicly available datasets that include the complete spatial video pipeline on consumer platforms, hindering reproducibility and comparative evaluation of emerging algorithms.
Mohammad Hossein Izadimehr, Milad Ghanbari, Guodong Chen 0004, Wei Zhou 0021, Xiaoshuai Hao, Mallesham Dasari, Christian Timmerer, Hadi Amirpour
ACM Multimedia6
2025 TVMC: Time-Varying Mesh Compression Using Volume-Tracked Reference Meshes
abstract
Time-varying meshes (TVMs), characterized by their varying connectivity and number of vertices, hold significant potential in AR/VR applications. However, their practical use is challenging due to their large file sizes and the complexity of time-varying topology. Many time-varying mesh compression methods attempted to exploit redundancy between consecutive meshes to compress TVMs more efficiently, however, most face difficulties in establishing stable vertex and surface correspondence between the frames of a TVM. We propose TVMC, a novel TVM compression method that leverages volume tracking and extracts high-quality reference meshes for inter-frame prediction. Specifically, we use as-rigid-as-possible volume tracking to align consecutive TVMs and track volume centers, followed by multidimensional scaling to refine reference centers. This allows us to precisely deform a group of frames to the reference space and extract the reference mesh which is then deformed to approximate each mesh in the group to get displacement fields for TVM compression. Extensive experiments show that TVMC outperforms state-of-the-art methods (e.g., Google Draco, V-DMC 4.0, etc.), with bitrates of 4-6 Mbps compared to 9--12 Mbps for Draco and 10-15 Mbps for V-DMC 4.0. It reduces the decoding time by 66.1% compared to Draco and enables an increased group of frames (up to 15) without significant distortion.
Guodong Chen 0004, Filip Hácha, Libor Vása, Mallesham Dasari
MMSys4
2025 Invited Talk: Time-Varying Mesh Compression
abstract
The proliferation of immersive applications such as telepresence, AR/VR streaming, and 3D digital twins demands efficient capture, compression, and delivery of dynamic 3D scenes. Time-varying meshes (TVMs), sequences of meshes with evolving geometry and topology, are a compact and expressive format for volumetric video, but their high data rates and irregular structure make real-time streaming challenging. This paper presents an overview of the compression of TVMs. The talk also includes two complementary contributions from our group: TVMC, a time-varying mesh compression framework for deformable objects utilizing volume-tracked references, and an extension designed for full, unbounded scene meshes with both static and dynamic parts. Together, these systems demonstrate high compression ratios, low decoding latencies, and robustness to real-world scene complexity, paving the way for scalable, low-bitrate 3D video streaming.
Guodong Chen 0004, Mallesham Dasari
VCIP2
2025 Grasp-HGN: Grasping the Unexpected
abstract
For transradial amputees, robotic prosthetic hands promise to regain the capability to perform daily living activities. To advance next-generation prosthetic hand control design, it is crucial to address current shortcomings in robustness to out of lab artifacts, and generalizability to new environments. Due to the fixed number of object to interact with in existing datasets, contrasted with the virtually infinite variety of objects encountered in the real world, current grasp models perform poorly on unseen objects, negatively affecting users’ independence and quality of life. To address this: (i) we define semantic projection, the ability of a model to generalize to unseen object types and show that conventional models like YOLO, despite 80% training accuracy, drop to 15% on unseen objects. (ii) We propose Grasp-LLaVA, a Grasp Vision Language Model enabling human-like reasoning to infer the suitable grasp type estimate based on the object’s physical characteristics resulting in a significant 50.2% accuracy over unseen object types compared to 36.7% accuracy of an SOTA grasp estimation model. Lastly, to bridge the performance-latency gap, we propose Hybrid Grasp Network (HGN), an edge-cloud deployment infrastructure enabling fast grasp estimation on edge and accurate cloud inference as a fail-safe, effectively expanding the latency vs. accuracy Pareto. HGN with confidence calibration (DC) enables dynamic switching between edge and cloud models, improving semantic projection accuracy by 5.6% (to 42.3%) with 3.5× speedup over the unseen object types. Over a real-world sample mix, it reaches 86% average accuracy (12.2% gain over edge-only), and 2.2× faster inference than Grasp-LLaVA alone.
Mehrshad Zandigohar, Mallesham Dasari, Gunar Schirner
ACM Trans. Embed. Comput. Syst.2
2024 Capacitive Sensing-based Eye Tracking for XR Glasses
abstract
Eye tracking technology has become increasingly vital as human-computer interaction utilizing extended reality (XR) technologies becomes more mainstream. Traditional eye tracking systems predominantly rely on cameras. While effective, these systems often suffer from high complexity, power consumption, and significant costs. This paper proposes a novel approach to eye tracking with capacitive sensing technology. Leveraging the principles of capacitance, we envision more accessible and efficient eye tracking solutions that can be integrated into various applications. This paper outlines the process of designing and analyzing a prototype of capacitance-based eye tracking.
Aidan Hanson, Amr Kassab, Mallesham Dasari
MobiCom3
2024 A First Look at Apple's Stereoscopic Video and its Potential in Live Video Streaming for XR Headsets
abstract
With the evolution of stereoscopic video technology, live video streaming stands at the threshold of a more engaging and immersive future. Apple's recent innovations in video and Extended Reality (XR) technologies have paved the way for immersive live video experiences. This work characterizes Apple's spatial video and explores its potential in live video streaming, providing insights into the future of video streaming applications.
Mingkun Liu, Mallesham Dasari, Dimitrios Koutsonikolas
MobiCom3
2024 MeshReduce: Scalable and Bandwidth Efficient 3D Scene Capture
abstract
3D video enables a remote viewer to observe a 3D scene from any angle or location. However, current 3D capture solutions incur high latency, consume significant bandwidth, and scale poorly with the number of depth sensors and size of scenes. These problems are largely caused by the current monolithic approach to 3D capture and the use of inefficient data representations for streaming. This paper introduces MeshReduce, a distributed scene capture, stream, and render system that advocates for the use of textured mesh data representation early in the 3D video capture and transmission process. Textured meshes are compact and can provide lower bitrates for the same quality compared to other 3D data representations. However, streaming textured meshes creates compute and memory challenges to achieve bandwidth efficiency. MeshReduce addresses these issues by using a pipeline that creates independent mesh reconstructions and incrementally merges them, rather than creating a single mesh directly from all sensor streams. While this enables a more efficient implementation, this approach requires optimal exchange of textured meshes across the network. MeshReduce also incorporates a novel approach for network rate control that divides bandwidth between texture and mesh for efficient, adaptive 3D video streaming. We demonstrate a real-time integrated embedded compute implementation of MeshReduce that can operate with commercial Azure Kinect depth cameras as well as a custom sensor front-end that uses LiDAR and 360° camera inputs to dramatically increase coverage.
Mallesham Dasari, Connor Smith, Kittipat Apicharttrisorn, Srinivasan Seshan, Anthony Rowe 0001
VR2
2024 StageAR: Markerless Mobile Phone Localization for AR in Live Events
abstract
Localizing mobile phone users precisely enough to provide AR content in theaters and concert venues is extremely challenging due to dynamic staging and variable lighting. Visual markers are often disruptive in terms of aesthetics, and static pre-defined feature maps are not robust to visual changes. In this paper, we study several techniques that leverage sparse fixed infrastructure to monitor and adapt to changes in the environment at runtime to enable robust AR quality pose tracking for large audiences. Our most basic technique uses one or more fixed cameras in the environment to prune away poor feature points due to motion and lighting from a static model. For more challenging environments, we propose transmitting dynamic 3D feature maps that adapt to changes in the scene in real-time. Users with a mobile phone camera can use these maps to accurately localize across highly dynamic environments without explicit markers. We show the performance trade-offs resulting from StageAR’s different reconstruction techniques, ranging from multiple stereo cameras to cameras paired with LiDAR. We evaluate each approach in our system across a wide variety of simulated and real environments at auditorium/theater scale and find that our most accurate technique can match the performance of large ($1.5 \times 1.5{\mathrm {m}}$) back-lit static markers without being visible to users.
Shengxi Wu, Mallesham Dasari, Kittipat Apicharttrisorn, Anthony Rowe 0001
VR3
2024 Fumos: Neural Compression and Progressive Refinement for Continuous Point Cloud Video Streaming
abstract
Point cloud video (PCV) offers watching experiences in photorealistic 3D scenes with six-degree-of-freedom (6-DoF), enabling a variety of VR and AR applications. The user's Field of View (FoV) is more fickle with 6-DoF movement than 3-DoF movement in 360-degree video. PCV streaming is extremely bandwidth-intensive. However, current streaming systems require hundreds of Mbps bandwidth, exceeding the bandwidth capabilities of commodity devices. To save bandwidth, FoV-adaptive streaming predicts a user's FoV and only downloads point cloud data falling in the predicted FoV. But it is difficult to accurately predict the user's FoV even 2-3 seconds before playback due to 6-DoF. Misprediction of FoV or network bandwidth dips results in frequent stalls. To avoid rebuffering, existing systems would cause incomplete FoV and degraded experience, deteriorating the user's quality of experience (QoE). In this paper, we describe Fumos, a novel system that preserves interactive experience by avoiding playback stalls while maintaining high perceptual quality and high compression rate. We find a research gap in inter-frame redundant utilization and progressive mechaism. Fumos has three crucial designs, including (1) Neural compression framework with inter-frame coding, namely N-PCC, which achieves both bandwidth efficiency and high fidelity. (2) Progressive refinement streaming framework that enables continuous playback by incrementally upgrading a fetched portion to a higher quality (3) System-level adaptation that employs Lyapunov optimization to jointly optimize the long-term user QoE. Experimental results demonstrate that Fumos significantly outperforms Draco, achieving an average decoding rate acceleration of over 260×. Moreover, the proposed compression framework N-PCC attains remarkable BD-Rate gains, averaging 91.7% and 51.7% against the state-of-the-art point cloud compression methods G-PCC and V-PCC, respectively.
Zhicheng Liang, Junhua Liu 0003, Mallesham Dasari, Fangxin Wang 0001
IEEE Trans. Vis. Comput. Graph.3
2023 RenderFusion: Balancing Local and Remote Rendering for Interactive 3D Scenes
abstract
Many modern-day XR devices (e.g. mobile headsets, phones, etc.) lack the computing resources required to render complex 3D scenes in real-time. Typically, to render a high-resolution scene on a lightweight XR device, 3D designers arduously decimate and fine-tune the objects. As an alternative, remote rendering systems can utilize powerful nearby servers to stream rendering results to a client. While this is a promising solution, it can introduce a variety of latency and reliability issues, especially under variable network conditions. In this paper, we present a distributed rendering system that combines both remote rendering and on-device, “local” rendering to add robustness to network fluctuations and device workloads. To maximize user QoE, our approach dynamically swaps an object’s rendering medium, adjusting for client workload, low frame rates, and several perceptual characteristics. To model these characteristics, we perform a study under simulated conditions to measure how users perceive latency and complexity differences between objects in a scene. Using the results of the study, we then provide an algorithm for choosing the optimal object rendering medium, based on rendering complexity as well as network and latency models, ensuring that a target frame rate will be met. Finally, we evaluate this algorithm on a prototype implementation that can provide cross-platform split rendering using web technologies.
Edward Lu, Sagar Bharadwaj, Mallesham Dasari, Connor Smith, Srinivasan Seshan, Anthony Rowe 0001
ISMAR3
2023 Scaling VR Video Conferencing
abstract
Virtual Reality (VR) telepresence platforms are being challenged to support live performances, sporting events, and conferences with thousands of users across seamless virtual worlds. Current systems have struggled to meet these demands which has led to high-profile performance events with groups of users isolated in parallel sessions. The core difference in scaling VR environments compared to classic 2D video content delivery comes from the dynamic peer-to-peer spatial dependence on communication. Users have many pair-wise interactions that grow and shrink as they explore spaces. In this paper, we discuss the challenges of VR scaling and present an architecture that supports hundreds of users with spatial audio and video in a single virtual environment. We leverage the property of spatial locality with two key optimizations: (1) a Quality of Service (QoS) scheme to prioritize audio and video traffic based on users' locality, and (2) a resource manager that allocates client connections across multiple servers based on user proximity within the virtual world. Through real-world deployments and extensive evaluations under real and simulated environments, we demonstrate the scalability of our platform while showing improved QoS compared with existing approaches.
Mallesham Dasari, Edward Lu, Michael W. Farb, Nuno Pereira 0001, Ivan Liang, Anthony Rowe 0001
VR1
2022 Swift: Adaptive Video Streaming with Layered Neural Codecs
Mallesham Dasari, Kumara Kahatapitiya, Samir Ranjan Das, Aruna Balasubramanian, Dimitris Samaras
NSDI1
2022 Live 3D Scene Capture for Virtual Teleportation
abstract
It has long been a goal of immersive telepresence to capture and stream 3D spaces such that a remote viewer can watch from any location or angle within the scene. This demonstration presents Mosaic, a new distributed 3D scene capture system that uses textured mesh data representation for streaming a 3D volumetric video of a space to remote viewers. Compared to more common point cloud based methods, we show that textured mesh data requires less bandwidth and yields the same visual quality. However, textured mesh reconstruction is compute and memory intensive, mesh simplification is not easily parallelizable, and texture maps lacks spatial and temporal coherence. Mosaic tackles these challenges by examining each computational stage and determines how they can be efficiently distributed across multiple compute nodes to reduce overall latency, minimize bandwidth, and maintain quality. We then provide an end-to-end latency and bandwidth breakdown that can be used to target future acceleration work.
Mallesham Dasari, Connor Smith, Kittipat Apicharttrisorn, Anthony Rowe 0001, Srinivasan Seshan
SenSys2
2022 Cyclops: an FSO-based wireless link for VR headsets
abstract
The ultimate goal of virtual reality (VR) is to create an experience indistinguishable from actual reality. To provide such a "life-like" experience, (i) the VR headset (VRH) should be wireless so that the user can move around freely, and (ii) the wireless link, connecting the VRH to a high-performance renderer, should support high data rates (tens to hundreds of Gbps). Industry is already pushing towards such wireless VRHs; however, these wireless links can only support a few Gbps rates. In general, current radio-frequency (RF) links (including mmWave) are not able to provide desired data rates. In this paper, we build a system, we call Cyclops, which uses free-space optical (FSO) technology to create a high-bandwidth VR wireless link. FSO links are capable of very high data rates (up to Tbps) due to the high frequencies of light waves and narrow beams. The main challenges in developing an effective FSO link are: (i) designing a link with sufficient movement tolerance, and (ii) developing a viable tracking and pointing (TP) mechanism which maintains the link while the VRH moves. As traditional TP approaches seem infeasible in our context, we develop a novel TP approach based on learning techniques, leveraging the VRH's inbuilt tracking system. We build robust 10 Gbps and 25Gbps link prototypes from commodity components, demonstrate their viability for expected movement speeds of a VRH, and show that, with certain custom-built components, we can support much higher movement speeds and bandwidths.
Himanshu Gupta 0001, Max Curran, Jon P. Longtin, Torin Rockwell, Kai Zheng 0017, Mallesham Dasari
SIGCOMM6
2021 dcSR: practical video quality enhancement using data-centric super resolution
abstract
With the next generation immersive video applications, network capacity is becoming a growing bottleneck to deliver a high quality video to end-users. Recent advances to tackle this challenge introduced super-resolution (SR) for video quality enhancement through neural computations by leveraging client-side compute capacity. However, the existing SR models are bulky, compute-, and memory-expensive, which makes it difficult to deploy them in practice. In this work, we present dcSR, a lightweight data-centric SR approach that enables a practical neural quality enhancement for videos. On the server-side, dcSR constructs micro SR models trained on a few selected frames from each video through a data-centric paradigm by employing a long term video scene understanding mechanism. On the client-side, dcSR integrates the micro SR models into the regular video decoder and enhances the video quality in real-time without compromising on quality enhancement. We evaluate dcSR and show its benefits by comparing it with previous methods.
Duin Baek, Mallesham Dasari, Samir Ranjan Das, Jihoon Ryoo
CoNEXT2
2021 L3BOU: Low Latency, Low Bandwidth, Optimized Super-Resolution Backhaul for 360-Degree Video Streaming
abstract
In recent years, streamed 360° videos have gained popularity within Virtual Reality (VR) and Augmented Reality (AR) applications. However, they are of much higher resolutions than 2D videos, causing greater bandwidth consumption when streamed. This increased bandwidth utilization puts tremendous strain on the network capacity of the cloud providers streaming these videos. In this paper, we introduce L3BOU, a novel, three-tier distributed software framework that reduces cloud-edge bandwidth in the backhaul network and lowers average end-to-end latency for 360° video streaming applications. The L3BOU framework achieves low bandwidth and low latency by leveraging edge-based, optimized upscaling techniques. L3BOU accomplishes this by utilizing down-scaled MPEG-DASH-encoded 360° video data, known as Ultra Low Resolution (ULR) data, that the L3BOU edge applies distributed super-resolution (SR) techniques on, providing a high quality video to the client. L3BOU is able to reduce the cloud-edge backhaul bandwidth by up to a factor of 24, and the optimized super-resolution multi-processing of ULR data provides a 10-fold latency decrease in super resolution upscaling at the edge.
Ayush Sarkar, John O. Murray, Mallesham Dasari, Michael Zink, Klara Nahrstedt
ISM3
2020 Streaming 360-Degree Videos Using Super-Resolution
abstract
360° videos provide an immersive experience to users, but require considerably more bandwidth to stream compared to regular videos. State-of-the-art 360° video streaming systems use viewport prediction to reduce bandwidth requirement, that involves predicting which part of the video the user will view and only fetching that content. However, viewport prediction is error prone resulting in poor user Quality of Experience (QoE). We design PARSEC, a 360° video streaming system that reduces bandwidth requirement while improving video quality. PARSEC trades off bandwidth for additional client-side computation to achieve its goals. PARSEC uses an approach based on super-resolution, where the video is significantly compressed at the server and the client runs a deep learning model to enhance the video to a much higher quality. PARSEC addresses a set of challenges associated with using super-resolution for 360° video streaming: large deep learning models, slow inference rate, and variance in the quality of the enhanced videos. To this end, PAR-SEC trains small micro-models over shorter video segments, and then combines traditional video encoding with super-resolution techniques to overcome the challenges. We evaluate PARSEC on a real WiFi network, over a broadband network trace released by FCC, and over a 4G/LTE network trace. PARSEC significantly outperforms the state-of-art 360° video streaming systems while reducing the bandwidth requirement.
Mallesham Dasari, Arani Bhattacharya, Santiago Vargas, Pranjal Sahu, Aruna Balasubramanian, Samir Ranjan Das
INFOCOM1
2019 S3'19 - Wireless of the Students, by the Students, and for the Students Workshop
abstract
The Wireless of the Students, by the Students, and for the Students (S3) Workshop provides a unique venue for graduate students around the world to present, discuss, and exchange ideas on cross-cutting research on mobile wireless networks. As its name suggests, the workshop is organized by students, and the technical sessions are given by student presenters. The workshop aims at fostering early-career development among students and exposing them to the workings of academic life. It provides a venue for students to learn about each other's' work and discover opportunities for collaboration. The workshop invites students to submit papers, posters, and demos. All submissions are peer-reviewed by the student program committee.
Mallesham Dasari, Elahe Soltanaghai, Chia-Yi Yeh
MobiCom1
2019 Advancing User Quality of Experience in 360-degree Video Streaming
abstract
Conventional streaming solutions for streaming 360-degree panoramic videos are inefficient in that they download the entire 360-degree panoramic scene, while the user views only a small sub-part of the scene called the viewport. This can waste over 80% of the network bandwidth. We develop a comprehensive approach called Mosaic that combines a powerful neural network-based viewport prediction with a rate control mechanism that assigns rates to different tiles in the 360-degree frame such that the video quality of experience is optimized subject to a given network capacity. We model the optimization as a multi-choice knapsack problem and solve it using a greedy approach. We also develop an end-to-end testbed using standards-compliant components and provide a comprehensive performance evaluation of Mosaic along with four other streaming techniques - two for conventional adaptive video streaming and two for 360-degree tile-based video streaming. Mosaic outperforms the best of the competition by as much as 50% in terms of median video quality.
Sohee Kim Park, Arani Bhattacharya, Zhibo Yang 0002, Mallesham Dasari, Samir Ranjan Das, Dimitris Samaras
Networking4
2019 Spectrum Protection from Micro-transmissions Using Distributed Spectrum Patrolling
Mallesham Dasari, Muhammad Bershgal Atique, Arani Bhattacharya, Samir Ranjan Das
PAM1
2019 A Lightweight Multi-Section CNN for Lung Nodule Classification and Malignancy Estimation
abstract
The size and shape of a nodule are the essential indicators of malignancy in lung cancer diagnosis. However, effectively capturing the nodule's structural information from CT scans in a computer-aided system is a challenging task. Unlike previous models that proposed computationally intensive deep ensemble models or three-dimensional CNN models, we propose a lightweight, multiple view sampling based multi-section CNN architecture. The model obtains a nodule's cross sections from multiple view angles and encodes the nodule's volumetric information into a compact representation by aggregating information from its different cross sections via a view pooling layer. The compact feature is subsequently used for the task of nodule classification. The method does not require the nodule's spatial annotation and works directly on the cross sections generated from volume enclosing the nodule. We evaluated the proposed method on lung image database consortium (LIDC) and image database resource initiative (IDRI) dataset. It achieved the state-of-the-art performance with a mean 93.18% classification accuracy. The architecture could also be used to select the representative cross sections determining the nodule's malignancy that facilitates in the interpretation of results. Because of being lightweight, the model could be ported to mobile devices, which brings the power of artificial intelligence (AI) driven application directly into the practitioner's hand.
Pranjal Sahu, Dantong Yu, Mallesham Dasari, Fei Hou 0001, Hong Qin 0001
IEEE J. Biomed. Health Informatics3
2018 Impact of Device Performance on Mobile Internet QoE
Mallesham Dasari, Santiago Vargas, Arani Bhattacharya, Aruna Balasubramanian, Samir Ranjan Das, Michael Ferdman
Internet Measurement Conference1
2018 Scalable Ground-Truth Annotation for Video QoE Modeling in Enterprise WiFi
abstract
Mobile video traffic is dominant in cellular and enterprise wireless networks. With the advent of myriads of applications from video telephony and streaming to virtual reality, network administrators face the challenge to provide high quality of experience (QoE) in the face of diverse wireless conditions and application contents. Yet, state-of-the-art networks lack analytics for QoE, as this requires support from the application or user feedback. While there are existing techniques to map quality of service (QoS) to QoE by training machine learning (ML) models without requiring user feedback, these techniques are limited to only few applications (e.g., Skype), due to insufficient QoE ground-truth annotation for ML. To address these limitations, we focus on video telephony applications and model key artefacts of spatial and temporal video QoE. Our key contribution is designing content- and device-independent metrics and training across diverse WiFi conditions. We show that our metrics achieve a median 90% accuracy by comparing with mean-opinion-score (MOS) from more than 200 users and 800 video samples. Our content-independent metrics significantly reduce the MOS prediction error of previous works and are validated over three popular video telephony applications - Skype, FaceTime and Google Hangouts.
Mallesham Dasari, Shruti Sanadhya, Christina Vlachou, Kyu-Han Kim, Samir Ranjan Das
IWQoS1
2017 Real time detection of MAC layer DoS attacks in IEEE 802.11 wireless networks
abstract
In this paper, a real time detection of Medium Access Control (MAC) layer attacks in IEEE 802.11 wireless networks is proposed. There can be different kinds of Denial of Service (DoS) attacks observed at the MAC layer such as misbehaviour and selfish attacks. The malicious nodes manipulate the MAC protocol parameters such as back-off time, network allocation vector value and short inter frame space, or flood the network with huge volume of dummy packets. With this, the attacker nodes capture entire network bandwidth causing legitimate nodes not communicate with other nodes, consequently decreasing the throughput of the nodes significantly. This paper gives an effective real time detection of these attacks with minimal detection delay. We collect the delay and throughput data and apply a change point detection algorithm to observe the change of distribution. To observe the effect of these attacks, two types of attacks: back-off manipulation and RTS flooding are simulated using Network Simulator (NS-3). The simulation results shows the efficiency of the detection algorithm in terms of delay and throughput results.
Mallesham Dasari
CCNC1