VLDB 2026 Research / reviewers in the wild / expert
Viswanathan (Vishy) Swaminathan
dblp:306/7338 · also Vishy Swaminathan, Viswanathan Swaminathan 0001
· DBLP profile ↗
65ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0002-1357-1502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Computer networks · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Offloading-based Power-Efficient Mobile VTuber Live StreamingabstractVirtual YouTuber (VTuber) live streaming, which renders and streams a virtual avatar of the actual streamer on top of the live camera view, has gained significant popularity recently. Despite the engaging user experience, the intensive and power-consuming computations required by VTuber applications, such as facial feature extraction and avatar rendering, pose significant challenges to the constrained battery life of mobile devices. We develop a power-efficient VTuber live streaming system by offloading the camera view and the computation-intensive operations from the mobile device to an edge server. Our approach not only reduces the power consumption of the mobile device but also enables larger-scale rendering of multiple avatars, which is infeasible in existing mobile VTuber systems. Furthermore, to reduce the bandwidth overhead caused by the camera view offloading, we develop an adaptive framerate control mechanism to dynamically adjust the framerate of the offloaded camera view based on the variations of inter-frame luminance, as well as resolution control to dynamically adjust the resolution of the offloaded camera view based on the number and size of the faces. Our evaluations on the end-to-end VTuber live streaming system demonstrate 26%-29% power savings with limited latency, bandwidth, and quality overhead. Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | VADER: Video Alignment Differencing and RetrievalabstractWe propose VADER, a spatio- temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer- based alignment module then refines the temporal localization of the query fragment within the matched video. A space- time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non- editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered. Code and data are available at https://github.com/AlexBlck/vader Alexander Black 0001, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, John P. Collomosse |
ICCV | 7 |
| 2023 | Active Context Modeling for Efficient Image and Burst CompressionabstractState-of-the-art compression frameworks usually contain a prediction module and an error context modeling module to reduce redundancy among pixels and improve compression performance. Modern compression algorithms are context adaptive. While adaptive compression algorithms improve over a static context model, they are computationally prohibitive, as the model has to be learned per image during encoding. In this work, we formulate the problem of active context modeling where we train an approximated error context model using an actively selected subset of pixels to significantly speedup the error context modeling while minimizing the impact on compression rate. We investigate the proposed active context modeling framework for both single image compression and burst image compression where the goal is to compress a set of images (usually 6 to 12) captured at a very short time interval between each other. We find that our active context modeling framework is significantly faster than the state-of-the-art while achieving a comparable compression rate. These results indicate the utility of the proposed active context modeling framework for image compression. Gang Wu 0013, Stefano Petrangeli, Ryan Rossi, Viswanathan (Vishy) Swaminathan |
ISM | 6 |
| 2023 | GPU-accelerated Lossless Image Compression with Massive ParallelizationabstractWith the rapid increase of digital content like images or videos nowadays, compression technology contributes more to saving storage or transferring time with large-scale data. While some existing methods already achieved a great compression ratio, they are not applicable to certain live applications under low efficiency. In this work, we use massive parallelization to speed up the SOTA baseline FLIF, including bitwise-equivalent speedup and learning-based speedup. Our method achieves $38.7 \times$ throughputs for encoding and $2.45 \times$ throughputs for decoding, compared to the baseline FLIF. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli, Tong Yu 0001 |
ISM | 3 |
| 2023 | Content-aware Progressive Image Compression and SyncingabstractProgressive image compression and syncing between devices is an important and challenging problem. When the users are collaboratively editing the same image online, they would expect the changes made by others to be instantly displayed on their side. Since such syncing can be very frequent and usually the image sizes are significantly larger than text data, image live co-editing cannot be easily achieved in the same way as those commonly seen in document co-editing tools. While previous compression techniques like PNG, JPEG and FLIF enable spatially progressive compression, they do not support content-aware compression. Thus, even though the image can be gradually displayed, users cannot prioritize the transmission and display of the most important bits of the image, and often times the resulting pixelation during syncing greatly hurts the user experience. Many existing works on saliency detection can be utilized to provide content awareness. However, many of those techniques are deep-learning-based and it would be computationally prohibitive to directly use them in a latency sensitive scenario like collaborative editing on client devices. In this work, we aim to find a middle ground between a good quality pixel prioritization strategy and extremely fast compression. We start with the pipeline proposed in FLIF and improve it with an entropy-based pixel prioritization strategy, which enables better progressive compression and syncing. Specifically, we modify the traditional Adam interlacing mode [1] to enable an arbitrary pixel transmission order avoiding spatial dependency issues. After constructing the MANIAC tree, we calculate entropy values for each leaf nodes and use them to determine the priority. In addition, we propose to use pixel masks of individual zoom levels to indicate the positions of the transmitted pixels. We further integrate the mask compression algorithm to reduce the communication cost. Through extensive experiments on over 2000 images, we show our proposed method outperforms the baseline methods. Junda Wu, Tong Yu 0001, Gang Wu 0013, Stefano Petrangeli, Handong Zhao, Sungchul Kim, Viswanathan (Vishy) Swaminathan |
ISM | 8 |
| 2023 | Power Efficient Mobile VTuber Live StreamingabstractVirtual YouTuber (VTuber) live streaming, which renders and streams a virtual avatar of the real-person streamer on top of the live camera view, has gained significant popularity recently. Despite the engaging user experience, the intensive and power-consuming computations required by VTuber, such as facial feature extraction and avatar rendering, pose significant challenges to the constrained battery life of the mobile device. We develop a power efficient VTuber live streaming system by offloading the camera view and the computation-intensive operations from the mobile device to an edge server, which not only significantly reduces the power consumption of the mobile device but also enables larger-scale rendering of multiple avatars that are not feasible in the existing mobile VTuber systems. Furthermore, to reduce the bandwidth overhead caused by the camera view offloading, we develop an adaptive framerate control mechanism to dynamically adjust the framerate of the offloaded camera view based on the variations of inter-frame luminance. Our evaluations on the end-to-end VTuber live streaming system demonstrate significant power savings with limited bandwidth, latency, and quality overhead. Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001 |
MMAsia | 3 |
| 2023 | Privacy Aware Experiments without CookiesabstractConsider two brands that want to jointly test alternate web experiences for their customers with an A/B test. Such collaborative tests are today enabled usingthird-party cookies, where each brand has information on the identity of visitors to another website, ensuring a consistent treatment experience. With the imminent elimination of third-party cookies, such A/B tests will become untenable. We propose a two-stage experimental design, where the two brands only need to agree on high-level aggregate parameters of the experiment to test the alternate experiences. Our design respects the privacy of customers. We propose an unbiased estimator of the Average Treatment Effect (ATE), and provide a way to use regression adjustment to improve this estimate. On real and simulated data, we show that the approach provides valid estimate of the ATE and is robust to the proportion of visitors overlapping across the brands. Our demonstration describes how a marketer can design such an experiment and analyze the results. Shiv Shankar, Ritwik Sinha, Saayan Mitra, Viswanathan (Vishy) Swaminathan, Sridhar Mahadevan, Moumita Sinha |
WSDM | 4 |
| 2023 | User Navigation Modeling, Rate-Distortion Analysis, and End-to-End Optimization for Viewport-Driven 360$^\circ $ Video StreamingabstractThe emerging technologies of Virtual Reality (VR) and 360$^\circ $video introduce new challenges for state-of-the-art video communication systems. Enormous data volume and spatial user navigation are unique characteristics of 360$^\circ$videos that necessitate a space-time effective allocation of the available network streaming bandwidth over the 360$^\circ$video content to maximize the Quality of Experience (QoE) delivered to the user. Towards this objective, we investigate a framework for viewport-driven rate-distortion optimized 360$^\circ$video streaming that integrates the user view navigation patterns and the spatiotemporal rate-distortion characteristics of the 360$^\circ$video content to maximize the delivered user viewport video quality, for the given network/system resources. The framework comprises a methodology for assigning dynamic navigation likelihoods over the 360$^\circ$video spatiotemporal panorama, induced by the user navigation patterns, an analysis and characterization of the 360$^\circ$video panorama's spatiotemporal rate-distortion characteristics that leverage preprocessed spatial tilling of the content, and an optimization problem formulation and solution that capture and aim to maximize the delivered expected viewport video quality, given a user's navigation patterns, the 360$^\circ$video encoding/streaming decisions, and the available system/network resources. We formulate a Markov model to capture the navigation patterns of a user over the 360$^\circ$video panorama and simultaneously extend our actual navigation datasets by synthesizing additional realistic navigation data. Moreover, we investigate the impact of using two different tile sizes for equirectangular tiling of the 360$^\circ$video panorama. Our experimental results demonstrate the advantages of our framework over the conventional approach of streaming a monolithic uniformly-encoded 360$^\circ$video and a state-of-the-art navigation-speed based reference method. Considerable average and instantaneous viewport video quality gains of up to 5 dB are demonstrated in the case of five popular 4 K 360$^\circ$videos. In addition, we explore the impact of two different popular 360$^\circ$video quality metrics applied to evaluate the streaming performance of our system framework and the two reference methods. Finally, we demonstrate that by exploiting the unequal rate-distortion characteristics of the different spatial sectors of the 360$^\circ$video panorama, we can enable spatially more uniform and temporally higher 360$^\circ$video viewport quality delivered to the user, relative to monolithic streaming. Jacob Chakareski, Xavier Corbillon, Gwendal Simon, Viswanathan (Vishy) Swaminathan |
IEEE Trans. Multim. | 4 |
| 2022 | Contextualized Styling of Images for Web Interfaces using Reinforcement LearningabstractContent personalization is one of the foundations of today’s digital marketing. Often the same image needs to be adapted for different design schemes for content that is created for different occasions, geographic locations or other aspects of the target population. We present a novel reinforcement learning (RL) based method for automatically stylizing images to complement the design scheme of media, e.g., interactive websites, apps, or posters. Our approach considers attributes related to the design of the media and adapts the style of the input image to match the context. We do so using a preferential reward system in the RL framework that learns a reward function using human feedback. We conducted several user studies to evaluate our approach and demonstrate that we are able to effectively adapt image styles to different design schemes. In user studies, images stylized through our approach were the most preferred variation across a majority of our experiments. Additionally, we also release a dataset consisting of perceptual associations of web context with the associated image style. Pooja Guhan, Saayan Mitra, Somdeb Sarkhel, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, Aniket Bera, Dinesh Manocha |
ISM | 6 |
| 2022 | Towards Efficient Video Super Resolution for Faster StreamingabstractLive Video Streaming contributes to a major portion of Internet traffic. Low network bandwidth conditions can negatively impact the quality of the live video stream, which should be as high as possible for a good user experience. An emergent solution to deal with this problem is to use Video Super Resolution (VSR) methods to only transmit a lower quality version of the video, and super resolve it to a higher quality version at the client side, therefore saving on the amount of data to transmit. While up-scaling incoming low resolution video frames, it is important to consider the tradeoff between quality and latency of the super resolved output video. This is necessary to ensure that the super resolution process does not negatively impact the latency of the live video stream. The goal of this work is to investigate this trade off and present our approach based on a combination of super resolution and video frames interpolation. We propose a new strategy to super resolve the input video by leveraging both VSR and Video Interpolation (VI) in order to speed up computation without impacting the quality of the final video. Preliminary experimental results confirm that, compared with standalone VSR models, our proposed pipelined VSR+VI approach achieves 20% performance speedup, with little impact on the the quality of the final output video. Sowmya Vasuki Jallepalli, Pratik Mulchandani, Chirag Trasikar, Chetan Manjesh, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli |
ISM | 6 |
| 2022 | Task-Oriented Near-Lossless Burst CompressionabstractUnlike single images, capturing bursts enables many possible downstream tasks (e.g. superresolution, HDR enhancement) due to the rich information preserved in the consecutive frames. Efficient compression of these bursts is therefore essential given the additional frames to store. In this paper, we propose a novel near-lossless compression method that can preserve the most relevant information in the burst to enable multiple downstream image enhancement tasks, while at the same time reducing the file size. Specifically, we propose a two-bitstream near-lossless compression pipeline that controls the image-space distortion at frame level, and introduce the Lipschitz condition to bound the task-space distortion at burst level. Experiments conducted on a real-world burst dataset confirm the benefit of the proposed solution in terms of rate-distortion both in the burst frame space and the superresolution task space, a popular downstream task in burst processing. Weixin Jiang, Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli, Ryan Rossi, Nedim Lipka |
ISM | 3 |
| 2022 | Towards Accurate Positioning in Multiuser Augmented Reality on Mobile DevicesabstractMultiuser Augmented Reality (MuAR) is essential to implementing the vision of Metaverse for its capability to provide immersive and interactive experiences. In such experiences, peer positions are critical to understand each other’s intentions and actions so as to guarantee the smooth cooperation among users. However, we find that the explicit peer positions provided by the current practice could be incomplete and/or inaccurate in some situations, which leads to the weakened spatial awareness. To achieve the accurate peer tracking in MuAR, we propose a novel multiple sensors information fusion method, CSA (Coordinate System Alignment), to detect and correct defective relative positions by the current practice. CSA firstly formulates problem of correcting erroneous positions into an overdetermined system, and then finds the solution by applying the simulated annealing algorithm to expedite the search process. The evaluation results show that CSA’s ability to reduce errors significantly (58.3% on average) under long-term error duration, especially its advantage in reducing the relative direction errors. The result confirms the potential of CSA to provide reliable peer tracking in MuAR. Meanwhile, it does not impose extra restrictions on users’ practice with current mobile devices in experiences. Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Fei Li 0001, Songqing Chen |
ISM | 4 |
| 2022 | Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head AttentionabstractWe propose a method to detect individualized highlights for users on given target videos based on their preferred highlight clips marked on previous videos they have watched. Our method explicitly leverages the contents of both the preferred clips and the target videos using pre-trained features for the objects and the human activities. We design a multi-head attention mechanism to adaptively weigh the preferred clips based on their object- and human-activity-based contents, and fuse them using these weights into a single feature representation for each user. We compute similarities between these per-user feature representations and the per-frame features computed from the desired target videos to estimate the user-specific highlight clips from the target videos. We test our method on a large-scale highlight detection dataset containing the annotated highlights of individual users. Compared to current baselines, we observe an absolute improvement of 2-4% in the mean average precision of the detected highlights. We also perform extensive ablation experiments on the number of preferred highlight clips associated with each user as well as on the object- and human-activity-based feature representations to validate that our method is indeed both content-based and user-specific. Uttaran Bhattacharya, Gang Wu 0013, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Dinesh Manocha |
ACM Multimedia | 4 |
| 2022 | A Reality Check of Positioning in Multiuser Mobile Augmented Reality: Measurement and AnalysisabstractMultiuser Augmented Reality (MuAR) is essential to implementing the vision of Metaverse. With the pervasive mobile devices, MuAR enables multiple devices to share a common AR experience. In such experiences, the peer positions are critical to understand peers' intentions and actions so as to achieve the smooth interaction in AR. Such a spacial awareness requirement poses new challenges to MuAR. Traditionally, in AR experiences designed for the single user, the SLAM algorithm is adopted to compute self positions. However, the computed positions cannot be directly used to compute the relative positions of peer devices in MuAR, because they are computed with respect to independent coordinate systems associated with participating devices. To fill in the gap, the industry has recently proposed to implement peer tracking with the help of built-in Ultra Wideband (UWB) chip. In this work, we aim to perform a reality check on the proposed support, with the Nearby Interaction (NI) framework developed for iOS mobile devices as an example. The goal of our study is to gain an in-depth understanding about the reliability of the proposed support and identify potential issues. Through extensive measurements, we discover the peer tracking solution is not reliable sometimes, in terms of availability and accuracy. Furthermore, with regard to erroneous position reports, we present a quantitative analysis, summarizing the error types (e.g., transient errors and permanent errors) and revealing their underlying reasons. We believe the preliminary findings could help to improve the spacial awareness and enhance user experiences in MuAR. Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Fei Li 0001, Songqing Chen |
MMAsia | 4 |
| 2022 | Tailor Me: An Editing Network for Fashion Attribute Shape ManipulationabstractFashion attribute editing aims to manipulate fashion images based on a user-specified attribute, while preserving the details of the original image as intact as possible. Recent works in this domain have mainly focused on direct manipulation of the raw RGB pixels, which only allows to perform edits involving relatively small shape changes (e.g., sleeves). The goal of our Virtual Personal Tailoring Network (VPTNet) is to extend the editing capabilities to much larger shape changes of fashion items, such as cloth length. To achieve this goal, we decouple the fashion attribute editing task into two conditional stages: shape-then-appearance editing. To this aim, we propose a shape editing network that employs a semantic parsing of the fashion image as an interface for manipulation. Compared to operating on the raw RGB image, our parsing map editing enables performing more complex shape editing operations. Second, we introduce an appearance completion network that takes the previous stage results and completes the shape difference regions to produce the final RGB image. Qualitative and quantitative experiments on the DeepFashion-Synthesis dataset confirm that VPTNet outperforms state-of-the-art methods for both small and large shape attribute editing. Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Viswanathan (Vishy) Swaminathan, Henry Fuchs |
WACV | 5 |
| 2022 | An overview of rate control techniques in HEVC and SHVC video encoding
Ishfaq Ahmad 0001, Viswanathan (Vishy) Swaminathan, Alex Aved, Saifullah Khalid 0002 |
Multim. Tools Appl. | 2 |
| 2021 | HighlightMe: Detecting Highlights from Human-Centric VideosabstractWe present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as poses and faces. We use an autoencoder network equipped with spatial-temporal graph convolutions to detect human activities and interactions based on these modalities. We train our network to map the activity- and interaction-based latent structural representations of the different modalities to per-frame highlight scores based on the representativeness of the frames. We use these scores to compute which frames to highlight and stitch contiguous frames to produce the excerpts. We train our network on the large-scale AVA-Kinetics action dataset and evaluate it on four benchmark video highlight datasets: DSH, TVSum, PHD2, and SumMe. We observe a 4–12% improvement in the mean average precision of matching the human-annotated highlights over state-of-the-art methods in these datasets, without requiring any user-provided preferences or dataset-specific fine-tuning. Uttaran Bhattacharya, Gang Wu 0013, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Dinesh Manocha |
ICCV | 4 |
| 2021 | OSCAR-Net: Object-centric Scene Graph Attention for Image AttributionabstractImages tell powerful stories but cannot always be trusted. Matching images back to trusted sources (attribution) enables users to make a more informed judgment of the images they encounter online. We propose a robust image hashing algorithm to perform such matching. Our hash is sensitive to manipulation of subtle, salient visual details that can substantially change the story told by an image. Yet the hash is invariant to benign transformations (changes in quality, codecs, sizes, shapes, etc.) experienced by images during online redistribution. Our key contribution is OSCAR-Net1(Object-centric Scene Graph Attention for Image Attribution Network); a robust image hashing model inspired by recent successes of Transformers in the visual domain. OSCAR-Net constructs a scene graph representation that attends to fine-grained changes of every object’s visual appearance and their spatial relationships. The network is trained via contrastive learning on a dataset of original and manipulated images yielding a state of the art image hash for content fingerprinting that scales to millions of images. Eric Nguyen, Tu Bui, Viswanathan (Vishy) Swaminathan, John P. Collomosse |
ICCV | 3 |
| 2021 | Open-Domain Trending Hashtag Recommendation for VideosabstractWe describe a novel algorithm for an open-domain trending hashtag recommendation task using zero-shot hashtag prediction in an online learning paradigm. Our method utilizes joint representation learning of latent embeddings for features extracted from long-form videos and semantic embeddings of hashtags trending on social platforms. In particular, we apply graph convolutional networks to a link prediction task using videos and hashtags as nodes in a heterogeneous graph. Comparing it to the existing models for closely related tasks, we demonstrate state-of-the-art results in trending hashtag recommendations for videos. The architecture is designed to be modular in a plug-and-play fashion to enable quick and easy incorporation of the latest advances in natural language understanding and image and video processing, with a practical view to its implementation in a real-time, online setting. Swapneel Mehta, Somdeb Sarkhel, Xiang Chen 0010, Saayan Mitra, Viswanathan (Vishy) Swaminathan, Ryan Rossi, Ali Aminian, Kshitiz Garg |
ISM | 5 |
| 2021 | BOhance: Bayesian Optimization for Content EnhancementabstractWe present BOhance, an efficient solution for optimizing digital content like images. Our approach enhances the standard and widely-used method for optimizing content, A/B testing, by using Bayesian Optimization. Our work effectively extends A/B testing in the continuous domain where A/B testing cannot efficiently test infinitely many variants. We test our approach on an image enhancement task where we use iterative human feedback on different variants of an image to arrive at the optimal variant. BOhance auto-generates candidate content variants to be tested based on the human feedback on prior variants. We demonstrate with user-studies conducted on Amazon Mechanical Turk that BOhance can be both time and cost-efficient; and a superior alternative to existing solutions. Furthermore, we conduct a Visual Turing Test to obtain human impressions on the optimum variants generated by BOhance. Our experiments show that given a human-enhanced image and an image generated by BOhance, 53% users think that the BOhance image was generated by a human expert. Trisha Mittal, Viswanathan (Vishy) Swaminathan, Somdeb Sarkhel, Ritwik Sinha, David T. Arbour, Saayan Mitra, Dinesh Manocha |
ISM | 2 |
| 2021 | Full UHD 360-Degree Video Dataset and Modeling of Rate-Distortion Characteristics and Head Movement NavigationabstractWe investigate the rate-distortion (R-D) characteristics of full ultra-high definition (UHD) 360° videos and capture corresponding head movement navigation data of virtual reality (VR) headsets. We use the navigation data to analyze how users explore the 360° look-around panorama for such content and formulate related statistical models. The developed R-D characteristics and modeling capture the spatiotemporal encoding efficiency of the content at multiple scales and can be exploited to enable higher operational efficiency in key use cases. The high quality expectations for next generation immersive media necessitate the understanding of these intrinsic navigation and content characteristics of full UHD 360° videos. Jacob Chakareski, Ridvan Aksu, Viswanathan (Vishy) Swaminathan, Michael Zink |
MMSys | 3 |
| 2021 | Heterogeneous Spatial Quality for Omnidirectional VideoabstractHeterogeneous spatial quality in 360-degree videos is characterized by regions of high and low quality appearing in the same video chunk. Encoding quality-varying omnidirectional videos enables bandwidth waste optimization without compromising the quality inside the client's viewport. Heterogeneous spatial quality has been implemented using either the concept of tiling or Facebook's offset projection. In this paper, we study the preparation of heterogeneous 360-degree videos with three main contributions. We propose two novel approaches that prepare heterogeneous quality versions of a 360-degree video based on two video decomposition techniques: the Gaussian pyramid and the Laplace pyramid. We also introduce a spherical filter for omnidirectional videos which allows to further reduce the video bit-rate by exploring the spherical properties of the 360-degree videos. We compare the performance of the proposed three methods to tiling and the offset projection in representative scenarios of heterogeneous spatial quality in 360-degree videos and highlight the key trade-off between bit-rate and video quality to consider when implementing these approaches. Hristina Hristova, Gwendal Simon, Xavier Corbillon, Alisa Devlic, Viswanathan (Vishy) Swaminathan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video
Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Eunbyung Park, Viswanathan (Vishy) Swaminathan, Henry Fuchs |
ECCV (4) | 6 |
| 2020 | SumBot: Summarize Videos Like a HumanabstractVideo currently accounts for 70% of all internet traffic and this number is expected to continue to grow. Each minute, more than 500 hours worth of videos are uploaded on YouTube. Generating engaging short videos out of the raw captured content is often a time-consuming and cumbersome activity for content creators. Existing ML- based video summarization and highlight generation approaches often neglect the fact that many summarization tasks require specific domain knowledge of the video content, and that human editors often follow a semistructured template when creating the summary (e.g. to create the highlights for a sport event). We therefore address in this paper the challenge of creating domain-specific summaries, by actively leveraging this editorial template. Particularly, we present an Inverse Reinforcement Learning (IRL)-based framework that can automatically learn the hidden structure or template followed by a human expert when generating a video summary for a specific domain. Particularly, we propose to formulate the video summarization task as a Markov Decision Process, where each state is a combination of the features of the video shots added to the summary, and the possible actions are to include/remove a shot from the summary or leave it as is. Using a set of domain-specific human-generated video highlights as examples, we employ a Maximum Entropy IRL algorithm to learn the implicit reward function governing the summary generation process. The learned reward function is then used to train an RL-agent that can produce video summaries for a specific domain, closely resembling what a human expert would create. Learning from expert demonstrations allows our approach to be applicable to any domain or editorial styles. To demonstrate the superior performance of our approach, we employ it to the task of soccer games highlight generation and show that it outperforms other state-of-the-art methods, both quantitatively and qualitatively. Hongxiang Gu, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan |
ISM | 3 |
| 2020 | Closing-the-Loop: A Data-Driven Framework for Effective Video SummarizationabstractToday, videos are the primary way in which information is shared over the Internet. Given the huge popularity of video sharing platforms, it is imperative to make videos engaging for the end-users. Content creators rely on their own experience to create engaging short videos starting from the raw content. Several approaches have been proposed in the past to assist creators in the summarization process. However, it is hard to quantify the effect of these edits on the end-user engagement. Moreover, the availability of video consumption data has opened the possibility to predict the effectiveness of a video before it is published. In this paper, we propose a novel framework to close the feedback loop between automatic video summarization and its data-driven evaluation. Our Closing-The-Loop framework is composed of two main steps that are repeated iteratively. Given an input video, we first generate a set of initial video summaries. Second, we predict the effectiveness of the generated variants based on a data-driven model trained on users' video consumption data. We employ a genetic algorithm to search the space of possible summaries (i.e., adding/removing shots to the video) in an efficient way, where only those variants with the highest predicted performance are allowed to survive and generate new variants in their place. Our results show that the proposed framework can improve the effectiveness of the generated summaries with minimal computation overhead compared to a baseline solution - 28.3% more video summaries are in the highest effectiveness class than those in the baseline. Ran Xu 0003, Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Saurabh Bagchi |
ISM | 4 |
| 2020 | Recommendation for video advertisements based on personality traits and companion contentabstractPeople encounter video ads every day when they access online content. While ads can be annoying or greeted with resistance, they can also be seen as informative and enjoyable. We asked the question, what might make an ad more enjoyable? And, do people with different personality traits prefer to watch different ads --- could it be possible to better match ads and people? To answer these questions, we conducted an online study where we asked people to watch video ads of different emotional sentiments. We also measured their personality traits through an online survey. We found that the sentiment of people's preferred video ads varies significantly based on their personality traits. Additionally, we investigated when these ads are accompanied by content, how the emotional state induced by accompanying content affects people's ad preferences. We found that there was a complex relationship between people's emotional state induced by accompanying content and their ad preference when an ad highlighted either an alertness or calmness sentiment. However, when an ad highlighted activeness and amusement, the relationship was not significant. Overall, our results show that people's personality traits and their emotional states are two key elements that predict the tone of their preferred video ads. Sanorita Dey, Brittany R. L. Duff, Niyati Chhaya, Wai Fu, Viswanathan (Vishy) Swaminathan, Karrie Karahalios |
IUI | 5 |
| 2020 | Rotationally-Consistent Novel View Synthesis for HumansabstractHuman novel view synthesis aims to synthesize target views of a human subject given input images taken from one or more reference viewpoints. Despite significant advances in model-free novel view synthesis, existing methods present two major limitations when applied to complex shapes like humans. First, these methods mainly focus on simple and symmetric objects, e.g., cars and chairs, limiting their performances to fine-grained and asymmetric shapes. Second, existing methods cannot guarantee visual consistency across different adjacent views of the same object. To solve these problems, we present in this paper a learning framework for the novel view synthesis of human subjects, which explicitly enforces consistency across different generated views of the subject. Specifically, we introduce a novel multi-view supervision and an explicit rotational loss during the learning process, enabling the model to preserve detailed body parts and to achieve consistency between adjacent synthesized views. To show the superior performance of our approach, we present qualitative and quantitative results on the Multi-View Human Action (MVHA) dataset we collected (consisting of 3D human models animated with different Mocap sequences and captured from 54 different viewpoints), the Pose-Varying Human Model (PVHM) dataset, and ShapeNet. The qualitative and quantitative results demonstrate that our approach outperforms the state-of-the-art baselines in both per-view synthesis quality, and in preserving rotational consistency and complex shapes (e.g. fine-grained details, challenging poses) across multiple adjacent views in a variety of scenarios, for both humans and rigid objects. Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Henry Fuchs, Viswanathan (Vishy) Swaminathan |
ACM Multimedia | 6 |
| 2020 | 3CPS: a novel supercompression for the delivery of 3D object texturesabstractThe growing popularity of applications based on 3D rendering, such as visual effects, gaming, augmented and virtual reality, calls for the development of new solutions for the delivery of 3D objects, in particular textures. The format of texture images, which capture the characteristics of materials, has to address two constraints. First, the delivery on the Internet imposes a reduction of the image size. Second, because of memory limitations, the processing by the rendering engine is done in the GPU by extracting small areas of the image only. The format of texture images should thus enable the random-access feature for independent processing of small blocks of the images, called texels, which negatively affects the texture compression performance and, therefore, the network delivery. Hristina Hristova, Gwendal Simon, Viswanathan (Vishy) Swaminathan, Stefano Petrangeli |
MMSys | 3 |
| 2020 | QuRate: power-efficient mobile immersive video streamingabstractSmartphones have recently become a popular platform for deploying the computation-intensive virtual reality (VR) applications, such as immersive video streaming (a.k.a., 360-degree video streaming). One specific challenge involving the smartphone-based head mounted display (HMD) is to reduce the potentially huge power consumption caused by the immersive video. To address this challenge, we first conduct an empirical power measurement study on a typical smartphone immersive streaming system, which identifies the major power consumption sources. Then, we develop QuRate, a quality-aware and user-centric frame rate adaptation mechanism to tackle the power consumption issue in immersive video streaming. QuRate optimizes the immersive video power consumption by modeling the correlation between the perceivable video quality and the user behavior. Specifically, QuRate builds on top of the user's reduced level of concentration on the video frames during view switching and dynamically adjusts the frame rate without impacting the perceivable video quality. We evaluate QuRate with a comprehensive set of experiments involving 5 smartphones, 21 users, and 6 immersive videos using empirical user head movement traces. Our experimental results demonstrate that QuRate is capable of extending the smartphone battery life by up to 1.24X while maintaining the perceivable video quality during immersive video streaming. Also, we conduct an Institutional Review Board (IRB)-approved subjective user study to further validate the minimum video quality impact caused by QuRate. Nan Jiang 0020, Yao Liu 0001, Tian Guo 0001, Wenyao Xu, Viswanathan (Vishy) Swaminathan, Lisong Xu, Sheng Wei 0001 |
MMSys | 5 |
| 2020 | Optimal Bidding Strategy without Exploration in Real-time BiddingabstractMaximizing utility with a budget constraint is the primary goal for advertisers in real-time bidding (RTB) systems. The policy maximizing the utility is referred to as the optimal bidding strategy. Earlier works on optimal bidding strategy apply model-based batch reinforcement learning methods which can not generalize to unknown budget and time constraint. Further, the advertiser observes a censored market price which makes direct evaluation infeasible on batch test datasets. Previous works ignore the losing auctions to alleviate the difficulty with censored states; thus significantly modifying the test distribution. We address the challenge of lacking a clear evaluation procedure as well as the error propagated through batch reinforcement learning methods in RTB systems. We exploit two conditional independence structures in the sequential bidding process that allow us to propose a novel practical framework using the maximum entropy principle to imitate the behavior of the true distribution observed in real-time traffic. Moreover, the framework allows us to train a model that can generalize to the unseen budget conditions than limit only to those observed in history. We compare our methods on two real-world RTB datasets with several baselines and demonstrate significantly improved performance under various budget settings. Aritra Ghosh 0001, Saayan Mitra, Somdeb Sarkhel, Viswanathan (Vishy) Swaminathan |
SDM | 4 |
| 2020 | Metadata Matters in User Engagement PredictionabstractPredicting user engagement (e.g., click-through rate, conversion rate) on the display ads plays a critical role in delivering the right ad to the right user in online advertising. Existing techniques spanning Logistic Regression to Factorization Machines and their derivatives, focus on modeling the interactions among handcrafted features to predict the user engagement. Little attention has been paid on how the ad fits with the context (e.g., hosted webpage, user demographics). In this paper, we propose to include the metadata feature, which captures the visual appearance of the ad, in the user engagement prediction task. In particular, given a data sample, we combine both the basic context features, which have been widely used in existing prediction models, and the metadata feature, which is extracted from the ad using a state-of-the-art deep learning framework, to predict user engagement. To demonstrate the effectiveness of the proposed metadata feature, we compare the performance of the widely used prediction models before and after integrating the metadata feature. Our experimental results on a real-world dataset demonstrate that the metadata feature is able to further improve the prediction performance. Xiang Chen 0010, Saayan Mitra, Viswanathan (Vishy) Swaminathan |
SIGIR | 3 |
| 2019 | A Scalable Data Augmentation and Training Pipeline for Logo DetectionabstractLogo detection in images is particularly challenging due to limited access to well-labelled data. Many existing logo detection methods are not scalable to larger datasets due to tedious bounding-box annotation work. As a result, with only a small number of logo classes and limited well-labelled images per class, their performance deteriorates on real-world applications. In this work, we propose a data augmentation and training pipeline to tackle these challenges. Specifically, we develop an incremental learning approach that starts training using synthetic data, followed by iteratively obtaining real training images from a given source and updating the current model with the newly-obtained data. To avoid model drift, we add a human curation step where incorrect detections (false-positives) are filtered out by simple-clicks using a User Interface, we designed. With this approach, we were able to generate a large (173,000 images of 173 logo classes) dataset termed Logo173 where all images are annotated with bounding-boxes. This image dataset can also be used to train a frame-by-frame baseline logo detector for videos. We demonstrate with extensive experiments that the proposed pipeline significantly saves time and effort for tedious data annotation and outperforms a current state-of-the-art logo detection method. Viswanathan (Vishy) Swaminathan, Saayan Mitra |
ISM | 2 |
| 2019 | Dynamic Adaptive Streaming for Augmented Reality ApplicationsabstractAugmented Reality (AR) superimposes digital content on top of the real world, to enhance it and provide a new generation of media experiences. To provide a realistic AR experience, objects in the scene should be delivered with both high photorealism and low latency. Current AR experiences are mostly delivered with a download-and-play strategy, where the whole scene is considered a monolithic entity for delivery. This approach results in high start-up latencies and therefore a poor user experience. A similar problem in the video domain has already been tackled with the HTTP Adaptive Streaming (HAS) principle, where the video is split into segments, and a rate adaptation heuristic dynamically adapts the video quality based on the available network resources. In this paper, we apply the adaptive streaming principle from the video to the AR domain, and propose a streaming framework for AR applications. In our proposed framework, the AR objects are available at different Levels-Of-Detail (LODs) and can be streamed independently from each other. An LOD adaptation heuristic is in charge of dynamically deciding what object should be fetched from the server and at what LOD level. Our proposed heuristic prioritizes content that is more likely to be viewed by the user and selects the best LOD to maximize the object's perceived visual quality. Moreover, the adaptation takes into account the available bandwidth resources to ensure a timely delivery of the AR objects. Experiments carried out over the Internet using an AR streaming prototype developed on an iOS device allow us to show the gains brought by the proposed framework. Particularly, our approach can decrease start-up latency up to 90% with respect to a download-and-play baseline, and decrease the amount of data needed to deliver the AR experience up to 79%, without sacrificing on the visual quality of the AR objects. Stefano Petrangeli, Gwendal Simon, Viswanathan (Vishy) Swaminathan |
ISM | 4 |
| 2019 | Generative Networks for Synthesizing Human Videos in Text-Defined OutfitsabstractGenerating a video from a textual input is a challenging research topic that would have a variety of applications in industries such as retail, e-commerce, online entertainment, education etc. In this paper, we discuss the application of generating videos of a human subject in a desired outfit using an input video of the subject. We present a two stage solution, wherein at the first stage a generative model is learned such that, given the subject's image and a textual description of the outfit, a corresponding image of the subject in the described outfit is synthesized. At the second stage, all the frames of the subject's video are individually processed by the stage 1 model to generate corresponding frames and an optical flow based post processing step is performed to maintain visual coherence across the generated frames. Towards the stage-1 objective, multiple supervised and unsupervised convolutional neural network (CNN) based generative models have been proposed. A novel approach to inject an external masking layer that maintains the structural integrity of the generated images is also presented. We train and test the different methods on the publicly available multi-view clothing image data-set and the performance in videos is showcased on a set of real-world commercial videos. The experiments show the efficacy of our approach in generating images/videos in both low (64 × 64) and high (256 × 256) resolutions. Akshay Malhotra, Viswanathan (Vishy) Swaminathan, Gang Wu 0013, Ioannis D. Schizas |
MMSP | 2 |
| 2019 | Scalable Bid Landscape Forecasting in Real-Time BiddingabstractIn programmatic advertising, ad slots are usually sold using second-price (SP) auctions in real-time. The highest bidding advertiser wins but pays only the second-highest bid (known as the winning price). In SP, for a single item, the dominant strategy of each bidder is to bid the true value from the bidder's perspective. However, in a practical setting, with budget constraints, bidding the true value is a sub-optimal strategy. Hence, to devise an optimal bidding strategy, it is of utmost importance to learn the winning price distribution accurately. Moreover, a demand-side platform (DSP), which bids on behalf of advertisers, observes the winning price if it wins the auction. For losing auctions, DSPs can only treat its bidding price as the lower bound for the unknown winning price. In literature, typically censored regression is used to model such partially observed data. A common assumption in censored regression is that the winning price is drawn from a fixed variance (homoscedastic) uni-modal distribution (most often Gaussian). However, in reality, these assumptions are often violated. We relax these assumptions and propose a heteroscedastic fully parametric censored regression approach, as well as a mixture density censored network. Our approach not only generalizes censored regression but also provides flexibility to model arbitrarily distributed real-world data. Experimental evaluation on the publicly available dataset for winning price estimation demonstrates the effectiveness of our method. Furthermore, we evaluate our algorithm on one of the largest demand-side platforms and significant improvement has been achieved in comparison with the baseline solutions. Aritra Ghosh 0001, Saayan Mitra, Somdeb Sarkhel, Jason Xie, Gang Wu 0013, Viswanathan (Vishy) Swaminathan |
ECML/PKDD (3) | 6 |
| 2019 | Streaming a Sequence of Textures for Adaptive 3D Scene DeliveryabstractDelivering rich, high quality 3D scenes over the internet is challenged by the size of the 3D objects in terms of geometry and textures. This paper proposes a new method for the delivery of textures, which are encoded and delivered as a video sequence, rather than independently. Implemented on the existing video delivery infrastructure, our method provides a fine-grained control on the quality of the resulting video sequence. Gwendal Simon, Stefano Petrangeli, Nathan Carr 0001, Viswanathan (Vishy) Swaminathan |
VR | 4 |
| 2019 | An Open Initiative for the Delivery of Infinitely Scalable and Animated 3D ScenesabstractPlanet-scale Augmented Reality (AR) and photorealistic Virtual Reality (VR) are two examples of applications that require the delivery of a rich, wide, and animated 3D scene. We extend this concept to infinitely-scalable and animated 3D scenes and we propose key concepts to launch an open initiative for the delivery of these 3D scenes. We draw on the lessons learned from large-scale implementation of video streaming to design a pull-based mechanism with segmented representations of objects and periodic map updates. Gwendal Simon, Viswanathan (Vishy) Swaminathan |
VR | 2 |
| 2018 | Viewport-Driven Rate-Distortion Optimized 360º Video StreamingabstractThe growing popularity of virtual and augmented reality communications and 360° video streaming is moving video communication systems into much more dynamic and resource-limited operating settings. The enormous data volume of 360° videos requires an efficient use of network bandwidth to maintain the desired quality of experience for the end user. To this end, we propose a framework for viewport-driven rate-distortion optimized 360° video streaming that integrates the user view navigation pattern and the spatiotemporal rate-distortion characteristics of the 360° video content to maximize the delivered user quality of experience for the given network/system resources. The framework comprises a methodology for constructing dynamic heat maps that capture the likelihood of navigating different spatial segments of a 360° video over time by the user, an analysis and characterization of its spatiotemporal rate-distortion characteristics that leverage preprocessed spatial tilling of the 360° view sphere, and an optimization problem formulation that characterizes the delivered user quality of experience given the user navigation patterns, 360° video encoding decisions, and the available system/network resources. Our experimental results demonstrate the advantages of our framework over the conventional approach of streaming a monolithic uniformly-encoded 360° video and a state-of-the-art reference method. Considerable video quality gains of 4 - 5 dB are demonstrated in the case of two popular 4K 360° videos. Jacob Chakareski, Ridvan Aksu, Xavier Corbillon, Gwendal Simon, Viswanathan (Vishy) Swaminathan |
ICC | 5 |
| 2018 | From Thumbnails to Summaries-A Single Deep Neural Network to Rule Them AllabstractVideo summaries come in many forms, from traditional single-image thumbnails, animated thumbnails, storyboards, to trailer-like video summaries. Content creators use the summaries to display the most attractive portion of their videos; the users use them to quickly evaluate if a video is worth watching. All forms of summaries are essential to video viewers, content creators, and advertisers. Often video content management systems have to generate multiple versions of summaries that vary in duration and presentational forms. We present a framework ReconstSum that utilizes LSTM-based autoencoder architecture to extract and select a sparse subset of video frames or keyshots that optimally represent the input video in an unsupervised manner. The encoder selects a subset from the input video while the decoder seeks to reconstruct the video from the selection. The goal is to minimize the difference between the original input video and the reconstructed video. Our method is easily extendable to generate a variety of applications including static video thumbnails, animated thumbnails, storyboards and “trailer-like” highlights. We specifically study and evaluate two most popular use cases: thumbnail generation and storyboard generation. We demonstrate that our methods generate better results than the state-of-the-art techniques in both use cases. Hongxiang Gu, Viswanathan (Vishy) Swaminathan |
ICME | 2 |
| 2018 | BAS-360°: Exploring Spatial and Temporal Adaptability in 360-degree Videos over HTTP/2abstractToday, 360-degree video streaming has become a popular Internet service with the rise of affordable virtual reality (VR) technologies. However, streaming 360-degree videos suffers from the prohibitive bandwidth demand. Existing bandwidth-efficient solutions mainly focus on exploiting the inherent spatial adaptability of 360-degree videos, delivering only video content (spatially-cut tiles) in the viewer's region of interest (ROI) with higher quality. Temporal adaptability, which has been widely leveraged in HTTP streaming, has not been well exploited to select proper quality for video segments according to the bandwidth variations. When these two dimensions of adaptability are jointly considered, bitrate selection for the tiles become more complicated and challenging. The importance of a tile with a spatial coordination played at a specific time should be quantified so that we can determine how to allocate bandwidth for improving the viewer's quality of experience. Furthermore, viewer's head orientation prediction is highly variable, which makes the determination of important tiles highly dynamic. In addition, network fluctuations are very common on the Internet. To overcome these challenges, we propose Bi-Adaptive Streaming for 360-degree videos (BAS-360°). In BAS-360°, both spatial and temporal adaptabilities are explored in the bitrate selection for different tiles. The objective is to minimize the bandwidth waste by allocating bandwidth to more important tiles (the tiles that are more likely to be watched). To tackle the high variability of visual region prediction and the unpredictable network fluctuations, we employ two features provided by HTT P /2: stream termination and stream priority, to efficiently organize tile delivery. Evaluation results show that BAS-360° outperforms naive tile-based 360-degree video streaming strategies when network fluctuations or errors in viewport predictions occur. Mengbai Xiao, Chao Zhou 0004, Viswanathan (Vishy) Swaminathan, Yao Liu 0001, Songqing Chen |
INFOCOM | 3 |
| 2018 | Content-Based Effectiveness Prediction of Video AdvertisementsabstractAdvertisements are an integral part of internet economics and culture, and video ads are the most popular and arguably the most entertaining form of advertisements. With the recent growth in digital marketing, video ads have seen unprecedented growth and are growing in importance as an advertising means. Video ads are expensive to create and are not always effective. The effectiveness of a video ad is usually not known before its deployment, which is non-ideal for creators, advertisers, and ad platforms. In this paper, we outline an idea to provide feedback before an ad is placed on its effectiveness based on the video along with the historical data about the effectiveness of other video ads. We propose a multi-modal mixture based algorithm to predict the effectiveness automatically. Specifically, we exploit rich textual information often found with an advertisement as well as visual information to learn a finite mixture model. Our experiments on a publicly available dataset show that our approach can outperform other baseline approaches. Qi Lou, Somdeb Sarkhel, Saayan Mitra, Viswanathan (Vishy) Swaminathan |
ISM | 4 |
| 2018 | Heterogeneous Spatial Quality for Omnidirectional VideoabstractA video with heterogeneous spatial quality is a video where some regions of the frame have a different quality than other regions (for instance, a better quality could mean more pixels and less encoding distortion). Such a quality-variable encoding is a key enabler of Virtual Reality application, with 360-degree videos. So far, the main technique that has been proposed to prepare spatially heterogeneous quality is based on the concept of tiling. More recently, Facebook has implemented another approach: the offset projection where more emphasis is put on a specific direction of the frame. In this paper, we study quality-variable 360-degree videos with two main contributions. First, we provide the theoretical analysis of the offset projection and show the impact of the parameter settings on the video quality. Second, we propose another approach which consists in preparing the 360-degree video from a Gaussian pyramid of downscaled and blurred versions of the video. We perform an evaluation of tiling, offset and Gaussian-based approaches in representative scenarios of heterogeneous spatial quality in 360-degree videos and highlight the main trade-off to consider when implementing these approaches. Hristina Hristova, Xavier Corbillon, Gwendal Simon, Viswanathan (Vishy) Swaminathan, Alisa Devlic |
MMSP | 4 |
| 2017 | Digital content recommendation system using implicit feedback dataabstractMost of existing digital content recommendation systems use explicit feedback data like user's ratings. While such systems rely on user input, those do not sense the context. In this paper, we propose a framework for digital content recommendation using only implicit feedback data (i.e., information collected from session usage without any direct feedback from user), which not only considers interactions among users and contents but also various other implicit information available during a video session. To capture interactions among such attributes, we choose Higher-Order Factorization Machines (HoFM) as our predictor and test our approach on real-world video usage data. In the experiments we explore different possible factors that may affect the performance of HoFM predictor. We observe that increasing the number of sessions of users considered to build the predictor significantly improves prediction accuracy, whereas increasing the order or depth of interactions may not. We also present an application of our work to a video recommendation system. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Saayan Mitra, Ratnesh Kumar 0001 |
IEEE BigData | 2 |
| 2017 | Context-aware video recommendation based on session progress predictionabstractIn the analysis of digital content consumption, session progress provides a good alternative to using manual ratings for measuring user engagement. A good prediction of session progress is useful for optimizing and personalizing the end-user experience. Most prevalent methods of predicting session progress are based on matrix completion and only consider the interaction among users and videos, while the associated contextual information is usually not used. In this paper, we present our approach for video recommendation, based on session progress prediction and incorporating the context. We test our approach on real-world session progress data, and observe considerable improvement in prediction accuracy achieved by incorporating selected context. Our experiments also show that proper context selection and the number of observed sessions for users are two key factors affecting the prediction accuracy. Gang Wu 0013, Viswanathan (Vishy) Swaminathan, Saayan Mitra, Ratnesh Kumar 0001 |
ICME | 2 |
| 2017 | Feature Selection for FM-Based Context-Aware Recommendation SystemsabstractContext-Aware Recommendation Systems has gained lots of attention in both industry and academic research. Factorization Machines (FM) based recommendation has been successfully used in sparse industrial datasets for user personalized video recommendations. FM is a collaborative filtering technique for predicting a target such as user rating, given observations of interaction between some users and items. The model can incorporate any available auxiliary information about the user, the item, or the interaction which serves as context or features of the data. In this paper, we propose a framework to automatically select features on FM-based recommender systems to improve the prediction quality. FM requires the input data as a one-hot encoded feature vector in a binary space domain. We use the values of the FM parameters in the binary space to determine the importance of the context. We consider the density of the important features in the binary space to rank and select the relevant features in the original data. Experiments on multiple datasets have been conducted to validate the efficiency and robustness of our method. Xueyu Mao 0001, Saayan Mitra, Viswanathan (Vishy) Swaminathan |
ISM | 3 |
| 2017 | User Segment Identification Based on Similarity in Content ConsumptionabstractWith the rapid growth of online content consumption, knowing end-users and having actionable content insights has become extremely important for any online content provider. Insights from user segment identification could help in developing a content recommendation as well as new content acquisition. For advertisers, identifying segments could assist in designing ad campaigns with greater target accuracy. In this paper, we propose a new approach of finding user segments based on similarity in content consumption. We have exploited content metadata such as genres for this purpose. However, as many videos have multiple genres, the relative importance of these genres for a movie is not known. To solve this problem, we propose a two-step clustering process. First, we identify movie clusters based on metadata-based similarity. Then, based on these movie clusters, user segments are identified. We also propose a segment based recommendation system. Finally, we demonstrate the effectiveness of our approach through experiments on a large online movie database. Somdeb Sarkhel, Wreetabrata Kar, Viswanathan (Vishy) Swaminathan |
ISM | 3 |
| 2017 | Personalized Video Recommendations for Shared AccountsabstractAccount sharing is a significant problem for online recommender systems to generate accurate personalized recommendations. To solve this problem, one not only has to identify whether an account is shared, but also needs to recognize the different users sharing that account. However, to generate relevant, personalized recommendations, the particular user under a shared account has to be correctly identified at the time of delivering the recommendations. In this paper, we address this problem by first identifying users behind each account using a projection based unsupervised method, and then learning a function which can predict a user's preference accurately based on their 'contextual' information. This approach allows us to generate personalized recommendations for each of the users sharing a single account. We empirically show that on real and synthetic data set our approach performs better than other state-of-the-art approaches. Shuo Yang 0004, Somdeb Sarkhel, Saayan Mitra, Viswanathan (Vishy) Swaminathan |
ISM | 4 |
| 2017 | An HTTP/2-Based Adaptive Streaming Framework for 360° Virtual Reality VideosabstractVirtual Reality (VR) devices are becoming accessible to a large public, which is going to increase the demand for 360° VR videos. VR videos are often characterized by a poor quality of experience, due to the high bandwidth required to stream the 360° video. To overcome this issue, we spatially divide the VR video into tiles, so that each temporal segment is composed of several spatial tiles. Only the tiles belonging to the viewport, the region of the video watched by the user, are streamed at the highest quality. The other tiles are instead streamed at a lower quality. We also propose an algorithm to predict the future viewport position and minimize quality transitions during viewport changes. The video is delivered using the server push feature of the HTTP/2 protocol. Instead of retrieving each tile individually, the client issues a single push request to the server, so that all the required tiles are automatically pushed back to back. This approach allows to increase the achieved throughput, especially in mobile, high RTT networks. In this paper, we detail the proposed framework and present a prototype developed to test its performance using real-world 4G bandwidth traces. Particularly, our approach can save bandwidth up to 35% without severely impacting the quality viewed by the user, when compared to a traditional non-tiled VR streaming solution. Moreover, in high RTT conditions, our HTTP/2 approach can reach 3 times the throughput of tiled streaming over HTTP/1.1, and consistently reduce freeze time. These results represent a major improvement for the efficient delivery of 360° VR videos over the Internet. Stefano Petrangeli, Viswanathan (Vishy) Swaminathan, Mohammad Hosseini 0002, Filip De Turck |
ACM Multimedia | 2 |
| 2017 | Improving Virtual Reality Streaming using HTTP/2abstractThe demand for 360° Virtual Reality (VR) videos is expected to grow in the near future, thanks to the diffusion of VR headsets. VR Streaming is however challenged by the high bandwidth requirements of 360° videos. To save bandwidth, we spatially tile the video using the H.265 standard and stream only tiles in view at the highest quality. The video is also temporally segmented, so that each temporal segment is composed of several spatial tiles. In order to minimize quality transitions when the user moves, an algorithm is developed to predict where the user is likely going to watch in the near future. Consequently, predicted tiles are also streamed at the highest quality. Finally, the server push in HTTP/2 is used to deliver the tiled video. Only one request is sent from the client; all the tiles of a segment are automatically pushed from the server. This approach results in a better bandwidth utilization and video quality compared to traditional streaming over HTTP/1.1, where each tile has to be requested independently by the client. We showcase the benefits of our framework using a prototype developed on a Samsung Galaxy S7 and a Gear VR, which supports both tiled and non-tiled videos and streaming over HTTP/1.1 and HTTP/2. Under limited bandwidth conditions, we demonstrate how our framework can improve the quality watched by the user compared to a non-tiled solution where all of the video is streamed at the same quality. This result represents a major improvement for the efficient streaming of VR videos. Stefano Petrangeli, Filip De Turck, Viswanathan (Vishy) Swaminathan, Mohammad Hosseini 0002 |
MMSys | 3 |
| 2017 | Power Evaluation of 360 VR Video Streaming on Head Mounted Display DevicesabstractVirtual reality (VR) video streaming with 360-degree views has become a trending video application recently. While providing the users with immersive video viewing experiences, the 360 video streaming introduces significantly higher overhead than the traditional 2D video streaming in both bandwidth and power consumption, due to the additional video bytes that must be transmitted and processed. While almost all the prior work in this domain has been focused on the bandwidth optimization, we for the first time investigate the power consequence of VR streaming on head mounted displays (HMDs). In particular, we build an end-to-end VR streaming system using DASH and WebVR technologies, which enables us to conduct empirical power measurements at runtime. In order to uncover the specific power impact caused by VR video streaming, we design eight controlled test cases with various streaming configurations and derive a quantitative power breakdown of the HMD through differential power analysis. Our evaluation and analysis results indicate that the VR streaming overhead accounts for 28.5% of the total power consumption on the HMD, with 18.4% for network transmission of the extra video bytes, 3.6% for the VR video decoding, and 6.5% for the VR view calculation, generation, and rendering. Our research findings quantify the room for improvement in VR video power consumption, based on which, we propose several power optimization strategies aiming to motivate further research in low power VR streaming. Nan Jiang 0020, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001 |
NOSSDAV | 2 |
| 2016 | Adaptive 360 VR Video Streaming: Divide and ConquerabstractWhile traditional multimedia applications such as games and videos are still popular, there has been a significant interest in the recent years towards new 3D media such as 3D immersion and Virtual Reality (VR) applications, especially 360 VR videos. 360 VR video is an immersive spherical video where the user can look around during playback. Unfortunately, 360 VR videos are extremely bandwidth intensive, and therefore are difficult to stream at acceptable quality levels. In this paper, we propose an adaptive bandwidth-efficient 360 VR video streaming system using a divide and conquer approach. We propose a dynamic view-aware adaptation technique to tackle the huge bandwidth demands of 360 VR video streaming. We spatially divide the videos into multiple tiles while encoding and packaging, use MPEG-DASH SRD to describe the spatial relationship of tiles in the 360-degree space, and prioritize the tiles in the Field of View (FoV). In order to describe such tiled representations, we extend MPEG-DASH SRD to the 3D space of 360 VR videos. We spatially partition the underlying 3D mesh, and construct an efficient 3D geometry mesh called hexaface sphere to optimally represent a tiled 360 VR video in the 3D space. Our initial evaluation results report up to 72% bandwidth savings on 360 VR video streaming with minor negative quality impacts compared to the baseline scenario when no adaptations is applied. Mohammad Hosseini 0002, Viswanathan (Vishy) Swaminathan |
ISM | 2 |
| 2016 | Adaptive 360 VR Video Streaming Based on MPEG-DASH SRDabstractWe demonstrate an adaptive bandwidth-efficient 360 VR video streaming system based on MPEG-DASH SRD. We extend MPEG-DASH SRD to the 3D space of 360 VR videos, and showcase a dynamic view-aware adaptation technique to tackle the high bandwidth demands of streaming 360 VR videos to wireless VR headsets. We spatially partition the underlying 3D mesh into multiple 3D sub-meshes, and construct an efficient 3D geometry mesh called hexaface sphere to optimally represent tiled 360 VR videos in the 3D space. We then spatially divide the 360 videos into multiple tiles while encoding and packaging, use MPEG-DASH SRD to describe the spatial relationship of tiles in the 3D space, and prioritize the tiles in the Field of View (FoV) for view-aware adaptation. Our initial evaluation results show that we can save up to 72% of the required bandwidth on 360 VR video streaming with minor negative quality impacts compared to the baseline scenario when no adaptations is applied. Mohammad Hosseini 0002, Viswanathan (Vishy) Swaminathan |
ISM | 2 |
| 2016 | Audience Validation from Demographic Mix and Insufficient Individual DataabstractAn advertiser usually provides a specific demographic target to a content publisher. But due to the limited profile information on each viewer, the publisher suffers from low accuracy in targeting. Once an advertisement is shown to viewers, a third party (like Nielsen, Comscore) validates how close the publisher was to the target audience. Publisher also receives the demographic mix of each show from the third party. Low accuracy in targeting is expensive as the publisher gets paid only for the impressions which were on target. In this work, we propose a new approach to incorporate user level latent features developed from the show-wise demographic mix provided by the third party, along with session level features of viewers to improve the demographic predictions. In congruence with current industry practice, we train our model on a small labeled demographic segment and establish the effectiveness of our approach over existing approaches through experiments. Wreetabrata Kar, Sarathkrishna Swaminathan, Viswanathan (Vishy) Swaminathan |
ISM | 3 |
| 2016 | DASH2M: Exploring HTTP/2 for Internet Streaming to Mobile DevicesabstractToday HTTP/1.1 is the most popular vehicle for delivering Internet content, including streaming video. Standardized in 2015 with a few new features, HTTP/2 is gradually replacing HTTP 1.1 to improve user experience. Yet, how HTTP/2 can help improve the video streaming delivery has not been thoroughly investigated. In this work, we set to investigate how to utilize the new features offered by HTTP/2 for video streaming over the Internet, focusing on the streaming delivery to mobile devices as, today, more and more users watch video on their mobile devices. For this purpose, we design DASH2M, Dynamic Adaptive Streaming over HTTP/2 to Mobile Devices. DASH2M deliberately schedules the streaming content delivery by comprehensively considering the user's Quality of Experience (QoE), the dynamics of the network resources, and the power efficiency on the mobile devices. Experiments based on an implemented prototype show that DASH2M can outperform prior strategies for users' QoE while minimizing the battery power consumption on mobile devices. Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen |
ACM Multimedia | 2 |
| 2016 | Two-way real time multimedia stream authentication using physical unclonable functionsabstractMultimedia authentication is an integral part of multimedia signal processing in many real-time and security sensitive applications, such as video surveillance. In such applications, a full-fledged video digital rights management (DRM) mechanism is not applicable due to the real time requirement and the difficulties in incorporating complicated license/key management strategies. This paper investigates the potential of multimedia authentication from a brand new angle by employing hardware-based security primitives, such as physical unclonable functions (PUFs). We show that the hardware security approach is not only capable of accomplishing the authentication for both the hardware device and the multimedia stream but, more importantly, introduce minimum performance, resource, and power overhead. We justify our approach using a prototype PUF implementation on Xilinx FPGA boards. Our experimental results on the real hardware demonstrate the high security and low overhead in multimedia authentication obtained by using hardware security approaches. Mehrdad Zaker Shahrak, Mengmei Ye, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001 |
MMSP | 3 |
| 2016 | Evaluating and improving push based video streaming with HTTP/2abstractThe sever-initiated push mechanism is one of the most prominent features in the next generation HTTP/2 protocol, having shown its capability on saving network traffic and improving the web page retrieval latency. Our prior work has investigated the server push-based mechanism for HTTP video streaming and proposed a k-push scheme, where the server pushes k video segments following the response to a request. In this study, we further conduct an analysis and evaluation of the k-push scheme in HTTP streaming. Our results uncover that the push mechanism can efficiently increase the network utilization (under certain conditions) compared to regular HTTP streaming. However the results also show that the k-push scheme deteriorates network adaptability and leads to the "over-push" problem, in which the pushed video content waste network resources due to user abandonment behaviors. To overcome these limitations, we propose a new " adaptive-push" scheme, which dynamically adjusts the parameter k to adapt to the runtime environment. To evaluate the performance of adaptive-push, we implemented a prototype system. The experimental results show that compared to k-push, adaptive-push can improve the network adaptability. Furthermore, our real-world trace based simulation results show that adaptive-push can effectively alleviate the over-push problem. Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen |
NOSSDAV | 2 |
| 2015 | Power efficient mobile video streaming using HTTP/2 server pushabstractThis paper proposes a power efficient video streaming mechanism on mobile devices over cellular networks. We first develop an analytical model to identify and quantify the power inefficiency in mobile video streaming, due to the mismatch between HTTP request schedule and the radio resource control schedule. Based on the analytical model, we develop a low power video streaming mechanism by employing the server push technology available in the HTTP/2 protocol. We implemented the server push-based low power streaming mechanism in an HTTP DASH video streaming prototype involving mobile devices and the 4G/LTE cellular network. Our experiments show significant battery power savings on mobile devices using our server push strategy. Sheng Wei 0001, Viswanathan (Vishy) Swaminathan, Mengbai Xiao |
MMSP | 2 |
| 2015 | Selection and Ordering of Linear Online Video AdsabstractThis paper studies the selection and ordering of in-stream ads in videos shown in online content publishers. We propose an allocation algorithm that uses a collective measure of price and quality for each ad and factors in slot-specific continuation probabilities to maximize publisher revenue. The algorithm is based on cascade models and uses a dynamic programming method to assign linear (video) ads to slots in an online video. The approach accounts for the negative externality created by lower quality ads placed in a video, leading to viewer exit and thereby preventing the publisher from showing the subsequent ads scheduled in that session. Our algorithm is scalable and suited for real-time applications. A large log of viewer activity from a video ad platform is used to empirically test the algorithm. A series of simulations show that our algorithm, when compared to other algorithms currently practiced in industry, generates more revenue for the publisher and increases viewer retention. Wreetabrata Kar, Viswanathan (Vishy) Swaminathan, Paulo Albuquerque |
RecSys | 2 |
| 2014 | Cost effective video streaming using server push over HTTP 2.0abstractThe Hypertext Transfer Protocol (HTTP) has been widely adopted and deployed as the key protocol for video streaming over the Internet. One of the consequences of leveraging traditional HTTP for video streaming is the significantly increased request overhead due to the segmentation of the video content into HTTP resources. The overhead becomes even more significant when non-multiplexed video and audio segments are deployed. In this paper, we investigate and address the request overhead problem by employing the server push technology in the new HTTP 2.0 protocol. In particular, we develop a set of push strategies that actively deliver video and audio content from the HTTP server without requiring a request for each individual segment. We evaluate our approach in a Dynamic Adaptive Streaming over HTTP (DASH) streaming system. We show that the request overhead can be significantly reduced by using our push strategies. Also, we validate that the server push based approach is compatible with the existing HTTP streaming features, such as adaptive bitrate switching. Sheng Wei 0001, Viswanathan (Vishy) Swaminathan |
MMSP | 2 |
| 2014 | Low Latency Live Video Streaming over HTTP 2.0abstractHypertext Transfer Protocol (HTTP) has been widely adopted as a scalable and efficient protocol for streaming video content over the Internet. HTTP streaming clients receive a manifest file, download the referred video segments over HTTP, and play them back seamlessly emulating video streaming. This introduces at least one segment duration latency making HTTP streaming unsuitable for live video streaming use cases that require low latencies. The straightforward solution to lower live latency that reduces segment duration leads to an explosion in the number of HTTP requests, as well as inefficient deployment of assets in HTTP caches. To solve this problem, we develop a low latency live video streaming technique over HTTP 2.0. In particular, we employ the new server push feature in HTTP 2.0 to stream the live video actively from the web server to the client, as soon as the video segments become available. We implement this server push based low latency mechanism in a MPEG Dynamic Adaptive Streaming over HTTP (DASH) prototype. Our experimental results indicate performance gains in live latency using the server push scheme. More importantly, by leveraging the server push feature in HTTP 2.0, we are able to avoid the request explosion problem while lowering latency by reducing the segment duration. Sheng Wei 0001, Viswanathan (Vishy) Swaminathan |
NOSSDAV | 2 |
| 2013 | Designing a universal format for encrypted mediaabstractIncreasingly, video delivery over the internet is being monetized through advertisement or paid services. Almost all monetized video is delivered encrypted to the clients. Clients are authorized to receive the encryption keys after watching advertisements or based on payments. Although the exact same audio and video compression standards are used, the way media is encrypted is very different in different eco-systems like Adobe Flash, Apple HTTP Live Streaming, MPEG Dynamic Adaptive Streaming over HTTP, etc., making them incompatible with each other. In some cases, it is sample (frame) based encryption while it is packet based in others. Even while using sample based encryption, different parts of a sample are selectively encrypted by different schemes. As encryption algorithms typically use different chaining modes (e.g., Cipher Block Chaining) there is some continuity from one encryption block to another. Although there is considerable overlap between encrypted data, the encryption chains are constructed differently across different formats. The goal of this paper is to define a single mezzanine file format to serve as a Universal Encryption Format that a client platform can implement to playback media from different encrypted video delivery systems. We use some basic encryption characteristics to identify and preserve the chains by storing minimal house keeping information about which part of the data is encrypted, where chains are broken, additional initialization vectors, etc. We design and propose an encryption map that is stored typically with the media sample headers. The map securely and efficiently stores information about encryption runs in terms of sizes, IVs, and offsets. We propose additional optimizations in the format to make the encryption maps compact and the decryption, efficient. Viswanathan (Vishy) Swaminathan, Saayan Mitra, Sheng Wei 0001 |
MMSP | 1 |
| 2013 | Are we in the middle of a video streaming revolution?abstractIt has been roughly 20 years since the beginning of video streaming over the Internet. Until very recently, video streaming experiences left much to be desired. Over the last few years, this has significantly improved making monetization of streaming, possible. Recently, there has been an explosion of commercial video delivery services over the Internet, sometimes referred to as over-the-top (OTT) delivery. All these services invariably use streaming technologies. Initially, streaming had all the promise, then for a long time, it was download and play, later progressive download for short content, and now it is streaming again. Did streaming win the download versus streaming contest? Did the best technology win? The improvement in streaming experience has been possible through a variety of new streaming technologies, some proprietary and others extensions to standard protocols. The primary delivery mechanism for entertainment video, both premium content like movies and user generated content (UGC), tends to be HTTP streaming. Is HTTP streaming the panacea for all problems? The goal of this article is to give an industry perspective of what fundamentally changed in video streaming that makes it commercially viable now. This article outlines how a blend of technology choices between download and streaming makes the current wave of ubiquitous streaming possible for entertainment video delivery. After identifying problems that still need to be solved, the article concludes with the lessons learnt from the video streaming evolution. Viswanathan (Vishy) Swaminathan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | An optimal client buffer model for multiplexing HTTP streamsabstractThe basic tenet of HTTP streaming is to deliver fragments of video and audio that are individually addressable chunks of content over HTTP. Some media players consume incoming video and audio data only in a time ordered multiplexed format. If alternate tracks need to be added post packaging of the media, it has to be repackaged that involves duplication resulting in multiple multiplexed files. Additionally for adaptive streaming, a set of all those files need to be added for each bitrate. Alternatively, it is more efficient to store component tracks separately, fetching only the required tracks and multiplexing audio and video in the client before sending the data to the decoder. To deliver an optimal viewing experience, the client has to take care of the seemingly conflicting constraints viz., handling the network jitter, minimizing the time to switch to an alternate track and minimizing the live latency. For instance, to absorb more network jitter more data should be available in the buffers but this would increase the switching latency. We introduce a formal buffer model for a client that gathers video and audio fragments and multiplexes them on the fly. This model uses separate video and audio buffers, a multiplexed buffer in the application, and decoding buffer associated with the decoder. We model the buffer sizes, their thresholds to request data from the network, and the rate of transfer of data between buffers. We show that these buffers can be designed varying these parameters to optimize for the above constraints. This buffer model can also be leveraged for deciding when to switch in adaptive bitrate streaming. We further validate these by experimental results from our implementation. Saayan Mitra, Viswanathan (Vishy) Swaminathan |
MMSP | 2 |
| 2011 | Low latency live video streaming using HTTP chunked encodingabstractHypertext transfer protocol (HTTP) based streaming solutions for live video and video on demand (VOD) applications have become available recently. However, the existing HTTP streaming solutions cannot provide a low latency experience due to the fact that inherently in all of them, latency is tied to the duration of the media fragments that are individually requested and obtained over HTTP. We propose a low latency HTTP streaming approach using HTTP chunked encoding, which enables the server to transmit partial fragments before the entire video fragment is published. We develop an analytical model to quantify and compare the live latencies in three HTTP streaming approaches. Then, we present the details of our experimental setup and implementation. Both the analysis and experimental results show that the chunked encoding approach is capable of reducing the live latency to one to two chunk durations and that the resulting live latency is independent of the fragment duration. Viswanathan (Vishy) Swaminathan, Sheng Wei 0001 |
MMSP | 1 |
| 2000 | MPEG-J: Java application engine in MPEG-4abstractMPEG-4 Systems version 1 defines the presentation engine for the MPEG-4 terminal. MPEG-J is the Java based application engine defined in the version 2 of the MPEG-4 standard. This facilitates programmatic control as opposed to the parametric control provided by the presentation engine. MPEG-J refers to the collection of Java APIs defined by the Systems subgroup of MPEG. These APIs are used to access and control the underlying MPEG-4 terminal. The MPEG-J framework also defines the delivery mechanism and the lifecycle of applications. These applications can be local or remote (MPEGlet). They use the MPEG-J APIs to control the MPEG-4 terminal. MPEG-J can be used to include behavior based on time-varying terminal conditions in the content. Viswanathan (Vishy) Swaminathan, Gerard Fernando |
ISCAS | 1 |