EDBT 2026 Demo / reviewers in the wild / expert
Wei Tsang Ooi
dblp:93/347
· DBLP profile ↗
146ranked-venue papers
4as first author
43since 2021 · last 2026
0000-0001-8994-1736ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 4 first-author · 23 since 2021Computer networks · 34 · 5 since 2021Human-computer interaction and ubiquitous computing · 13 · 9 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Systems, architecture and hardware · 8Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesabstractGenerative world models have become essential data engines for autonomous driving, yet most focus on videos or occupancy grids and overlook the unique challenges of LiDAR. Extending LiDAR generation to dynamic 4D modeling requires addressing controllability, temporal coherence, and standardized evaluation. We present LiDARCrafter, a unified framework for controllable 4D LiDAR generation and editing. Free-form language instructions are converted into ego-centric scene graphs that guide a tri-branch diffusion model to generate object geometry, motion, and structural priors. An autoregressive module further produces temporally coherent and stable LiDAR sequences with improved global consistency. To enable fair comparison, we introduce a comprehensive benchmark covering scene-, object-, and sequence-level metrics for rigorous and reproducible evaluation. Experiments on nuScenes show that LiDARCrafter achieves state-of-the-art fidelity, controllability, and temporal consistency, paving the way for scalable data augmentation and realistic simulation in diverse scenarios. Code have been publicly available at https://lidarcrafter.github.io. Alan Liang, Youquan Liu, Dongyue Lu, Lingdong Kong, Huaici Zhao, Wei Tsang Ooi |
AAAI | 8 |
| 2026 | Towards LLM-powered Assistive Drone for Blind and Low Vision UsersabstractDrones have gained traction as a versatile form of assistive robots for Blind and Low Vision (BLV) people. Nonetheless, novel interaction techniques are required to enable BLV people to communicate with drones naturally. In this work, we built an LLM-powered assistive drone for BLV users. We leverage an LLM to translate high-level user goals to step-by-step instructions for the drone and to extract visual information from the images. Through a formative study with BLV users (N=9), we identified envisioned use cases and desired interaction modalities. Then, we took a participatory and iterative approach to build a prototype, incorporating feedback received from 3 BLV users, as well as 5 domain experts. Finally, we conducted a user study with an additional 6 BLV participants to evaluate the iterated prototype, and received positive feedback. This work is contributing to a growing body of research on harnessing the power of LLMs to build a more inclusive world. Yize Wei, Ibnu Taimiyyah Bin Adam, Hanjun Wu, Moritz Messerschmidt, Wei Tsang Ooi, Christophe Jouffrais, Suranga Nanayakkara |
CHI | 5 |
| 2026 | P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and VideoabstractGaussian splatting has emerged as a competitive explicit representation for image and video reconstruction. In this work, we present P-GSVC, the first layered progressive 2D Gaussian splatting framework that provides a unified solution for scalable Gaussian representation in both images and videos. P-GSVC organizes 2D Gaussian splats into a base layer and successive enhancement layers, enabling coarse-to-fine reconstructions. To effectively optimize this layered representation, we propose a joint training strategy that simultaneously updates Gaussians across layers, aligning their optimization trajectories to ensure inter-layer compatibility and a stable progressive reconstruction. P-GSVC supports scalability in terms of both quality and resolution. Our experiments show that the joint training strategy can gain up to 1.9 dB improvement in PSNR for video and 2.6 dB improvement in PSNR for image when compared to methods that perform sequential layer-wise training. Longan Wang, Yuang Shi, Wei Tsang Ooi |
MMSys | 3 |
| 2026 | SEE4D: Pose-Free 4D Generation via Auto-Regressive Video InpaintingabstractAbstract Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video‐to‐4D methods typically rely on manually annotated camera poses, which are labor‐intensive and brittle for in‐the‐wild footage. Recent warp‐then‐inpaint approaches mitigate the need for pose labels by warping input frames along a novel camera trajectory and using an inpainting model to fill missing regions, thereby depicting the 4D scene from diverse viewpoints. However, this trajectory‐to‐trajectory formulation often entangles camera motion with scene dynamics and complicates both modeling and inference. We introduce S ee 4D , a pose‐free, trajectory‐to‐camera framework that replaces explicit trajectory prediction with rendering to a bank of fixed virtual cameras, thereby separating camera control from scene modeling. A view‐conditional video inpainting model is trained to learn a robust geometry prior by denoising realistically synthesized warped images and to inpaint occluded or missing regions across virtual viewpoints, eliminating the need for explicit 3D annotations. Building on this inpainting core, we design a spatiotemporal autoregressive inference pipeline that traverses virtual‐camera splines and extends videos with overlapping windows, enabling coherent generation at bounded per‐step complexity. We validate See4D on cross‐view video generation and sparse reconstruction benchmarks. Across quantitative metrics and qualitative assessments, our method achieves superior generalization and improved performance relative to pose‐ or trajectory‐conditioned baselines, advancing practical 4D world modeling from casual videos. Dongyue Lu, Ao Liang, Tianxin Huang, Baorui Ma, Liang Pan, Wei Yin 0006, Lingdong Kong, Wei Tsang Ooi, Ziwei Liu 0002 |
Comput. Graph. Forum | 10 |
| 2026 | Graph-Based Event and Sub-Event Grouping in User-Generated Videos with Distortion-Aware Keyframe ClusteringabstractUser-generated content (UGC) videos recorded in uncontrolled environments often exhibit blur, camera shake, lighting fluctuations, and large viewpoint differences, making it difficult to organize multiple recordings of the same event. This work proposes an integrated pipeline that groups UGC videos into coherent events and sub-events by combining distortion-aware keyframe selection, adaptive audio–visual fusion, and confidence-weighted graph construction. The method first filters and clusters segment-level representations to obtain reliable keyframes, then fuses audio and visual cues through a lightweight gating module to produce robust multimodal descriptors. These descriptors populate a similarity graph whose strong and weak edges reveal sub-event and event structure without requiring shot boundaries or manual segmentation. Although distortion modeling, keyframe extraction, and multimodal similarity have been studied separately, existing approaches do not integrate them for hierarchical UGC video grouping. Experiments on the JIKU dataset and a curated YouTube dataset show consistent improvements in fidelity, diversity, and clustering metrics, demonstrating the applicability of the approach to video summarization, multi-view organization, and other UGC analysis tasks. Malya Singh, Wei Tsang Ooi, Abdulmotaleb El Saddik, Mukesh Saini |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | LapisGS: Layered Progressive 3D Gaussian Splatting for Adaptive StreamingabstractThe rise of Extended Reality$(X R)$requires efficient streaming of 3D online worlds, challenging current 3DGS representations to adapt to bandwidth-constrained environments. This paper proposes LapisGS, a layered 3DGS that supports adaptive streaming and progressive rendering. Our method constructs a layered structure for cumulative representation, incorporates dynamic opacity optimization to maintain visual fidelity, and utilizes occupancy maps to efficiently manage Gaussian splats. This proposed model offers a progressive representation supporting a continuous rendering quality adapted for bandwidth-aware streaming. Extensive experiments validate the effectiveness of our approach in balancing visual fidelity with the compactness of the model, with up to 50.71 % improvement in SSIM, 286.53% improvement in LPIPS with 23% of the original model size, and shows its potential for bandwidth-adapted 3D streaming and rendering applications. Project page: https://yuang-ian.github.io/lapisgs/ Yuang Shi, Géraldine Morin, Simone Gasparini, Wei Tsang Ooi |
3DV | 4 |
| 2025 | Human Robot Interaction for Blind and Low Vision People: A Systematic Literature ReviewabstractInternational audience Yize Wei, Nathan Rocher, Chitralekha Gupta, Mia Huong Nguyen, Roger Zimmermann, Wei Tsang Ooi, Christophe Jouffrais, Suranga Nanayakkara |
CHI | 6 |
| 2025 | SafeSpect: Safety-First Augmented Reality Heads-up Display for Drone InspectionsabstractInternational audience Peisen Xu, Jérémie Garcia, Wei Tsang Ooi, Christophe Jouffrais |
CHI | 3 |
| 2025 | EventFly: Event Camera Perception from Ground to the SkyabstractCross-platform adaptation in event-based dense perception is crucial for deploying event cameras across diverse settings, such as vehicles, drones, and quadrupeds, each with unique motion dynamics, viewpoints, and class distributions. In this work, we introduce EventFly, a framework for robust cross-platform adaptation in event camera perception. Our approach comprises three key components: i) Event Activation Prior (EAP), which identifies high-activation regions in the target domain to minimize prediction entropy, fostering confident, domain-adaptive predictions; ii) EventBlend, a data-mixing strategy that integrates source and target event voxel grids based on EAP-driven similarity and density maps, enhancing feature alignment; and iii) EventMatch, a dual-discriminator technique that aligns features from source, target, and blended domains for better domain-invariant learning. To holistically assess cross-platform adaptation abilities, we introduce EXPo, a large-scale benchmark with diverse samples across vehicle, drone, and quadruped platforms. Extensive experiments validate our effectiveness, demonstrating substantial gains over popular adaptation methods. We hope this work can pave the way for more adaptive, high-performing event perception across diverse and complex environments. Lingdong Kong, Dongyue Lu, Xiang Xu 0009, Lai Xing Ng, Wei Tsang Ooi, Benoit Cottereau |
CVPR | 5 |
| 2025 | Perspective-Invariant 3D Object DetectionabstractWith the rise of robotics, LiDAR-based 3D object detection has garnered significant attention in both academia and industry. However, existing datasets and methods predominantly focus on vehicle-mounted platforms, leaving other autonomous platforms underexplored. To bridge this gap, we introduce Pi3DET, the first benchmark featuring LiDAR data and 3D bounding box annotations collected from multiple platforms: vehicle, quadruped, and drone, thereby facilitating research in 3D object detection for non-vehicle platforms as well as cross-platform 3D detection. Based on Pi3DET, we propose a novel cross-platform adaptation framework that transfers knowledge from the well-studied vehicle platform to other platforms. This framework achieves perspective-invariant 3D detection through robust alignment at both geometric and feature levels. Additionally, we establish a benchmark to evaluate the resilience and robustness of current 3D detectors in cross-platform scenarios, providing valuable insights for developing adaptive 3D perception systems. Extensive experiments validate the effectiveness of our approach on challenging cross-platform tasks, demonstrating substantial gains over existing adaptation methods. We hope this work paves the way for generalizable and unified 3D perception systems across diverse and complex environments. Our Pi3DET dataset, cross-platform benchmark suite, and annotation toolkit have been made publicly available. Ao Liang, Lingdong Kong, Dongyue Lu, Youquan Liu, Huaici Zhao, Wei Tsang Ooi |
ICCV | 7 |
| 2025 | Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and GuaranteesabstractTwo-stage Learning-to-Defer (L2D) enables optimal task delegation by assigning each input to either a fixed main model or one of several offline experts, supporting reliable decision-making in complex, multi-agent environments. However, existing L2D frameworks assume clean inputs and are vulnerable to adversarial perturbations that can manipulate query allocation—causing costly misrouting or expert overload. We present the first comprehensive study of adversarial robustness in two-stage L2D systems. We introduce two novel attack strategies—untargeted and targeted—which respectively disrupt optimal allocations or force queries to specific agents. To defend against such threats, we propose SARD, a convex learning algorithm built on a family of surrogate losses that are provably Bayes-consistent and $(\mathcal{R}, \mathcal{G})$-consistent. These guarantees hold across classification, regression, and multi-task settings. Empirical results demonstrate that SARD significantly improves robustness under adversarial attacks while maintaining strong clean performance, marking a critical step toward secure and trustworthy L2D deployment. Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi |
ICML | 4 |
| 2025 | A Two-Stage Learning-to-Defer Approach for Multi-Task LearningabstractThe Two-Stage Learning-to-Defer (L2D) framework has been extensively studied for classification and, more recently, regression tasks. However, many real-world applications require solving both tasks jointly in a multi-task setting. We introduce a novel Two-Stage L2D framework for multi-task learning that integrates classification and regression through a unified deferral mechanism. Our method leverages a two-stage surrogate loss family, which we prove to be both Bayes-consistent and $(\mathcal{G}, \mathcal{R})$-consistent, ensuring convergence to the Bayes-optimal rejector. We derive explicit consistency bounds tied to the cross-entropy surrogate and the $L_1$-norm of agent-specific costs, and extend minimizability gap analysis to the multi-expert two-stage regime. We also make explicit how shared representation learning—commonly used in multi-task models—affects these consistency guarantees. Experiments on object detection and electronic health record analysis demonstrate the effectiveness of our approach and highlight the limitations of existing L2D methods in multi-task scenarios. Yannis Montreuil, Yeo Shu Heng, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi |
ICML | 5 |
| 2025 | FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong, Roger Wattenhofer, Wei Tsang Ooi |
AAMAS | 5 |
| 2025 | Video Lecture Analysis Toolkit: An Open-Source Framework for Interactive LearningabstractThe growth of online educational content, particularly slide-based video lectures, has created a need for tools that enhance navigation, comprehension, and accessibility. Many existing systems for video analysis are closed-source, hindering reproducibility and extension. To address this, we present the Lecture Video Analysis Toolkit, an open-source, proof-of-concept application designed for the multimodal analysis of slide video lectures. The toolkit integrates a processing pipeline that includes scene detection, visual entity extraction, transcription, optical character recognition (OCR), and semantic linking between spoken and visual content using embeddings. A key contribution is its interactive interface, motivated by early user feedback, that allows for customization of the viewing experience to suit individual preferences. The entire system is openly available and serves as a research prototype for validating the potential of multimodal analysis in creating more inclusive and improved learning experiences. A live demo is accessible at https://travis-seng.fr/svla, and the source code is openly available at https://github.com/travisseng/svla-toolkit. Travis Seng, Axel Carlier, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2025 | LTS: A DASH Streaming System for Dynamic Multi-Layer 3D Gaussian Splatting ScenesabstractWe present a novel DASH-based streaming system for dynamic 3D Gaussian Splatting (3DGS) scenes, addressing the challenges of streaming large amounts of 3DGS data over diverse and dynamic networks. Our Layer, Tile, and Segment Adaptive streaming (LTS) system combines three key features: (i) multi-layer streaming, which adapts to diverse client capabilities while balancing visual quality and bandwidth usage, (ii) tiled streaming, which reduces unnecessary data transmission by focusing on the user's viewport, and (iii) segment streaming, which divides dynamic 3DGS scenes into segments, letting clients request them dynamically to handle network fluctuations. Our experimental results demonstrate that our LTS system achieves superior performance in both live and on-demand streaming of dynamic 3DGS scenes compared to the baselines. For example, in live streaming, LTS could achieve up to 99.70% reduction in missing frames on average and deliver a maximum PSNR (Peak Signal-to-Noise Ratio) improvement of 10.08 dB. In on-demand streaming, LTS could reduce the freeze time by up to 92.01%, and increase the synthesized view quality by up to 5.14 dB in PSNR and 0.11 in SSIM (Structural Similarity Index). Our source codes are available at: https://github.com/AIINS-NTHU/LTS-DASH-Streaming-System-for-3DGS. Yuan-Chun Sun, Yuang Shi, Cheng-Tse Lee, Mufeng Zhu, Wei Tsang Ooi, Yao Liu 0001, Chun-Ying Huang, Cheng-Hsin Hsu |
MMSys | 5 |
| 2025 | Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasabstractEvent cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven object grounding in event-based perception. Built from real-world driving data, Talk2Event provides over 30,000 validated referring expressions, each enriched with four grounding attributes -- appearance, status, relation to viewer, and relation to other objects -- bridging spatial, temporal, and relational reasoning. To fully exploit these cues, we propose EventRefer, an attribute-aware grounding framework that dynamically fuses multi-attribute representations through a Mixture of Event-Attribute Experts (MoEE). Our method adapts to different modalities and scene dynamics, achieving consistent gains over state-of-the-art baselines in event-only, frame-only, and event-frame fusion settings. We hope our dataset and approach will establish a foundation for advancing multimodal, temporally-aware, and language-driven perception in real-world robotics and autonomy. Lingdong Kong, Dongyue Lu, Alan Liang, Yuhao Dong, Tianshuai Hu, Lai Xing Ng, Wei Tsang Ooi, Benoit Cottereau |
NeurIPS | 8 |
| 2025 | FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational FrequenciesabstractEvent cameras offer unparalleled advantages for real-time perception in dynamic environments, thanks to the microsecond-level temporal resolution and asynchronous operation. Existing event detectors, however, are limited by fixed-frequency paradigms and fail to fully exploit the high-temporal resolution and adaptability of event data. To address these limitations, we propose FlexEvent, a novel framework that enables detection at varying frequencies. Our approach consists of two key components: FlexFuse, an adaptive event-frame fusion module that integrates high-frequency event data with rich semantic information from RGB frames, and FlexTune, a frequency-adaptive fine-tuning mechanism that generates frequency-adjusted labels to enhance model generalization across varying operational frequencies. This combination allows our method to detect objects with high accuracy in both fast-moving and static scenarios, while adapting to dynamic environments. Extensive experiments on large-scale event camera datasets demonstrate that our approach surpasses state-of-the-art methods, achieving significant improvements in both standard and high-frequency settings. Notably, our method maintains robust performance when scaling from 20 Hz to 90 Hz and delivers accurate detection up to 180 Hz, proving its effectiveness in extreme conditions. Our framework sets a new benchmark for event-based object detection and paves the way for more adaptable, real-time vision systems. Dongyue Lu, Lingdong Kong, Gim Hee Lee, Camille Simon 0001, Wei Tsang Ooi |
NeurIPS | 5 |
| 2025 | GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splattingabstract3D Gaussian splats have emerged as a revolutionary, effective, learned representation for static 3D scenes. This work explores using 2D Gaussian splats as a new primitive for representing videos. We propose GSVC, an approach to learning a set of 2D Gaussian splats that can effectively represent and compress video frames. GSVC incorporates the following techniques: (i) To exploit temporal redundancy among adjacent frames, which can speed up training and improve the compression efficiency, we predict the Gaussian splats of a frame based on its previous frame; (ii) To control the trade-offs between file size and quality, we remove Gaussian splats with low contribution to the video quality; (iii) To capture dynamics in videos, we randomly add Gaussian splats to fit content with large motion or newly-appeared objects; (iv) To handle significant changes in the scene, we detect key frames based on loss differences during the learning process. Experiment results show that GSVC achieves good rate-distortion trade-offs, comparable to state-of-the-art video codecs such as AV1 and VVC, and a rendering speed of 1500 fps for a 1920×1080 video. Project page: https://yuang-ian.github.io/gsvc/. Longan Wang, Yuang Shi, Wei Tsang Ooi |
NOSSDAV | 3 |
| 2025 | RBMark: Robust and blind video watermark in DT CWT domain
I-Chun Huang, Ji-Yan Wu, Wei Tsang Ooi |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous DrivingabstractEfficient data utilization is crucial for advancing 3D scene understanding in autonomous driving, where reliance on heavily human-annotated LiDAR point clouds challenges fully supervised methods. Addressing this, our study extends into semi-supervised learning for LiDAR semantic segmentation, leveraging the intrinsic spatial priors of driving scenes and multi-sensor complements to augment the efficacy of unlabeled datasets. We introduce LaserMix++, an evolved framework that integrates laser beam manipulations from disparate LiDAR scans and incorporates LiDAR-camera correspondences to further assist data-efficient learning. Our framework is tailored to enhance 3D scene consistency regularization by incorporating multi-modality, including 1) multi-modal LaserMix operation for fine-grained cross-sensor interactions; 2) camera-to-LiDAR feature distillation that enhances LiDAR feature learning; and 3) language-driven knowledge guidance generating auxiliary supervisions using open-vocabulary models. The versatility of LaserMix++ enables applications across LiDAR representations, establishing it as a universally applicable solution. Our framework is rigorously validated through theoretical analysis and extensive experiments on popular driving perception datasets. Results demonstrate that LaserMix++ markedly outperforms fully supervised alternatives, achieving comparable accuracy with five times fewer annotations and significantly improving the supervised-only baselines. This substantial advancement underscores the potential of semi-supervised approaches in reducing the reliance on extensive labeled data in LiDAR-based 3D scene understanding systems. Lingdong Kong, Xiang Xu 0009, Jiawei Ren 0001, Liang Pan, Kai Chen 0026, Wei Tsang Ooi, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Composing Error Concealment Pipelines for Dynamic 3D Point Cloud StreamingabstractDynamic 3D point clouds enable the immersive user experience and thus have become increasingly more popular in volumetric video streaming applications. When being streamed over best-effort networks, point cloud frames may suffer from lost or late packets, leading to non-trivial quality degradation. To solve this problem, we proposed the very first error concealment pipeline framework, which comprises five stages: pre-processing, matching, motion estimation, prediction, and post-processing. Alternative algorithms can be developed for each stage, while algorithms of different stages could be mixed and matched into pipelines for end-to-end performance evaluations. We discussed the design goal and proposed multiple algorithms for each stage. These algorithms were then quantitatively compared using dynamic 3D point cloud sequences with diverse characteristics. Based on the comparison results, we proposed four representative pipelines for: (i) diverse degrees of motion variance, i.e., minor versus significant, and (ii) different application requirements, i.e., high quality versus low overhead. Extensive end-to-end evaluations of our proposed pipelines demonstrated their superior concealed quality over the 3D frame-copy method in both: (i) 3D metrics, by up to 5.32 dB in GPSNR and 1.7 dB in CPSNR,and (ii) 2D metrics, by up to 2.22 dB in PSNR, 0.06 in SSIM, and 11.67 in VMAF. Adding to that, a user study with 15 subjects indicated that our best-performing pipeline achieved 100% preference winning rate over the state-of-the-art learning-based interpolation algorithms while consuming merely up to 8.55% of running time. I-Chun Huang, Yuang Shi, Yuan-Chun Sun, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | GlassMail: Towards Personalised Wearable Assistant for On-the-Go Email Creation on Smart GlassesabstractOptical See-through Head-Mounted Displays (OHMDs) offer new opportunities for completing complex information processing tasks on the go. We introduce GlassMail, a Large Language Models (LLMs)-based wearable assistant on OHMDs for mobile email creation. Our formative study identified two challenges of the LLM-based wearable email assistant: (i) achieving efficient and accurate understanding of user intentions, and (ii) ensuring effective information presentation for email processes. Through two empirical studies, we developed a "Single Turn with Optional Clarification " approach for accurate user intention recognition and a "Fade Context with Optional Audio " mode for effective email processing. An observation study then evaluated GlassMail ’s feasibility in composing formal and semi-formal emails, supporting the usefulness and effectiveness of GlassMail in simple scenarios and yielding insights into potential future improvements for complex scenarios. We further discuss the design implications for the future development of wearable AI-enabled assistants. Ashwin Ram 0002, Can Liu 0003, Yun Huang 0003, Wei Tsang Ooi, Shengdong Zhao 0001 |
Conference on Designing Interactive Systems | 8 |
| 2024 | Drones for all: Creating an Authentic Programming Experience for Students with Visual ImpairmentsabstractProgramming has become a highly sought-after skill in STEM-related studies and careers, but it has only reached a fraction of students with visual impairments. Therefore, there is a need to explore new methods for teaching and learning. This study aims to understand the potential of using drones to create an authentic learning environment to help students with visual impairments learn programming. Based on a month-long engagement with five students with visual impairments, we present insights on using drones to support programming education for students with visual impairments. Yize Wei, Maëlle Dubucq, Malsha de Zoysa, Christophe Jouffrais, Suranga Nanayakkara, Wei Tsang Ooi |
ASSETS | 6 |
| 2024 | OpenESS: Event-Based Semantic Scene Understanding with Open VocabulariesabstractEvent-based semantic segmentation (ESS) is a fundamental yet challenging task for event camera sensing. The difficulties in interpreting and annotating event data limit its scalability. While domain adaptation from images to event data can help to mitigate this issue, there exist data representational differences that require additional effort to resolve. In this work, for the first time, we synergize information from image, text, and event-data domains and introduce OpenESS to enable scalable ESS in an open-world, annotation-efficient manner. We achieve this goal by transferring the semantically rich CLIP knowledge from image-text pairs to event streams. To pursue better cross-modality adaptation, we propose a frame-to-event contrastive distillation and a text-to-event semantic consistency regularization. Experimental results on popular ESS benchmarks showed our approach outperforms existing methods. Notably, we achieve 53.93% and 43.31% mIoU on DDD17 and DSEC-Semantic without using either event or frame labels. Lingdong Kong, Youquan Liu, Lai Xing Ng, Benoit Cottereau, Wei Tsang Ooi |
CVPR | 5 |
| 2024 | SlideCraft: Synthetic Slides Generation for Robust Slide Analysis
Travis Seng, Axel Carlier, Thomas Forgione, Vincent Charvillat, Wei Tsang Ooi |
ICDAR (1) | 5 |
| 2024 | QV4: QoE-based Viewpoint-Aware V-PCC-encoded Volumetric Video StreamingabstractVolumetric videos allow six degrees of freedom (6DoF) movement for viewers, enabling numerous applications in domains such as entertainment, healthcare, and education. MPEG's Video-based Point Cloud Compression (V-PCC) is a recent new standard for volumetric video compression that achieves a considerable compression rate while maintaining the quality of the point cloud sequence. However, V-PCC is hard to fit into existing tiling-based volumetric video streaming framework due to the lack of proper user viewing adaptive techniques. In this paper, we propose QV4, a Quality-of-Experience (QoE) based streaming pipeline for viewpoint-aware V-PCC-encoded volumetric video. Specifically, we leverage the intermediate results produced by the V-PCC encoder to achieve effective and efficient viewpoint-aware tiling for V-PCC. We then build a QoE model and a 6DoF movement model based on real-world user data, to predict the users' viewing experience and behaviors, respectively. The proposed QoE model and 6DoF movement model are combined with viewpoint-aware V-PCC tiling to maximize the visual quality of volumetric videos. Extensive simulations show that by enabling viewpoint-aware adaptation and optimization for V-PCC-encoded volumetric videos, QV4 can achieve up to 14.67% improvement in structural similarity index (SSIM) and 7.39% improvement in video multi-method assessment fusion (VMAF) over highly dynamic viewing behaviors in a network with limited and fluctuating bandwidth. Yuang Shi, Bennett Clement, Wei Tsang Ooi |
MMSys | 3 |
| 2024 | A highly robust deep learning technique for overlap detection using audio fingerprinting
Akash Uikey, Anterpreet Kaur Bedi, Priyankar Choudhary, Wei Tsang Ooi, Mukesh Saini |
Multim. Tools Appl. | 4 |
| 2023 | Prompting a Large Language Model to Generate Diverse Motivational Messages: A Comparison with Human-Written MessagesabstractLarge language models (LLMs) are increasingly capable and prevalent, and can be used to produce creative content. The quality of content is influenced by the prompt used, with more specific prompts that incorporate examples generally producing better results. On from this, it could be seen that using instructions written for crowdsourcing tasks (that are specific and include examples to guide workers) could prove effective LLM prompts. To explore this, we used a previous crowdsourcing pipeline that gave examples to people to help them generate a collectively diverse corpus of motivational messages. We then used this same pipeline to generate messages using GPT-4, and compared the collective diversity of messages from: (1) crowd-writers, (2) GPT-4 using the pipeline, and (3 & 4) two baseline GPT-4 prompts. We found that the LLM prompts using the crowdsourcing pipeline caused GPT-4 to produce more diverse messages than the two baseline prompts. We also discuss implications from messages generated by both human writers and LLMs. Samuel Rhys Cox, Ashraf M. Abdul, Wei Tsang Ooi |
HAI | 3 |
| 2023 | "The Use of Deception in Dementia-Care Robots: Should Robots Tell \"White Lies\" to Limit Emotional Distress?"abstractWith projections of ageing populations and increasing rates of dementia, there is need for professional caregivers. Assistive robots have been proposed as a solution to this, as they can assist people both physically and socially. However, caregivers often need to use acts of deception (such as misdirection or white lies) in order to ensure necessary care is provided while limiting negative impacts on the cared-for such as emotional distress or loss of dignity. We discuss such use of deception, and contextualise their use within robotics. Samuel Rhys Cox, Grace Cheong, Wei Tsang Ooi |
HAI | 3 |
| 2023 | Comparing How a Chatbot References User Utterances from Previous Chatting Sessions: An Investigation of Users' Privacy Concerns and PerceptionsabstractChatbots are capable of remembering and referencing previous conversations, but does this enhance user engagement or infringe on privacy? To explore this trade-off, we investigated the format of how a chatbot references previous conversations with a user and its effects on a user’s perceptions and privacy concerns. In a three-week longitudinal between-subjects study, 169 participants talked about their dental flossing habits to a chatbot that either, (1-None): did not explicitly reference previous user utterances, (2-Verbatim): referenced previous utterances verbatim, or (3-Paraphrase): used paraphrases to reference previous utterances. Participants perceived Verbatim and Paraphrase chatbots as more intelligent and engaging. However, the Verbatim chatbot also raised privacy concerns with participants. To gain insights as to why people prefer certain conditions or had privacy concerns, we conducted semi-structured interviews with 15 participants. We discuss implications from our findings that can help designers choose an appropriate format to reference previous user utterances and inform in the design of longitudinal dialogue scripting. Samuel Rhys Cox, Yi-Chieh Lee, Wei Tsang Ooi |
HAI | 3 |
| 2023 | AdaptReview: Towards Effective Video Review Using Text Summaries and Concept Maps
Shan Zhang 0006, Yang Chen 0054, Nuwan Janaka, Chloe Dolma Si Ying Haigh, Shengdong Zhao 0001, Wei Tsang Ooi |
INTERACT (2) | 6 |
| 2023 | VOLVQAD: An MPEG V-PCC Volumetric Video Quality Assessment DatasetabstractWe present VOLVQAD, a volumetric video quality assessment dataset consisting 7,680 ratings on 376 video sequences from 120 participants. The volumetric video sequences are first encoded with MPEG V-PCC using 4 different avatar models and 16 quality variations, and then rendered into test videos for quality assessment using 2 different background colors and 16 different quality switching patterns. The dataset is useful for researchers who wish to understand the impact of volumetric video compression on subjective quality. Analysis of the collected data are also presented in this paper. Samuel Rhys Cox, May Lim, Wei Tsang Ooi |
MMSys | 3 |
| 2023 | Enabling Low Bit-Rate MPEG V-PCC-encoded Volumetric Video Streaming with 3D Sub-samplingabstractMPEG's Video-based Point Cloud Compression (V-PCC) is a recent new standard for volumetric video compression. By mapping a 3D dynamic point cloud to a 2D image sequence, V-PCC can rely on state-of-the-art video codecs to achieve high compression rate while maintaining the visual fidelity of the point cloud sequence. The quality of a compressed point cloud degrades steeply, however, below the operational bit-rate range of the video codec. In this work, we show that redundant information inherent in a 3D point cloud can be exploited to further extend the bit-rate range of the V-PCC codec, enabling it to operate in a low bit-rate scenario that is important in the context of volumetric video streaming. By simplifying the 3D point clouds through down-sampling and down-scaling during the encoding phase, and reversing the process during the decoding phase, we show that V-PCC could achieve up to 2.1 dB improvement in peak signal-to-noise ratio (PSNR), 7.1% improvement in structural similarity index (SSIM) and 14.8 improvement in video multimethod assessment fusion (VMAF) of the rendered point clouds at the same bit-rate and correspondingly up to 48.5% lower bit-rate at the same image quality. Yuang Shi, Pranav Venkatram, Wei Tsang Ooi |
MMSys | 4 |
| 2023 | A Dynamic 3D Point Cloud Dataset for Immersive ApplicationsabstractMotion estimation in a 3D point cloud sequence is a fundamental operation with many applications, including compression, error concealment, and temporal upscaling. While there have been multiple research contributions toward estimating the motion vector of points between frames, there is a lack of a dynamic 3D point cloud dataset with motion ground truth to benchmark against. In this paper, we present an open dynamic 3D point cloud dataset to fill this gap. Our dataset consists of synthetically generated objects with pre-determined motion patterns, allowing us to generate the motion vectors for the points. Our dataset contains nine objects in three categories (shape, avatar, and textile) with different animation patterns. We also provide semantic segmentation of each avatar object in the dataset. Our dataset can be used by researchers who need temporal information across frames. As an example, we present an evaluation of two motion estimation methods using our dataset. Yuan-Chun Sun, I-Chun Huang, Yuang Shi, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu |
MMSys | 4 |
| 2023 | RoboDepth: Robust Out-of-Distribution Depth Estimation under CorruptionsabstractDepth estimation from monocular images is pivotal for real-world visual perception systems. While current learning-based depth estimation models train and test on meticulously curated data, they often overlook out-of-distribution (OoD) situations. Yet, in practical settings -- especially safety-critical ones like autonomous driving -- common corruptions can arise. Addressing this oversight, we introduce a comprehensive robustness test suite, RoboDepth, encompassing 18 corruptions spanning three categories: i) weather and lighting conditions; ii) sensor failures and movement; and iii) data processing anomalies. We subsequently benchmark 42 depth estimation models across indoor and outdoor scenes to assess their resilience to these corruptions. Our findings underscore that, in the absence of a dedicated robustness evaluation framework, many leading depth estimation models may be susceptible to typical corruptions. We delve into design considerations for crafting more robust depth estimation models, touching upon pre-training, augmentation, modality, model capacity, and learning paradigms. We anticipate our benchmark will establish a foundational platform for advancing robust OoD depth estimation. Lingdong Kong, Shaoyuan Xie, Hanjiang Hu, Lai Xing Ng, Benoit Cottereau, Wei Tsang Ooi |
NeurIPS | 6 |
| 2023 | Quantitative Comparison of Point Cloud Compression Algorithms With PCC ArenaabstractWith the growth of Extended Reality (XR) and capturing devices, point cloud representation has become attractive to academics and industry. Point Cloud Compression (PCC) algorithms further promote numerous XR applications that may change our daily life. However, in the literature, PCC algorithms are often evaluated with heterogeneous datasets, metrics, and parameters, making the results hard to interpret. In this article, we propose an open-source benchmark platform called PCC Arena. Our platform is modularized in three aspects: PCC algorithms, point cloud datasets, and performance metrics. Users can easily extend PCC Arena in each aspect to fulfill the requirements of their experiments. To show the effectiveness of PCC Arena, we integrate seven PCC algorithms into PCC Arena along with six point cloud datasets. We then compare the algorithms on ten carefully selected metrics to evaluate the quality of the output point clouds. We further conduct a user study to quantify the user-perceived quality of rendered images that are produced by different PCC algorithms. Several novel insights are revealed in our comparison: (i) Signal Processing (SP)-based PCC algorithms are stable for different usage scenarios, but the trade-offs between coding efficiency and quality should be carefully addressed, (ii) Neural Network (NN)-based PCC algorithms have the potential to consume lower bitrates yet provide similar results to SP-based algorithms, (iii) NN-based PCC algorithms may generate artifacts and suffer from long running time, and (iv) NN-based PCC algorithms are worth more in-depth studies as the recently proposed NN-based PCC algorithms improve the quality and running time. We believe that PCC Arena can play an essential role in allowing engineers and researchers to better interpret and compare the performance of future PCC algorithms. Cheng-Hao Wu, Chih-Fan Hsu, Tzu-Kuan Hung, Carsten Griwodz, Wei Tsang Ooi, Cheng-Hsin Hsu |
IEEE Trans. Multim. | 5 |
| 2022 | Error Concealment of Dynamic 3D Point Cloud StreamingabstractRecently standardized MPEG Video-based Point Cloud Compression (V-PCC) codec has shown promise in achieving a good rate-distortion ratio of dynamic 3D point cloud compression. Current error concealment methods of V-PCC, however, lead to significantly distorted 3D point cloud frames under imperfect network conditions. To address this problem, we propose a general framework for concealing distorted and lost 3D point cloud frames due to packet loss. We also design, implement, and evaluate a suite of tools for each stage of our framework, which can be combined into multiple variants of error concealment algorithms. We conduct extensive experiments using seven dynamic 3D point cloud sequences with diverse characteristics to understand the strengths and limitations of our proposed error concealment algorithms. Our experiment results show that our algorithms outperform: (i) the method employed by V-PCC by at least 3.58 dB in Geometry Peak Signal-to-Noise Ratio (GPSNR) and 10.68 in Video Multi-Method Assessment Fusion (VMAF) and (ii) point cloud frame copy method by at most 5.8 dB in (3D) GPSNR and 12.0 in (2D) VMAF. Further, the proposed error concealment framework and algorithms work in the 3D domain, and thus are agnostic to the codecs and are applicable to future point cloud compression standards Tzu-Kuan Hung, I-Chun Huang, Samuel Rhys Cox, Wei Tsang Ooi, Cheng-Hsin Hsu |
ACM Multimedia | 4 |
| 2022 | Bandwidth-Efficient Multi-video Prefetching for Short Video StreamingabstractApplications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem. Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001 |
ACM Multimedia | 4 |
| 2022 | MultiLive: Adaptive Bitrate Control for Low-Delay Multi-Party Interactive Live StreamingabstractIn multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates tosend; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we designMultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We also deploy the algorithm on a commercial live streaming platform that provides such services for more than 2300 users. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-$5\times $improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing. Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015 |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | Multi-Camera Video Scene Graphs for Surveillance Videos Indexing and RetrievalabstractModern video surveillance systems often consist of multiple cameras capturing continuous videos. However, existing systems do not fully take advantage of shared scene semantics across multiple video streams, and hence querying the captured videos to find an event or object of interest can be a time-consuming task. In this paper, we propose a compact representation of objects and their relationships in the scene under surveillance, over time, and across multiple cameras, using a combined global spatio-temporal scene graph. The same objects that appear in multiple cameras are stored only once, reducing the storage required and speeding up the query when compared to storing the scene graph from each camera individually. Our experiments show that in a 5-camera system our proposed representation can speed up querying time by a factor of 3.9 times. Toshal Patel, Alvin Yan Hong Yao, Yu Qiang, Wei Tsang Ooi, Roger Zimmermann |
ICIP | 4 |
| 2021 | The ACM Multimedia 2021 Meet Deadline Requirements Grand ChallengeabstractDelay-sensitive multimedia streaming applications require their data to be delivered before a deadline to be useful. The data transmitted by these applications can usually be partitioned into blocks with different priorities, assigned based on the impact of a block on the Quality of Experience (QoE) if it misses its delivery deadline. Meet their deadline requirements is challenging due to the dynamics of the network and these applications' high demand on network resources. To encourage the research community to address this challenge, we organize the "Meet Deadline Requirements" Grand Challenge at ACM Multimedia 2021. This grand challenge provides a simulation platform onto which the participants can implement their block scheduler and bandwidth estimator and then benchmark against each other using a common set of application traces and network traces. Junjie Deng, Mowei Wang, Yong Cui 0001, Wei Tsang Ooi, Jiangchuan Liu, Xinyu Zhang 0003, Kai Zheng 0003, Yi Li 0015 |
ACM Multimedia | 5 |
| 2021 | Playing chunk-transferred DASH segments at low latency with QLiveabstractMore users have a growing interest in low latency over-the-top (OTT) applications such as online video gaming, video chat, online casino, sports betting, and live auctions. OTT applications face challenges in delivering low latency live streams using Dynamic Adaptive Streaming over HTTP (DASH) due to large playback buffer and video segment duration. A potential solution to this issue is the use of HTTP chunked transfer encoding (CTE) with the common media application format (CMAF). This combination allows the delivery of each segment in several chunks to the client, starting before the segment is fully available in real-time. However, CTE and CMAF alone are not sufficient as they do not address other limitations and challenges at the client-side, including inaccurate bandwidth measurement, latency control, and bitrate selection. Praveen Kumar Yadav, Abdelhak Bentaleb, May Lim, Junyi Huang, Wei Tsang Ooi, Roger Zimmermann |
MMSys | 5 |
| 2021 | Dynamic 3D point cloud streaming: distortion and concealmentabstractWe present a study on the impact of packet loss on dynamic 3D point cloud streaming, encoded with MPEG Video-based Point Cloud Compression (V-PCC) standard. We show the distortion when different channels of V-PCC bitstream are lost, with the loss of occupancy and geometry data impacting the quality most significantly. Our results point to the need for better error concealment techniques. We end the paper by presenting preliminary thoughts and experimental results of two naive error concealment techniques in the point cloud domain, for attributes and geometry data, respectively, and highlight the limitations of each. Cheng-Hao Wu, Xiner Li, Rahul Rajesh, Wei Tsang Ooi, Cheng-Hsin Hsu |
NOSSDAV | 4 |
| 2020 | MultiLive: Adaptive Bitrate Control for Low-delay Multi-party Interactive Live StreamingabstractIn multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates to send; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we design MultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We conduct extensive trace-driven simulations to evaluate the algorithm. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-5× improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing. Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015 |
INFOCOM | 5 |
| 2020 | Tile Rate Allocation for 360-Degree Tiled Adaptive Video Streamingabstract360-degree video streaming commonly encodes and transmits the video as independently-decodable tiles to conserve bandwidth of regions out of the viewer's field of view (FoV). The bitrate of the tiles, however, can vary significantly across the tiles, complicating the choice of the representation to download for each tile in each segment to adapt to the bandwidth dynamics. In this paper, we model the tile rate allocation problem as a multiclass knapsack problem with a dynamic profit function that is a function of the FoV and the buffer occupancy. Experiments show that our approach can reduce bandwidth wastage by up to 41%, the number of stalls by up to 31%, stall durations by up to 26.5%, switches in quality by up to 20%, without sacrificing the quality of the tiles within the FoV, even when there are significant head movement and changes in FoV during streaming. Praveen Kumar Yadav, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2020 | CloudyGame: Enabling cloud gaming on the edge with dynamic asset streaming and shared game instances
Anand Bhojan, Siang Ping Ng, Joel Ng, Wei Tsang Ooi |
Multim. Tools Appl. | 4 |
| 2020 | DQ-DASH: A Queuing Theory Approach to Distributed Adaptive Video StreamingabstractThe significant popularity of HTTP adaptive video streaming (HAS), such as Dynamic Adaptive Streaming over HTTP (DASH), over the Internet has led to a stark increase in user expectations in terms of video quality and delivery robustness. This situation creates new challenges for content providers who must satisfy the Quality-of-Experience (QoE) requirements and demands of their customers over a best-effort network infrastructure. Unlike traditional single server DASH, we developed a D istributed Q ueuing theory bitrate adaptation algorithm for DASH (DQ-DASH) that leverages the availability of multiple servers by downloading segments in parallel. DQ-DASH uses a M x /D/1/K queuing theory based bitrate selection in conjunction with the request scheduler to download subsequent segments of the same quality through parallel requests to reduce quality fluctuations. DQ-DASH facilitates the aggregation of bandwidth from different servers and increases fault-tolerance and robustness through path diversity. The resulting resilience prevents clients from suffering QoE degradations when some of the servers become congested. DQ-DASH also helps to fully utilize the aggregate bandwidth from the servers and download the imminently required segment from the server with the highest throughput. We have also analyzed the effect of buffer capacity and segment duration for multi-source video streaming. Abdelhak Bentaleb, Praveen Kumar Yadav, Wei Tsang Ooi, Roger Zimmermann |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Using 3D Bookmarks for Desktop and Mobile DASH-3D ClientsabstractNavigating in a 3D networked virtual environment with a six-degree of freedom on a mobile device can be disorientating and challenging to users. In this technical demonstration, we show how 3D bookmarks can be used to simplify such interactions. Our system integrates the 3D bookmarks into a DASH-based network virtual environment in a DASH-compliant manner and is available on both the desktop and mobile. Thomas Forgione, Axel Carlier, Géraldine Morin, Wei Tsang Ooi, Vincent Charvillat |
ACM Multimedia | 4 |
| 2019 | The ACM Multimedia 2019 Live Video Streaming Grand ChallengeabstractLive video streaming delivery over Dynamic Adaptive Video Streaming (DASH) is challenging as it requires low end-to-end latency, is more prone to stall, and the receiver has to decide online which representation at which bitrate to download and whether to adjust the playback speed to control the latency. To encourage the research community to come together to address this challenge, we organize the Live Video Streaming Grand Challenge at ACM Multimedia 2019. This grand challenge provides a simulation platform onto which the participants can implement their adaptive bitrate (ABR) logic and latency control algorithm, and then benchmark against each other using a common set of video traces and network traces. The ABR algorithms are evaluated using a common Quality-of- Experience (QoE) model that accounts for playback bitrate, latency constraint, frame-skipping penalty, and rebuffering penalty. Gang Yi, Abdelhak Bentaleb, Yi Li 0015, Kai Zheng 0003, Jiangchuan Liu, Wei Tsang Ooi, Yong Cui 0001 |
ACM Multimedia | 8 |
| 2018 | An Implementation of a DASH Client for Browsing Networked Virtual EnvironmentabstractWe demonstrate the use of DASH, a widely-deployed standard for streaming video content, for streaming 3D content in an NVE (Networked Virtual Environment) consisting of 3D geometry and associated textures. We have developed a DASH client for NVE to show how NVE benefits from the advantages of DASH: it offers a scalable, easy-to-deploy 3D streaming framework. In our system, the 3D content is first statically partitioned into compliant DASH data, and metadata is provided in order for the client to manage which data to download. Based on a proposed utility metric for geometry and texture at the different resolution, the client can choose the content to request depending on its viewpoint. We effectively provide a Web-based client to navigate through our sample 3D scene, while deriving the streaming requests from its computation of the necessary online parameters, in a receiver-driven manner. Thomas Forgione, Axel Carlier, Géraldine Morin, Wei Tsang Ooi, Vincent Charvillat, Praveen Kumar Yadav |
ACM Multimedia | 4 |
| 2018 | DASH for 3D Networked Virtual EnvironmentabstractDASH is now a widely deployed standard for streaming video content due to its simplicity, scalability, and ease of deployment. In this paper, we explore the use of DASH for a different type of media content -- networked virtual environment (NVE), with different properties and requirements. We organize a polygon soup with textures into a structure that is compatible with DASH MPD (Media Presentation Description), with a minimal set of view-independent metadata for the client to make intelligent decisions about what data to download at which resolution. We also present a DASH-based NVE client that uses a view-dependent and network dependent utility metric to decide what to download, based only on the information in the MPD file. We show that DASH can be used on NVE for 3D content streaming. Our work opens up the possibility of using DASH for highly interactive applications, beyond its current use in video streaming. Thomas Forgione, Axel Carlier, Géraldine Morin, Wei Tsang Ooi, Vincent Charvillat, Praveen Kumar Yadav |
ACM Multimedia | 4 |
| 2018 | Cloud Baking: Collaborative Scene Illumination for Dynamic Web3D ScenesabstractWe propose Cloud Baking, a collaborative rendering architecture for dynamic Web3D scenes. In our architecture, the cloud renderer renders the scene with the global illumination (GI) information in a GI map; the web-based client renderer renders the scene with ambient lighting only and blends it with the GI map received from the cloud for the final scene. This approach allows the users to interact with the web scene and change the scene dynamically through the web interface end, yet move the computationally heavy tasks of global illumination computation to the cloud. A challenge we face is the interaction delay that causes the frames rendered on the cloud and the client to go out of sync. We propose to use 3D warping and a hole-filling algorithm designed for GI map to predict the late GI map. We show both quantitatively and visually the quality of the GI map produced using our method. Our prediction algorithm allows us to further reduce the frequency at which the GI map is computed and sent from the server, reducing both computational needs and bandwidth usage. Chang Liu 0037, Wei Tsang Ooi, Jinyuan Jia 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | AltMM 2017 - 2nd International Workshop on Multimedia Alternate RealitiesabstractAltMM 2017, the 2nd International Workshop on Multimedia Alternate Realities at ACM Multimedia aims to provide a forum for researchers and practitioners concerned with multimedia that enables experiencing "alternate realities". Such experiences may allow us to access other worlds, to live other people's stories, to communicate with or experience alternate realities. Different spaces, times or situations can be entered thanks to multimedia contents and systems, which coexist with our current reality, and are sometimes so vivid and engaging that we feel we are living in them. Advances in multimedia are making it possible to create immersive experiences that may involve the user in a different or augmented world, as an alternate reality. Teresa Chambel, Rene Kaiser, Omar Niamut, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2017 | QUETRA: A Queuing Theory Approach to DASH Rate AdaptationabstractDASH, or Dynamic Adaptive Streaming over HTTP, relies on a rate adaptation component to decide on which representation to download for each video segment. A plethora of rate adaptation algorithms has been proposed in recent years. The decisions of which bitrate to download made by these algorithms largely depend on several factors: estimated network throughput, buffer occupancy, and buffer capacity. Yet, these algorithms are not informed by a fundamental relationship between these factors and the chosen bitrate, and as a result, we found that they do not perform consistently in all scenarios, and require parameter tuning to work well under different buffer capacity. In this paper, we model a DASH client as an M/D/1/K queue, which allows us to calculate the expected buffer occupancy given a bitrate choice, network throughput, and buffer capacity. Using this model, we propose QUETRA, a simple rate adaptation algorithm. We evaluated QUETRA under a diverse set of scenarios and found that, despite its simplicity, it leads to better quality of experience (7% - 140%) than existing algorithms. Praveen Kumar Yadav, Arash Shafiei 0001, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2017 | Guest Editorial: Interactive Media: Technology and Experience
Britta Meixner, Rene Kaiser, Joscha Jäger, Wei Tsang Ooi, Harald Kosch |
Multim. Tools Appl. | 4 |
| 2017 | Sprite tree: an efficient image-based representation for networked virtual environments
Minhui Zhu, Géraldine Morin, Vincent Charvillat, Wei Tsang Ooi |
Vis. Comput. | 4 |
| 2016 | JurCast: Joint user and rate allocation for video multicast over multiple APsabstractWireless multicast has been exploited to bridge the gap between the limited wireless bandwidth and the rapidly increasing mobile video traffic demand. Multicast of videos to a set of heterogeneous users over multiple wireless access points, however, is challenging because of the trade-offs between high transmission rate, load balancing, and multicast opportunities. In this paper, we present JurCast, a joint user and rate allocation scheme for video multicast over multiple APs. Our approach balances the trade-off between these factors by determining user to Access Points (APs) association, the video resolution version (quality) to be delivered for each session, and the transmission link rate for each video version. The aim of our solution is to maximize the overall received video quality over all users. We have implemented and evaluated our solution on a WiFi testbed as well as the simulation of a large scale deployment. The results indicate that our method considerably outperforms the baseline schemes and achieves up to 3dB and 55% improvements in terms of peak signal-to-noise ratio (PSNR) and goodput, respectively. Wei Tsang Ooi, Mun Choon Chan |
INFOCOM | 2 |
| 2016 | AltMM 2016: 1st International Workshop on Multimedia Alternate RealitiesabstractMultimedia experiences allow us to access other worlds, to live other people's stories, to communicate with or experience alternate realities. Different spaces, times or situations can be entered thanks to multimedia contents and systems, which coexist with our current reality, and are sometimes so vivid and engaging that we feel we are living in them. Advances in multimedia are making it possible to create immersive experiences that may involve the user in a different or augmented world, as an alternate reality. AltMM 2016, the 1st International Workshop on Multimedia Alternate Realities at ACM Multimedia, aims at exploring how the synergy between multimedia technologies and effects can foster the creation of alternate realities and make their access an enriching, valuable and real experience. The workshop program will contain a combination of oral and invited keynote presentations, and poster, demo and discussion sessions, altogether enabling interactive scientific sharing and discussion between practitioners and researchers. Teresa Chambel, Rene Kaiser, Omar Niamut, Wei Tsang Ooi, Judith Redi |
ACM Multimedia | 4 |
| 2016 | SuperStreamer: Enabling Progressive Content Streaming in a Game EngineabstractThis technical demonstration presents the SuperStreamer project, which enables progressive game assets streaming to players while games are played, reducing the startup time required to download and start playing a cloud-based game. SuperStreamer modifies a popular game engine, Unreal Engine 4, to support developing and playing games with progressive game asset streaming. With SuperStreamer, developers can mark the minimal set of files, containing only the game content essential to start playing the game. When a player plays the game, these minimal set of files will be downloaded to the player's device. SuperStreamer also generates low resolution textures automatically when a developer publishes a game, and these low resolution textures are transmitted first into the game client. As players move through a game level, high quality textures required for the game will be downloaded. In our demo game, we are able to decrease the time taken to startup and load the first game level by around 30%. Yong Xue Eu, Jermyn Tanu, Justin Jieting Law, Muhammad Hanif B. Ghazali, Shuan Siang Tay, Wei Tsang Ooi, Anand Bhojan |
ACM Multimedia | 6 |
| 2016 | Impact of 3D bookmarks on navigation and streaming in a networked virtual environmentabstractA 3D bookmark in a networked virtual environment (NVE) provides a navigation aid, allowing the user to move quickly from its current viewpoint to a bookmarked viewpoint by simply clicking on the bookmark. In this paper, we first validate the positive impact that 3D bookmarks have in easing navigation in a 3D scene. Then, we show that, in the context of a NVE that streams content on demand from server to client, navigating with bookmarks leads to lower rendering quality at the bookmarked viewpoint, due to lower locality of data. We then investigate into how prefetching the 3D data at the bookmarks and precomputation of visible faces at the bookmarks help to improve the rendering quality. Thomas Forgione, Axel Carlier, Géraldine Morin, Wei Tsang Ooi, Vincent Charvillat |
MMSys | 4 |
| 2016 | Throughput Estimation for Short Lived TCP Cubic FlowsabstractMobile devices are increasingly becoming the dominant device for Internet access. The network throughput achieved by a mobile device directly affects the performance and user experience. Throughput measurement techniques thus play an important role in predicting expected performance. Measurement techniques that require the transfer of large amounts of data can provide higher accuracy but incur large overhead. Further, since most mobile cellular plans impose usage quota, the overhead of such measurements over cellular networks can become quite high. Smaller data transfers have also been used to measure the throughput. Due to the conservative TCP slow start behaviour, however, these measurements often underestimate the achievable throughput. Considering these weaknesses in existing throughput measurement techniques, we propose a throughput estimation technique for TCP Cubic that uses 1 MB of data transfer to predict the throughput for prevalent large transfer sizes in mobile traffic such as 5 MB, 10 MB and 20 MB. Our evaluation shows that our approach can achieve high accuracy with low overhead, in predicting the achievable throughput. Girisha De Silva, Mun Choon Chan, Wei Tsang Ooi |
MobiQuitous | 3 |
| 2016 | Multimedia prefetching with optimal Markovian policies
Cezar Plesca, Vincent Charvillat, Wei Tsang Ooi |
J. Netw. Comput. Appl. | 3 |
| 2016 | Secret sharing approach for securing cloud-based pre-classification volume ray-casting
Manoranjan Mohanty, Wei Tsang Ooi, Pradeep K. Atrey |
Multim. Tools Appl. | 2 |
| 2016 | Neurite Tracing With Object ProcessabstractIn this paper we present a pipeline for automatic analysis of neuronal morphology: from detection, modeling to digital reconstruction. First, we present an automatic, unsupervised object detection framework using stochastic marked point process. It extracts connected neuronal networks by fitting special configuration of marked objects to the centreline of the neurite branches in the image volume giving us position, local width and orientation information. Semantic modeling of neuronal morphology in terms of critical nodes like bifurcations and terminals, generates various geometric and morphology descriptors such as branching index, branching angles, total neurite length, internodal lengths for statistical inference on characteristic neuronal features. From the detected branches we reconstruct neuronal tree morphology using robust and efficient numerical fast marching methods. We capture a mathematical model abstracting out the relevant position, shape and connectivity information about neuronal branches from the microscopy data into connected minimum spanning trees. Such digital reconstruction is represented in standard SWC format, prevalent for archiving, sharing, and further analysis in the neuroimaging community. Our proposed pipeline outperforms state of the art methods in tracing accuracy and minimizes the subjective variability in reconstruction, inherent to semi-automatic methods. Sreetama Basu, Wei Tsang Ooi, Daniel Racoceanu |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Modeling 3D synthetic view dissimilarity
Shanghong Zhao 0002, Wei Tsang Ooi |
Vis. Comput. | 2 |
| 2015 | A Video Timeline with Bookmarks and Prefetch State for Faster Video BrowsingabstractReducing seek latency by predicting what the users will access is important for user experience, particularly during video browsing, where users seek frequently to skim through a video. Much existing research strived to predict user access pattern more accurately to improve the prefetching hit rate. This paper proposed a different approach whereby the prefetch hit rate is improved by biasing the users to seek to prefetched content with higher probability, through changing the video player user interface. Through a user study, we demonstrated that our player interface can lead to up to 4$\times$ more seeks to bookmarked segments and reduce seek latency by 40\%, compared to a video player interface commonly used today. The user study also showed that the user experience and the understanding of the video content when browsing is not compromised by the changes in seek behavior. Axel Carlier, Vincent Charvillat, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2015 | Mitigating Unfairness Due to Physical Layer Capture in Practical 802.11 Mesh NetworksabstractIn this paper, we describeFairMesh, which is the first attempt at mitigating the unfairnessarising from physical layer capture (PLC)in 802.11 mesh networks. In the presence of PLC, which is surprisingly common in practical mesh networks, existing state-of-art solutions either fail to correctly identify the sender that needs to be throttled or are too aggressive in reducing the sending rate. FairMesh is able to accurately detect unfairness quickly and employs a simple$CW_{min}$adjustment algorithm to achieve approximate max-min fairness. Our key insight is that the nodes that cause an unfair situation to arise and can act to remedy it are often distinct from the ones that can accurately assess the degree of unfairness. To the best of our knowledge, we are the first to decouple the detection and assessment of unfairness from the remedial action. A key strength of our approach is itssimplicity, which makes it amenable for deployment in practical 802.11 mesh networks to allow an arbitrary number of flows to operate concurrently without modifications to the 802.11 MAC. We show via simulation and with experiments on a 20-node outdoor 802.11 wireless mesh testbed that FairMesh has many desirable properties. First, it is fully distributed and has negligible control overhead. Second, it achieves approximate max-min fairness, and can be modified to support a different notion of fairness (e.g., proportional fairness). Third, it can handle multiple (more than two) competing links and can scale up to mesh networks with tens of nodes. Fourth, it remains efficient under high data rates and high loss rates. Finally, FairMesh interacts well with TCP and maintains good fairness when a multi-hop flow competes with a single-hop flow. Wei Wang 0102, Ben Leong, Wei Tsang Ooi |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Introduction to the Special Issue on MMSys 2014 and NOSSDAV 2014abstractNo abstract available. Kuan-Ta Chen, Songqing Chen, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Automated Link Generation for Sensor-Enriched Smartphone ImagesabstractThe ubiquity of the smartphones makes them ideal platforms for generating in-situ content. In well-attended events, photos captured by attendees have diverse views that could be subjected to occlusion and abnormal lighting effects that could obscure the view. Such unstructured photo collections also have significant redundancy. Thus, a scene that is partially occluded or has bad contrast in one photo may be captured in another photo, possibly with higher details. We propose an application called Autolink that automatically establishes content-based links between sensor-annotated photos in unstructured photo collections captured using smartphones, such that users could navigate between high-context and high-detail images. This hierarchically structured image collection facilitates the design of applications for navigation and discovery, analytics about user photography patterns, user taste, and content/event popularity. Autolink includes a framework that constructs this hierarchy efficiently and with little content-specific training data by combining photo content processing with associated sensor logs obtained from multiple participants. We evaluated the performance of Autolink on two real-world sensor tagged photo datasets. The result shows that Autolink is able to efficiently cluster photos at 20 times faster than candidate algorithms, into the appropriate hierarchy with at least 70% precision and 37% better recall than candidate algorithms. Padmanabha Venkatagiri Seshadri, Mun Choon Chan, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Wireless Multicast for Zoomable Video StreamingabstractZoomable video streaming refers to a new class of interactive video applications, where users can zoom into a video stream to view a selected region of interest in higher resolutions and pan around to move the region of interest. The zoom and pan effects are typically achieved by breaking the source video into a grid of independently decodable tiles. Streaming the tiles to a set of heterogeneous users using broadcast is challenging, as users have different link rates and different regions of interest at different resolution levels. In this article, we consider the following problem: Given the subset of tiles that each user requested, the link rate of each user, and the available time slots, at which resolution should each tile be sent, to maximize the overall video quality received by all users. We design an efficient algorithm to solve this problem and evaluate the solution on a testbed using 10 mobile devices. Our method is able to achieve up to 12dB improvements over other heuristic methods. Mun Choon Chan, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2014 | BezelCopy: an efficient cross-application copy-paste technique for touchscreen smartphonesabstractCopy-Paste (CP) operations on touchscreen smartphones are not as easy to perform as compared with similar operations on desktop computers. The smaller screen size and input area make both text selection and application switching more difficult to perform. To enable faster copy-paste on touchscreen smartphones, we introduce BezelCopy, a copy-paste technique that uses a bezel-swipe gesture to determine a rough area of interest in the document. Chosen text is magnified in a new panel to enable fast and precise selection. With the new panel, users can perform easy tap-and-drag gestures to select the exact content, and tap the application icon on the bottom of the panel to paste it to the target application. Users can further adjust the location of the pasted text in the target application using drag and drop. We conducted two experiments to compare the performance of BezelCopy with alternative approaches, and our results show that BezelCopy outperform existing copy-paste techniques for a number of commonly performed copy-paste tasks. Simon T. Perrault, Shengdong Zhao 0001, Wei Tsang Ooi |
AVI | 4 |
| 2014 | What does computer vision say about face reading?abstractFace reading is an ancient Chinese practice that assesses a person's personality traits by applying a set of rules on facial features. In this paper, we take an image processing approach to validate these face reading rules. We collected facial features and personality traits from 54 human subjects, analyzed their correlation, and found only weak correlations between the facial features and personality traits. We also attempted to discover new, more complex, face reading rules based on the data we collected, but we found no new rules with significant correlation. In short, we found no evidence that supports this ancient Chinese practice. Terence Sim, Wei Tsang Ooi |
ICIP | 3 |
| 2014 | 3D Interest Maps From Simultaneous Video RecordingsabstractWe consider an emerging situation where multiple cameras are filming the same event simultaneously from a diverse set of angles. The captured videos provide us with the multiple view geometry and an understanding of the 3D structure of the scene. We further extend this understanding by introducing the concept of 3D interest map in this paper. As most users naturally film what they find interesting from their respective viewpoints, the 3D structure can be annotated with the level of interest, naturally crowdsourced from the users. A 3D interest map can be understood as an extension of saliency maps in the 3D space that captures the semantics of the scene. We evaluate the idea of 3D interest maps on two real datasets, taken from the environment or the cameras that are equipped enough to have an estimation of the poses of cameras and a reasonable synchronization between them. We study two aspects of the 3D interest maps in our evaluation. First, by projecting them into 2D, we compare them to state-of-the-art saliency maps. Second, to demonstrate the usefulness of the 3D interest maps, we apply them to a video mashup system that automatically produces an edited video from one of the datasets. Axel Carlier, Lilian Calvet, Duong-Trung-Dung Nguyen, Wei Tsang Ooi, Pierre Gurdjos, Vincent Charvillat |
ACM Multimedia | 4 |
| 2014 | Jiku director 2.0: a mobile video mashup system with zoom and pan using motion mapsabstractIn this demonstration, we show an automated mobile video mashup system that takes a set of videos filming the same scene as input, and generate an output mashup video consisting of temporally coherent clips selected from these input videos. The key difference of our system over the existing state-of-the-art is that it can generate virtual close-up shots and three camera operations: zooming in, zooming out, and panning, automatically. To achieve this, the system first computes the motion maps of the input videos and then determines a set of rectangles that correspond to highly interesting regions (in terms of motion). The choice of which shot types to use is done heuristically, ensuring diversity and coherency in the content presented in the mashup. Duong-Trung-Dung Nguyen, Axel Carlier, Wei Tsang Ooi, Vincent Charvillat |
ACM Multimedia | 3 |
| 2014 | Mixing Tile Resolutions in Tiled Video: A Perceptual Quality AssessmentabstractThe mismatch between increasingly large video resolution and constrained screen size of mobile devices has led to the proposal of zoomable video systems based on tiled video. In the current system, a tiled video frame is constructed from multiple tiles in a single resolution stream. In this paper, we explore the perceptual effect of mixed-resolution tiles in tiled video, in which tiles within a video frame could come from streams with different resolutions, with the aim to tradeoff bandwidth and perceptual video quality. To understand how users perceive the video quality of mixed-resolution tiled video, we conducted a psychophysical study with 50 participants on tiled videos where the tile resolutions are randomly chosen from two resolution levels with equal probability. The experiment results show that in many cases, we can mix tiles from HD (1920×1080p) stream and tiles from 1600×900p stream without being noticed by the viewers. Even when participants notice quality degradation in videos combined with tiles from HD stream and tiles from 960×540p stream, the majority of participants still accept the degradation when viewing videos with low and medium motion; and greater than 40% of participants accept the quality degradation when viewing video with dense motion. Vu-Thanh Nguyen, Wei Tsang Ooi, Mun Choon Chan |
NOSSDAV | 3 |
| 2014 | Guest editorial: special issue on network and systems support for games
Maha Abdallah, Khaled Boussetta, Wei Tsang Ooi |
Multim. Syst. | 3 |
| 2014 | Bandwidth adaptation for 3D mesh preview streamingabstractOnline galleries of 3D models typically provide two ways to preview a model before the model is downloaded and viewed by the user: (i) by showing a set of thumbnail images of the 3D model taken from representative views (or keyviews); (ii) by showing a video of the 3D model as viewed from a moving virtual camera along a path determined by the content provider. We propose a third approach called preview streaming for mesh-based 3D objects: by streaming and showing parts of the mesh surfaces visible along the virtual camera path. This article focuses on the preview streaming architecture and framework and presents our investigation into how such a system would best handle network congestion effectively. We present three basic methods: (a) stop-and-wait , where the camera pauses until sufficient data is buffered; (b) reduce-speed , where the camera slows down in accordance to reduce network bandwidth; and (c) reduce-quality , where the camera continues to move at the same speed but fewer vertices are sent and displayed, leading to lower mesh quality. We further propose two advanced methods: (d) keyview-aware , which trades off mesh quality and camera speed appropriately depending on how close the current view is to the keyviews, and (e) adaptive-zoom , which improves visual quality by moving the virtual camera away from the original path. A user study reveals that our keyview-aware method is preferred over the basic methods. Moreover, the adaptive-zoom scheme compares favorably to the keyview-aware method, showing that path adaptation is a viable approach to handling bandwidth variation. Shanghong Zhao 0002, Wei Tsang Ooi, Axel Carlier, Géraldine Morin, Vincent Charvillat |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Secure Cloud-Based Volume Ray-CastingabstractAdvances in cloud computing have allowed volume rendering tasks, typically done by volume ray-casting, to be outsourced to cloud data centers. The availability of volume data and rendered images (which can contain important information such as the disease information of a patient) to a third-party cloud provider, however, presents security and privacy challenges. This paper addresses these challenges by proposing a secure cloud-based volume ray-casting framework that distributes the rendering tasks among the data centers and hides the information that is exchanged between the server and a data center, between two data centers, and between a data center and the client by using Shamir's secret sharing, such that none of the data centers has enough information to know the secret data and/or rendered image. Experiments and analyses show that our framework is highly secure and requires low computation cost. Manoranjan Mohanty, Wei Tsang Ooi, Pradeep K. Atrey |
CloudCom (1) | 2 |
| 2013 | Scale me, crop me, knowme not: Supporting scaling and cropping in secret image sharingabstractSecret image sharing is a method for distributing a secret image amongst n data stores, each storing a shadow image of the secret, such that the original secret image can be recovered only if any k out of the n shares is available. Existing secret image sharing schemes, however, do not support scaling and cropping operations on the shadow image, which are useful for zooming on large images. In this paper, we propose an image sharing scheme that allows the user to retrieve a scaled or cropped version of the secret image by operating directly on the shadow images, therefore reducing the amount of data sent from the data stores to the user. Results and analyses show that our scheme is highly secure, requires low computational cost, and supports a large number of scale factors with arbitrary crop. Manoranjan Mohanty, Wei Tsang Ooi, Pradeep K. Atrey |
ICME | 2 |
| 2013 | A Stochastic Model for Automatic Extraction of 3D Neuronal Morphology
Sreetama Basu, Maria S. Kulikova, Elena A. Zhizhina, Wei Tsang Ooi, Daniel Racoceanu |
MICCAI (1) | 4 |
| 2013 | Jiku director: a mobile video mashup systemabstractIn this technical demonstration, we demonstrate a Web-based application called Jiku Director that automatically creates a mashup video from event videos uploaded by users. The system runs an algorithm that considers view quality (shakiness, tilt, occlusion), video quality (blockiness, contrast, sharpness, illumination, burned pixels), and spatial-temporal diversity (shot angles, shot lengths) to create a mashup video with smooth shot transitions while covering the event from different perspectives. Duong-Trung-Dung Nguyen, Mukesh Saini, Vu-Thanh Nguyen, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2013 | Joserlin: joint request and service scheduling for peer-to-peer non-linear media accessabstractA peer-to-peer non-linear media streaming system needs to schedule both on-demand and prefetch requests carefully so as to reduce the server load and ensure good user experience. In this work, we propose, Joserlin, a joint request and service scheduling solution that not only alleviates request contentions (requests compete for limited service capacity), but also schedules the prefetch requests by considering their contributions to potential reduction of server load. In particular, we propose a novel request binning algorithm to prevent self-contention among on-demand requests issued from the same peer. A service and rejection policy is devised to resolve contention among on-demand requests issued from different neighbors. More importantly, Joserlin employs a gain function to prioritize prefetch requests at both requesters and responders, and a prefetch request issuing algorithm to fully utilize available upload bandwidth. Evaluation with traces collected from a popular networked virtual environment shows that Joserlin leads to 20%~60% reduction in server load. Zhen Wei Zhao, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2013 | The jiku mobile video datasetabstractProliferation of mobile devices with video recording capability has lead to a tremendous growth in the amount of user-generated mobile videos. Researchers have embarked on developing new interesting applications and enhancement algorithms for mobile video. There is, however, no standard dataset with videos that could represent characteristics of mobile videos captured in realistic scenarios. In this paper, we present our effort to create one such dataset, consisting of videos simultaneously recorded using mobile devices in an unconstrained manner by multiple users attending performance events. Each video is accompanied by concurrent readings from accelerometer and compass sensors. At the time of writing, the dataset contains 473 video clips, with a total length of 30 hours 41 minutes and total size of 122.8 GB. We believe this dataset is useful as a common benchmark dataset for a variety of different research topics on mobile videos, including video analytics, video quality enhancement, and automatic video mashups. Mukesh Saini, Padmanabha Venkatagiri Seshadri, Wei Tsang Ooi, Mun Choon Chan |
MMSys | 3 |
| 2013 | 3D mesh preview streamingabstractPublishers of 3D models online typically provide two ways to preview a model before the model is downloaded and viewed by the user: (i) by showing a set of thumbnail images of the 3D model taken from representative views (or keyviews); (ii) by showing a video of the 3D model as viewed from a moving virtual camera along a path determined by the content provider. We propose a third approach called preview streaming for mesh-based 3D object: by streaming and showing parts of the mesh surfaces visible along the virtual camera path. This paper focuses on the preview streaming architecture and framework, and presents our investigation into how such a system would best handle network congestion effectively. We study three basic methods: (a) stop-and-wait, where the camera pauses until sufficient data is buffered; (b) reduce-speed, where the camera slows down in accordance to reduce network bandwidth; and (c) reduce-quality, where the camera continues to move at the same speed but fewer vertices are sent and displayed, leading to lower mesh quality. We further propose a keyview-aware method that trades off mesh quality and camera speed appropriately depending on how close the current view is to the keyviews. A user study reveals that our keyview-aware method is preferred over the basic methods. Shanghong Zhao 0002, Wei Tsang Ooi, Axel Carlier, Géraldine Morin, Vincent Charvillat |
MMSys | 2 |
| 2013 | Adaptive antenna adjustment for 3D urban wireless mesh networksabstractWe design and evaluate a new type of wireless mesh nodes called Dyntenna nodes that are equipped with steerable omnidirectional antenna. Designed for 3D wireless mesh networks, these nodes adaptively adjust the antenna orientation to increase throughput by improving the Received Signal Strength Indicator (RSSI) reading between nodes. We demonstrate the importance of being able to programmatically orient the antenna, by presenting the measurement results from our 3D urban mesh testbed. We propose a simple antenna adjustment algorithm that can improve the throughput for 26% of one-hop paths and 35% of multi-hop paths by a median value of 31% and 46%, respectively. Our algorithm converges quickly and typically probes less than 10% of all possible antenna orientations on average. Guoqing Yu, Wei Wang 0102, Kim Leng Yong, Ben Leong, Wei Tsang Ooi |
SECON | 5 |
| 2013 | APRICOD: An access-pattern-driven distributed caching middleware for fast content discovery of noncontinuous media accessabstractContent discovery is a major source of latency in peer-to-peer (P2P) media streaming systems, especially in the presence of noncontinuous user access, such as random seek in Video-on-Demand (VoD) streaming and teleportation in a Networked Virtual Environment (NVE). After the aforementioned user interactions, streaming systems often need to initiate the content discovery process to identify where to retrieve the requested media objects. Short content lookup latency is demanded to ensure smooth user experience. Existing content discovery systems based on either a Distributed Hash Table (DHT) or gossip mechanism cannot cope with noncontinuous access efficiently due to their long lookup latency. In this work, we propose an access-pattern-driven distributed caching middleware named APRICOD, which caters for fast and scalable content discovery in peer-to-peer media streaming systems, especially when user interactions are present. APRICOD exploits correlations among media objects accessed by users, and adapts to shift in the user access pattern automatically. We first present a general APRICOD design that can be used with any existing content discovery system. We then present an implementation of APRICOD on top of Pastry, which we use to evaluate APRICOD. Our evaluation in a 1024-node system, using a Second Life trace with 5,735 users and a VoD trace with 54 users, shows that APRICOD can effectively resolve all continuous access queries with a single hop deterministically with node failure as an exception, and resolve noncontinuous access queries with a single hop with high probability. Zhen Wei Zhao, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Modeling the effect of user interactions on mesh-based P2P VoD streaming systemsabstractUser interactions such as seeks and pauses are widely supported by existing Peer-to-Peer Video-on-Demand (P2P VoD) streaming systems. Their effect on the streaming system, however, has not been well studied. Seeks cause peers to skip part of the video, making them stay in the system for shorter time, and thus contribute less. On the other hand, only part of the video is downloaded due to seeks, reducing peers' demand from the system. It is unclear which factor dominates the effect of seeks on the streaming system. Pauses during playback, on one hand, allow peers to stay longer in the system and upload more content. When interleaved with seeks, however, long pauses may increase peers' demand unnecessarily as peers may download content that will eventually be skipped by subsequent forward seeks. The collective effect of seeks and pauses, together with the known random peer departure, is unintuitive and needs to be addressed properly so as to understand the effect of human factors on the streaming system performance. In this article, we develop an analytical model to both qualitatively and quantitatively study the effect of seeks and pauses on mesh-based P2P VoD streaming systems, in particular, the effect on the server cost. Our model can help in understanding how human factors such as seeks and pauses affect the streaming system performance, tuning a P2P VoD system towards better system performance and stability, and providing a framework for capacity planning. Zhen Wei Zhao, Sameer Samarth, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | AutoComPaste: auto-completing text as an alternative to copy-pasteabstractThe copy-paste command is a fundamental and widely used operation in daily computing. It is generally regarded as a simple task but the process can become tedious when frequent window switching is required to copy-paste across different documents. Auto-completion is another popular operation aimed at reducing users' typing effort. It contrasts to copy-paste by allowing for text completion without switching windows. However, the available content for completion is predefined. We introduce AutoComPaste, an enhanced autocompletion technique for cross-document copy-paste. AutoComPaste allows users to copy-paste different granularity of text from all opened documents without window switching. Our theoretical analysis and empirical study show that AutoComPaste nicely complements traditional copy-paste techniques and outperforms the traditional copy-paste techniques when users have knowledge of the content to be copied. Shengdong Zhao 0001, Fanny Chevalier, Wei Tsang Ooi, Chee Yuan Lee |
AVI | 3 |
| 2012 | El-pincel: a painter cloud service for greener web pagesabstractDue to their thin size, vivid colors, high contrast and power efficiency, OLED (Organic Light-Emitting Diode) display and its variants such as AMOLED (Active Matrix OLED) displays are increasingly replacing traditional LCD (Liquid Crystal Display) screens in smart phones. However, the power efficiency of OLED screens greatly depends on the luminance and colors of the displayed contents on the screen. Web browsing is one of the most widely used applications in mobile devices. In this paper, we present our cloud service, which intelligently re-paints the web pages in real-time with power efficient colors and HVS (Human Visual System) based tone mapping techniques, without adversely affecting the identity (brand color) of the web pages as well as the user's browsing experience. El-pincel helps to save up to 60% of OLED energy with color combinations that ensure good legibility and pleasing affective response to human eyes. Anand Bhojan, Lee Kee Chong, Ee-Chien Chang, Mun Choon Chan, Akkihebbal L. Ananda, Wei Tsang Ooi |
ACM Multimedia | 6 |
| 2012 | ACM multimedia 2012 workshop on crowdsourcing for multimediaabstractCrowdsourcing for multimedia involves exploiting both human intelligence and the combination of a large number of individual human contributions (i.e., the 'wisdom of the crowd') to develop techniques, systems and data sets that advance the state of the art. The ACM Multimedia 2012 Workshop on Crowdsourcing for Multimedia (CrowdMM 2012) provides a forum presenting crowdsourcing techniques for multimedia, as well as innovative ideas exemplifying how multimedia research can benefit from crowdsourcing. Through presented papers, invited talks and a panel, the workshop will promote interactive discussion on the scope and research potentials of crowdsourcing. The goal is to provide information to the multimedia research community on the principles of crowdsourcing and to inspire researchers to address the limitations of current studies by innovative use of human computation and collective intelligence. The workshop views crowdsourcing in the broad sense: it encompasses both unsolicited human contributions, e.g., tags assigned by users to images, and also solicited contributions, e.g., annotations gathered by making use of crowdsourcing platforms that micro-outsource tasks to a large pool of human workers. Kuan-Ta Chen, Wei-Ta Chu, Martha A. Larson, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2012 | Secure cloud-based medical data visualizationabstractOutsourcing the tasks of medical data visualization to cloud centers presents new security challenges. In this paper, we propose a framework for cloud-based remote medical data visualization that protects the security of data at the cloud centers. To achieve this, we integrate the cryptographic secret sharing with pre-classification volume ray-casting and propose a secure volume ray-casting pipeline that hides the color-coded information of the secret medical data during rendering at the data centers. Results and analysis show the utility of the proposed framework. Manoranjan Mohanty, Pradeep K. Atrey, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2012 | MoViMash: online mobile video mashupabstractWith the proliferation of mobile video cameras, it is becoming easier for users to capture videos of live performances and socially share them with friends and public. As an attendee of such live performances typically has limited mobility, each video camera is able to capture only from a range of restricted viewing angles and distance, producing a rather monotonous video clip. At such performances, however, multiple video clips can be captured by different users, likely from different angles and distances. These videos can be combined to produce a more interesting and representative mashup of the live performances for broadcasting and sharing. The earlier works select video shots merely based on the quality of currently available videos. In real video editing process, however, recent selection history plays an important role in choosing future shots. In this work, we present MoViMash, a framework for automatic online video mashup that makes smooth shot transitions to cover the performance from diverse perspectives. Shot transition and shot length distributions are learned from professionally edited videos. Further, we introduce view quality assessment in the framework to filter out shaky, occluded, and tilted videos. To the best of our knowledge, this is the first attempt to incorporate history-based diversity measurement, state-based video editing rules, and view quality in automated video mashup generations. Experimental results have been provided to demonstrate the effectiveness of MoViMash framework. Mukesh Saini, Raghudeep Gadde, Shuicheng Yan, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2012 | Jiku live: a live zoomable video streaming systemabstractWe present Jiku Live, a client-server system that supports zoom and pan operations in live video streaming from network cameras. The client is an Android mobile application that plays back live video from a selected camera and supports multi-touch zoom and pan interaction. The server acquires video streams from network cameras and transcodes the video feeds into one-second video segments at multiple resolutions. The transcoded video supports random access into any region-of-interest (RoI) within the video. Upon receiving zoom or pan requests, the server transmits the RoIs from the corresponding video segments to the client. Arash Shafiei 0001, Ngo Quang Minh Khiem, Guntur Ravindra, Mukesh Saini, Cong Pang, Wei Tsang Ooi |
ACM Multimedia | 6 |
| 2012 | Mobile-to-Mobile Video Recommendation
Padmanabha Venkatagiri Seshadri, Mun Choon Chan, Wei Tsang Ooi |
MobiQuitous | 3 |
| 2012 | On tile assignment for region-of-interest video streaming in a wireless LANabstractWe consider the following problem in this paper: A video is encoded as a set of tiles T and is streamed to multiple users via a one-hop wireless LAN. Each user selects a region-of-interest (RoI), represented as a subset of T, in the video to watch. The RoI selected by the users may overlap. Each tile may be multicast or unicast. We define the tile assignment problem as: which subset of tiles should be multicast such that every user receives, within a transmission deadline, the subset of tiles pertaining to the RoI the user selected, while minimizing the number of unwanted tiles received by users. We present and evaluate five tile assignment methods. We show that: (i) minimizing transmission delay can lead to significant wasteful reception in the multicast group, (ii) using tile access probability to assign tiles frequently leads to assignments that violate the deadline, and (iii) a fast, greedy, heuristic works well: it performs close to the optimal method and can always find an assignment within the deadline (as long as such assignment exists). Guntur Ravindra, Wei Tsang Ooi |
NOSSDAV | 2 |
| 2012 | Adaptive encoding of zoomable video streams based on user access pattern
Ngo Quang Minh Khiem, Guntur Ravindra, Wei Tsang Ooi |
Signal Process. Image Commun. | 3 |
| 2012 | Introduction to the special section of best papers of ACM multimedia 2011abstractNo abstract available. Daniel Gatica-Perez, Gang Hua 0001, Wei Tsang Ooi, Pål Halvorsen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Introduction to special section on 3D mobile multimediaabstract10.1145/2348816.2348820 Shervin Shirmohammadi, Mohamed Hefeeda, Wei Tsang Ooi, Romulus Grigoras |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2011 | Understanding and mitigating TCP starvation in 802.11 wireless mesh networksabstractIt is well known that the pervasive IEEE 802.11 MAC is intrinsically unfair. In particular, in the topology shown in Fig. 1(a), when links AB and CD both carry backlogged transmissions, the packets from sender A experience persistent collisions at node B while sender C enjoys collision-free transmission to D. Node A can transmit successfully only if it is able to "insert" its packets into the small inter-packet gaps of C's packets. Thus, we refer to the topology in Fig. 1(a) as the unfair topology and to C and A as the superior and inferior nodes respectively. Wei Wang 0102, Ben Leong, Wei Tsang Ooi |
ICNP | 3 |
| 2011 | SYMA: A Synchronous Multihop Architecture for Wireless Ad Hoc Multiplayer GamesabstractIn this paper, we propose a synchronous multihop architecture (SYMA) for wireless ad hoc multiplayer games. In SYMA, devices are assumed to have one-hop neighbor device information, as well as timers synchronized with good accuracy. A player, called game initiator, broadcasts an invitation message to invite other players to join a new game. A joining player issues a message along with its neighborhood information, including the expected transmission time (ETT) for each neighbor node, to reply to the invitation so that the initiator can derive the topology of the participating players. The initiator applies Floyd-Warshall shortest path algorithm, taking the topology and ETT values as input to construct the shortest path spanning tree rooted at each player (or node). It then calculates the expected broadcast time (EBT) for each tree, which is useful to estimate the time for the root node of a tree to broadcast a game state to all tree nodes. Finally, the tree with the smallest EBT is selected as the communication tree and its root is designated as the coordinator, which collects time stamped actions of each player, executes the game logic, and then broadcasts the new game state. By exploiting synchronized timers, the game state consistency is guaranteed, and each node can be scheduled to send/receive game messages without causing collisions. Yu-Tan Huang, Chih-Ming Lin, Jehn-Ruey Jiang, Wei Tsang Ooi, Maha Abdallah, Khaled Boussetta |
ICPADS | 4 |
| 2011 | Combining content-based analysis and crowdsourcing to improve user interaction with zoomable videoabstractThis paper introduces a new paradigm for interacting with zoomable video. Our interaction technique reduces the number of zooms and pans required by providing recommended viewports to the users, and replaces multiple zoom and pan actions with a simple click on the recommended viewport. The usefulness of our technique is visible in the quality of the recommended viewport, which needs to match the user intention, track movement in the scene, and properly frame the scene in the video. To this end, we propose a hybrid method where content analysis is complimented by the implicit feedback of a community of users in order to recommend viewports. We first compute preliminary sets of recommended viewports by analyzing the content of the video. These viewports allow tracking of moving objects in the scene, and are framed without violating basic aesthetic rules. To improve the relevance of the recommended viewports, we collect viewing statistics as users view a video, and use the viewports they select to reinforce the importance of certain recommendations and penalize others. New recommendations that are not previously recognized by content analysis may also emerge. The resulting recommended viewports converge towards the regions in the video that are relevant to users. A user study involving 70 participants shows that an user interface incorporating with our paradigm leads to more number of zooms, into more informative regions with fewer interactions. Axel Carlier, Guntur Ravindra, Vincent Charvillat, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2011 | COZI: crowdsourced and content-based zoomable video playerabstractWe present a new user interface designed to allow easy yet effective zooming and panning into high-definition videos for playback on low resolution displays. Our system first applies state-of-the-art video analysis algorithms to detect salient regions of interest and recommends them to users. These recommendations help users to quickly identify important regions in the video and zoom into the regions with a single mouse click. The salient regions may moves according to the movement of track objects, further reducing the need for users to manually pan to track an object of interests. To further improve the relevance of the recommended regions, users' interactions are logged and analyzed. The actual regions selected and viewed by users serve as a feedback and is integrated into the system to improve the recommendations. We have implemented a Web-based version of the user interface, running on modern browsers supporting HTML5. We describe the algorithms and optimizations used to implement and improve the system. Axel Carlier, Arash Shafiei 0001, Julien Badie, Salim Bensiali, Wei Tsang Ooi |
ACM Multimedia | 5 |
| 2011 | Towards synergy between the open source and the research multimedia communitiesabstractThis panel extends current efforts from the ACM Multimedia 2011 Organization Committee in taking an important step towards open source projects. The panelists include speakers who are among the leading figures from the open source community. The goal is to provide a shared space for discussion and interaction among consolidated and new open source projects and multimedia researchers. Pablo César, Wei Tsang Ooi, Ben Moskowitz, Zohar Babin, Dick C. A. Bulterman, Rainer Lienhart, Robert Richter |
ACM Multimedia | 2 |
| 2011 | Towards understanding user tolerance to network latency in zoomable video streamingabstractWe conducted a user study with 35 participants viewing 5 video clips to understand user tolerance to network latency when zooming and panning in zoomable video streams. With zooming or panning, unseen spatial regions in a frame are revealed and momentarily in an unknown state until data arrive from the server. To handle such unknown state, two common concealment schemes are used, namely Black scheme and Low-Res scheme. Black scheme renders the newly revealed region as black pixels, while Low-Res covers the unknown part with data from a low resolution video stream, which is additionally streamed by the server. In the context of these schemes, our study based on the simulation of delays shows that users are more tolerable to delay in Low-Res scheme. Up to 94% of participants can tolerate 1 second delay and 80% can tolerate up to a delay of 2 seconds in Low-Res scheme, while only 77% of participants can tolerate 1 second delay in Black scheme. The tolerable delay in zoomable video streaming is higher than thresholds found in some high interactive multimedia applications Ngo Quang Minh Khiem, Guntur Ravindra, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2011 | Peer-assisted texture streaming in metaversesabstractUser-extensible metaverses need an effective way to disseminate massive and dynamic 3D contents (i.e., meshes, textures, animations, etc.) to end users, and at the same time maintain a low consumption of server bandwidth. Peer-to-peer (or peer-assisted) technologies have been widely considered as a desirable complementary solution to efficaciously offload servers in large-scale media streaming applications. However, due to both the bandwidth constraints of heterogeneous users and unpredictable access patterns of latency-sensitive 3D contents, it is very challenging to cut the server bandwidth cost in metaverses. In this paper, we propose a peer-assisted texture streaming architecture to minimize the server bandwidth consumption without degrading the end-user satisfaction. We propose a game-theoretic peer selection strategy which achieves a good trade-off between performance and complexity. Our algorithm is light-weight, and can efficiently utilize the bandwidth of users in a fully decentralized manner by enabling each peer (i.e., user) to quickly select its content providers who can satisfy the requests of the peer within the latency constraint of the content. We evaluate our algorithm through an extensive comparison study based on simulations using realistic data (i.e., avatar mobility traces and textures) collected from Second Life. The simulation results show that the proposed algorithm can effectively reduce the server bandwidth consumption without degrading the user experience. Ke Liang 0001, Roger Zimmermann, Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2011 | APRICOD: a distributed caching middleware for fast content discovery of non-continuous media accessabstractWe propose an access pattern-driven distributed caching middleware named APRICOD, which can be built on top of any existing content discovery system. APRICOD caters for fast and scalable content discovery in peer-to-peer media streaming systems, especially when user interactions that leads to non-continuous media access (such as random seek in video and teleportation in virtual environment) are present. APRICOD caches query results based correlations among media objects accessed by users, reducing the content discovery time. Our evaluation using a VoD access trace shows that close to 70% of non-continuous access queries can be resolved with a single hop using APRICOD. Zhen Wei Zhao, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2011 | Towards peer-assisted rendering in networked virtual environmentsabstractThis paper introduces a new technique, called peer-assisted rendering, that aims to enable interactive navigation in a 3D networked virtual environment using a resource-constrained device, by speeding up the rendering. A resource-constrained client requests part of the rendered scenes from other peers with similar viewpoints within the virtual environment, and merges the rendered parts into its own view. This approach is more scalable than previous solutions based on server-based pre-rendering. The goal of this paper is to make a strong case for the feasibility of peer-assisted rendering through the following two messages. First, by analyzing a large number of user traces from a popular virtual world called Second Life, we show that there are surprisingly many users with similar viewpoints and encompass large number of common objects in their viewing areas, indicating that a client can potentially find multiple other peers that can assist in rendering. Second, by combining three different rendering methods, each contributing to rendering of different classes of objects in the scene, we show that it is possible for a client to render the scene efficiently with little visual artifacts. Minhui Zhu, Sebastien Mondet, Géraldine Morin, Wei Tsang Ooi |
ACM Multimedia | 4 |
| 2011 | Adaptive encoding of zoomable video streams based on user access patternabstractZoomable video allows users to selectively zoom and pan into regions of interest within the video for viewing at higher resolutions. Such interaction requires dynamic cropping of RoIs on the source video. In this paper, we consider how the bandwidth needed to transmit the RoIs can be reduced by carefully encoding the source video. The key idea is to exploit user access patterns to the RoIs, and encode different regions of the video with different encoding parameters based on the popularity of the region. We show that our encoding method can reduce the expected bandwidth by upto 27%. Ngo Quang Minh Khiem, Guntur Ravindra, Wei Tsang Ooi |
MMSys | 3 |
| 2011 | Modeling progressive mesh streaming: Does data dependency matter?abstract3D triangular meshes are becoming an increasingly prevalent data type in networked applications such as digital museums, online games, and virtual worlds. In these applications, a 3D mesh is typically coded progressively, yielding a multiresolution representation suitable for streaming. While such progressive coding allows incremental rendering for users while data is being transmitted, it introduces dependencies between data, causing delay in rendering when packets are lost. This article quantitatively analyzes the effects of such dependency by modeling the distribution of decoding time as a function of mesh properties and network parameters. We apply our model to study two extreme cases of dependency in progressive meshes and show that the effect of dependencies on decoded mesh quality diminishes with time. Our model provides the expected decoded mesh quality at the receiver at a given time. Based on this expected value, we propose a packetization strategy that improves the decoded mesh quality during the initial stage of streaming. We validate the accuracy of our model under a variety of network conditions, including bursty losses, fluctuating RTT, and varying sending rate. The values predicted from our model match the measured value reasonably well in all cases except when losses are too bursty. Wei Tsang Ooi, Sebastien Mondet, Romulus Grigoras, Géraldine Morin |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | Video quality for face detection, recognition, and trackingabstractMany distributed multimedia applications rely on video analysis algorithms for automated video and image processing. Little is known, however, about the minimum video quality required to ensure an accurate performance of these algorithms. In an attempt to understand these requirements, we focus on a set of commonly used face analysis algorithms. Using standard datasets and live videos, we conducted experiments demonstrating that the algorithms show almost no decrease in accuracy until the input video is reduced to a certain critical quality, which amounts to significantly lower bitrate compared to the quality commonly acceptable for human vision. Since computer vision percepts video differently than human vision, existing video quality metrics, designed for human perception, cannot be used to reason about the effects of video quality reduction on accuracy of video analysis algorithms. We therefore investigate two alternate video quality metrics, blockiness and mutual information, and show how they can be used to estimate the critical video qualities for face analysis algorithms. Pavel Korshunov, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2010 | Crowdsourced automatic zoom and scroll for video retargetingabstractScreen size and display resolution limit the experience of watching videos on mobile devices. The viewing experience can be improved by determining important or interesting regions within the video (called regions of interest, or ROIs) and displaying only the ROIs to the viewer. Previous work focuses on analyzing the video content using visual attention model to infer the ROIs. Such content-based technique, however, has limitations. In this paper, we propose an alternative paradigm to infer ROIs from a video. We crowdsource from a large number of users through their implicit viewing behavior using a zoom and pan interface, and infer the ROIs from their collective wisdom. A retargeted video, consisting of relevant shots determined from historical users behavior, can be automatically generated and replayed to subsequent users who would prefer a less interactive viewing experience. This paper presents how we collect the user traces, infer the ROIs and their dynamics, group the ROIs into shots, and automatically reframe those shots to improve the aesthetics of the video. A user study with 48 participants shows that our automatically retargeted video is of comparable quality to one handcrafted by an expert user Axel Carlier, Vincent Charvillat, Wei Tsang Ooi, Romulus Grigoras, Géraldine Morin |
ACM Multimedia | 3 |
| 2010 | Reducing Frame Rate for Object Tracking
Pavel Korshunov, Wei Tsang Ooi |
MMM | 2 |
| 2010 | Supporting zoomable video streams with dynamic region-of-interest croppingabstractStreaming of an arbitrary region of interest (RoI) from a high resolution video is essential to supporting zooming and panning within a video stream. This paper explores two methods for RoI-based streaming, referring to them as tiled streaming and monolithic streaming. Tiled streaming partitions video frames into grid of tiles and encodes each tile as an independently decodable stream. Monolithic streaming applies to video encoded using off-the-shelf encoder, and relies on a pre-computed dependency information to send the necessary bits for the RoI. We evaluated these two methods in terms of bandwidth efficiency, storage requirement, and computational costs under different video encoding parameters. Experimental results show that bandwidth efficiency of tiled streams for RoI-based streaming reduces when tile size increases, despite improvement in compression efficiency. In the case of monolithic streams, use of a larger motion vector range coupled with careful run-time optimization can still improve the bandwidth efficiency, despite an increase in motion vector dependency. Ngo Quang Minh Khiem, Guntur Ravindra, Axel Carlier, Wei Tsang Ooi |
MMSys | 4 |
| 2010 | Towards understanding user tolerance to network latency and data rate in remote viewing of progressive meshesabstractWe conducted experiments with 38 users who interacted with 3 progressively streamed and rendered 3D meshes in order to study their tolerance levels for network data rate and delay. Our study shows that over 90% of users can tolerate a data rate of 80 KBps and above (when the delay is 400ms) and over 95% of users can tolerate delay up to 1 second (when the data rate is 100 KBps). Our study shows that data rate and delay tolerance levels do not vary significantly among the three meshes we used. Ransi Nilaksha De Silva, Wei Tsang Ooi, Shengdong Zhao 0001 |
NOSSDAV | 3 |
| 2010 | Guest Editorial: "Special issue on selected papers from NOSSDAV 2009"
Wei Tsang Ooi, Dongyan Xu |
Multim. Syst. | 1 |
| 2009 | Group-Based Peer-to-Peer 3D Streaming AuthenticationabstractIn this paper, we present an authentication scheme for group-based peer-to-peer 3D streaming that takes advantage of secure group communication to reduce computation overheads of authentication. Users of the same interest first form a group, and one of the users is then elected to be the trusted leader, who is to download 3D contents, to verify their authenticity and integrity, and to send group members a checksum value for each 3D content piece via group secure channels. By the encrypted checksum values, a group member can authenticate 3D contents downloaded from any source. Since checksum is cheap to compute, much computation is saved. We have evaluated the computation saving of the proposed scheme for the case of progressive meshes based on the hash chain signature 3D streaming authentication scheme. We have also evaluated the rendering quality when embedding hash chain signatures into the least significant bits (LSBs) of the mesh data to make the signatures imperceptible. Mo-Che Chan, Jehn-Ruey Jiang, Chao-Wei Hung, Wei Tsang Ooi |
ICPADS | 4 |
| 2009 | Evaluation of the HyperVerse Avatar Management Scheme Based on the Analysis of Second Life TracesabstractMassive multiuser virtual environments (MMVEs) and the idea of a global scale 3D Web have grown popular in recent years. While commercial precursors of such environments for the most part rely on centralized client/server architectures, it is commonly accepted that a global scale virtual online world can only be realized in a distributed fashion. Within the HyperVerse project, we have developed and recently presented a two-tier Peer-to-Peer (P2P) architecture that incorporates a loosely structured P2P overlay of user peers and a highly structured overlay of server machines constituting a reliable backbone service. In such a distributed environment, an essential question is how avatars are tracked and interconnected in order to allow mutual rendering and interaction. We have previously proposed a hybrid avatar management scheme that utilizes the backbone service for avatar tracking if necessary, but handles tracking in a P2P fashion when peers can track each other to reduce the backbone load. This paper presents a detailed performance analysis of this algorithm under a realistic scenario, using traces from a large scale MMVE called Second Life. Moreover this paper presents and evaluates an optimization for the hybrid avatar tracking scheme that can be utilized under a weaker condition. Markus Esch, Wei Tsang Ooi, Ingo Scholtes |
ICPADS | 2 |
| 2009 | Peer-assisted view-dependent progressive mesh streamingabstractProgressive mesh streaming is increasingly used in 3D networked applications, such as online games, virtual worlds, and digital museums. To scale such applications to a large number of users without high infrastructure cost, we apply peer-to-peer techniques to mesh streaming. We consider two issues: how to partition a progressive mesh into chunks and how to lookup the provider of a chunk. For the latter issue, we investigated into two solutions, which trade off server overhead and response time. The first uses a simple centralized lookup service, while the second organizes peers into groups according to the hierarchical structure of the progressive meshes to take advantage of access pattern. Simulation results show that our proposed systems are robust under high churn rate, reduce the server overhead by more than 90%, keep control overhead below 10%, and achieve low average response time. Wei Tsang Ooi |
ACM Multimedia | 3 |
| 2009 | A demonstration of MobiTree: progressive 3D tree models streaming on mobile clientsabstractWe demonstrate MobiTree, a system we built that allows progressive streaming and rendering of 3D tree models on a mobile phone. MobiTree enables user to retrieve not only typical media describing a species (such as text and image), but also a 3D model that yields detail information about the structure of the branches and the foliage. MobiTree adopts our previous proposed progressive representation to speed up display of trees at the mobile client, trading off latency and quality. Progressivity also allows MobiTree to flexibly adopt the level of details to the capability of the mobile devices. Andra Doran, Sebastien Mondet, Romulus Grigoras, Géraldine Morin, Wei Tsang Ooi, Frédéric Boudon |
ACM Multimedia | 5 |
| 2009 | Towards characterizing user interaction with progressively transmitted 3D meshesabstract10.1145/1631272.1631438 Ransi Nilaksha De Silva, Wei Tsang Ooi, Shengdong Zhao 0001 |
ACM Multimedia | 4 |
| 2009 | Avatar mobility in user-created networked virtual worlds: measurements, analysis, and implications
Huiguang Liang, Ransi Nilaksha De Silva, Wei Tsang Ooi, Mehul Motani |
Multim. Tools Appl. | 3 |
| 2009 | Integrated Optimization of Video Server Resource and Streaming Quality Over Best-Effort NetworkabstractA video streaming server needs to adapt its source/channel encoding parameters (or configurations) to changes in network conditions and to differences in users' connection profiles. The adaptation can be achieved by adjusting parameters such as frame rate, error protection ratio, and resolution. Ideally, the server should adapt the serving configurations with respect to the current network and user conditions to improve received video quality. However, adaptations that optimize playable frame rate require intensive computation, and storing all possible configurations requires a tremendous amount of storage. This brings forth the issues of how to obtain good video quality and reduce server resources usage at the same time. We address this issue in this paper. Our approach is based on the observation that transcoding between certain configurations can be performed very efficiently. We propose a framework to compute a set of configurations to store on the server by considering two opposing goals: (a) maximizing expected received quality of the video, and (b) minimizing server resource usage by lowering transcoding cost and expected number of switches between configurations. The second objective also reduces the number of configurations, and therefore reduces the total storage required. Our framework models the relationship among different configurations in a partial order, formulates the search of a good set of configurations as an energy minimization problem, and we use techniques in image segmentation to solve the problem. Experimental results show that our framework relieves the server load and increases the number of clients served, while only slightly reducing the expected frame rate. Ee-Chien Chang, Wei Tsang Ooi, Mun Choon Chan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Compact and progressive plant models for streaming in networked virtual environmentsabstractJust as in the real world, plants are important objects in virtual worlds for creating pleasant and realistic environments, especially those involving natural scenes. As such, much effort has been made in realistic modeling of plants. As the trend moves towards networked and distributed virtual environments, however, the current models are inadequate as they are not designed for progressive transmissions. In this article, we fill in this gap by proposing a progressive representation for plants based on generalized cylinders. We model the shape and thickness of branches in a plant as Bézier curves, group the curves according to the similarity, and differentially code the curves to represent the plant in a compact and progressive manner. To facilitate the transmission of the plants, we quantify the visual contribution of each branch and use this weight in packet scheduling. We show the efficiency of our representations and the effectiveness of our packet scheduler through experiments over a wide area network. Sebastien Mondet, Géraldine Morin, Romulus Grigoras, Frédéric Boudon, Wei Tsang Ooi |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2008 | Textures in Second Life: Measurement and AnalysisabstractWe collected packet traces from second life client sessions and analyzed the packet contents. We observed that textures constitute a majority of the network traffic. We further characterized the textures from three selected regions in second life in terms of their size and spatial distributions. We found that textures in these regions exhibit a different size distribution from files on a file system or documents on the Web. We also verified the intuition that texture objects are spatially non-uniformly distributed. Surprisingly, we found that the selected second life regions can contain up to hundreds of megabytes of textures, and there exist locations in these regions that encompass a large portion of these textures within their area-of-interest. Our work motivates the need to manage textures carefully and efficiently in the design of networked virtual environments such as second life, and hints at the amount of storage and bandwidth required at a peer if peer-to-peer techniques are applied for texture caching. Our traces are useful for simulation studies and can lead to a model to generate realistic workload for networked virtual environments. Huiguang Liang, Mehul Motani, Wei Tsang Ooi |
ICPADS | 3 |
| 2008 | Streaming of plants in distributed virtual environmentsabstractInternational audience Sebastien Mondet, Géraldine Morin, Romulus Grigoras, Frédéric Boudon, Wei Tsang Ooi |
ACM Multimedia | 6 |
| 2008 | Receiver-driven view-dependent streaming of progressive meshabstractProgressive mesh streaming enables users to view 3D meshes over the network with increasing level of details, by sending coarse version of the meshes initially, followed by a series of refinements. To optimally increase the rendered mesh quality, refinements should be sent in descending order of their visual contributions based on the user's viewpoint. A common approach is to let the sender decide this sending order, but the computational cost of making this decision prohibits such sender-driven approach from scaling to large number of clients. To improve scalability, we propose a receiver-driven protocol, in which the receiver decides the sending order and explicitly requests the refinements, while the sender simply sends the data requested. The sending order is computed at the receiver by estimating the visibility and visual contributions of the refinements, even before receiving them, with the help of GPU. Experiments show that our protocol reduces the CPU cost of the sender by 24% and the outgoing traffic of the sender by 40%. Wei Tsang Ooi |
NOSSDAV | 2 |
| 2008 | Equilibrium analysis through separation of user and network behavior
Y. C. Tay, Dinh Nguyen Tran, Eric Yi Liu, Wei Tsang Ooi, Robert Morris 0005 |
Comput. Networks | 4 |
| 2008 | Multimedia Applications in Mobile/Wireless ContextabstractThe five articles in this special section focus on multimedia applications in mobile/wireless context. Two papers focus on video transmission over 802.11a/b and WiMax, the next discusses a visual sensor network in wireless environments, and the last two propose supporting technologies for two mobile applications that rely on multimedia interactions. Bo Shen 0003, Wei Tsang Ooi, Giacomo Morabito, Eckehard G. Steinbach |
IEEE Trans. Multim. | 2 |
| 2007 | Reducing Data-Memory Footprint of Multimedia Applications by Delay RedistributionabstractIt is now common for multimedia applications to be partitioned and mapped onto multiple processing elements of a system-on-chip architecture. An important design constraint in such architectures is that the FIFO buffers connecting the processing elements (in a pipelined fashion) should not overflow and the playout buffer should never underflow. To meet these constraints, an usual design practice is to increase the initial playout delay after which the output device starts reading from the playout buffer. Although implementing this technique is straightforward and involves only the the computation of an appropriate playout delay, it suffers from the downside of a large playout buffer being required. In this paper, instead of associating the playout delay solely with the output device, we propose to redistribute this delay among all the processing elements running the various tasks of the multimedia application. We show that this delay redistribution technique can signficantly reduce (up to 70%) the total on-chip memory required. Balaji Raman 0001, Samarjit Chakraborty, Wei Tsang Ooi, Santanu Dutta |
DAC | 3 |
| 2007 | Congestion Control in Distributed Media StreamingabstractDistributed media streaming, which uses multiple senders to collaboratively and simultaneously stream media content to a receiver, poses new challenges in congestion control. Such approach establishes multiple flows within a session. Since conventional congestion control only aims to make each of these flows TCP-friendly, selfish users can increase the number of flows to grab a larger share of the bandwidth, introducing more congestion and degrading the overall network performance. To address this issue, we propose the idea of task-level TCP-friendliness, which enforces TCP-friendliness upon a set of flows belonging to a task instead of upon individual flow. We design DMSCC, a congestion control scheme, to achieve task-level TCP-friendliness in distributed media streaming. By observing shared congestion, DMSCC identifies the set of flows experiencing congestion and dynamically adjusts those flows such that their combined throughput is TCP-friendly. To achieve this goal, DMSCC addresses two issues: (i) given abeta(betabeta-times the throughput of a TCP flow, and (ii) how to identify the set of flows that share a bottleneck. In our simulations, DMSCC can effectively regulate the throughput of flows on every bottleneck, resulting in a TCP-friendly combined throughput. Wei Tsang Ooi |
INFOCOM | 2 |
| 2007 | An analytical model for progressive mesh streamingabstract3D triangular mesh is becoming an increasingly important data type for networked applications such as digital museums, online games, and virtual worlds. In these applications, a multi-resolution representation is typically desired for streaming large 3D meshes, allowing for incremental rendering at the viewers while data is still being transmitted. Such progressive coding, however, introduces dependencies between data. This paper quantitatively analyzes the effects of such dependency on the intermediate decoded mesh quality when the progressive mesh is transmitted over a lossy network, by modeling the distribution of decoding time as a function of mesh properties and network parameters. To illustrate the usefulness of our analytical model, we describe three of its applications. First, we show how it can be used to analytically compute the expected decoded mesh quality. Second, we study two extreme cases of dependency in progressive mesh and show that the effect of dependencies on decoded mesh quality diminishes with time. Finally, based on the model, we propose a packetization strategy that improves the decoded mesh quality during the initial stage of streaming. Wei Tsang Ooi, Sebastien Mondet, Romulus Grigoras, Géraldine Morin |
ACM Multimedia | 2 |
| 2006 | Games are up for DVFSabstractGraphics-intensive computer games are no longer restricted to high-performance desktops, but are also available on a variety of portable devices ranging from notebooks to PDAs and mobile phones. Battery life has been a major concern in the design of both the hardware and the software for such devices. Towards this, dynamic voltage and frequency scaling (DVFS) has emerged as a powerful technique. However, the showcase application for DVFS algorithms so far has largely been video decoding, primarily because it is computationally expensive and its workload exhibits a high degree of variability. This paper investigates the possibility of applying DVFS to interactive computer games, which to the best of our knowledge has not been studied before. We show that the variability in the workload associated with a popular First Person Shooter game like Quake II is significantly higher than video decoding. Although this variability makes game applications an attractive candidate for DVFS, it is unclear if DVFS algorithms can be applied to games due to their interactive (and hence highly unpredictable) nature. In this paper, we show using detailed experiments that (surprisingly) interactive computer games are highly amenable to DVFS. Towards this we present a novel workload characterization of computer games, based on the game engine for Quake II. We believe that our findings might potentially lead to a number of innovative DVFS algorithms targeted towards game applications, exactly as video decoding has motivated a variety of schemes for DVFS. Yan Gu 0002, Samarjit Chakraborty, Wei Tsang Ooi |
DAC | 3 |
| 2006 | Analysis and implications of student contact patterns derived from campus schedulesabstractCharacterizing mobility or contact patterns in a campus environment is of interest for a variety of reasons. Existing studies of these patterns can be classified into two basic approaches - model based and measurement based. The model based approach involves constructing a mathematical model to generate movement patterns while the measurement based approach measures locations and proximity of wireless devices to infer mobility patterns. In this paper, we take a completely different approach. First we obtain the class schedules and class rosters from a university-wide Intranet learning portal, and use this information to infer contacts made between students. The value of our approach is in the population size involved in the study, where contact patterns among 22341 students are analyzed. This paper presents the characteristics of these contact patterns, and explores how these patterns affect three scenarios. We first look at the characteristics from the DTN perspective, where we study inter-contact time and time distance between pairs of students. Next, we present how these characteristics impact the spread of mobile computer viruses, and show that viruses can spread to virtually the entire student population within a day. Finally, we consider aggregation of information from a large number of mobile, distributed sources, and demonstrate that the contact patterns can be exploited to design efficient aggregation algorithms, in which only a small number of nodes (less than 0.5%) is needed to aggregate a large fraction (over 90%) of the data. Vikram Srinivasan, Mehul Motani, Wei Tsang Ooi |
MobiCom | 3 |
| 2005 | Approximate VCCs: a new characterization of multimedia workloads for system-level MpSoC designabstractSystem-level design methods specifically targeted towards multimedia applications have recently received a lot of attention. Multimedia workloads are known to have a high degree of variability. Therefore, designs based on a worstcase analysis of such workloads tend of be overly pessimistic. We address this issue by introducing a new concept called approximate variability characterization curves (or Approximate VCCs), to characterize the "average-case" behavior of multimedia workloads in a parameterized fashion. Since most multimedia applications only have soft real-time constraints, it is often possible to tolerate a small amount of performance degradation. By allowing such small degradations in the performance, large amounts of resource savings are possible. The concept of Approximate VCCs that we present in this paper allows a designer to quantitatively account for the performance degradation and the associated resource savings. We illustrate this using two typical system design cases. Samarjit Chakraborty, Wei Tsang Ooi |
DAC | 3 |
| 2005 | Packetization of 3D progressive meshes for streaming over lossy networksabstractIn this paper, we consider the problem of packetizing progressive 3D geometry models for streaming over a lossy network. We adopt a graph-theoretic approach to model packetization of progressive 3D models, with the goal of minimizing dependencies among packets. We show that this packetization problem is strongly NP-complete, and we propose two packing heuristics. Our experiments show that both heuristics perform better than the naive method. In particular, the greedy packing heuristic produces significant improvement in the number of rendered nodes when the network is lossy. Yan Gu 0002, Wei Tsang Ooi |
ICCCN | 2 |
| 2005 | Controlling Gossip Protocol Infection Pattern Using Adaptive FanoutabstractWe propose and evaluate a model for controlling infection patterns defined over rounds or real time in a gossip-based protocol using adaptive fanout. We model three versions of gossip-based protocols: the synchronous protocol, the pseudosynchronous protocol and the asynchronous protocol. Our objective is to ensure that the members of a group receive a desired message within a bounded latency with very high probability. We argue that the most important parameter that controls the latency of message delivery is the fanout used during gossiping, i.e., the number of gossip targets chosen in a particular instance of gossip. We formally analyze the three protocols and provide expressions for fanout. We introduce the idea of using variable fanouts in different rounds in the synchronous protocol. We define fanout as a function of time for the asynchronous protocol such that an expected infection pattern is observed with high probability. For a better understanding of the theoretical model, we develop a pseudosynchronous protocol to highlight the modelling done in order to derive time dependent fanout. We show that our protocols generate Theta(n log n) messages, which is optimal for gossip protocols. We aim to use the gossiping mechanism for large-scale group communication with soft real time constraints. This would alleviate the dependence on tree-based deterministic protocols which usually lack scalability Satish Verma, Wei Tsang Ooi |
ICDCS | 2 |
| 2005 | Critical video quality for distributed automated video surveillanceabstractLarge-scale distributed video surveillance systems pose new scalability challenges. Due to the large number of video sources in such systems, the amount of bandwidth required to transmit video streams for monitoring often strains the capability of the network. On the other hand, large-scale surveillance systems often rely on computer vision algorithms to automate surveillance tasks. We observe that these surveillance tasks present an opportunity for trade-off between the accuracy of the tasks and the bit rate of the video being sent. This paper shows that there exists a sweet spot, which we term critical video quality that can be used to reduce video bit rate without significantly affecting the accuracy of the surveillance tasks. We demonstrate this point by running extensive experiments on standard face detection and face tracking algorithms. Our experiments show that face detection works equally well even if the quality of compression is significantly reduced, and face tracking still works even if the frame rate is reduced to 6 frames per second. We further develop a prototype video surveillance system to demonstrate this idea. Our evaluation shows that we can achieve up to 29 times reduction in video bit rate when detecting faces and 16 times reduction when tracking faces. This paper also proposes a formal rate-accuracy optimization framework which can be used to determine appropriate encoding parameters in distributed video surveillance systems that are subjected to either bandwidth constraints or accuracy constraints. Pavel Korshunov, Wei Tsang Ooi |
ACM Multimedia | 2 |
| 2005 | Retransmission in distributed media streamingabstractThis paper considers the use of Automatic Request Request (ARQ) schemes in distributed media streaming. We analytically model three different ARQ schemes and derive effective packet loss rate and burst length of these schemes. Our model is verified through simulations and experiments over wide-area network. Our results show that retransmitting lost packet from senders other than the one who lost the packet could reduce effective loss rate and burst length. We also find that ARQ with a dedicated retransmitter outperforms other schemes if the retransmitter is chosen appropriately. Wei Tsang Ooi |
NOSSDAV | 2 |
| 2005 | Meeting CPU constraints by delaying playout of multimedia tasksabstractMultimedia applications today constitute a significant fraction of the workload running on portable devices such as mobile phones, PDAs and MP3 players. However, the processors in such devices are usually not powerful enough to support multiple concurrently executing multimedia tasks. In this context, different processor scheduling algorithms have attracted a lot of attention. This paper attempts to address the CPU constraint problem from a different perspective. It is based on the observation that by increasing the playout delay of a multimedia task, the minimum processor frequency required to run the task decreases. This is due to the high data-dependent variability in the execution requirements of multimedia tasks. We also present a framework, using which it is possible to compute the minimum processor frequency corresponding to any playout delay. Given a set of concurrently executing multimedia tasks, using our framework it is possible to compute the playout delays for each of these tasks, such that the sum of their corresponding processor cycle requirements do not exceed the maximum frequency supported by the processor. Balaji Raman 0001, Samarjit Chakraborty, Wei Tsang Ooi |
NOSSDAV | 3 |
| 2004 | Hierarchical, non-uniform locality sensitive hashing and its application to video identificationabstractSearching for similar video clips in large video database, or video identification, requires finding the nearest neighbor in high-dimensional feature space. Locality sensitive hashing, or LSH, is a well-known indexing method that allows us to efficiently find approximate nearest neighbor in such space. We address two weaknesses of LSH when applied to the video identification problem. We propose two enhancements to LSH, and show that our enhancements improve the performance of LSH significantly in terms of efficiency and accuracy. Zixiang Yang, Wei Tsang Ooi, Qibin Sun |
ICME | 2 |
| 2004 | Distributed construction of resource-efficient overlay tree by approximating MSTabstractThis work presents a distributed protocol called RESMO for constructing an overlay tree to support video streaming applications. RESMO reduces network resource usage by approximating minimum spanning tree and achieves low end-to-end latency between the sender and each receiver at the same time. The resulting overlay is a compromise between overlay minimum spanning tree and shortest path tree. We evaluated the tree constructed by RESMO through simulations, and found that RESMO gives significant improvement over existing protocols in terms of link stress, relative delay penalty and resource usage. Wei Tsang Ooi |
ICME | 2 |
| 2004 | Layered coding with good allocation outperforms multiple description coding over multiple pathsabstractPacket loss is a serious problem that severely affects the quality of multimedia streaming over error-prone networks. To reduce the variability of packet loss and delay, packets can be transmitted over different network paths (path diversity), after being coded by error-concealment source coding methods like multiple description coding (MDC) or layered coding (LC). Researches in this area lead to a common belief that MDC is better than LC when the network conditions (packet loss rate, bandwidth) are grave. However, We show that the decision of which packets to send over which paths can greatly affect the performance of LC and MDC, therefore the quality of the streams received. Particularly, using our analytical framework and polynomial algorithms for finding optimal packet allocations, we show that LC outperforms MDC under various critical network conditions. Vu-Thanh Nguyen, Ee-Chien Chang, Wei Tsang Ooi |
ICME | 3 |
| 2004 | Processor Frequency Selection for SoC Platforms for Multimedia ApplicationsabstractOf late, there has been a considerable interest in generic and configurable system-on-chip platforms specifically targeted towards implementing multimedia applications. A number of such platforms offer the possibility of including processor soft cores which are highly customizable. For voltage/frequency scaled processors, such customization includes the selection of appropriate voltage/frequency operating points which are tuned to the application set to be mapped onto the platform. In this context, we present an analytical framework that can guide a system designer in identifying the frequency ranges that should be supported by the different processors of a platform architecture. This framework can also be used to identify how such frequency ranges depend on the different parameters of the architecture (such as on-chip buffer sizes), and the performance impacts associated with selecting a particular frequency range. In the case of multimedia streaming applications, identifying such performance impacts and tradeoffs involved in customizing a platform architecture is especially difficult due to the bursty nature of on-chip traffic arising out of multimedia processing and the high variability in their execution requirements. The framework presented here is designed to precisely capture such characteristics and can be used in the design-space exploration of energy-aware platform architectures for multimedia processing. Alexander Maxiaguine, Samarjit Chakraborty, Wei Tsang Ooi |
RTSS | 4 |
| 2001 | Distributing media transformation over multiple media gatewaysabstractMedia gateways have been proposed as a solution to the network heterogeneity problem in media multicasting. Services on the gateways transform media streams as they flow through the gateways. In this paper, we present our work on composable services in media gateways. A user can request a computation to be performed on a set of media streams. The system then distributes the computation over multiple gateways for execution. We present an algorithm for decomposing the computation into sub-computations, and an application-level protocol that locates appropriate media gateways to run these sub-computations. Wei Tsang Ooi, Robbert van Renesse |
ACM Multimedia | 1 |
| 2000 | An adaptive protocol for localing programmable media gatewaysabstractWe describe a new control protocol called Adaptive Gateway Location Protocol (AGLP). In this protocol, a client requests a computation on a multimedia stream. AGLP discovers programmable Internet servers that process multimedia streams, and assigns the computation to one of these so-called gateways. AGLP continuously searches for alternate gateways, and, transparent to users, migrates computations between them to improve efficiency. The AGLP protocol uses soft-states for robustness and scale. Simulation results support that our protocol quickly locates gateways and migrates computations while keeping the load on the network low. We also outline planned enhancements to AGLP. Wei Tsang Ooi, Robbert van Renesse |
ACM Multimedia | 1 |
| 1998 | The Dali multimedia software libraryabstractThis paper presents a new approach for constructing libraries for building processing-intensive multimedia software. Such software is currently constructed either "from scratch" or by using high-level libraries. We have found that the second approach produces inefficient code, while the first approach is time-consuming. We therefore designed and implemented Dali, a set of reusable, high-performance primitives and abstractions that are at an intermediate point in this design space. By decomposing common multimedia data types and operations into thin abstractions and primitives, programs written using Ball are shorter and more reusable than hand-tuned C code, yet achieve competitive performance. This paper describes the design and implementation of Dali. Wei Tsang Ooi, Brian Christopher Smith, Sugata Mukhopadhyay, Haye Hsi Chan, Steve Weiss, Matthew Chiu, Jiesang Song |
MMSP | 1 |