VLDB 2026 Research / reviewers in the wild / expert
B. Prabhakaran 0001
dblp:p/BPrabhakaran · also Balakrishnan Prabhakaran 0001, Prabhakaran Balakrishnan 0001
· DBLP profile ↗
144ranked-venue papers
10as first author
12since 2021 · last 2026
0000-0003-0385-8662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 99 · 9 first-author · 6 since 2021Computer networks · 23 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 9Artificial intelligence and machine learning · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAOA - Completion-Assisted Object-CAD AlignmentabstractAccurately aligning CAD models to their corresponding objects in indoor RGB-D scans is a central challenge in 3D semantic reconstruction. The task requires estimating a 9-Degree-of-Freedom (DoF) pose-position, rotation, and scale along three axes-but is hindered by noisy and incomplete scans, as well as segmentation errors that cause geometric distortions. We present Completion-Assisted Ob-ject-CAD Alignment (CAOA), a method that integrates a semantically and contextually aware point cloud completion module with a symmetry-aware relative pose estimation algorithm, enabling precise alignment of CAD models to scanned objects. Existing completion methods are typically trained and evaluated on synthetic datasets, which often fail to generalize to real-world scans. To bridge this gap, we introduce a synthetic data generation strategy tailored to indoor scenes, significantly reducing the synthetic-to-real domain gap-validated through quantitative comparisons with widely used completion datasets. In addition, we release S 2 C-Completion, an expert-annotated dataset of over 8,500 object-CAD pairs from Scan2CAD, created for real-world indoor single-object completion and intended as a new benchmark for this task. For object-CAD alignment, we incorporate symmetry information via a symmetry-aware loss, improving robustness to symmetric ambiguities. On the Scan2CAD benchmark, CAOA achieves a 17 % accuracy improvement over state-of-the-art methods. All code, datasets, and annotation tools will be publicly available on GitHub. Hiranya Garbha Kumar, Minhas Kamal, B. Prabhakaran 0001 |
3DV | 3 |
| 2026 | CIDER: Collaborative Interactive Dynamic Environments for eXtended RealityabstractRemote collaboration systems based on physical environments face several critical challenges, including data-heavy virtual representations and high latencies during data acquisition, reconstruction, rendering, and transmission. Existing approaches often suffer from significant latency, making them unsuitable for real-time collaboration, rely on static scenes that limit interaction, and require multiple specialized hardware, restricting accessibility. To address these challenges, we present Collaborative Interactive Dynamic Environments for eXtended Reality (CIDER)—the first eXtended Reality (XR) platform to integrate Mixed Reality (MR) for co-located users and Virtual Reality (VR) for remote participants through a fully automated pipeline that replicates entire physical environments. CIDER dynamically transforms a user’s physical space into an interactive virtual environment, shareable with remote collaborators within seconds. It employs an efficient approach to represent, render, distribute, and synchronize virtual scenes, achieving interaction latencies of 0.22 seconds, about 10 times lower than comparable systems (2.4 seconds). We evaluate CIDER’s performance quantitatively with collaboration-oriented metrics in scenarios where participants are separated by up to 12,000 km. We also conducted a questionnaire-based user study with 17 participants to evaluate usability and overall user experience. Furthermore, CIDER allows collaborators to participate using a broad range of devices, including personal computers (via Unity emulators, functioning similarly to a MR/VR device), MR devices (e.g., HoloLens 2), and VR devices (e.g., Meta Quest 2 and 3), enhancing accessibility and usability for diverse user groups. Hung-Jui Guo, Hiranya Garbha Kumar, Yung-Jen Lin Guo, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | RobotFingerPrint: Unified Gripper Coordinate Space for Multi-Gripper Grasp Synthesis and TransferabstractWe introduce a novel grasp representation named the Unified Gripper Coordinate Space (UGCS) for grasp synthesis and grasp transfer. Our representation leverages spherical coordinates to create a shared coordinate space across different robot grippers, enabling it to synthesize and transfer grasps for both novel objects and previously unseen grippers. The strength of this representation lies in the ability to map palm and fingers of a gripper in the unified coordinate space. Grasp synthesis is formulated as predicting the unified spherical coordinates on object surface points via a conditional variational autoencoder. The predicted unified gripper coordinates establish exact correspondences between the gripper and object points, which is used to optimize grasp pose and joint values. Grasp transfer is facilitated through the point-to-point correspondence between any two (potentially unseen) grippers and solved via a similar optimization. Extensive simulation and real-world experiments showcase the efficacy of the unified grasp representation for grasp synthesis in generating stable and diverse grasps. Similarly, we showcase real-world grasp transfer from human demonstrations across different objects.1 Ninad Khargonkar, Luis Felipe Casas, B. Prabhakaran 0001, Yu Xiang 0001 |
IROS | 3 |
| 2024 | SceneReplica: Benchmarking Real-World Robot Manipulation by Creating Replicable ScenesabstractWe present a new reproducible benchmark for evaluating robot manipulation in the real world, specifically focusing on a pick-and-place task. Our benchmark uses the YCB object set, a commonly used dataset in the robotics community, to ensure that our results are comparable to other studies. Additionally, the benchmark is designed to be easily reproducible in the real world, making it accessible to researchers and practitioners. We also provide our experimental results and analyzes for model-based and model-free 6D robotic grasping on the benchmark, where representative algorithms are evaluated for object perception, grasping planning, and motion planning. We believe that our benchmark will be a valuable tool for advancing the field of robot manipulation. By providing a standardized evaluation framework, researchers can more easily compare different techniques and algorithms, leading to faster progress in developing robot manipulation methods.1 Ninad Khargonkar, Sai Haneesh Allu, Yangxiao Lu, Jishnu Jaykumar, B. Prabhakaran 0001, Yu Xiang 0001 |
ICRA | 5 |
| 2024 | MultiGripperGrasp: A Dataset for Robotic Grasping from Parallel Jaw Grippers to Dexterous HandsabstractWe introduce a large-scale dataset named MultiGripperGrasp for robotic grasping. Our dataset contains 30.4M grasps from 11 grippers for 345 objects. These grippers range from two-finger grippers to five-finger grippers, including a human hand. All grasps in the dataset are verified in the robot simulator Isaac Sim to classify them as successful and unsuccessful grasps. Additionally, the object fall-off time for each grasp is recorded as a grasp quality measurement. Furthermore, the grippers in our dataset are aligned according to the orientation and position of their palms, allowing us to transfer grasps from one gripper to another. The grasp transfer significantly increases the number of successful grasps for each gripper in the dataset. Our dataset is useful to study generalized grasp planning and grasp transfer across different grippers.1 Luis Felipe Casas Murillo, Ninad Khargonkar, B. Prabhakaran 0001, Yu Xiang 0001 |
IROS | 3 |
| 2024 | Room2XR: Virtual Interactive Collaboration in Real-world ScenesabstractThe rising prominence of Virtual Reality (VR), Mixed Reality (MR), and Extended Reality (XR) devices is transforming various industries by offering immersive and interactive experiences. Despite this, there is a notable gap in research and development of frameworks that seamlessly integrate real-world scenes into virtual environments for collaborative use. Current methods experience extended reconstruction times and latency issues, rendering them impractical for real-time collaborative applications. This paper introduces Room2XR, a framework designed to dynamically reconstruct real-world scenes into virtual semantically similar representations using CAD models, and share them for remote collaboration. Room2XR employs advanced scene understanding technologies to create detailed virtual environments, enabling multiple users to interact with and manipulate these spaces in real time. By leveraging efficient algorithms and data transmission methods, Room2XR ensures accessibility and usability even on low-bandwidth networks. The source code and detailed documentation are provided in the following two repositories- 1. https://github.com/HenryGuo2003/Room2XR-Unity and 2. https://hub.docker.com/r/kumarhiranya/vrrec Hung-Jui Guo, Hiranya Garbha Kumar, Minhas Kamal, B. Prabhakaran 0001 |
ACM Multimedia | 4 |
| 2024 | Optimizing Camera Setup for In-Home First-Person Rendering Mixed-Reality Gaming SystemabstractWith the advent of 3D humanoid reconstruction techniques, using a realistic 3D human avatar in a serious game has become popular. This realistic representation in the virtual environment could be achieved using a single low-cost RGB-D camera such as Kinect. However, properly setting up such a camera system for high-quality rendering can be challenging due to the relatively restricted in-home environment. In this paper, we address the challenge of finding optimized camera setup guidelines for an in-home first-person perspective mixed reality (MR) gaming system. We use an MR system with personalized humanoids to simulate the texture reconstruction for a user under a specific camera configuration. Then, a derivative-free optimization is leveraged as a black box approach to search for the optimized camera setup through iterative simulation. We also introduce a novel skeleton-based calibration to address the effects of physically varying the camera's position. For evaluation, two experiments are carried out to evaluate the correctness and effectiveness of the proposed calibration. Furthermore, we conduct a case study using simulation-based optimization for reconstructing lower limb amputees. This work can potentially help locate a proper camera setup for an MR system within the constraints of an in-home environment. Yu-Yen Chung, B. Prabhakaran 0001 |
IEEE Trans. Games | 2 |
| 2023 | Performance and User Experience Studies of HILLES: Home-based Immersive Lower Limb Exergame SystemabstractHead-Mounted Devices (HMDs) have become popular for home-based immersive gaming. However, using lower limb motion in the immersive virtual environment is still restricted. This work introduces an RGB-D camera-based motion capture system alongside a standalone HMD for Home-based Immersive Lower Limbs Exergame Systems (HILLES) in a seated pose. With the advance of neural network models, camera-based 3D body tracking accuracy is increasing. Nevertheless, the high demand for computing resources on model inference may compromise the game engine's performance. Accordingly, HILLES applies a distributed architecture to leverage the resources effectively. The system performances, such as frames per second and latency, are compared with a centralized system. For an immersive exergame, a pet walking around could raise safety issues. Hence, we also showcase that the camera system can provide an additional safety feature by combining an object detection model. Besides, another challenge in games focusing on lower limb interactions is the safe reachability of different virtual objects from a seated pose. Accordingly, in the user study, a stomping game with two reachability enhancements, including leg extension and seated navigation, is implemented based on the HILLES to evaluate and explore the gaming experience. The result shows that the system motivates the leg exercise, and the added enhancements may adjust the game difficulty. However, the enhancements may also distract users from focusing on leg exertion. The derived insight could benefit the lower limb exergame design in the future. Yu-Yen Chung, Thiru Annaswamy, B. Prabhakaran 0001 |
MMSys | 3 |
| 2023 | An augmented virtuality system facilitating learning through nature walk
Shanthi Vellingiri, Ryan P. McMahan, Vinu Johnson, B. Prabhakaran 0001 |
Multim. Tools Appl. | 4 |
| 2022 | Virtepex: Virtual Remote Tele-Physical Examination SystemabstractRemote strength assessment is critical for providing accessible rehabilitation, especially in the absence of in-person meetings due to the pandemic. In this paper, we introduce ”Virtepex”, an immersive exergame for remote strength assessment developed through participatory design principles. We bring out the design process starting with a needs assessment to highlight the challenges for physicians in telehealth, followed by the expert guidelines for iterative system refinement. Virtepex addresses the challenges for remote strength assessment through a marker-less and an easy-to-setup strength estimation pipeline. It utilizes an RGB-D camera for motion tracking and an inverse dynamics module for force estimation. The force estimates are used for VR object interaction and can be assessed by a physician synchronously or asynchronously for an objective evaluation. Validation by external experts shows that Virtepex produces reliable force estimates for upper body joints, indicating the potential of marker-less force estimation for future remote assessment designs. Ninad Khargonkar, Kevin Desai, B. Prabhakaran 0001, Thiru Annaswamy |
Conference on Designing Interactive Systems | 3 |
| 2022 | Dynamic X-Ray Vision in Mixed RealityabstractX-ray vision, a technique that allows users to see through walls and other obstacles, is a popular technique for Augmented Reality (AR) and Mixed Reality (MR). In this paper, we demonstrate a dynamic X-ray vision window that is rendered in real-time based on the user’s current position and changes with movement in the physical environment. Moreover, the location and transparency of the window are also dynamically rendered based on the user’s eye gaze. We build this X-ray vision window for a current state-of-the-art MR Head-Mounted Device (HMD) – HoloLens 2 [5] by integrating several different features: scene understanding, eye tracking, and clipping primitive. Hung-Jui Guo, Jonathan Z. Bakdash, Laura Marusich, B. Prabhakaran 0001 |
VRST | 4 |
| 2021 | CEFEs: A CNN Explainable Framework for ECG Signals
Barbara Mukami Maweu, Sagnik Dakshit, Rittika Shamsuddin, B. Prabhakaran 0001 |
Artif. Intell. Medicine | 4 |
| 2020 | SCeVE: A Component-based Framework to Author Mixed Reality ToursabstractAuthoring a collaborative, interactive Mixed Reality (MR) tour requires flexible design and development of various software modules for tasks such as managing geographically distributed participants, adaptable travel and virtual camera techniques, data logging for assessment of the incorporated techniques, as well as for evaluating the Quality of Experiences (QoE). In most cases, authors might have to develop all these software modules, instead of focusing only on the virtual environment design. In this article, we propose SCeVE, a component-based framework that supports flexible design and authoring of interactive MR tours by offering ease of access to four major design choices: (i) S ynchronization, (ii) C ollaborative e xploration, (iii) V isualization, and (iv) E valuation. Based on tour requirements, an author can access one or more components (or software libraries) of design choices via SCeVE’s API (Application Programming Interface) services , as demonstrated by the two case studies on group travel in a plant walk MR tour. SCeVE framework is innovative in the sense that it facilitates group travel in virtual environments involving “live” models of participants from geographically distributed sites. SCeVE empowers authors to focus only on the design of the required virtual environments. They can quickly build a diverse set of collaborative MR tours by utilizing the flexibility of SCeVE in terms of the various available options for traveling, rendering on multiple devices, and virtual camera viewpoint computation strategies. By providing data logs of various components, SCeVE facilitates performance evaluation of the various strategies used as well as the user experience in collaborative MR tours. SCeVE is designed in an extensible manner, allowing authors to add devices and software services as additional components. Shanthi Vellingiri, Ryan P. McMahan, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Using Mr. MAPP for Lower Limb Phantom Pain ManagementabstractPhantom pain is a chronic pain that is experienced as a vivid sensation stemming from the missing limb. From traditional mirror box to virtual reality-based approaches, a wide spectrum of treatments using mimic feedback of the amputated limb have been developed for alleviating phantom limb pain. In our previous work, Mixed reality-based framework for MAnaging Phantom Pain (Mr.MAPP) was presented and used to generate a virtual phantom upper limb, in real time, to manage the phantom pain. However, amputation of the lower limb is more common than that of the upper limb. Hence, in this paper, on top of demonstrating the reproducibility of the Mr.MAPP framework for upper limb, we extend it to manage lower limb phantom pain as well. Unlike an upper limb amputee, a patient with lower limb amputated is constrained to perform the training procedure in a sitting posture. Accordingly, virtual training games are designed for lower limb exercises with sitting posture such as knee flexion and extension, ankle dorsiflexion and tandem coordinated movement. Finally, the technical details of the system setup for playing the training games are introduced. Kanchan Bahirat, Yu-Yen Chung, Thiru Annaswamy, Gargi Raval, Kevin Desai, B. Prabhakaran 0001, Michael Riegler 0001 |
ACM Multimedia | 6 |
| 2019 | ADD-FAR: attacked driving dataset for forensics analysis and researchabstractFor ensuring safety and for handling emergency situations, most autonomous vehicles use remote human operators by communicating the road condition captured using LiDAR (Light Detection and Ranging) and stereo RGB cameras [10, 11]. Recent research [2, 3] has identified several possible scenarios of forgery attacks during such data communication between autonomous vehicles and human operators. Hence, there is a distinct requirement for forensics research on the multi-modal LiDAR and RGB camera data generated from autonomous vehicles. In this paper, we present a new dataset, ADD-FAR (Attacked Driving Dataset for Forensics Analysis and Research) that contains forged driving scenarios based on KITTI Vision Benchmark Suite [4, 5]. This dataset is created by identifying objects of interest using automated 3D object detection and carrying out the attacks with different levels of risk as defined in [3]. As part of the dataset, we are also providing the scripts used for automatically generating the different types of attacks. These scripts can easily be modified to work with other driving datasets apart from KITTI as well as for creating new forms of attacks. Kanchan Bahirat, Nidhi Vaishnav, Sandeep Sukumaran, B. Prabhakaran 0001 |
MMSys | 4 |
| 2019 | Improving the Security of Visual ChallengesabstractThis article proposes new tools to detect the tampering of video feeds from surveillance cameras. Our proposal illustrates the unique cyber-physical properties that sensor devices can leverage for their cyber-security. While traditional attestation algorithms exchange digital challenges between devices authenticating each other, our work instead proposes challenges that manifest physically in the field of view of the camera (e.g., a QR code in a display). This physical (challenge) and cyber (verification) attestation mechanism can help protect systems even when the sensors (cameras) and actuators (a display, infrared LEDs, color light bulbs) are compromised. In this article, we consider skillful adversaries that can capture the correct challenges (our system is sending) and can re-create them in the response to try fooling our verification system, and we propose new algorithms to detect these powerful attackers. Also, we introduce new visual challenges that make harder for anti-forensics attackers to succeed, and we present experimental results showing how our system is robust against a variety of attacks ranging from naive attacks to more sophisticated anti-forensics attackers. Junia Valente, Kanchan Bahirat, Kelly Venechanos, Alvaro A. Cárdenas, B. Prabhakaran 0001 |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2018 | ALERT: Adding a Secure Layer in Decision Support for Advanced Driver Assistance System (ADAS)abstractWith the ever-increasing popularity of LiDAR (Light Image Detection and Ranging) sensors, a wide range of applications such as vehicle automation and robot navigation are developed utilizing the 3D LiDAR data. Many of these applications involve remote guidance - either for safety or for the task performance - of these vehicles and robots. Research studies have exposed vulnerabilities of using LiDAR data by considering different security attack scenarios. Considering the security risks associated with the improper behavior of these applications, it has become crucial to authenticate the 3D LiDAR data that highly influence the decision making in such applications. In this paper, we propose a framework, ALERT (Authentication, Localization, and Estimation of Risks and Threats), as a secure layer in the decision support system used in the navigation control of vehicles and robots. To start with, ALERT tamper-proofs 3D LiDAR data by employing an innovative mechanism for creating and extracting a dynamic watermark. Next, when tampering is detected (because of the inability to verify the dynamic watermark), ALERT then carries out cross-modal authentication for localizing the tampered region. Finally, ALERT estimates the level of risk and threat based on the temporal and spatial nature of the attacks on LiDAR data. This estimation of risk and threats can then be incorporated into the decision support system used by ADAS (Advanced Driver Assistance System). We carried out several experiments to evaluate the efficacy of the proposed ALERT for ADAS and the experimental results demonstrate the effectiveness of the proposed approach. Kanchan Bahirat, Umang Shah, Alvaro A. Cárdenas, B. Prabhakaran 0001 |
ACM Multimedia | 4 |
| 2018 | Combining skeletal poses for 3D human model generation using multiple kinectsabstractRGB-D cameras, such as the Microsoft Kinect, provide us with the 3D information, color and depth, associated with the scene. Interactive 3D Tele-Immersion (i3DTI) systems use such RGB-D cameras to capture the person present in the scene in order to collaborate with other remote users and interact with the virtual objects present in the environment. Using a single camera, it becomes difficult to estimate an accurate skeletal pose and complete 3D model of the person, especially when the person is not in the complete view of the camera. With multiple cameras, even with partial views, it is possible to get a more accurate estimate of the skeleton of the person leading to a better and complete 3D model. In this paper, we present a real-time skeletal pose identification approach that leverages on the inaccurate skeletons of the individual Kinects, and provides a combined optimized skeleton. We estimate the Probability of an Accurate Joint (PAJ) for each joint from all of the Kinect skeletons. We determine the correct direction of the person and assign the correct joint sides for each skeleton. We then use a greedy consensus approach to combine the highly probable and accurate joints to estimate the combined skeleton. Using the individual skeletons, we segment the point clouds from all the cameras. We use the already computed PAJ values to obtain the Probability of an Accurate Bone (PAB). The individual point clouds are then combined one segment after another using the calculated PAB values. The generated combined point cloud is a complete and accurate 3D representation of the person present in the scene. We validate our estimated skeleton against two well-known methods by computing the error distance between the best view Kinect skeleton and the estimated skeleton. An exhaustive analysis is performed by using around 500000 skeletal frames in total, captured using 7 users and 7 cameras. Visual analysis is performed by checking whether the estimated skeleton is completely present within the human model. We also develop a 3D Holo-Bubble game to showcase the real-time performance of the combined skeleton and point cloud. Our results show that our method performs better than the state-of-the-art approaches that use multiple Kinects, in terms of objective error, visual quality and real-time user performance. Kevin Desai, B. Prabhakaran 0001, Suraj Raghuraman |
MMSys | 2 |
| 2018 | Skeleton-based continuous extrinsic calibration of multiple RGB-D kinect camerasabstractApplications involving 3D scanning and reconstruction & 3D Tele-immersion provide an immersive experience by capturing a scene using multiple RGB-D cameras, such as Kinect. Prior knowledge of intrinsic calibration of each of the cameras, and extrinsic calibration between cameras, is essential to reconstruct the captured data. The intrinsic calibration for a given camera rarely ever changes, so only needs to be estimated once. However, the extrinsic calibration between cameras can change, even with a small nudge to the camera. Calibration accuracy depends on sensor noise, features used, sampling method, etc., resulting in the need for iterative calibration to achieve good calibration. Kevin Desai, B. Prabhakaran 0001, Suraj Raghuraman |
MMSys | 2 |
| 2018 | Real-Time, Curvature-Sensitive Surface Simplification Using Depth ImagesabstractWith the rising popularity of handheld virtual reality (VR) devices and depth sensing RGB-D cameras, a variety of VR applications merging these two technologies has been suggested. However, immersive quality of experience in such VR applications is constrained mainly by the large data size and the hardware limitations to handle it. The depth data captured by RGB-D cameras provide a dense sampling of the surface, resulting in a high-poly mesh, which is difficult to be rendered on handheld VR devices due to their limited processing power. To improve the immersive VR experience, a sparse approximation of the depth data is needed. Traditional mesh and point cloud simplification methods are iterative and so are unsuitable for real-time applications. In this paper, we introduce a depth-imagebased approach that is capable of generating a good quality sparse mesh for visualization in real time. We propose a curvaturesensitive surface simplification-CS3operator that assigns an importance measure to each point in the depth image, based on the local curvature. Further, it applies an importance-order-based restrictive sampling to generate a sparse representation that retains the overall shape as well as the finer features of the object. We also modify the 2-D sweep-line-based constrained Delaunay triangulation to generate 3-D meshes from the sparse point sampling obtained using CS3. In addition, the proposed approach preserves key surface properties, such as texture coordinates and materials. We used three different datasets containing dense 3-D models with and without texture, which are scanned using various sensors to validate and compare the robustness, real-time performance, and accuracy of the proposed method over existing approaches. Based on the experimental results, we show that the proposed CS3operator and modified 2-D sweep-line-based triangulation generate sparse meshes from depth image in real time, performing significantly faster than current state-of-the-art methods while maintaining similar visual quality. Kanchan Bahirat, Suraj Raghuraman, B. Prabhakaran 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Editorial IEEE Transactions on Multimedia Special Section on Video Analytics: Challenges, Algorithms, and ApplicationsabstractThe papers in this special section focus on the topic of video analytics. Also known as video content analysis, video analytics refer to the capability of automatically analyzing video to extract knowledge/information and detect and determine temporal and spatial events. The algorithms designed for these analytics can be implemented as software on general-purpose machines, or as hardware in specialized video processing units. Video analytics is still an emerging technology with techniques that are continuously being developed to help make widespread implementation feasible in the years ahead. Such analytics has been typically used in semantic categorization and retrieval of video databases. A goal of this special issue is to focus on video analytics beyond categorization and retrieval. With increasing hardware capability and advances in algorithms used, real-time video analytics is now being used in a wide range of domains including entertainment, health-care, retail, automotive, transport, home automation, emotion analysis, aesthetics, inappropriate content detection, safety and security. For instance, video analytics are increasingly being deployed for real-time alerts in situation monitoring systems such as traffic surveillance (vehicle counting), counting people in lines (some hospitals are using this to get more nurses from a less busy department to serve thewaiting patients) and in manufacturing (for monitoring and counting). From the sensing aspect, 3-D cameras such as RGB-D and LiDAR (Light Detection and Ranging) cameras are becoming more and more affordable, enabling additional areas of research and applications, such as self-driving cars employing video analytics on LiDAR captured data for path planning as well as obstacle detection. B. Prabhakaran 0001, Yu-Gang Jiang 0001, Hari Kalva, Shih-Fu Chang |
IEEE Trans. Multim. | 1 |
| 2018 | Designing and Evaluating a Mesh Simplification Algorithm for Virtual RealityabstractWith the increasing accessibility of the mobile head-mounted displays (HMDs), mobile virtual reality (VR) systems are finding applications in various areas. However, mobile HMDs are highly constrained with limited graphics processing units (GPUs) and low processing power and onboard memory. Hence, VR developers must be cognizant of the number of polygons contained within their virtual environments to avoid rendering at low frame rates and inducing simulator sickness. The most robust and rapid approach to keeping the overall number of polygons low is to use mesh simplification algorithms to create low-poly versions of pre-existing, high-poly models. Unfortunately, most existing mesh simplification algorithms cannot adequately handle meshes with lots of boundaries or nonmanifold meshes, which are common attributes of many 3D models. In this article, we present QEM 4VR , a high-fidelity mesh simplification algorithm specifically designed for VR. This algorithm addresses the deficiencies of prior quadric error metric (QEM) approaches by leveraging the insight that the most relevant boundary edges lie along curvatures while linear boundary edges can be collapsed. Additionally, our algorithm preserves key surface properties, such as normals, texture coordinates, colors, and materials, as it preprocesses 3D models and generates their low-poly approximations offline. We evaluated the effectiveness of our QEM 4VR algorithm by comparing its simplified-mesh results to those of prior QEM variations in terms of geometric approximation error, texture error, progressive approximation errors, frame rate impact, and perceptual quality measures. We found that QEM 4VR consistently yielded simplified meshes with less geometric approximation error and texture error than the prior QEM variations. It afforded better frame rates than QEM variations with boundary preservation constraints that create unnecessary lower bounds on overall polygon count reduction. Our evaluation revealed that QEM 4VR did not fair well in terms of existing perceptual distance measurements, but human-based inspections demonstrate that these algorithmic measurements are not suitable substitutes for actual human perception. In turn, we present a user-based methodology for evaluating the perceptual qualities of mesh simplification algorithms. Kanchan Bahirat, Chengyuan Lai, Ryan P. McMahan, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | A study on lidar data forensicsabstract3D LiDAR (Light Imaging Detection and Ranging) data has recently been used in a wide range of applications such as vehicle automation and crime scene reconstruction. Decision making in such applications is highly dependent on LiDAR data. Thus, it becomes crucial to authenticate the data before using it. Though authentication of 2D digital images and video has been widely studied, the area of 3D data forensic is relatively unexplored. In this paper, we investigate and identify three possible attacks on the LiDAR data. We also propose two novel forensic approaches as a countermeasure for such attacks and study their effectiveness. The first forensic approach utilises the density consistency check while the second method leverages the occlusion effect for revealing the forgery. Experimental results demonstrate the effectiveness of the proposed forgery attacks and raise the awareness against unauthenticated use of LiDAR data. The performance analyses of the proposed forensic approaches indicate that the proposed methods are very efficient and provide the detection accuracy of more than 95% for certain kinds of forgery attacks. While the forensic approach is unable to handle all forgery attacks, the study motivates to explore more sophisticated forensic methods for LiDAR data. Kanchan Bahirat, B. Prabhakaran 0001 |
ICME | 2 |
| 2017 | Learning-based objective evaluation of 3D human open meshesabstractCurrent state-of-the-art mesh quality measures evaluate closed and complete meshes obtained after mesh postprocessing applications, such as mesh simplification or watermarking, and compare them against the corresponding reference mesh. Emerging 3D immersive VR/AR applications use noisy 3D point cloud, typically from single RGB-D camera (such as Microsoft's Kinect) to generate standalone (no reference) 3D human open mesh (with boundaries) in real time, that needs evaluation. A learning-based objective measure is proposed to rate the visual quality by emulating human perception of 3D human open mesh quality. 2-pronged objective evaluation is performed: (a) Global holistic score captures the efficacy of the mesh to represent the human model as a whole, by considering mesh completeness and mesh noise. (b) Local part-based score caters to the need of varying roughness in different parts of the human body, by finding the deviation in the face normals for all the adjacent triangles in that part (segment). Learning technique aligns the objective scores with the subjective user evaluation, in turn combining the concepts of white-box and black-box evaluation for 3D meshes. Experimental results for a database, specifically generated for the purpose proves the efficacy of the proposed method. Kevin Desai, Kanchan Bahirat, B. Prabhakaran 0001 |
ICME | 3 |
| 2017 | QoE Studies on Interactive 3D Tele-ImmersionabstractUsers' Quality of Experience (QoE) in Interactive 3D Tele-Immersion (i3DTI) systems is influenced by several factors such as the quality of the "live" 3D avatars of the users, network latency, rendering methodology (head mounted display or regular TV type of display), etc. Hence, it becomes important to answer the question: "Is Visual Quality (VQ) the only factor to be considered or do better immersion and faster interactions matter, in having good QoE?" To answer this question, in this paper, a highly optimized state-of-the-art i3DTI framework implementation is introduced along with a soccer-inspired penalty shootout game. This game allows users to experience various situations, view different angles, perceive delays, track virtual ball motion, and play naturally using their entire bodies. A head mounted display device - Oculus Rift allows the users to get completely immersed and perform better in the penalty shootout game compared to watching themselves play on a 3D TV. This scenario is obtained in a controlled lab setting with ultra-high-speed network that has ultra-low latency, high VQ, fast and realistic interactions. Such ideal conditions are not typically available in a wide area Internet. Hence, for faster interaction scenario, we used lesser RGB-D cameras for 3D reconstruction, thereby reducing the model quality significantly. The high responsiveness of the game masked the user's perception of quality, and resulted in them not noticing the lower VQ of the reconstructed 3D models. Based on the results from the user study that focused on immersion, interaction and VQ aspects; users felt the game to be visually appealing, intuitive, engaging, and highly entertaining in all of the scenarios. Kevin Desai, Suraj Raghuraman, Rong Jin 0003, B. Prabhakaran 0001 |
ISM | 4 |
| 2017 | Mr.MAPP: Mixed Reality for MAnaging Phantom PainabstractPhantom Limb Pain or simply, Phantom Pain is a severe chronic pain that is experienced as a vivid sensation of the pain in missing body part. Epidemiological studies obtained from a large samples indicate that the short-term incidence rate of the phantom limb pain is 72% [13], while long-term incidence rate (6 months after amputation) is 67%, [5, 13]. A wide spectrum of treatments developed for alleviating phantom limb pain includes the traditional mirror box therapy as well as recently developed virtual reality-based methods. Most of the virtual reality-based methods rely on 3D CAD models of the virtual limb, animating them using the motion data acquired either from patient's existing anatomical limb or myoelectric activity at patient's stump (of the amputated limb). Since motion activity is typically captured using body sensors (Electromyography, EMG, or inertial sensors), these methods are considered as invasive approaches. Further, in the case of virtual reality-based methods, the dependency on the pre-built 3D models degrades the immersive experience due to a mismatch in the skin color, clothes, artificial and rigid look and misalignment of the phantom limb. Kanchan Bahirat, Thiru Annaswamy, B. Prabhakaran 0001 |
ACM Multimedia | 3 |
| 2017 | H-TIME: Haptic-enabled Tele-Immersive Musculoskeletal ExaminationabstractThe current state-of-the-art tele-medicine applications only allow audiovisual communication between a doctor and the patient, necessitating a clinician to physically examine the patient. The doctor relies on the physical examination performed by the clinician, along with the audiovisual dialogue with the patient. In this paper, a Haptic-enabled Tele-Immersive Musculoskeletal Examination (H-TIME) system is introduced, that allows doctors to physically examine musculoskeletal conditions of the patients remotely, by looking at the 3D reconstructed model of the patient in the virtual world, and physically feeling the patient's range of mobility using a haptic device. The proposed bidirectional haptic rendering in H-TIME can allow the doctor to evaluate a patient who suffers from problems in their upper extremities, such as the shoulder, elbow, wrist, etc., and evaluate them remotely. Real world user study was performed, between the doctors and the patients, and it highlighted the potential of the proposed system. The study indicated a high degree of correlation between the in-person and H-TIME evaluations of the patient. Both the doctors and patients involved in the study, felt that the system could potentially replace in-person consultations, someday. Yuan Tian 0002, Suraj Raghuraman, Thiru Annaswamy, Aleksander Borresen, Klara Nahrstedt, B. Prabhakaran 0001 |
ACM Multimedia | 6 |
| 2017 | A Boundary and Texture Preserving Mesh Simplification Algorithm for Virtual RealityabstractWith the increasing accessibility of the mobile head-mounted displays (HMDs), mobile virtual reality (VR) systems are finding applications in various areas. However, mobile HMDs are highly constrained with limited graphics processing units (GPUs), low processing power and onboard memory. Hence, VR developers must be cognizant of the number of polygons contained within their virtual environments to avoid rendering at low frame rates and inducing simulator sickness. The most robust and rapid approach to keeping the overall number of polygons low is to use mesh simplification algorithms to create low-poly versions of preexisting, high-poly models. Unfortunately, most existing mesh simplification algorithms cannot adequately handle meshes with lots of boundaries or non-manifold meshes, which are common attributes of 3D models made with computer-aided design tools.; [email protected] this paper, we present a high-fidelity mesh simplification algorithm specifically designed for VR. This new algorithm, QEM4VR, addresses the deficiencies of prior quadric error metric (QEM) approaches by leveraging the insight that the most relevant boundary edges lie along curvatures while linear boundary edges can be collapsed. Additionally, our QEM4VR algorithm preserves key surface properties, such as normals, texture coordinates, colors, and materials. It pre-processes the 3D models and generate their low-poly approximations offline. We used six publicly available, high-poly models, with and without textures to compare the accuracy and fidelity of our QEM4VR algorithm to previous QEM variations. We also performed a frame rate analysis with original high-poly models and low-poly models obtained using QEM4VR and previous QEM variations. Our results indicate that QEM4VR creates low-poly, high-fidelity virtual environments for VR applications on devices that are constrained by the low number of polygons they can work with. Kanchan Bahirat, Chengyuan Lai, Ryan P. McMahan, B. Prabhakaran 0001 |
MMSys | 4 |
| 2017 | A Visual Latency Estimator for 3D Tele-Immersionabstract3D Tele-Immersion systems allow geographically distributed users to interact in a virtual world using their "live" 3D models. The capture, reconstruction, transfer, and rendering of these models introduce significant latency into the system. Implicit Latency (ℒ') can be estimated using system clocks to measure the time after the data was received from the RGB-D camera, till the request to render the result. The Observed Latency (ℒ) between a real world event and the event being rendered on the display, cannot be accurately represented by ℒ' since ℒ' ignores the time taken to capture, or update the display, etc. In this paper, a Visual Pattern based Latency Estimation (VPLE) approach is introduced to calculate the real world visual latency of a system without the need for any custom hardware. VPLE generates a constantly changing pattern that is captured and rendered by the 3DTI system. An external observer records both the pattern and the rendered results at high frame rates. ℒ is estimated by calculating the difference between the generated and rendered patterns. VPLE is extended to allow ℒ estimation between geographically distributed sites. Evaluations show that the accuracy of VPLE depends on the refresh rate of the pattern, and is within 4ms. ℒ of a distributed 3DTI system implemented on the GPU is significantly lower than the CPU implementation, and is comparable to video streaming. It is also shown that the ℒ' estimates for GPU based 3DTI implementations are off by almost 100% compared to the ℒ. Suraj Raghuraman, Kanchan Bahirat, B. Prabhakaran 0001 |
MMSys | 3 |
| 2017 | Real Time Stable Haptic Rendering Of 3D Deformable Streaming SurfaceabstractIn recent years, many researches are focusing on the haptic interaction with streaming data like RGBD video / point cloud stream captured by commodity depth sensors. Most previous methods use partial streaming data from depth sensors and only investigate haptic rendering of the rigid surface without complex physics simulation. Many virtual reality and tele-immersive applications such as medical training, and art designing require the complete scene and physics simulation. In this paper, we propose a stable haptic rendering method capable of interacting with streaming deformable surface in real-time. Our method applies KinectFusion for real-time reconstruction of real-world object surface instead of incomplete surface. While construction, it simultaneously uses hierarchical shape matching (HSM) method to simulate the surface deformation in haptic-enabled interaction. We have demonstrated how to combine the fusion and physics simulation of deformation together, and proposed a continuous collision detection method based on Truncated Signed Distance Function (TSDF). Furthermore, we propose a fast TSDF warping method to update the deformation to TSDF, and a proxy finding method to find the proxy position. The proposed method is able to simulate the haptic-enabled deformation of the 3D fusion surface. Therefore it provides a novel haptic interaction for virtual reality and 3D tele-immersive applications. Experimental results show that the proposed approach provides stable haptic rendering and fast simulation of 3D deformable surface. Yuan Tian 0002, Chao Li 0021, Xiaohu Guo, B. Prabhakaran 0001 |
MMSys | 4 |
| 2017 | Modeling User Quality of Experience (QoE) through Position Discrepancy in Multi-Sensorial, Immersive, Collaborative EnvironmentsabstractUsers' QoE (Quality of Experience) in Multi-sensorial, Immersive, Collaborative Environments (MICE) applications is mostly measured by psychometric studies. These studies provide a subjective insight into the performance of such applications. In this paper, we hypothesize that spatial coherence or the lack of it of the embedded virtual objects among users has a correlation to the QoE in MICE. We use Position Discrepancy (PD) to model this lack of spatial coherence in MICE. Based on that, we propose a Hierarchical Position Discrepancy Model (HPDM) that computes PD at multiple levels to derive the application/system-level PD as a measure of performance.; [email protected] results on an example task in MICE show that HPDM can objectively quantify the application performance and has a correlation to the psychometric study-based QoE measurements. We envisage HPDM can provide more insight on the MICE application without the need for extensive user study. Shanthi Vellingiri, B. Prabhakaran 0001 |
MMSys | 2 |
| 2017 | Developing a Low Dimensional Patient Class Profile in Accordance to Their Respiration-Induced Tumor MotionabstractTumor location displacement caused by respiration-induced motion reduces the efficacy of radiation therapy. Three medically relevant patterns are often observed in the respiration-induced motion signal: baseline shift, ES-Range shift , and D-Range shift. In this paper, for patients with lower body cancer, we develop class profiles (a low dimensional pattern frequency structure) that characterize them in terms of these three medically relevant patterns. We propose an adaptive segmentation technique that turns each respiration-induced motion signal into a multi-set of segments based on persistent variations within the signal. These multi-sets of segments is then probed for base behaviors. These base behaviors are then used to develop the group/class profiles using a modified version of the clustering technique described in [1]. Finally, via quantitative analysis, we provide a medical characterization for the class profiles, which can be used to explore breathing intervention technique. We show that, with i) carefully designed feature sets, ii) the proposed adaptive segmentation technique, iii) the reasonable modifications to an existing clustering algorithm for multi-sets, and iv) the proposed medical characterization methodology, it is possible to reduce the time series respiration-induced motion signals into a compact class profile. One of our co-authors is a medical physician and we used his expert opinion to verify the results. Rittika Shamsuddin, B. Prabhakaran 0001, Amit Sawant |
Proc. VLDB Endow. | 2 |
| 2017 | Motion Capture With Ellipsoidal Skeleton Using Multiple Depth CamerasabstractThis paper introduces a novel motion capturing framework which works by minimizing the fitting error between an ellipsoid based skeleton and the input point cloud data captured by multiple depth cameras. The novelty of this method comes from that it uses the ellipsoids equipped with the spherical harmonics encoded displacement and normal functions to capture the geometry details of the tracked object. This method is also integrated with a mechanism to avoid collisions of bones during the motion capturing process. The method is implemented parallelly with CUDA on GPU and has a fast running speed without dedicated code optimization. The errors of the proposed method on the data from Berkeley Multimodal Human Action Database (MHAD) are within a reasonable range compared with the ground truth results. Our experiment shows that this method succeeds on many challenging motions which are failed to be reported by Microsoft Kinect SDK and not tested by existing works. In the comparison with the state-of-art marker-less depth camera based motion tracking work our method shows advantages in both robustness and input data modality. Liang Shuai, Chao Li 0021, Xiaohu Guo, B. Prabhakaran 0001, Jinxiang Chai |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | International workshop on multimodal virtual and augmented reality (workshop summary)abstractVirtual reality (VR) and augmented reality (AR) are expected by many to become the next wave of computing with significant impacts on our daily lives. Motivated by this, we organized a workshop on “Multimodal Virtual and Augmented Reality (MVAR)” at the 18th ACM International Conference on Multimodal Interaction (ICMI 2016). While current VR and AR installations mostly focus on the visual domain, we expect multimodality to play a crucial role in future, next generation VR/AR systems. The submissions for this workshop reflect the potential of multimodality for VR and AR, illustrate interesting new directions, and pinpoint important issues. This paper gives a short motivation for the workshop and its aim, and summarizes the aforementioned trends and challenges identified from the submissions. Wolfgang Hürst, Daisuke Iwai, B. Prabhakaran 0001 |
ICMI | 3 |
| 2016 | On the "Face of Things"abstractFace is crucial for human identity, while face identification has become crucial to information security. It is important to understand and work with the problems and challenges for all different aspects of facial feature extraction and face identification. In this tutorial, we identify and discuss four research challenges in current Face Detection/Recognition research and related research areas: (1) Unavoidable Facial Feature Alterations, (2) Voluntary Facial Feature Alterations, (3) Uncontrolled Environments, and (4) Accuracy Control on Large-scale Dataset. We also direct several different applications (spin-offs) of facial feature studies in the tutorial. Ranran Feng, B. Prabhakaran 0001 |
ICMR | 2 |
| 2016 | Augmented reality-based exergames for rehabilitationabstractRehabilitation for stroke afflicted patients, through exercises tailored for individual needs, aims at relearning basic motor skills, especially in the extremities. Rehabilitation through Augmented Reality (AR) based games engage and motivate patients to perform exercises which, otherwise, maybe boring and monotonic. Also, mirror therapy allows users to observe one's own movements in the game providing them with good visual feedback. This paper presents an augmented reality based system for rehabilitation by playing four interactive, cognitive and fun Exergames (exercise and gaming). Kevin Desai, Kanchan Bahirat, Sudhir Ramalingam, B. Prabhakaran 0001, Thiru Annaswamy, Una E. Makris |
MMSys | 4 |
| 2016 | Region graph based method for multi-object detection and tracking using depth camerasabstractIn this paper, we propose a multi-object detection and tracking method using depth cameras. Depth maps are very noisy and obscure in object detection. We first propose a region-based method to suppress high magnitude noise which cannot be filtered using spatial filters. Second, the proposed method detect Region of Interests by temporal learning which are then tracked using weighted graph-based approach. We demonstrate the performance of the proposed method on standard depth camera datasets with and without object occlusions. Experimental results show that the proposed, method is able to suppress high magnitude noise in depth maps and detect/track the objects (with and without occlusion). Sachin Mehta, B. Prabhakaran 0001 |
WACV | 2 |
| 2016 | Special issue on collaborative haptic audio-visual environments and systems
Xiaohu Guo, B. Prabhakaran 0001, Abdulmotaleb El Saddik |
Multim. Syst. | 2 |
| 2016 | Scene-based fingerprinting method for traitor tracing
Sachin Mehta, Rajarathnam Nallusamy, B. Prabhakaran 0001 |
Multim. Syst. | 3 |
| 2015 | Evaluating the efficacy of RGB-D cameras for surveillanceabstractRGB-D cameras have enabled real-time 3D video processing for numerous computer vision applications, especially for surveillance type applications. In this paper, we first present a real-time anti-forensic 3D object stream manipulation framework to capture and manipulate live RBG-D data streams to create realistic images/videos showing individuals performing activities they did not actually do. The framework uses computer vision and graphics methods to render photorealistic animations of live mesh models captured using the camera. Next, we conducted a visual inspection of the manipulated RGB-D streams (just like security personnel would do) by users who are computer vision and graphics scientists. The study shows that it was significantly difficult to distinguish between the real or reconstructed rendering of such 3D video sequences, thus clearly showing the potential security risk involved. Finally, we investigate the efficacy of forensic approaches for detecting such manipulations. Suraj Raghuraman, Kanchan Bahirat, B. Prabhakaran 0001 |
ICME | 3 |
| 2015 | Network Adaptive Textured Mesh Generation for Collaborative 3D Tele-Immersionabstract3D Tele-Immersion (3DTI) has emerged as an efficient environment for virtual interactions and collaborations in a variety of fields like rehabilitation, education, gaming, etc. In 3DTI, geographically distributed users are captured using multiple cameras and immersed in a single virtual environment. The quality of experience depends on the available network bandwidth, quality of the 3D model generated and the time taken for rendering. In a collaborative environment, achieving high quality, high frame rate rendering by transmitting data to multiple sites having different bandwidth is challenging. In this paper we introduce a network adaptive textured mesh generation scheme to transmit varying quality data based on the available bandwidth. To reduce the volume of information transmitted, a visual quality based vertex selection approach is used to generate a sparse representation of the user. This sparse representation is then transmitted to the receiver side where a sweep-line based technique is used to generate a 3D mesh of the user. High visual quality is maintained by transmitting a high resolution texture image compressed using a lossy compression algorithm. In our studies users were unable to notice visual quality variations of the rendered 3D model even at 90% compression. Kevin Desai, Kanchan Bahirat, Suraj Raghuraman, B. Prabhakaran 0001 |
ISM | 4 |
| 2015 | MMT+AVR: enabling collaboration in augmented virtuality/reality using ISO's MPEG media transportabstractAugmented Reality (AR) and Augmented Virtuality (AV) systems have been used in various fields such as entertainment, broadcasting, gaming [1], etc. Collaborative AR or AV (CAR/CAV) systems are a special kind of such system in which the interaction happens through the exchange of multi-modal data between multiple users/sites. Multiple sensors capture the real objects and enable interaction with shared virtual objects in a customizable virtual environment. Haptic devices can be added to introduce force feedback when the virtual objects are manipulated. These applications are demanding in terms of network resources to support low latency media delivery and media source switching similar to broadcast applications. Enabling real time interaction with multiple modalities with high volume data requires an advanced media transport protocol that supports low latency media delivery and fast media source (channel) switching. To enable such collaboration over a stochastic network like the Internet requires a combination of technologies from data design, synchronization to real time media delivery. MPEG Media Transport (MMT) [ISO/IEC 23008-1] is a new standard suite of protocols designed to work with demanding, real-time interactive multimedia applications, typically in the context of one-to-one and one-to-many communication. In this paper, we identify the augmentations that are required for the many-to-many nature of CAR/CAV applications and propose MMT+AVR as a middle ware solution for use in CAV applications. Through an example CAV application implemented on top of MMT+AVR, we show how it provides efficient support for developing CAV applications with ease. Karthik Venkatraman, Yuan Tian 0002, Suraj Raghuraman, B. Prabhakaran 0001, Nhut Nguyen |
MMSys | 4 |
| 2015 | Distortion score based pose selection for 3D tele-immersionabstract3D Tele-Immersion (3DTI) systems capture and transmit large volumes of data per frame to enable virtual world interaction between geographically distributed people. Large delays/latencies introduced during the transmission of these large volumes of data can lead to poor quality of experience of the 3DTI systems. Such poor experiences can possibly be overcome by animating the previously received mesh using the current skeletal data (that is very small in size and hence experiences much lower communication delays). However, using just the previously transmitted mesh for animation is not ideal and could render inconsistent results. In this paper, we present a DIstortion Score based Pose SElection (DISPOSE) approach to render the person by using an appropriate mesh for a given pose. Unlike pose space animation methods that require manual or offline time consuming pose set creation, our distortion score based scheme can choose the mesh to be transmitted and update the pose set accordingly. DISPOSE works with partial meshes and does not require dense registration enabling real time pose space creation. With DISPOSE incorporated into 3DTI, the latency for rendering the mesh on the receiving side is limited by only the transmission delay of the skeletal data (which is only around 250 bytes). Our evaluations show the effectiveness of DISPOSE for generating good quality online animation faster than real time. Suraj Raghuraman, B. Prabhakaran 0001 |
VRST | 2 |
| 2015 | Real-Time Facial Expression Recognition on SmartphonesabstractTemporal segmentation of real time video is an important part for automatic facial expression recognition system. Many studies for facial expression recognition have been carried out under restricted experimental environment such as pre-segmented video set. In this paper, we present a real-time temporal video segmenting approach for automatic facial expression recognition applicable in a smartphone. The proposed system uses a Finite State Machine (FSM) for segmenting real time video into temporal phases from neutral expression to the peak of an expression. The FSM uses Lucas-Kanade's optical flow vector based scores for state transitions to adapt the varying speeds of facial expressions. While even HMM based or hybrid HMM model based approaches handling time series data require sampling times, the proposed system runs without any sampling time delay. The proposed system performs facial expression recognition with Support Vector Machines (SVM) on every apex state after automatic temporal segmentation. The mobile app with our approach runs on Samsung Galaxy S3 with 3.7 fps and the accuracy of real-time mobile emotion recognition is about 70.6% for 6 basic emotions by 5 subjects who are not professional actors. Myunghoon Suk, B. Prabhakaran 0001 |
WACV | 2 |
| 2015 | Multi-level sample importance ranking based progressive transmission strategy for time series body sensor dataabstractBody sensors have gained increasing interest during the past several years. With more applications deployed, it is imperative to ensure the success of data analysis, which largely depends on data transmission reliability as well as the importance of samples received. Traditional approaches focus on improving data reliability through various schemes such as prioritization of MAC access. In this paper, we analyzed the characteristics of time series body sensor data and propose to rank sample importance based on a multi-level approach. With this approach, samples are grouped into five levels, indicating their importance with regard to data analysis. Then, a progressive transmission strategy is designed to transmit samples in order of their importance so that the overall received data quality is maximized. Preliminary simulation results indicate that as much as 40-60% bandwidth saving can be achieved while meeting the requirements of data analysis algorithms. Ming Li 0007, Yu Cao 0002, B. Prabhakaran 0001 |
WOWMOM | 3 |
| 2015 | Stable haptic interaction based on adaptive hierarchical shape matchingabstractIn this paper, we present a framework allowing users to interact with geometrically complex 3D deformable objects using (multiple) haptic devices based on an extended shape matching approach. There are two major challenges for haptic-enabled interaction using the shape matching method. The first is how to obtain a rapid deformation propagation when a large number of shape matching clusters exist. The second is how to robustly handle the collision response when the haptic interaction point hits the particle-sampled deformable volume. Our framework extends existing multi-resolution shape matching methods, providing an improved energy convergence rate. This is achieved by using adaptive integration strategies to avoid insignificant shape matching iterations during the simulation. Furthermore, we present a new mechanism called stable constraint particle coupling which ensures consistent deformable behavior during haptic interaction. As demonstrated in our experimental results, the proposed method provides natural and smooth haptic rendering as well as efficient yet stable deformable simulation of complex models in real time. Yuan Tian 0002, Yin Yang 0002, Xiaohu Guo, B. Prabhakaran 0001 |
Comput. Vis. Media | 4 |
| 2014 | 3D content fingerprintingabstractFingerprint is a set of features that uniquely characterizes a video. The aim of content fingerprinting is to determine the duplicate videos over the Internet. In this paper, a method for content fingerprinting of Depth-Image-Based-Rendering (DIBR) 3D videos is proposed. The proposed method is two pronged approach: (i) histogram based global fingerprint and (ii) keypoint based local fingerprint. Though global fingerprint is fast and robust towards DIBR 3D pre-processing, it is not robust against severe distortions such as change in brightness. To make the proposed method robust against such severe distortions, we have complemented the global fingerprint with widely used keypoint based local fingerprint. Experimental results show that the proposed two pass method as robust as keypoint based local fingerprinting method. Additionally, the proposed method improves the video matching time of keypoint based local fingerprint method by 60% to 100%. Sachin Mehta, B. Prabhakaran 0001 |
ICIP | 2 |
| 2014 | Opti-speech: a real-time, 3d visual feedback system for speech trainingabstractWe describe an interactive 3D system to provide talkers with real-time information concerning their tongue and jaw movements during speech. Speech movement is tracked by a magnetometer system (Wave; NDI, Waterloo, Ontario, Canada). A customized interface allows users to view their current tongue position (represented as an avatar consisting of flesh-point markers and a modeled surface) placed in a synchronously moving, transparent head. Subjects receive augmented visual feedback when tongue sensors achieve the correct place of articulation. Preliminary data obtained for a group of adult talkers suggest this system can be used to reliably provide real-time feedback for American English consonant place of articulation targets. Future studies, including tests with communication disordered subjects, are described. William F. Katz, Thomas F. Campbell, Jun Wang 0037, Eric Farrar, Jessie Colette Eubanks, Arvind Balasubramanian, B. Prabhakaran 0001, Rob Rennaker |
INTERSPEECH | 7 |
| 2014 | Demonstration abstract: upper body motion capture system using inertial sensors
Jian Wu 0016, Zhanyu Wang, Suraj Raghuraman, B. Prabhakaran 0001, Roozbeh Jafari |
IPSN | 4 |
| 2014 | Quantifying and Improving User Quality of Experience in Immersive Tele-Rehabilitationabstract3D Tele-Immersion (3DTI) environments are emerging as a new medium for human interactions and collaborations in the areas of education, sports training, physical medicine and rehabilitation. By adding a tactile element to a visually centered 3DTI environment, such applications can be made even more engaging. But it also opens up a few challenges in terms of fusing the visual and tactile data streams in a synchronous way. In this paper we describe a 3DTI Tele-Rehabilitation system with Microsoft Kinect cameras and hap tic devices. We describe some of the challenges we face in providing as well as quantifying a good quality of experience (QoE) in this system. We propose a set of solutions that: (i) improve the user's QoE (by using multi-modal prediction for handling latencies, better synchronization that accounts for the global state of the system, etc.), (ii) quantify the QoE (by designing a controlled virtual environment and by defining appropriate user QoE metrics for immersive tele-rehabilitation). The experimental results show a marked improvement in the performance of the system, consequently improving the user-experience. This is also verified by the results of the user performance study. Karthik Venkatraman, Suraj Raghuraman, Yuan Tian 0002, B. Prabhakaran 0001, Klara Nahrstedt, Thiru Annaswamy |
ISM | 4 |
| 2014 | MPEG Media Transport (MMT) for 3D Tele-Immersion Systemsabstract3D Tele-Immersion (3DTI) environments are a new medium for highly interactive and immersive means of collaborations through a shared virtual 3D environment. They have many applications in the areas of education, entertainment, sports training, tele-medicine etc. The data in these systems are multi-modal, some high volume, some high frequency and all highly correlated. We identify three major challenges in a general 3DTI system, session management, synchronization and data format conversion. We discuss the shortcomings of some of the existing protocols/solutions to them. We describe some features relevant to 3DTI of the MPEG Media Transport (MMT) standard. In this paper we evaluate the use of MMT in a 3DTI application. We provide a feature comparison with the most popular protocols currently being used in such applications, RTP, RTSP, TCP and UDP etc. MPEG DASH was another protocol that was being considered, but that also fails to fully address some of the challenges that 3DTI applications face. Through this comparison study we advocate the use of MMT in 3DTI applications. Karthik Venkatraman, Shanthi Vellingiri, B. Prabhakaran 0001, Nhut Nguyen |
ISM | 3 |
| 2014 | 3D Immersive Cardiopulmonary Resuscitation (CPR) TrainerabstractCardiopulmonary resuscitation (CPR) plays a primary role in first-aid treatment. Instead of the traditional instructor-led training course, we propose a virtual reality system which provides an immersive 3D environment for CPR training with visual and haptic feedback. To simulate a real world CPR experience, our immersive trainer system enables a trainee to perform CPR compressions to a virtual human, inside the virtual world. During the training procedure, the trainee can not only watch his/her 3D image performing CPR, but also feel the force feedback from the chest compressions in real-time. To further enhance the visual fidelity, a haptic-enabled deformable model is applied to show the visual change of chest during compression. Yuan Tian 0002, Suraj Raghuraman, Yin Yang 0002, Xiaohu Guo, B. Prabhakaran 0001 |
ACM Multimedia | 5 |
| 2014 | A novel method for post-surgery face recognition using sum of facial parts recognitionabstractPlastic surgery is becoming more and more commonplace today due to its increasing acceptance in society and its cost-affordability. This in turn has led to the need for developing highly accurate post-surgery face recognition techniques, a problem space which differs significantly from traditional face recognition. In this paper we first conduct a statistical study to show that facial plastic surgery operations correlate with a desire to conform to a golden ratio with respect to the human face. We then apply this knowledge, with the notion of considering a face in terms of the sum of its parts, to propose a novel face recognition technique. The proposed technique is then evaluated against well known datasets, and as per our experiments achieves a recognition rate of 85.35%, which significantly outperforms other state of the art techniques. Ranran Feng, B. Prabhakaran 0001 |
WACV | 2 |
| 2014 | Editorial
Albert Banchs, B. Prabhakaran 0001 |
Pervasive Mob. Comput. | 2 |
| 2013 | Facilitating fashion camouflage artabstractArtists and fashion designers have recently been creating a new form of art -- Camouflage Art -- which can be used to prevent computer vision algorithms from detecting faces. This digital art technique combines makeup and hair styling, or other modifications such as facial painting to help avoid automatic face-detection. In this paper, we first study the camouflage interference and its effectiveness on several current state of art techniques in face detection/recognition; and then present a tool that can facilitate digital art design for such camouflage that can fool these computer vision algorithms. This tool can find the prominent or decisive features from facial images that constitute the face being recognized; and give suggestions for camouflage options (makeup, styling, paints) on particular facial features or facial parts. Testing of this tool shows that it can effectively aid the artists or designers in creating camouflage-thwarting designs. The evaluation on suggested camouflages applied on 40 celebrities across eight different face recognition systems (both non-commercial or commercial) shows that 82.5% ~ 100% of times the subject is unrecognizable using the suggested camouflage. Ranran Feng, B. Prabhakaran 0001 |
ACM Multimedia | 2 |
| 2013 | A 3D tele-immersion streaming approach using skeleton-based predictionabstract3D collaborative Tele-Immersive environments allow reconstruction of real world 3D scenes in the virtual world across multiple physical locations. This kind of reconstruction results in a lot of 3D data being transmitted over the internet in real time. The current systems allow for transmission at low frame rates due to the large volume of data and network bandwidth restrictions. In this paper we propose a prediction based approach that generates future frames by animating the live model based on few skeleton points. By doing so the magnitude of data transmitted is reduced to few hundred bytes. The prediction errors are corrected when an entire frame is received. This approach allows minimal amounts (few bytes) of data to be transmitted per frame, thus allowing for high frame rates and still maintain an acceptable visual quality of reconstruction at the receiver side. Suraj Raghuraman, Karthik Venkatraman, Zhanyu Wang, B. Prabhakaran 0001, Xiaohu Guo |
ACM Multimedia | 4 |
| 2013 | A multigrid approach for bandwidth and display resolution aware streaming of 3D deformationsabstractIn this paper, we propose a novel multimedia system adaptively streaming the animation according to display resolution and/or network bandwidth. A Multigrid-like technique is used in this framework to accelerate the converging rate of the optimization of the nonlinear deformation energy. The computation is performed from coarsest mesh at the top level to the finest mesh at the bottom level and then goes back to the top again. Such V-shape calculation provides great flexibility for the networked environment. Clients are able to receive the data streaming corresponding to its display resolution and network bandwidth. A more compact form of deformation data packaging is also used in this system such that a cube element only needs six parameters instead of 24 variables as used in regular mesh representation, which significantly reduces the network overhead for the streaming. Yuan Tian 0002, Yin Yang 0002, Xiaohu Guo, B. Prabhakaran 0001 |
ACM Multimedia | 4 |
| 2012 | Quantifying the Makeup Effect in Female Faces and Its Applications for Age EstimationabstractIn this paper, a comprehensive statistical study of makeup effect on facial parts (skin, eyes, and lip) is conducted first. According to the statistical study, a method to detect whether makeup is applied or not based on input facial image is proposed, then the makeup effect is further quantified as Young Index (YI) for female age estimation. An age estimator with makeup effect considered is presented in this paper. Results from the experiments find that with the makeup effect considered, the method proposed in this paper can improve accuracy by 0.9-6.7% in CS (Cumulative Score) and 0.26-9.76 in MAE (Mean of Absolute Errors between the estimated age and the ground truth age labeled or acquired from the data) comparing with other age estimation methods. Ranran Feng, B. Prabhakaran 0001 |
ISM | 2 |
| 2012 | FaceFetch: A User Emotion Driven Multimedia Content Recommendation System Based on Facial Expression RecognitionabstractRecognition of facial expressions of users allows researchers to build context-aware applications that adapt according to the users' emotional states. Facial expression recognition is an active area of research in the computer vision community. In this paper, we present Face Fetch, a novel context-based multimedia content recommendation system that understands a user's current emotional state (happiness, sadness, fear, disgust, surprise and anger) through facial expression recognition and recommends multimedia content to the user. Our system can understand a user's emotional state through a desktop as well as a mobile user interface and pull multimedia content such as music, movies and other videos of interest to the user from the cloud with near real time performance. Mahesh Babu Mariappan, Myunghoon Suk, B. Prabhakaran 0001 |
ISM | 3 |
| 2012 | Facial Expression Recognition Using Dual Layer Hierarchical SVM Ensemble ClassificationabstractIn this paper, we present our approach for automatic facial expression recognition. We use a feature extraction technique inspired by our empirical study on human recognition of facial expressions. We propose our dual-layer hierarchical SVM ensemble mechanism for classification. We also provide system architecture and system implementation details in this paper. Mahesh Babu Mariappan, Myunghoon Suk, B. Prabhakaran 0001 |
ISM | 3 |
| 2012 | SAMHIS: A Robust Motion Space for Human Activity RecognitionabstractIn recent years, many local descriptor based approaches have been proposed for human activity recognition, which perform well on challenging datasets. However, most of these approaches are computationally intensive, extract irrelevant background features and fail to capture global temporal information. We propose to overcome these issues by introducing a compact and robust motion space that can be used to extract both spatial and temporal aspects of activities using local descriptors. We present Speed Adapted Motion History Image Space (SAMHIS) that employs a variant of Motion History Image for representing motion. This space alleviates both self-occlusion as well as the speed-related issues associated with different kinds of motion. We go on to show using a standard bag of visual words model that extracting appearance based local descriptors from this space is very effective for recognizing activity. Our approach yields promising results on the KTH and Weizmann dataset. Suraj Raghuraman, B. Prabhakaran 0001 |
ISM | 2 |
| 2012 | Immersive multiplayer tennis with microsoft kinect and body sensor networksabstractWe present an immersive gaming demonstration using the minimum amount of wearable sensors. The game demonstrated is two-player tennis. We combine a virtual environment with real 3D representations of physical objects like the players and the tennis racquet (if available). The main objective of the game is to provide as real an experience of tennis as possible, while also being as less intrusive as possible. The game is played across a network, and this opens the possibility of two remote players playing a game together on a single virtual tennis pitch. The Microsoft Kinect sensors are used to obtain a 3D point cloud and a skeletal map representation of the player. This 3D point cloud is mapped on to the virtual tennis pitch. We also use a wireless wearable Attitude and Heading Reference System (AHRS) mote, which is strapped onto the wrist of the players. This mote gives us precise information about the movement (swing, rotation etc.) of the playing arm. This information along with the skeletal map is used to implement the physics of the game. Using this game we demonstrate our solutions for simultaneous data acquisition, 3D point-cloud mapping in a virtual space, use of the Kinect and AHRS sensors to calibrate real and virtual objects and for interaction of virtual objects with a 3D point cloud. Suraj Raghuraman, Karthik Venkatraman, Zhanyu Wang, Jian Wu 0016, Jacob Clements, Reza Lotfian, B. Prabhakaran 0001, Xiaohu Guo, Roozbeh Jafari, Klara Nahrstedt |
ACM Multimedia | 7 |
| 2012 | 2012 IEEE international symposium on a world of wireless, mobile, and multimedia networks WoWMoMabstractIt is our great pleasure to welcome you to the Thirteenth IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks, WoWMoM 2012. Over the years, WoWMoM has emerged to be a flagship forum that brings together researchers from academia, industry, and government laboratories who are involved in various aspects of mobile multimedia networking technologies, ranging from communications platforms to services and applications. Albert Banchs, B. Prabhakaran 0001 |
WOWMOM | 2 |
| 2012 | Multimedia data semantics: guest editors introduction
Raphaël Troncy, B. Prabhakaran 0001, Yu Cao 0002 |
Multim. Tools Appl. | 2 |
| 2012 | Analyzing and Visualizing Jump Performance Using Wireless Body SensorsabstractAdvancement in technology has led to the deployment of body sensor networks (BSN) to monitor and sense human activity in pervasive environments. Using multiple wireless on-body systems, such as physiological data monitoring and motion capture systems, body sensor network data consists of heterogeneous physiologic and motoric streams that form a multidimensional framework. In this article, we analyze such high-dimensional body sensor network data by proposing an efficient, multidimensional factor analysis technique for quantifying human performance and, at the same time, providing visualization for performances of participants in a low-dimensional space for easier interpretation. Gaurav N. Pradhan, B. Prabhakaran 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Spectral Watermarking for Parameterized SurfacesabstractThis paper presents a blind spectral two-way watermarking framework for 3-D models with parametric information. We introduce a spectral geometric watermarking technique based on Dirichlet Manifold Harmonic Transform to alter the geometric shape, while the spectral basis functions are computed from the parametric mesh as the analysis domain. This new geometric method embeds watermarks into small surface patches without introducing discontinuity across the patch boundary, while at the same time is robust against various spatial attacks. By manipulating part of the geometric shape on the intermediate model instead of the original model, this method gains robustness against connectivity changing and cropping attacks. By combining the new geometric method with the existing texture method into the two-way watermarking framework, we can withstand various attacks applied to either geometric mesh or parametric information. Theoretical analysis and experiments show that this new geometric method is robust against the majority of attacks and the two-way watermarking framework helps achieve better robustness. Yang Liu 0013, B. Prabhakaran 0001, Xiaohu Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | Video Human Motion Recognition Using a Knowledge-Based Hybrid Method Based on a Hidden Markov ModelabstractHuman motion recognition in video data has several interesting applications in fields such as gaming, senior/assisted-living environments, and surveillance. In these scenarios, we may have to consider adding new motion classes (i.e., new types of human motions to be recognized), as well as new training data (e.g., for handling different type of subjects). Hence, both the accuracy of classification and training time for the machine learning algorithms become important performance parameters in these cases. In this article, we propose a knowledge-based hybrid (KBH) method that can compute the probabilities for hidden Markov models (HMMs) associated with different human motion classes. This computation is facilitated by appropriately mixing features from two different media types (3D motion capture and 2D video). We conducted a variety of experiments comparing the proposed KBH for HMMs and the traditional Baum-Welch algorithms. With the advantage of computing the HMM parameter in a noniterative manner, the KBH method outperforms the Baum-Welch algorithm both in terms of accuracy as well as in reduced training time. Moreover, we show in additional experiments that the KBH method also outperforms the linear support vector machine (SVM). Myunghoon Suk, Ashok Ramadass, Yohan Jin, B. Prabhakaran 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2012 | Point-Based Manifold HarmonicsabstractThis paper proposes an algorithm to build a set of orthogonal Point-Based Manifold Harmonic Bases (PB-MHB) for spectral analysis over point-sampled manifold surfaces. To ensure that PB-MHB are orthogonal to each other, it is necessary to have symmetrizable discrete Laplace-Beltrami Operator (LBO) over the surfaces. Existing converging discrete LBO for point clouds, as proposed by Belkin et al., is not guaranteed to be symmetrizable. We build a new point-wisely discrete LBO over the point-sampled surface that is guaranteed to be symmetrizable, and prove its convergence. By solving the eigen problem related to the new operator, we define a set of orthogonal bases over the point cloud. Experiments show that the new operator is converging better than other symmetrizable discrete Laplacian operators (such as graph Laplacian) defined on point-sampled surfaces, and can provide orthogonal bases for further spectral geometric analysis and processing tasks. Yang Liu 0013, B. Prabhakaran 0001, Xiaohu Guo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | PicoLife: A Computer Vision-based Gesture Recognition and 3D Gaming System for Android Mobile DevicesabstractPico Life is envisioned to be an augmented reality game in which 3D characters will be controlled by hand gestures on Android smart phones. Pico Life is currently powered by two mobile optimized engines: (1) The computer vision engine that runs our advanced object tracking program for hand tracking and (2) The 3D engine that runs our 3D models for the characters in the game. In the near future, we will be adding yet another mobile optimized engine, namely, the augmented reality engine. In this paper, we will present our work on object tracking and 3D modeling for Pico Life and contrast the performances of the two engines on three different mobile platforms, namely, Texas Instruments' OMAP3630 (Motorola Droid X running Android Gingerbread), Qualcomm's MSM8660 Snapdragon (HTC Evo 3D running Android Gingerbread) and the Texas Instruments' OMAP4430 (Blaze Development platform running Android Gingerbread). Mahesh Babu Mariappan, Xiaohu Guo, B. Prabhakaran 0001 |
ISM | 3 |
| 2011 | Motion fault detection and isolation in Body Sensor NetworksabstractSignificant amount of research and development is being directed on monitoring activities of daily living of senior citizens who live alone as well as those affected with certain disorders such as Alzheimer's and Parkinson's. A combination of sophisticated inertial sensing, wireless communication and signal processing technologies have made such a pervasive and remote monitoring possible. Due to the nature of the sensing and communication mechanisms, these monitoring sensors are susceptible to errors and failures. In this paper, we address the issue of identifying and isolating faulty sensors in a Body Sensor Network that is used for remote monitoring of daily living activities. We identify three different types of fault isolation strategies and propose both history-based and non-history based approaches. Duk-jin Kim, B. Prabhakaran 0001 |
PerCom | 2 |
| 2011 | Receiver-based loss tolerance method for 3D progressive streaming
Ziying Tang, Xiaohu Guo, B. Prabhakaran 0001 |
Multim. Tools Appl. | 3 |
| 2011 | Motion fault detection and isolation in Body Sensor Networks
Duk-jin Kim, B. Prabhakaran 0001 |
Pervasive Mob. Comput. | 2 |
| 2011 | Knowledge discovery from 3D human motion streams through semantic dimensional reductionabstract3D human motion capture is a form of multimedia data that is widely used in entertainment as well as medical fields (such as orthopedics, physical medicine, and rehabilitation where gait analysis is needed). These applications typically create large repositories of motion capture data and need efficient and accurate content-based retrieval techniques. 3D motion capture data is in the form of multidimensional time-series data. To reduce the dimensions of human motion data while maintaining semantically important features, we quantize human motion data by extracting spatio-temporal features through SVD and translate them onto a symbolic sequential representation through our proposed sGMMEM (semantic Gaussian Mixture Modeling with EM). In order to handle variations in motion capture data due to human body characteristics and speed of motion, we transform the semantically quantized values into a histogram representation. This representation is used as a signature for classification and similarity-based retrieval. We achieved good classification accuracies for “coarse” human motion categories (such as walking 92.85%, run 91.42%, and jump 94.11%) and even for subtle categories (such as dance 89.47%, laugh 83.33%, basketball signal 85.71%, golf putting 80.00%). Experiments also demonstrated that the proposed approach outperforms earlier techniques such as the wMSV (weighted Motion Singular Vector) approach and LB_Keogh method. Yohan Jin, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | Architecture and protocol design for a pervasive robot swarm communication networksabstractAbstract There has been increasing interest in deploying a team of robots, or robot swarms, to fulfill certain complicate tasks such as surveillance. Since robot swarms may move to areas of far distance, it is important to have a pervasive networking environment for communications among robots, administrators, and mobile users. In this paper, we first propose a pervasive architecture to integrate wireless mesh networks and robot swarm networks to build a robot swarm communication network within the areas of special interest. Under the proposed architecture, one or more robots can get connected with a nearby mesh router and access the remote server, while a self‐organizing mobilead hocnetwork is formed within each swarm for communications among the robots. We then address and analyze many important issues and challenges. Finally, we describe our work to enable this architecture through a scalable algorithm for autonomous swarm deployment and ROBOTRAK, a socket‐based‐swarm monitoring and control toolkit. Extensive simulation results and demonstrations are presented to show the desirable features of the proposed algorithm and toolkit. Copyright © 2009 John Wiley & Sons, Ltd. Ming Li 0007, John Harris, Min Chen 0003, Shiwen Mao, Yang Xiao 0001, Walter Read, B. Prabhakaran 0001 |
Wirel. Commun. Mob. Comput. | 7 |
| 2010 | Feature extraction method for video based Human action recognitions: Extended Optical Flow algorithmabstractThis paper focuses on the issue of improving the quality of low level 2D feature extraction for human action recognition. For instance, existing algorithms such as the Optical Flow algorithm detects noisy and irrelevant features because of its lack of ground truth data sets for complex scenes. For these features, it is difficult to extract data such as coordinate positions of the features, velocity and the direction of the moving objects, and the differential data information between different frames. Extracting such low level feature data is one of the major steps involved in video based Human action recognition. The paper proposes an extended Optical Flow algorithm focusing on human actions. This uses a Frame Jump technique along with thresholding of unwanted features to overcome the problems due to complex scenes. Frame Jump restricts to detecting only useful features by removing other features detected by the existing Optical Flow algorithm. In addition to the above, it also elucidates the integration of the proposed technique with other feature extraction algorithms. Ashok Ramadass, Myunghoon Suk, B. Prabhakaran 0001 |
ICASSP | 3 |
| 2010 | Blind invisible watermarking for 3D meshes with texturesabstractWe propose to embed watermarks by modifying the texture mapping information of 3D models rather than modifying the geometry information or texture image as existing works do. We present a blind watermarking method based on spectral decomposition that incorporates the process of Texture Image Compensation (TIC) which ensures no visual distortion. We describe a Neighbor Couple Embedding (NCE) scheme that works on the Manifold Harmonics Transform (MHT) of the texture coordinate functions. Experiments show that this method is robust against common attacks such as adding noise attacks, uniform affine transformation attacks, local modification attacks and produces no visual distortion on the rendered 3D models. Our contributions include watermarking the texture mapping information with no visual distortion as well as a novel embedding method that is robust against various possible attacks. Yang Liu 0013, B. Prabhakaran 0001, Xiaohu Guo |
ICIP | 2 |
| 2010 | Video Human Motion Recognition Using Knowledge-Based Hybrid MethodabstractHuman motion recognition in video data has several interesting applications in fields such as gaming, senior/assisted living environments, and surveillance. In these scenarios, we might have to consider adding new motion classes (i.e. new types of human motions to be recognized) as well as new training data (say, for handling different type of subjects). Hence, both accuracy of classification and training time for the machine learning algorithms become important performance parameters in these cases. In this paper, we propose a Knowledge Based Hybrid (KBH) method that can compute the probabilities for Hidden Markov Models (HMMs) associated with different human motion classes. This computation is facilitated by appropriately mixing features from two different media types (3D motion capture and 2D video). We conducted a variety of experiments comparing the proposed KBH for HMMs and the traditional Baum-Welch algorithms. With the advantage of computing the HMMs parameters in a non-iterative manner, the KBH method outperforms the Baum-Welch algorithm both in terms of accuracy as well as reduced training time. Myunghoon Suk, Ashok Ramadass, Yohan Jin, B. Prabhakaran 0001 |
ISM | 4 |
| 2010 | A multimodal virtual environment for interacting with 3d deformable modelsabstractIn this video presentation, we introduce an immersive multimodal virtual environment which supports real-time interactions with 3D deformable model through a haptic device. We include a system called "FakeSpace" to imitate 3D environment, and a PHAMTOM device to simulate touching forces. Movements of 3D deformable models are simulated based on a spectral method, and forces are simulated as spring forces. We are able to real-time update both visual and haptic feedbacks, so that to provide a more realistic user interaction. In addition, with the help of stereoscopic display, we can present an immersive 3D experience. This video illustrates the settings of our environment and demonstrates how users real-time manipulate 3D models in this immersive system using some interactive examples. Our system is reconfigurable and is useful for different applications in the fields of education, entertainment, medical simulation and so on. Ziying Tang, Anant Patel, Xiaohu Guo, B. Prabhakaran 0001 |
ACM Multimedia | 4 |
| 2010 | Streaming 3D shape deformations in collaborative virtual environmentabstractCollaborative virtual environment has been limited on static or rigid 3D models, due to the difficulties of real-time streaming of large amounts of data that is required to describe motions of 3D deformable models. Streaming shape deformations of complex 3D models arising from a remote user's manipulations is a challenging task. In this paper, we present a framework based on spectral transformation that encodes surface deformations in a frequency format to successfully meet the challenge, and demonstrate its use in a distributed virtual environment. Our research contributions through this framework include: i) we reduce the data size to be streamed for surface deformations since we stream only the transformed spectral coefficients and not the deformed model; ii) we propose a mapping method to allow models with multi-resolutions to have the same deformations simultaneously; iii) our streaming strategy can tolerate loss without the need for special handling of packet loss. Our system guarantees real-time transmission of shape deformations and ensures the smooth motions of 3D models. Moreover, we achieve very effective performance over real Internet conditions as well as a local LAN. Experimental results show that we get low distortion and small delays even when surface deformations of large and complicated 3D models are streamed over lossy networks. Ziying Tang, Guodong Rong 0001, Xiaohu Guo, B. Prabhakaran 0001 |
VR | 4 |
| 2010 | Dirichlet Harmonic Shape Compression with Feature Preservation for Parameterized SurfacesabstractAbstract With the rapid advancement of 3D scanning devices, large and complicated 3D shapes are becoming ubiquitous, and require large amount of resources to store and transmit them efficiently. This makes shape compression a demanding technique in order for the user to reduce the data transmission latency. Existing shape compression methods could achieve very low bit‐rates by sacrificing shape quality. But none of them guarantees the preservation of salient feature lines that users care. In addition, many 3D shapes come with parametric information for texture mapping purposes. In this paper we describe a spectral method to compress the geometric shapes equipped with arbitrary valid parametric information. It guarantees to preserve user‐specified feature lines while achieving a high compression ratio. By applying the spectral shape analysis – Dirichlet Manifold Harmonics, in the 2D parametric domain, this method provides a progressive compression mechanism to trade‐off between bit‐rate and shape quality. Experiments show that this method provides very low bit‐rate with high shape‐quality and still guarantees the preservation of user‐specified feature lines. Yang Liu 0013, B. Prabhakaran 0001, Xiaohu Guo |
Comput. Graph. Forum | 2 |
| 2010 | A body sensor network with electromyogram and inertial sensors: multimodal interpretation of muscular activitiesabstractThe evaluation of the postural control system (PCS) has applications in rehabilitation, sports medicine, gait analysis, fall detection, and diagnosis of many diseases associated with a reduction in balance ability. Standing involves significant muscle use to maintain balance, making standing balance a good indicator of the health of the PCS. Inertial sensor systems have been used to quantify standing balance by assessing displacement of the center of mass, resulting in several standardized measures. Electromyogram (EMG) sensors directly measure the muscle control signals. Despite strong evidence of the potential of muscle activity for balance evaluation, less study has been done on extracting unique features from EMG data that express balance abnormalities. In this paper, we present machine learning and statistical techniques to extract parameters from EMG sensors placed on the tibialis anterior and gastrocnemius muscles, which show a strong correlation to the standard parameters extracted from accelerometer data. This novel interpretation of the neuromuscular system provides a unique method of assessing human balance based on EMG signals. In order to verify the effectiveness of the introduced features in measuring postural sway, we conduct several classification tests that operate on the EMG features and predict significance of different balance measures. Hassan Ghasemzadeh 0001, Roozbeh Jafari, B. Prabhakaran 0001 |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2010 | Blind robust watermarking of 3d motion dataabstractThe article addresses the problem of copyright protection for 3D motion-captured data by designing a robust blind watermarking mechanism. The mechanism segments motion capture data and identifies clusters of 3D points per segment. A watermark can be embedded and extracted within these clusters by using a proposed extension of 3D quantization index modulation. The watermarking scheme is blind in nature and the encoded watermarks are shown to be imperceptible, and secure. The resulting hiding capacity has bounds based on cluster size. The watermarks are shown to be robust against attacks such as uniform affine transformations (scaling, rotation, and translation), cropping, reordering, and noise addition. The time complexity for watermark embedding and extraction is estimated as O(nlogn) and O(n2logn), respectively. Parag Agarwal, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2010 | On supporting reliable QoS in multi-hop multi-rate mobile ad hoc networks
Ming Li 0007, B. Prabhakaran 0001 |
Wirel. Networks | 2 |
| 2010 | Dynamic priority re-allocation scheme for quality of service in IEEE 802.11e wireless networks
Ming Li 0007, B. Prabhakaran 0001 |
Wirel. Networks | 3 |
| 2009 | Association rule mining in multiple, multidimensional time series medical dataabstractTime series pattern mining (TSPM) finds correlations or dependencies in same series or in multiple time series. When the numerous instances of multiple time series data are associated with different quantitative attributes, they form a multiple multi-dimensional framework. In this paper, we consider real-life time series data of muscular activities of human participants obtained from multiple Electromyogram (EMG) sensors and discover patterns in these EMG data streams. Each EMG data stream is associated with quantitative attributes such as energy of the signal and onset time which are required to be mined along with EMG time series patterns. We propose a two-stage approach for this purpose: in the first stage, our emphasis is on discovering frequent patterns in multiple time series by doing sequential mining across time slices. And in the next stage, we focus on the quantitative attributes of only those time series that are present in the patterns discovered in the first stage. Our evaluation with large sets of time series data from multiple EMG sensors demonstrate that our two-stage approach speeds up the process of finding association rules in such multidimensional environment as compared to other methods and scales up linearly in terms of number of time series involved. Our approach is generic and applicable to any multiple time series dataset format. Gaurav N. Pradhan, B. Prabhakaran 0001 |
ICME | 2 |
| 2009 | A Comprehensive Approach for Streaming 3D Progressive MeshesabstractFast and efficient streaming of detailed 3D model over lossy network has long been a challenge, although progressive compression techniques were proposed long time ago. One reason is that packet loss occurring in unreliable networks is highly unpredictable, and leads to connectivity inconsistency and distortions. In this paper, we address this problem by proposing a receiver-based loss tolerance scheme based on a prediction technique. Our method works without introducing protection bits and retransmission. We stream mesh refinement data on reliable and unreliable networks separately so as to reduce the transmission delay as well as to obtain a satisfactory decompression result. The tests indicate that the decompression is completed quickly, suggesting that it is a practical solution. Moreover, the proposed prediction technique achieves a good approximation of the original mesh with low distortion. Ziying Tang, Xiaohu Guo, B. Prabhakaran 0001 |
ISM | 3 |
| 2009 | Multimedia aspects in health careabstractRecently, Body Sensor Networks (BSNs) are being deployed for monitoring and managing medical conditions as well as human performance in sports. These BSNs include various sensors such as accelerometers, gyroscopes, EMG (Electromyogram), EKG (Electro-cardiograms), and other sensors depending on the needs of the medical conditions. Data from these sensors are typically Time Series data and the data from multiple sensors form multiple, multidimensional time series data. Duk-jin Kim, B. Prabhakaran 0001 |
ACM Multimedia | 2 |
| 2009 | On Supporting High-Quality 3D Geometry Multicasting over IEEE 802.11 Wireless NetworksabstractWith significant improvements in both wireless technologies and computational capabilities of mobile devices, it is now possible to exchange and render 3D graphics over wireless networks on mobile devices such as PDAs and laptops. In this paper, we consider a typical scenario where users holding mobile devices of different display resolutions and rendering capabilities request the same 3D object in an IEEE 802.11 wireless LAN. Several schemes are proposed. First, to support high-quality 3D content multicasting, we analyze the characteristics of 3D data and choose the minimum data set for unicast in order to avoid excessive bandwidth consumption. Then, a transcoding algorithm is proposed to address the issue of multiuser diversity. In addition, to take advantage of the nature of broadcast in wireless medium and further mitigate the issue of serious resource usage due to large data size, we propose to broadcast certain less important refinement data. With the proposed hybrid unicast/broadcast transmission scheme and a packet overhearing mechanism, good scalability can be achieved. Finally, we schedule the unicast according to users' experienced link condition to handle user mobility. Simulation results show that the proposed schemes in combination can efficiently achieve the dual objectives of low transmission delay and small distortion. Ming Li 0007, B. Prabhakaran 0001 |
IEEE Trans. Computers | 3 |
| 2009 | Robust blind watermarking of point-sampled geometryabstractDigital watermarking for copyright protection of 3-D meshes cannot be directly applied to point clouds, since we need to derive consistent connectivity information, which might change due to attacks, such as noise addition and cropping. Schemes for point clouds operate only on the geometric data and, hence, are generic and applicable to mesh-based representations of 3-D models. For building generic copyright schemes for 3-D models, this paper presents a robust blind watermarking mechanism for 3-D point-sampled geometry. The basic idea is to find a cluster tree from clusters of 3-D points. Using the cluster tree, watermarks can be embedded and extracted by deriving an order among points at global (intracluster) and local levels (intercluster). The multiple bit watermarks are encoded/decoded inside each cluster based on an extension of the cluster structure-based 3-D quantization index modulation. The encoding mechanism makes the technique robust against uniform affine transformations (rotation, scaling, and transformation), reordering, cropping, simplification, and noise addition attacks. The technique when applied to 3-D meshes also achieves robustness against retriangulation and progressive compression techniques. Customization of the bit-encoding scheme achieves high hiding capacity with embedding rates that are equal to 4 b/point, while maintaining the imperceptibility of the watermark with low distortions. The estimated time complexity isO(nlogn), wherenis the number of 3-D points. Parag Agarwal, B. Prabhakaran 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Indexing 3-D Human Motion Repositories for Content-Based RetrievalabstractContent-based retrieval of the similar motions for the human joints has significant impact in the fields of physical medicine, biomedicine, rehabilitation, and motion therapy. In this paper, we propose an efficient indexing approach for 3-D human motion capture data, supporting queries involving both subbody motions as well as whole-body motions. Gaurav N. Pradhan, B. Prabhakaran 0001 |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2008 | Fault Detection Framework for Video Surveillance SystemsabstractWe consider cameras whose outputs do not reflect true scenes as faulty cameras. To build a fault detection video surveillance system without using additional hardware devices, we use the video outputs from cameras to do self-checking. We study two categories of faults: spatial faults and temporal faults, and reduce the two sub-problems into graph theoretical problems on two graphs (surveillance sharing graph (SSG) and surveillance partitioning graph (SPG)). We prove a theoretical upper bound for the spatial fault detection, and develop two algorithms for detecting the two types of faults respectively.Then, we integrate both in a framework, which is capable of the following: given the outputs from a video surveillance system, it isolates cameras which are faulty or suspected to be faulty. It gives warnings with types of faults, locations, detection confidence. Our experiments confirm the effectiveness of the framework's methodologies. Junqiang Zhou, Simeon C. Ntafos, B. Prabhakaran 0001 |
AVSS | 3 |
| 2008 | QOAR: Adaptive QoS Scheme in Multi-Rate Wireless LANsabstractWith the availability of multiple rates in IEEE 802.11 a/b/g wireless LANs, it is desirable to improve the network capacity and temporal fairness by sending multiple consecutive frames over high rate links, as proposed in opportunistic auto rate (OAR [1]). However, the basic OAR does not provide quality of service (QoS) guarantee and thus is not sufficient in supporting real time voice/video traffic. In this paper, we further enhance the OAR protocol with a set of QoS mechanisms. The proposed QOAR protocol consists of two protocols: (i) a traffic- differentiating flow weight adaptation protocol (FWA) that dynamically tunes both contention window and concatenation number per channel access; (ii) an admission control protocol (AC) that guarantees the bandwidth/delay requirements of multimedia services. Extensive simulation studies show that QOAR enables QoS for real-time traffic yet maximizes performance of best effort traffic. Ming Li 0007, Yang Xiao 0001, Imrich Chlamtac, B. Prabhakaran 0001 |
ICC | 5 |
| 2008 | Loss tolerance scheme for 3D progressive meshes streaming over networksabstractNowadays, the Internet provides a convenient medium for sharing complex 3D models online. However, transmitting 3D progressive meshes over networks may encounter the problem of packets loss that can lead to connectivity inconsistency and distortion of the reconstructed meshes. In this paper, we combine reliable and unreliable channels to reduce both time delay and mesh distortion, and we propose an error-concealment scheme for tolerating packet loss when the meshes are transmitted over unreliable network channels. When the loss of connectivity data occurs, the decoder can predict the geometry data and mesh connectivity information, and construct an approximation of the original mesh. Therefore, the proposed error-concealment scheme can significantly reduce the data size required to be transmitted over reliable channels. The results show that both the computational cost of our error-concealment scheme and the distortion introduced by our scheme are small. Ziying Tang, Xiaohu Guo, B. Prabhakaran 0001 |
ICME | 4 |
| 2008 | Adaptive Frame Concatenation Mechanisms for QoS in Multi-Rate Wireless Ad Hoc NetworksabstractProviding quality of service (QoS) to users in a wireless ad-hoc network is a key concern for service providers. With the availability of multiple rates in IEEE 802.11a/b/g wireless LANs, it is desirable to improve the network capacity and temporal fairness by sending multiple consecutive frames (also referred as frame concatenation mechanism) over high rate links, as proposed in opportunistic auto rate (OAR). However, OAR does not consider the effect of frame sizes and may yield unsatisfactory performance for high priority multimedia flows transmitting over low rate links. Therefore, a more appropriate frame concatenation strategy and a corresponding service differentiation scheme should be devised to provide better performance for high priority voice/video flows than low priority data flows, under various channel rate scenarios. We first analyze the effect of frame size on the performance of OAR. Then, we propose a general concatenation mechanism (GCM), a more accurate frame concatenation mechanism for multi-rate MAC with better fairness. Finally, we propose two mechanisms: adaptive weighted fair frame concatenation mechanism (AWFCM) and adaptive QoS aware frame concatenation mechanism (AQCM), for supporting service differentiation and QoS in multi-rate wireless ad hoc networks. The primary idea is to adjust the number of concatenated frames based on flow weights/priorities, frame sizes, link rates, and network traffic. Simulation results show that the proposed mechanisms achieve desirable performance on supporting multimedia applications in multi-rate wireless ad-hoc networks. Ming Li 0007, Yang Xiao 0001, Imrich Chlamtac, B. Prabhakaran 0001 |
INFOCOM | 5 |
| 2008 | Storage, retrieval, and communication of body sensor network dataabstractRecently, Body Sensor Networks (BSNs) are being deployed for monitoring and managing medical conditions as well as human performance in sports. These BSNs include various sensors such as accelerometers, gyroscopes, EMG (Electromyogram), EKG (Electro-cardiograms), and other sensors depending on the needs of the medical conditions. Data from these sensors are typically Time Series data and the data from multiple sensors form multiple, multidimensional time series data. Analyzing data from such multiple medical sensors pose several challenges: different sensors have different characteristics, different people generate different patterns through these sensors, and even for the same person the data can vary widely depending on time and environment.This tutorial describes the technologies that go behind BSNs - both in terms of the hardware infrastructure as well as the basic software. First, we outline the BSN hardware features and the related requirements. We then discuss the energy and communication choices for BSNs. Next, we discuss approaches for classification, data mining, visualization, and securing these data. We also show several demonstrations of body sensor networks as well as the software that aid in analyzing the data. Gaurav N. Pradhan, B. Prabhakaran 0001 |
ACM Multimedia | 2 |
| 2008 | Semantic Quantization of 3D Human Motion Capture Data Through Spatial-Temporal Feature Extraction
Yohan Jin, B. Prabhakaran 0001 |
MMM | 2 |
| 2008 | Content Based Querying and Searching for 3D Human Motions
Manoj M. Pawar, Gaurav N. Pradhan, Kang Zhang 0001, B. Prabhakaran 0001 |
MMM | 4 |
| 2008 | Partial query resolution for animation authoringabstractAnimations are a part of multimedia and techniques such as motion mapping and inverse kinematics aid in reusing models and motion sequences to create new animations. This reuse approach is facilitated by the use of content-based retrieval techniques that often require fuzzy query resolution. Most fuzzy query resolution approaches work on all the attributes of the query to minimize the database access cost thus resulting in an unsatisfactory result set. It turns out that the query resolution can be carried out in a partial manner to achieve user satisfactory results and aid in easy authoring. In this article, we present two partial fuzzy query resolution approaches, one that results in high-quality animations and the other that produces results with decreasing number of satisfied conditions in the query. Phani S. Kotharu, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2008 | Minimizing probable collision pairs searched in interactive animation authoring
Parag Agarwal, Srinivas Rajagopalan, B. Prabhakaran 0001 |
Vis. Comput. | 3 |
| 2007 | On supporting high quality 3D geometry multicasting over IEEE 802.11 wireless networksabstractTraditionally, rendering and display of 3D graphics require powerful workstations with specialized 3D graphics adapters. However, with significant improvements achieved in both wireless technologies and computational capabilities of mobile devices, it is now possible to exchange and render 3D graphics over wireless networks on mobile devices such as PDA and Laptop. In this paper, we consider a typical scenario where mobile users with different display resolutions and rendering capabilities request the same 3D object in an infrastructure based IEEE 802.11 wireless LAN. We propose to enable packet overhearing and design a feedback based unicast scheduling according to users’ experienced signal strength and delay requirements to handle user mobility. Extensive simulation results show that the proposed schemes in combination can efficiently achieve the dual objectives of maintaining low transmission delays and small distortion. Ming Li 0007, B. Prabhakaran 0001 |
BROADNETS | 3 |
| 2007 | Data Hiding based Compression Mechanism for 3D ModelsabstractDifferent compression methods (Jingliang Peng et al., 2005) such as progressive meshes and single refinement mesh compression exist for mesh representation for 3D models and improving them is a challenge. This paper is a step in this direction, where we show that data hiding methods can be used to significantly improve the compression achieved for such methods. 3D meshes are made up of connectivity (edge set E) and geometric (vertices set V) information. The data hiding method should provide a very high embedding rate in order to achieve a desirable compression ratio. This can be modeled mathematically, given 'V points or vertices in 3D space; for a given compression ratio we use 'AT' points and hide rest of the connectivity and geometric information inside it. These points are termed as the encoding points and the encoded information is termed is the compressed information. Parag Agarwal, B. Prabhakaran 0001 |
DCC | 3 |
| 2007 | Tamper Proofing 3D Motion Data Streams
Parag Agarwal, B. Prabhakaran 0001 |
MMM (1) | 2 |
| 2007 | Hierarchical Indexing Structure for 3D Human Motions
Gaurav N. Pradhan, Chuanjun Li, B. Prabhakaran 0001 |
MMM (1) | 3 |
| 2007 | Animation toolkit based on a database approach for reusing motions and models
Akanksha, B. Prabhakaran 0001, Conrado R. Ruiz Jr. |
Multim. Tools Appl. | 3 |
| 2007 | Segmentation and recognition of motion capture data stream by classification
Chuanjun Li, Punit R. Kulkarni, B. Prabhakaran 0001 |
Multim. Tools Appl. | 3 |
| 2007 | Segmentation and recognition of motion streams by similarity searchabstractFast and accurate recognition of motion data streams from gesture sensing and motion capture devices has many applications and is the focus of this article. Based on the analysis of the geometric structures revealed by singular value decompositions (SVD) of motion data, a similarity measure is proposed for simultaneously segmenting and recognizing motion streams. A direction identification approach is explored to further differentiate motions with similar data geometric structures. Experiments show that the proposed similarity measure can segment and recognize motion streams of variable lengths with high accuracy, without knowing beforehand the number of motions in a stream. Chuanjun Li, Si-Qing Zheng, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2006 | Motion Stream Segmentation and Recognition by ClassificationabstractThis paper proposes a classification-based approach to segmenting and recognizing patterns in motion signals. Feature vectors are extracted based on singular value decomposition (SVD) for classification. Multi-class support vector machine (SVM) classifiers with class probability estimates are explored for segmenting and recognizing motion streams. Experiments show that the proposed approach can find patterns in the multi-attribute motion streams with high accuracy Chuanjun Li, Punit R. Kulkarni, B. Prabhakaran 0001 |
ICASSP (5) | 3 |
| 2006 | A Novel Indexing Approach for Efficient and Fast Similarity Search of Captured Motions
Chuanjun Li, B. Prabhakaran 0001 |
PAKDD | 2 |
| 2006 | Visual Querying on Human Motion for the DisabledabstractThe development of visual query languages can ease the retrieval of human motion data. This paper describes a process that allows users to specify queries for human motions describing disabilities, medical conditions, testing criteria, or other domain requirements as an encoding of grammar rules Kevin L. Ates, Kang Zhang 0001, B. Prabhakaran 0001 |
VL/HCC | 3 |
| 2006 | Real-time classification of variable length multi-attribute motions
Chuanjun Li, Latifur Khan, B. Prabhakaran 0001 |
Knowl. Inf. Syst. | 3 |
| 2006 | Middleware for streaming 3D progressive meshes over lossy networksabstractStreaming 3D graphics have been widely used in multimedia applications such as online gaming and virtual reality. However, a gap exists between the zero-loss-tolerance of the existing compression schemes and the lossy network transmissions. In this article, we propose a generic 3D middleware between the 3D application layer and the transport layer for the transmission of triangle-based progressively compressed 3D models. Significant features of the proposed middleware include. 1) handling 3D compressed data streams from multiple progressive compression techniques. 2) considering end user hardware capabilities for effectively saving the data size for network delivery. 3) a minimum cost dynamic reliable set selector to choose the transport protocol for each sublayer based on the real-time network traffic. Extensive simulations with TCP/UDP and SCTP show that the proposed 3D middleware can achieve the dual objectives of maintaining low transmission delay and small distortion, and thus supporting high quality 3D streaming with high flexibility. Ming Li 0007, B. Prabhakaran 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2006 | End-to-end QoS framework for heterogeneous wired-cum-wireless networks
Ming Li 0007, Imrich Chlamtac, B. Prabhakaran 0001 |
Wirel. Networks | 4 |
| 2005 | Similarity measure for multi-attribute data [haptic data recognition]abstractEfficient recognition of haptic data such as 3D motion capture data and sign language sensory data can have wide applications in the interactive computer animation and sign language automatic translation areas. For this purpose, we propose a similarity measure for multi-attribute haptic data, a new form of multimedia signal. The proposed similarity measure, based on singular value decomposition, captures the most important features of the signal data, allows for different signal generating rates and reasonable variations in similar signals. Experiments with real life and synthetic data demonstrate that the proposed similarity measure can capture the similarities of motions with different speeds and different lengths and can have up to 100% recognition rates. Chuanjun Li, B. Prabhakaran 0001, Si-Qing Zheng |
ICASSP (2) | 2 |
| 2005 | MAC Layer Admission Control and Priority Re-allocation for Handling QoS Guarantees in Non-cooperative Wireless LANs
Ming Li 0007, B. Prabhakaran 0001 |
Mob. Networks Appl. | 2 |
| 2005 | Flexible Strategies for Disk Scheduling in Multimedia Presentation Servers
Sindhu Emilda, Lillykutty Jacob, Ovidiu Daescu, B. Prabhakaran 0001 |
Multim. Tools Appl. | 4 |
| 2004 | Accessing Documents via Audio: An Extensible Transcoder for HTML to VoiceXML Conversion
Narayan Annamalai, Gopal Gupta 0001, B. Prabhakaran 0001 |
ICCHP | 3 |
| 2004 | Segmentation and recognition of multi-attribute motion sequencesabstractIn this work, we focus on fast and efficient recognition of motions in multi-attribute continuous motion sequences. 3D motion capture data, animation motion data, and sensor data from gesture sensing devices are examples of multi-attribute continuous motion sequences. These sequences have multiple attributes rather than only one attribute as time series data has. Motions can have different rates and durations, and the resulting data can thus have different lengths. Also, motion data can have noises due to transitions between successive motions. Hence, traditional distance measuring approaches used for time series data (such as Euclidean distances or dynamic time-warped distances) are not suitable for recognition in multi-attribute motion sequences. Hence, we have defined a similarity measure based on the analysis of singular value decomposition (SVD) properties of similar multi-attribute motions. A five-phase algorithm has then been proposed that gives good pruning power by exploiting the proximity of continuous motion data. We experimented this algorithm with data from different sources: 3D motion capture devices, animation motions, and CyberGlove gesture sensing device. These experiments show that our algorithm can segment and recognize long motion streams with high accuracy and in real time without knowing beforehand the number of motions in a stream. Chuanjun Li, Peng Zhai, Si-Qing Zheng, B. Prabhakaran 0001 |
ACM Multimedia | 4 |
| 2004 | End-to-End Framework for QoS Guarantee in Heterogeneous Wired-cum-Wireless NetworksabstractWith information access becoming more and more ubiquitous, there is a need for providing QoS support for communication that spans wired and wireless networks. For the wired side, RSVP/SBM has been widely accepted as a flow reservation scheme in IEEE 802 style LANs. In this paper, we investigate the integration of RSVP and a RSVP-like flow reservation scheme in wireless LANs, as an end-to-end solution for QoS guarantee in wired-cum-wireless networks. We propose WRESV, an RSVP-like flow reservation and admission control scheme for IEEE 802.11 wireless LAN. Using WRESV, wired/wireless integration can be easily implemented by cross-layer interaction at the access point. Main components of the integration are RSVP-WRESV parameter mapping, and the initiation of new reservation messages, depending on where senders/receivers are located. In addition, we also propose various optimizations for supporting multicast session, mobility management, and admission control. Ming Li 0007, Sathish Sathyamurthy, Imrich Chlamtac, B. Prabhakaran 0001 |
QSHINE | 5 |
| 2003 | A Framework for Reuse From Animation Multi-Databases
N. Chokkareddy, B. Prabhakaran 0001, M. Vattikuti |
MMM | 3 |
| 2003 | On flow reservation and admission control for distributed scheduling strategies in IEEE802.11 wireless LANabstractProviding service differentiation in IEEE802.11 Wireless LANs [3] has been investigated by many researchers ([2], [5], [6], [7], [9], [13]). It has been shown [1] that some distributed schedulers such as DFS [9] and EDCF [5] can achieve high throughput and certain service differentiation comparing to DCF and PCF provided that the traffic load in the system is low or medium. However, those strategies do not support flow reservation and thus cannot guarantee QoS requirements of high priority real-time flows under overloading network traffics. In this paper, we present a MAC layer flow reservation and admission control scheme for distributed scheduling strategies in the aim of achieving QoS guarantee in IEEE802.11 wireless LANs. Our approach has several desirable features: (1) It can work with most of the distributed scheduling strategies like DCF, DFS, EDCF without modification of the underlying scheduling mechanism. (2) A dynamic priority re-allocation method is integrated with the admission control to further improve system throughput. (3) Misuse of priority can be easily handled. Simulation of our proposed reservation scheme upon various distributed scheduling strategies has been conducted, and results show that this scheme can achieve low collision rate, high throughput, and less delay. Ming Li 0007, B. Prabhakaran 0001, Sathish Sathyamurthy |
MSWiM | 2 |
| 2003 | Visualizing Animation DatabasesabstractWe consider a repository of animation models and motions that can be reused to generate new animation sequences. For instance, a user can retrieve an animation of a dog kicking its leg (in air) and manipulate the result to generate a new animation where the dog is kicking a ball. In this particular example, inverse kinematics technique can be used to retarget the kicking motion of a dog to a ball. This approach of reusing models and motions to generate new animation sequences can be facilitated by operations such as querying of animation databases for required models and motions, and manipulation of the query results to meet new constraints. However, manipulation operations such as motion retargeting are quite complex in nature. Hence, there is a need for visualizing the queries on animation databases as well as the manipulation operations on the query results. In this paper, we propose a visually interactive method for reusing motions and models, by adjusting the query results from animation databases for new situations while at the same time, keeping the desired properties of the original models and motions. Here, a user first queries for animation objects, i.e., geometric models and motions. Then, the user interactively makes new animations by visually manipulating the query results. Depending on the orders in which the GUIs (Graphical User Interfaces) are invoked and the parameters are changed, the system automatically generates a sequence of operations, a list of SQL-like syntax commands, and applies it to the query results of motions and models. With the help of visualization tools, the user can view the changes before accepting them. Akanksha, B. Prabhakaran 0001, Conrado R. Ruiz Jr. |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2003 | Experiences with an object-level scalable web framework
B. Prabhakaran 0001, Yuguang Tu |
J. Netw. Comput. Appl. | 1 |
| 2003 | Unified Read Requests
Eenjun Hwang, B. Prabhakaran 0001 |
Multim. Tools Appl. | 2 |
| 2003 | Application-Layer Protocol for Collaborative Multimedia Presentations
Eenjun Hwang, B. Prabhakaran 0001 |
Multim. Tools Appl. | 2 |
| 2002 | MAC protocol enhancements for QoS guarantee and fairness over the IEEE 802.11 wireless LANsabstractIn future wireless networks, different traffic classes will exhibit a large variety of characteristics and QoS requirements, such as transmission rate, maximum tolerable bit error rate and timeout specifications. However, currently there is no standard way of guaranteeing QoS in wireless access networks like wireless LAN based on IEEE 802.11. In this paper, we propose medium access control protocol enhancements and a distributed scheduler for QoS guarantees and fairness over IEEE 802.11 WLAN. The proposed scheme uses distributed scheduling at both the access point and the user terminals to schedule the transmission of packets according to their delay requirements. The algorithms used for both scheduling are the same. A flexible and fair resource allocation method among the traffic classes and the user terminals is provided by this scheme. Its performance has been evaluated using the UCB/LBNL/VINT Network Simulator and an implementation in the Linux kernel. Qiu Qiang, Lillykutty Jacob, R. Radhakrishna Pillai, B. Prabhakaran 0001 |
ICCCN | 4 |
| 2002 | Presentation Planning for Distributed VoD SystemsabstractA distributed video-on-demand (VoD) system is one where a collection of video data is located at dispersed sites across a computer network. In a single site environment, a local video server retrieves video data from its local storage device. However, in distributed VoD systems, when a customer requests a movie from the local server, the server may need to interact with other servers located across the network. In this paper, we present different types of presentation plans that a local server can construct in order to satisfy a customer request. Informally speaking, a presentation plan is a temporally synchronized sequence of steps that the local server must perform in order to present the requested movie to the customer. This involves obtaining commitments from other video servers, obtaining commitments from the network service provider, as well as making commitments of local resources, while keeping within the limitations of available bandwidth, available buffer, and customer data consumption rates. Furthermore, in order to evaluate the quality of a presentation plan, we introduce two measures of optimality for presentation plans: minimizing wait time for a customer and minimizing access bandwidth which, informally speaking, specifies how much network/disk bandwidth is used. We develop algorithms to compute three different optimal presentation plans that work at a block level, or at a segment level, or with a hybrid mix of the two, and compare their performance through simulation experiments. We have also mathematically proven effects of increased buffer or bandwidth and data replications for presentation plans which had previously been verified experimentally in the literature. Eenjun Hwang, B. Prabhakaran 0001, V. S. Subrahmanian |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | An On-Line Repository for Embedded SoftwareabstractThe use of off-the-shelf components (COTS) can significantly reduce the time and cost of developing large-scale software systems. However, there are some difficult problems with the component-based approach. First, the developers have to be able to effectively retrieve components. This requires the developers to have an extensive knowledge of available components and how to retrieve them. After identifying the components, the developers also face a steep learning curve to master the use of these components. We are developing an On-line Repository for Embedded Software (ORES) to facilitate component management and retrieval. In this paper, we address the issues of designing software repository systems to assist users in obtaining appropriate components and learning to understand and use the components efficiently. We use an ontology to construct an abstract view of the organization of the components in ORES. The ontology structure facilitates repository browsing and effective search. We also develop a set of tools to assist with component comprehension, including a tutorial manager and a component explorer. I-Ling Yen, Latifur Khan, B. Prabhakaran 0001, Farokh B. Bastani, John Linn |
ICTAI | 3 |
| 2001 | Multimedia Information Delivery Over Wireless Channels
B. Prabhakaran 0001 |
Multim. Tools Appl. | 1 |
| 2000 | A forward error recovery technique for real-time MPEG-2 video transport and its performance over wireless IEEE 802.11 LANabstractReal-time MPEG-2 video transport applications do not usually have the luxury of a reverse channel for recovering from any errors that might occur during communication. Degradation in quality of decoded video frames is immediately apparent in the presence of errors in headers. In this paper, we focus on protecting header information by replicating it in any free space that might be available in the defined MPEG-2 transport stream packets. We also present our implementation experience over wired ATM as well as wireless IEEE 802.11 LAN by incorporating this forward error recovery approach with a real-time MPEG-2 encoder. In our experiments, it is found that the free space available is generally more than adequate for replicating essential header information. R. Radhakrishna Pillai, B. Prabhakaran 0001, Qiu Qiang |
ICCCN | 2 |
| 2000 | Retrieval Scheduling for Collaborative Multimedia Presentations
Ping Bai, B. Prabhakaran 0001, Aravind Srinivasan |
Multim. Syst. | 2 |
| 2000 | Multimedia authoring and presentation techniques - guest editor's introduction
B. Prabhakaran 0001 |
Multim. Syst. | 1 |
| 2000 | Guest Editor's Introduction: Multimedia Authoring and Presentation Strategies, Tools, and Experiences
B. Prabhakaran 0001 |
Multim. Tools Appl. | 1 |
| 2000 | Adaptive Multimedia Presentation Strategies
B. Prabhakaran 0001 |
Multim. Tools Appl. | 1 |
| 1999 | Application-layer broker for scalable Internet services with resource reservationabstractin both directions: however. the armroach does not verv Scalability is a very important issue in providing Internet services, especially in view of the explosive growth in the Ping Bai, B. Prabhakaran 0001, Aravind Srinivasan |
ACM Multimedia (2) | 2 |
| 1999 | A forward error recovery technique for MPEG-II video transportabstractArticle Free AccessA forward error recovery technique for MPEG-II video transport Share on Authors: R. Radhakrishna Pillai Kent Ridge Digital Labs, 21 Heng Mui Keng Terrace, Singapore 119613 Kent Ridge Digital Labs, 21 Heng Mui Keng Terrace, Singapore 119613View Profile , B. Prabhakaran School of Computing, National University of Singapore, Singapore 119260 School of Computing, National University of Singapore, Singapore 119260View Profile , Qui Qiang School of Computing, National University of Singapore, Singapore 119260 School of Computing, National University of Singapore, Singapore 119260View Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999 Pages 59–61https://doi.org/10.1145/319878.319897Online:01 October 1999Publication History 3citation288DownloadsMetricsTotal Citations3Total Downloads288Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF R. Radhakrishna Pillai, B. Prabhakaran 0001, Qui Qiang |
ACM Multimedia (2) | 2 |
| 1999 | Collaborative Multimedia Presentations in Mobile Environments
B. Prabhakaran 0001 |
Multim. Tools Appl. | 1 |
| 1999 | Guest Editors' Introduction
B. Prabhakaran 0001, Mohsen Kavehrad |
Multim. Tools Appl. | 1 |
| 1998 | Distributed Video PresentationsabstractConsiders a distributed video server environment where video movies need not be stored entirely in one server. Blocks of a video movie are be distributed and replicated over multiple video servers. Customers are served by one video server. This video server, termed the originating server, might have to interact with other servers for downloading missing blocks of the requested movie. We present three types of presentation plans that an originating server can possibly construct for satisfying a customer's request. A presentation plan can be considered as a detailed (temporally synchronized) sequence of steps carried out by the originating server for presenting the requested movie to the customer. The creation of presentation plans involves obtaining commitments from other video servers and the network service provider, as well as making local resource commitments, within the limitations of available bandwidth, available buffer and customer consumption rates. For evaluating the goodness of a presentation plan, we introduce two measures of optimality for presentation plans: minimizing the waiting time for a customer and minimizing the access bandwidth. We present algorithms for computing optimal presentation plans and compare their performance experimentally. We have also mathematically proved certain results for the presentation plans. Eenjun Hwang, V. S. Subrahmanian, B. Prabhakaran 0001 |
ICDE | 3 |
| 1998 | Collaborative multimedia documents: Authoring and presentationabstractMultimedia documents are composed of different data types such as video, audio, text, and images. Authoring a multimedia document is a creative exercise. Unlike traditional computer supported collaborative work where documents are composed of static objects, multimedia documents have temporal and spatial requirements that must be supported by any collaborative multimedia platform. In this paper, we show that most requirements (including temporal and spatial) for collaborative multimedia authoring systems can be expressed in terms of a highly structured class of linear constraints called prioritized difference constraints. Based on our prioritized difference constraint-based characterization, we develop efficient, incremental algorithms for creating and modifying multimedia documents so as to satisfy the required temporal and spatial constraints. We further develop methods to identify inconsistent requirements, and show how such inconsistencies may be removed through constraint relaxation techniques. We also report on the collaborative heterogeneous interactive multimedia platform (CHIMP) system developed using the framework described. © 1998 John Wiley & Sons, Inc. K. Selçuk Candan, B. Prabhakaran 0001, V. S. Subrahmanian |
Int. J. Intell. Syst. | 2 |
| 1998 | Retrieval Schedules Based on Resource Availability and Flexible Presentation Specifications
K. Selçuk Candan, B. Prabhakaran 0001, V. S. Subrahmanian |
Multim. Syst. | 2 |
| 1996 | CHIMP: A Framework for Supporting Distributed Multimedia Document Authoring and PresentationabstractA multimedia document consists of different media objects that are to be sequenced and presented according to temporal and spatial specifications. Collaborative authoring helps in simultaneous editing and viewing of a multimedia document by multiple authors. However, it may cause the objects composing a multimedia document to be distributed over a computer network. In this paper, we propose a framework for distributed multimedia document authoring and presentation. The salient features of this framework are: flexible temporal specification based on difference constraints, system and user defined access filters, local editing, format conversions of media objects, and flexible object retrieval schedules for handling variations in system parameters such as network throughput and buffer resources. We propose shortestpath based algorithms for solving difference constraints. We show how the proposed algorithms can handle local editing and access filtering of multimedia documents. We also describe how the difference constraints based temporal specifications can help in deriving a flexible object retrieval schedule. K. Selçuk Candan, B. Prabhakaran 0001, V. S. Subrahmanian |
ACM Multimedia | 2 |
| 1996 | Synchronization Representation and Traffic Source Modeling in Orchestrated PresentationabstractMultimedia applications comprise several media streams, which are semantically synchronized at different time instants. The application behavior is stored along with the multimedia database using representation mechanisms such as OCPN (object composition Petri nets) or dynamic timed Petri nets (DTPN). It is imperative that one translates the application behavior to the corresponding schedulable entities, such as packets, so that the performance engineering of any system can be done, using the traffic model arising out of the (media related) application behavior as opposed to individual media level behavior. This requires that a function be defined, which takes the stored temporal representation as input and produces packets as output, preserving the semantic relationships among the streams. The authors propose a methodology based on probabilistic, attributed context free grammar (PACFG) to address this issue. They demonstrate the appropriateness of this methodology by applying it to the OCPN/DTPN representation of a typical multimedia application vis-a-vis orchestrated presentation. B. Prabhakaran 0001, Satish K. Tripathi |
IEEE J. Sel. Areas Commun. | 2 |
| 1994 | Synchronization Models for Multimedia Presentation with User Participation
B. Prabhakaran 0001 |
Multim. Syst. | 1 |
| 1993 | Synchronization Models for Multimedia Presentation with User ParticipationabstractThis paper addresses the key issue of providingflexible multimedia presentation with user participation and suggests synchronization models that can specify the user participation during the presentation. We study models like the Petrinet-based hypertext model and the object composition Petri nets (OCPN). We suggest adynamic timed Petri nets structure that can model pre-emptions and modifications to the temporal characteristics of the net. This structure can be adopted by the OCPN to facilitate modeling of multimedia synchronization characteristics with dynamic user participation. We show that the suggested enhancements for the dynamic timed Petri nets satisfy all the properties of the Petri net theory. We use the suggested enhancements to model typical scenarios in a multimedia presentation with user inputs. B. Prabhakaran 0001 |
ACM Multimedia | 1 |