Ivan Marsic

dblp:42/767 · DBLP profile ↗
← Back
92ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-1033-6865ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 21 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 since 2021Computer networks · 19Artificial intelligence and machine learning · 16 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2026 Multi-view real-time detection of patient arrival in trauma resuscitation using computer vision
abstract
Every minute delay in life-saving intervention increases mortality risk in injured patients. Given this relationship, quality measurement of trauma resuscitation includes the timing of provider decisions and interventions. The current approach of manual recording of patient arrival and other timestamps may be inaccurate due to providers underestimating elapsed time during resuscitation. We introduce a frame-based computer vision system to automatically detect and classify two specific phases of patient arrival: (1) when the patient enters the trauma resuscitation room, and (2) when the patient is moved onto the bed. The proposed system consists of two stages. The first stage uses a pattern-based method to detect when a patient enters the room. The second stage uses both side-view and top-view video feeds to determine when the patient has been moved to the bed, addressing occlusion issues and enhancing robustness and precision. To minimize labeling effort and accelerate detection, we introduce PA-YOLO, a frame-based classification model, in both stages of the system. We evaluated our system using 5-fold cross-validation on 60 trauma resuscitation cases. Our results show that this approach achieves an accuracy of 0.92 ± 0.01 for patient-at-door detection and 0.93 ± 0.03 for patient-on-bed detection, with an average detection delay of 2.15 ± 1.24 s. The proposed system outperforms SlowFast by up to 5 percentage points and our previous I3D-based system by up to 7 percentage points in detection accuracy, while reducing detection delay from over 5 s to about 2 s. Compared to baseline model YOLO11, our PA-YOLO improves detection performance by 2.0% while reducing floating point operations per second (FLOPS) by 51.5%. • Use of computer vision methods to detect different phases of patient arrival. • Modification of the YOLO model for fast scene classification and reconstruction of activities from the video stream. • Mitigation of occlusion problems by switching between different views based on occlusion detection. • Evaluation of the effectiveness of the system in actual resuscitations and comparison with existing methods.
Sifan Yuan, Mary S. Kim, Aaron H. Mun, Rebecca Cunningham, Ivan Marsic, Randall S. Burd
Comput. Vis. Image Underst.5
2026 SAFE: A Smart Adherence Detection Framework for Monitoring Personal Protective Equipment in Healthcare Settings
abstract
Personal protective equipment (PPE) is critical for infection control in healthcare, protecting workers and patients from infection risks. The COVID-19 pandemic further highlighted the importance of correct PPE use, yet adherence to U.S. Centers for Disease Control and Prevention guidelines remains inconsistent. Continuous human monitoring of PPE adherence is impractical because it is labor-intensive and may expose observers to infection risk. Automated monitoring is a promising alternative, but reliable PPE assessment in clinical videos remains difficult due to occlusion and subtle differences between adherence levels. To address these challenges, we propose SAFE - Smart Adherence detection Framework for PPE, a cascaded computer vision system for real-time monitoring of PPE wearing status, including complete, incomplete, and absent cases, with a focus on gowns and masks. SAFE uses a two-stage design: Stage 1 detects gown status and localizes head regions, and Stage 2 classifies mask status from head crops. We evaluate SAFE on R2PPE, a ceiling-view trauma-room simulation dataset with dense PPE annotations and complex scenes. SAFE improves overall average precision from 0.48 to 0.67 and increases mask-class average precision by 0.33 compared to a baseline one-stage detector. We further validate SAFE across modern detector backbones, including transformer-based detectors, and on real-case trauma-room data using class-level and alarm-level criteria, improving class-level mask accuracy from 0.59 to 0.65 while maintaining a high alarm-level recall of 0.98. SAFE could enhance PPE monitoring with minimal human intervention, providing a scalable solution for improving infection control in healthcare settings.
Wanzhao Yang, Beomseok Park, Mary S. Kim, Aleksandra Sarcevic, Syed Muhammad Anwar, Marius G. Linguraru, Ivan Marsic, Randall S. Burd
IEEE J. Biomed. Health Informatics7
2025 Understanding Personal Protective Equipment Use in Interdisciplinary Medical Settings: Design Explorations for Just-in-Time Compliance Alerts: Improving PPE Practices in Medical Settings Through Alert Design
abstract
We examine the use of personal protective equipment (PPE) in two interdisciplinary medical settings to inform the design of just-in-time alerts and reminders for correcting PPE noncompliance. We reviewed videos of 26 pediatric resuscitations occurring over the course of the COVID-19 pandemic at an urban pediatric teaching hospital. Through video review, we identified causes for PPE noncompliance, activities that were frequently performed without PPE, instances in which PPE was intentionally removed, and mechanisms by which healthcare providers corrected PPE noncompliance. We also interviewed 18 registered nurses working in the hospital's emergency department (ED) and intensive care unit (ICU) to better understand observed PPE behaviors and practices. Our results suggest that alert design will require considering the urgency of correcting PPE noncompliance against the urgency of tasks being performed. We discuss our findings through the lens of the COM-B framework and conclude by exploring design opportunities for just-in-time alerts and reminders for prompting PPE noncompliance corrections in dynamic medical work.
Aleksandra Sarcevic, Eleanor Wood, Katherine Ann Zellner, Christine Dodeye Ikponmwonba, Mary S. Kim, Ivan Marsic, Randall S. Burd
Conference on Designing Interactive Systems6
2025 Single-View Cotraining and Active Learning for Concurrent Medical Activity Labeling
abstract
Labeling medical activities from visual datasets requires extensive time and effort for practitioners. Deep learning methods have been widely used to train partially labeled datasets for predicting activity labels. Cotraining, a semi-supervised learning technique that utilizes predictions from two different models, is used in conjunction with active learning to mitigate sampling bias. There is a need to optimize manual labeling efforts with active learning while improving the prediction performance. In this paper, we developed the SCALE-TAG methodology, based on single-view cotraining and active learning, to label concurrent activities in video records cost-effectively. ScALE-TAG utilized the cosine similarity metric to enhance diversity within the subsets and selected informative samples based on prediction uncertainty for retraining the models. We evaluated the performance of scale-tag for six trauma resuscitation activities. During the training stage, the average labeling time required for selected samples was 10.93 % of the manual annotation time for all unlabeled samples. We then applied scale-tag to a test set. The F1 scores for test samples were over 0.7 for four activities, where the labeling time required for ambiguous predictions of test samples was 8 % for one of them, 17 % for another, and around 20 % for the other two. The improvement in prediction accuracy and decrease in labeling time we obtained with ScAleTAG are promising for developing deep learning-based methods to support healthcare providers with effective activity labeling.
Aydin Saribudak, Aaron H. Mun, Ivan Marsic
BIBE3
2025 SemiVisBooster: Boosting Semi-Supervised Learning for Fine-Grained Classification through Pseudo-Label Semantic Guidance
Chenyang Gao, Ivan Marsic
ICCV4
2025 ASELMAR: Active and semi-supervised learning-based framework to reduce multi-labeling efforts for activity recognition
abstract
Manual annotation of unlabeled data for model training is expensive and time-consuming, especially for visual datasets requiring domain-specific experience for multi-labeling, such as video records generated in hospital settings. There is a need to build frameworks to reduce human labeling efforts while improving training performance. Semi-supervised learning is widely used to generate predictions for unlabeled samples in a partially labeled datasets. Active learning can be used with semi-supervised learning to annotate unlabeled samples to reduce the sampling bias due to the label predictions. We developed the aselmar framework based on active and semi-supervised learning techniques to reduce the time and effort associated with multi-labeling of unlabeled samples for activity recognition. aselmar (i) categorizes the predictions for unlabeled data based on the confidence level in predictions using fixed and adaptive threshold settings, (ii) applies a label verification procedure for the samples with the ambiguous prediction, and (iii) retrains the model iteratively using samples with their high-confidence predictions or manual annotations. We also designed a software tool to guide domain experts in verifying ambiguous predictions. We applied aselmar to recognize eight selected activities from our trauma resuscitation video dataset and evaluated their performance based on the label verification time and the mean ap score metric. The label verification required by aselmar was 12.1% of the manual annotation effort for the unlabeled video records. The improvement in the mean ap score was 5.7% for the first iteration and 8.3% for the second iteration with the fixed threshold-based method compared to the baseline model . The p-values were below 0.05 for the target activities. Using an adaptive-threshold method, aselmar achieved a decrease in ap score deviation, implying an improvement in model robustness. For a speech-based case study , the word error rate decreased by 6.2%, and the average transcription factor increased 2.6 times, supporting the broad applicability of ASELMAR in reducing labeling efforts from domain experts.
Aydin Saribudak, Sifan Yuan, Chenyang Gao, Waverly Gestrich-Thompson, Zachary P. Milestone, Randall S. Burd, Ivan Marsic
Comput. Vis. Image Underst.7
2025 Comparative analysis of personal protective equipment nonadherence detection: computer vision versus human observers
abstract
OBJECTIVES: Human monitoring of personal protective equipment (PPE) adherence among healthcare providers has several limitations, including the need for additional personnel during staff shortages and decreased vigilance during prolonged tasks. To address these challenges, we developed an automated computer vision system for monitoring PPE adherence in healthcare settings. We assessed the system performance against human observers detecting nonadherence in a video surveillance experiment. MATERIALS AND METHODS: The automated system was trained to detect 15 classes of eyewear, masks, gloves, and gowns using an object detector and tracker. To assess how the system performs compared to human observers in detecting nonadherence, we designed a video surveillance experiment under 2 conditions: variations in video durations (20, 40, and 60 seconds) and the number of individuals in the videos (3 versus 6). Twelve nurses participated as human observers. Performance was assessed based on the number of detections of nonadherence. RESULTS: Human observers detected fewer instances of nonadherence than the system (parameter estimate -0.3, 95% CI -0.4 to -0.2, P < .001). Human observers detected more nonadherence during longer video durations (parameter estimate 0.7, 95% CI 0.4-1.0, P < .001). The system achieved a sensitivity of 0.86, specificity of 1, and Matthew's correlation coefficient of 0.82 for detecting PPE nonadherence. DISCUSSION: An automated system simultaneously tracks multiple objects and individuals. The system performance is also independent of observation duration, an improvement over human monitoring. CONCLUSION: The automated system presents a potential solution for scalable monitoring of hospital-wide infection control practices and improving PPE usage in healthcare settings.
Mary S. Kim, Beomseok Park, Genevieve J. Sippel, Aaron H. Mun, Wanzhao Yang, Kathleen H. McCarthy, Emely Fernandez, Marius George Linguraru, Aleksandra Sarcevic, Ivan Marsic, Randall S. Burd
J. Am. Medical Informatics Assoc.10
2025 Human intention recognition for trauma resuscitation: An interpretable deep learning approach for medical process data
Mary S. Kim, Sen Yang 0002, Genevieve J. Sippel, Aleksandra Sarcevic, Randall S. Burd, Ivan Marsic
J. Biomed. Informatics8
2024 ProcessGAN: Generating Privacy-Preserving Time-Aware Process Data with Conditional Generative Adversarial Nets
abstract
Process data constructed from event logs provides valuable insights into procedural dynamics over time. The confidential information in process data, together with the data's intricate nature, makes the datasets not sharable and challenging to collect. Consequently, research is limited using process data and analytics in the process mining domain. In this study, we introduced a synthetic process data generation task to address the limitation of sharable process data. We introduced a generative adversarial network, called ProcessGAN, to generate process data with activity sequences and corresponding timestamps. ProcessGAN consists of a transformer-based network as the generator, and a time-aware self-attention network as the discriminator. It can generate privacy-preserving process data from random noise. ProcessGAN considers the duration of the process and time intervals between activities to generate realistic activity sequences with timestamps. We evaluated ProcessGAN on five real-world datasets, two that are public and three collected in medical domains that are private. To evaluate the synthetic data, in addition to statistical metrics, we trained a supervised model to score the synthetic processes. We also used process mining to discover workflows for synthetic medical processes and had domain experts evaluate the clinical applicability of the synthetic workflows. ProcessGAN outperformed the existing generative models in generating complex processes with valid parallel pathways. The synthetic process data generated by ProcessGAN better represented the long-range dependencies between activities, a feature relevant to complicated medical and other processes. The timestamps generated by the ProcessGAN model showed similar distributions with the authentic timestamps. In addition, we trained a transformer-based network to generate synthetic contexts (e.g., patient demographics) that were associated with the synthetic processes. The synthetic contexts generated by our model outperformed the baseline models, with the distributions similar to the authentic contexts. We conclude that ProcessGAN can generate sharable synthetic process data indistinguishable from authentic data. Our source code is available in https://github.com/raaachli/ProcessGAN.
Sen Yang 0002, Travis M. Sullivan, Randall S. Burd, Ivan Marsic
ACM Trans. Knowl. Discov. Data5
2023 Supporting Awareness of Dynamic Data: Approaches to Designing and Capturing Data within Interactive Clinical Checklists
abstract
Automatically integrating data within interactive clinical checklists allows for enhanced dynamic displays, while also providing information needed for checklist adaptation to the context of the medical event. In this mixed-methods study, we used user-centered design sessions with clinicians to design a checklist interface that automatically captures and displays dynamic patient data. We compared the manual and automatic checklist versions during video-guided simulation sessions, evaluating the effects of automatic capture on clinicians' interactions with dynamic data and their situation awareness. Despite clinicians' concerns that automatic data capture would affect situation awareness, we found no significant difference in awareness scores. Participants preferred the automatic version, highlighting its improved accuracy and completeness. From our findings, we propose a framework for capturing dynamic data and designing dynamic data interfaces within interactive checklists. We conclude by discussing barriers and design opportunities for supporting awareness of data trends through checklists.
Angela Mastrianni, Aleksandra Sarcevic, Megan A. Krentsa, Travis M. Sullivan, Issa Zakeri, Ivan Marsic, Randall S. Burd
Conference on Designing Interactive Systems7
2023 Improving Label Assignments Learning by Dynamic Sample Dropout Combined with Layer-wise Optimization in Speech Separation
abstract
In supervised speech separation, permutation invariant training (PIT) is widely used to handle label ambiguity by selecting the best permutation to update the model. Despite its success, previous studies showed that PIT is plagued by excessive label assignment switching in adjacent epochs, impeding the model to learn better label assignments. To address this issue, we propose a novel training strategy, dynamic sample dropout (DSD), which considers previous best label assignments and evaluation metrics to exclude the samples that may negatively impact the learned label assignments during training. Additionally, we include layer-wise optimization (LO) to improve the performance by solving layer-decoupling. Our experiments showed that combining DSD and LO outperforms the baseline and solves excessive label assignment switching and layer-decoupling issues. The proposed DSD and LO approach is easy to implement, requires no extra training sets or steps, and shows generality to various speech separation tasks.
Chenyang Gao, Ivan Marsic
INTERSPEECH3
2023 Discovering interpretable medical process models: A case study in trauma resuscitation
Ivan Marsic, Aleksandra Sarcevic, Sen Yang 0002, Travis M. Sullivan, Peyton E. Tempel, Zachary P. Milestone, Karen J. O'Connell, Randall S. Burd
J. Biomed. Informatics2
2022 An Analysis of Speech during Life Saving Interventions to Inform the Design of a Computerized System for Delay Detection
Katherine Ann Zellner, Louis Jiorgio Villegas, Charles Neff, Waverly Gestrich-Thompson, Randall S. Burd, Ivan Marsic, Aleksandra Sarcevic
AMIA6
2022 TubeR: Tubelet Transformer for Video Action Detection
abstract
We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an offline actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21. Code will be available on GluonCV(https://cv.gluon.ai/).
Jiaojiao Zhao, Yanyi Zhang, Xinyu Li 0003, Hao Chen 0024, Bing Shuai, Chunhui Liu 0002, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, Ivan Marsic, Cees Snoek, Joseph Tighe
CVPR11
2022 A Speech-Based Model for Tracking the Progression of Activities in Extreme Action Teamwork
abstract
Designing computerized approaches to support complex teamwork requires an understanding of how activity-related information is relayed among team members. In this paper, we focus on verbal communication and describe a speech-based model that we developed for tracking activity progression during time-critical teamwork. We situated our study in the emergency medical domain of trauma resuscitation and transcribed speech from 104 audio recordings of actual resuscitations. Using the transcripts, we first studied the nature of speech during 34 clinically relevant activities. From this analysis, we identified 11 communicative events across three different stages of activity performance-before, during, and after. For each activity, we created sequential ordering of the communicative events using the concept of narrative schemas. The final speech-based model emerged by extracting and aggregating generalized aspects of the 34 schemas. We evaluated the model performance by using 17 new transcripts and found that the model reliably recognized an activity stage in 98% of activity-related conversation instances. We conclude by discussing these results, their implications for designing computerized approaches that support complex teamwork, and their generalizability to other safety-critical domains.
Swathi Jagannath, Neha Kamireddi, Katherine Ann Zellner, Randall S. Burd, Ivan Marsic, Aleksandra Sarcevic
Proc. ACM Hum. Comput. Interact.5
2021 Designing Interactive Alerts to Improve Recognition of Critical Events in Medical Emergencies
abstract
Vital sign values during medical emergencies can help clinicians recognize and treat patients with life-threatening injuries. Identifying abnormal vital signs, however, is frequently delayed and the values may not be documented at all. In this mixed-methods study, we designed and evaluated a two-phased visual alert approach for a digital checklist in trauma resuscitation that informs users about undocumented vital signs. Using an interrupted time series analysis, we compared documentation in the periods before (two years) and after (four months) the introduction of the alerts. We found that introducing alerts led to an increase in documentation throughout the post-intervention period, with clinicians documenting vital signs earlier. Interviews with users and video review of cases showed that alerts were ineffective when clinicians engaged less with the checklist or set the checklist down to perform another activity. From these findings, we discuss approaches to designing alerts for dynamic team-based settings.
Angela Mastrianni, Aleksandra Sarcevic, Lauren Chung, Issa Zakeri, Emily Alberto, Zachary P. Milestone, Ivan Marsic, Randall S. Burd
Conference on Designing Interactive Systems7
2021 Multi-Label Activity Recognition Using Activity-Specific Features and Activity Correlations
abstract
Multi-label activity recognition is designed for recognizing multiple activities that are performed simultaneously or sequentially in each video. Most recent activity recognition networks focus on single-activities, that assume only one activity in each video. These networks extract shared features for all the activities, which are not designed for multi-label activities. We introduce an approach to multi-label activity recognition that extracts independent feature descriptors for each activity and learns activity correlations. This structure can be trained end-to-end and plugged into any existing network structures for video classification. Our method outperformed state-of-the-art approaches on four multi-label activity recognition datasets. To better understand the activity-specific features that the system generated, we visualized these activity-specific features in the Charades dataset. The code will be released later.
Yanyi Zhang, Xinyu Li 0003, Ivan Marsic
CVPR3
2021 VidTr: Video Transformer Without Convolutions
abstract
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatiotemporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels, but with heavy memory usage. We then present VidTr which reduces the memory cost by 3.3× while keeping the same performance. To further optimize the model, we propose the standard deviation based topK pooling for attention (pooltopK_std), which reduces the computation by dropping non-informative features along temporal dimension. VidTr achieves state-of-the-art performance on five commonly used datasets with lower computational requirement, showing both the efficiency and effectiveness of our design. Finally, error analysis and visualization show that VidTr is especially good at predicting actions that require long-term temporal reasoning.
Yanyi Zhang, Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Biagio Brattoli, Hao Chen 0024, Ivan Marsic, Joseph Tighe
ICCV8
2021 Real-time medical phase recognition using long-term video understanding and progress gate method
Yanyi Zhang, Ivan Marsic, Randall S. Burd
Medical Image Anal.2
2020 Residual Recurrent Neural Network for Speech Enhancement
abstract
Most current speech enhancement models use spectrogram features that require an expensive transformation and result in phase information loss. Previous work has overcome these issues by using convolutional networks to learn the temporal correlations across high-resolution waveforms. These models, however, are limited by memory-intensive dilated convolution and aliasing artifacts from upsampling. We introduce an end-to-end fully recurrent neural network for single-channel speech enhancement. The network structured as an hourglass-shape that can efficiently capture long-range temporal dependencies by reducing the features resolution without information loss. Also, we use residual connections to prevent gradient decay over layers and improve the model generalization. Experimental results show that our model outperforms state-of-the-art approaches in six quantitative evaluation metrics.
Jalal Abdulbaqi, Shuhong Chen, Ivan Marsic
ICASSP4
2019 Mutual Correlation Attentive Factors in Dyadic Fusion Networks for Speech Emotion Recognition
abstract
Emotion recognition in dyadic communication is challenging because: 1. Extracting informative modality-specific representations requires disparate feature extractor designs due to the heterogenous input data formats. 2. How to effectively and efficiently fuse unimodal features and learn associations between dyadic utterances are critical to the model generalization in actual scenario. 3. Disagreeing annotations prevent previous approaches from precisely predicting emotions in context. To address the above issues, we propose an efficient dyadic fusion network that only relies on an attention mechanism to select representative vectors, fuse modality-specific features, and learn the sequence information. Our approach has three distinct characteristics: 1. Instead of using a recurrent neural network to extract temporal associations as in most previous research, we introduce multiple sub-view attention layers to compute the relevant dependencies among sequential utterances; this significantly improves model efficiency. 2. To improve fusion performance, we design a learnable mutual correlation factor inside each attention layer to compute associations across different modalities. 3. To overcome the label disagreement issue, we embed the labels from all annotators into a k-dimensional vector and transform the categorical problem into a regression problem; this method provides more accurate annotation information and fully uses the entire dataset. We evaluate the proposed model on two published multimodal emotion recognition datasets: IEMOCAP and MELD. Our model significantly outperforms previous state-of-the-art research by 3.8%-7.5% accuracy, using a more efficient model.
Xinyu Lyu, Weijia Sun, Weitian Li, Shuhong Chen, Xinyu Li 0003, Ivan Marsic
ACM Multimedia7
2018 Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment
abstract
Multimodal affective computing, learning to recognize and interpret human affect and subjective information from multiple data sources, is still challenging because:(i) it is hard to extract informative features to represent human affects from heterogeneous inputs; (ii) current fusion strategies only fuse different modalities at abstract levels, ignoring time-dependent interactions between modalities. Addressing such issues, we introduce a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data. Our introduced model outperforms state-of-the-art approaches on published datasets, and we demonstrate that our model's synchronized attention over modalities offers visual interpretability.
Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li 0003, Ivan Marsic
ACL (1)6
2018 Hybrid Attention based Multimodal Network for Spoken Language Classification
abstract
We examine the utility of linguistic content and vocal characteristics for multimodal deep learning in human spoken language understanding. We present a deep multimodal network with both feature attention and modality attention to classify utterance-level speech data. The proposed hybrid attention architecture helps the system focus on learning informative representations for both modality-specific feature extraction and model fusion. The experimental results show that our system achieves state-of-the-art or competitive results on three published multimodal datasets. We also demonstrated the effectiveness and generalization of our system on a medical speech dataset from an actual trauma scenario. Furthermore, we provided a detailed comparison and analysis of traditional approaches and deep learning methods on both feature extraction and fusion.
Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li 0003, Ivan Marsic
COLING6
2018 Deep Mul Timodal Learning for Emotion Recognition in Spoken Language
abstract
In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text and audio via a hybrid deep multimodal structure, which considers the spatial information from text, temporal information from audio, and high-level associations from low-level handcrafted features. Second, we fuse all features by using a three-layer deep neural network to learn the correlations across modalities and train the feature extraction and fusion modules together, allowing optimal global fine-tuning of the entire structure. We evaluated the proposed framework on the IEMOCAP dataset. Our result shows promising performance, achieving 60.4% in weighted accuracy for five emotion categories.
Shuhong Chen, Ivan Marsic
ICASSP3
2018 Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder
abstract
Human conversation analysis is challenging because the meaning can be expressed through words, intonation, or even body language and facial expression. We introduce a hierarchical encoder-decoder structure with attention mechanism for conversation analysis. The hierarchical encoder learns word-level features from video, audio, and text data that are then formulated into conversation-level features. The corresponding hierarchical decoder is able to predict different attributes at given time instances. To integrate multiple sensory inputs, we introduce a novel fusion strategy with modality attention. We evaluated our system on published emotion recognition, sentiment analysis, and speaker trait analysis datasets. Our system outperformed previous state-of-the-art approaches in both classification and regressions tasks on three datasets. We also outperformed previous approaches in generalization tests on two commonly used datasets. We achieved comparable performance in predicting co-existing labels using the proposed model instead of multiple individual models. In addition, the easily-visualized modality and temporal attention demonstrated that the proposed attention mechanism helps feature selection and improves model interpretability.
Xinyu Li 0003, Kaixiang Huang, Shiyu Fu, Kangning Yang, Shuhong Chen, Moliang Zhou, Ivan Marsic
ACM Multimedia8
2018 An approach to automatic process deviation detection in a time-critical clinical process
Sen Yang 0002, Aleksandra Sarcevic, Richard A. Farneth, Shuhong Chen, Omar Z. Ahmed, Ivan Marsic, Randall S. Burd
J. Biomed. Informatics6
2017 Exploring Design Opportunities for a Context-Adaptive Medical Checklist Through Technology Probe Approach
abstract
This paper explores the workflow and use of an interactive medical checklist for trauma resuscitation-an emerging technology developed for trauma team leaders to support decision making and task coordination among team members. We used a technology probe approach and ethnographic methods, including video review, interviews, and content analysis of checklist logs, to examine how team leaders use the checklist probe during live resuscitations. We found that team leaders of various experience levels use the technology differently. Some leaders frequently glance at the checklist and take notes during task performance, while others place the checklist on a stand and only interact with the checklist when checking items. We compared checklist timestamps to task activities and found that most items are checked off after tasks are performed. We conclude by discussing design implications and new design opportunities for a future dynamic, adaptive checklist.
Leah Kulp, Aleksandra Sarcevic, Richard A. Farneth, Omar Z. Ahmed, Dung Mai, Ivan Marsic, Randall S. Burd
Conference on Designing Interactive Systems6
2017 3D activity localization with multiple sensors: poster abstract
abstract
We present a deep learning framework for fast 3D activity localization and tracking in a dynamic and crowded real world setting. Our training approach reverses the traditional activity localization approach, which first estimates the possible location of activities and then predicts their occurrence. Instead, we first trained a deep convolutional neural network for activity recognition using depth video and RFID data as input, and then used the activation maps of the network to locate the recognized activity in the 3D space. Our system achieved around 20cm average localization error (in a 4m × 5m room) which is comparable to Kinect's body skeleton tracking error (10--20cm), but our system tracks activities instead of Kinect's location of people.
Xinyu Li 0003, Yanyi Zhang, Shuhong Chen, Richard A. Farneth, Ivan Marsic, Randall S. Burd
IPSN7
2017 CAR - a deep learning structure for concurrent activity recognition: poster abstract
abstract
We introduce the Concurrent Activity Recognizer (CAR) - an efficient deep learning structure that recognizes complex concurrent teamwork activities from multimodal data. We implemented the system in a challenging medical setting, where it recognizes 35 different activities using Kinect depth video and data from passive RFID tags on 25 types of medical objects. Our preliminary results showed our system achieved an 84% average accuracy with 0.20 F1-Score.
Yanyi Zhang, Xinyu Li 0003, Shuhong Chen, Moliang Zhou, Richard A. Farneth, Ivan Marsic, Randall S. Burd
IPSN7
2017 A Data-driven Process Recommender Framework
abstract
We present an approach for improving the performance of complex knowledge-based processes by providing data-driven step-by-step recommendations. Our framework uses the associations between similar historic process performances and contextual information to determine the prototypical way of enacting the process. We introduce a novel similarity metric for grouping traces into clusters that incorporates temporal information about activity performance and handles concurrent activities. Our data-driven recommender system selects the appropriate prototype performance of the process based on user-provided context attributes. Our approach for determining the prototypes discovers the commonly performed activities and their temporal relationships. We tested our system on data from three real-world medical processes and achieved recommendation accuracy up to an F1 score of 0.77 (compared to an F1 score of 0.37 using ZeroR) with 63.2% of recommended enactments being within the first five neighbors of the actual historic enactments in a set of 87 cases. Our framework works as an interactive visual analytic tool for process mining. This work shows the feasibility of data-driven decision support system for complex knowledge-based processes.
Sen Yang 0002, Xin Dong 0010, Leilei Sun, Richard A. Farneth, Hui Xiong 0001, Randall S. Burd, Ivan Marsic
KDD8
2017 Region-based Activity Recognition Using Conditional GAN
abstract
We present a method for activity recognition that first estimates the activity performer's location and uses it with input data for activity recognition. Existing approaches directly take video frames or entire video for feature extraction and recognition, and treat the classifier as a black box. Our method first locates the activities in each input video frame by generating an activity mask using a conditional generative adversarial network (cGAN). The generated mask is appended to color channels of input images and fed into a VGG-LSTM network for activity recognition. To test our system, we produced two datasets with manually created masks, one containing Olympic sports activities and the other containing trauma resuscitation activities. Our system makes activity prediction for each video frame and achieves performance comparable to the state-of-the-art systems while simultaneously outlining the location of the activity. We show how the generated masks facilitate the learning of features that are representative of the activity rather than accidental surrounding information.
Xinyu Li 0003, Yanyi Zhang, Yueyang Chen, Huangcan Li, Ivan Marsic, Randall S. Burd
ACM Multimedia6
2016 Checklist as a Memory Externalization Tool during a Critical Care Process
Aleksandra Sarcevic, Zhan Zhang 0008, Ivan Marsic, Randall S. Burd, Leah Kulp
AMIA3
2016 Privacy Preserving Dynamic Room Layout Mapping
Xinyu Li 0003, Yanyi Zhang, Ivan Marsic, Randall S. Burd
ICISP3
2016 Deep Learning for RFID-Based Activity Recognition
abstract
We present a system for activity recognition from passive RFID data using a deep convolutional neural network. We directly feed the RFID data into a deep convolutional neural network for activity recognition instead of selecting features and using a cascade structure that first detects object use from RFID data followed by predicting the activity. Because our system treats activity recognition as a multi-class classification problem, it is scalable for applications with large number of activity classes. We tested our system using RFID data collected in a trauma room, including 14 hours of RFID data from 16 actual trauma resuscitations. Our system outperformed existing systems developed for activity recognition and achieved similar performance with process-phase detection as systems that require wearable sensors or manually-generated input. We also analyzed the strengths and limitations of our current deep learning architecture for activity recognition from RFID data.
Xinyu Li 0003, Yanyi Zhang, Ivan Marsic, Aleksandra Sarcevic, Randall S. Burd
SenSys3
2016 Enhancing semi-supervised learning through label-aware base kernels
Qiaojun Wang, Kai Zhang 0001, Zhengzhang Chen, Dequan Wang, Guofei Jiang, Ivan Marsic
Neurocomputing6
2016 Passive RFID for Object and Use Detection during Trauma Resuscitation
abstract
We evaluated passive radio-frequency identification (RFID) technology for detecting the use of objects and related activities during trauma resuscitation. Our system consists of RFID tags and antennas, optimally placed for object detection, as well as algorithms for processing RFID data to infer object use. To evaluate our approach, we tagged 81 objects in the resuscitation room and recorded RFID signal strength during 32 simulated resuscitations performed by trauma teams. We then analyzed RFID data to identify cues for recognizing resuscitation activities. Using these cues, we extracted descriptive features and applied machine-learning techniques to monitor interactions with objects. Our results show that an instance of a used object can be detected with accuracy rates greater than 90 percent in a crowded and fast-paced medical setting using off-the-shelf RFID equipment, and the time and duration of use can be identified with up to 83 percent accuracy. We conclude with insights into the limitations of passive RFID and areas in which RFID needs to be complemented with other sensing technologies.
Siddika Parlak, Ivan Marsic, Aleksandra Sarcevic, Waheed U. Bajwa, Lauren J. Waterhouse, Randall S. Burd
IEEE Trans. Mob. Comput.2
2015 From Categorical to Numerical: Multiple Transitive Distance Learning and Embedding
abstract
Categorical data are ubiquitous in real-world databases. However, due to the lack of an intrinsic proximity measure, many powerful algorithms for numerical data analysis may not work well on their categorical counterparts, making it a bottleneck in practical applications. In this paper, we propose a novel method to transform categorical data to numerical representations, so that abundant numerical learning methods can be exploited in categorical data mining. Our key idea is to learn a pairwise dissimilarity among categorical symbols, henceforth a continuous embedding, which can then be used for subsequent numerical treatment. There are two important criteria for learning the dissimilarities. First, it should capture the important “transitivity” which has shown to be particularly useful in measuring the proximity relation in categorical data. Second, the pairwise sample geometry arising from the learned symbol distances should be maximally consistent with prior knowledge (e.g., class labels) to obtain a good generalization performance. We achieve them through multiple transitive distance learning and embedding. Encouraging results are observed on a number of benchmark classification tasks against state-of-the-art.
Kai Zhang 0001, Qiaojun Wang, Zhengzhang Chen, Ivan Marsic, Vipin Kumar 0001, Guofei Jiang, Jie Zhang 0012
SDM4
2014 Improving Semi-Supervised Target Alignment via Label-Aware Base Kernels
Qiaojun Wang, Kai Zhang 0001, Guofei Jiang, Ivan Marsic
AAAI4
2014 Balancing design tensions: iterative display design to support ad hoc and multidisciplinary medical teamwork
abstract
In this paper, we describe how we developed an information display prototype for trauma resuscitation teams based on design ideas and feedback from clinicians. Our approach is grounded in participatory design, emphasizing the importance of gaining long-term commitment from clinicians in system development. Through a series of participatory design workshops, heuristic evaluation, and simulated resuscitation sessions, we identified the main information features to include on our display. Our results focus on how we balanced the design tensions that emerged when addressing the ad hoc, hierarchical, and multidisciplinary nature of trauma teamwork. We discuss the implications of balancing role-based differences for each information feature, as well as two major design tensions: process-based vs. state-based designs and role-based vs. team-based displays.
Diana S. Kusunoki, Aleksandra Sarcevic, Nadir Weibel, Ivan Marsic, Zhan Zhang 0008, Genevieve Tuveson, Randall S. Burd
CHI4
2014 Sparse semi-supervised learning on low-rank kernel
Kai Zhang 0001, Qiaojun Wang, Liang Lan, Yu Sun 0076, Ivan Marsic
Neurocomputing5
2014 Design and Evaluation of RFID Deployments in a Trauma Resuscitation Bay
abstract
We examined configuring a radio frequency identification (RFID) equipment for the best object use detection in a trauma bay. Unlike prior work on RFID, we 1) optimized the accuracy of object use detection rather than just object detection; and 2) quantitatively assessed antenna placement while addressing issues specific to tag placement likely to occur in a trauma bay. Our design started with an analysis of the environment requirements and constraints. We designed several antenna setups with different number of components (RFID tags or antennas) and their orientations. Setups were evaluated under scenarios simulating a dynamic medical setting. We used three metrics with increasing complexity and bias: read rate, received signal strength indication distribution distance, and target application performance. Our experiments showed that antennas above the regions with high object density are most suitable for detecting object use. We explored tagging strategies for challenging objects so that sufficient readout rates are obtained for computing evaluation metrics. Among the metrics, distribution distance was correlated with target application performance, and also less biased and simpler to calculate, which made it an excellent metric for context-aware applications. We present experimental results obtained in the real trauma bay to validate our findings.
Siddika Parlak, Shriniwas Ayyer, Ying Yu Liu, Ivan Marsic
IEEE J. Biomed. Health Informatics4
2013 Covariate Shift in Hilbert Space: A Solution via Sorrogate Kernels
abstract
Covariate shift is a unconventional learning scenario in which training and testing data have different distributions. A general principle to solve the problem is to make the training data distribution similar to the test one, such that classifiers computed on the former generalizes well to the latter. Current approaches typically target on the sample distribution in the input space, however, for kernel-based learning methods, the algorithm performance depends directly on the geometry of the kernel-induced feature space. Motivated by this, we propose to match data distributions in the Hilbert space, which, given a pre-defined empirical kernel map, can be formulated as aligning kernel matrices across domains. In particular, to evaluate similarity of kernel matrices defined on arbitrarily different samples, the novel concept of surrogate kernel is introduced based on the Mercer's theorem. Our approach caters the model adaptation specifically to kernel-based learning mechanism, and demonstrates promising results on several real-world applications.
Kai Zhang 0001, Vincent Wenchen Zheng, Qiaojun Wang, James T. Kwok, Qiang Yang 0001, Ivan Marsic
ICML (3)6
2012 Introducing RFID technology in dynamic and time-critical medical settings: Requirements and challenges
Siddika Parlak, Aleksandra Sarcevic, Ivan Marsic, Randall S. Burd
J. Biomed. Informatics3
2012 Teamwork Errors in Trauma Resuscitation
abstract
Human errors in trauma resuscitation can have cascading effects leading to poor patient outcomes. To determine the nature of teamwork errors, we conducted an observational study in a trauma center over a two-year period. While eventually successful in treating the patients, trauma teams had problems tracking and integrating information in a longitudinal trajectory, which resulted in inefficiencies and near-miss errors. As an initial step in system design to support trauma teams, we proposed a model of teamwork and a novel classification of team errors. Four types of team errors emerged from our analysis: communication errors, vigilance errors, interpretation errors, and management errors. Based on these findings, we identified key information structures to support team cognition and decision making. We believe that displaying these information structures will support distributed cognition of trauma teams. Our findings have broader applicability to other collaborative and dynamic work settings that are prone to human error.
Aleksandra Sarcevic, Ivan Marsic, Randall S. Burd
ACM Trans. Comput. Hum. Interact.2
2010 Monitoring Interactions with RFID Tagged Objects Using RSSI
Siddika Parlak, Ivan Marsic
MobiQuitous2
2010 Performance analysis of the IEEE 802.11 DCF in the presence of the hidden stations
Fu-Yi Hung, Ivan Marsic
Comput. Networks2
2010 MAC-layer proactive mixing for network coding in multi-hop wireless networks
Jian Zhang 0066, Yuanzhu Peter Chen, Ivan Marsic
Comput. Networks3
2009 Improved delayed ACK for TCP over multi-hop wireless networks
abstract
TCP performance in contention-based multi-hop wireless networks is shaped by two main factors, which are unlike the wired network case. First, the maximum throughput for a given topology and flow pattern is reached for a specific congestion window. This window is very difficult to detect under dynamically changing network traffic. Second, the excessive control traffic consumes channel bandwidth more severely than in the wired case. Our analysis and simulations show that the smaller TCP ACK packets consume channel resource comparable to the much longer TCP DATA packets, over high-speed connections. Motivated by this observation, we propose a new approach to improve TCP performance by further lowering the number of control packets compared with the known methods. Extensive simulations show that our strategy improves the TCP throughput up to 205% compared with the regular TCP. Although our simulation is based on 802.11, the same idea works in other networks using a contention-based MAC design.
Beizhong Chen, Ivan Marsic, Huai-Rong Shao, Ray Miller
WCNC2
2008 Transactive memory in trauma resuscitation
abstract
This paper describes an ethnographic study conducted to explore the possibilities for future design and development of technological support for trauma teams. We videotaped 10 trauma resuscitations and transcribed each event. Using a framework that we developed, we coded each transcript to allow qualitative and quantitative analysis of the trauma teams' collaborative processes. We analyzed teams' tasks, interactions, and communication patterns that support information acquisition and sharing. Our results showed the importance of team transactive memory, but also pointed to inefficiencies in communication processes, which enable the functioning of this collective memory system. Based on quantitative and qualitative observations of trauma teamwork, we present opportunities for technological solutions that may reduce the cognitive effort needed for maintaining the working memory of trauma teams.
Aleksandra Sarcevic, Ivan Marsic, Michael E. Lesk, Randall S. Burd
CSCW2
2008 Improving TCP Performance Over Multi-Hop Wireless Networks
abstract
In multi-hop wireless networks, unlike the wired ones, two main factors affect the TCP performance. First, the maximum throughput is reached for a specific congestion window that is very difficult to detect under dynamically changing network traffic. Second, the excessive control traffic consumes channel bandwidth more severely than in the wired case. Our analysis and simulations show that, over high-speed connections, the much shorter ACKtcppackets consume the channel capacity comparable to the much longer data packets. Motivated by this insight, we first propose using different routes for data and ACK flows in a grid sensor network which achieves 60~100% performance gain. After that, we reformulate and improve an existing idea for TCP performance improvement by further lowering the number of control packets compared to the known methods. Extensive simulations show that this improves the TCP throughput up to 200% compared to the regular TCP, in long-hop wireless networks.
Beizhong Chen, Ivan Marsic, Ray Miller
VTC Fall2
2008 Network Coding via Opportunistic Forwarding in Wireless Mesh Networks
abstract
Network coding has been used to increase transportation capabilities in wireless mesh networks. In mesh networks, the coding opportunities depend on the co-location of multiple traffic flows. With fixed routes given by a routing protocol, the coding opportunities are limited. This paper presents a new protocol called BEND, which combines the features of network coding and opportunistic forwarding in 802.11-based mesh networks to create more coding opportunities in the network. Taking advantage of redundancy of packets among the forwarder candidates, our protocol bends the routes locally and dynamically to attain better coding opportunities. This higher coding gain is verified using a network simulator.
Jian Zhang 0066, Yuanzhu Peter Chen, Ivan Marsic
WCNC3
2008 Launching an E-Learning System in a School - Cross-European e-/m-Learning System UNITE: A Case Study
Maja Cukusic, Andrina Granic, Ivan Marsic
WEBIST (1)3
2007 SYNG: A middleware for statefull groupware in mobile environments
abstract
Computer supported collaboration systems, or groupware, are being used more and more in the real life. In the recent years, we are witnessing an increasing demand for supporting such systems in mobile environments. In this paper we address the following questions: ldquoHow are the limited and highly dynamic resources of mobile clients, like network connection, energy supply or display size, influencing the design and deployment of groupware systems?rdquo and ldquoWhat type of policie the users need to specify in order to be able to collaborate in such environments?rdquo. We present a middleware system, called SYNG, that allows a statefull model, suitable for mobile environments. In our system, each user can define a state composed of a number of variables. Examples of such variables include battery usage, quality of network connection, display size, etc. Based on this state, the participant in a collaborative session specifies its own policy for receiving messages from the other participants. Experimental results show good performance and scalability of our approach.
Mihail Ionescu, Ivan Marsic
CollaborateCom2
2007 Access Delay Analysis of IEEE 802.11 DCF in the Presence of Hidden Stations
abstract
In this paper, we present an analytical model to evaluate the hidden station effect on the access delay of the IEEE 802.11 distributed coordination function (DCF) in both non-saturation and saturation condition. DCF is a random channel-access scheme based on carrier sense multiple access with collision avoidance (CSMA/CA) method and the exponential backoff procedure to reduce packet collisions. However, hidden stations still cause many collisions under CSMA/CA method because stations cannot sense each other's transmission and often send packets concurrently, resulting in significant performance degradation. Prior research has built accurate access delay model for 802.11 DCF. However, the hidden station effect on the performance has not been adequately studied. Our model generalizes the existing work on access delay modeling of 802.11 DCF for both non-saturation and saturation conditions, under the hidden-station effect. The performance of our model is evaluated by comparison with ns-2 simulations and they are found to agree.
Fu-Yi Hung, Ivan Marsic
GLOBECOM2
2007 Persistent Pseudo-Clearance Problem in IEEE802.11 Mesh Networks and its Multicast Based Solutions
abstract
Wireless mesh networks are flexible solutions to extend services from wireless LANs. The current IEEE 802.11 Specification, however, needs to be modified in various ways to be a suitable technology for this purpose. In particular, in order to handle the well-known hidden node problem (HNP), the Specification adopts MACAW by employing an RTS/CTS/DATA/ACK 4-way handshake. Some flaws of this scheme have been noticed, e.g. the Masked Node Problem (MNP). In this work, we identify a critical problem of the Specification's 4-way handshake, called persistent pseudo-clearance (PPC). PPC occurs when for two sender/receiver pairs a CTS from one pair's receiver collides with the DATA frames of the other pair. This logjam can persist for a period of time despite of the random backoff the senders employ. The persistent frame losses in PPC can cause more serious problems. The effect of giving up a frame transfer after reaching the maximum number of retries can propagate to upper layers, causing routing errors or TCP sender backoff. Multicast RTS (or MRTS) provides a good solution framework to break the cycle of losses and retransmissions between such peers. With minimal modification to MRTS, we provide an effective and efficient solution to PPC. Our experiments show that MRTS breaks the logjam of PPC while fully utilizing the network capacity.
Jian Zhang 0066, Yuanzhu Peter Chen, Ivan Marsic
LANMAN3
2007 Analysis of Non-Saturation and Saturation Performance of IEEE 802.11 DCF in the Presence of Hidden Stations
abstract
In this paper, we propose an analytical model to evaluate the hidden station effect on both non-saturation and saturation performance of the IEEE 802.11 Distributed Coordination Function (DCF). DCF is a random channel-access scheme based on Carrier Sense Multiple Access with Collision Avoidance (CSMACA) method and the binary slotted exponential backoff procedure to reduce the packet collision. Hidden stations cause most collisions because stations cannot sense each other's transmission and often send packets concurrently, resulting in significant degradation of the network performance. The proposed model generalizes the existing work on 802.11 DCF performance modeling for both non-saturation and saturation conditions, under the hidden-station effect. The performance of our model is evaluated by comparison with NS-2 simulations and found to agree with the analytic model.
Fu-Yi Hung, Ivan Marsic
VTC Fall2
2007 Effectiveness of Physical and Virtual Carrier Sensing in IEEE 802.11 Wireless Ad Hoc Networks
abstract
IEEE 802.11 defines physical and virtual carrier sensing mechanisms to avoid interference in wireless local area networks for the kind of interference originating from within the receiving range of a receiver. However, in wireless ad hoc networks most interference comes from outside of this range. So, the effectiveness of IEEE 802.11 carrier sensing mechanism in ad hoc networks has attracted many studies. Prior research has attempted to evaluate effectiveness from a spatial viewpoint only, using an analytical model to estimate the size of the interference area of an ongoing communication based on the transmitter-receiver distance. Unlike this, the temporal effectiveness of the carrier sensing mechanism has been ignored. This paper proposed an analysis combining spatial and temporal viewpoint to study the effect of interference on the performance of IEEE 802.11 protocol in ad hoc networks. The authors also compare the effectiveness of physical and virtual carrier sensing mechanisms, known as RTS/CTS mechanism, in wireless ad hoc networks.
Fu-Yi Hung, Ivan Marsic
WCNC2
2006 Link Quality and Signal-to-Noise Ratio in 802.11 WLAN with Fading: A Time-Series Analysis
abstract
It is known that for multipath fading channels individual points or average signal-to-noise ratio (SNR) alone do not adequately describe the wireless channel quality. This paper uses time-series modeling to investigate the relationship between SNR and 802.11 link bandwidth which we use to define the link quality. Two models, one linear and another nonlinear, are constructed and fitted to time-series of SNR as input and link bandwidth as output. Their performance is measured in terms of the accuracy of prediction of the link bandwidth. By the linear Auto- Regressive Moving Average exogenous variables (ARMAX) model, we show the existence of nonlinearity in the input data series, which results in high prediction error. The prediction performance is significantly improved by the nonlinear Echo State Network (ESN) model due to its ability of nonlinear re-expression of the SNR time-series and associating them with the correct link bandwidth.
Jian Zhang 0066, Ivan Marsic
VTC Fall2
2005 An optimization approach to group coupling in heterogeneous collaborative systems
abstract
Recent proliferation of computing devices has brought attention to heterogeneous collaborative systems, where key challenges arise from the resource limitations and disparities. Sharing data across disparate devices makes it necessary to employ mechanisms for adapting the original data and presenting it to the user in the best possible way. However, this could represent a major problem for effective collaboration, since users may find it difficult to reach consensus with everyone working with individually tailored data. This paper presents a novel approach to controlling the coupling of heterogeneous collaborative systems by combining concepts from complex systems and data adaptation techniques. The key idea is that data must be adapted to each individual's preferences and resource capabilities. To support and promote collaboration this adaptation must be interdependent, and adaptation performed by one individual should influence the adaptation of the others. These influences are defined according to the user's roles and collaboration requirements. We model the problem as a distributed optimization problem, so that the most useful data--both for the individual and the group as a whole--is scheduled for each user, while satisfying their preferences, their resource limitations, and their mutual influences. We show how this approach can be applied in a collaborative 3D design application and how it can be extended to other applications.
Carlos D. Correa, Ivan Marsic
GROUP2
2004 A Simplification Architecture for Exploring Navigation Tradeoffs in Mobile VR
Carlos D. Correa, Ivan Marsic
VR2
2004 Software Framework for Managing Heterogeneity in Mobile Collaborative Systems
Carlos D. Correa, Ivan Marsic
Comput. Support. Cooperative Work.2
2004 Hierarchical Routing Overhead in Mobile Ad Hoc Networks
abstract
Hierarchical techniques have long been known to afford scalability in networks. By summarizing topology detail via a hierarchical map of the network topology, network nodes are able to conserve memory and link resources. Extensive analysis of the memory requirements of hierarchical routing was undertaken in the 1970s. However, there has been little published work that assesses analytically the communication overhead incurred in hierarchical routing. This paper assesses the scalability, with respect to increasing node count, of hierarchical routing in mobile ad hoc networks (MANETs). The performance metric of interest is the number of control packet transmissions per second per node (/spl Phi/). To derive an expression for /spl Phi/, the components of hierarchical routing that incur overhead as a result of hierarchical cluster formation and location management are identified. It is shown here that /spl Phi/ is only polylogarithmic in the node count.
John Sucec, Ivan Marsic
IEEE Trans. Mob. Comput.2
2004 "Who's in charge here?" communicating across unequal computer platforms
abstract
People use personal data assistants in the field to collect data and to communicate with others both in the field and office. The individual in the office invariably has a laptop or a high-end personal workstation and thus, significantly more computing power, more screen real estate, and higher volume input devices, such as a mouse and keyboard. These differences give the high-end user the ability to represent and manipulate collaborative tasks more effectively. It is therefore useful to know what impact these differences have on work performance and work communications. Four different platform combinations involving a PC and a PDA were used to examine the effect of communicating via heterogeneous computer platforms. The PC platform used a mouse, a keyboard, and a 3-dimensional screen display. The PDA platform used a stylus, soft buttons, and a 2-dimensional screen display. A variation of the Tetris wall-building game called Slow Tetris was used as the subjects' collaborative task. A second factor in the experiment was role asymmetry. One subject was arbitrarily put in charge of the task solution in all of the combinations. An analysis of the solution times found that subjects with mixed platforms worked slower than their homogeneous counterparts, that is, a person in charge with a PC worked faster if his partner had a PC. An in-depth analysis of the communication patterns found significant differences in the exchanges between heterogeneous and homogenous combinations. The PC-to-PDA combination (with the person on the PC in charge of the solution) took significantly more time than the PC-to-PC combination. This extra time appears to come from the disadvantage of having a partner on the PDA who is unable to help in solving the problems. The PDA-to-PC combination took approximately the same amount of time as the PDA-to-PDA combination despite having one team member with a better representation. This member was, unfortunately, not in charge of the solution. The PDA-to-PC heterogeneous combination exhibited more direction giving, less one-sided collaboration, and more takeover attempts than any of the other combinations. Overall, roles were maintained in the partnerships except for the person with the PDA directing the person with the PC.
Maria C. Velez, Marilyn Tremaine, Aleksandra Sarcevic, Bogdan Dorohonceanu, Allan Meng Krebs, Ivan Marsic
ACM Trans. Comput. Hum. Interact.6
2003 Software framework for managing heterogeneity in mobile collaborative systems
abstract
Heterogeneity aspects in mobile collaborative systems, such as differences in user's interest, semantic conflicts across different domains and representations, and disparate device capabilities, cause difficulties in developing software applications. One of the key problems for collaborative applications is maintaining a consistent shared state.In this paper, we describe a framework that manages several aspects of heterogeneity to maintain consistency across the collaborating sites. We assume graph data structure for application state representation. Our framework is based on structural and semantic mappings between graph structures. The mapping can be customized to meet different requirements through user-defined policies and rules. An important constraint is efficient use of scarce system resources.We describe several applications built using the framework to collaboratively share XML documents. The XML documents in our case are 2D/3D representations of virtual worlds. We also show the performance results of our framework which demonstrate its feasibility for mobile scenarios.
Carlos D. Correa, Ivan Marsic
GROUP2
2003 A framework for rapid development of multimodal interfaces
abstract
Despite the availability of multimodal devices, there are very few commercial multimodal applications available. One reason for this may be the lack of a framework to support development of multimodal applications in reasonable time and with limited resources. This paper describes a multimodal framework enabling rapid development of applications using a variety of modalities and methods for ambiguity resolution, featuring a novel approach to multimodal fusion. An example application is studied that was created using the framework.
Frans Flippo, Allan Meng Krebs, Ivan Marsic
ICMI3
2003 Publish-Subscribe for Mobile Environments
Mihail F. Ionescu, Ivan Marsic
ICWE2
2003 Modeling and prediction of session throughput of constant bit rate streams in wireless data networks
abstract
This paper presents an approach to modeling and prediction of session throughput of constant bit rate streams in wireless data networks. A stable traffic generator is used to generate smooth data streams that are transmitted across various types of wireless connections in real-world wireless data networks, including wireless LANs and wireless cellular WANs. The throughput values of the data streaming sessions are recorded. Based on the analysis of statistical properties of the collected data, linear time series analysis is used to models and predict the session throughput. Autoregressive (AR) models are selected from a number of linear time series models since they can be fit to data in deterministic amount of time. The performance of AR models for prediction is compared to simpler models for prediction is compared to simpler models such as MEAN and window mean (WM) models, and our study shows that successful models, such as AR and WM models, have similar performance in predicting the session throughput of wireless data networks. The main contribution of our research is that by statistical study, it shows that session throughputs in wireless data networks can be modeled and predicted to a useful degree from past values by using linear time series analysis such as AR and WM models.
Liang Cheng 0001, Ivan Marsic
WCNC2
2003 Capacity compatible 2-level link state routing for ad hoc networks with mobile clusterheads
abstract
The throughput of a mobile ad hoc network (MANET) is determined by the transceiver link capacity available at each node and the type of traffic pattern that is prevalent in the network. In order for a routing protocol to be scalable, its control overhead must not exceed transceiver link capacity. To achieve capacity compatible routing, hierarchical techniques may be employed. This paper describes how link state routing, with a single layer of hierarchy, provides sufficient scalability for MANETs where the traffic pattern consists of unicast communication between arbitrary pairs of nodes.
John Sucec, Ivan Marsic
WCNC2
2003 Tree-Based Concurrency Control in Distributed Groupware
Mihail F. Ionescu, Ivan Marsic
Comput. Support. Cooperative Work.2
2003 Flexible User Interfaces for Group Collaboration
abstract
Flexible user interfaces that can be customized to meet the needs of the task at hand are particularly important for telecollaboration. This article presents the design and implementation of a user interface for DISCIPLE, a platform-independent telecollaboration framework. DISCIPLE supports sharing of Java components that are imported into the shared workspace at run-time and can be interconnected into more complex components. As a result, run-time interconnection of various components allows user tailoring of the human-computer interface. Software architecture for customization of both a group-level and application-level interfaces is presented, with interface components that are loadable on demand. The architecture integrates the sensory modalities of speech, sight, and touch. Instead of imposing one "right" solution onto users, the framework lets users tailor the user interface that best suits their needs. Finally, laboratory experience with DISCIPLE tested on a variety of applications with the framework is discussed along with future research directions.
Ivan Marsic, Bogdan Dorohonceanu
Int. J. Hum. Comput. Interact.1
2003 A Query Scope Agent for Flood Search Routing Protocols
John Sucec, Ivan Marsic
Wirel. Networks2
2002 Clustering Overhead for Hierarchical Routing in Mobile Ad hoc Networks
abstract
Numerous clustering algorithms have been proposed that can support routing in mobile ad hoc networks (MANET). However, there is very little formal analysis that considers the communication overhead incurred by these procedures. Further, there is no published investigation of the overhead associated with the recursive application of clustering algorithms to support hierarchical routing. This paper provides a theoretical upper bound on the communication overhead incurred by a particular clustering algorithm for hierarchical routing in MANET. It is demonstrated that, given reasonable assumptions, the average clustering overhead generated per node per second is only polylogarithmic in the node count. To derive this result, novel techniques to assess cluster maintenance overhead are employed.
John Sucec, Ivan Marsic
INFOCOM2
2002 Handling Heterogeneity in Networked Virtual Environments
abstract
The availability of inexpensive and powerful graphics cards as well as fast Internet connections make Networked Virtual Environments viable for millions of users and many new applications. It is therefore necessary to cope with the growing heterogeneity that arises from differences in computing power, network speed and users' preferences. This paper describes an architecture that accommodates the heterogeneity mentioned above while allowing a manager to define system-wide policies. Policies and users' preferences can be expressed as simple linear equations forming a mathematical model that describes the system as a whole as well as its individual components. When solutions to this model are mapped back to the problem domain, viable solutions that accommodate heterogeneity and system policies are obtained. The results of our experiments with a proof-of-concept system are described.
Helmuth Trefftz, Ivan Marsic, Michael Zyda
VR2
2002 Location management for hierarchically organized mobile ad hoc networks
abstract
A geography-based grid location service (GLS), proposed elsewhere, has resulted in a scalable location management service for mobile ad hoc networks (MANETs) where packet forwarding decisions are based on geographic position. A similarly scalable location management strategy has been devised for MANETs that employ hierarchical link state routing. Both approaches employ hierarchical principles to facilitate scalability. However, currently proposed approaches for hierarchical link state routing rely on a designated subset of nodes for location management. Such nodes represent potential sites of hot spot contention. In this paper, it is proposed that by applying the distributed database selection technique of GLS, a hierarchical location management scheme may be realized for MANETs based on link state routing that equitably distributes location server functionality among network nodes. Second, it is shown that location registration overhead per node for hierarchical location management is only logarithmic in the node count.
John Sucec, Ivan Marsic
WCNC2
2002 Accurate bandwidth measurement in xDSL service networks
Liang Cheng 0001, Ivan Marsic
Comput. Commun.2
2002 Piecewise Network Awareness Service for Wireless/Mobile Pervasive Computing
Liang Cheng 0001, Ivan Marsic
Mob. Networks Appl.2
2001 Charging for QoS in internetworks
abstract
Pricing is an effective regulatory tool to provide proper incentives so that users' self-interest will lead them to modify their usage according to their needs. This leads to better overall network utilization and enhanced users' satisfaction. In this work, a scalable pricing framework for QoS capable networks supporting real time, adjustable real time, and non-real time traffic is studied. The scheme, which belongs to usage-based methods, is independent of the underlying network and the mechanisms for QoS provisioning. The framework is credit-based ensuring the fairness, comprehensibility, and predictability of usage cost. On the other hand, it provides a means for the network providers to ensure, with high probability, cost recovery and profit, competitiveness of prices, and encouragement of client behavior that will enhance the network's efficiency. This is achieved by appropriate charging mechanisms and suitable incentives. Simulation results suggest that users have better overall satisfaction; providers are able to recover costs; better network utilization is achieved while reduced call blocking probability is observed. The implementation and usage costs of the framework are low.
Safiullah Faizullah, Ivan Marsic
GLOBECOM2
2001 An Application of Parameter Estimation to Route Discovery by On-Demand Routing Protocols
abstract
To discover a route to a peer node, an on-demand routing protocol may initiate a flood-search procedure known as route discovery. By selecting the correct query radius, the number of packet transmissions required for route discovery can be minimized. This paper presents methods to estimate the geographic radius (R/sub G/) and the number of currently active pairs of communicating nodes (P) in a mobile ad-hoc network. The methods are entirely distributed and incur little communication overhead. Network nodes can apply the estimated parameters to predict the probability mass function (PMF) of the route discovery hop distance. An accurate prediction of the PMF aids the selection of an appropriate query radius for the route discovery process. A computationally lightweight procedure to select an appropriate query radius, based only on an estimate of P, is also proposed. Simulation results show that this procedure facilitates a sensible tradeoff between the route request packet overhead and the route reply delay.
John Sucec, Ivan Marsic
ICDCS2
2001 An Architecture for Heterogeneous Groupware Applications
abstract
The proliferation of wireless networks and small portable computing devices raises the need for applications that are adaptable to heterogeneous computing and communication environments and the contexts in which they are used. However, most current groupware systems as well as other software applications are not well prepared to handle the heterogeneity. The Manifold framework presented provides a software architecture for synchronous groupware applications to deal with heterogeneity. The framework's main characteristic is data centricity. The users collaborate on and exchange data, and the data is dynamically transformed to adapt to the particular computing/network platform. The design is based on a multi-tier architecture and uses eXtensible Markup Language (XML) as a generic means for information exchange. The resulting design is simple yet very powerful and scalable. Manifold is implemented and tested by developing several complex groupware applications.
Ivan Marsic
ICSE1
2001 ViBE: virtual biology experiments
abstract
Article Share on ViBE: virtual biology experiments Authors: Rajaram Subramanian Department of Electrical and Computer Engineering and the CAIP Center, Rutgers -- The State University of New Jersey, Piscataway, NJ Department of Electrical and Computer Engineering and the CAIP Center, Rutgers -- The State University of New Jersey, Piscataway, NJView Profile , Ivan Marsic Department of Electrical and Computer Engineering and the CAIP Center, Rutgers -- The State University of New Jersey, Piscataway, NJ Department of Electrical and Computer Engineering and the CAIP Center, Rutgers -- The State University of New Jersey, Piscataway, NJView Profile Authors Info & Claims WWW '01: Proceedings of the 10th international conference on World Wide WebMay 2001 Pages 316–325https://doi.org/10.1145/371920.372076Online:01 April 2001Publication History 14citation844DownloadsMetricsTotal Citations14Total Downloads844Last 12 Months17Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Rajaram Subramanian, Ivan Marsic
WWW2
2000 Message caching for local and global resource optimization in shared virtual environments
abstract
The use of Shared Virtual Environments is growing in areas such as multi-player video games, military and industrial training, and collaborative design and engineering. As a result, different mixes of computing power and graphics capabilities of the participating computers arise naturally as the variety of people/organizations sharing a virtual environment grows. This paper presents an adaptive mechanism to reduce bandwidth usage and to optimize the use of computing resources of heterogeneous computers mixes utilized in a shared virtual environment. The mechanism is based on caching of both outgoing- and incoming-messages. We also report the results of implementing the proposed scheme in a simple shared virtual environment.
Helmuth Trefftz, Ivan Marsic
VRST2
2000 Natural communication with information systems
abstract
Pervasive networking and sophisticated computing open opportunities for collaborative information processing independent of time and space. In this instance the information system becomes an enhancer of human intellect, as well as a mediator for communication among participants. The human user favors the sensory dimensions of sight, sound, and touch as primary channels of communication. Machines that can accommodate these modes promise flexibilities and functionalities that transcend the traditional mouse and keyboard. The paper describes research to establish human-computer interfaces that capture attributes of natural face-to-face communication. An experimental multimodal system is developed to study several aspects of natural style human-computer communication. While as yet primitive, the technologies of image and gaze processing, hands-free conversation, and force feedback tactile transduction are combined and used simultaneously for manipulating objects in a shared workspace. Software agents fuse the sensory signals to estimate and interpret user intent. Current areas of experimental application include disaster relief/crisis management, telemedicine/rehabilitation, and mobile office/wearable computers.
Ivan Marsic, Attila Medl, James L. Flanagan
Proc. IEEE1
1999 A Desktop Design for Synchronous Collaboration
Bogdan Dorohonceanu, Ivan Marsic
Graphics Interface2
1999 Collaboration transparency in the DISCIPLE framework
abstract
Sharing single-user software applications is a major goal of synchronous groupware particularly because the majority of applications continues to be developed for single users. We present a mechanism for sharing collaboration-transparent single-user applications in our DISCIPLE collaboration framework. DISCIPLE is the equivalent of a Web browser that allows sharing applets (Java components, both transparent and aware of collaboration). It allows users with no programming background to quickly assemble arbitrary collaborative applications. Even though the presented solutions are specific to Java, many apply to other platforms as well. We introduce a novel concept of resource servers to solve the problem of resource access in collaborationtransparent applications. We also discuss the limitations of the framework in particular and of sharing collaboration-transparent applications in general. The framework has been implemented and tested on a variety of applications. Preliminary experiment...
Weicong Wang, Ivan Marsic
GROUP3
1999 An Advanced Communication Toolkit for Implementing the Broker Pattern
abstract
The Broker pattern is a powerful solution when building middleware communication systems. Existing toolkits, such as BAST, GTS, and ACE, although useful, are insufficient to implement the Broker pattern architecture. These systems concentrate on wrappers for communication protocols, and on implementing auxiliary communication patterns, but address only some aspects of object communication. In this work we demonstrate how the Broker pattern can be easily implemented by using an Advanced Communication Toolkit (ACT). ACT model defines four layers according to the increasing degree of abstraction of exchanged information. The resulting systems are highly customizable, extensible, portable, and can communicate at any of the four layers independently. ACT supports various high-level communication protocols (e.g., HTTP, IIOP, SMTP) and can be used to implement Broker-based systems such as OMG CORBA, Java RMI, or Microsoft DCOM.
Cristian Francu, Ivan Marsic
ICDCS2
1999 View-based object recognition using saliency maps
Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson
Image Vis. Comput.2
1998 View-Based Object Matching
abstract
We introduce a novel view-based object representation, called the saliency map graph (SMG), which captures the salient regions of an object view at multiple scales using a wavelet transform. This compact representation is highly invariant to translation, rotation (image and depth), and scaling, and offers the locality of representation required for occluded object recognition. To compare two saliency map graphs, we introduce two graph similarity algorithms. The first computes the topological similarity between two SMG's, providing a coarse-level matching of two graphs. The second computes the geometrical similarity between two SMG's, providing a fine-level matching of two graphs. We test and compare these two algorithms on a large database of model object views.
Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson
ICCV2
1998 A system for medical consultation and education using multimodal human/machine communication
abstract
Recent developments in networking and computing have enabled collaborative biomedical engineering research by geographically separated participants. One of the most promising goals is to use these technologies to extend human intellectual capabilities in medical decision making. These emerging technologies are poised to drastically reduce healthcare cost by providing service at remote locations. This also increases diagnosis capacity since information is made available to experts at any location. In this paper, we propose a novel application of a recently developed interactive and distributed system in medical consultation and education. Our approach builds on the notion that interactive and distributive capabilities of the system are crucial for medical consultation and education. The presented application uses a multiuser, collaborative environment with multimodal human/machine communication in the dimensions of sight, sound, and touch. The experimental setup, consisting of two user stations, and the multimodal interfaces, including sight (eye-tracking), sound (automatic speech), and touch (microbeam pen), were tested and evaluated. The system uses a collaborative workspace as a common visualization space. Users communicate with the application through a fusion agent by eye-tracking, speech, and microbeam pen. The audio/video teleconferencing is also included to help the radiologists to communicate with each other simultaneously while they are working on the mammograms. The system used in this study has three software agents: a fusion agent, a conversational agent, and an analytic agent. The fusion agent interprets multimodal commands by integrating the multimodal inputs. The conversational agent answers the user's questions and detects human-related or semantic errors and notifies the user about the results of the image analysis. The analytic agent enhances the digitized images using the wavelet denoising algorithm if requested by the user. To show how well the system performs in practice, we used the system for medical consultation on mammograms. Results also show that the relevant information about the region of interest (ROI) of the mammograms chosen by the users is extracted automatically and used to enhance the mammograms.
Metin Akay, Ivan Marsic, Attila Medl, Guangming Bu
IEEE Trans. Inf. Technol. Biomed.2
1997 Issues in measuring the benefits of multimodal interfaces
abstract
Multimedia interfaces are rapidly evolving to facilitate human/machine communication. Most of the technologies on which they are based are, as yet, imperfect. But, the interfaces do begin to allow information exchange in ways familiar and comfortable to the human-principally through natural actions in the sensory dimensions of sight, sound and touch. Further, as digital networking becomes ubiquitous, the opportunity grows for collaborative work through conferenced computing. In this context the machine takes on the role of mediator in human/machine/human communication-the ideal being to extend the intellectual abilities of humans through access to distributed information resources and collective decision making. The challenge is how to design machine mediation so that it extends, not impedes, human abilities. This report describes evolving work to incorporate multimodal interfaces into a networked system for collaborative distributed computing. It also addresses strategies for quantifying the synergies that may be gained.
James L. Flanagan, Ivan Marsic
ICASSP2
1997 Saliency-Based Visual Representation for Compression
abstract
This paper presents a representation for video with application to coding. In recent years, scalability of image compression systems has gained in popularity. Our approach is scalable in the normal sense, but instead of allowing the overall image frame quality to degrade with lower bandwidth, salient portions of the image frame maintain a high quality while less important portions are allowed to degrade. The judgment of importance is made by a region-of-interest (ROI) finder. The use of prioritized ROIs help to eliminate redundancy and can be organized under a motion compensation ruling.
Thomas E. Slowe, Ivan Marsic
ICIP (2)2
1997 Compression guidelines for diagnostic telepathology
abstract
As the healthcare community has begun to rely increasingly upon digital technologies for acquisition, storage, and transmission of pictorial data, image compression has become an indispensable tool. We have investigated the feasibility of lossy compression in a well-defined task domain, the clinical assessment of digitized images of chromatic microscopic pathology specimens. The effect of compression was measured under two distinct perceptual criteria, just noticeable difference (j.n.d.) and largest tolerable distortion (l.t.d.), differing in the involvement required from subjects, who were experts in pathology. For standard JPEG compressed images it was found that when the experiment is performed under the l.t.d. criterion, a significantly larger compression ratio is reported as satisfactory. It is concluded that lossy compression holds promise for diagnostic telepathology.
David J. Foran, Peter Meer, Thomas V. Papathomas, Ivan Marsic
IEEE Trans. Inf. Technol. Biomed.4
1996 Establishing perceptual criteria on image quality in diagnostic telepathology
abstract
The potential of lossy image compression is investigated in a well defined task domain, specifically, clinical assessment of chromatic, surgical and hematopathology specimens. Two criteria were employed, just noticeable difference and largest tolerable distortion. Compression tolerances differed among pathologists, but conformed to well defined upper and lower limits. The level of tolerable compression was significantly larger for the second criterion. It is concluded that lossy image compression is feasible for diagnostic pathology and may hold promise for telepathology applications.
David J. Foran, Peter Meer, Thomas V. Papathomas, Ivan Marsic, Leiguang Gong, Casimir A. Kulikowski, R. L. Trelstad
ICIP (1)4