VLDB 2026 Research / reviewers in the wild / expert
Margrit Betke
dblp:b/MargritBetke
· DBLP profile ↗
88ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0002-4491-6868ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 16 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTyabstractDifferent forms of customized 2D avatars are widely used in gaming applications, virtual communication, education, and content creation. However, existing approaches often fail to capture fine-grained facial expressions and struggle to preserve identity across different expressions. We propose Gen-AFFECT, a novel framework for personalized avatar generation that generates expressive and identity-consistent avatars with a diverse set of facial expressions. Our framework proposes conditioning a multimodal diffusion transformer on an extracted identity-expression representation. This enables identity preservation and representation of a wide range of facial expressions. Gen-AFFECT additionally employs consistent attention at inference for information sharing across the set of generated expressions, enabling the generation process to maintain identity consistency over the array of generated fine-grained expressions. Gen-AFFECT demonstrates superior performance compared to previous state-of-the-art methods on the basis of the accuracy of the generated expressions, the preservation of the identity and the consistency of the target identity across an array of fine-grained facial expressions. Hao Yu 0014, Rupayan Mallick, Margrit Betke, Sarah Adel Bargal |
WACV | 3 |
| 2026 | FourierMIL: Fourier Filtering-based Multiple Instance Learning for Whole Slide Image AnalysisabstractRecent advancements in computer vision, including convolutional neural networks, multilayer perceptrons, graph-based methods and transformer architectures, have significantly improved image classification. However, applying these techniques to digital pathology, particularly gigapixel whole-slide images (WSIs), presents unique challenges due to their vast size and heterogeneity. We introduce FourierMIL, a multiple instance learning framework that leverages the discrete Fourier transform to efficiently capture global and local dependencies in WSIs. Unlike conventional approaches, FourierMIL is adaptable to diverse digital stains and pathology tasks. To evaluate its versatility, we tested FourierMIL on three distinct challenges using public and private datasets. (1) Metastasis detection in hematoxylin and eosin (H&E)- stained lymph node WSIs from CAncer MEtastases in LYmph nOdes challeNge (CAMELYON16) dataset. (2) Lung cancer classification (adenocarcinoma versus squamous cell carcinoma) using The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC) datasets. (3) Alzheimer's disease pathology identification in phospho-tau monoclonal antibody (AT8)- stained WSIs from the Understanding Neurologic Injury and Traumatic Encephalopathy (UNITE), the Framingham Heart Study (FHS), and the Boston University Alzheimer's Disease Research Center (ADC) cohorts. FourierMIL outperformed state-of-the-art methods across all tasks, demonstrating its robustness as an attention-free solution for diverse applications in digital pathology. Yi Zheng 0006, Margrit Betke, Jonathan D. Cherry, Jesse B. Mez, Jennifer E. Beane, Vijaya B. Kolachalama |
Int. J. Comput. Vis. | 3 |
| 2025 | Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token PruningabstractLarge models achieve strong performance on Vision-and-Language Navigation (VLN) tasks, but are costly to run in resource-limited environments.Token pruning offers appealing tradeoffs for efficiency with minimal performance loss by reducing model input size, but prior work overlooks VLN-specific challenges.For example, information loss from pruning can effectively increase computational cost due to longer walks.Thus, the inability to identify uninformative tokens undermines the supposed efficiency gains from pruning.To address this, we propose Navigation-Aware Pruning (NAP), which uses navigation-specific traits to simplify the pruning process by pre-filtering tokens into foreground and background.For example, image views are filtered based on whether the agent can navigate in that direction.We also extract navigation-relevant instructions using a Large Language Model.After filtering, we focus pruning on background tokens, minimizing information loss.To further help avoid increases in navigation length, we discourage backtracking by removing low-importance navigation nodes.Experiments on standard VLN benchmarks show NAP significantly outperforms prior work, preserving higher success rates while saving more than 50% FLOPS 1 . Wenda Qin, Andrea Burns, Bryan A. Plummer, Margrit Betke |
EMNLP | 4 |
| 2025 | Generative adversarial networks for handwriting image generation: a reviewabstractAbstract Handwriting synthesis, the task of automatically generating realistic images of handwritten text, has gained increasing attention in recent years, both as a challenge in itself, as well as a task that supports handwriting recognition research. The latter task is to synthesize large image datasets that can then be used to train deep learning models to recognize handwritten text without the need for human-provided annotations. While early attempts at developing handwriting generators yielded limited results [1], more recent works involving generative models of deep neural network architectures have been shown able to produce realistic imitations of human handwriting [2–19]. In this review, we focus on one of the most prevalent and successful architectures in the field of handwriting synthesis, the generative adversarial network (GAN). We describe the capabilities, architecture specifics, and performance of the GAN-based models that have been introduced to the literature since 2019 [2–14]. These models can generate random handwriting styles, imitate reference styles, and produce realistic images of arbitrary text that was not in the training lexicon. The generated images have been shown to contribute to improving handwriting recognition results when augmenting the training samples of recognition models with synthetic images. The synthetic images were often hard to expose as non-real, even by human examiners, but also could be implausible or style-limited. The review includes a discussion of the characteristics of the GAN architecture in comparison with other paradigms in the image-generation domain and highlights the remaining challenges for handwriting synthesis. Randa I. Elanwar, Margrit Betke |
Vis. Comput. | 2 |
| 2024 | Affect Behavior Prediction: Using Transformers and Timing Information to Make Early Predictions of Student Exercise Outcome
Hao Yu 0014, Danielle Allessio, William Rebelsky, Tom Murray 0001, John J. Magee, Ivon Arroyo, Beverly P. Woolf, Sarah Adel Bargal, Margrit Betke |
AIED (2) | 9 |
| 2024 | Demonstration of CameraMouseAI: A Head-Based Mouse-Control System for People with Severe Motor DisabilitiesabstractWe propose the mouse control system CameraMouseAI that includes real-time facial feature detection and new ways to map facial feature movements to mouse clicks. In addition to selecting items on the screen with the traditional dwell-time mechanism, users can perform mouse commands on the screen by opening their mouth or raising their eyebrows. The modular and open source design of CameraMouseAI enables developers to update the platform with new computer vision and machine learning models and extend its functionality. User experiments with the CameraMouseAI involved target selection, web browsing, and typing. A group of 12 adults without motor impairments and four adults with severe motor impairments were able to use the application to complete both target selection and web browsing tasks. Farid Karimli, Hao Yu 0014, Srishti Jain, Emmanuel Sarpong Akosah, Margrit Betke, Wenxin Feng 0001 |
ASSETS | 5 |
| 2024 | Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text GenerationabstractPredicting emotions elicited by news headlines can be challenging as the task is largely influenced by the varying nature of people’s interpretations and backgrounds. Previous works have explored classifying discrete emotions directly from news headlines. We provide a different approach to tackling this problem by utilizing people’s explanations of their emotion, written in free-text, on how they feel after reading a news headline. Using the dataset BU-NEmo+ (Gao et al., 2022), we found that for emotion classification, the free-text explanations have a strong correlation with the dominant emotion elicited by the headlines. The free-text explanations also contain more sentimental context than the news headlines alone and can serve as a better input to emotion classification models. Therefore, in this work we explored generating emotion explanations from headlines by training a sequence-to-sequence transformer model and by using pretrained large language model, ChatGPT (GPT-4). We then used the generated emotion explanations for emotion classification. In addition, we also experimented with training the pretrained T5 model for the intermediate task of explanation generation before fine-tuning it for emotion classification. Using McNemar’s significance test, methods that incorporate GPT-generated free-text emotion explanations demonstrated significant improvement (P-value < 0.05) in emotion classification from headlines, compared to methods that only use headlines. This underscores the value of using intermediate free-text explanations for emotion prediction tasks with headlines. Ge Gao 0006, Jongin Kim, Sejin Paik, Ekaterina Novozhilova, Sarah Bonna, Margrit Betke, Derry Wijaya |
LREC/COLING | 7 |
| 2024 | Graph Attention-Based Fusion of Pathology Images and Gene Expression for Prediction of Cancer SurvivalabstractMultimodal machine learning models are being developed to analyze pathology images and other modalities, such as gene expression, to gain clinical and biological insights. However, most frameworks for multimodal data fusion do not fully account for the interactions between different modalities. Here, we present an attention-based fusion architecture that integrates a graph representation of pathology images with gene expression data and concomitantly learns from the fused information to predict patient-specific survival. In our approach, pathology images are represented as undirected graphs, and their embeddings are combined with embeddings of gene expression signatures using an attention mechanism to stratify tumors by patient survival. We show that our framework improves the survival prediction of human non-small cell lung cancers, outperforming existing state-of-the-art approaches that leverage multimodal data. Our framework can facilitate spatial molecular profiling to identify tumor heterogeneity using pathology images and gene expression data, complementing results obtained from more expensive spatial transcriptomic and proteomic technologies. Yi Zheng 0006, Regan D. Conrad, Emily J. Green, Eric J. Burks, Margrit Betke, Jennifer E. Beane, Vijaya B. Kolachalama |
IEEE Trans. Medical Imaging | 5 |
| 2023 | The Affective Nature of AI-Generated News Images: Impact on Visual JournalismabstractThis study explores the affective responses and newsworthiness perceptions of generative AI for visual journalism. While generative AI offers advantages for newsrooms in terms of producing unique images and cutting costs, the potential misuse of AI-generated news images is a cause for concern. For our study, we designed a 3-part news image codebook for affect-labeling news images based on journalism ethics and photography guidelines. We collected 200 news headlines and images retrieved from a variety of U.S. news sources on the topics of gun violence and climate change, generated corresponding news images from DALL-E 2 and asked annotators their emotional responses to the human-selected and AI-generated news images following the codebook. We also examined the impact of modality on emotions by measuring the effects of visual and textual modalities on emotional responses. The findings of this study provide insights into the quality and emotional impact of generative news images produced by humans and AI. Further, results of this work can be useful in developing technical guidelines as well as policy measures for the ethical use of generative AI systems in journalistic production. The codebook, images and annotations are made publicly available to facilitate future research in affective computing, specifically tailored to civic and public-interest journalism. Sejin Paik, Sarah Bonna, Ekaterina Novozhilova, Ge Gao 0006, Jongin Kim, Derry Wijaya, Margrit Betke |
ACII | 7 |
| 2023 | Age-constrained Ear Recognition: The EICZA Dataset and SASE Baseline ModelabstractUsing the ear as a biometric identifier, particularly for children in healthcare settings, has an important advantage over using the face - the privacy of the person can be protected better. However, aging and the resulting appearance differences, known to be challenges for face recognition models, have not been addressed for ear recognition yet. To address this limitation, we curated a publicly available dataset, which we call Ears of Infant Cohort in Zambia with Aging (EICZA)1. The dataset contains 3,330 ear images of 177 subjects, each photographed multiple times between the ages of 6 days and 9 months, when ear growth is most significant. For the task of age-constrained ear recognition, i.e., recognizing a person who has aged since the model was trained, we propose a new ear recognition model, called SASE for Self-Attention-based Sequential Ear image analysis. The model takes a sequence of ear images at early but different ages as input (instead of a single image) and processes them with a feature extraction network and a Transformer encoder. Trained with a large margin cosine loss function, the model is encouraged to learn a feature representation that distinguishes subjects from each other: Our experiments show that accounting for age enables our model to outperform other models that do not in recognizing ears that have grown and look different in later time periods. Wenda Qin, Lauren P. Etter, Alinani Simukanga, Christopher J. Gill, Margrit Betke |
IJCB | 5 |
| 2023 | CDAC: Cross-domain Attention Consistency in Transformer for Domain Adaptive Semantic SegmentationabstractWhile transformers have greatly boosted performance in semantic segmentation, domain adaptive transformers are not yet well explored. We identify that the domain gap can cause discrepancies in self-attention. Due to this gap, the transformer attends to spurious regions or pixels, which deteriorates accuracy on the target domain. We propose Cross-Domain Attention Consistency (CDAC), to perform adaptation on attention maps using cross-domain attention layers that share features between source and target domains. Specifically, we impose consistency between predictions from cross-domain attention and self-attention modules to encourage similar distributions across domains in both the attention and output of the model, i.e., attention-level and output-level alignment. We also enforce consistency in attention maps between different augmented views to further strengthen the attention-based alignment. Combining these two components, CDAC mitigates the discrepancy in attention maps across domains and further boosts the performance of the transformer under unsupervised domain adaptation settings. Our method is evaluated on various widely used benchmarks and outperforms the state-of-the-art baselines, including GTAV-to-Cityscapes by 1.3 and 1.5 percent point (pp) and Synthia-to-Cityscapes by 0.6 pp and 2.9 pp when combining with two competitive Transformer-based backbones, respectively. Our code will be publicly available at https://github.com/wangkaihong/CDAC. Kaihong Wang, Donghyun Kim 0006, Rogério Feris, Margrit Betke |
ICCV | 4 |
| 2023 | COVES: A Cognitive-Affective Deep Model that Personalizes Math Problem Difficulty in Real Time and Improves Student Engagement with an Online TutorabstractA key to personalized online learning is presenting content at an appropriate difficulty level; content that is too difficult can cause frustration and content that is too easy may result in boredom. Appropriate content can improve students' engagement and learning outcome. In this research, we propose a computer vision enhanced problem selector (COVES), a deep learning model to select a personalized difficulty level for each student. A combination of visual information and traditional log data is used to predict student-problem interactions, which are then used to guide problem difficulty selection in real time. COVES was trained on a dataset of fifty-one sixth-grade students interacting with the online math tutor MathSpring. Once COVES was integrated into the tutor, its effectiveness was tested with twenty-two seventh-grade students in controlled experiments. Students who received problems at an appropriate difficulty level, based on real-time predictions of their performance, demonstrated improved engagement with the math tutor. Results indicate that COVES leads to higher mastery of math concepts, better timing, and higher scores, thus providing a positive learning experience for the participants. Hao Yu 0014, Danielle Allessio, William Lee 0002, William Rebelsky, Frank Sylvia, Tom Murray 0001, John J. Magee, Ivon Arroyo, Beverly P. Woolf, Sarah Adel Bargal, Margrit Betke |
ACM Multimedia | 11 |
| 2023 | Animal Pose Tracking: 3D Multimodal Dataset and Token-based Pose OptimizationabstractAbstract Accurate tracking of the 3D pose of animals from video recordings is critical for many behavioral studies, yet there is a dearth of publicly available datasets that the computer vision community could use for model development. We here introduce the Rodent3D dataset that records animals exploring their environment and/or interacting with each other with multiple cameras and modalities (RGB, depth, thermal infrared). Rodent3D consists of 200 min of multimodal video recordings from up to three thermal and three RGB-D synchronized cameras (approximately 4 million frames). For the task of optimizing estimates of pose sequences provided by existing pose estimation methods, we provide a baseline model called OptiPose. While deep-learned attention mechanisms have been used for pose estimation in the past, with OptiPose, we propose a different way by representing 3D poses as tokens for which deep-learned context models pay attention to both spatial and temporal keypoint patterns. Our experiments show how OptiPose is highly robust to noise and occlusion and can be used to optimize pose sequences provided by state-of-the-art models for animal pose estimation. Mahir Patel, Yiwen Gu, Lucas C. Carstensen, Michael E. Hasselmo, Margrit Betke |
Int. J. Comput. Vis. | 5 |
| 2022 | A Unified Framework for Domain Adaptive Pose Estimation
Donghyun Kim 0006, Kaihong Wang, Kate Saenko, Margrit Betke, Stan Sclaroff |
ECCV (33) | 4 |
| 2022 | BU-NEmo: an Affective Dataset of Gun Violence NewsabstractGiven our society’s increased exposure to multimedia formats on social media platforms, efforts to understand how digital content impacts people’s emotions are burgeoning. As such, we introduce a U.S. gun violence news dataset that contains news headline and image pairings from 840 news articles with 15K high-quality, crowdsourced annotations on emotional responses to the news pairings. We created three experimental conditions for the annotation process: two with a single modality (headline or image only), and one multimodal (headline and image together). In contrast to prior works on affectively-annotated data, our dataset includes annotations on the dominant emotion experienced with the content, the intensity of the selected emotion and an open-ended, written component. By collecting annotations on different modalities of the same news content pairings, we explore the relationship between image and text influence on human emotional response. We offer initial analysis on our dataset, showing the nuanced affective differences that appear due to modality and individual factors such as political leaning and media consumption habits. Our dataset is made publicly available to facilitate future research in affective computing. Carley Reardon, Sejin Paik, Ge Gao 0006, Meet Parekh, Lei Guo 0017, Margrit Betke, Derry Wijaya |
LREC | 7 |
| 2022 | A Graph-Transformer for Whole Slide Image ClassificationabstractDeep learning is a powerful tool for whole slide image (WSI) analysis. Typically, when performing supervised deep learning, a WSI is divided into small patches, trained and the outcomes are aggregated to estimate disease grade. However, patch-based methods introduce label noise during training by assuming that each patch is independent with the same label as the WSI and neglect overall WSI-level information that is significant in disease grading. Here we present a Graph-Transformer (GT) that fuses a graph-based representation of an WSI and a vision transformer for processing pathology images, called GTP, to predict disease grade. We selected 4,818 WSIs from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), the National Lung Screening Trial (NLST), and The Cancer Genome Atlas (TCGA), and used GTP to distinguish adenocarcinoma (LUAD) and squamous cell carcinoma (LSCC) from adjacent non-cancerous tissue (normal). First, using NLST data, we developed a contrastive learning framework to generate a feature extractor. This allowed us to compute feature vectors of individual WSI patches, which were used to represent the nodes of the graph followed by construction of the GTP framework. Our model trained on the CPTAC data achieved consistently high performance on three-label classification (normal versus LUAD versus LSCC: mean accuracy = 91.2 ± 2.5%) based on five-fold cross-validation, and mean accuracy = 82.3 ± 1.0% on external test data (TCGA). We also introduced a graph-based saliency mapping technique, called GraphCAM, that can identify regions that are highly associated with the class label. Our findings demonstrate GTP as an interpretable and effective deep learning framework for WSI-level classification. Yi Zheng 0006, Rushin H. Gindra, Emily J. Green, Eric J. Burks, Margrit Betke, Jennifer E. Beane, Vijaya B. Kolachalama |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Consistency Regularization with High-dimensional Non-adversarial Source-guided Perturbation for Unsupervised Domain Adaptation in SegmentationabstractUnsupervised domain adaptation for semantic segmentation has been intensively studied due to the low cost of the pixel-level annotation for synthetic data. The most common approaches try to generate images or features mimicking the distribution in the target domain while preserving the semantic contents in the source domain so that a model can be trained with annotations from the latter. However, such methods highly rely on an image translator or feature extractor trained in an elaborated mechanism including adversarial training, which brings in extra complexity and instability in the adaptation process. Furthermore, these methods mainly focus on taking advantage of the labeled source dataset, leaving the unlabeled target dataset not fully utilized. In this paper, we propose a bidirectional style-induced domain adaptation method, called BiSIDA, that employs consistency regularization to efficiently exploit information from the unlabeled target domain dataset, requiring only a simple neural style transfer model. BiSIDA aligns domains by not only transferring source images into the style of target images but also transferring target images into the style of source images to perform high-dimensional perturbation on the unlabeled target images, which is crucial to the success in applying consistency regularization in segmentation tasks. Extensive experiments show that our BiSIDA achieves new state-of-the-art on two commonly-used synthetic-to-real domain adaptation benchmarks: GTA5-to-CityScapes and SYNTHIA-to-CityScapes. Code and pretrained style transfer model are available at: https://github.com/wangkaihong/BiSIDA. Kaihong Wang, Chenhongyi Yang, Margrit Betke |
AAAI | 3 |
| 2021 | Leveraging Affect Transfer Learning for Behavior Prediction in an Intelligent Tutoring SystemabstractIn this work, we propose a video-based transfer learning approach for predicting problem outcomes of students working with an intelligent tutoring system (ITS). By analyzing a student's face and gestures, our method predicts the outcome of a student answering a problem in an ITS from a video feed. Our work is motivated by the reasoning that the ability to predict such outcomes enables tutoring systems to adjust interventions, such as hints and encouragement, and to ultimately yield improved student learning. We collected a large labeled dataset of student interactions with an intelligent online math tutor consisting of 68 sessions, where 54 individual students solved 2,749 problems. We will release this dataset publicly upon publication of this paper. It will be available at https://www.cs.bu.edu/faculty/betke/research/learning/. Working with this dataset, our transfer-learning challenge was to design a representation in the source domain of pictures obtained “in the wild” for the task of facial expression analysis, and transferring this learned representation to the task of human behavior prediction in the domain of webcam videos of students in a classroom environment. We developed a novel facial affect representation and a user-personalized training scheme that unlocks the potential of this representation. We designed several variants of a recurrent neural network that models the temporal structure of video sequences of students solving math problems. Our final model, named ATL-BP for Affect Transfer Learning for Behavior Prediction, achieves a relative increase in mean F -score of 50 % over the state-of-the-art method on this new dataset. Nataniel Ruiz, Hao Yu 0014, Danielle Allessio, Mona Jalal, Ajjen Joshi, Tom Murray 0001, John J. Magee, Jacob Whitehill, Vitaly Ablavsky, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke |
FG | 13 |
| 2021 | Semantic-Based Sentence Recognition in Images Using Bimodal Deep LearningabstractThe accuracy of computer vision systems that understand sentences in images with text can be improved when semantic information about the text is utilized. Nonetheless, the semantic coherence within a region of text in natural or document images is typically ignored by state-of-the-art systems, which identify isolated words or interpret text word by word. However, when analyzed together, seemingly isolated words may be easier to recognize. On this basis, we propose a novel “Semantic-based Sentence Recognition” (SSR) deep learning model that reads text in images with the help of understanding context. SSR consists of a Word Ordering and Grouping Algorithm (WOGA) to find sentences in images and a Sequence-to-Sequence Recognition Correction (SSRC) model to extract semantic information in these sentences to improve their recognition. We present experiments with three notably distinct datasets, two of which we created ourselves. They respectively contain scanned catalog images of interior designs and photographs of protesters with hand-written signs. Our results show that SSR statistically significantly outperforms baseline methods that use state-of-the-art single-word-recognition techniques. By successfully combining both computer vision and natural language processing methodologies, we reveal the important opportunity that bi-modal deep learning can provide in addressing a task which was previously considered a single-modality computer vision task. Yi Zheng 0006, Qitong Wang 0001, Margrit Betke |
ICIP | 3 |
| 2021 | Extracting text from scanned Arabic books: a large-scale benchmark dataset and a fine-tuned Faster-R-CNN model
Randa I. Elanwar, Wenda Qin, Margrit Betke, Derry Wijaya |
Int. J. Document Anal. Recognit. | 3 |
| 2020 | Multi-Label and Multilingual News Framing AnalysisabstractNews framing refers to the practice in which aspects of specific issues are highlighted in the news to promote a particular interpretation.In NLP, although recent works have studied framing in English news, few have studied how the analysis can be extended to other languages and in a multi-label setting.In this work, we explore multilingual transfer learning to detect multiple frames from just the news headline in a genuinely low-resource context where there are few/no frame annotations in the target language.We propose a novel method that can leverage elementary resources consisting of a dictionary and few annotations to detect frames in the target language.Our method performs comparably or better than translating the entire target language headline to the source language for which we have annotated data.This work opens up an exciting new capability of scaling up frame analysis to many languages, even those without existing translation technologies.Lastly, we apply our method to detect frames on the issue of U.S. gun violence in multiple languages and obtain exciting insights on the relationship between different frames of the same problem across different countries with different languages. Afra Feyza Akyürek, Lei Guo 0017, Randa I. Elanwar, Prakash Ishwar, Margrit Betke, Derry Wijaya |
ACL | 5 |
| 2020 | Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes
Chenhongyi Yang, Vitaly Ablavsky, Kaihong Wang, Qi Feng 0004, Margrit Betke |
ECCV (18) | 5 |
| 2020 | LAL: Linguistically Aware Learning for Scene Text RecognitionabstractScene text recognition is the task of recognizing character sequences in images of natural scenes. The considerable diversity in the appearance of text in a scene image and potentially highly complex backgrounds make text recognition challenging. Previous approaches employ character sequence generators to analyze text regions and, subsequently, compare the candidate character sequences against a language model. In this work, we propose a bimodal framework that simultaneously utilizes visual and linguistic information to enhance recognition performance. Our linguistically aware learning (LAL) method effectively learns visual embeddings using a rectifier, encoder, and attention decoder approach, and linguistic embeddings, using a deep next-character prediction model. We present an innovative way of combining these two embeddings effectively. Our experiments on eight standard benchmarks show that our method outperforms previous methods by large margins, particularly on rotated, foreshortened, and curved text. We show that the bimodal approach has a statistically significant impact. We also contribute a new dataset, and show robust performance when LAL is combined with a text detector in a pipelined text spotting framework. Yi Zheng 0006, Wenda Qin, Derry Wijaya, Margrit Betke |
ACM Multimedia | 4 |
| 2020 | Visual complexity analysis using deep intermediate-layer featuresabstractIn this paper, we focus on visual complexity, an image attribute that humans can subjectively evaluate based on the level of details in the image. We explore unsupervised information extraction from intermediate convolutional layers of deep neural networks to measure visual complexity. We derive an activation energy metric that combines convolutional layer activations to quantify visual complexity. To show the effectiveness of our proposed metric for various applications, we introduce Savoias, a visual complexity dataset that compromises of more than 1,400 images from seven diverse image categories (e.g., advertisement and interior design). We demonstrate high correlations of our deep neural network-based measure of visual complexity with human-curated ground-truth (GT) scores on various widely used network architectures, e.g., VGG16, ResNet-v2-152, and EfficientNet, and in networks trained on two classification tasks (object and scene classification). This result reveals that intermediate convolutional layers of deep neural networks carry information about the complexity of images that is meaningful to people. Furthermore, we show that our method of measuring visual complexity outperforms traditional methods on Savoias and two other state-of-the-art benchmark datasets. Moreover, we perform extensive analysis on the performance difference between our unsupervised method and a supervised method trained on the feature map, and show that by supervision, we can improve the prediction. Finally, we demonstrate that, within the context of a category, visually more complex images are also more memorable to human observers. Elham Saraee, Mona Jalal, Margrit Betke |
Comput. Vis. Image Underst. | 3 |
| 2019 | Detecting Frames in News Headlines and Its Application to Analyzing News Framing Trends Surrounding U.S. Gun ViolenceabstractDifferent news articles about the same topic often offer a variety of perspectives: an article written about gun violence might emphasize gun control, while another might promote 2nd Amendment rights, and yet a third might focus on mental health issues.In communication research, these different perspectives are known as "frames", which, when used in news media will influence the opinion of their readers in multiple ways.In this paper, we present a method for effectively detecting frames in news headlines.Our training and performance evaluation is based on a new dataset of news headlines related to the issue of gun violence in the United States.This Gun Violence Frame Corpus (GVFC) was curated and annotated by journalism and communication experts.Our proposed approach sets a new state-of-the-art performance for multiclass news frame detection, significantly outperforming a recent baseline by 35.9% absolute difference in accuracy.We apply our frame detection approach in a large scale study of 88k news headlines about the coverage of gun violence in the U.S. between 2016 and 2018. Lei Guo 0017, Kate K. Mays, Margrit Betke, Derry Wijaya |
CoNLL | 4 |
| 2019 | Affect-driven Learning Outcomes Prediction in Intelligent Tutoring SystemsabstractEquipping an Intelligent Tutoring System (ITS) with the ability to interpret affective signals from students could potentially improve the learning experience of students by enabling the tutor to monitor the students' progress and provide timely interventions as well as present appropriate affective reactions via a virtual tutor. Most ITSs equipped with affect modeling capabilities attempt to predict the emotional state of users. However, the focus in this work is instead on trying to directly predict the learning outcomes of students from a stream of video capturing the students faces as they work on a set of math problems. Using facial features extracted from a video stream, we train classifiers to directly predict the success or failure of a student's attempt to answer a question while the student has just begun to work on the problem. In this work, we first introduce a novel dataset of student interactions with MathSpring, a popular ITS. We provide an exploratory analysis of the different problem outcome classes using typical facial action unit activations. We develop baseline models to predict the problem outcome labels of students solving math problems and discuss how early problem outcome labels can be forecasted and utilized to provide possible interventions. Ajjen Joshi, Danielle Allessio, John J. Magee, Jacob Whitehill, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke |
FG | 8 |
| 2019 | Predicting How to Distribute Work Between Algorithms and Humans to Segment an Image Batch
Danna Gurari, Suyog Dutt Jain, Margrit Betke, Kristen Grauman |
Int. J. Comput. Vis. | 4 |
| 2018 | Context-Sensitive Prediction of Facial Expressivity Using Multimodal Hierarchical Bayesian Neural NetworksabstractObjective automated affect analysis systems can be applied to quantify the progression of symptoms in neurodegenerative diseases such as Parkinson's Disease (PD). PD hampers the ability of patients to emote by decreasing the mobility of their facial musculature, a phenomenon known as ``facial masking.'' In this work, we focus on building a system that can predict an accurate score of active facial expressivity in people suffering from Parkinson's disease using features extracted from both video and audio. An ideal automated system should be able to mimic the ability of human experts to take into account contextual information while making these predictions. For example, patients exhibit different emotions with varying intensities when speaking about positive and negative experiences. We utilize a hierarchical Bayesian neural network framework to enable the learning of model parameters that subtly adapt to pre-defined notions of context, such as the gender of the patient or the valence of the expressed sentiment. We evaluate our formulation on a dataset of 772 20-second video clips of Parkinson's disease patients and demonstrate that training a context-specific hierarchical Bayesian framework yields an improvement in model performance in both multiclass classification and regression settings compared to baseline models trained on all data pooled together. Ajjen Joshi, Soumya Ghosh, Sarah Gunnery, Linda Tickle-Degnen, Stan Sclaroff, Margrit Betke |
FG | 6 |
| 2018 | Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)
Danna Gurari, Kun He 0003, Jianming Zhang 0001, Mehrnoosh Sameki, Suyog Dutt Jain, Stan Sclaroff, Margrit Betke, Kristen Grauman |
Int. J. Comput. Vis. | 8 |
| 2018 | Making scanned Arabic documents machine accessible using an ensemble of SVM classifiers
Randa I. Elanwar, Wenda Qin, Margrit Betke |
Int. J. Document Anal. Recognit. | 3 |
| 2018 | Exploration of Assistive Technologies Used by People with Quadriplegia Caused by Degenerative Neurological DiseasesabstractVarious assistive devices and interfaces to access the computer have been developed for people with severe motor impairments. This article explores how effective these technologies are for individuals with quadriplegia caused by degenerative neurological diseases. The following questions are studied: (1) What activities are performed? (2) What tools are used? (3) What are the advantages and limitations of the tools? (4) How do users learn about and choose assistive technologies? (5) Why are some technologies abandoned? Results of a qualitative study with 15 participants indicate that study participants have strong needs for efficient text entry and communication that are not met. A lack of information about technology options limits the choices of several of the study participants. The study revealed that automated interface personalization and adaptation to disease progression should be important design goals for future assistive technologies that support users with quadriplegia caused by degenerative neurological diseases. Wenxin Feng 0001, Mehrnoosh Sameki, Margrit Betke |
Int. J. Hum. Comput. Interact. | 3 |
| 2017 | Personalizing Gesture Recognition Using Hierarchical Bayesian Neural NetworksabstractBuilding robust classifiers trained on data susceptible to group or subject-specific variations is a challenging pattern recognition problem. We develop hierarchical Bayesian neural networks to capture subject-specific variations and share statistical strength across subjects. Leveraging recent work on learning Bayesian neural networks, we build fast, scalable algorithms for inferring the posterior distribution over all network weights in the hierarchy. We also develop methods for adapting our model to new subjects when a small number of subject-specific personalization data is available. Finally, we investigate active learning algorithms for interactively labeling personalization data in resource-constrained scenarios. Focusing on the problem of gesture recognition where inter-subject variations are commonplace, we demonstrate the effectiveness of our proposed techniques. We test our framework on three widely used gesture recognition datasets, achieving personalization performance competitive with the state-of-the-art. Ajjen Joshi, Soumya Ghosh, Margrit Betke, Stan Sclaroff, Hanspeter Pfister |
CVPR | 3 |
| 2017 | Crowd-O-Meter: Predicting if a Person Is Vulnerable to Believe Political ClaimsabstractSocial media platforms have been criticized for promoting false information during the 2016 U.S. presidential election campaign. Our work is motivated by the idea that a platform could reduce the circulation of false information if it could estimate whether its users are vulnerable to believing political claims. We here explore whether such a vulnerability could be measured in a crowdsourcing setting. We propose Crowd-O-Meter, a framework that automatically predicts if a crowd worker will be consistent in his/her beliefs about political claims; i.e., consistently believes the claims are true or consistently believes the claims are not true. Crowd-O-Meter is a user-centered approach which interprets a combination of cues characterizing the user's implicit and explicit opinion bias. Experiments on 580 quotes from PolitiFact's fact checking corpus of 2016 U.S. presidential candidates show that Crowd-O-Meter is precise and accurate for two news modalities: text and video. Our analysis also reveals which are the most informative cues of a person's vulnerability. Mehrnoosh Sameki, Linli Ding, Margrit Betke, Danna Gurari |
HCOMP | 4 |
| 2017 | Salient Object Subitizing
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
Int. J. Comput. Vis. | 5 |
| 2017 | Comparing random forest approaches to segmenting and classifying gestures
Ajjen Joshi, Camille Monnier, Margrit Betke, Stan Sclaroff |
Image Vis. Comput. | 3 |
| 2016 | EyeSwipe: Dwell-free Text Entry Using Gaze PathsabstractText entry using gaze-based interaction is a vital communication tool for people with motor impairments. Most solutions require the user to fixate on a key for a given dwell time to select it, thus limiting the typing speed. In this paper we introduce EyeSwipe, a dwell-time-free gaze-typing method. With EyeSwipe, the user gaze-types the first and last characters of a word using the novel selection mechanism "reverse crossing." To gaze-type the characters in the middle of the word, the user only needs to glance at the vicinity of the respective keys. We compared the performance of EyeSwipe with that of a dwell-time-based virtual keyboard. EyeSwipe afforded statistically significantly higher typing rates and more comfortable interaction in experiments with ten participants who reached 11.7 words per minute (wpm) after 30 min typing with EyeSwipe. Andrew T. N. Kurauchi, Wenxin Feng 0001, Ajjen Joshi, Carlos Hitoshi Morimoto, Margrit Betke |
CHI | 5 |
| 2016 | Pull the Plug? Predicting If Computers or Humans Should Segment ImagesabstractForeground object segmentation is a critical step for many image analysis tasks. While automated methods can produce high-quality results, their failures disappoint users in need of practical solutions. We propose a resource allocation framework for predicting how best to allocate a fixed budget of human annotation effort in order to collect higher quality segmentations for a given batch of images and automated methods. The framework is based on a proposed prediction module that estimates the quality of given algorithm-drawn segmentations. We demonstrate the value of the framework for two novel tasks related to "pulling the plug" on computer and human annotators. Specifically, we implement two systems that automatically decide, for a batch of images, when to replace 1) humans with computers to create coarse segmentations required to initialize segmentation tools and 2) computers with humans to create final, fine-grained segmentations. Experiments demonstrate the advantage of relying on a mix of human and computer efforts over relying on either resource alone for segmenting objects in three diverse datasets representing visible, phase contrast microscopy, and fluorescence microscopy images. Danna Gurari, Suyog Dutt Jain, Margrit Betke, Kristen Grauman |
CVPR | 3 |
| 2016 | Investigating the Influence of Data Familiarity to Improve the Design of a Crowdsourcing Image Annotation SystemabstractCrowdsourced demarcations of object boundaries in images (segmentations) are important for many vision-based applications. A commonly reported challenge is that a large percentage of crowd results are discarded due to concerns about quality. We conducted three studies to examine (1) how does the quality of crowdsourced segmentations differ for familiar everyday images versus unfamiliar biomedical images?, (2) how does making familiar images less recognizable (rotating images upside down) influence crowd work with respect to the quality of results, segmentation time, and segmentation detail?, and (3) how does crowd workers’ judgments of the ambiguity of the segmentation task, collected by voting, differ for familiar everyday images and unfamiliar biomedical images? We analyzed a total of 2,525 segmentations collected from 121 crowd workers and 1,850 votes from 55 crowd workers. Our results illustrate the potential benefit of explicitly accounting for human familiarity with the data when designing computer interfaces for human interaction. Danna Gurari, Mehrnoosh Sameki, Margrit Betke |
HCOMP | 3 |
| 2016 | Discovering useful parts for pose estimation in sparsely annotated datasetsabstractOur work introduces a novel way to increase pose estimation accuracy by discovering parts from unannotated regions of training images. Discovered parts are used to generate more accurate appearance likelihoods for traditional part-based models like Pictorial Structures and its derivatives. Our experiments on images of a hawkmoth in flight show that our proposed approach significantly improves over existing work for this application, while also being more generally applicable. Our proposed approach localizes landmarks at least twice as accurately as a baseline based on a Mixture of Pictorial Structures (MPS) model. Our unique High-Resolution Moth Flight (HRMF) dataset is made publicly available with annotations. Mikhail Breslav, Tyson L. Hedrick, Stan Sclaroff, Margrit Betke |
WACV | 4 |
| 2016 | Global optimization for coupled detection and data association in multiple object tracking
Zheng Wu 0003, Margrit Betke |
Comput. Vis. Image Underst. | 2 |
| 2015 | Salient Object SubitizingabstractPeople can immediately and precisely identify that an image contains 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this problem, we propose a new image dataset annotated using an online crowdsourcing marketplace. We show that a proposed subitizing technique using an end-to-end Convolutional Neural Network (CNN) model achieves significantly better than chance performance in matching human labels on our dataset. It attains 94% accuracy in detecting the existence of salient objects, and 42–82% accuracy (chance is 20%) in predicting the number of salient objects (1, 2, 3, and 4+), without resorting to any object localization process. Finally, we demonstrate the usefulness of the proposed subitizing technique in two computer vision applications: salient object detection and object proposal. Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
CVPR | 5 |
| 2015 | Predicting Quality of Crowdsourced Image Segmentations from Crowd BehaviorabstractQuality control (QC) is an integral part of many crowd- sourcing systems. However, popular QC methods, such as aggregating multiple annotations, filtering workers, or verifying the quality of crowd work, introduce additional costs and delays. We propose a complementary paradigm to these QC methods based on predicting the quality of submitted crowd work. In particular, we pro- pose to predict the quality of a given crowd drawing directly from a crowd worker’s drawing time, number of user clicks, and average time per user click. We focus on the task of drawing the boundary of a single object in an image. To train and test our prediction models, we collected a total of 2,025 crowd-drawn segmentations for 405 familiar everyday images and unfamiliar biomedical images from 90 unique crowd workers. We first evaluated five prediction models learned using different combinations of the three worker behavior cues for all images. Experiments revealed that time per number of user clicks was the most effective cue for predicting segmentation quality. We next inspected the predictive power of models learned using crowd annotations collected for familiar and unfamiliar data independently. Prediction models were significantly more effective for estimating the segmentation quality from crowd worker behavior for familiar image content than unfamiliar image content. Mehrnoosh Sameki, Danna Gurari, Margrit Betke |
HCOMP | 3 |
| 2015 | How to Collect Segmentations for Biomedical Images? A Benchmark Evaluating the Performance of Experts, Crowdsourced Non-experts, and AlgorithmsabstractAnalyses of biomedical images often rely on demarcating the boundaries of biological structures (segmentation). While numerous approaches are adopted to address the segmentation problem including collecting annotations from domain-experts and automated algorithms, the lack of comparative benchmarking makes it challenging to determine the current state-of-art, recognize limitations of existing approaches, and identify relevant future research directions. To provide practical guidance, we evaluated and compared the performance of trained experts, crowd sourced non-experts, and algorithms for annotating 305 objects coming from six datasets that include phase contrast, fluorescence, and magnetic resonance images. Compared to the gold standard established by expert consensus, we found the best annotators were experts, followed by non-experts, and then algorithms. This analysis revealed that online paid crowd sourced workers without domain-specific backgrounds are reliable annotators to use as part of the laboratory protocol for segmenting biomedical images. We also found that fusing the segmentations created by crowd sourced internet workers and algorithms yielded improved segmentation results over segmentations created by single crowd sourced or algorithm annotations respectively. We invite extensions of our work by sharing our data sets and associated segmentation annotations (http://www.cs.bu.edu/~betke/Biomedical Image Segmentation). Danna Gurari, Diane H. Theriault, Mehrnoosh Sameki, Brett Isenberg, Tuan A. Pham 0002, Alberto Purwada, Patricia Solski, Matthew L. Walker, Chentian Zhang, Joyce Y. Wong, Margrit Betke |
WACV | 11 |
| 2014 | 3D pose estimation of bats in the wildabstractVision-based methods have gained popularity as a tool for helping to analyze the behavior of bats. Though, for bats in the wild, there are still no tools capable of estimating and subsequently analyzing articulated 3D bat pose. We propose a model-based multi-view articulated 3D bat pose estimation framework for this novel problem. Key challenges include the large search space associated with articulated 3D pose, the ambiguities that arise from 2D projections of 3D bodies, and the low resolution image data we have available. Our method uses multi-view camera geometry and temporal constraints to reduce the state space of possible articulated 3D bat poses and finds an optimal set using a Markov Random Field based model. Our experiments use real video data of flying bats and gold-standard annotations by a bat biologist. Our results show, for the first time in the literature, articulated 3D pose estimates being generated automatically for video sequences of bats flying in the wild. The average differences in body orientation and wing joint angles, between estimates produced by our method and those based on gold-standard annotations, ranged from 16° - 21° (i.e., ≈ 17% - 23%) for orientation and 14° - 26° (i.e., ≈ 7%- 14%) for wing joint angles. Mikhail Breslav, Nathan W. Fuller, Stan Sclaroff, Margrit Betke |
WACV | 4 |
| 2014 | Using kernels for a video-based mouse-replacement interfaceabstractSome people cannot use their hands to control a computer mouse due to conditions such as cerebral palsy or multiple sclerosis. For these individuals, there are various mouse-replacement solutions. One approach is to enable them to control the mouse pointer using head motions captured with a web camera. One such system, the Camera Mouse, uses an optical flow approach to track a manually-selected small patch of the subject’s face, such as the nostril or the edge of the eyebrow. The optical flow tracker may lose the facial feature when the tracked image patch drifts away from the initially-selected feature or when a user makes a rapid head movement. To address the problem of feature loss, we developed and incorporated the Kernel-Subset-Tracker into the Camera Mouse. The Kernel-Subset-Tracker is an exemplar-based method that uses a training set of representative images to produce online templates for positional tracking. We designed the augmented Camera Mouse so that it can compute these templates in real time, employing kernel techniques traditionally used for classification. We propose three versions of the Kernel-Subset-Tracker , each using a different kernel, and compared their performance to the optical-flow tracker under five different experimental conditions. Our experiments with test subjects show that augmenting the Camera Mouse with the Kernel-Subset-Tracker improves communication bandwidth statistically significantly. Tracking of facial features was accurate, without feature drift, even during rapid head movements and extreme head orientations. We conclude by describing how the Camera Mouse augmented with the Kernel-Subset-Tracker enabled a stroke-victim with severe motion impairments to communicate via an on-screen keyboard. Samuel Epstein 0001, Eric S. Missimer, Margrit Betke |
Pers. Ubiquitous Comput. | 3 |
| 2014 | Assistive environments for the disabled and the senior citizens: theme issue of PETRA 2010 and 2011 conferences
Ilias Maglogiannis, Margrit Betke, Grammati E. Pantziou, Fillia Makedon |
Pers. Ubiquitous Comput. | 2 |
| 2013 | Online Motion Agreement TrackingabstractThis paper proposes a fast online multi-target tracking method, called motion agreement algorithm, which dynamically selects stable object regions to track. The appearance of each object, here pedestrians, is represented by multiple local patches. For each patch, the algorithm computes a local estimate of the direction of motion. By fusion of the agreements between a global estimate of the object motion and each local estimate, the algorithm identifies the object stable regions and enables robust tracking. The proposed patch-based appearance model was integrated into an efficient online tracking system that uses bipartite matching for data association. The experiments on recent pedestrian tracking benchmark sequences show that the proposed method achieves competitive results compared to state-of-the-art methods, including several offline tracking techniques. Zheng Wu 0003, Jianming Zhang 0001, Margrit Betke |
BMVC | 3 |
| 2013 | Randomized Ensemble TrackingabstractWe propose a randomized ensemble algorithm to model the time-varying appearance of an object for visual tracking. In contrast with previous online methods for updating classifier ensembles in tracking-by-detection, the weight vector that combines weak classifiers is treated as a random variable and the posterior distribution for the weight vector is estimated in a Bayesian manner. In essence, the weight vector is treated as a distribution that reflects the confidence among the weak classifiers used to construct and adapt the classifier ensemble. The resulting formulation models the time-varying discriminative ability among weak classifiers so that the ensembled strong classifier can adapt to the varying appearance, backgrounds, and occlusions. The formulation is tested in a tracking-by-detection implementation. Experiments on 28 challenging benchmark videos demonstrate that the proposed method can achieve results comparable to and often better than those of state-of-the-art approaches. Qinxun Bai, Zheng Wu 0003, Stan Sclaroff, Margrit Betke, Camille Monnier |
ICCV | 4 |
| 2013 | SAGE: An approach and implementation empowering quick and reliable quantitative analysis of segmentation qualityabstractFinding the outline of an object in an image is a fundamental step in many vision-based applications. It is important to demonstrate that the segmentation found accurately represents the contour of the object in the image. The discrepancy measure model for segmentation analysis focuses on selecting an appropriate discrepancy measure to compute a score that indicates how similar a query segmentation is to a gold standard segmentation. Observing that the score depends on the gold standard segmentation, we propose a framework that expands this approach by introducing the consideration of how to establish the gold standard segmentation. The framework shows how to obtain project-specific performance indicators in a principled way that links annotation tools, fusion methods, and evaluation algorithms into a unified model we call SAGE. We also describe a freely available implementation of SAGE that enables quick segmentation validation against either a single annotation or a fused annotation. Finally, three studies are presented to highlight the impact of annotation tools, an-notators, and fusion methods on establishing trusted gold standard segmentations for cell and artery images. Danna Gurari, Suele Ki Kim, Brett Isenberg, Tuan A. Pham 0002, Alberto Purwada, Patricia Solski, Matthew L. Walker, Joyce Y. Wong, Margrit Betke |
WACV | 10 |
| 2013 | The Kernel Semi-Least Squares Method for Sparse Distance ApproximationabstractWe extend the semi-least squares problem defined by Rao and Mitra ( 1971 ) to the kernel semi-least squares problem. We introduce subset projection, a technique that produces a solution to this problem. We show how the results of subset projection can be used to approximate a computationally expensive distance metric. Samuel Epstein 0001, Margrit Betke |
Neural Comput. | 2 |
| 2012 | Coupling detection and data association for multiple object trackingabstractWe present a novel framework for multiple object tracking in which the problems of object detection and data association are expressed by a single objective function. The framework follows the Lagrange dual decomposition strategy, taking advantage of the often complementary nature of the two subproblems. Our coupling formulation avoids the problem of error propagation from which traditional “detection-tracking approaches” to multiple object tracking suffer. We also eschew common heuristics such as “nonmaximum suppression” of hypotheses by modeling the joint image likelihood as opposed to applying independent likelihood assumptions. Our coupling algorithm is guaranteed to converge and can handle partial or even complete occlusions. Furthermore, our method does not have any severe scalability issues but can process hundreds of frames at the same time. Our experiments involve challenging, notably distinct datasets and demonstrate that our method can achieve results comparable to those of state-of-art approaches, even without a heavily trained object detector. Zheng Wu 0003, Ashwin Thangali, Stan Sclaroff, Margrit Betke |
CVPR | 4 |
| 2012 | Hierarchical Partial Matching and Segmentation of Interacting Cells
Zheng Wu 0003, Danna Gurari, Joyce Y. Wong, Margrit Betke |
MICCAI (1) | 4 |
| 2012 | Cell morphology classification and clutter mitigation in phase-contrast microscopy images using machine learning
Diane H. Theriault, Matthew L. Walker, Joyce Y. Wong, Margrit Betke |
Mach. Vis. Appl. | 4 |
| 2011 | Click control: improving mouse interaction for people with motor impairmentsabstractCamera-based mouse-replacement systems allow people with motor impairments to control the mouse pointer with head movements if they are unable to use their hands. To address the difficulties of accidental clicking and usable simulation of a real computer mouse, we developed Click Control, a tool to augment the functionality of these systems. When a user attempts to click, Click Control displays a form that allows him or her to cancel the click if it was accidental, or send different types of clicks with an easy-to-use gesture interface. Initial studies of a prototype with users with motor impairments showed that Click Control improved their mouse control experiences. Christopher Kwan, Isaac Paquette, John J. Magee, Paul Y. Lee, Margrit Betke |
ASSETS | 5 |
| 2011 | Enhancing social connections through automatically-generated online social network messagesabstractSocial isolation and loneliness are important challenges faced by people with certain physical disabilities. Technical and complexity issues may prevent some people from participating in online social networks that otherwise may address some of these issues. We propose to generate social network messages automatically from within assistive technology and augmentative and alternative communication software. These messages will help users post some of their daily activities with the software to online social networks. Based on our initial user studies, the inclusion of social networking connections may help improve engagement and interaction between users with disabilities and their friends, families and caregivers, and this increased interest can lead to the desire to use the assistive technology more fully. John J. Magee, Christopher Kwan, Margrit Betke, Fletcher Hietpas |
ASSETS | 3 |
| 2011 | Efficient track linking methods for track graphs using network-flow and set-cover techniquesabstractThis paper proposes novel algorithms that use network-flow and set-cover techniques to perform occlusion reasoning for a large number of small, moving objects in single or multiple views. We designed a track-linking framework for reasoning about short-term and long-term occlusions. We introduce a two-stage network-flow process to automatically construct a “track graph” that describes the track merging and splitting events caused by occlusion. To explain short-term occlusions, when local information is sufficient to distinguish objects, the process links trajectory segments through a series of optimal bipartite-graph matches. To resolve long-term occlusions, when global information is needed to characterize objects, the linking process computes a logarithmic approximation solution to the set cover problem. If multiple views are available, our method builds a track graph, independently for each view, and then simultaneously links track segments from each graph, solving a joint set cover problem for which a logarithmic approximation also exists. Through experiments on different datasets, we show that our proposed linear and integer optimization techniques make the track graph a particularly useful tool for tracking large groups of individuals in images. Zheng Wu 0003, Thomas H. Kunz, Margrit Betke |
CVPR | 3 |
| 2010 | Adaptive mappings for mouse-replacement interfacesabstractUsers of mouse-replacement interfaces may have difficulty conforming to the motion requirements of their interface system. We have observed users with severe motor disabilities who controlled the mouse pointer with a head tracking interface. Our analysis shows that some users may be able to move in some directions easier than other directions. We propose several mouse pointer mappings that adapt to the user's movement abilities. These mappings will take into account the user's motions in two-or three-dimensions to move the mouse pointer in the intended direction. John J. Magee, Samuel Epstein 0001, Eric S. Missimer, Margrit Betke |
ASSETS | 4 |
| 2010 | Customizable keyboardabstractCustomizable Keyboard is an on-screen keyboard designed to be flexible and expandable. Instead of giving the user a keyboard layout Customizable Keyboard allows the user to create a layout that is accommodating to the user's needs. Customizable Keyboard also allows the user to select from a variety of ways to interact with the keyboard including but not limited to using the mouse pointer to select keys and different types of scan based systems. Customizable Keyboard provides more functionality than a typical onscreen keyboard including the ability to control infrared devices such as TVs and send Twitter® Tweets. Eric S. Missimer, Samuel Epstein 0001, John J. Magee, Margrit Betke |
ASSETS | 4 |
| 2010 | HAIL: Hierarchical Adaptive Interface Layout
John J. Magee, Margrit Betke |
ICCHP (1) | 2 |
| 2010 | An Information Fusion Approach for Multiview Feature TrackingabstractWe propose an information fusion approach to tracking objects from different viewpoints that can detect and recover from tracking failures. We introduce a reliability measure that is a combination of terms associated with correlation-based template matching and the epipolar geometry of the cameras. The measure is computed to evaluate the performance of 2D trackers in each camera view and detect tracking failures. The 3D object trajectory is constructed using stereoscopy and evaluated to predict the next 3D position of the object. In case of track loss in one camera view, the projection of the predicted 3D position onto the image plane of this view is used to reinitialize the lost 2D tracker. We conducted experiments with 34 subjects to evaluate our proposed system on videos of facial feature movements during human-computer interaction. The system successfully detected feature loss and gave promising results on accurate re-initialization of the feature. Esra Ataer Cansizoglu, Margrit Betke |
ICPR | 2 |
| 2009 | Tracking a large number of objects from multiple viewsabstractWe propose a multi-object multi-camera framework for tracking large numbers of tightly-spaced objects that rapidly move in three dimensions. We formulate the problem of finding correspondences across multiple views as a multidimensional assignment problem and use a greedy randomized adaptive search procedure to solve this NP-hard problem efficiently. To account for occlusions, we relax the one-to-one constraint that one measurement corresponds to one object and iteratively solve the relaxed assignment problem. After correspondences are established, object trajectories are estimated by stereoscopic reconstruction using an epipolar-neighborhood search. We embedded our method into a tracker-to-tracker multi-view fusion system that not only obtains the three-dimensional trajectories of closely-moving objects but also accurately settles track uncertainties that could not be resolved from single views due to occlusion. We conducted experiments to validate our greedy assignment procedure and our technique to recover from occlusions. We successfully track hundreds of flying bats and provide an analysis of their group behavior based on 150 reconstructed 3D trajectories. Zheng Wu 0003, Nickolay I. Hristov, Tyson L. Hedrick, Thomas H. Kunz, Margrit Betke |
ICCV | 5 |
| 2008 | Tracking with Dynamic Hidden-State Shape Models
Zheng Wu 0003, Margrit Betke, Jingbin Wang, Vassilis Athitsos, Stan Sclaroff |
ECCV (1) | 2 |
| 2008 | Detecting Objects of Variable Shape Structure With Hidden State Shape ModelsabstractThis paper proposes a method for detecting object classes that exhibit variable shape structure in heavily cluttered images. The term "variable shape structure" is used to characterize object classes in which some shape parts can be repeated an arbitrary number of times, some parts can be optional, and some parts can have several alternative appearances. Hidden State Shape Models (HSSMs), a generalization of Hidden Markov Models (HMMs), are introduced to model object classes of variable shape structure using a probabilistic framework. A polynomial inference algorithm automatically determines object location, orientation, scale and structure by finding the globally optimal registration of model states with the image features, even in the presence of clutter. Experiments with real images demonstrate that the proposed method can localize objects of variable shape structure with high accuracy. For the task of hand shape localization and structure identification, the proposed method is significantly more accurate than previously proposed methods based on chamfer-distance matching. Furthermore, by integrating simple temporal constraints, the proposed method gains speed-ups of more than an order of magnitude, and produces highly accurate results in experiments on non-rigid hand motion tracking. Jingbin Wang, Vassilis Athitsos, Stan Sclaroff, Margrit Betke |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2008 | A Human-Computer Interface Using Symmetry Between Eyes to Detect Gaze DirectionabstractIn the cases of paralysis so severe that a person's ability to control movement is limited to the muscles around the eyes, eye movements or blinks are the only way for the person to communicate. Interfaces that assist in such communication are often intrusive, require special hardware, or rely on active infrared illumination. A nonintrusive communication interface system called EyeKeys was therefore developed, which runs on a consumer-grade computer with video input from an inexpensive Universal Serial Bus camera and works without special lighting. The system detects and tracks the person's face using multiscale template correlation. The symmetry between left and right eyes is exploited to detect if the person is looking at the camera or to the left or right side. The detected eye direction can then be used to control applications such as spelling programs or games. The game ldquoBlockEscaperdquo was developed to evaluate the performance of EyeKeys and compare it to a mouse substitution interface. Experiments with EyeKeys have shown that it is an easily used computer input and control device for able-bodied people and has the potential to become a practical tool for people with severe paralysis. John J. Magee, Margrit Betke, James Gips, Matthew R. Scott, Benjamin N. Waber |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2007 | Tracking Large Variable Numbers of Objects in ClutterabstractWe propose statistical data association techniques/or visual tracking of enormously large numbers of objects. We do not assume any prior knowledge about the numbers involved, and the objects may appear or disappear anywhere in the image frame and at any time in the sequence. Our approach combines the techniques of multitarget track initiation, recursive Bayesian tracking, clutter modeling, event analysis, and multiple hypothesis filtering. The original multiple hypothesis filter addresses an NP-hard problem and is thus not practical. We propose two cluster-based data association approaches that are linear in the number of detections and tracked objects. We applied the method to track wildlife in infrared video. We have successfully tracked hundreds of thousands of bats which were flying at high speeds and in dense formations. Margrit Betke, Diane H. Theriault, Angshuman Bagchi, Nickolay I. Hristov, Nicholas C. Makris, Thomas H. Kunz |
CVPR | 1 |
| 2007 | Block-Based MAP Disparity Estimation Under Alpha-Channel ConstraintsabstractDisparity estimation belongs to the most important, but difficult, problems in image processing and computer vision. Its importance stems from a wide range of applications, while its difficulty is related to ill-posedness. To date, numerous disparity estimation algorithms have been developed. In this paper, we consider a particular case of disparity estimation based on two views and a known alpha channel partitioning each view into foreground and background. The main idea is to use this partitioning in order to enhance disparity estimation in the foreground object close to its boundary. We propose a block-based disparity model with two alpha-channel constraints: a photometric one, disabling invalid intensity/color matches, and a geometric one, preventing disparity smoothing between foreground and background. We incorporate these constraints into a Bayesian framework using the maximum a posteriori probability criterion. We experimentally demonstrate improvements in the estimated disparities at foreground object boundaries, and show examples of image relighting using these disparities Peter J. McNerney, Janusz Konrad, Margrit Betke |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Detecting Instances of Shape Classes That Exhibit Variable Structure
Vassilis Athitsos, Jingbin Wang, Stan Sclaroff, Margrit Betke |
ECCV (1) | 4 |
| 2006 | Pulmonary fissure segmentation on CT
Jingbin Wang, Margrit Betke, Jane P. Ko |
Medical Image Anal. | 2 |
| 2005 | MosaicShape: Stochastic Region Grouping with Shape PriorabstractA method that combines shape-based object recognition and image segmentation is proposed for shape retrieval from images. Given a shape prior represented in a multi-scale curvature form, the proposed method identifies the target objects in images by grouping oversegmented image regions. The problem is formulated in a unified probabilistic framework, and object segmentation and recognition are accomplished simultaneously by a stochastic Markov Chain Monte Carlo (MCMC) mechanism. Within each sampling move during the simulation process, probabilistic region grouping operations are influenced by both the image information and the shape similarity constraint. The latter constraint is measured by a partial shape matching process. A generalized cluster sampling algorithm is presented in A. Barbu and S. Zhu (2003), combined with a large sampling jump and other implementation improvements, and greatly speeds up the overall stochastic process. The proposed method supports the segmentation and recognition of multiple occluded objects in images. Experimental results are provided for both synthetic and real images. Jingbin Wang, Erdan Gu, Margrit Betke |
CVPR (1) | 3 |
| 2005 | Tracking, Analysis, and Recognition of Human Gestures in VideoabstractAn overview of research in automated gesture spotting, tracking and recognition by the Image and Video Computing Group at Boston University is given. Approaches for localization and tracking human hands in video, estimation of hand shape and upper body pose; tracking head and facial motion, as well as efficient spotting and recognition of specific gestures in video streams are summarized. Methods for efficient dimensionality reduction of gesture time series, boosting of classifiers for nearest neighbor search in pose space, and model-based pruning of gesture alignment hypotheses are described. Algorithms are demonstrated in three domains: American sign language, hand signals like those employed by flight-directors on airport runways, and gesture-based interfaces for severely disabled users. The methods described are general and can be applied in other domains that require efficient detection and analysis of patterns in time-series, images or video. Stan Sclaroff, Margrit Betke, George Kollios, Jonathan Alon, Vassilis Athitsos, Rui Li 0053, John J. Magee, Tai-Peng Tian |
ICDAR | 2 |
| 2004 | Real-Time 4D Tumor Tracking and Modeling from Internal and External Fiducials in Fluoroscopy
Johanna Brewer, Margrit Betke, David P. Gierga, George T. Y. Chen |
MICCAI (2) | 2 |
| 2004 | Shape-Based Curve Growing Model and Adaptive Regularization for Pulmonary Fissure Segmentation in CT
Jingbin Wang, Margrit Betke, Jane P. Ko |
MICCAI (1) | 2 |
| 2003 | Landmark detection in the chest and registration of lung surfaces with an application to nodule registration
Margrit Betke, Harrison Hong, Deborah Thomas, Chekema Prince, Jane P. Ko |
Medical Image Anal. | 1 |
| 2002 | Evaluation of Tracking Methods for Human-Computer InteractionabstractTracking methods are evaluated in a real-time feature tracking system used for human-computer interaction (HCI). The Camera Mouse, a HCI system for people with severe disabilities that interprets video input to manipulate the mouse pointer, was improved and used as the test platform for this study. Tracking methods tested are the Lucas-Kanade tracker and a tracker based on normalized correlation. Both methods are evaluated with and without multidimensional Kalman filters. Two-, four-, and six-dimensional filters are tested to model feature location, velocity, and acceleration. The various tracker and filter combinations are evaluated for accuracy, computational efficiency, and practicality. The normalized correlation coefficient tracker without Kalman filtering is found to be the tracker best suited for a variety of HCI tasks. Christopher Fagiani, Margrit Betke, James Gips |
WACV | 2 |
| 2001 | Necessary Conditions to Attain Performance Bounds on Structure and Motion Estimates of Rigid ObjectsabstractAnalytic conditions that are necessary for the maximum likelihood estimate to become asymptotically unbiased and attain minimum variance are derived for estimation problems in computer vision. In particular, problems of estimating the parameters that describe the 3D structure of rigid objects or their motion are investigated. It is common practice to compute Cramer-Rao lower bounds (CRLB) to approximate the mean-square error in parameter estimation problems, but the CRLB is not guaranteed to be a tight bound and typically underestimates the true mean-square error. The necessary conditions for the Cramer-Rao lower bound to be a good approximation of the mean-squareerror are derived. The tightness of the bound depends on the noise level, the number of pixels on the surface of the object, and the texture of the surface. We examine our analytical results experimentally using polyhedral objects that consist of planar surface patches with various textures that move in 3D space. We provide necessary conditions for the CRLB to be attained that depend on the size, texture, and noise level of the surface patch. Margrit Betke, Eran Naftali, Nicholas C. Makris |
CVPR (2) | 1 |
| 2001 | Communication via Eye Blinks - Detection and Duration Analysis in Real TimeabstractA method for a real-time vision system that automatically detects a user's eye blinks and accurately measures their durations is introduced. The system is intended to provide an alternate input modality to allow people with severe disabilities to access a computer. Voluntary long blinks trigger mouse clicks, while involuntary short blinks are ignored. The system enables communication using "blink patterns:" sequences of long and short blinks which are interpreted as semiotic messages. The location of the eyes is determined automatically through the motion of the user's initial blinks. Subsequently, the eye is tracked by correlation across time, and appearance changes are automatically analyzed in order to classify the eye as either open or closed at each frame. No manual initialization, special lighting, or prior face detection is required. The system has been tested with interactive games and a spelling program. Results demonstrate overall detection accuracy of 95.61% and an average rate of 28 frames per second. Kristen Grauman, Margrit Betke, James Gips, Gary R. Bradski |
CVPR (1) | 2 |
| 2001 | Automatic 3D Registration of Lung Surfaces in Computed Tomography Scans
Margrit Betke, Harrison Hong, Jane P. Ko |
MICCAI | 1 |
| 2001 | Recognition, Resolution, and Complexity of Objects Subject to Affine Transformations
Margrit Betke, Nicholas C. Makris |
Int. J. Comput. Vis. | 1 |
| 2000 | Real-time multiple vehicle detection and tracking from a moving vehicle
Margrit Betke, Esin Haritaoglu, Larry Davis 0001 |
Mach. Vis. Appl. | 1 |
| 1999 | Detection of Pulmonary Nodules on Ct and Volumetric Assessment of Change over Time
Margrit Betke, Jane P. Ko |
MICCAI | 1 |
| 1999 | Piecemeal Graph Exploration by a Mobile RobotabstractWe study how a mobile robot can learn an unknown environment in a piecemeal manner. The robot's goal is to learn a complete map of its environment, while satisfying the constraint that it must return every so often to its starting position (for refueling, say). The environment is modeled as an arbitrary, undirected graph, which is initially unknown to the robot. We assume that the robot can distinguish vertices and edges that it has already explored. We present a surprisingly efficient algorithm for piecemeal learning an unknown undirected graph G=(V, E) in which the robot explores every vertex and edge in the graph by traversing at most O(E+V1+o(1)) edges. This nearly linear algorithm improves on the best previous algorithm, in which the robot traverses at most O(E+V2) edges. We also give an application of piecemeal learning to the problem of searching a graph for a “treasure.” Baruch Awerbuch, Margrit Betke, Ronald L. Rivest, Mona Singh 0001 |
Inf. Comput. | 2 |
| 1998 | Information-Conserving Object RecognitionabstractFollowing the theory of statistical estimation, the problem of recognizing objects imaged in complex real-world scenes is examined from a parametric perspective. A scalar measure of an object's complexity, which is invariant under affine transformation and changes in image noise level, is extracted from the object's Fisher information. The volume of Fisher information is shown to provide an overall statistical measure of the object's recognizability in a particular image, while the complexity provides an intrinsically physical measure that characterizes the object in any image. An information-conserving method is then developed for recognizing an object imaged in a complex scene. Here the term information-conserving means that the method uses all the measured data pertinent to the object's recognizability, attains the theoretical lower bound on estimation error for any unbiased estimate, and therefore is statistically optimal. This method is then successfully applied to finding objects imaged in thousands of complex real-world scenes. Margrit Betke, Nicholas C. Makris |
ICCV | 1 |
| 1997 | Mobile robot localization using landmarksabstractWe describe an efficient method for localizing a mobile robot in an environment with landmarks. We assume that the robot can identify these landmarks and measure their bearings relative to each other. Given such noisy input, the algorithm estimates the robot's position and orientation with respect to the map of the environment. The algorithm makes efficient use of our representation of the landmarks by complex numbers. The algorithm runs in time linear in the number of landmarks. We present results of simulations and propose how to use our method for robot navigation. Margrit Betke, Leonid Gurvits |
IEEE Trans. Robotics Autom. | 1 |
| 1995 | Piecemeal Graph Exploration by a Mobile Robot (Extended Abstract)abstract) Baruch Awerbuch y Margrit Betke Ronald L. Rivest Mona Singh Laboratory for Computer Science Massachusetts Institute of Technology Cambridge, MA 02139 Abstract We study the problem of learning a graph by piecemeal exploration, in which a mobile robot must return every so often to its starting point (for refueling, say). We assume that the robot can distinguish vertices and edges which it has already explored. We present an algorithm for piecemeal learning an unknown undirected graph G = (V; E) in which the robot explores every vertex and edge in G by traversing at most O(E + V 1+o(1) ) edges. This nearly linear algorithm improves on the best previous algorithm, in which the robot traverses at most O(E + V 2 ) edges. We also address the related problem of searching a graph for a particular distinguished location or treasure. If this location or treasure is known to be near the starting point, then the robot should search in a breadth-first manner from the starting point. We gi... Baruch Awerbuch, Margrit Betke, Ronald L. Rivest, Mona Singh 0001 |
COLT | 2 |
| 1995 | Fast Object Recognition in Noisy Images Using Simulated AnnealingabstractA fast simulated annealing algorithm is developed for automatic object recognition. The object recognition problem is addressed as the problem of best describing a match between a hypothesized object and an image. The normalized correlation coefficient is used as a measure of the match. Templates are generated on-line during the search by transforming model images. Simulated annealing reduces the search time by orders of magnitude with respect to an exhaustive search. The algorithm is applied to the problem of how landmarks, e.g., traffic signs, can be recognized by a navigating robot. We illustrate the performance of our algorithm with real-world images of complicated scenes with traffic signs. False positive matches occur only for templates with very small information content. To avoid false positive matches, we propose a method to select model images for robust object recognition by measuring the information content of the model images. The algorithm works well in noisy images for model images with high information content.> Margrit Betke, Nicholas C. Makris |
ICCV | 1 |
| 1995 | Piecemeal Learning of an Unknown Environment
Margrit Betke, Ronald L. Rivest, Mona Singh 0001 |
Mach. Learn. | 1 |
| 1994 | Mobile robot localization using landmarksabstractWe describe an efficient algorithm for localizing a mobile robot in an environment with landmarks. We assume that the robot has a camera and maybe other sensors that enable it to both identify landmarks and measure the angles subtended by these landmarks. We show how to estimate the robot's position using a new technique that involves a complex number representation of the landmarks. Our algorithm runs in time linear in the number of landmarks. We present results of our simulations and propose how to use our method for robot navigation.> Margrit Betke, Leonid Gurvits |
IROS | 1 |
| 1993 | Piecemeal Learning of an Unknown EnvironmentabstractWe introduce a new learning problem: learning a graph by piecemeal search, in which the learner must return every so often to its starting point (for refueling, say).We present two linear-time piecemeal-search algorithms for learning city-block graphs: grid graphs with rectangular obstacles. Margrit Betke, Ronald L. Rivest, Mona Singh 0001 |
COLT | 1 |