VLDB 2026 Research / reviewers in the wild / expert
Tanya Y. Berger-Wolf
dblp:b/TYBergerWolf
· DBLP profile ↗
66ranked-venue papers
7as first author
24since 2021 · last 2025
0000-0001-7610-1412ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 2 first-author · 17 since 2021Databases, data management, data science and information retrieval · 28 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Theory of computation · 8 · 2 first-author · 1 since 2021Systems, architecture and hardware · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained AnalysisabstractWe present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pretrained ViTs, such as DINO, have demonstrated remarkable capabilities in extracting localized, discriminative features. However, saliency maps like Grad-CAM often fail to identify these traits, producing blurred, coarse heatmaps that highlight entire objects instead. We propose a novel approach, Prompt Class Attention Map (Prompt-CAM), to address this limitation. Prompt-CAM learns class-specific prompts for a pre-trained ViT and uses the corresponding outputs for classification. To correctly classify an image, the true-class prompt must attend to unique image patches not present in other classes’ images (i.e., traits). As a result, the true class’s multi-head attention maps reveal traits and their locations. Implementation-wise, Prompt-CAM is almost a "free lunch," requiring only a modification to the prediction head of Visual Prompt Tuning (VPT). This makes Prompt-CAM easy to train and apply, in stark contrast to other interpretable methods that require designing specific models and training processes. Extensive empirical studies on a dozen datasets from various domains (e.g., birds, fishes, insects, fungi, flowers, food, and cars) validate the superior interpretation capability of Prompt-CAM. The source code and demo are available at https://github.com/Imageomics/Prompt_CAM. Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu, Kazi Sajeed Mehrab, Elizabeth G. Campolongo, Daniel I. Rubenstein, Charles V. Stewart, Anuj Karpatne, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao |
CVPR | 11 |
| 2025 | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from ImagesabstractWe introduce Fish-Visual Trait Analysis (Fish-Vista), the first organismal image dataset designed for the analysis of visual traits of aquatic species directly from images using machine learning and computer vision methods. Fish-Vista contains 69,269 annotated images spanning 4,316 fish species, curated and organized to serve three downstream tasks: species classification, trait identification, and trait segmentation. Our work makes two key contributions. First, we provide a fully reproducible data processing pipeline to process fish images sourced from various museum collections, contributing to the advancement of AI in biodiversity science. We annotate the images with carefully curated labels from biological databases and manual annotations to create an AI-ready dataset of visual traits. Second, our work offers fertile grounds for researchers to develop novel methods for a variety of problems in computer vision such as handling long-tailed distributions, out-of-distribution generalization, learning with weak labels, explainable AI, and segmenting small objects. Dataset and code for Fish-Vista are available at https://github.com/Imageomics/Fish-Vista Kazi Sajeed Mehrab, M. Maruf, Arka Daw, Abhilash Neog, Harish Babu Manogaran, Mridul Khurana, Zhenyang Feng, Bahadir Altintas, Yasin Bakis, Elizabeth G. Campolongo, Matthew J. Thompson, Hilmar Lapp, Tanya Y. Berger-Wolf, Paula M. Mabee, Henry L. Bart Jr., Wei-Lun Chao, Wasila M. Dahdul, Anuj Karpatne |
CVPR | 14 |
| 2025 | Finer-CAM: Spotting the Difference Reveals Finer Details for Visual ExplanationabstractClass activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often struggles to identify discriminative regions that distinguish visually similar fine-grained classes. Prior efforts address this limitation by introducing more sophisticated explanation processes, but at the cost of extra complexity. In this paper, we propose Finer-CAM, a method that retains CAM’s efficiency while achieving precise localization of discriminative regions. Our key insight is that the deficiency of CAM lies not in "how" it explains, but in "what" it explains. Specifically, previous methods attempt to identify all cues contributing to the target class’s logit value, which inadvertently also activates regions predictive of visually similar classes. By explicitly comparing the target class with similar classes and spotting their differences, Finer-CAM suppresses features shared with other classes and emphasizes the unique, discriminative details of the target class. Finer-CAM is easy to implement, compatible with various CAM methods, and can be extended to multi-modal models for accurate localization of specific concepts. Additionally, Finer-CAM allows adjustable comparison strength, enabling users to selectively highlight coarse object contours or fine discriminative details. Quantitatively, we show that masking out the top 5% of activated pixels by Finer-CAM results in a larger relative confidence drop compared to baselines. The source code and demo are available at https://github.com/Imageomics/Finer-CAM. Jianyang Gu, Arpita Chowdhury, Zheda Mai, David Carlyn, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao |
CVPR | 6 |
| 2025 | What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary TraitsabstractA grand challenge in biology is to discover evolutionary traits---features of organisms common to a group of species with a shared ancestor in the tree of life (also referred to as phylogenetic tree). With the growing availability of image repositories in biology, there is a tremendous opportunity to discover evolutionary traits directly from images in the form of a hierarchy of prototypes. However, current prototype-based methods are mostly designed to operate over a flat structure of classes and face several challenges in discovering hierarchical prototypes, including the issue of learning over-specific prototypes at internal nodes. To overcome these challenges, we introduce the framework of Hierarchy aligned Commonality through Prototypical Networks (HComP-Net). The key novelties in HComP-Net include a novel over-specificity loss to avoid learning over-specific prototypes, a novel discriminative loss to ensure prototypes at an internal node are absent in the contrasting set of species with different ancestry, and a novel masking module to allow for the exclusion of over-specific prototypes at higher levels of the tree without hampering classification performance. We empirically show that HComP-Net learns prototypes that are accurate, semantically consistent, and generalizable to unseen species in comparison to baselines. Our code is publicly accessible at Imageomics Institute Github site: https://github.com/Imageomics/HComPNet. Harish Babu Manogaran, M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Caleb Charpentier, Josef C. Uyeda, Wasila M. Dahdul, Matthew J. Thompson, Elizabeth G. Campolongo, Kaiya Provost, Wei-Lun Chao, Tanya Y. Berger-Wolf, Paula M. Mabee, Hilmar Lapp, Anuj Karpatne |
ICLR | 12 |
| 2025 | Edge-Native, Behavior-Adaptive Drone System for Wildlife MonitoringabstractWildlife monitoring with drones must balance competing demands: approaching close enough to capture behaviorally-relevant video while avoiding stress responses that compromise animal welfare and data validity. Human operators face a fundamental attentional bottleneck: they cannot simultaneously control drone operations and monitor vigilance states across entire animal groups. By the time elevated vigilance becomes obvious, an adverse flee response by the animals may be unavoidable. To solve this challenge, we present an edge-native, behavior-adaptive drone system for wildlife monitoring. This configurable decision-support system augments operator expertise with automated group-level vigilance monitoring. Our system continuously tracks individual behaviors using YOLOv11m detection and YOLO-Behavior classification, aggregates vigilance states into a real-time group stress metric, and provides graduated alerts (alert vigilance → flee response) with operator-tunable thresholds for context-specific calibration. We derive service-level objectives (SLOs) from video frame rates and behavioral dynamics: to monitor 30fps video streams in real-time, our system must complete detection and classification within 33ms per frame. Our edge-native pipeline achieves 23.8ms total inference on GPU-accelerated hardware, meeting this constraint with a substantial margin. Retrospective analysis of seven wildlife monitoring missions demonstrates detection capability and quantifies the cost of reactive control: manual piloting results in 14 seconds average adverse behavior duration with 71.9% usable frames. Our analysis reveals operators could have received actionable alerts 51s before animals fled in 57% of missions. Simulating 5-second operator intervention yields a projected performance of 82.8% usable frames with 1-second adverse behavior duration, a 93% reduction compared to manual piloting. Jenna Kline, Rugved Katole, Tanya Y. Berger-Wolf, Christopher Stewart |
SEC | 3 |
| 2025 | BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive LearningabstractFoundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space. Jianyang Gu, Samuel Stevens 0001, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Jiaman Wu, Andrei Kopanev, Zheda Mai, Alexander E. White, James P. Balhoff, Wasila M. Dahdul, Daniel I. Rubenstein, Hilmar Lapp, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001 |
NeurIPS | 14 |
| 2025 | OmniMesh: Addressing Findability Challenges in Distributed Nature Data Repositories
Arnab Nandi 0001, Wei-Lun Chao, Rongjun Qin, Carl Boettiger, Hilmar Lapp, Tanya Y. Berger-Wolf |
SSDBM | 6 |
| 2025 | Adapting the Re-ID Challenge for Static SensorsabstractABSTRACT The Grévy's zebra, an endangered species native to Kenya and southern Ethiopia, has been the target of sustained conservation efforts in recent years. Accurately monitoring Grévy's zebra populations is essential for ecologists to evaluate ongoing conservation initiatives. Recently, in both 2016 and 2018, a full census of the Grévy's zebra population was enabled by the Great Grévy's Rally (GGR), a citizen science event that combines teams of volunteers to capture data with computer vision algorithms that help experts estimate the number of individuals in the population. A complementary, scalable, cost‐effective and long‐term Grévy's population monitoring approach involves deploying a network of camera traps, which we have done at the Mpala Research Centre in Laikipia County, Kenya. In both scenarios, a substantial majority of the images of zebras are not usable for individual identification due to ‘in‐the‐wild’ imaging conditions—occlusions from vegetation or other animals, oblique views, low image quality and animals that appear in the far background and are thus too small to identify. Camera trap images, without an intelligent human photographer to select the framing and focus on the animals of interest, are of even poorer quality, with high rates of occlusion and high spatiotemporal similarity within image bursts. We employ an image filtering pipeline incorporating animal detection, species identification, viewpoint estimation, quality evaluation and temporal subsampling to compensate for these factors and obtain individual crops from camera trap and GGR images of suitable quality for re‐ID. We then employ the local clusterings and their alternatives (LCA) algorithm, a hybrid computer vision and graph clustering method for animal re‐ID, on the resulting high‐quality crops. Our method processed images taken during GGR‐16 and GGR‐18 in Meru County, Kenya, into 4142 highly comparable annotations, requiring only 120 contrastive same‐vs‐different‐individual decisions from a human reviewer to produce a population estimate of 349 individuals (within 4.6 of the ground truth count in Meru County). Our method also efficiently processed 8.9M unlabelled camera trap images from 70 camera traps at Mpala over 2 years into 685 encounters of 173 unique individuals, requiring only 331 contrastive decisions from a human reviewer. Avirath Sundaresan, Jason Parham, Jonathan P. Crall, Rosemary Warungu, Timothy Muthami, Jackson Miliko, Margaret Mwangi, Jason Holmberg, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Charles V. Stewart, Sara Beery |
IET Comput. Vis. | 9 |
| 2025 | BaboonLand Dataset: Tracking Primates in the Wild and Automating Behaviour Recognition from Drone Videos
Isla Duporge, Maksim Kholiavchenko, Roi Harel, Scott Wolf, Daniel I. Rubenstein, Margaret Crofoot, Tanya Y. Berger-Wolf, Stephen J. Lee, Julie Barreau, Jenna Kline, Michelle Ramirez, Charles V. Stewart |
Int. J. Comput. Vis. | 7 |
| 2025 | Correction: BaboonLand Dataset: Tracking Primates in the Wild and Automating Behaviour Recognition from Drone Videos
Isla Duporge, Maksim Kholiavchenko, Roi Harel, Scott Wolf, Daniel I. Rubenstein, Margaret Crofoot, Tanya Y. Berger-Wolf, Stephen J. Lee, Julie Barreau, Jenna Kline, Michelle Ramirez, Charles V. Stewart |
Int. J. Comput. Vis. | 7 |
| 2025 | Deep dive into KABR: a dataset for understanding ungulate behavior from in-situ drone video
Maksim Kholiavchenko, Jenna Kline, Maksim Kukushkin, Otto Brookes, Samuel Stevens 0001, Isla Duporge, Alec Sheets, Reshma Ramesh Babu, Namrata Banerji, Elizabeth G. Campolongo, Matthew J. Thompson, Nina Van Tiel, Jackson Miliko, Eduardo Bessa, Majid Mirmehdi, Thomas Schmid 0003, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Tilo Burghardt, Charles V. Stewart |
Multim. Tools Appl. | 17 |
| 2024 | Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
Vardaan Pahuja, Weidi Luo, Yu Gu 0016, Cheng-Hao Tu 0001, Hong-You Chen, Tanya Y. Berger-Wolf, Charles V. Stewart, Song Gao 0001, Wei-Lun Chao, Yu Su 0001 |
CIKM | 6 |
| 2024 | BioCLIP: A Vision Foundation Model for the Tree of LifeabstractImages of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an ex-plosion of computational methods and tools, particularly computer vision, for extracting biologically relevant information from images for science and conservation. Yet most of these are bespoke approaches designed for a specific task and are not easily adaptable or extendable to new questions, contexts, and datasets. A vision model for general or-ganismal biology questions on images is of timely need. To approach this, we curate and release Tree Of Life-10m, the largest and most diverse ML-ready dataset of biology images. We then develop Bioclip, a foundation model for the tree of life, leveraging the unique properties of bi-ology captured by Treeoflife-10m, namely the abun-dance and variety of images of plants, animals, and fungi, together with the availability of rich structured biological knowledge. We rigorously benchmark our approach on di-verse fine-grained biology classification tasks and find that BloCLIP consistently and substantially outperforms existing baselines (by 16% to 17% absolute). Intrinsic evaluation reveals that BloCLIP has learned a hierarchical representation conforming to the tree of life, shedding light on its strong generalizability.11imageomics.github.io/bioclip has models, data and code. Samuel Stevens 0001, Jiaman Wu, Matthew J. Thompson, Elizabeth G. Campolongo, Chan Hee Song, David Carlyn, Wasila M. Dahdul, Charles V. Stewart, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001 |
CVPR | 10 |
| 2024 | Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species EvolutionabstractAbstract A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion ) Mridul Khurana, Arka Daw, M. Maruf, Josef C. Uyeda, Wasila M. Dahdul, Caleb Charpentier, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Anuj Karpatne |
ECCV (89) | 14 |
| 2024 | A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisabstractWe present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn ''class-specific'' queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via ''multi-head'' cross-attention, INTR could identify different ''attributes'' of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR. Dipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang, David Carlyn, Samuel Stevens 0001, Kaiya Provost, Anuj Karpatne, Bryan Carstens, Daniel I. Rubenstein, Charles V. Stewart, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao |
ICLR | 12 |
| 2024 | Characterizing and Modeling AI-Driven Animal Ecology Studies at the EdgeabstractPlatforms that run artificial intelligence (AI) pipelines on edge computing resources are transforming the fields of animal ecology and biodiversity, enabling novel wildlife studies in animals' natural habitats. With emerging remote sensing hardware, e.g., camera traps and drones, and sophisticated AI models in situ, edge computing will be more significant in future AI-driven animal ecology (ADAE) studies. However, the study's objectives, the species of interest, its behaviors, range, and habitat, and camera placement affect the demand for edge resources at runtime. If edge resources are under-provisioned, studies can miss opportunities to adapt the settings of camera traps and drones to improve the quality and relevance of captured data. This paper presents salient features of ADAE studies that can be used to model latency, throughput objectives, and provision edge resources. Drawing from studies that span over fifty animal species, four geographic locations, and multiple remote sensing methods, we characterized common patterns in ADAE studies, revealing increasingly complex workflows involving various computer vision tasks with strict service level objectives (SLO). ADAE workflow demands will soon exceed individual edge devices' compute and memory resources, requiring multiple networked edge devices to meet performance demands. We developed a framework to scale traces from prior studies and replay them offline on representative edge platforms, allowing us to capture throughput and latency data across edge configurations. We used the data to calibrate queuing and machine learning models that predict performance on unseen edge configurations, achieving errors as low as 19%. Jenna Kline, Austin O'Quinn, Tanya Y. Berger-Wolf, Christopher Stewart |
SEC | 3 |
| 2024 | AI for Nature: From Science to ImpactabstractComputation has fundamentally changed the way we study nature.New data collection technologies, such as GPS, high-definition cameras, autonomous vehicles under water, on the ground, and in the air, genotyping, acoustic sensors, and crowdsourcing, are generating data about life on the planet that are orders of magnitude richer than any previously collected.Yet, our ability to extract insight from this data lags substantially behind our ability to collect it.The need for understanding is more urgent than ever and the challenges are great.We are in the middle of the 6th extinction, losing the planet's biodiversity at an unprecedented rate and scale.In many cases, we do not even have the basic numbers of what species we are losing, which impacts our ability to understand biodiversity loss drivers, predict the impact on ecosystems, and implement policy.From the basic science perspective, the new data opens the possibility of understanding function of traits of organisms and ecosystems, which is critical for biologists to predict effects of environmental change or genetic manipulation and to understand the significance of patterns in the four-billion-year evolutionary history of life.The key to unlocking the potential of this data are machine learning (ML) and artificial intelligence (AI) methods, which are already beginning to have significant impacts on research across ecology and conservation.AI can turn data into high resolution information source about living organisms, enabling scientific inquiry, conservation, and policy decisions.The talk introduces a new field of science, imageomics, and presents a vision and examples of AI as a trustworthy partner both in science and biodiversity conservation, discussing opportunities and challenges. Tanya Y. Berger-Wolf |
KDD | 1 |
| 2024 | Fine-Tuning is Fine, if CalibratedabstractFine-tuning is arguably the most straightforward way to tailor a pre-trained model (e.g., a foundation model) to downstream applications, but it also comes with the risk of losing valuable knowledge the model had learned in pre-training. For example, fine-tuning a pre-trained classifier capable of recognizing a large number of classes to master a subset of classes at hand is shown to drastically degrade the model's accuracy in the other classes it had previously learned. As such, it is hard to further use the fine-tuned model when it encounters classes beyond the fine-tuning data. In this paper, we systematically dissect the issue, aiming to answer the fundamental question, "What has been damaged in the fine-tuned model?" To our surprise, we find that the fine-tuned model neither forgets the relationship among the other classes nor degrades the features to recognize these classes. Instead, the fine-tuned model often produces more discriminative features for these other classes, even if they were missing during fine-tuning! What really hurts the accuracy is the discrepant logit scales between the fine-tuning classes and the other classes, implying that a simple post-processing calibration would bring back the pre-trained model's capability and at the same time unveil the feature improvement over all classes. We conduct an extensive empirical study to demonstrate the robustness of our findings and provide preliminary explanations underlying them, suggesting new directions for future theoretical analysis. Zheda Mai, Arpita Chowdhury, Ping Zhang 0016, Cheng-Hao Tu 0001, Hong-You Chen, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 7 |
| 2024 | VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological ImagesabstractImages are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of $12$ state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of $469K$ question-answer pairs involving $30K$ images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images. M. Maruf, Arka Daw, Kazi Sajeed Mehrab, Harish Babu Manogaran, Abhilash Neog, Medha Sawhney, Mridul Khurana, James P. Balhoff, Yasin Bakis, Bahadir Altintas, Matthew J. Thompson, Elizabeth G. Campolongo, Josef C. Uyeda, Hilmar Lapp, Henry L. Bart Jr., Paula M. Mabee, Yu Su 0001, Wei-Lun Chao, Charles V. Stewart, Tanya Y. Berger-Wolf, Wasila M. Dahdul, Anuj Karpatne |
NeurIPS | 20 |
| 2023 | Discovering Novel Biological Traits From Images Using Phylogeny-Guided Neural NetworksabstractDiscovering evolutionary traits that are heritable across species on the tree of life (also referred to as a phylogenetic tree) is of great interest to biologists to understand how organisms diversify and evolve. However, the measurement of traits is often a subjective and labor-intensive process, making trait discovery a highly label-scarce problem. We present a novel approach for discovering evolutionary traits directly from images without relying on trait labels. Our proposed approach, Phylo-NN, encodes the image of an organism into a sequence of quantized feature vectors -or codes- where different segments of the sequence capture evolutionary signals at varying ancestry levels in the phylogeny. We demonstrate the effectiveness of our approach in producing biologically meaningful results in a number of downstream tasks including species image generation and species-to-species image translation, using fish species as a target example Mohannad Elhamod, Mridul Khurana, Harish Babu Manogaran, Josef C. Uyeda, Meghan A. Balk, Wasila M. Dahdul, Yasin Bakis, Henry L. Bart Jr., Paula M. Mabee, Hilmar Lapp, James P. Balhoff, Caleb Charpentier, David Carlyn, Wei-Lun Chao, Charles V. Stewart, Daniel I. Rubenstein, Tanya Y. Berger-Wolf, Anuj Karpatne |
KDD | 17 |
| 2023 | Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target DataabstractWe propose a learning problem involving adapting a pre-trained source model to the target domain for classifying all classes that appeared in the source data, using target data that covers only a partial label space. This problem is practical, as it is unrealistic for the target end-users to collect data for all classes prior to adaptation. However, it has received limited attention in the literature. To shed light on this issue, we construct benchmark datasets and conduct extensive experiments to uncover the inherent challenges. We found a dilemma --- on the one hand, adapting to the new target domain is important to claim better performance; on the other hand, we observe that preserving the classification accuracy of classes missing in the target adaptation data is highly challenging, let alone improving them. To tackle this, we identify two key directions: 1) disentangling domain gradients from classification gradients, and 2) preserving class relationships. We present several effective solutions that maintain the accuracy of the missing classes and enhance the overall performance, establishing solid baselines for holistic transfer of pre-trained models with partial target data. Cheng-Hao Tu 0001, Hong-You Chen, Zheda Mai, Jike Zhong, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 6 |
| 2023 | Maximizing Coverage While Ensuring Fairness: A Tale of Conflicting ObjectivesabstractEnsuring fairness in computational problems has emerged as a $key$ topic during recent years, buoyed by considerations for equitable resource distributions and social justice. It $is$ possible to incorporate fairness in computational problems from several perspectives, such as using optimization, game-theoretic or machine learning frameworks. In this paper we address the problem of incorporation of fairness from a $combinatorial$ $optimization$ perspective. We formulate a combinatorial optimization framework, suitable for analysis by researchers in approximation algorithms and related areas, that incorporates fairness in maximum coverage problems as an interplay between $two$ conflicting objectives. Fairness is imposed in coverage by using coloring constraints that $minimizes$ the discrepancies between number of elements of different colors covered by selected sets; this is in contrast to the usual discrepancy minimization problems studied extensively in the literature where (usually two) colors are $not$ given $a$ $priori$ but need to be selected to minimize the maximum color discrepancy of $each$ individual set. Our main results are a set of randomized and deterministic approximation algorithms that attempts to $simultaneously$ approximate both fairness and coverage in this framework. Abolfazl Asudeh, Tanya Y. Berger-Wolf, Bhaskar DasGupta, Anastasios Sidiropoulos |
Algorithmica | 2 |
| 2021 | Understanding the Dynamics between Vaping and Cannabis Legalization Using Twitter Opinions
Shishir Adhikari, Akshay Uppal, Robin Mermelstein, Tanya Y. Berger-Wolf, Elena Zheleva |
ICWSM | 4 |
| 2021 | Variable-lag Granger Causality and Transfer Entropy for Time Series AnalysisabstractGranger causality is a fundamental technique for causal inference in time series data, commonly used in the social and biological sciences. Typical operationalizations of Granger causality make a strong assumption that every time point of the effect time series is influenced by a combination of other time series with a fixed time delay. The assumption of fixed time delay also exists in Transfer Entropy, which is considered to be a non-linear version of Granger causality. However, the assumption of the fixed time delay does not hold in many applications, such as collective behavior, financial markets, and many natural phenomena. To address this issue, we develop Variable-lag Granger causality and Variable-lag Transfer Entropy, generalizations of both Granger causality and Transfer Entropy that relax the assumption of the fixed time delay and allow causes to influence effects with arbitrary time delays. In addition, we propose methods for inferring both Variable-lag Granger causality and Transfer Entropy relations. In our approaches, we utilize an optimal warping path of Dynamic Time Warping to infer variable-lag causal relations. We demonstrate our approaches on an application for studying coordinated collective behavior and other real-world casual-inference datasets and show that our proposed approaches perform better than several existing methods in both simulated and real-world datasets. Our approaches can be applied in any domain of time series analysis. The software of this work is available in the R-CRAN package: VLTimeCausality. C. Amornbunchornvej, Elena Zheleva, Tanya Y. Berger-Wolf |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | Framework for Inferring Following Strategies from Time Series of Movement DataabstractHow do groups of individuals achieve consensus in movement decisions? Do individuals follow their friends, the one predetermined leader, or whomever just happens to be nearby? To address these questions computationally, we formalize C oordination S trategy I nference P roblem . In this setting, a group of multiple individuals moves in a coordinated manner toward a target path. Each individual uses a specific strategy to follow others (e.g., nearest neighbors, pre-defined leaders, and preferred friends). Given a set of time series that includes coordinated movement and a set of candidate strategies as inputs, we provide the first methodology (to the best of our knowledge) to infer whether each individual uses local-agreement system or dictatorship-like strategy to achieve movement coordination at the group level. We evaluate and demonstrate the performance of the proposed framework by predicting directions of movement of an individual in a group in both simulated datasets as well as in two real-world datasets: a school of fish and a troop of baboons. Moreover, since there is no prior methodology for inferring individual-level strategies, we compare our framework with the state-of-the-art approach for the task of classification of group-level-coordination models. Results show that our approach is highly accurate in inferring correct strategies in simulated datasets even in complicated mixed strategy settings, which no existing method can infer. In the task of classification of group-level-coordination models, our framework performs better than the state-of-the-art approach in all datasets. Animal data experiments show that fish, as expected, follow their neighbors, while baboons have a preference to follow specific individuals. Our methodology generalizes to arbitrary time series data of real numbers, beyond movement data. C. Amornbunchornvej, Tanya Y. Berger-Wolf |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Predictive temporal embedding of dynamic graphsabstractIn recent years, substantial effort has been devoted to learning to represent the static graphs and their substructures. A few studies explored utilizing temporal information available in a dynamic setting in order to address the node representation learning. However, the representation learning problem for the entire graph in a dynamic context is yet to be addressed. In this paper, we propose an unsupervised encoder-decoder framework that projects a dynamic graph at each time step into a d-dimensional space, taking into account both the graph's topology and dynamics. We investigate two different strategies. First, we address the representation learning problem by auto-encoding the graph dynamics. Second, we formulate a graph prediction problem and enforce the encoder to learn the representation that an autoregressive decoder then uses to predict the future of a dynamic graph. Gated graph neural networks (GGNNs) are incorporated to learn the topology of the graph at each time step and Long short-term memory networks (LSTMs) are leveraged to propagate the temporal information among the nodes through time. We demonstrate the efficacy of our approach with a graph classification task using two real-world datasets of animal behaviour and brain networks. Aynaz Taheri, Tanya Y. Berger-Wolf |
ASONAM | 2 |
| 2019 | Variable-Lag Granger Causality for Time Series AnalysisabstractGranger causality is a fundamental technique for causal inference in time series data, commonly used in the social and biological sciences. Typical operationalizations of Granger causality make a strong assumption that every time point of the effect time series is influenced by a combination of other time series with a fixed time delay. However, the assumption of the fixed time delay does not hold in many applications, such as collective behavior, financial markets, and many natural phenomena. To address this issue, we develop variable-lag Granger causality, a generalization of Granger causality that relaxes the assumption of the fixed time delay and allows causes to influence effects with arbitrary time delays. In addition, we propose a method for inferring variable-lag Granger causality relations. We demonstrate our approach on an application for studying coordinated collective behavior and show that it performs better than several existing methods in both simulated and real-world datasets. Our approach can be applied in any domain of time series analysis. C. Amornbunchornvej, Elena Zheleva, Tanya Y. Berger-Wolf |
DSAA | 3 |
| 2019 | Three-D Safari: Learning to Estimate Zebra Pose, Shape, and Texture From Images "In the Wild"abstractWe present the first method to perform automatic 3D pose, shape and texture capture of animals from images acquired in-the-wild. In particular, we focus on the problem of capturing 3D information about Grevy's zebras from a collection of images. The Grevy's zebra is one of the most endangered species in Africa, with only a few thousand individuals left. Capturing the shape and pose of these animals can provide biologists and conservationists with information about animal health and behavior. In contrast to research on human pose, shape and texture estimation, training data for endangered species is limited, the animals are in complex natural scenes with occlusion, they are naturally camouflaged, travel in herds, and look similar to each other. To overcome these challenges, we integrate the recent SMAL animal model into a network-based regression pipeline, which we train end-to-end on synthetically generated images with pose, shape, and background variation. Going beyond state-of-the-art methods for human shape and pose estimation, our method learns a shape space for zebras during training. Learning such a shape space from images using only a photometric loss is novel, and the approach can be used to learn shape in other settings with limited 3D supervision. Moreover, we couple 3D pose and shape prediction with the task of texture synthesis, obtaining a full texture map of the animal from a single image. We show that the predicted texture map allows a novel per-instance unsupervised optimization over the network features. This method, SMALST (SMAL with learned Shape and Texture) goes beyond previous work, which assumed manual keypoints and/or segmentation, to regress directly from pixels to 3D animal shape, pose and texture. Silvia Zuffi, Angjoo Kanazawa, Tanya Y. Berger-Wolf, Michael J. Black |
ICCV | 3 |
| 2018 | Mining and Modeling Complex Leadership Dynamics of Movement dataabstractLeadership is an essential part of collective decision and organization in social animals, including humans. In nature, leadership is dynamic and varies with context or temporal factors. Understanding dynamics of leadership, such as how leaders change, emerge, or converge, allows scientists to gain more insight into group decision-making and collective behavior in general. However, given only data of individual activities, it is challenging to infer these dynamic leadership events. In this paper, we focus on mining and modeling frequent patterns of leadership dynamics. We formalize a new computational problem, Mining Patterns Of Leadership Dynamics, as well as propose a framework as a solution of this problem. Our framework can be used to address several questions regarding leadership dynamics of group movement. We use the leadership inference framework, mFLICA, to infer the time series of leaders from movement datasets, then propose the approach to mine and model frequent patterns of leadership dynamics. We evaluate our framework performance by using several simulated datasets, as well as using the real-world dataset of baboon movement to demonstrate the application of our framework. There are no existing methods to address this problem, thus, we modify and extend the existing leadership inference framework to provide a non-trivial baseline. Our framework performs better than this baseline in all datasets. Moreover, we also propose a method to perform statistical significance tests, comparing inferred frequent patterns of leadership dynamics with our proposed null hypotheses. Our framework opens the opportunities for scientists to generate scientific hypotheses that can be tested statistically regarding dynamics of leadership in movement data. C. Amornbunchornvej, Tanya Y. Berger-Wolf |
ASONAM | 2 |
| 2018 | A Game-Theoretic Adversarial Approach to Dynamic Network Prediction
Vena Jia Li, Brian D. Ziebart, Tanya Y. Berger-Wolf |
PAKDD (3) | 3 |
| 2018 | Framework for Inferring Leadership Dynamics of Complex Movement from Time SeriesabstractLeadership plays a key role in social animals, including humans, decision-making and coalescence in coordinated activities such as hunting, migration, sport, diplomatic negotiation etc. In these coordinated activities, leadership is a process that organizes interactions among members to make a group achieve collective goals. Understanding initiation of coordinated activities allows scientists to gain more insight into social species behaviors. However, by using only time series of activities data, inferring leadership as manifested by the initiation of coordinated activities faces many challenging issues. First, coordinated activities are dynamic and are changing over time. Second, several different coordinated activities might occur simultaneously among subgroups. Third, there is no fundamental concept to describe these activities computationally. In this paper, we formalize Faction Initiator Inference Problem and propose a leadership inference framework as a solution of this problem. The framework makes no assumption about the characteristics of a leader or the parameters of the coordination process. The framework performs better than our non-trivial baseline in both simulated and biological datasets (schools of fish). Moreover, we demonstrate the application of our framework as a tool to study group merging and splitting dynamics on another biological dataset of trajectories of wild baboons. In addition, our problem formalization and framework enable opportunities for scientists to analyze coordinated activities and generate scientific hypotheses about collective behaviors that can be tested statistically and in the field. C. Amornbunchornvej, Tanya Y. Berger-Wolf |
SDM | 2 |
| 2018 | A Biology-themed Introductory CS Course at a Large, Diverse Public UniversityabstractWe present the curriculum and evaluation of a pilot Biology-themed CS1 course offering at a large public university. Inspired by Harvey Mudd's CS 5 Green, we adapt CS1 + Bio to fit the needs of our student body, which is much more typical for those US institutions that produce the bulk of the nation's CS undergraduate degrees. This course was team-taught by a computer science professor and a biology professor, and combined typical CS1 topics with relevant biology content. Our initial offering attracted students who would not otherwise have taken CS1, and was the only one of our three CS1 courses where more students reported planning to major in CS after the course than before it. Tanya Y. Berger-Wolf, Boris Igic, Cynthia Bagier Taylor, Robert H. Sloan, Rachel Poretsky |
SIGCSE | 1 |
| 2018 | An Animal Detection Pipeline for IdentificationabstractThis paper proposes a 5-component detection pipeline for use in a computer vision-based animal recognition system. The end result of our proposed pipeline is a collection of novel annotations of interest (AoI) with species and view-point labels. These AoIs, for example, could be fed as the focused input data into an appearance-based animal identification system. The goal of our method is to increase the reliability and automation of animal censusing studies and to provide better ecological information to conservationists. Our method is able to achieve a localization mAP of 81.67%, a species and viewpoint annotation classification accuracy of 94.28% and 87.11%, respectively, and an AoI accuracy of 72.75% across 6 animal species of interest. We also introduce the Wildlife Image and Localization Dataset (WILD), which contains 5,784 images and 12,007 labeled annotations across 28 classification species and a variety of challenging, real-world detection scenarios. Jason R. Parham, Charles V. Stewart, Jonathan P. Crall, Daniel I. Rubenstein, Jason Holmberg, Tanya Y. Berger-Wolf |
WACV | 6 |
| 2018 | Coordination Event Detection and Initiator Identification in Time Series DataabstractBehavior initiation is a form of leadership and is an important aspect of social organization that affects the processes of group formation, dynamics, and decision-making in human societies and other social animal species. In this work, we formalize the C oordination I nitiator I nference P roblem and propose a simple yet powerful framework for extracting periods of coordinated activity and determining individuals who initiated this coordination, based solely on the activity of individuals within a group during those periods. The proposed approach, given arbitrary individual time series, automatically (1) identifies times of coordinated group activity, (2) determines the identities of initiators of those activities, and (3) classifies the likely mechanism by which the group coordination occurred, all of which are novel computational tasks. We demonstrate our framework on both simulated and real-world data: trajectories tracking of animals as well as stock market data. Our method is competitive with existing global leadership inference methods but provides the first approaches for local leadership and coordination mechanism classification. Our results are consistent with ground-truthed biological data and the framework finds many known events in financial data which are not otherwise reflected in the aggregate NASDAQ index. Our method is easily generalizable to any coordinated time series data from interacting entities. C. Amornbunchornvej, Ivan Brugere, Ariana Strandburg-Peshkin, Damien Farine, Margaret Crofoot, Tanya Y. Berger-Wolf |
ACM Trans. Knowl. Discov. Data | 6 |
| 2017 | Identifying Traits of Leaders in Movement InitiationabstractHow do leaders lead? Are individuals with influence always at the front of their group? Do they initiate travel in new directions or are they first to start moving? Which attempts to initiate movement translate to leadership? In this paper we present a computational method to characterize and classify the types of leaders in movement initiation. We adapt a leadership inference framework, FLICA, to extract information about which individuals act as leaders. We then propose a framework for ranking leaders according to their position, velocity, and heading relative to the group and perform hypothesis testing of correlations between target features and leadership ranking. We use a time series of GPS positions of wild olive baboons (Papio anubis) as an application of our approach. Our results demonstrate that there is no correlation between leadership and early movement, there is negative correlation between leadership and new directions, while leadership and new area exploration are positively correlated. Thus, as an example, in baboons, our approach shows that while leaders are not the first to move, they are typically at the front and move in a new area with everybody immediately aligning in the direction of leader. Our simple scheme is flexible to be applied to other data sets and sets of traits to characterize leadership. C. Amornbunchornvej, Margaret Crofoot, Tanya Y. Berger-Wolf |
ASONAM | 3 |
| 2016 | Adversarial Sequence Tagging
Vena Jia Li, Kaiser Asif, Brian D. Ziebart, Tanya Y. Berger-Wolf |
IJCAI | 5 |
| 2016 | Optimization techniques for sparse matrix-vector multiplication on GPUs
Marco Maggioni, Tanya Y. Berger-Wolf |
J. Parallel Distributed Comput. | 2 |
| 2015 | Data Driven Science: SIGKDD PanelabstractThe panel session 'Data Driven Science' discusses application and use of knowledge discovery, machine learning and data analytics in science disciplines; in natural, physical, medical and social science; from physics to geology, and from neuroscience to population health. Knowledge discovery methods are finding broad application in all areas of scientific endeavor, to explore experimental data, to discover new models, to propose new scientific theories and ideas. In addition, the availability of ever larger scientific data sets is driving a new data-driven paradigm for modeling of complex phenomena in physical, natural and social sciences. Katharina Morik, Hugh F. Durrant-Whyte, Gary C. Hill, R. Dietmar Müller, Tanya Y. Berger-Wolf |
KDD | 5 |
| 2015 | Column-Generation Framework of Nonlinear Similarity Model for Reconstructing Sibling GroupsabstractEstablishing family relationships, such as parentage and sibling relationships, is fundamental in biological research, especially in wild species, as they are often important to understanding evolutionary, ecological, and behavioral processes. Because it is commonly impossible to determine familial relationships from field observations alone, the reconstruction of sibling relationships often depends on informative genetic markers coupled with accurate sibling reconstruction algorithms. Most studies in the literature reconstruct sibling relationships using methods that are based on either statistical analyses (i.e., likelihood estimation) or combinatorial concepts (i.e., Mendelian inheritance laws) of genetic data. We present a novel computational framework that integrates both combinatorial concepts and statistical analyses into one sibling reconstruction optimization model. To solve this integrated model, we propose a column-generation approach with a branch-and-price method. Under the assumption of parsimonious reconstruction, the master problem is to find the minimum set of sibling groups to cover the tested population. Pricing subproblems, which include both statistical similarity and combinatorial concepts of genetic data, are iteratively solved to generate high-quality sibling group candidates. Tested on real biological data sets, our approach efficiently provides reconstruction results that are more accurate than those provided by other state-of-the-art reconstruction algorithms. Chun-An Chou, Zhe Liang, W. Art Chaovalitwongse, Tanya Y. Berger-Wolf, Bhaskar DasGupta, Saad I. Sheikh, Mary V. Ashley, Isabel C. Caballero |
INFORMS J. Comput. | 4 |
| 2014 | Balancing the exploration and exploitation in an adaptive diversity guided genetic algorithmabstractExploration and exploitation are the two cornerstones which characterize Evolutionary Algorithms (EAs) capabilities. Maintaining the reciprocal balance of the explorative and exploitative power is the key to the success of EA applications. Accordingly, this work is concerned with proposing a diversity-guided genetic algorithm with a new mutation scheme that is capable of exploring the unseen regions of the search space, as well as exploiting the already-found promising elements. The proposed mutation operator specifies different mutation rates for different sites of an encoded solution. These site-specific rates are carefully derived based on the underlying pattern of highly-fit solutions, adjusted to every single individual, and adapted throughout the evolution to retain a good ratio between exploration and exploitation. Furthermore, in order to more directly monitor the exploration vs. exploitation balance, the proposed method is augmented with a diversity control process assuring that the search process does not lose the required balance between the two forces. Fatemeh Vafaee, György Turán, Peter C. Nelson, Tanya Y. Berger-Wolf |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | Among-site rate variation: adaptation of genetic algorithm mutation rates at each single siteabstractThis paper is concerned with proposing an elitist genetic algorithm which makes use of a new mutation scheme aimed to tackle both explorative and exploitative responsibilities of genetic operators. The proposed mutation scheme follows an approach similar to motif representation in biology, to derive the underlying pattern of highly-fit solutions discovered so far. This pattern is then used to derive mutation rates specified for every site along the encoded solutions. The site-specific rates are amended for every individual to balance the required explorative and exploitative power. The Markov chain model of the proposed method is also derived and used to analyze its convergence properties. Fatemeh Vafaee, György Turán, Peter C. Nelson, Tanya Y. Berger-Wolf |
GECCO | 4 |
| 2014 | Expansion and decentralized search in complex networks
Arun S. Maiya, Tanya Y. Berger-Wolf |
Knowl. Inf. Syst. | 2 |
| 2013 | AdELL: An Adaptive Warp-Balancing ELL Format for Efficient Sparse Matrix-Vector Multiplication on GPUsabstractThe sparse matrix-vector multiplication (SpMV) is a fundamental computational kernel used in science and engineering. As a result, the performance of a large number of applications depends on the efficiency of the SpMV. This kernel is, in fact, a bandwidth-limited operation and poses a challenge for optimization when the matrix has an irregular structure. The literature on implementing SpMV on throughput-oriented many core processors is extensive and mostly focuses on matrix formats, proposing different ideas to adapt matrix sparsity to the underlying architecture. In this paper, we propose a novel ELL-based matrix format called Adaptive ELL (AdELL) to improve the state-of-the-art of the SpMV on Graphic Processing Units (GPUs). The AdELL format is based on the idea of distributing working threads to rows according to their computational load, creating balanced hardware-level blocks (warps) that take full advantage of the vectorized execution on Streaming Multiprocessors (SMs). The AdELL data structure is created using a novel warp-balancing heuristic designed to smooth the workload among warps without the need of tuning any parameters. AdELL provides an efficient warp-level synchronization (as opposed to block-level) but can also use atomic operations to distribute very skewed rows over multiple warps. Moreover, we introduce a loop unrolling heuristic that optimizes the SpMV performance by selecting the best unrolling factor based on the warp workload. We tested the proposed AdELL sparse format on a set of conventional benchmarks from heterogeneous application domains. The results show substantial and consistent performance improvements for double-precision calculations, outperforming the state-of-the-art ensemble framework clSpMV. We could observe speedup peaks up to 1.94 and a 25% (geometric) average improvement, which can be potentially increased to 43% introducing a simple 1x2 blocking strategy. Marco Maggioni, Tanya Y. Berger-Wolf |
ICPP | 2 |
| 2013 | Computational Behavioral Ecology
Tanya Y. Berger-Wolf |
ISBRA | 1 |
| 2013 | HotSpotter - Patterned species instance recognitionabstractWe present HotSpotter, a fast, accurate algorithm for identifying individual animals against a labeled database. It is not species specific and has been applied to Grevy's and plains zebras, giraffes, leopards, and lionfish. We describe two approaches, both based on extracting and matching keypoints or “hotspots”. The first tests each new query image sequentially against each database image, generating a score for each database image in isolation, and ranking the results. The second, building on recent techniques for instance recognition, matches the query image against the database using a fast nearest neighbor search. It uses a competitive scoring mechanism derived from the Local Naive Bayes Nearest Neighbor algorithm recently proposed for category recognition. We demonstrate results on databases of more than 1000 images, producing more accurate matches than published methods and matching each query image in just a few seconds. Jonathan P. Crall, Charles V. Stewart, Tanya Y. Berger-Wolf, Daniel I. Rubenstein, Siva R. Sundaresan |
WACV | 3 |
| 2012 | A high performance multiple sequence alignment system for pyrosequencing reads from multiple reference genomes
Fahad Saeed, Alan Perez-Rathke, Jaroslaw Gwarnicki, Tanya Y. Berger-Wolf, Ashfaq Khokhar 0001 |
J. Parallel Distributed Comput. | 4 |
| 2011 | Finding Communities in Dynamic Social NetworksabstractCommunities are natural structures observed in social networks and are usually characterized as "relatively dense" subsets of nodes. Social networks change over time and so do the underlying community structures. Thus, to truly uncover this structure we must take the temporal aspect of networks into consideration. Previously, we have represented framework for finding dynamic communities using the social cost model and formulated the corresponding optimization problem [33], assuming that partitions of individuals into groups are given in each time step. We have also presented heuristics and approximation algorithms for the problem, with the same assumption [32]. In general, however, dynamic social networks are represented as a sequence of graphs of snapshots of the social network and the assumption that we have partitions of individuals into groups does not hold. In this paper, we extend the social cost model and formulate an optimization problem of finding community structure from the sequence of arbitrary graphs. We propose a semi definite programming formulation and a heuristic rounding scheme. We show, using synthetic data sets, that this method is quite accurate on synthetic data sets and present its results on a real social network. Chayant Tantipathananandh, Tanya Y. Berger-Wolf |
ICDM | 2 |
| 2011 | Benefits of bias: towards better characterization of network samplingabstractFrom social networks to P2P systems, network sampling arises in many settings. We present a detailed study on the nature of biases in network sampling strategies to shed light on how best to sample from networks. We investigate connections between specific biases and various measures of structural representativeness. We show that certain biases are, in fact, beneficial for many applications, as they "push" the sampling process towards inclusion of desired properties. Finally, we describe how these sampling biases can be exploited in several, real-world applications including disease outbreak detection and market research. Arun S. Maiya, Tanya Y. Berger-Wolf |
KDD | 2 |
| 2011 | Biometric animal databases from field photographs: identification of individual zebra in the wildabstractWe describe an algorithmic and experimental approach to a fundamental problem in field ecology: computer-assisted individual animal identification. We use a database of noisy photographs taken in the wild to build a biometric database of individual animals differentiated by their coat markings. A new image of an unknown animal can then be queried by its coat markings against the database to determine if the animal has been observed and identified before. Our algorithm, called StripeCodes, efficiently extracts simple image features and uses a dynamic programming algorithm to compare images. We test its accuracy against two different classes of methods: Eigenface, which is based on algebraic techniques, and matching multi-scale histograms of differential image features, an approach from signal processing. StripeCodes performs better than all competing methods for our dataset, and scales well with database size. Mayank Lahiri, Chayant Tantipathananandh, Rosemary Warungu, Daniel I. Rubenstein, Tanya Y. Berger-Wolf |
ICMR | 5 |
| 2011 | Visualizing the Evolution of Community Structures in Dynamic Social NetworksabstractAbstract Social network analysis is the study of patterns of interaction between social entities. The field is attracting increasing attention from diverse disciplines including sociology, epidemiology, and behavioral ecology. An important sociological phenomenon that draws the attention of analysts is the emergence of communities, which tend to form, evolve, and dissolve gradually over a period of time. Understanding this evolution is crucial to sociologists and domain scientists, and often leads to a better appreciation of the social system under study. Therefore, it is imperative that social network visualization tools support this task. While graph‐based representations are well suited for investigating structural properties of networks at a single point in time, they appear to be significantly less useful when used to analyze gradual structural changes over a period of time. In this paper, we present an interactive visualization methodology for dynamic social networks. Our technique focuses on revealing the community structure implied by the evolving interaction patterns between individuals. We apply our visualization to analyze the community structure in the US House of Representatives. We also report on a user study conducted with the participation of behavioral ecologists working with social network datasets that depict interactions between wild animals. Findings from the user study confirm that the visualization was helpful in providing answers to sociological questions as well as eliciting new observations on the social organization of the population under study. Khairi Reda, Chayant Tantipathananandh, Andrew E. Johnson 0001, Jason Leigh, Tanya Y. Berger-Wolf |
Comput. Graph. Forum | 5 |
| 2010 | Expansion and search in networksabstractBorrowing from concepts in expander graphs, we study the expansion properties of real-world, complex networks (e.g. social networks, unstructured peer-to-peer or P2P networks) and the extent to which these properties can be exploited to understand and address the problem of decentralized search. We first produce samples that concisely capture the overall expansion properties of an entire network, which we collectively refer to as the expansion signature. Using these signatures, we find a correspondence between the magnitude of maximum expansion and the extent to which a network can be efficiently searched. We further find evidence that standard graph-theoretic measures, such as average path length, fail to fully explain the level of "searchability" or ease of information diffusion and dissemination in a network. Finally, we demonstrate that this high expansion can be leveraged to facilitate decentralized search in networks and show that an expansion-based search strategy outperforms typical search methods. Arun S. Maiya, Tanya Y. Berger-Wolf |
CIKM | 2 |
| 2010 | Online Sampling of High Centrality Individuals in Social Networks
Arun S. Maiya, Tanya Y. Berger-Wolf |
PAKDD (1) | 2 |
| 2010 | Discovering Kinship through Small Subsets
Dan Brown 0001, Tanya Y. Berger-Wolf |
WABI | 2 |
| 2010 | Sampling community structureabstractWe propose a novel method, based on concepts from expander graphs, to sample communities in networks. We show that our sampling method, unlike previous techniques, produces subgraphs representative of community structure in the original network. These generated subgraphs may be viewed as stratified samples in that they consist of members from most or all communities in the network. Using samples produced by our method, we show that the problem of community detection may be recast into a case of statistical relational learning. We empirically evaluate our approach against several real-world datasets and demonstrate that our sampling method can effectively be used to infer and approximate community affiliation in the larger network. Arun S. Maiya, Tanya Y. Berger-Wolf |
WWW | 2 |
| 2010 | New Optimization Model and Algorithm for Sibling Reconstruction from Genetic MarkersabstractWith improved tools for collecting genetic data from natural and experimental populations, new opportunities arise to study fundamental biological processes, including behavior, mating systems, adaptive trait evolution, and dispersal patterns. Full use of the newly available genetic data often depends upon reconstructing genealogical relationships of individual organisms, such as sibling reconstruction. This paper presents a new optimization framework for sibling reconstruction from single generation microsatellite genetic data. Our framework is based on assumptions of parsimony and combinatorial concepts of Mendel's inheritance rules. Here, we develop a novel optimization model for sibling reconstruction as a large-scale mixed-integer program (MIP), shown to be a generalization of the set covering problem. We propose a new heuristic approach to efficiently solve this large-scale optimization problem. We test our approach on real biological data as presented in other studies as well as simulated data, and compare our results with other state-of-the-art sibling reconstruction methods. The empirical results show that our approaches are very efficient and outperform other methods while providing the most accurate solutions for two benchmark data sets. The results suggest that our framework can be used as an analytical and computational tool for biologists to better study ecological and evolutionary processes involving knowledge of familial relationships in a wide variety of biological systems. W. Art Chaovalitwongse, Chun-An Chou, Tanya Y. Berger-Wolf, Bhaskar DasGupta, Saad I. Sheikh, Mary V. Ashley, Isabel C. Caballero |
INFORMS J. Comput. | 3 |
| 2010 | Periodic subgraph mining in dynamic networks
Mayank Lahiri, Tanya Y. Berger-Wolf |
Knowl. Inf. Syst. | 2 |
| 2009 | On Approximating an Implicit Cover Problem in Biology
Mary V. Ashley, Tanya Y. Berger-Wolf, W. Art Chaovalitwongse, Bhaskar DasGupta, Ashfaq Khokhar 0001, Saad I. Sheikh |
AAIM | 2 |
| 2009 | Constant-factor approximation algorithms for identifying dynamic communitiesabstractWe propose two approximation algorithms for identifying communities in dynamic social networks. Communities are intuitively characterized as "unusually densely knit" subsets of a social network. This notion becomes more problematic if the social interactions change over time. Aggregating social networks over time can radically misrepresent the existing and changing community structure. Recently, we have proposed an optimization-based framework for modeling dynamic community structure. Also, we have proposed an algorithm for finding such structure based on maximum weight bipartite matching. In this paper, we analyze its performance guarantee for a special case where all actors can be observed at all times. In such instances, we show that the algorithm is a small constant factor approximation of the optimum. We use a similar idea to design an approximation algorithm for the general case where some individuals are possibly unobserved at times, and to show that the approximation factor increases twofold but remains a constant regardless of the input size. This is the first algorithm for inferring communities in dynamic networks with a provable approximation guarantee. We demonstrate the general algorithm on real data sets. The results confirm the efficiency and effectiveness of the algorithm in identifying dynamic communities. Chayant Tantipathananandh, Tanya Y. Berger-Wolf |
KDD | 2 |
| 2009 | On approximating four covering and packing problems
Mary V. Ashley, Tanya Y. Berger-Wolf, Piotr Berman, W. Art Chaovalitwongse, Bhaskar DasGupta, Ming-Yang Kao |
J. Comput. Syst. Sci. | 2 |
| 2008 | Mining Periodic Behavior in Dynamic Social NetworksabstractSocial interactions that occur regularly typically correspond to significant yet often infrequent and hard to detect interaction patterns. To identify such regular behavior, we propose a new mining problem of finding periodic or near periodic subgraphs in dynamic social networks. We analyze the computational complexity of the problem, showing that, unlike any of the related subgraph mining problems, it is polynomial. We propose a practical, efficient and scalable algorithm to find such subgraphs that takes imperfect periodicity into account. We demonstrate the applicability of our approach on several real-world networks and extract meaningful and interesting periodic interaction patterns. Mayank Lahiri, Tanya Y. Berger-Wolf |
ICDM | 2 |
| 2007 | Structure Prediction in Temporal Networks using Frequent SubgraphsabstractThere are several types of processes which can be modeled explicitly by recording the interactions between a set of actors over time. In such applications, a common objective is, given a series of observations, to predict exactly when certain interactions will occur in the future. We propose a representation for this type of temporal data and a generic, streaming, adaptive algorithm to predict the pattern of interactions at any arbitrary point in the future. We test our algorithm on predicting patterns in e-mail logs, correlations between stock closing prices, and social grouping in herds of Plains zebras. Our algorithm averages over 85% accuracy in predicting a set of interactions at any unseen timestep. To the best of our knowledge, this is the first algorithm that predicts interactions at the finest possible time grain Mayank Lahiri, Tanya Y. Berger-Wolf |
CIDM | 2 |
| 2007 | A framework for community identification in dynamic social networksabstractWe propose frameworks and algorithms for identifying communities in social networks that change over time. Communities are intuitively characterized as "unusually densely knit" subsets of a social network. This notion becomes more problematic if the social interactions change over time. Aggregating social networks over time can radically misrepresent the existing and changing community structure. Instead, we propose an optimization-based approach for modeling dynamic community structure. We prove that finding the most explanatory community structure is NP-hard and APX-hard, and propose algorithms based on dynamic programming, exhaustive search, maximum matching, and greedy heuristics. We demonstrate empirically that the heuristics trace developments of community structure accurately for several synthetic and real-world examples. Chayant Tantipathananandh, Tanya Y. Berger-Wolf, David Kempe 0001 |
KDD | 2 |
| 2006 | A framework for analysis of dynamic social networksabstractFinding patterns of social interaction within a population has wide-ranging applications including: disease modeling, cultural and information transmission, and behavioral ecology. Social interactions are often modeled with networks. A key characteristic of social interactions is their continual change. However, most past analyses of social networks are essentially static in that all information about the time that social interactions take place is discarded. In this paper, we propose a new mathematical and computational framework that enables analysis of dynamic social networks and that explicitly makes use of information about when social interactions occur. Tanya Y. Berger-Wolf, Jared Saia |
KDD | 1 |
| 2004 | Online Consensus and Agreement of Phylogenetic Trees
Tanya Y. Berger-Wolf |
WABI | 1 |
| 2002 | Index assignment for multichannel communication under failureabstractWe consider the problem of constructing multiple description scalar quantizers and describing the achievable rate-distortion tuples in that setting. We model this as a combinatorial optimization problem of number arrangements in a matrix. This approach gives a general technique for deriving lower bounds on the distortion at given channel rates. This technique is constructive, thus allowing an algorithm that gives an upper bound. For the case of two communication channels with equal rates, the bounds coincide, thus giving the precise lowest achievable distortion at fixed rates. The bounds are within a small constant for a higher number of channels. To the best of our knowledge, this is the first result involving systems with more than two communication channels. Tanya Y. Berger-Wolf, Edward M. Reingold |
IEEE Trans. Inf. Theory | 1 |
| 1999 | Optimal Multichannel Communication Under Failure
Tanya Y. Berger-Wolf, Edward M. Reingold |
SODA | 1 |