Vitaly Ablavsky

dblp:17/2109 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0003-2703-7666ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Image recognition and object detection · 28% Video understanding and tracking · 24% Segmentation and scene understanding · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
1.352022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes · ECCV (18) 2020
Learning a Family of Detectors via Multiplicative Kernels · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Segmentation and scene understanding
object segmentation
0.722022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Multiplicative kernels: Object detection, segmentation and pose estimation · CVPR 2008
Computer vision › Video understanding and tracking
object tracking
0.622021
Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers · CVPR 2021
Layered graphical models for tracking partially-occluded objects · CVPR 2008
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Computer vision › Vision and language
multimodal fusion
0.512021
Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers · CVPR 2021
Computer vision › Video understanding and tracking › object tracking › vision-language tracking
tracking by natural language specification
0.512021
Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers · CVPR 2021
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
instance decoupling
0.412020
Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes · ECCV (18) 2020
Machine learning › Efficient and distributed learning › inference efficiency
cost-aware inference
0.412019
Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations · ICCV 2019
Machine learning › Efficient and distributed learning › edge computing
edge inference
0.412019
Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations · ICCV 2019
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.412019
Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations · ICCV 2019
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.222011
Layered Graphical Models for Tracking Partially Occluded Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Layered graphical models for tracking partially-occluded objects · CVPR 2008
Computer vision › 3D vision
pose estimation
0.222011
Learning a Family of Detectors via Multiplicative Kernels · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Multiplicative kernels: Object detection, segmentation and pose estimation · CVPR 2008
Computer vision › Video understanding and tracking › object tracking
person tracking
0.122011
Layered Graphical Models for Tracking Partially Occluded Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Layered graphical models for tracking partially-occluded objects · CVPR 2008
Computer vision › Video understanding and tracking
action recognition
0.112011
Learning parameterized histogram kernels on the simplex manifold for image and action classification · ICCV 2011
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112011
Layered Graphical Models for Tracking Partially Occluded Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Image recognition and object detection
image classification
0.112011
Learning parameterized histogram kernels on the simplex manifold for image and action classification · ICCV 2011
Computer vision › 3D vision › 3d scene modeling
scene representation
0.012011
Layered Graphical Models for Tracking Partially Occluded Objects · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection
0.012011
Learning a Family of Detectors via Multiplicative Kernels · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Mathematical optimization › constrained optimization
margin maximization
0.012011
Learning parameterized histogram kernels on the simplex manifold for image and action classification · ICCV 2011
Computer vision › Face, body and person analysis
face detection
0.012007
Parameter Sensitive Detectors · CVPR 2007

Methods — techniques the papers use, named apart from their topics

semantic segmentation · 1.1instance segmentation · 1.1siamese network · 0.5region proposal network · 0.5instance separation · 0.4foveation · 0.4deep reinforcement learning · 0.4DDPG · 0.4generalization error minimization · 0.2multiplicative kernel · 0.2support vector machine · 0.1histogram kernel learning · 0.1
YearPublicationVenuePosition
2024 SSP-GNN: Learning to Track via Bilevel Optimization
abstract
We propose a graph-based tracking formulation for multi-object tracking (MOT) where target detections contain kinematic information and re-identification features (attributes). Our method applies a successive shortest paths (SSP) algorithm to a tracking graph defined over a batch of frames. The edge costs in this tracking graph are computed via message-passing network, a graph neural network (GNN) variant. The parameters of the GNN, and hence, the tracker, are learned end-to-end on a training set of example ground-truth tracks and detections. Specifically, learning takes the form of bilevel optimization guided by our novel loss function. We evaluate our algorithm on simulated scenarios to understand its sensitivity to scenario aspects and model hyperparameters. Across varied scenario complexities, our method compares favorably to a strong baseline.
Griffin Golias, Masa Nakura-Fan, Vitaly Ablavsky
FUSION3
2022 ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
abstract
Less than 35% of recyclable waste is being actually recycled in the US [2], which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/
Dina Bashkirova, Mohamed Abdelfattah, Ziliang Zhu, James Akl, Fadi M. Alladkani, Ping Hu 0001, Vitaly Ablavsky, Berk Çalli, Sarah Adel Bargal, Kate Saenko
CVPR7
2021 Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers
abstract
We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) descriptions task. The proposed SNLT is applicable to a wide range of Siamese trackers, providing a new class of baselines for the tracking by NL task and promising future improvements from the advancements of Siamese trackers. The carefully designed architecture of the Siamese Natural Language Region Proposal Network (SNL-RPN), together with the Dynamic Aggregation of vision and language modalities, is introduced to perform the tracking by NL task. Empirical results over tracking benchmarks with NL annotations show that the proposed SNLT improves Siamese trackers by 3 to 7 percentage points with a slight tradeoff of speed. The proposed SNLT outperforms all NL trackers to-date and is competitive among state-of-the-art real-time trackers on LaSOT benchmarks while running at 50 frames per second on a single GPU. Code for this work is available at https://github.com/fredfung007/snlt.
Qi Feng 0004, Vitaly Ablavsky, Qinxun Bai, Stan Sclaroff
CVPR2
2021 Leveraging Affect Transfer Learning for Behavior Prediction in an Intelligent Tutoring System
abstract
In this work, we propose a video-based transfer learning approach for predicting problem outcomes of students working with an intelligent tutoring system (ITS). By analyzing a student's face and gestures, our method predicts the outcome of a student answering a problem in an ITS from a video feed. Our work is motivated by the reasoning that the ability to predict such outcomes enables tutoring systems to adjust interventions, such as hints and encouragement, and to ultimately yield improved student learning. We collected a large labeled dataset of student interactions with an intelligent online math tutor consisting of 68 sessions, where 54 individual students solved 2,749 problems. We will release this dataset publicly upon publication of this paper. It will be available at https://www.cs.bu.edu/faculty/betke/research/learning/. Working with this dataset, our transfer-learning challenge was to design a representation in the source domain of pictures obtained “in the wild” for the task of facial expression analysis, and transferring this learned representation to the task of human behavior prediction in the domain of webcam videos of students in a classroom environment. We developed a novel facial affect representation and a user-personalized training scheme that unlocks the potential of this representation. We designed several variants of a recurrent neural network that models the temporal structure of video sequences of students solving math problems. Our final model, named ATL-BP for Affect Transfer Learning for Behavior Prediction, achieves a relative increase in mean F -score of 50 % over the state-of-the-art method on this new dataset.
Nataniel Ruiz, Hao Yu 0014, Danielle Allessio, Mona Jalal, Ajjen Joshi, Tom Murray 0001, John J. Magee, Jacob Whitehill, Vitaly Ablavsky, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke
FG9
2021 Distillation Multiple Choice Learning for Multimodal Action Recognition
abstract
In this work, we address the problem of learning an ensemble of specialist networks using multimodal data, while considering the realistic and challenging scenario of possible missing modalities at test time. Our goal is to leverage the complementary information of multiple modalities to the benefit of the ensemble and each individual network. We introduce a novel Distillation Multiple Choice Learning framework for multimodal data, where different modality networks learn in a cooperative setting from scratch, strengthening one another. The modality networks learned using our method achieve significantly higher accuracy than if trained separately, due to the guidance of other modalities. We evaluate this approach on three video action recognition benchmark datasets. We obtain state-of-the-art results in comparison to other approaches that work with missing modalities at test time.
Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff
WACV3
2020 Learning to Separate: Detecting Heavily-Occluded Objects in Urban Scenes
Chenhongyi Yang, Vitaly Ablavsky, Kaihong Wang, Qi Feng 0004, Margrit Betke
ECCV (18)2
2020 Real-time Visual Object Tracking with Natural Language Description
abstract
In this work, we argue that conditioning on the natural language (NL) description of a target provides information for longer-term invariance, and thus helps cope with typical tracking challenges. However, deriving a formulation to combine the strengths of appearance-based tracking with the language modality is not straightforward. Therefore, we propose a novel deep tracking-by-detection formulation that can take advantage of NL descriptions. Regions that are related to the given NL description are generated by a proposal network during the detection stage of the tracker. Our LSTM based tracker then predicts the update of the target from regions proposed by the NL based detection stage. Our method runs at over 30 fps on a single GPU. In benchmarks, our method is competitive with state of the art trackers that employ bounding boxes for initialization, while it outperforms all other trackers on targets given unambiguous and precise language annotations. When conditioned on NL descriptions only, our model doubles the performance of the previous best attempt [25].
Qi Feng 0004, Vitaly Ablavsky, Qinxun Bai, Guorong Li, Stan Sclaroff
WACV2
2020 DIPNet: Dynamic Identity Propagation Network for Video Object Segmentation
abstract
Many recent methods for semi-supervised Video Object Segmentation (VOS) have achieved good performance by exploiting the annotated first frame via one-shot fine-tuning or mask propagation. However, heavily relying on the first frame may weaken the robustness for VOS, since video objects can show large variations through time. In this work, we propose a Dynamic Identity Propagation Network (DIPNet) that adaptively propagates and accurately segments the video objects over time. To achieve this, DIPNet factors the VOS task at each time step into a dynamic propagation phase and a spatial segmentation phase. The former utilizes a novel identity representation to adaptively propagate objects’ reference information over time, which enhances the robustness to videos’ temporal variations. The segmentation phase uses the propagated information to tackle the object segmentation as an easier static image problem that can be optimized via light-weight fine-tuning on the first frame, thus reducing the computational cost. As a result, by optimizing these two components to complement each other, we can achieve a robust system for VOS. Evaluations on four benchmark datasets show that DIPNet provides state-of-the-art performance with time efficiency.
Ping Hu 0001, Jun Liu 0036, Gang Wang 0012, Vitaly Ablavsky, Kate Saenko, Stan Sclaroff
WACV4
2019 Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
abstract
We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with fine-grained classification, we adopt the perspective of sequential fixation with a foveated field-of-view to model cloud-edge interactions. We propose a novel deep reinforcement learning-based foveation model, DRIFT, that sequentially generates and recognizes mixed-acuity images. Training of DRIFT requires only image-level category labels and encourages fixations to contain task-relevant information, while maintaining data efficiency. Specifically, we train a foveation actor network with a novel Deep Deterministic Policy Gradient by Conditioned Critic and Coaching(DDPGC3) algorithm. In addition, we propose to shape the reward to provide informative feedback after each fixation to better guide RL training. We demonstrate the effectiveness of DRIFT on this task by evaluating on five fine-grained classification benchmark datasets, and show that the proposed approach achieves state-of-the-art performance with over 3X reduction in transmitted pixels.
Hanxiao Wang 0001, Venkatesh Saligrama, Stan Sclaroff, Vitaly Ablavsky
ICCV4
2014 Take your eyes off the ball: Improving ball-tracking by focusing on team play
Xinchao Wang, Vitaly Ablavsky, Horesh Ben Shitrit, Pascal Fua
Comput. Vis. Image Underst.2
2011 Learning parameterized histogram kernels on the simplex manifold for image and action classification
abstract
State-of-the-art image and action classification systems often employ vocabulary-based representations. The classification accuracy achieved with such vocabulary-based representations depends significantly on the chosen histogram-distance. In particular, when the decision function is a support-vector-machine (SVM), the classification accuracy depends on the chosen histogram kernel. In this paper we focus on smoothly-parameterized kernels in the space of histograms, such as, but not limited to, kernels that are derived from smoothly-parameterized histogram-distance functions. We learn parameters of histogram kernels so that the SVM accuracy is improved. This is accomplished by simultaneously maximizing the SVM's geometric margin and minimizing an estimate of its generalization error. We validate our approach on a previously-published two-class synthetic dataset and three real-world multi-class datasets: Oxford5K, KTH, and UCF. On these datasets our approach yields results that compare favorably to or exceed the state of the art.
Vitaly Ablavsky, Stan Sclaroff
ICCV1
2011 Layered Graphical Models for Tracking Partially Occluded Objects
abstract
We propose a representation for scenes containing relocatable objects that can cause partial occlusions of people in a camera's field of view. In many practical applications, relocatable objects tend to appear often; therefore, models for them can be learned offline and stored in a database. We formulate an occluder-centric representation, called a graphical model layer, where a person's motion in the ground plane is defined as a first-order Markov process on activity zones, while image evidence is aggregated in 2D observation regions that are depth-ordered with respect to the occlusion mask of the relocatable object. We represent real-world scenes as a composition of depth-ordered, interacting graphical model layers, and account for image evidence in a way that handles mutual overlap of the observation regions and their occlusions by the relocatable objects. These layers interact: Proximate ground-plane zones of different model instances are linked to allow a person to move between the layers, and image evidence is shared between the observation regions of these models. We demonstrate our formulation in tracking pedestrians in the vicinity of parked vehicles. Our results compare favorably with a sprite-learning algorithm, with a pedestrian tracker based on deformable contours, and with pedestrian detectors.
Vitaly Ablavsky, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Learning a Family of Detectors via Multiplicative Kernels
abstract
Object detection is challenging when the object class exhibits large within-class variations. In this work, we show that foreground-background classification (detection) and within-class classification of the foreground class (pose estimation) can be jointly learned in a multiplicative form of two kernel functions. Model training is accomplished via standard SVM learning. When the foreground object masks are provided in training, the detectors can also produce object segmentations. A tracking-by-detection framework to recover foreground state in video sequences is also proposed with our model. The advantages of our method are demonstrated on tasks of object detection, view angle estimation, and tracking. Our approach compares favorably to existing methods on hand and vehicle detection tasks. Quantitative tracking results are given on sequences of moving vehicles and human faces.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Layered graphical models for tracking partially-occluded objects
abstract
Partial occlusions are commonplace in a variety of real world computer vision applications: surveillance, intelligent environments, assistive robotics, autonomous navigation, etc. While occlusion handling methods have been proposed, most methods tend to break down when confronted with numerous occluders in a scene. In this paper, a layered image-plane representation for tracking people through substantial occlusions is proposed. An image-plane representation of motion around an object is associated with a pre-computed graphical model, which can be instantiated efficiently during online tracking. A global state and observation space is obtained by linking transitions between layers. A reversible jump Markov chain Monte Carlo approach is used to infer the number of people and track them online. The method outperforms two state-of-the-art methods for tracking over extended occlusions, given videos of a parking lot with numerous vehicles and a laboratory with many desks and workstations.
Vitaly Ablavsky, Ashwin Thangali, Stan Sclaroff
CVPR1
2008 Multiplicative kernels: Object detection, segmentation and pose estimation
abstract
Object detection is challenging when the object class exhibits large within-class variations. In this work, we show that foreground-background classification (detection) and within-class classification of the foreground class (pose estimation) can be jointly learned in a multiplicative form of two kernel functions. One kernel measures similarity for foreground-background classification. The other kernel accounts for latent factors that control within-class variation and implicitly enables feature sharing among foreground training samples. Detector training can be accomplished via standard SVM learning. The resulting detectors are tuned to specific variations in the foreground class. They also serve to evaluate hypotheses of the foreground state. When the foreground parameters are provided in training, the detectors can also produce parameter estimate. When the foreground object masks are provided in training, the detectors can also produce object segmentation. The advantages of our method over past methods are demonstrated on data sets of human hands and vehicles.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
CVPR3
2007 Parameter Sensitive Detectors
abstract
Object detection can be challenging when the object class exhibits large variations. One commonly-used strategy is to first partition the space of possible object variations and then train separate classifiers for each portion. However, with continuous spaces the partitions tend to be arbitrary since there are no natural boundaries (for example, consider the continuous range of human body poses). In this paper, a new formulation is proposed, where the detectors themselves are associated with continuous parameters, and reside in a parameterized function space. There are two advantages of this strategy. First, a-priori partitioning of the parameter space is not needed; the detectors themselves are in a parameterized space. Second, the underlying parameters for object variations can be learned from training data in an unsupervised manner. In profile face detection experiments, at a fixed false alarm number of 90, our method attains a detection rate of 75% vs. 70% for the method of Viola-Jones. In hand shape detection, at a false positive rate of 0.1%, our method achieves a detection rate of 99.5% vs. 98% for partition based methods. In pedestrian detection, our method reduces the miss detection rate by a factor of three at a false positive rate of 1%, compared with the method of Dalal-Triggs.
Ashwin Thangali, Vitaly Ablavsky, Stan Sclaroff
CVPR3
2005 Sequential Correction of Perspective Warp in Camera-based Documents
abstract
Documents captured with hand-held devices, such as digital cameras often exhibit perspective warp artifacts. These artifacts pose problems for OCR systems which at best can only handle in-plane rotation. We propose a method for recovering the planar appearance of an input document image by examining the vertical rate of change in scale of features in the document. Our method makes fewer assumptions about the document structure than do previously published algorithms.
Camille Monnier, Steve Holden, Magnús Snorrason, Vitaly Ablavsky
ICDAR4
2003 Automatic Feature Selection with Applications to Script Identification of Degraded Documents
abstract
Current approaches to script identification rely onhand-selected features and often require processing a significantpart of the document to achieve reliable identification.We present an approach that applies a large pool ofimage features to a small training sample and uses subsetfeature selection techniques to automatically select a subsetwith the most discriminating power. At run time we usea classifier coupled with an evidence accumulation engineto report a script label once a preset likelihood thresholdhas been reached. We apply the system to a diverse corpusof printed Russian and English documents that suffer fromcommon degradation problems. Our validation studyshows promising results both in terms of the script identificationaccuracy and the ability to identify script on thescale of individual words and text lines.
Vitaly Ablavsky, Mark R. Stevens
ICDAR1
2003 Background models for tracking objects in water
abstract
This paper presents a novel background analysis technique to enable robust tracking of objects in water- based scenarios. Current pixel-wise statistical background models support automatic change detection in many outdoor situations, but are limited to background changes which can be modeled via a set of per-pixel spatially uncorrelated processes. In water-based scenarios, waves caused by wind or by moving vessels (wakes) form highly correlated moving patterns that confuse traditional background analysis models. In this work we introduce a framework that explicitly models this type of background variation. The framework combines the output of a statistical background model with localized optical flow analysis to produce two motion maps. In the final stage we apply object-level fusion to filter out moving regions that are most likely caused by wave clutter. A tracking algorithm can now handle the resulting set of objects.
Vitaly Ablavsky
ICIP (3)1
2002 Real-time autonomous video enhancement system (RAVE)
abstract
The ability to autonomously enhance low-quality or corrupted streaming video data is essential in a number of important civilian and defense scenarios. Applications include visual surveillance, motion picture restoration, and remote control of unmanned aerial vehicles. We have developed a prototype of RAVE: real-time autonomous video enhancement system. It consists of a suite of video artifact detection algorithms and corresponding correction algorithms. The system is autonomously controlled by an intelligent software agent. Our prototype has been successfully validated on several video sequences from different application domains and is being matured into a fully-functional, real-time embedded system.
Vitaly Ablavsky, Magnús Snorrason, C. J. Taylor
ICIP (2)1