Ali Athar

dblp:187/5650 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
abstract
Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically involve category-guided or referral-based object segmentation. Although both directions are essential for developing models with human-level video comprehension, they have largely evolved separately, with distinct benchmarks and architectures. This paper aims to unify these efforts by introducing ViCaS, a new dataset containing thousands of challenging videos, each annotated with detailed, human-written captions and temporally consistent, pixel-accurate masks for multiple objects with phrase grounding. Our benchmark evaluates models on both holistic/high-level understanding and language-guided, pixel-precise segmentation. We also present carefully validated evaluation measures and propose an effective model architecture that can tackle our benchmark.
Ali Athar, Xueqing Deng, Liang-Chieh Chen
CVPR1
2025 COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
abstract
This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome limitations in existing image-text datasets that often lack detailed, scene-comprehensive descriptions. The COCONut-PanCap dataset incorporates fine-grained, region-level captions grounded in panoptic segmentation masks, ensuring consistency and improving the detail of generated captions.Through human-edited, densely annotated descriptions, COCONut-PanCap supports improved training of vision-language models (VLMs) for image understanding and generative models for text-to-image tasks.Experimental results demonstrate that COCONut-PanCap significantly boosts performance across understanding and generation tasks, offering complementary benefits to large-scale datasets. This dataset sets a new benchmark for evaluating models on joint panoptic segmentation and grounded captioning tasks, addressing the need for high-quality, detailed image-text annotations in multi-modal learning.
Xueqing Deng, Qihang Yu, Ali Athar, Xiaojie Jin 0004, Xiaohui Shen, Liang-Chieh Chen
NeurIPS4
2023 TarViS: A Unified Approach for Target-Based Video Segmentation
abstract
The general domain of video segmentation is currently fragmented into different tasks spanning multiple benchmarks. Despite rapid progress in the state-of-the-art, current methods are overwhelmingly task-specific and cannot conceptually generalize to other tasks. Inspired by recent approaches with multi-task capability, we propose TarViS: a novel, unified network architecture that can be applied to any task that requires segmenting a set of arbitrarily defined ‘targets' in video. Our approach is flexible with respect to how tasks define these targets, since it models the latter as abstract ‘queries' which are then used to predict pixel-precise target masks. A single TarViS model can be trained jointly on a collection of datasets spanning different tasks, and can hot-swap between tasks during inference without any task-specific retraining. To demonstrate its effectiveness, we apply TarViS to four different tasks, namely Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), Video Object Segmentation (VOS) and Point Exemplar-guided Tracking (PET). Our unified, jointly trained model achieves state-of-the-art performance on 5/7 benchmarks spanning these four tasks, and competitive performance on the remaining two. Code and model weights are available at: https://github.com/Ali2500/TarViS
Ali Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan, Bastian Leibe
CVPR1
2023 BURST: A Benchmark for Unifying Object Recognition, Segmentation and Tracking in Video
abstract
Multiple existing benchmarks involve tracking and segmenting objects in video e.g., Video Object Segmentation (VOS) and Multi-Object Tracking and Segmentation (MOTS), but there is little interaction between them due to the use of disparate benchmark datasets and metrics (e.g. $\mathcal{J}\& {\mathcal{F}}$, mAP, sMOTSA). As a result, published works usually target a particular benchmark, and are not easily comparable to each another. We believe that the development of generalized methods that can tackle multiple tasks requires greater cohesion among these research sub-communities. In this paper, we aim to facilitate this by proposing BURST, a dataset which contains thousands of diverse videos with high-quality object masks, and an associated benchmark with six tasks involving object tracking and segmentation in video. All tasks are evaluated using the same data and comparable metrics, which enables researchers to consider them in unison, and hence, more effectively pool knowledge from different methods across different tasks. Additionally, we demonstrate several baselines for all tasks and show that approaches for one task can be applied to another with a quantifiable and explainable performance difference. Dataset annotations are available at: https://github.com/Ali2500/BURST-benchmark.
Ali Athar, Jonathon Luiten, Paul Voigtlaender, Tarasha Khurana, Achal Dave, Bastian Leibe, Deva Ramanan
WACV1
2022 HODOR: High-level Object Descriptors for Object Re-segmentation in Video Learned from Static Images
abstract
Existing state-of-the-art methods for Video Object Segmentation (VOS) learn low-level pixel-to-pixel correspondences between frames to propagate object masks across video. This requires a large amount of densely annotated video data, which is costly to annotate, and largely redundant since frames within a video are highly correlated. In light of this, we propose HODOR: a novel method that tackles VOS by effectively leveraging annotated static images for understanding object appearance and scene context. We encode object instances and scene information from an image frame into robust high-level descriptors which can then be used to re-segment those objects in different frames. As a result, HODOR achieves state-of-the-art performance on the DAVIS and YouTube-VOS benchmarks compared to existing methods trained without video annotations. With out any architectural modification, HODOR can also learn from video context around single annotated video frames by utilizing cyclic consistency, whereas other methods rely on dense, temporally consistent annotations. Source code: https://github.com/Ali2500/HODOR.
Ali Athar, Jonathon Luiten, Alexander Hermans, Deva Ramanan, Bastian Leibe
CVPR1
2022 D2Conv3D: Dynamic Dilated Convolutions for Object Segmentation in Videos
abstract
Despite receiving significant attention from the research community, the task of segmenting and tracking objects in monocular videos still has much room for improvement. Existing works have simultaneously justified the efficacy of dilated and deformable convolutions for various image-level segmentation tasks. This gives reason to believe that 3D extensions of such convolutions should also yield performance improvements for video-level segmentation tasks. However, this aspect has not yet been explored thoroughly in existing literature. In this paper, we propose Dynamic Dilated Convolutions (D2Conv3D): a novel type of convolution which draws inspiration from dilated and deformable convolutions and extends them to the 3D (spatio-temporal) domain. We experimentally show that D2Conv3D can be used to improve the performance of multiple 3D CNN architectures across multiple video segmentation related benchmarks by simply employing D2Conv3D as a drop-in replacement for standard convolutions. We further show that D2Conv3D out-performs trivial extensions of existing dilated and deformable convolutions to 3D. Lastly, we set a new state-of-the-art on the DAVIS 2016 Unsupervised Video Object Segmentation benchmark. Code is made publicly available at https://github.com/Schmiddo/d2conv3d.
Christian Schmidt 0029, Ali Athar, Sabarinath Mahadevan, Bastian Leibe
WACV2
2020 Making a Case for 3D Convolutions for Object Segmentation in Videos
Sabarinath Mahadevan, Ali Athar, Aljosa Osep, Laura Leal-Taixé, Bastian Leibe, Sebastian Hennen
BMVC2
2020 STEm-Seg: Spatio-Temporal Embeddings for Instance Segmentation in Videos
Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixé, Bastian Leibe
ECCV (11)1
2016 Whole-body motion planning for humanoid robots with heuristic search
abstract
The task of whole-body motion planning for humanoid robots is challenging due to its high-DOF nature, stability constraints, and the need for obstacle avoidance and movements that are efficient. Over the years, various approaches have been adopted to solve this problem such as bounding-box models and jacobian-based techniques. More commonly though, sampling-based algorithms are employed for this task since they perform admirably well in high-dimensional spaces. As an alternative, search-based planners offer improvements in terms of optimality and consistency of the solution. However, they are normally considered impractical for high-dimensional motion planning. In this paper, we present a heuristic search-based motion planning framework for humanoid robots that circumvents the drawbacks traditionally associated with search-based planners while catering to the specific requirements of humanoid motion planning. This is achieved primarily through a combination of informative yet computationally inexpensive heuristics, carefully crafted motion primitives as atomic actions, and a whole body inverse kinematics solver for achieving desired end effector orientations. The experimental results show the ability of our framework to perform complex motion planning tasks quickly and efficiently.
Ali Athar, Abdul Moeed Zafar, Rizwan Asif, Armaghan Ahmad Khan, Yasar Ayaz, Osman Hasan
IROS1