Rohan Sarkar

dblp:151/9488 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-3256-7277ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 28% Robot navigation and mapping · 28% Representation and self-supervised learning · 28%

Topics — the 4 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
embedding learning
0.812024
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval · CVPR 2024
Robotics › Robot navigation and mapping
object search
0.812024
Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval · CVPR 2024
Robotics › Robot manipulation › nonprehensile manipulation
dynamic manipulation
0.212014
Resonance-driven dynamic manipulation: Dribbling and juggling with elastic beam · ICRA 2014
Robotics › Motion planning and robot control
robot control
0.212014
Resonance-driven dynamic manipulation: Dribbling and juggling with elastic beam · ICRA 2014

Methods — techniques the papers use, named apart from their topics

ranking loss · 0.8contrastive learning · 0.8attention · 0.8vibration control · 0.2resonant mode analysis · 0.2
YearPublicationVenuePosition
2026 A Dataset and Framework for Learning State-invariant Object Representations
abstract
We introduce state invariance - robustness to changes in an object’s structural form (as for a folded umbrella or for crumpled clothing) - complementing other common invariances to learn object representations for recognition and retrieval tasks. For this, we present ObjectsWithStateChange (OWSC), a novel dataset that captures variations in object appearance arising from state changes, along with pose, viewpoint, and illumination variations, to advance research in fine-grained 3D object recognition and retrieval. A key challenge is that objects within and across categories may look visually similar under certain state changes, making discrimination difficult. To address this, we propose a curriculum learning based mining strategy that progressively samples harder object pairs based on inter-object distances in the learned embedding space after each epoch, gradually sampling harder-to-distinguish examples of visually similar objects from within and across categories during training. Our ablation shows that this curriculum learning strategy enhances the model’s ability to learn discriminative invariant features for fine-grained tasks, improving object recognition accuracy by 7.9% and retrieval mAP by 9.2% over prior methods on our new OWSC dataset and three other multi-view datasets, such as ModelNet40, ObjectPI, FG3D. Our OWSC dataset is available at https://github.com/sarkar-rohan/ObjectsWithStateChange.
Rohan Sarkar, Avinash C. Kak
WACV1
2024 Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
abstract
In the context of pose-invariant object recognition and retrieval, we demonstrate that it is possible to achieve significant improvements in performance if both the category-based and the object-identity-based embed-dings are learned simultaneously during training. In hind-sight, that sounds intuitive because learning about the cat-egories is more fundamental than learning about the indi-vidual objects that correspond to those categories. How-ever, to the best of what we know, no prior work in pose-invariant learning has demonstrated this effect. This paper presents an attention-based dual-encoder architecture with specially designed loss functions that optimize the inter-and intra-class distances simultaneously in two different embedding spaces, one for the category embeddings and the other for the object level embeddings. The loss functions we have proposed are pose-invariant ranking losses that are designed to minimize the intra-class distances and maximize the inter-class distances in the dual representation spaces. We demonstrate the power of our approach with three challenging multi-view datasets, Model Net-40, ObjectPI, and FG3D. With our dual approach, for single-view object recognition, we outperform the previous best by 20.0% on ModelNet40, 2.0% on ObjectPI, and 46.5% on FG3D. On the other hand, for single-view object retrieval, we outperform the previous best by 33.7% on ModelNet40, 18.8% on ObjectPI, and 56.9% on FG3D.
Rohan Sarkar, Avinash C. Kak
CVPR1
2023 OutfitTransformer: Learning Outfit Representations for Fashion Recommendation
abstract
Learning an effective outfit-level representation is critical for predicting the compatibility of items in an outfit, and retrieving complementary items for a partial outfit. We present a framework, OutfitTransformer, that uses the pro-posed task-specific tokens and leverages the self-attention mechanism to learn effective outfit-level representations en-coding the compatibility relations between all items in the entire outfit for addressing both compatibility prediction and complementary item retrieval. For compatibility pre-diction, we design an outfit token to capture a global out-fit representation and train the framework using a classification loss. For complementary item retrieval, we design a target item token that additionally takes the target item specification (in the form of a category or text description) into consideration. We train our framework using a pro-posed set-wise outfit ranking loss to generate a target item embedding given an outfit, and a target item specification as inputs. The generated target item embedding is then used to retrieve compatible items that match the rest of the out-fit. Additionally, we adopt a pre-training approach and a curriculum learning strategy to improve retrieval performance. Experiments show that our approach outperforms state-of-the-art methods on compatibility prediction, fill-in-the-blank, and complementary item retrieval tasks.
Rohan Sarkar, Navaneeth Bodla, Mariya I. Vasileva, Yen-Liang Lin, Anurag Beniwal, Alan Lu, Gérard G. Medioni
WACV1
2014 Resonance-driven dynamic manipulation: Dribbling and juggling with elastic beam
abstract
This paper presents a new device and a method for dynamic manipulation. The device consists of a planar robotic arm and an elastic beam as an end-effector. Using it the elastic end-effector will tend to increase performance and energy efficiency while executing dynamic and repetitive tasks. Through the control of the beam vibration and resonant modes, we modify the state of manipulated objects. For lightweight objects the control is provided through the intermittent contacts without changing dynamics of the beam. However, we show that by using proper synchronization technique continuous-phase contacts are also possible. Juggling and dribbling of a ball are considered to be an alternating non-prehensile catching and throwing task. Such alternating decelerating and accelerating impacts on the ball and the curvature of the beam at the time of impact will stabilize the cyclic orbit of the ball. By proper analysis of continuous-time contact and dynamics of the beam we establish a rhythmic movement of the system. With the variation of frequency and amplitude of the beam it is possible to switch between different dynamic actions such as juggling, dribbling, throwing, catching and balancing.
Alexander Pekarovskiy, Kunal Saluja, Rohan Sarkar, Martin Buss
ICRA3