James Akl

dblp:292/2967 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-1025-404XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 40% Segmentation and scene understanding · 21% 3D vision · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
object pose estimation
0.912025
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation · ICRA 2025
Robotics › Robot manipulation › tactile sensing › tactile perception
visuo-tactile perception
0.912025
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation · ICRA 2025
Robotics › Robot manipulation
grasping
0.822025
ECNNs: Ensemble Learning Methods for Improving Planar Grasp Quality Estimation · ICRA 2021
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation · ICRA 2025
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Computer vision › Image recognition and object detection
object detection
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Computer vision › Segmentation and scene understanding
object segmentation
0.612022
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes · CVPR 2022
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.512021
ECNNs: Ensemble Learning Methods for Improving Planar Grasp Quality Estimation · ICRA 2021
Robotics › Robot manipulation › grasping
grasp quality evaluation
0.512021
ECNNs: Ensemble Learning Methods for Improving Planar Grasp Quality Estimation · ICRA 2021
Machine learning › Deep learning architectures and training
convolutional neural network
0.112021
ECNNs: Ensemble Learning Methods for Improving Planar Grasp Quality Estimation · ICRA 2021

Methods — techniques the papers use, named apart from their topics

semantic segmentation · 1.1instance segmentation · 1.1visual foundation model · 0.9test-time optimization · 0.9spring-mass model · 0.9mixture of experts · 0.5ensemble learning · 0.5convolutional neural network · 0.5
YearPublicationVenuePosition
2025 ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
abstract
Object 6D pose estimation is a critical challenge in robotics, particularly for manipulation tasks. While prior research combining visual and tactile (visuotactile) information has shown promise, these approaches often struggle with generalization due to the limited availability of visuotactile data. In this paper, we introduce ViTa-Zero, a zero-shot visuotactile pose estimation framework. Our key innovation lies in leveraging a visual model as its backbone and performing feasibility checking and test-time optimization based on physical constraints derived from tactile and proprioceptive observations. Specifically, we model the gripper-object interaction as a spring-mass system, where tactile sensors induce attractive forces, and proprioception generates repulsive forces. We validate our framework through experiments on a real-world robot setup, demonstrating its effectiveness across representative visual backbones and manipulation scenarios, including grasping, object picking, and bimanual handover. Compared to the visual models, our approach overcomes some drastic failure modes while tracking the in-hand object pose. In our experiments, our approach shows an average increase of 55% in AUC of ADD-S and 60% in ADD, along with an 80% lower position error compared to FoundationPose.
Hongyu Li 0003, James Akl, Srinath Sridhar 0002, Tye Brady, Taskin Padir
ICRA2
2024 Feature-Driven Next View Planning for Cutting Path Generation in Robotic Metal Scrap Recycling
abstract
Metal recycling in scrapyards, where workers cut decommissioned structures using gas torches, is labor-intensive, difficult, and dangerous. As global metal scrap recycling demands are rising, robotics and automation technologies could play a significant role to address this demand. However, the unstructured nature of the scrap cutting problem—due to highly variable object shapes and environments—poses significant challenges to integrate robotic solutions. We propose a novel collaborative workflow for robotic metal cutting that combines worker expertise with robot autonomy. In this workflow, the skilled worker studies the scene, determines an appropriate cutting reference, and marks it on the object with spray paint. The robot, then, autonomously explores the surface of the object for identifying and reconstructing the drawn reference, converts it to a cutting trajectory, and finally executes the cut. This paper focuses on the surface exploration and cutting reference reconstruction tasks, which require appropriate next view planning (NVP) algorithms. We devise three NVP algorithms enabling the robot to explore and extract desired features from the scene,i.e., the drawn reference, without requiring anya prioriobject model. Contrasting with global or feature-agnostic NVP algorithms, our approaches guide the robot via desired local features to increase the efficiency of the exploration. We evaluate our NVP algorithms against six categories of objects both in simulation and in physical experiments.Note to Practitioners—This work is motivated by the need of extracting a desired cutting reference determined and drawn on the object by scrap yard workers. From the robot’s perspective, it must explore and reconstruct the drawing, starting from an unknown scene containing an unknown object featuring an unknown drawing. We assume that an RGB-D camera is attached to the tool-tip of the robot, and the color of the drawn path is significantly different from the object’s color. The goal of the robotic system is to explore the object surface to uncover the drawn path entirely without colliding with the object. The exploration algorithm must overcome complex object shapes and must be fast enough for practical use in scrap yards. This means conventional exploration (active vision) techniques are insufficient since they focus on exploring the entirety of the object, which is unnecessary for our task and is time-consuming. Our methods exploit the drawing information to guide the exploration for quickly determining a suitable viewpoint, which results in an efficient extraction of the entire cutting reference, without needing to explore the entire object’s surface. Our algorithms are robust against adversarial features such as discontinuous, non-smooth, or self-occluded object surfaces. Our feature-driven strategies are not limited to robotic scrap cutting as they are applicable to any viewpoint planning problem requiring high performance while extracting the local features in the scene.
James Akl, Fadi M. Alladkani, Berk Çalli
IEEE Trans Autom. Sci. Eng.1
2023 Vision-Based Oxy-Fuel Torch Control for Robotic Metal Cutting
abstract
The automation of key processes in metal cutting would substantially benefit many industries such as manufacturing and metal recycling. We present a vision-based control scheme for automated metal cutting with oxy-fuel torches, an established cutting medium in industry. The system consists of a robot equipped with a cutting torch and an eye-in-hand camera observing the scene behind a tinted visor. We develop a vision-based control algorithm to servo the torch's motion by visually observing its effects on the metal surface. As such, the vision system processes the metal surface's heat pool and computes its associated features, specifically pool convexity and intensity, which are then used for control. The operating conditions of the control problem are defined within which the stability is proven. In addition, metal cutting experiments are performed using a physical 1-DOF robot and oxy-fuel cutting equipment. Our results demonstrate the successful cutting of metal plates across three different plate thicknesses, relying purely on visual information without a priori knowledge of the thicknesses.
James Akl, Yash Patil, Chinmay Todankar, Berk Çalli
IROS1
2022 ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
abstract
Less than 35% of recyclable waste is being actually recycled in the US [2], which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/
Dina Bashkirova, Mohamed Abdelfattah, Ziliang Zhu, James Akl, Fadi M. Alladkani, Ping Hu 0001, Vitaly Ablavsky, Berk Çalli, Sarah Adel Bargal, Kate Saenko
CVPR4
2021 ECNNs: Ensemble Learning Methods for Improving Planar Grasp Quality Estimation
abstract
We present an ensemble learning methodology that combines multiple existing robotic grasp synthesis algorithms and obtain a success rate that is significantly better than the individual algorithms. The methodology treats the grasping algorithms as "experts" providing grasp "opinions". An Ensemble Convolutional Neural Network (ECNN) is trained using a Mixture of Experts (MOE) model that integrates these opinions and determines the final grasping decision. The ECNN introduces minimal computational cost overhead, and the network can virtually run as fast as the slowest expert. We test this architecture using open-source algorithms in the literature by adopting GQCNN 4.0, GGCNN and a custom variation of GGCNN as experts and obtained a 6% increase in the grasp success on the Cornell Dataset compared to the best-performing individual algorithm. The performance of the method is also demonstrated using a Franka Emika Panda arm.
Fadi M. Alladkani, James Akl, Berk Çalli
ICRA2