Alen Alempijevic

dblp:83/4673 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-1769-8041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 5 since 2021Systems, architecture and hardware · 17 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2024 Decentralized multi-phase formation control for cattle herding
abstract
Herding is performed by people or trained animals to control the movement of livestock under the desired direction of an operator. This paper presents a novel decentralized control strategy for a group of robots to herd animals which consists of two phases, a surrounding phase and a driving phase. In the surrounding phase, a custom artificial potential field is employed to simultaneously guide the robots to encircle the herd by tracking the outmost animals and maintaining a safe distance from other neighboring robots. Once the encirclement is complete, the robots transition to drive the animals toward a designated goal by simply maintaining their initial formation and traversing to it. Unlike existing works on herding using flocking control, local observations of the nearest animals and communication with other robots within the sensing range are the only requirements for the robots to surround and herd the animals effectively. Moreover, the animal-robot behavior model resembles the interaction of livestock in the presence of an external predatory threat, where robots act as predators. An analytical proof and empirical results collected from different simulators demonstrate that the proposed control enables the robots to converge around the boundary of the animals and guide them toward the designated goal.
Dac Dang Khoa Nguyen, Gavin Paul, Alen Alempijevic
ICRA3
2024 Sim2real Cattle Joint Estimation in 3D point clouds
abstract
Understanding the well-being of cattle is crucial in various agricultural contexts. Cattle’s body shape and joint articulation carry significant information about their welfare, yet acquiring comprehensive datasets for 3D body pose estimation presents a formidable challenge. This study delves into the construction of such a dataset specifically tailored for cattle. Leveraging the expertise of digital artists, we use a single animated 3D model to represent diverse cattle postures. To address the disparity between virtual and real-world data, we augment the 3D model’s shape to encompass a range of potential body appearances, thereby narrowing the "sim2real" gap. We use these annotated models to train a deep-learning framework capable of estimating internal joints solely based on external surface curvature. Our contribution is specifically the use of geodesic distance over the surface manifold, coupled with multilateration to extract joints in a semantic keypoint detection encoder-decoder architecture. We demonstrate the robustness of joint extraction by comparing the link lengths extracted on real cattle mobbing and walking within a race. Furthermore, inspired by the established allometric relationship between bone length and the overall height of mammals, we utilise the estimated joints to predict hip height within a real cattle dataset, extending the utility of our approach to offer insights into improving cattle monitoring practices.
Mohammad Okour, Raphael Falque, Alen Alempijevic
IROS3
2023 Semantic Keypoint Extraction for Scanned Animals using Multi-Depth-Camera Systems
abstract
Keypoint annotation in pointclouds is an important task for 3D reconstruction, object tracking and alignment, in particular in deformable or moving scenes. In the context of agriculture robotics, it is a critical task for livestock automation to work toward condition assessment or behaviour recognition. In this work, we propose a novel approach for semantic keypoint annotation in pointclouds, by reformulating the keypoint extraction as a regression problem of the distance between the keypoints and the rest of the pointcloud. We use the distance on the pointcloud manifold mapped into a radial basis function (RBF), which is then learned using an encoder-decoder architecture. Special consideration is given to the data augmentation specific to multi-depth-camera systems by considering noise over the extrinsic calibration and camera frame dropout. Additionally, we investigate computationally efficient non-rigid deformation methods that can be applied to animal pointclouds. Our method is tested on data collected in the field, on moving beef cattle, with a calibrated system of multiple hardware-synchronised RGB-D cameras.
Raphael Falque, Teresa Vidal-Calleja, Alen Alempijevic
ICRA3
2023 Skirting Line Estimation Using Sparse to Dense Deformation
abstract
Automating the process of fleece contaminant removal has the potential to drastically improve the quality of wool leaving the farm gate. Towards this goal, we present a method to automatically extract skirting lines, i.e., the separations between clean and contaminated wool of a fleece using RGB images. We propose a learning-based sparse-to-dense approach for estimating the non-rigid deformation of fleeces in order to estimate the skirting lines. Our method is bootstrapped from a set of sparse inlier feature correspon-dences, which are heavily filtered through a set of strict criteria. The inlier correspondences are then greedily expanded by adding correspondences from a denser set through a filtering process. This process is based on a learning approach that takes as inputs the pixel similarity and the consistency with their inlier neighbours. Each greedy iteration is initialised with a non-rigid deformation using as-rigid-as-possible as a prior to the filtering process. The proposed method outperforms both a rigid deformation baseline and optic flow deep learning approach, as evidenced by the quantitative evaluation of pixel location error in controlled experiments. To further prove its practicality, we demonstrate qualitative results comparing the predicted skirting line from various methods on images of skirted fleeces collected from several wool sheds.
Daniel Perez Banuelos, Raphael Falque, Tim Patten, Alen Alempijevic
IROS4
2021 Probabilistic Dynamic Crowd Prediction for Social Navigation
abstract
In this paper, we present a novel approach that predicts spatially and temporally crowd behaviour for robotic social navigation. Integrating mobile robots into human society involves the fundamental problem of navigation in crowds. A robot should attempt to navigate in a way that is minimally invasive to the humans in its environment. However, planning in a dynamic environment is difficult as the environment must be predicted into the future. This problem has been thoroughly studied considering the behaviour of pedestrians at the level of individuals. Instead, we represent a pedestrian crowd by its macroscopic properties over space, such as density and velocity. With this spatial representation, we propose to learn a convolutional recurrent model to predict these properties into the future. The key design of a probabilistic loss function capturing the crowd's macroscopic properties empowers the spatio-temporal crowd prediction. Using a social invasiveness metric defined on these properties predicted by our convolutional recurrent model, we develop a framework that produces globally-optimal plans in expectation. Extensive results using a realistic pedestrian simulator show the validity and performance of the proposed social navigation approach.
Stefan H. Kiss, Kavindie Katuwandeniya, Alen Alempijevic, Teresa Vidal-Calleja
ICRA3
2021 Learning Image-Based Contaminant Detection in Wool Fleece from Noisy Annotations
Tim Patten, Alen Alempijevic, Robert Fitch
ICVS2
2018 Socially Constrained Tracking in Crowded Environments Using Shoulder Pose Estimates
abstract
Detecting and tracking people is a key requirement in the development of robotic technologies intended to operate in human environments. In crowded environments such as train stations this task is particularly challenging due the high numbers of targets and frequent occlusions. In this paper we present a framework for detecting and tracking humans in such crowded environments in terms of 2D pose ( x, y, θ). The main contributions are a method for extracting pose from the most visible parts of the body in a crowd, the head and shoulders, and a tracker which leverages social constraints regarding peoples orientation, movement and proximity to one another, to improve robustness in this challenging environment. The framework is evaluated on two datasets: one captured in a lab environment with ground truth obtained using a motion capture system, and the other captured in a busy inner city train station. Pose errors are reported against the ground truth and the tracking results are then compared with a state-of-the-art person tracking framework.
Alexander Virgona, Alen Alempijevic, Teresa Vidal-Calleja
ICRA2
2018 Continuous Optimization Framework for Depth Sensor Viewpoint Selection
Behnam Maleki, Alen Alempijevic, Teresa Vidal-Calleja
WAFR2
2016 Exploring in 3D with a climbing robot: Selecting the next best base position on arbitrarily-oriented surfaces
abstract
This paper presents an approach for selecting the next best base position for a climbing robot so as to observe the highest information gain about the environment. The robot is capable of adhering to and moving along and transitioning to surfaces with arbitrary orientations. This approach samples known surfaces, and takes into account the robot kinematics, to generate a graph of valid attachment points from which the robot can either move to other positions or make observations of the environment. The information value of nodes in this graph are estimated and a variant of A* is used to traverse the graph and discover the most worthwhile node that is reachable by the robot. This approach is demonstrated in simulation and shown to allow a 7 degree-of-freedom inchworm-inspired climbing robot to move to positions in the environment from which new information can be gathered about the environment.
Phillip Quin, Gavin Paul, Alen Alempijevic, Dikai Liu
IROS3
2013 Bootstrapping navigation and path planning using human positional traces
abstract
Navigating and path planning in environments with limited a priori knowledge is a fundamental challenge for mobile robots. Robots operating in human-occupied environments must also respect sociocontextual boundaries such as personal workspaces. There is a need for robots to be able to navigate in such environments without having to explore and build an intricate representation of the world. In this paper, a method for supplementing directly observed environmental information with indirect observations of occupied space is presented. The proposed approach enables the online inclusion of novel human positional traces and environment information into a probabilistic framework for path planning. Encapsulation of sociocontextual information, such as identifying areas that people tend to use to move through the environment, is inherently achieved without supervised learning or labelling. Our method bootstraps navigation with indirectly observed sensor data, and leverages the flexibility of the Gaussian process (GP) for producing a navigational map that sampling based path planers such as Probabilistic Roadmaps (PRM) can effectively utilise. Empirical results on a mobile platform demonstrate that a robot can efficiently and socially-appropriately reach a desired goal by exploiting the navigational map in our Bayesian statistical framework.
Alen Alempijevic, Robert Fitch, Nathan Kirchner
ICRA1
2013 Efficient neighbourhood-based information gain approach for exploration of complex 3D environments
abstract
This paper presents an approach for exploring a complex 3D environment with a sensor mounted on the end effector of a robot manipulator. In contrast to many current approaches which plan as far ahead as possible using as much environment information as is available, our approach considers only a small set of poses (vector of joint angles) neighbouring the robot's current pose in configuration space. Our approach is compared to an existing exploration strategy for a similar robot. Our results demonstrate a significant decrease in the number of information gain estimation calculations that need to be performed, while still gathering an equivalent or increased amount of information about the environment.
Phillip Quin, Gavin Paul, Alen Alempijevic, Dikai Liu, Gamini Dissanayake
ICRA3
2013 A robot centric perspective on the HRI paradigm
abstract
The industrial revolution undoubtedly defined the role of machines in our society, and it directly shaped the paradigm for human machine interaction - a paradigm which was inherited by the field of Human Robot Interaction (HRI) as the machines became robots. This paper argues that, for a foreseeable set of interactions, reshaping this paradigm would result in more effective and more often successful interactions. This paper presents our Robot Centric paradigm for HRI. Evidence in the form of summaries of relevant literature and our past efforts in developing social-robotics enabling technology is presented to support our paradigm. A definition and a set of recommendations for designing the key enabling component, sociocontextual cues, of our paradigm are presented. Finally, empirical evidence generated through a number of experiments and field studies (N = 456 and N = 320) demonstrates our paradigm is both feasibly incorporated into HRI and moreover, yields significant contributions to the successfulness of a set of HRIs.
Nathan Kirchner, Alen Alempijevic
J. Hum. Robot Interact.2
2012 Head-to-shoulder signature for person recognition
abstract
Ensuring that an interaction is initiated with a particular and unsuspecting member of a group is a complex task. As a first step the robot must effectively, expediently and reliably recognise the humans as they carry on with their typical behaviours (in situ). A method for constructing a scale and viewing angle robust feature vector (from analysing a 3D pointcloud) designed to encapsulate the inter-person variations in the size and shape of the people's head to shoulder region (Head-to-shoulder signature - HSS) is presented. Furthermore, a method for utilising said feature vector as the basis of person recognition via a Support-Vector Machine is detailed. An empirical study was performed in which person recognition was attempted on in situ data collected from 25 participants over 5 days in a office environment. The results report a mean accuracy over the 5 days of 78.15% and a peak accuracy 100% for 9 participants. Further, the results show a considerably better-than-random (1/23 = 4.5%) result for when the participants were: in motion and unaware they were being scanned (52.11%), in motion and face directly away from the sensor (36.04%), and post variations in their general appearance. Finally, the results show the HSS has considerable ability to accommodate for a person's head, shoulder and body rotation relative to the sensor - even in cases where the person is faced directly away from the robot.
Nathan Kirchner, Alen Alempijevic, Alexander Virgona
ICRA2
2012 Towards robust vision-based self-localization of vehicles in dense urban environments
abstract
Self-localization of ground vehicles in densely populated urban environments poses a significant challenge. The presence of tall buildings in close proximity to traversable areas limits the use of GPS-based positioning techniques in such environments. This paper presents an approach to global localization on a hybrid metric-topological map using a monocular camera and wheel odometry. The global topology is built upon spatially separated reference places represented by local image features. In contrast to other approaches we employ a feature selection scheme ensuring a more discriminative representation of reference places while simultaneously rejecting a multitude of features caused by dynamic objects. Through fusion with additional local cues the reference places are assigned discrete map positions allowing metric localization within the map. The self-localization is carried out by associating observed visual features with those stored for each reference place. Comprehensive experiments in a dense urban environment covering a time difference of about 9 months are carried out. This demonstrates the robustness of our approach in environments subjected to high dynamic and environmental changes.
Marian Himstedt, Alen Alempijevic, Liang Zhao 0003, Shoudong Huang, Hans-Joachim Böhme
IROS2
2012 A robust RGB-D SLAM algorithm
abstract
Recently RGB-D sensors have become very popular in the area of Simultaneous Localisation and Mapping (SLAM). The major advantage of these sensors is that they provide a rich source of 3D information at relatively low cost. Unfortunately, these sensors in their current forms only have a range accuracy of up to 4 metres. Many techniques which perform SLAM using RGB-D cameras rely heavily on the depth and are restrained to office type and geometrically structured environments. In this paper, a switching based algorithm is proposed to heuristically choose between RGB-BA and RGBD-BA based local maps building. Furthermore, a low cost and consistent optimisation approach is used to join these maps. Thus the potential of both RGB and depth image information are exploited to perform robust SLAM in more general indoor cases. Validation of the proposed algorithm is performed by mapping a large scale indoor scene where traditional RGB-D mapping techniques are not possible.
Gibson Hu, Shoudong Huang, Liang Zhao 0003, Alen Alempijevic, Gamini Dissanayake
IROS4
2011 Nonverbal robot-group interaction using an imitated gaze cue
abstract
Ensuring that a particular and unsuspecting member of a group is the recipient of a salient-item hand-over is a complicated interaction. The robot must effectively, expediently and reliably communicate its intentions to advert any tendency within the group towards antinormative behaviour. In this paper, we study how a robot can establish the participant roles of such an interaction using imitated social and contextual cues. We designed two gaze cues, the first was designed to discourage antinormative behaviour through individualising a particular member of the group and the other to the contrary. We designed and conducted a field experiment (456 participants in 64 trials) in which small groups of people (between 3 and 20 people) assembled in front of the robot, which then attempted to pass a salient object to a particular group member by presenting a physical cue, followed by one of two variations of a gaze cue. Our results showed that presenting the individualising cue had a significant (z=3.733, p=0.0002) effect on the robot's ability to ensure that an arbitrary group member did not take the salient object and that the selected participant did.
Nathan Kirchner, Alen Alempijevic, Gamini Dissanayake
HRI2
2011 Learning navigational maps by observing human motion patterns
abstract
Observing human motion patterns is informative for social robots that share the environment with people. This paper presents a methodology to allow a robot to navigate in a complex environment by observing pedestrian positional traces. A continuous probabilistic function is determined using Gaussian process learning and used to infer the direction a robot should take in different parts of the environment. The approach learns and filters noise in the data producing a smooth underlying function that yields more natural movements. Our method combines prior conventional planning strategies with most probable trajectories followed by people in a principled statistical manner, and adapts itself online as more observations become available. The use of learning methods are automatic and require minimal tuning as compared to potential fields or spline function regression. This approach is demonstrated testing in cluttered office and open forum environments using laser and vision sensing modalities. It yields paths that are similar to the expected human behaviour without any a priori knowledge of the environment or explicit programming.
Simon Timothy O'Callaghan, Surya P. N. Singh, Alen Alempijevic, Fabio Ramos 0001
ICRA3
2009 Cross-modal localization through mutual information
abstract
Relating information originating from disparate sensors observing a given scene is a challenging task, particularly when an appropriate model of the environment or the behaviour of any particular object within it is not available. One possible strategy to address this task is to examine whether the sensor outputs contain information which can be attributed to a common cause. In this paper, we present an approach to localise this embedded common information through an indirect method of estimating mutual information between all signal sources. Ability of L1regularization to enforce sparseness of the solution is exploited to identify a subset of signals that are related to each other, from among a large number of sensor outputs. As opposed to the conventional L2regularization, the proposed method leads to faster convergence with much reduced spurious associations. Simulation and experimental results are presented to validate the findings.
Alen Alempijevic, Sarath Kodagoda, Gamini Dissanayake
IROS1
2007 Towards improving driver situation awareness at intersections
abstract
Providing safety critical information to the driver is vital in reducing road accidents, especially at intersections. Intersections are complex to deal with due to the presence of large number of vehicle and pedestrian activities, and possible occlusions. Information available from only the sensors onboard a vehicle has limited value in this scenario. In this paper, we propose to utilize sensors on-board the vehicle of interest as well as the sensors that are mounted on nearby vehicles to enhance the driver situation awareness. The resulting major research challenge of sensor registration with moving observers is solved using a mutual information based technique. The response of the sensors to common causes are identified and exploited for computing their unknown relative locations. Experimental results, for a mock up traffic intersection in which mobile robots equipped with laser range finders are used, are presented to demonstrate the efficacy of the proposed technique.
Sarath Kodagoda, Alen Alempijevic, Stephan Sehestedt, Gamini Dissanayake
IROS2
2007 Robust lane detection in urban environments
abstract
Most of the lane marking detection algorithms reported in the literature are suitable for highway scenarios. This paper presents a novel clustered particle filter based approach to lane detection, which is suitable for urban streets in normal traffic conditions. Furthermore, a quality measure for the detection is calculated as a measure of reliability. The core of this approach is the usage of weak models, i.e. the avoidance of strong assumptions about the road geometry. Experiments were carried out in Sydney urban areas with a vehicle mounted laser range scanner and a ccd camera. Through experimentations, we have shown that a clustered particle filter can be used to efficiently extract lane markings.
Stephan Sehestedt, Sarath Kodagoda, Alen Alempijevic, Gamini Dissanayake
IROS3
2006 Sensor Registration and Calibration using Moving Targets
abstract
Multimodal sensor registration and calibration are crucially important aspects in distributed sensor fusion. Unknown relationships of sensors and joint probability distribution between sensory signals make the sensor fusion nontrivial. In this paper, we adopt a Mutual Information (MI) based approach for sensor registration and calibration. It is based on unsupervised learning of a nonparametric sensing model by maximizing mutual information between signal streams. Experiments were carried out in an office like environment with two laser sensors capturing arbitrarily moving people. Attributes of the moving targets are used. Problems due to target occlusions are alleviated by the multiple model tracker. The registration and calibration methodology does not require any artificially generated patterns or motions unlike other calibration methodologies
Sarath Kodagoda, Alen Alempijevic, James Patrick Underwood, Gamini Dissanayake
ICARCV2
2006 Mutual Information based Sensor Registration and Calibration
abstract
Knowledge of calibration, that defines the location of sensors relative to each other, and registration, that relates sensor response due to the same physical phenomena, are essential in order to be able to fuse information from multiple sensors. In this paper, a mutual information (MI) based approach for automatic sensor registration and calibration is presented. Unsupervised learning of a nonparametric sensing model by maximizing mutual information between signal streams is used to relate information from different sensors, allowing unknown sensor registration and calibration to be determined. Experiments conducted in an office environment are used to illustrate the effectiveness of the proposed technique. Two laser sensors are used to capture people mobbing in an arbitrarily manner in the environment and MI from a number of attributes of the motion are used for relating the signal streams from the sensors. Thus the sensor registration and calibration is achieved without using artificial patterns or pre-specified motions
Alen Alempijevic, Sarath Kodagoda, James Patrick Underwood, Gamini Dissanayake
IROS1