VLDB 2026 Research / reviewers in the wild / expert
Jorge Dias 0001
dblp:159/8105 · also Jorge Manuel Miranda Dias
· DBLP profile ↗
112ranked-venue papers
4as first author
33since 2021 · last 2025
0000-0002-2725-8867ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 3 first-author · 16 since 2021Systems, architecture and hardware · 49 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 since 2021Databases, data management, data science and information retrieval · 4Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Online Risk-Bounded Graph-Based Local Planning for Autonomous Driving With Theoretical GuaranteesabstractRisk-bounded motion planning in dynamic environments for autonomous driving presents complex challenges, particularly in solving the nonconvex problem of ensuring continuous, safe, and real-time navigation towards a destination. This paper introduces an online graph-based local planning approach constrained by a user-defined driving style in terms of a risk budget$\Delta$for the entire mission. Our online approach assigns a risk bound to each motion planning decision, ensuring that the total risk consumed remains within$\Delta$. First, we construct a spatial lattice graph that adheres to the vehicle's curvature constraints. Then, the trajectory planning problem is reformulated as an online optimization problem, where decisions must be made sequentially without prior knowledge of future events. Therefore, we propose a reduction to the problem to be online multiple-choice knapsack problem (ON-MCKP), where the knapsack items are candidate paths generated by solving constrained shortest-path problems. To solve the ON-MCKP, we deploy online algorithms that offer theoretical guarantees on the risk allocation throughout the entire mission. The effectiveness of our method is demonstrated empirically, showing significant improvements in the objective without violating safety constraints. Abdulrahman Ahmad, Majid Khonji, Khaled M. Elbassioni, Jorge Dias 0001, Ameena Saad Al-Sumaiti |
ICRA | 4 |
| 2025 | PMIL: A Topology Module to Improve MIL-based WSI ClassificationabstractDeep learning models have achieved remarkable success in pathology image analysis. However, they still face challenges in effectively modeling fine-grained, object-level features. Topological Data Analysis (TDA) has shown promise for addressing these issues but remains underexplored, particularly for whole-slide pathology applications. Additionally, the effectiveness of TDA has yet to be firmly established, as current studies largely use small-scale datasets. In this work, we address these gaps by introducing Persistent Homology in Multiple Instance Learning (PMIL), the first adaptable TDA-based module within the MIL framework. We validate our approach on a large-scale classification dataset, benchmarking against multiple state-of-the-art methods. Ahmad Obeid 0001, Anabia Sohail, Said Boumaraf, Xiabi Liu, Sajid Javed, Hasan Almarzouqi, Jorge Dias 0001, Mohammed Bennamoun, Naoufel Werghi, Ibrahim M. Elfadel |
ISCAS | 7 |
| 2025 | RobMOT: 3D Multi-Object Tracking Enhancement Through Observational Noise and State Estimation Drift Mitigation in LiDAR Point CloudsabstractThis paper addresses key limitations in recent 3D tracking-by-detection methods, focusing on the challenges of identifying legitimate trajectories and mitigating state estimation drift in the Kalman filter. Current methods rely heavily on threshold-based detection score filtering approaches to reduce false positives and prevent ghost trajectories. However, these approaches fail for distant and partially occluded objects, where detection scores drop, and false positives surpass that threshold. Additionally, many existing methods assume that detections provide precise localization, overlooking the inherent noise that affects localization accuracy and causes state drift for occluded objects, as demonstrated in this work. To this end, a novel track validity mechanism, combined with a multi-stage observational gating process, is proposed that significantly reduces ghost tracks and improves tracking performance. Our method achieves 29.47% enhancement in Multi-Object tracking accuracy (MOTA) on the KITTI validation dataset with the Second detector. Furthermore, a refined Kalman filter term mitigates localization noise, ensuring robust state estimation for objects that are occluded and superior recovery during prolonged occlusions. This results in higher-order tracking accuracy (HOTA) improving by 4.8% on the KITTI validation dataset with the PV-RCNN detector. The proposed online framework, RobMOT, outperforms state-of-the-art methods, including deep learning approaches, across multiple detectors, with HOTA improvements of up to 3.92% on the KITTI testing dataset and 8.7% on the KITTI validation dataset while achieving the lowest identity switch (IDSW) scores of 7 and 0, respectively. RobMOT excels under challenging scenarios, such as tracking distant objects and handling prolonged occlusions, surpassing state-of-the-art methods on the Waymo Open testing dataset with a 1.77% improvement in MOTA for objects at distances exceeding 50 meters. RobMOT achieves a groundbreaking runtime of 3221 FPS using a single CPU, establishing itself as a highly efficient and scalable solution for real-time multi-object tracking. Mohamed Nagy, Naoufel Werghi, Bilal Hassan, Jorge Dias 0001, Majid Khonji |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Graph-Based Local Planning with Spatiotemporal Risk Assessment for Risk-Bounded and Prediction-Aware Autonomous DrivingabstractRisk-bounded motion planning for autonomous driving in dynamic environments presents significant research challenges. Ensuring continuous navigation towards a destination while making real-time decisions is a nonconvex problem. This paper presents a graph-based local planning method constrained by user-specific driving preference, represented as a risk-bound criterion for motion planning. First, we propose a lattice graph construction method that adheres to the vehicle's curvature constraints. Then, we formulate the trajectory planning problem as an integer-linear programming task, addressed by our novel risk-bounded and prediction-aware constrained shortest path. Our solution accounts for both static and dynamic obstacles in urban settings, adhering to traffic regulations. At the core of our approach is a conservative spatiotemporal risk assessment mechanism, which evaluates collisions considering the uncertain delay from speed control of the ego vehicle and predicted trajectories of dynamic obstacles. We implemented our solution using the CARLA simulator and the ROS2 platform, within a comprehensive framework encompassing global planning, local planning, and vehicle control. The effectiveness of our approach is demonstrated through notable collision avoidance, improved path-tracking, and enhanced risk-bounded planning capabilities. Abdulrahman Ahmad, Majid Khonji, Ameena Saad Al-Sumaiti, Jorge Dias 0001, Khaled M. Elbassioni |
ICARCV | 4 |
| 2024 | Embodied Neuromorphic Artificial Intelligence for Robotics: Perspectives, Challenges, and Research Development StackabstractRobotic technologies have been an indispensable part for improving human productivity since they have been helping humans in completing diverse, complex, and intensive tasks in a fast yet accurate and efficient way. Therefore, robotic technologies have been deployed in a wide range of applications, ranging from personal to industrial use-cases. However, current robotic technologies and their computing paradigm still lack embodied intelligence to efficiently interact with operational environments, respond with correct/expected actions, and adapt to changes in the environments. Toward this, recent advances in neuromorphic computing with Spiking Neural Networks (SNN) have demonstrated the potential to enable the embodied intelligence for robotics through bio-plausible computing paradigm that mimics how the biological brain works, known as “neuromorphic artificial intelligence (AI)”. However, the field of neuromorphic AI-based robotics is still at an early stage, therefore its development and deployment for solving real-world problems expose new challenges in different design aspects, such as accuracy, adaptability, efficiency, reliability, and security. To address these challenges, this paper will discuss how we can enable embodied neuromorphic AI for robotic systems through our perspectives: (P1) Embodied intelligence based on effective learning rule, training mechanism, and adaptability; (P2) Cross-layer optimizations for energy-efficient neuromorphic computing; (P3) Representative and fair benchmarks; (P4) Low-cost reliability and safety enhancements; (P5) Security and privacy for neuromorphic computing; and (P6) A synergistic development for energy-efficient and robust neuromorphic-based robotics. Furthermore, this paper identifies research challenges and opportunities, as well as elaborates our vision for future research development toward embodied neuromorphic AI for robotics. Rachmad Vidya Wicaksana Putra, Alberto Marchisio, Fakhreddine Zayer, Jorge Dias 0001, Muhammad Shafique 0001 |
ICARCV | 4 |
| 2024 | BVE + EKF: A Viewpoint Estimator for the Estimation of the Object's Position in the 3D Task Space Using Extended Kalman Filters
Sandro Magalhães 0001, António Paulo Moreira, Filipe Neves dos Santos, Jorge Dias 0001 |
ICINCO (2) | 4 |
| 2024 | SMO-CLIP: Enhancing Anomalous Smoke Density Assessment Using A Hybrid LLM-VLM ApproachabstractFlare stacks are among the crucial components in the safety and emission control of petrochemical plants. However, due to the imperceptibility of smoke and contaminants, analyzing these released particles during flare stack operation is one of the top challenges. To stress the problem, our work presents a novel solution called SMO-CLIP that can hybridize knowledge from Vision-Language Models (VLMs), specifically the Contrastive Language Image Pretraining (CLIP) model, with extra insights derived from GPT-4 Large Language Model (LLM). Furthermore, two new tasks, Finegrained Smoke Density Recognition (FSDR) and Coarsegrained Smoke Density Recognition (CSDR) are investigated in this paper to accurately detect and evaluate varying smoke intensities. Notable advancements over current approaches are observed through extensive experiments, demonstrating the superior performance of the proposed approach against state-of-the-art models. Muaz Al Radi, Mahmoud Said Elmezain, Abdelfatah Hassan Ahmed, Abderrahmene Boudiaf, Said Boumaraf, Jorge Dias 0001, Hamad Karki, Sajid Javed, Khalid Yousef Al Awadhi, Naoufel Werghi |
ICIP | 7 |
| 2024 | TerrainSense: Vision-Driven Mapless Navigation for Unstructured Off-Road EnvironmentsabstractNavigating autonomous vehicles efficiently across unstructured and off-road terrains remains a formidable challenge, often requiring intricate mapping or multi-step pipelines. However, these conventional approaches struggle to adapt to dynamic environments. This paper presents TerrainSense, an end-to-end framework that overcomes these limitations. By utilizing a transformers, TerrainSense detects lane semantics and topology from camera images, enabling mapless path planning without the reliance on highly detailed maps. The efficacy of TerrainSense was rigorously assessed on six diverse datasets, evaluating its efficacy in detection, segmentation, and path prediction using various metrics. Notably, it outperforms the other state-of-the-art methods by 9.32% in precisely predicting the path with 18.28% faster inference time. Bilal Hassan, Arjun Sharma, Nadya Abdel Madjid, Majid Khonji, Jorge Dias 0001 |
ICRA | 5 |
| 2024 | Efficient Hybrid Neuromorphic-Bayesian Model for Olfaction Sensing: Detection and ClassificationabstractOlfaction sensing in autonomous robotics faces challenges in dynamic operations, energy efficiency, and edge processing. It necessitates a machine learning algorithm capable of managing real-world odor interference, ensuring resource efficiency for mobile robotics, and accurately estimating gas features for critical tasks such as odor mapping, localization, and alarm generation. This paper introduces a hybrid approach that exploits neuromorphic computing in combination with probabilistic inference to address these demanding requirements. Our approach implements a combination of a convolutional spiking neural network for feature extraction and a Bayesian spiking neural network for odor detection and identification. The developed algorithm is rigorously tested on a dataset for sensor drift compensation for robustness evaluation. Additionally, for efficiency evaluation, we compare the energy consumption of our model with a non-spiking machine learning algorithm under identical dataset and operating conditions. Our approach demonstrates superior efficiency alongside comparable accuracy outcomes. Rizwana Kausar, Fakhreddine Zayer, Jaime Viegas, Jorge Dias 0001 |
ICRA | 4 |
| 2024 | Deep Learning-based Delay Compensation Framework For Teleoperated Wheeled Rovers on Soft TerrainsabstractThe difficulties posed by terrain-induced slippage for wheeled rovers traversing soft terrains are critical to ensuring safe and precise mobility. While bilateral teleoperation systems offer a promising solution to this issue, the inherent network-induced delays hinder the fidelity of the closed-loop integration, potentially compromising teleoperator system controls, and resulting in poor command-tracking performance. This work introduces a new model-free predictor framework based on deep learning designed to improve prediction performance and effectively compensate for large network delays in teleoperated wheeled rovers. Our approach employs the Recurrent Neural Network (RNN) to achieve a significant improvement in modeling complexity and prediction accuracy. Particularly, our framework consists of two distinct predictors, each tailored to the forward and backward coupling variables of the teleoperated wheeled rover. Human-in-the-loop experiments were conducted to validate the effectiveness of the developed framework in compensating for the delays encountered by teleoperated wheeled rovers coupled with terrain-induced slippage. The results confirm the improved prediction accuracy of the framework. This improvement is evidenced by improved performance and transparency metrics, which lead to better command-tracking performance. A supplementary video is available at https://youtu.be/-06UGumQ0tA. Ahmad Abubakar, Yahya Zweiri, Mubarak Yakubu, Ruqayya Alhammadi, Mohammed Basheer Mohiuddin, Abdel Gafoor Haddad, Jorge Dias 0001, Lakmal D. Seneviratne |
IROS | 7 |
| 2024 | PathFormer: A Transformer-Based Framework for Vision-Centric Autonomous Navigation in Off-Road EnvironmentsabstractThe efficient navigation of autonomous vehicles across rugged and unstructured terrains remains a significant challenge. Most existing research in this area emphasizes the need for complex mappings or intricate multi-step methodologies. However, these traditional approaches often struggle to adapt to dynamic changes in environmental conditions. In this paper, we introduce PathFormer, an end-to-end framework designed specifically to address these challenges. PathFormer utilizes transformers to decode free-space semantics and configurations directly from camera images, enabling efficient path planning without the reliance on detailed, pre-existing maps. The performance of PathFormer was rigorously evaluated across diverse datasets, where it demonstrated superior capabilities, outperforming other state-of-the-art methods by 3.68% in precisely segmenting free-space regions and showing a 13.65% improvement in correctly predicting traversable paths. Bilal Hassan, Nadya Abdel Madjid, Fatima Kashwani, Mohamad Alansari, Majid Khonji, Jorge Dias 0001 |
IROS | 6 |
| 2024 | Evaluation of Predictive Display for Teleoperated Driving Using CARLA SimulatorabstractBefore the world-wide deployment of autonomous vehicles, it is essential to implement intermediate solutions with partial autonomy. One such solution is the use of vehicle teleoperation, the act of controlling a vehicle from a distance. In real time applications of teleoperation, it is often pertinent to use augmented reality components within the teleoperator view, which are referred to as a predictive display. In this work, we evaluate our predictive display method, which is a guiding path based on the free space in the environment. The path is generated based on our Dual Transformer Network (DTNet), which uses both object detection and lane semantic segmentation to define the free space in the environment. While the model has previously performed well on image data, it is necessary to observe its accuracy in the presence of time delay and packet loss, to assess its performance in a real-time setting. Thus, in this work, we use CARLA simulator to compare the detected free space on the teleoperator side to the true free space on the vehicle side across different values of time delay and packet loss. Under optimal network conditions, our model yielded a remarkable 87.9% DSC score and 81.3% IoU score. Defining our minimum performance threshold as 80% DSC and 70% IoU, we conclude that our model can effectively mitigate the challenges of time delay below 100ms and packet loss below 1%, both of which represent substantial tolerances. Fatima Kashwani, Bilal Hassan, Peng Yong Kong, Majid Khonji, Jorge Dias 0001 |
IROS | 5 |
| 2024 | Efficient and lightweight in-memory computing architecture for hardware security
Hala Ajmi, Fakhreddine Zayer, Amira Hadj Fredj, Belgacem Hamdi, Baker Mohammad, Naoufel Werghi, Jorge Dias 0001 |
J. Parallel Distributed Comput. | 7 |
| 2024 | Programmable broad learning system for baggage threat recognition
Muhammad Shafay, Abdelfatah Hassan Ahmed, Taimur Hassan, Jorge Dias 0001, Naoufel Werghi |
Multim. Tools Appl. | 4 |
| 2024 | Incremental convolutional transformer for baggage threat detection
Taimur Hassan, Bilal Hassan, Muhammad Owais, Divya Velayudhan, Jorge Dias 0001, Mohammed Ghazal, Naoufel Werghi |
Pattern Recognit. | 5 |
| 2024 | Neural Graph Refinement for Robust Recognition of Nuclei Communities in Histopathological LandscapeabstractAccurate classification of nuclei communities is an important step towards timely treating the cancer spread. Graph theory provides an elegant way to represent and analyze nuclei communities within the histopathological landscape in order to perform tissue phenotyping and tumor profiling tasks. Many researchers have worked on recognizing nuclei regions within the histology images in order to grade cancerous progression. However, due to the high structural similarities between nuclei communities, defining a model that can accurately differentiate between nuclei pathological patterns still needs to be solved. To surmount this challenge, we present a novel approach, dubbed neural graph refinement, that enhances the capabilities of existing models to perform nuclei recognition tasks by employing graph representational learning and broadcasting processes. Based on the physical interaction of the nuclei, we first construct a fully connected graph in which nodes represent nuclei and adjacent nodes are connected to each other via an undirected edge. For each edge and node pair, appearance and geometric features are computed and are then utilized for generating the neural graph embeddings. These embeddings are used for diffusing contextual information to the neighboring nodes, all along a path traversing the whole graph to infer global information over an entire nuclei network and predict pathologically meaningful communities. Through rigorous evaluation of the proposed scheme across four public datasets, we showcase that learning such communities through neural graph refinement produces better results that outperform state-of-the-art methods. Taimur Hassan, Zhu Li 0001, Sajid Javed, Jorge Dias 0001, Naoufel Werghi |
IEEE Trans. Image Process. | 4 |
| 2024 | Center-Focused Affinity Loss for Class Imbalance Histology Image ClassificationabstractEarly-stage cancer diagnosis potentially improves the chances of survival for many cancer patients worldwide. Manual examination of Whole Slide Images (WSIs) is a time-consuming task for analyzing tumor-microenvironment. To overcome this limitation, the conjunction of deep learning with computational pathology has been proposed to assist pathologists in efficiently prognosing the cancerous spread. Nevertheless, the existing deep learning methods are ill-equipped to handle fine-grained histopathology datasets. This is because these models are constrained via conventional softmax loss function, which cannot expose them to learn distinct representational embeddings of the similarly textured WSIs containing an imbalanced data distribution. To address this problem, we propose a novel center-focused affinity loss (CFAL) function that exhibits 1) constructing uniformly distributed class prototypes in the feature space, 2) penalizing difficult samples, 3) minimizing intra-class variations, and 4) placing greater emphasis on learning minority class features. We evaluated the performance of the proposed CFAL loss function on two publicly available breast and colon cancer datasets having varying levels of imbalanced classes. The proposed CFAL function shows better discrimination abilities as compared to the popular loss functions such as ArcFace, CosFace, and Focal loss. Moreover, it outperforms several SOTA methods for histology image classification across both datasets. Taslim Mahbub, Ahmad Obeid 0001, Sajid Javed, Jorge Dias 0001, Taimur Hassan, Naoufel Werghi |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | DFR-FastMOT: Detection Failure Resistant Tracker for Fast Multi-Object Tracking Based on Sensor FusionabstractPersistent multi-object tracking (MOT) allows autonomous vehicles to navigate safely in highly dynamic environments. One of the well-known challenges in MOT is object occlusion when an object becomes unobservant for subsequent frames. The current MOT methods store objects information, such as trajectories, in internal memory to recover the objects after occlusions. However, they retain short-term memory to save computational time and avoid slowing down the MOT method. As a result, they lose track of objects in some occlusion scenarios, particularly long ones. In this paper, we propose DFR-FastMOT, a light MOT method that uses data from a camera and LiDAR sensors and relies on an algebraic formulation for object association and fusion. The formulation boosts the computational time and permits long-term memory that tackles more occlusion scenarios. Our method shows outstanding tracking performance over recent learning and non-learning benchmarks with about 3% and 4% margin in MOTA, respectively. Also, we conduct extensive experiments that simulate occlusion phenomena by employing detectors with various distortion levels. The proposed solution enables superior performance under various distortion levels in detection over current state-of-art methods. Our framework processes about 7,763 frames in 1.48 seconds, which is seven times faster than recent benchmarks. The framework will be available at https://github.com/MohamedNagyMostafa/DFR-FastMOT. Mohamed Nagy, Majid Khonji, Jorge Dias 0001, Sajid Javed |
ICRA | 3 |
| 2023 | Multi-view Inspection of Flare Stacks Operation Using a Vision-controlled Autonomous UAVabstractFlare stacks are crucial safety control components in petrochemical plants that required efficient monitoring and inspection. In this work, an Unmanned Aerial Vehicle (UAV)-based multi-view operation inspection system for monitoring and assessing the operation of flare stacks is proposed. Image-Based Visual Servoing (IBVS) control is used to guide the autonomous UAV for multi-view visual data collection. Afterwards, the collected visual data is analyzed using a new Multi-View Convolutional Neural Network (MV-CNN) deep learning model to obtain useful conclusions on the system's operation and classify the current state of the observed system. The proposed system's performance was validated in a simulated petrochemical plant environment with operational flare stacks and the results showed superior performance of the proposed MV-CNN model compared to a conventional single-view CNN model. Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001 |
IECON | 6 |
| 2023 | Adapting Behavior and Persistence via Reinforcement and Self-Emotion Mediated Exploration in a Social RobotabstractAdaptability and behavioral diversity are core components of social interactions between humans. Naturally, these are traits research should strive to achieve in social robotics so agents may be better accepted and engage with their user peers. In this paper, we propose a novel activity modulation to increase behavioral diversity, based on a surprise-exploration correlation model, in a social robot undergoing behavioral optimization to user state and preference. This framework was tested with 21 participants to assess preferences as well as the impact that action variability and persistence would have on user perception of the robot. Results indicate a positive effect of persistence and variability over robot likability as well as user engagement, contributing insight for future research in social robotics. Gustavo Assunção, Alessandra Sorrentino, Jorge Dias 0001, Miguel Castelo-Branco, Paulo Menezes 0001, Filippo Cavallo |
RO-MAN | 3 |
| 2023 | Benchmarking edge computing devices for grape bunches and trunks detection using accelerated object detection single shot multibox deep learning modelsabstractVisual perception enables robots to perceive the environment. Visual data is processed using computer vision algorithms that are usually time-expensive and require powerful devices to process the visual data in real-time, which is unfeasible for open-field robots with limited energy. This work benchmarks the performance of different heterogeneous platforms for object detection in real-time. This research benchmarks three architectures: embedded GPU—Graphical Processing Units (such as NVIDIA Jetson Nano 2 GB and 4 GB, and NVIDIA Jetson TX2), TPU—Tensor Processing Unit (such as Coral Dev Board TPU), and DPU—Deep Learning Processor Unit (such as in AMD/Xilinx ZCU104 Development Board, and AMD/Xilinx Kria KV260 Starter Kit). The authors used the RetinaNet ResNet-50 fine-tuned using the natural VineSet dataset. After the trained model was converted and compiled for target-specific hardware formats to improve the execution efficiency. The platforms were assessed in terms of performance of the evaluation metrics and efficiency (time of inference). Graphical Processing Units (GPUs) were the slowest devices, running at 3 FPS to 5 FPS, and Field Programmable Gate Arrays (FPGAs) were the fastest devices, running at 14 FPS to 25 FPS. The efficiency of the Tensor Processing Unit (TPU) is irrelevant and similar to NVIDIA Jetson TX2. TPU and GPU are the most power-efficient, consuming about 5 W. The performance differences, in the evaluation metrics, across devices are irrelevant and have an F1 of about 70 % and mean Average Precision (mAP) of about 60 %. Sandro Magalhães 0001, Filipe Neves dos Santos, António Paulo Moreira, Jorge Dias 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Cascaded structure tensor for robust baggage threat detection
Taimur Hassan, Samet Akcay, Bilal Hassan, Mohammed Bennamoun, Salman Khan 0001, Jorge Dias 0001, Naoufel Werghi |
Neural Comput. Appl. | 6 |
| 2023 | Robot-Person Tracking in Uniform Appearance Scenarios: A New Dataset and ChallengesabstractPerson-tracking robots have many applications including security, surveillance, and autonomous driving. Despite the abundance of uniform appearance in many contexts and the challenges they exhibit, there is a lack of video datasets dedicated to benchmarking tracking algorithms in such contexts. In this article, we propose a new high-quality RGB-D benchmark called PTUA for robot–person tracking in uniform appearance scenarios. PTUA is recorded using an RGB-D sensor on top of a moving robot and consists of 45 sequences containing more than 85 K frames. Each frame is manually annotated with a bounding box and attributes, making PTUA the largest and the most challenging person tracking RGB-D dataset. To the best of our knowledge, such a densely annotated and properly synchronized RGB-D tracking benchmark does not exist in the literature. Each sequence comprises various challenges deriving from real-life scenarios where the target person appears highly similar to the background or distractors. By releasing PTUA, we expect to provide the community with a large-scale challenging RGB-D benchmark with high quality for the robust evaluation of trackers on uniform appearance scenarios for autonomous robots. We also present a rigorous experimental evaluation of the state-of-the-art trackers on the PTUA dataset with a comprehensive analysis. The findings evidence the challenges of person tracking in a uniform appearance scenario for both target tracking and robot–person tracking, and the need to bridge the performance gap. In addition, we propose a new RGB-D tracker that extracts features from RGB-D frames and it achieves the best performance on each challenging scenario of PTUA. Xiaoxiong Zhang 0001, Adarsh Ghimire, Sajid Javed, Jorge Dias 0001, Naoufel Werghi |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | UTB180: A High-Quality Benchmark for Underwater Tracking
Basit Alawode, Mehnaz Ummar, Naoufel Werghi, Jorge Dias 0001, Ajmal Mian, Sajid Javed |
ACCV (5) | 5 |
| 2022 | An Evaluation of Direct Image Based Visual Tracking System for Autonomous ManipulationabstractIn this paper, we implement and evaluate a direct image-based visual tracking system for autonomous manipulation applications. The direct image-based visual tracking method is developed to relax complex image processing tasks from traditional image and position-based visual tracking methods. The visual signals are constructed by using multiresolution coefficients. The method compares the multiresolution image transformation of the current and desired images and uses the mismatch between them to drive the system. The method needs to design a multiresolution interaction matrix with half and details images. The matrix connects the multiresolution image transformation coefficients with the velocity of the manipulator and controller. To illustrate the effectiveness, the direct image-based tracking system design is implemented and evaluated to locate the position of the valve stem and set of wrenches for autonomous grasping and manipulation for fire fighting applications. A. Saleh, Jorge Dias 0001, Anderson Sunda-Meya |
IECON | 3 |
| 2022 | Vision-based Inspection of Flare Stacks Operation Using a Visual Servoing Controlled Autonomous Unmanned Aerial Vehicle (UAV)abstractThe inspection of flare stacks’ operation is a challenging task that requires technical expertise and human effort. Flare stack systems undergo various types of faults that need to be monitored in a timely manner to avoid costly and dangerous accidents. Automating this process via the application of autonomous robotic systems for collecting comprehensive data of the flare stack’s operation is a promising solution for minimizing the involved hazards and costs. In this work, a novel Unmanned Aerial Vehicle (UAV)-based autonomous inspection system for flare stacks performance monitoring is proposed. The system employs a deep learning detection network that was trained for detection of flame and smoke for vision-based flaring performance analysis. A visual servoing control technique was used for guiding the UAV’s movement throughout the inspection mission for collecting comprehensive visual inspection data. Simulations in a simulated petrochemical plant environment with flare stacks were performed for validating the performance of the proposed system. The proposed UAV system was able to collect the required data successfully and analysis of the obtained data returned useful information about the flare stack’s operation. Muaz Al Radi, Hamad Karki, Naoufel Werghi, Sajid Javed, Jorge Dias 0001 |
IECON | 5 |
| 2022 | Hybrid Machine-Learning-Based Spectrum Sensing and Allocation With Adaptive Congestion-Aware Modeling in CR-Assisted IoV NetworksabstractUnlicensed cognitive-radio (CR)-assisted Internet of Vehicles (IoV) users can access licensed providers’ radio spectrum and concurrently utilize the dedicated channel for data transmission in vehicular communication. Optimizing channel access in cognitive IoV networks can help maximize available spectrum resources. This article proposes a novel sensing and communication integrated framework, dubbed as the CR-assisted IoV network (CRAV-Net), using a cluster-based hybrid optimization approach with adaptive congestion-aware modeling for dynamic high-mobility vehicular networks in an urban city context. In CRAV-Net, intelligent hybrid learning spectrum agents are introduced, which perform spectrum sensing (SS) using a deep learning (DL) model. It dynamically learns the multilevel spatial and temporal graphical features from input spectrograms through layer-by-layer propagation. It efficiently predicts the spectrum occupancy in the primary spectrum, without a priori knowledge of the radio environment. Then, to assign the vacant channels to the secondary vehicles, a support vector machine classifier is trained based on several learning features, including the vehicle stay time, vehicle density, and network capacity, to select the optimal resource route. The proposed framework achieves an overall accuracy of 99.74% in SS using the custom data set, outperforming state of the art by 12.60% at −25-dB signal-to-noise ratio. In addition, it brings a performance gain of 0.81% in SS accuracy when evaluated on real-world signals. Furthermore, in optimal network node allocation, the proposed framework achieves a mean accuracy of 98.45%, outperforming the existing methods by 0.63% and 18.32% in terms of accuracy and allocation time, respectively. Ramsha Ahmed, Yueyun Chen, Bilal Hassan, Liping Du, Taimur Hassan, Jorge Dias 0001 |
IEEE Internet Things J. | 6 |
| 2022 | Hierarchical Spatiotemporal Graph Regularized Discriminative Correlation Filter for Visual Object TrackingabstractVisual object tracking is a fundamental and challenging task in many high-level vision and robotics applications. It is typically formulated by estimating the target appearance model between consecutive frames. Discriminative correlation filters (DCFs) and their variants have achieved promising speed and accuracy for visual tracking in many challenging scenarios. However, because of the unwanted boundary effects and lack of geometric constraints, these methods suffer from performance degradation. In the current work, we propose hierarchical spatiotemporal graph-regularized correlation filters for robust object tracking. The target sample is decomposed into a large number of deep channels, which are then used to construct a spatial graph such that each graph node corresponds to a particular target location across all channels. Such a graph effectively captures the spatial structure of the target object. In order to capture the temporal structure of the target object, the information in the deep channels obtained from a temporal window is compressed using the principal component analysis, and then, a temporal graph is constructed such that each graph node corresponds to a particular target location in the temporal dimension. Both spatial and temporal graphs span different subspaces such that the target and the background become linearly separable. The learned correlation filter is constrained to act as an eigenvector of the Laplacian of these spatiotemporal graphs. We propose a novel objective function that incorporates these spatiotemporal constraints into the DCFs framework. We solve the objective function using alternating direction methods of multipliers such that each subproblem has a closed-form solution. We evaluate our proposed algorithm on six challenging benchmark datasets and compare it with 33 existing state-of-the art trackers. Our results demonstrate an excellent performance of the proposed algorithm compared to the existing trackers. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Lakmal D. Seneviratne, Naoufel Werghi |
IEEE Trans. Cybern. | 3 |
| 2021 | Distributed Adaptive Protocol for Asymptotic Consensus for a Networked Euler-Lagrange Systems with UncertaintyabstractThis paper investigates distributed asymptotic consensus protocol for a group of cloud connected Euler-Lagrange nonlinear systems with the presence of bounded uncertainty. The consensus protocol is designed by combining linear sliding surface vectors with robust adaptive learning algorithms. The sliding surface is designed by comprising position and velocity signals of the leader and neighboring follower Lagrange systems. Adaptive learning algorithm uses to learn and compensate bounded uncertainty associated with parameters and other external disturbance uncertainty. Lyapunov and sliding mode control theory uses to design and illustrate the convergence of the closed loop system under proposed protocol. The convergence analysis has three parts. In first part, it proves that the position and velocity consensus error states are bounded provided that the parameter estimates are continuous and bounded by positive constant. The second part guarantees that the sliding mode motion occurs for each Lagrange system in finite-time. The third part ensures that the states for a group of follower Lagrange systems can achieve asymptotic consensus tracking provided that the interaction communication topology has a directed spanning tree. This analysis shows asymptotic consensus property of the position and velocity consensus error states on the sliding mode surface. The design and implementation of the proposed asymptotic consensus protocol is easier as it does not use the exact bound of the uncertainty. Jorge Dias 0001, Gurdial Arora, Anderson Sunda-Meya |
IECON | 2 |
| 2021 | Distributed Cooperative LFC Protocols for Regulation Synchronization for Networked Multi-area Power Grid NetworksabstractIn this paper, we propose consensus based distributed cooperative LFC schemes for leader-less networked multi-area power grid network systems in the presence of uncertainty. The LFC schemes design combine local states with the states of the neighboring area with directed communication topology. We propose two distributed cooperative LFC schemes. First, robust LFC schemes are designed by assuming that the bounds of the uncertainty associated with the power network dynamics are available. Then, we remove the demand of the bound on the uncertainty from LFC schemes by designing robust adaptive learning algorithm. Robust adaptive control terms uses to deal with the presence of uncertainty associated with the power networks and external fault disturbance. Lyapunov and graph theory uses to show that the proposed distributed cooperative design can reach an agreement with control areas and solve regulation synchronization problem. Analysis shows that the state of the control areas can reach an agreement and ensure both finite-time and asymptotic consensus property. Evaluation results on a four-area interconnected power grid networks are presented to show the effectiveness of the proposed consensus based distributed LFC algorithm for real-time applications. Jorge Dias 0001, Anderson Sunda-Meya |
IECON | 2 |
| 2021 | Distributed Tracking Synchronization Protocol for a Networked of Leader-follower Unmanned Aerial Vehicles with UncertaintyabstractThis work investigates robust asymptotic consensus tracking problems for a group of cloud-connected leader-follower unmanned aerial vehicles with uncertainty. The protocols for attitude and position subsystems dynamics are constructed by using the states of the local and neighboring vehicles provided that they are connected by local area networks. Robust adaptive learning algorithms are also integrated with both protocols to learn and adapt to the modeling errors and external disturbance uncertainties. Lyapunov method and Graph theory use to prove that the proposed protocol allows the vehicles to reach an agreement with follower vehicles and track the states of the leader vehicle asymptotically. Convergence analysis shows that consensus protocol can force the states of the follower MAVs to track the state of the leader MAV asymptotically. The protocol designs are simple and easy to implement as they do not need the exact bound of the uncertainty that appears from external disturbances and the modeling errors. The design does not require the bound of the input of the leader vehicle. The protocol design can ensure faster and robust consensus in the presence of uncertainty as opposed to the convergence of other asymptotic consensus designs. Jorge Dias 0001, Anderson Sunda-Meya |
IECON | 2 |
| 2021 | On the Design and Development of Vision-Based Autonomous Mobile ManipulationabstractThis paper investigates image-feature based visual tracking systems for autonomous mobile manipulation applications. First, we briefly present various visual tracking methods and their components for autonomous tracking applications. Second, we introduce the development process for the most popular image-feature based autonomous visual tracking system for grasping and mobile manipulation applications. Then, the application scenario for evaluation is provided with the detailed software and hardware components for the fire-fighting application. Finally, the evaluation results on a 6-DOF UR5 mobile robot manipulator arms are presented for autonomous grasping and manipulation for valve turning applications. Jorge Dias 0001, Anderson Sunda-Meya |
IECON | 2 |
| 2021 | Spatially Constrained Context-Aware Hierarchical Deep Correlation Filters for Nucleus Detection in Histology Images
Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi, Nasir M. Rajpoot |
Medical Image Anal. | 3 |
| 2020 | Deep Bidirectional Correlation Filters for Visual Object TrackingabstractVisual Object Tracking (VOT) is an essential task for many computer vision applications. VOT becomes challenging when a target object faces severe occlusion, drastic illumination changes, and scale variation problems. In the literature, Discriminative Correlation Filters (DCFs)-based tracking methods have achieved promising results in terms of accuracy and efficiency in many complex VOT scenarios. A plethora of DCFs trackers have been proposed which exploit information observed in past frames to create and update DCFs for VOT. To adapt to target appearance variations, the DCFs are enhanced by incorporating spatial and temporal consistency constraints. Nevertheless, the performance degradation is observed for these methods because of the aforementioned limitations. To address these issues, we propose a novel algorithm based on bidirectional DCFs for VOT. In this algorithm, we propose the original idea of leveraging information from both past and future frames. The proposed algorithm first tracks the target object forward in the video sequence and then its uses the predicted location of the last window frame and track the target object backward towards the current frame. We design an appearance consistency loss function by taking the$L_{2}$norm between the regression target of the forward tracking and response map of the backward tracking to obtain the resulting response map. Our proposed algorithm realizes a highly accurate DCFs because forward and backward tracking information are fused together for consistent VOT. Although, a result will be output with some small delay because information is taken from a future to the present period, our proposed algorithm has the merit of addressing the drastic appearance variations VOT challenges. We evaluate our proposed tracker using deep features on three publicly available challenging datasets. Our results demonstrate the superior performance of the proposed tracker compared to the existing state-of-the-art trackers. Sajid Javed, Xiaoxiong Zhang 0001, Lakmal D. Seneviratne, Jorge Dias 0001, Naoufel Werghi |
FUSION | 4 |
| 2020 | CS-RPCA: Clustered Sparse RPCA for Moving Object DetectionabstractMoving object detection (MOD) is an important step for many computer vision applications. In the last decade, it is evident that RPCA has shown to be a potential solution for MOD and achieved a promising performance under various challenging background scenes. However, because of the lack of different types of features, RPCA still shows degraded performance in many complicated background scenes such as dynamic backgrounds, cluttered foreground objects, and camouflage. To address these problems, this paper presents a Clustered Sparse RPCA (CS-RPCA) for MOD under challenging environments. The proposed algorithm extracts multiple features from video sequences and then employs RPCA to get the low-rank and sparse component from each representation. The sparse subspaces are then emerged into a common sparse component using Grassmann manifold. We proposed a novel objective function which computes the composite sparse component from multiple representations and it is solved using non-negative matrix factorization method. The proposed algorithm is evaluated on two challenging datasets for MOD. Results demonstrate excellent performance of the proposed algorithm as compared to existing state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
ICIP | 3 |
| 2020 | Gender Recognition on RGB-D ImageabstractIn this paper, we propose a deep-learning approach for human gender classification on RGB-D images. Unlike most of the existing methods, which use hand-crafted features from the human face, we exploit local information from the head and global information from the whole body to classify people's gender. A head detector is fine-tuned on YOLO to detect the head regions on the images automatically. Two gender classifiers are trained using head images and whole-body images separately. The final prediction is made by fusing the two classifiers' results. The presented method outperforms the state-of-art with an improvement in the accuracy of 2.6%, 7.6%, and 8.4% on three different test data of a challenging gender dataset which includes human standing, walking, and interacting scenarios. Xiaoxiong Zhang 0001, Sajid Javed, Ahmad Obeid 0001, Jorge Dias 0001, Naoufel Werghi |
ICIP | 4 |
| 2020 | Robust Structural Low-Rank TrackingabstractVisual object tracking is an essential task for many computer vision applications. It becomes very challenging when the target appearance changes especially in the presence of occlusion, background clutter, and sudden illumination variations. Methods, that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when facing the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking in complex scenarios. In the proposed algorithm, we consider spatial and temporal appearance consistency constraints, among the particles in the low-rank subspace, embedded in four different graphs. The resulting objective function encoding these constraints is novel and it is solved using linearized alternating direction method with adaptive penalty both in batch fashion as well as in online fashion. Our proposed objective function jointly learns the spatial and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on four challenging datasets demonstrate excellent performance of the proposed algorithm as compared to current state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
IEEE Trans. Image Process. | 3 |
| 2019 | Structural Low-Rank TrackingabstractVisual object tracking is an important step for many computer vision applications. The task becomes very challenging when the target undergoes heavy occlusion, background clutters, and sudden illumination variations. Methods that incorporate sparse representation and low-rank assumptions on the target particles have achieved promising results. However, because of the lack of structural constraints, these methods show performance degradation when an object faces the aforementioned challenges. To alleviate these limitations, we propose a new structural low-rank modeling algorithm for robust object tracking. In the proposed algorithm, we enforce local spatial, global spatial and temporal appearance consistency among the particles in the low-rank subspace by constructing three graphs. The Laplacian matrices of these graphs are incorporated into the novel low-rank objective function which is solved using linearized alternating direction method with an adaptive penalty. Our proposed objective function jointly learns the spatial, global, and temporal structure of the target particles in consecutive frames and makes the proposed tracker consistent against many complex tracking scenarios. Results on two challenging benchmark datasets show the superiority of the proposed algorithm as compared to current state-of-the-art methods. Sajid Javed, Arif Mahmood, Jorge Dias 0001, Naoufel Werghi |
AVSS | 3 |
| 2019 | Cooperative and Social Robots: Understanding Human Activities and Intentions
Rebeca Marfil, Jorge Dias 0001, Antonio Bandera, George Azzopardi |
Pattern Recognit. Lett. | 2 |
| 2019 | αPOMDP: POMDP-based user-adaptive decision-making for social robots
Gonçalo S. Martins, Hend Al Tair, Luís Santos 0001, Jorge Dias 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Toward a Context-Aware Human-Robot Interaction Framework Based on Cognitive DevelopmentabstractThe purpose of this paper was to understand how an agent's performance is affected when interaction workflows are incorporated in its information model and decision-making process. Our expectation was that this incorporation could reduce errors and faults on agent's operation, improving its interaction performance. We based this expectation on the existing challenges in designing and implementing artificial social agents, where an approach based on predefined user scenarios and action scripts is insufficient to account for uncertainty in perception or unclear expectations from the user. Therefore, we developed a framework that captures the expected behavior of the agent into descriptive scenarios and then translated these into the agent's information model and used the resulting representation in probabilistic planning and decision making to control interaction. Our results indicated an improvement in terms of specificity while maintaining precision and recall, suggesting that the hypothesis being proposed in our approach is plausible. We believe the presented framework will contribute to the field of cognitive robotics, e.g., by improving the usability of artificial social companions, thus overcoming the limitations imposed by approaches that use predefined static models for an agent's behavior resulting in non-natural interaction. João Quintas, Gonçalo S. Martins, Luís Santos 0001, Paulo Menezes 0001, Jorge Dias 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2018 | Coverage Path Planning with Adaptive Viewpoint Sampling to Construct 3D Models of Complex Structures for the Purpose of InspectionabstractIn this paper, we introduce a coverage path planning algorithm with adaptive viewpoint sampling to construct accurate 3D models of complex large structures using Unmanned Aerial Vehicle (UAV). The developed algorithm, Adaptive Search Space Coverage Path Planner (ASSCPP), utilizes an existing 3D reference model of the complex structure and the onboard sensors' noise models to generate paths that are evaluated based on the traveling distance and the quality of the model. The algorithm generates a set of viewpoints by performing adaptive sampling that directs the search towards areas with low accuracy and low coverage. The algorithm predicts the coverage percentage obtained by following the generated coverage path using the reference model. A set of experiments were conducted in real and simulated environments with structures of different complexities to test the validity of the proposed algorithm. Randa Almadhoun, Tarek Taha, Dongming Gan, Jorge Dias 0001, Yahya Zweiri, Lakmal D. Seneviratne |
IROS | 4 |
| 2018 | An Extended Bayesian User Model (BUM) for Capturing Cultural Attributes with a Social RobotabstractIn this work we propose a Bayesian User Model which is able capture a unified representation of cultural attributes from heterogeneous information in the context of Human-Robot Interaction. Despite the latest advances in robotic technologies, virtually no robots are able to cope with the specificities of the “modus vivendi” of different cultures. We start by proposing Bayesian classifiers to capture unitary attributes of different users, clustering them in a n-dimensional semantic attribute space, aggregating groups of persons that share similar attributes. Results show a highly accurate classification framework, both capable of detecting specific subtleties in user's properties, and generalizing them into representative profiles. We then discuss its application towards adapting the actions of a robot and its potential impact on culture-awareness, demonstrating how the proposed framework can enable culture-awareness, exploring this new frontier in social robotics. Luís Santos 0001, Gonçalo S. Martins, Jorge Dias 0001 |
IROS | 3 |
| 2018 | Discrete Cosserat Approach for Multisection Soft Manipulator DynamicsabstractNowadays, the most adopted model for the design and control of soft robots is the piecewise constant curvature model, with its consolidated benefits and drawbacks. In this work, an alternative model for multisection soft manipulator dynamics is presented based on a discrete Cosserat approach, in which the continuous Cosserat model is discretized by assuming a piecewise constant strain along the soft arm. As a consequence, the soft manipulator state is described by a finite set of constant strains. This approach has several advantages with respect to the existing models. First, it takes into account shear and torsional deformations, which are both essential to cope with out-of-plane external loads. Furthermore, it inherits desirable geometrical and mechanical properties of the continuous Cosserat model, such as intrinsic parameterization and greater generality. Finally, this approach allows to extend to soft manipulators, the recursive composite-rigid-body and articulated-body algorithms, whose performances are compared through a cantilever beam simulation. The soundness of the model is demonstrated through extensive simulation and experimental results. Federico Renda, Frédéric Boyer, Jorge Dias 0001, Lakmal D. Seneviratne |
IEEE Trans. Robotics | 3 |
| 2017 | A Simulation Environment for Active Endoscopic CapsulesabstractThe best way for researchers to test their algorithms and design concepts, before experimenting on real human beings, is to create simulation environment platforms. In this paper, a virtual simulator for active endoscopic capsules is proposed. The simulator intends to provide researchers with an environment to test their vision and navigation algorithms applied to endoscopic capsule applications. The proposed simulation was created using Gazebo simulator, , a robust physics engine under Robotic Operating System (ROS) environment. It consists of three main software modules: (i) capsule model, (ii) capsule control, and (iii) Gazebo customized plugins. The current version of the simulator can provide three main functions: lumen tracking, capsule tele-operation and haptic feedback for capsule navigation. Yasmeen Abu-Kheil, Lakmal D. Seneviratne, Jorge Dias 0001 |
CBMS | 3 |
| 2017 | Convolutional neural networkasa feature extractor for automatic polyp detectionabstractColorectal cancer is one of the major causes of cancer deaths worldwide. To achieve early cancer screening, detecting the presence of polyps in the colon tract is the preferred technique. In this paper, a deep learning approach for identifying polyps in colonoscopy images is proposed. The novelty of our technique stems from the fact that it fully employs a pre-trained Convolutional Neural Network (CNN) architecture as a feature extractor. Contrary to the conventional methods which either perform fine-tuning or train the CNN from scratch, we utilize the CNN output features as an input to train the Support Vector Machine (SVM) Classifier. The efficiency of the presented framework is demonstrated on the public CVC ColonDB, in which the experimental results indicate that our methodology significantly outperforms other competitive paradigms. Bilal Taha, Jorge Dias 0001, Naoufel Werghi |
ICIP | 2 |
| 2017 | BUM: Bayesian user model for distributed social robotsabstractIn this work we present a Bayesian User Model for inferring the characteristics and inter-user patterns of a population users. The model can receive evidence gathered by various interactive devices, such as social robots or wearable devices. The system is modular, with each module being responsible for gathering information and observations from persons present in the system's operation scenario. This information enables each module to determine a single characteristic of the person. New observations and measurements received by the system are fused with previous knowledge by a sub-process based on an information theory technique. This allows the system to be implemented in diverse heterogeneous distributed system topologies, extending beyond robotics. We have conducted experiments involving a team of social robots and simulated user population. Our experiments have shown that the system is able to learn and classify the persons' characteristics, and to find relevant user groups via clustering. This system can potentially be used to gather information on a large set of persons, as well as to be an information source for user-adaptive applications in areas such as Robotics, Ambient Assisted Living (AAL) and Internet of Things. Gonçalo S. Martins, Luís Santos 0001, Jorge Dias 0001 |
RO-MAN | 3 |
| 2017 | Speaking robots: The challenges of acceptance by the ageing societyabstractThe ability of robots to dialogue with humans appears as one critical Human-Machine Interaction feature when it comes to transferring robots into society. This ability gains additional importance when it comes to elderly people, since they find it more comfortable and natural to interact using voice, due to possible natural physical impairments that hinder the usage of some of the interaction modalities (e.g. touch screens). Challenges like recognition accuracy, distant speech, the idiosyncrasies of elderly voices (fading, muffled pronunciation, etc.), the effects of surrounding environment noise or the expressiveness of the robot when speaking, become highly relevant in the acceptance and usability of service robots by the ageing population. In this paper, we present the results, challenges and solutions developed during a nine-month iterative evaluation process that took place within the GrowMeUp project, with focus on speech recognition and synthesis. The paper concludes with an identification of open scientific and technological problems, based on our interpretation of results, which we identify as critical for the acceptance and usability of robots by an ageing society. Gonçalo S. Martins, Ana Luísa Jegundo, Carina Dantas, Cindy Wings, Luís Santos 0001, Jorge Dias 0001, Fernando Perdigão |
RO-MAN | 7 |
| 2017 | Interoperability in cloud robotics - Developing and matching knowledge information models for heterogenous multi-robot systemsabstractEvery file, document, database and digital information is now going through the Cloud. Leveraged by the developments in information systems, Cloud Robotics is evolving at a steady pace and raised attention in the past 5 years. This recent field of Robotics is allowing engineers to envisage new and exciting applications for robots in the near future. This work proposes Cloud Robotics as a mean to integrate semantic reasoning in a multi-robot system, using self-created knowledge bases in each robot, in order to perform the coordination of complex task allocation. An auction-based coordination method and a knowledge matching algorithm were implemented to study this subject. The obtained results demonstrated that, the coordination of a large multi-robot system and the knowledge matching process can be computationally demanding, thus making them perfect candidate features to be “cloudyfied”. João Quintas, Paulo Menezes 0001, Jorge Dias 0001 |
RO-MAN | 3 |
| 2017 | A Bayesian hierarchy for robust gaze estimation in human-robot interaction
Pablo Lanillos, João Filipe Ferreira, Jorge Dias 0001 |
Int. J. Approx. Reason. | 3 |
| 2017 | Integration of touch attention mechanisms to improve the robotic haptic exploration of surfacesabstractThis text presents the integration of touch attention mechanisms to improve the efficiency of the action-perception loop, typically involved in active haptic exploration tasks of surfaces by robotic hands . The progressive inference of regions of the workspace that should be probed by the robotic system uses information related with haptic saliency extracted from the perceived haptic stimulus map ( exploitation ) and a “curiosity”-inducing prioritisation based on the reconstruction's inherent uncertainty and inhibition-of-return mechanisms ( exploration ), modulated by top-down influences stemming from current task objectives, updated at each exploration iteration. This work also extends the scope of the top-down modulation of information presented in a previous work, by integrating in the decision process the influence of shape cues of the current exploration path. The Bayesian framework proposed in this work was tested in a simulation environment. A scenario made of three different materials was explored autonomously by a robotic system . The experimental results show that the system was able to perform three different haptic discontinuity following tasks with a good structural accuracy, demonstrating the selectivity and generalization capability of the attention mechanisms. These experiments confirmed the fundamental contribution of the haptic saliency cues to the success and accuracy of the execution of the tasks. Ricardo Martins 0002, João Filipe Ferreira, Miguel Castelo-Branco, Jorge Dias 0001 |
Neurocomputing | 4 |
| 2017 | Information Model and Architecture Specification for Context Awareness Interaction Decision Support in Cyber-Physical Human-Machine SystemsabstractThis paper aims to contribute to situation, activity, and goal awareness in cyber-physical human-machine systems (HMS) by presenting a new information model and specifications for a decision-making component that can be integrated in current system architectures. The objective of this work is to improve the efficacy, acceptance, adaptability, and overall performance of HMS and human-system interaction (HSI) applications using a context-based approach. Our hypothesis is that we can enhance current interaction functionalities by integrating context and interaction information models into a decision-making component that behaves as a supervision process for controlling interaction. In HSI, we aim to define a general human model that may lead to principles and algorithms, allowing more natural and effective interaction between humans and artificial agents. The approach was implemented and tested targeting application in the domain of active and assisted living. The challenge of user acceptance is of vital importance for future solutions and is still one of the major reasons for reluctance to adopt cyber-physical systems in this domain. João Quintas, Paulo Menezes 0001, Jorge Dias 0001 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2016 | Bayesian inference implemented on FPGA with stochastic bitstreams for an autonomous robotabstractThis paper presents an FPGA implementation of a machine performing exact Bayesian inference using stochastic bitstreams. We revisited stochastic computing, not to perform better computations with unreliable hardware, but to perform approximate computations with less hardware. The underlying trade-off is between precision and computation time. An automatic design of probabilistic machines that compute soft inferences with an arithmetic based on stochastic bitstreams is presented. The computation tree provided by a Bayesian inference software is used to define the stochastic circuit. Tests were performed and results presented concerning accuracy and resource usage of the stochastic computing implementation of Bayesian machines performing exact inference. An application example is given of a Bayesian sensorimotor system that performs obstacle avoidance for an autonomous robot, fully implemented on an FPGA. Some conclusions were drawn on the followed approach, providing insights for future implementations. Hugo Fernandes, M. Awais Aslam, Jorge Lobo 0002, João Filipe Ferreira, Jorge Dias 0001 |
FPL | 5 |
| 2016 | Modeling, design & characterization of a novel Passive Variable Stiffness Joint (pVSJ)abstractIn this paper we present the design and characterization of a novel Passive Variable Stiffness Joint (pVSJ). pVSJ is the proof of concept of a passive revolute joint with controllable variable stiffness. The current design is intended to be a bench-test for future development towards applications in haptic teleoperation purposed exoskeletons. The main feature of the pVSJ is its capability of varying the stiffness with infinite range based on a simple mechanical system. Moreover, the joint can rotate freely at the zero stiffness case without any limitation. The stiffness varying mechanism consists of two torsional springs, mounted with an offset from the pVSJ rotation center and coupled with the joint shaft by an idle roller. The position of the roller between the pVSJ rotation center and the spring's center is controlled by a linear sliding actuator fitted on the chassis of the joint. The variation of the output stiffness is obtained by changing the distance from the roller-springs contact point to the joint rotation center (effective arm). If this effective arm is null, the stiffness of the joint will be zero. The stiffness increases to reach high stiffness values when the effective arm approaches its maximum value, bringing the roller close to the torsional springs' center. The experimental results matched with the physical-based modeling of the pVSJ in terms of stiffness variation curve, stiffness dependency upon the springs' elasticity, joint deflection and the spring's deflection. Mohammad I. Awad, Dongming Gan, Marco Cempini, Mario Cortese, Nicola Vitiello, Jorge Dias 0001, Paolo Dario, Lakmal D. Seneviratne |
IROS | 6 |
| 2016 | Discrete Cosserat approach for soft robot dynamics: A new piece-wise constant strain model with torsion and shearsabstractModeling and control of soft robots is an up-to-date and exciting area of research which has been tackled with complementary approaches so far. In this paper, we modify the existing continuum Cosserat approach optimizing it for soft robot arms which can be discretized in a finite number of sections and degrees of freedom. The resulting new piece-wise constant strain model extends the existing piece-wise constant curvature model by allowing torsion and shears strains which are fundamental to cope with out-of-the-plane external forces as appearing for example during ground locomotion. A first experimental comparison has been also conducted using one fluidic actuated leg of the soft crawler FASTT. Federico Renda, Vito Cacucciolo, Jorge Dias 0001, Lakmal D. Seneviratne |
IROS | 3 |
| 2016 | Torque reflecting coordination control for bilateral shared autonomous system over open communication networksabstractIn this paper, torque reflection based coordination control algorithm is designed for network-based bilateral shared autonomous system over open communication networks. The control algorithm for master and slave manipulator is designed by combining delayed position and velocity signal with the delayed reflected torques from the interaction between human and master and between slave and environment. Robust and adaptive control technique is used to deal with uncertainty associated with the gravity, unmodeled dynamic and other external input disturbance. The convergence of the closed loop system is shown by using Lyapunov method. In contrast with existing force reflection based design, the proposed controller can deal with uncertainty associated with the gravity, unmodeled dynamic and external input disturbance. Compared with other methods, the proposed design uses reflected torques from the interaction between master and human and between slave and environment so as to improve the transparency of the bilateral shared autonomous system. Finally, evaluation results are presented to demonstrate the validity of the proposed design for real-time applications. Jorge Dias 0001, Lakmal D. Seneviratne |
SMC | 2 |
| 2016 | Context-based decision system for human-machine interaction applicationsabstractIn this paper we present a decision process to auto-adapt and improve human-machine interaction, simplifying the integration of algorithms and functionalities. The decision process is part of an innovative approach that integrates contextual information to orchestrate behaviours of an interactive system (i.e. perception and actuation features involved during interaction). Classical approaches focus on designing and implementing algorithms that take into account several environment features (e.g. light, pose, etc.) to adapt its performance obtaining accurate results. An advantage of these approaches is to concentrate complexity in one algorithm leading to simple system architectures. In the other hand, a disadvantage of such approaches is their limitation to adapt to conditions under different scenarios, which typically requires manual adjustments to compensate changes of environment features. Our hypothesis is that, we can improve the overall performance of human-machine interaction process if a decision process is introduced, which is responsible for selecting the most adequate actions/algorithms, with maximum performance, that achieve a certain goal under a given context. The results from exploratory simulations validate the proposed approach to be more effective in attaining specific goals in the interaction process, resorting to algorithms with low complexity. João Quintas, Paulo Menezes 0001, Jorge Dias 0001 |
SMC | 3 |
| 2016 | Advances and trends in visual crowd analysis: A systematic survey and evaluation of crowd modelling techniques
M. Sami Zitouni, Harish Bhaskar, Jorge Dias 0001, Mohammed E. Al-Mualla |
Neurocomputing | 3 |
| 2016 | Probabilistic Social Behavior Analysis by Exploring Body Motion-Based PatternsabstractUnderstanding human behavior through nonverbal-based features, is interesting in several applications such as surveillance, ambient assisted living and human-robot interaction. In this article in order to analyze human behaviors in social context, we propose a new approach which explores interrelations between body part motions in scenarios with people doing a conversation. The novelty of this method is that we analyze body motion-based features in frequency domain to estimate different human social patterns: Interpersonal Behaviors (IBs) and a Social Role (SR). To analyze the dynamics and interrelations of people's body motions, a human movement descriptor is used to extract discriminative features, and a multi-layer Dynamic Bayesian Network (DBN) technique is proposed to model the existent dependencies. Laban Movement Analysis (LMA) is a well-known human movement descriptor, which provides efficient mid-level information of human body motions. The mid-level information is useful to extract the complex interdependencies. The DBN technique is tested in different scenarios to model the mentioned complex dependencies. The study is applied for obtaining four IBs (Interest, Indicator, Empathy and Emphasis) to estimate one SR (Leading).The obtained results give a good indication of the capabilities of the proposed approach for people interaction analysis with potential applications in human-robot interaction. Kamrad Khoshhal, Urbano Nunes 0001, Jorge Dias 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Robust adaptive control of quadrotor unmanned aerial vehicle with uncertaintyabstractIn this paper, we deal with the stability and tracking control problem of quadrotor unmanned aerial vehicle (UAV) in the presence of the modeling error and disturbance uncertainty. The flight tracking control system combines classical proportional-derivative (PD) like term with robust and adaptive control term. Lyapunov method is used to design and show the asymptotic behavior of the linear and angular states of the vehicle. In contrast with other existing adaptive backstepping design, the proposed design is very simple and easy to implement as it does not require multiple design steps without using augmented signals and known bound of the uncertainty. Various experimental results on quadrotor UAV system are presented to demonstrate the effectiveness of the proposed design for real-time application. Muhammad Faraz Faraz, Reem Ashour, Jorge Dias 0001, Lakmal D. Seneviratne |
ICRA | 4 |
| 2015 | Designing an artificial attention system for social robotsabstractIn this paper, we introduce the main components comprising the action-perception loop of an overarching framework implementing artificial attention, designed to fulfil the requirements of social interaction (i.e., reciprocity, and awareness), with strong inspiration on current theories in functional neuroscience. We demonstrate the potential of our framework, by showing how it exhibits coherent behaviour without any inbuilt prior expectations regarding the experimental scenario. Current research in cognitive systems for social robots has suggested that automatic attention mechanisms are essential to social interaction. In fact, we hypothesise that enabling artificial cognitive systems with middleware implementing these mechanisms will empower robots to perform adaptively and with a higher degree of autonomy in complex and social environments. However, this type of assumption is yet to be convincingly and systematically put to the test. The ultimate goal will be to test our working hypothesis and the role of attention in adaptive, social robotics. Pablo Lanillos, João Filipe Ferreira, Jorge Dias 0001 |
IROS | 3 |
| 2015 | A Multi-soft-body Dynamic Model for Underwater Soft Robots
Federico Renda, Francesco Giorgio-Serchi, Frédéric Boyer, Cecilia Laschi, Jorge Dias 0001, Lakmal D. Seneviratne |
ISRR (1) | 5 |
| 2015 | Hierarchical Crowd Detection and Representation for Big Data Analytics in Visual SurveillanceabstractIn this paper, a motion and appearance saliency combined detection framework for hierarchical representation of targets from groups to individuals in crowded scenes of surveillance videos is proposed. Big data analytic solutions within surveillance often require compact representations for target (s)- of-interest that allows simultaneous micro (individualistic) and macro (holistic) levels of inference on visual information. The target detection method proposed in this paper combines the estimation of motion saliency through dynamic texture (DT) based Gaussian Mixture Model (GMM) and appearance saliency through person detection using combined Histogram of Oriented Gradient (HOG) and Local Binary Patterns (LBP) feature sets. The saliency models are tightly integrated such that initially motion information is used to update and improve detection within an appearance framework, which in turn compliments the motion segmentation for accurate localization of people in groups. The improved people detection thus proposed is capable of eliminating false detections and can accurately delineate individuals within groups. The quantitative and qualitative results of experiments conducted on benchmark datasets have proven the validity and robustness of the proposed technique. M. Sami Zitouni, Jorge Dias 0001, Mohammed E. Al-Mualla, Harish Bhaskar |
SMC | 2 |
| 2015 | Recognition and action for scene understanding
Rebeca Marfil, Jorge Dias 0001, Francisco Escolano |
Neurocomputing | 2 |
| 2015 | Trajectory-based human action segmentation
Luís Santos 0001, Kamrad Khoshhal, Jorge Dias 0001 |
Pattern Recognit. | 3 |
| 2014 | Frequency-domain flight dynamics model identification of MAVs -miniature quad-rotor aerial vehiclesabstractIn this paper, a complete system identification process for identifying flight dynamics model of MAVs (miniature aerial vehicles) is presented. CIFER identification toolkit, which is developed by NASA Ames research center and particularly suitable for identifying rotary-wing aircraft dynamics, is adopted. The modeling procedure is detailed by addressing the following four key steps: 1) data collection and processing based on frequency-sweep input form, 2) parameter identification in frequency domain by minimizing the cost function, 3) result analysis based on the frequency responses matching, the Cramer-Rao Bound, and the Insensitivity, and 4) model fidelity validation in time domain. Hind Al Mehairi, Hanan Al-Hosani, Jorge Dias 0001, Lakmal D. Seneviratne |
IROS | 4 |
| 2014 | Touch attention Bayesian models for robotic active haptic exploration of heterogeneous surfacesabstractThis work contributes to the development of active haptic exploration strategies of surfaces using robotic hands in environments with an unknown structure. The architecture of the proposed approach consists two main Bayesian models, implementing the touch attention mechanisms of the system. The model πperperceives and discriminates different categories of materials (haptic stimulus) integrating compliance and texture features extracted from haptic sensory data. The model πtaractively infers the next region of the workspace that should be explored by the robotic system, integrating the task information, the permanently updated saliency and uncertainty maps extracted from the perceived haptic stimulus map, as well as, inhibition-of-return mechanisms. The experimental results demonstrate that the Bayesian model πpercan be used to discriminate 10 different classes of materials with an average recognition rate higher than 90%. The generalization capability of the proposed models was demonstrated experimentally. The ATLAS robot, in the simulation, was able to perform the following of a discontinuity between two regions made of different materials with a divergence smaller than 1cm (30 trials). The tests were performed in scenarios with 3 different configurations of the discontinuity. The Bayesian models have demonstrated the capability to manage the uncertainty about the structure of the surfaces and sensory noise to make correct motor decisions from haptic percepts. Ricardo Martins 0002, João Filipe Ferreira, Jorge Dias 0001 |
IROS | 3 |
| 2014 | Be the robot: Human embodiment in tele-operation driving tasksabstractThis paper proposes a new interaction mechanism for tele-operating a mobile robot. The approach explores the notion of telepresence and physical embodiment to create what may be called tele-embodiment. Its principle is that the operator will see himself at the remote site and this will enable him/her to better operate the robot. Four interaction styles were experimentally compared, from the traditional joystick approaches to more innovative based on natural body posture intentions. The environment perception is provided by the visual feedback, according to head pose behaviour. The results show that the gesture and body based methods improves the user dexterity performing this kind of task. Moreover, the present study suggests that, when a person is focused on the task, achieving the ownership illusion towards remote body, there are autonomic responses that correspond to what would be expected in events that take place in reality (like avoiding collisions). Luís Almeida 0002, Bruno Patrão, Paulo Menezes 0001, Jorge Dias 0001 |
RO-MAN | 4 |
| 2013 | Context-aware cooperation between human and robotic teams in catastrophic incidentsabstractThe study of cooperative interaction between multi-party multi-agent teams that include humans and robots is a recent scientific challenge. Preliminary empiric results about this interaction, in the scope of search and rescue applications, demonstrate the need for deeper studies on how humans should interact with teams of autonomous mobile robots and on how to establish a mutual beneficial interaction. This work presents the work in progress in the scope of CHOPIN1project, which aims to address some of these issues and will focus on devising new methods for collaborative context awareness and context sharing between teams of humans and teams of robots. João Quintas, Paulo Menezes 0001, Jorge Dias 0001 |
RO-MAN | 3 |
| 2013 | Context-based perception and understanding of human intentionsabstractThis work focus in the importance of context awareness and intention understanding capabilities in modern robots when faced with different situations. The objective is to be capable of providing new features for robots, which enable new real-world applications, and extend their autonomy, in terms of self-management and cooperation with humans or other systems. Gaze estimation and gesture interpretation are modalities, closely related with context-dependent human intention understanding, that are addressed in this work. João Quintas, Paulo Menezes 0001, Jorge Dias 0001 |
RO-MAN | 3 |
| 2013 | Mapping for Unknown Environments Using Triangulated MapsabstractIn this paper we present a new geometrical mapping structure that captures both the geometry and the connectivity of the environment. A robot's ability to successfully complete a required task is bound by its knowledge about the operation environment. Thus, the robot must be able to collect information from its surrounding and map it accurately to create a correct and complete representation of the environment. The solution in this paper uses the Gap-Navigation Tree as the underlying structure for the proposed Triangulation-Based exploration which maps the environment using the Dynamic Triangulation Tree structure DTT developed in this study. The efficiency of the proposed strategy is validated experimentally through simulations. The DTT does not only embed the geometry of the environment but also provides a direct mapping of the connectivity of its free space. The proposed algorithm is tested in simulations using various scenarios for exhaustive validation to prove its main advantages, namely ease of construction, compactness and completeness. Amna AlDahak AlShamsi, Lakmal D. Seneviratne, Jorge Dias 0001 |
SMC | 3 |
| 2013 | Improved Semantic-Based Human Interaction Understanding Using Context-Based KnowledgeabstractThis paper proposes a descriptive approach for context-based human activity analysis through an hierarchical framework in a scene understanding application. Each human movement with respect to himself, others and scene, can arise different layers of human activities analysis, which usually investigated separately depend on the application. Human behaviour can not be analysed properly, since the all different layers of information were not considered. The effect of using the different layers of information to increase the accuracy of the analysis is presented in the study. The contributions are, using different information layers such as human body parts movement and human-object interaction, in 3D space, to improve human activity analysis, and proposing a probabilistic and descriptive model, based on a well-known human movement descriptor and Bayesian Network (BN) approach. Thus, based on the mentioned framework, the model is generalizable and flexible which are necessary for having such an applicable system. The capability of the proposed approach is presented in the experiment's section. Kamrad Khoshhal, Jorge Dias 0001 |
SMC | 2 |
| 2013 | Probabilistic human interaction understanding: Exploring relationship between human body motion and the environmental context
Kamrad Khoshhal, Jorge Dias 0001 |
Pattern Recognit. Lett. | 2 |
| 2013 | Scene understanding and behaviour analysis
Antonio Bandera, Jorge Dias 0001, Francisco Escolano |
Pattern Recognit. Lett. | 2 |
| 2013 | A Bayesian Framework for Active Artificial PerceptionabstractIn this paper, we present a Bayesian framework for the active multimodal perception of 3-D structure and motion. The design of this framework finds its inspiration in the role of the dorsal perceptual pathway of the human brain. Its composing models build upon a common egocentric spatial configuration that is naturally fitting for the integration of readings from multiple sensors using a Bayesian approach. In the process, we will contribute with efficient and robust probabilistic solutions for cyclopean geometry-based stereovision and auditory perception based only on binaural cues, modeled using a consistent formalization that allows their hierarchical use as building blocks for the multimodal sensor fusion framework. We will explicitly or implicitly address the most important challenges of sensor fusion using this framework, for vision, audition, and vestibular sensing. Moreover, interaction and navigation require maximal awareness of spatial surroundings, which, in turn, is obtained through active attentional and behavioral exploration of the environment. The computational models described in this paper will support the construction of a simultaneously flexible and powerful robotic implementation of multimodal active perception to be used in real-world applications, such as human-machine interaction or mobile robot navigation. João Filipe Ferreira, Jorge Lobo 0002, Pierre Bessière, Miguel Castelo-Branco, Jorge Dias 0001 |
IEEE Trans. Cybern. | 5 |
| 2012 | A Multi-criteria Sorting Approach for Diagnosing Mental Disabilities
Paulo Freitas, Carlos Henggeler Antunes, Jorge Dias 0001 |
ICORES | 3 |
| 2012 | Context-based understanding of interaction intentionsabstractThis paper focus in the importance of context awareness and intention understanding capabilities in modern robots when faced with different situations. The inclusion of such requirements in robot design aim for more intelligent robots capable to adapt its behaviours to the faced situations. Gaze estimation and gesture interpretation are modalities, closely related with context-depent human intention understanding, that are addressed in this work. João Quintas, Luís Almeida 0002, Miguel Brito, Gustavo Quintela, Paulo Menezes 0001, Jorge Dias 0001 |
RO-MAN | 6 |
| 2011 | Towards human motion capture from a camera mounted on a mobile robot
Paulo Menezes 0001, Frédéric Lerasle, Jorge Dias 0001 |
Image Vis. Comput. | 3 |
| 2010 | Human silhouette volume reconstruction using a gravity-based virtual camera network
Hadi Aliakbarpour, Jorge Dias 0001 |
FUSION | 2 |
| 2010 | Crowd behavior analysis under cameras network fusion using probabilistic methods
Paulo L. J. Drews-Jr, João Quintas, Jorge Dias 0001, Maria Andersson, Jonas Nygårds, Joakim Rydell |
FUSION | 3 |
| 2010 | Probabilistic LMA-based classification of human behaviour understanding using Power Spectrum technique
Kamrad Khoshhal, Hadi Aliakbarpour, João Quintas, Paulo L. J. Drews-Jr, Jorge Dias 0001 |
FUSION | 5 |
| 2010 | Novelty detection and 3D shape retrieval using superquadrics and multi-scale sampling for autonomous mobile robotsabstractThere are several applications for which it is important to both detect and communicate changes in data models. For instance, in some mobile robotics applications (e.g. surveillance) a robot needs to detect significant changes in the environment (e.g. a layout change) which it may achieve by comparing current data provided by its sensors with previously acquired data (e.g. map) of the environment. This often constitutes an extremely challenging task due to the large amounts of data that must be compared in real-time. This paper proposes a framework to detect, and represent changes through a compact model. The main steps of the procedure are: multi-scale sampling to reduce the computation burden; change detection based on Gaussian mixture models; fitting superquadrics to detected changes; and refinement and optimization using the split and merge paradigm. Experimental results in various real and simulated scenarios demonstrate the approach's feasibility and robustness with large datasets. Paulo L. J. Drews-Jr, Pedro Núñez Trujillo, Rui P. Rocha, Mario Fernando Montenegro Campos, Jorge Dias 0001 |
ICRA | 5 |
| 2010 | Probabilistic representation of 3D object shape by in-hand explorationabstractThis work presents a representation of 3D object shape using a probabilistic volumetric map derived from in-hand exploration. The exploratory procedure is based on contour following through the fingertip movements on the object surface. We first consider the simple case of having single hand exploration of a static object. The cumulative pose data provides a 3D point cloud that is quantized to the probabilistic volumetric map. For each voxel we have a probability distribution for the occupancy percentage. This is then extended to in-hand exploration of non-static objects. Since the object is moving during the in-hand exploration, and we also consider the use of the other hand for re-grasping, object pose has to be tracked. By keeping track of object motion we can register data to the initial pose to build a consistent object representation. An object centered representation is implemented using the computed object center of mass to define its frame of reference. Results are presented for in-hand exploration of both static and non-static objects that show that valid models can be obtained. The 3D object probabilistic representation can be used in several applications related with grasp generation tasks. Diego R. Faria, Ricardo Martins 0002, Jorge Lobo 0002, Jorge Dias 0001 |
IROS | 4 |
| 2010 | Change detection in 3D environments based on Gaussian Mixture Model and robust structural matching for autonomous robotic applicationsabstractThe ability to detect perceptions which were never experienced before, i.e. novelty detection, is an important component of autonomous robots working in real environments. It is achieved by comparing current data provided by its sensors with a previously known map of the environment. This often constitutes an extremely challenging task due to the large amounts of data that must be compared in real-time. With respect to previously proposed approaches, this paper detects changes in 3D environment based on probabilistic models, the Gaussian Mixture Model, and a fast and robust combined constraint matching algorithm. The matching allows to represent the scene view as a graph which emerges from the comparison between Mixtures of Gaussians. Finding the largest set of mutually consistent matches is equivalent to find the maximum clique on a graph. The proposed approach has been tested for mobile robotics purposes in real environments and compared to other matching algorithms. Experimental results demonstrate the performance of the proposal. Pedro Núñez Trujillo, Paulo L. J. Drews-Jr, Antonio Bandera, Rui P. Rocha, Mario Fernando Montenegro Campos, Jorge Dias 0001 |
IROS | 6 |
| 2009 | Multiclass brain computer interface based on visual attention
Rolando Grave de Peralta Menendez, Jorge Dias 0001, José Augusto Prado, Hadi Aliakbarpour, Sara González Andino |
ESANN | 2 |
| 2009 | 3D hand trajectory segmentation by curvatures and hand orientation for classification through a probabilistic approachabstractIn this work we present the segmentation and classification of 3D hand trajectory. Curvatures features are acquired by (r, ¿, h) and the hand orientation is acquired by approximating the hand plane in 3D space. The 3D positions of the hand movement are acquired by markers of a magnetic tracking system. Observing humans movements we perform a learning phase using histogram techniques. Based on the learning phase is possible classify reach-to-grasp movements applying Bayes rule to recognize the way that a human grasps an object by continuous classification based on multiplicative updates of beliefs. We are classifying the hand trajectory by its curvatures and by hand orientation along the trajectory individually. Both results are compared after some trials to verify the best classification between these two kinds of segmentation. Using entropy as confidence level, we can give weights for each kind of classification to combine both, acquiring a new classification for results comparison. Using these techniques we developed an application to estimate and classify two possible types of grasping by the reach-to-grasp movements performed by humans. These reported steps are important to understand some human behaviors before the object manipulation and can be used to endow a robot with autonomous capabilities (e.g. reaching objects for handling). Diego R. Faria, Jorge Dias 0001 |
IROS | 2 |
| 2009 | Human Robot interaction studies on laban human movement analysis and dynamic background segmentationabstractHuman movement analysis through vision sensing systems is an important subject regarding Human-Robot interaction. This is a growing area of research, with wide range of applications fields. The ability to recognize human actions using passive sensing modalities, is a decisive factor for machine interaction. In mobile platforms, image processing is regarded as a problem, due to constant changes. We propose an approach, based on Horopter technique, to extract Regions Of Interest (ROI) delimiting human contours. This fact will allow tracking algorithms to provide faster and accurate responses to human feature extraction. The key features are head and both hand positions, that will be tracked within image context. Posterior to feature acquisition, they will be contextualized within a technique, Laban Movement Analysis (LMA) and will be used to provide sets of classifiers. The implementation of the LMA technique will be based on Bayesian Networks. We will use these Bayesian classifiers to label/classify human emotion within the context of expressive movements. Compared to full image tracking, results improved with the implemented approach, the horopter and consequently so did classification results. Luís Santos 0001, José Augusto Prado, Jorge Dias 0001 |
IROS | 3 |
| 2009 | Novelty detection and 3D shape retrieval based on Gaussian Mixture Models for autonomous surveillance roboticsabstractThis paper describes an efficient method for retrieving the 3-dimensional shape associated to novelties in the environment of an autonomous robot, which is equipped with a laser range finder. First, changes are detected over the point clouds using a combination of theGaussian mixture model(GMM) and theearth mover's distance(EMD) algorithms. Next, the shape retrieval is achieved using two different algorithms. First, new samplings are generated from each Gaussian function, followed by arandom sampling consensus(RANSAC) algorithm to retrieve geometric primitives. Furthermore, a new algorithm is developed to directly retrieve the shape according to the mathematical space of Gaussian mixture. In this paper, the set of geometric primitives has been limited to the setC = {sphere, cylinder, plane}. The two shape retrieval methods are compared in terms of computational cost and accuracy. Experimental results in various real and simulated scenarios demonstrate the feasibility of the approach. Pedro Núñez Trujillo, Paulo L. J. Drews-Jr, Rui P. Rocha, Mario Fernando Montenegro Campos, Jorge Dias 0001 |
IROS | 5 |
| 2009 | A technique for dynamic background segmentation using a robotic stereo vision headabstractHuman-robot interaction approaches like face detection, face recognition, pedestrian detection are widely known in robotics field; however often they lead to performance problems. Additionally, false positive and false negative problems are commonly associated to bad illumination and strong featured images. Moreover background segmentation approaches are frequently used to solve this problem on static camera surveillance. Though all these approaches are unable to effectively deal with the constant background changes that certainly happens when the camera sensor is installed on a mobile robot. Hence, in this work we propose a stereo vision dynamic background segmentation solution to this problem. José Augusto Prado, Luís Santos 0001, Jorge Dias 0001 |
RO-MAN | 3 |
| 2008 | Robotic implementation of biological Bayesian models for visuo-inertial image stabilization and gaze controlabstractRobotic implementations of gaze control and image stabilization have been previously proposed, that rely on fusing inertial and visual sensing modalities. Human and biological system also combine the two sensing modalities for the same goal. In this work we build upon these previous results and, with the contribution of psychophysical studies, attempt a more bio-inspired approach to the robotic implementation. Since Bayesian models have been successfully used to explain psychophysical experimental findings, we propose a robotic implementation using Bayesian inference. Jorge Lobo 0002, João Filipe Ferreira, José Augusto Prado, Jorge Dias 0001 |
IROS | 4 |
| 2008 | Laban Movement Analysis for multi-ocular systemsabstractWe present as a contribution to the field of human-machine interaction a system that analyzes human movements online through multiple observers, based on the concept of Laban Movement Analysis (LMA). The implementation uses a Bayesian model for learning and classification, while the results are presented for the application to analyze expressive movements. In sports like Karate four judges are placed in the corners to observe the fight to ensure that the overall judgment is correct. In this paper we propose a multi-ocular system where each sub-system observes a movement from a different monocular perspective. The sub-systems send continuously guesses in form of probability distributions to the central system. The central system fuses the evidences and presents the final result. We present the Laban Movement Analysis as a concept to identify useful features of human movements to classify human actions. The movements are extracted using both, vision and magnetic tracker. The descriptor opens possibilities towards expressiveness and emotional content. To solve the problem of classification we use the Bayesian framework as it offers an intuitive approach to learning and classification. The presented work targets applications like social robots, smart houses and surveillance. Jörg Rett, Luís Santos 0001, Jorge Dias 0001 |
IROS | 3 |
| 2008 | Multi-robot complete exploration using hill climbing and topological recoveryabstractThis article addresses the problem of autonomous map building and exploration of an unknown environment with mobile robots. The proposed method assumes that mobile robots use occupancy grid maps as the main representation model for the built maps and a hill climbing local search algorithm for exploring the environment without any kind of human intervention. It is demonstrated that hill climbing based exploration may recover from local minima and cover completely any environment, if a topological representation of the environment is created incrementally along the mapping and exploration mission. The approach is devised for either a single mobile robot or multiple cooperative mobile robots. Rui P. Rocha, João Filipe Ferreira, Jorge Dias 0001 |
IROS | 3 |
| 2007 | Trajectory recovery and 3D mapping from rotation-compensated imagery for an airshipabstractOn this paper, inertial orientation measurements are exploited to compensate the rotational degrees of freedom for an aerial vehicle carrying a perspective camera, taking a sequence of images of the ground plane. It is known that, on the pure translation case, full homographies are reduced to planar homologies, and the relative scene depth of two points equals the reciprocal ratio of their image distances to the the FOE. The first part of this paper covers trajectory recovery for an airship carrying a perspective camera taking a sequence of images of the ground plane, as a series of relative poses between successive camera poses. This is commonly named “Visual Odometry“. Previous results showed that the ratio of heights over the ground plane on two views can be calculated more accurately, and thus the altitude component of the trajectory, and here these results are extended by recovering the full 3D camera trajectory. In the second part, the same rotation-compensated imagery is exploited on the mapping domain: from pixel correspondences between successive images the height of points over the ground plane can be recovered, and placed on a DEM grid, performing 3D mapping from monocular aerial images. These results may be useful on the SLAM context. Luiz G. B. Mirisola, Jorge Dias 0001, Aníbal T. de Almeida |
IROS | 2 |
| 2006 | Visual Tracking Modalities for a Companion RobotabstractThis article presents the development of a human-robot interaction mechanism based on vision. The functionalities required for such mechanism range from user detection and recognition, to gesture tracking. Particle filters, which are extensively described in the literature, are well suited to this context as they enable a straight combination of several visual cues like colour, shape or motion. Additionally, different algorithms can be considered for a better handling of the particles depending of the context. This article presents the visual functionalities developed namely user recognition and following, and 3D gestures tracking. The challenge is to find which algorithms and visual cues fulfil the best, the requirements of the considered functionalities for our companion robot. The employed methods to attain these required functionalities and their results are presented Paulo Menezes 0001, Frédéric Lerasle, Jorge Dias 0001 |
IROS | 3 |
| 2005 | Multimodal functional and morphological nonrigid image registrationabstractThe aim of this work is to present the registration of two different and complementary imaging modalities of the human eye fundus. One is a widely used modality consisting of a color photograph of the human retina and the other consists of a functional imaging modality that assesses, in vivo, the human blood-retinal barrier function. The need for the nonrigid registration of these two modalities is demonstrated and is due to the acquisition mode of the functional modality. The two modalities here with registered are from two different devices, have different resolutions and different fields-of-view, i.e., the smaller is only a subset of the larger one. In this paper, all these problems were addressed and a solution proposed. Rui Bernardes, Pedro Baptista, Jorge Dias 0001, José Cunha-Vaz |
ICIP (1) | 3 |
| 2005 | Cooperative Multi-Robot Systems A study of Vision-based 3-D Mapping using Information TheoryabstractBuilding cooperatively 3-D maps of unknown environments is one of the application fields of multi-robot systems. This article introduces a distributed architecture, within a probabilistic framework for vision-based 3-D mapping, whereby each robot is committed to cooperate with other robots through information sharing. An entropy-based measure of information utility is defined, which a robot uses for communicating to its teammates the most useful measurements, thus preventing the robot to overwhelm communication resources with redundant information. Experiments with real robots, equipped with stereo-vision, yielded important conclusions about the way robots should cooperate on sharing information. Rui P. Rocha, Jorge Dias 0001, Adriano Carvalho 0001 |
ICRA | 2 |
| 2005 | A low-level framework for a probabilistic treatment of the topological description of a robot missionabstractThis article describes a mathematical basis required to integrate features obtained for perception for topological navigation. It is intended for application to navigation in an environment that is not mapped, but in which a mission is described in the form of a semantic description of the perception stimulus that the robot is expected to encounter. The need to integrate features from different sensors led to the use of an uncertainty estimate employed in information theory; binary entropy. By using entropy, the features are ranked in order of decreasing uncertainty. This article describes the state of the work in an as yet preliminary stage, but appears promising for application to navigation using topological information. It also offers interesting perspectives on commonly used sensory data such as local intensity image features. João Filipe Ferreira, Vítor M. F. Santos, Jorge Dias 0001 |
IROS | 3 |
| 2005 | Exploring information theory for vision-based volumetric mappingabstractThis article presents an innovative probabilistic approach for building volumetric maps of unknown environments with autonomous mobile robots, which is based on information theory. Each mobile robot uses an entropy gradient-based exploration strategy, with the aim of maximizing information gain when building and improving a 3D map upon measurements yielded by an on-board stereo-vision sensor. The proposed framework was validated through experiments with a real mobile robot equipped with stereo-vision, in order to be further used on cooperative volumetric mapping with teams of mobile robots. Rui P. Rocha, Jorge Dias 0001, Adriano Carvalho 0001 |
IROS | 2 |
| 2004 | Human-robot Interaction based on Haar-like Features and EigenfacesabstractThis paper describes a machine learning approach for visual object detection and recognition which is capable of processing images rapidly and achieving high detection and recognition rates. This framework is demonstrated on, and in part motivated by, the task of human-robot interaction. There are three main parts on this framework. The first is the person's face detection used as a preprocessing system to the second stage which is the recognition of the face of the person interacting with the robot, and the third one is the hand detection. The detection technique is based on Haar-like features introduced by Viola et al. and then improved by Lienhart et al. The eigenimages and PCA are used in the recognition stage of the system. Used in real-time human-robot interaction applications the system is able to detect and recognise faces at 10.9 frames per second in a PIV 2.2 GHz equipped with a USB camera. Jose Barreto, Paulo Menezes 0001, Jorge Dias 0001 |
ICRA | 3 |
| 2003 | Registration and Segmentation for 3D Map Building: A Solution Based on Stereo Vision and Inertial SensorsabstractThis article presents a technique for registration and segmentation of dense depth maps provided by a stereo vision system. The vision system uses inertial sensors to give a reference for camera pose. The maps are registered using a modified version of the ICP - iterative closet point algorithm to register dense depth maps obtained from a stereo vision system. The proposed technique explores the integration of inertial sensor data for dense map registration. Depth maps obtained by vision systems, are very point of view dependent, providing discrete layers of detected depth aligned with the camera. In this work we use inertial sensors to recover camera pose, and rectify the maps to a reference ground plane, enabling the segmentation of vertical and horizontal geometric features and map registration. We propose a real-time methodology segmentation of structures, object recognition, robot navigation or any other task that requires a three-dimensional representation of the physical environment. The aim of this work is a fast real-time system, which can be applied to autonomous robotic systems or to automated car driving systems, for modelling the road, identifying obstacles and roadside features in real-time. Jorge Lobo 0002, Luís Almeida 0002, Jorge Dias 0001 |
ICRA | 4 |
| 2003 | Vision and Inertial Sensor Cooperation Using Gravity as a Vertical ReferenceabstractThis paper explores the combination of inertial sensor data with vision. Visual and inertial sensing are two sensory modalities that can be explored to give robust solutions on image segmentation and recovery of 3D structure from images, increasing the capabilities of autonomous robots and enlarging the application potential of vision systems. In biological systems, the information provided by the vestibular system is fused at a very early processing stage with vision, playing a key role on the execution of visual movements such as gaze holding and tracking, and the visual cues aid the spatial orientation and body equilibrium. In this paper, we set a framework for using inertial sensor data in vision systems, and describe some results obtained. The unit sphere projection camera model is used, providing a simple model for inertial data integration. Using the vertical reference provided by the inertial sensors, the image horizon line can be determined. Using just one vanishing point and the vertical, we can recover the camera's focal distance and provide an external bearing for the system's navigation frame of reference. Knowing the geometry of a stereo rig and its pose from the inertial sensors, the collineations of level planes can be recovered, providing enough restrictions to segment and reconstruct vertical features and leveled planar patches. Jorge Lobo 0002, Jorge Dias 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Segmentation of dense depth maps using inertial data a real-time implementationabstractWe propose a real-time system that extracts information from dense relative depth maps. This method enables the integration of depth cues on higher level processes including segmentation of structures, object recognition, robot navigation or any other task that requires a 3D representation of the physical environment. Inertial sensors coupled to a vision system can provide important inertial cues for the ego-motion and system pose. In this work we explore the integration of inertial sensor data in vision systems. Depth maps obtained by vision systems, are very point of view dependant, providing discrete layers of detected depth aligned with the camera. We use inertial sensors to recover the camera pose, and rectify the maps to a reference ground plane, enabling the segmentation of vertical and horizontal geometric features. The aim of this work is a fast real-time system, so that it can be applied to autonomous robotic systems or to automated car driving systems, for modelling the road, identifying obstacles and roadside features in real-time. Jorge Lobo 0002, Luís Almeida 0002, Jorge Dias 0001 |
IROS | 3 |
| 1998 | Ground plane detection using visual and inertial data fusionabstractActive vision systems can be used in robotic systems for navigation. The active vision system provides data on the robot's environment. In mobile systems the position and attitude of the cameras relative to the world can be hard to determine. Inertial sensors coupled to the active vision system can provide valuable information to aid the image processing task. In human and other animals the vestibular system plays a similar role. In this article, we explain our recent steps in the integration of inertial data with an active vision system. The active vision system has a set of stereo cameras capable of vergence, with a common baseline, pan and tilt. A process of visual fixation has already been implemented, enabling symmetric vergence on any selected point. An inertial system prototype, based on low-cost sensors was built. It is used to keep track of the gravity vector, allowing the identification of the vertical in the images. By performing visual fixation of a ground plane point, and knowing the 3D vector normal to level ground we can determine the ground plane. The image can therefore be segmented, and the ground plane along which the robot can move identified. For on-the-fly visualisation of the segmented images and the detected points a VRML viewer is used. Jorge Lobo 0002, Jorge Dias 0001 |
IROS | 2 |
| 1998 | Exploring spherical image properties for robot navigationabstractAddresses the problem of performing navigation with a mobile robot using active vision exploring the image sphere properties in a stereo divergent configuration. The navigational process is supported by the control of the robot's steering and forward movements just using visual information as feedback. The steering control solution is based on the difference between signals of visual motion flow computed in images on different positions of a virtual image sphere. The majority of solutions based on motion flow and proposed until now, were usually very unstable because they normally compute other parameters from the motion flow. In our case the control is based directly on the different between motion flow signals on different images. Those multiple images are obtained by small mirrors, that simulate cameras positioned in different positions on the image sphere. The control algorithm described in this work is based on a discrete-event approach to generate a control feedback signal for navigation of an autonomous robot with an active vision system. Inácio Fonseca, Jorge Dias 0001 |
IROS | 2 |
| 1998 | Simulating pursuit with machine experiments with robots and artificial visionabstractThis article describes one solution for the problem of pursuit of objects moving on a plane by using a mobile robot and an active vision system. The solution deals with the interaction of different control systems using visual feedback and it is accomplished by the implementation of a visual gaze holding process interacting cooperatively with the control of the trajectory of a mobile robot. These two systems are integrated to follow a moving object at constant distance and orientation with respect to the mobile robot. The orientation and the position of the active vision system running a gaze holding process give the feedback signals to the control used to pursuit the target in real-time. The paper addresses the problems of visual fixation, visual smooth pursuit, navigation using visual feedback and compensation for system's movements. The algorithms for visual processing and control are described in the article. The mechanisms of cooperation between the different control and visual algorithms are also described. The final solution is a system able to operate at approximately human walking rates as the experimental results show in the paper. Jorge Dias 0001, Carlos Paredes, Inácio Fonseca, Helder Araújo, Jorge P. Batista, Aníbal T. de Almeida |
IEEE Trans. Robotics Autom. | 1 |
| 1997 | Avoiding obstacles using a connectionist networkabstractIn this article, visual data obtained by a binocular active vision system is integrated, together with ultrasonic range measurements, in the development of a obstacle detection and avoidance system based on a connectionist grid. The traditional notion of probabilistic occupation grid is extended through the use of a three-layer structure of connectionist networks which allows the integration of several sensorial modalities (in this case ultrasonic sensor readings and stereo vision information) in a probabilistic environment representation. The connectionist nature of the network also allows us to deal with obstacle avoidance by using a mechanism similar to potential field over a discrete set of the robot's configuration space with each grid node representing a possible configuration. The value in each grid node gives us a measure of the configuration occupancy probability and can also be used to guide the robot to a predefined goal configuration simulating a simple gradient descending technique. Finally we present experimental results obtained with the implementation of the above method in a mobile platform which also provides the support for the sensing devices described throughout the article. A. Silva, Paulo Menezes 0001, Jorge Dias 0001 |
IROS | 3 |
| 1996 | Pursuit control in a binocular active vision system using optical flowabstractAn active vision system must be able to perform reactive visual processes in real time. In this paper we discuss a number of issues related to the implementation of a real-time control architecture and describe the architecture we use with camera heads. Another important issue of the operation of active vision binocular heads is their integration into more complex robotic systems. Higher levels of autonomy and integration can be obtained by designing the control system architecture based on the concept of purposive behavior. At the lower levels we consider the vision as a sensor and integrate it in the control systems (both feedforward and servo loops); and several visual processes are implemented in parallel, computing relevant measures for the control process. At higher levels the architecture is modelled as a state transition system. Finally, we show how this architecture can be used to implement a pursuit behavior using optical flow. Simultaneously vergence control can also be performed using the same visual processes. Helder Araújo, Jorge P. Batista, Paulo Peixoto, Jorge Dias 0001 |
ICPR | 4 |
| 1995 | Simulating Pursuit with Machines Experiments with Robots and Artificial VisionabstractThis article describes one solution for the problem of pursuit by using a mobile robot and an active vision system. The solution deals with the interaction of different control systems using visual feedback and it is accomplished by the implementation of a visual gaze holding process interacting cooperatively with the control of the trajectory of a mobile robot. These two systems are integrated to follow a moving object at constant distance and orientation with respect to the mobile robot. The orientation and the position of the active vision system running a gaze holding process give the feedback signals to the control used to pursuit the target in real-time. The mechanisms of cooperation between the different control and visual algorithms are described. The final solution is a system able to operate at approximately human walking rates. Jorge Dias 0001, Carlos Paredes, Inácio Fonseca, Helder Araújo, Jorge P. Batista, Aníbal T. de Almeida |
ICRA | 1 |
| 1993 | Monoplanar Camera Calibration
Jorge P. Batista, Jorge Dias 0001, Helder Araújo, Aníbal T. de Almeida |
BMVC | 2 |
| 1991 | Depth recovery using active focus in robotics [vision]abstractIn practice focusing can be obtained by displacing the sensor plate with respect to the image plane, by moving the lens or by moving the object with respect to the optical system. Moving the lens or sensor plate with respect to each other, causes changes of the magnification and corresponding changes on the object coordinates. In order to overcome these problems, the authors propose a technique to vary the degree of focusing by moving the camera with respect to the object position. The camera is attached to the tool of a manipulator in a hand-eye configuration with the position of the camera always known. This approach ensures that the focused areas of the image are always subjected to the same magnification. To measure the focus quality operators are used to evaluate the quantity of high-frequency components on the image. Different types of these operators was tested and the results compared.> Jorge Dias 0001, Aníbal T. de Almeida, Helder Araújo |
IROS | 1 |
| 1991 | Improving camera calibration by using multiple frames in hand-eye robotic systemsabstractAddresses the geometric modelling aspects related with the use of a pair of video cameras mounted on a six degrees-of-freedom manipulator in an hand-eye configuration. The information obtained by these sensors shall be used to generate the depth map of the environment where the manipulator works. The procedure to determining the image formation parameters is known by camera calibration and a method to optimize these calibration parameters is presented. After presenting the calibration technique used, the authors develop a calibration matrix invariant to movement. The optimization process is applied to this matrix to obtain a better performance in the numerical data.> Jorge Dias 0001, Aníbal T. de Almeida, Helder Araújo, Jorge P. Batista |
IROS | 1 |
| 1989 | Image processing system with multiple DSPs
Luí de Sá, Jorge Dias 0001, Vítor Silva 0001 |
Microprocess. Microprogramming | 2 |