EDBT 2026 Demo / reviewers in the wild / expert
Shan Luo 0001
dblp:93/622-1
· DBLP profile ↗
41ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0003-4760-0372ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 4 first-author · 21 since 2021Systems, architecture and hardware · 18 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective Robotic Cloth Grasping Through Suppressing False DiscoveriesabstractEnabling robots to grasp disorganized cloth for efficient storage is valuable in robot-assisted room organization. Diverse deformations of cloth and the stacking of multiple items limit grasping-pose estimation that relies on annotations. This necessitates segmenting each cloth item in an unsupervised manner before estimating the grasping position. However, existing segmentation methods primarily focus on improving metrics such as Intersection-over-Union and Pixel Accuracy, which cannot effectively measure the segmentation errors of the cloth area and thus lead to failure grasping position estimation. To address this challenge, we use False Discovery Rate (FDR) as a novel measure of segmentation errors and analyze its impact on grasping success. Our preliminary study reveals a negative correlation between segmentation FDR and grasping success rate, highlighting the need for more reliable segmentation in cluttered cloth scenarios. Therefore, we propose an unsupervised cloth segmentation network based on feature distance-weighted constraints, designed to reduce the false discovery rate in cloth area perception without requiring expensive pixel-level manual annotations. Additionally, to estimate the grasping position on the perceived cloth area, we introduce a strategy based on cloth surface wrinkle analysis, which operates without the need for annotations or training. By integrating the proposed segmentation network and grasping strategy, we develop a robotic system capable of sequentially grasping cluttered cloth from a table. Extensive real-world robotic experiments demonstrate the effectiveness of our approach, outperforming multiple baseline methods in segmentation FDR and grasping success rate. Xingyu Zhu 0014, Zhiwen Tu, Yan Wu 0002, Shan Luo 0001, Hechang Chen, Yixing Gao 0001 |
AAAI | 4 |
| 2026 | TransFace++: Rethinking the Face Recognition Paradigm With a Focus on Accuracy, Efficiency, and SecurityabstractFace Recognition (FR) technology has made significant strides with the emergence of deep learning. Typically, most existing FR models are built upon Convolutional Neural Networks (CNN) and take RGB face images as the model's input. In this work, we take a closer look at existing FR paradigms from high-efficiency, security, and precision perspectives, and identify the following three problems: (i) CNN frameworks are vulnerable in capturing global facial features and modeling the correlations between local facial features. (ii) Selecting RGB face images as the model's input greatly degrades the model's inference efficiency, increasing the extra computation costs. (iii) In the real-world FR system that operates on RGB face images, the integrity of user privacy may be compromised if hackers successfully penetrate and gain access to the input of this model. To solve these three issues, we propose two novel FR frameworks, i.e., TransFace and TransFace++, which successfully explore the feasibility of applying ViTs and image bytes to FR tasks, respectively. Firstly, as revealed from our observations, we find that ViTs perform vulnerably when applied to FR scenarios with extremely large datasets. We investigate the reasons for this phenomenon and discover that the existing data augmentation approaches and hard sample mining strategies are incompatible with ViTs-based FR backbone due to the lack of tailored consideration on preserving face structural information and leveraging each local token information. To remedy these problems, we first propose a superior FR model called TransFace, which contains a patch-level data augmentation strategy named Dominant Patch Amplitude Perturbation (DPAP) and a hard sample mining strategy named Entropy-guided Hard Sample Mining (EHSM). Furthermore, to improve inference efficiency and user privacy protection, we investigate the intrinsic property of image bytes and propose a superior FR model termed TransFace++. The proposed model is trained directly on image bytes, presenting a novel approach to address the aforementioned issues. Specifically, considering the importance of local correlations in bytes, an image bytes compression strategy named Topology-based Image Bytes Compression (TIBC) is introduced to extract prominent features from the raw bytes and integrate these features with byte embeddings, effectively mitigating information loss during the bytes mapping process. Moreover, to strengthen the model's perception on geometric information encoded in image bytes, a novel cross-attention module named Structure Information-guided Cross-Attention (SICA) is designed to inject structure information into byte tokens for information interaction, significantly improving the model's generalization ability. Experiments on popular face benchmarks demonstrate the superiority of our TransFace and TransFace++. Jun Dan, Yang Liu 0356, Baigui Sun, Jiankang Deng, Shan Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Can Vision Feel Touch? Tactile-Aware Visual Grasping for Transparent ObjectsabstractTransparent object manipulation has long posed a significant challenge in robotic grasping tasks. Existing methods for transparent object grasping rely heavily on visual sensors, aiming to extract relevant features from raw visual data to facilitate grasp execution. However, transparent objects often possess unreliable visual properties, while tactile contact reliably captures their physical properties. Moreover, these visual-based methods often overlook variations in object standing type (OST) and weight, limiting the precise grasping of transparent objects in different physical states. In contrast, humans naturally form memory associations between visual and tactile information and adjust grip force based on tactile feedback. Building on this foundation, we propose tactile-enhanced visual grasping (TEVG)—a novel method that augments robotic visual capabilities with tactile information to enable precise grasping of transparent objects with unknown OST and weight. The TEVG framework comprises two key components: pre-grasp enhancement (PE) and in-hand enhancement (IE). During the pre-grasp phase, PE embeds tactile features into the visual encoder to predict physical properties in advance, facilitating explicit identification of OST and accurate grasp pose prediction through the tactile-enhanced visual (TEV) encoder. IE enables real-time adaptive adjustment of grasp force during contact manipulation, allowing the system to handle objects with unknown weight effectively. Experimental results on two different robotic platforms demonstrate that TEVG significantly enhances the accuracy and stability of grasping transparent objects. The experiment video and project are publicly available at: https://sites.google.com/view/cvft1. Ling Tong 0006, Kun Qian 0005, Zhaokun Yue, Shan Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | ViTaDex: Vision-Tactile Fusion for 6-D Object-in-Hand Pose Estimation in Dexterous Anthropomorphic ManipulationabstractObject-in-hand 6-D pose estimation is critical for dexterous manipulation and assembly tasks. Existing object-in-hand pose estimation methods neglect the varying significance of tactile information from distinct hand regions during feature extraction and employ simplistic concatenation for vision–tactile fusion, which fails to capture the complementary relationships between these modalities. Moreover, prevailing object-in-hand datasets predominantly rely on predefined motion trajectories and single-point fingertip contact information, which inadequately represent real-world scenarios. To address these limitations, we propose the vision–tactile adaptive fusion network. Specifically, we design a novel tactile-gated adaptive graph convolution to dynamically model multiregional tactile features across the hand. We also introduce a vision–tactile cross-attention gated fusion module that facilitates efficient integration of multimodal sensory information. In addition, we construct the Vision–Tactile Object-in-Hand Dexterous Manipulation dataset based on human demonstration for 6-D pose estimation. The dataset employs a teleoperation system that combines an exoskeleton hand with a pose tracking system to replicate human dexterous behaviors and encompasses high-density tactile data spanning the entire hand. Extensive experiments demonstrate the robustness of our framework against occlusions. Ling Tong 0006, Kun Qian 0005, Zhaokun Yue, Shan Luo 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | EleTac: Elephant Trunk Tip-Inspired Soft Gripper With Vision-Based Tactile Sensing and ProprioceptionabstractSoft grippers offer gentle interaction with objects, significantly reducing the risk of damage. Their compliance in both material and structure allows them to adapt to a wide variety of object shapes and sizes. However, the deformable nature of soft grippers poses challenges for integrating precise proprioception and tactile sensing, especially when aiming for large-area high-resolution tactile perception. In nature, the elephant's trunk exemplifies an ideal combination of compliance and tactile sensitivity, enabling it to delicately manipulate diverse objects without causing damage while also exploring its surroundings through touch. Inspired by this, we present EleTac, a soft vision-based tactile gripper that enables safe object grasping via a pinch-like motion and delivers high-resolution full-surface tactile feedback for integrated proprioceptive and exteroceptive sensing. Experimental results demonstrate that even with a simple control strategy, EleTac can robustly grasp and lift a variety of objects, exhibiting strong adaptability and generalization. Furthermore, the seamless integration of grasping and tactile sensing capabilities facilitates practical applications, including exploration and excavation in granular media, as well as adaptive surface following. Overall, EleTac validates the 'manipulator-as-sensor' design philosophy, achieving high-quality tactile feedback without requiring additional sensing modules. Tuan Tai Nguyen, Quan Khanh Luu, Shan Luo 0001, Van Anh Ho |
IEEE Trans. Robotics | 4 |
| 2025 | Fractal Calibration for Long-tailed Object DetectionabstractReal-world datasets follow an imbalanced distribution, which poses significant challenges in rare-category object detection. Recent studies tackle this problem by developing re-weighting and re-sampling methods, that utilise the class frequencies of the dataset. However, these techniques focus solely on the frequency statistics and ignore the distribution of the classes in image space, missing important information. In contrast to them, we propose FRActal CALibration (FRACAL): a novel post-calibration method for long-tailed object detection. FRACAL devises a logit adjustment method that utilises the fractal dimension to estimate how uniformly classes are distributed in image space. During inference, it uses the fractal dimension to inversely down-weight the probabilities of uniformly spaced class predictions achieving balance in two axes: between frequent and rare categories, and between uniformly spaced and sparsely spaced classes. FRACAL is a post-processing method and it does not require any training, also it can be combined with many off-the-shelf models such as one-stage sigmoid detectors and two-stage instance segmentation models. FRACAL boosts the rare class performance by up to 8.6% and surpasses all previous methods on LVIS dataset, while showing good generalisation to other datasets such as COCO, V3Det and OpenImages. We provide the code at https://github.com/kostas1515/FRACAL. Konstantinos Panagiotis Alexandridis, Ismail Elezi, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
CVPR | 5 |
| 2025 | TransForce: Transferable Force Prediction for Vision-Based Tactile Sensors with Sequential Image TranslationabstractVision-based tactile sensors (VBTSs) provide highresolution tactile images crucial for robot in-hand manipulation. However, force sensing in VBTSs is underutilized due to the costly and time-intensive process of acquiring paired tactile images and force labels. In this study, we introduce a transferable force prediction model, TransForce, designed to leverage collected image-force paired data for new sensors under varying illumination colors and marker patterns while improving the accuracy of predicted forces, especially in the shear direction. Our model effectively achieves translation of tactile images from the source domain to the target domain, ensuring that the generated tactile images reflect the illumination colors and marker patterns of the new sensors while accurately aligning the elastomer deformation observed in existing sensors, which is beneficial to force prediction of new sensors. As such, a recurrent force prediction model trained with generated sequential tactile images and existing force labels is employed to estimate higher-accuracy forces for new sensors with lowest average errors of 0.69 N (5.8 % in full work range) in$x$-axis, 0.70 N(5.8%) in$y$-axis, and 1.11 N(6.9%) in$z$-axis compared with models trained with single images. The experimental results also reveal that pure marker modality is more helpful than the RGB modality in improving the accuracy of force in the shear direction, while the RGB modality show better performance in the normal direction. Zhuo Chen 0016, Ni Ou, Shan Luo 0001 |
ICRA | 4 |
| 2025 | ConViTac: Aligning Visual-Tactile Fusion with Contrastive RepresentationsabstractVision and touch are two fundamental sensory modalities for robots, offering complementary information that enhances perception and manipulation tasks. Previous research has attempted to jointly learn visual-tactile representations to extract more meaningful information. However, these approaches often rely on direct combination, such as feature addition and concatenation, for modality fusion, which tend to result in poor feature integration. In this paper, we propose ConViTac, a visual-tactile representation learning network designed to enhance the alignment of features during fusion using contrastive representations. Our key contribution is a Contrastive Embedding Conditioning (CEC) mechanism that leverages a contrastive encoder pretrained through self-supervised contrastive learning to project visual and tactile inputs into unified latent embeddings. These embeddings are used to couple visual-tactile feature fusion through cross-modal attention, aiming at aligning the unified representations and enhancing performance on downstream tasks. We conduct extensive experiments to demonstrate the superiority of ConViTac in real world over current state-of-the-art methods and the effectiveness of our proposed CEC mechanism, which improves accuracy by up to 12.0% in material classification and grasping prediction tasks. More details are on our project website. Shan Luo 0001 |
IROS | 3 |
| 2025 | Soft Contact Simulation and Manipulation Learning of Deformable Objects With Vision-Based Tactile SensorabstractDeformable object manipulation is a challenging problem due to its complex deformable properties. With the development of artificial intelligence, learning-based methods have shown outstanding performance in robotic manipulation. Previous works have investigated the manipulation of deformable objects via Reinforcement Learning (RL) in simulation. However, they approximate object deformation with particles, using particle states as observations, which are unavailable in reality. To address these issues, we utilize Vision-Based Tactile Sensors (VBTSs) as the end-effector to manipulate and observe the deformable objects. In this work, we develop a new contact simulation environment for deformable objects, including elastic, plastic, and elastoplastic. We utilize RL strategies and expert demonstrations to train agents in the simulation. Finally, we build a real experimental platform to complete the sim-to-real tasks and robustness testing. Our work introduces an innovative strategy that utilizes high-resolution VBTSs for contact simulation and manipulation of deformable objects. The experimental results show superior performances of deformable object manipulation with the proposed method. Shixin Zhang, Zixi Chen 0002, Zirong Shen, Fuchun Sun 0001, Cesare Stefanini, Di Guo 0002, Shan Luo 0001, Jianwei Zhang 0001, Jianhua Shan, Bin Fang 0003 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | RoTipBot: Robotic Handling of Thin and Flexible Objects Using Rotatable Tactile SensorsabstractThis paper introduces RoTipBot, a novel robotic system for handling thin, flexible objects. Different from previous works that are limited to singulating them using suction cups or soft grippers, RoTipBot can count multiple layers and then grasp them simultaneously in a single grasp closure. Specifically, we first develop a vision-based tactile sensor named RoTip that can rotate and sense contact information around its tip. Equipped with two RoTip sensors, RoTipBot rolls and feeds multiple layers of thin, flexible objects into the centre between its fingers, enabling effective grasping. Moreover, we design a tactile-based grasping strategy that uses RoTip's sensing ability to ensure both fingers maintain secure contact with the object while accurately counting the number of fed objects. Extensive experiments demonstrate the efficacy of the RoTip sensor and the RoTipBot approach. The results show that RoTipBot not only achieves a higher success rate but also grasps and counts multiple layers simultaneously – capabilities not possible with previous methods. Furthermore, RoTipBot operates up to three times faster than state-of-the-art methods. The success of RoTipBot paves the way for future research in object manipulation using mobilised tactile sensors. All the materials used in this paper are available athttps://sites.google.com/view/rotipbot. Daniel Fernandes Gomes, Thanh-Toan Do, Shan Luo 0001 |
IEEE Trans. Robotics | 5 |
| 2025 | Tactile Robotics: An Outlook
Shan Luo 0001, Nathan F. Lepora, Wenzhen Yuan 0001, Kaspar Althoefer, Gordon Cheng, Ravinder S. Dahiya |
IEEE Trans. Robotics | 1 |
| 2025 | Guest EditorialSpecial Collection on Tactile RoboticsabstractTHE sense of touch is an indispensable requirement for humans to effectively interact with the physical world around them and perform dexterous tasks. Similarly, this should be no different for robots. Imagine, for example, a robot that can open a bottle of medicine and dispense pills to an elderly person. Although this might seem a straightforward task for a human, it remains a significant challenge for a robot. Critically, the completion of the task depends on tactile sensing: the robot needs to receive and interpret the feedback from interacting with the bottle, determine the appropriate force based on the size and hardness of the pills, and adjust its pose to safely dispense them. Each step involves contact-rich interactions that can only be effectively deciphered through tactile sensing. Typically, tactile sensing works in conjunction with other modalities, such as vision, enabling the robot to adjust its actions dynamically and complete the task. In response to this vision of robots interacting with the physical world through touch, tactile robotics has now emerged as a key research area. Tactile robots can be defined as intelligent systems equipped with tactile sensors that can extract and process tactile data to guide their operations and interactions. The development of tactile robots presents scientific challenges, ranging from the design and fabrication of tactile sensors to methodologies for processing tactile data, integrating tactile feedback into task execution, and combining it with other sensory modalities to improve robot perception. As a result, tactile robotics demands collaborative efforts across several disciplines, involving material and data scientists … Mark Yim, Shan Luo 0001, Nathan F. Lepora, Wenzhen Yuan 0001, Kaspar Althoefer, Gordon Cheng, Julio Rogelio Guadarrama-Olvera, Ravinder S. Dahiya |
IEEE Trans. Robotics | 2 |
| 2025 | Design and Benchmarking of a Multimodality Sensor for Robotic Manipulation With GAN-Based Cross-Modality InterpretationabstractIn this paper, we present the design and benchmark of an innovative sensor, ViTacTip, which fulfills the demand for advanced multi-modal sensing in a compact design. A notable feature of ViTacTip is its transparent skin, which incorporates a ‘see-through-skin’ mechanism. This mechanism aims at capturing detailed object features upon contact, significantly improving both vision-based and proximity perception capabilities. In parallel, the biomimetic tips embedded in the sensor's skin are designed to amplify contact details, thus substantially augmenting tactile and derived force perception abilities. To demonstrate the multi-modal capabilities of ViTacTip, we developed a multi-task learning model that enables simultaneous recognition of hardness, material, and textures. To assess the functionality and validate the versatility of ViTacTip, we conducted extensive benchmarking experiments, including object recognition, contact point detection, pose regression, and grating identification. To facilitate seamless switching between various sensing modalities, we employed a Generative Adversarial Network (GAN)-based approach. This method enhances the applicability of the ViTacTip sensor across diverse environments by enabling cross-modality interpretation. Dandan Zhang 0001, Wen Fan 0001, Jialin Lin, Haoran Li 0013, Qingzheng Cong, Weiru Liu, Nathan F. Lepora, Shan Luo 0001 |
IEEE Trans. Robotics | 8 |
| 2024 | Adaptive Parametric Activation
Konstantinos Panagiotis Alexandridis, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
ECCV (54) | 4 |
| 2024 | ViTacTip: Design and Verification of a Novel Biomimetic Physical Vision-Tactile Fusion SensorabstractTactile sensing is significant for robotics since it can obtain physical contact information during manipulation. To capture multimodal contact information within a compact framework, we designed a novel sensor called ViTacTip, which seamlessly integrates both tactile and visual perception capabilities into a single, integrated sensor unit. ViTacTip features a transparent skin to capture fine features of objects during contact, which can be known as the see-through-skin mechanism. In the meantime, the biomimetic tips embedded in ViTacTip can amplify touch motions during tactile perception. For comparative analysis, we also fabricated a ViTac sensor devoid of biomimetic tips, as well as a TacTip sensor with opaque skin. Furthermore, we develop a Generative Adversarial Network (GAN)-based approach for modality switching between different perception modes, effectively alternating the emphasis between vision and tactile perception modes. We conducted a performance evaluation of the proposed sensor across three distinct tasks: i) grating identification, ii) pose regression, iii) contact localization and force estimation. In the grating identification task, ViTacTip demonstrated an accuracy of 99.72%, surpassing TacTip, which achieved 94.60%. It also exhibited superior performance in both pose and force estimation tasks with the minimum error of 0.08 mm and 0.03N, respectively, in contrast to ViTac’s 0.12 mm and 0.15N. Results indicate that ViTacTip outperforms single-modality sensors. Wen Fan 0001, Haoran Li 0013, Weiyong Si, Shan Luo 0001, Nathan F. Lepora, Dandan Zhang 0001 |
ICRA | 4 |
| 2024 | Multi-class Road Defect Detection and Segmentation using Spatial and Channel-wise Attention for Autonomous Road RepairingabstractRoad pavement detection and segmentation are critical for developing autonomous road repair systems. However, developing an instance segmentation method that simultaneously performs multi-class defect detection and segmentation is challenging due to the textural simplicity of road pavement image, the diversity of defect geometries, and the morphological ambiguity between classes. We propose a novel end-to-end method for multi-class road defect detection and segmentation. The proposed method comprises multiple spatial and channel-wise attention blocks available to learn global representations across spatial and channel-wise dimensions. Through these attention blocks, more globally generalised representations of morphological information (spatial characteristics) of road defects and colour and depth information of images can be learned. To demonstrate the effectiveness of our framework, we conducted various ablation studies and comparisons with prior methods on a newly collected dataset annotated with nine road defect classes. The experiments show that our proposed method outperforms existing state-of-the-art methods for multi-class road defect detection and segmentation methods. Jongmin Yu, Chen Bene Chi, Sebastiano Fichera, Paolo Paoletti, Devansh Mehta, Shan Luo 0001 |
ICRA | 6 |
| 2024 | Deep Domain Adaptation Regression for Force Calibration of Optical Tactile SensorsabstractOptical tactile sensors provide robots with rich force information for robot grasping in unstructured environments. The fast and accurate calibration of three-dimensional contact forces holds significance for new sensors and existing tactile sensors which may have incurred damage or aging. However, the conventional neural-network-based force calibration method necessitates a large volume of force-labeled tactile images to minimize force prediction errors, with the need for accurate Force/Torque measurement tools as well as a time-consuming data collection process. To address this challenge, we propose a novel deep domain-adaptation force calibration method, designed to transfer the force prediction ability from a calibrated optical tactile sensor to uncalibrated ones with various combinations of domain gaps, including marker presence, illumination condition, and elastomer modulus. Experimental results show the effectiveness of the proposed unsupervised force calibration method, with lowest force prediction errors of 0.102N (3.4% in full force range) for normal force, and 0.095N (6.3%) and 0.062N (4.1%) for shear forces along the x-axis and y-axis, respectively. This study presents a promising, general force calibration methodology for optical tactile sensors. Zhuo Chen 0016, Ni Ou, Shan Luo 0001 |
IROS | 4 |
| 2024 | A Case Study on Visual-Audio-Tactile Cross-Modal RetrievalabstractCross-Modal Retrieval (CMR), which retrieves relevant items from one modality (e.g., audio) given a query in another modality (e.g., visual), has undergone significant advancements in recent years. This capability is crucial for robots to integrate and interpret information across diverse sensory inputs. However, the retrieval space in existing robotic CMR approaches often consists of only one modality, which limits the performance of the robot. In this paper, we propose a novel CMR model that incorporates three different modalities, i.e., visual, audio, and tactile, for enhanced multi-modal object retrieval, referred to as VAT-CMR. In this model, multi-modal representations are first fused to provide a holistic view of object features. Then, to mitigate the semantic gaps between representations of different modalities, a dominant modality is selected during the classification training phase to improve the distinctiveness of the representations and enhance the retrieval performance. To evaluate our proposed approach, we conducted a case study and the results demonstrate that our VAT-CMR model surpasses competing approaches. Further, our proposed dominant modality selection significantly enhances cross-retrieval accuracy. Jagoda Wojcik, Shan Luo 0001 |
IROS | 4 |
| 2024 | TopoFR: A Closer Look at Topology Alignment on Face RecognitionabstractThe field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks has demonstrated the effectiveness of data structure information. Considering that the FR task can leverage large-scale training data, which intrinsically contains significant structure information, we aim to investigate how to encode such critical structure information into the latent space. As revealed from our observations, directly aligning the structure information between the input and latent spaces inevitably suffers from an overfitting problem, leading to a structure collapse phenomenon in the latent space. To address this problem, we propose TopoFR, a novel FR model that leverages a topological structure alignment strategy called PTSA and a hard sample mining strategy named SDE. Concretely, PTSA uses persistent homology to align the topological structures of the input and latent spaces, effectively preserving the structure information and improving the generalization performance of FR model. To mitigate the impact of hard samples on the latent space structure, SDE accurately identifies hard samples by automatically computing structure damage score (SDS) for each sample, and directs the model to prioritize optimizing these samples. Experimental results on popular face benchmarks demonstrate the superiority of our TopoFR over the state-of-the-art methods. Code and models are available at: https://github.com/modelscope/facechain/tree/main/face_module/TopoFR. Jun Dan, Yang Liu 0356, Jiankang Deng, Haoyu Xie 0002, Siyuan Li 0002, Baigui Sun, Shan Luo 0001 |
NeurIPS | 7 |
| 2024 | Road Surface Defect Detection - From Image-Based to Non-Image-Based: A SurveyabstractEnsuring traffic safety is crucial, which necessitates the detection and prevention of road surface defects. As a result, there has been a growing interest in the literature on the subject, leading to the development of various road surface defect detection methods. The methods for detecting road defects can be categorised in various ways depending on the input data types or training methodologies. The predominant approach involves image-based methods, which analyse pixel intensities and surface textures to identify defects. Despite popularity, image-based methods share the distinct limitation of vulnerability to weather and lighting changes. To address this issue, researchers have explored the use of additional sensors, such as laser scanners or LiDARs, providing explicit depth information to enable the detection of defects in terms of scale and volume. However, the exploration of data beyond images has not been sufficiently investigated. In this survey paper, we provide a comprehensive review of road surface defect detection studies, categorising them based on input data types and methodologies used. Additionally, we review recently proposed non-image-based methods and discuss several challenges and open problems associated with these techniques. Jongmin Yu, Sebastiano Fichera, Paolo Paoletti, Lisa Layzell, Devansh Mehta, Shan Luo 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Vis2Hap: Vision-based Haptic Rendering by Cross-modal GenerationabstractTo assist robots in teleoperation tasks, haptic rendering which allows human operators access a virtual touch feeling has been developed in recent years. Most previous haptic rendering methods strongly rely on data collected by tactile sensors. However, tactile data is not widely available for robots due to their limited reachable space and the restrictions of tactile sensors. To eliminate the need for tactile data, in this paper we propose a novel method named as Vis2Hap to generate haptic rendering from visual inputs that can be obtained from a distance without physical interaction. We take the surface texture of objects as key cues to be conveyed to the human operator. To this end, a generative model is designed to simulate the roughness and slipperiness of the object's surface. To embed haptic cues in Vis2Hap, we use height maps from tactile sensors and spectrograms from friction coefficients as the intermediate outputs of the generative model. Once Vis2Hap is trained, it can be used to generate height maps and spectrograms of new surface textures, from which a friction image can be obtained and displayed on a haptic display. The user study demonstrates that our proposed Vis2Hap method enables users to access a realistic haptic feeling similar to that of physical objects. The proposed vision-based haptic rendering has the potential to enhance human operators' perception of the remote environment and facilitate robotic manipulation. Guanqun Cao, Ningtao Mao, Danushka Bollegala, Min Li 0003, Shan Luo 0001 |
ICRA | 6 |
| 2023 | Multi-source Domain Adaptation for Unsupervised Road Defect SegmentationabstractThe performance of road defect segmentation (a.k.a. pixel-level road defect detection) has been improved alongside with remarkable achievement of deep learning. Those improvements need a large-scale and well-constructed dataset. However, road surface materials or designs vary from country to country, and the patterns of defects are hard to pre-define. In this paper, we propose a novel multi-source domain adaptation method to boost the performance of road defect segmentation on an unlabelled dataset. The proposed method generates multi-source ensembled labels using transferred information from models trained with multiple labelled source domains, which are utilised as supervisory signals for the unlabelled target domain. Furthermore, to reduce the domain gap between each source domain and a target domain, these domains are re-aligned with outlier repositioning to improve the defect segmentation performance. We demonstrate the effectiveness of our proposed method on Cracktree200, CRACK500, CFD, and Crack360 datasets. Experimental results show that the proposed method outperforms the existing unsupervised road defect segmentation methods and achieves competitive performance compared with recent supervised methods. The source code is publicly available on https://github.com/andreYoo/MSDA_RDS.git. Jongmin Yu, Hyeontaek Oh, Sebastiano Fichera, Paolo Paoletti, Shan Luo 0001 |
ICRA | 5 |
| 2023 | Learn from Incomplete Tactile Data: Tactile Representation Learning with Masked AutoencodersabstractThe missing signal caused by the objects being occluded or an unstable sensor is a common challenge during data collection. Such missing signals will adversely affect the results obtained from the data, and this issue is observed more frequently in robotic tactile perception. In tactile perception, due to the limited working space and the dynamic environment, the contact between the tactile sensor and the object is frequently insufficient and unstable, which causes the partial loss of signals, thus leading to incomplete tactile data. The tactile data will therefore contain fewer tactile cues with low information density. In this paper, we propose a tactile representation learning method, named TacMAE, based on Masked Autoencoder to address the problem of incomplete tactile data in tactile perception. In our framework, a portion of the tactile image is masked out to simulate the missing contact regions. By reconstructing the missing signals in the tactile image, the trained model can achieve a high-level understanding of surface geometry and tactile properties from limited tactile cues. The experimental results of tactile texture recognition show that TacMAE can achieve a high recognition accuracy of 71.4% in the zero-shot transfer and 85.8% after fine-tuning, which are 15.2% and 8.2% higher than the results without using masked modeling. The extensive experiments on YCB objects demonstrate the knowledge transferability of our proposed method and the potential to improve efficiency in tactile exploration. Guanqun Cao, Danushka Bollegala, Shan Luo 0001 |
IROS | 4 |
| 2023 | Development of a Robot-assisted Virtual Rehabilitation System with Haptic FeedbackabstractIn this paper, the issue of integrating haptic technology of a robotic system and virtual environment is discussed. A robot-assisted virtual rehabilitation system is developed for bilateral training and telerehabilitation. In the system, we designed a training scene in the Unity game engine for these two rehabilitation modes, the proposed framework between two robots and virtual environment can allow the end-effectors to perform tasks as hand avatars in Unity and provide force feedback from virtual environment, the haptic rendering is generated on robot’s end-effector using a task-space impedance controller. Besides providing the feeling of interaction with virtual objects, we proposed a robot-assisted strategy to provide assistance force when the patient is unable to finish the task in virtual training scene, the assistance force can guide the patient to training path. We provide experiment results to demonstrate the performance of proposed system. Yan-Bo Liou, Shan Luo 0001, Yen-Chen Liu |
RO-MAN | 2 |
| 2023 | Inverse Image Frequency for Long-Tailed Image RecognitionabstractThe long-tailed distribution is a common phenomenon in the real world. Extracted large scale image datasets inevitably demonstrate the long-tailed property and models trained with imbalanced data can obtain high performance for the over-represented categories, but struggle for the under-represented categories, leading to biased predictions and performance degradation. To address this challenge, we propose a novel de-biasing method named Inverse Image Frequency (IIF). IIF is a multiplicative margin adjustment transformation of the logits in the classification layer of a convolutional neural network. Our method achieves stronger performance than similar works and it is especially useful for downstream tasks such as long-tailed instance segmentation as it produces fewer false positive detections. Our extensive experiments show that IIF surpasses the state of the art on many long-tailed benchmarks such as ImageNet-LT, CIFAR-LT, Places-LT and LVIS, reaching 55.8% top-1 accuracy with ResNet50 on ImageNet-LT and 26.3% segmentation AP with MaskRCNN ResNet50 on LVIS. Code available at https://github.com/kostas1515/iif. Konstantinos Panagiotis Alexandridis, Shan Luo 0001, Anh Nguyen 0003, Jiankang Deng, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 2 |
| 2022 | Long-Tailed Instance Segmentation Using Gumbel Optimized Loss
Konstantinos Panagiotis Alexandridis, Jiankang Deng, Anh Nguyen 0003, Shan Luo 0001 |
ECCV (10) | 4 |
| 2022 | Reducing Tactile Sim2Real Domain Gaps via Deep Texture Generation NetworksabstractRecently simulation methods have been developed for optical tactile sensors to enable the Sim2Real learning, i.e., first training models in simulation before deploying them on a real robot. However, some artefacts in real objects are unpredictable, such as imperfections caused by fabrication processes, or scratches by natural wear and tear, and thus cannot be represented in the simulation, resulting in a significant gap between the simulated and real tactile images. To address this Sim2Real gap, we propose a novel texture generation network to map the simulated images into photorealistic tactile images that resemble a real sensor contacting a real imperfect object. Each simulated tactile image is first divided into two types of regions: areas that are in contact with the object and areas that are not. The former is applied with generated textures learned from real textures in the real tactile images, whereas the latter maintains its appearance as when the sensor is not in contact with any object. This makes sure that the artefacts are only applied to deformed regions of the sensor. Our extensive experiments show that the proposed texture generation network can generate realistic artefacts on the deformed regions of the sensor, while avoiding leaking the textures into areas of no contact. Quantitative experiments further reveal that when using the adapted images generated by our proposed network for a Sim2Real classification task, the drop in accuracy caused by the Sim2Real gap is reduced from 38.43% to merely 0.81%. As such, this work has potential to accelerate the Sim2Real learning for robotic tasks requiring tactile sensing. Tudor Jianu, Daniel Fernandes Gomes, Shan Luo 0001 |
ICRA | 3 |
| 2022 | Logic Rules Meet Deep Learning: A Novel Approach for Ship Type Classification (Extended Abstract)abstractThe shipping industry is an important component of the global trade and economy. In order to ensure law compliance and safety, it needs to be monitored. In this paper, we present a novel ship type classification model that combines vessel transmitted data from the Automatic Identification System, with vessel imagery. The main components of our approach are the Faster R-CNN Deep Neural Network and a Neuro-Fuzzy system with IF-THEN rules. We evaluate our model using real world data and showcase the advantages of this combination while also compare it with other methods. Results show that our model can increase prediction scores by up to 15.4% when compared with the next best model we considered, while also maintaining a level of explainability as opposed to common black box approaches. Manolis Pitsikalis, Thanh-Toan Do, Alexei Lisitsa 0001, Shan Luo 0001 |
IJCAI | 4 |
| 2022 | Visual-Tactile Multimodality for Following Deformable Linear Objects Using Reinforcement LearningabstractManipulation of deformable objects is a challenging task for a robot. It would be problematic to use a single sensory input to track the behaviour of such objects: vision can be subjected to occlusions, whereas tactile inputs cannot capture the global information that is useful for the task. In this paper, we study the problem of using vision and tactile inputs together to complete the task of following deformable linear objects, for the first time. We create a Reinforcement Learning agent using different sensing modalities and investigate how its behaviour can be boosted using visual-tactile fusion, compared to using a single sensing modality. To this end, we developed a benchmark in simulation for manipulating the deformable linear objects using multimodal sensing inputs. The policy of the agent uses distilled information, e.g., the pose of the object in both visual and tactile perspectives, instead of the raw sensing signals, so that it can be directly transferred to real environments. In this way, we disentangle the perception system and the learned control policy. Our extensive experiments show that the use of both vision and tactile inputs, together with proprioception, allows the agent to complete the task in up to 92% of cases, compared to 77% when only one of the signals is given. Our results can provide valuable insights for the future design of tactile sensors and for deformable objects manipulation. Code and videos can be found at: https://github.com/lpecyna/SoftSlidingGym. Leszek Pecyna, Siyuan Dong, Shan Luo 0001 |
IROS | 3 |
| 2022 | End-to-end weakly supervised semantic segmentation with reliable region mining
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Kaizhu Huang, Shan Luo 0001, Yao Zhao 0001 |
Pattern Recognit. | 5 |
| 2021 | Representation and Processing of Instantaneous and Durative Temporal Phenomena
Manolis Pitsikalis, Alexei Lisitsa 0001, Shan Luo 0001 |
LOPSTR | 3 |
| 2021 | Monoscopic vs. Stereoscopic Views and Display Types in the Teleoperation of Unmanned Ground Vehicles for Object AvoidanceabstractVirtual reality (VR) head-mounted displays (HMD) have recently been used to provide an immersive, first-person vision/view in real-time for manipulating remotely-controlled unmanned ground vehicles (UGV). The teleoperation of UGV can be challenging for operators when it is done in real time. One big challenge is for operators to perceive quickly and rapidly the distance of objects that are around the UGV while it is moving. In this research, we explore the use of monoscopic and stereoscopic views and display types (immersive and non-immersive VR) for operating vehicles remotely. We conducted two user studies to explore their feasibility and advantages. Results show a significantly better performance when using an immersive display with stereoscopic view for dynamic, real-time navigation tasks that require avoiding both moving and static obstacles. The use of stereoscopic view in an immersive display in particular improved user performance and led to better usability. Jialin Wang 0002, Hai-Ning Liang, Shan Luo 0001, Eng Gee Lim |
RO-MAN | 4 |
| 2020 | Spatio-temporal Attention Model for Tactile Texture RecognitionabstractRecently, tactile sensing has attracted great interest in robotics, especially for facilitating exploration of unstructured environments and effective manipulation. A detailed understanding of the surface textures via tactile sensing is essential for many of these tasks. Previous works on texture recognition using camera based tactile sensors have been limited to treating all regions in one tactile image or all samples in one tactile sequence equally, which includes much irrelevant or redundant information. In this paper, we propose a novel Spatio-Temporal Attention Model (STAM) for tactile texture recognition, which is the very first of its kind to our best knowledge. The proposed STAM pays attention to both spatial focus of each single tactile texture and the temporal correlation of a tactile sequence. In the experiments to discriminate 100 different fabric textures, the spatially and temporally selective attention has resulted in a significant improvement of the recognition accuracy, by up to 18.8%, compared to the non-attention based models. Specifically, after introducing noisy data that is collected before the contact happens, our proposed STAM can learn the salient features efficiently and the accuracy can increase by 15.23% on average compared with the CNN based baseline approach. The improved tactile texture perception can be applied to facilitate robot tasks like grasping and manipulation. Guanqun Cao, Yi Zhou 0019, Danushka Bollegala, Shan Luo 0001 |
IROS | 4 |
| 2020 | GelTip: A Finger-shaped Optical Tactile Sensor for Robotic ManipulationabstractSensing contacts throughout the fingers is an essential capability for a robot to perform manipulation tasks in cluttered environments. However, existing tactile sensors either only have a flat sensing surface or a compliant tip with a limited sensing area. In this paper, we propose a novel optical tactile sensor, the GelTip, that is shaped as a finger and can sense contacts on any location of its surface. The sensor captures high-resolution and color-invariant tactile image that can be exploited to extract detailed information about the end-effector's interactions against manipulated objects. Our extensive experiments show that the GelTip sensor can effectively localise the contacts on different locations its finger-shaped body, with a small localisation error of approximately 5 mm, on average, and under 1 mm in the best cases. The obtained results show the potential of the GelTip sensor in facilitating dynamic manipulation tasks with its all-round tactile sensing capability. The sensor models and further information about the GelTip sensor can be found at http://danfergo.github.io/geltip. Daniel Fernandes Gomes, Zhonglin Lin, Shan Luo 0001 |
IROS | 3 |
| 2019 | "Touching to See" and "Seeing to Feel": Robotic Cross-modal Sensory Data Generation for Visual-Tactile PerceptionabstractThe integration of visual-tactile stimulus is common while humans performing daily tasks. In contrast, using unimodal visual or tactile perception limits the perceivable dimensionality of a subject. However, it remains a challenge to integrate the visual and tactile perception to facilitate robotic tasks. In this paper, we propose a novel framework for the cross-modal sensory data generation for visual and tactile perception. Taking texture perception as an example, we apply conditional generative adversarial networks to generate pseudo visual images or tactile outputs from data of the other modality. Extensive experiments on the ViTac dataset of cloth textures show that the proposed method can produce realistic outputs from other sensory inputs. We adopt the structural similarity index to evaluate similarity of the generated output and real data and results show that realistic data have been generated. Classification evaluation has also been performed to show that the inclusion of generated data can improve the perception performance. The proposed framework has potential to expand datasets for classification tasks, generate sensory outputs that are not easy to access, and also advance integrated visual-tactile perception. Jet-Tsyn Lee, Danushka Bollegala, Shan Luo 0001 |
ICRA | 3 |
| 2018 | ViTac: Feature Sharing Between Vision and Tactile Sensing for Cloth Texture RecognitionabstractVision and touch are two of the important sensing modalities for humans and they offer complementary information for sensing the environment. Robots could also benefit from such multi-modal sensing ability. In this paper, addressing for the first time (to the best of our knowledge) texture recognition from tactile images and vision, we propose a new fusion method named Deep Maximum Covariance Analysis (DMCA) to learn a joint latent space for sharing features through vision and tactile sensing. The features of camera images and tactile data acquired from a GelSight sensor are learned by deep neural networks. But the learned features are of a high dimensionality and are redundant due to the differences between the two sensing modalities, which deteriorates the perception performance. To address this, the learned features are paired using maximum covariance analysis. Results of the algorithm on a newly collected dataset of paired visual and tactile data relating to cloth textures show that a good recognition performance of greater than 90% can be achieved by using the proposed DMCA framework. In addition, we find that the perception performance of either vision or tactile sensing can be improved by employing the shared representation space, compared to learning from unimodal data. Shan Luo 0001, Wenzhen Yuan 0001, Edward H. Adelson, Anthony G. Cohn 0001, Raul A. Fuentes-Samaniego |
ICRA | 1 |
| 2017 | Knock-Knock: Acoustic object recognition by using stacked denoising autoencoders
Shan Luo 0001, Leqi Zhu, Kaspar Althoefer, Hongbin Liu 0001 |
Neurocomputing | 1 |
| 2016 | Iterative Closest Labeled Point for tactile object shape recognitionabstractTactile data and kinesthetic cues are two important sensing sources in robot object recognition and are complementary to each other. In this paper, we propose a novel algorithm named Iterative Closest Labeled Point (iCLAP) to recognize objects using both tactile and kinesthetic information. The iCLAP first assigns different local tactile features with distinct label numbers. The label numbers of the tactile features together with their associated 3D positions form a 4D point cloud of the object. In this manner, the two sensing modalities are merged to form a synthesized perception of the touched object. To recognize an object, the partial 4D point cloud obtained from a number of touches iteratively matches with all the reference cloud models to identify the best fit. An extensive evaluation study with 20 real objects shows that our proposed iCLAP approach outperforms those using either of the separate sensing modalities, with a substantial recognition rate improvement of up to 18%. Shan Luo 0001, Wenxuan Mou, Kaspar Althoefer, Hongbin Liu 0001 |
IROS | 1 |
| 2015 | Localizing the object contact through matching tactile features with visual mapabstractThis paper presents a novel framework for integration of vision and tactile sensing by localizing tactile readings in a visual object map. Intuitively, there are some correspondences, e.g., prominent features, between visual and tactile object identification. To apply it in robotics, we propose to localize tactile readings in visual images by sharing same sets of feature descriptors through two sensing modalities. It is then treated as a probabilistic estimation problem solved in a framework of recursive Bayesian filtering. Feature-based measurement model and Gaussian based motion model are thus built. In our tests, a tactile array sensor is utilized to generate tactile images during interaction with objects and the results have proven the feasibility of our proposed framework. Shan Luo 0001, Wenxuan Mou, Kaspar Althoefer, Hongbin Liu 0001 |
ICRA | 1 |
| 2013 | Fiber optics tactile array probe for tissue palpation during minimally invasive surgeryabstractThis paper presents a novel fiber optic tactile probe designed for tissue palpation during minimally invasive surgery (MIS). The probe consists of 3×4 tactile sensing elements at 2.6mm spacing with a dimension of 12×18×8 mm3allowing its application via a 25mm surgical port. Each tactile element converts the applied pressure values into a circular image pattern. The image patterns of all the sensing elements are captured by a camera attached at the proximal end of the sensor system. Processing the intensity and the area of these circular patterns allows the computation of the applied pressure across the sensing array. Validation tests show that each sensing element of the tactile probe can measure forces from 0 to 1N with a resolution of 0.05 N. The proposed sensing concept is low cost, lightweight, sterilizable, easy to be miniaturized and compatible for magnetic resonance (MR) environments. Experiments using the developed sensor for tissue abnormality detection were conducted. Results show that the proposed tactile probe can accurately and effectively detect nodules embedded inside soft tissue, demonstrating the promising application of this probe for surgical palpation during MIS. Hui Xie 0006, Hongbin Liu 0001, Shan Luo 0001, Lakmal D. Seneviratne, Kaspar Althoefer |
IROS | 3 |
| 2013 | Haptics for Multi-fingered PalpationabstractDuring open surgery, surgeons can perceive the locations of tumors inside soft-tissue organs using their fingers. Palpating an organ, surgeons acquire distributed pressure (tactile) information that can be interpreted as stiffness distribution across the organ -an important aid in detecting buried tumors in otherwise healthy tissue. Previous research has focused on haptic systems to feedback the tactile sensation experienced during palpation to the surgeon during minimally invasive. However, the control complexity and high cost of tactile actuators limits its current application. This paper describes a pneumatic multi-fingered haptic feedback system for robot-assisted minimally invasive surgery. It simulates soft tissue stiffness by changing the pressure of an air balloon and recreates the deformation of fingers as experienced during palpation. The pneumatic haptic feedback actuator is validated by using finite element analysis. The results prove that the interaction stress between the fingertip and the soft tissue as well as the deformation of fingertips during palpation can be recreated by using our pneumatic multi-fingered haptic feedback method. Min Li 0003, Shan Luo 0001, Lakmal D. Seneviratne, D. P. Thrishantha Nanayakkara, Kaspar Althoefer, Prokar Dasgupta |
SMC | 2 |