EDBT 2026 Demo / reviewers in the wild / expert
Baoru Huang
dblp:238/1618
· DBLP profile ↗
27ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-4421-652XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 17 since 2021Systems, architecture and hardware · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generative Aspect-Based Sentiment Quadruple Prediction Based on Multi-Order PromptingabstractRecently, generative aspect-level sentiment quadruple prediction (ASQP) methods based on pre-trained language models have made significant progress. However, some challenges remain in extracting and recognizing complex sentiment elements from semantically rich sentences, limiting the generalization and adaptability of unidirectional generative models in aspect-level sentiment analysis. To overcome this limitation, this article proposes a Generative Aspect-Based Sentiment Quadruple Prediction Model based on Multi-Order Prompting (GenMOP). The model draws on the concept of prompt learning and introduces a multi-order prompting strategy, which breaks the traditional framework of a single generative order and enhances the flexibility and adaptability of the model. Furthermore, we integrate a quadruple quantity-aware module and a multi-view uncertainty-aware module based on a basic generative architecture, not only providing the model with more fine-grained information about the quadruple quantity but also improving the prediction accuracy through uncertainty estimation. The extensive experiments show that the GenMOP method achieves excellent performance in the ASQP task. On the four benchmark datasets including Rest15, Rest16, Rest and Lap, our model achieves F1 score improvements of 1.26%, 0.23%, 0.28%, and 2.07%, respectively, compared to existing state-of-the-arts, demonstrating its effectiveness and superiority in dealing with the joint extraction of multiple sentiment elements of the ASQP model. Rui Wang 0077, Muyao He, Yixue Hao, Long Hu, Min Chen 0003, Baoru Huang |
ACM Trans. Inf. Syst. | 7 |
| 2025 | EndoWave: 4D Gaussian Splatting with Rational Wavelet for Endoscopic ReconstructionabstractIn robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric inconsistencies, non-rigid tissue motion, and view-dependent highlights. Most 3DGS-based methods that rely solely on appearance constraints for optimizing 3DGS are often insufficient in this context, as these dynamic visual artifacts can mislead the optimization process and lead to inaccurate reconstructions. To address these limitations, we present EndoWave, a unified spatiotemporal Gaussian Splatting framework by incorporating an optical flow-based geometric constraint and a multi-resolution rational wavelet supervision. First, we adopt a unified spatiotemporal Gaussian representation that directly optimizes primitives in a 4D domain. Second, we propose a geometric constraint derived from optical flow to enhance temporal coherence and effectively constrain the 3D structure of the scene. Third, we propose a multi-resolution rational orthogonal wavelet as a constraint, which can effectively separate the details of the endoscope and enhance the rendering performance. Extensive evaluations on two real surgical datasets, EndoNeRF [1] and StereoMIS [2], demonstrate that our method EndoWave achieves state-of-theart reconstruction quality and visual accuracy compared to the baseline method. Taoyu Wu, Yiyi Miao, Sihang Zhao, Zhuoxiao Li, Baoru Huang, Limin Yu |
BIBM | 8 |
| 2025 | EgoMusic-Driven Human Dance Motion Estimation with Skeleton Mamba
Nhat Le, Baoru Huang, Minh Nhat Vu, Chengcheng Tang, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICCV | 3 |
| 2025 | FedEFM: Federated Endovascular Foundation Model with Unseen DataabstractIn endovascular surgery, the precise identification of catheters and guidewires in X-ray images is essential for reducing intervention risks. However, accurately segmenting catheter and guidewire structures is challenging due to the limited availability of labeled data. Foundation models offer a promising solution by enabling the collection of similar-domain data to train models whose weights can be fine-tuned for downstream tasks. Nonetheless, large-scale data collection for training is constrained by the necessity of maintaining patient privacy. This paper proposes a new method to train a foundation model in a decentralized federated learning setting for endovascular intervention. To ensure the feasibility of the training, we tackle the unseen data issue using differentiable Earth Mover's Distance within a knowledge distillation frame-work. Once trained, our foundation model's weights provide valuable initialization for downstream tasks, thereby enhancing task-specific performance. Intensive experiments show that our approach achieves new state-of-the-art results, contributing to advancements in endovascular intervention and robotic-assisted endovascular surgery, while addressing the critical issue of data sharing in the medical domain. Tuong KL. Do, Nghia Vu, Tudor Jianu, Baoru Huang, Minh Nhat Vu, Jionglong Su, Erman Tjiputra, Quang D. Tran, Te-Chuan Chiu, Anh Nguyen 0003 |
ICRA | 4 |
| 2025 | Robotic-CLIP: Fine-Tuning CLIP on Action Data for Robotic ApplicationsabstractVision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely on static images paired with text prompts and has not yet been fully adapted for robotic tasks involving dynamic actions. In this paper, we introduce Robotic-CLIP to enhance robotic perception capabilities. We first gather and label large-scale action data, and then build our Robotic-CLIP by fine-tuning CLIP on 309,433 videos (≈ 7.4 million frames) of action data using contrastive learning. By leveraging action data, Robotic-CLIP inherits CLIP's strong image performance while gaining the ability to understand actions in robotic contexts. Intensive experiments show that our Robotic-CLIP outperforms other CLIP-based models across various language-driven robotic tasks. Additionally, we demonstrate the practical effectiveness of Robotic-CLIP in real-world grasping applications. Minh Nhat Vu, Tung D. Ta, Baoru Huang, Thieu Vo, T. Hoang Ngan Le, Anh Nguyen 0003 |
ICRA | 4 |
| 2025 | Tracking Everything in Robotic-Assisted SurgeryabstractAccurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and tools. Traditional keypoint-based sparse tracking is limited by featured points, while flow-based dense two-view matching suffers from long-term drifts. Recently, the Tracking Any Point (TAP) algorithm was proposed to overcome these limitations and achieve dense accurate long-term tracking. However, its efficacy in surgical scenarios remains untested, largely due to the lack of a comprehensive surgical tracking dataset for evaluation. To address this gap, we introduce a new annotated surgical tracking dataset for benchmarking tracking methods for surgical scenarios, comprising real-world surgical videos with complex tissue and instrument motions. We extensively evaluate state-of-the-art (SOTA) TAP-based algorithms on this dataset and reveal their limitations in challenging surgical scenarios, including fast instrument motion, severe occlusions, and motion blur, etc. Furthermore, we propose a new tracking method, namely SurgMotion, to solve the challenges and further improve the tracking performance. Our proposed method outperforms most TAP-based algorithms in surgical instruments tracking, and especially demonstrates significant improvements over baselines in challenging medical videos. Our code and dataset are available at https://github.com/zhanbh1019/SurgicalMotion. Bohan Zhan, Yi Fang 0006, Francisco Vasconcelos 0001, Danail Stoyanov, Daniel S. Elson, Baoru Huang |
ICRA | 8 |
| 2025 | Hybrid Deep Reinforcement Learning for Radio Tracer Localisation in Robotic-Assisted Radioguided SurgeryabstractRadioguided surgery, such as sentinel lymph node biopsy, relies on the precise localization of radioactive targets by non-imaging gamma/beta detectors. Manual radioactive target detection based on visual display or audible indication of gamma level is highly dependent on the ability of the surgeon to track and interpret the spatial information. This paper presents a learning-based method to realize the autonomous radiotracer detection in robot-assisted surgeries by navigating the probe to the radioactive target. We proposed novel hybrid approach that combines deep reinforcement learning (DRL) with adaptive robotic scanning. The adaptive grid-based scanning could provide initial direction estimation while the DRL-based agent could efficiently navigate to the target utilising historical data. Simulation experiments demonstrate a 95% success rate, and improved efficiency and robustness compared to conventional techniques. Real-world evaluation on the da Vinci Research Kit (dVRK) further confirms the feasibility of the approach, achieving an 80% success rate in radiotracer detection. This method has the potential to enhance consistency, reduce operator dependency, and improve procedural accuracy in radioguided surgeries. Hanyi Zhang, Kaizhong Deng, Zhaoyang Jacopo Hu, Baoru Huang, Daniel S. Elson |
ICRA | 4 |
| 2025 | SplineFormer: An Explainable Transformer Network for Autonomous Endovascular NavigationabstractRobot-assisted endovascular navigation provides significant advantages, including reduced radiation exposure for surgeons and improved patient safety. However, a major challenge is to control curvilinear instruments like guidewires precisely for smooth and accurate navigation while adapting to anatomical variations and external forces. Traditional segmentation-based approaches struggle with real-time prediction of the guidewire’s evolving shape, limiting their effectiveness in navigation tasks. In this paper, we propose SplineFormer, an explainable transformer network that predicts the continuous, structured representation of the guidewire as a B-spline. This formulation enables a compact, smooth, and explainable state representation that facilitates downstream navigation. By leveraging SplineFormer’s predictions within an imitation learning framework, our system successfully performs autonomous endovascular navigation. Experimental results show that SplineFormer achieves a 50% success rate when fully autonomously cannulating the Brachiocephalic Artery in a real robotic setup, demonstrating its potential for improved autonomous navigation in endovascular interventions. Tudor Jianu, Shayan Doust, Mengyun Li, Baoru Huang, Tuong KL. Do, Hoan Nguyen, Karl Bates, Tung D. Ta, Sebastiano Fichera, Pierre Berthet-Rayne, Anh Nguyen 0003 |
IROS | 4 |
| 2025 | GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent SystemabstractLanguage-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges. First, they often struggle to interpret complex text instructions or operate ineffectively in densely cluttered environments. Second, most methods require a training or fine-tuning step to adapt to new domains, limiting their generation in real-world applications. In this paper, we introduce GraspMAS, a new multi-agent system framework for language-driven grasp detection. GraspMAS is designed to reason through ambiguities and improve decision-making in real-world scenarios. Our framework consists of three specialized agents: Planner, responsible for strategizing complex queries; Coder, which generates and executes source code; and Observer, which evaluates the outcomes and provides feedback. Intensive experiments on two large-scale datasets demonstrate that our GraspMAS significantly outperforms existing baselines. Additionally, robot experiments conducted in both simulation and real-world settings further validate the effectiveness of our approach. Our project page is available at https://zquang2202.github.io/GraspMAS. Thieu Vo, Tung D. Ta, Baoru Huang, Minh Nhat Vu, Anh Nguyen 0003 |
IROS | 6 |
| 2025 | SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
Jialei Chen 0007, Mobarak I. Hoque, Francisco Vasconcelos 0001, Danail Stoyanov, Daniel S. Elson, Baoru Huang |
MICCAI (11) | 7 |
| 2025 | Learning Human Motion with Temporally Conditional MambaabstractLearning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal patterns of conditioning inputs. Existing methods typically rely on cross-attention mechanisms to fuse the condition with motion. However, this approach primarily captures global interactions and struggles to maintain step-by-step temporal alignment. To address this limitation, we introduce Temporally Conditional Mamba, a new mamba-based model for human motion generation. Our approach integrates conditional information into the recurrent dynamics of the Mamba block, enabling better temporally aligned motion. To validate the effectiveness of our method, we evaluate it on a variety of human motion tasks. Extensive experiments demonstrate that our model significantly improves temporal alignment, motion realism, and condition consistency over state-of-the-art approaches. Our project page is available at https://zquang2202.github.io/TCM. Baoru Huang, Minh Nhat Vu, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
SIGGRAPH Asia | 3 |
| 2024 | Guide3D: A Bi-planar X-ray Dataset for 3D Shape Reconstruction
Tudor Jianu, Baoru Huang, Hoan Nguyen, Binod Bhattarai, Tuong KL. Do, Erman Tjiputra, Quang D. Tran, Pierre Berthet-Rayne, T. Hoang Ngan Le, Sebastiano Fichera, Anh Nguyen 0003 |
ACCV (5) | 2 |
| 2024 | Language-driven Grasp DetectionabstractGrasp detection is a persistent and intricate challenge with various industrial applications. Recently, many meth-ods and datasets have been proposed to tackle the grasp detection problem. However, most of them do not consider using natural language as a condition to detect the grasp poses. In this paper, we introduce Grasp-Anything++, a new language-driven grasp detection dataset featuring 1M samples, over 3M objects, and upwards of 10M grasping in-structions. We utilize foundation models to create a large-scale scene corpus with corresponding images and grasp prompts. We approach the language-driven grasp detection task as a conditional generation problem. Drawing on the success of diffusion models in generative tasks and given that language plays a vital role in this task, we propose a new language-driven grasp detection method based on dif-fusion models. Our key contribution is the contrastive training objective, which explicitly contributes to the denoising process to detect the grasp pose given the language instructions. We illustrate that our approach is theoretically sup-portive. The intensive experiments show that our method outperforms state-of-the-art approaches and allows real-world robotic grasping. Finally, we demonstrate our large-scale dataset enables zero-short grasp detection and is a challenging benchmark for future work. Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Thieu Vo, Anh Nguyen 0003 |
CVPR | 3 |
| 2024 | Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ECCV (19) | 3 |
| 2024 | Grasp-Anything: Large-scale Grasp Dataset from Foundation ModelsabstractFoundation models such as ChatGPT have made significant strides in robotic tasks due to their universal representation of real-world domains. In this paper, we leverage foundation models to tackle grasp detection, a persistent challenge in robotics with broad industrial applications. Despite numerous grasp datasets, their object diversity remains limited compared to real-world figures. Fortunately, foundation models possess an extensive repository of real-world knowledge, including objects we encounter in our daily lives. As a consequence, a promising solution to the limited representation in previous grasp datasets is to harness the universal knowledge embedded in these foundation models. We present Grasp-Anything, a new large-scale grasp dataset synthesized from foundation models to implement this solution. Grasp-Anything excels in diversity and magnitude, boasting 1M samples with text descriptions and more than 3M objects, surpassing prior datasets. Empirically, we show that Grasp-Anything successfully facilitates zero-shot grasp detection on vision-based tasks and real-world robotic experiments. Our dataset and code are available at https://airvlab.github.io/grasp-anything/. Vuong Dinh An, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Andreas Kugi, Anh Nguyen 0003 |
ICRA | 4 |
| 2024 | Language-Conditioned Affordance-Pose Detection in 3D Point CloudsabstractAffordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding affordance task. Previous methods for affodance-pose joint learning are limited to a predefined set of affordances, thus limiting the adaptability of robots in real-world environments. In this paper, we propose a new method for language-conditioned affordance-pose joint learning in 3D point clouds. Given a 3D point cloud object, our method detects the affordance region and generates appropriate 6-DoF poses for any unconstrained affordance label. Our method consists of an open-vocabulary affordance detection branch and a language-guided diffusion model that generates 6-DoF poses based on the affordance text. We also introduce a new high-quality dataset for the task of language-driven affordance-pose joint learning. Intensive experimental results demonstrate that our proposed method works effectively on a wide range of open-vocabulary affordances and outperforms other baselines by a large margin. In addition, we illustrate the usefulness of our method in real-world robotic applications. Our code and dataset are publicly available at https://3DAPNet.github.io. Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Tuan Van Vo, Vy Truong, T. Hoang Ngan Le, Thieu Vo, Bac Le, Anh Nguyen 0003 |
ICRA | 3 |
| 2024 | Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point CorrelationabstractAffordance detection presents intricate challenges and has a wide range of robotic applications. Previous works have faced limitations such as the complexities of 3D object shapes, the wide range of potential affordances on real-world objects, and the lack of open-vocabulary support for affordance understanding. In this paper, we introduce a new open-vocabulary affordance detection method in 3D point clouds, leveraging knowledge distillation and text-point correlation. Our approach employs pre-trained 3D models through knowledge distillation to enhance feature extraction and semantic understanding in 3D point clouds. We further introduce a new text-point correlation method to learn the semantic links between point cloud features and open-vocabulary labels. The intensive experiments show that our approach outperforms previous works and adapts to new affordance labels and unseen objects. Notably, our method achieves the improvement of 7.96% mIOU score compared to the baselines. Furthermore, it offers real-time inference which is well-suitable for robotic manipulation applications. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, Toan Nguyen 0004, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
ICRA | 3 |
| 2024 | Lightweight Language-driven Grasp Detection using Conditional Consistency ModelabstractLanguage-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. This work presents a new approach for language-driven grasp detection that leverages lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability. Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 3 |
| 2024 | Language-driven Grasp Detection with Mask-guided AttentionabstractGrasp detection is an essential task in robotics with various industrial applications. However, traditional methods often struggle with occlusions and do not utilize language for grasping. Incorporating natural language into grasp detection remains a challenging task and largely unexplored. To address this gap, we propose a new method for language-driven grasp detection with mask-guided attention by utilizing the transformer attention mechanism with semantic segmentation features. Our approach integrates visual data, segmentation mask features, and natural language instructions, significantly improving grasp detection accuracy. Our work introduces a new framework for language-driven grasp detection, paving the way for language-driven robotic applications. Intensive experiments show that our method outperforms other recent baselines by a clear margin, with a 10.0% success score improvement. We further validate our method in real-world robotic experiments, confirming the effectiveness of our approach. Tuan Van Vo, Minh Nhat Vu, Baoru Huang, An Vuong, T. Hoang Ngan Le, Thieu Vo, Anh Nguyen 0003 |
IROS | 3 |
| 2024 | HabiCrowd: A High Performance Simulator for Crowd-Aware Visual NavigationabstractVisual navigation, a foundational aspect of Embodied AI (E-AI) and robotics has been extensively studied in the past few years. While many 3D simulators have been introduced for the visual navigation tasks, scarcely works have combined human dynamics, creating the gap between simulation and real-world applications. Furthermore, current 3D simulators incorporating human dynamics have several limitations, particularly in terms of computational efficiency, which is a promise of modern simulators. To overcome these issues, we introduce HabiCrowd, the new standard benchmark for crowd-aware visual navigation that includes a crowd dynamics model with diverse human settings into photorealistic environments. Empirical evaluations demonstrate that our proposed human dynamics model achieves state-of-the-art performance in collision avoidance while exhibiting superior computational efficiency compared to its counterparts. We leverage HabiCrowd to conduct several comprehensive studies on crowd-aware visual navigation tasks and human-robot interactions. The source code and data can be found at https://habicrowd.github.io/. An Vuong, Toan Nguyen 0004, Minh Nhat Vu, Baoru Huang, Huynh Thi Thanh Binh, Thieu Vo, Anh Nguyen 0003 |
IROS | 4 |
| 2024 | Spatial-temporal graph feature learning driven by time-frequency similarity assessment for robust fault diagnosis of rotating machinery
Fuchen Xie, Baoru Huang |
Adv. Eng. Informatics | 5 |
| 2024 | Residual Aligner-based Network (RAN): Motion-separable structure for coarse-to-fine discontinuous deformable registrationabstractDeformable image registration, the estimation of the spatial transformation between different images, is an important task in medical imaging. Deep learning techniques have been shown to perform 3D image registration efficiently. However, current registration strategies often only focus on the deformation smoothness, which leads to the ignorance of complicated motion patterns (e.g., separate or sliding motions), especially for the intersection of organs. Thus, the performance when dealing with the discontinuous motions of multiple nearby objects is limited, causing undesired predictive outcomes in clinical usage, such as misidentification and mislocalization of lesions or other abnormalities. Consequently, we proposed a novel registration method to address this issue: a new Motion Separable backbone is exploited to capture the separate motion, with a theoretical analysis of the upper bound of the motions' discontinuity provided. In addition, a novel Residual Aligner module was used to disentangle and refine the predicted motions across the multiple neighboring objects/organs. We evaluate our method, Residual Aligner-based Network (RAN), on abdominal Computed Tomography (CT) scans and it has shown to achieve one of the most accurate unsupervised inter-subject registration for the 9 organs, with the highest-ranked registration of the veins (Dice Similarity Coefficient (%)/Average surface distance (mm): 62%/4.9mm for the vena cava and 34%/7.9mm for the portal and splenic vein), with a smaller model structure and less computation compared to state-of-the-art methods. Furthermore, when applied to lung CT, the RAN achieves comparable results to the best-ranked networks (94%/3.0mm), also with fewer parameters and less computation. Jian-Qing Zheng, Baoru Huang, Ngee Han Lim, Bartlomiej Wladyslaw Papiez |
Medical Image Anal. | 3 |
| 2023 | Detecting the Sensing Area of a Laparoscopic Probe in Minimally Invasive Cancer Surgery
Baoru Huang, Anh Nguyen 0003, Stamatia Giannarou, Daniel S. Elson |
MICCAI (9) | 1 |
| 2023 | Language-driven Scene Synthesis using Multi-conditional Diffusion ModelabstractScene synthesis is a challenging problem with several industrial applications. Recently, substantial efforts have been directed to synthesize the scene using human motions, room layouts, or spatial graphs as the input. However, few studies have addressed this problem from multiple modalities, especially combining text prompts. In this paper, we propose a language-driven scene synthesis task, which is a new task that integrates text prompts, human motion, and existing objects for scene synthesis. Unlike other single-condition synthesis tasks, our problem involves multiple conditions and requires a strategy for processing and encoding them into a unified space. To address the challenge, we present a multi-conditional diffusion model, which differs from the implicit unification approach of other diffusion literature by explicitly predicting the guiding points for the original data distribution. We demonstrate that our approach is theoretically supportive. The intensive experiment results illustrate that our method outperforms state-of-the-art benchmarks and enables natural scene editing applications. The source code and dataset can be accessed at https://lang-scene-synth.github.io/. Vuong Dinh An, Minh Nhat Vu, Toan Nguyen 0004, Baoru Huang, Dzung Nguyen, Thieu Vo, Anh Nguyen 0003 |
NeurIPS | 4 |
| 2022 | Self-supervised Depth Estimation in Laparoscopic Image Using 3D Geometric Consistency
Baoru Huang, Jian-Qing Zheng, Anh Nguyen 0003, Ioannis Gkouzionis, Kunal Vyas, David Tuch, Stamatia Giannarou, Daniel S. Elson |
MICCAI (8) | 1 |
| 2021 | Self-supervised Generative Adversarial Network for Depth Estimation in Laparoscopic Images
Baoru Huang, Jian-Qing Zheng, Anh Nguyen 0003, David Tuch, Kunal Vyas, Stamatia Giannarou, Daniel S. Elson |
MICCAI (4) | 1 |
| 2021 | Dual-arm Coordinated Manipulation for Object Twisting with Human IntelligenceabstractRobotic dual-arm twisting is a common but very challenging task in both industrial production and daily services, as it often requires dexterous collaboration, a large scale of end-effector rotating, and good adaptivity for object manipulation. Meanwhile, safety and efficiency are primary concerns for robotic dual-arm coordinated manipulation. Thus, the normally adopted fully automated task execution approaches based on environmental perception and motion planning techniques are still inadequate and problematic for the arduous twisting tasks. To this end, this paper presents a novel strategy of the dual-arm coordinated control for twisting manipulation based on the combination of optimized motion planning for one arm and real-time telecontrol with human intelligence for the other. The analysis and simulation results showed it can achieve collision and singularity free for dual arms with enhanced dexterity, safety, and efficiency. Weibang Bai, Ningshan Zhang, Baoru Huang, Ziwei Wang 0001, Francesco Cursi, Ya-Yen Tsai, Bo Xiao 0002, Eric M. Yeatman |
SMC | 3 |