EDBT 2026 Demo / reviewers in the wild / expert
Tianwei Zhang 0002
dblp:77/7902-2
· DBLP profile ↗
14ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-1462-5402ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 5 since 2021Systems, architecture and hardware · 9 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human-Robot Cooperative Heavy Payload Manipulation based on Whole-Body Model Predictive ControlabstractHuman-robot collaborative manipulation with mobile, multiple manipulators is crucial for expanding robotic applications, requiring precise handling of coupled force-position constraints between partners. Current systems, however, exhibit end-effector oscillations and instability during dynamic interactions. To overcome these limitations, this work develops a collaborative framework integrating a collaborative controller and a whole-body controller. The collaborative controller employs the object’s center-of-mass dynamics model with real-time contact forces and motion states to predict trajectories while coordinating with an attitude stabilization controller to adjust the desired end-effector poses. The whole-body controller utilizes model predictive control to generate coordinated motions that strictly follow pose commands from the collaborative controller, ensuring stable transportation. Simulation and physical experiments validate the proposed framework’s effectiveness in real-world scenarios. Tin Lun Lam, Tianwei Zhang 0002 |
IROS | 4 |
| 2024 | Vision-Language Model-based Physical Reasoning for Robot Liquid PerceptionabstractThere is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an open challenge as they can only process textual input. Recent advancements in large vision-language models (LVLMs) have enabled a more comprehensive understanding of the physical world by incorporating visual input, which provides richer contextual information than language alone. In this work, we proposed a novel paradigm that leveraged GPT-4V(ision), the state-of-the-art LVLM by OpenAI, to enable embodied agents to perceive liquid objects via image-based environmental feedback. Specifically, we exploited the physical understanding of GPT-4V to interpret the visual representation (e.g., time-series plot) of non-visual feedback (e.g., F/T sensor data), indirectly enabling multimodal perception beyond vision and language using images as proxies. We evaluated our method using 10 common household liquids with containers of various geometry and material. Without any training or fine-tuning, we demonstrated that our method can enable the robot to indirectly perceive the physical response of liquids and estimate their viscosity. We also showed that by jointly reasoning over the visual and physical attributes learned through interactions, our method could recognize liquid objects in the absence of strong visual cues (e.g., container labels with legible text or symbols), increasing the accuracy from 69.0%—achieved by the best-performing vision-only variant—to 86.0%. Wenqiang Lai, Tianwei Zhang 0002, Tin Lun Lam, Yuan Gao 0024 |
IROS | 2 |
| 2023 | Mouth Cavity Visual Analysis Based on Deep Learning for Oropharyngeal Swab Robot SamplingabstractThe visual analysis of the mouth cavity plays a significant role in the pathogen specimen sampling and disease diagnosis of the mouth cavity. Aiming at performance defects of general detectors based on deep learning in detecting mouth cavity components, this article proposes a mouth cavity analysis network (MCNet), which is an instance segmentation method with spatial features, and a mouth cavity dataset (MCData), which is the first available dataset for mouth cavity detecting and segmentation. First, given the lack of a mouth cavity image dataset, the MCData for detecting and segmenting key parts in the mouth cavity was developed for model training and testing. Second, the MCNet was designed based on the mask region-based convolutional neural network. To improve the performance of feature extraction, a parallel multiattention module was designed. Besides, to solve low detection accuracy of small-sized objects, a multiscale region proposal network structure was designed. Then, the mouth cavity spatial structure features were introduced, and the detection confidence could be refined to increase the detection accuracy. The MCNet achieved 81.5% detection accuracy and 78.1% segmentation accuracy (intersection over union = 0.50:0.95) on the MCData. Comparative experiments with the MCData showed that the proposed MCNet outperformed state-of-the-art approaches with the task of mouth cavity instance segmentation. In addition, the MCNet has been used in an oropharyngeal swab robot for COVID-19 oropharyngeal sampling. Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Tianwei Zhang 0002, Yuquan Leng |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | Fast and Comfortable Interactive Robot-to-Human Object HandoverabstractTransferring tools and objects to human hands is an important ability of collaborative robots. Most of the existing approaches focus on handover affordance, however, the comfort of receiving objects with human hands is often neglected. In this paper, we use advanced deep learning models to pre-generate handover target configurations that are convenient for human grasping based on the characteristics of the objects and tools, and then the robot grasps and passes the objects to the human. Experimental results on a mobile collaborative robot show that our proposed framework can robustly and efficiently deliver different shapes and types of objects to a human hand of any pose within the robot's field of view in a target pose that is convenient for grasping and can quickly deliver objects to a new target location even after the human hand moves to a new position. Chongxi Meng, Tianwei Zhang 0002, Tin Lun Lam |
IROS | 2 |
| 2021 | AcousticFusion: Fusing Sound Source Localization to Visual SLAM in Dynamic EnvironmentsabstractDynamic objects in the environment, such as people and other agents, lead to challenges for existing simultaneous localization and mapping (SLAM) approaches. To deal with dynamic environments, computer vision researchers usually apply some learning-based object detectors to remove these dynamic objects. However, these object detectors are computationally too expensive for mobile robot on-board processing. In practical applications, these objects output noisy sounds that can be effectively detected by on-board sound source localization. The directional information of the sound source object can be efficiently obtained by direction of sound arrival (DoA) estimation, but the depth estimation is difficult. Therefore, in this paper, we propose a novel audio-visual fusion approach that fuses sound source direction into the RGB-D image and thus removes the effect of dynamic obstacles on the multi-robot SLAM system. Experimental results of multirobot SLAM in different dynamic environments show that the proposed method uses very small computational resources to obtain very stable self-localization results. Tianwei Zhang 0002, Huayan Zhang, Xiaofei Li 0001, Tin Lun Lam, Sethu Vijayakumar |
IROS | 1 |
| 2021 | PoseFusion2: Simultaneous Background Reconstruction and Human Shape Recovery in Real-timeabstractDynamic environments that include unstructured moving objects pose a hard problem for Simultaneous Localization and Mapping (SLAM) performance. The motion of rigid objects can be typically tracked by exploiting their texture and geometric features. However, humans moving in the scene are often one of the most important, interactive targets – they are very hard to track and reconstruct robustly due to non-rigid shapes. In this work, we present a fast, learning-based human object detector to isolate the dynamic human objects and realise a real-time dense background reconstruction framework. We go further by estimating and reconstructing the human pose and shape. The final output environment maps not only provide the dense static backgrounds but also contain the dynamic human meshes and their trajectories. Our Dynamic SLAM system runs at around 26 frames per second (fps) on GPUs, while additionally turning on accurate human pose estimation can be executed at up to 10 fps. Huayan Zhang, Tianwei Zhang 0002, Tin Lun Lam, Sethu Vijayakumar |
IROS | 2 |
| 2020 | Learning to Optimize Non-Rigid TrackingabstractOne of the widespread solutions for non-rigid tracking has a nested-loop structure: with Gauss-Newton to minimize a tracking objective in the outer loop, and Preconditioned Conjugate Gradient (PCG) to solve a sparse linear system in the inner loop. In this paper, we employ learnable optimizations to improve tracking robustness and speed up solver convergence. First, we upgrade the tracking objective by integrating an alignment data term on deep features which are learned end-to-end through CNN. The new tracking objective can capture the global deformation which helps Gauss-Newton to jump over local minimum, leading to robust tracking on large non-rigid motions. Second, we bridge the gap between the preconditioning technique and learning method by introducing a ConditionNet which is trained to generate a preconditioner such that PCG can converge within a small number of steps. Experimental results indicate that the proposed learning method converges faster than the original PCG by a large margin. Yang Li 0143, Aljaz Bozic, Tianwei Zhang 0002, Yanli Ji, Tatsuya Harada, Matthias Nießner |
CVPR | 3 |
| 2020 | FlowFusion: Dynamic Dense RGB-D SLAM Based on Optical FlowabstractDynamic environments are challenging for visual SLAM since the moving objects occlude the static environment features and lead to wrong camera motion estimation. In this paper, we present a novel dense RGB-D SLAM solution that simultaneously accomplishes the dynamic/static segmentation and camera ego-motion estimation as well as the static background reconstructions. Our novelty is using optical flow residuals to highlight the dynamic semantics in the RGB-D point clouds and provide more accurate and efficient dynamic/static segmentation for camera tracking and background reconstruction. The dense reconstruction results on public datasets and real dynamic scenes indicate that the proposed approach achieved accurate and efficient performances in both dynamic and static environments compared to state-of-the-art approaches. Tianwei Zhang 0002, Huayan Zhang, Yang Li 0143, Yoshihiko Nakamura, Lei Zhang 0079 |
ICRA | 1 |
| 2020 | SplitFusion: Simultaneous Tracking and Mapping for Non-Rigid ScenesabstractWe present SplitFusion, a novel dense RGB-D SLAM framework that simultaneously performs tracking and dense reconstruction for both rigid and non-rigid components of the scene. SplitFusion first adopts deep learning based semantic instant segmentation technique to split the scene into rigid or non-rigid surfaces. The split surfaces are independently tracked via rigid or non-rigid ICP and reconstructed through incremental depth map fusion. Experimental results show that the proposed approach can provide not only accurate environment maps but also well-reconstructed non-rigid targets, e.g., the moving humans. Yang Li 0143, Tianwei Zhang 0002, Yoshihiko Nakamura, Tatsuya Harada |
IROS | 2 |
| 2020 | Gait planning and control method for humanoid robot using improved target positioning
Lei Zhang 0079, Huayan Zhang, Tianwei Zhang 0002, Guibin Bian |
Sci. China Inf. Sci. | 4 |
| 2016 | Human activity prediction by mapping grouplets to recurrent Self-Organizing Map
Qianru Sun, Hong Liu 0008, Tianwei Zhang 0002 |
Neurocomputing | 4 |
| 2016 | A novel hierarchical Bag-of-Words model for compact action representation
Qianru Sun, Hong Liu 0008, Liqian Ma, Tianwei Zhang 0002 |
Neurocomputing | 4 |
| 2012 | Hierarchical RRT for humanoid robot footstep planning with multiple constraints in complex environmentsabstractHumanoid robots have abilities of stepping over or onto obstacles, which is different from wheeled robots. However it may be difficult to apply the ordinary motion planning methods such as Rapidly-exploring Random Trees (RRT) to humanoid robots directly. Because these kinds of methods only consider to circumvent obstacles and ignore the constraint of balance. Aiming at dealing with these problems in one frame, a novel approach based on hierarchical RRT is used to plan the footstep for humanoid robots. It is designed according to three basic constraints: a transition model based gait generator, an inverted pendulum based balance controller and a collision detection based path planner. First, a set of layered transition model is utilized to revise the footstep according to the terrain condition, which is able to take full use of the motion ability as well as improve the efficiency. Then, a hierarchical strategy is exploited to select the feasible foot location to be added in the random tree based on the results of collision checking and balance control. Finally, a dynamic RRT method is introduced in our work to revise paths in changing environments. Different experiments are given to verify the feasibility and performance of the proposed approach in complicated environments with both dynamic and static obstacles. Hong Liu 0008, Tianwei Zhang 0002 |
IROS | 3 |
| 2012 | A "capacitor" bridge builder based safe path planner for difficult regions identification in changing environmentsabstractFinding paths in difficult regions of C-space, such as narrow passages and configuration obstacle boundaries, is a rather challenging problem for path planning in changing environments. When obstacles move in W-space, these regions in C-space will change their edge points from free to collision or on the contrary, for which a “Capacitor” Bridge Builder (CBB) is proposed in this paper to identify their changing characteristics. Specifically, a “Capacitor” bridge is built between positive and negative toggled points in C-Space, which looks like capacitors stuck between narrow passages or boundary regions. Through CBB, the back boundary of an obstacle, which is less likely to be occupied immediately, is marked as a temporary safe region. Furthermore, a Half Bridge Strategy (HBS) is novelly proposed to boost samples inside these regions. Eventually, highly safe paths are revealed by predicting moving directions of obstacles, then replanning times and total planning times will be decreased significantly. Effectiveness of the proposed method has been verified by experiments with two manipulators in difficult changing environments. Hong Liu 0008, Tianwei Zhang 0002, Chuangqi Wang |
IROS | 2 |