Zhengtao Zhang

dblp:37/8914 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Multistep Intent Estimation Guided Adaptive Passive Control for Safety-Aware Physical Human-Robot Collaboration
abstract
physical human-robot collaboration (pHRC) requires strict safety and efficiency guarantees, imposing heightened demands on accurate human intent estimation and adaptive control in a stable manner. To address these challenges, we propose a novel two-loop adaptive passive control framework guided by multistep human intent estimation to reduce human-robot disagreement and improve robot assistance level, facilitating safety-aware efficient pHRC. In the framework, outer loop's intent estimation guides the inner loop's adaptive passive controller, ensuring real-time robot behavior adjustment based on multistep intention. Specifically, the outer loop incorporates a transformer-based human intent estimator (THIE) that integrates the Transformer with a conditional variational autoencoder (CVAE) for multistep predictions, accurately estimating motion and force to guide the robot. The inner loop incorporates a goal-oriented reinforcement learning (GoRL)-based adaptive impedance control, which constructs multistep rewards based on prediction and probability from THIE to adjust impedance parameters, thereby balancing disagreement and assistance, and promoting locally optimal robot behaviors. Furthermore, an energy tank-based passive model predictive control (ET-PMPC) is employed to limit robot stored energy, avoiding the impact of variable impedance on safety. Experiments validate that our framework outperforms state-of-the-art (SOTA) methods, significantly improving intent estimation accuracy, robot assistance level, and safety, highlighting its potential to advance pHRC.
Zhengtao Zhang, Yuchuang Tong, Zhaojie Ju
IEEE Trans. Cybern.2
2026 SdaPS*: A Novel Source-Free Domain Adaption Method for Point Cloud Primitive Segmentation
Shaohu Wang, Yuchuang Tong, Rongtao Xu, Zhengtao Zhang
IEEE Trans. Ind. Informatics4
2025 Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection
abstract
Recently, vision-language models (e.g. CLIP) have demonstrated remarkable performance in zero-shot anomaly detection (ZSAD). By leveraging auxiliary data during training, these models can directly perform cross-category anomaly detection on target datasets, such as detecting defects on industrial product surfaces or identifying tumors in organ tissues. Existing approaches typically construct text prompts through either manual design or the optimization of learnable prompt vectors. However, these methods face several challenges: 1) handcrafted prompts require extensive expert knowledge and trial-and-error; 2) single-form learnable prompts struggle to capture complex anomaly semantics; and 3) an unconstrained prompt space limits generalization to unseen categories. To address these issues, we propose Bayesian Prompt Flow Learning (Bayes-PFL), which models the prompt space as a learnable probability distribution from a Bayesian perspective. Specifically, a prompt flow module is designed to learn both image-specific and image-agnostic distributions, which are jointly utilized to regularize the text prompt space and improve the model's generalization on unseen categories. These learned distributions are then sampled to generate diverse text prompts, effectively covering the prompt space. Additionally, a residual cross-model attention (RCA) module is introduced to better align dynamic text embeddings with fine-grained image features. Extensive experiments on 15 industrial and medical datasets demonstrate our method's superior performance. The code is available at https://github.com/xiaozhen228/Bayes-PFL.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Qiyu Chen 0002, Zhengtao Zhang, Guiguang Ding
CVPR6
2025 DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
abstract
Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Fei Shen 0002, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding
ICCV8
2025 DTRT: Enhancing Human Intent Estimation and Role Allocation for Physical Human-Robot Collaboration
abstract
In physical Human-Robot Collaboration (pHRC), accurate human intent estimation and rational human-robot role allocation are crucial for safe and efficient assistance. Existing methods that rely on short-term motion data for intention estimation lack multi-step prediction capabilities, hindering their ability to sense intent changes and adjust human-robot assignments autonomously, resulting in potential discrepancies. To address these issues, we propose a Dual Transformer-based Robot Trajectron (DTRT) featuring a hierarchical architecture, which harnesses human-guided motion and force data to rapidly capture human intent changes, enabling accurate trajectory predictions and dynamic robot behavior adjustments for effective collaboration. Specifically, human intent estimation in DTRT uses two Transformer-based Conditional Variational Autoencoders (CVAEs), incorporating robot motion data in obstacle-free case with human-guided trajectory and force for obstacle avoidance. Additionally, Differential Cooperative Game Theory (DCGT) is employed to synthesize predictions based on human-applied forces, ensuring robot behavior align with human intention. Compared to state-of-the-art (SOTA) methods, DTRT incorporates human dynamics into long-term prediction, providing an accurate understanding of intention and enabling rational role allocation, achieving robot autonomy and maneuverability. Experiments demonstrate DTRT's accurate intent estimation and superior collaboration performance.
Yuchuang Tong, Zhengtao Zhang
ICRA3
2025 Fine-Grained Region Perception Network for Few-Shot Defect Classification of IC Package Substrates: Benchmark Methodology and Dataset
abstract
As the core of the modern electronics industry, integrated circuits (IC) involve highly complex design and manufacturing processes, with the design and fabrication of the package substrates particularly impacting the circuit’s performance and reliability. Therefore, defect detection and classification of integrated circuits package substrates (ICPS) are crucial in IC production. Addressing issues such as the scarcity of data and the challenges in data perception for ICPS, we propose a Fine-grained Region Perception Network (FRPNet) to achieve multi-view perception and precise few-shot classification of ICPS. Specifically, FRPNet consists of three modules: the Category-Perceptive Interaction Module, responsible for feature aggregation perception during class simulation changes; the Fine-Grained Region Aggregation Module, which observes the regions of interest from multiple views and ensures intra-class connectivity; and the Localization Refinement Module, which enhances positional information to ensure the stability of features from local to global scales. Additionally, we construct a CPS2D-FSC dataset comprising single-layer and multi-layer ICPS. We conducted extensive experiments in CPS2D-FSC to validate FRPNet, including comparisons with SOTA algorithms and ablation studies, demonstrating the superiority of our algorithm and the effectiveness of each module.
Haoyuan Li 0003, Ruiyun Yu, Bingyang Guo, Zhengtao Zhang
IECON4
2025 Teacher Motion Priors: Enhancing Robot Locomotion over Challenging Terrain
abstract
Achieving robust locomotion on complex terrains remains a challenge due to high-dimensional control and environmental uncertainties. This paper introduces a teacher-prior framework based on the teacher-student paradigm, integrating imitation and auxiliary task learning to improve learning efficiency and generalization. Unlike traditional paradigms that strongly rely on encoder-based state embeddings, our framework decouples the network design, simplifying the policy network and deployment. A high-performance teacher policy is first trained using privileged information to acquire generalizable motion skills. The teacher’s motion distribution is transferred to the student policy, which relies only on noisy proprioceptive data, via a generative adversarial mechanism to mitigate performance degradation caused by distributional shifts. Additionally, auxiliary task learning enhances the student policy’s feature representation, speeding up convergence and improving adaptability to varying terrains. The framework is validated on a humanoid robot, showing a great improvement in locomotion stability on dynamic terrains and significant reductions in development costs. This work provides a practical solution for deploying robust locomotion strategies in humanoid robots.
Fangcheng Jin, Peixin Ma, En Li 0001, Zhengtao Zhang
IROS7
2025 IDAGC: Adaptive Generalized Human-Robot Collaboration via Human Intent Estimation and Multimodal Policy Learning
abstract
In Human-Robot Collaboration (HRC), which encompasses physical interaction and remote cooperation, accurate estimation of human intentions and seamless switching of collaboration modes to adjust robot behavior remain paramount challenges. To address these issues, we propose an Intent-Driven Adaptive Generalized Collaboration (IDAGC) framework that leverages multimodal data and human intent estimation to facilitate adaptive policy learning across multi-tasks in diverse scenarios, thereby facilitating autonomous inference of collaboration modes and dynamic adjustment of robotic actions. This framework overcomes the limitations of existing HRC methods, which are typically restricted to a single collaboration mode and lack the capacity to identify and transition between diverse states. Central to our framework is a predictive model that captures the interdependencies among vision, language, force, and robot state data to accurately recognize human intentions with a Conditional Variational Autoencoder (CVAE) and automatically switch collaboration modes. By employing dedicated encoders for each modality and integrating extracted features through a Transformer decoder, the framework efficiently learns multi-task policies, while force data optimizes compliance control and intent estimation accuracy during physical interactions. Experiments highlights our framework’s practical potential to advance the comprehensive development of HRC.
Yuchuang Tong, Guanchen Liu, Zhaojie Ju, Zhengtao Zhang
IROS5
2025 Multi-component controllable diversified augmentation of industrial images based on feature disentanglement
Qingfeng Shi, Fei Shen 0002, Zhengtao Zhang
Eng. Appl. Artif. Intell.4
2025 Adaptive few-shot image augmentation for fine-grained industrial defects based on region-level modeling
Qingfeng Shi, Zhengtao Zhang, Fei Shen 0002, Huiyuan Luo
Eng. Appl. Artif. Intell.3
2025 A surface defect detection instrument for large aperture spherical optical elements
Yali Shi, Zhengtao Zhang, Xian Tao, Xiuqin Shang
Neural Comput. Appl.3
2025 Progressive Boundary Guided Anomaly Synthesis for Industrial Anomaly Detection
abstract
Unsupervised anomaly detection methods can identify surface defects in industrial images by leveraging only normal samples for training. Due to the risk of overfitting when learning from a single class, anomaly synthesis strategies are introduced to enhance detection capability by generating artificial anomalies. However, existing strategies heavily rely on anomalous textures from auxiliary datasets. Moreover, their limitations in the coverage and directionality of anomaly synthesis may result in a failure to capture useful information and lead to significant redundancy. To address these issues, we propose a novel Progressive Boundary-guided Anomaly Synthesis (PBAS) strategy, which can directionally synthesize crucial feature-level anomalies without auxiliary textures. It consists of three core components: Approximate Boundary Learning (ABL), Anomaly Feature Synthesis (AFS), and Refined Boundary Optimization (RBO). To make the distribution of normal samples more compact, ABL first learns an approximate decision boundary by center constraint, which improves the center initialization through feature alignment. AFS then directionally synthesizes anomalies with more flexible scales guided by the hypersphere distribution of normal features. Since the boundary is so loose that it may contain real anomalies, RBO refines the decision boundary through the binary classification of artificial anomalies and normal features. Experimental results show that our method achieves state-of-the-art performance and the fastest detection speed on three widely used industrial datasets, including MVTec AD, VisA, and MPDD. The code will be available at:https://github.com/cqylunlun/PBAS.
Qiyu Chen 0002, Huiyuan Luo, Han Gao 0008, Chengkan Lv, Zhengtao Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2025 Center-Aware Residual Anomaly Synthesis for Multiclass Industrial Anomaly Detection
abstract
Anomaly detection plays a vital role in the inspection of industrial images. Most existing methods require separate models for each category, resulting in multiplied deployment costs. This highlights the challenge of developing a unified model for multiclass anomaly detection. However, the significant increase in interclass interference leads to severe missed detections. Furthermore, the intraclass overlap between normal and abnormal samples, particularly in synthesis-based methods, cannot be ignored and may lead to over-detection. To tackle these issues, we propose a novel center-aware residual anomaly synthesis (CRAS) method for multiclass anomaly detection. CRAS leverages center-aware residual learning to couple samples from different categories into a unified center, mitigating the effects of interclass interference. To further reduce intraclass overlap, CRAS introduces distance-guided anomaly synthesis that adaptively adjusts noise variance based on normal data distribution. Experimental results on diverse datasets and real-world industrial applications demonstrate the superior detection accuracy and competitive inference speed of CRAS.
Qiyu Chen 0002, Huiyuan Luo, Haiming Yao, Zhen Qu, Chengkan Lv, Zhengtao Zhang
IEEE Trans. Ind. Informatics7
2025 OV-BIS: Open-Vocabulary Boundary Guide Zero-Shot 3D Instance Segmentation
abstract
Open vocabulary 3D instance segmentation aims to align 3D instance segmentation results with natural language text, thereby achieving semantic prediction without relying on predefined class labels for specific scenes, which has been widely used in the field of multimedia. Current open vocabulary 3D instance segmentation methods mainly rely on 2D masks provided by various 2D segmentation foundation models. However, in complex scenes, the calculation of 2D masks often struggles to balance over-segmentation of large objects and under-segmentation of small objects. In this paper, we introduce OV-BIS, a novel zero-shot open vocabulary 3D instance segmentation method that leverages instance boundary information to improve 3D semantic segmentation performance. The key insight of our method is that the edge map as 3D boundary projection is suitable for multi-scale tasks and capable of compensating for the weakness of 2D masks in multi-scale adaptability for complex scenes. Our method aggregates multiview edge maps and 2D masks, iteratively guiding the merging of over-segmented point clouds with regions growing to cluster 3D primitives into distinct 3D instances. By projecting 3D instances onto images and using CLIP to calculate semantic features from multiple perspectives with an outliers filter, 3D semantic instance segmentation has been achieved. Experiments on multiple datasets demonstrate the superiority of our method.
Tinghao Yi, Shaohu Wang, Zhengtao Zhang, Changwei Wang 0001, Dong-Ming Yan 0001, Rongtao Xu, Enhong Chen
IEEE Trans. Multim.3
2025 Human-Inspired Adaptive Optimal Control Framework for Robot-Environment Interaction
abstract
Enabling robots with uncertain dynamics to perform human-like adaptive operations in unknown environments remains a significant challenge in robotics research. Drawing inspiration from the dynamic modification of human arm muscles, we propose an innovative adaptive optimal control framework to address this issue. The framework integrates a variable optimal impedance adaptation (VOIA) method and an adaptive bias broad fuzzy neural network (ABBFNN) controller, facilitating adaptive manipulation behaviors in robot-environment interaction tasks. It can adaptively learn the impedance gain of unknown environment in the presence of uncertain robot dynamic model based on different task properties, simultaneously keeping the tracking error and interaction force optimized and minimized. The ABBFNN controller combines adaptive node increments to approximate uncertain dynamic model and introduces additional global bias and adaptive gain adjustment to improve the approximation accuracy and the rate of convergence significantly. VOIA seamlessly integrates a finely tuned proportional-integral-derivative (PID) variable target stiffness and an impact compensator, ensuring accurate responses to varying environmental conditions and improved disturbance rejection. Moreover, a momentum-based force observer is utilized within the framework for interaction force estimation, eliminating the need for force sensors and simplifying the system. Simulations and experiments validate the effectiveness and practicality of the proposed optimal interaction control framework, demonstrating its potential to propel robots toward interactions with unknown environments.
Yuchuang Tong, Zhengtao Zhang
IEEE Trans. Syst. Man Cybern. Syst.3
2024 A Unified Anomaly Synthesis Strategy with Gradient Ascent for Industrial Anomaly Detection and Localization
Qiyu Chen 0002, Huiyuan Luo, Chengkan Lv, Zhengtao Zhang
ECCV (67)4
2024 VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation
Zhen Qu, Xian Tao, Mukesh Prasad, Fei Shen 0002, Zhengtao Zhang, Xinyi Gong, Guiguang Ding
ECCV (69)5
2024 Few-Shot Defect Image Generation Based on Consistency Modeling
Qingfeng Shi, Fei Shen 0002, Zhengtao Zhang
ECCV (76)4
2024 Orthogonal Latent Compression for Streaming Anomaly Detection in Industrial Vision
Han Gao 0008, Huiyuan Luo, Fei Shen 0002, Zhengtao Zhang
ICPR (9)4
2024 Multi-Confidence Guided Source-Free Domain Adaption Method for Point Cloud Primitive Segmentation
abstract
Point cloud primitive segmentation aims to segment the surface point cloud into various geometric types of primitives, which plays a vital role in robot operation and industrial automation. However, differences in object structures and shapes across industrial datasets create domain shift issues, compounded by privacy concerns preventing dataset sharing. To address these challenges, we propose a novel source-free domain adaptation method for point cloud primitive segmentation, which follows the popular pseudo-label based self-training framework. Unlike previous works using single-model uncertainty to refine pseudo labels, our method leverages multi-confidence, including transformation consistency, task confidence, and geometric saliency to provide more informative guidance. Specifically, the transformation consistency is first utilized to vote pseudo-labels and task confidences. Furthermore, to filter out high-confident noises and obtain more reliable pseudo-labels, we investigate the geometric curvature properties of primitives and propose a geometric saliency guided dynamic prototype matching and label graph aggregation strategies for pseudo-label reassignment with different task confidence. For this novel task, we construct several datasets and verify the effectiveness of the proposed methods through a series of experiments.
Shaohu Wang, Yuchuang Tong, Xiuqin Shang, Zhengtao Zhang
ICRA4
2024 ALMRR: Anomaly Localization Mamba on Industrial Textured Surface with Feature Reconstruction and Refinement
Shichen Qu, Xian Tao, Zhen Qu, Xinyi Gong, Zhengtao Zhang, Mukesh Prasad
PRCV (9)5
2024 RGR-Net: Refined Graph Reasoning Network for multi-height hotspot defect detection in photovoltaic farms
Shenshen Zhao, Haiyong Chen, Yatong Zhou, Zhengtao Zhang
Expert Syst. Appl.5
2024 Produce Once, Utilize Twice for Anomaly Detection
abstract
Visual anomaly detection aims at classifying and locating the regions that deviate from the normal appearance. Embedding-based methods and reconstruction-based methods are two main approaches for this task. The embedding-based methods typically predict the anomaly by measuring the distances between the deep representations of the test samples and a limited number of nominal samples, which enables these methods to be efficient but struggle in providing a fine-grained pixel-level anomaly location. The reconstruction-based methods rely on the pixel-level reconstruction errors to locate the anomaly, thereby the anomaly predictions are fine-grained. However, there are repetitive feature extractions and usually extra modules to guarantee the quality of the reconstructed images, resulting in unsatisfactory detection efficiency. In a nutshell, the prior methods are either not efficient or not precise enough for the industrial detection. To deal with this problem, we derive POUTA (Produce Once Utilize Twice for Anomaly detection), which improves both the accuracy and efficiency by reusing the discriminant information potential in the reconstructive network. We observe that the encoder and decoder representations of the reconstructive network are able to stand for the features of the original and reconstructed image respectively. And the discrepancies between the symmetric reconstructive representations provides roughly accurate anomaly information. To refine this information, a coarse-to-fine process is proposed in POUTA, which calibrates the semantics of each discriminative layer by the high-level representations and supervision loss. Equipped with the above modules, POUTA is endowed with the ability to provide a more precise anomaly location than the prior arts. Besides, the representation reusage also enables to exclude the feature extraction process in the discriminative network, which reduces the parameters and improves the efficiency. Extensive experiments show that, POUTA is superior or comparable to the prior methods with even less cost. Furthermore, POUTA also achieves better performance than the state-of-the-art few-shot anomaly detection methods without any special design, showing that POUTA has strong ability to learn representations inherent in the training data.
Shuyuan Wang, Qi Li 0005, Huiyuan Luo, Chengkan Lv, Zhengtao Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Prediction of MiRNA-Disease Association Based on Higher-Order Graph Convolutional Networks
Zhengtao Zhang, Pengyong Han, Zhengwei Li 0001, Ru Nie
ICIC (2)1
2021 CADN: A weakly supervised learning-based category-aware object detection network for surface defect detection
Jiabin Zhang, Hu Su, Xinyi Gong, Zhengtao Zhang, Fei Shen 0002
Pattern Recognit.5
2021 Quality Inspection Based on Quadrangular Object Detection for Deep Aperture Component
abstract
This article focuses on automatic inspection for the commonly used component called spring-wire socket. An automatic inspection system is built that adopts an endoscope to improve the imaging quality. To detect the low contrast targets in complex background, we adopt the pipeline of Faster R-CNN but with several improvements. The improved network specifies the targets in the form of quadrangular bounding box as opposed to previous methods that specify them by rectangular bounding box. With the quadrangular representation, additional shape and pose information is provided and unexpected overlaps and background disturbance could be avoided. In the network, an eight-dimensional (8-D) vector is designed to represent the quadrangular bounding box followed by the improved anchor mechanism in the region proposal network. And also, a novel overlap score calculation method is proposed. On the basis of the detection result, rules arisen from expertise are provided to determine the quality of the component in which way automatic quality inspection could be accomplished. The superiority of our detection network over existing ones is sufficiently demonstrated with the detection result of irregular targets in images captured in the industrial scenario. Meanwhile, successful inspection result proves that the system meets the industrial requirements in terms of both accuracy and speed and thus is of practical significance to industrial applications.
Jiabin Zhang, Zhengtao Zhang, Hu Su, Xinyi Gong, Feng Zhang 0006
IEEE Trans. Syst. Man Cybern. Syst.2
2017 Partially Decoupled Image-Based Visual Servoing Using Different Sensitive Features
abstract
A new image-based visual servoing method based on sensitive features is presented to separately realize the position control and orientation control. Line features are used for the orientation control because of their sensitivities to rotational motions. Point features and area size features are employed to realize the position control since area size features are very sensitive to the objects' depths. The translations resulting from rotational motions are introduced into the position control as the compensation in order to eliminate the influence of the camera's motions on the point features. The depths for all active features are estimated via interaction matrices, features variations, and the executed camera motions. The proposed method can keep the tracked objects in the camera's field of view in the visual servoing process. In addition, the determination methods of the interaction matrices for point, line, and area size features are proposed. Comparing to the traditional method, the proposed determination method of the interaction matrix for line is independent from the parameters of the plane containing the line. Experimental results verify the effectiveness of the proposed methods.
De Xu, Jinyan Lu, Peng Wang 0024, Zhengtao Zhang, Zi-ze Liang
IEEE Trans. Syst. Man Cybern. Syst.4
2016 High Precision Automatic Assembly Based on Microscopic Vision and Force Information
abstract
An automatic system is developed to realize high precision assembly of two components in the size of mm level with an interference fit in 3-dimensional (3-D) space with 6-degree-of- freedoms (DOF), which consists of a manipulator, an adjusting platform, a sensing system and a computer. The manipulator is employed to align component B to the component A in position. The adjusting platform aligns the component A to component B in orientations and inserts A into B. The sensing system includes three microscopes and a force sensor. The three microscopes are mounted approximately orthogonal to observe components from different directions in the aligning stage. The force sensor is introduced to detect the contact force in assembly process. In the aligning stage, a pose control method based on image Jacobian matrix is proposed. In the insertion stage, a position control method based on the contact force is proposed. The calibration of image Jacobian matrix is also presented. Experimental results demonstrate the effectiveness of the proposed system and methods.
Song Liu 0003, De Xu, Zhengtao Zhang
IEEE Trans Autom. Sci. Eng.4
2010 Control system design for a 5-DOF table tennis robot
abstract
A visual control system is designed for a table tennis robot with five degrees of Freedom (DOFs). It consists of four parts such as ball sensing, trajectory predicting, motion planning, and motion control. A highvelocity stereo vision system with parallel architecture is developed to sense the motions of table tennis ball. The striking parameters including position, velocity, and time are predicted according to the predicted trajectory of the ball based on several measured positions. The motion computer receives the predicted striking parameters and performs motion planning for the robot. A motion control card embedded in the motion computer receives the planning results and controls the motions of the robot via the servo drivers for X and Y axes. A microprocessor is designed to produce pulses to control the motions of the rest three axes via the drivers. Experiments are well conducted to verify the effectiveness of the developed robot and control system.
De Xu, Zhengtao Zhang
ICARCV4