VLDB 2026 Research / reviewers in the wild / expert
Hu Su
dblp:89/7698
· DBLP profile ↗
19ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-0551-3193ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Systems, architecture and hardware · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shape-aware defect detection network
Hu Su, Hongxuan Ma, Xingkun Li, Ningbo Cheng |
Pattern Recognit. | 2 |
| 2025 | A Light-Weight Framework for Open-Set Object Detection with Decoupled Feature Alignment in Joint SpaceabstractOpen-set object detection (OSOD) is highly desirable for robotic manipulation in unstructured environments. However, existing OSOD methods often fail to meet the requirements of robotic applications due to their high computational burden and complex deployment. To address this issue, this paper proposes a light-weight framework called Decoupled OSOD (DOSOD), which is a practical and highly efficient solution to support real-time OSOD tasks in robotic systems. Specifically, DOSOD builds upon the YOLO-World pipeline by integrating a vision-language model (VLM) with a detector. A Multilayer Perceptron (MLP) adaptor is developed to transform text embeddings extracted by the VLM into a joint space, within which the detector learns the region representations of classagnostic proposals. Cross-modality features are directly aligned in the joint space, avoiding the complex feature interactions and thereby improving computational efficiency. DOSOD operates like a traditional closed-set detector during the testing phase, effectively bridging the gap between closed-set and openset detection. Compared to the baseline YOLO-World, the proposed DOSOD significantly enhances real-time performance while maintaining comparable accuracy. The slight DOSODS model achieves a Fixed AP of 26.7 %, compared to 26.2 % for YOLO-World-v1-S and 22.7 % for YOLO-World-v2-S, using similar backbones on the LVIS minival dataset. Meanwhile, the FPS of DOSOD-S is 57.1 % higher than YOLO-World-v1S and 29.6 % higher than YOLO-World-v2-S. Meanwhile, we demonstrate that the DOSOD model facilitates the deployment of edge devices. The codes and models are publicly available at https://github.com/D-Robotics-AI-Lab/DOSOD. Yonghao He, Hu Su, Haiyong Yu, Wei Sui |
ICRA | 2 |
| 2025 | IoU-Aware Clustering for Anchor Configuration Determination in Efficient Defect DetectionabstractDeep-learning-based object detection has gained widespread application in surface defect inspection, with anchor-based detectors achieving remarkable success by utilizing dense anchors to align with defects. Determining the optimal anchor configuration, i.e., sizes and aspect ratios of anchor boxes, remains a critical challenge, particularly when addressing defects with significant shape variations. While previous studies have predominantly focused on developing more efficient network architectures and learning strategies, the problem of anchor configuration determination has not been thoroughly explored. To address this gap, this paper proposes the IoU-Aware Clustering (IAC) algorithm, which autonomously learns suitable anchor configurations by extracting shape priors from diverse defects. IAC takes the training bounding boxes as potential clustering centers and selects a subset that aligns with the shape distribution of the training samples. The algorithm involves only a single hyper-parameter, the anchor number k, making it highly adaptable to various scenarios. Experimental results demonstrate that IAC can effectively generate anchor configurations tailored to defect shapes, significantly improving the mean Average Precision (mAP) by 6.9% and 14.4% on two industrial defect datasets with substantial shape variations. Hongxuan Ma, Hu Su, Song Liu 0003 |
IROS | 5 |
| 2025 | A review of deep-learning-based super-resolution: From methods to applications
Hu Su |
Pattern Recognit. | 1 |
| 2025 | Template Matching-Based Nanoscale Visual Tracking for Out-of-Plane Rotations Inside SEMabstractVisual tracking is crucial in nanomanipulation inside scanning electron microscopy (SEM), especially for complex 3D manipulation tasks. However, tracking the micro- and nanoscale objects and manipulators under different rotational angles, especially out-of-plane rotation, remains challenging due to significant changes in their appearance in the image space. In this paper, we propose a template matching-based method for nanoscale tracking, particularly addressing challenges from out-of-plane rotation. By leveraging the image Jacobian matrix, we establish the relationships between image and Cartesian coordinates, enabling dynamic generation of templates for specified rotation angles. Then, a visual tracking pipeline is proposed, consisting of an offline preparation stage and an online tracking stage. Based on the proposed template generation method, the pipeline dynamically generates appropriate templates as rotational angles vary and performs accurate tracking using template matching. Further, the templates can be conveniently generated for specified magnifications using the image Jacobian matrix, enabling adaptation to changes in SEM magnification. Extensive experiments, including tracking under various SEM magnifications, complex trajectories, and different types of end-effectors, are conducted to demonstrate the effectiveness of the proposed method. Comparisons with widely used tracking approaches further highlight its superiority in both tracking accuracy and real-time performance. Finally, the successful deployment of the proposed method in a real nanomanipulation task confirms its practical applicability. Ying Li 0062, Yanqin Ma, Hu Su, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Vision-Based Closed-Loop Control With Spatiotemporal Multiplexing Strategy for Noncontact Trapping of Multiple Micro-ParticlesabstractNoncontact trapping of micro objects has great application potential in fields like material science and biomedical engineering due to its label-freeness and biocompatibility. In this paper, an automated acoustic micro-particle trapping system implemented with phased transducer array (PTA) is prototyped. The system is incorporated with a stereo vision to provide visual feedback benefited from localization of the invisible acoustic field through hydrophone scanning. Binocular vision calibration and stereo matching are realized using image Jacobian matrix. An efficient phase modulation algorithm is proposed for the calculation of desired PTA phase profile in real-time and a spatiotemporal multiplexing control strategy is adopted to dynamically generate multiple trappings. Experimental results well demonstrated that the stable trapping of multiple particles can be robustly realized by the system, leading to the improvements of robotic noncontact manipulation with invisible acoustic end-effector. Note to Practitioners—This paper is motivated by the problem that previous classic acoustic trapping was achieved as a physical phenomenon that particles within the trapping zone would be automatically trapped and thus required people to place the particle into the invisible trapping zone, which is neither precision nor efficient. Such problem is a crucial factor that limits acoustic tweezer to be further readily usable in bioengineering, surface manufacturing, and quantitative micromechanical characterization. In this work, automated acoustic trapping is presented in the context of robotics, as grasping task in conventional industrial robots, that can generate the acoustic trap exactly in the location where particles are detected (by microscopic vision, or micro-CT or acoustic imaging, etc.). This paper proposes a full pipeline to automatically trap multiple particles using ultrasonic transducer array and binocular microscopic vision. The experiments verified the ability of proposed method in simultaneously trapping three micro particles with opposite acoustic properties. Such trapping method is the foundational technology for further acoustic manipulation such as arraying and sorting, which will be the tasks in our future work. Jiaqi Li 0029, Chengxi Zhong, Teng Li 0017, Zhenhuan Sun, Youfu Li 0001, Hu Su, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2024 | Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source LocalizationabstractAudio-Visual Source Localization (AVSL) is the task of identifying specific sounding objects in the scene given audio cues. In our work, we focus on semi-supervised AVSL with pseudo-labeling. To address the issues with vanilla hard pseudo-labels including bias accumulation, noise sensitivity, and instability, we propose a novel method named Cross Pseudo-Labeling (XPL), wherein two models learn from each other with the cross-refine mechanism to avoid bias accumulation. We equip XPL with two effective components. Firstly, the soft pseudo-labels with sharpening and pseudolabel exponential moving average mechanisms enable models to achieve gradual self-improvement and ensure stable training. Secondly, the curriculum data selection module adaptively selects pseudo-labels with high quality during training to mitigate potential bias. Experimental results demonstrate that XPL significantly outperforms existing methods, achieving state-of-the-art performance while effectively mitigating confirmation bias and ensuring training stability. Shijie Ma, Hu Su |
ICASSP | 4 |
| 2024 | NanoNeRF: Robot-assisted Nanoscale 360° reconstruction with neural radiance field under scanning electron microscopeabstractThe pursuit of 3D reconstruction from 2D images for nanomanipulation under scanning electron microscopy stands as a critical research endeavor. Previous methods either necessitates additional lighting which is difficult in standard SEM devices or relies on feature matching with low resolution and precision, further constraining reconstruction performance. In this paper, we propose a novel robot-assisted nanoscale 360° reconstruction approach, which simplifies SEM setups and maximizes the utilization of robot motion and feedback. By harnessing a nanorobotic system, we capture 360°multi-view images automatically with precise mapping information and camera postures. Sequentially, neural radiance field reconstruct the pixel-wise structure and synthesizing images from diverse perspectives. Experimental results using two real datasets demonstrates our approach’s efficacy, achieving PSNR of 28.1 and SSIM of 0.93 for nanotube reconstruction, and PSNR of 32.8 and SSIM of 0.98 for AFM cantilever reconstruction. These results validate the reliability and robustness of our proposed robot-assisted reconstruction method. Haojian Lu, Jiaqi Li 0029, Youfu Li 0001, Hu Su, Song Liu 0003 |
IROS | 7 |
| 2024 | Binary Amplitude-Only Hologram Generation for Acoustic End-Effector Design by Physics-based deep learningabstractAcoustic holography has emerged as a cutting-edge technique for constructing a micro-robot acoustic end-effector for non-contact manipulation. As one of typical implementations of acoustic holography, Binary Amplitude-Only Hologram (BAOH) featured with a simple structure provides an efficient alternative for modulating acoustic fields that support micro-robotic manipulation. In the present study, we propose a deep learning based BAOH generation method for constructing precise and high-resolution end-effector based on acoustic field. Specifically, we model the BAOH generation problem into an optimization framework. The framework combines an acoustic wave propagation model with the deep neural network, in favor of bypassing the laborious collection of labeled data and facilitating the model to learn the inverse mapping. Additionally, to address the issues of gradient invalidation and information loss caused by binarization, the framework uses an adaptive binarization layer consisting of differentiable binarization and adaptive threshold automatically learned during training, which facilitates to realize end-to-end optimization and increase the non-linear capacity of the model. The simulation experiments show that the proposed method is capable to predict BAOH that supports precise, robust, versatile and real-time construction of acoustic end-effector, enjoying broad prospects in various applications related to micro-robotic manipulation. Qing Liu 0025, Hu Su, Jiaqi Li 0029, Youfu Li 0001, Song Liu 0003 |
IROS | 2 |
| 2024 | Real-Time Acoustic Holography With Physics-Based Deep Learning for Robotic ManipulationabstractAcoustic holography (AH) is a promising technique for precise noncontact micro-nano robotic manipulation. It encodes a three-dimensional (3D) acoustic field acting as a virtual end-effector into a two-dimensional (2D) hologram, whereby the desired acoustic field reconstruction is made possible. Most traditional methods to implement AH, such as 3D printed holographic lens and phased array of transducers (PAT), have limitations of dynamic and dexterous manipulation. Furthermore, existing iterative optimization algorithms to calculate 2D holograms have inadequate accuracy and real-time performance. To address these issues, this paper proposes a physics-based deep learning method with a novel training framework for phase-only hologram (POH) calculation enabling further pushing forward the PAT-based AH for noncontact robotic manipulation. By implementing independent control of each channel on PAT referring real-time calculated POH by a well-trained network, the desired acoustic field can be reconstructed in real-time with high fidelity. The results both on a simulated dataset and a real dataset demonstrate that our method supports accurate and dynamic reconstruction of desired acoustic field with distinct morphologies, with an average reconstruction error of 0.085 and average POH computing time of 47 milliseconds on GPU. Indeed, this work shows the future potential of AH in the field of noninvasive medical therapy, exogenous material delivery, and miniaturized industrial assembly.Note to Practitioners—This paper addresses the challenge of noncontact micro-nano robotic manipulation by PAT-based AH, an intriguing technique in bioengineering, micro-assembly, and material characterization. However, existing approaches have limited precision and real-time performance. To overcome these limitations, this paper proposes a physics-based deep learning method with a novel training framework. Our method achieves excellent accuracy and real-time performance, enabling efficient reconstruction of various complicated acoustic field morphologies for precise and dynamic acoustic manipulation. Experimental results demonstrate its high manipulation flexibility due to the independent modulation of each channel of PAT and real-time precise control due to the ultrafast calculation of the proposed deep learning method, though the method has not yet been deployed into an acoustic manipulation system and tested in practice. Future research will focus on designing physical experiments for further evaluation. Overall, the proposed method provides a novel and promising basis for desired acoustic field generation. Chengxi Zhong, Jiaqi Li 0029, Zhenhuan Sun, Teng Li 0017, Yao Guo 0002, David C. Jeong, Hu Su, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2023 | Noncontact Particle Manipulation on Water Surface with Ultrasonic Phased Array System and Microscopic VisionabstractNoncontact particle manipulation (NPM) shows great application potential than its conventional counterpart particularly in terms of non-invasiveness, and thus has significantly extended robotic manipulation capacity into bio- medical engineering, material science, etc. As NPM by means of electric, magnetic, and optical field has successfully demonstrated powerful strength in both academia and industry, NPM boosted by acoustic field, however, still faces staggering challenges. It is indeed in the very recent years that controllable dynamic airborne or waterborne acoustic field modulation technology emerged in academia. In this paper, we report our latest research regarding dexterous and dynamic noncontact micro-particle manipulation on water surface effected by acoustic field in terms of automated trapping, closed-loop positioning, and real-time motion planning, which can be applied to scenarios such as parallel 3D printing, cell assembly, etc. The main contribution of this work is we demonstrated the feasibility of objective-oriented and fully automated acoustic manipulation of micro-particle in precision scale based on robotic approach in 2D plane. Experiment results showed that the repetitive positioning accuracy can reach as high as 16 μm, which is essentially the pixel scale factor. Yexin Zhang, Jiaqi Li 0029, Yuyu Jia, Teng Li 0017, Yang Wang 0063, David C. Jeong, Hu Su, Song Liu 0003 |
ICRA | 7 |
| 2023 | Real-time Acoustic Holography with Iterative Unsupervised Learning for Acoustic Robotic ManipulationabstractPhase-only acoustic holography is a fundamental and promising technique for contactless robotic manipulation. Through independently controlling phase-only hologram (POH) of phase array of transducers (PAT) and simultaneously driving each channel by sophisticated circuits, a certain acoustic field is dynamically generated in working medium (e.g., air, water or biological tissues) at certain moment. The phase profile of PAT is required dynamically and precisely as per arbitrary expected acoustic field for the sake of versatile and stable robotic manipulation. However, the most conventional methods rely on iterative optimization algorithms which are inevitably time-consuming and probably non-convergent, moreover hindering versatility and fidelity of acoustic robotic manipulation. To address these issues, this paper reports a real-time phase-only acoustic holography algorithm by virtue of iterative unsupervised learning. Using a physics model to construct two queues, which we refer to as experience pools, data pairs consisting of a target acoustic amplitude hologram in expected acoustic field and corresponding POH of PAT are collected on-the-fly, circumventing costly preparation of annotated dataset in advance. With iterative learning between neural network training and experience pools update, both the solution of objective inverse mapping and the adaptation for arbitrary desired acoustic field are mutually enhanced. The experiments and results validated that the proposed approach surpasses previous algorithms in terms of real time and precision. Chengxi Zhong, Zhenhuan Sun, Teng Li 0017, Hu Su, Song Liu 0003 |
ICRA | 4 |
| 2023 | Ultrafast Acoustic Holography with Physics-Reinforced Self-Supervised Learning for Precise Robotic ManipulationabstractUltrafast acoustic holography (AH) enabling dynamic contactless micro-nano robotic manipulation has recently attracted wide attention. As an advanced technique, AH encodes specific three-dimensional (3D) acoustic field on a two-dimensional (2D) hologram whereby realizing holographic reconstruction with high fidelity. However, current approaches face the limitation of encoding time, accuracy and flexibility, thus, leading to inapplicability for dynamic and precise robotic manipulation. Here, we develop an approach to overcome these issues. Its basic idea is to use a convolutional neural network trained in a self-supervised manner with iterative interaction with virtual physical environment. Energy conservation is incorporated to access the physical constrain during wave propagation. The experimental results demonstrate that the proposed method circumvents laborious annotated dataset preparation and boosts the reinforcement from physics model. By the validation and comparison on distinct acoustic fields with various patterns, the accuracy and real-time performance of the proposed method are confirmed supporting dynamic and precise robotic manipulation. Qingyi Lu, Chengxi Zhong, Qing Liu 0025, Teng Li 0017, Hu Su, Song Liu 0003 |
IROS | 5 |
| 2023 | Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source LocalizationabstractAudio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in effectively utilizing the abundance of unlabeled audio-visual pairs. In this paper, we propose a novel Semi-Supervised Learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The optimal utilization of both labeled and unlabeled data combined with this unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of $\textbf{90.4\%}$ and $\textbf{48.8\%}$ on Flickr-SoundNet and VGG-Sound Source, obtaining $\textbf{8.9\%}$ and $\textbf{9.6\%}$ improvements respectively, given only $3\%$ of data positional-annotated. We also extend our framework to some existing AVSL methods and consistently boost their performance. Our code is publicly available at https://github.com/gyx-gloria/DMT. Shijie Ma, Hu Su, Siyang Sun |
NeurIPS | 3 |
| 2023 | Weakly Supervised Instance Segmentation via Category-aware Centerness Learning with Localization Supervision
Jiabin Zhang, Hu Su, Yonghao He |
Pattern Recognit. | 2 |
| 2022 | DSLA: Dynamic smooth label assignment for efficient anchor-free object detection
Hu Su, Yonghao He, Jiabin Zhang |
Pattern Recognit. | 1 |
| 2021 | CADN: A weakly supervised learning-based category-aware object detection network for surface defect detection
Jiabin Zhang, Hu Su, Xinyi Gong, Zhengtao Zhang, Fei Shen 0002 |
Pattern Recognit. | 2 |
| 2021 | Quality Inspection Based on Quadrangular Object Detection for Deep Aperture ComponentabstractThis article focuses on automatic inspection for the commonly used component called spring-wire socket. An automatic inspection system is built that adopts an endoscope to improve the imaging quality. To detect the low contrast targets in complex background, we adopt the pipeline of Faster R-CNN but with several improvements. The improved network specifies the targets in the form of quadrangular bounding box as opposed to previous methods that specify them by rectangular bounding box. With the quadrangular representation, additional shape and pose information is provided and unexpected overlaps and background disturbance could be avoided. In the network, an eight-dimensional (8-D) vector is designed to represent the quadrangular bounding box followed by the improved anchor mechanism in the region proposal network. And also, a novel overlap score calculation method is proposed. On the basis of the detection result, rules arisen from expertise are provided to determine the quality of the component in which way automatic quality inspection could be accomplished. The superiority of our detection network over existing ones is sufficiently demonstrated with the detection result of irregular targets in images captured in the industrial scenario. Meanwhile, successful inspection result proves that the system meets the industrial requirements in terms of both accuracy and speed and thus is of practical significance to industrial applications. Jiabin Zhang, Zhengtao Zhang, Hu Su, Xinyi Gong, Feng Zhang 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2011 | Trajectory prediction of spinning ball for ping-pong player robotabstractAn analytic flying model that can well represent the physical behavior is derived, where the ball's self-rotational velocity changes along with the flying velocity. Based on the least square method, a rebound model that represents the relation between the velocities before and after rebound is established. The initial trajectory is fitted to three second order polynomials of the flying time with the measured positions of the ball. The initial velocities of the ball in the analytic flying model, including the flying velocity and the self-rotational velocity, are computed from the polynomials. The ball's landing position and velocity is predicted with the model. The velocities after rebound are determined with the rebound model. By taking the velocities after rebound as new initial ones, the flying trajectory after rebound is described with the model again. In other words, the ball's trajectory is predicted. Experimental results verify the effectiveness of the proposed method. De Xu, Min Tan 0001, Hu Su |
IROS | 4 |