Xingyu Chen 0002

dblp:59/7651-2 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-3627-0371ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
abstract
In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training images, SegDINO3D is designed to fully leverage 2D representation from a pre-trained 2D detection model, including both image-level and object-level features, for improving 3D representation. SegDINO3D takes both a point cloud and its associated 2D images as input. In the encoder stage, it first enriches each 3D point by retrieving 2D image features from its corresponding image views and then leverages a 3D encoder for 3D context fusion. In the decoder stage, it formulates 3D object queries as 3D anchor boxes and performs cross-attention from 3D queries to 2D object queries obtained from 2D images using the 2D detection model. These 2D object queries serve as a compact object-level representation of 2D images, effectively avoiding the challenge of keeping thousands of image feature maps in the memory while faithfully preserving the knowledge of the pre-trained 2D model. The introducing of 3D box queries also enables the model to modulate cross-attention using the predicted boxes for more precise querying. SegDINO3D achieves the state-of-the-art performance on the ScanNetV2 and ScanNet200 3D instance segmentation benchmarks. Notably, on the challenging ScanNet200 dataset, SegDINO3D significantly outperforms prior methods by +8.7 and +6.8 mAP on the validation and hidden test sets, respectively, demonstrating its superiority.
Jinyuan Qu, Hongyang Li 0003, Xingyu Chen 0002, Shilong Liu 0004, Yukai Shi, Tianhe Ren, Ruitao Jing, Lei Zhang 0001
AAAI3
2025 HandOS: 3D Hand Reconstruction in One Stage
abstract
Existing approaches of hand reconstruction predominantly adhere to a multi-stage framework, encompassing detection, left-right classification, and pose estimation. This paradigm induces redundant computation and cumulative errors. In this work, we propose HandOS, an end-to-end framework for 3D hand reconstruction. Our central motivation lies in leveraging a frozen detector as the foundation while incorporating auxiliary modules for 2D and 3D keypoint estimation. In this manner, we integrate the pose estimation capacity into the detection framework, while at the same time obviating the necessity of using the left-right category as a prerequisite. Specifically, we propose an interactive 2D-3D decoder, where 2D joint semantics is derived from detection cues while 3D representation is lifted from those of 2D joints. Furthermore, hierarchical attention is designed to enable the concurrent modeling of 2D joints, 3D vertices, and camera translation. Consequently, we achieve an end-to-end integration of hand detection, 2D pose estimation, and 3D mesh reconstruction within a one-stage framework, so that the above multi-stage drawbacks are overcome. Meanwhile, the HandOS reaches state-of-the-art performances on public benchmarks, e.g., 5.0 PA-MPJPE on FreiHand and 64.6% [email protected] on HInt-Ego4D.
Xingyu Chen 0002, Zhuheng Song, Xiaoke Jiang, Yaoqing Hu, Junzhi Yu 0001, Lei Zhang 0001
CVPR1
2025 Adjusting Distributed Cameras for Robust Moving Object Pose Estimation
abstract
Robust moving object pose estimation is crucial in fine manipulation tasks, such as surgical instrument tracking. This paper presents a distributed-camera system with robotic adjustments to maintain consistent tracking of moving objects, thus avoiding tracking failures. An integrated framework for camera adjustment and pose estimation is developed for this distributed-camera system. In each detection cycle, the camera exhibiting the largest deviation with the object is adjusted by a visual servoing technique. After adjustment, the camera extrinsics are re-calibrated in the following detection cycles. For the unadjusted cameras, an online extrinsic optimization method based on multi-frame detection results is proposed to refine the camera extrinsics. Based on the refined camera extrinsics and detection results from multiple cameras, the pose of moving objects relative to the principal camera can be robustly estimated. We test the performance of this system in both simulation environments and real-world scenarios. The results indicate that our system achieves higher pose estimation accuracy and exhibits strong resistance to limited field-of-view (FoV) compared to conventional equivalent fixed multi-camera systems.
Yaoqing Hu, Shaoan Wang, Xingyu Chen 0002, Mingzhu Zhu, Zhanhua Xin, Junzhi Yu 0001
IEEE Trans Autom. Sci. Eng.4
2024 Binary Similarity Few-Shot Object Detection With Modeling of Hard Negative Samples
abstract
For few-shot object detection, this work proposes a binary similarity detector (BSDet), which realizes a novel similarity-based multiple binary classification and enhances the feature margin between positive and hard negative samples. First, we revisit the classification paradigm, concluding that multiple binary classification paradigm is more suitable than multi-class classification paradigm for the few-shot task. Hence, we propose a binary similarity head (BSH) by posing the classification task as multiple binary similarity measurements rather than a multi-class prediction. Second, focusing on the hard negative samples, we propose a feature enhancement module (FEM). During training phase, the FEM can push the features of positive and hard negative samples far away from each other, and thus effectively suppresses false positives. Abundant experiments and visualizations indicate that our method achieves state-of-the-art performances on few-shot object detection tasks.
Xingyu Chen 0002, Zhengxing Wu, Min Tan 0001, Junzhi Yu 0001
IEEE Trans. Multim.2
2023 Decoupled Metric Network for Single-Stage Few-Shot Object Detection
abstract
Within the last few years, great efforts have been made to study few-shot learning. Although general object detection is advancing at a rapid pace, few-shot detection remains a very challenging problem. In this work, we propose a novel decoupled metric network (DMNet) for single-stage few-shot object detection. We design a decoupled representation transformation (DRT) and an image-level distance metric learning (IDML) to solve the few-shot detection problem. The DRT can eliminate the adverse effect of handcrafted prior knowledge by predicting objectness and anchor shape. Meanwhile, to alleviate the problem of representation disagreement between classification and location (i.e., translational invariance versus translational variance), the DRT adopts a decoupled manner to generate adaptive representations so that the model is easier to learn from only a few training data. As for a few-shot classification in the detection task, we design an IDML tailored to enhance the generalization ability. This module can perform metric learning for the whole visual feature, so it can be more efficient than traditional DML due to the merit of parallel inference for multiobjects. Based on the DRT and IDML, our DMNet efficiently realizes a novel paradigm for few-shot detection, called single-stage metric detection. Experiments are conducted on the PASCAL VOC dataset and the MS COCO dataset. As a result, our method achieves state-of-the-art performance in few-shot object detection. The codes are available at https://github.com/yrqs/DMNet.
Xingyu Chen 0002, Zhengxing Wu, Junzhi Yu 0001
IEEE Trans. Cybern.2
2023 HybrUR: A Hybrid Physical-Neural Solution for Unsupervised Underwater Image Restoration
abstract
Robust vision restoration of underwater images remains a challenge. Owing to the lack of well-matched underwater and in-air images, unsupervised methods based on the cyclic generative adversarial framework have been widely investigated in recent years. However, when using an end-to-end unsupervised approach with only unpaired image data, mode collapse could occur, and the color correction of the restored images is usually poor. In this paper, we propose a data- and physics-driven unsupervised architecture to perform underwater image restoration from unpaired underwater and in-air images. For effective color correction and quality enhancement, an underwater image degeneration model must be explicitly constructed based on the optically unambiguous physics law. Thus, we employ the Jaffe-McGlamery degeneration theory to design a generator and use neural networks to model the process of underwater visual degeneration. Furthermore, we impose physical constraints on the scene depth and degeneration factors for backscattering estimation to avoid the vanishing gradient problem during the training of the hybrid physical-neural model. Experimental results show that the proposed method can be used to perform high-quality restoration of unconstrained underwater images without supervision. On multiple benchmarks, the proposed method outperforms several state-of-the-art supervised and unsupervised approaches. We demonstrate that our method yields encouraging results in real-world applications.
Shuaizheng Yan, Xingyu Chen 0002, Zhengxing Wu, Min Tan 0001, Junzhi Yu 0001
IEEE Trans. Image Process.2
2022 UC-OWOD: Unknown-Classified Open World Object Detection
Xingyu Chen 0002, Zhengxing Wu, Liwen Kang, Junzhi Yu 0001
ECCV (10)3
2022 CycleHand: Increasing 3D Pose Estimation Ability on In-the-wild Monocular Image through Cyclic Flow
abstract
Current methods for 3D hand pose estimation fail to generalize well to in-the-wild new scenarios due to varying camera viewpoints, self-occlusions, and complex environments. To address this problem, we propose CycleHand to improve the generalization ability of the model in a self-supervised manner. Our motivation is based on an observation: if one globally rotates the whole hand and reversely rotates it back, the estimated 3D poses of fingers should keep consistent before and after the rotation because the wrist-relative hand poses stay unchanged during global 3D rotation. Hence, we propose arbitrary-rotation self-supervised consistency learning to improve the model's robustness for varying viewpoints. Another innovation of CycleHand is that we propose a high-fidelity texture map to render the photorealistic rotated hand with different lighting conditions, backgrounds, and skin tones to further enhance the effectiveness of our self-supervised task. To reduce the potential negative effects brought by the domain shift of synthetic images, we use the idea of contrastive learning to learn a synthetic-real consistent feature extractor in extracting domain-irrelevant hand representations. Experiments show that CycleHand can largely improve the hand pose estimation performance in both canonical datasets and real-world applications.
Daiheng Gao, Xindi Zhang 0003, Xingyu Chen 0002, Andong Tan, Bang Zhang, Ping Tan 0002
ACM Multimedia3
2022 A novel robotic visual perception framework for underwater operation
abstract
Underwater robotic operation usually requires visual perception (e.g., object detection and tracking), but underwater scenes have poor visual quality and represent a special domain which can affect the accuracy of visual perception. In addition, detection continuity and stability are important for robotic perception, but the commonly used static accuracy based evaluation (i.e., average precision) is insufficient to reflect detector performance across time. In response to these two problems, we present a design for a novel robotic visual perception framework. First, we generally investigate the relationship between a quality-diverse data domain and visual restoration in detection performance. As a result, although domain quality has an ignorable effect on within-domain detection accuracy, visual restoration is beneficial to detection in real sea scenarios by reducing the domain shift. Moreover, non-reference assessments are proposed for detection continuity and stability based on object tracklets. Further, online tracklet refinement is developed to improve the temporal performance of detectors. Finally, combined with visual restoration, an accurate and stable underwater robotic visual perception framework is established. Small-overlap suppression is proposed to extend video object detection (VID) methods to a single-object tracking task, leading to the flexibility to switch between detection and tracking. Extensive experiments were conducted on the ImageNet VID dataset and real-world robotic tasks to verify the correctness of our analysis and the superiority of our proposed approaches. The codes are available at https://github.com/yrqs/VisPerception .
Xingyu Chen 0002, Zhengxing Wu, Junzhi Yu 0001
Frontiers Inf. Technol. Electron. Eng.2
2021 Joint Anchor-Feature Refinement for Real-Time Accurate Object Detection in Images and Videos
abstract
Object detection has been vigorously investigated for years but fast accurate detection for real-world scenes remains a very challenging problem. Overcoming drawbacks of single-stage detectors, we take aim at precisely detecting objects for static and temporal scenes in real time. Firstly, as a dual refinement mechanism, a novel anchor-offset detection is designed, which includes an anchor refinement, a feature location refinement, and a deformable detection head. This new detection mode is able to simultaneously perform two-step regression and capture accurate object features. Based on the anchor-offset detection, a dual refinement network (DRNet) is developed for high-performance static detection, where a multi-deformable head is further designed to leverage contextual information for describing objects. As for temporal detection in videos, temporal refinement networks (TRNet) and temporal dual refinement networks (TDRNet) are developed by propagating the refinement information across time. We also propose a soft refinement strategy to temporally match object motion with the previous refinement. Our proposed methods are evaluated on PASCAL VOC, COCO, and ImageNet VID datasets. Extensive comparisons on static and temporal detection verify the superiority of DRNet, TRNet, and TDRNet. Consequently, our developed approaches run in a fairly fast speed, and in the meantime achieve a significantly enhanced detection accuracy, i.e., 84.4% mAP on VOC 2007, 83.6% mAP on VOC 2012, 69.4% mAP on VID 2017, and 42.4% AP on COCO. Ultimately, producing encouraging results, our methods are applied to online underwater object detection and grasping with an autonomous system. Codes are publicly available at https://github.com/SeanChenxy/TDRN.
Xingyu Chen 0002, Junzhi Yu 0001, Shihan Kong, Zhengxing Wu
IEEE Trans. Circuits Syst. Video Technol.1
2020 Temporally Identity-Aware SSD With Attentional LSTM
abstract
Temporal object detection has attracted significant attention, but most popular detection methods cannot leverage rich temporal information in videos. Very recently, many algorithms have been developed for video detection task, yet very few approaches can achieve real-time online object detection in videos. In this paper, based on the attention mechanism and convolutional long short-term memory (ConvLSTM), we propose a temporal single-shot detector (TSSD) for real-world detection. Distinct from the previous methods, we take aim at temporally integrating pyramidal feature hierarchy using ConvLSTM, and design a novel structure, including a low-level temporal unit as well as a high-level one for multiscale feature maps. Moreover, we develop a creative temporal analysis unit, namely, attentional ConvLSTM, in which a temporal attention mechanism is specially tailored for background suppression and scale suppression, while a ConvLSTM integrates attention-aware features across time. An association loss and a multistep training are designed for temporal coherence. Besides, an online tubelet analysis (OTA) is exploited for identification. Our framework is evaluated on ImageNet VID dataset and 2DMOT15 dataset. Extensive comparisons on the detection and tracking capability validate the superiority of the proposed approach. Consequently, the developed TSSD-OTA achieves a fast speed and an overall competitive performance in terms of detection and tracking. Finally, a real-world maneuver is conducted for underwater object grasping.
Xingyu Chen 0002, Junzhi Yu 0001, Zhengxing Wu
IEEE Trans. Cybern.1
2020 Toward a Maneuverable Miniature Robotic Fish Equipped With a Novel Magnetic Actuator System
abstract
Most existing robotic fish have a large body size driven by servo motor system, while conventional small-sized actuators hardly generate a high swimming performance. This paper reports a miniature untethered robotic fish, whose body length is 69 mm. In particular, a newly designed magnetic actuator system (MAS) is equipped, which guarantees both small-sized dimension and flexibility of the robot. More specifically, the magnetic field generated by a permanent magnet is first investigated based on Biot-Savart law. Then, a novel tail-beating rhythm called magnetically actuated pulse width modulation (MAPWM) is modeled for the new actuator system. Further, an MAPWM-based control method is presented, in which the duty ratio of MAPWAM is innovatively utilized to realize the turning maneuvers for the first time. In addition, Lagrangian method is employed to establish the dynamic model to assess the MAPWM-based control method and the turning performance of the robotic fish. To further improve the maneuverability, the effect of a shape-variable caudal fin is analyzed based on computational fluid dynamics and the built dynamic model. Finally, combined with the MAS, the MAPWM-based control method, and the optimally selected caudal fin, extensive aquatic experiments are conducted on the robotic prototype. The results indicate that the developed miniature robotic fish achieves a considerably higher level of maneuverability in terms of turning radius when compared to swimming robots with equivalent dimensions.
Xingyu Chen 0002, Junzhi Yu 0001, Zhengxing Wu, Shihan Kong
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Dual Refinement Network for Single-Shot Object Detection
abstract
Object detection methods fall into two categories, i.e., two-stage and single-stage detectors. The former is characterized by high detection accuracy while the latter usually has a considerable inference speed. Hence, it is imperative to fuse their merits for a better accuracy vs. speed trade-off. To this end, we propose a dual refinement network (DRN) to boost the performance of the single-stage detector. Inheriting from the advantages of two-stage approaches (i.e., two-step regression and accurate features for detection), anchor refinement and feature offset refinement are conducted in a novel anchor-offset detection, where the detection head is comprised of deformable convolutions. Moreover, to leverage contextual information for describing objects, we design a multi-deformable head, in which multiple detection paths with different receptive field sizes devote themselves to detecting objects. Extensive experiments on PASCAL VOC and ImageNet VID datasets are conducted, and we achieve a state-of-the-art detection performance in terms of both accuracy and inference speed.
Xingyu Chen 0002, Xiyuan Yang, Shihan Kong, Zhengxing Wu, Junzhi Yu 0001
ICRA1
2018 TSSD: Temporal Single-Shot Detector Based on Attention and LSTM
abstract
Temporal object detection has attracted significant attention, but most popular methods can not leverage the rich temporal information in video or robotic vision. Although many different algorithms have been developed for video detection task, real-time online approaches are frequently deficient. In this paper, based on attention mechanism and convolutional long short-term memory (ConvLSTM), we propose a temporal single-shot detector (TSSD)for robotic vision. Distinct from previous methods, we aim to temporally integrate pyramidal feature hierarchy using ConvLSTM, and design a novel structure including a high-level ConvLSTM unit as well as a low-level one (HL-LSTM)for multi-scale feature maps. Moreover, we develop a creative temporal analysis unit, namely, ConvLSTM-based attention and attention-based ConvLSTM (A&CL), in which the ConvLSTM-based attention is specially tailored for background suppression and scale suppression while the attention-based ConvLSTM temporally integrates attention-aware features. Finally, our method is evaluated on ImageNet VID dataset. Extensive comparisons on detection performance confirm the superiority of the proposed approach, and the developed TSSD achieves a considerably enhanced accuracy vs. speed trade-off, i.e., 64.8% mAP vs. 27 FPS.
Xingyu Chen 0002, Zhengxing Wu, Junzhi Yu 0001
IROS1