EDBT 2026 Demo / reviewers in the wild / expert
Yuxiang Yang 0001
dblp:59/5778-1
· DBLP profile ↗
28ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-8613-7822ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-Guided Feature Calibration for Multi-modal Object Re-identification with Missing Modalities
Bingyu Hu, Jufeng Zhao, Yuxiang Yang 0001 |
ICIC (19) | 5 |
| 2026 | Urban scene reconstruction using Geometry-aware Gaussian primitives
Xuepu Zeng, Jinlong Fan 0001, Zhekang Dong, Jing Zhang 0037, Yuxiang Yang 0001 |
Neural Networks | 5 |
| 2026 | 2D-Slice and 3D-Cube Mamba Network for Snapshot Spectral Compressive ImagingabstractHyperspectral image (HSI) reconstruction algorithms are fundamental to coded aperture snapshot spectral imaging (CASSI) systems. Recently, deep unfolding networks (DUNs) have emerged as a dominant solution, seamlessly combining traditional optimization frameworks with the strengths of deep learning. Among these, Mamba stands out as a prominent method for modeling long-range dependencies. However, its reliance on one-dimensional (1D) spatial scanning often compromises spectral consistency and spatial coherence, leading to misalignment of neighboring pixels within sequences. To address these limitations, we propose a novel multi-view framework based on 2D-slice modeling, which ensures spatial-spectral continuity in 1D sequences while maintaining computational efficiency. Furthermore, motivated by the need for precise local patch modeling in 2D images, we develop a 3D-cube Mamba model for HSI reconstruction. By integrating the UNet architecture, this model enhances spatial and spectral detail representation through multi-scale receptive field modeling, using fixed cube sizes to dynamically adjust pixel distances. These advancements are incorporated into the A-HQS-accelerated deep unfolding framework, synergistically combining the strengths of 2D-slice and 3D-cube MambaNet to achieve state-of-the-art HSI reconstruction performance. Experimental evaluations on simulated and real-world CASSI datasets demonstrate the efficacy of the proposed approach, achieving superior spectral fidelity and detailed feature representation. The source code is available at: https://github.com/fengyuchao97/SCM-DUN. Yuchao Feng, Zongliang Wu, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | UAWTrack: Universal 3D Single Object Tracking in Adverse Weatherabstract3D single object tracking (3D SOT) in LiDAR point clouds is essential for autonomous driving. Most existing 3D SOT methods focus on clear weather, where point clouds are more defined. However, adverse weather conditions lead to sparser and noisier point clouds, significantly degrading tracking performance and posing safety risks. In this study, we introduce UAWTrack, a universal 3D SOT model designed to perform effectively across diverse real-world weather conditions. UAWTrack comprises three key modules: 1) Voxel Feature Extraction, which mitigates the perturbations in point clouds caused by adverse weather; 2) Motion-centric Spatial-temporal Aggregation and Motion-guided Feature Fusion, capturing motion clues and sampling dense BEV motion features to address the issue of sparsity; and 3) Weather-Specific Tracker, which efficiently handles tracking in various weather conditions. To fill the gap of lacking benchmarks for 3D SOT in adverse weather, we simulate physically valid adverse weather conditions on the KITTI and NuScenes datasets, creating two benchmarks: KITTI-A and NuScenes-A. Extensive experiments demonstrate that UAWTrack achieves state-of-the-art performance under all weather conditions. Yuxiang Yang 0001, Hongjie Gu, Yingqi Deng, Zhekang Dong, Zhiwei He 0001, Jing Zhang 0037 |
AAAI | 1 |
| 2025 | DDPA-3DVG: Vision-Language Dual-Decoupling and Progressive Alignment for 3D Visual Groundingabstract3D visual grounding aims to localize target objects in point clouds based on free-form natural language, which often describes both target and reference objects. Effective alignment between visual and text features is crucial for this task. However, existing two-stage methods that rely solely on object-level features can yield suboptimal accuracy, while one-stage methods that align only point-level features can be prone to noise. In this paper, we propose DDPA-3DVG, a novel framework that progressively aligns visual locations and language descriptions at multiple granularities. Specifically, we decouple natural language descriptions into distinct representations of target objects, reference objects, and their mutual relationships, while disentangling 3D scenes into object-level, voxel-level, and point-level features. By progressively fusing these dual-decoupled features from coarse to fine, our method enhances cross-modal alignment and achieves state-of-the-art performance on three challenging benchmarks—ScanRefer, Nr3D, and Sr3D. The code will be released at https://github.com/HDU-VRLab/DDPA-3DVG. Hongjie Gu, Jinlong Fan 0001, Jing Zhang 0037, Yuxiang Yang 0001 |
IJCAI | 5 |
| 2025 | BEVTrack: A Simple and Strong Baseline for 3D Single Object Tracking in Bird's-Eye Viewabstract3D Single Object Tracking (SOT) is a fundamental task in computer vision and plays a critical role in applications like autonomous driving. However, existing algorithms often involve complex designs and multiple loss functions, making model training and deployment challenging. Furthermore, their reliance on fixed probability distribution assumptions (e.g., Laplacian or Gaussian) hinders their ability to adapt to diverse target characteristics such as varying sizes and motion patterns, ultimately affecting tracking precision and robustness. To address these issues, we propose BEVTrack, a simple yet effective motion-based tracking method. BEVTrack directly estimates object motion in Bird's-Eye View (BEV) using a single regression loss. To enhance accuracy for targets with diverse attributes, it learns adaptive likelihood functions tailored to individual targets, avoiding the limitations of fixed distribution assumptions in previous methods. This approach provides valuable priors for tracking and significantly boosts performance. Comprehensive experiments on KITTI, NuScenes, and Waymo Open Dataset demonstrate that BEVTrack achieves state-of-the-art results while operating at 200 FPS, enabling real-time applicability. The code will be released at https://github.com/xmm-prio/BEVTrack. Yuxiang Yang 0001, Yingqi Deng, Mian Pan, Zhengjun Zha, Jing Zhang 0037 |
IJCAI | 1 |
| 2025 | A multiple aging factor interactive learning framework for lithium-ion battery state-of-health estimation
Zhengyi Bao, Tingting Luo, Mingyu Gao 0002, Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | MTSBV Algorithm Design for Real-Time Battery Data SegmentationabstractWith the increasing number of new energy vehicles, the safety of power batteries has gained significant attention. A prerequisite for accurate estimation of battery state is effective segmentation of battery data. In response to this issue, this paper develops a frequency-domain feature-based real-time segmentation methodology for battery data analysis, specifically proposing a Multi-Threshold Spectral Band Variance (MTSBV) approach. To verify the effectiveness and robustness of the approach in real-world driving applications, we used two cycles of data, the Federal Urban Driving Program (FUDS) and US06, from publicly available battery datasets at the University of Maryland. In addition to analyzing the performance of the algorithm under ideal noise-free conditions, we also evaluated its performance when the signal is corrupted by additive white Gaussian noise (AWGN), salt pepper noise (SPN), and colored noise (CN). The performance of the algorithm was analyzed from two perspectives: the algorithm’s performance under different Signal-to-Noise Ratios (SNRs) and its comparison with other commonly used segmentation algorithms. Experimental results show that MTSBV algorithm has a good effect on solving the problem of battery data segmentation under low SNR conditions. Under FUDS and UD06 conditions, when SNR=15dB, Root Mean Square Error (RMSE) is reduced by 32.1% and 21.8% compared with other optimal conditions, respectively. Ping Li 0032, Yuxiang Yang 0001, Zhiwei He 0001, Mingyu Gao 0002 |
IEEE Internet Things J. | 3 |
| 2025 | Rethinking Semantic-Level Building Change Detection: Ensemble Learning and Dynamic InteractionabstractBuilding change detection (BCD) of multi-temporal images plays a significant role in urban expansion and area internal change analysis. However, current BCD methods remain stagnant at binary-level predictions due to the scarcity of detectable changes and the imbalance between new constructions and demolitions. To advance semantic-level BCD, we propose a dynamic interaction ensemble learning network (DIELNet) using a collaborative training paradigm across multiple datasets and tasks. Firstly, we create a simulated BCD dataset, Inria-CD, derived from the building segmentation dataset. It features complex structures, large scale, and balanced ratios with both binary- and semantic-level labels. Importantly, we shift the traditional single-dataset and single-task BCD learning paradigm by introducing ensemble learning. This mechanism feeds multiple datasets into the model to obtain binary- and semantic-level predictions through a single training process, accommodating partial samples without semantic-level labels. In addition, our DIELNet incorporates bitemporal dynamic interactions during data processing and feature extraction. The former generates progressive sequences by swapping mutual high-frequency components during the Fourier transformation, while the latter is achieved through Mamba-structure modules, which integrate local convolution with dynamic-static kernels and long-range dependencies via state space models. Numerical and visual comparisons demonstrate the superiority of DIELNet. Moreover, existing algorithms can also benefit significantly from our ensemble learning approach. Datasets and codes are available at: https://github.com/fengyuchao97/DIELNet. Yuchao Feng, Yuxiang Yang 0001, Junhua Gao, Xin Yuan 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | CDRP3: Cascade Deep Reinforcement Learning for Urban Driving Safety With Joint Perception, Prediction, and PlanningabstractSafe urban driving is challenging due to the high density of traffic flow and various potential hazards, such as the sudden appearance of unknown objects. Traditional rule-based approaches and imitation learning methods struggle to address the diverse driving scenarios encountered in urban environments. Reinforcement learning (RL), which adapts to a wide range of driving scenarios through continuous interaction with the environment, has demonstrated success in autonomous driving. Making safety decisions when driving in urban environments necessitates a comprehensive perception of the current scene and the ability to predict the evolution of the dynamic scene. In this paper, we present a novel cascade deep reinforcement learning framework, CDRP3, designed to enhance the safety decision-making capabilities of self-driving vehicles in complex scenarios and emergencies. We leverage a multi-modal spatio-temporal perception (MmSTP) module to fuse multi-modal sensor data and introduce temporal perception to capture spatio-temporal information about dynamic driving environments, and a future state prediction (FSP) module to model complex interactions between different traffic participants and explicitly predict their future states. Subsequently, in the PPO-based planning module, we use the comprehensive environmental information obtained from perception and prediction to decode an optimized driving strategy using a lateral and longitudinal separated multi-branch network structure guided by a customized reward function. This approach enables knowledge transfer from the perception and prediction components to planning, and planning-oriented enhancement of safety decision-making capabilities to improve driving safety. Our experiments demonstrate that CDRP3 outperforms state-of-the-art methods, providing superior driving safety in complex urban environments. Yuxiang Yang 0001, Fenglong Ge, Jinlong Fan 0001, Jufeng Zhao, Zhekang Dong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | SAR-SLAM: Self-Attentive Rendering-based SLAM with Neural Point Cloud EncodingabstractNeural implicit representations have recently revolutionized simultaneous localization and mapping (SLAM), giving rise to a groundbreaking paradigm known as NeRF-based SLAM. However, existing methods often fall short in accurately estimating poses and reconstructing scenes. This limitation largely stems from their reliance on volume rendering techniques, which oversimplify the modeling process. In this paper, we introduce a novel neural implicit SLAM system named SAR-SLAM to address these shortcomings. Our approach reconstructs Neural Radiance Fields (NeRFs) using a self-attentive architecture and represents scenes through neural point cloud encoding. Unlike previous NeRF-based SLAM methods, which depend on traditional volume rendering equations for scene representation and view synthesis, our method employs a self-attentive rendering framework with the Transformer architecture during mapping and tracking stages. To enable incremental mapping, we anchor scene features within a neural point cloud, striking a balance between estimation accuracy and computational cost. Experimental results on three challenging datasets show the superior performance and robustness of our SAR-SLAM compared to recent NeRF-based SLAM systems. The code will be released. Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Jing Zhang 0037 |
ACM Multimedia | 3 |
| 2024 | MSF-SLAM: Multi-Sensor-Fusion-Based Simultaneous Localization and Mapping for Complex Dynamic EnvironmentsabstractWe proposed a multi-sensor fusion-based localization and scene reconstruction method for a complex dynamic scene. The multi-level fusion between multiple sensors was implemented by fusing data collected from different sensors in different system modules. In the front-end of the system, the camera and the LiDAR assisted each other. The LiDAR point clouds provided 3D information for the feature points in the image. The moving objects elimination method based on the image can remove the points on the moving objects in the LiDAR point clouds for localization accuracy improvement and static 3D scene reconstruction. To further improve the localization accuracy, a combination of visual loop closure detection and LiDAR loop closure detection was utilized to ensure the global consistency of scene reconstruction. At the system’s back-end, the observation model of different sensors was integrated to construct a multiple constraint factor graph with nonlinear optimization to obtain the optimal system states. Experimental results demonstrated that the proposed multi-sensor fusion-based localization and scene reconstruction algorithm could operate robustly in multiple complex dynamic scenes. Zhiwei He 0001, Yuxiang Yang 0001, Jiahao Nie 0001, Zhekang Dong, Shuo Wang 0030, Mingyu Gao 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | GLT-T: Global-Local Transformer Voting for 3D Single Object Tracking in Point CloudsabstractCurrent 3D single object tracking methods are typically based on VoteNet, a 3D region proposal network. Despite the success, using a single seed point feature as the cue for offset learning in VoteNet prevents high-quality 3D proposals from being generated. Moreover, seed points with different importance are treated equally in the voting process, aggravating this defect. To address these issues, we propose a novel global-local transformer voting scheme to provide more informative cues and guide the model pay more attention on potential seed points, promoting the generation of high-quality 3D proposals. Technically, a global-local transformer (GLT) module is employed to integrate object- and patch-aware prior into seed point features to effectively form strong feature representation for geometric positions of the seed points, thus providing more robust and accurate cues for offset learning. Subsequently, a simple yet effective training strategy is designed to train the GLT module. We develop an importance prediction branch to learn the potential importance of the seed points and treat the output weights vector as a training constraint term. By incorporating the above components together, we exhibit a superior tracking method GLT-T. Extensive experiments on challenging KITTI and NuScenes benchmarks demonstrate that GLT-T achieves state-of-the-art performance in the 3D single object tracking task. Besides, further ablation studies show the advantages of the proposed global-local transformer voting scheme over the original VoteNet. Code and models will be available at https://github.com/haooozi/GLT-T. Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Jing Zhang 0037 |
AAAI | 3 |
| 2023 | OSP2B: One-Stage Point-to-Box Network for 3D Siamese TrackingabstractTwo-stage point-to-box network acts as a critical role in the recent popular 3D Siamese tracking paradigm, which first generates proposals and then predicts corresponding proposal-wise scores. However, such a network suffers from tedious hyper-parameter tuning and task misalignment, limiting the tracking performance. Towards these concerns, we propose a simple yet effective one-stage point-to-box network for point cloud-based 3D single object tracking. It synchronizes 3D proposal generation and center-ness score prediction by a parallel predictor without tedious hyper-parameters. To guide a task-aligned score ranking of proposals, a center-aware focal loss is proposed to supervise the training of the center-ness branch, which enhances the network's discriminative ability to distinguish proposals of different quality. Besides, we design a binary target classifier to identify target-relevant points. By integrating the derived classification scores with the center-ness scores, the resulting network can effectively suppress interference proposals and further mitigate task misalignment. Finally, we present a novel one-stage Siamese tracker OSP2B equipped with the designed network. Extensive experiments on challenging benchmarks including KITTI and Waymo SOT Dataset show that our OSP2B achieves leading performance with a considerable real-time speed. Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Zhengyi Bao, Mingyu Gao 0002, Jing Zhang 0037 |
IJCAI | 3 |
| 2023 | DMCL: Robot Autonomous Navigation via Depth Image Masked Contrastive LearningabstractAchieving high performance in deep reinforcement learning relies heavily on the ability to obtain good state representations from pixel inputs. However, learning an observation-space-to-action-space mapping from high-dimensional inputs is challenging in reinforcement learning, particularly when dealing with consecutive depth images as input states. In addition, we observe that the consecutive inputs of depth images are highly correlated for the autonomous navigation of a mobile robot, which inspires us to capture temporal correlations between consecutive inputs and infer scene change relationships. To this end, we propose a novel end-to-end robot vision navigation method dubbed DMCL, which obtains good spatial-temporal state representation via Depth image Masked Contrastive Learning. It reconstructs the latent representation from consecutive depth images masked in both spatial and temporal dimensions, resulting in a complete environment state representation. To obtain the optimal navigation policy, we leverage the Soft Actor-Critic reinforcement learning in conjunction with the above representation learning. Extensive experiments demonstrate that the proposed DMCL outperforms representative state-of-the-art methods. The source code will be made publicly available. Ping Li 0032, Yuxiang Yang 0001 |
IROS | 4 |
| 2023 | Learning Localization-Aware Target Confidence for Siamese Visual TrackingabstractSiamese tracking paradigm has achieved great success, providing effective appearance discrimination and size estimation by classification and regression. While such a paradigm typically optimizes the classification and regression independently, leading to task misalignment (accurate prediction boxes have no high target confidence scores). In this paper, to alleviate this misalignment, we propose a novel tracking paradigm, called SiamLA. Within this paradigm, a series of simple, yet effective localization-aware components are introduced to generate localization-aware target confidence scores. Specifically, with the proposedlocalization-aware dynamic label(LADL) loss andlocalization-aware label smoothing(LALS) strategy, collaborative optimization between the classification and regression is achieved, enabling classification scores to be aware of location state, not just appearance similarity. Besides, we propose a separatelocalization-aware quality prediction(LAQP) branch to produce location quality scores to further modify the classification scores. To guide a more reliable modification, a novellocalization-aware feature aggregation(LAFA) module is designed and embedded into this branch. Consequently, the resulting target confidence scores are more discriminative for the location state, allowing accurate prediction boxes tend to be predicted as high scores. Extensive experiments are conducted on six challenging benchmarks, including GOT-10 k, TrackingNet, LaSOT, TNL2K, OTB100 and VOT2018. Our SiamLA achieves competitive performance in terms of both accuracy and efficiency. Furthermore, a stability analysis reveals that our tracking paradigm is relatively stable, implying that the paradigm is potential for real-world applications. Jiahao Nie 0001, Zhiwei He 0001, Yuxiang Yang 0001, Mingyu Gao 0002, Zhekang Dong |
IEEE Trans. Multim. | 3 |
| 2022 | ISNet: Shape Matters for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) refers to extracting small and dim targets from blurred backgrounds, which has a wide range of applications such as traffic management and marine rescue. Due to the low signal-to-noise ratio and low contrast, infrared targets are easily submerged in the background of heavy noise and clutter. How to detect the precise shape information of infrared targets remains challenging. In this paper, we propose a novel infrared shape network (ISNet), where Taylor finite difference (TFD) -inspired edge block and two-orientation attention aggregation (TOAA) block are devised to address this problem. Specifically, TFD-inspired edge block aggregates and enhances the comprehensive edge information from different levels, in order to improve the contrast between target and background and also lay a foundation for extracting shape information with mathematical interpretation. TOAA block calculates the lowlevel information with attention mechanism in both row and column directions and fuses it with the high-level information to capture the shape characteristic of targets and suppress noises. In addition, we construct a new benchmark consisting of 1, 000 realistic images in various target shapes, different target sizes, and rich clutter backgrounds with accurate pixel-level annotations, called IRSTD-1k. Experiments on public datasets and IRSTD-1 k demonstrate the superiority of our approach over representative state-of-the-art IRSTD methods. The dataset and code are available at github.com/RuiZhang97/ISNet. Mingjin Zhang, Rui Zhang 0124, Yuxiang Yang 0001, Haichen Bai, Jing Zhang 0037, Jie Guo 0009 |
CVPR | 3 |
| 2022 | SAR-to-Optical Image Translation via Neural Partial Differential EquationsabstractSynthetic Aperture Radar (SAR) becomes prevailing in remote sensing while SAR images are challenging to interpret by human visual perception due to the active imaging mechanism and speckle noise. Recent researches on SAR-to-optical image translation provide a promising solution and have attracted increasing attentions, though still suffering from low optical image quality with geometric distortion due to the large domain gap. In this paper, we mitigate this issue from a novel perspective, i.e., neural partial differential equations (PDE). First, based on the efficient numerical scheme for solving PDE, i.e., Taylor Central Difference (TCD), we devise a basic TCD residual block to build the backbone network, which promotes the extraction of useful information in SAR images by aggregating and enhancing features from different levels. Furthermore, inspired by the Perona-Malik Diffusion (PMD), we devise a PMD neural module to implement feature diffusion through layers, aiming at removing the noises in smooth regions while preserving the geometric structures. Assembling them together, we propose a novel SAR-to-Optical image translation network named S2O-NPDE, which delivers optical images with finer structures and less noise while enjoying an explainability advantage from explicit mathematical derivation. Experiments on the popular SEN1-2 dataset show that our model outperforms state-of-the-art methods in terms of both objective metrics and visual quality. Mingjin Zhang, Chengyu He, Jing Zhang 0037, Yuxiang Yang 0001, Xiaoqi Peng, Jie Guo 0009 |
IJCAI | 4 |
| 2022 | APT-36K: A Large-scale Benchmark for Animal Pose Estimation and TrackingabstractAnimal pose estimation and tracking (APT) is a fundamental task for detecting and tracking animal keypoints from a sequence of video frames. Previous animal-related datasets focus either on animal tracking or single-frame animal pose estimation, and never on both aspects. The lack of APT datasets hinders the development and evaluation of video-based animal pose estimation and tracking methods, limiting the applications in real world, e.g., understanding animal behavior in wildlife conservation. To fill this gap, we make the first step and propose APT-36K, i.e., the first large-scale benchmark for animal pose estimation and tracking. Specifically, APT-36K consists of 2,400 video clips collected and filtered from 30 animal species with 15 frames for each video, resulting in 36,000 frames in total. After manual annotation and careful double-check, high-quality keypoint and tracking annotations are provided for all the animal instances. Based on APT-36K, we benchmark several representative models on the following three tracks: (1) supervised animal pose estimation on a single frame under intra- and inter-domain transfer learning settings, (2) inter-species domain generalization test for unseen animals, and (3) animal pose estimation with animal tracking. Based on the experimental results, we gain some empirical insights and show that APT-36K provides a useful animal pose estimation and tracking benchmark, offering new challenges and opportunities for future research. The code and dataset will be made publicly available at https://github.com/pandorgan/APT-36K. Yuxiang Yang 0001, Yufei Xu, Jing Zhang 0037, Long Lan, Dacheng Tao |
NeurIPS | 1 |
| 2022 | CODON: On Orchestrating Cross-Domain Attentions for Depth Super-Resolution
Yuxiang Yang 0001, Jing Zhang 0037, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2020 | Deep time-frequency representation and progressive decision fusion for ECG classification
Jing Zhang 0037, Yang Cao 0010, Yuxiang Yang 0001, Xiaobin Xu 0002 |
Knowl. Based Syst. | 4 |
| 2019 | An Automatic Detection and Sorting System for Valve Core Based on Machine VisionabstractThe air-conditioning energy consumption will indirectly be affected by the machining accuracy of the valve core in throttle valve used to control refrigerant, so each valve core needs to be tested for its intermediate aperture before leaving the factory. Moreover, the repeatability of the general electronic pneumatic measuring instrument cannot meet the requirements. This paper introduces an automatic detection system of the valve core aperture based on machine vision, which is mainly composed of mechanical movement part, visual part, and control part. The system is based on a method of measuring the valve core aperture by pneumatic measuring instrument, and automatically discriminates the values of pneumatic measuring instrument and barometer by machine vision. Then, the system cooperates with mechanical movement part to complete valve core loading and automatic sorting, thus realizing online automatic detection and sorting of the micro-apertures of the air-conditioning valve core. The experimental results show that the online detection system meets the production requirements of enterprises. Mingyu Gao 0002, Zhiping Zhan, Yuxiang Yang 0001, Zouchao Deng, Jiye Huang, Zhekang Dong |
IECON | 3 |
| 2019 | Adaptive anchor box mechanism to improve the accuracy in the object detection system
Mingyu Gao 0002, Yujie Du, Yuxiang Yang 0001, Jing Zhang 0037 |
Multim. Tools Appl. | 3 |
| 2018 | An automatic aperture detection system for LED cup based on machine vision
Yuxiang Yang 0001, Yanting Lou, Mingyu Gao 0002, Guojin Ma |
Multim. Tools Appl. | 1 |
| 2016 | A machine vision based sealing rings automatic grabbing and putting systemabstractIn order to allow the robot to automatically grab the front sealing rings, in this paper, a recognition and grabbing system of sealing rings based on machine vision is presented. Machine vision technologies are first utilized to help locate the sealing rings and also determine their orientations, an industrial robot is then utilized to accomplish the grabbing of them. The whole system consists of a light source, a camera, an image processing machine and a 4-degrees of freedom industrial robot. Specifically, the system uses the method of the non calibration to accurately locate the positions of sealing rings. Then, sealing rings with a front view are automatically identified by the Hough transform algorithm. Finally, the industrial robot grabs the one of the sealing rings and put it at the proper position of a battery lid. Experimental results show that the system can recognize, grab and put the sealing rings successfully. The proposed system can improve the efficiency for the assembly line of the battery manufacturing and enhance the flexibility and adaptability of the robot. Guojin Ma, Yanting Lou, Mingyu Gao 0002, Yuxiang Yang 0001, Zhiwei He 0001, Hongjuan Zhu |
INDIN | 5 |
| 2016 | A robust vision inspection system for detecting surface defects of film capacitors
Yuxiang Yang 0001, Zhengjun Zha, Mingyu Gao 0002, Zhiwei He 0001 |
Signal Process. | 1 |
| 2015 | Depth map super-resolution using stereo-vision-assisted model
Yuxiang Yang 0001, Mingyu Gao 0002, Jing Zhang 0037, Zhengjun Zha, Zengfu Wang |
Neurocomputing | 1 |
| 2014 | A Stereo-Vision-Assisted model for depth map super-resolutionabstractIn this paper, we propose a novel Stereo-Vision-Assisted (SVA) model for depth map super-resolution. Given a low-resolution depth map as input, we investigate to enhance its resolution or quality using the registered and potentially highresolution color stereo image pair. First, based on the mutual benefits between raw depth map and features of highresolution color image, we model the relationship with two constraint terms of local and non-local priors which sufficiently explore their complementary nature. Then by considering reliable disparity pixels calculated from stereo matching algorithm, we formulate a stereo disparity regularization term to further reinforce the preservation of fine depth detail. In addition, we employ an efficient iterative algorithm to optimize the objective function. Experimental results demonstrate that our approach can achieve high-quality depth map in terms of both spatial resolution and depth precision. Yuxiang Yang 0001, Zhengjun Zha, Mingyu Gao 0002, Qi Tian 0001 |
ICME | 1 |