Jin Wang 0015

dblp:92/1375-15 · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-3106-021XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 3D modeling from a single sketch with multifaceted semantic understanding
Jin Wang 0015, Senyun Jia, Guodong Lu
Expert Syst. Appl.2
2026 GasSeg: A lightweight real-time infrared gas segmentation network for edge devices
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Guodong Lu, Shengfeng He
Pattern Recognit.2
2026 Hybrid Force-Velocity Model Predictive Control Framework for Coordinated Manipulation of Humanoid Dual-Arm Robots
abstract
This article presents a model predictive control (MPC) method for humanoid dual-arm robots (DARs) to realize hybrid force-velocity control in the task space. The proposed method, named hybrid force-velocity MPC (HFV-MPC), enables simultaneous tracking of velocity and external force in the task space, as well as internal force regulation. First, a synchronous decoupling model of internal force, external force, and velocity is developed to address the limitations of conventional methods that decouple internal and external forces solely in the end-effector space and fail to meet the independent regulation requirements of external force and velocity in the task space. Furthermore, a generalized velocity model that integrates task-space velocity, external force, and internal force was established, and a HFV-MPC framework is constructed to enhance system robustness and achieve precise dual-arm force–velocity tracking. Subsequently, a comprehensive MPC cost function is designed, incorporating a dual-arm motion coordination coefficient to suppress undesired redundant joint motions and improve whole-body motion stability. In addition, a feedforward-linearized incremental system is derived, where the incremental model and force prediction jointly construct a generalized state-space model for MPC, to balance modeling accuracy and computational efficiency. Finally, the proposed method is validated in terms of effectiveness and robustness through both simulation and physical robotic experiments.
Jin Wang 0015, Haiyun Zhang, Xiao-Fei Li, Jianwei Niu 0002, Guodong Lu
IEEE Trans Autom. Sci. Eng.2
2026 An Intelligent Multitask Framework for Industrial Gas Leak Detection and Analysis With Infrared Optical Gas Imaging
abstract
Infrared (IR) optical gas imaging (OGI) is widely adopted in industrial environments for detecting fugitive gas emissions. However, conventional IR OGI systems rely heavily on manual inspection, lacking capabilities for active leak localization and in-depth analysis, which increases labor costs and risks of human error. To address these challenges, we present LeakHunter, an intelligent multitask framework designed for industrial gas leak monitoring and decision support. LeakHunter integrates seamlessly with IR cameras and can be deployed on edge computing devices, enabling real-time, on-site leak detection in harsh industrial settings. At the core of LeakHunter is a novel keypoint detection paradigm tailored for IR OGI, capable of localizing both leak sources and diffusion endpoints to enable effective spatiotemporal trend analysis. The framework also estimates critical leak attributes, including plume morphology and flow rate, supporting rapid and informed response. To further enhance detection accuracy, we introduce a biomimetic attention module that improves gas-background separation under complex thermal conditions, and a collaborative multitask head for efficient cross-task feature sharing. In addition, two benchmark datasets are proposed, one of which is a field-test set collected in real industrial scenarios. Experiments demonstrate that LeakHunter achieves state-of-the-art performance across multiple tasks, with an F2 score of 94.8% for gas segmentation and 97.9% for leak keypoint localization, while running at 28.4 FPS on a portable IR OGI device. These results highlight its potential as a deployable, intelligent solution for enhancing industrial safety and automation.
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Zaixing He, Guodong Lu
IEEE Trans. Ind. Informatics2
2026 Purified Zero-Shot Sketch-Based Image Retrieval
abstract
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler that generates linguistic masks based on semantic labels to obscure semantically irrelevant image features. Our novel contribution is the concept of purified masked matching (PMM), which comprises two processes: (1)reconstruction, which compels the image encoder to reconstruct the masked image feature, and (2)interaction, which involves a transformer decoder that processes both sketch and masked image features to investigate cross-domain relationships for effective matching. Evaluated on the TU-Berlin, Sketchy, and QuickDraw datasets, PuXIM sets new benchmarks in terms of performance. Importantly, the distraction-agnostic nature of the matching process renders PuXIM more conducive to training, enabling efficient adaptation to zero-shot scenarios with reduced data requirements and low data quality.
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Shengfeng He
IEEE Trans. Multim.3
2026 GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction
abstract
Predicting 3D Semantic Scene Graphs (3DSSG) is vital for understanding complex scenes by constructing structured representations. Current methods struggle with significant granularity discrepancies among instances, often relying on features at a single scale, which hampers their ability to perceive and interact with differently sized instances. To tackle this challenge, we introduce GranSSG, a novel approach that integrates volumetric granular awareness into 3DSSG prediction. Central to GranSSG is the Volumetric Pooling block, which aggregates features from multiple instance volumes, enhancing the representation of instance patterns across different granularities. Complementing this, the Granularity Transformer block dynamically directs attention to instance features across various network layers, ensuring precise perception of instances regardless of their granularity. Furthermore, the Cross-Granularity Correlation Transformer block mitigates performance degradation in instance pair relationship prediction by adaptively fusing hybrid features from different granularities, providing a comprehensive representation of instance pairs. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that GranSSG significantly enhances prediction performance, setting a new state-of-the-art in 3DSSG prediction.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Jiao Yi, Guodong Lu, Shengfeng He
IEEE Trans. Vis. Comput. Graph.3
2025 Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch Grounding
Jin Wang 0015, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
ACM Multimedia2
2025 Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding
abstract
3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights using a Corroborative Refinement strategy, which adaptively reorganizes information to correct ambiguities arising from biased or incomplete observations. Through this two-stage deliberation, Refer-Judge significantly enhances the reliability of data judgments. Extensive experiments demonstrate that our framework not only achieves human-level discrimination at the scene level but also improves the performance of baseline algorithms via data purification. Code is available at https://github.com/Hermione-HKX/Refer_Judge.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Huan Yu 0002, Guodong Lu, Shengfeng He
NeurIPS3
2025 A lightweight and robust detection network for diverse glass surface defects via scale- and shape-aware feature extraction
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Yiming Liang, Zhan Wang 0002, Haiyan He, Guodong Lu
Eng. Appl. Artif. Intell.2
2025 Sketch-SparseNet: Sparse convolution framework for sketch recognition
Jingru Yang, Jin Wang 0015, Guodong Lu, Huan Yu 0002, Heming Fang, Shengfeng He
Pattern Recognit.2
2024 Incremental accelerated gradient descent and adaptive fine-tuning heuristic performance optimization for robotic motion planning
Shengjie Li 0004, Jin Wang 0015, Haiyun Zhang, Yichang Feng, Guodong Lu, Anbang Zhai
Expert Syst. Appl.2
2024 Cross-Modal Pixel-and-Stroke representation aligning networks for free-hand sketch recognition
Jin Wang 0015, Jingru Yang, Ping Ni, Guodong Lu, Heming Fang, Huan Yu 0002, Kaixiang Huang
Expert Syst. Appl.2
2024 IEFM and IDS: Enhancing 3D environment perception via information encoding in indoor point cloud semantic segmentation
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Guodong Lu, Huan Yu 0002
Neurocomputing2
2024 MsVFE and V-SIAM: Attention-based multi-scale feature interaction and fusion for outdoor LiDAR semantic segmentation
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Huan Yu 0002, Wenming Zou
Neurocomputing2
2024 Granular3D: Delving into multi-granularity 3D scene graph prediction
abstract
This paper addresses the significant challenges in 3D Semantic Scene Graph (3DSSG) prediction, essential for understanding complex 3D environments. Traditional approaches, primarily using PointNet and Graph Convolutional Networks , struggle with effectively extracting multi-grained features from intricate 3D scenes , largely due to a focus on global scene processing and single-scale feature extraction. To overcome these limitations, we introduce Granular3D, a novel approach that shifts the focus towards multi-granularity analysis by predicting relation triplets from specific sub-scenes. One key is the Adaptive Instance Enveloping Method (AIEM), which establishes an approximate envelope structure around irregular instances, providing shape-adaptive local point cloud sampling, thereby comprehensively covering the contextual environments of instances. Moreover, Granular3D incorporates a Hierarchical Dual-Stage Network (HDSN), which differentiates and processes features of instances and their pairs at varying scales, leading to a targeted prediction of instance categories and their relationships. To advance the perception of sub-scene in HDSN, we design a Gather Point Transformer structure (GaPT) that enables the combinatorial interaction of local information from multiple point cloud sets, achieving a more comprehensive local contextual feature extraction. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that our methods provide substantial improvements, establishing a new state-of-the-art in 3DSSG prediction, boosting the top-50 triplet accuracy by +2.8%.
Kaixiang Huang, Jingru Yang, Jin Wang 0015, Shengfeng He, Zhan Wang 0002, Haiyan He, Guodong Lu
Pattern Recognit.3
2023 Roller-Quadrotor: A Novel Hybrid Terrestrial/Aerial Quadrotor with Unicycle-Driven and Rotor-Assisted Turning
abstract
The Roller-Quadrotor is a novel quadrotor that combines the maneuverability of aerial drones with the endurance of ground vehicles. This work focuses on the design, modeling, and experimental validation of the Roller-Quadrotor. Flight capabilities are achieved through a quadrotor config-uration, with four thrust-providing actuators. Additionally, rolling motion is facilitated by a unicycle-driven and rotor-assisted turning structure. By utilizing terrestrial locomotion, the vehicle can overcome rolling and turning resistance, thereby conserving energy compared to its flight mode. This innovative approach not only tackles the inherent challenges of traditional rotorcraft but also enables the vehicle to roll through narrow gaps and overcome obstacles by taking advantage of its aerial mobility. We develop comprehensive models and controllers for the Roller-Quadrotor and validate their performance through experiments. The results demonstrate its seamless transition between aerial and terrestrial locomotion, as well as its ability to safely roll through gaps half the size of its diameter. Moreover, the terrestrial range of the vehicle is approximately 2.8 times greater, while the operating time is about 41.2 times longer compared to its aerial capabilities. These findings underscore the feasibility and effectiveness of the proposed structure and control mechanisms for efficient rolling through challenging terrains while conserving energy.
Jin Wang 0015, Yuze Wu, Qifeng Cai, Huan Yu 0002, Ruibin Zhang, Jie Tu, Jun Meng, Guodong Lu, Fei Gao 0011
IROS2
2023 A distributed variable density path search and simplification method for industrial manipulators with end-effector's attitude constraints
abstract
In many robot operation scenarios, the end-effector’s attitude constraints of movement are indispensable for the task process, such as robotic welding, spraying, handling, and stacking. Meanwhile, the inverse kinematics, collision detection, and space search are involved in the path planning procedure under attitude constraints, making it difficult to achieve satisfactory efficiency and effectiveness in practice. To address these problems, we propose a distributed variable density path planning method with attitude constraints (DVDP-AC) for industrial robots. First, a position–attitude constraints reconstruction (PACR) approach is proposed in the inverse kinematic solution. Then, the distributed signed-distance-field (DSDF) model with single-step safety sphere (SSS) is designed to improve the efficiency of collision detection. Based on this, the variable density path search method is adopted in the Cartesian space. Furthermore, a novel forward sequential path simplification (FSPS) approach is proposed to adaptively eliminate redundant path points considering path accessibility. Finally, experimental results verify the performance and effectiveness of the proposed DVDP-AC method under end-effector’s attitude constraints, and its characteristics and advantages are demonstrated by comparison with current mainstream path planning methods.
Jin Wang 0015, Shengjie Li 0004, Haiyun Zhang, Guodong Lu, Yichang Feng, Jituo Li
Frontiers Inf. Technol. Electron. Eng.1
2023 ContourPose: Monocular 6-D Pose Estimation Method for Reflective Textureless Metal Parts
abstract
Pose estimation is an essential technology for industrial robots to perform precise gripping and assembly. The state-of-the-art deep learning-based approach uses an indirect strategy, i.e., first finding local correspondence between the 2-D image and 3-D model, and then using the perspective-n-point and RANSAC methods to calculate the poses of ordinary objects. However, the metal parts in industry are reflective and textureless, making it difficult to identify distinguishable point features to establish 2-D–3-D correspondences. To address this problem, in this article, we propose a novel deep learning based two-stage method for pose estimation of reflective textureless metal parts, which accurately estimates the target pose using monocular red green blue (RGB) images. Since contours play an important role in both keypoints prediction and pose estimation stages, our method is named ContourPose. First, an additional contour decoder is adopted to implicitly constrain the keypoints prediction in the former stage, which improves the accuracy of the keypoints prediction. Then, the predicted contour of the previous stage is taken as geometric prior that is used to iteratively solve for the optimal pose. Experiments indicate that the proposed approach for reflective textureless metal parts has a significant improvement over the state-of-the-art approaches.
Zaixing He, Quanzhi Li, Xinyue Zhao, Jin Wang 0015, Huarong Shen, Shuyou Zhang 0001, Jianrong Tan
IEEE Trans. Robotics4
2022 Reconstruction of Colored Soft Deformable Objects Based on Self-Generated Template
Jituo Li, Xinqi Liu, Haijing Deng, Guodong Lu, Jin Wang 0015
Comput. Aided Des.6
2022 Real-time skeletonization for sketch-based modeling
Jin Wang 0015, Jituo Li
Comput. Graph.2
2021 Indirect adaptive fuzzy-regulated optimal control for unknown continuous-time nonlinear systems
abstract
We present a novel indirect adaptive fuzzy-regulated optimal control scheme for continuous-time nonlinear systems with unknown dynamics, mismatches, and disturbances. Initially, the Hamilton-Jacobi-Bellman (HJB) equation associated with its performance function is derived for the original nonlinear systems. Unlike existing adaptive dynamic programming (ADP) approaches, this scheme uses a special non-quadratic variable performance function as the reinforcement medium in the actor-critic architecture. An adaptive fuzzy-regulated critic structure is correspondingly constructed to configure the weighting matrix of the performance function for the purpose of approximating and balancing the HJB equation. A concurrent self-organizing learning technique is designed to adaptively update the critic weights. Based on this particular critic, an adaptive optimal feedback controller is developed as the actor with a new form of augmented Riccati equation to optimize the fuzzy-regulated variable performance function in real time. The result is an online indirect adaptive optimal control mechanism implemented as an actor-critic structure, which involves continuous-time adaptation of both the optimal cost and the optimal control policy. The convergence and closed-loop stability of the proposed system are proved and guaranteed. Simulation examples and comparisons show the effectiveness and advantages of the proposed method.
Haiyun Zhang, Deyuan Meng, Jin Wang 0015, Guodong Lu
Frontiers Inf. Technol. Electron. Eng.3
2021 A sub-region one-to-one mapping (SOM) detection algorithm for glass passivation parts wafer surface low-contrast texture defects
Jin Wang 0015, Zhiyong Yu 0003, Zhizhao Duan, Guodong Lu
Multim. Tools Appl.1
2020 Pattern understanding and synthesis based on layout tree descriptor
Jin Wang 0015, Guodong Lu
Vis. Comput.2
2019 Optimized self-adapting contrast enhancement algorithm for wafer contour extraction
Zhiyong Yu 0003, Jin Wang 0015, Guodong Lu
Multim. Tools Appl.2
2019 A Plane Projection Based Method for Base Frame Calibration of Cooperative Manipulators
abstract
Base frame calibration is the foundation for the cooperative work of manipulators. The commonly applied method usually constructed a matrix equation. The contact-mode approach is limited by the insufficient accuracy and efficiency while the high cost impedes the application of methods with the noncontact mode. In this paper, a projection-based method is proposed. All the transformation parameters can be solved based on the geometric constraints. When cooperative manipulators form the closed-chain asynchronously, the topological structure is projected into a particular plane. The rotation angle and translation parameters can be determined with utilization of trigonometric functions and a simple calculation process. The greatest novel feature of this analysis is that only two calibration points are required. The experiment result shows that both the accuracy and efficiency can be much better in the asynchronous calibration.
Jin Wang 0015, Wei Wang 0193, Chao-Hua Wu, Si-Lu Chen 0001, Jianhui Fu, Guo-Dong Lu
IEEE Trans. Ind. Informatics1
2019 Hand-drawn grayscale image colorful colorization based on natural image
Liyang Fang, Jin Wang 0015, Guodong Lu, Jianhui Fu
Vis. Comput.2
2018 Adaptive robust neural control of a two-manipulator system holding a rigid object with inaccurate base frame parameters
abstract
The problem of self-tuning control with a two-manipulator system holding a rigid object in the presence of inaccurate translational base frame parameters is addressed. An adaptive robust neural controller is proposed to cope with inaccurate translational base frame parameters, internal force, modeling uncertainties, joint friction, and external disturbances. A radial basis function neural network is adopted for all kinds of dynamical estimation, including undesired internal force. To validate the effectiveness of the proposed approach, together with simulation studies and analysis, the position tracking errors are shown to asymptotically converge to zero, and the internal force can be maintained in a steady range. Using an adaptive engine, this approach permits accurate online calibration of the relative translational base frame parameters of the involved manipulators. Specialized robust compensation is established for global stability. Using a Lyapunov approach, the controller is proved robust in the face of inaccurate base frame parameters and the aforementioned uncertainties.
Jin Wang 0015, Guo-Dong Lu
Frontiers Inf. Technol. Electron. Eng.2
2014 Sketch2Jewelry: Semantic feature modeling for sketch-based jewelry design
Long Zeng 0001, Yong-Jin Liu 0001, Jin Wang 0015, Matthew M. F. Yuen
Comput. Graph.3
2011 Sketch based garment modeling on an arbitrary view of a 3D virtual human model
abstract
This paper presents a new approach for modeling a virtual garment intuitively and simply by sketching garment style lines. The user sketches directly onto the surface of 3D virtual human from arbitrary viewing directions, and the 3D garment suited to the virtual human can be created automatically. First, a distance field based allocation algorithm is proposed to find the 3D point which has the shortest given distance to the virtual human along the view direction. Then, the 3D style lines are generated by transforming from the 2D strokes on the human model and all the garment pieces are recognized from the 3D style lines. Finally, the 3D garment model is constructed by using the angle and offset based interpolation and Delaunay triangulation. In addition, we propose a body feature based template reusing method to fit the 3D garment to different virtual human models. The method can be adapted to designer habits and improve the usefulness of garment design. Examples show that the method is useful and efficient.
Yu-Lei Geng, Jin Wang 0015, Guo-Dong Lu
J. Zhejiang Univ. Sci. C2
2009 Interactive 3D garment design with constrained contour curves and style curves
Jin Wang 0015, Guo-Dong Lu, Wei-Long Li, Yoshiyuki Sakaguti
Comput. Aided Des.1