Wenhu Qin

dblp:42/5698 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Optimization-Based Variational Bayesian Filter for Nonlinear State Estimation
Guoqiang Mao, Keyin Wang, Baoqi Huang, Tianxuan Fu, Wei Xiang 0001, Wenhu Qin
IEEE Internet Things J.6
2026 Musculoskeletal Motion Control and Generation Based on Muscle Synergies
abstract
ABSTRACT This paper presents a unified two‐stage framework for physics‐based musculoskeletal motion control and generation. To tackle the challenges posed by high dimensionality and redundancy, we first employ an autoencoder to learn a low‐dimensional muscle synergy space from activation data. Policies trained in this space make the character faithfully reproduce motions while generating physiologically plausible muscle activations. We then leverage these expert trajectories to train a Conditional VAE, encoding skills into a continuous latent space for task‐agnostic motion synthesis and downstream control. Experiments show our method achieves high motion imitation accuracy and generation diversity, ensures control stability, and maintains physiological realism, offering an effective solution for generalizing control of complex musculoskeletal characters.
Libo Sun 0001, Jiwei Wen, Wenhu Qin
Comput. Animat. Virtual Worlds3
2026 KPGBeltNet: in-vehicle seatbelt detection algorithm based on human keypoint-guided sampling and local-global attention
Jiacheng Tang, Libo Sun 0001, Zeyun Zhang, Wenhu Qin
Vis. Comput.4
2025 A Method for Pig Body Condition Scoring Based on MobileSAM
abstract
The body condition scoring (BCS) of pigs comprehensively reflects the health status of pigs by evaluating their fat and muscle reserves. Traditional evaluation methods rely on manual measurement of pig backfat thickness, which is inefficient and can causes severe stress reactions in pigs. In recent years, the development of computer vision has provided new ideas for pig BCS evaluation. This paper proposes a pig body condition score evaluation method based on the MobileSAM. The method first detects pigs from a rear-view perspective, then uses the resulting bounding boxes as prompts for an improved MobileSAM to perform instance segmentation. We enhanced MobileSAM's segmentation accuracy by fine-tuning the model, adapting MobileSAM from general instance segmentation tasks to the specific task of pig instance segmentation. Finally, we calculated the aspect ratio (LWR) of the minimum bounding rectangle of the mask from the pig's rear-view perspective. A decision tree model is employed to perform a non-linear mapping of the LWR value, classifying it into one of five predefined body condition score levels to determine the final grade. Our method effectively overcomes the limitations of traditional manual measurements. It features low computational cost, excellent realtime performance, and supports deployment on edge computing terminals in environments such as pig farms. Experimental results show that our method has significant advantages over other image classification and object detection methods in terms of accuracy and interpretability.
Zeyun Zhang, Libo Sun 0001, Xiaoyu Chang, Weipeng Shi, Wenhu Qin
CW6
2025 QuantBEVFusion: A Fully Quantized Framework for LiDAR-camera 3D Object Detection
abstract
3D object detection is essential for robust environmental perception in autonomous driving and robotics. While LiDAR-camera fusion methods offer high accuracy, their computational complexity hinders deployment on resource-constrained edge devices. To address this, we introduce QuantBEVFusion, a fully quantized 3D object detection framework that prioritizes both quantization and operator optimization. Our approach tackles the inherent asymmetry between LiDAR and camera data in pillar-based models by incorporating a novel pillar bird’s-eye view (BEV) encoder, significantly boosting performance. Furthermore, we introduce 1) an optimized LiDAR input processing method that filters noise and enables per-tensor quantization; 2) an improved sparse feature quantization process with log-histogram balancing, adaptive bin widths, and distillation loss for enhanced accuracy; and 3) a deployment-friendly 3D-to-2D transformation operator facilitating fixed-point implementation. Extensive experiments demonstrate that QuantBEVFusion achieves state-of-the-art quantization performance while maintaining accuracy suitable for real-time applications on edge devices.
Xubin Wen, Ming Shao, Libo Sun 0001, Wenhu Qin, Si-Yu Xia
IJCNN4
2025 Unsupervised Retinex Exposure Control: A Novel Approach to Image Enhancement
abstract
ABSTRACT In domains such as autonomous driving and remote sensing, images often suffer from challenging lighting conditions, including low‐light, backlighting and overexposure, which hinder the recognition of pedestrians, vehicles and traffic signs. While numerous methods have been proposed to address poor image exposure, they often struggle with images containing both low‐light and overexposed regions. This paper presents an unsupervised learning‐based exposure control method, providing a novel approach to improving image quality under diverse lighting conditions. Leveraging the inherent properties of Retinex theory, we introduce a novel yet simple formula that adjusts image exposure to produce visually pleasing results without requiring paired training data. Experiments on diverse image datasets validate the effectiveness of our approach in addressing various exposure challenges while preserving critical visual details. Our framework not only simplifies the exposure control process but also achieves state‐of‐the‐art performance, highlighting its potential for real‐world applications in computer vision and image processing.
Yukun Yang 0004, Libo Sun 0001, Weipeng Shi, Wenhu Qin
IET Image Process.4
2025 SP-Det: Anchor-based lane detection network with structural prior perception
Libo Sun 0001, Hangyu Zhu, Wenhu Qin
Pattern Recognit. Lett.3
2025 PEPillar: a point-enhanced pillar network for efficient 3D object detection in autonomous driving
Libo Sun 0001, Wenhu Qin
Vis. Comput.3
2025 Adaptive formation control and transformation in virtual crowds via deep reinforcement learning
Libo Sun 0001, Yongchun Qiu, Wenhu Qin
Vis. Comput.3
2025 The crowd cooperation approach for formation maintenance and collision avoidance using multi-agent deep reinforcement learning
Libo Sun 0001, Jiahui Yan, Yongchun Qiu, Wenhu Qin
Vis. Comput.4
2024 Physical based motion reconstruction from videos using musculoskeletal model
abstract
Abstract We propose a novel method that combines human pose estimation and physical simulation of character animation. Our approach allows characters to learn from the actor's skills captured in videos and subsequently reconstruct the motions with high fidelity in a physically simulated environment. Firstly, we model the character based on the human musculoskeletal system and build a complete dynamics model of the proposed system using the Lagrange equations of motion. Next, we employ the pose estimation method to process the input video and generate human reference motion. Finally, we design a hierarchical control framework consisting of a trajectory tracking layer and a muscle control layer. The trajectory tracking layer aims to minimize the difference between the reference motion pose and the actual output pose, while the muscle control layer aims to minimize the difference between the target torque and the actual output muscle force. The two layers interact by passing parameters through a proportional differential controller until the desired learning objective is achieved. A series of complex experimental results demonstrate that our proposed method can learn to produce comparable high‐quality motions with high similarity from videos of different complexity levels and remains stable in the presence of muscle contracture weakness perturbations.
Libo Sun 0001, Wenhu Qin
Comput. Animat. Virtual Worlds3
2024 A language-directed virtual human motion generation approach based on musculoskeletal models
abstract
Abstract The development of the systems capable of synthesizing natural and life‐like motions for virtual characters has long been a central focus in computer animation. It needs to generate high‐quality motions for characters and provide users with a convenient and flexible interface for guiding character motions. In this work, we propose a language‐directed virtual human motion generation approach based on musculoskeletal models to achieve interactive and higher‐fidelity virtual human motion, which lays the foundation for the development of language‐directed controllers in physics‐based character animation. First, we construct a simplified model of musculoskeletal dynamics for the virtual character. Subsequently, we propose a hierarchical control framework consisting of a trajectory tracking layer and a muscle control layer, obtaining the optimal control policy for imitating the reference motions through the training. We design a multi‐policy aggregation controller based on large language models, which selects the motion policy with the highest similarity to user text commands from the action‐caption data pool, facilitating natural language‐based control of virtual character motions. Experimental results demonstrate that the proposed approach not only generates high‐quality motions highly resembling reference motions but also enables users to effectively guide virtual characters to perform various motions via natural language instructions.
Libo Sun 0001, Wenhu Qin
Comput. Animat. Virtual Worlds3
2024 AG-YOLO: Attention-guided network for real-time object detection
Hangyu Zhu, Libo Sun 0001, Wenhu Qin, Feng Tian 0006
Multim. Tools Appl.3
2023 Bidirectional temporal feature for 3D human pose and shape estimation from a video
abstract
Abstract 3D human pose and shape estimation is the foundation of analyzing human motion. However, estimating accurate and temporally consistent 3D human motion from a video remains a challenge. By now, most of the video‐based methods for estimating 3D human pose and shape rely on unidirectional temporal features and lack more comprehensive information. To solve this problem, we propose a novel model “bidirectional temporal feature for human motion recovery” (BTMR), which consists of a human motion generator and a discriminator. The transformer‐based generator effectively captures the forward and reverse temporal features to enhance the temporal correlation between frames and reduces the loss of spatial information. The motion discriminator based on Bi‐LSTM can distinguish whether the generated pose sequences are consistent with the realistic sequences of the AMASS dataset. In the process of continuous generation and discrimination, the model can learn more realistic and accurate poses. We evaluate our BTMR on 3DPW and MPI‐INF‐3DHP datasets. Without the training set of 3DPW, BTMR outperforms VIBE by 2.4 mm and 14.9 mm/s2 in PA‐MPJPE and Accel metrics and outperforms TCMR by 1.7 mm in PA‐MPJPE metric on 3DPW. The results demonstrate that our BTMR produces better accurate and temporal consistent 3D human motion.
Libo Sun 0001, Ting Tang, Yuke Qu, Wenhu Qin
Comput. Animat. Virtual Worlds4
2023 Path planning for multiple agents in an unknown environment using soft actor critic and curriculum learning
abstract
Abstract Path planning can guarantee that agents reach their goals without colliding with obstacles and other agents in an optimal way and it is a very important component in the research of crowd simulation. In this article, we propose a novel path planning approach for multiple agents which combines soft actor critic (SAC) algorithm and curriculum learning to solve the problems of single policy, slow convergence of the policy in an unknown environment with sparse rewards. The path planning task is set as lessons from easy to difficult, and the neural network of the SAC algorithm is arranged to learn in sequence, and finally the neural network can be fully competent for the path planning task. We also stack the state information to address the problems caused by limited observation for policy learning, and design a comprehensive reward function to make agents reach their goals successfully and avoid collisions with static obstacles and other agents. The experimental results demonstrate that our approach can plan smooth and natural paths for multiple agents, and furthermore, our model has a certain generalization ability and a better adaptability to the changes in a dynamic environment.
Libo Sun 0001, Jiahui Yan, Wenhu Qin
Comput. Animat. Virtual Worlds3
2023 Semantic Representation Fusion-Based Network for Robust Land Cover Classification in Foggy Conditions
abstract
A precise and robust classification of land cover is crucial for land use estimation. A robust model that can provide rich semantic information is imperative for the challenging task of land cover classification in foggy conditions. We propose Semantic Representation Enhancement (SRE) and Semantic Representation Aggregation (SRA) modules for the fusion of semantic representation. The Dense Depthwise Separable Atrous Spatial Pyramid Pooling (DDS-ASPP) module in SRE possesses a large receptive field, which covers an extensive scale range. Enhanced asymmetric convolution module (EACM) in SRE focus on features of various directions. DDS-ASPP and EACM generate the class-based and pixel-based representation respectively. By means of SRA and dual representations, we model the relationship between global context and coarse class regions to capture long-range correlation. Moreover, evaluated on Potsdam, Vaihingen and custom real-world datasets under fog, we demonstrate that our work is competitive with state-of-the-art models in terms of robustness. Code will be available at https://github.com/bowenroom/Robust-land-cover-classification.
Weipeng Shi, Wenhu Qin, Zhonghua Yun, Yukun Yang 0004
IEEE Trans. Geosci. Remote. Sens.2
2022 A Pig Pose Estimation Model for Measuring Pig's Body Size
Yukun Yang 0004, Wenhu Qin, Libo Sun 0001, Weipeng Shi
CGI2
2022 Muscle-driven virtual human motion generation approach based on deep reinforcement learning
abstract
Abstract We propose a muscle‐driven motion generation approach to realize virtual human motion with user interaction and higher fidelity, which can address the problem that the joint‐driven fails to reflect the motion process of the human body. First, a simplified virtual human musculoskeletal model is built based on human biomechanics. Then, a hierarchical policy learning framework is constructed including motion tracking layer, SPD controller and muscle control layer. The motion tracking layer is responsible for mimicking reference motion and completing control command, using proximal policy optimization to train the policy; the muscle control layer is aimed to minimize muscle energy consumption and train the policy based on supervised learning; the SPD controller acts as a link between the two layers. At the same time, we integrate the curriculum learning to improve the efficiency and success rate of policy training. Simulation experiments show that the proposed approach can use motion capture data and pose estimation data as reference motions to generate better and more adaptable motions. Furthermore, the virtual human has the ability to respond to the user control command during the motion, and can complete the target task successfully.
Wenhu Qin, Libo Sun 0001, Kaiyue Dong
Comput. Animat. Virtual Worlds1
2022 Crowd navigation in an unknown and complex environment based on deep reinforcement learning
abstract
Abstract We propose a virtual crowd navigation approach based on deep reinforcement learning to improve the adaptability of virtual crowds in an unknown and complex environment. To address the problem of local optimum or slow iteration or even failure to converge due to sparse rewards in complex environments, we integrate the curiosity‐driven mechanism, the key navigation points acquisition and the failure path penalty method in addition to combining long short‐term memory networks, dynamic obstacle collision prediction with proximal policy optimization algorithms, which realizes the crowd navigation in a complex environment. The experimental results show that the proposed approach can simulate the motions of virtual crowds in various dynamic and complex scenarios without the environment modeling. The use of continuous action space also ensures that the movement trajectories of the virtual crowds are more realistic and natural. Furthermore, our approach can provide analysis and demonstration tools for a variety of independent collaborations such as competition and cooperation of group intelligence in an open and dynamic environment.
Libo Sun 0001, Yuke Qu, Wenhu Qin
Comput. Animat. Virtual Worlds3
2022 Research on target recognition and tracking in mobile augmented reality assisted maintenance
abstract
Abstract The recognition and tracking of maintenance targets is the basis of augmented reality assisted maintenance. To address the problems of the current recognition and tracking algorithms, such as high time complexity, low pose tracking accuracy, and high requirements for hardware equipment, this paper studies a light weight maintenance target recognition algorithm based on YOLOv5s, and then the VI ORB‐SLAM algorithm is adopted to track the maintenance target. In addition, the ORB feature extraction and the visual inertial initialization are improved for the VI ORB‐SLAM algorithm. Finally, the combined algorithm is deployed on the mobile phone. Taking the augmented reality assisted vehicle maintenance as an example, it is verified that the proposed approach is practical feasible and effective in the actual assisted maintenance scene.
Libo Sun 0001, Wenhu Qin
Comput. Animat. Virtual Worlds3
2022 Land Cover Classification in Foggy Conditions: Toward Robust Models
abstract
Robust semantic labeling of high-resolution remote sensing images in foggy conditions is crucial for automatic monitoring of land covers. This remains a challenging task owing to the low inter-class differentiation yet high intra-class variance and geometric size diversity. Although conventional Convolutional Neural Networks have demonstrated state of the art performance in semantic segmentation, most networks are primarily concerned with standard accuracy, while the influence on robustness is rarely explored. This letter proposes a reliable framework which is evaluated across various severity levels of fog corruptions. Utilizing HRNet as the backbone to maintain high-resolution representations, we develop a multimodal fusion module to exploit the complementary information of lidar and multispectral data. Based on the evaluation experiment on fog corrupted ISPRS 2D datasets, our model demonstrates promising performance with an average mIoU on the clean along with the corrupted datasets exceeding 80% and 56% respectively.
Weipeng Shi, Wenhu Qin, Zhonghua Yun, Allshine Chen
IEEE Geosci. Remote. Sens. Lett.2
2021 AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
Zhe Zhang 0045, Chunyu Wang 0001, Weichao Qiu, Wenhu Qin, Wenjun Zeng 0001
Int. J. Comput. Vis.4
2020 Fusing Wearable IMUs With Multi-View Images for Human Pose Estimation: A Geometric Approach
abstract
We propose to estimate 3D human pose from multi-view images and a few IMUs attached at person's limbs. It operates by firstly detecting 2D poses from the two signals, and then lifting them to the 3D space. We present a geometric approach to reinforce the visual features of each pair of joints based on the IMUs. This notably improves 2D pose estimation accuracy especially when one joint is occluded. We call this approach Orientation Regularized Network (ORN). Then we lift the multi-view 2D poses to the 3D space by an Orientation Regularized Pictorial Structure Model (ORPSM) which jointly minimizes the projection error between the 3D and 2D poses, along with the discrepancy between the 3D pose and IMU orientations. The simple two-step approach reduces the error of the state-of-the-art by a large margin on a public dataset. Our code will be released at https://github.com/microsoft/imu-human-pose-estimation-pytorch.
Zhe Zhang 0045, Chunyu Wang 0001, Wenhu Qin, Wenjun Zeng 0001
CVPR3
2015 Planning Feasible and Smooth Paths for Simulating Realistic Crowd
Libo Sun 0001, Wenhu Qin
ICIC (1)3
2015 A Closed-Loop Speed Advisory Model With Driver's Behavior Adaptability for Eco-Driving
abstract
Providing drivers with speed advisories is an effective eco-driving method at signalized intersections. However, all current research on speed advisory models has excluded the driver's behavior factor. In this paper, we focus on developing a speed advisory model that is able to adapt to the driver's behavior for eco-driving. First, we propose a closed-loop speed advisory framework, with simulation results to show that the current model could not fit in the closed-loop implementation. Next, the continuous acceleration with explicit high velocity boundary (CAEHV) model is established to address the issues when the existing model is used. However, the simulation results for the CAEHV model are not fully satisfactory due to the existence of oscillations in actual speed trajectories. Third, the CAEHV with coasting (CAEHV-C) model is established, in which the vehicle coasting is applied to supplement cruising to avoid oscillations. Simulation results show that the fuel economy performance of the CAEHV-C model is improved by 4% when compared with the CAEHV model. It also shows that CAEHV-C performs the best in terms of the driver's behavior adaptability.
Xuehai Xiang, Wei-Bin Zhang, Wenhu Qin, Qingzhou Mao
IEEE Trans. Intell. Transp. Syst.4
2014 Research on a DSRC-Based Rear-End Collision Warning Model
abstract
Dedicated short-range communication (DSRC) is an emerging technology that allows vehicles to communicate with each other. The rear-end collision warning system based on DSRC has its unique advantages. However, there are problems (e.g., high rates of false alarms and missing alarms in emergency warnings) in the system due to uncertain measurement errors. In this paper, we propose to address the problems by establishing a robust rear-end collision warning model without using expensive high-end devices. Simulations have shown that high rates (up to 56%) of missing alarms occur in the vehicle kinematics (VK) model, as well as false alarms (most of which exceed 70%) in the VK model with maximum compensation (VK-MC). Pertaining to these rates, a novel model based on the neural network (NN) approach is implemented. Through training and validation, the NN model is able to provide emergency warnings with an improved performance of false alarm probability under 20% and the missing alarm probability under 10% for all test cases.
Xuehai Xiang, Wenhu Qin, Binfu Xiang
IEEE Trans. Intell. Transp. Syst.2
2013 Simulating realistic crowd based on agent trajectories
abstract
ABSTRACT This paper presents a model for simulating realistic crowd behaviors at low computation cost. The proposed model is inspired by video data. In our approach, we first classify the crowd into two categories: main and background characters. Whether the agents are main characters or not is influenced by two factors, one is the agent's trajectories and the other one is the change of the environment. In the second stage, we adopt two approaches to simulate the behaviors of main and background characters. Main characters are intelligent agents with the perception, the memory, the planning, and the psychology so that they can make decisions themselves. Background characters are informed of the behavior options for execution by the “smart environment.” Finally, we simulate the road‐crossing scenario in a three‐dimensional virtual environment. The experimental results demonstrate that our approach not only well reflects the characteristics of agent behaviors but also reduces the computation complexity of simulating realistic crowd. Copyright © 2013 John Wiley & Sons, Ltd.
Libo Sun 0001, Wenhu Qin
Comput. Animat. Virtual Worlds3
2012 Animating synthetic dyadic conversations with variations based on context and agent attributes
abstract
ABSTRACT Conversations between two people are ubiquitous in many inhabited contexts. The kinds of conversations that occur depend on several factors, including the time, the location of the participating agents, the spatial relationship between the agents, and the type of conversation in which they are engaged. The statistical distribution of dyadic conversations among a population of agents will therefore depend on these factors. In addition, the conversation types, flow, and duration will depend on agent attributes such as interpersonal relationships, emotional state, personal priorities, and socio‐cultural proxemics. We present a framework for distributing conversations among virtual embodied agents in a real‐time simulation. To avoid generating actual language dialogues, we express variations in the conversational flow by using behavior trees implementing a set of conversation archetypes. The flow of these behavior trees depends in part on the agents' attributes and progresses based on parametrically estimated transitional probabilities. With the participating agents' state, a ‘smart event’ model steers the interchange to different possible outcomes as it executes. Example behavior trees are developed for two conversation archetypes: buyer–seller negotiations and simple asking–answering; the model can be readily extended to others. Because the conversation archetype is known to participating agents, they can animate their gestures appropriate to their conversational state. The resulting animated conversations demonstrate reasonable variety and variability within the environmental context. Copyright © 2012 John Wiley & Sons, Ltd.
Libo Sun 0001, Alexander Shoulson, Nicole Nelson, Wenhu Qin, Ani Nenkova, Norman I. Badler
Comput. Animat. Virtual Worlds5
2010 Smart Events and Primed Agents
Catherine Stocker, Libo Sun 0001, Wenhu Qin, Jan M. Allbeck, Norman I. Badler
IVA4