VLDB 2026 Research / reviewers in the wild / expert
Libo Sun 0001
dblp:37/7606-1
· DBLP profile ↗
26ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0002-7838-9410ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 12 first-author · 19 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Musculoskeletal Motion Control and Generation Based on Muscle SynergiesabstractABSTRACT This paper presents a unified two‐stage framework for physics‐based musculoskeletal motion control and generation. To tackle the challenges posed by high dimensionality and redundancy, we first employ an autoencoder to learn a low‐dimensional muscle synergy space from activation data. Policies trained in this space make the character faithfully reproduce motions while generating physiologically plausible muscle activations. We then leverage these expert trajectories to train a Conditional VAE, encoding skills into a continuous latent space for task‐agnostic motion synthesis and downstream control. Experiments show our method achieves high motion imitation accuracy and generation diversity, ensures control stability, and maintains physiological realism, offering an effective solution for generalizing control of complex musculoskeletal characters. Libo Sun 0001, Jiwei Wen, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2026 | SDR-GAIN: A High Real-Time Occluded Pedestrian Pose Completion Method for Autonomous DrivingabstractWith the advancement of vision-based autonomous driving technology, pedestrian detection have become an important component for improving traffic safety and driving system robustness. Nevertheless, in complex traffic scenarios, conventional pose estimation approaches frequently fail to accurately reconstruct occluded keypoints, primarily due to obstructions caused by vehicles, vegetation, or architectural elements. To address this issue, we propose a novel real-time occluded pedestrian pose completion framework termed Separation and Dimensionality Reduction-based Generative Adversarial Imputation Nets (SDR-GAIN). Unlike previous approaches that train visual models to distinguish occlusion patterns, SDR-GAIN aims to learn human pose directly from the numerical distribution of keypoint coordinates and interpolate missing positions. It employs a self-supervised adversarial learning paradigm to train lightweight generators with residual structures for the imputation of missing pose keypoints. Additionally, it integrates multiple pose standardization techniques to alleviate the difficulty of the learning process. Experiments conducted on the COCO and JAAD datasets demonstrate that SDR-GAIN surpasses conventional machine learning and Transformer-based missing data interpolation algorithms in accurately recovering occluded pedestrian keypoints, while simultaneously achieving microsecond-level real-time inference. Honghao Fu, Yongli Gu, Yidong Yan, Yilang Shen, Libo Sun 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | KPGBeltNet: in-vehicle seatbelt detection algorithm based on human keypoint-guided sampling and local-global attention
Jiacheng Tang, Libo Sun 0001, Zeyun Zhang, Wenhu Qin |
Vis. Comput. | 2 |
| 2025 | A Method for Pig Body Condition Scoring Based on MobileSAMabstractThe body condition scoring (BCS) of pigs comprehensively reflects the health status of pigs by evaluating their fat and muscle reserves. Traditional evaluation methods rely on manual measurement of pig backfat thickness, which is inefficient and can causes severe stress reactions in pigs. In recent years, the development of computer vision has provided new ideas for pig BCS evaluation. This paper proposes a pig body condition score evaluation method based on the MobileSAM. The method first detects pigs from a rear-view perspective, then uses the resulting bounding boxes as prompts for an improved MobileSAM to perform instance segmentation. We enhanced MobileSAM's segmentation accuracy by fine-tuning the model, adapting MobileSAM from general instance segmentation tasks to the specific task of pig instance segmentation. Finally, we calculated the aspect ratio (LWR) of the minimum bounding rectangle of the mask from the pig's rear-view perspective. A decision tree model is employed to perform a non-linear mapping of the LWR value, classifying it into one of five predefined body condition score levels to determine the final grade. Our method effectively overcomes the limitations of traditional manual measurements. It features low computational cost, excellent realtime performance, and supports deployment on edge computing terminals in environments such as pig farms. Experimental results show that our method has significant advantages over other image classification and object detection methods in terms of accuracy and interpretability. Zeyun Zhang, Libo Sun 0001, Xiaoyu Chang, Weipeng Shi, Wenhu Qin |
CW | 2 |
| 2025 | QuantBEVFusion: A Fully Quantized Framework for LiDAR-camera 3D Object Detectionabstract3D object detection is essential for robust environmental perception in autonomous driving and robotics. While LiDAR-camera fusion methods offer high accuracy, their computational complexity hinders deployment on resource-constrained edge devices. To address this, we introduce QuantBEVFusion, a fully quantized 3D object detection framework that prioritizes both quantization and operator optimization. Our approach tackles the inherent asymmetry between LiDAR and camera data in pillar-based models by incorporating a novel pillar bird’s-eye view (BEV) encoder, significantly boosting performance. Furthermore, we introduce 1) an optimized LiDAR input processing method that filters noise and enables per-tensor quantization; 2) an improved sparse feature quantization process with log-histogram balancing, adaptive bin widths, and distillation loss for enhanced accuracy; and 3) a deployment-friendly 3D-to-2D transformation operator facilitating fixed-point implementation. Extensive experiments demonstrate that QuantBEVFusion achieves state-of-the-art quantization performance while maintaining accuracy suitable for real-time applications on edge devices. Xubin Wen, Ming Shao, Libo Sun 0001, Wenhu Qin, Si-Yu Xia |
IJCNN | 3 |
| 2025 | Unsupervised Retinex Exposure Control: A Novel Approach to Image EnhancementabstractABSTRACT In domains such as autonomous driving and remote sensing, images often suffer from challenging lighting conditions, including low‐light, backlighting and overexposure, which hinder the recognition of pedestrians, vehicles and traffic signs. While numerous methods have been proposed to address poor image exposure, they often struggle with images containing both low‐light and overexposed regions. This paper presents an unsupervised learning‐based exposure control method, providing a novel approach to improving image quality under diverse lighting conditions. Leveraging the inherent properties of Retinex theory, we introduce a novel yet simple formula that adjusts image exposure to produce visually pleasing results without requiring paired training data. Experiments on diverse image datasets validate the effectiveness of our approach in addressing various exposure challenges while preserving critical visual details. Our framework not only simplifies the exposure control process but also achieves state‐of‐the‐art performance, highlighting its potential for real‐world applications in computer vision and image processing. Yukun Yang 0004, Libo Sun 0001, Weipeng Shi, Wenhu Qin |
IET Image Process. | 2 |
| 2025 | SP-Det: Anchor-based lane detection network with structural prior perception
Libo Sun 0001, Hangyu Zhu, Wenhu Qin |
Pattern Recognit. Lett. | 1 |
| 2025 | PEPillar: a point-enhanced pillar network for efficient 3D object detection in autonomous driving
Libo Sun 0001, Wenhu Qin |
Vis. Comput. | 1 |
| 2025 | Adaptive formation control and transformation in virtual crowds via deep reinforcement learning
Libo Sun 0001, Yongchun Qiu, Wenhu Qin |
Vis. Comput. | 1 |
| 2025 | The crowd cooperation approach for formation maintenance and collision avoidance using multi-agent deep reinforcement learning
Libo Sun 0001, Jiahui Yan, Yongchun Qiu, Wenhu Qin |
Vis. Comput. | 1 |
| 2025 | ACL-SAR: model agnostic adversarial contrastive learning for robust skeleton-based action recognition
Jiaxuan Zhu, Ming Shao, Libo Sun 0001, Si-Yu Xia |
Vis. Comput. | 3 |
| 2024 | Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action RecognitionabstractHuman action video recognition has recently attracted more attention in applications such as video security and sports posture correction. Popular solutions, including graph convolutional networks (GCNs) that model the human skeleton as a spatiotemporal graph, have proven very effective. GCNs-based methods with stacked blocks usually utilize top-layer semantics for classification/annotation purposes. Although the global features learned through the procedure are suitable for the general classification, they have difficulty capturing fine-grained action change across adjacent frames - decisive factors in sports actions. In this paper, we propose a novel “Cross-block Fine-grained Semantic Cascade (CFSC)” module to overcome this challenge. In summary, the proposed CFSC progressively integrates shallow visual knowledge into high-level blocks to allow networks to focus on action details. In particular, the CFSC module utilizes the GCN feature maps produced at different levels, as well as aggregated features from proceeding levels to consolidate fine-grained features. In addition, a dedicated temporal convolution is applied at each level to learn short-term temporal features, which will be carried over from shallow to deep layers to maximize the leverage of low-level details. This cross-block feature aggregation methodology, capable of mitigating the loss of fine-grained information, has resulted in improved performance. Last, FD-7, a new action recognition dataset for fencing sports, was collected and will be made publicly available. Experimental results and empirical analysis on public benchmarks (FSD-10) and self-collected (FD-7) demonstrate the advantage of our CFSC module on learning discriminative patterns for action classification over others. Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia |
FG | 4 |
| 2024 | Sketch3D: Style-Consistent Guidance for Sketch-to-3D Generation
Wangguandong Zheng, Haifeng Xia, Libo Sun 0001, Ming Shao, Si-Yu Xia, Zhengming Ding |
ACM Multimedia | 4 |
| 2024 | Physical based motion reconstruction from videos using musculoskeletal modelabstractAbstract We propose a novel method that combines human pose estimation and physical simulation of character animation. Our approach allows characters to learn from the actor's skills captured in videos and subsequently reconstruct the motions with high fidelity in a physically simulated environment. Firstly, we model the character based on the human musculoskeletal system and build a complete dynamics model of the proposed system using the Lagrange equations of motion. Next, we employ the pose estimation method to process the input video and generate human reference motion. Finally, we design a hierarchical control framework consisting of a trajectory tracking layer and a muscle control layer. The trajectory tracking layer aims to minimize the difference between the reference motion pose and the actual output pose, while the muscle control layer aims to minimize the difference between the target torque and the actual output muscle force. The two layers interact by passing parameters through a proportional differential controller until the desired learning objective is achieved. A series of complex experimental results demonstrate that our proposed method can learn to produce comparable high‐quality motions with high similarity from videos of different complexity levels and remains stable in the presence of muscle contracture weakness perturbations. Libo Sun 0001, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2024 | A language-directed virtual human motion generation approach based on musculoskeletal modelsabstractAbstract The development of the systems capable of synthesizing natural and life‐like motions for virtual characters has long been a central focus in computer animation. It needs to generate high‐quality motions for characters and provide users with a convenient and flexible interface for guiding character motions. In this work, we propose a language‐directed virtual human motion generation approach based on musculoskeletal models to achieve interactive and higher‐fidelity virtual human motion, which lays the foundation for the development of language‐directed controllers in physics‐based character animation. First, we construct a simplified model of musculoskeletal dynamics for the virtual character. Subsequently, we propose a hierarchical control framework consisting of a trajectory tracking layer and a muscle control layer, obtaining the optimal control policy for imitating the reference motions through the training. We design a multi‐policy aggregation controller based on large language models, which selects the motion policy with the highest similarity to user text commands from the action‐caption data pool, facilitating natural language‐based control of virtual character motions. Experimental results demonstrate that the proposed approach not only generates high‐quality motions highly resembling reference motions but also enables users to effectively guide virtual characters to perform various motions via natural language instructions. Libo Sun 0001, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2024 | AG-YOLO: Attention-guided network for real-time object detection
Hangyu Zhu, Libo Sun 0001, Wenhu Qin, Feng Tian 0006 |
Multim. Tools Appl. | 2 |
| 2023 | Bidirectional temporal feature for 3D human pose and shape estimation from a videoabstractAbstract 3D human pose and shape estimation is the foundation of analyzing human motion. However, estimating accurate and temporally consistent 3D human motion from a video remains a challenge. By now, most of the video‐based methods for estimating 3D human pose and shape rely on unidirectional temporal features and lack more comprehensive information. To solve this problem, we propose a novel model “bidirectional temporal feature for human motion recovery” (BTMR), which consists of a human motion generator and a discriminator. The transformer‐based generator effectively captures the forward and reverse temporal features to enhance the temporal correlation between frames and reduces the loss of spatial information. The motion discriminator based on Bi‐LSTM can distinguish whether the generated pose sequences are consistent with the realistic sequences of the AMASS dataset. In the process of continuous generation and discrimination, the model can learn more realistic and accurate poses. We evaluate our BTMR on 3DPW and MPI‐INF‐3DHP datasets. Without the training set of 3DPW, BTMR outperforms VIBE by 2.4 mm and 14.9 mm/s2 in PA‐MPJPE and Accel metrics and outperforms TCMR by 1.7 mm in PA‐MPJPE metric on 3DPW. The results demonstrate that our BTMR produces better accurate and temporal consistent 3D human motion. Libo Sun 0001, Ting Tang, Yuke Qu, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2023 | Path planning for multiple agents in an unknown environment using soft actor critic and curriculum learningabstractAbstract Path planning can guarantee that agents reach their goals without colliding with obstacles and other agents in an optimal way and it is a very important component in the research of crowd simulation. In this article, we propose a novel path planning approach for multiple agents which combines soft actor critic (SAC) algorithm and curriculum learning to solve the problems of single policy, slow convergence of the policy in an unknown environment with sparse rewards. The path planning task is set as lessons from easy to difficult, and the neural network of the SAC algorithm is arranged to learn in sequence, and finally the neural network can be fully competent for the path planning task. We also stack the state information to address the problems caused by limited observation for policy learning, and design a comprehensive reward function to make agents reach their goals successfully and avoid collisions with static obstacles and other agents. The experimental results demonstrate that our approach can plan smooth and natural paths for multiple agents, and furthermore, our model has a certain generalization ability and a better adaptability to the changes in a dynamic environment. Libo Sun 0001, Jiahui Yan, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2022 | A Pig Pose Estimation Model for Measuring Pig's Body Size
Yukun Yang 0004, Wenhu Qin, Libo Sun 0001, Weipeng Shi |
CGI | 3 |
| 2022 | Muscle-driven virtual human motion generation approach based on deep reinforcement learningabstractAbstract We propose a muscle‐driven motion generation approach to realize virtual human motion with user interaction and higher fidelity, which can address the problem that the joint‐driven fails to reflect the motion process of the human body. First, a simplified virtual human musculoskeletal model is built based on human biomechanics. Then, a hierarchical policy learning framework is constructed including motion tracking layer, SPD controller and muscle control layer. The motion tracking layer is responsible for mimicking reference motion and completing control command, using proximal policy optimization to train the policy; the muscle control layer is aimed to minimize muscle energy consumption and train the policy based on supervised learning; the SPD controller acts as a link between the two layers. At the same time, we integrate the curriculum learning to improve the efficiency and success rate of policy training. Simulation experiments show that the proposed approach can use motion capture data and pose estimation data as reference motions to generate better and more adaptable motions. Furthermore, the virtual human has the ability to respond to the user control command during the motion, and can complete the target task successfully. Wenhu Qin, Libo Sun 0001, Kaiyue Dong |
Comput. Animat. Virtual Worlds | 3 |
| 2022 | Crowd navigation in an unknown and complex environment based on deep reinforcement learningabstractAbstract We propose a virtual crowd navigation approach based on deep reinforcement learning to improve the adaptability of virtual crowds in an unknown and complex environment. To address the problem of local optimum or slow iteration or even failure to converge due to sparse rewards in complex environments, we integrate the curiosity‐driven mechanism, the key navigation points acquisition and the failure path penalty method in addition to combining long short‐term memory networks, dynamic obstacle collision prediction with proximal policy optimization algorithms, which realizes the crowd navigation in a complex environment. The experimental results show that the proposed approach can simulate the motions of virtual crowds in various dynamic and complex scenarios without the environment modeling. The use of continuous action space also ensures that the movement trajectories of the virtual crowds are more realistic and natural. Furthermore, our approach can provide analysis and demonstration tools for a variety of independent collaborations such as competition and cooperation of group intelligence in an open and dynamic environment. Libo Sun 0001, Yuke Qu, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2022 | Research on target recognition and tracking in mobile augmented reality assisted maintenanceabstractAbstract The recognition and tracking of maintenance targets is the basis of augmented reality assisted maintenance. To address the problems of the current recognition and tracking algorithms, such as high time complexity, low pose tracking accuracy, and high requirements for hardware equipment, this paper studies a light weight maintenance target recognition algorithm based on YOLOv5s, and then the VI ORB‐SLAM algorithm is adopted to track the maintenance target. In addition, the ORB feature extraction and the visual inertial initialization are improved for the VI ORB‐SLAM algorithm. Finally, the combined algorithm is deployed on the mobile phone. Taking the augmented reality assisted vehicle maintenance as an example, it is verified that the proposed approach is practical feasible and effective in the actual assisted maintenance scene. Libo Sun 0001, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Planning Feasible and Smooth Paths for Simulating Realistic Crowd
Libo Sun 0001, Wenhu Qin |
ICIC (1) | 1 |
| 2013 | Simulating realistic crowd based on agent trajectoriesabstractABSTRACT This paper presents a model for simulating realistic crowd behaviors at low computation cost. The proposed model is inspired by video data. In our approach, we first classify the crowd into two categories: main and background characters. Whether the agents are main characters or not is influenced by two factors, one is the agent's trajectories and the other one is the change of the environment. In the second stage, we adopt two approaches to simulate the behaviors of main and background characters. Main characters are intelligent agents with the perception, the memory, the planning, and the psychology so that they can make decisions themselves. Background characters are informed of the behavior options for execution by the “smart environment.” Finally, we simulate the road‐crossing scenario in a three‐dimensional virtual environment. The experimental results demonstrate that our approach not only well reflects the characteristics of agent behaviors but also reduces the computation complexity of simulating realistic crowd. Copyright © 2013 John Wiley & Sons, Ltd. Libo Sun 0001, Wenhu Qin |
Comput. Animat. Virtual Worlds | 1 |
| 2012 | Animating synthetic dyadic conversations with variations based on context and agent attributesabstractABSTRACT Conversations between two people are ubiquitous in many inhabited contexts. The kinds of conversations that occur depend on several factors, including the time, the location of the participating agents, the spatial relationship between the agents, and the type of conversation in which they are engaged. The statistical distribution of dyadic conversations among a population of agents will therefore depend on these factors. In addition, the conversation types, flow, and duration will depend on agent attributes such as interpersonal relationships, emotional state, personal priorities, and socio‐cultural proxemics. We present a framework for distributing conversations among virtual embodied agents in a real‐time simulation. To avoid generating actual language dialogues, we express variations in the conversational flow by using behavior trees implementing a set of conversation archetypes. The flow of these behavior trees depends in part on the agents' attributes and progresses based on parametrically estimated transitional probabilities. With the participating agents' state, a ‘smart event’ model steers the interchange to different possible outcomes as it executes. Example behavior trees are developed for two conversation archetypes: buyer–seller negotiations and simple asking–answering; the model can be readily extended to others. Because the conversation archetype is known to participating agents, they can animate their gestures appropriate to their conversational state. The resulting animated conversations demonstrate reasonable variety and variability within the environmental context. Copyright © 2012 John Wiley & Sons, Ltd. Libo Sun 0001, Alexander Shoulson, Nicole Nelson, Wenhu Qin, Ani Nenkova, Norman I. Badler |
Comput. Animat. Virtual Worlds | 1 |
| 2010 | Smart Events and Primed Agents
Catherine Stocker, Libo Sun 0001, Wenhu Qin, Jan M. Allbeck, Norman I. Badler |
IVA | 2 |