VLDB 2026 Research / reviewers in the wild / expert
Ronghui Li
dblp:52/9777
· DBLP profile ↗
30ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Safety-Critical Pursuit-Evasion Game of Multiple Autonomous Surface Vehicles Based on Min-Max Optimization and Neural Network Dynamic ControlabstractThis article investigates the pursuit–evasion problem of multiple underactuated autonomous surface vehicles (ASVs) under velocity and collision avoidance constraints. A safety–critical pursuit–evasion game (PEG) method based on min–max optimization and neural network dynamic control is proposed. Specifically, an allocation strategy is designed at first based on the position information of the pursuing and evading ASVs to achieve a rational and efficient allocation of pursuit targets by minimizing pursuit distances. Next, a nominal PEG guidance law is proposed by combining model predictive control (MPC) with min–max optimization methods. Then, the nominal guidance law is optimized based on a heading-constrained control barrier function (CBF) such that a safety–critical guidance law for collision avoidance can be achieved. Finally, a predicator-based neural network is developed to estimate the uncertainty and external disturbance, and a dynamic control law is proposed to track the guidance signals without using any model parameters. It is proven that the closed-loop system is input–to–state stable (ISS), and the ASV system is safe. A robot-operating-system (ROS)-based simulation results demonstrate the effectiveness of the proposed safety–critical PEG method based on min–max optimization and neural network dynamic control. Ronghui Li, Nan Gu, Dan Wang 0001, Zhouhua Peng, Weidong Zhang 0004 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | Harmonious Music-driven Group Choreography with Trajectory-Controllable DiffusionabstractCreating group choreography from music is crucial in cultural entertainment and virtual reality, with a focus on generating harmonious movements. Despite growing interest, recent approaches often struggle with two major challenges: multi-dancer collisions and single-dancer foot sliding. To address these challenges, we propose a Trajectory-Controllable Diffusion (TCDiff) framework, which leverages non-overlapping trajectories to ensure coherent and aesthetically pleasing dance movements. To mitigate collisions, we introduce a Dance-Trajectory Navigator that generates collision-free trajectories for multiple dancers, utilizing a distance-consistency loss to maintain optimal spacing. Furthermore, to reduce foot sliding, we present a footwork adaptor that adjusts trajectory displacement between frames, supported by a relative forward-kinematic loss to further reinforce the correlation between movements and trajectories. Experiments demonstrate our method's superiority. Yuqin Dai, Wanlu Zhu, Ronghui Li, Zeping Ren, Xiangzheng Zhou, Jixuan Ying, Jun Li 0027, Jian Yang 0003 |
AAAI | 3 |
| 2025 | AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision RewardabstractRecently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex relationship between textual prompts and desired motion outcomes. To address this, we introduce AToM, a framework that enhances the alignment between generated motion and text prompts by leveraging reward from GPT-4Vision. AToM comprises three main stages: Firstly, we construct a dataset MotionPreferthat pairs three types of event-level textual prompts with generated motions, which cover the integrity, temporal relationship and frequency of motion. Secondly, we design a paradigm that utilizes GPT-4Vision for detailed motion annotation, including visual data formatting, task-specific instructions and scoring rules for each sub-task. Finally, we fine-tune an existing text-to-motion model using reinforcement learning guided by this paradigm. Experimental results demonstrate that AToM significantly improves the event-level alignment quality of text-to-motion generation. Project page is available at https://atom-motion.github.io/. Haonan Han, Xiangzuo Wu, Huan Liao, Zunnan Xu, Zhongyuan Hu, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
CVPR | 6 |
| 2025 | Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
Ronghui Li, Shukai Fang, Shuzhao Xie, Jiaqing Zhou, Junkun Peng |
ICCV | 2 |
| 2025 | A Plug-And-Play Physical Motion Restoration Approach for In-The-Wild High-Difficulty MotionsabstractExtracting physically plausible 3D human motion from videos is a critical task. Although existing simulation-based motion imitation methods can enhance the physical quality of daily motions estimated from monocular video capture, extending this capability to high-difficulty motions remains an open challenge. This can be attributed to some flawed motion clips in video-based motion capture results and the inherent complexity in modeling high-difficulty motions. Therefore, sensing the advantage of segmentation in localizing human body, we introduce a mask-based motion correction module (MCM) that leverages motion context and video mask to repair flawed motions, producing imitation-friendly motions; and propose a physics-based motion transfer module (PTM), which employs a pretrain and adapt approach for motion imitation, improving physical plausibility with the ability to handle in-the-wild and challenging motions. Our approach is designed as a plug-and-play module to physically refine the video motion capture results, including high-difficulty in-the-wild motions. Finally, to validate our approach, we collected a challenging in-the-wild test set to establish a benchmark, and our method has demonstrated effectiveness on both the new benchmark and existing public datasets.https://physicalmotionrestoration.github.io Youliang Zhang, Ronghui Li, Yachao Zhang 0001, Liang Pan, Yebin Liu, Xiu Li 0001 |
ICCV | 2 |
| 2025 | A Motion is Worth a Hybrid Sentence: Taming Language Model for Unified Motion Generation by Fine-grained PlanningabstractExisting LLM-based motion models fail to fully leverage large models' planning capabilities for motion-related tasks, exhibiting poor generalization, limited text-motion alignment, and an inability to perform multimodal condition joint driven motion generation. We argue that these issues arise from the modality gap and the highly coupled nature of motion tokens. To address this, we proposed the hybrid motion sentence, which is consistant of fine-grained motion decription and atomic body-part motion token that can bridge the gap between motion and text. To obtain a large corpus of hybrid motion sentences, we introduced a novel motion-to-text generation method that combines atomic motion operators with GPT-4o, resulting in 68.2 million fine-grained textual descriptions across diverse modalities. To reconstruct high-quality motion from hybrid sentences and make better motion-text alignment, we introduce Semantic-Aware Decoupled Motion Tokenization. Furthermore, we propose MotionUPG based on LLaMA, leveraging MotionWords dataset for both pretraining and instruction tuning. Our method achieves strong fine-grained text-motion alignment, impressive zero-shot motion generation, and is the first to support multimodal condition joint driven motion generation tasks. Ronghui Li, Lingxiao Han, Shi Shu, Yueyao Liu, Yukang Lin, Yue Ma 0016, Ziwei Liu 0002, Xiu Li 0001 |
ACM Multimedia | 1 |
| 2025 | InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong 0001, Zunnan Xu, Xindi Li, Chuanbiao Song, Ronghui Li, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Xiu Li 0001 |
ACM Multimedia | 7 |
| 2025 | Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
Zihao Liu 0006, Mingwen Ou, Zunnan Xu, Jiaqi Huang 0003, Haonan Han, Ronghui Li, Xiu Li 0001 |
ACM Multimedia | 6 |
| 2025 | Large multimodal models evaluation: a survey
Farong Wen, Yijin Guo, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Jiahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun 0029, Zijian Chen 0001, Kaiwei Zhang, Yuqin Cao, Yue Zhou 0005, Xuemei Zhou, Juntai Cao, Wei Zhou 0021, Jinyu Cao, Ronghui Li, Yuan Tian 0017, Chunyi Li 0001, Haoning Wu 0001, Xiaohong Liu 0001, Junjun He, Yu Zhou 0016, Zesheng Wang 0004, Huiyu Duan, Yingjie Zhou 0003, Xiongkuo Min, Dongzhan Zhou, Jiezhang Cao, Xue Yang 0005, Junzhi Yu 0001, Songyang Zhang 0001, Haodong Duan, Guangtao Zhai |
Sci. China Inf. Sci. | 26 |
| 2025 | FS-Net collaborative flattening transformation for high performance 3D-to-2D vessel segmentation in OCTA images
Jiwei Xing, Songyi Jiang, Dongbei Guo, Jinde Zhang, Peishan Li, Linyan Xue, Valery V. Tuchin, Xiufei Gu, Ronghui Li, Qingliang Zhao |
Neurocomputing | 11 |
| 2025 | CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis With Multimodal DiffusionabstractAlthough remarkable progress has been made in image style transfer, style is just one of the components of artistic paintings. Directly transferring extracted style features to natural images often results in outputs with obvious synthetic traces. This is because key painting attributes including layout, perspective, shape, and semantics often cannot be conveyed and expressed through style transfer. Large-scale pretrained text-to-image generation models have demonstrated their capability to synthesize a vast amount of high-quality images. However, even with extensive textual descriptions, it is challenging to fully express the unique visual properties and details of paintings. Moreover, generic models often disrupt the overall artistic effect when modifying specific areas, making it more complicated to achieve a unified aesthetic in artworks. Our main novel idea is to integrate multimodal semantic information as a synthesis guide into artworks, rather than transferring style to the real world. We also aim to reduce the disruption to the harmony of artworks while simplifying the guidance conditions. Specifically, we propose an innovative multi-task unified framework called CreativeSynth, based on the diffusion model with the ability to coordinate multimodal inputs. CreativeSynth combines multimodal features with customized attention mechanisms to seamlessly integrate real-world semantic content into the art domain through Cross-Art-Attention for aesthetic maintenance and semantic fusion. We demonstrate the results of our method across a wide range of different art categories, proving that CreativeSynth bridges the gap between generative models and artistic expression. Nisha Huang, Weiming Dong, Yuxin Zhang 0006, Fan Tang, Ronghui Li, Chongyang Ma, Xiu Li 0001, Tong-Yee Lee, Changsheng Xu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlabstractThis study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance. Zunnan Xu, Yachao Zhang 0001, Ronghui Li, Xiu Li 0001 |
AAAI | 4 |
| 2024 | Cross-Modal Match for Language Conditioned 3D Object GroundingabstractLanguage conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion mechanism or bridging the gap between detection and matching. However, several mismatches are ignored, i.e., mismatch in local visual representation and global sentence representation, and mismatch in visual space and corresponding label word space. In this paper, we propose crossmodal match for 3D grounding from mitigating these mismatches perspective. Specifically, to match local visual features with the global description sentence, we propose BEV (Bird’s-eye-view) based global information embedding module. It projects multiple object proposal features into the BEV and the relations of different objects are accessed by the visual transformer which can model both positions and features with long-range dependencies. To circumvent the mismatch in feature spaces of different modalities, we propose crossmodal consistency learning. It performs cross-modal consistency constraints to convert the visual feature space into the label word feature space resulting in easier matching. Besides, we introduce label distillation loss and global distillation loss to drive these matches learning in a distillation way. We evaluate our method in mainstream evaluation settings on three datasets, and the results demonstrate the effectiveness of the proposed method. Yachao Zhang 0001, Runze Hu, Ronghui Li, Yanyun Qu, Yuan Xie 0006, Xiu Li 0001 |
AAAI | 3 |
| 2024 | Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance PrimitivesabstractWe propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method. Code, models, and demonstrative video results are available at: https://li-ronghui.github.io/lodge Ronghui Li, Yuxiang Zhang 0006, Yachao Zhang 0001, Hongwen Zhang 0001, Yan Zhang 0002, Yebin Liu, Xiu Li 0001 |
CVPR | 1 |
| 2024 | Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable AttributeabstractGenerating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address these issues, we propose Text2Avatar, which can generate realistic-style 3D avatars based on the coupled text prompts. Text2Avatar leverages a discrete codebook as an intermediate feature to establish a connection between text and avatars, enabling the disentanglement of features. Furthermore, to alleviate the scarcity of realistic style 3D human avatar data, we utilize a pre-trained unconditional 3D human avatar generation model to obtain a large amount of 3D avatar pseudo data, which allows Text2Avatar to achieve realistic style generation. Experimental results demonstrate that our method can generate realistic 3D avatars from coupled textual data, which is challenging for other existing methods in this field. Chaoqun Gong, Yuqin Dai, Ronghui Li, Achun Bao, Jun Li 0027, Jian Yang 0003, Yachao Zhang 0001, Xiu Li 0001 |
ICASSP | 3 |
| 2024 | Exploring Multi-Modal Control in Music-Driven Dance GenerationabstractExisting music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal control, including genre control, semantic control, and spatial control. First, we decouple the dance generation network from the dance control network, thereby avoiding the degradation in dance quality when adding additional control information. Second, we design specific control strategies for different control information and integrate them into a unified framework. Experimental results show that the proposed dance generation framework outperforms state-of-the-art methods in terms of motion quality and controllability. Ronghui Li, Yuqin Dai, Yachao Zhang 0001, Jun Li 0027, Jian Yang 0003, Xiu Li 0001 |
ICASSP | 1 |
| 2024 | MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsabstractGesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality.
Recent advancements have utilized the diffusion model to improve gesture synthesis.
However, the high computational complexity of these techniques limits the application in reality.
In this study, we explore the potential of state space models (SSMs).
Direct application of SSMs in gesture synthesis encounters difficulties, which stem primarily from the diverse movement dynamics of various body parts.
The generated gestures may also exhibit unnatural jittering issues.
To address these, we implement a two-stage modeling strategy with discrete motion priors to enhance the quality of gestures.
Built upon the selective scan mechanism, we introduce MambaTalk, which integrates hybrid fusion modules, local and global scans to refine latent space representations.
Subjective and objective experiments demonstrate that our method surpasses the performance of state-of-the-art models. Our project is publicly available at~\url{https://kkakkkka.github.io/MambaTalk/}. Zunnan Xu, Yukang Lin, Haonan Han, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
NeurIPS | 5 |
| 2024 | EPRD: Exploiting prior knowledge for evidence-providing automatic rumor detection
Jiawen Li 0003, Ronghui Li, Shiwen Ni, Hung-Yu Kao |
Neurocomputing | 2 |
| 2024 | A Novel Adaptive Control Design for a Class of Nonstrict-Feedback Discrete-Time Systems via Reinforcement LearningabstractIn this article, an adaptive reinforcement learning (RL) control problem is explored for a class of nonstrict-feedback discrete-time systems. First, different from the existing results, considering the noncausal problem which may exist in the backstepping design procedure, a universal system transformation method is first proposed for a class of nonstrict-feedback discrete-time systems. Second, by defining a compensation term to compensate the controller and utilizing the property of radial-basis-function neural network (RBFNN), an RL-based direct adaptive control strategy is developed via a backstepping method to achieve optimal control, and the multigradient recursive (MGR) algorithm is employed to estimate the weight vector. Finally, the stability of the control system is guaranteed and all signals in the closed-loop system are semiglobal uniformly ultimately bounded (SGUUB) on the basis of the Lyapunov theory. In addition, a universal system transformation is first proposed which breaks through the limitations on the controller design for the discrete-time nonstrict-feedback nonlinear system by using the traditional method. The validity of this strategy is verified by two simulation examples that include a course keeping system of the marine vessel. Weiwei Bai, Tieshan Li 0001, Yue Long 0002, C. L. Philip Chen, Yang Xiao 0001, Wenjiang Li, Ronghui Li |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2023 | FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationabstractGenerating full-body and multi-genre dance sequences from given music is a challenging task, due to the limitations of existing datasets and the inherent complexity of the fine-grained hand motion and dance genres. To address these problems, we propose FineDance, which contains 14.6 hours of music-dance paired data, with fine-grained hand motions, fine-grained genres (22 dance genres), and accurate posture. To the best of our knowledge, FineDance is the largest music-dance paired dataset with the most dance genres. Additionally, to address monotonous and unnatural hand movements existing in previous methods, we propose a full-body dance generation network, which utilizes the diverse generation capabilities of the diffusion model to solve monotonous problems, and use expert nets to solve unreal problems. To further enhance the genre-matching and long-term stability of generated dances, we propose a Genre&Coherent aware Retrieval Module. Besides, we propose a novel metric named Genre Matching Score to evaluate the genre-matching degree between dance and music. Quantitative and qualitative experiments demonstrate the quality of FineDance, and the state-of-the-art performance of FineNet. The FineDance Dataset and more qualitative samples can be found at website. Ronghui Li, Junfan Zhao, Yachao Zhang 0001, Mingyang Su, Zeping Ren, Yansong Tang, Xiu Li 0001 |
ICCV | 1 |
| 2023 | RICH: Robust Implicit Clothed Humans Reconstruction from Multi-scale Spatial Cues
Yukang Lin, Ronghui Li, Kedi Lyu, Yachao Zhang 0001, Xiu Li 0001 |
PRCV (2) | 2 |
| 2023 | Competitive advantage assessment for container shipping liners using a novel hybrid method with intuitionistic fuzzy linguistic variables
Junzhong Bao, Yuanzi Zhou, Ronghui Li |
Neural Comput. Appl. | 3 |
| 2023 | Event-triggered fixed-time adaptive neural formation control for underactuated ASVs with connectivity constraints and prescribed performance
Haitao Liu 0004, Jianfei Lin, Ronghui Li, Xuehong Tian, Qingqun Mai |
Neural Comput. Appl. | 3 |
| 2022 | Sign language recognition and translation network based on multi-view data
Ronghui Li, Lu Meng |
Appl. Intell. | 1 |
| 2022 | Detection of cell markers from single cell RNA-seq with sc2markerabstractBACKGROUND: Single-cell RNA sequencing (scRNA-seq) allows the detection of rare cell types in complex tissues. The detection of markers for rare cell types is useful for further biological analysis of, for example, flow cytometry and imaging data sets for either physical isolation or spatial characterization of these cells. However, only a few computational approaches consider the problem of selecting specific marker genes from scRNA-seq data. RESULTS: Here, we propose sc2marker, which is based on the maximum margin index and a database of proteins with antibodies, to select markers for flow cytometry or imaging. We evaluated the performances of sc2marker and competing methods in ranking known markers in scRNA-seq data of immune and stromal cells. The results showed that sc2marker performed better than the competing methods in accuracy, while having a competitive running time. Ronghui Li, Bella Banjanin, Rebekka K. Schneider, Ivan G. Costa |
BMC Bioinform. | 1 |
| 2019 | Design of Optimal Ship Steering Active Disturbance Rejection Controller Based on Adaptive Particle Swarm OptimizationabstractIn this paper, a novel active disturbance rejection controller (ADRC) of ship steering control is designed under the internal parameter uncertainties and external disturbances. The ship heading control model is rather a complex model with large time delay, uncertainties and serious external disturbances, which may result in poor control performance. Therefore, ADRC is applied to control the steering of the ship. However, manually tuning the parameters of the ADRC is timeconsuming and labor-intensive. An optimization method based on adaptive particle swarm optimization (APSO) to adjust the ship’s steering ADRC parameters is proposed. In the developed APSO, considering the control performance, the fitness function is modified in the light of the integrated time absolute error (ITAE) criteria. Then the inertia weight of the previous velocity in the current velocity update is varied based on the fitness value of each individual at the previous moment. Finally, the effectiveness of the proposed adaptive optimization method is proved by simulation experiments. Junhai Cao, Yimin Zhou 0001, Ronghui Li, Juncheng Zhu |
CEC | 3 |
| 2016 | An adaptive neural network approach for ship roll stabilization via fin control
Ronghui Li, Tieshan Li 0001, Weiwei Bai, Xian Du |
Neurocomputing | 1 |
| 2013 | Active Disturbance Rejection Control on Path Following for Underactuated Ships
Ronghui Li, Tieshan Li 0001, Qingling Zheng, Xiaori Gao |
ISNN (2) | 1 |
| 2012 | Adaptive neural control of nonlinear MIMO systems with unknown time delays
Tieshan Li 0001, Ronghui Li, Dan Wang 0001 |
Neurocomputing | 2 |
| 2011 | Decentralized adaptive neural control of nonlinear interconnected large-scale systems with unknown time delays and input saturation
Tieshan Li 0001, Ronghui Li, Junfang Li |
Neurocomputing | 2 |