Wenqi Zhang 0001

dblp:16/5404-1 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0002-8312-0184ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Reality vs Counterfactual: Multi-World Contrastive Reinforcement Learning for Enhancing MLLM's Theory of Mind in Egocentric Videos
abstract
Theory of Mind (ToM) refers to the ability to infer others' mental states, which is an essential capability for embodied AI agents to effectively collaborate and interact with humans. While improving Large Language Models' ability to reason about characters' mental states in text-based stories/dialogues has been extensively studied, enhancing Multimodal Large Language Models' ToM capabilities, particularly in egocentric video from an embodied perspective, remains unexplored. In this paper, we propose a contrastive Reinforcement Learning (RL) paradigm that explicitly encourages models to leverage temporal and causal evolutionary patterns in user action sequences to infer user's mental states (goals, beliefs, and potential next actions). Evaluation results on in-domain and out-of-domain demonstrate that our method achieves performance improvements of (+30.00%, +2.00%) and (+5.83%, +5.00%) compared to the backbone model and vanilla Group Relative Policy Optimization (GRPO) model, respectively. Additionally, we compare the performance of two post-training paradigms (Supervise Fine-Tuning and RL) and systematically analyze the reasoning trajectories across the base model, vanilla GRPO model, and our proposed method.
Guiyang Hou, Yihui Fu, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001
AAAI6
2026 GUI-G²: Gaussian Reward Modeling for GUI Grounding
abstract
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
AAAI8
2026 UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
abstract
Zhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001
ACL (1)7
2026 CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
abstract
Teng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang, Weiming Lu, Jun Xiao, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Teng Pan, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yongliang Shen 0001
ACL (1)6
2026 Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
abstract
Wenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenqi Zhang 0001, Mengna Wang, Gangao Liu, Huixin Xu, Yongliang Shen 0001, Guiyang Hou, Xin Li 0056, Weiming Lu 0001, Peng Li 0031, Yueting Zhuang
ACL (1)1
2025 STaR-SQL: Self-Taught Reasoner for Text-to-SQL
abstract
Generating step-by-step "chain-of-thought" rationales has proven effective for improving the performance of large language models on complex reasoning tasks.However, applying such techniques to structured tasks, such as text-to-SQL, remains largely unexplored.In this paper, we introduce Self-Taught Reasoner for text-to-SQL (STaR-SQL), a novel approach that reframes SQL query generation as a reasoningdriven process.Our method prompts the LLM to produce detailed reasoning steps for SQL queries and fine-tunes it on rationales that lead to correct outcomes.Unlike traditional methods, STaR-SQL dedicates additional test-time computation to reasoning, thereby positioning LLMs as spontaneous reasoners rather than mere prompt-based agents.To further scale the inference process, we incorporate an outcomesupervised reward model (ORM) as a verifier, which enhances SQL query accuracy.Experimental results on the challenging Spider benchmark demonstrate that STaR-SQL significantly improves text-to-SQL performance, achieving an execution accuracy of 86.6%.This surpasses a few-shot baseline by 31.6% and a baseline fine-tuned to predict answers directly by 18.0%.Additionally, STaR-SQL outperforms agent-like prompting methods that leverage more powerful yet closed-source models such as GPT-4.These findings underscore the potential of reasoning-augmented training for structured tasks and open the door to extending self-improving reasoning models to text-to-SQL generation and beyond.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
ACL (1)3
2025 ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
abstract
The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code is available at https://github.com/RhDang/ECBench.
Ronghao Dang, Yuqian Yuan, Wenqi Zhang 0001, Yifei Xin, Boqiang Zhang, Liuyi Wang, Qinyang Zeng, Xin Li 0056, Lidong Bing
CVPR3
2025 AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
abstract
Xuan Zhang, Yongliang Shen, Zhe Zheng, Linjuan Wu, Wenqi Zhang, Yuchen Yan, Qiuying Peng, Jun Wang, Weiming Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yongliang Shen 0001, Linjuan Wu, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
EMNLP5
2025 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
abstract
Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook.
Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing
ICCV1
2025 SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
abstract
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius.
Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia9
2025 Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
abstract
Large language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from **Thought Leaps** due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called **ScaleQM+**, based on the structured ScaleQuestMath dataset, and trained **CoT-Bridge** to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87\% on NuminaMath. Our approach effectively enhances distilled data (+3.02\%) and provides better starting points for reinforcement learning (+3.1\%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrates improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits.
Haolei Xu, Yongliang Shen 0001, Wenqi Zhang 0001, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS4
2025 Let LRMs Break Free from Overthinking via Self-Braking Tuning
abstract
Large reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking. Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions. In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point. Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models.
Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS5
2024 Learning Global Controller in Latent Space for Parameter-Efficient Fine-Tuning
abstract
Zeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong, Wenqi Zhang, Jian Shao, Weiming Lu, Yueting Zhuang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zeqi Tan, Yongliang Shen 0001, Xiaoxia Cheng, Chang Zong, Wenqi Zhang 0001, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang
ACL (1)5
2024 Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
abstract
Wenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001
ACL (1)1
2024 Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
abstract
Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Mengna Wang, Yongliang Shen 0001, Guiyang Hou, Zeqi Tan, Peng Li 0031, Yueting Zhuang, Weiming Lu 0001
ACL (1)1
2024 Advancing Process Verification for Large Language Models via Tree-Based Preference Learning
abstract
Large Language Models (LLMs) have demonstrated remarkable potential in handling complex reasoning tasks by generating step-by-step rationales.Some methods have proven effective in boosting accuracy by introducing extra verifiers to assess these paths.However, existing verifiers, typically trained on binarylabeled reasoning paths, fail to fully utilize the relative merits of intermediate steps, thereby limiting the effectiveness of the feedback provided.To overcome this limitation, we propose Tree-based Preference Learning Verifier (Tree-PLV), a novel approach that constructs reasoning trees via a best-first search algorithm and collects step-level paired data for preference training.Compared to traditional binary classification, step-level preferences more finely capture the nuances between reasoning steps, allowing for a more precise evaluation of the complete reasoning path.We empirically evaluate Tree-PLV across a range of arithmetic and commonsense reasoning tasks, where it significantly outperforms existing benchmarks.For instance, Tree-PLV achieved substantial performance gains over the Mistral-7B selfconsistency baseline on GSM8K (67.55% → 82.79%), MATH (17.00% → 26.80%), CSQA (68.14% → 72.97%), and StrategyQA (82.86% → 83.25%).Additionally, our study explores the appropriate granularity for applying preference learning, revealing that step-level guidance provides feedback that better aligns with the evaluation of the reasoning process.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Zeqi Tan, Weiming Lu 0001
EMNLP3
2024 Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
abstract
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenqi Zhang 0001, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen 0001, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu 0001, Yueting Zhuang
EMNLP1
2024 TaskBench: Benchmarking Large Language Models for Task Automation
abstract
In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents.
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Kan Ren, Weiming Lu 0001, Dongsheng Li 0002, Yueting Zhuang
NeurIPS4
2024 Specialized Mathematical Solving by a Step-By-Step Expression Chain Generation
abstract
Math Solving requires both semantic understanding and relation reasoning. Most current approaches treat it as a translation task from natural language to mathematical symbols, generating tokens one by one. However, token-level generation is usually vulnerable when confronted with diverse annotations and complex reasoning. We consider the equation is an ordered combination of multiple sub-expressions, and math reasoning should be performed on the sub-expression level rather than the token level. We treat sub-expression as a minimum generative unit and minimum reasoning node. At each step, candidate sub-expression nodes are generated in parallel, and the whole reasoning chain is deduced by combining multiple nodes in order. Besides, we can obtain multiple valid reasoning chains by sub-expression searching, further improving interpretability and precision. Experiments on multilingual datasets show our method significantly outperforms the baselines. Additionally, our approach is more stable and efficient when faced with the challenges of diverse annotation and complex reasoning with limited resources. Moreover, our approach can be seamlessly integrated with large language models (LLMs), enhancing LLM's mathematical reasoning capabilities at a minimal cost. Experiments show a synergistic collaboration between general-purpose LLM and our specialized model yields superior performance.
Wenqi Zhang 0001, Yongliang Shen 0001, Guiyang Hou, Kuangyi Wang, Weiming Lu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 PromptNER: Prompt Locating and Typing for Named Entity Recognition
abstract
Yongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang, Rongsheng Zhang, Yadong Xi, Weiming Lu, Yueting Zhuang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yongliang Shen 0001, Zeqi Tan, Shuhui Wu, Wenqi Zhang 0001, Yadong Xi, Weiming Lu 0001, Yueting Zhuang
ACL (1)4
2023 An Expression Tree Decoding Strategy for Mathematical Equation Generation
abstract
Generating mathematical equations from natural language requires an accurate understanding of the relations among math expressions.Existing approaches can be broadly categorized into token-level and expression-level generation.The former treats equations as a mathematical language, sequentially generating math tokens.Expression-level methods generate each expression one by one.However, each expression represents a solving step, and there naturally exist parallel or dependent relations between these steps, which are ignored by current sequential methods.Therefore, we integrate tree structure into the expression-level generation and advocate an expression tree decoding strategy.To generate a tree with expression as its node, we employ a layer-wise parallel decoding strategy: we decode multiple independent expressions (leaf nodes) in parallel at each layer and repeat parallel decoding layer by layer to sequentially generate these parent node expressions that depend on others.Besides, a bipartite matching algorithm is adopted to align multiple predictions with annotations for each layer.Experiments show our method outperforms other baselines, especially for these equations with complex structures.
Wenqi Zhang 0001, Yongliang Shen 0001, Qingpeng Nong, Zeqi Tan, Yanna Ma, Weiming Lu 0001
EMNLP1
2022 Query-based Instance Discrimination Network for Relational Triple Extraction
abstract
Joint entity and relation extraction has been a core task in the field of information extraction.Recent approaches usually consider the extraction of relational triples from a stereoscopic perspective, either learning a relation-specific tagger or separate classifiers for each relation type.However, they still suffer from error propagation, relation redundancy and lack of highlevel connections between triples.To address these issues, we propose a novel query-based approach to construct instance-level representations for relational triples.By metric-based comparison between query embeddings and token embeddings, we can extract all types of triples in one step, thus eliminating the error propagation problem.In addition, we learn the instance-level representation of relational triples via contrastive learning.In this way, relational triples can not only enclose rich classlevel semantics but also access to high-order global connections.Experimental results show that our proposed method achieves the state of the art on five widely used benchmarks.
Zeqi Tan, Yongliang Shen 0001, Xuming Hu, Wenqi Zhang 0001, Xiaoxia Cheng, Weiming Lu 0001, Yueting Zhuang
EMNLP4
2022 A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation
abstract
Reliable navigation systems have a wide range of applications in robotics and autonomous driving. Current approaches employ an open-loop process that converts sensor inputs directly into actions. However, these open-loop schemes are challenging to handle complex and dynamic real-world scenarios due to their poor generalization. Imitating human navigation, we add a reasoning process to convert actions back to internal latent states, forming a two-stage closed loop of perception, decision-making, and reasoning. Firstly, VAE-Enhanced Demonstration Learning endows the model with the understanding of basic navigation rules. Then, two dual processes in RL-Enhanced Interaction Learning generate reward feedback for each other and collectively enhance obstacle avoidance capability. The reasoning model can substantially promote generalization and robustness, and facilitate the deployment of the algorithm to real-world robots without elaborate transfers. Experiments show our method is more adaptable to novel scenarios compared with state-of-the-art approaches.
Wenqi Zhang 0001, Peng Li 0031, Yongliang Shen 0001, Yanna Ma, Weiming Lu 0001
IJCAI1
2021 Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots
abstract
Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of hexapod robots is formulated as a Markov Decision Process (MDP) with a specified reward function. Second, a transition feasibility model is proposed for hexapod robots, which describes the feasibility of the state transition under the condition of satisfying kinematics and dynamics, and in turn determines the rewards. Third, the footholds and Center-of-Mass (CoM) sequences are sampled from a diagonal Gaussian distribution and the sequences are optimized through learning the optimal policies using the designed DRL algorithm. Both of the simulation and experimental results on physical systems demonstrate the feasibility and efficiency of the proposed method. Videos are shown at https://videoviewpage.wixsite.com/mcrl.
Huiqiao Fu, Kaiqiang Tang, Peng Li 0031, Wenqi Zhang 0001, Xinpeng Wang 0006, Guizhou Deng, Tao Wang 0004, Chunlin Chen 0001
IJCAI4
2021 Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert Approach
abstract
Despite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in a VUCA environment. With a heuristic exploration mechanism considering target location, path cost, and safety level, the upper layer performs simultaneous map exploration and route-planning to avoid trapping in a blind alley, similar to the cerebrum in the CNS. Using a local adaptive model fusing multiple discrepant strategies, the lower layer pursuits a balance between collision-avoidance and go-straight strategies, acting as the cerebellum in the CNS. We conduct simulation and real-world experiments on multiple platforms, including legged and wheeled robots. Experimental results demonstrate our algorithm outperforms the existing methods in terms of task achievement, time efficiency, and security. A video of our results is available at https://youtu.be/lAnW4QIWDoU.
Wenqi Zhang 0001, Peng Li 0031, Faping Ye, Weijie Jiang 0003, Huiqiao Fu, Tao Wang 0004
IROS1