Guojun Xie

dblp:216/9151 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-3787-2983ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 70% Representation and self-supervised learning · 30%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning
model-based multi-agent reinforcement learning
0.912025
Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios · AAAI 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios · AAAI 2025
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.912025
Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios · AAAI 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
model-based planning
0.312025
Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios · AAAI 2025

Methods — techniques the papers use, named apart from their topics

teammate modeling · 0.9heuristic policy optimization · 0.9contrastive variational bound · 0.9
YearPublicationVenuePosition
2025 Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative Scenarios
abstract
Learning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods.
Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang
AAAI4
2024 A framework for formal verification of robot kinematics
Guojun Xie, Huanhuan Yang
J. Log. Algebraic Methods Program.1
2023 CoqMatrix: Formal matrix library with multiple models in Coq
ZhengPu Shi, Guojun Xie
J. Syst. Archit.2
2022 Self-supervised representations for multi-view reinforcement learning
abstract
Learning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, such a learning paradigm is sample-inefficiency and sensitive to hyper-parameters when supervised merely by the reward signals. Based on this, we present Self-Supervised Representations (S2R) for multi-view reinforcement learning, a sample-efficient representation learning method for learning features from high-dimensional images. In S2R, we introduce a representation learning framework and define a novel multi-view auxiliary objective based on the multi-view image states and Conditional Entropy Bottleneck (CEB) principle. We integrate S2R with the deep RL agent to learn robust representations that preserve task-relevant information while discarding task-irrelevant information and find optimal policies that maximize the expected return. Empirically, we demonstrate the effectiveness of S2R in the visual DeepMind Control (DMControl) suite and show its better performance on the default DMControl tasks and their variants by replacing the tasks’ default background with a random image or natural video.
Huanhuan Yang, Dian-xi Shi, Guojun Xie, Yingxuan Peng, Yantai Yang, Shaowu Yang
UAI3
2021 CIExplore: Curiosity and Influence-based Exploration in Multi-Agent Cooperative Scenarios with Sparse Rewards
abstract
Learning in a sparse-reward setting is a well-known challenge in RL (Reinforcement Learning). In the single-agent domain, this challenge can be addressed by introducing exploration bonuses driven by intrinsic motivation to encourage agents to visit unseen states. However, naively applying these methods in MARL (Multi-Agent Reinforcement Learning) cooperative settings with sparse rewards results in some inevitable problems: misunderstanding environmental knowledge and lack of collaboration among agents, etc. Based on this, in this paper, we propose the Curiosity and Influence-based Explore (CIExplore) method, which includes a new form of intrinsic reward and an internal counterfactual advantage function. Concretely, the intrinsic reward is a combination of joint curiosity reward and influence reward. The former is the variance of outputs across an ensemble of prediction models that take joint observations and actions of all agents as inputs to predict the next time's joint observations. And the latter quantifies the influence of one agent's behavior on other agents' state-value functions. Given that the joint curiosity reward is shared by all agents, we compute an internal counterfactual advantage function to address this intrinsic reward assignment problem. We demonstrate the efficacy of CIExplore in the multi-agent grid-world environments and show that it is compatible with both on-policy and off-policy MARL algorithms and be scalable to complex settings where agents' number or environment randomness increases.
Huanhuan Yang, Dian-xi Shi, Chenran Zhao, Guojun Xie, Shaowu Yang
CIKM4
2021 Investigating the Effectiveness of Virtual Reality for Culture Learning
abstract
People who are to live, study and work abroad will face more challenges in the new cultural environment and suffer more acculturative stress. Virtual Reality (VR), by which an immersive learning environment can be built, may help them adapt to a foreign culture at a lower cost of time and money. In order to work out a design method for culture learning in VR, we have designed a VR application so that learners can experience and learn the typical western festival culture – Christmas culture – in an immersive environment. To evaluate the effectiveness of the VR method, 50 EFL Chinese university students were enrolled in our experiments and randomly assigned to the VR group and the non-VR group, the data was drawn from cultural knowledge questionnaire, behavior test and Intercultural Sensitivity Scale (ISS). The ANCOVA revealed no major effect for group factor on knowledge learning. Similarly, the Mixed ANOVA identified no major effect for group factor on behavior learning and attitude learning. There was no interaction effect between time and group in all experiments. Our results show that the VR method is preferred by most of the participants, but it shows no remarkable advantage over the non-VR method. Moreover, regression analysis between the culture learning and the sense of presence in VR shows that presence has the potential to improve the performance of intercultural interaction engagement. Our findings are of practical value for culture learning in VR.
Lei Gao 0007, Bo Wan 0002, Gang Liu 0006, Guojun Xie, Jiayang Huang, Guanglan Meng
Int. J. Hum. Comput. Interact.4