Guoxi Zhang

dblp:211/5754 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Game-Theoretica Negotiation Framework for Cross-Cultural Consensus
Guoxi Zhang, Tianzhuo Yang, Jiaming Ji, Yaodong Yang 0001, Juntao Dai
ACL (1)1
2025 SYNERGAI: Perception Alignment for Human-Robot Collaboration
abstract
Recently, large language models (LLMs) have shown strong potential in facilitating human-robotic interaction and collaboration. However, existing LLM-based systems often overlook the misalignment between human and robot perceptions, which hinders their effective communication and real-world robot deployment. To address this issue, we introduce SYNERGAI, a unified system designed to achieve both perceptual alignment and human-robot collaboration. At its core, SYNERGAI employs 3D Scene Graph (3DSG) as its explicit and innate representation. This enables the system to leverage LLM to break down complex tasks and allocate appropriate tools in intermediate steps to extract relevant information from the 3DSG, modify its structure, or generate responses. Importantly, SYNERGAI incorporates an automatic mechanism that enables perceptual misalignment correction with users by updating its 3DSG with online interaction. SYNERGAI achieves comparable performance with the data-driven models in ScanQA in a zero-shot manner. Through comprehensive experiments across 10 real-world scenes, SYNERGAI demonstrates its effectiveness in establishing common ground with humans, realizing a success rate of 61.9 % in alignment tasks. It also significantly improves the success rate from 3.7% to 45.68 % on novel tasks by transferring the knowledge acquired during alignment.
Yixin Chen 0003, Guoxi Zhang, Yaowei Zhang, Hongming Xu 0003, Peiyuan Zhi, Qing Li 0003, Siyuan Huang 0001
ICRA2
2024 End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations
abstract
Neuro-symbolic reinforcement learning (NS-RL) has emerged as a promising paradigm for explainable decision-making, characterized by the interpretability of symbolic policies. NS-RL entails structured state representations for tasks with visual observations, but previous methods cannot refine the structured states with rewards due to a lack of efficiency. Accessibility also remains an issue, as extensive domain knowledge is required to interpret symbolic policies. In this paper, we present a neuro-symbolic framework for jointly learning structured states and symbolic policies, whose key idea is to distill the vision foundation model into an efficient perception module and refine it during policy learning. Moreover, we design a pipeline to prompt GPT-4 to generate textual explanations for the learned policies and decisions, significantly reducing users’ cognitive load to understand the symbolic policies. We verify the efficacy of our approach on nine Atari tasks and present GPT-generated explanations for policies and decisions.
Lirui Luo, Guoxi Zhang, Hongming Xu 0003, Yaodong Yang 0001, Cong Fang 0001, Qing Li 0003
ICML2
2024 Treatment Effect Estimation Under Unknown Interference
Xiaofeng Lin 0001, Guoxi Zhang, Xiaotian Lu, Hisashi Kashima
PAKDD (2)2
2024 VickreyFeedback: Cost-Efficient Data Construction for Reinforcement Learning from Human Feedback
Guoxi Zhang, Jiuding Duan
PRIMA1
2024 Learning state importance for preference-based reinforcement learning
Guoxi Zhang, Hisashi Kashima
Mach. Learn.1
2023 Behavior Estimation from Multi-Source Data for Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) have received rising interest due to its appealing data efficiency. The present study addresses behavior estimation, a task that aims at estimating the data-generating policy. In particular, this work considers a scenario where data are collected from multiple sources. Neglecting data heterogeneity, existing approaches cannot provide good estimates and impede policy learning. To overcome this drawback, the present study proposes a latent variable model and a model-learning algorithm to infer a set of policies from data, which allows an agent to use as behavior policy the policy that best describes a particular trajectory. To illustrate the benefit of such a fine-grained characterization for multi-source data, this work showcases how the proposed model can be incorporated into an existing offline RL algorithm. Lastly, with extensive empirical evaluation this work confirms the risks of neglecting data heterogeneity and the efficacy of the proposed model.
Guoxi Zhang, Hisashi Kashima
AAAI1
2023 Estimating Treatment Effects Under Heterogeneous Interference
Xiaofeng Lin 0001, Guoxi Zhang, Xiaotian Lu, Han Bao 0002, Koh Takeuchi 0001, Hisashi Kashima
ECML/PKDD (1)2
2022 Improving Pairwise Rank Aggregation via Querying for Rank Difference
abstract
Pairwise rank aggregation (PRA) aims at learning a ranking from pairwise comparisons between objects that specify their relative ordering. The present study proposes the use of rank difference information for PRA, which characterizes the extent winners in paired comparisons beat their opponents. While such information can be effortlessly recognized by annotators, to our knowledge, it has not been utilized for PRA before. The challenge is three-fold: how to solicit such information, how to utilize it in rank aggregation, and how to overcome the noise from heterogeneous annotators. This study proposes a new query for soliciting information about rank difference that imposes limited cognitive burden on annotators. As prior methods for PRA abounds, it is of interest to empower them with information on rank difference. To this end, this study proposes a conservative learning objective that can be combined seamlessly with many existing PRA algorithms. The third contribution is a new method for PRA called mixture of exponentials (MoE). Annotators from a heterogeneous population might have diverse views concerning rank difference. For example, an annotator might be good at recognizing rank difference only for a subset of items but not the rest. This means that information about rank difference is likely to be perturbed. Unfortunately, such an object-dependent error pattern cannot be modeled with existing approaches. MoE assumes that each annotator uses a mixture of ranking functions in generating answers, and the mixture components can capture object-related patterns in data. The present study evaluates the proposals with extensive experiments on both real and synthetic datasets. The results confirm the efficacy of the proposals and shed light on their practical usage.
Guoxi Zhang, Jiyi Li, Hisashi Kashima
DSAA1
2022 Batch Reinforcement Learning from Crowds
Guoxi Zhang, Hisashi Kashima
ECML/PKDD (4)1
2018 On Reducing Dimensionality of Labeled Data Efficiently
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima
PAKDD (3)1
2017 Robust Multi-view Topic Modeling by Incorporating Detecting Anomalies
Guoxi Zhang, Tomoharu Iwata, Hisashi Kashima
ECML/PKDD (2)1