Yixiang Shan

dblp:331/0031 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 56% Generative modeling · 32% Representation and self-supervised learning · 12%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Machine learning › Reinforcement learning
offline reinforcement learning
1.622025
ContraDiff: Planning Towards High Return States via Contrastive Learning · ICLR 2025
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Reinforcement learning
partial observability
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning
state representation learning
0.912025
Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency · ICLR 2025
Machine learning › Generative modeling › diffusion model
diffusion-based data augmentation
0.812024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024
Machine learning › Reinforcement learning › offline reinforcement learning
trajectory stitching
0.812024
DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching · ICML 2024

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.6state reconstruction · 0.9contrastive learning · 0.9agent-wise attention · 0.9data augmentation · 0.8
YearPublicationVenuePosition
2026 Improving Question Recommendation Through Oracle Recommendation Imitation
Ting Long, Yixiang Shan, Yi Chang 0001
DASFAA (1)3
2025 Reconstruction-Guided Policy: Enhancing Decision-Making through Agent-Wise State Consistency
abstract
An important challenge in multi-agent reinforcement learning is partial observability, where agents cannot access the global state of the environment during execution and can only receive observations within their field of view. To address this issue, previous works typically use the dimensional-wise state, which is obtained by applying MLP or dimensional-based attention on the global state, for decision-making during training and relying on a reconstructed dimensional-wise state during execution. However, dimensional-wise states tend to divert agent attention to specific features, neglecting potential dependencies between agents, making it difficult to make optimal decisions. Moreover, the inconsistency between the states used in training and execution further increases additional errors. To resolve these issues, we propose a method called Reconstruction-Guided Policy (RGP) to reconstruct the agent-wise state, which represents the information of inter-agent relationships, as input for decision-making during both training and execution. This not only preserves the potential dependencies between agents but also ensures consistency between the states used in training and execution. We conducted extensive experiments on both discrete and continuous action environments to evaluate RGP, and the results demonstrates its superior effectiveness. Our code is public in https://anonymous.4open.science/r/RGP-9F79
Qifan Liang, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001, Yuan Tian 0016
ICLR2
2025 ContraDiff: Planning Towards High Return States via Contrastive Learning
abstract
The performance of offline reinforcement learning (RL) is sensitive to the proportion of high-return trajectories in the offline dataset. However, in many simulation environments and real-world scenarios, there are large ratios of low-return trajectories rather than high-return trajectories, which makes learning an efficient policy challenging. In this paper, we propose a method called Contrastive Diffuser (ContraDiff) to make full use of low-return trajectories and improve the performance of offline RL algorithms. Specifically, ContraDiff groups the states of trajectories in the offline dataset into high-return states and low-return states and treats them as positive and negative samples correspondingly. Then, it designs a contrastive mechanism to pull the planned trajectory of an agent toward high-return states and push them away from low-return states. Through the contrast mechanism, trajectories with low returns can serve as negative examples for policy learning, guiding the agent to avoid areas associated with low returns and achieve better performance. Through the contrast mechanism, trajectories with low returns provide a ``counteracting force'' guides the agent to avoid areas associated with low returns and achieve better performance. Experiments on 27 sub-optimal datasets demonstrate the effectiveness of our proposed method. Our code is publicly available at https://github.com/Looomo/contradiff.
Yixiang Shan, Zhengbang Zhu, Ting Long, Qifan Liang, Yi Chang 0001, Weinan Zhang 0001
ICLR1
2024 DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory Stitching
abstract
In offline reinforcement learning (RL), the performance of the learned policy highly depends on the quality of offline datasets. However, the offline dataset contains very limited optimal trajectories in many cases. This poses a challenge for offline RL algorithms, as agents must acquire the ability to transit to high-reward regions. To address this issue, we introduce Diffusionbased Trajectory Stitching (DiffStitch), a novel diffusion-based data augmentation pipeline that systematically generates stitching transitions between trajectories. DiffStitch effectively connects low-reward trajectories with high-reward trajectories, forming globally optimal trajectories and thereby mitigating the challenges faced by offline RL algorithms in learning trajectory stitching. Empirical experiments conducted on D4RL datasets demonstrate the effectiveness of our pipeline across RL methodologies. Notably, DiffStitch demonstrates substantial enhancements in the performance of one-step methods(IQL), imitation learning methods(TD3+BC) and trajectory optimization methods(DT). Our code is publicly available at https://github.com/guangheli12/DiffStitch
Guanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long, Weinan Zhang 0001
ICML2
2024 GL-GNN: Graph learning via the network of graphs
Yixiang Shan, Jielong Yang, Yixing Gao 0001
Knowl. Based Syst.1
2023 GLAE: A graph-learnable auto-encoder for single-cell RNA-seq analysis
Yixiang Shan, Jielong Yang, Xiangtao Li, Xionghu Zhong, Yi Chang 0001
Inf. Sci.1
2023 Towards fidelity of graph data augmentation via equivariance
Bai Zhang, Yixing Gao 0001, Linbo Xie, Xiaofeng Cao 0002, Yixiang Shan, Jielong Yang
Knowl. Based Syst.6