Zhuoran Chen

dblp:05/6278 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 41% 3D vision · 38% Generative modeling · 17%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.422024
Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024
Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023
Computer vision › 3D vision
3d scene understanding
1.012026
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026
Computer vision › 3D vision › 3d scene modeling › scene representation
semantic scene representation
1.012026
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026
Computer animation and physical simulation › motion synthesis
human motion synthesis
1.012026
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026
Computer animation and physical simulation › motion synthesis › human motion synthesis
scene-aware motion synthesis
1.012026
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.912025
ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization · ICCV 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent exploration
0.812024
Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning · AAAI 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
0.212023
Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning · IJCAI 2023
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.012003
Assessing sequence comparison methods with the average precision criterion · Bioinform. 2003
Information retrieval › web search
mobile search
0.012002
A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002
Information retrieval › document retrieval › spoken document retrieval
spoken query retrieval
0.012002
A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.012003
Assessing sequence comparison methods with the average precision criterion · Bioinform. 2003
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.012002
A system for spoken query information retrieval on mobile devices · IEEE Trans. Speech Audio Process. 2002

Methods — techniques the papers use, named apart from their topics

tri-plane decomposition · 2.0scene semantic occupancy · 2.0CLIP encoding · 2.0vision-language model · 1.7preference alignment · 1.7LoRA · 1.7transformer sequence modeling · 0.8intrinsic reward · 0.8curriculum learning · 0.8acyclicity constraint · 0.7time-AP plot · 0.0average precision · 0.0speech recognition · 0.0information retrieval · 0.0
YearPublicationVenuePosition
2026 Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
abstract
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take an unified Scene Semantic Occupancy (SSO) for scene representation, termed SSOMotion. We design a bi-directional tri-plane decomposition to derive a compact version of the SSO, and scene semantics are mapped to an unified feature space via CLIP encoding and shared linear dimensionality reduction. Such strategy can derive the fine-grained scene semantic structures while significantly reduce redundant computations. We further take these scene hints and movement direction derived from instructions for motion control via frame-wise scene query. Extensive experiments and ablation studies conducted on cluttered scenes using ShapeNet furniture, as well as scanned scenes from PROX and Replica datasets, demonstrate its cutting-edge performance while validating its effectiveness and generalization ability.
Jingyu Gong, Kunkun Tong, Zhuoran Chen, Chuanhan Yuan, Mingang Chen, Zhizhong Zhang 0001, Xin Tan 0002, Yuan Xie 0006
AAAI3
2025 ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization
abstract
We introduce ImageGem, a dataset for studying generative models that understand fine-grained individual preferences. We posit that a key challenge hindering the development of such a generative model is the lack of in-the-wild and fine-grained user preference annotations. Our dataset features real-world interaction data from 57K users, who collectively have built 242K customized LoRAs, written 3M text prompts, and created 5M generated images. With user preference annotations from our dataset, we were able to train better preference alignment models. In addition, leveraging individual user preference, we investigated the performance of retrieval models and a vision-language model on personalized image retrieval and generative model recommendation. Finally, we propose an end-to-end framework for editing customized diffusion models in a latent weight space to align with individual user preferences. Our results demonstrate that the ImageGem dataset enables, for the first time, a new paradigm for generative model personalization.
Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetztein, Hongyi Wen
ICCV3
2025 Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
AAMAS5
2025 AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose Estimation
abstract
Recent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominantly concentrate on local features surrounding the keypoints of the support image, they often overlook the importance of global features, leading to potential misalignment between the support and query image. To address the inherent conflicts between the two images, we propose AlignCAPE, a novel approach designed to mitigate such misalignment and enhance the model performance. Our method formulates a two-stage pipeline, generating initial proposals in the first stage, followed by another stage to refine iteratively. Specifically, we introduce two modules, Feature Alignment Module(FAM) and Keypoint Perception Module(KPM). FAM utilizes bidirectional cross-attention operation to align the support image feature and query image feature, thereby compensating for the limitations of previous methods. KPM employs self-attention mechanism to capture the interactions among keypoints, facilitating to localize keypoints in the query image. Experiments on MP-100 benchmark demonstrate that our method outperforms the widely-used baseline model in CAPE by 0.68% in [email protected] metric under 1-shot setting.
Zhuoran Chen, Shaojie Zhang 0004, Jianqin Yin
IROS1
2025 RoboCam: Model-Based Robotic Visual Sensing for Precise Inspection of Mesh Screens
abstract
The 3D-printed mesh screen with dense penetrating pores is a new structure for massive manufacturing of molded pulp package products. However, some of the pores may be clogged by the printing material powder during the printing process. Such defects negatively affect the quality of the pulp packages produced using the mesh screen mold. To pinpoint the defects, we design a model-based robotic visual sensing system, called RoboCam, which uses a robotic arm to carry a high-resolution camera for full inspection of a mold consisting of joined mesh screens. To inspect the entire mold, RoboCam plans the camera poses to capture multiple images of the mold and render synthesized images as references for identifying the clogged pores. In particular, we propose novel designs to rectify the inherent run-time pose errors of the robotic system for ensuring the reference quality and to accelerate the reference rendering for reducing inspection latency. Extensive evaluation shows that RoboCam’s design outperforms various baselines, including three existing computer vision and convolution neural network-based inspection systems. RoboCam achieves a recall rate of 94.95% within 528 seconds latency for inspecting an entire mold with 13,000 designed pores.
Duc Van Le, Linshan Jiang, Zhuoran Chen, Xiaohua Peng, Daren Ho, Jianmin Zheng, Rui Tan 0001
ACM Trans. Sens. Networks4
2024 Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning
abstract
Effective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models.
Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan
AAAI4
2023 Deep Hierarchical Communication Graph in Multi-Agent Reinforcement Learning
abstract
Sharing intentions is crucial for efficient cooperation in communication-enabled multi-agent reinforcement learning. Recent work applies static or undirected graphs to determine the order of interaction. However, the static graph is not general for complex cooperative tasks, and the parallel message-passing update in the undirected graph with cycles cannot guarantee convergence. To solve this problem, we propose Deep Hierarchical Communication Graph (DHCG) to learn the dependency relationships between agents based on their messages. The relationships are formulated as directed acyclic graphs (DAGs), where the selection of the proper topology is viewed as an action and trained in an end-to-end fashion. To eliminate the cycles in the graph, we apply an acyclicity constraint as intrinsic rewards and then project the graph in the admissible solution set of DAGs. As a result, DHCG removes redundant communication edges for cost improvement and guarantees convergence. To show the effectiveness of the learned graphs, we propose policy-based and value-based DHCG. Policy-based DHCG factorizes the joint policy in an auto-regressive manner, and value-based DHCG factorizes the joint value function to individual value functions and pairwise payoff functions. Empirical results show that our method improves performance across various cooperative multi-agent tasks, including Predator-Prey, Multi-Agent Coordination Challenge, and StarCraft Multi-Agent Challenge.
Zeyang Liu 0001, Lipeng Wan 0003, Xue Sui, Zhuoran Chen, Kewu Sun, Xuguang Lan
IJCAI4
2003 Assessing sequence comparison methods with the average precision criterion
abstract
MOTIVATION: Comprehensive performance assessment is important for improving sequence database search methods. Sensitivity, selectivity and speed are three major yet usually conflicting evaluation criteria. The average precision (AP) measure aims to combine the sensitivity and selectivity features of a search algorithm. It can be easily visualized and extended to analyze results from a set of queries. Finally, the time-AP plot can clearly show the overall performance of different search methods. RESULTS: Experiments are performed based on the SCOP database. Popular sequence comparison algorithms, namely Smith-Waterman (SSEARCH), FASTA, BLAST and PSI-BLAST are evaluated. We find that (1) the low-complexity segment filtration procedure in BLAST actually harms its overall search quality; (2) AP scores of different search methods are approximately in proportion of the logarithm of search time; and (3) homologs in protein families with many members tend to be more obscure than those in small families. This measure may be helpful for developing new search algorithms and can guide researchers in selecting most suitable search methods. AVAILABILITY: Test sets and source code of this evaluation tool are available upon request.
Zhuoran Chen
Bioinform.1
2002 A system for spoken query information retrieval on mobile devices
abstract
With the proliferation of handheld devices, information access on mobile devices is a topic of growing relevance. This paper presents a system that allows the user to search for information on mobile devices using spoken natural-language queries. We explore several issues related to the creation of this system, which combines state-of-the-art speech-recognition and information-retrieval technologies. This is the first work that we are aware of which evaluates spoken query based information retrieval on a commonly available and well researched text database, the Chinese news corpus used in the National Institute of Standards and Technology (NIST)s TREC-5 and TREC-6 benchmarks. To compare spoken-query retrieval performance for different relevant scenarios and recognition accuracies, the benchmark queries-read verbatim by 20 speakers-were recorded simultaneously through three channels: headset microphone, PDA microphone, and cellular phone. Our results show that for mobile devices with high-quality microphones, spoken-query retrieval based on existing technologies yields retrieval precisions that come close to that for perfect text input (mean average precision 0.459 and 0.489, respectively, on TREC-6).
Eric Chang, Frank Seide, Helen M. Meng, Zhuoran Chen, Yu Shi 0001, Yuk-Chi Li
IEEE Trans. Speech Audio Process.4