Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhuoyang Liu

dblp:330/0206 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Robot manipulation · 35% Motion planning and robot control · 35% Vision and language · 23%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
diffusion policy
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › Vision and language › multimodal reasoning
embodied reasoning
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Robotics › Robot manipulation
mobile manipulation
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Robotics › Motion planning and robot control
robot control
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.912025
Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning · NeurIPS 2025
Robotics › Motion planning and robot control
whole-body control
0.912025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › 3D vision
multimodal perception
0.312025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025
Computer vision › 3D vision › point cloud processing
point cloud and image fusion
0.312025
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

parameter sharing · 0.9multimodal conditioning · 0.9mobility-to-body conditioning · 0.9dual-system architecture · 0.9dual-aware co-training · 0.9diffusion transformer · 0.9
YearPublicationVenuePosition
2025 MACRec: A Meta-learning Enhanced Model for Academic Collaborator Recommendation
abstract
Identifying potential academic collaborators has become increasingly crucial to promote scientific development. Existing methods of recommending collaborators primarily focus on analyzing publication content or integrating both content and academic network information. However, these methods often suffer from the cold-start problem due to sparse interaction data among researchers. To this end, we propose a novel meta-learning enhanced model on heterogeneous academic networks for academic collaborator recommendation, named MACRec. This model considers the academic collaborator recommendation in both data and model levels to address the cold-start problem. We introduce a multi-view task constructor that uses academic data and meta-paths to capture semantic contexts within an academic network at the data level, and we also include an academic meta-learner that performs semantic-wise and task-wise adaptations, allowing for fast adaptation to new researcher tasks with few data at the model level. Extensive experiments on two real-world datasets show that MACRec performs significantly better than state-of-the-art baseline methods.
Jingya Zhou, Zhuoyang Liu
CSCWD3
2025 MDN: Modality Decomposition Network for Multimodal Recommendation
abstract
With the rapid growth of multimedia applications and content, multimodal recommendation systems have garnered significant attention due to their ability to leverage diverse data types for personalized recommendations. Existing methods, which primarily focus on extracting common features across modalities, encounter two critical limitations: (1) they often overlook modality-unique features that carry distinct and valuable information, and (2) they fail to effectively capture cooperative interactions between modalities, which are essential for comprehensive understanding. To address these challenges, we propose the Modality Decomposition Network for Multimodal Recommendation (MDN). MDN introduces a novel Multimedia Knowledge Decomposition module that systematically separates modality representations into three key components: common features, unique features, and cooperative features. This decomposition enables our model to learn richer and more comprehensive representations by explicitly modeling the interplay between shared and modality-unique information. Additionally, MDN incorporates a Multimodal Information Encoder to enhance item feature representation by integrating diverse data sources. Furthermore, a Multimodal Contrastive Enhancement Layer is designed to refine user and item representations through contrastive learning, ensuring robust and discriminative recommendations. Extensive experiments conducted on benchmark datasets demonstrate that MDN consistently outperforms existing state-of-the-art methods, achieving superior performance.
Zhuoyang Liu, Weihai Lu
ICMR1
2025 Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
abstract
Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency. To mitigate this dilemma, dual-system approaches have been proposed to leverage a VLM-based System 2 module for handling high-level decision-making, and a separate System 1 action module for ensuring real-time control. However, existing designs maintain both systems as separate models, limiting System 1 from fully leveraging the rich pretrained knowledge from the VLM-based System 2. In this work, we propose Fast-in-Slow (FiS), a unified dual-system vision-language-action (VLA) model that embeds the System 1 execution module within the VLM-based System 2 by partially sharing parameters. This innovative paradigm not only enables high-frequency execution in System 1, but also facilitates coordination between multimodal reasoning and execution components within a single foundation model of System 2. Given their fundamentally distinct roles within FiS-VLA, we design the two systems to incorporate heterogeneous modality inputs alongside asynchronous operating frequencies, enabling both fast and precise manipulation. To enable coordination between the two systems, a dual-aware co-training strategy is proposed that equips System 1 with action generation capabilities while preserving System 2’s contextual understanding to provide stable latent conditions for System 1. For evaluation, FiS-VLA outperforms previous state-of-the-art methods by 8% in simulation and 11% in real-world tasks in terms of average success rate, while achieving a 117.7 Hz control frequency with action chunk set to eight. Project web page: https://fast-in-slow.github.io.
Hao Chen 0193, Jiaming Liu 0003, Chenyang Gu, Zhuoyang Liu, Renrui Zhang, Xiaoqi Li 0020, Yandong Guo, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng
NeurIPS4
2025 AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
abstract
Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicitly model the influence of the mobile base on manipulator control, which easily leads to error accumulation under high degrees of freedom. On the other hand, they treat the entire mobile manipulation process with the same visual observation modality (e.g., either all 2D or all 3D), overlooking the distinct multimodal perception requirements at different stages during mobile manipulation. To address this, we propose the Adaptive Coordination Diffusion Transformer (AC-DiT), which enhances mobile base and manipulator coordination for end-to-end mobile manipulation. First, since the motion of the mobile base directly influences the manipulator's actions, we introduce a mobility-to-body conditioning mechanism that guides the model to first extract base motion representations, which are then used as context prior for predicting whole-body actions. This enables whole-body control that accounts for the potential impact of the mobile base’s motion. Second, to meet the perception requirements at different stages of mobile manipulation, we design a perception-aware multimodal conditioning strategy that dynamically adjusts the fusion weights between various 2D visual images and 3D point clouds, yielding visual features tailored to the current perceptual needs. This allows the model to, for example, adaptively rely more on 2D inputs when semantic information is crucial for action prediction, while placing greater emphasis on 3D geometric information when precise spatial understanding is required. We empirically validate AC-DiT through extensive experiments on both simulated and real-world mobile manipulation tasks, demonstrating superior performance compared to existing methods.
Sixiang Chen, Jiaming Liu 0003, Siyuan Qian, Han Jiang 0003, Zhuoyang Liu, Chenyang Gu, Xiaoqi Li 0009, Chengkai Hou, Pengwei Wang 0004, Zhongyuan Wang 0006, Renrui Zhang, Shanghang Zhang
NeurIPS5
2023 Sparse Non-Contact Multiple People Localization and Vital Signs Monitoring Via FMCW Radar
abstract
Non-contact vital signs monitoring (NCVSM) of multiple people is becoming a necessity in healthcare due to increasing morbidity and manpower shortage. In meeting these requirements, frequency modulated continuous wave (FMCW) radars have shown great potential. However, current techniques present difficulties in locating and monitoring humans in noisy environments containing multiple objects. In this work, we first develop a model for NCVSM of multiple people via FMCW radar, based on a single-input-multiple-output setup. By considering the sparse nature of the modeled signals along with human-typical cardiopulmonary characteristics, we provide a joint-sparse recovery mechanism to accurately localize targets in a clutter-rich scenario where existing techniques struggle. Then, we present a robust method for NCVSM of the found individuals, with improved performance results when compared to current NCVSM techniques using several statistical metrics. Our approach offers excellent performance in a medical application where high accuracy is required.
Yonathan Eder, Zhuoyang Liu, Yonina C. Eldar
ICASSP2
2022 Principle and Application of Physics-Inspired Neural Networks for Electromagnetic Problems
abstract
The interpretability and generalizability of neural networks are well-aware issues in traditional deep learning due to the black-box nature of pure data-driven neural networks. While the physics-inspired neural networks (PINNs) can achieve a more generalized supervised model under few-shot learning by taking the physical principles as prior information into the network design. Especially in the electromagnetic field, the applications and realization of the PINN based on electromagnetic information are important. It can help solve the few-shot learning problems and improve the generalization of deep learning for electromagnetic data. This paper introduces several PINN models for electromagnetic problems, which can significantly reduce the reliance on the training sample size under the same level of networks parameters.
Zhuoyang Liu, Feng Xu 0001
IGARSS1