Xinyang Tong

dblp:354/5795 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot manipulation · 26% Reinforcement learning · 24% Efficient and distributed learning · 18%

Topics — the 7 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
2.732026
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model · AAAI 2026
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025
Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning · ICRA 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning · ICRA 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025
Machine learning › Reinforcement learning › value function estimation
q-function learning
0.912025
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025
Robotics › Legged, aerial and field robots › legged robots
quadruped robot
0.912025
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.912025
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models · ICRA 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.0bridge attention · 1.0reinforcement learning · 0.9multimodal large language model · 0.9mixture of experts · 0.9low-rank adaptation · 0.9group relative policy optimization · 0.9graph of thoughts · 0.9fine-tuning · 0.9action chunk discretization · 0.9
YearPublicationVenuePosition
2026 VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
abstract
Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how to effectively bridge vision-language (VL) representations to action (A). We introduce VLA-Adapter, a novel paradigm designed to reduce the reliance of VLA models on large-scale VLMs and extensive pre-training. To this end, we first systematically analyze the effectiveness of various VL conditions and present key findings on which conditions are essential for bridging perception and action spaces. Based on these insights, we propose a lightweight Policy module with Bridge Attention, which autonomously injects the optimal condition into the action space. In this way, our method achieves high performance using only a 0.5B-parameter backbone, without any robotic data pre-training. Extensive experiments on both simulated and real-world robotic benchmarks show that VLA-Adapter not only achieves state-of-the-art level performance, but also offers the fast inference speed reported to date. Furthermore, thanks to the proposed advanced bridging paradigm, VLA-Adapter enables the training of a powerful VLA model on a single consumer-grade GPU, greatly lowering the barrier to deploying VLA model.
Yihao Wang 0006, Pengxiang Ding, Can Cui 0008, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao 0008, Pengxu Hou, Siteng Huang, Ru Zhang 0002
AAAI6
2025 Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning
abstract
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUARTOnline, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference at 50 Hz in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65 %. Our project page is https://quart-online.github.io.
Xinyang Tong, Pengxiang Ding, Yiguo Fan, Can Cui 0008, Han Zhao 0008, Hongyin Zhang 0001, Yonghao Dang, Siteng Huang, Shangke Lyu
ICRA1
2025 MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
abstract
Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinforcement learning (RL) for fine-tuning large-scale VLA models with a large amount of mixed-quality data. MoRE integrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model. This design enables the model to effectively adapt to a wide array of downstream tasks. Moreover, we employ a reinforcement learning-based training objective to train our model as a Q-function after deeply exploring the structural properties of our tasks. Effective learning from automatically collected mixed-quality data enhances data efficiency and model performance. Extensive experiments demonstrate that MoRE outperforms all baselines across six different skills and exhibits superior generalization capabilities in out-of-distribution scenarios. We further validate our method in real-world scenarios, confirming the practicality of our approach and laying a solid foundation for future research on multi-task learning in quadruped robots.
Han Zhao 0008, Wenxuan Song, Xinyang Tong, Pengxiang Ding, Xuelian Cheng, ZongYuan Ge
ICRA4
2025 A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
abstract
While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view scenarios.To bridge this gap, we introduce A2Seek (Aerial Anomaly Seek), a large-scale, reasoning-centric benchmark dataset for aerial anomaly understanding. This dataset covers various scenarios and environmental conditions, providing high-resolution real-world aerial videos with detailed annotations, including anomaly categories, frame-level timestamps, region-level bounding boxes, and natural language explanations for causal reasoning. Building on this dataset, we propose A2Seek-R1, a novel reasoning framework that generalizes R1-style strategies to aerial anomaly understanding, enabling a deeper understanding of “Where” anomalies occur and “Why” they happen in aerial frames.To this end, A2Seek-R1 first employs a graph-of-thought (GoT)-guided supervised fine-tuning approach to activate the model's latent reasoning capabilities on A2Seek. Then, we introduce Aerial Group Relative Policy Optimization (A-GRPO) to design rule-based reward functions tailored to aerial scenarios. Furthermore, we propose a novel “seeking” mechanism that simulates UAV flight behavior by directing the model's attention to informative regions.Extensive experiments demonstrate that A2Seek-R1 achieves up to a 22.04\% improvement in AP for prediction accuracy and a 13.9\% gain in mIoU for anomaly localization, exhibiting strong generalization across complex environments and out-of-distribution scenarios. Our dataset and code are released at https://2-mo.github.io/A2Seek/.
Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng, Jiankang Zheng, Haosheng Chen 0001, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001
NeurIPS2
2024 ABCF: An Adaptive Balanced Multimodal Website Classification Framework
abstract
Websites are a crucial medium for conveying multimedia information. However, malicious websites can lurk among them, posing a threat to the security of users’ sensitive data and personal privacy. Hence, it is significant to identify and categorize such harmful sites. Various solutions have been proposed by researchers for this task, including extracting features from URLs and analyzing web page content using classification algorithms. To tackle these challenges, we suggest the Adaptive Balanced multimodal website Classification Framework (ABCF), which is designed for web page classification tasks. The framework includes a FAModule and FEAModule for better multimodal data fusion on websites. We also employ incremental learning to continuously improve our model and adapt to the changing network environment. Experiments on three website datasets confirm the effectiveness of our modules in capturing website features and demonstrate the superiority of our model compared to past malicious website classification systems. Code will be released at github.com/lzyy2435/ABCF/.
Zhiyuan Liu 0001, Wei Liu 0161, Xinyang Tong, Hengrui Hu
CSCWD3
2023 Temporal Semantic Attention Network for Aspect-Based Sentiment Analysis
Bin Yang 0038, Xinyang Tong, Huiying Zhao, Zhipu Xie
DEXA (2)2