Yu Song 0008

dblp:54/1216-8 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-4660-5430ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Robot manipulation · 33% Knowledge representation and reasoning · 33% Motion planning and robot control · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control
robot learning
1.012026
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal reasoning
1.012026
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents · ACL (1) 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

temporal memory · 1.0parallel action decoding · 1.0mutual information objective · 1.0
YearPublicationVenuePosition
2026 MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
abstract
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control.However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding.In this work, we propose MIRTH, a unified framework designed to address these challenges.MIRTH augments a pretrained VLA backbone with three key innovations: (1) dualscale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutualinformation objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding scheme that replaces autoregressive generation with vector-wise prediction to maximize control throughput.Extensive evaluations on the LIBERO simulation benchmark and a real-world LeRobot platform demonstrate that MIRTH achieves state-of-the-art performance and exhibiting emergent error recovery capabilities.The codes and collected datasets are released at http://github.com/kiva12138/mirth.
Hao Sun 0013, Yu Song 0008, Shiyu Teng, Ziwei Niu, Yen-Wei Chen 0001
ACL (1)2
2026 One framework to rule them all: Unifying multimodal tasks with LLM neural-tuning
Hao Sun 0013, Yu Song 0008, Jiaqing Liu, Jihong Hu, Yen-Wei Chen 0001, Lanfen Lin
Pattern Recognit.2
2025 Accurate Tracking of Arabidopsis Root Cortex Cell Nuclei in 3D Time-Lapse Microscopy Images Based on Genetic Algorithm
abstract
Arabidopsis is a widely used model plant to study physiology and development. Live imaging is an important technique to visualize and quantify processes in plant growth and cell division, where accurate cell tracking is essential. The commonly used software TrackMate adopts a tracking-by-detection approach, applying Laplacian of Gaussian (LoG) for blob detection and a Linear Assignment Problem (LAP) tracker for tracking. However, its performance declines when cells are densely arranged. To overcome this limitation, we propose an accurate tracking method based on a Genetic Algorithm (GA) that incorporates knowledge of Arabidopsis root cellular patterns and spatial relationships among volumes. Our method follows a coarse-to-fine strategy: first performing relatively simple line-level tracking of nuclei, then refining associations based on the linear arrangement of cell files and their spatial relationships. We evaluated the method on long-term live imaging datasets of Arabidopsis root tips, and with minor manual correction, it achieved accurate nuclear tracking. To the best of our knowledge, this represents the first successful attempt to address a long-standing problem in time-lapse microscopy of the root meristem by providing an accurate tracking method for Arabidopsis root nuclei.
Yu Song 0008, Tatsuaki Goh, Yinhao Li 0002, Jiahua Dong 0001, Shunsuke Miyashima, Yutaro Iwamoto, Yohei Kondo, Keiji Nakajima, Yenwei Chen
IEEE Trans. Comput. Biol. Bioinform.1
2024 LGA: A Language Guide Adapter for Advancing the SAM Model's Capabilities in Medical Image Segmentation
Jihong Hu, Yinhao Li 0002, Hao Sun 0013, Yu Song 0008, Chujie Zhang, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (12)4
2021 A Teacher-Student Learning Based On Composed Ground-Truth Images For Accurate Cephalometric Landmark Detection
abstract
Computer-aided automatic cephalometric landmark localization has been a hot topic since last century. Recent proposed deep learning-based methods have made great contributions to this research topic. Among them, convolutional neural networks (CNN)-based regression is widely used, where ground-truth (GT) information is mainly used in the calculation of loss function, thus, mimics the difference between the predicted landmarks ' locations and the ground-truth locations through backpropagation. However, considering the limited number of annotated cephalometric data, we believe the performance can be better improved by better utilizing ground-truth information. In this paper, we propose a teacher-student learning method using GT images for accurate cephalometric detection. We first use images composed with GT landmarks as input images to train a detection model, which is treated as a teacher model. Then the teacher model is used to guide a student model, which is trained by original images, by transferring useful features. We believe the features between GT images and original images have similar domain distribution since they both represent same structure. We validate our method on public grand challenge dataset. Our method achieves better performance compared with state-of-the-art methods.
Yu Song 0008, Xu Qiao, Yutaro Iwamoto, Yen-Wei Chen 0001
ICIP1