Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zilong Song

dblp:231/8345 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 39% Trustworthy machine learning · 35% Efficient and distributed learning · 26%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 62% Human-robot interaction · 19% Immersive interaction · 19%

Topics — the 8 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.912025
AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis · ACM Multimedia 2025
Health and well-being technologies › rehabilitation technology
rehabilitation robotics
0.912025
Gaze and Go: Harnessing Visual Attention Valence in Upper-Limb Robotic Rehabilitation With Tailored Gamification and Eye Tracking for Neuroplasticity · ICRA 2025
Machine learning › Efficient and distributed learning
federated learning
0.812024
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning · AAAI 2024
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning · AAAI 2024
Machine learning › Efficient and distributed learning › federated learning › robust federated learning
noisy label federated learning
0.812024
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning · AAAI 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning · AAAI 2024
Machine learning › Efficient and distributed learning › federated learning
personalized federated learning
0.212024
FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning · AAAI 2024

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 0.9mixture of experts · 0.9human-aligned scoring · 0.9hilbert-schmidt independence criterion · 0.9eye tracking · 0.9adaptive control · 0.9personalized model · 0.8distance regularizer · 0.8confidence regularizer · 0.8
YearPublicationVenuePosition
2025 F2TEval: Human-Aligned Multi-Dimensional Evaluation for Figure-to-Text Task
abstract
Figure -to-Text (F2T) tasks aim to convert structured figure information into natural language text, serving as a bridge between visual perception and language understanding.However, existing evaluation methods remain limited: 1) Reference-based methods can only capture shallow semantic similarities and rely on costly labeled reference text; 2) Reference-free methods depend on multimodal large language models, which suffer from low efficiency and instruction sensitivity; 3) Existing methods provide only sample-level evaluations, lacking interpretability and alignment with expert-level multi-dimensional evaluation criteria.Accordingly, we propose F2TEval, a five-dimensional reference-free evaluation method aligned with expert criteria, covering faithfulness, completeness, conciseness, logicality, and analysis, to support fine-grained evaluation.We design a lightweight mixture-of-experts model that incorporates independent scoring heads and applies the Hilbert-Schmidt Independence Criterion to optimize the disentanglement of scoring representations across dimensions.Furthermore, we construct F2TBenchmark, a humanannotated benchmark dataset covering 21 chart types and 35 application domains, to support research on F2T evaluation.Experimental results demonstrate our model's superior performance and efficiency, outperforming Gemini-2.0and Claude-3.5 with only 0.9B parameters.
Tan Yue, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001
EMNLP3
2025 Gaze and Go: Harnessing Visual Attention Valence in Upper-Limb Robotic Rehabilitation With Tailored Gamification and Eye Tracking for Neuroplasticity
abstract
Therapeutic robotic systems have emerged as reliable tools for physical rehabilitation, providing variableintensity movement assistance to patients with motor impairments. Robot-assisted rehabilitation facilitates the restoration mobility and dexterity, promotes functional neuroplasticity and potentially enables workforce reentry through training-induced cognitive and motor learning. To boost participant engagement and visuomotor coordination, we propose ArmGuider Pro, an advanced upper-limb training system that integrates hand-eye collaboration and gazetriggered assistance within rehabilitation-tailored serious games. The system implements intuitive eye-tracking and visualtriggering strategies to align therapeutic interventions with participants' intentional focus, incorporating immersive gaming elements and adaptive control algorithms. Experimental validation demonstrates significant activation in motor and cognitive cerebral cortex regions, enhanced visual attention concentration in desired target areas (25.92 % improvement), and improved trajectory adherence across sequential sessions (27.27 % improvement). By harnessing visual attention valence, our proposed system could encourage neuroplasticity, supporting its viability for clinical application and widespread adaption in rehabilitation regimens.
Daomiao Wang, Peidong He, Yixi Wang 0001, Zhuo Jian, Zilong Song, Qihan Hu, Fanfu Fang, Cuiwei Yang, Daoyu Wang
ICRA5
2025 AnaFig: A Human-Aligned Dataset for Scientific Figure Analysis
abstract
Scientific Figure Analysis (SFA) aims to derive analytical insights from figures while incorporating background instructions. Unlike conventional tasks such as figure captioning or description generation, which focus on extracting surface-level information from the sole visual modality, SFA requires an intelligent system to summarize key patterns, infer implications, and contextualize scientific findings from visual and textual inputs. It demands not only visual recognition but also the integration of scientific knowledge, multimodal understanding, and contextual reasoning. In this work, we introduce an SFA dataset, AnaFig, comprising 2,000 high-quality samples across 56 domains. All samples are evaluated by using human-aligned five-dimensional scoring criteria, resulting 10,000 human-annotated score labels. The AnaFig dataset facilitates the assessment of three critical capabilities of multimodal large language models (MLLMs): adherence to complex instructions, multimodal perception, and analytical summarization. By building a new benchmark with widely used MLLMs, this study contributes to scientific knowledge discovery and reasoning, fostering the alignment of MLLMs and human experts in scientific analysis.
Tan Yue, Xuzhao Shi, Rui Mao 0010, Zilong Song, Zonghai Hu, Dongyan Zhao 0001
ACM Multimedia4
2025 LiteYOLO-GHG: a lightweight YOLOv8-based algorithm for transformer bushing fault detection
Senyue Xiao, Zeming Pan, Shaoze Wang, Zilong Song, Anni Fan
J. Supercomput.6
2024 FedFixer: Mitigating Heterogeneous Label Noise in Federated Learning
abstract
Federated Learning (FL) heavily depends on label quality for its performance. However, the label distribution among individual clients is always both noisy and heterogeneous. The high loss incurred by client-specific samples in heterogeneous label noise poses challenges for distinguishing between client-specific and noisy label samples, impacting the effectiveness of existing label noise learning approaches. To tackle this issue, we propose FedFixer, where the personalized model is introduced to cooperate with the global model to effectively select clean client-specific samples. In the dual models, updating the personalized model solely at a local level can lead to overfitting on noisy data due to limited samples, consequently affecting both the local and global models’ performance. To mitigate overfitting, we address this concern from two perspectives. Firstly, we employ a confidence regularizer to alleviate the impact of unconfident predictions caused by label noise. Secondly, a distance regularizer is implemented to constrain the disparity between the personalized and global models. We validate the effectiveness of FedFixer through extensive experiments on benchmark datasets. The results demonstrate that FedFixer can perform well in filtering noisy label samples on different clients, especially in highly heterogeneous label noise scenarios.
Xinyuan Ji, Zhaowei Zhu, Wei Xi 0003, Olga Gadyatskaya, Zilong Song, Yang Liu 0018
AAAI5
2024 SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
Zilong Song, Yang Gao 0032, Sijia Dai, Shuai Li 0001, Aimin Hao, Shoulong Zhang
ICXR1
2021 Reliable Reputation Review and Secure Energy Transaction of Microgrid Community Based on Hybrid Blockchain
abstract
A growing number of prosumers have entered the local power market in response to an increase in the number of residential users who can afford to install distributed energy resources. The traditional microgrid trading platform has many problems, such as low transaction efficiency, the high cost of market maintenance, opaque transactions, and the difficulty of ensuring user privacy, which are not conducive to encouraging users to participate in local electricity trading. A blockchain‐based mechanism of microgrid transactions can solve these problems, but the common single‐blockchain framework cannot manage user identity. This study thus proposes a mechanism for secure microgrid transactions based on the hybrid blockchain. A hybrid framework consisting of private blockchain and consortium blockchain is first proposed to complete market transactions. The private blockchain stores the identifying information of users and a review of their transactions, while the consortium blockchain is responsible for storing transaction information. The block digest of the private blockchain is stored in the consortium blockchain to prevent information on the private blockchain from being tampered with by the central node. A reputation evaluation algorithm based on user behavior is then developed to evaluate user reputation, which affects the results of the access audit on the private blockchain. The higher a user’s reputation score is, the more benefits he/she can obtain in the transaction process. Finally, an identity‐based proxy signcryption algorithm is proposed to help the intelligent management device with limited computing power obtain signcryption information in the transaction process to protect the transaction information. A system analysis showed that the secure transaction mechanism of the microgrid based on the hybrid blockchain boasts many security features, such as privacy, transparency, and imtamperability. The proposed reputation evaluation algorithm can objectively reflect all users’ behaviors through their reputation scores, and the identity‐based proxy signcryption algorithm is practical.
Zilong Song, Xiaohong Zhang 0004, Miaomiao Liang
Wirel. Commun. Mob. Comput.1