Siyuan Meng

dblp:303/9123 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mosaic: Data-free knowledge distillation via mixture-of-experts for heterogeneous distributed environments
Yanting Gao, Siyuan Meng, Aoqi Wu, Yirong Chen, Shiping Wen 0001
Knowl. Based Syst.4
2026 Embedded Grating Sensing and Compensation Enabling Cross-Scale Nanopositioning
abstract
Grating displacement sensing is regarded as one of the key technologies for achieving cross-scale nanopositioning. This paper proposes a real-time grating sensing correction and compensation technology to enhance the performance of the embedded grating displacement sensor in Cross-scale piezoelectric actuators (CSPAs), thereby enabling nano-scale motion positioning of CSPAs. Firstly, based on the principle of diffracted image reflection, a miniaturized grating sensing unit that can be monolithically integrated with the CSPA structure is designed. Secondly, an online self-correction algorithm based on amplitude iteration is proposed to dynamically eliminate DC offset and amplitude imbalance errors in the signals. Furthermore, a real-time error compensation strategy is constructed to compensate for inherent periodic errors and measurement lag errors of the system induced by stick-slip effects. Experimental results demonstrate that, with the proposed technology, the embedded grating displacement sensor can achieve a detection resolution of 0.9 nm within the full stroke. The CSPA integrated with this sensor achieved a positioning accuracy within ±1.3 nm over its scanning range, and a full-stroke bidirectional positioning consistency of 3.093 ± 1.358 nm.
Siyuan Meng, Jiankang Jiang, Fubo Wang, Dongmei Wu, Wei Dong 0004, Changhai Ru
IEEE Trans Autom. Sci. Eng.1
2025 Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
abstract
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially mitigate due to their modality isolation. While Multimodal Knowledge Graphs (MMKGs) promise enhanced cross-modal understanding, their practical construction is impeded by semantic narrowness of manual text annotations and inherent noise in visual-semantic entity linkages. In this paper, we propose Vision-align-to-Language integrated Knowledge Graph (VaLiK), a novel approach for constructing MMKGs that enhances LLMs reasoning through cross-modal information supplementation. Specifically, we cascade pre-trained Vision-Language Models (VLMs) to align image features with text, transforming them into descriptions that encapsulate image-specific information. Furthermore, we developed a cross-modal similarity verification mechanism to quantify semantic consistency, effectively filtering out noise introduced during feature alignment. Even without manually annotated image captions, the refined descriptions alone suffice to construct the MMKG. Compared to conventional MMKGs construction paradigms, our approach achieves substantial storage efficiency gains while maintaining direct entity-to-image linkage capability. Experimental results on multimodal reasoning tasks demonstrate that LLMs augmented with VaLiK outperform previous state-of-the-art models. Our code is published at https://github.com/Wings-Of-Disaster/VaLiK.
Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang 0001, Botian Shi
ICCV2
2025 HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
abstract
While Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge, conventional single-agent RAG remains fundamentally limited in resolving complex queries demanding coordinated reasoning across heterogeneous data ecosystems. We present HM-RAG, a novel Hierarchical Multi-agent Multimodal RAG framework that pioneers collaborative intelligence for dynamic knowledge synthesis across structured, unstructured, and graph-based data. The framework is composed of a three-tiered architecture with specialized agents: a Decomposition Agent that dissects complex queries into contextually coherent sub-tasks via semantic-aware query rewriting and schema-guided context augmentation; Multi-source Retrieval Agents that carry out parallel, modality-specific retrieval using plug-and-play modules designed for vector, graph, and web-based databases; and a Decision Agent that uses consistency voting to integrate multi-source answers and resolve discrepancies in retrieval results through Expert Model Refinement. This architecture attains comprehensive query understanding by combining textual, graph-relational, and web-derived evidence, resulting in a remarkable 12.95% improvement in answer accuracy and a 3.56% boost in question classification accuracy over baseline RAG systems on the ScienceQA and CrisisMMD benchmarks. Notably, HM-RAG establishes state-of-the-art results in zero-shot settings on both datasets. Its modular architecture ensures seamless integration of new data modalities while maintaining strict data governance, marking a significant advancement in addressing the critical challenges of multimodal reasoning and knowledge synthesis in RAG systems.
Ruoyu Yao, Siyuan Meng, Ding Wang 0001, Jun Ma 0008
ACM Multimedia5
2025 Adaptive Sub-Nanometer Control of a Piezoelectric Positioning Platform
abstract
This paper reports an asymmetric Bouc-Wen (ABW) hysteresis model and a hybrid control algorithm based on multi-modal Bayesian gradient optimization (MBGO) for trajectory tracking in the micro-positioning phase of a piezoelectric positioning platform. First, a system-level dynamic model capable of expressing hysteresis nonlinearity is established based on the asymmetric Bouc-Wen model. Second, an MBGO parameter identification algorithm based on Particle Swarm Optimization (PSO) is proposed to improve the characterization capability of the hysteresis model. Subsequently, a feedforward adaptive fuzzy PID (FF-AFPID) composite controller is designed by compensating the hysteresis nonlinearity through the ABW inverse model while dynamically adjusting PID parameters with adaptive fuzzy rules. Through triangular and sinusoidal trajectory tracking experiments, the root mean square errors were reduced to 0.112 nm and 0.103 nm by the FF-AFPID, with an improvement of 74.944%, 57.088% (triangular) and 77.511%, 57.083% (sinusoidal) over the FF-PID and FF-FPID algorithms, respectively. The results demonstrate that the trajectory tracking of performance the positioning platform in micro-positioning phase was significantly enhanced by FF-AFPID, with the maximum error being suppressed to sub-nanometer levels.
Siyuan Meng, Jiankang Jiang, Qian Ju, Dongmei Wu, Wei Dong 0004, Ming Pang, Changhai Ru
IEEE Trans Autom. Sci. Eng.1
2025 Automated Nanomanipulation for Repairing Defects on Nanoimprint Lithography Masters
abstract
The fabrication of Nanoimprint Lithography (NIL) masters serves as the initial process in the manufacturing of devices such as silicon photonic chips using NIL. Repairing defects and modifying structures on an NIL master accurately is critical for NIL manufacturing. This study proposes an innovative scanning electron microscope (SEM)-based in-situ nanomanipulation technique to address the macro-micro-nano cross-scale nanopositioning issues necessary for repair functions. A macro/micro closed-loop control system was designed, which includes a frequency/voltage (f/u) proportional controller, a real-time direct inverse hysteresis compensation feedforward controller, a grating displacement sensor, and a cross-scale nanopositioning platform. The performance testing of the nano-manipulation approach resulted in a repair range of 22.18×20.92×10.06 mm3, meeting the repair requirements for most silicon photo chip NIL masters in terms of size. The system achieved a repair accuracy at 4.946 nm (X-axis), 4.663 nm (Y-axis), and 4.679 nm (Z-axis). As a demonstrate, the repair functionality testing confirmed the ability of the system to remove target structures such as cantilever beams and detach adhered particles from a master surface.
Siyuan Meng, Jiankang Jiang, Fubo Wang, Qianjun Zhang, Wei Dong 0004, Changhai Ru
IEEE Trans Autom. Sci. Eng.1
2025 A Novel Contouring Control Method Based on Optimal Vector-Referenced Moving Frame for 3D Trajectory With Zero Curvature
abstract
Contouring control of 3D trajectory is critical in multi-axial machine tool, scanning stage and other precision automation systems. Currently, most contouring controllers are based on Frenet frames, thus their limited applicability to nonzero-curvature 3D trajectories rather than zero-curvature ones which exist ubiquitously in multi-axial motion systems. This paper proposes an optimal vector-referenced moving frame based contouring controller (OVRMFCC) suitable for contouring control of 3D zero-curvature trajectories. Firstly, the optimal vector-referenced moving frame (OVRMF) capable of framing arbitrary finite-length smooth 3D trajectory regardless of its curvature was proposed. Then, a contouring controller (OVRMFCC) based on OVRMF was designed, followed by derivation of its analytical form and proof of its convergence. Finally, this controller was deployed to an FPGA-based controller target with comprehensive comparison experiments on a triaxial system. Experimental results indicates that OVRMFCC reduces at least 46.6% maximum contour error, and 25.0% root-mean-square contour error compared to cross-coupled controller. Besides, OVRMFCC achieves almost the same precision on trajectories with nonzero curvature or curvature singularities compared to TCF. It still maintains high-precision contour tracking on trajectories with continuous zero-curvature segments or planned discrete trajectories with sharp curvature changes, while TCF crashes or leads to several times larger contour error. Note to Practitioners—This work is motivated by the increasing need of contouring control of 3D trajectory in precision automation systems. The mainstream 3D contouring controllers, like task coordinated frame method and model predicted contouring controller, are invalid for zero-curvature 3D trajectory which ubiquitously exists in motion system since they are based on Frenet frame which fails to be defined where curvature is zero. Although cross-coupled controller can tackle zero-curvature 3D trajectory, it proves inefficient in reducing contour error as it is commonly model-free and not specifically designed. To solve this problem, we proposed an optimal vector referenced moving frame (OVRMF) for arbitrary infinite-length smooth 3D trajectory framing and proves its existence strictly. The main advantage of OVRMF lies in its independence of curvature, thus its existence everywhere as long as the curve is$C^{1}$continuous. And based on OVRMF a contouring controller (OVRMFCC) is designed which combines the benefit of both TCF and OVRMF. With this, OVRMFCC can track any infinite-length$C^{3}$-continuous 3D trajectory, which fill the gap of traditional 3D contouring controller. Experimental results reveal that OVRMFCC maintains high precision regardless of trajectory curvature, which outperforms both cross-coupled controller and TCF.
Qianjun Zhang, Yongzhuo Gao, Siyuan Meng, Changhai Ru, Wei Dong 0004
IEEE Trans Autom. Sci. Eng.3
2024 LinkThief: Combining Generalized Structure Knowledge with Node Similarity for Link Stealing Attack against GNN
abstract
Graph neural networks (GNNs) have a wide range of applications in multimedia. Recent studies have shown that Graph neural networks (GNNs) are vulnerable to link stealing attacks, which infers the existence of edges in the target GNN's training graph. Existing attacks are usually based on the assumption that links exist between two nodes that share similar posteriors; however, they fail to focus on links that do not hold under this assumption. To this end, we propose LinkThief, an improved link stealing attack that combines generalized structure knowledge with node similarity, in a scenario where the attackers' background knowledge contains partially leaked target graph and shadow graph. Specifically, to equip the attack model with insights into the link structure spanning both the shadow graph and the target graph, we introduce the idea of creating a Shadow-Target Bridge Graph and extracting edge subgraph structure features from it. Through theoretical analysis from the perspective of privacy theft, we first explore how to implement the aforementioned ideas. Building upon the findings, we design the Bridge Graph Generator to construct the Shadow-Target Bridge Graph. Then, the subgraph around the link is sampled by the Edge Subgraph Preparation Module. Finally, the Edge Structure Feature Extractor is designed to obtain generalized structure knowledge, which is combined with node similarity to form the features provided to the attack model. Extensive experiments validate the correctness of theoretical analysis and demonstrate that LinkThief still effectively steals links without extra assumptions. Our code is available at https://github.com/octopusStar218/LinkThief-MM2024.
Siyuan Meng, Chunchun Chen, Mengyao Peng, Hongyan Gu, Xinli Huang
ACM Multimedia2
2024 Structure-Information-Based Reasoning over the Knowledge Graph: A Survey of Methods and Applications
abstract
The knowledge graph (KG) is an efficient form of knowledge organization and expression, providing prior knowledge support for various downstream tasks, and has received extensive attention in natural language processing. However, existing large-scale KGs have many hidden facts that need to be discovered. How to effectively use the structure information of KG is an important research direction of knowledge reasoning. Structure-Information-based reasoning over the KG is a technique used to find the missing facts by the structure information of KG. This survey summarizes the methods and applications of Structure-Information-based reasoning and hopes to be helpful to the research in this field. First, we introduced the definition of knowledge reasoning and the conceptual description of related tasks. Then, we reviewed the methods of Structure-Information-based reasoning. Specifically, we categorized them into four representative classes: PRA-based reasoning, Path-Embedding-based reasoning, RL-based reasoning, and GNN-based reasoning. We compared the motivations and details between practices in the same category. After that, we described the application of Structure-Information-based knowledge reasoning in the KG Completion, Question Answering System, Recommendation System, and other fields. Finally, we discussed the future research directions of Structure-Information-based reasoning.
Siyuan Meng, Jie Zhou 0017, Fengyuan Lu, Xinli Huang
ACM Trans. Knowl. Discov. Data1
2023 SimulE: A novel convolution-based model for knowledge graph embedding
abstract
Knowledge graph embedding technique is one of the mainstream methods to handle the link prediction task, which learns embedding representations for each entity and relation to predict missing links in knowledge graphs. In general, previous convolution-based models apply convolution filters on the reshaped input feature maps to extract expressive features. However, existing convolution-based models cannot extract the interaction information of entities and relations among the same and different dimensional entries simultaneously. To overcome this problem, we propose a novel convolution-based model (SimulE), which utilizes two paths simultaneously to capture the rich interaction information of entities and relations. One path uses 1D convolution filters on 2D reshaped input maps, which maintains the translation properties of the triplets and has the ability to extract interaction information of entities and relations among the same dimensional entries. Another path employs 3D convolution filters on the 3D reshaped input maps, which is suitable for capturing the interaction information of entities and relations among the different dimensional entries. Experimental results show that SimulE can effectively model complex relation types and achieve state-of-the-art performance in almost all metrics on three benchmark datasets. In particular, compared with baseline ConvE, SimulE outperforms it in MRR by 2.9%, 9.8% and 2.8% on FB15k-237, YAGO3-10 and DB100K respectively.
Chaoyi Yan, Xinli Huang, Hongyan Gu, Siyuan Meng
CSCWD4
2021 Adaptive Channel Attention and Feature Super-Resolution for Remote Sensing Images Spatiotemporal Fusion
abstract
CNN-based Spatiotemporal image fusion (STIF) methods have achieved better performance than traditional researches. However, most CNN-based methods fail to make full use of hierarchical features, and ignore the quality and distribution characteristics of feature maps in fine-grained STIF. In this paper, we propose a network with channel attention and feature super-resolution for STIF (CAFSRNet). First, our method uses the low resolution time-domain changing images as input to extract changes more accurately and simplify computational overhead. Second, channel attention mechanism is introduced into Cross-spatial Resolution Mapping module to make the network pay more attention to informative features. Third, by adding feature super-resolution into the supervision process, we enhance the distribution of feature maps and the quality of mapping results. The qualitative and quantitative experimental results on various datasets demonstrate the superiority of our proposed method over the state-of-the-art methods.
Shuai Fang, Siyuan Meng, Yang Cao 0010, Jing Zhang 0037, Weikai Shi
IGARSS2