EDBT 2026 Demo / reviewers in the wild / expert
Chao Yang 0035
dblp:00/5867-35
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-6220-0660ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AnchorMap: A Multi-Agent Pipeline For Variable Standardisation In Maritime EngineeringabstractMaritime co-simulation and collaborative design require different organisations to exchange model data, and equivalent physical quantities are often labeled differently across Original Equipment Manufacturer (OEM) datasheets, simulation models, and regulatory specifications because each organization conceptualizes its domain independently. This causes variable-name mismatches that hinder co-simulation. This paper presents AnchorMap, a schema-driven multi-agent pipeline for standardising heterogeneous variable labels to a canonical vocabulary. A JSON for Linked Data (JSON-LD) schema aligned with the Vessel Information Structure (VIS) defines canonical variable identifiers with metadata and provenance. AnchorMap combines embedding-based candidate retrieval with conditional large language model (LLM) reasoning and applies a confidence-based routing policy to accept reliable mappings, abstain when no justified match exists, or defer ambiguous cases to human review. The approach is evaluated on 144 variable labels from electric motor catalogues and research publications. At the selected operating point, the system achieves a precision of 0.787, recall of 0.842, and F1-score of 0.814, while routing uncertain cases to human review. Results demonstrate that confidence-aware routing enables controlled automation of variable standardisation, and highlight that LLM-based reasoning may select plausible candidates even when no justified match exists, which motivates improved abstention and calibration mechanisms in future work. Saad Ahmed Rana, Haitham Al-Shami, Riku Ala-Laurinaho, Chao Yang 0035, Raine Viitala |
ECMS | 4 |
| 2026 | Enhancing knowledge graph interactions: A comprehensive Text-to-Cypher pipeline with large language modelsabstractKnowledge Graphs (KGs) store structured information but typically require specialized query languages, such as Cypher for Neo4j, creating accessibility challenges for users unfamiliar with graph syntax. Large Language Models (LLMs) offer a solution by translating natural language into Cypher queries. However, existing models—including large-scale LLMs (e.g., ChatGPT) and smaller open-source models (e.g., Llama-7B, 8B) often struggle with accurately generating domain-specific queries due to inadequate alignment with KG schemas and limited domain-specific training data. To address these limitations, we propose a training pipeline tailored specifically for domain-aligned Cypher query generation, emphasizing usability for smaller-scale models. Our method integrates template-based synthetic data generation for diverse, high-quality training samples. We combine supervised fine-tuning with preference learning to enhance domain knowledge and Cypher syntax understanding. Additionally, our approach includes a context-aware retrieval mechanism that dynamically incorporates relevant schema elements at inference, improving alignment with domain-specific knowledge. We evaluated our method on the Hetionet biomedical KG using a benchmark dataset of 240 queries across three complexity levels. Our results show that our context-aware prompting achieves a substantial improvement, increasing component matching accuracy by 23.6% for ChatGPT-4o over the vanilla prompt baseline. When applying our full training pipeline to smaller-scale models, CodeLlama-13B* achieves an execution accuracy of 69.2%, nearly matching ChatGPT-4o’s 72.1%. Importantly, our approach significantly narrows the performance gap, enabling smaller models to effectively manage complex, domain-specific tasks previously dominated by larger models. These findings demonstrate that our method is scalable, computationally efficient, and robust for practical Cypher query generation applications. Chao Yang 0035, Changyi Li, Xiaodu Hu, Hao Yu 0013, Jinzhi Lu 0001 |
Inf. Process. Manag. | 1 |
| 2025 | TwinFlow: Empowering industrial material flow with data-sovereignty through digital twinsabstractIn the era of digital transformation and increasing data-centric operations, efficient and secure management of the supply chain remains a critical challenge. This article identifies the research gap in leveraging emerging technologies to enhance data-sovereign collaboration in the supply chain for manufacturers. To address this, we introduce TwinFlow, a novel architecture designed to facilitate the sharing of material flow data and information among manufacturers in the supply chain, following the principles of IDS (International Data Spaces) and ecosystems like Gaia-X while applying the digital twin methodology. TwinFlow enables knowledge representation of in-plant logistics through ontology modeling and fosters collaboration among manufacturers and relevant stakeholders through a shared data ecosystem. The proof-of-concept implementation of the proposed TwinFlow architecture further validates its efficacy in managing in-plant logistics operations. This study paves the way for a data-sovereign, interoperable, and real-time monitoring-enabled approach to optimizing industrial material flow, contributing significantly to the discourse on digital transformation in supply chain management. Chao Yang 0035, Xinyi Tu 0001, Riku Ala-Laurinaho, Joel Mattila, Jari Juhanko, Kari Tammi, Stefan Vogt, Paul Patolla, Dirk Reichelt |
INDIN | 1 |
| 2024 | Towards Human-Centric Manufacturing: Leveraging Digital Twin for Enhanced Industrial ProcessesabstractHuman-centric industrial processes, such as logistics, inspection, maintenance, and complex assembly, heavily rely on human expertise and judgment. In today’s dynamic and complex manufacturing environments, enhancing operator perception is crucial for timely and accurate decision-making. To facilitate effective communication between human workers and the complex factory ecosystem, this research proposes a system framework leveraging Digital Twin (DT) and semantic technologies to manage industrial heterogeneous data and provide operators with real-time insights. The system architecture comprises three primary layers: the Field Layer, the Information and Service Layer, and the Application Layer. The Information Layer integrates four core engines: Knowledge Engine for managing process-specific knowledge, Data Engine for handling streaming data, Artificial Intelligence (AI) Engine for incorporating advanced machine learning models, and 3D Engine for virtual representation and simulation. This paper presents a detailed implementation of the proposed system framework and validates it through a practical in-plant logistics transport use case. Results demonstrate the framework’s effectiveness in enhancing operator perception and decision-making by providing intuitive interfaces and timely insights. Chao Yang 0035, Hao Yu 0013, Riku Ala-Laurinaho, Lei Feng 0002, Kari Tammi |
IECON | 1 |
| 2024 | Knowledge-Enhanced Digital Twin for Industrial Production ProcessabstractThe manufacturing domain relies on Digital Twins (DTs) to mirror physical systems digitally, facilitating simulation, monitoring, and optimization. However, existing DTs may fail to capture the rich contextual knowledge essential for decision-making in complex manufacturing processes. The evolution to knowledge-enhanced DTs is essential, as it integrates domain-specific knowledge models, enabling a profound understanding of processes. To address this gap, this research introduces a knowledge-enhanced DT framework for the production process. This framework utilizes the ontology-based approach to aid the knowledge integration with the manufacturing DTs. The designed framework consists of three essential layers: The source layer, the Streaming data and knowledge coupling layer, and the Service layer. The proposed framework was further implemented in a lab-scale manufacturing setting and validated through several tests. The results demonstrated the seamless integration of knowledge and streaming data in the production process. Chao Yang 0035, Yuan Hua, Riku Ala-Laurinaho, Udayanto Dwi Atmojo, Kari Tammi |
INDIN | 1 |
| 2023 | Ontology-based knowledge representation of industrial production workflowabstractIndustry 4.0 is helping to unleash a new age of digitalization across industries, leading to a data-driven, interoperable, and decentralized production process. To achieve this major transformation, one of the main requirements is to achieve interoperability across various systems and multiple devices. Ontologies have been used in numerous industrial projects to tackle the interoperability challenge in digital manufacturing. However, there is currently no semantic model in the literature that can be used to represent the industrial production workflow comprehensively while also integrating digitalized information from a variety of systems and contexts. To fill this gap, this paper proposed industrial production workflow ontologies (InPro) for formalizing and integrating production process information. We implemented the 5 M model (manpower, machine, material, method, and measurement) for InPro partitioning and module extraction. The InPro comprises seven main domain ontology modules including Entities, Agents, Machines, Materials, Methods, Measurements, and Production Processes. The Machines ontology module was developed leveraging the OPC Unified Architecture (OPC UA) information model. The presented InPro ontology was further evaluated by a hybrid combination of approaches. Additionally, the InPro ontology was implemented with practical use cases to support production planning and failure analysis by retrieving relevant information via SPARQL queries. The validation results also demonstrated that using the proposed InPro ontology allows for efficiently formalizing, integrating, and retrieving information within the industrial production process context. Chao Yang 0035, Xinyi Tu 0001, Riku Ala-Laurinaho, Juuso Autiosalo, Olli Seppänen, Kari Tammi |
Adv. Eng. Informatics | 1 |