Christodoulos Constantinides

dblp:295/6559 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 53% Multi-agent systems · 25% Representation and self-supervised learning · 22%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
1.012026
Diversity Meets Relevancy: Multi-Agent Knowledge Probing for Industry 4.0 Applications · AAAI 2026
Natural language and speech › Language models and text generation
prompting
1.012026
Diversity Meets Relevancy: Multi-Agent Knowledge Probing for Industry 4.0 Applications · AAAI 2026
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.912025
FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes · NeurIPS 2025
Natural language and speech › Language models and text generation › natural language understanding › question answering
multiple-choice question answering
0.912025
FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes · NeurIPS 2025
Computational science and engineering › manufacturing automation › manufacturing
industry 4.0
0.312026
Diversity Meets Relevancy: Multi-Agent Knowledge Probing for Industry 4.0 Applications · AAAI 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large language model · 2.0information diversity metric · 2.0ensembling · 2.0react agent · 0.9perturbation-uncertainty-complexity analysis · 0.9
YearPublicationVenuePosition
2026 Diversity Meets Relevancy: Multi-Agent Knowledge Probing for Industry 4.0 Applications
abstract
Industrial data scientists require deep domain understanding to model asset conditions effectively, yet traditional sources such as Subject Matter Experts (SMEs) and Failure Modes and Effects Analysis (FMEA) documents are often unavailable or incomplete. We present a deployed Multi-Agent System (MAS) that leverages Large Language Models (LLMs) to automatically generate and refine domain-relevant questions, improving modeling decisions across industrial projects. The system addresses two key challenges—ensuring linguistic diversity and maintaining high relevance—by combining established information diversity metrics with a grounded relevancy classifier. We evaluate its effectiveness through diversity benchmarks, compare against direct prompting and AutoAgents baselines, knowledge coverage on downstream FMEA tasks, and controlled user studies. Deployed in real-world projects, the MAS has improved multiple stages of the CRISP-DM methodology, resulting in measurable savings in cost and man-hours.
Christodoulos Constantinides, Dhaval Patel 0002, Scott Kimbleton, Nishu Garg, Muhammad Paracha
AAAI1
2025 FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
abstract
We introduce FailureSensorIQ, a novel Multi-Choice Question-Answering (MCQA) benchmarking system designed to assess the ability of Large Language Models (LLMs) to reason and understand complex, domain-specific scenarios in Industry 4.0. Unlike traditional QA benchmarks, our system focuses on multiple aspects of reasoning through failure modes, sensor data, and the relationships between them across various industrial assets. Through this work, we envision a paradigm shift where modeling decisions are not only data-driven using statistical tools like correlation analysis and significance tests, but also domain-driven by specialized LLMs which can reason about the key contributors and useful patterns that can be captured with feature engineering.We evaluate the Industrial knowledge of over a dozen LLMs including GPT-4, Llama, and Mistral on FailureSensorIQ from different lens using Perturbation-Uncertainty-Complexity analysis, Expert Evaluation study, Asset-Specific Knowledge Gap analysis, ReAct agent using external knowledge-bases.Even though closed-source models with strong reasoning capabilities approach expert-level performance, the comprehensive benchmark reveals a significant drop in performance that is fragile to perturbations, distractions, and inherent knowledge gaps in the models.We also provide a real-world case study of how LLMs can drive the modeling decisions on 3 different failure prediction datasets related to various assets.We release: (a) expert-curated MCQA for various industrial assets, (b) FailureSensorIQ benchmark and Hugging Face leaderboard based on MCQA built from non-textual data found in ISO documents, and (c) ``LLMFeatureSelector'', an LLM-based feature selection scikit-learn pipeline. The software is available at https://github.com/IBM/FailureSensorIQ.
Christodoulos Constantinides, Dhaval Patel 0002, Shuxin Lin, Claudio Guerrero, Sunil Dagajirao Patil, Jayant Kalagnanam
NeurIPS1
2021 A Field Dependence-Independence Perspective on Eye Gaze Behavior within Affective Activities
Christos Fidas, Marios Belk, Christodoulos Constantinides, Argyris Constantinides, Andreas Pitsillides
INTERACT (1)3