Vitali Hirsch

dblp:257/8685 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 76% Data integration and cleaning · 24%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling
classification
1.122023
Exploiting domain knowledge to address class imbalance and a heterogeneous feature space in multi-class classification · VLDB J. 2023
Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data · Proc. VLDB Endow. 2020
Data mining › predictive modeling › classification
class imbalance
1.122023
Exploiting domain knowledge to address class imbalance and a heterogeneous feature space in multi-class classification · VLDB J. 2023
Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data · Proc. VLDB Endow. 2020
Data integration and cleaning
data preprocessing
1.122023
Exploiting domain knowledge to address class imbalance and a heterogeneous feature space in multi-class classification · VLDB J. 2023
Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data · Proc. VLDB Endow. 2020
Data mining › predictive modeling › classification
multiclass classification
1.122023
Exploiting domain knowledge to address class imbalance and a heterogeneous feature space in multi-class classification · VLDB J. 2023
Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data · Proc. VLDB Endow. 2020

Methods — techniques the papers use, named apart from their topics

taxonomy-based data preparation · 0.7domain knowledge exploitation · 0.4data preparation · 0.4
YearPublicationVenuePosition
2023 Exploiting domain knowledge to address class imbalance and a heterogeneous feature space in multi-class classification
abstract
Abstract Real-world data of multi-class classification tasks often show complex data characteristics that lead to a reduced classification performance. Major analytical challenges are a high degree of multi-class imbalance within data and a heterogeneous feature space, which increases the number and complexity of class patterns. Existing solutions to classification or data pre-processing only address one of these two challenges in isolation. We propose a novel classification approach that explicitly addresses both challenges of multi-class imbalance and heterogeneous feature space together. As main contribution, this approach exploits domain knowledge in terms of a taxonomy to systematically prepare the training data. Based on an experimental evaluation on both real-world data and several synthetically generated data sets, we show that our approach outperforms any other classification technique in terms of accuracy. Furthermore, it entails considerable practical benefits in real-world use cases, e.g., it reduces rework required in the area of product quality control.
Vitali Hirsch, Peter Reimann 0002, Dennis Treder-Tschechlov, Holger Schwarz, Bernhard Mitschang
VLDB J.1
2020 Exploiting Domain Knowledge to address Multi-Class Imbalance and a Heterogeneous Feature Space in Classification Tasks for Manufacturing Data
abstract
Classification techniques are increasingly adopted for quality control in manufacturing, e.g., to help domain experts identify the cause of quality issues of defective products. However, real-world data often imply a set of analytical challenges, which lead to a reduced classification performance. Major challenges are a high degree of multi-class imbalance within data and a heterogeneous feature space that arises from the variety of underlying products. This paper considers such a challenging use case in the area of End-of-Line testing, i.e., the final functional test of complex products. Existing solutions to classification or data pre-processing only address individual analytical challenges in isolation. We propose a novel classification system that explicitly addresses both challenges of multi-class imbalance and a heterogeneous feature space together. As main contribution, this system exploits domain knowledge to systematically prepare the training data. Based on an experimental evaluation on real-world data, we show that our classification system outperforms any other classification technique in terms of accuracy. Furthermore, we can reduce the amount of rework required to solve a quality issue of a product.
Vitali Hirsch, Peter Reimann 0002, Bernhard Mitschang
Proc. VLDB Endow.1
2019 Data-Driven Fault Diagnosis in End-of-Line Testing of Complex Products
abstract
Machine learning approaches may support various use cases in the manufacturing industry. However, these approaches often do not address the inherent characteristics of the real manufacturing data at hand. In fact, real data impose analytical challenges that have a strong influence on the performance and suitability of machine learning methods. This paper considers such a challenging use case in the area of End-of-Line testing, i.e., the final functional check of complex products after the whole assembly line. Here, classification approaches may be used to support quality engineers in identifying faulty components of defective products. For this, we discuss relevant data sources and their characteristics, and we derive the resulting analytical challenges. We have identified a set of sophisticated data-driven methods that may be suitable to our use case at first glance, e.g., methods based on ensemble learning or sampling. The major contribution of this paper is a thorough comparative study of these methods to identify whether they are able to cope with the analytical challenges. This comprises the discussion of both fundamental theoretical aspects and major results of detailed experiments we have performed on the real data of our use case.
Vitali Hirsch, Peter Reimann 0002, Bernhard Mitschang
DSAA1