VLDB 2026 Research / reviewers in the wild / expert
Sumika Arima
dblp:02/11296
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-1853-7215ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Objective Multi-Factory Scheduling with Robust Optimization for Inter-Factory Routing
Yurika Suzuki, Sumika Arima |
ICORES | 2 |
| 2025 | Interaction Modeling for High-Dimensional Mixed Data Considering Distribution and CorrelationabstractThis study aimed to develop an interaction modeling method for high-dimensional mixed data with numerical and categorical variables, assuming application to industrial data. In particular, we focused on the development of a method for the selection of interactions consisting of numerical variables and categorical variables, which have not yet been discussed in previous research. Specifically, we proposed an advanced integration form of Adaptive Sparse Factorization Machines (SFM1A) and Safe Pruning, and confirmed the effect through verification of data simulating real industrial data. As a background, the conventional Factorization Machines based system allows incorrect selection due to a manner of recommendation systems, and tens of thousands of incorrect interactions remain even in SFM1A, which reduces the number of incorrect interactions by 97%, so it has been a fatal limit in the application phase to real data where the correct variables are unknown. Therefore, in this study, we proposed SFM1A_SPC in which the SPC criterion of Safe Pruning, which is an effective screening method for interactions between categorical variables, is applied as a pretreatment to SFM1A. As a result, the selection of incorrect interactions was reduced by 99.7% for data with only categorical variables, and the practical application target (a few hundred or less) was expected to be achieved. On the other hand, the method of converting numerical variables to categorical variables by the simple approach of 0-1 normalization described in the original SPC paper for mixed data with numerical and categorical variables could not be reduced to the target level for practical use. Therefore, as an approach to mixed data with numerical and categorical variables, we propose SFM1A_NCmix in which numerical variables are appropriately categorized and normalized before SPC criterion are applied with data-driven reasonable relaxation. Numerical evaluations are performed on mixed data with numerical and categorical variables of any three types that consider correlations and distributions. As a result, the F1 score was improved by several times (0.90 or more) when using the SFM1A_NCmix for all datasets. In particular, for the mixed data with numerical and categorical variables that follow a multivariate normal distribution that consider correlations, it achieves the incorrect interactions ( FP int ) level for practical applications, besides the superiority over SFM1A and SFM1A_SPC. Taiki Ito, Takuya Matsuzawa, Sumika Arima |
KES | 3 |
| 2025 | The reduction of false positives in Sparse Factorization MachinesabstractIn the semiconductor field, which is the primary focus of this study, the increasing demand has led to a supply shortage, necessitating the expansion of production capacity and the reinforcement of the manufacturing infrastructure. Among the various challenges, the issue of defective products is particularly critical, highlighting the importance of feature analysis for identifying defective factors. However, due to the large scale and complexity of semiconductor manufacturing data, feature analysis poses significant challenges. In this study, we focus on Sparse Factorization Machines (SFM) as an interaction modeling approach for high-dimensional data, characterized by high computational efficiency and strong robustness to missing data. Our previous paper proposed an advanced SFM method denoted by SFM1A, which enhances both selections of main factors and interactions by employing a regularization and a new adaptive technique to SFM with Triangle inequality upper boundary. While SFM1A has demonstrated a significant reduction in false positives (FPs) (main: -99%, interaction: -97%), however, it is not yet at a practical level (e.g. target is less than a few hundred selections including FPs when input data dimension is 10000). To address the issue, this study aims to develop an advanced method that minimizes incorrect interaction selections without compromising selection accuracy for true main factors and interactions. The first proposed method, SFM1A_SPC is to combine SPC criteria mathematically proved in the safe pruning method to SFM1A. SFM1A_SPC successfully reduces FPs of interactions. The second proposal, SFM1A_SPC_FP method, further improves the selection accuracy of main factors by applying a sequential selection method (FP) based on the properties of submodular optimization to perform dimensionality reduction. Numerical evaluation of SFM1A_SPC_FP confirms that it maintains the interaction selection accuracy of SFM1A_SPC while further enhancing the categorical main factor selection. Finally, SFM1A_SPC_FP method achieves the target required for the practical application. Takuya Matsuzawa, Taiki Ito, Sumika Arima |
KES | 3 |
| 2024 | Interaction Modeling for High-Dimensional Mixed Data with Numerical and Categorical FeaturesabstractThis study aimed to develop an interaction modeling method for high-dimensional industrial data with sparsity. Particularly, we discussed the potential and limitations of Sparse Factorization Machines (SFM) with feature selection capabilities after examining the applicability of Factorization Machines (FMs) to numerical and categorical mixed data like industrial data. FMs has been already a major recommendation engine outperforms SVM because of robustness to the sparse data. However, conventional FMs and SFM based on the L2 norm regularization tolerate huge False-Positives (FPs), which is fatal in the application to real data to which the oracle model is unknown. Therefore, in this study, we focus on the way to automatically reduce the several millions of FPs in interactions while keeping high True-Positives (TPs). For the purpose, SFM with trigonometric inequality (TI) upper boundaries (Atarashi et al., 2021) is improved by two directions. The first is the development of TI_SFM (L1) with an L1 norm for selection of main factors, particularly for FPs reduction of the main factors. The second is the application of adaptive technique for reducing FPs of interactions (combinatorial features). We newly developed “Adaptive SFM” with adaptive technique to introduce data-driven penalty of the interaction term. As the result of numerical evaluations using a mass production oracle interaction model and several simulation data, False-Positives of the main factors ( F P main ) and the interactions ( F P int ) are significantly reduced, while keeping high level of True-Positives of the main factor (TP main ) and the interactions (TP int ). Concretely, our proposed Adaptive SFM (L1) outperforms the original TI_SFM (L2) as much reducing F P main and F P int by over 99% when applying our proposed penalty considering not only the relationship between the explanatory variable X i and the objective variable Y, as basic adaptive technique, but also the factor loading indicating the relationship between the latent vector internally optimized and the explanatory variable X i . Our contribution is to lead the possibility for applications to actual industrial data with uncertain main factors and interactions, beyond the applications as recommendation engines. Sara Hoshino, Taiki Ito, Takuya Matsuzawa, Haruki Ozawa, Sumika Arima |
KES | 5 |
| 2024 | Interaction modeling of high-correlated data based on Sparse Factorization Machines and CHANOLabstractHuge amount of data such as 20TB including more than several millions of variables per day is accumulated in recent mega-fabs, however, it is known that only smaller part of the variables influences the final product quality defects, and it is like a treasure hunting in the sea of massive data. In addition, the complex interaction effects are sometimes designed and also appear as the result of complex and the large industrial process. Therefore, automatic selection and estimation of the effects of main factors and interactions can greatly contribute to improving productivity. However, interaction modeling methods for highly-correlated data as well as high-dimensional data are difficult and still developing. This study proposed a new method by applying Convex Hull Approximation of Nearly Optimal Lasso Solutions (CHANOL) to modified Sparse Factorization Machines (SFM) to improve its feature selections and effect estimations for highly-correlated high-dimensional data by comparing with existing methods. First, we aimed to respond to highly correlated data by screening main factors appears only in less solutions/models in a set of CHANOL solutions/models, while responding to high-dimensional sparse data by using an advanced form of SFM. The performance of the proposed CHANOL-SFM was evaluated by using simulated data based on the volume and quality of variables’ correlation similar of variables of an actual semiconductor mass production. As the numerical result, it was confirmed that the incorrect variables which do not influence the final quality are rejected in higher accuracy than the existing method, and the interaction effect estimation is more accurate than the existing method as in selections and estimations of interaction effects. Haruki Ozawa, Sara Hoshino, Sumika Arima |
KES | 3 |
| 2023 | Interaction Modelling of High Dimensional Production DataabstractThe final goal of this study is to present a practical interaction modeling methodology for big and complex industrial data in which there are mixed difficulty of high dimension, sparsity, interaction, multicollinearity, series bias, low signal noise ratio, nonlinearity, and so on. Particularly, this paper focuses on the issues of high dimension, sparsity, and interactions. In addition, it is necessary to estimate unknown factors and interactions in order to improve overall efficiency in complex and diverse situations of industry. The main purpose here is to summarize the applicability and challenges of Factorization Machines (FMs) to manufacturing data. FMs has possibility because of the ad-vantage of to identify unknown factors and its robustness to data with more than 95% missing values which has already been reported for recommendation problems. Manufacturing data assumed here incudes the inspected quality of the final products, and a set of various variables (settings, sensing results, etc.) of the series production process which may relate to the quality. Common characteristics of manufacturing data include mixture of numerical and categorical variables, high dimensionality, many missing values, and so on. Therefore, this paper particularly focuses on sparse high-dimensional type-mix data. So far, while interaction modeling of high-dimensional data has been discussed and verified, mainly for screening methods, still there is very few studies to apply FMs to the manufacturing data with many numerical variables. There-fore, we report the evaluation results using some synthesis data and simulation data from actual oracle model to summarize FMs applicability as the first step. The second step is to compare the performance of sparse FMs which has feature selection function, beside our proposal. Sara Hoshino, Kyo Watanabe, Sumika Arima |
KES | 3 |
| 2023 | Image-Multimodal Data Analysis for Defect Classification: Case Study of Industrial Printing
Hiroki Itou, Kyo Watanabe, Sumika Arima |
KES-IDT | 3 |
| 2023 | Image-Multimodal Data Analysis for Defect Classification: Case Study of Semiconductor Defect Patterns
Daisuke Takada, Hiroki Itou, Ryo Ohta, Takumi Maeda, Kyo Watanabe, Sumika Arima |
KES-IDT | 6 |
| 2021 | Variable Selection for Correlated High-Dimensional Data with Infrequent Categorical Variables: Based on Sparse Sample Regression and Anomaly Detection Technology
Yuhei Kotsuka, Sumika Arima |
KES-IDT | 2 |
| 2019 | Applications of Sparse Modelling and Principle Component Analysis for the Virtual Metrology of Comprehensive Multi-dimensional QualityabstractThis paper discussed the virtual metrology (VM) modelling of multi-class quality to describe the relationship between the variables of a production machine's condition and the estimated/forecasted product quality soon after finishing the machine processing. Applications of PCA and LASSO technique of the Sparse modelling were introduced to define the multi-dimensional quality. Because the high accuracy and quick computations are required for the VM modelling, in this study, the PCA-LASSO combination was applied before building the VM models based on the kernel SVM (kSVM), particularly the linear kernel for real-time use. As the result of evaluation of a CVD (Chemical vapor deposition) process in an actual semiconductor factory, LASSO and linear-SVM could reduce the scale of the machine variable's set and calculation time by almost 57% and 95% without deterioration of accuracy even without PCA. In addition, as the PCA-LASSO, the multi-dimensional quality was rotated to the orthogonality space by PCA to summarize the extracted variables responding to the primary independent hyperspace. As the result of the PCA-LASSO combination, the scale of machine variables extracted was improved by 83%, besides the accuracy of the linear-SVM is 98%. It is also effective as the pre-process of Partial Least Square (PLS). Sumika Arima, Takuya Nagata, Huizhen Bu, Satsuki Shimada |
ICORES | 1 |
| 2012 | Development of Sequential Association Rules for Preventing Minor-stoppages in Semi-conductor Manufacturing
Sumika Arima, Ushio Sumita, Jun Yoshii |
ICORES | 1 |