EDBT 2026 Demo / reviewers in the wild / expert
Shujian Yu
dblp:154/5763
· DBLP profile ↗
7ranked-venue papers in the field
1as first author
5since 2021 · last 2023
0000-0002-6385-1705ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5 (1 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Gated information bottleneck for generalization in sequential environments
Francesco Alesiani, Shujian Yu |
Knowl. Inf. Syst. | 2 |
| 2023 | Selective Imputation for Multivariate Time Series Datasets With Missing ValuesabstractMultivariate time series often contain missing values for reasons such as failures in data collection mechanisms. Since these missing values can complicate the analysis of time series data, imputation techniques are typically used to deal with this issue. However, the quality of the imputation directly affects the performance of downstream tasks. In this paper, we propose a selective imputation method that identifies a subset of timesteps with missing values to impute in a multivariate time series dataset. This selection, which will result in shorter and simpler time series, is based on both reducing the uncertainty of the imputations and representing the original time series as good as possible. In particular, the method uses multi-objective optimization techniques to select the optimal set of points, and in this selection process, we leverage the beneficial properties of the Multi-task Gaussian Process (MGP). The method is applied to different datasets to analyze the quality of the imputations and the performance obtained in downstream tasks, such as classification or anomaly detection. The results show that much shorter and simpler time series are able to maintain or even improve both the quality of the imputations and the performance of the downstream tasks. Ane Blázquez-García, Kristoffer Wickstrøm, Shujian Yu, Karl Øyvind Mikalsen, Ahcène Boubekki, Angel Conde, Usue Mori, Robert Jenssen, José Antonio Lozano 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Modular-Relatedness for Continual Learning
Ammar Shaker, Francesco Alesiani, Shujian Yu |
IDA | 3 |
| 2022 | Causality detection with matrix-based transfer entropy
Wanqi Zhou, Shujian Yu, Badong Chen |
Inf. Sci. | 2 |
| 2021 | Gated Information Bottleneck for Generalization in Sequential EnvironmentsabstractDeep neural networks suffer from poor generalization to unseen environments when the underlying data distribution is different from that in the training set. By learning minimum sufficient representations from training data, the information bottleneck (IB) approach has demonstrated its effectiveness to improve generalization in different AI applications. In this work, we propose a new neural network-based IB approach, termed gated information bottleneck (GIB), that dynamically drops spurious correlations and progressively selects the most task-relevant features across different environments by a trainable soft mask (on raw features). GIB enjoys a simple and tractable objective, without any variational approximation or distributional assumption. We empirically demonstrate the superiority of GIB over other popular neural network-based IB approaches in adversarial robustness and out-of-distribution (OOD) detection. Meanwhile, we also establish the connection between IB theory and invariant causal representation learning, and observed that GIB demonstrates appealing performance when different environments arrive sequentially, a more practical scenario where invariant risk minimization (IRM) fails. Francesco Alesiani, Shujian Yu |
ICDM | 2 |
| 2020 | Towards Interpretable Multi-task Learning Using Bilevel Programming
Francesco Alesiani, Shujian Yu, Ammar Shaker, Wenzhe Yin |
ECML/PKDD (2) | 2 |
| 2017 | Concept Drift Detection with Hierarchical Hypothesis TestingabstractWhen using statistical models (such as a classifier) in a streaming environment, there is often a need to detect and adapt to concept drifts to mitigate any deterioration in the model's predictive performance over time. Unfortunately, the ability of popular concept drift approaches in detecting these drifts in the relationship of the response and predictor variable is often dependent on the distribution characteristics of the data streams, as well as its sensitivity on parameter tuning. This paper presents Hierarchical Linear Four Rates (HLFR), a framework that detects concept drifts for different data stream distributions (including imbalanced data) by leveraging a hierarchical set of hypothesis tests in an online setting. The performance of HLFR is compared to benchmark approaches using both simulated and real-world datasets spanning the breadth of concept drift types. HLFR significantly outperforms benchmark approaches in terms of accuracy, G-mean, recall, delay in detection and adaptability across the various datasets. Shujian Yu, Zubin Abraham |
SDM | 1 |