VLDB 2026 Research / reviewers in the wild / expert
Ege Beyazit
dblp:201/4232
· DBLP profile ↗
9ranked-venue papers
5as first author
2since 2021 · last 2023
0000-0001-5731-7621ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Learning theory · 27% Generative modeling · 16% Representation and self-supervised learning · 15% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
online learning |
0.8 | 2 | 2019 | Online Learning from Capricious Data Streams: A Generative Approach · IJCAI 2019 Online Learning from Data Streams with Varying Feature Spaces · AAAI 2019 |
Machine learning › Learning theory
inductive bias |
0.7 | 1 | 2023 | An Inductive Bias for Tabular Deep Learning · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
tabular deep learning |
0.7 | 1 | 2023 | An Inductive Bias for Tabular Deep Learning · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.4 | 1 | 2020 | Learning Interpretable Representations with Informative Entanglements · IJCAI 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | Learning Interpretable Representations with Informative Entanglements · IJCAI 2020 |
Machine learning › Generative modeling › generative adversarial network
interpretable GANs |
0.4 | 1 | 2020 | Learning Interpretable Representations with Informative Entanglements · IJCAI 2020 |
Machine learning › Trustworthy machine learning › interpretability
interpretable representation learning |
0.4 | 1 | 2020 | Learning Interpretable Representations with Informative Entanglements · IJCAI 2020 |
Machine learning › Representation and self-supervised learning
feature space |
0.4 | 1 | 2019 | Online Learning from Data Streams with Varying Feature Spaces · AAAI 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
generative graphical models |
0.4 | 1 | 2019 | Online Learning from Capricious Data Streams: A Generative Approach · IJCAI 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.4 | 1 | 2019 | Online Learning from Capricious Data Streams: A Generative Approach · IJCAI 2019 |
Machine learning › Time series and sequential data
streaming data |
0.4 | 1 | 2019 | Online Learning from Data Streams with Varying Feature Spaces · AAAI 2019 |
Methods — techniques the papers use, named apart from their topics
spectral analysis · 0.7feature scaling · 0.7feature ranking · 0.7maximum likelihood estimation · 0.4bayesian network · 0.4universal feature space construction · 0.4feature sparsity · 0.4empirical risk minimization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | An Inductive Bias for Tabular Deep LearningabstractDeep learning methods have achieved state-of-the-art performance in most modeling tasks involving images, text and audio, however, they typically underperform tree-based methods on tabular data. In this paper, we hypothesize that a significant contributor to this performance gap is the interaction between irregular target functions resulting from the heterogeneous nature of tabular feature spaces, and the well-known tendency of neural networks to learn smooth functions. Utilizing tools from spectral analysis, we show that functions described by tabular datasets often have high irregularity, and that they can be smoothed by transformations such as scaling and ranking in order to improve performance. However, because these transformations tend to lose information or negatively impact the loss landscape during optimization, they need to be rigorously fine-tuned for each feature to achieve performance gains. To address these problems, we propose introducing frequency reduction as an inductive bias. We realize this bias as a neural network layer that promotes learning low-frequency representations of the input features, allowing the network to operate in a space where the target function is more regular. Our proposed method introduces less computational complexity than a fully connected layer, while significantly improving neural network performance, and speeding up its convergence on 14 tabular datasets. Ege Beyazit, Jonathan Kozaczuk, Vanessa Wallace, Bilal Fadlallah |
NeurIPS | 1 |
| 2021 | Toward Mining Capricious Data Streams: A Generative ApproachabstractLearning with streaming data has received extensive attention during the past few years. Existing approaches assume that the feature space is fixed or changes by following explicit regularities, limiting their applicability in real-time applications. For example, in a smart healthcare platform, the feature space of the patient data varies when different medical service providers use nonidentical feature sets to describe the patients' symptoms. To fill the gap, we in this article propose a novel learning paradigm, namely, Generative Learning With Streaming Capricious (GLSC) data, which does not make any assumption on the feature space dynamics. In other words, GLSC handles the data streams with a varying feature space, where each arriving data instance can arbitrarily carry new features and/or stop carrying partial old features. Specifically, GLSC trains a learner on a universal feature space that establishes relationships between old and new features, so that the patterns learned in the old feature space can be used in the new feature space. The universal feature space is constructed by leveraging the relatednesses among features. We propose a generative graphical model to model the construction process, and show that learning from the universal feature space can effectively improve the performance with theoretical guarantees. The experimental results demonstrate that GLSC achieves conspicuous performance on both synthetic and real data sets. Yi He 0007, Baijun Wu, Di Wu 0056, Ege Beyazit, Sheng Chen 0008, Xindong Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Online Learning to Accelerate Neural Network Inference with Traveling ClassifiersabstractDeep neural networks trained on millions of instances can recognize a wide variety of patterns. It is common to use these pre-trained deep networks in applications where the domain specific training data is not readily available. Once a pre-trained network is deployed to such applications, some of the information contained in the network may be irrelevant due to the difference between the training set and the application data distributions. As a result, parts of the neural network become redundant and slow down inference. This redundancy is unknown until the model is deployed and input data is received. Therefore, it can only be identified and avoided in real-time. Existing works on neural network acceleration can not exploit such redundancy during offline training when the domain-specific datasets are unavailable. In this paper, we study online learning to accelerate neural network inference. We propose traveling classifiers that continuously learn from the activations of two consecutive network layers to accelerate inference in real-time. Traveling classifiers model class conditional probabilities to generate early predictions and bypass unnecessary computation of network layers. The classifiers also adaptively switch the layers they learn from by measuring the feature space differences between the activations. This traveling mechanism automatically adjusts the aggressiveness of the acceleration without sacrificing prediction accuracy. We demonstrate the performance of the proposed algorithm on the ImageNet dataset [10] using the state-of-the-art ResNet-50, ResNet-152 [18] and VGG-16 [38] architectures. Experiments demonstrate that our method significantly outperforms baseline approaches. Ege Beyazit, Yi He 0007, Nian-Feng Tzeng, Xindong Wu 0001 |
ECAI | 1 |
| 2020 | Learning Interpretable Representations with Informative EntanglementsabstractLearning interpretable representations in an unsupervised setting is an important yet a challenging task. Existing unsupervised interpretable methods focus on extracting independent salient features from data. However they miss out the fact that the entanglement of salient features may also be informative. Acknowledging these entanglements can improve the interpretability, resulting in extraction of higher quality and a wider variety of salient features. In this paper, we propose a new method to enable Generative Adversarial Networks (GANs) to discover salient features that may be entangled in an informative manner, instead of extracting only disentangled features. Specifically, we propose a regularizer to punish the disagreement between the extracted feature interactions and a given dependency structure while training. We model these interactions using a Bayesian network, estimate the maximum likelihood parameters and calculate a negative likelihood score to measure the disagreement. Upon qualitatively and quantitatively evaluating the proposed method using both synthetic and real-world datasets, we show that our proposed regularizer guides GANs to learn representations with disentanglement scores competing with the state-of-the-art, while extracting a wider variety of salient features. Ege Beyazit, Doruk Tuncel, Xu Yuan 0001, Nian-Feng Tzeng, Xindong Wu 0001 |
IJCAI | 1 |
| 2019 | Online Learning from Data Streams with Varying Feature SpacesabstractWe study the problem of online learning with varying feature spaces. The problem is challenging because, unlike traditional online learning problems, varying feature spaces can introduce new features or stop having some features without following a pattern. Other existing methods such as online streaming feature selection (Wu et al. 2013), online learning from trapezoidal data streams (Zhang et al. 2016), and learning with feature evolvable streams (Hou, Zhang, and Zhou 2017) are not capable to learn from arbitrarily varying feature spaces because they make assumptions about the feature space dynamics. In this paper, we propose a novel online learning algorithm OLVF to learn from data with arbitrarily varying feature spaces. The OLVF algorithm learns to classify the feature spaces and the instances from feature spaces simultaneously. To classify an instance, the algorithm dynamically projects the instance classifier and the training instance onto their shared feature subspace. The feature space classifier predicts the projection confidences for a given feature space. The instance classifier will be updated by following the empirical risk minimization principle and the strength of the constraints will be scaled by the projection confidences. Afterwards, a feature sparsity method is applied to reduce the model complexity. Experiments on 10 datasets with varying feature spaces have been conducted to demonstrate the performance of the proposed OLVF algorithm. Moreover, experiments with trapezoidal data streams on the same datasets have been conducted to show that OLVF performs better than the state-of-the-art learning algorithm (Zhang et al. 2016). Ege Beyazit, Jeevithan Alagurajah, Xindong Wu 0001 |
AAAI | 1 |
| 2019 | Online Learning from Capricious Data Streams: A Generative ApproachabstractLearning with streaming data has received extensive attention during the past few years. Existing approaches assume the feature space is fixed or changes by following explicit regularities, limiting their applicability in dynamic environments where the data streams are described by an arbitrarily varying feature space. To handle such capricious data streams, we in this paper develop a novel algorithm, named OCDS (Online learning from Capricious Data Streams), which does not make any assumption on feature space dynamics. OCDS trains a learner on a universal feature space that establishes relationships between old and new features, so that the patterns learned in the old feature space can be used in the new feature space. Specifically, the universal feature space is constructed by leveraging the relatednesses among features. We propose a generative graphical model to model the construction process, and show that learning from the universal feature space can effectively improve performance with theoretical analysis. The experimental results demonstrate that OCDS achieves conspicuous performance on synthetic and real datasets. Yi He 0007, Baijun Wu, Di Wu 0056, Ege Beyazit, Sheng Chen 0008, Xindong Wu 0001 |
IJCAI | 4 |
| 2019 | Cost-Efficient Cloud-Based Video Streaming Through Measuring HotnessabstractVideo streaming providers generally have to store several formats of the same video and stream the appropriate format based on the characteristics of the viewer’s device. This approach, called pre-transcoding, incurs a significant cost to the stream providers that rely on cloud services. Furthermore, pre-transcoding proven to be inefficient due to the long-tail access pattern to video streams. To reduce the incurred cost, we propose to pre-transcode only frequently accessed videos (called hot videos) and partially pre-transcode others, depending on their hotness degree. Therefore, we need to measure video stream hotness. Accordingly, we first provide a model to measure the hotness of video streams. Then, we develop methods that operate based on the hotness measure and determine how to pre-transcode videos to minimize the cost of stream providers. The partial pre-transcoding methods operate at different granularity levels to capture different patterns in accessing videos. Particularly, one of the methods operates faster but cannot partially pre-transcode videos with the non-long-tail access pattern. Experimental results show the efficacy of our proposed methods, specifically, when a video stream repository includes a high percentage of the Frequently Accessed Video Streams and a high percentage of videos with the non-long-tail accesses pattern. Mahmoud Darwich, Mohsen Amini Salehi, Ege Beyazit, Magdy A. Bayoumi |
Comput. J. | 3 |
| 2018 | Learning Simplified Decision Boundaries from Trapezoidal Data Streams
Ege Beyazit, Matin Hosseini, Anthony S. Maida, Xindong Wu 0001 |
ICANN (1) | 1 |
| 2018 | Supervised Data Synthesizing and Evolving - A Framework for Real-World Traffic Crash Severity ClassificationabstractTraffic crashes have threatened properties and lives for more than thirty years. Thanks to the recent proliferation of traffic data, the machine learning techniques have been broadly expected to make contributions in the traffic safety community due to their triumphs in many other domains. Among these contributions, the most cited method is to classify traffic crashes in different severities since they have significantly unequal occurrences and costs. However, considering the complexity of transportation system, the traffic data are usually highly imbalanced and lowly separable (HILS), so that few proposed works report satisfactory results. In this paper, we propose a novel framework to deal with the HILS traffic crash data. The framework comprises two parts. In part I, a novel Supervised Data Synthesizing and Evolving algorithm is proposed, which can properly represent the HILS data into a more balanced and separable form without altering the original data distribution. In part II, the details of a customized Multi-Layer Perceptron (MLP) are presented, serving the purpose of learning from the represented data with fast convergence and high accuracy. A real-world traffic crash dataset, as a benchmark, is employed to evaluate the classification performances of our framework and three state-of-the-art imbalanced learning algorithms. The experimental results validate that our framework significantly outperforms the other algorithms. Moreover, the impacts of various parameter settings are studied and discussed Yi He 0007, Di Wu 0056, Ege Beyazit, Xiaoduan Sun, Xindong Wu 0001 |
ICTAI | 3 |