EDBT 2026 Demo / reviewers in the wild / expert
Sepehr Bakhshi
dblp:304/2718
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-2292-6130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-OFA: On-the-Fly Adaptation of Large Language Models to Address Temporal Drift Across Two Decades of NewsabstractWe investigate the problem of on-the-fly adaptation (OFA) with online feedback for large language models (LLMs) in the context of temporally evolving data. In this setting, each incoming instance-or a small batch- is first processed for inference, and its true label is revealed immediately after prediction, allowing the model to be updated in a sequential, single-pass manner. While pre-trained LLMs achieve state-of-the-art results across NLP tasks, they often struggle to generalize under dynamic distribution shifts-particularly in continuously evolving environments. Despite the importance of this problem, existing research on online adaptation of LLMs remains limited, and there is a lack of large-scale benchmarks for evaluating such methods. To address these gaps, we introduce 1M-News, a large-scale benchmark of one million New York Times headlines spanning two decades, and benchmark six state-of-the-art LLMs by fine-tuning them on the first 10 years and applying OFA on the following 10 years. To improve adaptation performance, we develop Adaptimizer, the first optimizer specifically designed for OFA, enabling rapid and stable model updates under temporal distribution shift. Adaptimizer maintains two sets of weights-fast and slow-balancing rapid adaptation with long-term stability and generalization across the stream. Our experiments demonstrate that OFA with Adaptimizer achieves consistent improvements over static baselines. All code and data are publicly available at https://github.com/pouyaghahramanian/LLM-OFA. Pouya Ghahramanian, Sepehr Bakhshi, Fazli Can |
CIKM | 2 |
| 2024 | Prioritized Binary Transformation Method for Efficient Multi-label Classification of Data Streams with Many LabelsabstractReal-time data processing systems generate huge amounts of data that need to be classified. The volume, variety, velocity, and veracity (uncertainty) of this data necessitate new approaches and the adaptation of existing classification methods. Moreover, the arriving data can belong to more than one class at the same time. As the number of labels grows larger, a significant portion of the multi-label data stream classification methods become computationally inefficient. We propose a novel online approach: the Prioritized Binary Transformation (PBT) method, which can classify data with large numbers of labels by ordering the labels using Principal Component Analysis (PCA) within a fixed-size window. This order is then used to transform the label vectors for classification. We perform an empirical analysis on 12 datasets and compare PBT to four prominent baselines using four evaluation metrics. PBT achieves the best average ranking in three of the four evaluation metrics. Moreover, we investigate efficiency under average execution time per data item and memory consumption where PBT achieves second and first average rankings, respectively. Onur Yildirim, Sepehr Bakhshi, Fazli Can |
CIKM | 2 |
| 2024 | Balancing efficiency vs. effectiveness and providing missing label robustness in multi-label stream classification
Sepehr Bakhshi, Fazli Can |
Knowl. Based Syst. | 1 |
| 2024 | A Novel Neural Ensemble Architecture for On-the-fly Classification of Evolving Text StreamsabstractWe study on-the-fly classification of evolving text streams in which the relation between the input data and target labels changes over time—i.e., “concept drift.” These variations decrease the model’s performance, as predictions become less accurate over time and they necessitate a more adaptable system. While most studies focus on concept drift detection and handling with ensemble approaches, the application of neural models in this area is relatively less studied. We introduce Adaptive Neural Ensemble Network ( AdaNEN ), a novel ensemble-based neural approach, capable of handling concept drift in data streams. With our novel architecture, we address some of the problems neural models face when exploited for online adaptive learning environments. Most current studies address concept drift detection and handling in numerical streams, and the evolving text stream classification remains relatively unexplored. We hypothesize that the lack of public and large-scale experimental data could be one reason. To this end, we propose a method based on an existing approach for generating evolving text streams by introducing various types of concept drifts to real-world text datasets. We provide an extensive evaluation of our proposed approach using 12 state-of-the-art baselines and 13 datasets. We first evaluate concept drift handling capability of AdaNEN and the baseline models on evolving numerical streams; this aims to demonstrate the concept drift handling capabilities of our method on a general spectrum and motivate its use in evolving text streams. The models are then evaluated in evolving text stream classification. Our experimental results show that AdaNEN consistently outperforms the existing approaches in terms of predictive performance with conservative efficiency. Pouya Ghahramanian, Sepehr Bakhshi, Hamed R. Bonab, Fazli Can |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | DynED: Dynamic Ensemble Diversification in Data Stream ClassificationabstractEnsemble methods are commonly used in classification due to their remarkable performance. Achieving high accuracy in a data stream environment is a challenging task considering disruptive changes in the data distribution, also known as concept drift. A greater diversity of ensemble components is known to enhance prediction accuracy in such settings. Despite the diversity of components within an ensemble, not all contribute as expected to its overall performance. This necessitates a method for selecting components that exhibit high performance and diversity. We present a novel ensemble construction and maintenance approach based on MMR (Maximal Marginal Relevance) that dynamically combines the diversity and prediction accuracy of components during the process of structuring an ensemble. The experimental results on both four real and 11 synthetic datasets demonstrate that the proposed approach (DynED) provides a higher average mean accuracy compared to the five state-of-the-art baselines. Soheil Abadifard, Sepehr Bakhshi, Sanaz Gheibuni, Fazli Can |
CIKM | 2 |