VLDB 2026 Research / reviewers in the wild / expert
George Michalopoulos
dblp:157/6081
· DBLP profile ↗
6ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-3958-3499ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 40% Information extraction and text analysis · 34% Representation and self-supervised learning · 26% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research · EMNLP 2025 |
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation |
0.6 | 1 | 2022 | LexSubCon: Integrating Knowledge from Lexical Resources into Contextual Embeddings for Lexical Substitution · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis › lexical semantics
lexical substitution |
0.6 | 1 | 2022 | LexSubCon: Integrating Knowledge from Lexical Resources into Contextual Embeddings for Lexical Substitution · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis
lexical resources |
0.2 | 1 | 2022 | LexSubCon: Integrating Knowledge from Lexical Resources into Contextual Embeddings for Lexical Substitution · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
benchmark evaluation · 1.7sentence similarity · 0.6mix-up embedding · 0.6fine-tuning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchabstractShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang, Daoyang Li, Liqiang Jing, Kaiyu He, Peilin Wu, Juntong Ni, George Michalopoulos, Yue Zhang, Ziyang Zhang, Mian Zhang, Zhiyu Chen, Xinya Du. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziming Luo, Daoyang Li, Liqiang Jing, Kaiyu He, Juntong Ni, George Michalopoulos, Zhiyu Chen 0002, Xinya Du |
EMNLP | 10 |
| 2022 | LexSubCon: Integrating Knowledge from Lexical Resources into Contextual Embeddings for Lexical SubstitutionabstractLexical substitution is the task of generating meaningful substitutes for a word in a given textual context.Contextual word embedding models have achieved state-of-the-art results in the lexical substitution task by relying on contextual information extracted from the replaced word within the sentence.However, such models do not take into account structured knowledge that exists in external lexical databases.We introduce LexSubCon, an end-to-end lexical substitution framework based on contextual embedding models that can identify highly-accurate substitute candidates.This is achieved by combining contextual information with knowledge from structured lexical resources.Our approach involves: (i) introducing a novel mix-up embedding strategy to the target word's embedding through linearly interpolating the pair of the target input embedding and the average embedding of its probable synonyms; (ii) considering the similarity of the sentence-definition embeddings of the target word and its proposed candidates; and, (iii) calculating the effect of each substitution on the semantics of the sentence through a fine-tuned sentence similarity model.Our experiments show that LexSubCon outperforms previous state-of-the-art methods by at least 2% over all the official lexical substitution metrics on LS07 and CoInCo benchmark datasets that are widely used for lexical substitution tasks. George Michalopoulos, Ian McKillop, Alexander Wong, Helen H. Chen |
ACL (1) | 1 |
| 2021 | UmlsBERT: Clinical Domain Knowledge Augmentation of Contextual Embeddings Using the Unified Medical Language System MetathesaurusabstractGeorge Michalopoulos, Yuanxin Wang, Hussam Kaka, Helen Chen, Alexander Wong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. George Michalopoulos, Yuanxin Wang 0001, Hussam Kaka, Helen H. Chen, Alexander Wong |
NAACL-HLT | 1 |
| 2020 | Revealing Common and Rare Patterns for Peritoneal Dialysis Eligibility Decisions with Association Discovery and DisentanglementabstractPeritoneal dialysis (PD) removes waste products from blood when the kidney is malfunctioned. Since there is no clear criterion for PD recommendation for patients with kidney disease, existing machine learning models (ML), which rely on credible decision criterion, are ineffective in making PD eligibility decisions, especially when the correlated traits or indicators (patterns) inherent in the PD data are diverse and subtle. Furthermore, the lack of interpretable transparency in traditional ML also weakens the credibility of the decision they produce. Hence, an in-depth knowledge of the patients' characteristics is needed to render a clearer picture of the decision-making process and model to detect the rare PD eligibility cases. In this paper, we extend our previous work (Attribute-Value-Association Discovery and Disentanglement (ADD)), to an extended ADD for PD data analysis (PD-ADD) to overcome these problems. We show that PD-ADD is able to discover association patterns of patient profiles and symptoms to reveal PD characteristics and detect eligible rare cases. Experimental results show that PDADD is much superior to existing unsupervised clustering (with accuracy of 89.87% vs 73.37% of K-Means). It also enables straightforward interpretation of the underlying relations of patient characteristics in an unsupervised setting. Peiyuan Zhou, Andrew K. C. Wong, George Michalopoulos, Robert R. Quinn, Matthew J. Oliver, Zahid A. Butt, Helen H. Chen |
BIBM | 3 |
| 2016 | Engineering Oracles for Time-Dependent Road NetworksabstractWe implement and experimentally evaluate landmark-based oracles for min-cost paths in two different types of road networks with time-dependent arc-cost functions, based on distinct real-world historic traffic data: the road network for the metropolitan area of Berlin, and the national road network of Germany. Our first contribution is a significant improvement on the implementation of the FLAT oracle, which was proposed and experimentally tested in previous works. Regarding the implementation, we exploit parallelism to reduce preprocessing time and real-time responsiveness to live-traffic reports. We also adopt a lossless compression scheme that severely reduces preprocessing space and time requirements. As for the experimentation, apart from employing the new data set of Germany, we also construct several refinements and hybrids of the most prominent landmark sets for the city of Berlin. A significant improvement to the speedup of FLAT is observed: For Berlin, the average query time can now be as small as 83μsec, achieving a speedup (against the time-dependent variant of Dijkstra's algorithm) of more than 1, 119 in absolute running times and more than 1, 570 in Dijkstra-ranks, with worst-case observed stretch less than 0.781%. For Germany, our experimental findings are analogous: The average query-response time can be 1.269msec, achieving a speedup of more than 902 in absolute running times, and 1, 531 in Dijkstra-ranks, with worst-case stretch less than 1.534%. Our second contribution is the implementation and experimental evaluation of a novel hierarchical oracle (HORN). It is based on a hierarchy of landmarks, with a few “global” landmarks at the top level possessing travel-time information for all possible destinations, and many more “local” landmarks at lower levels possessing travel-time information only for a small neighborhood of destinations around them. As it was previously proved, the advantage of HORN over FLAT is that it achieves query times sublinear, not just in the size of the network, but in the Dijkstra-rank of the query at hand, while requiring asymptotically similar preprocessing space and time. Our experimentation of HORN in Berlin indeed demonstrates improvements in query times (more than 30.37%), Dijkstra-ranks (more than 39.66%), and also worst-case error (more than 35.89%), at the expense of a small blow-up in space. Finally, we implement and experimentally test a dynamic scheme to provide responsiveness to live-traffic reports of incidents with a small timelife (e.g., a temporary blockage of a road segment due to an accident). Our experiments also indicate that the traffic-related information can be updated in seconds. Spyros C. Kontogiannis, George Michalopoulos, Georgia Papastavrou, Andreas Paraskevopoulos, Dorothea Wagner, Christos D. Zaroliagis |
ALENEX | 2 |
| 2015 | Analysis and Experimental Evaluation of Time-Dependent Distance OraclesabstractUrban road networks are represented as directed graphs, accompanied by a metric which assigns cost functions (rather than scalars) to the arcs, e.g. representing time-dependent arc-traversal-times. In this work, we present oracles for providing time-dependent min-cost route plans, and conduct their experimental evaluation on a real-world data set (city of Berlin). Our oracles are based on precomputing all landmark-to-vertex shortest travel-time functions, for properly selected landmark sets. The core of this preprocessing phase is based on a novel, quite efficient and simple one-to-all approximation method for creating approximations of shortest travel-time functions. We then propose three query algorithms, including a PTAS, to efficiently provide min-cost route plan responses to arbitrary queries. Apart from the purely algorithmic challenges, we deal also with several implementation details concerning the digestion of raw traffic data, and we provide heuristic improvements of both the preprocessing phase and the query algorithms. We conduct an extensive, comparative experimental study with all query algorithms and six landmark sets. Our results are quite encouraging, achieving remarkable speedups (at least by two orders of magnitude) and quite small approximation guarantees, over the time-dependent variant of Dijkstra's algorithm. Spyros C. Kontogiannis, George Michalopoulos, Georgia Papastavrou, Andreas Paraskevopoulos, Dorothea Wagner, Christos D. Zaroliagis |
ALENEX | 2 |