VLDB 2026 Research / reviewers in the wild / expert
Khotso Selialia
dblp:339/6775
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-0710-6794ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 27% Trustworthy machine learning · 27% Machine translation · 23% | |
| Computer networks
1 paper |
Edge and fog computing · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
1.0 | 1 | 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
positional encoding |
1.0 | 1 | 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine Translation · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
0.6 | 1 | 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-Study · SenSys 2022 |
Machine learning › Trustworthy machine learning
fairness |
0.6 | 1 | 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-Study · SenSys 2022 |
Machine learning › Efficient and distributed learning › federated learning › data heterogeneity
feature heterogeneity |
0.6 | 1 | 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-Study · SenSys 2022 |
Machine learning › Efficient and distributed learning
federated learning |
0.6 | 1 | 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-Study · SenSys 2022 |
Edge and fog computing › edge devices
heterogeneous edge devices |
0.2 | 1 | 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-Study · SenSys 2022 |
Methods — techniques the papers use, named apart from their topics
normalization · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Tokenization-Induced Distance Distortion in Long-Context Multilingual Machine TranslationabstractMultilingual neural machine translation (MNMT) models degrade in performance as input context length increases, causing positional encoding schemes to misinterpret token distances.Existing absolute and relative positional encodings rely on fixed token indices and implicitly assume uniform semantic density, which breaks down for long-context inputs.We introduce DCARPE, a tokenization-aware adaptive positional encoding that conditions relative positional bias on inputlevel sequence length and fragmentation statistics, allowing the model to reinterpret positional distance when tokenization-induced inflation arises rather than semantic factors.Evaluations on JW300 and out-of-distribution FLORES-200 demonstrate consistent improvements in long-context robustness, achieving gains of up to +10.81 ChrF++ and +8.00 BLEU over baselines. Khotso Selialia, Antoine Nzeyimana, Fatima M. Anwar 0001 |
ACL (1) | 1 |
| 2024 | A Neurosymbolic Approach to Adaptive Feature Extraction in SLAMabstractAutonomous robots, autonomous vehicles, and humans wearing mixed-reality headsets require accurate and reliable tracking services for safety-critical applications in dynamically changing real-world environments. However, the existing tracking approaches, such as Simultaneous Localization and Mapping (SLAM), do not adapt well to environmental changes and boundary conditions despite extensive manual tuning. On the other hand, while deep learning-based approaches can better adapt to environmental changes, they typically demand substantial data for training and often lack flexibility in adapting to new domains. To solve this problem, we propose leveraging the neurosymbolic program synthesis approach to construct adaptable SLAM pipelines that integrate the domain knowledge from traditional SLAM approaches while leveraging data to learn complex relationships. While the approach can synthesize end-to-end SLAM pipelines, we focus on synthesizing the feature extraction module. We first devise a domain-specific language (DSL) that can encapsulate domain knowledge on the essential attributes for feature extraction and the real-world performance of various feature extractors. Our neurosymbolic architecture then undertakes adaptive feature extraction, optimizing parameters via learning while employing symbolic reasoning to select the most suitable feature extractor. Our evaluations demonstrate that our approach, neurosymbolic Feature EXtraction (nFEX), yields higher-quality features. It also reduces the pose error observed for the state-of-the-art baseline feature extractors ORB and SIFT by up to 90% and up to 66%, respectively, thereby enhancing the system’s efficiency and adaptability to novel environments. Yasra Chandio, Momin Ahmad Khan, Khotso Selialia, Luis Garcia 0001, Joseph DeGol, Fatima M. Anwar 0001 |
IROS | 3 |
| 2022 | Federated Learning Biases in Heterogeneous Edge-Devices: A Case-StudyabstractCritical machine learning applications (medical image guidance, task prediction, anomaly detection) require large amounts of data that could not be sufficiently supplied from a single entity, so multiple edge devices collaboratively train their collected data. But this raises privacy and overhead concerns. Federated learning (FL) can be a promising solution to enable these applications while preserving data privacy and mitigating communication overhead. However, an FL model originating from edge deployments with heterogeneous resources may be biased towards a set of devices. We observe that existing bias mitigation techniques in FL focus mainly on the bias that originates from label heterogeneity (due to the skewed distribution of data). We argue that sample feature heterogeneity due to different feature representations at devices is a major contributor to bias in FL. In this paper, we present an analysis of the bias that arises from sampling feature heterogeneity, and analyze the potential of existing performance enhancing techniques (normalization) to overcome bias. Our results demonstrate that normalization techniques do not eliminate bias and motivate the need for dedicated bias mitigation techniques in FL. Khotso Selialia, Yasra Chandio, Fatima M. Anwar 0001 |
SenSys | 1 |