VLDB 2026 Research / reviewers in the wild / expert
Samiul Alam
dblp:222/1821
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-8458-4642ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MemoLens: Empowering Augmented Reality Glasses with Super MemoryabstractSpatial computing empowered by augmented reality (AR) glasses emerges as a new computing paradigm. One of its most transformative capabilities is super memory - the super power to accurately recall objects and people an individual has been paying attention to or interacting with in the physical world. In this work, we propose MemoLensthat empowers AR glasses with super memory. Achieving such capability is not trivial: it requires storing vast amounts of visual data as compact digital memory and retrieving relevant information from the memory efficiently when being prompted. MemoLens achieves this via two key innovations. First, MemoLens leverages eye gaze information captured by the AR glasses to detect what the individual is paying attention to, and introduces a spatio-temporal token compression technique to generate highly compact visual memory. Second, MemoLens introduces a hierarchical search scheme that enables efficient retrieval of relevant memory based on user prompts. We implemented MemoLens using Meta's Aria AR glasses, and evaluated its performance on more than 100 hours of egocentric videos collected in real-world settings. Our results show that MemoLens is able to achieve accurate memory retrieval in real time. Given its promising performance, we believe MemoLens represents a significant step towards realizing always-on super memory in next-generation AR glasses. The project's homepage is https://aiot-mlsys-lab.github.io/memolens.github.io/. Samiul Alam, Shakhrul Iman Siam, Mi Zhang 0002 |
MobiSys | 1 |
| 2026 | GeoFL: A Framework for Efficient Geo-Distributed Cross-Device Federated LearningabstractIn this paper, GeoFL develops a hierarchical federated learning (FL) framework to address the unique challenges in large-scale geo-distributed scenarios. The key idea is to deploy multiple aggregators to geo-distributed clients and aggregate the local model and the global model efficiently and effectively. By assigning each aggregator as a relay layer, GeoFL can elaborately aggregate the geo-distributed clients and systematically determine when to upload the model to the central server based on bandwidth to efficiently update the global model under inadequate and heterogeneous WAN bandwidth constraints. GeoFL designs three key components to optimize the inefficient model aggregation and cope with the non-importance model updates. It further addresses the statistical heterogeneity across geo-distributed aggregators by considering the clients’ graph relationship, delivering an end-to-end clien-taggregator- server architecture for large-scale clients. Compared with existing works, our results on large-scale real-life datasets show that GeoFL speeds up the training process by 1.4×–8× and reduces 6%–80% unnecessary communication rounds between the aggregator and the central server. Maolin Gan, Lanpeng Li, Samiul Alam, Li Liu 0048, Mi Zhang 0002, Huacheng Zeng, Zhichao Cao 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | GeoFL: A Framework for Efficient Geo-Distributed Cross-Device Federated Learning
Maolin Gan, Lanpeng Li, Samiul Alam, Li Liu 0048, Mi Zhang 0002, Zhichao Cao 0001 |
INFOCOM | 3 |
| 2025 | SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model CompressionabstractXin Wang, Samiul Alam, Zhongwei Wan, Hui Shen, Mi Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xin Wang 0120, Samiul Alam, Zhongwei Wan, Hui Shen 0008, Mi Zhang 0002 |
NAACL (Long Papers) | 2 |
| 2025 | Position: Benchmarking is Broken - Don't Let AI be Its Own JudgeabstractThe meteoric rise of Artificial Intelligence (AI), with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as current benchmarks increasingly reveal critical vulnerabilities. Issues like data contamination and selective reporting by model developers fuel hype, while inadequate data quality control can lead to biased evaluations that, even if unintentionally, may favor specific approaches. As a flood of participants enters the AI space, this "Wild West" of assessment makes distinguishing genuine progress from exaggerated claims exceptionally difficult. Such ambiguity blurs scientific signals and erodes public confidence, much as unchecked claims would destabilize financial markets reliant on credible oversight from agencies like Moody's.In high-stakes human examinations (e.g., SAT, GRE), substantial effort is devoted to ensuring fairness and credibility; why settle for less in evaluating AI, especially given its profound societal impact? This position paper argues that a laissez-faire approach is untenable. For true and sustainable AI advancement, we call for a paradigm shift to a unified, live, and quality-controlled benchmarking framework—robust by construction rather than reliant on courtesy or goodwill. Accordingly, we dissect the systemic flaws undermining today’s evaluation ecosystem and distill the essential requirements for next-generation assessments. To concretize this position, we introduce the idea of PeerBench, a community-governed, proctored evaluation blueprint that seeks to improve security and credibility through sealed execution, item banking with rolling renewal, and delayed transparency. PeerBench is presented as a complementary, certificate-grade layer alongside open benchmarks, not a replacement. We discuss trade-offs and limits and call for further research on mechanism design, governance, and reliability guarantees. Our goal is to lay the groundwork for evaluations that restore integrity and deliver genuinely trustworthy measures of AI progress. Zerui Cheng, Stella Wohnig, Ruchika Gupta, Samiul Alam, Tassallah Abdullahi, João Alves Ribeiro, Christian Nielsen-Garcia, Saif Mir, Jason Orender, Seyed Ali Bahrainian, Daniel Kirste, Aaron Gokaslan, Carsten Eickhoff, Ruben Wolff |
NeurIPS | 4 |
| 2025 | Reading Recognition in the WildabstractTo enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public. Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004 |
NeurIPS | 2 |
| 2025 | Artificial Intelligence of Things: A SurveyabstractThe integration of the Internet of Things (IoT) and modern Artificial Intelligence (AI) has given rise to a new paradigm known as the Artificial Intelligence of Things (AIoT). In this survey, we provide a systematic and comprehensive review of AIoT research. We examine AIoT literature related to sensing, computing, and networking & communication, which form the three key components of AIoT. In addition to advancements in these areas, we review domain-specific AIoT systems that are designed for various important application domains. We have also created an accompanying GitHub repository, where we compile the papers included in this survey: https://github.com/AIoT-MLSys-Lab/AIoT-Survey. This repository will be actively maintained and updated with new research as it becomes available. As both IoT and AI become increasingly critical to our society, we believe that AIoT is emerging as an essential research field at the intersection of IoT and modern AI. It is our hope that this survey will serve as a valuable resource for those engaged in AIoT research and act as a catalyst for future explorations to bridge gaps and drive advancements in this exciting field. Shakhrul Iman Siam, Hyunho Ahn, Li Liu 0048, Samiul Alam, Hui Shen 0008, Zhichao Cao 0001, Ness Shroff, Bhaskar Krishnamachari, Mani Srivastava 0001, Mi Zhang 0002 |
ACM Trans. Sens. Networks | 4 |
| 2023 | FedAudio: A Federated Learning Benchmark for Audio TasksabstractFederated learning (FL) has gained substantial attention in recent years due to data privacy concerns related to the pervasiveness of consumer devices that continuously collect data from users. While a number of FL benchmarks have been developed to facilitate FL research, none of them include audio data and audio-related tasks. In this paper, we fill this critical gap by introducing a new FL benchmark for audio tasks which we refer to as FedAudio. FedAudio includes four representative and commonly used audio datasets from three important audio tasks that are well aligned with FL use cases. In particular, a unique contribution of FedAudio is the introduction of data noises and label errors to the datasets to emulate challenges when deploying FL systems in real-world settings. FedAudio also includes the benchmark results of the datasets and a PyTorch library with the objective of facilitating researchers to fairly compare their algorithms. We hope FedAudio could act as a catalyst to inspire new FL research for audio tasks and thus benefit the acoustic and speech research community. The datasets and benchmark results can be accessed at https://github.com/zhang-tuo-pdf/FedAudio. Tiantian Feng, Samiul Alam, Sunwoo Lee 0001, Mi Zhang 0002, Shri Narayanan, Amir Salman Avestimehr |
ICASSP | 3 |
| 2022 | FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model ExtractionabstractMost cross-device federated learning (FL) studies focus on the model-homogeneous setting where the global server model and local client models are identical. However, such constraint not only excludes low-end clients who would otherwise make unique contributions to model training but also restrains clients from training large models due to on-device resource bottlenecks. In this work, we propose FedRolex, a partial training (PT)-based approach that enables model-heterogeneous FL and can train a global server model larger than the largest client model. At its core, FedRolex employs a rolling sub-model extraction scheme that allows different parts of the global server model to be evenly trained, which mitigates the client drift induced by the inconsistency between individual client models and server model architectures. Empirically, we show that FedRolex outperforms state-of-the-art PT-based model-heterogeneous FL methods (e.g. Federated Dropout) and reduces the gap between model-heterogeneous and model-homogeneous FL, especially under the large-model large-dataset regime. In addition, we provide theoretical statistical analysis on its advantage over Federated Dropout. Lastly, we evaluate FedRolex on an emulated real-world device distribution to show that FedRolex can enhance the inclusiveness of FL and boost the performance of low-end devices that would otherwise not benefit from FL. Our code is available at: https://github.com/AIoT-MLSys-Lab/FedRolex. Samiul Alam, Ming Yan 0006, Mi Zhang 0002 |
NeurIPS | 1 |
| 2022 | FedSEA: A Semi-Asynchronous Federated Learning Framework for Extremely Heterogeneous DevicesabstractFederated learning (FL) has attracted increasing attention as a promising technique to drive a vast number of edge devices with artificial intelligence. However, it is very challenging to guarantee the efficiency of a FL system in practice due to the heterogeneous computation resources on different devices. To improve the efficiency of FL systems in the real world, asynchronous FL (AFL) and semi-asynchronous FL (SAFL) methods are proposed such that the server does not need to wait for stragglers. However, existing AFL and SAFL systems suffer from poor accuracy and low efficiency in realistic settings where the data is non-IID distributed across devices and the on-device resources are extremely heterogeneous. In this work, we propose FedSEA - a semi-asynchronous FL framework for extremely heterogeneous devices. We theoretically disclose that the unbalanced aggregation frequency is a root cause of accuracy drop in SAFL. Based on this analysis, we design a training configuration scheduler to balance the aggregation frequency of devices such that the accuracy can be improved. To improve the efficiency of the system in realistic settings where the devices have dynamic on-device resource availability, we design a scheduler that can efficiently predict the arriving time of local updates from devices and adjust the synchronization time point according to the devices' predicted arriving time. We also consider the extremely heterogeneous settings where there exist extremely lagging devices that take hundreds of times as long as the training time of the other devices. In the real world, there might be even some extreme stragglers which are not capable of training the global model. To enable these devices to join in training without impairing the systematic efficiency, Fed-SEA enables these extreme stragglers to conduct local training on much smaller models. Our experiments show that compared with status quo approaches, FedSEA improves the inference accuracy by 44.34% and reduces the systematic time cost and local training time cost by 87.02× and 792.9×. FedSEA also reduces the energy consumption of the devices with extremely limited resources by 752.9×. Jingwei Sun 0002, Ang Li 0005, Lin Duan, Samiul Alam, Xuliang Deng, Xin Guo 0008, Haiming Wang 0002, Maria Gorlatova, Mi Zhang 0002, Hai Li 0001, Yiran Chen 0001 |
SenSys | 4 |
| 2021 | A Large Multi-target Dataset of Common Bengali Handwritten Graphemes
Samiul Alam, Tahsin Reasat, Asif Shahriyar Sushmit, Sadi Mohammad Siddiquee, Fuad Rahman 0001, Mahady Hasan, Ahmed Imtiaz Humayun |
ICDAR (4) | 1 |