VLDB 2026 Research / reviewers in the wild / expert
Chan-Tong Lam
dblp:13/4632
· DBLP profile ↗
96ranked-venue papers
6as first author
82since 2021 · last 2026
0000-0002-8022-7744ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 19 since 2021Computer networks · 18 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Human-computer interaction and ubiquitous computing · 11 · 10 since 2021Security and privacy · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward a Unified Architecture for Smart Home Energy Monitoring: Requirements, Design, and Use-Case ValidationabstractThe increasing deployment of smart devices in residential environments opens new opportunities for intelligent energy management. However, existing platforms often fall short in providing intuitive interfaces, zone-level control, and advanced predictive analytics accessible to non-expert users. This paper presents the design of a modular smart home energy management system that integrates real-time monitoring, consumption forecasting, and intelligent assistance via Large Language Models (LLMs). The system features an interactive floor plan interface, multi-user support, threshold-based alerting, and detailed historical analytics. Additionally, it introduces LLM-powered agents that guide users in configuring smart devices and adopting more efficient consumption behaviors. This architecture emphasizes accessibility, adaptability, and extensibility, aiming to empower users with actionable insights and seamless device management. The proposed solution addresses current gaps in existing platforms and lays the groundwork for intelligent, personalized, and proactive home energy systems. Manuel Andruccioli, Kelvin Olaiya, Alex Testa, Salvatore Bennici, Rares Vasiliu, Cui Congwen, Lin Jingzhe, Lou Kuok Keon, Bao Rui, Wang Taoyuan, Cheng Xinyuan, Paola Salomoni, Vittorio Ghini, Chan-Tong Lam, Su-Kit Tang, Giovanni Delnevo |
CCNC | 15 |
| 2026 | Language-Driven Autonomy for Sustainable Consumer Robotics: Toward Energy- and Data-Efficient LLM ReasoningabstractLarge Language Models (LLMs) are rapidly being embedded in consumer and service robots, enabling richer human–robot interaction, multimodal reasoning, and language-driven autonomy. However, the computational and lifecycle costs of training, inference, and continuous upgrade cycles raise urgent digital sustainability concerns: energy consumption, network dependency, privacy exposure, and hardware obsolescence. In this conceptual paper, we introduce the Sustainable Language-Driven Autonomy Framework (SLAF), a modular architecture and set of operational policies that align multimodal LLM reasoning with sustainability goals. SLAF decomposes the intelligence stack into Perception & Preprocessing, Local Cognition (Edge), High-level Reasoning (LLM), and Control & Execution layers, mediated by an Adapter responsible for compact semantic encoding, adaptive triggers, caching, and energy budgets. We propose quantitative primitives and trade-off models (e.g., energy-per-inference Einf, calls-per-mission Ncalls, mission energy Emission) and an evaluation protocol to make sustainability claims comparable and auditable. Finally, we map how SLAF addresses four research questions on zero-shot generalization, energy-efficient architectures, software-first lifespan extension, and cloud/on-device trade-offs. We conclude with a roadmap for empirical validation, lifecycle analysis, and user-centered studies to operationalize sustainable, language-enabled robotics. Kelvin Olaiya, Chan-Tong Lam, Silvia Mirri, Giovanni Pau 0001, Paola Salomoni |
CCNC | 2 |
| 2026 | Motivation in Programming Education: A Comparative Analysis between Students from Portugal and Macao
Anabela Jesus Gomes, Tânia Garbin, Carlos Alberto Dainese, Calana Chan, Philip Lei, Chan-Tong Lam, Ana Rosa Pereira Borges, Fernanda Brito Correia, António J. Mendes |
CSEDU (3) | 6 |
| 2026 | From Centralized Learning to Federated Setting: Keeping Reliability on Track
Junjian Yan, Paulo Carvalho 0001, Jorge Henriques, João Loureiro, Chan-Tong Lam, Henrique Madeira |
DSN | 5 |
| 2026 | Generating pivot Gray codes for spanning trees of complete graphs in constant amortized timeabstractWe present the first known pivot Gray code for spanning trees of complete graphs, listing all spanning trees such that consecutive trees differ by pivoting a single edge around a vertex. This pivot Gray code thus addresses an open problem posed by Knuth in The Art of Computer Programming, Volume 4 (Exercise 101, Section 7.2.1.6, [Knuth 2011]), rated at a difficulty level of 46 out of 50, and imposes stricter conditions than existing revolving-door or edge-exchange Gray codes for spanning trees of complete graphs. Our recursive algorithm generates each spanning tree in constant amortized time using \(O(n^2)\) space. In addition, we provide a novel proof of Cayley’s formula, \(n^{n-2}\), for the number of spanning trees in a complete graph, derived from our recursive approach. We extend the algorithm to generate edge-exchange Gray codes for general graphs with \(n\) vertices, achieving \(O(n^2)\) time per tree using \(O(n^2)\) space. For specific graph classes, the algorithm can be optimized to generate edge-exchange Gray codes for spanning trees in constant amortized time per tree for complete bipartite graphs, \(O(n)\)-amortized time per tree for fan graphs, and \(O(n)\)-amortized time per tree for wheel graphs, all using \(O(n^2)\) space. Bowie Liu, Dennis Wong, Chan-Tong Lam, Sio Kei Im |
SODA | 3 |
| 2026 | A noise-assistant network for tampering detection via inconspicuous feature enhancement and multi-perspective perception
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang, Nuno Lourenço 0002 |
Expert Syst. Appl. | 3 |
| 2026 | Complementarity in software code complexity metrics
Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira |
J. Syst. Softw. | 5 |
| 2026 | Which domain fits best? domain similarity measures for two-step heterogeneous transfer learning for early laryngeal cancer diagnosisabstractHeterogeneous transfer learning is an effective approach for medical imaging problems with limited data and scarce public homogeneous resources, yet selecting the optimal domain for feature extraction remains an open, often intuition-driven challenge. This study proposes and validates a set of quantitative domain similarity measurements to a priori identify the most suitable intermediate domain for early laryngeal cancer detection within a two-step heterogeneous transfer learning (THTL) framework, thereby avoiding computationally expensive trial-and-error training. We introduce eight domain similarity measurements to access the similarity between intermediate domains and the target domain. Multiple common medical imaging modalities, including angiography, chest radiographs, lung computed tomography (CT), brain magnetic resonance imaging (MRI), pathological section images, diabetic retinopathy fundus images, skin lesion images, and gastroenteroscopy, are served as candidate intermediate domains. The resulting similarity scores are ranked and compared with actual THTL performance rankings. Finally, we employ normalized discounted cumulative gain (NDCG) to determine the most predictive measurement. Our findings reveal that Earth Mover’s Distance (EMD) is the most effective domain similarity measurement for grayscale images, while cosine similarity based on global features extracted from a convolutional neural network (CNN) is optimal for RGB images. Using these measurements, angiography and skin lesion images are identified as the most beneficial intermediate domains. This work establishes a validated, data-driven methodology that enables future researchers to replace subjective intuition in domain selection, thereby saving substantial computational resources while improving model performance. Xinyi Fang, Yuqi Luo, Kei Long Wong, Benjamin K. Ng, Chan-Tong Lam, Marco Simões |
Knowl. Based Syst. | 5 |
| 2026 | SEM-UCSNet: A Novel Semantic Maps-Guided Compressive Sensing Framework for Underwater ImagesabstractUnderwater images (UWIs) captured by underwater detectors are essential for underwater detection and exploration. The compressive sensing theory (CS) provides a method for recovering images from few measurements, and it has been proven to be suitable for underwater environments with narrow bandwidth and limited communication channel resources, which may have a significant negative impact on the quality of captured UWIs. However, most existing state-of-art CS methods do not take the characteristics of UWIs into account, so their performance is limited in underwater applications. Compared with on-land images, UWIs have the following characteristics: 1) UWIs contain relatively few semantics, with a large amount of similar feature within the same semantics; 2) The importance of different semantics in UWIs is closely related to the underwater imaging model. In this paper, we combine the underwater imaging model and semantic of UWIs with CS task and propose a novel semantic maps-guided CS framework for UWIs, dubbed SEM-UCSNet, which can improve the performance of sampling and reconstruction, especially under extremely low sampling rate. In the sampling stage, a semantic importance analysis module combined with the imaging model is designed to guide the sampling. In the reconstruction process, a graph-based reconstruction strategy guided by semantic maps is proposed to model all features under the same semantic and mine complementarity between them to improve the reconstruction quality. Simultaneously, we introduce GAN into the underwater CS reconstruction task and use sampled features as conditions to make the reconstructed UWIs have richer details. Experimental results on some real-world UWIs datasets have demonstrated the superiority of our SEM-UCSNet on both objective and subjective metrics. Lihao Zhuang, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | BRPDNet: A BioRegion Prompt Distillation Network for Physiological MonitoringabstractPhysiological signal extraction from video data is challenging in dynamic and occluded environments, requiring both accuracy and real-time performance. Existing methods struggle to balance accuracy with model efficiency, particularly under partial facial occlusion or redundant signals. We propose BRPDNet, a novel framework for efficient physiological signal extraction which includes a BioRegion Prompt module for adaptive convolution and a Hyper Distillation module to reduce signal redundancy, ensuring high accuracy and robustness, especially in dynamic and occluded environments. Additionally, the teacher-student network structure enhances the model's adaptability to occlusions and reduces computational complexity without relying on explicit segmentation. Experimental results show that BRPDNet outperforms state-of-the-art models in accuracy, robustness, and efficiency across multiple datasets. For instance, BRPDNet achieves an Mean Absolute Error (MAE) of 1.55 beats per minute (bpm) and a Pearson Correlation Coefficient (PCC) of 0.76 on PURE and UBFC-rPPG datasets with fewer parameters than existing models, ensuring efficient real-time performance. Zhengxuan Chen, Bin Huang 0014, Kangyang Cao, Tao Tan 0002, Bingsheng Huang, Chan-Tong Lam, Yue Sun 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | HRMamba: Fusing Luminance Information for Remote Physiological Measurement in Varied Lighting ConditionsabstractCamera-based photoplethysmography (cbPPG) represents a non-invasive technique for capturing physiological parameters through facial videos, enabling the extraction of vital signs such as heart rate, respiration rate, and blood oxygen saturation without direct physical contact. Existing deep learning methods face two core challenges when dealing with cbPPG: firstly, extracting weak PPG signals from video segments with large spatial and temporal redundancy and understanding their periodic patterns in long contexts; secondly, accurately extracting PPG signals in complex lighting environments, especially in low-light conditions. To address these issues, this paper proposes an end-to-end method based on Mamba, named HRMamba. This method employs temporal difference mamba to process temporal signals and combines bidirectional state space to enable Mamba to robustly understand the scene and learn the periodic patterns of PPG. Furthermore, a luminance post-processing module is designed to extract luminance information from the video without enhancing lighting or altering the original video data, and embed it into the PPG signal. Experimental results demonstrate that HRMamba achieves state-of-the-art performance, and the designed luminance post-processing module can be applied in various lighting environments, significantly enhancing the performance in dark environments without degrading the performance in normal light scenes. Nuoer Long, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002, Zitong Yu, Yue Sun 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | An Advanced Index Modulation Scheme for V2I Communications: Reflection Modulation-Assisted Adaptive Space Shift KeyingabstractIn Intelligent Transportation Systems (ITS), reliable Vehicle-to-Infrastructure (V2I) communication is critical for real-time traffic management, collision avoidance, and environmental monitoring in high-mobility urban scenarios. To support real-time intersection collision-warning and traffic-signal-priority in ITS, we propose an RIS-assisted index modulation scheme that simultaneously increases road-safety message reliability and reduces roadside unit deployment density, termed Reflection Modulation-Assisted Adaptive Space Shift Keying (RM-ASSK). The RM-ASSK scheme leverages reconfigurable intelligent surface (RIS) technology to enhance the efficiency and reliability of wireless communication within transportation systems. By integrating RIS with adaptive space shift keying, the scheme aims to improve spectral efficiency and reduce bit-error rate performance. Meanwhile, we present theoretical analyses of the scheme’s spectral efficiency, average BER upper bounds, detection complexity, and ergodic channel capacity lower bounds. These analyses demonstrate the RM-ASSK scheme’s superior performance in terms of spectral efficiency, reduced complexity, and improved BER. Besides, to further enhance the scheme’s effectiveness and flexibility, we propose a novel antenna-combination selection algorithm, mixture optimized antenna combination selection, and a reflection-element subset design. These innovations optimize BER performance while balancing the trade-off between performance and complexity. Simulations show that RM-ASSK reduces data error rates by up to 25% in Nakagami-$m$and Rayleigh fading channels typical of urban V2I, enabling 15-20% faster dissemination of traffic alerts and improving vehicle throughput by supporting 30% more connected devices without spectral congestion. Chaorong Zhang, Benjamin K. Ng, Chan-Tong Lam |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | TSNN: A Non-Parametric and Interpretable Framework for Traffic Time Series ForecastingabstractAlthough many complex models were proposed to analyze time series data, some studies have demonstrated remarkable performance with simpler structures. A recent study proposed a non-parametric framework for 3D point cloud classification, which has the potential to be adapted for time series forecasting and enable interpretability. Inspired by the previous works, we present TSNN, a non-parametric and interpretable framework for traffic time series forecasting. TSNN consists of multiple layers that decouple the time series by matching the entries in a memory bank, where the memory bank is constructed using a similar matching process within the training set. It leverages the periodicity in traffic data to enhance forecasting accuracy while maintaining a simple model architecture. The proposed model operates without trainable parameters, preserving its inherent interpretability. In the experiments, TSNN achieves competitive performance compared to the typical deep learning models in four real-world traffic flow datasets. We also visualize the decoupling process to show the effectiveness of the components. Finally, we demonstrate the interpretability of the model and illustrate the contribution of each time step within the memory bank. Our code is available athttps://github.com/pzzzzzm/TSNN_release. Bowie Liu, Haijian Lai, Chan-Tong Lam, Junhao Dong 0004, Benjamin K. Ng, Wei Ke 0001, Sio Kei Im |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior RecognitionabstractBehavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset. Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001 |
IEEE Trans. Multim. | 8 |
| 2025 | SHIELDNet: Multi-Region Fusion and Denoising for Enhanced rPPG Signal Extraction in Healthcare MonitoringabstractAccurate extraction of Remote Photoplethysmography (rPPG) signals from video data is critical for medical applications such as remote patient monitoring. However, the process is hindered by significant challenges, including noise interference, occlusions, and multi bio-region signal processing. To address these, we propose SHIELDNet, an efficient and robust framework for real-time extraction of rPPG signals from multiple anatomical regions, incorporating advanced noise reduction mechanisms. SHIELDNet integrates a novel Differential Attention (DA) module, which adaptively focuses on multiple anatomical regions, enabling the model to effectively handle dynamic real-world conditions. Additionally, the network leverages an advanced Efficient Space Attention Module (ESAM) to enhance spatial feature extraction and multi bio-region signal fusion. BioRegion Prompt Module (BRPM) is further introduced to prioritize region-specific features, reducing the model's dependence on facial features alone. Futhermore, we introduce M-rPPG dataset, a comprehensive multi bio-region reference for BioRegion-based studies with full-body details at higher resolution than existing datasets. Extensive evaluations on multiple public datasets demonstrate significant improvements in Mean Absolute Error (MAE$=\mathbf{5. 2 8} \mathbf{~ b p m}$) and Pearson Correlation Coefficient ($\mathbf{P C C} \boldsymbol{=} \mathbf{0. 8 0}$), outperforming current state-of-the-art models. SHIELDNet provides an effective solution for noncontact, multi bio-region rPPG monitoring. We will release our code upon acceptance. Zhengxuan Chen, Tao Tan 0002, Chan-Tong Lam, Yue Sun 0001 |
BIBM | 5 |
| 2025 | Experiments of Crowd Detection for Crowd Digital TwinsabstractThe development of a crowd digital twin offers significant potential for enhancing public safety, urban planning, and event management. A key challenge in creating such a digital twin lies in the efficient and accurate acquisition of crowd-related data, particularly through object detection models deployed on resource-constrained devices. Through a series of experiments, we compare TinyML and Edge approaches in terms of detection accuracy, inferencing time, and resource utilization. Our findings highlight the trade-offs inherent in selecting detection models for crowd digital twin applications, underscoring the importance of aligning model choice with specific deployment needs. Kuan Pok Chong, Chon Hou Lai, Weibo Ling, Zhuoqian Lu, Yanjun Yu, Alex Testa, Chan-Tong Lam, Su-Kit Tang, Giovanni Delnevo, Roberto Casadei, Roberto Girau, Silvia Mirri |
CCNC | 9 |
| 2025 | Generating a Cyclic 2-Gray Code for Lucas Words in Constant Amortized Time
Bowie Liu, Dennis Wong, Chan-Tong Lam, Sio Kei Im |
CPM | 3 |
| 2025 | MoEdit: On Learning Quantity Perception for Multi-object Image EditingabstractMulti-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliaryfree multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github.com/Tear-kitty/MoEdit. Ka-Hou Chan, Yue Sun 0001, Chan-Tong Lam, Tong Tong 0001, Zitong Yu, Keren Fu, Xiaohong Liu 0001, Tao Tan 0002 |
CVPR | 4 |
| 2025 | NRevisit: A Cognitive Behavioral Metric for Code Understandability AssessmentabstractMeasuring code understandability is both highly relevant and exceptionally challenging. This paper proposes a dynamic code understandability assessment method, which estimates a personalized code understandability score from the perspective of the specific programmer handling the code. The method consists of dynamically dividing the code unit under development or review in code regions (invisible to the programmer) and using the number of revisits (NRevisit) to each region as the primary feature for estimating the code understandability score. This approach removes the uncertainty related to the concept of a "typical programmer" assumed by static software code complexity metrics and can be easily implemented using a simple, low-cost, and non-intrusive desktop eye tracker or even a standard computer camera. This metric was evaluated using cognitive load measured through electroencephalography (EEG) in a controlled experiment with 35 programmers. Results show a very high correlation ranging from rs = 0.9067 to rs = 0.9860 (with p nearly 0) between the scores obtained with different alternatives of NRevisit and the ground truth represented by the EEG measurements of programmers’ cognitive load, demonstrating the effectiveness of our approach in reflecting the cognitive effort required for code comprehension. The paper also discusses possible practical applications of NRevisit, including its use in the context of AI-generated code, which is already widely used today. Hao Gao 0002, Haytham Hijazi, Júlio Medeiros, João Durães, Chan-Tong Lam, Paulo Carvalho 0001, Henrique Madeira |
EASE | 5 |
| 2025 | Comparison of Data Imputation Performance in Deep Generative Models for Educational Tabular Missing Data
Wan-Chong Choi, Chan-Tong Lam, António J. Mendes |
EDM | 2 |
| 2025 | SAM-FE: Segment Anything Model Guided Feature Enhancement for Semantic Change Detection of Remote Sensing ImagesabstractSemantic change detection (SCD) is a crucial research topic in remote sensing. To achieve high-precision semantic segmentation results, Segment Anything Model-Guided Feature Enhancement (SAM-FE) is proposed. SAM-FE utilizes Mobile-SAM to extract features from bi-temporal remote sensing images (RSIs). In addition, the cross-temporal feature aggregation module (CTFA), the multiscale contextual information fusion module (MCIF), and the change feature enhancement module (CFE) are utilized to enhance the general features and the representation of the change information of the RSIs, thus improving the accuracy of change detection. Experimental results indicate that SAM-FE significantly outperforms the existing methods in both the Second datasets and MusSCD, with F1 of 0.6241 and 0.8316, respectively. Meanwhile, SAM-FE maintains lower parameters, demonstrating its superiority and practicality. Junqing Huang, Tong Liu 0021, Chan-Tong Lam, Xiaochen Yuan |
ICME | 3 |
| 2025 | Exploring the Capabilities and Limitations of Large Language Models for Zero-Shot Human-Robot InteractionabstractHuman-robot interaction (HRI) is an evolving field with a growing emphasis on enabling robots to understand and perform tasks based on natural language commands. Recently, Large Language Models (LLMs) have emerged as a promising tool for such tasks, offering the potential to enable zero-shot learning and flexible interaction without task-specific training. In this paper, we explore the use of LLMs for zero-shot navigation and exploration tasks in robotic systems, specifically evaluating their performance with the PR2 Clearpath and Khepera IV robots in a simulated environment. Our findings demonstrate promising results, particularly in the LLM’s ability to exhibit exploratory behavior and iterative reasoning when faced with ambiguous or incomplete visual input. These capabilities suggest a strong potential for LLMs in human-robot interaction. However, challenges were also identified, such as difficulties with target recognition, object misidentification, hallucination of information, and issues with movement execution, highlighting the need for improvements in these areas for real-world applications. Kelvin Olaiya, Giovanni Delnevo, Chan-Tong Lam, Giovanni Pau 0001, Paola Salomoni |
ISCC | 3 |
| 2025 | A Systematic Literature Review of Explainable Artificial Intelligence (XAI) for Interpreting Student Performance Prediction in Computer Science and STEM EducationabstractEducational Data Mining (EDM) supports early detection of learning difficulties by predicting student performance. However, machine learning models often operate as black boxes. Explainable Artificial Intelligence (XAI) helps to explain why black-box models produce specific predictions. This paper systematically reviews the past five years of research on XAI applications for interpreting student performance prediction in Computer Science and STEM education. We found that behavioral and academic performance data were the most commonly used features, with the main prediction goals focused on course failure risk or grades. This study also examined the application areas of XAI, revealing that the most common uses were global feature importance analysis, individual prediction explanations, and supporting interventions and decision-making. Moreover, we found that SHapley Additive exPlanations (SHAP) were the most frequently utilized XAI technique, predominantly applied at the global level, with limited use at the individual level. Furthermore, a research gap was identified in utilizing XAI to support course improvements, customize visualizations, and generate personalized recommendations. Addressing this gap could enable educators to provide personalized, data-driven guidance to better support individual students. Wan-Chong Choi, Chan-Tong Lam, Patrick Pang 0001, António J. Mendes |
ITiCSE (1) | 2 |
| 2025 | FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002 |
MICCAI (3) | 8 |
| 2025 | RefineNet: Elevating Medical Foundation Models Through Quality-Centric Data Curation by MLLM-Annotated Proxy Distillation
Ningyi Zhang, Xin Wang 0121, Ka-Hou Chan, Jian Wu 0033, Chan-Tong Lam, Shanshan Wang 0010, Yue Sun 0001, Sio Kei Im, Tao Tan 0002 |
MICCAI (11) | 6 |
| 2025 | Bandwidth-Aware Adaptive Gradient Quantization for Cross-Organization Federated Learning
Hong Shen 0001, Chan-Tong Lam, Ka Lun Eddie Law |
Networking | 3 |
| 2025 | MIFNet: Mamba-Based Information Fusion Network for Remote Sensing Change Detection
Yichen Cui, Hong Shen 0001, Chan-Tong Lam |
PDCAT | 3 |
| 2025 | Blockchain-Assisted Lightweight Secure Aggregation in Federated Learning via Trust-Aware Client Selection
Hong Shen 0001, Ka Lun Eddie Law, Chan-Tong Lam |
PDCAT | 4 |
| 2025 | Blockchain-Assisted Lightweight Secure Aggregation in Federated Learning via Trust-Aware Client Selection
Hong Shen 0001, Ka Lun Eddie Law, Chan-Tong Lam |
PDCAT | 4 |
| 2025 | Radar Signal Recognition Based on DAVG-GRN Network
Zeyu Tang 0008, Hong Shen 0001, Chan-Tong Lam |
PDCAT | 3 |
| 2025 | Dual-Scale Motion Extraction for Enhanced Human Action Recognition Based on RGB and Skeleton Modalities
Hong Shen 0001, Chan-Tong Lam |
PDCAT | 3 |
| 2025 | When Transformer Meets CSI Feedback in mMIMO Systems: A Lightweight CsiMobileViT ApproachabstractAccurate channel state information (CSI) feedback is essential in frequency division duplex massive multiple-input multiple-output systems, but increasing antennas cause the CSI matrix to grow exponentially, leading to significant feedback overhead. Inspired by the success of Transformers in natural language processing, recent Transformer-based CSI feedback methods have achieved excellent performance, though often with high computational costs that hinder real-time deployment on terminal devices. To address this challenge, in this paper, we present CsiMobileVit, a lightweight network that lowers computational complexity while maintaining reconstruction accuracy. The method achieves a good balance between simplicity and accuracy, making it practical for resource-limited devices. Extensive experiments confirm the effectiveness of this network. Xiangyu Cen, Chan-Tong Lam, Benjamin K. Ng, Ke Wang 0059 |
VTC2025-Fall | 2 |
| 2025 | RIS-Assisted Received Adaptive Spatial Modulation for Wireless CommunicationsabstractA novel wireless transmission scheme, as named the reconfigurable intelligent surface (RIS)-assisted received adaptive spatial modulation (RASM) scheme, is proposed in this paper. In this scheme, the adaptive spatial modulation (ASM)-based antennas selection works at the receiver by employing the characteristics of the RIS in each time slot, where the signal-to-noise ratio at specific selected antennas can be further enhanced with near few powers. Besides for the bits from constellation symbols, the extra bits can be mapped into the indices of receive antenna combinations and conveyed to the receiver through the ASM-based antenna-combination selection, thus providing higher spectral efficiency. To explicitly present the RASM scheme, the analytical performance of bit error rate of it is discussed in this paper. As a trade-off selection, the proposed scheme shows higher spectral efficiency and remains the satisfactory error performance. Simulation and analytical results demonstrate the better performance and exhibit more potential to apply in practical wireless communication. Chaorong Zhang, Benjamin K. Ng, Chan-Tong Lam, Ke Wang 0059 |
WCNC | 4 |
| 2025 | Maximum Channel Coding Rate of Finite Block Length MIMO Faster-Than-Nyquist SignalingabstractThe pursuit of higher data rates and efficient spectrum utilization in modern communication technologies necessitates novel solutions. In order to provide insights into improving spectral efficiency and reducing latency, this study investigates the maximum channel coding rate (MCCR) of finite block length (FBL) multiple-input multiple-output faster-than-Nyquist (FTN) channels. By optimizing power allocation, we derive the system's MCCR expression. Simulation results are compared with the existing literature to reveal the benefits of FTN in FBL transmission. Melda Yuksel, Halim Yanikomeroglu, Benjamin K. Ng, Chan-Tong Lam |
WCNC | 5 |
| 2025 | Reparameterization convolutional neural networks for handling imbalanced datasets in solar panel fault classification
Jielong Guo, Chak Fong Chong, Pedro H. Abreu, Chao Mao, Jiaxuan Li 0003, Chan-Tong Lam, Benjamin K. Ng |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | An algorithm for improving lower bounds in dynamic time warping
Yuqi Luo, Xinyi Fang, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Luís Paquete |
Expert Syst. Appl. | 4 |
| 2025 | Bi-branch bidirectional coupled interaction fusion network for multi-retinal diseases diagnosis
Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Xiayu Xu, Yanwu Xu 0004, Chan-Tong Lam, Yue Sun 0001 |
Knowl. Based Syst. | 8 |
| 2025 | An orchestration learning framework for ultrasound imaging: Prompt-Guided Hyper-Perception and Attention-Matching Downstream Synchronization
Shuo Li 0001, Shanshan Wang 0010, Zhifan Gao, Yue Sun 0001, Chan-Tong Lam, Xindi Hu, Xin Yang 0009, Dong Ni 0001, Tao Tan 0002 |
Medical Image Anal. | 6 |
| 2025 | RIS-assisted differential transmitted spatial modulation design
Chaorong Zhang, Benjamin K. Ng, Chan-Tong Lam |
Signal Process. | 4 |
| 2025 | Point-FCW: Transposed-FCW Graph Representation for Point Cloud Classification Using TDAabstractDual challenges of computational efficiency and representation effectiveness exist in processing point clouds. Inspired by the TDA (Topological Data Analysis), we propose to convert the point cloud to a transposed fully connected and weighted (t-FCW) graph in order to significantly decrease the computational complexity in the following processing steps. We design a TDA pipeline called Point-FCW with a series of vectorization techniques for the 3D object point cloud feature extraction, which is plugged into the non-parametric classification head. Our experimental results demonstrate that Point-FCW achieves 75.28% accuracy on the ModelNet40 dataset with 512 points, providing a tiny, consistent, and effective representation for TDA. Furthermore, when integrated with the state-of-the-art non-parametric network Point-NN, the mixture model performs better, with an improvement of 4.47% in the OBJ-BG split of the ScanObjectNN dataset. Similarly, when integrating Point-FCW into the parametric network, PointMLP yields a performance improvement of 3.54% in the PB-T50-RS split of the ScanObjectNN dataset. The proposed Point-FCW can serve as a complementary enhancement feature when integrated into the Point-NN and PointMLP models. Moreover, the t-FCW graph representation can be efficiently converted at a rate of 3739 samples/second. Our code is available inhttps://github.com/hawkinglai/Point-FCW. Haijian Lai, Bowie Liu, Chan-Tong Lam, Benjamin K. Ng, Sio Kei Im |
IEEE Signal Process. Lett. | 3 |
| 2025 | Recursive and iterative approaches to generate rotation Gray codes for stamp foldings and semi-meanders
Bowie Liu, Dennis Wong, Chan-Tong Lam, Marcus Im |
Theor. Comput. Sci. | 3 |
| 2025 | TransHFC: Joints Hypergraph Filtering Convolution and Transformer Framework for TemporalForgery LocalizationabstractThe authenticity of audio-visual content is being challenged by advanced multimedia editing technologies inspired by Artificial Intelligence-Generated Content (AIGC). Temporal forgery localization aims to detect suspicious contents by locating forged segments. So far, most of the existing methods are based on Convolutional Neural Networks (CNNs) or Transformers, yet neither of them has fully considered the complex relationships within forged audio-visual content. To address this issue, in this paper, we propose a novel method, named TransHFC, which innovatively introduces hypergraphs to model group relationships among segments while considering point-to-point relationships through Transformers. Through its dual hypergraph filtering convolution branch, TransHFC captures both temporal and spatial level group relationships, enhancing the representation of forged segment features. Furthermore, we propose a new hypergraph filtering convolution Auto-Encoder that uses a multi-frequency filter bank for adaptive signal capture. This design compensates for the limitation of a single hypergraph filter. Our extensive experiments on Lav-DF, TVIL, Psynd, and HAD datasets demonstrate that TransHFC achieves state-of-the-art performance. Xiaochen Yuan, Chan-Tong Lam, Sio Kei Im, Fangyuan Lei, Xiuli Bi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PFL-ALP: Personalized Federated Learning Against Backdoor Attacks via Attention-Based Local PurificationabstractFederated learning (FL) enables collaborative model training with local data privacy preserving, but is vulnerable to backdoor attacks from malicious clients. These attacks can manipulate the global model to produce malicious output when encountering specific triggers. Existing defenses, categorized as server-side and client-side approaches, have limitations such as reliance on auxiliary data availability, susceptibility to inference attacks, and instability under non-independent and identically distributed (Non-IID) data. In response to these challenges, we propose a Personalized Federated Learning via Attention-based Local Purification (PFL-ALP) algorithm, a hybrid defense mechanism integrating server-side dynamic clustering and client-side purification enhanced with personalized model knowledge. This approach effectively mitigates bias introduced by Non-IID data on the server side and further purifies the backdoored model on the client side. Specifically, we employ neural attention distillation (NAD) for model purification and enhance it with personalized model knowledge, extending the effectiveness of NAD in Non-IID FL settings. This design makes PFL-ALP compatible with privacy protocols to mitigate inference attacks. Moreover, we establish a convergence guarantee for PFL-ALP and experimentally validate its superior performance in defending against various backdoor attacks compared to multiple state-of-the-art (SOTA) defenses across three datasets. The results show that even with malicious rates ranging from 30% to 90%, PFL-ALP can reduce the attack success rate by more than 69.4 percentage points, with the reduction in main task accuracy less than 12.4 percentage points. Yifeng Jiang 0008, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Symmetric Self-Embedding Mechanism for High-Fidelity Image Recovery Against TamperingabstractDigital images are inherently fragile and vulnerable to malicious tampering, significantly compromising their authenticity and integrity. Image recovery is crucial for restoring altered content and preserving the reliability of digital images. Traditional fragile watermarking methods achieve high-quality recovery but fail under post-processing attacks, while existing deep learning-based approaches offer some robustness, yet often produce lower-quality recovered images, typically with a PSNR of around 28 dB. To address these challenges, we propose a novel Symmetric Self-embedding Mechanism for High-Fidelity Image Recovery against tampering (SSEM-HIR), which is capable of restoring tampered images with high quality while maintaining some robustness against common attacks. Unlike existing methods that use the fragility of watermarking solely for tampering localization, SSEM-HIR is the first work to integrate fragility with spatial symmetry, enabling high-quality tampering recovery. Specifically, our SSEM-HIR employs a hierarchical watermark embedding module to embed an inverted version of the original image, utilizing spatial symmetry to retrieve lost information from the extracted watermark. To further improve recovery quality, we design a Dual-branch Region-based Self-Recovery module, where a Spatial-based Watermark Extraction block restores tampered regions using embedded watermark information, while a Frequency-assisted Image Repair block compensates for quality degradation in the untampered area. Extensive experiments show that our method achieves an average PSNR of 34.14 dB under common attack scenarios, including noise addition, image scaling, Gaussian blurring, and no post-processing. This represents an improvement of over 5 dB and 18% in recovered image quality compared to state-of-the-art approaches. Tong Liu 0021, Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im, Pedro Martins 0003 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | 3MT-Net: A Multi-Modal Multi-Task Model for Breast Cancer and Pathological Subtype Classification Based on a Multicenter StudyabstractBreast cancer poses a significant threat to women's health, and ultrasound plays a critical role in the assessment of breast lesions. This study introduces a prospective deep learning architecture, termed the "Multi-modal Multi-task Network" (3MT-Net), which integrates clinical data with B-mode and color Doppler ultrasound images. Specifically, an AM-CapsNet is employed to extract key features from ultrasound images, while a cascaded cross-attention mechanism is utilized to fuse clinical data. Moreover, an ensemble learning approach with an optimization algorithm is adopted to dynamically assign weights to different modalities, accommodating both high-dimensional and low-dimensional data. The 3MT-Net performs binary classification of benign versus malignant lesions and further classifies the pathological subtypes. Data were retrospectively collected from nine medical centers to ensure the broad applicability of the 3MT-Net. Two separate testsets were created and extensive experiments were conducted. Comparative analyses demonstrated that the AUC of the 3MT-Net outperforms the industry-standard computer-aided detection product, S-Detect, by 1.4% to 3.8%. Yaofei Duan, Patrick Pang 0001, Rongsheng Wang 0004, Yue Sun 0001, Chuntao Liu, Xirong Yuan, Pengjie Song, Chan-Tong Lam, Ligang Cui, Tao Tan 0002 |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | Hierarchical Split Federated Learning: Convergence Analysis and System OptimizationabstractAs AI models expand in size, it has become increasingly challenging to deploy federated learning (FL) on resource-constrained edge devices. To tackle this issue,split federated learning(SFL) has emerged as an FL framework with reduced workload on edge devices via model splitting; it has received extensive attention from the research community in recent years. Nevertheless, most prior works on SFL focus only on a two-tier architecture without harnessing multi-tier cloud-edge computing resources. In this paper, we intend to analyze and optimize the learning performance of SFL under multi-tier systems. Specifically, we propose the hierarchical SFL (HSFL) framework and derive its convergence bound. Based on the theoretical results, we formulate a joint optimization problem for model splitting (MS) and model aggregation (MA). To solve this rather hard problem, we then decompose it into MS and MA sub-problems that can be solved via an iterative descending algorithm. Simulation results demonstrate that the tailored algorithm can effectively optimize MS and MA in multi-tier systems and significantly outperform existing schemes. Zheng Lin 0001, Wei Wei 0054, Zhe Chen 0015, Chan-Tong Lam, Xianhao Chen, Yue Gao 0001, Jun Luo 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Multimodal Interface for Games: A Case Study with TinyMLabstractMultimodal interfaces go beyond the traditional interaction through keyboard and mouse by incorporating multiple modes of interaction, such as touch, voice, gesture, and even gaze, to create more intuitive and immersive user experiences. This paper investigates how TinyML can be employed for multimodal interfaces in the context of games. An endless game in which the character has to avoid obstacles and fight monsters to advance has been developed. An Arduino Nano 33 BLE Sense is then used as the input device for the game by recognizing the hand gestures of the players. Haoxuan Xie, Lam Chi Hou, Lap Tou Chau, Lei Ka Weng, Xichen Wang, Giovanni Delnevo, Chiara Ceccarini, Chan-Tong Lam, Su-Kit Tang |
CCNC | 9 |
| 2024 | How Various Educational Features Influence Programming Performance in Primary School EducationabstractIn the digital age, programming education has become increasingly important, even in primary schools. However, introducing programming at such an early stage presents unique challenges, given the need for students to grasp mathematical concepts, abstract thinking, and the intricacies of programming syntax. Educational Data Mining (EDM) offers a potential contribution by predicting learning performance, facilitating the optimization of the learning processes, and providing real-time guidance. A notable gap in the current literature about EDM in programming education is its predominant emphasis on the university level. Our research objectives were to identify features influencing primary school students' programming capabilities. A more comprehensive dataset was introduced, incorporating psychometric data and highlighting features such as learning motivation and attitude, computational thinking data, and other potentially influential variables, which set our study apart from previous studies. We found that the strongest predictor was academic performance in Information Technology, followed by psychometric data on students' learning attitudes and motivation. Computational thinking also emerged as a significant feature in predicting programming performance. It's worth highlighting that involvement in extra-curricular activities, like Olympic Mathematics training, showed a significant association, underscoring the importance of mathematical logic and reasoning in programming. This is further bolstered by the evident correlation with academic performance in Mathematics, confirming its pivotal role in shaping programming abilities. Interestingly, the correlation of academic performance in Chinese is also significant, indicating that the language medium of instruction can notably influence success. Wan-Chong Choi, Chan-Tong Lam, António J. Mendes |
EDUCON | 2 |
| 2024 | Learning Sequencing with Bee-Bot: A Study on Improving Computational Thinking and Motivation for Young Learners in Programming EducationabstractThis Research-to-Practice full paper presents an exploratory study investigating the impact of using a Bee-Bot educational robot simulator to enhance learning sequencing concepts and student motivation among Macao primary school students. Sequencing in computational thinking (CT) is understanding and applying the logical order of steps in problem-solving processes. We introduced a Bee-Bot computer simulator for children to learn sequencing. Our study adopted a pretest-posttest method involving 35 grade two students. The Computational Thinking Test for Beginners (BCTt) was used to assess CT abilities, and the Instructional Materials Motivation Survey (IMMS) was utilized to measure learning motivation. We found a significant improvement in sequencing ability and more advanced CT concepts (loops and conditions) and a significant correlation between those concepts. Departing from the existing literature, we delved deeper into how Bee-Bot's influence on sequencing extended to more advanced CT concepts. Moreover, considering the ARCS motivation model, this study examined how Bee-Bot affects learning motivation at the primary education level. After the intervention, the findings revealed that the students showed significantly higher learning motivation, meaning that the different learning activities using the Bee-Bot simulator positively influenced various sub-dimensions of the ARCS model: attention, relevance, confidence, and satisfaction. The correlation between the IMMS scores and the BCTt outcomes further suggested that enhanced motivation positively correlated with better CT abilities. Wan-Chong Choi, Iek Chong Choi, Chan-Tong Lam, António J. Mendes |
FIE | 3 |
| 2024 | Enhance Learning Performance Predictions with Explainable Machine LearningabstractThis Research Full Paper focuses on predicting learning performance using machine learning algorithms and interpreting the results using Explainable Machine Learning (EML) techniques. The study compared a comprehensive set of machine learning algorithms, including Logistic Regression, Decision Trees, AdaBoost, XGBoost, SVM, and KNN. The performance of these algorithms in predicting students' final grades in a course was accessed using various evaluation metrics. Our study used feature selection to identify the most relevant predictors to enhance predictive accuracy, implemented the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance, and performed hyperparameter optimization to find the most effective model settings. This comprehensive approach improved the predictive accuracy of our models over previous studies. Additionally, the importance of early prediction in identifying at-risk students was explored, with models demonstrating promising accuracy at the first checkpoint of the course. Departing from traditional machine learning research that often focused on model performance, our study integrated the EML technique of Shapley Additive exPlanations (SHAP), which is grounded on the theoretical framework of Game Theory, to facilitate the interpretation of the predictive outcomes. This approach offered an explanatory perspective on the key factors influencing model decisions. By contributing to the predictability and interpretability of student performance, this research enriched the field of Educational Data Mining (EDM) and enhanced the understanding of student learning trajectories. Wan-Chong Choi, Chan-Tong Lam, António J. Mendes |
FIE | 2 |
| 2024 | Evolution of Motivational Factors During an Introductory Programming CourseabstractThis research-to-practice paper describes a study of motivational factors in introductory programming learning. Learning to program is challenging, as students need to develop multiple skills and competencies. Motivation drives students to confront complex challenges, persevere despite obstacles, and continuously strive for improvement. However, motivation is a complex interplay of internal and external factors. Analyzing the factors that can stimulate student motivation is essential for educators when planning and implementing learning activities and contexts. Therefore, we conducted a study to a) identify factors influencing the motivation of programming students and b) analyze the evolution of students' motivation during the different phases of a programming course. The study involved 137 students enrolled in a Programming I course at a Macao higher education institution. It used the motivation section of the Motivated Strategies for Learning Questionnaire (MSLQ), which comprises 31 statements grouped into six components (Intrinsic Goal Orientation (IGO), Extrinsic Goal Orientation (EGO), Value of Activity (VAT), Control of Learning (COL), Learning Self-Efficacy (LSE), and Test Anxiety (TAX)). These components can be organized into three factors (Value Components, Expectancy Components, and Affective Components). The students were asked to answer the questionnaire in three different moments: the initial phase of the course (3–4 weeks after its start), after knowing the results of the mid-term exam, and at the end of the course. For the analysis, only the answers of the 92 students who completed the questionnaire in the three phases were considered. We applied Principal Component Analysis (PCA) to identify the evolution of the different components and factors during the course. Based on this analysis, it is possible to highlight significant variations between the various phases of the study, especially concerning the factor of Value Components. In Phase 1, participants expressed a more positive perception of the importance of the course contents, as evidenced by the VAT component. In Phase 2, a change in focus was noticed, with the prioritization of obtaining a good grade, as reflected by the EGO component. Finally, in Phase 3, there was again a reorientation of value components, with students demonstrating appreciation for the course topic, as indicated again by the VAT component. Given these results, it is possible to conclude that changes occurred in the different phases of the study, suggesting an evolutionary dynamic in the interests of participants over time. Tânia Garbin, Carlos Alberto Dainese, Calana Chan, Philip Lei, Chan-Tong Lam, Anabela Jesus Gomes, António J. Mendes |
FIE | 5 |
| 2024 | A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image ClassificationsabstractAlthough current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propose a parameterized GAN (ParaGAN) that effectively controls the changes of synthetic samples among domains and highlights the attention regions for downstream classification. Specifically, ParaGAN incorporates projection distance parameters in cyclic projection and projects the source images to the decision boundary to obtain the class-difference maps. Our experiments show that ParaGAN can consistently outperform the existing augmentation methods with explainable classification on two small-scale medical datasets. Xiangyu Xiong, Yue Sun 0001, Xiaohong Liu 0001, Chan-Tong Lam, Tong Tong 0001, Hao Chen 0037, Qinquan Gao, Wei Ke 0001, Tao Tan 0002 |
ICASSP | 4 |
| 2024 | Quantum Robust Coding for Quantum Image Watermarking
Xiaochen Yuan, Chan-Tong Lam |
ICIC (10) | 3 |
| 2024 | MSFGNet: Multi-Scale Features Gathering Network for Change Detection of Remote Sensing ImagesabstractChange detection is an important research area in remote sensing. To achieve accurate results, it is essential to extract multi-scale spatial information from images while filtering out noise. However, existing models lack this capability. Therefore, Multi-Scale Feature Gathering Network (MSFGNet) is proposed. Within MSFGNet, Bi-Temporal Image Multi-Level Fusion Module (BMF) is utilized to fuse bi-temporal remote sensing images. Additionally, Multi-Receptive Field Features Extraction Module (MRFE) is utilized to extract deep features. Within MRFE, Large Receptive Field Features Extraction Module (LRFE) and Multi-Scale Information Fusion Module (MSIF) are designed, which use large kernel convolution and dilated convolution respectively to capture spatial information with large receptive fields. Furthermore, Cross-Dimension Feature Sifting Fusion Module (CDFSF) is designed to sift noise from various dimensions, fusing valuable information. Across multiple public datasets, MSFGNet consistently achieves the best experimental results. The code can be accessed at https://github.com/juncyan/msfgnet.git. Junqing Huang, Xiaochen Yuan, Chan-Tong Lam, Wei Ke 0001 |
ICME | 3 |
| 2024 | Dual Hypergraph Convolution Networks for Image Forgery Localization
Xiaochen Yuan, Wei Ke 0001, Chan-Tong Lam |
ICPR (22) | 4 |
| 2024 | Defending Against Backdoor Attacks with Feature Activation-Based Detection and Model RecoveryabstractThe characteristics of Federated Learning (FL) make FL highly susceptible to malicious poisoning attacks from adversaries. Existing FL defense methods can detect attacks of fixed poisoning patterns, hence are lack of flexibility. Additionally, they typically remove the malicious node models upon detection, leading to a certain degree of data loss. To address these issues, we propose a novel defense method against backdoor attacks in FL systems, effectively enhancing their robustness to malicious poisoning attacks. Specifically, we introduce a malicious update detection method based on feature activation matrices. This method compares the distribution differences of updates from different clients on the same validation data and detects malicious clients based on the outlier rates of their updates. Furthermore, to mitigate the data loss caused by the removal of malicious clients, the server assesses the distance between the distribution of feature activation matrices from the client’s historical updates and the overall model distribution in the current iteration. Based on this distance, the server performs model recovery to a certain extent. Extensive experiments on two benchmark datasets demonstrate that our method accurately detects malicious clients under various state-of-the-art model poisoning attacks. Additionally, the model recovery method provides a notable improvement to the system, ensuring the robustness and performance of the FL system. Hong Shen 0001, Chan-Tong Lam |
NCA | 3 |
| 2024 | DCAFNet: An Efficient Change Detection Structure for Remote Sensing Images
Yichen Cui, Hong Shen 0001, Chan-Tong Lam |
PDCAT | 3 |
| 2024 | Multi-scale TFT-Net Time-Frequency Representation for Multi-component Radar Signal Recognition
Zeyu Tang 0008, Hong Shen 0001, Chan-Tong Lam |
PDCAT | 3 |
| 2024 | Complex-Valued Neural Network Detection for RIS-Assisted Generalized Spatial ModulationabstractApplying the deep learning in signal processing for communication systems, several models based on Real-Valued Deep Neural Network (RVDNN) and convolutional neural network (RVCNN) have been previously proposed to detect signals of generalized spatial modulation. This paper proposes a complex-valued deep neural network (CVDNN) and a complex-valued convolutional neural network (CVCNN) as detectors for reconfigurable intelligent surface (RIS)-assisted generalized spatial modulation. Contrary to previous models, the complex-valued signals are directly fed into the neural network, which requires few feature vector generators and therefore has a simpler structure. Simulation results show that the proposed CVNN detectors exhibit improved error performance and stability for various modulation schemes compared with other traditional detection schemes over Nakagami-m fading channels. The results are shown to be approaching that of maximum likelihood detection, while outperforming existing RVNN detectors. Chaorong Zhang, Benjamin K. Ng, Chan-Tong Lam |
VTC Fall | 4 |
| 2024 | How Phase Errors Influence Phase-Dependent Amplitudes in Near-Field RISs?abstractNear-field reconfigurable intelligent surfaces (RISs) are unlocking promising potentials for the next generation of communications. Different from prior works that separately address phase shifts with errors and phase-dependent amplitudes (PDAs) in the RIS pixel hardware, this paper jointly studies power losses (PLs) caused by these two impairment factors. We propose three different pixel reflection models to accommodate different practical scenarios and derive their approximated upper bounds on the PL. It is important to note that neglecting uncertainties in the PDA may lead to an overestimation of the performance improvement offered by the RIS, thereby explaining the discrepancy between analytical and measurement results in several previous studies. Numerical simulations verify the correctness of the theoretical results. Ke Wang 0059, Rongbin Chen, Chan-Tong Lam, Benjamin K. Ng, Chaorong Zhang |
VTC Fall | 3 |
| 2024 | Definition and implementation of the Cloud Infrastructure for the integration of the Human Digital Twin in the Social Internet of ThingsabstractWith the integration of virtualization technologies, the Internet of Things (IoT) is expanding its capabilities and quickly becoming a complex ecosystem of networked devices. The Social Internet of Things (SIoT), where intelligent things include social properties that improve functioning and user engagement, is the result of this progress. The SIoT still has issues with scalability, data management, and user-centric operations, despite tremendous progress. In order to overcome these obstacles, a strong architecture is needed that can handle the enormous number of IoT devices while simultaneously streamlining the user interface. This study provides a unique architecture for the IoT that uses containerization to efficiently deploy and manage services while integrating Virtual Users (VUs) and Social Virtual Objects (SVOs) into a scalable Cloud/Edge infrastructure. These innovative aspects collectively advance previous works presented in literature and focused on novel SIoT architectures and implementations, by addressing key challenges in scalability, efficiency, and automation within the SIoT. The proposed method presents an extensible, modular architecture that lets VUs self-manage IoT services, making user administration easier and improving system security and scalability. Important parts of the design include a host controller for container orchestration, a deployer for automated service deployment, and user clusters for aggregating VUs, SVOs, and apps to provide secured and efficient data sharing. We show through experimental assessment that the architecture can manage high-volume installations and operating needs, exceeding the conventional platform based on Google App Engine in terms of system overhead and deployment timeframes. The obtained results highlight how our suggested architecture, which provides an easy-to-use, scalable, and secure foundation for IoT deployments, has the potential to advance the SIoT landscape. Roberto Girau, Matteo Anedda, Roberta Presta, Silvia Corpino, Pietro Ruiu, Mauro Fadda, Chan-Tong Lam, Daniele D. Giusto |
Comput. Networks | 7 |
| 2024 | EEBA: Efficient and ergonomic Big-Arm for distant object manipulation in VR
Jian Wu 0033, Lili Wang 0006, Sio Kei Im, Chan-Tong Lam |
Int. J. Hum. Comput. Stud. | 4 |
| 2024 | An accurate slicing method for dynamic time warping algorithm and the segment-level early abandoning optimization
Yuqi Luo, Wei Ke 0001, Chan-Tong Lam, Sio Kei Im |
Knowl. Based Syst. | 3 |
| 2024 | Recognition of score words in freestyle kayaking using improved DTW matching
Xiaochen Yuan, Chan-Tong Lam |
Multim. Tools Appl. | 3 |
| 2024 | MMQW: Multi-Modal Quantum Watermarking SchemeabstractTo address the problem that existing quantum image watermarking schemes have only a single watermarking mode with weak robustness, in this paper we propose a novel multi-modal quantum watermarking (MMQW) scheme using the generalized model of novel enhanced quantum representation. Our scheme provides four quantum watermarking modes (G_G, G_C, C_C, C_G), covering both types of grayscale and color images for the watermark and the carrier image. To enhance the robustness, we propose the Block Bit-plane Centrosymmetric Expansion (BBCE) method, which utilizes controlled quantum gates to extend the watermark, making our method resistant to noise and geometric attacks. Moreover, we propose a Brightness-based Watermarking Mechanism (BWM) for embedding and extraction. By uniform embedding, BWM not only minimizes the impact on the carrier image but also reduces the visual distortion of the extracted watermark. In the proposed MMQW, we implement three adaptive embedding strategies using controlled quantum gates, each of which is adaptively triggered according to the corresponding modalities. Detailed quantum circuits for quantum computing are provided. To evaluate imperceptibility and robustness of the MMQW, we conduct experiments using high-resolution images from the USC-SIPI dataset. The results show that PSNR of the watermarked image ranges from 36 dB to 56 dB, indicating the high visual quality. The PSNR of the extracted watermark is about 34 dB when the noise density is 0.05, while the PSNR is higher than 48 dB under common quantum rotation attacks, which indicate the high robustness against noise addition and geometric attacks. In addition, the proposed MMQW can resist to cropping attack with cropping percentage up to 55%. A comprehensive comparison with existing state-of-the-art works shows that our method has significant advantages. Chan-Tong Lam, Xiaochen Yuan, Sio Kei Im, Penousal Machado |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Characterization and Optimization of Coding Performance in Downlink NOMA With Finite-Alphabet Inputs and Finite BlocklengthabstractIn this paper, the channel coding performance with finite-alphabet inputs and finite blocklength in a two-user downlink non-orthogonal-multiple-access (NOMA) system is characterized from an information-theoretic perspective, and optimized with proper power allocation and constellation design. While previous works in NOMA were mainly focused on either infinite blocklength performance or finite blocklength performance with Gaussian inputs, we obtain the information-theoretic achievable rates accurate up to second-order as a function of the blocklength for both NOMA users employing finite-alphabet inputs subject to the average total power constraint. Taking advantage of the theoretical results, we formulate the rate and error-rate performance optimization problems in the finite blocklength regime for finite-alphabet inputs, from which the optimal power allocation and constellation-rotation can be derived. Furthermore, polar codes and LDPC are employed to demonstrate how close their performances are from the second-order achievable bound when the blocklength is short. Our results are important in machine-type or IoT communications where lightweight modulations and short blocklength are more relevant compared with traditional Gaussian-inputs assumption. Benjamin K. Ng, Chan-Tong Lam |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Real-scene-constrained virtual scene layout synthesis for mixed reality
Runze Fan, Lili Wang 0006, Xinda Liu, Sio Kei Im, Chan-Tong Lam |
Vis. Comput. | 5 |
| 2024 | ARGA-Unet: Advanced U-net segmentation model using residual grouped convolution and attention mechanism for brain tumor MRI image segmentationabstractMagnetic resonance imaging (MRI) has played an important role in the rapid growth of medical imaging diagnostic technology, especially in the diagnosis and treatment of brain tumors owing to its non-invasive characteristics and superior soft tissue contrast. However, brain tumors are characterized by high non-uniformity and non-obvious boundaries in MRI images because of their invasive and highly heterogeneous nature. In addition, the labeling of tumor areas is time-consuming and laborious. To address these issues, this study uses a residual grouped convolution module, convolutional block attention module, and bilinear interpolation upsampling method to improve the classical segmentation network U-net. The influence of network normalization, loss function, and network depth on segmentation performance is further considered. In the experiments, the Dice score of the proposed segmentation model reached 97.581%, which is 12.438% higher than that of traditional U-net, demonstrating the effective segmentation of MRI brain tumor images. In conclusion, we use the improved U-net network to achieve a good segmentation effect of brain tumor MRI images. Siyi Xun, Sixu Duan, Tong Tong 0001, Qinquan Gao, Chan-Tong Lam, Menghan Hu, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 8 |
| 2023 | A Systematic Literature Review on Performance Prediction in Learning Programming Using Educational Data MiningabstractProgramming education has become an essential skill for the digital generation. However, it presents a unique set of challenges that can be difficult for beginners. Educational data mining (EDM) has been increasingly utilized in programming education to enhance learning outcomes and understand students' learning behavior. By collecting and analyzing data from various sources, such as students' learning activities, interactions with learning resources, and assessment results, EDM can provide valuable insights into students' learning performance and potential areas for improvement. This paper presents a systematic literature review of recent literature (last five years) and reports on state of the art and trends in using EDM for student performance prediction in programming courses. It provides a comprehensive analysis of the input data used in previous work, exploring the different types of datasets used and the features that affect student performance. In addition, it addresses the predictive objectives and target variables for performance prediction in programming courses. On the other hand, it explores the most common prediction approaches, data pre-processing procedures, cross-validation methods, and evaluation metrics used to describe the performance of prediction algorithms. In addition, we discuss the limitations and challenges of various prediction approaches and provide valuable insights and directions for future research. Wan-Chong Choi, Chan-Tong Lam, António J. Mendes |
FIE | 2 |
| 2023 | MIMO Asynchronous MAC with Faster-than-Nyquist (FTN) SignalingabstractFaster-than-Nyquist (FTN) signaling is a non-orthogonal transmission technique, which brings in intentional inter-symbol interference. This way it can significantly enhance spectral efficiency for practical pulse shapes such as the root raised cosine pulses. This paper proposes an achievable rate region for the multiple antenna (MIMO) asynchronous multiple access channel (aMAC) with FTN signaling. The scheme applies waterfilling in the spatial domain and precoding in time. Water-filling in space provides better power allocation and precoding helps mitigate inter-symbol interference due to asynchronous transmission and FTN. The results show that the gains due to asynchronous transmission and FTN are more emphasized in MIMO aMAC than in single antenna aMAC. Moreover, FTN improves single-user rates, and asynchronous transmission improves the sum-rate, due to better inter-user interference management. Melda Yuksel, Halim Yanikomeroglu, Benjamin K. Ng, Chan-Tong Lam |
GLOBECOM | 5 |
| 2023 | A Noise Convolution Network for Tampering Detection
Zhiyao Xie, Xiaochen Yuan, Chan-Tong Lam, Guoheng Huang |
ICANN (10) | 3 |
| 2023 | Rapid APT Detection in Resource-Constrained IoT Devices Using Global Vision Federated Learning (GV-FL)
Han Zhu 0005, Chan-Tong Lam, Liyazhou Hu, Benjamin K. Ng, Kai Fang 0001 |
ICONIP (7) | 3 |
| 2023 | Locomotion-aware Foveated RenderingabstractOptimizing rendering performance improves the user's immersion in virtual scene exploration. Foveated rendering uses the features of the human visual system (HVS) to improve rendering performance without sacrificing perceptual visual quality. We collect and analyze the viewing motion of different locomotion methods, and describe the effects of these viewing motions on HVS's sensitivity, as well as the advantages of these effects that may bring to foveated rendering. Then we propose the locomotion-aware foveated rendering method (LaFR) to further accelerate foveated rendering by leveraging the advantages. In LaFR, we first introduce the framework of LaFR. Secondly, we propose an eccentricity-based shading rate controller that provides the shading rate control of the given region in foveated rendering. Thirdly, we propose a locomotion-aware log-polar mapping method, which controls the foveal average shading rate, the peripheral shading rate decrease speed, and the overall shading quantity with the locomotion-aware coefficients based on the eccentricity-based shading rate controller. LaFR achieves similar perceptual visual quality as the conventional foveated rendering while achieving up to 1.6× speedup. Compared with the full resolution rendering, LaFR achieves up to 3.8× speedup. Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Wei Ke 0001, Chan-Tong Lam |
VR | 5 |
| 2023 | How Long Can RIS Work Effectively: An Electronic Reliability PerspectiveabstractIn this paper, from an electronic reliability perspective, non-residual stochastic hardware aging (HA) effects are introduced to reconfigurable intelligent surfaces (RISs) aided communication systems, for characterizing the life cycle of the RIS. Different from traditional residual impairment factors such as RIS phase imperfections and transceiver noises, the impact of the stochastic HA effect on the RIS is related to RIS runtimes and lifetimes. Given this background, we first propose a Rician channel model for the RIS-aided system with the stochastic HA effect. Then, the definition for the lifetime of the RIS is mathematically given as the time at which 63.2% of the elements expire. Besides, closed-form achievable rate expression is also derived. Analytical and simulated results unveil an important insight that throughout the life cycle of the RIS, the residual impairment dominates when the runtime is shorter than the lifetime, otherwise the stochastic HA effect should be paid more attention to. This work can be regarded as the first guideline for evaluating and predicting the whole life cycle performance of the RIS-assisted system. Ke Wang 0059, Chan-Tong Lam, Benjamin K. Ng |
VTC Fall | 2 |
| 2023 | Wearable Real-time Air-writing System Employing KNN and Constrained Dynamic Time WarpingabstractIn the digital world, gesture recognition plays a crucial role in human-computer interaction (HCI). In this paper, we propose an innovative wearable air-writing system that allows users to write the English alphabet and Arabic numerals in free space without using any predefined gestures or rules. Based on an Inertial Measurement Unit (IMU), the proposed air-writing wearable device uses the constrained dynamic time warping (cDTW) algorithm for the distance measure and K-nearest neighbors (KNN) as the classifier. In addition, to increase the recognition accuracy and meet HCI requirements, we develop a novel method that allows users to rapidly switch to correct recognition results when the initial results are erroneous. In the experiment, the accuracy rate is 88.9% for the alphabet and 10 decimal digits in the user-dependent condition, and the recognition is in real-time, consuming only 0.427s for each character, which is superior to many other approaches that employ classic DTW or FastDTW. With the proposed HCI design, character input accuracy of over 95% can be obtained in about 1 second. We also simulated the application scenarios of Parkinson’s disease patients and obtained a high accuracy rate of 85.4%. Besides, we explored the variety of K values in KNN and w values in cDTW, and propose a multi-template system that gives new optimization directions for the KNN-cDTW algorithm. Yuqi Luo, Wei Ke 0001, Chan-Tong Lam |
WCNC | 3 |
| 2023 | Tampering localization and self-recovery using block labeling and adaptive significance
Xiaochen Yuan, Tong Liu 0021, Chan-Tong Lam, Guoheng Huang, Di Lin 0002, Ping Li 0016 |
Expert Syst. Appl. | 4 |
| 2023 | Bit-Interleaved Multiple Access: Improved Fairness, Reliability, and Latency for Massive IoT NetworksabstractInternet of Things (IoT) networks require massive connections in dense areas. Therefore, a resource-efficient multiple access scheme seems inevitable to enable immense connectivity where multiple devices have to share the same resource block (RB). Nonorthogonal multiple access (NOMA) has been considered as the strongest candidate in recent years. However, in this article, by considering the practical implementation, we first provide a true power allocation (PA) constraint with finite alphabet inputs for conventional downlink NOMA and demonstrate that it cannot support massive connections in practical systems. To this end, we propose the bit-interleaved multiple access (BIMA) scheme in downlink IoT networks. The proposed BIMA scheme implements bitwise multiaccess interleaving and deinterleaving at the transceiver ends and there are no strict PA constraints, unlike conventional NOMA, thus allowing a high number of devices in the same RB. We provide a comprehensive analytical framework for BIMA by investigating all key performance indicators (KPIs) to present both information-theoretic [i.e., ergodic capacity (EC) and outage probability (OP)] and finite alphabet inputs [i.e., bit error rate (BER)] performance metrics with both instantaneous and statistical channel ordering. In addition, we define Jain’s fairness index and proportional fairness index (PFI) in terms of all KPIs. Based on the extensive computer simulations, we reveal that BIMA outperforms conventional NOMA significantly, with a performance gain of up to 20–30 dB in terms of KPIs in some scenarios. In other words, compared to conventional NOMA schemes, the same KPIs are met in BIMA with 20–30 dB less transmit power, which is quite promising for energy-limited use cases. Moreover, this performance gain becomes greater when more IoT devices are supported. BIMA provides a full diversity order for all IoT devices and enables the implementation of an arbitrary number of devices and modulation orders, which is crucial for IoT networks where a huge number of devices should be supported in a single RB in dense areas. In addition to the overall performance gain, BIMA guarantees a fairness system where none of the devices gets a severely degraded performance and the sum-rate is shared in a fair manner among devices. It guarantees QoS satisfaction for all devices. Finally, we provide an intense complexity and latency analysis for BIMA and demonstrate that it provides lower latency compared to conventional NOMA receivers, since it allows parallel computation at the receivers and no iterative operations are required. We show that compared to conventional NOMA receivers, BIMA reduces latency by up to 350% for specific IoT devices and 170% on average. Ferdi Kara, Hakan Kaya, Halim Yanikomeroglu, Benjamin K. Ng, Chan-Tong Lam |
IEEE Internet Things J. | 5 |
| 2023 | PCTMF-Net: heart sound classification with parallel CNNs-transformer and second-order spectral analysisabstractHeart disease is a common condition worldwide and has become one of the leading causes of death worldwide. The electrocardiogram (PCG) is a safe, painless, and non-invasive test that captures bioacoustic information reflecting the function of the heart by capturing the acoustic signal of the patient’s heart. Nowadays, based on biosignal processing and artificial intelligence technologies, automated heart sound classification is playing an increasingly important role in clinical applications. In this paper, we propose a new parallel CNNs-transformer network with multi-scale feature context aggregation (PCTMF-Net). It combines the advantages of CNNs and transformer to achieve efficient heart sound classification. In PCTMF-Net, firstly, the heart tone signal features are extracted using the second-order spectral analysis, and a transformer-based MHTE-4 (multi-head transformer encoder with four attention heads) is designed to encode and aggregate the contextual information, and then, two CNNs feature extractors are designed in parallel with MHTE-4 to capture the hierarchical features. Finally, the feature vectors obtained from CNNs and MHTE-4 through feature fusion in PCTMF-Net will be fed into the fully connected layer for predicting the classification results of heart sounds. In addition, we perform validation based on two publicly available mutually exclusive heart sound datasets and conduct extensive experiments and comparisons of existing algorithms under different metrics. The experimental results show that our proposed method achieves 99.36% accuracy on the Yaseen dataset and 93% accuracy on the PhysioNet dataset. It surpasses current algorithms in terms of accuracy, recall and F 1-score metrics. The aim of this study is to apply these new techniques and methods to improve the diagnostic accuracy and validity of heart disease for clinical use. Rongsheng Wang 0004, Yaofei Duan, Dashun Zheng, Xiaohong Liu 0001, Chan-Tong Lam, Tao Tan 0002 |
Vis. Comput. | 6 |
| 2022 | Local perspective based synthesis for vehicle re-identification: A transformation state adversarial method
Yanbing Chen, Wei Ke 0001, Hong Lin 0006, Chan-Tong Lam, Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Exploring the Association between Self-Regulation of Learning and Programming Learning: A Multinational InvestigationabstractThis Research Full Paper presents a collection of evidences about the association between self-regulation variables and programming learning. Researchers have been investigating this thematic and despite the apparent benefits, it is necessary to summarize the published evidence and provide a new collection of them, which this study seeks to contribute. An observational investigation was performed in two countries with fifty-nine students, who had their SRL and programming learning metrics collected and correlated. Moreover, a systematic literature review was also conducted, and the existing evidence summarized. The results support an association between metacognitive and motivational regulatory strategies with programming learning but do not support for cognitive strategies. Our analysis shows a need for more studies to provide a solid body of knowledge on this thematic. Leonardo S. Silva, António J. Mendes, Anabela Jesus Gomes, Gabriel Fortes Cavalcanti de Macêdo, Chan-Tong Lam, Calana Chan |
FIE | 5 |
| 2021 | IRS-aided Predictable High-Mobility Vehicular Communication with Doppler Effect MitigationabstractIn this paper, we propose a novel anti-Doppler spread technique in IRS-aided predictable high-mobility vehicular communication. The predictable information of vehicle (e.g., speed, trajectory, and timetable) can be used to design a suboptimal phase shift set by solving a multi-objective problem, which can make IRS capable of maximizing instantaneous signal-to-noise ratio (SNR), minimizing Doppler spread, and keeping delay spread to a relatively low range simultaneously. The technique we proposed in this paper is more cost-effective, more practical, and of higher energy efficiency since the phase shift set by every element of the IRS for one vehicle pass can be designed and stored in advance, instead of processing in real-time. Simulation results show the effectiveness of the proposed phase shift set as compared to benchmark schemes. Ke Wang 0059, Chan-Tong Lam, Benjamin K. Ng |
VTC Spring | 2 |
| 2021 | Airtime Aware Dynamic Network Slicing for Heterogeneous IoT Services in IEEE 802.11ahabstractAssuring an efficient Quality of Service (QoS) for heterogeneous Internet of Things services is a challenge, mainly due to spectrum scarcity and bandwidth limitations on the radio access network. To solve these problems, efficient management of airtime per station (STA) is necessary. In this work, we propose a scheduler to perform dynamic network slicing in IEEE 802.11ah Networks. The proposed scheduler is based on virtualization technologies to assure QoS restrictions per slice and can be deployed in the Access Point (AP) or in a virtual machine connected to the AP. Network metrics are used to verify QoS violations per slice over time. Once a QoS restriction is detected, the scheduler performs a reallocation of resources re-configuring the Restricted Access Windows parameters that compose a slice, adjusting airtime per STA. To evaluate the proposed scheduler, we consider a dynamic policy. Simulation results in a smart city scenario show that the dynamic policy is able to scale the network and satisfies QoS slice requirements. Pedro Paulo Libório, Chan-Tong Lam, Benjamin K. Ng, Daniel L. Guidoni, Marília Curado, Leandro A. Villas |
WCNC | 2 |
| 2019 | Pedestrian Similarity Extraction to Improve People Counting AccuracyabstractCurrent state-of-the-art single shot object detection pipelines, composed by an object detector such as Yolo, generate multiple detections for each object, requiring a post-processing Non-Maxima Suppression (NMS) algorithm to remove redundant detections. However, this pipeline struggles to achieve high accuracy, particularly in object counting applications, due to a trade-off between precision and recall rates. A higher NMS threshold results in fewer detections suppressed and, consequently, in a higher recall rate, as well as lower precision and accuracy. In this paper, we have explored a new pedestrian detection pipeline which is more flexible, able to adapt to different scenarios and with improved precision and accuracy. A higher NMS threshold is used to retain all true detections and achieve a high recall rate for different scenarios, and a Pedestrian Similarity Extraction (PSE) algorithm is used to remove redundant detentions, consequently improving counting accuracy. The PSE algorithm significantly reduces the detection accuracy volatility and its dependency on NMS thresholds, improving the mean detection accuracy for different input datasets. Xu Yang 0010, José Gaspar, Wei Ke 0001, Chan-Tong Lam, Yanwei Zheng, Weng Hong Lou, Yapeng Wang 0001 |
ICPRAM | 4 |
| 2019 | Network Slicing in IEEE 802.11ahabstractRecently, Network Slicing in Radio Access Technologies (RATs) has been proposed as an approach to divide the wireless network infrastructure into isolated logical slices which are defined following their requirements and features in a service-driven way. To the best of our knowledge, the application of Network Slicing has not been explored in IEEE 802.11ah Networks. In this work, we propose a Virtual Network Slicing Broker (VNSB), a Virtual Network Function (VNF) instantiated inside a Software Defined Network (SDN) controller which communicates with an IEEE 802.11ah network Access Point (AP) via a southbound Application Program Interface (API). Based on information obtained by the northbound API, the slicing broker gets the information contained in the slicing templates, which describes the services features and respective Quality of Service (QoS) restrictions. With these data, the slicing broker makes logical slices in the context of the Restricted Access Windows (RAW) which are periodically sent by the AP to the stations (STA) via the Raw Parameter Set (RPS). Since there is no hardware with the IEEE 802.11ah standard available on the market, we validate the proposed solution through extensive simulations in a typical smart city scenario. Results showed the broker capacity to build and manage logical slices per service, respecting the QoS restrictions of each offered service. Moreover, the proposed slicing broker makes use of available resources of neighbor slices reducing delay, packet loss, and efficiently maximizing throughput per slice. Pedro Paulo Libório, Chan-Tong Lam, Benjamin K. Ng, Daniel L. Guidoni, Marília Curado, Leandro A. Villas |
NCA | 2 |
| 2018 | Student motivation towards learning to programabstractThis Research to Practice Full Paper presents a study on student's motivation towards learning to program. Motivation is a key factor in learning. Hence, stimulating student motivation strategies should be present in any pedagogical approach. This is particularly true in courses where a very active student attitude is fundamental. Introductory programming courses in higher education are a good example, which are known to be difficult for many students. To be successful students need to be motivated, as effort and commitment are necessary to overcome the difficulties many of them experience. In our study we analyzed several motivational aspects separately and then we correlated that information with the marks students obtained in introductory programming courses. We used two questionnaires. The Course Interest Survey (CIS) and the Instructional Materials Motivation Survey (IMMS). We could find some interesting correlations that confirm the importance of different motivational aspects to learning. We found other issues that demand more investigation, in order to create the best context to promote student motivation and learning. Anabela Jesus Gomes, Wei Ke 0001, Chan-Tong Lam, Maria José Marcelino, António J. Mendes |
FIE | 3 |
| 2017 | Fast spectrum sensing using a large number of receive antennasabstractIn this paper, we propose a novel spectrum sensing method to detect the presence of a primary user in a cognitive radio network when the secondary user is employing a large number of receive antennas. The new method is a suboptimal approach based on the generalized likelihood ratio test (GLRT), but unlike the previous GLRT-based methods which rely on large number of received samples obtained over time, we take advantage of the large number of samples available at the receive antennas. By exploiting the properties of the samples obtained at the receive antennas, effective spectrum sensing is achieved using mainly the spatial samples and therefore the sensing time is significantly reduced. With the suitable channel models and prior knowledge about the signal structure, the simulation results show that the proposed method offers performance comparable to well-known methods but with much smaller sensing time. Benjamin K. Ng, Chan-Tong Lam |
WoWMoM | 2 |
| 2013 | Effects of channel estimation errors on BER performance of DFT-precoded OFDM systemsabstractWe study the effects of channel estimation errors on bit error rate (BER) performance of a DFT-precoded OFDM system with frequency domain linear minimum mean square error (MMSE) equalizer, by analytically studying the effects of channel estimation errors on the mean square error (MSE) of the equalizer output, which can be related to the BER performance using a Gaussian approximation. We first obtain an approximate expression for the BER performance in terms of the MSE of the equalizer output with channel estimation errors, which is found to be a sum of the equalizer output MSE for the known channel frequency response and the scaled MSE of the channel frequency estimates. For a fixed MSE of channel frequency estimates, the degradation of the equalizer output signal-to-noise ratio (SNR) increases as Es/Noincreases. We also propose a semi-analytical approach to obtain the BER performance, by modeling the frequency channel estimates as the actual channel estimates plus a Gaussian random variable with zero mean and a variance of the averaged MSE of the channel frequency estimates. Simulation results show that both proposed approaches can accurately estimate the BER performance of DFT-precoded OFDM systems with channel estimation errors. Chan-Tong Lam, David D. Falconer |
IWCMC | 1 |
| 2013 | On maximizing the performance of SC-FDMA systems with rotated constellationabstractIn this paper, we consider the optimization of Single-Carrier FDMA (SC-FDMA) performance using rotated constellation. Using maximum-likelihood (ML) detection, it was previously shown that the diversity order of SC-FDMA is upper-bounded by the number of multi-paths or the number of subcarriers employed. As in OFDM, channel coding is usually required in SC-FDMA system in order to maximize the frequency diversity performance, especially when number of subcarriers is small. By using rotated constellation and the proposed angles of rotation, we show that the SC-FDMA system may attain the maximum diversity and coding performance without a reduction in bandwidth-efficiency or a significant change in the peak-to-average power ratio. Furthermore, when the number of subcarriers allocated in a SC-FDMA system is small, it becomes feasible to employ high-complexity detection methods such as ML or QR-decomposition method. In particular, to handle resource blocks of various sizes, we propose that the block of subcarriers be divided into pairs so that rotation can be separately applied in each pair over the block. As such, ML detection and flexible allocation of subcarriers can be facilitated for short or long resource block in the proposed SC-FDMA system. Simulations have shown that the proposed SC-FDMA system may outperform the ordinary SC-FDMA system and OFDM system in situations where coding rate is high. Benjamin K. Ng, Chan-Tong Lam |
IWCMC | 2 |
| 2008 | Joint Frequency-Domain Equalization and Channel Estimation Using Superimposed PilotsabstractSC modulations (single-carrier) with FDE (frequency-domain equalization) are promising candidates for future broadband wireless systems. These modulations allow excellent performances in severely time-dispersive channels, provided that accurate channel estimates are available at the receiver. For this purpose, pilot symbols and/or training sequences are usually multiplexed with data symbols. Since this leads to overhead and some spectral degradation, the use of superimposed pilots (i.e., pilots added to data) was recently proposed. In this paper we consider SC-FDE systems where the channel estimation is based on superimposed pilots. Since the interference levels between data and pilots might be very high, we propose an iterative receiver with joint equalization and channel estimation. Our performance results show that the use of superimposed pilots, combined with the proposed receiver, allows performances close to the case with perfect channel estimation, even for severely time-dispersive channels. Rui Dinis 0001, Chan-Tong Lam, David D. Falconer |
WCNC | 2 |
| 2008 | Iterative frequency domain channel estimation for dft-precoded ofdm systems using in-band pilotsabstractWe consider two techniques of in-band frequency domain multiplexed (FDM) pilots using interleaved frequency domain multiple access (IFDMA) signal with a Chu sequence for DFT-precoded OFDM (or single-carrier (SC)) systems. One, called frequency domain superimposed pilot technique (FDSPT), superimposes pilot tones onto scaled or deleted data tones, which preserves spectral efficiency at the expense of a slight performance loss. The other, called frequency expanding technique (FET), multiplexes pilot tones by displacing data tones, which slightly reduces spectral efficiency. Using FDM pilots in SC systems facilitates flexible and efficient assignment of signals to available spectrum. We propose an iterative frequency domain decision-directed interference cancellation technique to reduce the intersymbol interference level of SC signals with FDSPT pilots (resulting from the suppression of data tones). Moreover, we propose a low complexity frequency domain iterative decision-directed channel estimation (IDDCE) technique for SC systems using FDM pilots. Using IDDCE, the frame error rate (FER) performance for coded SC systems using FET and FDSPT pilots with interference cancellation is found to be about 0.2 dB and about 0.5 dB, respectively, away from the FER performance with known channel frequency response at FER=10-2. FDSPT pilots can also be used for OFDM systems with channel coding. It is found that an extra 1 dB of SNR is required at FER=10-2,compared with that using the conventional FET pilots for OFDM systems. Chan-Tong Lam, David D. Falconer, Florence Danilo-Lemoine |
IEEE J. Sel. Areas Commun. | 1 |
| 2007 | Design of Time and Frequency Domain Pilots for Generalized Multicarrier SystemsabstractBy the generalized multi-carrier (GMC) principle a unified framework to describe various multi-carrier as well as single carrier approaches is established. In this paper the GMC principle is extended to pilot design and channel estimation, so that a unified description of pilot aided channel estimation (PACE) by interpolation in time and frequency is established. This applies to frequency domain pilots which are embedded in the GMC signal, as well as pilots sequences time multiplexed with data-bearing GMC blocks. A comparative performance evaluation of the two different GMC variants OFDM and single carrier with various pilot allocation schemes is carried out, also taking into account practical constraints such as the peak to average power ratio (PAPR) and required power amplifier back-off. It is found that a serial modem with time or frequency multiplexed pilots could be designed with a power amplifier with about 2 dB lower maximum power rating than that of a corresponding OFDM modem. Chan-Tong Lam, Gunther Auer, Florence Danilo-Lemoine, David D. Falconer |
ICC | 1 |
| 2007 | A Low Complexity Frequency Domain Iterative Decision-Directed Channel Estimation Technique for Single-Carrier SystemsabstractA low complexity frequency domain iterative decision-directed channel estimation (FD-IDDCE) technique for single-carrier (SC) systems with frequency domain multiplexed (FDM) pilots is proposed. The tentative hard decisions from either the frequency domain equalizer output or the Viterbi decoder output are used as extra pilots to further improve the initial frequency channel estimates using a cascaded 2 times 1D Wiener filter. The issue of noise enhancement when finding the least square (LS) estimates of the channel frequency response (CFR) using the tentative decisions in the frequency domain is overcome by the so called frequency replacement algorithm, which replaces the noise enhanced LS estimates of the CFR with the corresponding channel frequency estimates in the previous iteration, by comparing with a threshold. Using the proposed FD-IDDCE, the frame error rate (FER) performance for coded SC systems using FDM pilots with frequency expanding technique and frequency domain superimposed pilot technique was found to be about 0.4 dB and about 1 dB away from the FER performance with known CFR at FER=10-2. Chan-Tong Lam, David D. Falconer, Florence Danilo-Lemoine |
VTC Spring | 1 |
| 2007 | PAPR Reduction using Frequency Domain Multiplexed Pilot SequencesabstractWe investigate the feasibility of applying the peak-to-average power ratio (PAPR) reduction method using pilot sequences, originally proposed for orthogonal frequency division multiplexing (OFDM) signals, for single-carrier (SC) signals with frequency domain multiplexed (FDM) pilots. The idea is to select the FDM pilot sequence with which the transmitted signal produces the lowest PAPR. We also investigate the applicability of the sum of square error (SSE) selection rule for high order modulation of SC signals. The SSE rule selects the pilot sequence which produces the minimum SSE between the transmitted signal and a pre-defined threshold, proportional to the saturation level of a high power amplifier (HPA). It is found that for both SC and OFDM systems, the PAPR reduction capabilities of orthogonal Walsh-Hadamard (W-H) sequences and cyclic shifted Chu (CS-Chu) sequences are similar for small block size, but not for large block size. Using CS-Chu sequences produces better PAPR reduction capability. With an appropriate choice of the value of input backoff power of a HPA, the SSE selection rule produces similar results as that of the minimum PAPR selection rule. The effects of out-of-band radiation for SC and OFDM signals with PAPR reduction using FDM pilot sequences after HPA depends on the amount of non-linearity portion of the HPA. The out-of-band radiation improvement is obvious for a HPA that approximates a linear clipper. Chan-Tong Lam, David D. Falconer, Florence Danilo-Lemoine |
WCNC | 1 |
| 2007 | Turbo frequency domain equalization for single-carrier broadband wireless systemsabstractIn this paper, a new class of equalization and channel estimation techniques, using the turbo frequency domain equalization (TFDE), is presented as a promising low-complexity detection method for single-carrier broadband wireless transmissions. Serial modulation (SM), being a direct counterpart of the well-known OFDM modulation, is receiving considerable attention recently, owing to the fact that it delivers comparable performance as OFDM while avoiding the problem of high peak-to-average power ratio. When frequency domain equalization (FDE) is applied, the complexity requirement is low and it becomes feasible to employ iterative processing which relies on decision feedback. This paper considers the turbo principle applied jointly to FDE, channel decoding and channel estimation. The result of this work is a set of effective iterative algorithms which may bring about 2-3 dB improvement over the linear FDE method. Furthermore, it is shown that they can provide performance comparable to the time-domain turbo equalization methods but with lower complexity. We then reach the conclusion that SM-based transmission with TFDE is a suitable technology for next generation wireless systems Benjamin K. Ng, Chan-Tong Lam, David D. Falconer |
IEEE Trans. Wirel. Commun. | 2 |
| 2006 | Channel Estimation for SC-FDE Systems Using Frequency Domain Multiplexed PilotsabstractWe investigate channel estimation for single-carrier-frequency domain equalization (SC-FDE) system using the techniques typically used for an orthogonal frequency domain multiplexing (OFDM) system. Two techniques of frequency domain multiplexed (FDM) pilot insertion using interleaved frequency domain multiple access (IFDMA) signal with a Chu sequence are considered. One called frequency domain superimposed pilot technique (FDSPT) scales data-carrying tones and then superimposes them with pilot tones. This technique preserves spectral efficiency at the expense of performance loss. The other, called frequency expanding technique (FET), shifts groups of data frequencies for multiplexing of pilot tones at the expense of spectral efficiency. Our results show that both techniques increase peak to average power ratio (PAPR) although it is still lower than that of an OFDM system. The application of FDSPT is limited by the pilot overhead ratio, resulting from the removal of data frequencies for pilot frequencies. It is shown that channel estimation using conventional time domain multiplexed pilots and FET pilot tones produce the same BER, while the FDSPT requires about 1.5 dB more power for the same performance. Using FDM pilots in SC system facilitates flexible and efficient assignment of signals to available spectrum. Chan-Tong Lam, David D. Falconer, Florence Danilo-Lemoine, Rui Dinis 0001 |
VTC Fall | 1 |
| 2004 | A multiple access scheme for the uplink of broadband wireless systemsabstractWe present a multiple access scheme for the uplink of broadband wireless systems. We consider SC (single carrier) based block transmission, with all users transmitting continuously, regardless of their data-rate. The transmitted signals can have very low envelope fluctuations. Moreover, the different users remain orthogonal, even for severe time-dispersive channels. By employing IB-DFE (iterative block decision feedback equalization) techniques we can have detection performances close to the matched filter bound, and even for fully loaded systems and severe time-dispersive channels. Rui Dinis 0001, David D. Falconer, Chan-Tong Lam, Maryam Sabbaghian |
GLOBECOM | 3 |