Jacqueline C. K. Lam

dblp:16/9987 · also Jacqueline Chi Kei Lam · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-8805-3574ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Databases, data management, data science and information retrieval · 5Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Unravelling Causal Genetic Biomarkers of Alzheimer's Disease via Neuron to Gene-token Backtracking in Neural Architecture: A Groundbreaking Reverse-Gene-Finder Approach
abstract
Alzheimer’s Disease (AD) affects over 55 million people globally, yet the key genetic contributors remain poorly understood. Leveraging recent advancements in genomic foundation models, we present the innovative Reverse-Gene-Finder technology, a ground-breaking neuron-to-gene-token backtracking approach in a neural network architecture to elucidate the novel causal genetic biomarkers driving AD onset. Reverse-Gene-Finder comprises three key innovations. Firstly, we exploit the observation that genes with the highest probability of causing AD, defined as the most causal genes (MCGs), must have the highest probability of activating those neurons with the highest probability of causing AD, defined as the most causal neurons (MCNs). Secondly, we utilize a gene token representation at the input layer to allow each gene (known or novel to AD) to be represented as a discrete and unique entity in the input space. Lastly, in contrast to the existing neural network architectures, which track neuron activations from the input layer to the output layer in a feed-forward manner, we develop an innovative backtracking method to track backwards from the MCNs to the input layer, identifying the Most Causal Tokens (MCTs) and the corresponding MCGs. Reverse-Gene-Finder is highly interpretable, generalizable, and adaptable, providing a promising avenue for application in other disease scenarios.
Victor O. K. Li, Yang Han 0006, Jacqueline C. K. Lam
AAAI3
2025 DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
abstract
Alzheimer’s Disease (AD) is an irreversible neurodegenerative disease affecting 50 million people worldwide. Low-cost, accurate identification of key markers of AD is crucial for timely diagnosis and intervention. Language impairment is one of the earliest signs of cognitive decline, which can be used to discriminate AD patients from normal control (NC) individuals. Patient-interviewer dialogues may be used to detect such impairments, but they are often mixed with ambiguous, noisy, and irrelevant information, making the AD detection task difficult. Moreover, the limited availability of AD speech samples and variability in their speech styles pose significant challenges in developing robust speech-based AD detection models. To address these challenges, we propose DECT, a novel speech-based domain-specific approach leveraging large language models (LLMs) for fine-grained linguistic analysis and label-switched label-preserved (LSLP) data generation. Our study presents four novelties: (1) We harness the summarizing capabilities of LLMs to identify and distill key Cognitive-Linguistic (CL) information (atoms) from noisy speech transcripts, effectively filtering irrelevant information. (2) We leverage the inherent linguistic knowledge of LLMs to extract linguistic markers from unstructured and heterogeneous audio transcripts. (3) We exploit the compositional ability of LLMs to generate LSLP AD speech transcripts consisting of diverse linguistic patterns to overcome the speech data scarcity challenge and enhance the robustness of AD detection models. (4) We use the augmented AD textual speech transcript dataset and a more fine-grained representation of AD textual speech transcript data to fine-tune the AD detection model. The results have shown that DECT, an integrated, LLM-assisted, speech-based AD detection model demonstrates superior model performance with an 11% improvement in AD detection accuracy on the datasets from DementiaBank compared to the baselines.
Tingyu Mo, Jacqueline C. K. Lam, Victor O. K. Li, Lawrence Y. L. Cheung
AAAI2
2025 CTDI: CNN-Transformer-Based Spatial-Temporal Missing Air Pollution Data Imputation
abstract
Accurate and comprehensive air pollution data is essential for understanding and addressing environmental challenges. Missing data can impair accurate analysis and decision-making. This study presents a novel approach, named CNN-Transformer-based Spatial-Temporal Data Imputation (CTDI), for imputing missing air pollution data. Data pre-processing incorporates observed air pollution data and related urban data to produce 24-hour period tensors as input samples. 1-by-1 CNN layers capture the interaction between different types of input data. Deep learning transformer architecture is employed in a spatial-temporal (S-T) transformer module to capture long-range dependencies and extract complex relationships in both spatial and temporal dimensions. Hong Kong air pollution data is statistically analyzed and used to evaluate CTDI in its recovery of generated and actual patterns of missing data. Experimental results show that CTDI consistently outperforms existing imputation methods across all evaluated scenarios, including cases with higher rates of missing data, thereby demonstrating its robustness and effectiveness in enhancing air quality monitoring. Additionally, ablation experiments reveal that each component significantly contributes to the model's performance, with the temporal transformer proving particularly crucial under varying rates of missing data.
Yangwen Yu, Victor O. K. Li, Jacqueline C. K. Lam, Kelvin Chan, Qi Zhang 0121
IEEE Trans. Big Data3
2023 Personalized Ambient Pollution Estimation Based on Stationary-Camera-Taken Images Under Cross-Camera Information Sharing in Smart City
abstract
Timely and high-density air quality monitoring is essential for the development of future smart cities. The images captured from widely deployed stationary-cameras can be transferred quickly via the Internet of Things (IoT) to facilitate ambient pollution estimation anytime anywhere. Image-based air pollution estimation is normally formulated as a supervised learning problem, relying on an extended number of image samples. However, individual stationary-cameras can offer only very limited samples and scenes, while locally trained estimation models can easily overfit. A global method was proposed to address this challenge. The global model was trained via images captured from different cameras. However, such a model is less effective in extracting local features from scenes. A personalized method is therefore proposed to improve not only the generalization of the estimation model but also to preserve the local characteristics of individual cameras. Our personalized method consists of a two-stage architecture: 1) images from different cameras are used to train the global estimation model to avoid overfitting due to fixed scenes and small sample size and 2) the global model is further refined by images captured from individual cameras separately for adapting local characteristics. To evaluate our proposed personalized method, a large data set was constructed, based on stationary-camera-taken images captured in Hong Kong, consisting of different pollution measurements, including PM2.5, PM10, NO2, and O3. As compared to the local model, our proposed personalized model has reduced average MAE by 5.68% and average SMAPE by 6.82%, and improved average$r$by 4.69%.
Shiguang Song, Victor O. K. Li, Jacqueline C. K. Lam, Yi Wang 0022
IEEE Internet Things J.3
2023 Distilling Region-Wise and Channel-Wise Deep Structural Facial Relationships for FAU (DSR-FAU) Intensity Estimation
abstract
Facial emotions are expressed through a combination of facial muscle movements, namely, the Facial Action Units (FAUs). FAU intensity estimation aims to estimate the intensity of a set of structurally dependent FAUs. Contrary to the existing works that focus on improving FAU intensity estimation performance, this study investigates how knowledge distillation (KD) incorporated into a training model can improve FAU intensity estimation efficiency while achieving the comparable level of performance. Given the intrinsic structural characteristics of FAU, it is desirable to distill deep structural relationships, namely, DSR-FAU, using heatmap regression. Our methodology is as follows: First, a feature map-level distillation loss is applied to ensure that the student network and the teacher network share similar feature distributions. Second, the region-wise and channel-wise relationship distillation loss functions are introduced to penalize the difference in structural relationships. Specifically, the region-wise relationship can be represented by the structural correlations across the facial features, whereas the channel-wise relationship is represented by the implicit FAU co-occurrence dependencies. Third, we compare the model performance of DSR-FAU with the state-of-the-art models, based on two benchmarking datasets. It is shown that our model achieves comparable performance, with a lower number of model parameters and lower computation complexities.
Yingruo Fan, Jacqueline C. K. Lam, Victor O. K. Li
IEEE Trans. Affect. Comput.2
2023 Hierarchical Recovery of Missing Air Pollution Data via Improved Long-Short Term Context Encoder Network
abstract
Due to equipment and transmission failures, data loss presents a key challenge to air quality monitoring. This paper attempts to recover missing air quality data from an air quality database. Leveraging adaptive updating convolutional neural networks (CNNs), we propose a novel Long-short term context encoder (ILSCE) model, which can simultaneously capture any temporal-spatial correlation and periodic variation identified from an air quality dataset. In addition, our model applies a new mechanism to automatically update both the air quality data and their corresponding masks in every single layer of CNN. Our proposed method presents three novelties. First, it hierarchically recovers any missing air quality values. Second, domain specific weekday/weekend and seasonal information are incorporated into the training model. Third, model performance is enhanced by an additional regularization term that captures the correlation between different air pollutants, thereby considering both background ambient pollution and local emissions. Our experimental study shows these three newly proposed features allow the ILSCE model to significantly outperform existing state-of-the-art imputation methods in air pollution data recovery. Furthermore, as data loss becomes more severe, with more missing data and more consecutively missing data, the superior recovery performance and greater robustness of our model become more prominent.
Yangwen Yu, Victor O. K. Li, Jacqueline C. K. Lam
IEEE Trans. Big Data3
2023 GCN-ST-MDIR: Graph Convolutional Network-Based Spatial-Temporal Missing Air Pollution Data Pattern Identification and Recovery
abstract
Missing data pattern identification and recovery (MDIR) is vital for accurate air pollution monitoring. To recover the missing air pollution data, GCN-ST-MDIR, a Graph Convolutional Network (GCN)-based MDIR framework, is proposed to identify daily missing data patterns and automatically select the best recovery method. GCN-ST-MDIR presents four novelties: (1) A new graph construction is developed to improve GCN data representation for MDIR using S-T similarity matrix and domain-specific knowledge (e.g., weekend/weekday). (2) A TL component is used to pre-train LSCE and ILSCE models. (3) A GCN structure outputs a selection indicator to determine the dominant missing pattern for daily input. The pre-trained data recovery model's accuracy is incorporated into the GCN loss function to penalize the wrong indicator. (4) The output of the GCN structure is used as a score to combine LSCE and ILSCE. Results show that the domain-specific S-T regularity and irregularity can be used as the prior information for both GCN and ILSCE/LSCE to enhance feature extraction. Our model considerably improves the recovery performance as compared to the baselines. GCN-ST-MDIR has achieved an accuracy of 88.48% for general missing data recovery with consecutively and sporadically missing data. GCN-ST-MDIR can be extended to many other S-T MDIR challenges.
Yangwen Yu, Victor O. K. Li, Jacqueline C. K. Lam, Kelvin Chan
IEEE Trans. Big Data3
2022 Facial Expression Recognition With Deeply-Supervised Attention Network
abstract
Facial expression recognition (FER) is crucial for social communication. However, current studies present limitations when addressing facial expression difference due to demographic variation, such as race, gender, and age, etc. In this article, we first propose a deeply-supervised attention network (DSAN) to recognize human emotions based on facial images automatically. Based on DSAN, a two-stage training scheme is designed, taking full advantage of the race/gender/age-related information. In our DSAN framework, multi-scale features are leveraged to capture more discriminative information from the deep layers to the shallow layers. Furthermore, we adopt the attention block to highlight the essential local facial characteristics; it performs well when it is incorporated into the deeply-supervised framework. Finally, we combine the complementary characteristics of multiple convolutional layers in deeply-supervised manner and ensemble the intermediate predicted scores. Our experimental results have shown that our proposed framework can (i) effectively integrate demographic information in improving the performance of a variety of FER tasks, (ii) learn informative feature representations with a visual explanation by capturing the regions of interests (ROI), (iii) achieve superior performance for both the posed and the spontaneous FER databases, each containing pictures of human facial expressions varied in gender, age or race.
Yingruo Fan, Victor O. K. Li, Jacqueline C. K. Lam
IEEE Trans. Affect. Comput.3
2022 A Domain-Specific Bayesian Deep-Learning Approach for Air Pollution Forecast
abstract
Predicting air pollution concentration is crucial and beneficial for public health. This study proposes a domain-specific Bayesian deep-learning model for long-term air pollution forecast in China and the United Kingdom. Our proposed model carries three novelties: First, a domain-specific knowledge is integrated to take into account the strong statistical relationship between PM$_{2.5}$and PM$_{10}$as a regularization term; Second, an attention layer is included to capture the influential historical feature and the recursive temporal correlation of air quality data; Third, results generated from different multi-step forecast strategies are combined based on corresponding uncertainty measures to improve our model’s performance. Our model outperforms other baseline models. Results show that incorporating Bayesian and domain-specific knowledge into the deep learning model can reduce the prediction errors by a maximum of 3.7% and 12.4%, for Beijing and London, respectively. Specifically, incorporating domain-specific knowledge into the Bayesian deep-learning model reduces prediction errors whilst the integration of Bayesian techniques allows the fusion of different forecast strategies to improve prediction accuracy. In future, additional influential domain-specific features can be added to further improve our deep-learning model’s prediction accuracy and interpretability.
Yang Han 0006, Jacqueline C. K. Lam, Victor O. K. Li, Qi Zhang 0121
IEEE Trans. Big Data2
2022 Missing Air Pollution Data Recovery Based on Long-Short Term Context Encoder
abstract
Air pollution has become a global challenge, and obtaining real-time air quality information is urgently needed. Although the governments have been trying their best in delivering accurate air quality reports, missing air pollution data remains a key challenge. Based on the temporal-spatial correlation of the data, we propose a novel long-short term context encoder (LSCE) structure for recovering missing air pollution data. The original context encoder approach based on image completion focuses on reconstructing rectangular missing regions. Differing from traditional methods, our fully convolutional neural network architecture enjoys the following novelties. First, LSCE can recover irregular missing data patterns. Second, we devise two data pre-processing strategies to produce two types of context encoders, namely, the long-short term cutting context encoder (LSCCE) and the long-short term sliding context encoder (LSSCE). Compared with LSCCE, LSSCE increases the number of training data matrixes. Finally, we investigate the significance of adaptive training in addressing different types of missing data. Our simulation results have demonstrated that our approach, especially, LSSCE, can outperform existing missing data recovery methods. Besides, our techniques can be widely applicable for recovering other temporally and spatially correlated missing data, such as vehicular traffic or meteorology data.
Yangwen Yu, Victor O. K. Li, Jacqueline C. K. Lam
IEEE Trans. Big Data3
2020 Facial Action Unit Intensity Estimation via Semantic Correspondence Learning with Dynamic Graph Convolution
abstract
The intensity estimation of facial action units (AUs) is challenging due to subtle changes in the person's facial appearance. Previous approaches mainly rely on probabilistic models or predefined rules for modeling co-occurrence relationships among AUs, leading to limited generalization. In contrast, we present a new learning framework that automatically learns the latent relationships of AUs via establishing semantic correspondences between feature maps. In the heatmap regression-based network, feature maps preserve rich semantic information associated with AU intensities and locations. Moreover, the AU co-occurring pattern can be reflected by activating a set of feature channels, where each channel encodes a specific visual pattern of AU. This motivates us to model the correlation among feature channels, which implicitly represents the co-occurrence relationship of AU intensity levels. Specifically, we introduce a semantic correspondence convolution (SCC) module to dynamically compute the correspondences from deep and low resolution feature maps, and thus enhancing the discriminability of features. The experimental results demonstrate the effectiveness and the superior performance of our method on two benchmark datasets.
Yingruo Fan, Jacqueline C. K. Lam, Victor O. K. Li
AAAI2
2018 Unsupervised Domain Adaptation with Generative Adversarial Networks for Facial Emotion Recognition
abstract
Cross-dataset facial emotion recognition (FER) aims to reduce the discrepancy between the source and the target facial database. The topic is very challenging in FER, where facial features differ across different domains, such as ethnicity, age, gender and environmental condition. In practice, the labels of target facial expression database may be unavailable, making it impossible to fine-tune a pre-trained model via supervised transfer learning. To address this issue, we propose an unsupervised domain adaptation framework with adversarial learning for cross-dataset FER. We perform cross-dataset FER on three well-known publicly available facial expression databases, viz. CK+, Oulu-CASIA, and RAF-DB, showcasing the efficiency of our proposed approach.
Yingruo Fan, Jacqueline C. K. Lam, Victor O. K. Li
IEEE BigData2
2018 A Bayesian LSTM Model to Evaluate the Effects of Air Pollution Control Regulations in China
abstract
Rapid socio-economic development and urbanization have resulted in serious deterioration in air-quality in many world cities, including Beijing, China. This preliminary study is the first attempt to examine the effectiveness of air pollution control regulations implemented in Beijing during 2013 - 2017 through a data-driven regulatory intervention analysis. Our proposed machine-learning model utilizes proxy data including Aerosol Optical Depth (AOD) and meteorology; it can explain 80% of the PM2.5variability. Our preliminary results show that air pollution control regulatory measures introduced in China and Beijing have reduced PM2.5pollution in Beijing by 23% on average.
Yang Han 0006, Jacqueline C. K. Lam, Victor O. K. Li
IEEE BigData2
2018 In Search of A Better Land: Would People Move to A Country with Better Air Quality? A Global Survey Based on Twitter Data
abstract
This study examines the statistical relationship between people's international movements and air quality. Utilizing randomized geo-tagged tweets obtained from Twitter Streaming API, we extract international movements that Twitter users have actually made. Coupling these movements with the air quality data across the world, we verify that Twitter users tend to move to countries of better air qualities. We also find that in most countries, the out-movement rate tends to be positively associated with the monthly air pollution concentration, and the number of movements to destination countries is negatively associated with the decrease in air pollution concentration in the origin countries.
Zhiyi Lu, Jacqueline C. K. Lam, Victor O. K. Li, Yang Han 0006, Zafar Gilani
IEEE BigData2
2018 Optimization of Urban Heating Network Design Using Genetic Algorithm
abstract
As the main energy source is coal burning, district heating in Northern China is an important driver of air pollution. Optimization of the performance of District Heating Network (DHN) carries both social and economic benefits. This study proposes an approach for optimizing urban heating network design based on Genetic Algorithm. Our case study shows that DHN can meet the users' requirements and achieve minimum cost in parallel.
Andong Wang, Victor O. K. Li, Jacqueline C. K. Lam
IEEE BigData3
2018 Multi-region Ensemble Convolutional Neural Network for Facial Expression Recognition
Yingruo Fan, Jacqueline C. K. Lam, Victor O. K. Li
ICANN (1)2
2018 Video-based Emotion Recognition Using Deeply-Supervised Neural Networks
abstract
Emotion recognition (ER) based on natural facial images/videos has been studied for some years and considered a comparatively hot topic in the field of affective computing. However, it remains a challenge to perform ER in the wild, given the noises generated from head pose, face deformation, and illumination variation. To address this challenge, motivated by recent progress in Convolutional Neural Network (CNN), we develop a novel deeply supervised CNN (DSN) architecture, taking the multi-level and multi-scale features extracted from different convolutional layers to provide a more advanced representation of ER. By embedding a series of side-output layers, our DSN model provides class-wise supervision and integrates predictions from multiple layers. Finally, our team ranked 3rd at the EmotiW 2018 challenge with our model achieving an accuracy of 61.1%.
Yingruo Fan, Jacqueline C. K. Lam, Victor O. K. Li
ICMI2
2018 Task Allocation in Spatial Crowdsourcing: Current State and Future Directions
abstract
Spatial crowdsourcing (SC) is an emerging paradigm of crowdsourcing, which commits workers to move to some particular locations to perform spatio-temporal-relevant tasks (e.g., sensing and activity organization). Task allocation or worker selection is a significant problem that may impact the quality of completion of SC tasks. Based on a conceptual model and generic framework of SC task allocation, this paper first gives a review of the current state of research in this field, including single task allocation, multiple task allocation, low-cost task allocation, and quality-enhanced task allocation. We further investigate the future trends and open issues of SC task allocation, including skill-based task allocation, group recommendation and collaboration, task composition and decomposition, and privacypreserving task allocation. Finally, we discuss the practical issues on real-world deployment as well as the challenges for large-scale user study in SC task allocation.
Bin Guo 0001, Yan Liu 0045, Leye Wang, Victor O. K. Li, Jacqueline C. K. Lam, Zhiwen Yu 0001
IEEE Internet Things J.5
2018 CrowdTracker: Optimized Urban Moving Object Tracking Using Mobile Crowd Sensing
abstract
This paper proposes CrowdTracker, a novel object tracking system based on mobile crowd sensing (MCS). Different from traditional video-based object tracking approaches, CrowdTracker recruits people to collaboratively take photographs of the object to achieve object movement prediction and tracking. The optimization objective of CrowdTracker is to effectively track the moving object in real time and minimize the cost on user incentives. Specifically, the incentive is determined by the number of workers assigned and the total distance that workers move to complete the task. In order to achieve the objective, we propose the movement prediction (MPRE) model for object movement prediction and two other algorithms for task allocation, namely, T-centric and P-centric. T-centric selects workers in a task-centric way, while P-centric allocates tasks in a peoplecentric manner. By analyzing a large number of historical vehicle trajectories, MPRE builds a model to predict the object's next position. In the predicted regions, CrowdTracker selects workers by utilizing T-centric or P-centric. We evaluate the algorithms over a large-scale real-world dataset. Experimental results indicate that CrowdTracker can effectively track the object with a low incentive cost.
Yao Jing, Bin Guo 0001, Zhu Wang 0001, Victor O. K. Li, Jacqueline C. K. Lam, Zhiwen Yu 0001
IEEE Internet Things J.5
2017 Low-rank singular value thresholding for recovering missing air quality data
abstract
With the increasing awareness of the harmful impacts of urban air pollution, air quality monitoring stations have been deployed in many metropolitan areas. These stations provide air quality data to the public. However, due to sampling device failures and data processing errors, missing data in air quality measurements is common. Data integrity becomes a critical challenge when such data are employed for public services. In this paper, we investigate the mathematical property of air quality measurements, and attempt to recover the missing data. First, we empirically study the low rank property of these measurements. Second, we formulate the low rank matrix completion (LRMC) optimization problem to reconstruct the missing air quality data. The problem is transformed using duality theory, and singular value thresholding (SVT) is employed to develop sub-optimal solutions. Third, to evaluate the performance of our methodology, we conduct a series of case studies including different types of missing data patterns. The simulation results demonstrate that the proposed SVT methodology can effectively recover missing air quality data, and outperform the existing Interpolation. Finally, we investigate the parameter sensitivity of SVT. Our study can serve as a guideline for missing data recovery in the real world.
Yangwen Yu, James Jian Qiao Yu, Victor O. K. Li, Jacqueline C. K. Lam
IEEE BigData4
2017 Deep Learning Model to Estimate Air Pollution Using M-BP to Fill in Missing Proxy Urban Data
abstract
Air quality has deteriorated rapidly in Hong Kong and China in the past two decades, with NO2and PM2.5levels frequently exceeding WHO safety guidelines. While poor air quality has clear public health impacts, there are very limited air quality monitoring (AQM) stations, severely constraining evidence-based air quality decision-making, leading to severe criticisms about the utility of the current official Air Quality Health Index to the public. Since air pollution is highly location-dependent, a city-wide deployment of traditional, highly sophisticated air quality monitors would be prohibitively expensive. In this paper, we propose a deep learning model to estimate air pollution throughout the city, utilizing the readily available urban data as proxy data. As with many big data driven approaches, the proxy data may be sparse/missing. We propose the M-BP algorithm to recover/fill in such missing data. Our results show that the proposed model gives better estimates compared with existing big data approaches.
Victor O. K. Li, Jacqueline C. K. Lam, Yun Chen 0007, Jiatao Gu
GLOBECOM2