Kenneth Lai

dblp:156/3037 · also Kenneth Kam Fai Lai · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0003-1513-5509ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Localizing Events in Space: Comparing Humans and AI Models
Derrick Eui Gyu Kim, Kenneth Lai, James Pustejovsky
LREC2
2026 Distributed Partial Information Puzzles: Examining Common Ground Construction under Epistemic Asymmetry
Yifan Zhu 0014, Mariah Bradford, Kenneth Lai, Timothy Obiso, Videep Venkatesha, James Pustejovsky, Nikhil Krishnaswamy
LREC3
2025 Speech Is Not Enough: Interpreting Nonverbal Indicators of Common Knowledge and Engagement
abstract
Our goal is to develop an AI Partner that can provide support for group problem solving and social dynamics. In multi-party working group environments, multimodal analytics is crucial for identifying non-verbal interactions of group members. In conjunction with their verbal participation, this creates an holistic understanding of collaboration and engagement that provides necessary context for the AI Partner. In this demo, we illustrate our present capabilities at detecting and tracking nonverbal behavior in student task-oriented interactions in the classroom, and the implications for tracking common ground and engagement.
Derek Palmer, Yifan Zhu 0014, Kenneth Lai, Hannah VanderHoeven, Mariah Bradford, Ibrahim Khebour, Carlos Mabrey, Jack Fitzgerald, Nikhil Krishnaswamy, Martha Palmer, James Pustejovsky
AAAI3
2024 Building a Broad Infrastructure for Uniform Meaning Representations
abstract
This paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word senses, aspectuality of events, as well as person and number information for entities. The document-level graph represents coreferential, temporal, and modal relations that go beyond sentence boundaries. UMR is designed to capture the commonalities and variations across languages and this is done through the use of a common set of abstract concepts, relations, and attributes as well as concrete concepts derived from words from invidual languages. This UMR release includes annotations for six languages (Arapaho, Chinese, English, Kukama, Navajo, Sanapana) that vary greatly in terms of their linguistic properties and resource availability. We also describe on-going efforts to enlarge this data set and extend it to other genres and modalities. We also briefly describe the available infrastructure (UMR annotation guidelines and tools) that others can use to create similar data sets.
Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft 0001, Lukas Denk, Sijia Ge, Jan Hajic 0001, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdenka Uresová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao 0009
LREC/COLING9
2024 Common Ground Tracking in Multimodal Dialogue
abstract
Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on “dialogue state tracking” (DST), which is the ability to update the representations of the speaker’s needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is “common ground tracking” (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ”questions under discussion” (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task.
Ibrahim Khebour, Kenneth Lai, Mariah Bradford, Yifan Zhu 0014, Richard Brutti, Christopher Tam, Jingxuan Tu, Benjamin Ibarra, Nathaniel Blanchard, Nikhil Krishnaswamy, James Pustejovsky
LREC/COLING2
2024 Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR
abstract
Abstract Meaning Representation (AMR) is a general-purpose meaning representation that has become popular for its clear structure, ease of annotation and available corpora, and overall expressiveness. While AMR was designed to represent sentence meaning in English text, recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture. In this paper, we present an annotated corpus of multimodal (speech and gesture) AMR in a task-based setting. Our corpus is multilayered, containing temporal alignments to both the speech signal and to descriptions of gesture morphology. We also capture coreference relationships across modalities, enabling fine-grained analysis of how the semantics of gesture and natural language interact. We discuss challenges that arise when identifying cross-modal coreference and anaphora, as well as in creating and evaluating multimodal corpora in general. Although we find AMR’s abstraction away from surface form (in both language and gesture) occasionally too coarse-grained to capture certain cross-modal interactions, we believe its flexibility allows for future work to fill in these gaps. Our corpus and annotation guidelines are available at https://github.com/klai12/encoding-gesture-multimodal-dialogue.
Kenneth Lai, Richard Brutti, Lucia Donatelli, James Pustejovsky
LREC/COLING1
2024 Causal Inference in Deep Learning Forecasting: A Bayesian Approach to Analyzing the Relationship between Employment Rate and Immigration
abstract
Mass migration, caused by events such as climate change or global conflicts significantly impacts every aspect of society including social infrastructures, public service delivery, education, healthcare, and security to list a few which in turn can affect the employment market. This paper presents an innovative methodology to illustrate the causal relationship between migration and employment rates in Canada. We propose a combination of training recurrent neural networks and temporal convolutional networks for employment trend forecasting and populating Bayesian networks to reveal interrelationships across various causal networks. Applying our method shows how changes in demographics such as sex, age, education, and the country of origin of displaced individuals influence the employment rate of the host country.
Kenneth Lai, Gregor Wolbring, Svetlana N. Yanushkevich
IJCNN1
2023 Assessing Upper Limb Motor Function in the Immediate Post-Stroke Period using Accelerometry
abstract
Recent advancements in machine learning have enabled the use of long-term accelerometry data collection and machine learning algorithms to quickly and accurately detect upper limb weakness. Although accelerometry-derived measurements are commonly used in long-term rehabilitation studies, this study aimed to determine whether similar techniques could be used to detect short-term changes in upper limb motor function in patients who were hospitalized soon after experiencing a stroke. Six binary classification models were created by training on variable data window times of paretic upper limb accelerometer feature data, and four preliminary visualizations were proposed to provide health professionals with information on the duration, intensity, symmetry, and variability of upper limb activity. The models were evaluated using Area Under the Curve (AUC) scores to classify the data into two classes: severe or moderately severe motor function. The AUC scores ranged from 0.72 to 0.94, with higher scores indicating better model performance. While this study provides a preliminary assessment of the efficacy of using accelerometry and machine learning to characterize upper limb motor function immediately following a stroke, the results suggest that further investigation is warranted.
Mackenzie Wallich, Kenneth Lai, Svetlana N. Yanushkevich
SMC2
2022 Hand Gesture Classification on Praxis Dataset: Trading Accuracy for Expense
abstract
In this paper, we investigate hand gesture classifiers that rely upon the abstracted 'skeletal' data recorded using the RGB-Depth sensor. We focus on 'skeletal' data represented by the body joint coordinates, from the Praxis dataset. The PRAXIS dataset contains recordings of patients with cortical pathologies such as Alzheimer's disease, performing a Praxis test under the direction of a clinician. In this paper, we propose hand gesture classifiers that are more effective with the PRAXIS dataset than previously proposed models. Body joint data offers a compressed form of data that can be analyzed specifically for hand gesture recognition. Using a combination of windowing techniques with deep learning architecture such as a Recurrent Neural Network (RNN), we achieved an overall accuracy of 70.8% using only body joint data. In addition, we investigated a long-short-term-memory (LSTM) to extract and analyze the movement of the joints through time to recognize the hand gestures being performed and achieved a gesture recognition rate of 74.3% and 67.3% for static and dynamic gestures, respectively. The proposed approach contributed to the task of developing an automated, accurate, and inexpensive approach to diagnosing cortical pathologies for multiple healthcare applications.
Rahat Islam, Kenneth Lai, Svetlana N. Yanushkevich
IJCNN2
2022 Abstract Meaning Representation for Gesture
abstract
This paper presents Gesture AMR, an extension to Abstract Meaning Representation (AMR), that captures the meaning of gesture. In developing Gesture AMR, we consider how gesture form and meaning relate; how gesture packages meaning both independently and in interaction with speech; and how the meaning of gesture is temporally and contextually determined. Our case study for developing Gesture AMR is a focused human-human shared task to build block structures. We develop an initial taxonomy of gesture act relations that adheres to AMR’s existing focus on predicate-argument structure while integrating meaningful elements unique to gesture. Pilot annotation shows Gesture AMR to be more challenging than standard AMR, and illustrates the need for more work on representation of dialogue and multimodal meaning. We discuss challenges of adapting an existing meaning representation to non-speech-based modalities and outline several avenues for expanding Gesture AMR.
Richard Brutti, Lucia Donatelli, Kenneth Lai, James Pustejovsky
LREC3
2021 Capturing causality and bias in human action recognition
Kenneth Lai, Svetlana N. Yanushkevich, Vlad P. Shmerko, Ming Hou 0002
Pattern Recognit. Lett.1
2020 A Two-Level Interpretation of Modality in Human-Robot Dialogue
abstract
We analyze the use and interpretation of modal expressions in a corpus of situated human-robot dialogue and ask how to effectively represent these expressions for automatic learning and dynamic interpretation in context.We present a two-level annotation scheme for modality that captures both content and intent, integrating a logic-based, semantic representation and a task-oriented, pragmatic representation that maps to our robot's capabilities.Data from our annotation task reveals that the interpretation of modal expressions in human-robot dialogue is quite diverse, yet highly constrained by the physical environment and asymmetrical speaker/addressee relationship.We sketch a formal model of human-robot common ground in which modality can be grounded and dynamically interpreted relative to speaker role, temporal constraints, and physical environment.
Lucia Donatelli, Kenneth Lai, James Pustejovsky
COLING2
2020 An Ensemble of Knowledge Sharing Models for Dynamic Hand Gesture Recognition
abstract
The focus of this paper is dynamic gesture recognition in the context of the interaction between humans and machines. We propose a model consisting of two sub-networks, a transformer and an ordered-neuron long-short-term-memory (ON-LSTM) based recurrent neural network (RNN). Each sub-network is trained to perform the task of gesture recognition using only skeleton joints. Since each sub-network extracts different types of features due to the difference in architecture, the knowledge can be shared between the sub-networks. Through knowledge distillation, the features and predictions from each sub-network are fused together into a new fusion classifier. In addition, a cyclical learning rate can be used to generate a series of models that are combined in an ensemble, in order to yield a more generalizable prediction. The proposed ensemble of knowledge-sharing models exhibits an overall accuracy of 86.11% using only skeleton information, as tested using the Dynamic Hand Gesture-14/28 dataset.
Kenneth Lai, Svetlana N. Yanushkevich
IJCNN1
2020 Decision Support for Video-based Detection of Flu Symptoms
abstract
The development of decision support systems is a growing domain that can be applied in the area of disease control and diagnostics. Using video-based surveillance data, skeleton features are extracted to perform action recognition, specifically the detection and recognition of coughing and sneezing motions. Providing evidence of flu-like symptoms, a decision support system based on causal networks is capable of providing the operator with vital information for decision-making. A modified residual temporal convolutional network is proposed for action recognition using skeleton features. This paper addresses the capability of using results from a machine-learning model as evidence for a cognitive decision support system. We propose risk and trust measures as a metric to bridge between machine-learning and machine-reasoning. We provide experiments on evaluating the performance of the proposed network and how these performance measures can be combined with risk to generate trust.
Kenneth Lai, Svetlana N. Yanushkevich
SMC1
2020 Reliability of Decision Support in Cross-spectral Biometric-enabled Systems
abstract
This paper addresses the evaluation of the performance of face and facial expression biometrics through a decision support system. The evaluation criteria include capturing the risk of the system, estimating the reliability of decision, and predicting the change in the perceived operator's trust in the decision. The relevant applications include human behavior monitoring and stress detection in individuals and teams, and in situational awareness system. Using an available database of cross-spectral videos of faces and facial expressions, we conducted a series of experiments to 1) demonstrate the phenomenon of biases in biometrics that affect the evaluated measures of the performance in human-machine systems, 2) explore the overall risk of the system caused by error rates such as false match and false nonmatch rates, and 3) calculate the reliability of cross-spectral and emotion-varying face identification.
Kenneth Lai, Svetlana N. Yanushkevich, Vlad P. Shmerko
SMC1
2020 Instance Segmentation of Personal Protective Equipment using a Multi-stage Transfer Learning Process
abstract
This paper focuses on the instance segmentation of soft attributes on humans such as clothing and personal protective equipment at a hazardous workplace. We propose the use of soft biometric object classes from the Open Images V5 and DeepFashion2 datasets to pre-train a mask segmentation network to detect and segment personal protective equipment in the workplace. Preliminary results of our proposed model achieves a mean average precision, mAP50, of 61.7% with minimal optimization, resulting in very good segmentation of construction helmets, high visibility vests, welding masks, and ear protection in the workplace. Applications of the results from this paper include improving workplace safety in hazardous industries by providing a tool to ensure proper personal protective equipment usage while maintaining worker anonymity.
Thomas Truong, Aakash Bhatt, Leonardo Queiroz, Kenneth Lai, Svetlana N. Yanushkevich
SMC4
2019 Dog Identification using Soft Biometrics and Neural Networks
abstract
This paper addresses the problem of biometric identification of animals, specifically dogs. We apply advanced machine learning models such as deep neural network on the photographs of pets in order to determine the pet identity. In this paper, we explore the possibility of using different types of "soft" biometrics, such as breed, height, or gender, in fusion with "hard" biometrics such as photographs of the pet's face. We apply the principle of transfer learning on different Convolutional Neural Networks, in order to create a network designed specifically for breed classification. The proposed network is able to achieve an accuracy of 90.80% and 91.29% when differentiating between the two dog breeds, for two different datasets. Without the use of "soft" biometrics, the identification rate of dogs is 78.09% but by using a decision network to incorporate "soft" biometrics, the identification rate can achieve an accuracy of 84.94%.
Kenneth Lai, Xinyuan Tu, Svetlana N. Yanushkevich
IJCNN1
2018 CNN+RNN Depth and Skeleton based Dynamic Hand Gesture Recognition
abstract
Human activity and gesture recognition is an important component of rapidly growing domain of ambient intelligence, in particular in assisting living and smart homes. In this paper, we propose to combine the power of two deep learning techniques, the convolutional neural networks (CNN) and the recurrent neural networks (RNN), for automated hand gesture recognition using both depth and skeleton data. Each of these types of data can be used separately to train neural networks to recognize hand gestures. While RNN were reported previously to perform well in recognition of sequences of movement for each skeleton joint given the skeleton information only, this study aims at utilizing depth data and apply CNN to extract important spatial information from the depth images. Together, the tandem CNN+RNN is capable of recognizing a sequence of gestures more accurately. As well, various types of fusion are studied to combine both the skeleton and depth information in order to extract temporal-spatial information. An overall accuracy of 85.46% is achieved on the dynamic hand gesture-14/28 dataset.
Kenneth Lai, Svetlana N. Yanushkevich
ICPR1
2018 Mass Evidence Accumulation and Traveler Risk Scoring Engine in e-Border Infrastructure
abstract
This paper is concerned with mass evidence accumulation and risk assessment in a particular component of transportation systems and e-borders. We outline the challenges faced by contemporary border control technology and conduct a series of demonstrative experiments that cover critical scenarios, tasks, and states of both evidence accumulation and the traveler risk scoring engine. Using technology gap navigator methodology, this paper suggests an approach to traveler risk estimation based on a unified inference platform, such as a causal graphical model with various incorporated metrics of uncertainty.
Kenneth Lai, Shawn Eastwood, Warren Adam Shier, Svetlana N. Yanushkevich, Vlad P. Shmerko
IEEE Trans. Intell. Transp. Syst.1
2017 Multi-Scale histogram tone mapping algorithm enables better object detection in wide dynamic range images
abstract
In this paper, we present a novel tone mapping algorithm based on multi-scale histograms and fusion (MS-Hist), for displaying wide dynamic range (WDR) images and better detection of objects such as human faces. The proposed algorithm tone maps pixels based on multiple scale local histograms, where small scales are used to preserve local contrast and large scales allow to maintain the global brightness consistency. A database of WDR images of humans depicted in high-contrast light conditions was created to validate and compare the performance of various algorithms for face detection in tasks such as biometric based identification. Our experimental results show that the proposed MS-Hist algorithm preserves image detail, brightness and high local contrast, and can benefit tasks such as face detection in WDR images.
Jie Yang 0033, Alain Horé, Ulian Shahnovich, Kenneth Lai, Svetlana N. Yanushkevich, Orly Yadid-Pecht
AVSS4
2017 Risk assessment in the face-based watchlist screening in e-borders
abstract
This paper concerns with facial-based watch list technology as a component of automated border control machines deployed in e-borders. The key task of the watch list technology is to mitigate effects of mis-identification and impersonation. To address this problem, we developed a novel cost-based model of traveler risk assessment and proved its efficiency via intensive experiments using large-scale facial databases. The results of this study are applicable to any biometric modality to be used in watch list technology.
Kenneth Lai, Svetlana N. Yanushkevich, Vlad P. Shmerko
IJCB1