VLDB 2026 Research / reviewers in the wild / expert
Tom Gedeon
dblp:g/TamasDGedeon · also Tamas O. Gedeon, Tamás D. Gedeon, Tamás Domonkos Gedeon
· DBLP profile ↗
199ranked-venue papers
12as first author
51since 2021 · last 2026
0000-0001-8356-4909ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 137 · 7 first-author · 34 since 2021Human-computer interaction and ubiquitous computing · 31 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data ServerabstractWe explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usually exhibits distinct modes (i.e., semantic clusters representing data distribution), if the training set does not contain these target modes, the model performance would be compromised. While prior existing works improve algorithms iteratively, our research explores the often-overlooked potential of optimizing the structure of the data server. Inspired by the hierarchical nature of web search engines, we introduce a hierarchical data server, together with a bipartite mode matching algorithm (BMM) to align source and target modes. For each target mode, we look in the server data tree for the best mode match, which might be large or small in size. Through bipartite matching, we aim for all target modes to be optimally matched with source modes in a one-on-one fashion. Compared with existing training set search algorithms, we show that the matched server modes constitute training sets that have consistently smaller domain gaps with the target domain across object re-identification (re-ID) and detection tasks. Consequently, models trained on our searched training sets have higher accuracy than those trained otherwise. BMM allows data-centric unsupervised domain adaptation (UDA) orthogonal to existing model-centric UDA methods. By combining the BMM with existing UDA methods like pseudo-labeling, further improvement is observed. Yue Yao 0001, Ruining Yang, Tom Gedeon |
AAAI | 3 |
| 2026 | DermEVAL: A Dermatologist-Reviewed Benchmark for Multimodal Large Language ModelsabstractClinical photographs play a vital role in conversational computer-aided diagnosis, particularly in dermatology. However, existing skin disease benchmarks contain limitations like insufficient dataset size, the sole presence of categorical labels, the lack of expert inspections, and limited diversity in annotations. To address these shortcomings, we introduce DermEVAL, a large-scale benchmark specifically designed to evaluate the performance of Multimodal Large Language Models (MLLMs) in dermatology. Our benchmark includes image-text pairs depicting 16 distinct skin diseases, featuring a total of 11,347 representative images drawn from various dermatological datasets, carefully selected and annotated with the guidance of dermatologists. DermEVAL enables two primary tasks: visual question answering (VQA) and medical report generation (MRG), designed to simulate real-world medical diagnostics. We evaluate the performance of MLLMs in dermatology using multiple metrics, including traditional metrics and GPT-4V-based assessments. Our results indicate that accurately diagnosing skin diseases remains challenging for state-of-the-art MLLMs. We also demonstrate that fine-tuning MLLMs using DermEVAL significantly improves their performance on dermatology-related image–text tasks. Hongjin Zhao, Zhenyue Qin, Ge-Peng Ji, Tom Gedeon, Nick Barnes |
WACV | 6 |
| 2026 | Representation-centric survey of supervised skeletal action recognition and the new benchmarkabstract3D skeletal action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational efficiency, and enhanced privacy. Despite remarkable progress, current research remains fragmented across diverse input representations and lacks evaluation under scenarios that reflect real-world challenges. This paper presents a representation-centric review of supervised skeletal action recognition, systematically categorizing state-of-the-art methods by their input feature types: joint coordinates, bone vectors, motion flows, and extended representations, and analyzing how these choices influence spatiotemporal modeling strategies. Building on the insights from this review, we introduce ANUBIS, a large-scale, challenging dataset designed to address critical gaps in existing benchmarks. ANUBIS incorporates multi-view recordings with back-view perspectives, complex multi-person interactions, fine-grained and violent actions, and contemporary social behaviors. We benchmark a diverse set of state-of-the-art models on ANUBIS and conduct an in-depth analysis of how different feature types affect recognition performance across 102 action categories. Our results show strong action-feature dependencies, highlight the limitations of naïve multi-representational fusion, and point toward the need for task-aware, semantically aligned integration strategies. This work offers both a comprehensive foundation and a practical benchmarking resource, aiming to guide the next generation of robust, generalizable skeleton-based action recognition systems for complex real-world scenarios. The dataset, benchmarking framework, and code are available at https://yliu1082.github.io/ANUBIS/ . Yang Liu 0249, Jiyao Yang, Madhawa Perera, Pan Ji, Dongwoo Kim 0002, Min Xu 0009, Tianyang Wang 0004, Saeed Anwar, Tom Gedeon, Lei Wang 0108, Zhenyue Qin |
Pattern Recognit. | 9 |
| 2026 | Thought graph traversal for test-time scaling in chest X-ray VLLMs
Yue Yao 0001, Zelin Wen, Xuqing Li, Dongliang Xu, Tom Gedeon |
Pattern Recognit. | 8 |
| 2025 | Unsupervised Search for Ethnic Minorities' Medical Segmentation Training SetabstractThis paper investigates the critical issue of dataset bias in medical imaging, with a particular emphasis on racial disparities caused by uneven population distribution in dataset collection. Our analysis reveals that medical segmentation datasets are significantly biased, primarily influenced by the demographic composition of their collection sites. For instance, Scanning Laser Ophthalmoscopy (SLO) fundus datasets collected in the United States predominantly feature images of White individuals, with minority racial groups underrepresented. This imbalance can result in biased model performance and inequitable clinical outcomes, particularly for minority populations. To address this challenge, we propose a novel training set search strategy aimed at reducing these biases by focusing on underrepresented racial groups. Our approach utilizes existing datasets and employs a simple greedy algorithm to identify source images that closely match the target domain distribution. By selecting training data that aligns more closely with the characteristics of minority populations, our strategy improves the accuracy of medical segmentation models on specific minorities, i.e., Black. Our experimental results demonstrate the effectiveness of this approach in mitigating bias. We also discuss the broader societal implications, highlighting how addressing these disparities can contribute to more equitable healthcare outcomes. Our code is available at https://github.com/yorkeyao/SnP. Yue Yao 0001, Ruining Yang, Ashu Gupta, Tom Gedeon |
ICASSP | 6 |
| 2025 | TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention MapsabstractAccurately detecting and tracking high-speed, small objects, such as balls in sports videos, is challenging due to factors like motion blur and occlusion. Although recent deep learning frameworks like TrackNetV1, V2, and V3 have advanced tennis ball and shuttlecock tracking, they often struggle in scenarios with partial occlusion or low visibility. This is primarily because these models rely heavily on visual features without explicitly incorporating motion information, which is crucial for precise tracking and trajectory prediction. In this paper, we introduce an enhancement to the TrackNet family by fusing high-level visual features with learnable motion attention maps through a motion-aware fusion mechanism, effectively emphasizing the moving ball’s location and improving tracking performance. Our approach uses frame differencing maps, modulated by a motion prompt layer, to highlight key motion regions over time. Experimental results on the tennis ball and shuttlecock datasets show that our method enhances the tracking performance of both TrackNetV2 and V3. We refer to our lightweight, plug-and-play solution, built on top of the existing TrackNets, as TrackNetV4. Arjun Raj, Lei Wang 0108, Tom Gedeon |
ICASSP | 3 |
| 2025 | Learnable Expansion of Graph Operators for Multi-Modal Feature FusionabstractIn computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance, especially with the availability of powerful pre-trained models like vision-language models. However, common fusion methods, such as concatenation, element-wise operations, and non-linear techniques, often fail to capture structural relationships, deep feature interactions, and suffer from inefficiency or misalignment of features across domains or modalities. In this paper, we shift from high-dimensional feature space to a lower-dimensional, interpretable graph space by constructing relationship graphs that encode feature relationships at different levels, e.g., clip, frame, patch, token, etc. To capture deeper interactions, we expand graphs through iterative graph relationship updates and introduce a learnable graph fusion operator to integrate these expanded relationships for more effective fusion. Our approach is relationship-centric, operates in a homogeneous space, and is mathematically principled, resembling element-wise relationship score aggregation via multilinear polynomials. We demonstrate the effectiveness of our graph-based fusion method on video anomaly detection, showing strong performance across multi-representational, multi-modal, and multi-domain feature fusion tasks. Dexuan Ding, Lei Wang 0108, Liyun Zhu, Tom Gedeon, Piotr Koniusz |
ICLR | 4 |
| 2025 | When Robots Listen: Predicting Empathy Valence from Multimodal Storytelling Data
Himadri Shekhar Mondal, Tom Gedeon |
ICMI | 3 |
| 2025 | Ranked from Within: Ranking Large Multimodal Models Without LabelsabstractCan the relative performance of a pre-trained large multimodal model (LMM) be predicted without access to labels? As LMMs proliferate, it becomes increasingly important to develop efficient ways to choose between them when faced with new data or tasks. The usual approach does the equivalent of giving the models an exam and marking them. We opt to avoid marking and the associated labor of determining the ground-truth answers. Instead, we explore other signals elicited and ascertain how well the models know their own limits, evaluating the effectiveness of these signals at unsupervised model ranking. We evaluate 47 state-of-the-art LMMs (e.g., LLaVA) across 9 visual question answering benchmarks, analyzing how well uncertainty-based metrics can predict relative model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust and consistent basis for ranking models across various tasks. This facilitates the ranking of LMMs on unlabeled data, providing a practical approach for selecting models for diverse target domains without requiring manual annotation. Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao 0005, Jiyang Zheng, Tom Gedeon, Tongliang Liu |
ICML | 6 |
| 2025 | MRAC 2025: 3rd International Workshop on Multimodal, Generative and Responsible Affective ComputingabstractMultimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing. Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Göcke, Jianhua Tao 0001, Tom Gedeon |
ACM Multimedia | 10 |
| 2025 | AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World PerturbationsabstractThe rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025. Zhixi Cai, Kartik Kuckreja, Shreya Ghosh 0001, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall |
ACM Multimedia | 7 |
| 2025 | Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
Shreya Ghosh 0001, Tom Gedeon, Thanh-Toan Do, Abhinav Dhall |
ACM Multimedia | 3 |
| 2025 | MIP-GAF: A MLLM-Annotated Benchmark for Most Important Person Localization and Group Context UnderstandingabstractEstimating the Most Important Person (MIP) in any social event setup is a challenging problem mainly due to contextual complexity and scarcity of labeled data. Moreover, the causality aspects of MIP estimation are quite subjective and diverse. To this end, we aim to address the problem by annotating a large-scale ‘in-the-wild’ dataset for iden-tifying human perceptions about the ‘Most Important Person (MIP)‘ in an image. The paper provides a thorough description of our proposed Multimodal Large Language Model (MLLM) based data annotation strategy, and a thor-ough data quality analysis. Further, we perform a comprehensive benchmarking of the proposed dataset utilizing state-of-the-art MIP localization methods, indicating a significant drop in performance compared to existing datasets. The performance drop shows that the existing MIP localization algorithms must be more robust with respect to ‘in-the-wild’ situations. We believe the proposed dataset will play a vital role in building the next-generation social situation understanding methods. The dataset and associated code will be made available for research purposes. Surbhi Madan, Shreya Ghosh 0001, Lownish Rai Sookha, M. A. Ganaie 0001, Subramanian Ramanathan, Abhinav Dhall, Tom Gedeon |
WACV | 7 |
| 2025 | Toward a Holistic Evaluation of Robustness in CLIP ModelsabstractContrastive Language-Image Pre-training (CLIP) models have shown significant potential, particularly in zero-shot classification across diverse distribution shifts. Building on existing evaluations of overall classification robustness, this work aims to provide a more comprehensive assessment of CLIP by introducing several new perspectives. First, we investigate their robustness to variations in specific visual factors. Second, we assess two critical safety objectives-confidence uncertainty and out-of-distribution detection-beyond mere classification accuracy. Third, we evaluate the finesse with which CLIP models bridge the image and text modalities. Fourth, we extend our examination to 3D awareness in CLIP models, moving beyond traditional 2D image understanding. Finally, we explore the interaction between vision and language encoders within modern large multimodal models (LMMs) that utilize CLIP as the visual backbone, focusing on how this interaction impacts classification robustness. In each aspect, we consider the impact of six factors on CLIP models: model architecture, training distribution, training set size, fine-tuning, contrastive loss, and test-time prompts. Our study uncovers several previously unknown insights into CLIP. For instance, the architecture of the visual encoder in CLIP plays a significant role in their robustness against 3D corruption. CLIP models tend to exhibit a bias towards shape when making predictions. Moreover, this bias tends to diminish after fine-tuning on ImageNet. Vision-language models like LLaVA, leveraging the CLIP vision encoder, could exhibit benefits in classification performance for challenging categories over CLIP alone. Our findings are poised to offer valuable guidance for enhancing the robustness and reliability of CLIP models. Weijie Tu, Weijian Deng, Tom Gedeon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Empathy Detection From Text, Audiovisual, Audio or Physiological Signals: A Systematic Review of Task Formulations and Machine Learning MethodsabstractEmpathy indicates an individual's ability to understand others. Over the past few years, empathy has drawn attention from various disciplines, including but not limited to Affective Computing, Cognitive Science, and Psychology. Detecting empathy has potential applications in society, healthcare and education. Despite being a broad and overlapping topic, the avenue of empathy detection leveraging Machine Learning remains underexplored from a systematic literature review perspective. We collected 849 papers from 10 well-known academic databases, systematically screened them and analysed the final 82 papers. Our analyses reveal several prominent task formulations – including empathy on localised utterances or overall expressions, unidirectional or parallel empathy, and emotional contagion – in monadic, dyadic and group interactions. Empathy detection methods are summarised based on four input modalities – text, audiovisual, audio and physiological signals – thereby presenting modality-specific network architecture design protocols. We discuss challenges, research gaps and potential applications in the Affective Computing-basedempathydomain, which can facilitate new avenues of exploration. We further enlist the public availability of datasets and codes. This paper, therefore, provides a structured overview of recent advancements and remaining challenges towards developing a robust empathy detection system that could meaningfully contribute to enhancing human well-being. Md. Rakibul Hasan 0001, Shreya Ghosh 0001, Aneesh Krishna, Tom Gedeon |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Visual and Textual Prompts in VLLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others’ emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs’ video emotion recognition capabilities. Zhifeng Wang 0004, Qixuan Zhang, Peter Zhang, Wenjia Niu, Kaihao Zhang, Ramesh S. Sankaranarayana, Sabrina B. Caldwell, Tom Gedeon |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Position-Sensing Graph Neural Networks: Proactively Learning Nodes Relative PositionsabstractMost existing graph neural networks (GNNs) learn node embeddings using the framework of message passing and aggregation. Such GNNs are incapable of learning relative positions between graph nodes within a graph. To empower GNNs with the awareness of node positions, some nodes are set as anchors. Then, using the distances from a node to the anchors, GNNs can infer relative positions between nodes. However, position-aware GNNs (P-GNNs) arbitrarily select anchors, leading to compromising position awareness and feature extraction. To eliminate this compromise, we demonstrate that selecting evenly distributed and asymmetric anchors is essential. On the other hand, we show that choosing anchors that can aggregate embeddings of all the nodes within a graph is NP-complete. Therefore, devising efficient optimal algorithms in a deterministic approach is practically not feasible. To ensure position awareness and bypass NP-completeness, we propose position-sensing GNNs (PSGNNs), learning how to choose anchors in a backpropagatable fashion. Experiments verify the effectiveness of PSGNNs against state-of-the-art GNNs, substantially improving performance on various synthetic and real-world graph datasets while enjoying stable scalability. Specifically, PSGNNs on average boost area under the curve (AUC) more than 14% for pairwise node classification and 18% for link prediction over the existing state-of-the-art position-aware methods. Our source code is publicly available at: https://github.com/ZhenyueQin/PSGNN. Zhenyue Qin, Saeed Anwar, Dongwoo Kim 0002, Yang Liu 0249, Pan Ji, Tom Gedeon |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Motion meets Attention: Video Motion Prompts
Qixiang Chen, Lei Wang 0108, Piotr Koniusz, Tom Gedeon |
ACML | 4 |
| 2024 | Adaptive Multi-head Contrastive Learning
Lei Wang 0108, Piotr Koniusz, Tom Gedeon, Liang Zheng 0001 |
ECCV (69) | 3 |
| 2024 | Visual Prompting in LLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing.Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped.Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context.To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely.SoV improves accuracy in face count and emotion categorization while preserving the enriched image context.Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments.Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. Question:How many visible faces are there in the image?Analyze each face's emotion and group the face's emotions. Answer:In the image, there are 22 visible faces.Emotions can be grouped into three main categories. Neutral Emotion: Approximately 16Faces.This neutral expression is typical in settings requiring attention to a speaker or event, reflecting an engaged yet reserved emotional state. Mildly Positive Emotion: Approximately 5Faces.This could be a response to something interesting or mildly amusing happening at the event, but their reactions are subdued, in keeping with the formal setting.Happy: 1 Face.This expression stands out as more joyful or enthusiastic about the proceedings. Qixuan Zhang, Zhifeng Wang 0004, Dylan Zhang, Wenjia Niu, Sabrina B. Caldwell, Tom Gedeon, Yang Liu 0003, Zhenyue Qin |
EMNLP | 6 |
| 2024 | TCA-NET: Triplet Concatenated-Attentional Network for Multimodal Engagement EstimationabstractHuman social interactions involve intricate social signals that artificial intelligence and machine learning models aim to decipher, particularly in the context of artificial mediators that can enhance human interactions across domains like education and healthcare. Engagement, a key aspect of these interactions, relies heavily on multimodal information like facial expressions, voice and posture. Recently, many deep learning methods have been deployed in engagement estimation. Still, they often focus on unimodality or bimodality, leading to the results lacking robustness and adaptability due to factors like noise and varying individual responses. To address this challenge, we introduce a novel modality fusion framework named Triplet Concatenated-Attentional Net (TCA-Net). This framework takes three distinct types of data modality (video, audio and Kinect) as inputs and delivers a prediction score as output. Within this network, a specially designed concatenated-attention fusion mechanism serves the purpose of modality fusion and preserves the intra-modal features. Experimental results validate the efficiency of our TCA-Net in enhancing the accuracy and reliability of engagement estimation across diverse scenarios, with a test set Concordance Correlation Coefficient (CCC) of 0.75. We release our code at https://github.com/Daming-W/Multimodal_Engagement_Estimation. Hongyuan He, Md. Rakibul Hasan 0001, Tom Gedeon |
ICIP | 4 |
| 2024 | Taylor Videos for Action RecognitionabstractEffectively extracting motions from video is a critical and long-standing problem for action recognition. This problem is very challenging because motions (i) do not have an explicit form, (ii) have various concepts such as displacement, velocity, and acceleration, and (iii) often contain noise caused by unstable pixels. Addressing these challenges, we propose the Taylor video, a new video format that highlights the dominate motions (e.g., a waving hand) in each of its frames named the Taylor frame. Taylor video is named after Taylor series, which approximates a function at a given point using important terms. In the scenario of videos, we define an implicit motion-extraction function which aims to extract motions from video temporal block. In this block, using the frames, the difference frames, and higher-order difference frames, we perform Taylor expansion to approximate this function at the starting frame. We show the summation of the higher-order terms in the Taylor series gives us dominant motion patterns, where static objects, small and unstable motions are removed. Experimentally we show that Taylor videos are effective inputs to popular architectures including 2D CNNs, 3D CNNs, and transformers. When used individually, Taylor videos yield competitive action recognition accuracy compared to RGB videos and optical flow. When fused with RGB or optical flow videos, further accuracy improvement is achieved. Additionally, we apply Taylor video computation to human skeleton sequences, resulting in Taylor skeleton sequences that outperform the use of original skeletons for skeleton-based action recognition. Lei Wang 0108, Xiuyuan Yuan, Tom Gedeon, Liang Zheng 0001 |
ICML | 3 |
| 2024 | An Empirical Study Into What Matters for Calibrating Vision-Language ModelsabstractVision-Language Models (VLMs) have emerged as the dominant approach for zero-shot recognition, adept at handling diverse scenarios and significant distribution changes. However, their deployment in risk-sensitive areas requires a deeper understanding of their uncertainty estimation capabilities, a relatively uncharted area. In this study, we explore the calibration properties of VLMs across different architectures, datasets, and training strategies. In particular, we analyze the uncertainty estimation performance of VLMs when calibrated in one domain, label set or hierarchy level, and tested in a different one. Our findings reveal that while VLMs are not inherently calibrated for uncertainty, temperature scaling significantly and consistently improves calibration, even across shifts in distribution and changes in label set. Moreover, VLMs can be calibrated with a very small set of examples. Through detailed experimentation, we highlight the potential applications and importance of our insights, aiming for more reliable and effective use of VLMs in critical, real-world scenarios. Weijie Tu, Weijian Deng, Dylan Campbell, Stephen Gould, Tom Gedeon |
ICML | 5 |
| 2024 | AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetabstractThe detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small segments of audio-visual manipulations embedded in real videos. In this research, we emulate the process of such content generation and propose the AV-Deepfake1M dataset. The dataset contains content-driven (i) video manipulations, (ii) audio manipulations, and (iii) audio-visual manipulations for more than 2K subjects resulting in a total of more than 1M videos. The paper provides a thorough description of the proposed data generation pipeline accompanied by a rigorous analysis of the quality of the generated data. The comprehensive benchmark of the proposed dataset utilizing state-of-the-art deepfake detection and localization methods indicates a significant drop in performance compared to previous datasets. The proposed dataset will play a vital role in building the next-generation deepfake localization methods. The dataset and associated code are available at https://github.com/ControlNet/AV-Deepfake1M. Zhixi Cai, Shreya Ghosh 0001, Aman Pankaj Adatia, Munawar Hayat, Abhinav Dhall, Tom Gedeon, Kalin Stefanov |
ACM Multimedia | 6 |
| 2024 | C3-PO: A Convolutional Neural Network for COVID Onset Prediction from Cough Sounds
Md Ayshik Rahman Khan, Md. Rakibul Hasan 0001, Tom Gedeon |
MMM (3) | 4 |
| 2024 | Advancing Video Anomaly Detection: A Concise Review and a New DatasetabstractVideo Anomaly Detection (VAD) finds widespread applications in security surveillance, traffic monitoring, industrial monitoring, and healthcare. Despite extensive research efforts, there remains a lack of concise reviews that provide insightful guidance for researchers. Such reviews would serve as quick references to grasp current challenges, research trends, and future directions. In this paper, we present such a review, examining models and datasets from various perspectives. We emphasize the critical relationship between model and dataset, where the quality and diversity of datasets profoundly influence model performance, and dataset development adapts to the evolving needs of emerging approaches. Our review identifies practical issues, including the absence of comprehensive datasets with diverse scenarios. To address this, we introduce a new dataset, Multi-Scenario Anomaly Detection (MSAD), comprising 14 distinct scenarios captured from various camera views. Our dataset has diverse motion patterns and challenging variations, such as different lighting and weather conditions, providing a robust foundation for training superior models. We conduct an in-depth analysis of recent representative models using MSAD and highlight its potential in addressing the challenges of detecting anomalies across diverse and evolving surveillance scenarios. Liyun Zhu, Lei Wang 0108, Arjun Raj, Tom Gedeon, Chen Chen 0001 |
NeurIPS | 4 |
| 2024 | Meet JEANIE: A Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint AlignmentabstractAbstract Video sequences exhibit significant nuisance variations (undesired effects) of speed of actions, temporal locations, and subjects’ poses, leading to temporal-viewpoint misalignment when comparing two sets of frames or evaluating the similarity of two sequences. Thus, we propose Joint tEmporal and cAmera viewpoiNt alIgnmEnt (JEANIE) for sequence pairs. In particular, we focus on 3D skeleton sequences whose camera and subjects’ poses can be easily manipulated in 3D. We evaluate JEANIE on skeletal Few-shot Action Recognition (FSAR), where matching well temporal blocks (temporal chunks that make up a sequence) of support-query sequence pairs (by factoring out nuisance variations) is essential due to limited samples of novel classes. Given a query sequence, we create its several views by simulating several camera locations. For a support sequence, we match it with view-simulated query sequences, as in the popular Dynamic Time Warping (DTW). Specifically, each support temporal block can be matched to the query temporal block with the same or adjacent (next) temporal index, and adjacent camera views to achieve joint local temporal-viewpoint warping. JEANIE selects the smallest distance among matching paths with different temporal-viewpoint warping patterns, an advantage over DTW which only performs temporal alignment. We also propose an unsupervised FSAR akin to clustering of sequences with JEANIE as a distance measure. JEANIE achieves state-of-the-art results on NTU-60, NTU-120, Kinetics-skeleton and UWA3D Multiview Activity II on supervised and unsupervised FSAR, and their meta-learning inspired fusion. Lei Wang 0108, Jun Liu 0036, Liang Zheng 0001, Tom Gedeon, Piotr Koniusz |
Int. J. Comput. Vis. | 4 |
| 2024 | Attribute Descent: Simulating Object-Centric Datasets on the Content Level and BeyondabstractThis article aims to use graphic engines to simulate a large number of training data that have free annotations and possibly strongly resemble to real-world data. Between synthetic and real, a two-level domain gap exists, involving content level and appearance level. While the latter is concerned with appearance style, the former problem arises from a different mechanism, i.e., content mismatch in attributes such as camera viewpoint, object placement and lighting conditions. In contrast to the widely-studied appearance-level gap, the content-level discrepancy has not been broadly studied. To address the content-level misalignment, we propose an attribute descent approach that automatically optimizes engine attributes to enable synthetic data to approximate real-world data. We verify our method on object-centric tasks, wherein an object takes up a major portion of an image. In these tasks, the search space is relatively small, and the optimization of each attribute yields sufficiently obvious supervision signals. We collect a new synthetic asset VehicleX, and reformat and reuse existing the synthetic assets ObjectX and PersonX. Extensive experiments on image classification and object re-identification confirm that adapted synthetic data can be effectively used in three scenarios: training with synthetic data only, training data augmentation and numerically understanding dataset content. Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Napthade, Tom Gedeon |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Fusing Higher-Order Features in Graph Neural Networks for Skeleton-Based Action RecognitionabstractSkeleton sequences are lightweight and compact and thus are ideal candidates for action recognition on edge devices. Recent skeleton-based action recognition methods extract features from 3-D joint coordinates as spatial-temporal cues, using these representations in a graph neural network for feature fusion to boost recognition performance. The use of first- and second-order features, that is, joint and bone representations, has led to high accuracy. Nonetheless, many models are still confused by actions that have similar motion trajectories. To address these issues, we propose fusing higher-order features in the form of angular encoding (AGE) into modern architectures to robustly capture the relationships between joints and body parts. This simple fusion with popular spatial-temporal graph neural networks achieves new state-of-the-art accuracy in two large benchmarks, including NTU60 and NTU120, while employing fewer parameters and reduced run time. Our source code is publicly available at: https://github.com/ZhenyueQin/Angular-Skeleton-Encoding. Zhenyue Qin, Yang Liu 0249, Pan Ji, Dongwoo Kim 0002, Lei Wang 0108, Robert I. McKay, Saeed Anwar, Tom Gedeon |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | A Bag-of-Prototypes Representation for Dataset-Level ApplicationsabstractThis work investigates dataset vectorization for two dataset-level tasks: assessing training set suitability and test set difficulty. The former measures how suitable a training set is for a target domain, while the latter studies how challenging a test set is for a learned model. Central to the two tasks is measuring the underlying relationship between datasets. This needs a desirable dataset vectorization scheme, which should preserve as much discriminative dataset information as possible so that the distance between the resulting dataset vectors can reflect dataset-to-dataset similarity. To this end, we propose a bag-of-prototypes (BoP) dataset representation that extends the image-level bag consisting of patch descriptors to dataset-level bag consisting of semantic prototypes. Specifically, we develop a codebook consisting of K prototypes clustered from a ref-erence dataset. Given a dataset to be encoded, we quan-tize each of its image features to a certain prototype in the codebook and obtain a K -dimensional histogram. Without assuming access to dataset labels, the BoP representation provides rich characterization of the dataset semantic distribution. Furthermore, BoP representations cooperate well with Jensen-Shannon divergence for measuring dataset-to-dataset similarity. Although very simple, BoP consistently shows its advantage over existing representations on a series of benchmarks for two dataset-level tasks. Weijie Tu, Weijian Deng, Tom Gedeon, Liang Zheng 0001 |
CVPR | 3 |
| 2023 | Large-scale Training Data Search for Object Re-identificationabstractWe consider a scenario where we have access to the target domain, but cannot afford on-the-fly training data annotation, and instead would like to construct an alternative training set from a large-scale data pool such that a competitive model can be obtained. We propose a search and pruning (SnP) solution to this training data search problem, tailored to object re-identification (re-ID), an application aiming to match the same object captured by different cameras. Specifically, the search stage identifies and merges clusters of source identities which exhibit similar distributions with the target domain. The second stage, subject to a budget, then selects identities and their images from the Stage I output, to control the size of the resulting training set for efficient training. The two steps provide us with training sets 80% smaller than the source pool while achieving a similar or even higher re-ID accuracy. These training sets are also shown to be superior to a few existing search methods such as random sampling and greedy sampling under the same budget on training data size. If we release the budget, training sets resulting from the first stage alone allow even higher re-ID accuracy. We provide interesting discussions on the specificity of our method to the re-ID problem and particularly its role in bridging the re-ID domain gap. The code is available at https://github.com/yorkeyao/SnP Yue Yao 0001, Tom Gedeon, Liang Zheng 0001 |
CVPR | 2 |
| 2023 | A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)abstractContrastive Language-Image Pre-training (CLIP) models have demonstrated remarkable generalization capabilities across multiple challenging distribution shifts. However, there is still much to be explored in terms of their robustness to the variations of specific visual factors. In real-world applications, reliable and safe systems must consider other safety measures beyond classification accuracy, such as predictive uncertainty. Yet, the effectiveness of CLIP models on such safety-related objectives is less-explored. Driven by the above, this work comprehensively investigates the safety measures of CLIP models, specifically focusing on three key properties: resilience to visual factor variations, calibrated uncertainty estimations, and the ability to detect anomalous inputs. To this end, we study $83$ CLIP models and $127$ ImageNet classifiers. They are diverse in architecture (pre)training distribution and training strategies. We consider $10$ visual factors (\emph{e.g.}, shape and pattern), $5$ types of out-of-distribution data, and $8$ natural and challenging test conditions with different shift types, such as texture, style, and perturbation shifts. Our study has unveiled several previously unknown insights into CLIP models. For instance, they are not consistently more calibrated than other ImageNet models, which contradicts existing findings. Additionally, our analysis underscores the significance of training source design by showcasing its profound influence on the three key properties. We believe our comprehensive study can shed light on and help guide the development of more robust and reliable CLIP models. Weijie Tu, Weijian Deng, Tom Gedeon |
NeurIPS | 3 |
| 2023 | Personality Perception Using Scenario Based Stimulation and Physiological SignalsabstractPrevious studies on automatic personality perception have primarily focused on a limited number of personality traits. However, in real-world situations, humans exhibit a wide range of personality traits. To overcome this limitation, a new methodology for automatic personality perception is proposed in this paper. This revised approach can predict various personality traits (17 traits) with satisfactory performance by utilizing physiological signals. The underlying concept is to stimulate participants with different emotional stimuli to elicit physiological responses in a specific scenario. Biomarkers such as Electroencephalogram (EEG), Skin Conductance, Blood Volume Pulse, and Pupil Dilation reflect an individual's personality traits. Two experiments are conducted with different scenarios, including Image/Video Stimulation and Driving Simulation, to support this study. Based on the collection of data and validation of supervised learning models, Naive Bayes outperforms other classifiers explored in this research. EEG is the most effective signal for predicting personality, although combining other signals may produce similar results. Our method accurately predicts the 17 personality traits, demonstrating significant potential for clinical research. Amrijit Biswas, Fahimul Hoque Shubho, Tom Gedeon, Shafin Rahman |
SMC | 4 |
| 2023 | Glitch in the matrix: A large scale benchmark for content driven audio-visual forgery detection and localizationabstractMost deepfake detection methods focus on detecting spatial and/or spatio-temporal changes in facial attributes and are centered around the binary classification task of detecting whether a video is real or fake. This is because available benchmark datasets contain mostly visual-only modifications present in the entirety of the video. However, a sophisticated deepfake may include small segments of audio or audio-visual manipulations that can completely change the meaning of the video content. To addresses this gap, we propose and benchmark a new dataset, Localized Audio Visual DeepFake (LAV-DF), consisting of strategic content-driven audio, visual and audio-visual manipulations. The proposed baseline method, Boundary Aware Temporal Forgery Detection (BA-TFD), is a 3D Convolutional Neural Network-based architecture which effectively captures multimodal manipulations. We further improve (i.e. BA-TFD+) the baseline method by replacing the backbone with a Multiscale Vision Transformer and guide the training process with contrastive, frame classification, boundary matching and multimodal boundary matching loss functions. The quantitative analysis demonstrates the superiority of BA-TFD+ on temporal forgery localization and deepfake detection tasks using several benchmark datasets including our newly proposed dataset. The dataset, models and code are available at https://github.com/ControlNet/LAV-DF. Zhixi Cai, Shreya Ghosh 0001, Abhinav Dhall, Tom Gedeon, Kalin Stefanov, Munawar Hayat |
Comput. Vis. Image Underst. | 4 |
| 2023 | Interpretation of Depression Detection Models via Feature Selection MethodsabstractGiven the prevalence of depression worldwide and its major impact on society, several studies employed artificial intelligence modelling to automatically detect and assess depression. However, interpretation of these models and cues are rarely discussed in detail in the AI community, but have received increased attention lately. In this study, we aim to analyse the commonly selected features using a proposed framework of several feature selection methods and their effect on the classification results, which will provide an interpretation of the depression detection model. The developed framework aggregates and selects the most promising features for modelling depression detection from 38 feature selection algorithms of different categories. Using three real-world depression datasets, 902 behavioural cues were extracted from speech behaviour, speech prosody, eye movement and head pose. To verify the generalisability of the proposed framework, we applied the entire process to depression datasets individually and when combined. The results from the proposed framework showed that speech behaviour features (e.g. pauses) are the most distinctive features of the depression detection model. From the speech prosody modality, the strongest feature groups were F0, HNR, formants, and MFCC, while for the eye activity modality they were left-right eye movement and gaze direction, and for the head modality it was yaw head movement. Modelling depression detection using the selected features (even though there are only 9 features) outperformed using all features in all the individual and combined datasets. Our feature selection framework did not only provide an interpretation of the model, but was also able to produce a higher accuracy of depression detection with a small number of features in varied datasets. This could help to reduce the processing time needed to extract features and creating the model. Sharifa Alghowinem, Tom Gedeon, Roland Göcke, Jeffrey F. Cohn, Gordon Parker |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Search Interfaces for Biomedical Searching: How do Gaze, User Perception, Search Behaviour and Search Performance Relate?abstractThe objective of this controlled information retrieval (IR) user experiment is to gain an understanding of domain experts’ interactions with novel search interfaces within the context of biomedical information search, with a goal of better search interface design. In this paper, we examine the relationships among user perception, gaze and search behaviour and user search performance. An eye-tracking study of biomedical domain experts’ interactions with novel search interfaces was conducted. A total of thirty-two users participated and searched for documents answering eight complex exploratory search tasks, using four different search interfaces. The findings suggest that gaze behaviour in terms of fixation durations based measures of areas of interest (AOI), i.e., visual attention to the elements of title, author, abstract and MeSH (Medical Subject Headings) terms in document surrogates is correlated with search performance. Users are more likely to achieve better search performance by precision-based measures when 1) search tasks are perceived as difficult; 2) users attend to the element of abstract; and 3) users can recall using the per-query suggestions during the search processes. More importantly, our findings suggest that a user search interface design that displays contextual information between the suggested keywords and the document may better support users reformulating their queries for complex search tasks in the biomedical domain. We discuss implications for the design of search user interfaces for biomedical searching. Ying-Hsang Liu, Paul Thomas 0001, Tom Gedeon, Nicolay Rusnachenko |
CHIIR | 3 |
| 2022 | How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset?
Yuchi Liu, Zhongdao Wang, Tom Gedeon, Liang Zheng 0001 |
ECCV (8) | 3 |
| 2022 | S2FGAN: Semantically Aware Interactive Sketch-to-Face TranslationabstractInteractive facial image manipulation attempts to edit single and multiple face attributes using a photo-realistic face and/or semantic mask as input. In the absence of the photo-realistic image (only sketch/mask available), previous methods only retrieve the original face but ignore the potential of aiding model controllability and diversity in the translation process. This paper proposes a sketch-to-image generation framework called S2FGAN, aiming to improve users’ ability to interpret and flexibility of face attribute editing from a simple sketch. First, to restore a vivid face from a sketch, we propose semantic level perceptual loss to increase the translation quality. Second, we dedicate the theoretic analysis of attribute editing and build attribute mapping networks with latent semantic loss to modify latent space semantics of Generative Adversarial Networks (GANs). The users can command the model to retouch the generated images by involving the semantic information in the generation process. In this way, our method can manipulate single or multiple face attributes by only specifying attributes to be changed. Extensive experimental results on the CelebAMask-HQ dataset empirically show our superior performance and effectiveness on this task. Our method successfully outperforms state-of-the-art sketch-to-image generation and attribute manipulation methods by exploiting greater control of attribute intensity. Yan Yang 0011, Tom Gedeon, Shafin Rahman |
WACV | 3 |
| 2022 | Resolving Anomalies in the Behaviour of a Modularity-Inducing Problem Domain with Distributional Fitness EvaluationabstractDiscrete gene regulatory networks (GRNs) play a vital role in the study of robustness and modularity. A common method of evaluating the robustness of GRNs is to measure their ability to regulate a set of perturbed gene activation patterns back to their unperturbed forms. Usually, perturbations are obtained by collecting random samples produced by a predefined distribution of gene activation patterns. This sampling method introduces stochasticity, in turn inducing dynamicity. This dynamicity is imposed on top of an already complex fitness landscape. So where sampling is used, it is important to understand which effects arise from the structure of the fitness landscape, and which arise from the dynamicity imposed on it. Stochasticity of the fitness function also causes difficulties in reproducibility and in post-experimental analyses. We develop a deterministic distributional fitness evaluation by considering the complete distribution of gene activity patterns, so as to avoid stochasticity in fitness assessment. This fitness evaluation facilitates repeatability. Its determinism permits us to ascertain theoretical bounds on the fitness, and thus to identify whether the algorithm has reached a global optimum. It enables us to differentiate the effects of the problem domain from those of the noisy fitness evaluation, and thus to resolve two remaining anomalies in the behaviour of the problem domain of Espinosa-Soto and A. Wagner (2010). We also reveal some properties of solution GRNs that lead them to be robust and modular, leading to a deeper understanding of the nature of the problem domain. We conclude by discussing potential directions toward simulating and understanding the emergence of modularity in larger, more complex domains, which is key both to generating more useful modular solutions, and to understanding the ubiquity of modularity in biological systems. Zhenyue Qin, Tom Gedeon, Robert I. McKay |
Artif. Life | 2 |
| 2022 | Automatic Prediction of Group Cohesiveness in ImagesabstractThis article discusses the prediction of cohesiveness of a group of people in images. The cohesiveness of a group is an essential indicator of the emotional state, structure, and success of the group. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. To identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. A capsule network is explored for analyzing the contribution of facial expressions of the group members on predicting the Group Cohesion Score (GCS). We add GCS to the Group Affect database and propose the ‘GAF-Cohesion database’. The proposed model performs well on the database and achieves near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated. Further, we investigate the effect of face-level similarity, body pose and subset of a group on the task of automatic cohesion perception. Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | An Agile New Research Framework for Hybrid Human-AI Teaming: Trust, Transparency, and TransferabilityabstractWe propose a new research framework by which the nascent discipline of human-AI teaming can be explored within experimental environments in preparation for transferal to real-world contexts. We examine the existing literature and unanswered research questions through the lens of an Agile approach to construct our proposed framework. Our framework aims to provide a structure for understanding the macro features of this research landscape, supporting holistic research into the acceptability of human-AI teaming to human team members and the affordances of AI team members. The framework has the potential to enhance decision-making and performance of hybrid human-AI teams. Further, our framework proposes the application of Agile methodology for research management and knowledge discovery. We propose a transferability pathway for hybrid teaming to be initially tested in a safe environment, such as a real-time strategy video game, with elements of lessons learned that can be transferred to real-world situations. Sabrina B. Caldwell, Penny Kyburz, Nicholas O'Donnell, Matthew James Knight, Matthew Aitchison, Tom Gedeon, Daniel Johnson 0001, Margot Brereton, Marcus Gallagher, David Conroy |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2021 | Invertible Denoising Network: A Light Solution for Real Noise RemovalabstractInvertible networks have various benefits for image de-noising since they are lightweight, information-lossless, and memory-saving during back-propagation. However, applying invertible models to remove noise is challenging because the input is noisy, and the reversed output is clean, following two different distributions. We propose an invertible denoising network, InvDN, to address this challenge. InvDN transforms the noisy input into a low-resolution clean image and a latent representation containing noise. To discard noise and restore the clean image, InvDN replaces the noisy latent representation with another one sampled from a prior distribution during reversion. The de-noising performance of InvDN is better than all the existing competitive models, achieving a new state-of-the-art result for the SIDD dataset while enjoying less run time. Moreover, the size of InvDN is far smaller, only having 4.2% of the number of parameters compared to the most recently proposed DANet. Further, via manipulating the noisy latent representation, InvDN is also able to generate noise more similar to the original one. Our code is available at: https://github.com/Yang-Liu1082/InvDN.git. Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Pan Ji, Dongwoo Kim 0002, Sabrina B. Caldwell, Tom Gedeon |
CVPR | 7 |
| 2021 | Skeletons on the Stairs: Are They Deceptive?
Yang Liu 0249, Zhenyue Qin, Xuanying Zhu, Sabrina B. Caldwell, Tom Gedeon |
ICONIP (6) | 6 |
| 2021 | Rethinking Binary Hyperparameters for Deep Transfer Learning
Jo Plested, Xuyang Shen, Tom Gedeon |
ICONIP (2) | 3 |
| 2021 | Examining Transfer Learning with Neural Network and Bidirectional Neural Network on Thermal Imaging for Deception Recognition
Zishan Qin, Xuanying Zhu, Tom Gedeon |
ICONIP (6) | 3 |
| 2021 | A Lightweight Multi-scale Feature Fusion Network for Real-Time Semantic Segmentation
Tanmay Singha, Duc-Son Pham 0001, Aneesh Krishna, Tom Gedeon |
ICONIP (2) | 4 |
| 2021 | EEG Feature Significance Analysis
Yue Yao 0001, Shafin Rahman, Tom Gedeon |
ICONIP (6) | 5 |
| 2021 | Exploring Biases and Prejudice of Facial Synthesis via Semantic Latent SpaceabstractDeep learning (DL) models are widely used to provide a more convenient and smarter life. However, biased algorithms will negatively influence us. For instance, groups targeted by biased algorithms will feel unfairly treated and even fearful of negative consequences of these biases. This work targets biased generative models' behaviors, identifying the cause of the biases and eliminating them. We can (as expected) conclude that biased data causes biased predictions of face frontalization models. Varying the proportions of male and female faces in the training data can have a substantial effect on behavior on the test data: we found that the seemingly obvious choice of 50:50 proportions was not the best for this dataset to reduce biased behavior on female faces, which was 71% unbiased as compared to our top unbiased rate of 84%. Failure in generation and generating incorrect gender faces are two behaviors of these models. In addition, only some layers in face frontalization models are vulnerable to biased datasets. Optimizing the skip-connections of the generator in face frontalization models can make models less biased. We conclude that it is likely to be impossible to eliminate all training bias without an unlimited size dataset, and our experiments show that the bias can be reduced and quantified. We believe the next best to a perfect unbiased predictor is one that has minimized the remaining known bias. Xuyang Shen, Jo Plested, Sabrina B. Caldwell, Tom Gedeon |
IJCNN | 4 |
| 2021 | Detecting Lies: Finding the Degree of Falsehood from Observers' Physiological ResponsesabstractLying is a common act in daily life and may have various degrees of falsehood. Deception detection has always been a fascinating area of research in which many studies have been conducted using subjects’ facial, verbal or bodily cues to spot potential deceit. However, none of the studies have investigated the physiological responses of observers in response to misleading statements with various degrees of falsehood. In this paper, we investigated this problem by first conducting designed experiments to collect participants’ physiological signals while they were watching stimulus videos with various falsehood levels. Then, the data was analysed using machine learning or deep learning models. Various challenges including relatively small amounts of training data and imbalanced classes have been addressed by implementing data augmentation. The results show that deep learning models, such as ResNet and VAE-LSTM, can predict the degree of falsehood with an F1-measure up to 0.83 from observers’ reactions when compared to the stimuli ground truth. This was attained when the model was trained with the most useful physiological signal in this study, Electrodermal Activity (EDA). This result indicates that observers’ physiological signals can be used as an indicator to determine the degree of falsehood for misleading statements. In the future, this system may be applied to provide an objective evaluation for deception detection. Ruimin Chu, Jessica Sharmin Rahman, Sabrina B. Caldwell, Xuanying Zhu, Tom Gedeon |
SMC | 5 |
| 2021 | Towards Self-Guided Remote User Studies - Feasibility of Gesture Elicitation using Immersive Virtual RealityabstractGesture Elicitation Studies (GES) are a widely used empirical method to develop gesture vocabularies, interaction models and methods for gesture-based systems in different contexts. While GES show great promise to identify user-defined gestures, there are inherent problems with current methods used for GES. Especially during the ongoing pandemic, it has been nearly impossible to conduct in-person, in-lab GES, while ensuring the safety and well-being of the participants, and complying with social distancing regulations. Further, with prevailing experiment designs, increasing the number of participants is time consuming, while in-lab environments also limit ecological validity. This study explores an intuitive way of conducting self-guided GES using immersive Virtual Reality (VR), utilizing its capability to simulate various contexts to enhance ecological validity. We present a methodology and a tool set that use an immersive VR environment to conduct ecologically valid GES (as a use case) while requiring minimal involvement by the investigator. We evaluate our method using the case of a smart home environment and measure participant acceptance and discuss opportunities and challenges involved in this method. We believe that this study will help HCI research to move forward with participatory design research, even when lab experiments are difficult to conduct. Madhawa Perera, Tom Gedeon, Matt Adcock, Armin Haller |
SMC | 2 |
| 2021 | Predicting Visual Search Task Success from Eye Gaze Data as a Basis for User-Adaptive Information Visualization SystemsabstractInformation visualizations are an efficient means to support the users in understanding large amounts of complex, interconnected data; user comprehension, however, depends on individual factors such as their cognitive abilities. The research literature provides evidence that user-adaptive information visualizations positively impact the users’ performance in visualization tasks. This study attempts to contribute toward the development of a computational model to predict the users’ success in visual search tasks from eye gaze data and thereby drive such user-adaptive systems. State-of-the-art deep learning models for time series classification have been trained on sequential eye gaze data obtained from 40 study participants’ interaction with a circular and an organizational graph. The results suggest that such models yield higher accuracy than a baseline classifier and previously used models for this purpose. In particular, a Multivariate Long Short Term Memory Fully Convolutional Network shows encouraging performance for its use in online user-adaptive systems. Given this finding, such a computational model can infer the users’ need for support during interaction with a graph and trigger appropriate interventions in user-adaptive information visualization systems. This facilitates the design of such systems since further interaction data like mouse clicks is not required. Moritz Spiller, Ying-Hsang Liu, Tom Gedeon, Julia Geißler, Andreas Nürnberger |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2020 | RealSmileNet: A Deep End-to-End Network for Spontaneous and Posed Smile Recognition
Yan Yang 0011, Tom Gedeon, Shafin Rahman |
ACCV (5) | 3 |
| 2020 | Simulating Content Consistent Vehicle Datasets with Attribute Descent
Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Naphade, Tom Gedeon |
ECCV (6) | 5 |
| 2020 | EmotiW 2020: Driver Gaze, Group Emotion, Student Engagement and Physiological Signal based ChallengesabstractThis paper introduces the Eighth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW is a benchmarking effort run as a grand challenge of the 22nd ACM International Conference on Multimodal Interaction 2020. It comprises of four tasks related to automatic human behavior analysis: a) driver gaze prediction; b) audio-visual group-level emotion recognition; c) engagement prediction in the wild; and d) physiological signal based emotion recognition. The motivation of EmotiW is to bring researchers in affective computing, computer vision, speech processing and machine learning to a common platform for evaluating techniques on a test data. We discuss the challenge protocols, databases and their associated baselines. Abhinav Dhall, Roland Göcke, Tom Gedeon |
ICMI | 4 |
| 2020 | Identifying Real and Posed Smiles from Observers' Galvanic Skin Response and Blood Volume Pulse
Renshang Gao, Atiqul Islam, Tom Gedeon |
ICONIP (1) | 3 |
| 2020 | A Token-Wise CNN-Based Method for Sentence Compression
Weiwei Hou, Hanna Suominen, Piotr Koniusz, Sabrina B. Caldwell, Tom Gedeon |
ICONIP (1) | 5 |
| 2020 | A Genetic Feature Selection Based Two-Stream Neural Network for Anger Veracity Recognition
Chaoxing Huang, Xuanying Zhu, Tom Gedeon |
ICONIP (1) | 3 |
| 2020 | Are Deep Neural Architectures Losing Information? Invertibility is Indispensable
Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Sabrina B. Caldwell, Tom Gedeon |
ICONIP (3) | 5 |
| 2020 | Disguising Personal Identity Information in EEG Signals
Shiya Liu, Yue Yao 0001, Chaoyue Xing, Tom Gedeon |
ICONIP (5) | 4 |
| 2020 | Pairwise-GAN: Pose-Based View Synthesis Through Pair-Wise Training
Xuyang Shen, Jo Plested, Yue Yao 0001, Tom Gedeon |
ICONIP (4) | 4 |
| 2020 | MultiTune: Adaptive Integration of Multiple Fine-Tuning Models for Image Classification
Jo Plested, Tom Gedeon |
ICONIP (4) | 3 |
| 2020 | Exploring the Correlation Between Random Convolutional Architectures and the Trained EquivalentabstractIn this paper we explore the correlation between Convolutional Neural Network (CNN) architectures with random weights in the convolutional layers to the same architectures with trained weights. We show that this correlation extends to deep CNN architectures of up to 10 or even 12 layers to the extent that untrained model accuracy could be a useful proxy for trained model accuracy. We also find that for models with fewer layers much of this relationship comes from the strong correlation between the number of features output from the final CNN layer and final accuracy. With 10 and 12 layers there is a moderate correlation even when the size of the fully connected layer is held constant. We anticipate our findings in extending these correlations to deeper networks will be useful in designing faster Neural Architecture Search (NAS) models. Analytically solving for the weights of the final prediction layer is orders of magnitude faster than training the weights via backpropagation. Nicholas Evans, Jo Plested, Tom Gedeon |
IJCNN | 3 |
| 2020 | Brain Melody Informatics: Analysing Effects of Music on Brainwave PatternsabstractRecently, researchers in the field of affective neuroscience have taken a keen interest in identifying patterns in brain activities that correspond to specific emotions. The relationship between music stimuli and brain waves has been of particular interest due to music's disputed effects on brain activity. While music can have an anticonvulsant effect on the brain and act as a therapeutic stimulus, it can also have proconvulsant effects such as triggering epileptic seizures. In this paper, we take a computational approach to understand the effects of different types of music on the human brain; we analyse the effects of 3 different genres of music in participants electroencephalograms (EEGs). Brain activity was recorded using a 14-channel headset from 24 participants while they listened to different music stimuli. Statistical features were extracted from the signals and useful features and channels were identified using various feature selecting techniques. Using these features we built classification models based on K-nearest Neighbour (KNN), Support Vector Machine (SVM) and Neural Network (NN). Our analysis shows that NN, along with Genetic Algorithm (GA) feature selection, can reach the highest accuracy of 97.5% in classifying the 3 music genres. The model also reaches 98.6% accuracy in classifying music based on participants' subjective rating of emotion. Additionally, the recorded brain waves identify different gamma wave levels, which are crucial in detecting epileptic seizures. Our results show that these computational techniques are effective in distinguishing music genres based on their effects on human brains. Jessica Sharmin Rahman, Tom Gedeon, Sabrina B. Caldwell, Richard Jones 0002 |
IJCNN | 2 |
| 2020 | Deceit Detection: Identification of Presenter's Subjective Doubt Using Affective Observation Neural Network AnalysisabstractWe live in a world surrounded with `fake news' and manipulated information, so a system assisting people with knowing what information to trust would be beneficial. Our research investigates situations where the presenters themselves have doubts about the information they are delivering, and we detect this via advanced affective computing techniques. To this end we examine the physiological foundations for observer recognition of the doubt effect: the subjective belief or disbelief of a presenter in some information he or she is presenting. Firstly, we construct stimulus videos that display presenters delivering information about which we manipulate their degree of doubt. We then show these stimuli to observers, and record four of their physiological signals. We find that a generalised neural network trained with physiological features is more accurate in differentiating the presenters' doubt/manipulated belief when compared with the same observers' own conscious judgments. The affective recognition performance improves when we analyse the physiological signals using multi-task learning techniques to train personalised and group personalised neural networks. The ability to recognise this doubt effect derives from observers' fundamental emotional reactions to the viewed stimuli, reflected in their physiological responses, and learnt by our neural networks. We believe this system using observer physiological signals collected in real life could reveal accurate and hidden audience distrust, which could in turn lead to enhanced truthfulness in future public- presented statements. Xuanying Zhu, Tom Gedeon, Sabrina B. Caldwell, Richard Jones 0002, Xiaohan Gu |
SMC | 2 |
| 2020 | Information-preserving feature filter for short-term EEG signals
Yue Yao 0001, Jo Plested, Tom Gedeon |
Neurocomputing | 3 |
| 2020 | Using Temporal Features of Observers' Physiological Measures to Distinguish Between Genuine and Fake SmilesabstractFuture affective computing research could be enhanced by enabling the computer to recognise a displayer's mental state from an observer's reaction (measured by physiological signals), using this information to improve recognition algorithms, and eventually to computer systems which are more responsive to human emotions. In this paper, an observer's physiological signals are analysed to distinguish displayers' genuine from fake smiles. Overall, thirty smile videos were collected from four benchmark database and classified as showing genuine or fake smiles. Overall, forty observers viewed videos. We generally recorded four physiological signals: pupillary response (PR), electrocardiogram (ECG), galvanic skin response (GSR), and blood volume pulse (BVP). A number of temporal features were extracted after a few processing steps, and minimally correlated features between genuine and fake smiles were selected using the NCCA (canonical correlation analysis with neural network) system. Finally, classification accuracy was found to be as high as 98.8 percent from PR features using a leave-one-observer-out process. In comparison, the best current image processing technique [1] on the same video data was 95 percent correct. Observers were 59 percent (on average) to 90 percent (by voting) correct by their conscious choices. Our results demonstrate that humans can non-consciously (or emotionally) recognise the quality of smiles 4 percent better than current image processing techniques and 9 percent better than the conscious choices of groups. Tom Gedeon, Ramesh S. Sankaranarayana |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | A Neural Micro-Expression RecognizerabstractRecognizing micro-expressions underpins significant and critical research and significant application. We speculate that this problem requires the understanding of the subtle face movement, integration of face structures and a solution of limited training data. In this paper, we build an effective micro-expression recognition system that leverages techniques stemming from these speculations. First, we introduce an optical flow method based on the onset frame and the apex frame to encode the subtle face motion. This has already been validated by prior research. Second, to obtain discriminative representations from the rigid face structures, part-based average pooling is proposed to inject structure priors to the network. Finally, because the system suffers from small training sets, we propose to transfer domain knowledge from macro-expression recognition tasks to micro-expression recognition. Specifically, we adopt two domain adaptation techniques including adversarial training and expression magnification and reduction (EMR). Through experiment, we show that the proposed system achieves very competitive results on the 2ndMicro-Expression Grand Challenge (MEGC). Yuchi Liu, Heming Du, Liang Zheng 0001, Tom Gedeon |
FG | 4 |
| 2019 | Spotting Visual Keywords from Temporal Sliding WindowsabstractVisual Keyword Spotting (KWS), as a newly proposed task deriving from visual speech recognition, has plenty of room for improvements. This paper details our Visual Keyword Spotting system used in the first Mandarin Audio-Visual Speech Recognition Challenge (MAVSR 2019). With the assumption that the vocabularies of target dataset are a subset of the vocabulary of the training set, we proposed a simple and scalable classification based strategy that achieves 19.0% mean average precision (mAP) on this challenge. Our method is based on the idea of using sliding windows to bridge between the word-level dataset and the sentence-level dataset, showing that a strong word level classifier can be directly used in building sentence embedding, thereby making it possible to build a KWS system. Yue Yao 0001, Heming Du, Liang Zheng 0001, Tom Gedeon |
ICMI | 5 |
| 2019 | Improving Student Forum Responsiveness: Detecting Duplicate Questions in Educational Forums
Manal Mohania, Liyuan Zhou, Tom Gedeon |
ICONIP (3) | 3 |
| 2019 | An Analysis of the Interaction Between Transfer Learning Protocols in Deep Neural Networks
Jo Plested, Tom Gedeon |
ICONIP (1) | 2 |
| 2019 | Predicting Group Cohesiveness in ImagesabstractThe cohesiveness of a group is an essential indicator of the emotional state, structure and success of a group of people. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. In order to identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. For analyzing the contribution of facial expressions of the group members for predicting the Group Cohesion Score (GCS), a capsule network is explored. We add GCS to the Group Affect database and propose the `GAF-Cohesion database'. The proposed model performs well on the database and is able to achieve near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated. Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon |
IJCNN | 4 |
| 2019 | Generalized Alignment for Multimodal Physiological Signal LearningabstractRevealing the correspondences and relationships between physiological signals is attractive for bioinformatics and human-computer interaction. Time alignment is a straightforward way to figure out correspondences between time sequential data. However, alignment between multimodal physiological signals is hard to achieve because the similarity metrics are difficult to define if the two physiological signals being investigated are non-linearly correlated, misaligned or quite different in morphology. In this paper, we propose a generalized time alignment method for multimodal physiological signals which (i) learns the feature extractions on physiological signals in a generalized way, and (ii) enables learned features to be in a coordinated space where the similarity between sub-components from two signals can be defined. Furthermore, we applied our alignment based multimodal feature fusion on an evaluation model to perform emotion recognition tasks on the DEAP multimodal physiological signal dataset. The experimental results show that the alignment based feature fusion outperforms the non-aligned feature fusion in most cases. Yuchi Liu, Yue Yao 0001, Jo Plested, Tom Gedeon |
IJCNN | 5 |
| 2019 | Melodious Micro-frissons: Detecting Music Genres From Skin ResponseabstractThe relationship between music and human physiological signals has been a topic of interest among researchers for many years. Understanding this relationship can not only lead to more enhanced music therapy methods, but it may also help in finding a cure to mental disorders and epileptic seizures that are triggered by certain music. In this paper, we investigate the effects of 3 different genres of music in participants' Electrodermal Activity (EDA). Signals were recorded from 24 participants while they listened to 12 music stimuli. Various feature selection methods were applied to a number of features which were extracted from the signals. A simple neural network using Genetic Algorithm (GA) feature selection can reach as high as 96.8% accuracy in classifying 3 different music genres. Classification based on participants' subjective rating of emotion reaches 98.3% accuracy with the Statistical Dependency (SD) / Minimal Redundancy Maximum Relevance (MRMR) feature selection technique. This shows that human emotion has a strong correlation with different types of music. In the future this system can be used to distinguish music based on their positive of negative effect on human mental health. Jessica Sharmin Rahman, Tom Gedeon, Sabrina B. Caldwell, Richard Jones 0002, Xuanying Zhu |
IJCNN | 2 |
| 2019 | Improved Techniques for Building EEG Feature FiltersabstractRecent advances in the generative adversarial network (GAN) based image translation have shown its potential of being an image style transformer. Similarly, defined as a style transformer for physiological signals, a feature filter is used to filter privacy-related features while still keeping useful features. However, existing feature filter techniques have three problems: (1) the privacy-related features cannot be filtered out to the extent we need through a simple Conv-Deconv generator structure, and (2) the generator cannot control the semantics (maintain desired features) of given physiological signals. To address these problems, we utilize deeper neural networks and adopt techniques from domain adaptation. This includes semantic loss and a GAN based model structure with two generators, two discriminators and a classifier to form a game of five. Our results on the UCI EEG dataset demonstrate that our model can simultaneously (1) achieve the state-of-the-art accuracy removal for the privacy-related feature, (2) reduce the desired feature removal accuracy drop, and (3) make the filtered signals can be interpreted or visually checked. Yue Yao 0001, Jo Plested, Tom Gedeon, Yuchi Liu |
IJCNN | 3 |
| 2019 | Your Eyes Say You're Lying: An Eye Movement Pattern Analysis for Face Familiarity and Deceptive CognitionabstractEye movement patterns reflect human latent internal cognitive activities. We aim to discover eye movement patterns during face recognition under different conditions of information concealment. These conditions include the degrees of face familiarity and deception or not, namely telling the truth when observing familiar and unfamiliar faces, and deceiving in front of familiar and unfamiliar faces. We apply Hidden Markov models with Gaussian emission to generalise regions and trajectories of eye fixation points under the above four conditions. Our results show that both eye movement patterns and eye gaze regions become significantly different during deception compared with truth-telling. We show the feasibility of detecting deception and further cognitive activity classification using eye movement patterns. Jiaxu Zuo, Tom Gedeon, Zhenyue Qin |
IJCNN | 2 |
| 2019 | Observers' physiological measures in response to videos can be used to detect genuine smiles
Tom Gedeon |
Int. J. Hum. Comput. Stud. | 2 |
| 2019 | The effects of perceived chronic pressure and time constraint on information search behaviors and experience
Chang Liu 0007, Ying-Hsang Liu, Tom Gedeon |
Inf. Process. Manag. | 3 |
| 2019 | Health Informatics: Applications of Mobile and Wireless Technologies
Milos Stojmenovic, Tom Gedeon, Heng Qi, Seyed M. Buhari |
Wirel. Commun. Mob. Comput. | 2 |
| 2018 | EmotiW 2018: Audio-Video, Student Engagement and Group-Level Affect PredictionabstractThis paper details the sixth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW 2018 is a grand challenge in the ACM International Conference on Multimodal Interaction 2018, Colarado, USA. The challenge aims at providing a common platform to researchers working in the affective computing community to benchmark their algorithms on 'in the wild' data. This year EmotiW contains three sub-challenges: a) Audio-video based emotion recognition; b) Student engagement prediction; and c) Group-level emotion recognition. The databases, protocols and baselines are discussed in detail. Abhinav Dhall, Amanjot Kaur, Roland Göcke, Tom Gedeon |
ICMI | 4 |
| 2018 | An Independent Approach to Training Classifiers on Physiological Data: An Example Using Smiles
Tom Gedeon |
ICONIP (2) | 2 |
| 2018 | Neural Networks Assist Crowd Predictions in Discerning the Veracity of Emotional Expressions
Zhenyue Qin, Tom Gedeon, Sabrina B. Caldwell |
ICONIP (6) | 2 |
| 2018 | Artificial Neural Networks Can Distinguish Genuine and Acted Anger by Synthesizing Pupillary Dilation Signals from Different Participants
Zhenyue Qin, Tom Gedeon, Xuanying Zhu |
ICONIP (5) | 2 |
| 2018 | Neural Causality Detection for Multi-dimensional Point Processes
Christian J. Walder, Tom Gedeon |
ICONIP (4) | 3 |
| 2018 | Deep Feature Learning and Visualization for EEG Recording Using Autoencoders
Yue Yao 0001, Jo Plested, Tom Gedeon |
ICONIP (7) | 3 |
| 2018 | A Feature Filter for EEG Using Cycle-GAN Structure
Yue Yao 0001, Jo Plested, Tom Gedeon |
ICONIP (7) | 3 |
| 2018 | Detecting the Doubt Effect and Subjective Beliefs Using Neural Networks and Observers' Pupillary Responses
Xuanying Zhu, Zhenyue Qin, Tom Gedeon, Richard Jones 0002, Sabrina B. Caldwell |
ICONIP (4) | 3 |
| 2017 | What Snippet Size is Needed in Mobile Web Search?abstractA snippet (content summary for a web page) is one of the main elements in a search result page. Search engines have been improved to reduce users' effort in web search, e.g., providing flexible snippet sizes by considering the purpose of the search and suggesting predicted answers. In most cases, search engines for mobile devices present two or three lines of snippet for each result link. Some studies suggest that long snippets provide a better search experience on desktop screens, but this may not be true for mobile devices because of the smaller screen. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
CHIIR | 4 |
| 2017 | From individual to group-level emotion recognition: EmotiW 5.0abstractResearch in automatic affect recognition has come a long way. This paper describes the fifth Emotion Recognition in the Wild (EmotiW) challenge 2017. EmotiW aims at providing a common benchmarking platform for researchers working on different aspects of affective computing. This year there are two sub-challenges: a) Audio-video emotion recognition and b) group-level emotion recognition. These challenges are based on the acted facial expressions in the wild and group affect databases, respectively. The particular focus of the challenge is to evaluate method in `in the wild' settings. `In the wild' here is used to describe the various environments represented in the images and videos, which represent real-world (not lab like) scenarios. The baseline, data, protocol of the two challenges and the challenge participation are discussed in detail in this paper. Abhinav Dhall, Roland Göcke, Shreya Ghosh 0001, Jyoti Joshi, Jesse Hoey, Tom Gedeon |
ICMI | 6 |
| 2017 | Effect of Parameter Tuning at Distinguishing Between Real and Posed Smiles from Observers' Physiological Features
Tom Gedeon |
ICONIP (4) | 2 |
| 2016 | Pagination versus Scrolling in Mobile Web SearchabstractVertical scrolling is the standard method of exploring search results pages. For touch-enabled mobile devices that are not equipped with a mouse or keyboard, we adopt other methods of controlling the viewport with the aim of investigating user interaction. From the intuition that people are used to reading books by turning pages horizontally, we conducted a user experiment to investigate the effects of horizontal and vertical control types (pagination versus scrolling) on a touch-enabled mobile phone. Our findings suggest that participants using pagination were more likely to find relevant documents, especially those over the fold; spent more time attending to relevant results; and were faster to click while spending less time on the search result pages overall. We also found that the main reason for the difference in search speed is the time taken for the scroll itself. We conclude that search engines need to provide different viewport controls to allow better search experiences on touch-enabled mobile devices. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
CIKM | 4 |
| 2016 | Emotion recognition in the wild challenge 2016abstractThe fourth Emotion Recognition in the Wild (EmotiW) challenge is a grand challenge in the ACM International Conference on Multimodal Interaction 2016, Tokyo. EmotiW is a series of benchmarking and competition effort for researchers working in the area of automatic emotion recognition in the wild. The fourth EmotiW has two sub-challenges: Video based emotion recognition (VReco) and Group-level emotion recognition (GReco). The VReco sub-challenge is being run for the fourth time and GReco is a new sub-challenge this year. Abhinav Dhall, Roland Göcke, Jyoti Joshi, Tom Gedeon |
ICMI | 4 |
| 2016 | EmotiW 2016: video and group-level emotion recognition challengesabstractThis paper discusses the baseline for the Emotion Recognition in the Wild (EmotiW) 2016 challenge. Continuing on the theme of automatic affect recognition `in the wild', the EmotiW challenge 2016 consists of two sub-challenges: an audio-video based emotion and a new group-based emotion recognition sub-challenges. The audio-video based sub-challenge is based on the Acted Facial Expressions in the Wild (AFEW) database. The group-based emotion recognition sub-challenge is based on the Happy People Images (HAPPEI) database. We describe the data, baseline method, challenge protocols and the challenge results. A total of 22 and 7 teams participated in the audio-video based emotion and group-based emotion sub-challenges, respectively. Abhinav Dhall, Roland Göcke, Jyoti Joshi, Jesse Hoey, Tom Gedeon |
ICMI | 5 |
| 2016 | Mitigating distractions during online reading: An explorative studyabstractReading online can be difficult due to the distractions of digital environments. In this paper we present a user study in which participants' eye gaze was recorded as they read text in a visually distracting environment. We explore two distraction mitigation signals using real-time eye gaze data to investigate whether the effects help reduce distraction rate as well as aid recovery from distractions. These signals involved adding a signal to the last word read before a distraction occurred to show the reader where they were up to. We compared these experimental conditions on both first (L1) and second (L2) English language readers and for easy and difficult to read texts. The results demonstrate that the mitigation signals helped recovery from a distraction by drawing participants' attention back to the text as well as indicating from where to recommence reading. We conclude with recommendations on implementing distraction mitigation signals in text. Leana Copeland, Tom Gedeon, Sabrina B. Caldwell |
SMC | 2 |
| 2016 | Automatic clustering of eye gaze data for machine learningabstractEye gaze patterns or scanpaths of subjects looking at art while answering questions related to the art have been used to decode those tasks with the use of certain classifiers and machine learning techniques. Some of these techniques require the artwork to be divided into several Areas or Regions of Interest. In this paper, two ways of clustering the static visual stimuli - k-means and the density based clustering algorithm called OPTICS - were used for this purpose. These algorithms were used to cluster the gaze points before classification. The classification success rates were then compared. While it was observed that both k-means and OPTICS gave better success rates than manual clustering, which is itself higher than chance level, OPTICS consistently gave higher success rates than k-means given the right parameter settings. OPTICS also formed clusters that look more intuitive and consistent with the heat map readings than k-means, which formed clusters that look unintuitive and less consistent with the heat map. Khushnood Z. Naqshbandi, Tom Gedeon, Umran Azziz Abdulla |
SMC | 2 |
| 2016 | Understanding eye movements on mobile devices for better presentation of search resultsabstractCompared to the early versions of smart phones, recent mobile devices have bigger screens that can present more web search results. Several previous studies have reported differences in user interaction between conventional desktop computer and mobile device‐based web searches, so it is imperative to consider the differences in user behavior for web search engine interface design on mobile devices. However, it is still unknown how the diversification of screen sizes on hand‐held devices affects how users search. In this article, we investigate search performance and behavior on three different small screen sizes: early smart phones, recent smart phones, and phablets. We found no significant difference with respect to the efficiency of carrying out tasks, however participants exhibited different search behaviors: less eye movement within top links on the larger screen, fast reading with some hesitation before choosing a link on the medium, and frequent use of scrolling on the small screen. This result suggests that the presentation of web search results for each screen needs to take into account differences in search behavior. We suggest several ideas for presentation design for each screen size. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2015 | Accuracy and awareness of image veracity in human perceptions of manipulated and unmanipulated images
Sabrina B. Caldwell, Tom Gedeon, Richard Jones 0002, Leana Copeland |
CogSci | 2 |
| 2015 | Conservation of relative fuzziness: Retrospective and triangular extensionabstractFuzzy rule interpolation is one of the tools for reducing computational complexity of fuzzy systems, and can be used when there are gaps in the knowledge base. These gaps can be natural, due to cost, or due to rule base reduction. The fuzzy interpolation methods are all descendent techniques of Kóczy and Hirota's linear interpolation. In this paper we provide a retrospective on the development of these techniques, and then focus on an early technique of conservation of fuzziness which has advantages in interpolation in hierarchical fuzzy systems as only near flank information is meant to be used and this allows the interpolation between different levels in the fuzzy rule base hierarchy. We point out an error and rectify it using a triangular extension which restores the intuitive, philosophical and practical nature of the approach. Tom Gedeon |
FUZZ-IEEE | 1 |
| 2015 | Fighting the Symmetries: The Structure of Cryptographic Boolean Function SpacesabstractWe explore the problem space of maximum nonlinearity problems for balanced Boolean functions, examining the symmetry structure and fitness landscapes in the most common (bit string) representation. We present theoretical analyses of well understood aspects, together with detailed enumeration of the 4-bit problem, sampling of the 6-bit problem based on known optima, and sampling of the 8-bit problem based on its fittest known solutions. Stjepan Picek, Robert I. McKay, Roberto Santana 0001, Tom Gedeon |
GECCO | 4 |
| 2015 | Video and Image based Emotion Recognition Challenges in the Wild: EmotiW 2015abstractThe third Emotion Recognition in the Wild (EmotiW) challenge 2015 consists of an audio-video based emotion and static image based facial expression classification sub-challenges, which mimics real-world conditions. The two sub-challenges are based on the Acted Facial Expression in the Wild (AFEW) 5.0 and the Static Facial Expression in the Wild (SFEW) 2.0 databases, respectively. The paper describes the data, baseline method, challenge protocol and the challenge results. A total of 12 and 17 teams participated in the video based emotion and image based expression sub-challenges, respectively. Abhinav Dhall, O. V. Ramana Murthy, Roland Göcke, Jyoti Joshi, Tom Gedeon |
ICMI | 5 |
| 2015 | Extreme Learning Machines with Simple CascadesabstractWe compare extreme learning machines with cascade correlation on a standard benchmark dataset for
comparing cascade networks along with another commonly used dataset. We introduce a number of hybrid
cascade extreme learning machine topologies ranging from simple shallow cascade ELM networks to full
cascade ELM networks. We found that the simplest cascade topology provided surprising benefit with a
cascade correlation style cascade for small extreme learning machine layers. Our full cascade ELM
architecture achieved high performance with even a single neuron per ELM cascade, suggesting that our
approach may have general utility, though further work needs to be done using more datasets. We suggest
extensions of our cascade ELM approach, with the use of network analysis, addition of noise, and
unfreezing of weights. Tom Gedeon, Anthony Oakden |
SIMULTECH | 1 |
| 2015 | Eye-tracking analysis of user behavior and performance in web search on large and small screensabstractIn recent years, searching the web on mobile devices has become enormously popular. Because mobile devices have relatively small screens and show fewer search results, search behavior with mobile devices may be different from that with desktops or laptops. Therefore, examining these differences may suggest better, more efficient designs for mobile search engines. In this experiment, we use eye tracking to explore user behavior and performance. We analyze web searches with 2 task types on 2 differently sized screens: one for a desktop and the other for a mobile device. In addition, we examine the relationships between search performance and several search behaviors to allow further investigation of the differences engendered by the screens. We found that users have more difficulty extracting information from search results pages on the smaller screens, although they exhibit less eye movement as a result of an infrequent use of the scroll function. However, in terms of search performance, our findings suggest that there is no significant difference between the 2 screens in time spent on search results pages and the accuracy of finding answers. This suggests several possible ideas for the presentation design of search results pages on small devices. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2015 | Automatic Group Happiness Intensity AnalysisabstractThe recent advancement of social media has given users a platform to socially engage and interact with a larger population. Millions of images and videos are being uploaded everyday by users on the web from different events and social gatherings. There is an increasing interest in designing systems capable of understanding human manifestations of emotional attributes and affective displays. As images and videos from social events generally contain multiple subjects, it is an essential step to study these groups of people. In this paper, we study the problem of happiness intensity analysis of a group of people in an image using facial expression analysis. A user perception study is conducted to understand various attributes, which affect a person's perception of the happiness intensity of a group. We identify the challenges in developing an automatic mood analysis system and propose three models based on the attributes in the study. An `in the wild' image-based database is collected. To validate the methods, both quantitative and qualitative experiments are performed and applied to the problem of shot selection, event summarisation and album creation. The experiments show that the global and local attributes defined in the paper provide useful information for theme expression analysis, with results close to human perception results. Abhinav Dhall, Roland Göcke, Tom Gedeon |
IEEE Trans. Affect. Comput. | 3 |
| 2014 | Emotion Recognition In The Wild Challenge 2014: Baseline, Data and ProtocolabstractThe Second Emotion Recognition In The Wild Challenge (EmotiW) 2014 consists of an audio-video based emotion classification challenge, which mimics the real-world conditions. Traditionally, emotion recognition has been performed on data captured in constrained lab-controlled like environment. While this data was a good starting point, such lab controlled data poorly represents the environment and conditions faced in real-world situations. With the exponential increase in the number of video clips being uploaded online, it is worthwhile to explore the performance of emotion recognition methods that work `in the wild'. The goal of this Grand Challenge is to carry forward the common platform defined during EmotiW 2013, for evaluation of emotion recognition methods in real-world conditions. The database in the 2014 challenge is the Acted Facial Expression In Wild (AFEW) 4.0, which has been collected from movies showing close-to-real-world conditions. The paper describes the data partitions, the baseline method and the experimental protocol. Abhinav Dhall, Roland Göcke, Jyoti Joshi, Karan Sikka, Tom Gedeon |
ICMI | 5 |
| 2014 | Fuzzy Output Error as the Performance Function for Training Artificial Neural Networks to Predict Reading Comprehension from Eye Gaze
Leana Copeland, Tom Gedeon, B. Sumudu U. Mendis |
ICONIP (1) | 2 |
| 2014 | Fuzzy Signature Neural Networks for Classification: Optimising the Structure
Tom Gedeon, Xuanying Zhu, Kun He 0009, Leana Copeland |
ICONIP (1) | 1 |
| 2014 | Modeling observer stress for typical real environments
Nandita Sharma, Tom Gedeon |
Expert Syst. Appl. | 2 |
| 2013 | Modeling Stress Using Thermal Facial Patterns: A Spatio-temporal ApproachabstractStress is a serious concern facing our world today, motivating the development of better objective understanding using non-intrusive means for stress recognition. The aim for the work was to use thermal imaging of facial regions to detect stress automatically. The work uses facial regions captured in videos in thermal (TS) and visible (VS) spectrums and introduces our database ANU StressDB. It describes the experiment conducted for acquiring TS and VS videos of observers of stressed and not-stressed films for the ANU StressDB. Further, it presents an application of local binary patterns on three orthogonal planes (LBP-TOP) on VS and TS videos for stress recognition. It proposes a novel method to capture dynamic thermal patterns in histograms (HDTP) to utilize thermal and spatio-temporal characteristics associated in TS videos. Individual-independent support vector machine classifiers were developed for stress recognition. Results show that a fusion of facial patterns from VS and TS videos produced significantly better stress recognition rates than patterns from only VS or TS videos with p <; 0.01. The best stress recognition rate was 72% and it was obtained from HDTP features fused with LBP-TOP features for TS and VS videos respectively. Nandita Sharma, Abhinav Dhall, Tom Gedeon, Roland Göcke |
ACII | 3 |
| 2013 | A comparative study of different classifiers for detecting depression from spontaneous speechabstractAccurate detection of depression from spontaneous speech could lead to an objective diagnostic aid to assist clinicians to better diagnose depression. Little thought has been given so far to which classifier performs best for this task. In this study, using a 60-subject real-world clinically validated dataset, we compare three popular classifiers from the affective computing literature - Gaussian Mixture Models (GMM), Support Vector Machines (SVM) and Multilayer Perceptron neural networks (MLP) - as well as the recently proposed Hierarchical Fuzzy Signature (HFS) classifier. Among these, a hybrid classifier using GMM models and SVM gave the best overall classification results. Comparing feature, score, and decision fusion, score fusion performed better for GMM, HFS and MLP, while decision fusion worked best for SVM (both for raw data and GMM models). Feature fusion performed worse than other fusion methods in this study. We found that loudness, root mean square, and intensity were the voice features that performed best to detect depression in this dataset. Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Tom Gedeon, Michael Breakspear, Gordon Parker |
ICASSP | 5 |
| 2013 | Emotion recognition in the wild challenge (EmotiW) challenge and workshop summaryabstractThe Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenge, which mimics real-world conditions. In total, 27 teams participated in the challenge. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions. Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon |
ICMI | 5 |
| 2013 | Emotion recognition in the wild challenge 2013abstractEmotion recognition is a very active field of research. The Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenges, which mimics real-world conditions. Traditionally, emotion recognition has been performed on laboratory controlled data. While undoubtedly worthwhile at the time, such laboratory controlled data poorly represents the environment and conditions faced in real-world situations. The goal of this Grand Challenge is to define a common platform for evaluation of emotion recognition methods in real-world conditions. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions. Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon |
ICMI | 5 |
| 2013 | Distance Metrics for Time-Series Data with Concentric Multi-Sphere Self Organizing Maps
Tom Gedeon, Lachlan Paget, Dingyun Zhu |
ICONIP (2) | 1 |
| 2013 | Classification of Physiological Sensor Signals Using Artificial Neural Networks
Nandita Sharma, Tom Gedeon |
ICONIP (2) | 2 |
| 2013 | Wands Are Magic: A Comparison of Devices Used in 3D Pointing Interfaces
Martin Henschke, Tom Gedeon, Richard Jones 0002, Sabrina B. Caldwell, Dingyun Zhu |
INTERACT (3) | 2 |
| 2013 | Computational Models of Stress in Reading Using Physiological and Physical Sensor Data
Nandita Sharma, Tom Gedeon |
PAKDD (1) | 2 |
| 2012 | Comparing User Performance on an iPad to a 17-inch BackPadabstractWhat will a truly large iPad be like? Will it have a touchscreen at the front, or will some other changes be forced by the sheer sizeof the device? We mocked up a working device using a 17-inch Macbook laptop screen. The device size was too large for us to comfortably hold with one hand while using the other hand for touch input, so we placed the touch pad at the back. Hence, wecall our device a BackPad. In the first experiment, we compared user performance with our 17-inch BackPad and a normal iPad in game and typing tasks. The results on the game completion time and score were similar, and users liked our large screen,while time but not spelling errors were different in the BackPad versus the iPad. For the second experiment, we compared the front touchscreen versus the back trackpad user performance on same sized devices. Similar results to the first experiment were found on game completing time and score. Fateme Rajabiyazdi, Tom Gedeon |
CISIS | 2 |
| 2012 | Hand Grip Strength on a Large PDA: Holding While Reading Is Different from a Functional TaskabstractSeveral studies have been done measuring preferred hand grip strength, but none of them has measured preferred hand strength on a PDA or similar device when it is held and used. We measured dominant hand strength in two conditions similar to real PDA use, resting fore-arms on a table and holding the PDA without table support. We found that adult participants squeeze the device with their preferred hand significantly more than with their nonpreferred hand while holding. In addition, we examined users' hand strength while they were tapping on the back of the device with their right and left index fingers. Our results were different than expected from previous studies, as we found that there was no significant difference in dominant and non dominant hand strength during back tapping. Also participants' non preferred hand strength was not significantly different with their preferred hand when they tap on the back of the device. The results show that in such functional use during tapping, the dominant and non-dominant hands are used similarly which will contribute to future designs for PDAs and their interfaces. Our results may also contribute to design for more comfortable devices for users with hand disabilities. Fateme Rajabiyazdi, Tom Gedeon |
CISIS | 2 |
| 2012 | Artificial Neural Network Classification Models for Stress in Reading
Nandita Sharma, Tom Gedeon |
ICONIP (4) | 2 |
| 2012 | Modeling the Mental Differentiation Task with EEG
Tan Vo, Tom Gedeon, Dat Tran 0001 |
ICONIP (2) | 2 |
| 2012 | Complex Structured Decision Making Model: A hierarchical frame work for complex structured data
B. Sumudu U. Mendis, Tom Gedeon |
Inf. Sci. | 2 |
| 2011 | Exploring camera viewpoint control models for a multi-tasking setting in teleoperationabstractControl of camera viewpoint plays a vital role in many teleoperation activities, as watching live video streams is still the fundamental way for operators to obtain situational awareness from remote environments. Motivated by a real-world industrial setting in mining teleoperation, we explore several possible solutions to resolve a common multi-tasking situation where an operator is required to control a robot and simultaneously perform remote camera operation. Conventional control interfaces are predominantly used in such teleoperation settings, but could overload an operator's hand-operation capability, and require frequent attention switches and thus could decrease productivity. We report on an empirical user study in a model multi-tasking teleoperation setting where the user has a main task which requires their attention. We compare three different camera viewpoint control models: (1) dual manual control, (2) natural interaction (combining eye gaze and head motion) and (3) autonomous tracking. The results indicate the advantages of using the natural interaction model, while the manual control model performed the worst. Dingyun Zhu, Tom Gedeon, Ken Taylor |
CHI | 2 |
| 2011 | Emotion recognition using PHOG and LPQ featuresabstractWe propose a method for automatic emotion recognition as part of the FERA 2011 competition. The system extracts pyramid of histogram of gradients (PHOG) and local phase quantisation (LPQ) features for encoding the shape and appearance information. For selecting the key frames, K-means clustering is applied to the normalised shape vectors derived from constraint local model (CLM) based face tracking on the image sequences. Shape vectors closest to the cluster centers are then used to extract the shape and appearance features. We demonstrate the results on the SSPNET GEMEP-FERA dataset. It comprises of both person specific and person independent partitions. For emotion classification we use support vector machine (SVM) and largest margin nearest neighbour (LMNN) and compare our results to the pre-computed FERA 2011 emotion challenge baseline. Abhinav Dhall, Akshay Asthana, Roland Göcke, Tom Gedeon |
FG | 4 |
| 2011 | Performance enhancement of Hierarchical Document Signature: A comprehensive studyabstractHierarchical Document Signature (HDS) has been successfully applied in document computing to find similarity between different pieces of text [1], [2], [3]; for example sentence sentence similarity, sentence-phrase similarity. HDS is application specific, it is dependent on different features at different levels. This paper hence presents a comprehensive study of enhancement of the performance of HDS to find semantic sentence similarity by tuning some of its significant features. The experimental results support this and show the optimal conditions at which HDS performs similarly to humans. Sukanya Manna, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2011 | Document Classification on Relevance: A Study on Eye Gaze Patterns for Reading
Daniel Fahey, Tom Gedeon, Dingyun Zhu |
ICONIP (2) | 2 |
| 2011 | Stress Classification for Gender Bias in Reading
Nandita Sharma, Tom Gedeon |
ICONIP (3) | 2 |
| 2011 | Reading Your Mind: EEG during Reading Task
Tan Vo, Tom Gedeon |
ICONIP (1) | 2 |
| 2011 | Comparison between two mixed reality environments as a teleoperation interfaceabstractAn important aspect of teleoperation is situational awareness through visualization. The actual operation and control of a remote machine must be supported by an interface which provides enough information through visualization from a remote location to complete a task. This can be achieved with a Mixed Reality (MR) environment. The concept is to combine information from the real world and a virtual world. An experiment was conducted to assess the differences between two platforms and to determine interface features required to maximize operator performance and satisfaction. The result indicates that both mixed reality environments tested were suitable for teleoperation where sufficient information to perform the task could be modeled in the virtual world. However, one of the environments turned out to be superior where the task required information in the video but not modeled in the virtual environment. The preferred environment provided overlays on the video that were updated live as the model was manipulated where the other environment updated video overlays on completion of the manipulation. Ida Bagus Kerthyayana Manuaba, Ken Taylor, Tom Gedeon |
ICRA | 3 |
| 2011 | "Moving to the centre": A gaze-driven remote camera control for teleoperationabstractIn general, conventional control interfaces such as joysticks, switches, and wheels are predominantly used in teleoperation. However, operators normally have to control multiple complex devices simultaneously. For example, controlling a rock breaker and a remote camera at the same time in mining teleoperation. This overloads the operator’s control capability of using hands, increases workload and reduces productivity. We present a novel gaze-driven remote camera control with an implemented prototype, which follows a simple and natural design principle: “Whatever you look at on the screen, it moves to the centre!”. A user study of modeled hands-busy experiment has been conducted, comparing the performance of using gaze-driven control and traditional joystick control through both objective measures and subjective measures. The experimental results clearly show the gaze-driven control significantly outperformed the conventional joystick control. Dingyun Zhu, Tom Gedeon, Ken Taylor |
Interact. Comput. | 2 |
| 2010 | Improvements in Sugeno-Yasukawa modelling algorithmabstractA modified version of Sugeno-Yasukawa (SY) modelling algorithm is presented. We have employed a new method for parameter identification phase based on genetic algorithms (GA). Moreover, we have modified the modelling sequence by applying parameter identification on intermediate models. Models created with this method had lower mean square errors (MSE) compared to original algorithm. A case study on breast cancer survival prediction is also presented that demonstrates a thorough comparison of the new modelling algorithm with several other methods such as SVM, C5 decision tree, ANFIS and the original SY method. The modified SY method had the highest average of accuracies among all models. Moreover, it had significantly higher accuracy compared to the original SY method and ANFIS. 10-fold cross validation approach was employed for all evaluations. Amir Hossein Hadad, B. Sumudu U. Mendis, Tom Gedeon |
FUZZ-IEEE | 3 |
| 2010 | Semantic Hierarchical Document Signature for determining sentence similarityabstractIn this paper, we present a new approach that incorporates semantic information from a document, in the form of Hierarchical Document Signature (HDS), to measure semantic similarity between sentences. Due to variability of expressions of natural language, it is very essential to exploit the semantic properties of a document to accurately identify semantically similar sentences since sentences conveying the same fact or concept may be composed lexically and syntactically different. Inversely, sentences which are lexically common may not necessarily convey the same meaning. This poses a significant impact on many text mining applications performance where sentence-level judgment is involved. Our HDS uses the natural hierarchy of the document and represents it in a modularized form of document level to sentence level, sentence to word level; aggregating similarity components at the lower levels and propagating them to the next higher level to produce the final similarity between sentences. The evaluation of our HDS model has shown that it resembles the decision making process as done by human to a greater extent than different vector space models which only uses `bag of words' concept. Sukanya Manna, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2010 | Polymorphic fuzzy signaturesabstractThe fuzzy signature approach is aimed at finding a hierarchically decomposed solutions by adding new elements to Zadeh's approach. It tackles the problem by splitting the problem into hierarchically organized local sub-models and by applying more complex and heterogenous descriptors, more fit for the identification of extremely complex models. However, the computational time complexity still affects the fuzzy signatures as we were attempt to create an atomic fuzzy signature for each data point we get. Importantly, the atomic fuzzy signatures we store has properties we can make use of to make search over this structure computationally efficient. In this paper we introduce a new approach that uses the metadata about a set of fuzzy signatures to extract a Polymorphic Fuzzy Signature. Productively, a polymorphic fuzzy signature represents its base set of fuzzy signatures in a higher meta level which also allows search/inference, and so can reduce the computational time complexity of the inference process. B. Sumudu U. Mendis, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2010 | Enhancement of Subjective Logic for Semantic Document Analysis Using Hierarchical Document Signature
Sukanya Manna, Tom Gedeon, B. Sumudu U. Mendis |
ICONIP (1) | 2 |
| 2010 | Brain Computer Interfaces: A Recurrent Neural Network Approach
Gareth Oliver, Tom Gedeon |
ICONIP (2) | 2 |
| 2010 | Gaze Pattern and Reading Comprehension
Tan Vo, B. Sumudu U. Mendis, Tom Gedeon |
ICONIP (2) | 3 |
| 2010 | Estimation of Possibility-Probability Distributions
B. Sumudu U. Mendis, Tom Gedeon |
IPMU (1) | 2 |
| 2010 | An Enhanced Framework of Subjective Logic for Semantic Document Analysis
Sukanya Manna, B. Sumudu U. Mendis, Tom Gedeon |
MDAI | 3 |
| 2010 | Capture of Evidence for Summarization: An Application of Enhanced Subjective Logic
Sukanya Manna, B. Sumudu U. Mendis, Tom Gedeon |
PAKDD (2) | 3 |
| 2009 | Learning-based Face Synthesis for Pose-Robust Recognition from Single ImageabstractFace recognition in real-world conditions requires the ability to deal with a number of conditions, such as variations in pose, illumination and expression. In this paper, we focus on variations in head pose and use a computationally efficient regression-based approach for synthesising face images in different poses, which are used to extend the face recognition training set. In this data-driven approach, the correspondences between facial landmark points in frontal and non-frontal views are learnt offline from manually annotated training data via Gaussian Process Regression. We then use this learner to synthesise non-frontal face images from any unseen frontal image. To demonstrate the utility of this approach, two frontal face recognition systems (the commonly used PCA and the recent Multi-Region Histograms) are augmented with synthesised non-frontal views for each person. This synthesis and augmentation approach is experimentally validated on the FERET dataset, showing a considerable improvement in recognition rates for ±40° and ±60° views, while maintaining high recognition rates for ±15° and ±25° views. Akshay Asthana, Conrad Sanderson, Tom Gedeon, Roland Göcke |
BMVC | 3 |
| 2009 | Learning based automatic face annotation for arbitrary poses and expressions from frontal images onlyabstractStatistical approaches for building non-rigid deformable models, such as the active appearance model (AAM), have enjoyed great popularity in recent years, but typically require tedious manual annotation of training images. In this paper, a learning based approach for the automatic annotation of visually deformable objects from a single annotated frontal image is presented and demonstrated on the example of automatically annotating face images that can be used for building AAMs for fitting and tracking. This approach employs the idea of initially learning the correspondences between landmarks in a frontal image and a set of training images with a face in arbitrary poses. Using this learner, virtual images of unseen faces at any arbitrary pose for which the learner was trained can be reconstructed by predicting the new landmark locations and warping the texture from the frontal image. View-based AAMs are then built from the virtual images and used for automatically annotating unseen images, including images of different facial expressions, at any random pose within the maximum range spanned by the virtually reconstructed images. The approach is experimentally validated by automatically annotating face images from three different databases. Akshay Asthana, Roland Göcke, Novi Quadrianto, Tom Gedeon |
CVPR | 4 |
| 2009 | Motion control and communication of cooperating intelligent robots by fuzzy signaturesabstractThis paper presents two examples of usage of fuzzy signatures in the field of mobile robotics. The first shows a complex lateral drift control method base on fuzzy signatures. This method inspects the motion system of the robot as a whole, unlike as simple parts of a complex system. The state space is written down by fuzzy signatures which add up flexibility, adaptability and learning ability to the system. In the second experiment a new communication approach is investigated for intelligent cooperation of autonomous mobile robots. Effective, fast and compact communication is one of the most important cornerstones of a high-end cooperating system. In this paper we propose a fuzzy communication system where the codebooks are built up by fuzzy signatures. We use cooperating autonomous mobile robots to solve some logistic problems. Áron Ballagi, László T. Kóczy, Tom Gedeon |
FUZZ-IEEE | 3 |
| 2009 | Finding input sub-spaces for Polymorphic Fuzzy SignaturesabstractA significant feature of fuzzy signatures is its applicability for complex and sparse data. To create polymorphic fuzzy signatures (PFS) for sparse data, sparse input sub-spaces (ISSs) should be considered. Finding the optimal ISSs manually is not a simple task as it is time consuming; moreover, some knowledge about the dataset is necessary. Fuzzy c-means (FCM) clustering employed with a trapezoidal approximation method is needed to find ISSs automatically. Furthermore, dealing with sparse data, we should be mindful about choosing a reliable trapezoidal approximation method. This facilitates the optimal ISS creation for the data. In our experiment, two trapezoidal approximation methods were used to find optimal ISSs. The results demonstrate that our version of trapezoidal approximation for creating ISSs result in an PFS with lower mean square error compared to the original trapezoidal approximation method. Amir Hossein Hadad, Tom Gedeon, B. Sumudu U. Mendis |
FUZZ-IEEE | 2 |
| 2009 | Hierarchical document signature: A specialized application of fuzzy signature for document computingabstractWe develop document computing procedures for the analysis of discourse structures within a document, represented by hierarchical document signatures. A signature is a string of data characterizing a certain case (e.g. characteristics of a sentence in case of a document). The place of the individual data is fixed within the string, it holds a local value semantics. Fuzzy granulation is a semantic background technique for all kinds of information which originates from human estimation or recorded by human valuation of numerical data. For analysis of such data the development of special procedures is suggested, different from the usual statistical methods. We used a form of fuzzy signature, called hierarchical document signature to modularize an unstructured document in a hierarchical manner, from Document level to sentence level, sentence level to attribute level and then to word level. We used occurrence of words as the information of the lowest module to find the similarity among the next higher module by aggregating the signature values giving sentence pair coherence. Sukanya Manna, B. Sumudu U. Mendis, Tom Gedeon |
FUZZ-IEEE | 3 |
| 2009 | Keyboard before Head Tracking Depresses User Success in Remote Camera Control
Dingyun Zhu, Tom Gedeon, Ken Taylor |
INTERACT (2) | 2 |
| 2009 | Implications of resource limitations for a conscious machine
L. Andrew Coward, Tom Gedeon |
Neurocomputing | 2 |
| 2008 | A comparison: Fuzzy signatures and Choquet IntegralabstractFuzzy signatures are hierarchical multi aggregative descriptors of objects. They have reduced computational complexity compared to formal fuzzy rule based systems. Weighted relevance aggregation enhances the performance of hierarchical fuzzy signatures. Thus, they are very robust and flexible under perturbed input data. On the other hand the Choquet integral, which is based on fuzzy measures, is a powerful aggregation tool in multi-criteria decision making. We compared fuzzy signatures and the Choquet integral as practical applications for hierarchical and non-hierarchical data aggregation/organization methods. B. Sumudu U. Mendis, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2008 | Generalisation Performance vs. Architecture Variations in Constructive Cascade Networks
Suisin Khoo, Tom Gedeon |
ICONIP (2) | 2 |
| 2008 | A Hybrid Fuzzy Approach for Human Eye Gaze Pattern Recognition
Dingyun Zhu, B. Sumudu U. Mendis, Tom Gedeon, Akshay Asthana, Roland Göcke |
ICONIP (2) | 3 |
| 2008 | Fuzzy Logic for Cooperative Robot Communication
Dingyun Zhu, Tom Gedeon |
KES-AMSTA | 2 |
| 2008 | Pattern Trees Induction: A New Machine Learning MethodabstractFuzzy classification is one of the most important applications in fuzzy set and fuzzy-logic-related research. Its goal is to find a set of fuzzy rules that form a classification model. Most of the existing fuzzy rule induction methods (e.g., the fuzzy decision trees (FDTs) induction method) focus on searching rules consisting of triangular norms (t-norms) (i.e., and) only, but not triangular conorms (t-conorms) (or) explicitly. This may lead to the omission of generating important rules that involve t-conorms explicitly. This paper proposes a type of tree termed pattern trees (PTs) that makes use of different aggregations, including both t-norms and t-conorms. Like decision trees, PTs are an effective tool for classification applications. This paper discusses the difference between decision trees and PTs, and also shows that the subsethood-based method (SBM) and the weighted-subsethood-based method (WSBM) are two specific cases of PT induction. A novel PT induction method is proposed using similarity measure and fuzzy aggregations. The comparison to other classification methods including SBM, WSBM, C4.5, nearest neighbor, support vector machine, and FDT induction shows that: 1) PTs can obtain high accuracy rates in classifications; 2) PTs are robust to overfltting; and 3) PTs, especially simple pattern trees (SPTs), maintain compact tree structures. Zhiheng Huang, Tom Gedeon, Masoud Nikravesh |
IEEE Trans. Fuzzy Syst. | 2 |
| 2007 | Neural Network for Modeling Esthetic Selection
Tom Gedeon |
ICONIP (2) | 1 |
| 2007 | Weighted Pattern Trees: A Case Study with Customer Satisfaction Dataset
Zhiheng Huang, Masoud Nikravesh, Ben Azvine, Tom Gedeon |
IFSA (1) | 4 |
| 2007 | Fuzzy Signature and Cognitive Modelling for Complex Decision Model
Kevin Kok Wai Wong, Tom Gedeon, László T. Kóczy |
IFSA (2) | 2 |
| 2006 | Separated Antecedent and Consequent Learning for Takagi-Sugeno Fuzzy SystemsabstractIn this paper a new algorithm for the learning of Takagi-Sugeno fuzzy systems is introduced. In the algorithm different learning techniques are applied for the antecedent and the consequent parameters of the fuzzy system. We propose a hybrid method for the antecedent parameters learning based on the combination of the bacterial evolutionary algorithm (BEA) and the Levenberg-Marquardt (LM) method. For the linear parameters in fuzzy systems appearing in the rule consequents the least squares (LS) and the recursive least squares (RLS) techniques are applied, which will lead to a global optimal solution of linear parameter vectors in the least squares sense. Therefore a better performance can be guaranteed than with a complete learning by BEA and LM. The paper is concluded by evaluation results based on high-dimensional test data. These evaluation results compare the new method with some conventional fuzzy training methods with respect to approximation accuracy and model complexity. János Botzheim, Edwin Lughofer, Erich-Peter Klement, László T. Kóczy, Tom Gedeon |
FUZZ-IEEE | 5 |
| 2006 | Pattern TreesabstractThis paper proposes a new type of tree termed pattern trees. Like decision trees, pattern trees are an effective tool for classification applications. This paper discusses the difference between decision trees and pattern trees, and also shows that the subsethood based method and the weighted subsethood based method are two specific cases of pattern trees. A novel pattern tree induction method is proposed. The comparison to other classification methods including fuzzy decision tree induction shows that pattern trees can obtain higher accuracy rates in classifications. In addition, pattern trees are capable of generating patterns with good generality, while decision trees can easily fall into the trap of over-fitting. Zhiheng Huang, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2006 | Efficient Fuzzy Cognitive Modeling for Unstructured InformationabstractThis paper presents an efficient fuzzy cognitive modeling which can handle granulation, organisation and causation. This cognitive modeling technique consists of multiple levels where the lowest level includes details required to make a decision or to transfer to the next stage. This fuzzy cognitive modeling will enhance the usability of fuzzy theory in modeling complex systems as well as facilitating complex decision making process based on ill structured or missing information or data. Kevin Kok Wai Wong, Tom Gedeon, László T. Kóczy |
FUZZ-IEEE | 2 |
| 2006 | Is That Possible? (Or is it Probable?)
Tom Gedeon |
HIS | 1 |
| 2006 | Learning Generalized Weighted Relevance Aggregation Operators Using Levenberg-Marquardt Method
B. Sumudu U. Mendis, Tom Gedeon, László T. Kóczy |
HIS | 2 |
| 2006 | Uncertainty in Mineral Prospectivity Prediction
Pawalai Kraipeerapun, Lance Chun Che Fung, Warick Brown, Kevin Kok Wai Wong, Tom Gedeon |
ICONIP (2) | 5 |
| 2005 | Fuzzy Pseudo-Thesaurus Based Clustering of a Folkloristic CorpusabstractAutomatic thesaurus extraction is essential for modern information retrieval. We develop a method for fuzzy pseudo-thesaurus based on word pair co-occurrence in documents. In this study it is presented, that considering the word frequency degree counted on the whole corpus makes the obtained pseudo-thesaurus usable. Such parameters were found with which most of the obtained pairs of words were validated to be related by human expert. Among the extracted pairs and groups of words the relationship is often looser than synonymy, but they identify the frequently repeated topics of the corpus. We suggest the use of groups of closely related words for the definition of different topics and based on this clustering of the documents were performed (Chakrabarty, et al. (1999)) Sandor Szaszko, László T. Kóczy, Tom Gedeon |
FUZZ-IEEE | 3 |
| 2005 | Automatic generating detail-on-demand hypervideo using MPEG-7 and SMILabstractDetail-on-demand hypervideo will provide a powerful mechanism to allow viewers to see additional information of video segments through hyperlinks. A large number of tools are devoted to the identification of selectable video objects and the synchronizatio Tina T. Zhou, Tom Gedeon, Jesse S. Jin |
ACM Multimedia | 2 |
| 2005 | Fuzzy rule interpolation for multidimensional input spaces with applications: a case studyabstractFuzzy rule based systems have been very popular in many engineering applications. However, when generating fuzzy rules from the available information, this may result in a sparse fuzzy rule base. Fuzzy rule interpolation techniques have been established to solve the problems encountered in processing sparse fuzzy rule bases. In most engineering applications, the use of more than one input variable is common, however, the majority of the fuzzy rule interpolation techniques only present detailed analysis to one input variable case. This paper investigates characteristics of two selected fuzzy rule interpolation techniques for multidimensional input spaces and proposes an improved fuzzy rule interpolation technique to handle multidimensional input spaces. The three methods are compared by means of application examples in the field of petroleum engineering and mineral processing. The results show that the proposed fuzzy rule interpolation technique for multidimensional input spaces can be used in engineering applications. Kevin Kok Wai Wong, Domonkos Tikk, Tom Gedeon, László T. Kóczy |
IEEE Trans. Fuzzy Syst. | 3 |
| 2004 | Learning complex combinations of operations in a hybrid architectureabstractThe reasons why machine learning appears limited to the relatively simple control problems are analyzed. A primary issue is that, any condition detected by a learning system acquires multiple behavioural meanings. As the learning continues, the need to preserve these meanings severely constrains the architectural form of the system. A hybrid architecture called the recommendation architecture in which the preservation of such meanings is explicitly managed is compared with a wide range of alternative learning approaches. It is concluded that systems with this recommendation architecture have the capability to learn to solve the complex control problems. L. Andrew Coward, Tom Gedeon, Uditha Ratnayake |
FUZZ-IEEE | 2 |
| 2004 | Construction of fuzzy signature from data: an example of SARS pre-clinical diagnosis systemabstractThere are many areas where objects with very complex and sometimes interdependent features are to be classified; similarities and dissimilarities are to be evaluated. This makes a complex decision model difficult to construct effectively. Fuzzy signatures are introduced to handle complex structured data and interdependent feature problems. Fuzzy signatures can also be used in cases where data is missing. This work presents the concept of a fuzzy signature and how its flexibility can be used to quickly construct a medical pre-clinical diagnosis system. A severe acute respiratory syndrome (SARS) pre-clinical diagnosis system using fuzzy signatures is constructed as an example to show many advantages of the fuzzy signature. With the use of this fuzzy signature structure, complex decision models in the medical field should be able to be constructed more effectively. Kevin Kok Wai Wong, Tom Gedeon, László T. Kóczy |
FUZZ-IEEE | 2 |
| 2004 | Intelligent data mining and personalisation for customer relationship managementabstractCustomer relationship management (CRM) initiatives have gained much attention in recent years. With the aid of data mining technology, businesses can formulate specific strategies for different customer bases more precisely. Additionally, personalisation is another important issue in CRM - especially when a company has a huge product range. This paper presents a case model and investigates the use of computational intelligent techniques for CRM. These techniques allow the complex functions of relating customer behaviour to internal business processes to be learned more easily and the industry expertise and experience from business managers to be integrated into the modelling framework directly. Hence, they can be used in the CRM framework to enhance the creation of targeted strategies for specific customer bases. Kevin Kok Wai Wong, Lance Chun Che Fung, Tom Gedeon, Douglas Chai |
ICARCV | 3 |
| 2004 | Managing Interference Between Prior and Later Learning
L. Andrew Coward, Tom Gedeon, Uditha Ratnayake |
ICONIP | 2 |
| 2004 | A generalized concept for fuzzy rule interpolationabstractThe concept of fuzzy rule interpolation in sparse rule bases was introduced in 1993. It has become a widely researched topic in recent years because of its unique merits in the topic of fuzzy rule base complexity reduction. The first implemented technique of fuzzy rule interpolation was termed as /spl alpha/-cut distance based fuzzy rule base interpolation. Despite its advantageous properties in various approximation aspects and in complexity reduction, it was shown that it has some essential deficiencies, for instance, it does not always result in immediately interpretable fuzzy membership functions. This fact inspired researchers to develop various kinds of fuzzy rule interpolation techniques in order to alleviate these deficiencies. This paper is an attempt into this direction. It proposes an interpolation methodology, whose key idea is based on the interpolation of relations instead of interpolating /spl alpha/-cut distances, and which offers a way to derive a family of interpolation methods capable of eliminating some typical deficiencies of fuzzy rule interpolation techniques. The proposed concept of interpolating relations is elaborated here using fuzzy- and semantic-relations. This paper presents numerical examples, in comparison with former approaches, to show the effectiveness of the proposed interpolation methodology. Péter Baranyi, László T. Kóczy, Tom Gedeon |
IEEE Trans. Fuzzy Syst. | 3 |
| 2003 | Sparse fuzzy systems generation and fuzzy rule interpolation: a practical approachabstractIn this paper, we explore the use of a sparse fuzzy system generation technique in conjunction with simple projection-based fuzzy rule interpolation, to generate sparse fuzzy systems with relatively few rules whilst still achieving reasonable system accuracy. Through setting a parameter value, the user is able to control, to some extent, the number of rules generated by the rule extraction technique. The rule interpolation approach enables the sparse fuzzy system to maintain a reasonable accuracy. The effectiveness of this approach is validated experimentally. Alex Chong, Tom Gedeon, Szilveszter Kovács, László T. Kóczy |
FUZZ-IEEE | 2 |
| 2003 | A survey on universal approximation and its limits in soft computing techniques
Domonkos Tikk, László T. Kóczy, Tom Gedeon |
Int. J. Approx. Reason. | 3 |
| 2003 | Rainfall prediction model using soft computing technique
Kevin Kok Wai Wong, Patrick M. Wong, Tom Gedeon, Lance Chun Che Fung |
Soft Comput. | 3 |
| 2002 | Subspace clustering for hierarchical fuzzy system constructionabstractHierarchical fuzzy systems are proposed to deal with the rule explosion problem of traditional fuzzy systems. The inference operations of the fuzzy systems are well established. The next step is to tackle the problem of finding subspaces for automated hierarchical fuzzy system construction. We propose a clustering technique designed specifically for this purpose. It is both theoretically and experimentally confirmed that the algorithm has reasonable accuracy and scalability. Alex Chong, Tom Gedeon |
FUZZ-IEEE | 2 |
| 2002 | Fuzzy tolerance relations and relational maps applied to information retrieval
László T. Kóczy, Tom Gedeon, Judit A. Kóczy |
Fuzzy Sets Syst. | 2 |
| 2002 | Stability of interpolative fuzzy KH controllers
Domonkos Tikk, István Joó, László T. Kóczy, Péter Várlaki, Bernhard Moser 0001, Tom Gedeon |
Fuzzy Sets Syst. | 6 |
| 2002 | Improvements and critique on Sugeno's and Yasukawa's qualitative modelingabstractInvestigates Sugeno's and Yasukawa's (1993) qualitative fuzzy modeling approach. We propose some easily implementable solutions for the unclear details of the original paper, such as trapezoid approximation of membership functions, rule creation from sample data points, and selection of important variables. We further suggest an improved parameter identification algorithm to be applied instead of the original one. These details are crucial concerning the method's performance as it is shown in a comparative analysis and helps to improve the accuracy of the built-up model. Finally, we propose a possible further rule base reduction which can be applied successfully in certain cases. This improvement reduces the time requirement of the method by up to 16% in our experiments. Domonkos Tikk, György Biró, Tom Gedeon, László T. Kóczy, Jae Dong Yang |
IEEE Trans. Fuzzy Syst. | 3 |
| 2002 | Confidence bounds of petrophysical predictions from conventional neural networksabstractNeural networks are powerful tools for solving the complex regression problems which abound in geosciences. Estimation of prediction confidence from neural networks is an important area. Many procedures are available to date, but it is often tedious for practitioners to implement such procedures without significant modification of the existing learning algorithms. In many cases, the procedures are also computationally intensive. This paper presents a practical solution using conventional backpropagation networks with simple data pre-processing and post-processing algorithms. The methodology involves conversion of the target outputs into linguistic variables (classes) prior to learning. When the classification network converges, minimum and maximum predictions are derived from the output activations using a simple averaging algorithm. Two examples from petroleum reservoirs are used to demonstrate the proposed methodology. The results show that the confidence bounds of the petrophysical predictions are realistic in both cases. The proposed methodology is generally useful, and can be implemented in simple spreadsheets without altering any existing neural network code. Patrick M. Wong, Alexander G. Bruce, Tom Gedeon |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2001 | A Histogram-based Rule Extraction Technique for Fuzzy SystemsabstractWe propose a histogram-based rule extraction technique using straightforward histogram-based clustering that produces trapezoidal clusters that are well suited for the rule extraction purpose. Two experiments were carried out to validate the feasibility and effectiveness of the proposed technique and show that the rule base generated by the proposed technique is reasonably accurate. Alex Chong, Tom Gedeon, Kevin Kok Wai Wong, László T. Kóczy |
FUZZ-IEEE | 2 |
| 2001 | Constructing Heirarchical Fuzzy Rule Bases for ClassificationabstractFuzzy rule based systems have been very popular in many control applications. However, when fuzzy control systems are used in real problems, many rules may be required. The number of rules required depends on the number of inputs and the number of fuzzy linguistic terms used. This exponential explosion of fuzzy rules can take too much computing time to solve any but the simplest problems. This paper proposes a hierarchical fuzzy system that partitions a problem for more efficient computation. The hierarchical fuzzy rule base algorithm constructs rules from data for the purpose of performing fuzzy classification. Illustration examples are also generated and the results show that this hierarchical fuzzy system can be successfully used for classification applications. Tom Gedeon, Kevin Kok Wai Wong, Domonkos Tikk |
FUZZ-IEEE | 1 |
| 2001 | Application of the Recommendation Architecture to Telecommunications Network ManagementabstractThe recommendation architecture has been proposed as a system architecture which can enable a system to learn to perform a complex combination of interrelated functions. The capability of a system with the recommendation architecture to learn to manage complex telecommunication backbone networks has been investigated. A network model with a number of nodes and links and carrying realistic but randomly generated traffic was used as the target for the management system. Traffic data taken from the model was used as input to the recommendation architecture system. The traffic data was organized into inputs once every 5 minutes, and the management system organized these inputs into a hierarchy of repetition similarity. It was demonstrated that the outputs of this hierarchy provided information on the condition of the network. This output information was a compressed version of the inputs which correlated with major network conditions. L. Andrew Coward, Tom Gedeon, William Kenworthy |
Int. J. Neural Syst. | 2 |
| 2000 | α-cut interpolation technique in the space of regular conclusionabstractThe first published method of fuzzy rule interpolation is the /spl alpha/-cut distance based fuzzy interpolation. Its stable-approximation property has been deemed important not only in fuzzy theory, but in numerical analysis as well. However, its applicability is restricted by its interpretability in fuzzy theory. Many other fuzzy interpolation methods have been proposed to alleviate these restrictions, but they are rather complicated from a practical point of view. As a result the simplicity of /spl alpha/-cut fuzzy interpolation requiring small computational effort is lost. In this paper, a modified /spl alpha/-cut based fuzzy interpolation method is proposed which eliminates the problem of abnormal conclusion while maintains low computational capacity. Péter Baranyi, Domonkos Tikk, Tom Gedeon, László T. Kóczy |
FUZZ-IEEE | 3 |
| 2000 | On a stable and always applicable interpolation methodabstractThis paper aims to complete the analysis of a our recently proposed /spl alpha/-cut based interpolation technique (1997) originated from the KH interpolation. Our goal is to investigate its stability behaviour. As it was shown in Joo et al. (1997) and Tikk et al. (1997) the original KH interpolation is stable in the sense that if inputs change slightly the output also changes slightly. The main result of this paper shows that this significant feature of the KH interpolation can be carried over for the proposed method. A possible generalization of the proposed method is also presented. Domonkos Tikk, Péter Baranyi, László T. Kóczy, Tom Gedeon |
FUZZ-IEEE | 4 |
| 2000 | Fuzzy Hydrocyclone Modelling for Particle Separation Using Fuzzy Rule Interpolation
Kevin Kok Wai Wong, Lance Chun Che Fung, Tom Gedeon |
IDEAL | 3 |
| 2000 | Academic Articles on the Web: Reading Patterns and FormatsabstractThis article explores user reading activities and user preferences in the formats of Web-based academic articles by using the data from 2 online surveys. Researchers use the Web as a resource for academic articles. Despite this popular use, no generally agreed format exists on the Web. The Web environments of distributed users encourage the use of online remote evaluation. We applied an e-mail-based survey and a Web-based survey to the evaluation of some concepts for Web-based academic articles. The participants of the surveys were researchers in information technology and related areas. Our survey results show that readers take an overview of a Web-based academic article from the screen, print it out, and then read the printed article. The results also show that the formats employed by most of the Web sites for academic articles are against readers' preferences. The simple 2-frame format among the 5 given formats was most preferred by 47% of our respondents, but the cascaded page-windows format was regarded as the worst by 65% because of its high visual complexity on the screen. An interesting result is that 26% of the respondents regarded the paperlike format as the worst, but this format is widely used for Web-based articles. In addition, the importance of interactive examples embedded in a Web-based questionnaire was revealed from the 2 consecutive surveys. Details are discussed in this article. In the online remote surveys, the issues of Web-based academic articles were successfully addressed. The methods used in the surveys would be useful for usability tests of various concepts of other Web genres at an early design or redesign stage. Young J. Rho, Tom Gedeon |
Int. J. Hum. Comput. Interact. | 2 |
| 2000 | Managing heterogeneous information systems through discovery and retrieval of generic conceptsabstractAutonomy of operations combined with decentralized management of data gives rise to a number of heterogeneous databases or information systems within an enterprise. These systems are often incompatible in structure as well as content and, hence, difficult to integrate. Depsite heterogeneity, the unity of overall purpose within a common application domain, nevertheless, provides a degree of semantic similarity that manifests itself in the form of similar data structures and common usage patterns of existing information systems. This article introduces a conceptual integration approach that exploits the similarity in metalevel information in existing systems and performs metadata mining on database objects to discover a set of concepts that serve as a domain abstraction and provide a conceptual layer above existing legacy systems. This conceptual layer is further utilized by an information reengineering framework that customizes and packages information to reflect the unique needs of different user groups within the application domain. The architecture of the information reengineering framework is based on an object-oriented model that represents the discovered concepts as customized application objects for each distinct user group. Uma Srinivasan 0001, Anne H. H. Ngu, Tom Gedeon |
J. Am. Soc. Inf. Sci. | 3 |
| 1999 | A Pattern Adaptive Technique to Handle Data Quality Variation
Patrick M. Wong, Tom Gedeon |
Neural Process. Lett. | 2 |
| 1999 | Reducing the Dimensions of Texture Features for Image Retrieval Using Multi-layer Neural Networks
Jose Antonio Catalan, Jesse S. Jin, Tom Gedeon |
Pattern Anal. Appl. | 3 |
| 1999 | Exploring constructive cascade networksabstractConstructive algorithms have proved to be powerful methods for training feedforward neural networks. An important property of these algorithms is generalization. A series of empirical studies were performed to examine the effect of regularization on generalization in constructive cascade algorithms. It was found that the combination of early stopping and regularization resulted in better generalization than the use of early stopping alone. A cubic penalty term that greatly penalizes large weights was shown to be beneficial for generalization in cascade networks. An adaptive method of setting the regularization magnitude in constructive algorithms was introduced and shown to produce generalization results similar to those obtained with a fixed, user-optimized regularization setting. This adaptive method also resulted in the construction of smaller networks for more complex problems. The acasper algorithm, which incorporates the insights obtained from the empirical studies, was shown to have good generalization and network construction properties. This algorithm was compared to the cascade correlation algorithm on the Proben 1 and additional regression data sets. Nick K. Treadgold, Tom Gedeon |
IEEE Trans. Neural Networks | 2 |
| 1998 | Adaptive Regularization in a Constructive Cascade Network
Nick K. Treadgold, Tom Gedeon |
ICONIP | 2 |
| 1998 | Specialised neural network for learning synonyms and related concepts in large document collectionsabstractFor very large document collections or high volume streams of documents, finding relevant documents is a major information filtering problem. One of the main types of information retrieval systems produces a word frequency measure estimated by some important parts of the document using neural network approaches. This paper reports a new network structure for this task. It is specialised considering the main difficulties of these kinds of applications, namely, the calculation time complexity. It will be pointed out that the calculation, hence, the learning time is much reduced applying the new algorithm, however, the result is significantly improved compared to the former approaches, which offer a possibility to increase the number of considered words, hence, improve the effectiveness of information filtering systems. Péter Baranyi, P. Aradi, László T. Kóczy, Tom Gedeon |
KES (1) | 4 |
| 1998 | Intelligent information retrieval using fuzzy approachabstractOne of the main types of information retrieval systems produces a word frequency measure estimated by some important parts of the document using neural network approaches. This paper reports a fuzzy logic algorithm for this task. It is specialised considering the main difficulties of these kinds of applications, namely, the calculation time complexity. It will be pointed out that the calculation, hence, the learning time is much reduced applying the new algorithm, however, the result is significantly improved compared to the former approaches, which offer a possibility to increase the number of considered words, hence, improve the effectiveness of information filtering systems. Péter Baranyi, Tom Gedeon, László T. Kóczy |
SMC | 2 |
| 1998 | Fuzzy rule base interpolation based on semantic revisionabstractSometimes it is not possible to have a full dense rule base as there are gaps in the information. Furthermore, it is often necessary to deal with a sparse rule base to reduce the size and the inference/control time. In such sparse rule bases classic algorithms such as the CRI of Zadeh (1973) and the Mamdani method do not function for observations hitting gaps between rules. A linear fuzzy rule interpolation technique (KH-interpolation) has been introduced that is suitable for dealing with sparse bases. However, this method often results in conclusions which are not directly interpretable. In this paper an interpolation technique is proposed that is based on the interpolation of the semantics and interrelation of rules. This method guarantees the direct interpretability of the conclusion. Péter Baranyi, Sándor Mizik, László T. Kóczy, Tom Gedeon, István Nagy 0001 |
SMC | 4 |
| 1998 | Stochastic bidirectional trainingabstractWe consider connectionist compression schemes using auto-associative networks, demonstrate the advantages gained by imposing different constraints on allowed network weights, and give a comparison with pruning of the unconstrained auto-associative network. In this paper we demonstrate the advantages for generalisation performance of constraining weights symmetrically using weight sharing, and by constraining functional symmetry by the use of enhanced backpropagation networks trained bidirectionally. In the process, we derive the stochastic bidirectional training algorithm. Tom Gedeon |
SMC | 1 |
| 1998 | Hierarchical co-occurence relationsabstractWe introduce a method using fuzzy similarity (equivalence) and tolerance (compatibility) relations, that allows the "concentric" extension of searches based on the hierarchical co-occurrence of words and phrases. This is to solve the problem of automatic indexing and retrieval of documents where user queries may not include any words occurring in the documents that should be retrieved. Various methods are proposed and illustrated, with the intention of real application in legal document collections. Tom Gedeon, László T. Kóczy |
SMC | 1 |
| 1998 | Increased generalization through selective decay in a constructive cascade networkabstractDetermining the optimum amount of regularization to obtain the best generalization performance in feedforward neural networks is a difficult problem, and is a form of the bias-variance dilemma. This problem is addressed in the CasPer algorithm, a constructive cascade algorithm that uses weight decay. Previously the amount of weight decay used by this algorithm was set by a parameter prior to training, often by trial and error. This is overcome through the use of a pool of neurons which are candidates for insertion into the network. Each neuron in the pool has an associated decay level, and the one which produces the best generalization on a validation set is added to the network. This not only removes the need for the user to select a decay value, but results in better generalization compared to networks with fixed, user optimized, decay values. Nick K. Treadgold, Tom Gedeon |
SMC | 2 |
| 1998 | Simulated annealing and weight decay in adaptive learning: the SARPROP algorithmabstractA problem with gradient descent algorithms is that they can converge to poorly performing local minima. Global optimization algorithms address this problem, but at the cost of greatly increased training times. This work examines combining gradient descent with the global optimization technique of simulated annealing (SA). Simulated annealing in the form of noise and weight decay is added to resiliant backpropagation (RPROP), a powerful gradient descent algorithm for training feedforward neural networks. The resulting algorithm, SARPROP, is shown through various simulations not only to be able to escape local minima, but is also able to maintain, and often improve the training times of the RPROP algorithm. In addition, SARPROP may be used with a restart training phase which allows a more thorough search of the error surface and provides an automatic annealing schedule. Nick K. Treadgold, Tom Gedeon |
IEEE Trans. Neural Networks | 2 |
| 1997 | Genetic Algorithms Applied to University Exam Scheduling
N. Jenkins, Tom Gedeon |
ICONIP (2) | 2 |
| 1997 | Extending CasPer: A Regression Survey
Nick K. Treadgold, Tom Gedeon |
ICONIP (1) | 2 |
| 1997 | Data Mining of Inputs: Analysing Magnitude and Functional MeasuresabstractThe problem of data encoding and feature selection for training back-propagation neural networks is well known. The basic principles are to avoid encrypting the underlying structure of the data, and to avoid using irrelevant inputs. This is not easy in the real world, where we often receive data which has been processed by at least one previous user. The data may contain too many instances of some class, and too few instances of other classes. Real data sets often include many irrelevant or redundant input fields. This paper examines the use of weight matrix analysis techniques and functional measures using two real (and hence noisy) data sets. The first part of this paper examines the use of the weight matrix of the trained neural network itself to determine which inputs are significant. A new technique is introduced and compared with two other techniques from the literature. We present our experience and results on some satellite data augmented by a terrain model. The task was to predict the forest supra-type based on the available information. A brute force technique eliminating randomly selected inputs was used to validate our approach. The second part of this paper examines the use of measures to determine the functional contribution of inputs to outputs. Inputs which include minor but unique information to the network are more significant than inputs with higher magnitude contribution but providing redundant information, which is also provided by another input. A comparison is made to sensitivity analysis, where the sensitivity of outputs to input perturbation is used as a measure of the significance of inputs. This paper presents a novel functional analysis of the weight matrix based on a technique developed for determining the behavioral significance of hidden neurons. This is compared with the application of the same technique to the training and test data. Finally, a novel aggregation technique is introduced. Tom Gedeon |
Int. J. Neural Syst. | 1 |
| 1996 | The Cyclic Towers of Hanoi: An Iterative Solution Produced by TransformationabstractAn iterative solution to the Cyclic Towers of Hanoi puzzle is produced by largely automatic program transformation from a recursive solution. The result compares favourably with the best published, manually produced iterative algorithm, both in terms of comprehensibility in its own right, and in efficiency. Tom Gedeon |
Comput. J. | 1 |
| 1995 | An improved technique in porosity prediction: a neural network approachabstractGenetic reservoir characterization is important in developing, for a given petroleum reservoir, an improved understanding of the total amount and fluid flow properties of hydrocarbon reserves. Application of genetic concepts involves the classification of well log data into different lithofacies groups, followed by a facies-by-facies description of rock properties such as porosity and permeability. This work contrasts the genetic and nongenetic approaches in predicting porosity values of an oil well using backpropagation neural network methods. The performance of both methods are critically evaluated. A systematic technique to optimise the network configuration using weight visualization curves is proposed, thereby enabling the amount of training time to be significantly reduced. In the example problem, the genetic approach provides superior porosity estimates to that based on a nongenetic approach.> Patrick M. Wong, Tom Gedeon, Ian J. Taggart |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 1992 | The Reve's Puzzle: An Iterative Solution Produced by TransformationabstractAn iterative solution to the Reve's puzzle is produced by largely automatic program transformation from a recursive solution. The result compares favourably with published, manually produced iterative algorithms, both in terms of comprehensibility in its own right, and efficiency. Tom Gedeon |
Comput. J. | 1 |
| 1986 | The Reve's PuzzleabstractThe Towers of Hanoi problem, explored in many recent Journal papers, is extended to consider a related problem, The Reve's Puzzle. It is shown that The Reve's Puzzle is a generalisation of the Towers of Hanoi problem and a number of questions concerning the elegance of the algorithms are posed. Jeffrey S. Rohl, Tom Gedeon |
Comput. J. | 2 |