EDBT 2026 Demo / reviewers in the wild / expert
Gary K. L. Tam
dblp:68/5808 · also Gary Kwok-Leung Tam
· DBLP profile ↗
36ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-7387-5180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | R²D-LPCC: Relevance-Ranking Guided Region-Adaptive Dynamic LiDAR Point Cloud CompressionabstractDynamic LiDAR point cloud compression (LPCC) is crucial for the efficient transmission and storage of large-scale three-dimensional data in applications such as autonomous driving. However, many existing methods, which primarily focus on compressing geometric or motion information, face a fundamental limitation: they treat all points as equally important. This approach neglects the semantic priorities of a scene, resulting in inefficient bit allocation and particularly compromising the reconstruction quality of safety-critical regions, such as pedestrians and vehicles, which are vital to downstream perception tasks. To address these limitations, we propose R²D-LPCC, a relevance-ranking framework for region adaptive LPCC that prioritizes fidelity in semantically important regions. Central to our approach is the Adaptive Relevance Learning (ARL) module, which integrates semantic context with uncertainty to evaluate regional significance and guide compression. We also introduce a Multi-scale Region-Adaptive Transform (MRAT) module to enhance semantic feature modeling and preserve fine-grained details in key areas. Additionally, we develop an adaptive multi-modal motion estimation module to improve motion prediction in complex three-dimensional environments. Extensive experiments conducted on the SemanticKITTI benchmark demonstrate that R²D-LPCC significantly surpasses ten recent state-of-the-art methods, achieving a 45.48% BD-rate gain over the previous leading method, Unicorn, and a 98.58% gain over the GPCC standard, while ensuring superior reconstruction quality in semantically important regions. Fangzhe Nan, Frederick W. B. Li, Gary K. L. Tam, Zhaoyi Jiang, Bailin Yang, Jingke Cui, Changshuo Wang 0001 |
AAAI | 3 |
| 2026 | Video Mirror Detection with the Motion-in-Depth CueabstractDetecting mirror regions in RGB videos is essential for scene understanding in applications such as scene reconstruction and robotic navigation. Existing video mirror detectors typically rely on cues like inside-outside mirror correspondences and 2D motion inconsistencies. However, these methods often yield noisy or incomplete predictions when confronted with complex real-world video scenes, especially in areas with occlusion or limited visual features and motions. We observe that human perceive and navigate 3D occluded environments with remarkable ease, owing to Motion-in-Depth (MiD) perception. MiD integrates information from visual appearance (image colors and textures), the way objects move around us in 3D space (3D motions), and their relative distance from us (depth) to determine if something is approaching or receding and to support navigation. Motivated by this neuroscience mechanism, we introduce MiD-VMD, the first approach to explicitly model MiD for video mirror detection. MiD-VMD jointly utilizes contrastive 3D motion, depth, and image features through two novel modules based on a combinational QKV transformer architecture. The Motion-in-Depth Attention Learning (MiD-AL) module captures complementary relationships across these modalities with combinatorial attention and enforces a compact encoding to represent global 3D transformations, resulting in more accurate mirror detection and reduced motion artifacts. The Motion-in-Depth Boundary Detection (MiD-BD) module further sharpens mirror boundaries by leveraging cross-modal attention on 3D motion and depth features. Extensive experiments show that MiD-VMD outperforms current SOTAs. Alex Warren, Ke Xu 0010, Xin Tian 0015, Gary K. L. Tam, Benjamin W. Wah, Rynson W. H. Lau |
AAAI | 4 |
| 2026 | OctMamba: Mamba-based octree context entropy model for point cloud geometry compressionabstractExisting learned point cloud compression frameworks face two major limitations: (1) they focus almost exclusively on spatial redundancy and (2) rely on architectures built around local-global transformers or global Mamba blocks. Transformers incur quadratic complexity, while global Mamba lacks the granularity to capture structured correlations across multiple dimensions. We propose OctMamba, the first unified framework to jointly exploit spatial, channel, and topological redundancies, dimensions previously overlooked in point cloud geometry compression. Our approach introduces a new architectural principle: embedding Mamba modules within specialized subcomponents rather than applying them globally, challenging existing design paradigms. OctMamba combines two modules: Spatial-Channel Coupled Grouping Mamba (SCCGM) for spatial-channel fusion and Local Graph CNN-Mamba (LGCM) for topological encoding. This design enables efficient long-range modeling with linear complexity, delivering a smaller model and faster decoding while outperforming transformer-based and global Mamba baselines. On SemanticKITTI, OctMamba reduces bitrate by 60.2% over GPCC (D1 PSNR) and achieves state-of-the-art performance across LiDAR and dynamic human point cloud benchmarks with practical speed and scalability. By introducing multi-dimensional redundancy modeling, OctMamba has the potential to influence future research on efficient point cloud compression. The code is available at https://github.com/ZjgsVMC/OctMamba . Zhaoyi Jiang, Frederick W. B. Li, Gary K. L. Tam, Chao Song 0001, Bailin Yang |
Pattern Recognit. | 4 |
| 2026 | Robust and quality preserving 4D data watermarking for copyright protectionabstractThe use of 4D (3D dynamic mesh sequences) data, such as rendering objects and faces, has grown significantly with advancements in imaging, 3D printing, and V/AR technologies. However, there is currently a lack of 4D watermarking methods to protect copyright and prevent unauthorized distribution. Existing 3D mesh watermarking techniques are inadequate for 4D data, as they fail to address dynamic attacks, such as sequence reordering. To fill this gap, we propose a novel 4D watermarking technique for 4D face data, comprising two main components: watermarking for 3D mesh sequences and watermarking for texture images. Each component uses two watermarks: one to protect individual mesh/texture copyright and another to verify sequence order and detect reordering attacks. The Artificial Bee Colony algorithm optimizes embedding parameters to ensure watermark imperceptibility. Experimental results demonstrate the robustness and imperceptibility of our method, validating its effectiveness in protecting 4D data from unauthorized use and distribution. Ertugrul Gul, Ahmet Nusret Toprak, Gary K. L. Tam |
Vis. Comput. | 3 |
| 2025 | CymruFluency - A Fusion Technique and a 4D Welsh Dataset for Welsh Fluency AnalysisabstractWelsh is a linguistically rich yet under-resourced minority language. Despite its cultural significance, automated fluency assessment remains largely unexplored due to limited datasets and tools. Existing models focus on high-resource languages, leaving Welsh without sufficient multi-modal resources. To address this, we introduce CymruFluency, the first 4D dataset for Welsh fluency assessment, capturing both audio and 3D lip movements with expert-annotated fluency scores. Building on this, we propose a multi-modal fluency classification framework that combines audio features (mel spectrograms) and manually annotated 3D lip landmarks. Our fusion approach significantly improves fluency prediction over unimodal models, emphasizing the critical role of 3D lip dynamics in Welsh learning. This research advances minority language processing by integrating articulatory features into fluency evaluation, offering a powerful tool for Welsh language learning, assessment, and preservation. Project page: https://github.com/arvinsingh/CymruFluency . Arvinder Pal Singh Bali, Gary K. L. Tam, Avishek Siris, Gareth Andrews, Yukun Lai, Bernard Tiddeman, Gwenno Ffrancon |
ACIVS | 2 |
| 2025 | Pretraining Techniques for Steel Surface Roughness Prediction with Long Thin Spatial Industrial DataabstractMachine learning offers promising advancements in industrial processes, yet collecting labeled samples during production remains challenging. In steel production, the surface roughness $$R_a$$ parameter of steel coils is crucial, but on-line labeled data collection, with our apparatus, is infeasible, while off-line methods are time-consuming and imperfect. However, unlabeled samples are readily available from on-line production. This paper examines pretraining on a large, unlabeled dataset and its impacts on performance after fine-tuning on a smaller labeled dataset. We use three techniques: (1) contrastive learning, (2) Autoencoder, and (3) Classification of coil ID. We address the challenges posed by the unique structure of the data, comprising 2-dimensional, long and thin arrays. Our results show that our classification pretraining approach improves regression performance and outperforms the baseline. Alexander J. M. Milne, Xianghua Xie, Gary K. L. Tam |
ACIVS | 3 |
| 2025 | Multi-modal Dynamic Point Cloud Geometric Compression Based on Bidirectional Recurrent Scene FlowabstractDeep learning methods have recently shown significant promise in compressing the geometric features of point clouds. However, challenges arise when consecutive point clouds contain holes, resulting in incomplete information that complicates motion estimation. To our knowledge, most existing dynamic point cloud compression methods have largely overlooked this critical issue. Moreover, these methods typically employ a multi-scale single-pass approach for motion estimation, performing only one estimation at each scale. This limits accuracy and adversely impacts compression performance. To address these challenges, we propose a dynamic point cloud compression model called M2BR-DPCC (Multi-Modal Multi-Scale Bidirectional Recursion for Dynamic Point Cloud Compression). Our method introduces two key innovations. First, we integrate both point cloud and image data as inputs, leveraging a multi-modal feature representation completion (MFRepC) approach to align information across modalities. This addresses the issue of missing data in point clouds by using complementary information from images. Second, we implement a multi-scale bidirectional recursive (MSBR) motion estimation method. This module iteratively refines motion flows in both forward and backward directions, progressively enhancing point cloud features and improving motion estimation accuracy. Experimental results on widely used datasets, including MVUB and 8iVFB, demonstrate the effectiveness of our approach. Compared to existing methods, M2BR-DPCC achieves superior performance, with an average BD-rate improvement of 95.23% over V-PCC, 12.92% over D-DPCC, and 16.16% over patchDPCC. These results underscore the potential of leveraging multi-modal data and bidirectional refinement for dynamic point cloud compression. Fangzhe Nan, Frederick W. B. Li, Zhuoyue Wang, Gary K. L. Tam, Zhaoyi Jiang, DongZheng DongZheng, Bailin Yang |
ICASSP | 4 |
| 2025 | Denoising-While-Completing Network (DWCNet): Robust point cloud completion under corruptionabstractPoint cloud completion is crucial for 3D computer vision tasks in autonomous driving, augmented reality, and robotics. However, obtaining clean and complete point clouds from real-world environments is challenging due to noise and occlusions. Consequently, most existing completion networks – trained on synthetic data – struggle with real-world degradations. In this work, we tackle the problem of completing and denoising highly corrupted partial point clouds affected by multiple simultaneous degradations. To benchmark robustness, we introduce the Corrupted Point Cloud Completion Dataset (CPCCD), which highlights the limitations of current methods under diverse corruptions. Building on these insights, we propose DWCNet (Denoising-While-Completing Network), a completion framework enhanced with a Noise Management Module (NMM) that leverages contrastive learning and self-attention to suppress noise and model structural relationships. DWCNet achieves state-of-the-art performance on both clean and corrupted, synthetic and real-world datasets. The dataset and code will be publicly available at https://github.com/keneniwt/DWCNET-Robust-Point-Cloud-Completion-against-Corruptions . Keneni W. Tesema, Lyndon Hill, Mark W. Jones 0001, Gary K. L. Tam |
Comput. Graph. | 4 |
| 2025 | Talking Face Generation With Lip and Identity PriorsabstractABSTRACT Speech‐driven talking face video generation has attracted growing interest in recent research. While person‐specific approaches yield high‐fidelity results, they require extensive training data from each individual speaker. In contrast, general‐purpose methods often struggle with accurate lip synchronization, identity preservation, and natural facial movements. To address these limitations, we propose a novel architecture that combines an alignment model with a rendering model. The rendering model synthesizes identity‐consistent lip movements by leveraging facial landmarks derived from speech, a partially occluded target face, multi‐reference lip features, and the input audio. Concurrently, the alignment model estimates optical flow using the occluded face and a static reference image, enabling precise alignment of facial poses and lip shapes. This collaborative design enhances the rendering process, resulting in more realistic and identity‐preserving outputs. Extensive experiments demonstrate that our method significantly improves lip synchronization and identity retention, establishing a new benchmark in talking face video generation. Frederick W. B. Li, Gary K. L. Tam, Bailin Yang, Fangzhe Nan, Jia Pan 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2025 | Survey: 3D watermarking techniquesabstractIn today’s world, 3D multimedia data is widely utilized in diverse fields such as military, medical, and remote sensing. The advancement of multimedia technologies, however, exposes 3D multimedia content to an increasing risk of malicious interventions. It has become highly essential to implement security measures ensuring the authenticity and copyright protection of 3D multimedia content. Watermarking is considered one of the most reliable and practical approaches for this purpose. This work provides an in-depth and up-to-date overview of various 3D watermarking methods, covering different data forms, including 3D images, 3D videos, 3D meshes, point clouds, and NeRF. We have categorized these methods from multiple perspectives, comparing their respective advantages and disadvantages. The study also identifies attacks based on data type and discusses metrics for evaluating the methods according to their intended use and data type. We further present observations, research issues, challenges, and future directions for 3D watermarking. This includes strength factor optimization, copyright concerns related to 3D printed objects, detection and recovery of tampered areas, and watermarking in 4D (3D Dynamic) and NeRF domains. Ertugrul Gul, Gary K. L. Tam |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Survey on 3D Reconstruction Techniques: Large-Scale Urban City Reconstruction and Requirementsabstract3D representations of large-scale and urban scenes are crucial across various industries, including autonomous driving, urban planning, natural resource supervision and many more. Large-scale industrial reconstructions are inherently complex and multifaceted. Many existing surveys primarily focus on academic progressions and often neglect the intricate and diverse needs of industry. This survey aims to bridge this gap by providing a comprehensive analysis of 3D reconstruction methods, with a focus on industrial requirements such as scalability and integration of human interaction. Our approach involves utilizing Affinity Diagramming to systematically analyze qualitative data gathered from industrial partners. This methodology enables us to gain deep insights into how recent literature addresses these specific industrial needs. The survey encompasses various aspects, including input and reconstruction modalities, architectural models, datasets, evaluation metrics, and the incorporation of prior knowledge. We further discuss practical implications derived from our analysis, highlighting key considerations for future advancements in 3D reconstruction methods tailored for large-scale applications. Andreas Christodoulides, Gary K. L. Tam, James Clarke, Jon Horgan, Nicholas Micallef, Jeremy G. Morley, Nelly Villamizar, Sean Peter Walton |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Effective Video Mirror Detection with Inconsistent Motion CuesabstractImage-based mirror detection has recently undergone rapid research due to its significance in applications such as robotic navigation, semantic segmentation and scene re-construction. Recently, VMD-Net was proposed as the first video mirror detection technique, by modeling dual correspondences between the inside and outside of the mirror both spatially and temporally. However, this approach is not reliable, as correspondences can occur completely inside or outside of the mirrors. In addition, the proposed dataset VMD-D contains many small mirrors, limiting its applicability to real-world scenarios. To address these problems, we developed a more challenging dataset that includes mirrors of various shapes and sizes at different locations of the frames, providing a better reflection of real-world scenarios. Next, we observed that the motions between the inside and outside of the mirror are often in-consistent. For instance, when moving in front of a mirror, the motion inside the mirror is often much smaller than the motion outside due to increased depth perception. With these observations, we propose modeling inconsistent motion cues to detect mirrors, and a new network with two novel modules. The Motion Attention Module (MAM) ex-plicitly models inconsistent motions around mirrors via optical flow, and the Motion-Guided Edge Detection Module (MEDM) uses motions to guide mirror edge feature learning. Experimental results on our proposed dataset show that our method outperforms state-of-the-arts. The code and dataset are available at ht tps: // gi th ub. com/ AlexAnthonyWarren/MG-VMD. Alex Warren, Ke Xu 0010, Jiaying Lin 0001, Gary K. L. Tam, Rynson W. H. Lau |
CVPR | 4 |
| 2024 | Inferring Attention Shifts for Salient Instance RankingabstractAbstract The human visual system has limited capacity in simultaneously processing multiple visual inputs. Consequently, humans rely on shifting their attention from one location to another. When viewing an image of complex scenes, psychology studies and behavioural observations show that humans prioritise and sequentially shift attention among multiple visual stimuli. In this paper, we propose to predict the saliency rank of multiple objects by inferring human attention shift. We first construct a new large-scale salient object ranking dataset, with the saliency rank of objects defined by the order that an observer attends to these objects via attention shift. We then propose a new deep learning-based model to leverage both bottom-up and top-down attention mechanisms for saliency rank prediction. Our model includes three novel modules: Spatial Mask Module (SMM), Selective Attention Module (SAM) and Salient Instance Edge Module (SIEM). SMM integrates bottom-up and semantic object properties to enhance contextual object features, from which SAM learns the dependencies between object features and image features for saliency reasoning. SIEM is designed to improve segmentation of salient objects, which helps further improve their rank predictions. Experimental results show that our proposed network achieves state-of-the-art performances on the salient object ranking task across multiple datasets. Code and data are available at https://github.com/SirisAvishek/Attention_Shift_Ranks . Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
Int. J. Comput. Vis. | 3 |
| 2024 | Point Cloud Completion: A SurveyabstractPoint cloud completion is the task of producing a complete 3D shape given an input of a partial point cloud. It has become a vital process in 3D computer graphics, vision and applications such as autonomous driving, robotics, and augmented reality. These applications often rely on the presence of a complete 3D representation of the environment. Over the past few years, many completion algorithms have been proposed and a substantial amount of research has been carried out. However, there are not many in-depth surveys that summarise the research progress in such a way that allows users to make an informed choice of what algorithms to employ given the type of data they have, the end result they want, the challenges they may face and the possible strategies they could use. In this study, we present a comprehensive survey and classification of articles on point cloud completion untill August 2023 based on the strategies, techniques, inputs, outputs, and network architectures. We will also cover datasets, evaluation methods, and application areas in point cloud completion. Finally, we discuss challenges faced by the research community and future research directions. Keneni W. Tesema, Lyndon Hill, Mark W. Jones 0001, Muneeb Imtiaz Ahmad, Gary K. L. Tam |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | C2SPoint: A classification-to-saliency network for point cloud saliency detectionabstractPoint cloud saliency detection is an important technique that support downstream tasks in 3D graphics and vision, like 3D model simplification, compression, reconstruction and viewpoint selection. Existing approaches often rely on hand-crafted features and are only applicable to specific datasets. In this paper, we propose a novel weakly supervised classification network, called C2SPoint, which directly performs saliency detection on the point clouds. Unlike previous methods that require per-point saliency annotations, C2SPoint only requires category labels of the point clouds during training. The network consists of two branches: a Classification branch and a Saliency branch. The former branch is composed of two Adaptive Set Abstraction layers for feature extraction and a Saliency Transform layer for learning saliency knowledge from the classification network. The latter branch introduces a multi-scale point-cluster similarity matrix for propagating the cluster saliency to each point within it, resulting in the prediction of point-level saliency. Experimental results demonstrate the effectiveness of our method in point cloud saliency detection, with improvements of 2% in both AUC and NSS compared to state-of-the-art methods. Zhaoyi Jiang, Luyun Ding, Gary K. L. Tam, Chao Song 0001, Frederick W. B. Li, Bailin Yang |
Comput. Graph. | 3 |
| 2022 | A Deep Learning Driven Active Framework for Segmentation of Large 3D Shape Collections
David George 0001, Xianghua Xie, Yukun Lai, Gary K. L. Tam |
Comput. Aided Des. | 4 |
| 2021 | Scene Context-Aware Salient Object DetectionabstractSalient object detection identifies objects in an image that grab visual attention. Although contextual features are considered in recent literature, they often fail in real-world complex scenarios. We observe that this is mainly due to two issues: First, most existing datasets consist of simple foregrounds and backgrounds that hardly represent real-life scenarios. Second, current methods only learn contextual features of salient objects, which are insufficient to model high-level semantics for saliency reasoning in complex scenes. To address these problems, we first construct a new large-scale dataset with complex scenes in this paper. We then propose a context-aware learning approach to explicitly exploit the semantic scene contexts. Specifically, two modules are proposed to achieve the goal: 1) a Semantic Scene Context Refinement module to enhance contextual features learned from salient objects with scene context, and 2) a Contextual Instance Transformer to learn contextual relations between objects and scene context. To our knowledge, such high-level semantic contextual information of image scenes is under-explored for saliency detection in the literature. Extensive experiments demonstrate that the proposed approach outperforms state-of-the-art techniques in complex scenarios for saliency detection, and transfers well to other existing datasets. The code and dataset are available at https://github.com/SirisAvishek/Scene_Context_Aware_Saliency. Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
ICCV | 3 |
| 2020 | Inferring Attention Shift Ranks of Objects for Image SaliencyabstractPsychology studies and behavioural observation show that humans shift their attention from one location to another when viewing an image of a complex scene. This is due to the limited capacity of the human visual system in simultaneously processing multiple visual inputs. The sequential shifting of attention on objects in a non-task oriented viewing can be seen as a form of saliency ranking. Although there are methods proposed for predicting saliency rank, they are not able to model this human attention shift well, as they are primarily based on ranking saliency values from binary prediction. Following psychological studies, in this paper, we propose to predict the saliency rank by inferring human attention shift. Due to the lack of such data, we first construct a large-scale salient object ranking dataset. The saliency rank of objects is defined by the order that an observer attends to these objects based on attention shift. The final saliency rank is an average across the saliency ranks of multiple observers. We then propose a learning-based CNN to leverage both bottom-up and top-down attention mechanisms to predict the saliency rank. Experimental results show that the proposed network achieves state-of-the-art performances on salient object rank prediction. Code and dataset are available at https://github.com/SirisAvishek/Attention_Shift_Ranks. Avishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie, Rynson W. H. Lau |
CVPR | 3 |
| 2020 | Graph convolutional neural network for multi-scale feature learning
Michael Edwards, Xianghua Xie, Robert Ieuan Palmer, Gary K. L. Tam, Rob Alcock, Carl Roobottom |
Comput. Vis. Image Underst. | 4 |
| 2019 | Non-rigid registration under anisotropic deformationsabstractNon-rigid registration of deformed 3D shapes is a challenging and fundamental task in geometric processing, which aims to non-rigidly deform a source shape into alignment with a target shape. Current state-of-the-art methods assume deformations to be near-isometric. This assumption does not reflect real-world conditions, for example in large-scale deformation, where moderate anisotropic deformations (e.g., stretches) are common. In this paper we propose two significant changes to a typical registration pipeline to address such challenging deformations. First, we introduce a method to estimate anisotropic non-isometric deformations and incorporate this into an iterative non-rigid registration pipeline. Second, we compute additional correspondences in non-isometrically deforming regions using reliable correspondences as landmarks and prune inconsistent correspondences. We compare the performance of our proposed algorithm to several state-of-the-art methods using existing benchmarks. Experimental results show that our method outperforms existing methods. Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Gary K. L. Tam |
Comput. Aided Geom. Des. | 4 |
| 2019 | Consistent segment-wise matching with multi-layer graphs
Taiwei Wang, David George 0001, Yukun Lai, Xianghua Xie, Gary K. L. Tam |
Comput. Aided Geom. Des. | 5 |
| 2018 | 3D mesh segmentation via multi-branch 1D convolutional neural networks
David George 0001, Xianghua Xie, Gary K. L. Tam |
Graph. Model. | 3 |
| 2017 | Recognition, Tracking, and Optimisation
Xianghua Xie, Mark W. Jones 0001, Gary K. L. Tam |
Int. J. Comput. Vis. | 3 |
| 2017 | An Analysis of Machine- and Human-Analytics in ClassificationabstractIn this work, we present a study that traces the technical and cognitive processes in two visual analytics applications to a common theoretic model of soft knowledge that may be added into a visual analytics process for constructing a decision-tree model. Both case studies involved the development of classification models based on the "bag of features" approach. Both compared a visual analytics approach using parallel coordinates with a machine-learning approach using information theory. Both found that the visual analytics approach had some advantages over the machine learning approach, especially when sparse datasets were used as the ground truth. We examine various possible factors that may have contributed to such advantages, and collect empirical evidence for supporting the observation and reasoning of these factors. We propose an information-theoretic model as a common theoretic basis to explain the phenomena exhibited in these two case studies. Together we provide interconnected empirical and theoretical evidence to support the usefulness of visual analytics. Gary K. L. Tam, Vivek Kothari, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Shape Retrieval of Non-rigid 3D Human Modelsabstract3D models of humans are commonly used within computer graphics and vision, and so the ability to distinguish between body shapes is an important shape retrieval problem. We extend our recent paper which provided a benchmark for testing non-rigid 3D shape retrieval algorithms on 3D human models. This benchmark provided a far stricter challenge than previous shape benchmarks. We have added 145 new models for use as a separate training set, in order to standardise the training data used and provide a fairer comparison. We have also included experiments with the FAUST dataset of human scans. All participants of the previous benchmark study have taken part in the new tests reported here, many providing updated results using the new data. In addition, further participants have also taken part, and we provide extra analysis of the retrieval results. A total of 25 different shape retrieval methods are compared. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhouhui Lian, Masaki Aono, A. Ben Hamza, Alexander M. Bronstein, Michael M. Bronstein, S. Bu, Umberto Castellani, S. Cheng, Valeria Garro, Andrea Giachetti 0001, Afzal Godil, Luca Isaia, Henry Johan, Long Lai, Bo Li 0013, Chenfeng Li, Hai-Sheng Li 0002, Roee Litman, Yijuan Lu, Li Sun 0004, Gary K. L. Tam, Atsushi Tatsuma, Jianbo Ye |
Int. J. Comput. Vis. | 29 |
| 2015 | Automatic Aortic Root Segmentation with Shape Constraints and Mesh RegularisationabstractFully automated 3D segmentation is not only challenging due to, for instance, ambiguities in appearance, but it is also computationally demanding.We present a fullyautomatic, learning-based deformable modelling method for segmenting the aortic root in CT images using a two-stage mesh deformation: a non-iterative boundary segmentation with a statistical shape model for shape constraint, followed by an iterative boundary refinement process.At both stages, we introduce a B-spline mesh regularisation technique to avoid mesh entanglement during deformation.The initialisation of the deformable model is achieved through efficient detection and localisation of the aortic root using marginal space learning, which carries out similarity parameter estimation in an incremental fashion.Quantitative comparisons are carried out against a state-of-the-art deformable model-based approach and an active shape model based segmentation.The proposed method achieves both a lower average mesh error of 1.39 ± 0.29mm, and Hausdorff distance of 6.75 ± 2.05mm.Compared to these two approaches, it results in much more regularised mesh surfaces with no tangled mesh faces. Robert Ieuan Palmer, Xianghua Xie, Gary K. L. Tam |
BMVC | 3 |
| 2014 | An Efficient Approach to Correspondences between Multiple Non-Rigid PartsabstractAbstract Identifying multiple deformable parts on meshes and establishing dense correspondences between them are tasks of fundamental importance to computer graphics, with applications to e.g. geometric edit propagation and texture transfer. Much research has considered establishing correspondences between non‐rigid surfaces, but little work can both identify similar multiple deformable partsandhandle partial shape correspondences. This paper addresses two related problems, treating them as a whole: (i) identifying similar deformable parts on a mesh, related by anon‐rigidtransformation to a given query part, and (ii) establishing dense point correspondences automatically between such parts. We show that simple and efficient techniques can be developed if we make the assumption that these parts locally undergo isometric deformation. Our insight is that similar deformable parts are suggested by large clusters of point correspondences that are isometrically consistent. Once such parts are identified,densepoint correspondences can be obtained by an iterative propagation process. Our techniques are applicable to models with arbitrary topology. Various examples demonstrate the effectiveness of our techniques. Gary K. L. Tam, Ralph R. Martin, Paul L. Rosin, Yukun Lai |
Comput. Graph. Forum | 1 |
| 2014 | Facial expression recognition in dynamic sequences: An integrated approach
Hui Fang 0003, Neil Mac Parthaláin, Andrew J. Aubrey, Gary K. L. Tam, Rita Borgo, Paul L. Rosin, Phil W. Grant, David Marshall 0001, Min Chen 0001 |
Pattern Recognit. | 4 |
| 2014 | Diffusion pruning for rapidly and robustly selecting global correspondences using local isometryabstractFinding correspondences between two surfaces is a fundamental operation in various applications in computer graphics and related fields. Candidate correspondences can be found by matching local signatures, but as they only consider local geometry, many are globally inconsistent. We provide a novel algorithm to prune a set of candidate correspondences to those most likely to be globally consistent. Our approach can handle articulated surfaces, and ones related by a deformation which is globally nonisometric, provided that the deformation is locally approximately isometric. Our approach uses an efficient diffusion framework, and only requires geodesic distance calculations in small neighbourhoods, unlike many existing techniques which require computation of global geodesic distances. We demonstrate that, for typical examples, our approach provides significant improvements in accuracy, yet also reduces time and memory costs by a factor of several hundred compared to existing pruning techniques. Our method is furthermore insensitive to holes, unlike many other methods. Gary K. L. Tam, Ralph R. Martin, Paul L. Rosin, Yukun Lai |
ACM Trans. Graph. | 1 |
| 2013 | Visualizing Natural Image StatisticsabstractNatural image statistics is an important area of research in cognitive sciences and computer vision. Visualization of statistical results can help identify clusters and anomalies as well as analyze deviation, distribution, and correlation. Furthermore, they can provide visual abstractions and symbolism for categorized data. In this paper, we begin our study of visualization of image statistics by considering visual representations of power spectra, which are commonly used to visualize different categories of images. We show that they convey a limited amount of statistical information about image categories and their support for analytical tasks is ineffective. We then introduce several new visual representations, which convey different or more information about image statistics. We apply ANOVA to the image statistics to help select statistically more meaningful measurements in our design process. A task-based user evaluation was carried out to compare the new visual representations with the conventional power spectra plots. Based on the results of the evaluation, we made further improvement of visualizations by introducing composite visual representations of image statistics. Hui Fang 0003, Gary K. L. Tam, Rita Borgo, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, Christian Wallraven, Douglas W. Cunningham, David Marshall 0001, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Registration of 3D Point Clouds and Meshes: A Survey from Rigid to NonrigidabstractThree-dimensional surface registration transforms multiple three-dimensional data sets into the same coordinate system so as to align overlapping components of these sets. Recent surveys have covered different aspects of either rigid or nonrigid registration, but seldom discuss them as a whole. Our study serves two purposes: 1) To give a comprehensive survey of both types of registration, focusing on three-dimensional point clouds and meshes and 2) to provide a better understanding of registration from the perspective of data fitting. Registration is closely related to data fitting in which it comprises three core interwoven components: model selection, correspondences and constraints, and optimization. Study of these components 1) provides a basis for comparison of the novelties of different techniques, 2) reveals the similarity of rigid and nonrigid registration in terms of problem representations, and 3) shows how overfitting arises in nonrigid registration and the reasons for increasing interest in intrinsic techniques. We further summarize some practical issues of registration which include initializations and evaluations, and discuss some of our own observations, insights and foreseeable research trends. Gary K. L. Tam, Zhi-Quan Cheng, Yukun Lai, Frank C. Langbein, Yonghuai Liu, David Marshall 0001, Ralph R. Martin, Xianfang Sun, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | Embedding Retrieval of Articulated Geometry ModelsabstractDue to the popularity of computer games and animation, research on 3D articulated geometry model retrieval has attracted a lot of attention in recent years. However, most existing works extract high-dimensional features to represent models and suffer from practical limitations. First, misalignment in high-dimensional features may produce unreliable euclidean distances and affect retrieval accuracy. Second, the curse of dimensionality also degrades efficiency. In this paper, we propose an embedding retrieval framework to improve the practicability of these methods. It is based on a manifold learning technique, the Diffusion Map (DM). We project all pairwise distances onto a low-dimensional space. This improves retrieval accuracy because intercluster distances are exaggerated. Then we adapt the Density-Weighted Nyström extension and further propose a novel step to locally align the Nyström embedding to the eigensolver embedding so as to reduce extension error and preserve retrieval accuracy. Finally, we propose a heuristic to handle disconnected manifolds by augmenting the kernel matrix with multiple similarity measures and shortcut edges, and further discuss the choice of DM parameters. We have incorporated two existing matching algorithms for testing. Our experimental results show improvement in precision at high recalls and in speed. Our work provides a robust retrieval framework for the matching of multimedia data that lie on manifolds. Gary K. L. Tam, Rynson W. H. Lau |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Visualization of Time-Series Data in Parameter Space for Understanding Facial DynamicsabstractAbstract Over the past decade, computer scientists and psychologists have made great efforts to collect and analyze facial dynamics data that exhibit different expressions and emotions. Such data is commonly captured as videos and are transformed into feature‐based time‐series prior to any analysis. However, the analytical tasks, such as expression classification, have been hindered by the lack of understanding of the complex data space and the associated algorithm space. Conventional graph‐based time‐series visualization is also found inadequate to support such tasks. In this work, we adopt a visual analytics approach by visualizing the correlation between the algorithm space and our goal – classifying facial dynamics. We transform multiple feature‐based time‐series for each expression in measurement space to a multi‐dimensional representation in parameter space. This enables us to utilize parallel coordinates visualization to gain an understanding of the algorithm space, providing a fast and cost‐effective means to support the design of analytical algorithms. Gary K. L. Tam, Hui Fang 0003, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, David Marshall 0001, Min Chen 0001 |
Comput. Graph. Forum | 1 |
| 2007 | Motion Retrieval Based on Energy MorphingabstractMatching and retrieval of motion sequences has become an important research area in recent years, due to the increasing availability and popularity of motion capture data. The main challenge in matching two motion sequences is the diversity of the captured motions, including variable length, local shifting, local and global scaling. Most existing methods employ Dynamic Time Warping (DTW) or Uniform Scaling to handle these problems. In this paper, we propose a novel content-based method for matching of this human motion captured data. We convert the matching problem of motion capture data into a transportation problem. To solve this problem efficiently, we employ Earth Mover's Distance (EMD) as the matching framework. To penalize any strayed matching, we provide a ground distance that works similar to Sakoe- Chiba band of DTW. Empirical results obtained are encouraging. Gary K. L. Tam, Qingzheng Zheng, Mark Corbyn, Rynson W. H. Lau |
ISM | 1 |
| 2007 | Deformable Model Retrieval Based on Topological and Geometric SignaturesabstractWith the increasing popularity of 3D applications such as computer games, a lot of 3D geometry models are being created. To encourage sharing and reuse, techniques that support matching and retrieval of these models are emerging. However, only a few of them can handle deformable models, that is, models of different poses, and these methods are generally very slow. In this paper, we present a novel method for efficient matching and retrieval of 3D deformable models. Our research idea stresses using both topological and geometric features at the same time. First, we propose Topological Point Ring (TPR) analysis to locate reliable topological points and rings. Second, we capture both local and global geometric information to characterize each of these topological features. To compare the similarity of two models, we adapt the Earth Mover Distance (EMD) as the distance function and construct an indexing tree to accelerate the retrieval process. We demonstrate the performance of the new method, both in terms of accuracy and speed, through a large number of experiments. Gary K. L. Tam, Rynson W. H. Lau |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2004 | Deformable Object Model Matching by Topological and Geometric SimilarityabstractWe present a novel method for efficient 3D model comparison. The method is designed to match highly deformed models through capturing two types of information. First, we propose a feature point extraction algorithm, which is based on "Level Set Diagram ", to reliably capture the topological points of a general 3D model. These topological points represent the skeletal structure of the model. Second, we also capture both spatial and curvature information, which describes the global surface of a 3D model. This is different from traditional topological 3D matching methods that use only low-dimension local features. Our method can accurately distinguish different types of 3D models even if they have similar topology. By applying the bipartite graph matching technique, our method can achieve a high precision of 0.54 even at a recall rate of 1.0 as demonstrated in our experimental results. Gary K. L. Tam, Rynson W. H. Lau, Chong-Wah Ngo |
Computer Graphics International | 1 |