Frederick W. B. Li

dblp:39/193 · DBLP profile ↗
← Back
80ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0002-4283-4228ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 17 · 15 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-temporal 3D Gaussian Splatting for Sparse-view Guided Scene Update
abstract
Maintaining consistent 3D scene representations over time is a significant challenge in computer vision. Updating 3D scenes from sparse-view observations is crucial for various real-world applications, including urban planning, disaster assessment, and historical site preservation, where dense scans are often unavailable or impractical. In this paper, we propose Cross-Temporal 3D Gaussian Splatting (Cross-Temporal 3DGS), a novel framework for efficiently reconstructing and updating 3D scenes across different time periods, using sparse images and previously captured scene priors. Our approach comprises three stages: 1) Cross-temporal camera alignment for estimating and aligning camera poses across different timestamps; 2) Interference-based confidence initialization to identify unchanged regions between timestamps, thereby guiding updates; and 3) Progressive cross-temporal optimization, which iteratively integrates historical prior information into the 3D scene to enhance reconstruction quality. Our method supports non-continuous capture, enabling not only updates using new sparse views to refine existing scenes, but also recovering past scenes from limited data with the help of current captures. Furthermore, we demonstrate the potential of this approach to achieve temporal changes using only sparse images, which can later be reconstructed into detailed 3D representations as needed. Experimental results show significant improvements over baseline methods in reconstruction quality and data efficiency, making this approach a promising solution for scene versioning, cross-temporal digital twins, and long-term spatial documentation.
Zeyuan An, Yanghang Xiao, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
AAAI4
2026 R²D-LPCC: Relevance-Ranking Guided Region-Adaptive Dynamic LiDAR Point Cloud Compression
abstract
Dynamic LiDAR point cloud compression (LPCC) is crucial for the efficient transmission and storage of large-scale three-dimensional data in applications such as autonomous driving. However, many existing methods, which primarily focus on compressing geometric or motion information, face a fundamental limitation: they treat all points as equally important. This approach neglects the semantic priorities of a scene, resulting in inefficient bit allocation and particularly compromising the reconstruction quality of safety-critical regions, such as pedestrians and vehicles, which are vital to downstream perception tasks. To address these limitations, we propose R²D-LPCC, a relevance-ranking framework for region adaptive LPCC that prioritizes fidelity in semantically important regions. Central to our approach is the Adaptive Relevance Learning (ARL) module, which integrates semantic context with uncertainty to evaluate regional significance and guide compression. We also introduce a Multi-scale Region-Adaptive Transform (MRAT) module to enhance semantic feature modeling and preserve fine-grained details in key areas. Additionally, we develop an adaptive multi-modal motion estimation module to improve motion prediction in complex three-dimensional environments. Extensive experiments conducted on the SemanticKITTI benchmark demonstrate that R²D-LPCC significantly surpasses ten recent state-of-the-art methods, achieving a 45.48% BD-rate gain over the previous leading method, Unicorn, and a 98.58% gain over the GPCC standard, while ensuring superior reconstruction quality in semantically important regions.
Fangzhe Nan, Frederick W. B. Li, Gary K. L. Tam, Zhaoyi Jiang, Bailin Yang, Jingke Cui, Changshuo Wang 0001
AAAI2
2026 OctMamba: Mamba-based octree context entropy model for point cloud geometry compression
abstract
Existing learned point cloud compression frameworks face two major limitations: (1) they focus almost exclusively on spatial redundancy and (2) rely on architectures built around local-global transformers or global Mamba blocks. Transformers incur quadratic complexity, while global Mamba lacks the granularity to capture structured correlations across multiple dimensions. We propose OctMamba, the first unified framework to jointly exploit spatial, channel, and topological redundancies, dimensions previously overlooked in point cloud geometry compression. Our approach introduces a new architectural principle: embedding Mamba modules within specialized subcomponents rather than applying them globally, challenging existing design paradigms. OctMamba combines two modules: Spatial-Channel Coupled Grouping Mamba (SCCGM) for spatial-channel fusion and Local Graph CNN-Mamba (LGCM) for topological encoding. This design enables efficient long-range modeling with linear complexity, delivering a smaller model and faster decoding while outperforming transformer-based and global Mamba baselines. On SemanticKITTI, OctMamba reduces bitrate by 60.2% over GPCC (D1 PSNR) and achieves state-of-the-art performance across LiDAR and dynamic human point cloud benchmarks with practical speed and scalability. By introducing multi-dimensional redundancy modeling, OctMamba has the potential to influence future research on efficient point cloud compression. The code is available at https://github.com/ZjgsVMC/OctMamba .
Zhaoyi Jiang, Frederick W. B. Li, Gary K. L. Tam, Chao Song 0001, Bailin Yang
Pattern Recognit.3
2026 Uncertainty-aware calibrated 3D human motion forecasting with latent conformal prediction
Yue Ma 0035, Frederick W. B. Li, Xiaohui Liang 0001
Pattern Recognit.2
2026 Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes
abstract
Scenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world models, we introduce Dyn-HSI, the first cognitive architecture for dynamic human-scene interaction, which endows virtual humans with three humanoid components. (1) Vision (human eyes): we equip the virtual human with a Dynamic Scene-Aware Navigation, which continuously perceives changes in the surrounding environment and adaptively predicts the next waypoint. (2) Memory (human brain): we equip the virtual human with a Hierarchical Experience Memory, which stores and updates experiential data accumulated during training. This allows the model to leverage prior knowledge during inference for context-aware motion priming, thereby enhancing both motion quality and generalization. (3) Control (human body): we equip the virtual human with Human-Scene Interaction Diffusion Model, which generates high-fidelity interaction motions conditioned on multimodal inputs. To evaluate performance in dynamic scenes, we extend the existing static human-scene interaction datasets to construct a dynamic benchmark, Dyn-Scenes. We conduct extensive qualitative and quantitative experiments to validate Dyn-HSI, showing that our method consistently outperforms existing approaches and generates high-quality human-scene interaction motions in both static and dynamic settings.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.4
2025 Multi-modal Dynamic Point Cloud Geometric Compression Based on Bidirectional Recurrent Scene Flow
abstract
Deep learning methods have recently shown significant promise in compressing the geometric features of point clouds. However, challenges arise when consecutive point clouds contain holes, resulting in incomplete information that complicates motion estimation. To our knowledge, most existing dynamic point cloud compression methods have largely overlooked this critical issue. Moreover, these methods typically employ a multi-scale single-pass approach for motion estimation, performing only one estimation at each scale. This limits accuracy and adversely impacts compression performance. To address these challenges, we propose a dynamic point cloud compression model called M2BR-DPCC (Multi-Modal Multi-Scale Bidirectional Recursion for Dynamic Point Cloud Compression). Our method introduces two key innovations. First, we integrate both point cloud and image data as inputs, leveraging a multi-modal feature representation completion (MFRepC) approach to align information across modalities. This addresses the issue of missing data in point clouds by using complementary information from images. Second, we implement a multi-scale bidirectional recursive (MSBR) motion estimation method. This module iteratively refines motion flows in both forward and backward directions, progressively enhancing point cloud features and improving motion estimation accuracy. Experimental results on widely used datasets, including MVUB and 8iVFB, demonstrate the effectiveness of our approach. Compared to existing methods, M2BR-DPCC achieves superior performance, with an average BD-rate improvement of 95.23% over V-PCC, 12.92% over D-DPCC, and 16.16% over patchDPCC. These results underscore the potential of leveraging multi-modal data and bidirectional refinement for dynamic point cloud compression.
Fangzhe Nan, Frederick W. B. Li, Zhuoyue Wang, Gary K. L. Tam, Zhaoyi Jiang, DongZheng DongZheng, Bailin Yang
ICASSP2
2025 Uncertainty-aware Probabilistic 3D Human Motion Forecasting via Invertible Networks
abstract
3D human motion forecasting aims to enable autonomous applications. Estimating uncertainty for each prediction (i.e., confidence based on probability density or quantile) is essential for safety-critical contexts like human-robot collaboration to minimize risks. However, existing diverse motion fore-casting approaches struggle with uncertainty quantification due to implicit probabilistic representations hindering uncertainty modeling. We propose ProbHMI, which introduces invertible networks to parameterize poses in a disentangled latent space, enabling probabilistic dynamics modeling. A forecasting module then explicitly predicts future latent distributions, allowing effective uncertainty quantification. Evaluated on benchmarks, ProbHMI achieves strong performance for both deterministic and diverse prediction while validating uncertainty calibration, critical for risk-aware decision making.
Yue Ma 0035, Kanglei Zhou, Fuyang Yu, Frederick W. B. Li, Xiaohui Liang 0001
ICRA4
2025 User Evaluation of the Digital Algebra Game (DAG) "Algebra Adventure: Learning Through Storytelling"
Kubra Kaymakci Ustuner, Effie Lai-Chong Law, Frederick W. B. Li, Nazli Ozer
INTERACT (3)3
2025 3D data augmentation and dual-branch model for robust face forgery detection
abstract
We propose Dual-Branch Network (DBNet), a novel deepfake detection framework that addresses key limitations of existing works by jointly modeling 3D-temporal and fine-grained texture representations. Specifically, we aim to investigate how to (1) capture dynamic properties and spatial details in a unified model and (2) identify subtle inconsistencies beyond localized artifacts through temporally consistent modeling. To this end, DBNet extracts 3D landmarks from videos to construct temporal sequences for an RNN branch, while a Vision Transformer analyzes local patches. A Temporal Consistency-aware Loss is introduced to explicitly supervise the RNN. Additionally, a 3D generative model augments training data. Extensive experiments demonstrate our method achieves state-of-the-art performance on benchmarks, and ablation studies validate its effectiveness in generalizing to unseen data under various manipulations and compression.
Changshuang Zhou, Frederick W. B. Li, Chao Song 0001, Bailin Yang
Graph. Model.2
2025 WDFSR: Normalizing Flow Based on the Wavelet-Domain for Super-Resolution
abstract
We propose a normalizing flow based on the wavelet framework for super-resolution (SR) called WDFSR. It learns the conditional distribution mapping between low-resolution images in the RGB domain and high-resolution images in the wavelet domain to simultaneously generate high-resolution images of different styles. To address the issue of some flow-based models being sensitive to datasets, which results in training fluctuations that reduce the mapping ability of the model and weaken generalization, we designed a method that combines a T-distribution and QR decomposition layer. Our method alleviates this problem while maintaining the ability of the model to map different distributions and produce higher-quality images. Good contextual conditional features can promote model training and enhance the distribution mapping capabilities for conditional distribution mapping. Therefore, we propose a Refinement layer combined with an attention mechanism to refine and fuse the extracted condition features to improve image quality. Extensive experiments on several SR datasets demonstrate that WDFSR outperforms most general CNN- and flow-based models in terms of PSNR value and perception quality. We also demonstrated that our framework works well for other low-level vision tasks, such as low-light enhancement. The pretrained models and source code with guidance for reference are available at https://github.com/Lisbegin/WDFSR.
Chao Song 0001, Shaobang Li, Frederick W. B. Li, Bailin Yang
Comput. Vis. Media3
2025 Geometric visual fusion graph neural networks for multi-person human-object interaction recognition in videos
abstract
Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture appearance context, while geometric features provide structural patterns. Effectively fusing these multimodal features without compromising their unique characteristics remains challenging. We observe that establishing robust, entity-specific representations before modeling interactions helps preserve the strengths of each modality. Therefore, we hypothesize that a bottom-up approach is crucial for effective multimodal fusion. Following this insight, we propose the Geometric Visual Fusion Graph Neural Network (GeoVis-GNN), which uses dual-attention feature fusion combined with interdependent entity graph learning. It progressively builds from entity-specific representations toward high-level interaction understanding. To advance HOI recognition to real-world scenarios, we introduce the Concurrent Partial Interaction Dataset (MPHOI-120). It captures dynamic multi-person interactions involving concurrent actions and partial engagement. This dataset helps address challenges like complex human-object dynamics and mutual occlusions. Extensive experiments demonstrate the effectiveness of our method across various HOI scenarios. These scenarios include two-person interactions, single-person activities, bimanual manipulations, and complex concurrent partial interactions. Our method achieves state-of-the-art performance.
Tanqiu Qiao, Ruochen Li 0002, Frederick W. B. Li, Yoshiki Kubotani, Shigeo Morishima, Hubert P. H. Shum
Expert Syst. Appl.3
2025 Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
Int. J. Comput. Vis.5
2025 CLPFusion: A Latent Diffusion Model Framework for Realistic Chinese Landscape Painting Style Transfer
abstract
ABSTRACT This study focuses on transforming real‐world scenery into Chinese landscape painting masterpieces through style transfer. Traditional methods using convolutional neural networks (CNNs) and generative adversarial networks (GANs) often yield inconsistent patterns and artifacts. The rise of diffusion models (DMs) presents new opportunities for realistic image generation, but their inherent noise characteristics make it challenging to synthesize pure white or black images. Consequently, existing DM‐based methods struggle to capture the unique style and color information of Chinese landscape paintings. To overcome these limitations, we propose CLPFusion, a novel framework that leverages pre‐trained diffusion models for artistic style transfer. A key innovation is the Bidirectional State Space Models‐CrossAttention (BiSSM‐CA) module, which efficiently learns and retains the distinct styles of Chinese landscape paintings. Additionally, we introduce two latent space feature adjustment methods, Latent‐AdaIN and Latent‐WCT, to enhance style modulation during inference. Experiments demonstrate that CLPFusion produces more realistic and artistic Chinese landscape paintings than existing approaches, showcasing its effectiveness and uniqueness in the field.
Jiahui Pan 0001, Frederick W. B. Li, Bailin Yang, Fangzhe Nan
Comput. Animat. Virtual Worlds2
2025 Talking Face Generation With Lip and Identity Priors
abstract
ABSTRACT Speech‐driven talking face video generation has attracted growing interest in recent research. While person‐specific approaches yield high‐fidelity results, they require extensive training data from each individual speaker. In contrast, general‐purpose methods often struggle with accurate lip synchronization, identity preservation, and natural facial movements. To address these limitations, we propose a novel architecture that combines an alignment model with a rendering model. The rendering model synthesizes identity‐consistent lip movements by leveraging facial landmarks derived from speech, a partially occluded target face, multi‐reference lip features, and the input audio. Concurrently, the alignment model estimates optical flow using the occluded face and a static reference image, enabling precise alignment of facial poses and lip shapes. This collaborative design enhances the rendering process, resulting in more realistic and identity‐preserving outputs. Extensive experiments demonstrate that our method significantly improves lip synchronization and identity retention, establishing a new benchmark in talking face video generation.
Frederick W. B. Li, Gary K. L. Tam, Bailin Yang, Fangzhe Nan, Jia Pan 0001
Comput. Animat. Virtual Worlds2
2025 PHI: Bridging Domain Shift in Long-Term Action Quality Assessment via Progressive Hierarchical Instruction
abstract
Long-term Action Quality Assessment (AQA) aims to evaluate the quantitative performance of actions in long videos. However, existing methods face challenges due to domain shifts between the pre-trained large-scale action recognition backbones and the specific AQA task, thereby hindering their performance. This arises since fine-tuning resource-intensive backbones on small AQA datasets is impractical. We address this by identifying two levels of domain shift: task-level, regarding differences in task objectives, and feature-level, regarding differences in important features. For feature-level shifts, which are more detrimental, we propose Progressive Hierarchical Instruction (PHI) with two strategies. First, Gap Minimization Flow (GMF) leverages flow matching to progressively learn a fast flow path that reduces the domain gap between initial and desired features across shallow to deep layers. Additionally, a temporally-enhanced attention module captures long-range dependencies essential for AQA. Second, List-wise Contrastive Regularization (LCR) facilitates coarse-to-fine alignment by comprehensively comparing batch pairs to learn fine-grained cues while mitigating domain shift. Integrating these modules, PHI offers an effective solution. Experiments demonstrate that PHI achieves state-of-the-art performance on three representative long-term AQA datasets, proving its superiority in addressing the domain shift for long-term AQA.
Kanglei Zhou, Hubert P. H. Shum, Frederick W. B. Li, Xingxing Zhang 0001, Xiaohui Liang 0001
IEEE Trans. Image Process.3
2025 MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
abstract
We introduce MOST, a novel MOtion diffuSion model via Temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.4
2024 MAGR: Manifold-Aligned Graph Regularization for Continual Action Quality Assessment
Kanglei Zhou, Xingxing Zhang 0001, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
ECCV (11)5
2024 From Category to Scenery: An End-to-End Framework for Multi-person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Ruochen Li 0002, Frederick W. B. Li, Hubert P. H. Shum
ICPR (15)3
2024 Multi-style cartoonization: Leveraging multiple datasets with generative adversarial networks
abstract
Abstract Scene cartoonization aims to convert photos into stylized cartoons. While generative adversarial networks (GANs) can generate high‐quality images, previous methods focus on individual images or single styles, ignoring relationships between datasets. We propose a novel multi‐style scene cartoonization GAN that leverages multiple cartoon datasets jointly. Our main technical contribution is a multi‐branch style encoder that disentangles representations to model styles as distributions over entire datasets rather than images. Combined with a multi‐task discriminator and perceptual losses optimizing across collections, our model achieves state‐of‐the‐art diverse stylization while preserving semantics. Experiments demonstrate that by learning from inter‐dataset relationships, our method translates photos into cartoon images with improved realism and abstraction fidelity compared to prior arts, without iterative re‐training for new styles.
Jianlu Cai, Frederick W. B. Li, Fangzhe Nan, Bailin Yang
Comput. Animat. Virtual Worlds2
2024 Color Theme Evaluation through User Preference Modeling
abstract
Color composition (or color theme) is a key factor to determine how well a piece of art work or graphical design is perceived by humans. Despite a few color harmony models have been proposed, their results are often less satisfactory since they mostly neglect the variations of aesthetic cognition among individuals and treat the influence of all ratings equally as if they were all rated by the same anonymous user. To overcome this issue, in this article we propose a new color theme evaluation model by combining a back propagation neural network and a kernel probabilistic model to infer both the color theme rating and the user aesthetic preference. Our experiment results show that our model can predict more accurate and personalized color theme ratings than state of the art methods. Our work is also the first-of-its-kind effort to quantitatively evaluate the correlation between user aesthetic preferences and color harmonies of five-color themes, and study such a relation for users with different aesthetic cognition.
Bailin Yang, Tianxiang Wei, Frederick W. B. Li, Xiaohui Liang 0001, Zhigang Deng 0001, Yili Fang
ACM Trans. Appl. Percept.3
2024 Laplacian Projection Based Global Physical Prior Smoke Reconstruction
abstract
We present a novel framework for reconstructing fluid dynamics in real-life scenarios. Our approach leverages sparse view images and incorporates physical priors across long series of frames, resulting in reconstructed fluids with enhanced physical consistency. Unlike previous methods, we utilize a differentiable fluid simulator (DFS) and a differentiable renderer (DR) to exploit global physical priors, reducing reconstruction errors without the need for manual regularization coefficients. We introduce divergence-free Laplacian eigenfunctions (div-free LE) as velocity bases, improving computational efficiency and memory usage. By employing gradient-related strategies, we achieve better convergence and superior results. Extensive experiments demonstrate the effectiveness of our method, showcasing improved reconstruction quality and computational efficiency compared to existing approaches. We validate our approach using both synthetic and real data, highlighting its practical potential.
Shibang Xiao, Chao Tong 0001, Qifan Zhang 0003, Yunchi Cen, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Multi-Task Spatial-Temporal Graph Auto-Encoder for Hand Motion Denoising
abstract
In many human-computer interaction applications, fast and accurate hand tracking is necessary for an immersive experience. However, raw hand motion data can be flawed due to issues such as joint occlusions and high-frequency noise, hindering the interaction. Using only current motion for interaction can lead to lag, so predicting future movement is crucial for a faster response. Our solution is the Multi-task Spatial-Temporal Graph Auto-Encoder (Multi-STGAE), a model that accurately denoises and predicts hand motion by exploiting the inter-dependency of both tasks. The model ensures a stable and accurate prediction through denoising while maintaining motion dynamics to avoid over-smoothed motion and alleviate time delays through prediction. A gate mechanism is integrated to prevent negative transfer between tasks and further boost multi-task performance. Multi-STGAE also includes a spatial-temporal graph autoencoder block, which models hand structures and motion coherence through graph convolutional networks, reducing noise while preserving hand physiology. Additionally, we design a novel hand partition strategy and hand bone loss to improve natural hand motion generation. We validate the effectiveness of our proposed method by contributing two large-scale datasets with a data corruption algorithm based on two benchmark datasets. To evaluate the natural characteristics of the denoised and predicted hand motion, we propose two structural metrics. Experimental results show that our method outperforms the state-of-the-art, showcasing how the multi-task framework enables mutual benefits between denoising and prediction.
Kanglei Zhou, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Multi-feature fusion enhanced monocular depth estimation with boundary awareness
Chao Song 0001, Qingjie Chen, Frederick W. B. Li, Zhaoyi Jiang, Yuliang Shen, Bailin Yang
Vis. Comput.3
2023 DrawGAN: Multi-view Generative Model Inspired by the Artist's Drawing Method
Bailin Yang, Frederick W. B. Li, Haoqiang Sun, Jianlu Cai
CGI3
2023 Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model
abstract
Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text description. In this work, we propose a fine-grained method for generating high-quality, conditional human motion sequences supporting precise text description. Our approach consists of two key components: 1) a linguistics-structure assisted module that constructs accurate and complete language feature to fully utilize text information; and 2) a context-aware progressive reasoning module that learns neighborhood and overall semantic linguistics features from shallow and deep graph neural networks to achieve a multi-step inference. Experiments show that our approach outperforms text-driven motion generation methods on HumanML3D and KIT test sets and generates better visually confirmed motion to the text conditions.
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001
ICCV3
2023 HSE: Hybrid Species Embedding for Deep Metric Learning
abstract
Deep metric learning is crucial for finding an embedding function that can generalize to training and testing data, including unknown test classes. However, limited training samples restrict the model’s generalization to downstream tasks. While adding new training samples is a promising solution, determining their labels remains a significant challenge. Here, we introduce Hybrid Species Embedding (HSE), which employs mixed sample data augmentations to generate hybrid species and provide additional training signals. We demonstrate that HSE outperforms multiple state-of-the-art methods in improving the metric Recall@K on the CUB-200 , CAR-196 and SOP datasets, thus offering a novel solution to deep metric learning’s limitations.
Bailin Yang, Haoqiang Sun, Frederick W. B. Li, Jianlu Cai, Chao Song 0001
ICCV3
2023 Digital Educational Games with Storytelling for Students to Learn Algebra
Kubra Kaymakci Ustuner, Effie Lai-Chong Law, Frederick W. B. Li
INTERACT (4)3
2023 A Mixed Reality Training System for Hand-Object Interaction in Simulated Microgravity Environments
abstract
As human exploration of space continues to progress, the use of Mixed Reality (MR) for simulating microgravity environments and facilitating training in hand-object interaction holds immense practical significance. However, hand-object interaction in microgravity presents distinct challenges compared to terrestrial environments due to the absence of gravity. This results in heightened agility and inherent unpredictability of movements that traditional methods struggle to simulate accurately. To this end, we propose a novel MR-based hand-object interaction system in simulated microgravity environments, leveraging physics-based simulations to enhance the interaction between the user’s real hand and virtual objects. Specifically, we introduce a physics-based hand-object interaction model that combines impulse-based simulation with penetration contact dynamics. This accurately captures the intricacies of hand-object interaction in microgravity. By considering forces and impulses during contact, our model ensures realistic collision responses and enables effective object manipulation in the absence of gravity. The proposed system presents a cost-effective solution for users to simulate object manipulation in microgravity. It also holds promise for training space travelers, equipping them with greater immersion to better adapt to space missions. The system reliability and fidelity test verifies the superior effectiveness of our system compared to the state-of-the-art CLAP system.
Kanglei Zhou, Yue Ma 0035, Zhiying Leng, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
ISMAR6
2023 C2SPoint: A classification-to-saliency network for point cloud saliency detection
abstract
Point cloud saliency detection is an important technique that support downstream tasks in 3D graphics and vision, like 3D model simplification, compression, reconstruction and viewpoint selection. Existing approaches often rely on hand-crafted features and are only applicable to specific datasets. In this paper, we propose a novel weakly supervised classification network, called C2SPoint, which directly performs saliency detection on the point clouds. Unlike previous methods that require per-point saliency annotations, C2SPoint only requires category labels of the point clouds during training. The network consists of two branches: a Classification branch and a Saliency branch. The former branch is composed of two Adaptive Set Abstraction layers for feature extraction and a Saliency Transform layer for learning saliency knowledge from the classification network. The latter branch introduces a multi-scale point-cluster similarity matrix for propagating the cluster saliency to each point within it, resulting in the prediction of point-level saliency. Experimental results demonstrate the effectiveness of our method in point cloud saliency detection, with improvements of 2% in both AUC and NSS compared to state-of-the-art methods.
Zhaoyi Jiang, Luyun Ding, Gary K. L. Tam, Chao Song 0001, Frederick W. B. Li, Bailin Yang
Comput. Graph.5
2023 A Differential Diffusion Theory for Participating Media
abstract
Abstract We present a novel approach to differentiable rendering for participating media, addressing the challenge of computing scene parameter derivatives. While existing methods focus on derivative computation within volumetric path tracing, they fail to significantly improve computational performance due to the expensive computation of multiply‐scattered light. To overcome this limitation, we propose a differential diffusion theory inspired by the classical diffusion equation. Our theory enables real‐time computation of arbitrary derivatives such as optical absorption, scattering coefficients, and anisotropic parameters of phase functions. By solving derivatives through the differential form of the diffusion equation, our approach achieves remarkable speed gains compared to Monte Carlo methods. This marks the first differentiable rendering framework to compute scene parameter derivatives based on diffusion approximation. Additionally, we derive the discrete form of diffusion equation derivatives, facilitating efficient numerical solutions. Our experimental results using synthetic and realistic images demonstrate the accurate and efficient estimation of arbitrary scene parameter derivatives. Our work represents a significant advancement in differentiable rendering for participating media, offering a practical and efficient solution to compute derivatives while addressing the limitations of existing approaches.
Yunchi Cen, Frederick W. B. Li, Bailin Yang, Xiaohui Liang 0001
Comput. Graph. Forum3
2023 A Video-Based Augmented Reality System for Human-in-the-Loop Muscle Strength Assessment of Juvenile Dermatomyositis
abstract
As the most common idiopathic inflammatory myopathy in children, juvenile dermatomyositis (JDM) is characterized by skin rashes and muscle weakness. The childhood myositis assessment scale (CMAS) is commonly used to measure the degree of muscle involvement for diagnosis or rehabilitation monitoring. On the one hand, human diagnosis is not scalable and may be subject to personal bias. On the other hand, automatic action quality assessment (AQA) algorithms cannot guarantee 100% accuracy, making them not suitable for biomedical applications. As a solution, we propose a video-based augmented reality system for human-in-the-loop muscle strength assessment of children with JDM. We first propose an AQA algorithm for muscle strength assessment of JDM using contrastive regression trained by a JDM dataset. Our core insight is to visualize the AQA results as a virtual character facilitated by a 3D animation dataset, so that users can compare the real-world patient and the virtual character to understand and verify the AQA results. To allow effective comparisons, we propose a video-based augmented reality system. Given a feed, we adapt computer vision algorithms for scene understanding, evaluate the optimal way of augmenting the virtual character into the scene, and highlight important parts for effective human verification. The experimental results confirm the effectiveness of our AQA algorithm, and the results of the user study demonstrate that humans can more accurately and quickly assess the muscle strength of children using our system.
Kanglei Zhou, Ruizhi Cai, Yue Ma 0035, Qingqing Tan, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.8
2023 IAACS: Image Aesthetic Assessment Through Color Composition And Space Formation
abstract
Judging how an image is visually appealing is a complicated and subjective task. This highly motivates having a machine learning model to automatically evaluate image aesthetic by matching the aesthetics of general public. Although deep learning methods have been successfully learning good visual features from images, correctly assessing image aesthetic quality is still challenging for deep learning. To tackle this, we propose a novel multi-view convolutional neural network to assess image aesthetic by analyzing image color composition and space formation (IAACS). Specifically, from different views of an image, including its key color components with their contributions, the image space formation and the image itself, our network extracts their corresponding features through our proposed feature extraction module (FET) and the ImageNet weight-based classification model. By fusing the extracted features, our network produces an accurate prediction score distribution of image aesthetic. Experiment results have shown that we have achieved a superior performance.
Bailin Yang, Changrui Zhu, Frederick W. B. Li, Tianxiang Wei, Xiaohui Liang 0001, Qingxu Wang
Virtual Real. Intell. Hardw.3
2022 Geometric Features Informed Multi-person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Qianhui Men, Frederick W. B. Li, Yoshiki Kubotani, Shigeo Morishima, Hubert P. H. Shum
ECCV (4)3
2022 STIT: Spatio-Temporal Interaction Transformers for Human-Object Interaction Recognition in Videos
abstract
Recognizing human-object interactions is challenging due to their spatio-temporal changes. We propose the Spatio-Temporal Interaction Transformer-based (STIT) network to reason such changes. Specifically, spatial transformers learn humans and objects context at specific frame time. Temporal transformer then learns the relations at a higher level between spatial context representations at different time steps, capturing long-term dependencies across frames. We further investigate multiple hierarchy designs in learning human interactions. We achieved superior performance on Charades, Something-Something v1 and CAD-120 datasets, comparing to baseline models without learning human-object relations, or with prior graph-based networks. We also achieved state-of-the-art accuracy of 95.93% on CAD-120 dataset [1] by employing RGB data only.
Muna Almushyti, Frederick W. B. Li
ICPR2
2022 Distillation of human-object interaction contexts for action recognition
abstract
Abstract Modeling spatial‐temporal relations is imperative for recognizing human actions, especially when a human is interacting with objects, while multiple objects appear around the human differently over time. Most existing action recognition models focus on learning overall visual cues of a scene but disregard a holistic view of human–object relationships and interactions, that is, how a human interacts with respect to short‐term task for completion and long‐term goal. We therefore argue to improve human action recognition by exploiting both the local and global contexts of human–object interactions (HOIs). In this paper, we propose the Global‐Local Interaction Distillation Network (GLIDN), learning human and object interactions through space and time via knowledge distillation for holistic HOI understanding. GLIDN encodes humans and objects into graph nodes and learns local and global relations via graph attention network. The local context graphs learn the relation between humans and objects at a frame level by capturing their co‐occurrence at a specific time step. The global relation graph is constructed based on the video‐level of human and object interactions, identifying their long‐term relations throughout a video sequence. We also investigate how knowledge from these graphs can be distilled to their counterparts for improving HOI recognition. Finally, we evaluate our model by conducting comprehensive experiments on two datasets including Charades and CAD‐120. Our method outperforms the baselines and counterpart approaches.
Muna Almushyti, Frederick W. B. Li
Comput. Animat. Virtual Worlds2
2021 STGAE: Spatial-Temporal Graph Auto-Encoder for Hand Motion Denoising
abstract
Hand object interaction in mixed reality (MR) relies on the accurate tracking and estimation of human hands, which provide users with a sense of immersion. However, raw captured hand motion data always contains errors such as joints occlusion, dislocation, high-frequency noise, and involuntary jitter. Denoising and obtaining the hand motion data consistent with the user’s intention are of the utmost importance to enhance the interactive experience in MR. To this end, we propose an end-to-end method for hand motion denoising using the spatial-temporal graph auto-encoder (STGAE). The spatial and temporal patterns are recognized simultaneously by constructing the consecutive hand joint sequence as a spatial-temporal graph. Considering the complexity of the articulated hand structure, a simple yet effective partition strategy is proposed to model the physic-connected and symmetry-connected relationships. Graph convolution is applied to extract structural constraints of the hand, and a self-attention mechanism is to adjust the graph topology dynamically. Combining graph convolution and temporal convolution, a fundamental graph encoder or decoder block is proposed. We finally establish the hourglass residual auto-encoder to learn a manifold projection operation and a corresponding inverse projection through stacking these blocks. In this work, the proposed framework has been successfully used in hand motion data denoising with preserving structural constraints between joints. Extensive quantitative and qualitative experiments show that the proposed method has achieved better performance than the state-of-the-art approaches.
Kanglei Zhou, Zhiyuan Cheng 0004, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
ISMAR4
2021 Stable Hand Pose Estimation under Tremor via Graph Neural Network
abstract
Hand pose estimation, which predicts the spatial location of hand joints, is a fundamental task in VR/AR applications. Although existing methods can recover hand pose competently, the tremor issue occurring in hand motion has not been completely solved. Tremor is an involuntary motion accompanied by a desired gesture or hand motion, leading to hand pose that deviates from user's intentions. Considering the characteristic of tremor motion, we present a novel Graph Neural Network for stable 3D hand pose estimation. The input is depth images. The constraint adjacency matrix is devised in Graph Neural Network for dynamically adjusting the topology of a hand graph during message passing and aggregation. Firstly, since there are rich potential constraints among hand joints, we utilize the constraint adjacency matrix to mine the suitable topology, modeling spatial-temporal constraints of joints and outputting the precise tremor hand pose as the pre-estimation result. Then, for obtaining a stable hand pose, we provide a tremor compensation module based on the constraint adjacency matrix, which exploits the constraint between control points and tremor hand pose. Concretely, the control points represented the voluntary motion are employed as constraints to edit the tremor hand pose. Our extensive quantitative and qualitative experiments show that the proposed method has achieved decent performance for 3D tremor hand pose estimation.
Zhiying Leng, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001
VR4
2021 Facial reshaping operator for controllable face beautification
Shanfeng Hu, Hubert P. H. Shum, Xiaohui Liang 0001, Frederick W. B. Li, Nauman Aslam
Expert Syst. Appl.4
2020 Perceptual Quality Assessment On DIBR Synthesized Videos With Composite Distortions
abstract
DIBR (depth-image-based rendering) synthesized videos has become increasingly popular recently in supporting 3D-TV, Free-viewpoint video (FVV) and interactive graphic applications. Due to the synthesis process, it contains composite distortions, i.e., DCT-like distortions plus geometric distortions. It has led to new challenges for subjective and objective video quality assessment (VQA) community, where such composite distortions have not been taken into consideration. In this work, we construct a new database composed of DIBR synthesized videos, the reference view video is compressed by H.264 and HEVC encoders, which is subsequently synthesized to the virtual view with different DIBR algorithms. We also propose a novel no-reference (NR) synthesized video quality metric with multi-modal feature pooling and evaluate it against representative full reference (FR) and NR objective VQA models on the proposed database. The database and the proposed metric will be made available to the public.
Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001
ICIP4
2020 Cumuliform cloud formation control using parameter-predicting convolutional neural network
Yue Ma 0035, Frederick W. B. Li, Hubert P. H. Shum, Bailin Yang, Xiaohui Liang 0001
Graph. Model.4
2020 Sparse metric-based mesh saliency
Shanfeng Hu, Xiaohui Liang 0001, Hubert P. H. Shum, Frederick W. B. Li, Nauman Aslam
Neurocomputing4
2020 Example-based image recoloring in an indoor environment
abstract
Abstract Color structure of a home scene image closely relates to the material properties of its local regions. Existing color migration methods typically fail to fully infer the correlation between the coloring of local home scene regions, leading to a local blur problem. In this paper, we propose a color migration framework for home scene images. It picks the coloring from a template image and transforms such coloring to a home scene image through a simple interaction. Our framework comprises three main parts. First, we carry out an interactive segmentation to divide an image into local regions and extract their corresponding colors. Second, we generate a matching color table by sampling the template image according to the color structure of the original home scene image. Finally, we transform colors from the matching color table to the target home scene image with the boundary transition maintained. Experimental results show that our method can effectively transform the coloring of a scene matching with the color composition of a given natural or interior scenery.
Xianxuan Lin, Xun Wang 0007, Frederick W. B. Li, Jinyu Li 0002, Bailin Yang, Tianxiang Wei
Comput. Animat. Virtual Worlds3
2020 Target-driven cloud evolution using position-based fluids
abstract
Abstract To effectively control particle‐based cloud evolution without imposing strict position constraints, we propose a novel method integrating a control force field and a phase transition control into the position‐based fluids (PBF) framework. To produce realistic cloud simulation, we incorporate both fluid dynamics and thermodynamics to govern cloud particle movement. The fluid dynamics is simulated through our novel driving and damping force terms. As these terms are only formulated based on cloud particle density and position, they simplify the inputs and make our method free from artificial positional constraints. The thermodynamics is implemented by our phase transition control, which can effectively simulate cloud evolution between discrepant initial and target shapes, producing plausible results. Uniquely, our method can also support target shape change during cloud simulation. Experiment results have demonstrated our method surpasses existing methods.
Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001
Comput. Animat. Virtual Worlds4
2020 A Unified Deep Metric Representation for Mesh Saliency Detection and Non-Rigid Shape Matching
abstract
In this paper, we propose a deep metric for unifying the representation of mesh saliency detection and non-rigid shape matching. While saliency detection and shape matching are two closely related and fundamental tasks in shape analysis, previous methods approach them separately and independently, failing to exploit their mutually beneficial underlying relationship. In view of the existing gap between saliency and matching, we propose to solve them together using a unified metric representation of surface meshes. We show that saliency and matching can be rigorously derived from our representation as the principal eigenvector and the smoothed Laplacian eigenvectors respectively. Learning the representation jointly allows matching to improve the deformation-invariance of saliency while allowing saliency to improve the feature localization of matching. To parameterize the representation from a mesh, we also propose a deep recurrent neural network (RNN) for effectively integrating multi-scale shape features and a soft-thresholding operator for adaptively enhancing the sparsity of saliency. Results show that by jointly learning from a pair of saliency and matching datasets, matching improves the accuracy of detected salient regions on meshes, which is especially obvious for small-scale saliency datasets, such as those having one to two meshes. At the same time, saliency improves the accuracy of shape matchings among meshes with reduced matching errors on surfaces.
Shanfeng Hu, Hubert P. H. Shum, Nauman Aslam, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Multim.4
2019 Deep Blind Synthesized Image Quality Assessment with Contextual Multi-Level Feature Pooling
abstract
Blind image quality metrics have achieved significant improvement on traditional 2D image dataset, yet still being insufficient for evaluating synthesized images generated from depth-image-based rendering. The geometric distortions in synthesized image are non-uniform, which is challenging for feature representation and pooling. To address this, we propose an end-to-end deep blind synthesized image quality metric SIQA-CFP. We particularly design a contextual multilevel feature pooling module to encode low- and high-level features, which are extracted by a deep pre-trained ResNet. Experimental results on IRCCyN/IVC DIBR dataset show that our method outperforms state-of-the-art synthesized image quality metrics. Our method also achieves competitive performance on traditional 2D image datasets like LIVE Challenge and TID2013.
Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001
ICIP4
2019 Conceptual Framework for Virtual Field Trip Games
Amani Alsaqqaf, Frederick W. B. Li
iLRN2
2019 A Color-Pair Based Approach for Accurate Color Harmony Estimation
abstract
Abstract Harmonious color combinations can stimulate positive user emotional responses. However, a widely open research question is: how can we establish a robust and accurate color harmony measure for the public and professional designers to identify the harmony level of a color theme or color set. Building upon the key discovery that color pairs play an important role in harmony estimation, in this paper we present a novel color‐pair based estimation model to accurately measure the color harmony. It first takes a two‐layer maximum likelihood estimation (MLE) based method to compute an initial prediction of color harmony by statistically modeling the pair‐wise color preferences from existing datasets. Then, the initial scores are refined through a back‐propagation neural network (BPNN) with a variety of color features extracted in different color spaces, so that an accurate harmony estimation can be obtained at the end. Our extensive experiments, including performance comparisons of harmony estimation applications, show the advantages of our method in comparison with the state of the art methods.
Bailin Yang, Tianxiang Wei, Xianyong Fang, Zhigang Deng 0001, Frederick W. B. Li, Xun Wang 0007
Comput. Graph. Forum5
2019 No-reference synthetic image quality assessment with convolutional neural network and local image saliency
abstract
Depth-image-based rendering (DIBR) is widely used in 3DTV, free-viewpoint video, and interactive 3D graphics applications. Typically, synthetic images generated by DIBR-based systems incorporate various distortions, particularly geometric distortions induced by object dis-occlusion. Ensuring the quality of synthetic images is critical to maintaining adequate system service. However, traditional 2D image quality metrics are ineffective for evaluating synthetic images as they are not sensitive to geometric distortion. In this paper, we propose a novel no-reference image quality assessment method for synthetic images based on convolutional neural networks, introducing local image saliency as prediction weights. Due to the lack of existing training data, we construct a new DIBR synthetic image dataset as part of our contribution. Experiments were conducted on both the public benchmark IRCCyN/IVC DIBR image dataset and our own dataset. Results demonstrate that our proposed metric outperforms traditional 2D image quality metrics and state-of-the-art DIBR-related metrics.
Xiaohui Liang 0001, Bailin Yang, Frederick W. B. Li
Comput. Vis. Media4
2019 Compressed dynamic mesh sequence for progressive streaming
abstract
Abstract Dynamic mesh sequence (DMS) is a simple and accurate representation for precisely recording a 3D animation sequence. Despite its simplicity, this representation is typically large in data size, making storage and transmission expensive. This paper presents a novel framework that allows effective DMS compression and progressive streaming by eliminating spatial and temporal redundancy. To explore temporal redundancy, we propose a temporal frame‐clustering algorithm to organize DMS frames by their motion trajectory changes, eliminating intracluster redundancy by principal component analysis dimensionality reduction. To eliminate spatial redundancy, we propose an algorithm to transform the coordinates of mesh vertex trajectory into a decorrelated trajectory space, generating a new spatially nonredundant trajectory representation. We finally apply a spectral graph wavelet transform with color set partitioning embedded block encoding to turn the resultant DMS into a multiresolution representation to support progressive streaming. Experiment results show that our method outperforms several existing methods in terms of storage requirement and reconstruction quality.
Bailin Yang, Zhaoyi Jiang, Jiantao Shangguan, Frederick W. B. Li, Chao Song 0001, Yibo Guo, Mingliang Xu 0001
Comput. Animat. Virtual Worlds4
2019 Motion-Aware Compression and Transmission of Mesh Animation Sequences
abstract
With the increasing demand in using 3D mesh data over networks, supporting effective compression and efficient transmission of meshes has caught lots of attention in recent years. This article introduces a novel compression method for 3D mesh animation sequences, supporting user-defined and progressive transmissions over networks. Our motion-aware approach starts with clustering animation frames based on their motion similarities, dividing a mesh animation sequence into fragments of varying lengths. This is done by a novel temporal clustering algorithm, which measures motion similarity based on the curvature and torsion of a space curve formed by corresponding vertices along a series of animation frames. We further segment each cluster based on mesh vertex coherence, representing topological proximity within an object under certain motion. To produce a compact representation, we perform intra-cluster compression based on Graph Fourier Transform (GFT) and Set Partitioning In Hierarchical Trees (SPIHT) coding. Optimized compression results can be achieved by applying GFT due to the proximity in vertex position and motion. We adapt SPIHT to support progressive transmission and design a mechanism to transmit mesh animation sequences with user-defined quality. Experimental results show that our method can obtain a high compression ratio while maintaining a low reconstruction error.
Bailin Yang, Luhong Zhang, Frederick W. B. Li, Xiaoheng Jiang, Zhigang Deng 0001, Meng Wang 0001, Mingliang Xu 0001
ACM Trans. Intell. Syst. Technol.3
2018 DOOBNet: Deep Object Occlusion Boundary Detection from an Image
Guoxia Wang, Frederick W. B. Li, Xiaohui Liang 0001
ACCV (6)3
2017 Modeling Cumulus Cloud Scenes from High-resolution Satellite Images
abstract
Abstract We present a reconstruction framework, which fits physically‐based constraints to model large‐scale cloud scenes from satellite images. Applications include weather phenomena visualization, flight simulation, and weather spotter training. In our method, the cloud shape is assumed to be composed of a cloud top surface and a nearly flat cloud base surface. Based on this, an effective method of multi‐spectral data processing is developed to obtain relevant information for calculating the cloud base height and the cloud top height, including ground temperature, cloud top temperature and cloud shadow. A lapse rate model is proposed to formulate cloud shape as an implicit function of temperature lapse rate and cloud base temperature. After obtaining initial cloud shapes, we enrich the shapes by a fractal method and represent reconstructed clouds by a particle system. Experiment results demonstrate the capability of our method in generating physically sound large‐scale cloud scenes from high‐resolution satellite images.
Xiaohui Liang 0001, Chunqiang Yuan, Frederick W. B. Li
Comput. Graph. Forum4
2016 Visual saliency guided textured model simplification
Bailin Yang, Frederick W. B. Li, Xun Wang 0007, Mingliang Xu 0001, Xiaohui Liang 0001, Zhaoyi Jiang, Yanhui Jiang
Vis. Comput.2
2014 Failure rates in introductory programming revisited
abstract
Whilst working on an upcoming meta-analysis that synthesized fifty years of research on predictors of programming performance, we made an interesting discovery. Despite several studies citing a motivation for research as the high failure rates of introductory programming courses, to date, the majority of available evidence on this phenomenon is at best anecdotal in nature, and only a single study by Bennedsen and Caspersen has attempted to determine a worldwide pass rate of introductory programming courses.
Christopher Watson 0001, Frederick W. B. Li
ITiCSE2
2014 No tests required: comparing traditional and dynamic predictors of programming success
abstract
Research over the past fifty years into predictors of programming performance has yielded little improvement in the identification of at-risk students. This is possibly because research to date is based upon using static tests, which fail to reflect changes in a student's learning progress over time. In this paper, the effectiveness of 38 traditional predictors of programming performance are compared to 12 new data-driven predictors, that are based upon analyzing directly logged data, describing the programming behavior of students. Whilst few strong correlations were found between the traditional predictors and performance, an abundance of strong significant correlations based upon programming behavior were found. A model based upon two of these metrics (Watwin score and percentage of lab time spent resolving errors) could explain 56.3% of the variance in coursework results. The implication of this study is that a student's programming behavior is one of the strongest indicators of their performance, and future work should continue to explore such predictors in different teaching contexts.
Christopher Watson 0001, Frederick W. B. Li, Jamie L. Godwin
SIGCSE2
2014 A Fine-Grained Outcome-Based Learning Path Model
abstract
A learning path (or curriculum sequence) comprises steps for guiding a student to effectively build up knowledge and skills. Assessment is usually incorporated at each step for evaluating student learning progress. SCORM and IMS-LD have been established to define data structures for supporting systematic learning path construction. Although IMS-LD includes the concept of learning activity, no facilities are offered to help define its semantics, and pedagogy cannot be properly formulated. In addition, most existing work for learning path generation is content-based. They only focus on what learning content is delivered at each learning path step, and pedagogy is not incorporated. Such modeling limits the assessment of student learning outcome only by the mastery level of learning content. Other forms of assessments, such as generic skills, cannot be supported. In this paper, we propose a fine-grained outcome-based learning path model allowing learning activities and their assessment criteria to be formulated by Bloom's Taxonomy. Therefore, pedagogy can be explicitly defined and reused. Our model also supports the assessment of both subject content and generic skills related learning outcomes, providing more comprehensive student progress guidance and evaluation.
Fan Yang 0020, Frederick W. B. Li, Rynson W. H. Lau
IEEE Trans. Syst. Man Cybern. Syst.2
2014 Recent development in multimedia e-learning technologies
Rynson W. H. Lau, Neil Y. Yen, Frederick W. B. Li, Benjamin W. Wah
World Wide Web3
2013 Predicting Performance in an Introductory Programming Course by Logging and Analyzing Student Programming Behavior
abstract
The high failure rates of many programming courses means there is a need to identify struggling students as early as possible. Prior research has focused upon using a set of tests to assess the use of a student's demographic, psychological and cognitive traits as predictors of performance. But these traits are static in nature, and therefore fail to encapsulate changes in a student's learning progress over the duration of a course. In this paper we present a new approach for predicting a student's performance in a programming course, based upon analyzing directly logged data, describing various aspects of their ordinary programming behavior. An evaluation using data logged from a sample of 45 programming students at our University, showed that our approach was an excellent early predictor of performance, explaining 42.49% of the variance in coursework marks - double the explanatory power when compared to the closest related technique in the literature.
Christopher Watson 0001, Frederick W. B. Li, Jamie L. Godwin
ICALT2
2013 Utilizing Massive Spatiotemporal Samples for Efficient and Accurate Trajectory Prediction
abstract
Trajectory prediction is widespread in mobile computing, and helps support wireless network operation, location-based services, and applications in pervasive computing. However, most prediction methods are based on very coarse geometric information such as visited base transceiver stations, which cover tens of kilometers. These approaches undermine the prediction accuracy, and thus restrict the variety of application. Recently, due to the advance and dissemination of mobile positioning technology, accurate location tracking has become prevalent. The prediction methods based on precise spatiotemporal information are then possible. Although the prediction accuracy can be raised, a massive amount of data gets involved, which is undoubtedly a huge impact on network bandwidth usage. Therefore, employing fine spatiotemporal information in an accurate prediction must be efficient. However, this problem is not addressed in many prediction methods. Consequently, this paper proposes a novel prediction framework that utilizes massive spatiotemporal samples efficiently. This is achieved by identifying and extracting the information that is beneficial to accurate prediction from the samples. The proposed prediction framework circumvents high bandwidth consumption while maintaining high accuracy and being feasible. The experiments in this study examine the performance of the proposed prediction framework. The results show that it outperforms other popular approaches.
Addison Chan, Frederick W. B. Li
IEEE Trans. Mob. Comput.2
2012 Feature-varying skeletonization - Intuitive control over the target feature size and output skeleton topology
Chris G. Willcocks, Frederick W. B. Li
Vis. Comput.2
2011 The third ACM international workshop on multimedia technologies for distance learning (MTDL 2011)
abstract
The MTDL 2011 workshop in its third edition aims to continue in the contribution and evaluation of the impact of multimedia technologies to e-Learning. This workshop is held in conjunction with the ACM Multimedia 2011 Conference in Scottsdale, Arizona, U.S.A. As a cover paper of this workshop, we briefly summarize important issues to be addressed in e-learning in the first section, followed by a discussion of important issues proposed in the 5 papers accepted to the workshop (among the 9 submissions), plus 3 invited papers.
Rynson W. H. Lau, Timothy K. Shih, Frederick W. B. Li, Neil Y. Yen
ACM Multimedia3
2011 Game-on-demand: : An online game engine based on geometry streaming
abstract
In recent years, online gaming has become very popular. In contrast to stand-alone games, online games tend to be large-scale and typically support interactions among users. However, due to the high network latency of the Internet, smooth interactions among the users are often difficult. The huge and dynamic geometry data sets also make it difficult for some machines, such as handheld devices, to run those games. These constraints have stimulated some research interests on online gaming, which may be broadly categorized into two areas: technological support and user-perceived visual quality . Technological support concerns the performance issues while user-perceived visual quality concerns the presentation quality and accuracy of the game. In this article, we propose a game-on-demand engine that addresses both research areas. The engine distributes game content progressively to each client based on the player's location in the game scene. It comprises a two-level content management scheme and a prioritized content delivery scheme to help identify and deliver relevant game content at appropriate quality to each client dynamically. To improve the effectiveness of the prioritized content delivery scheme, it also includes a synchronization scheme to minimize the location discrepancy of avatars (game players). We demonstrate the performance of the proposed engine through numerous experiments.
Frederick W. B. Li, Rynson W. H. Lau, Danny Kilis, Lewis W. F. Li
ACM Trans. Multim. Comput. Commun. Appl.1
2011 A mobile environment for sketching-based skeleton generation
Qingzheng Zheng, Frederick W. B. Li
World Wide Web2
2008 An Effective Error Resilient Packetization Scheme for Progressive Mesh Transmission over Unreliable Networks
Bailin Yang, Frederick W. B. Li, Xun Wang 0007
J. Comput. Sci. Technol.2
2008 Technology supports for distributed and collaborative learning over the internet
abstract
With the advent of Internet and World Wide Web (WWW) technologies, distance education (e-learning or Web-based learning) has enabled a new era of education. There are a number of issues that have significant impact on distance education, including those from educational, sociological, and psychological perspectives. Rather than attempting to cover exhaustively all the related perspectives, in this survey article, we focus on the technological issues. A number of technology issues are discussed, including distributed learning, collaborative learning, distributed content management, mobile and situated learning, and multimodal interaction and augmented devices for e-learning. Although we have tried to include the state-of-the-art technologies and systems here, it is anticipated that many new ones will emerge in the near future. As such, we point out several emerging issues and technologies that we believe are promising, for the purpose of highlighting important directions for future research.
Qing Li 0001, Rynson W. H. Lau, Timothy K. Shih, Frederick W. B. Li
ACM Trans. Internet Techn.4
2006 Efficient rendering of deformable objects for real-time applications
abstract
Abstract Deformable objects can be used to model soft objects such as clothing, human faces and animal characters. They are important as they can improve the realism of the applications. However, most existing hardware accelerators cannot render deformable objects directly. A tessellation process is often used to convert a deformable object into polygons so that the hardware graphics accelerator may render them. Unfortunately, this tessellation process is computationally very expensive. While the object is deforming, the tessellation process needs to be performed repeatedly to convert the deforming objects into polygons. As a result, deformable objects are seldom used in real‐time applications such as virtual environments and computer games. Since trimmed NURBS surfaces are often used to represent deformable objects, in this paper we present an efficient method for incremental rendering of deformable trimmed NURBS surfaces. A trimmed NURBS surface typically deforms through the deformation of the trimmed NURBS surface and/or the trimming curve. Our method handles both trimmed surface deformation as well as trimming curve deformation. Experimental results show that our method performs significantly faster than the method used in OpenGL and can be used in real‐time applications, such as computer games. Copyright © 2006 John Wiley & Sons, Ltd.
Gary K. L. Cheung, Rynson W. H. Lau, Frederick W. B. Li
Comput. Animat. Virtual Worlds3
2006 A Trajectory-Preserving Synchronization Method for Collaborative Visualization
abstract
In the past decade, a lot of research work has been conducted to support collaborative visualization among remote users over the networks, allowing them to visualize and manipulate shared data for problem solving. There are many applications of collaborative visualization, such as oceanography, meteorology and medical science. To facilitate user interaction, a critical system requirement for collaborative visualization is to ensure that remote users will perceive a synchronized view of the shared data. Failing this requirement, the user's ability in performing the desirable collaborative tasks will be affected. In this paper, we propose a synchronization method to support collaborative visualization. It considers how interaction with dynamic objects is perceived by application participants under the existence of network latency, and remedies the motion trajectory of the dynamic objects. It also handles the false positive and false negative collision detection problems. The new method is particularly well designed for handling content changes due to unpredictable user interventions or object collisions. We demonstrate the effectiveness of our method through a number of experiments.
Lewis W. F. Li, Frederick W. B. Li, Rynson W. H. Lau
IEEE Trans. Vis. Comput. Graph.2
2005 Multiserver support for large-scale distributed virtual environments
abstract
CyberWalk is a distributed virtual walkthrough system that we have developed. It allows users at different geographical locations to share information and interact within a shared virtual environment (VE) via a local network or through the Internet. In this paper, we illustrate that as the number of users exploring the VE increases, the server will quickly become the bottleneck. To enable good performance, CyberWalk utilizes multiple servers and employs an adaptive region partitioning technique to dynamically partition the whole VE into regions. All objects within each region will be managed by one server. Under normal circumstances, when a viewer is exploring a region, the server of that region will be responsible for serving all requests from the viewer. When a viewer is crossing the boundary of two or more regions, the servers of all the regions involved will be serving requests from the viewer since the viewer might be able to view objects within all these regions. This is analogous to evaluating a database query using a parallel database server, which could improve the performance of serving a viewer's request tremendously. We evaluate the performance of this multiserver architecture of CyberWalk via a detail simulation model.
Beatrice Ng, Rynson W. H. Lau, Antonio Si, Frederick W. B. Li
IEEE Trans. Multim.4
2004 Supporting continuous consistency in multiplayer online games
abstract
Multiplayer online games have become very popular in recent years. However, they generally suffer from network latency problem. If a player changes its states, it will take some time before the changes are reflected to other concurrent players. This significantly affects the interactivity of the game. Sometimes, it may even cause disputes among the players. In this paper, we present a continuous consistency control mechanism to support collaborative game applications. Specifically, we propose a relaxed consistency control model for continuous events. Based on this model, we have developed a method to provide a global-wise continuous synchronization on the states of dynamic game objects presented among concurrent game players. We show the performance of the proposed method through some experiments.
Frederick W. B. Li, Lewis W. F. Li, Rynson W. H. Lau
ACM Multimedia1
2004 GameOD: an internet based game-on-demand framework
abstract
Multiplayer online 3D games are becoming very popular in recent years. However, existing games require the complete game content to be installed prior to game playing. Since the content is usually large in size, it may be difficult to run these games on a PDA or other handheld devices. It also pushes game companies to distribute their games as CDROMs/DVDROMs rather than online downloading. On the other hand, due to network latency, players may perceive discrepant status of some dynamic game objects. In this paper, we present a game-on-demand (GameOD) framework to distribute game content progressively in an on-demand manner. It allows critical contents to be available at the players' machines in a timely fashion. We present a simple distributed synchronization method to allow concurrent players to synchronize their perceived game status. Finally, we show some performance results of the proposed framework.
Frederick W. B. Li, Rynson W. H. Lau, Danny Kilis
VRST1
2003 Incremental rendering of deformable trimmed NURBS surfaces
abstract
Trimmed NURBS surfaces are often used to model smooth and complex objects. Unfortunately, most existing hardware graphics accelerators cannot render them directly. Although there are a lot of methods proposed to accelerate the rendering of such surfaces, majority of them are based on tessellation, which is developed primarily for handling non-deforming objects. For an object that may deform in run-time, such as clothing, facial expression, human and animal character, the tessellation process will need to be performed repeatedly while the object is deforming. However, as the tessellation process is very time consuming, interactive display of deforming objects is difficult. This explains why deformable objects are rarely used in virtual reality applications. In this paper, we present a efficient method for incremental rendering of deformable trimmed NURBS surfaces. This method can handle both trimmed surface deformation and trimming curve deformation. Experimental results show that our method performs significantly faster than the method used in OpenGL.
Gary K. L. Cheung, Rynson W. H. Lau, Frederick W. B. Li
VRST3
2003 A performance study on multi-server DVE systems
Beatrice Ng, Frederick W. B. Li, Rynson W. H. Lau, Antonio Si, Angus M. K. Siu
Inf. Sci.2
2003 VSculpt : a distributed virtual sculpting environment for collaborative design
abstract
A collaborative virtual sculpting system supports a team of geographically separated designers/engineers connected by networks to participate in designing three-dimensional (3D) virtual engineering tools or sculptures. It encourages international collaboration at a minimal cost. However, in order for the system to be useful, two factors need to be addressed: intuitiveness and real-time interaction. Although a lot of effort has been put into developing virtual sculpting environments, only limited work addresses collaborative virtual sculpting. This is because in order to support real-time collaborative virtual sculpting, many challenging issues need to be addressed. We propose a collaborative virtual sculpting framework, called VSculpt. Through adapting some techniques we developed earlier and integrating them with some techniques developed here, the proposed framework provides a real-time intuitive environment for collaborative design. In particular, it addresses issues on efficient rendering and transmission of deformable objects, intuitive object deformation using the CyberGlove and concurrent object deformation by multiple clients. We demonstrate and evaluate the performance of the proposed framework through a number of experiments.
Frederick W. B. Li, Rynson W. H. Lau, Frederick F. C. Ng
IEEE Trans. Multim.1
2002 Adaptive partitioning for multi-server distributed virtual environments
abstract
A distributed virtual environment (DVE) allows users at different geographical locations to share information and interact within a common virtual environment (VE) via a local network or through the Internet. However, when the number of users exploring the VE increases, the server will quickly become the bottleneck. To enable good performance, we are currently developing a multi-server DVE prototype. In this paper, we describe an adaptive data partitioning technique to dynamically partition the whole VE into regions. All objects within each region will be managed by a single server. As the loading of the servers changes, we show how it can be redistributed while minimizing the communication cost. Our initial results show that the proposed adaptive partitioning technique significantly improves the performance of the overall system.
Rynson W. H. Lau, Beatrice Ng, Antonio Si, Frederick W. B. Li
ACM Multimedia4
2002 LARGE a collision detection framework for deformable objects
abstract
Many collision detection methods have been proposed. Most of them can only be applied to rigid objects. In general, these methods precompute some geometric information of each object, such as bounding boxes, to be used for run-time collision detection. However, if the object deforms, the precomputed information may not be valid anymore and hence needs to be recomputed in every frame while the object is deforming. In this paper, we presents an efficient collision detection framework for deformable objects, which considers both inter-collisions and self-collisions of deformable objects modeled by NURBS surfaces. Towards the end of the paper, we show some experimental results to demonstrate the performance of the new method.
Rynson W. H. Lau, Oliver Chan, Mo Luk, Frederick W. B. Li
VRST4
2002 A multi-server architecture for distributed virtual walkthrough
abstract
CyberWalk is a distributed virtual walkthrough system that we have developed. It allows users at different geographical locations to share information and interact within a common virtual environment (VE) via a local network or through the Internet. In this paper, we illustrate that when the number of users exploring the VE increases, the server will quickly become the bottleneck. To enable good performance, CyberWalk utilizes multiple servers and employs an adaptive data partitioning techniques to dynamically partition the whole VE into regions. All objects within each region will be managed by one server. Under normal circumstances, when a viewer is exploring a region, the server of that region will be responsible for serving all requests from the viewer. When a viewer is crossing the boundary of two or more regions, the servers of all the regions involved will be serving requests from the viewer since the viewer might be able to view objects within all those regions. We evaluate the performance of this multi-server architecture of CyberWalk via a detail simulation model.
Beatrice Ng, Antonio Si, Rynson W. H. Lau, Frederick W. B. Li
VRST4
2001 NURBS Streams
abstract
Deformable objects are important in modeling clothing, facial expressions, animal characters and other soft objects. However, they are rarely used in distributed virtual environments. This is because both the transmission and the tessellation of deformable objects are too time-consuming to run in real time. Part of the problem can be solved by our earlier (1997, 1999) method on the real-time rendering of deformable NURBS (non-uniform rational B-spline) surfaces. In this paper, we present a technique to organize the data structures that are used to store the pre-computed polygon models and the deformation coefficients in supporting the rendering of deformable NURBS surfaces from a hierarchical form into a linear form, called NURBS streams. The NURBS streams not only support the progressive transmission and rendering of deformable NURBS surfaces, but also simplify the software and hardware implementation of the rendering method. To further speed up the rendering of these NURBS streams, we also introduce the idea of a virtual rendering list to keep track of incremental node changes between successive frames.
Frederick W. B. Li, Rynson W. H. Lau
Computer Graphics International1
2001 Collaborative Distributed Virtual Sculpting
abstract
A lot of effort is now being put into developing collaborative distributed virtual environments. However very few projects address collaborative virtual sculpting in which the shapes of the target objects are likely changing continuously. Some major issues including user interaction, data transmission, concurrent object editing by multiple clients and rendering of deforming objects must be addressed in a real-time context. We propose a framework for collaborative virtual sculpting in a distributed virtual environment. The system is based on a hybrid model which merges the client-server and the peer-to-peer architectures to allow fast data replication. To support real-time deformation and rendering of deformable objects, we model each of these objects using NURBS surfaces and render them using the real-time deformable NURBS rendering method that we have developed. We present a data structure for the transmission of these deformable objects. We also introduce the idea of editing region and the corresponding locking mechanism for simultaneous editing of the same object by multiple clients. Toward the end of the paper we show some performance results of the prototype system.
Frederick W. B. Li, Rynson W. H. Lau, Frederick F. C. Ng
VR1
1999 Real-time rendering of deformable parametric free-form surfaces
abstract
Deformable objects are required to improve the realism of virtual reality applications. They are particularly useful in modeling clothes, facial expression, human and animal characters. A common method to render these objects is by tessellation. However, the tessellation process is computationally very expensive. If the object deforms, we need to retessellate the surface every frame, as its shape changes from one frame to the next. This computational burden poses a significant challenge to the real-time rendering of deformable objects. Consequently, deformable objects are seldom incorporated in existing virtual reality systems. In this paper, we present an incremental method for rendering deformable objects modeled by parametric free-form surfaces. We also introduce two new frame coherence techniques for crack prevention and parameter caching. Finally, we present a single hierarchical data structure which provides a multi-resolution representation of the object model.
Frederick W. B. Li, Rynson W. H. Lau
VRST1
1997 Interactive Rendering of Deforming NURBS Surfaces
abstract
Non‐uniform rational B‐splines (NURBS) has been widely accepted as a standard tool for geometry representation and design. Its rich geometric properties allow it to represent both analytic shapes and free‐form curves and surfaces precisely. Moreover, a set of tools is available for shape modification or more implicitly, object deformation. Existing NURBS rendering methods include de Boor algorithm, Oslo algorithm, Shantz’s adaptive forward differencing algorithm and Silbermann’s high speed implementation of NURBS. However, these methods consider only speeding up the rendering process of individual frames. Recently, Kumar et al. proposed an incremental method for rendering NURBS surfaces, but it is still limited to static surfaces. In real‐time applications such as virtual reality, interactive display is needed. If a virtual environment contains a lot of deforming objects, these methods cannot provide a good solution. In this paper, we propose an efficient method for interactive rendering of deformable objects by maintaining a polygon model of each deforming NURBS surface and adaptively refining the resolution of the polygon model. We also look at how this method may be applied to multi‐resolution modelling.
Frederick W. B. Li, Rynson W. H. Lau, Mark Green 0001
Comput. Graph. Forum1