VLDB 2026 Research / reviewers in the wild / expert
Howard Leung
dblp:60/4971
· DBLP profile ↗
49ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-2633-2965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Low-Light Image Enhancement via Diffusion Models With Semantic Priors of Any Region
Lingyu Zhu 0006, Wenhan Yang, Howard Leung, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Unfolding High-Order Correlations for Interpretable Multi-Contrast MRI Super-ResolutionabstractDeep unfolding network has gained significant attention for magnetic resonance imaging super-resolution (MRI SR) due to its performance and interpretability. However, 1) existing methods predominantly focus on cross-contrast correlations while neglecting high-order correlations embedded within spatially adjacent slices in volumetric MRI data. 2) Their degradation models are optimized via the proximal gradient algorithm (PGA) that relies on manually designed hyperparameters (e.g., step size), often leading to overshooting or suboptimal solutions. To solve these limitations, we propose HocMRI, a deep unfolding multi-contrast MRI SR framework, which seamlessly integrates dual-prior modeling and hyperparameter-free PGA for enhanced reconstruction. Specifically, we first design a novel degradation model based on the dual-prior mechanism: an explicit prior based on low-rank tensor factorization to capture intra- and inter-slice dependencies, and an implicit prior leveraging a Mamba-based network with a novel 3D scanning strategy to further exploit high-order correlations across slices. Then, we derive a hyperparameter-free PGA to boost the traditional PGA, which employs a hyperbolic tangent function to dynamically control the gradient descent step, eliminating manual tuning while ensuring stable convergence with theoretical proofs. Based on the hyperparameter-free PGA, we develop an efficient iterative optimization algorithm to solve the degradation model and unfold it into a multi-stage deep network. Numerous experimental results from widely used MRI datasets demonstrate that our HocMRI achieves superior performance with enhanced efficiency compared to the state-of-the-art methods. Qiangqiang Shen, Xuanqi Zhang, Peilin Chen 0001, Zhiwei Zhong 0001, Howard Leung, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Deep Spatio-Temporal Network for Low-SNR Cryo-EM Movie Frame EnhancementabstractCryo-EM in single particle analysis is known to have low SNR and requires to utilize several frames of the same particle sample to restore one high-quality image for visualizing that particle. However, the low SNR of cryo-EM movie and motion caused by beam striking make the task very challenging. Video enhancement algorithms in computer vision shed new light on tackling such tasks by utilizing deep neural networks. However, they are designed for natural images with high SNR. Meanwhile, the lack of ground truth in cryo-EM movie seems to be one major limiting factor of the progress. Hence, we present a synthetic cryo-EM movie generation pipeline, which can produce realistic diverse cryo-EM movie datasets with low-SNR movie frames and multiple ground truth values. Then we propose a deep spatio-temporal network (DST-Net) for cryo-EM movie frame enhancement trained on our synthetic data. Spatial and temporal features are first extracted from each frame. Spatio-temporal fusion and high-resolution re-constructor are designed to obtain the enhanced output. For evaluation, we train our model on seven synthetic cryo-EM movie datasets and infer on real cryo-EM data. The experimental results show that DST-Net can achieve better enhancement performance both quantitatively and qualitatively compared with others. Xiaoya Chong, Howard Leung, Qing Li 0001, Jianhua Yao 0001, Niyun Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Video2mesh: 3D human pose and shape recovery by a temporal convolutional transformer networkabstractAbstract From a 2D video of a person in action, human mesh recovery aims to infer the 3D human pose and shape frame by frame. Despite progress on video‐based human pose and shape estimation, it is still challenging to guarantee high accuracy and smoothness simultaneously. To tackle this problem, we propose a Video2mesh, a temporal convolutional transformer (TConvTransformer) based temporal network which is able to recover accurate and smooth human mesh from 2D video. The temporal convolution block achieves the sequence‐level smoothness by aggregating image features from adjacent frames. The subsequent multi‐attention transformer improves the accuracy due to its multi‐subspace for better middle‐frame feature representation. Meanwhile, we add a TConvTransformer discriminator which is trained together with our 3D human mesh temporal encoder. This TConvTransformer discriminator further improves the accuracy and smoothness by restricting the pose and shape in a more reliable space based on the AMASS dataset. We conduct extensive experiments on three standard benchmark datasets and show that our proposed Video2mesh outperforms other state‐of‐the‐art methods in both accuracy and smoothness. Xianjin Chao, Zhipeng Ge, Howard Leung |
IET Comput. Vis. | 3 |
| 2023 | Focalized contrastive view-invariant learning for self-supervised skeleton-based action recognition
Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
Neurocomputing | 4 |
| 2022 | NoiseFlow: Learning Optical Flow from Low SNR Cryo-EM MovieabstractCryo-EM movie in single particle analysis has extremely low SNR and requires aligning multiple frames to achieve signal enhancement. Currently, signal processing technique is adopted to estimate the motion vector between a pair of cryo-EM movie frames at patch-level, and the estimated motion vector is used as the reference for frame alignment, whose accuracy will determine the resolution of the reconstructed 3D structure of the particle. The patch-level motion may not well represent the beam-induced motion of particles since particles in a patch move towards different directions due to beam striking. However, the low SNR of cryo-EM movie makes it difficult to estimate the motion of particles at pixel-level. Meanwhile, existing optical flow estimation models only consider the ideal case where high-quality videos are provided, which fail to obtain optical flow from cryo-EM movie. In this paper, we diminish this limitation by proposing a model called NoiseFlow, a deep learning network for optical flow estimation from low SNR cryo-EM movie. NoiseFlow makes use of the multi-frame stacking module and the denoising module to extract noise-invariant features, and then computes the correlation volume from noise-invariant features to learn optical flow. For evaluation, we train our model on two synthetic cryo-EM movie datasets and infer on real cryo-EM data. The experimental results illustrate that NoiseFlow achieves state-of-the-art performance on both synthetic and real cryo-EM datasets. Xiaoya Chong, Niyun Zhou, Qing Li 0001, Howard Leung |
ICPR | 4 |
| 2022 | GAN-based reactive motion synthesis with class-aware discriminators for human-human interaction
Qianhui Men, Hubert P. H. Shum, Edmond S. L. Ho, Howard Leung |
Comput. Graph. | 4 |
| 2022 | MP-NeRF: Neural Radiance Fields for Dynamic Multi-person synthesis from Sparse ViewsabstractAbstract Multi‐person novel view synthesis aims to generate free‐viewpoint videos for dynamic scenes of multiple persons. However, current methods require numerous views to reconstruct a dynamic person and only achieve good performance when only a single person is present in the video. This paper aims to reconstruct a multi‐person scene with fewer views, especially addressing the occlusion and interaction problems that appear in the multi‐person scene. We propose MP‐NeRF, a practical method for multi‐person novel view synthesis from sparse cameras without the pre‐scanned template human models. We apply a multi‐person SMPL template as the identity and human motion prior. Then we build a global latent code to integrate the relative observations among multiple people, so we could represent multiple dynamic people into multiple neural radiance representations from sparse views. Experiments on multi‐person dataset MVMP show that our method is superior to other state‐of‐the‐art methods. Xianjin Chao, Howard Leung |
Comput. Graph. Forum | 2 |
| 2021 | A Quadruple Diffusion Convolutional Recurrent Network for Human Motion PredictionabstractRecurrent neural network (RNN) has become popular for human motion prediction thanks to its ability to capture temporal dependencies. However, it has limited capacity in modeling the complex spatial relationship in the human skeletal structure. In this work, we present a novel diffusion convolutional recurrent predictor for spatial and temporal movement forecasting, with multi-step random walks traversing bidirectionally along an adaptive graph to model interdependency among body joints. In the temporal domain, existing methods rely on a single forward predictor with the produced motion deflecting to the drift route, which leads to error accumulations over time. We propose to supplement the forward predictor with a forward discriminator to alleviate such motion drift in the long term under adversarial training. The solution is further enhanced by a backward predictor and a backward discriminator to effectively reduce the error, such that the system can also look into the past to improve the prediction at early frames. The two-way spatial diffusion convolutions and two-way temporal predictors together form a quadruple network. Furthermore, we train our framework by modeling the velocity from observed motion dynamics instead of static poses to predict future movements that effectively reduces the discontinuity problem at early prediction. Our method outperforms the state of the arts on both 3D and 2D datasets, including the Human3.6M, CMU Motion Capture and Penn Action datasets. The results also show that our method correctly predicts both high-dynamic and low-dynamic moving trends with less motion drift. Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Adversarial Refinement Network for Human Motion Prediction
Xianjin Chao, Yanrui Bin, Wenqing Chu, Xuan Cao, Yanhao Ge, Chengjie Wang 0001, Feiyue Huang, Howard Leung |
ACCV (2) | 9 |
| 2020 | A Two-Stream Recurrent Network for Skeleton-based Human Interaction RecognitionabstractThis paper addresses the problem of recognizing human-human interaction from skeletal sequences. Existing methods are mainly designed to classify single human action. Many of them simply stack the movement features of two characters to deal with human interaction, while neglecting the abundant relationships between characters. In this paper, we propose a novel two-stream recurrent neural network by adopting the geometric features from both single actions and interactions to describe the spatial correlations with different discriminative abilities. The first stream is constructed under pairwise joint distance (PJD) in a fully-connected mesh to categorize the interactions with explicit distance patterns. To better distinguish similar interactions, in the second stream, we combine PJD with the spatial features from individual joint positions using graph convolutions to detect the implicit correlations among joints, where the joint connections in the graph are adaptive for flexible correlations. After spatial modeling, each stream is fed to a bi-directional LSTM to encode two-way temporal properties. To take advantage of the diverse discriminative power of the two streams, we come up with a late fusion algorithm to combine their output predictions concerning information entropy. Experimental results show that the proposed framework achieves state-of-the-art performance on 3D and comparable performance on 2D interaction datasets. Moreover, the late fusion results demonstrate the effectiveness of improving the recognition accuracy compared with single streams. Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
ICPR | 4 |
| 2020 | Hierarchical Visual-aware Minimax Ranking Based on Co-purchase Data for Personalized RecommendationabstractPersonalized recommendation aims at ranking a set of items according to the learnt preferences of the user. Existing methods optimize the ranking function by considering an item that the user has not bought yet as a negative item and assuming that the user prefers the positive item that he has bought to the negative item. The strategy is to exclude irrelevant items from the dataset to narrow down the set of potential positive items to improve ranking accuracy. It conflicts with the goal of recommendation from the seller’s point of view, which aims to enlarge that set for each user. In this paper, we diminish this limitation by proposing a novel learning method called Hierarchical Visual-aware Minimax Ranking (H-VMMR), in which a new concept of predictive sampling is proposed to sample items in a close relationship with the positive items (e.g., substitutes, compliments). We set up the problem by maximizing the preference discrepancy between positive and negative items, as well as minimizing the gap between positive and predictive items based on visual features. We also build a hierarchical learning model based on co-purchase data to solve the data sparsity problem. Our method is able to enlarge the set of potential positive items as well as true negative items during ranking. The experimental results show that our H-VMMR outperforms the state-of-the-art learning methods. Xiaoya Chong, Qing Li 0001, Howard Leung, Qianhui Men, Xianjin Chao |
WWW | 3 |
| 2019 | Multi-view depth-based pairwise feature learning for person-person interaction recognition
Meng Li 0021, Howard Leung |
Multim. Tools Appl. | 2 |
| 2019 | Action recognition from depth sequence using depth motion maps-based local ternary patterns and CNN
Zhifei Li 0002, Zhonglong Zheng, Feilong Lin, Howard Leung, Qing Li 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Retrieval of spatial-temporal motion topics from 3D skeleton data
Qianhui Men, Howard Leung |
Vis. Comput. | 2 |
| 2019 | Self-feeding frequency estimation and eating action recognition from skeletal representation using Kinect
Qianhui Men, Howard Leung, Yang Yang 0046 |
World Wide Web | 2 |
| 2018 | High-quality compatible triangulations and their application in interactive animation
Liuyang Zhou, Howard Leung, Hubert P. H. Shum |
Comput. Graph. | 3 |
| 2017 | Martial Arts, Dancing and Sports dataset: A challenging stereo and multi-view dataset for 3D human pose estimation
Liuyang Zhou, Howard Leung, Antoni B. Chan |
Image Vis. Comput. | 4 |
| 2017 | Graph-based approach for 3D human skeletal action recognition
Meng Li 0021, Howard Leung |
Pattern Recognit. Lett. | 2 |
| 2016 | Human action recognition via skeletal and depth based feature fusionabstractThis paper addresses the problem of recognizing human actions captured with depth cameras. Human action recognition is a challenging task as the articulated action data is high dimensional in both spatial and temporal domains. An effective approach to handle this complexity is to divide human body into different body parts according to human skeletal joint positions, and performs recognition based on these part-based feature descriptors. Since different types of features could share some similar hidden structures, and different actions may be well characterized by properties common to all features (sharable structure) and those specific to a feature (specific structure), we propose a joint group sparse regression-based learning method to model each action. Our method can mine the sharable and specific structures among its part-based multiple features meanwhile imposing the importance of these part-based feature structures by joint group sparse regularization, in favor of discriminative part-based feature structure selection. To represent the dynamics and appearance of the human body parts, we employ part-based multiple features extracted from skeleton and depth data respectively. Then, using the group sparse regularization techniques, we have derived an algorithm for mining the key part-based features in the proposed learning framework. The resulting features derived from the learnt weight matrices are more discriminative for multi-task classification. Through extensive experiments on three public datasets, we demonstrate that our approach outperforms existing methods. Meng Li 0021, Howard Leung, Hubert P. H. Shum |
MIG | 2 |
| 2016 | An interactive human morphing system with self-occlusion enhancementabstractPlanar shape morphing methods offer solutions to blend two shapes with different silhouettes. A naive method to solve the shape morphing problem is to linearly interpolate the coordinates of each corresponding vertex pair between the source and the target polygons. However, simple linear interpolation sometimes creates intermediate polygons that contain self-intersection, resulting in geometrically incorrect transformations. Howard Leung, Hubert P. H. Shum |
MIG | 2 |
| 2016 | 3D human motion retrieval using graph kernels based on adaptive graph construction
Meng Li 0021, Howard Leung, Liuyang Zhou |
Comput. Graph. | 2 |
| 2016 | Graph-based representation learning for automatic human motion segmentation
Meng Li 0021, Howard Leung |
Multim. Tools Appl. | 2 |
| 2016 | Multiview Skeletal Interaction Recognition Using Active Joint Interaction GraphabstractThis paper addresses the problem of recognizing human skeletal interactions using multiview data captured from depth sensors. The interactions among people are important cues for group and crowd human behavior analysis. In this paper, we focus on modeling the person-person skeletal interactions for human activity recognition. First, we propose a novel graph model in each single-view case to encode class-specific person-person interaction patterns. Particularly, we model each person-person interaction by an attributed graph, which is designed to preserve the complex spatial structure among skeletal joints according to their activity levels as well as the spatio-temporal joint features. Then, combining the graph models for each single-view case, we propose the multigraph model to characterize each multiview interaction. Finally, we apply a general multiple kernel learning method to determine the optimal kernel weights for the proposed multigraph model while the optimal classifier is jointly learned. We evaluate the proposed approach on the M2I dataset, the SBU Kinect interaction dataset, and our interaction dataset. The experimental results show that our proposed approach outperforms several existing interaction recognition methods. Meng Li 0021, Howard Leung |
IEEE Trans. Multim. | 2 |
| 2016 | Kinect Posture Reconstruction Based on a Local Mixture of Gaussian Process ModelsabstractDepth sensor based 3D human motion estimation hardware such as Kinect has made interactive applications more popular recently. However, it is still challenging to accurately recognize postures from a single depth camera due to the inherently noisy data derived from depth images and self-occluding action performed by the user. In this paper, we propose a new real-time probabilistic framework to enhance the accuracy of live captured postures that belong to one of the action classes in the database. We adopt the Gaussian Process model as a prior to leverage the position data obtained from Kinect and marker-based motion capture system. We also incorporate a temporal consistency term into the optimization framework to constrain the velocity variations between successive frames. To ensure that the reconstructed posture resembles the accurate parts of the observed posture, we embed a set of joint reliability measurements into the optimization framework. A major drawback of Gaussian Process is its cubic learning complexity when dealing with a large database due to the inverse of a covariance matrix. To solve the problem, we propose a new method based on a local mixture of Gaussian Processes, in which Gaussian Processes are defined in local regions of the state space. Due to the significantly decreased sample size in each local Gaussian Process, the learning time is greatly reduced. At the same time, the prediction speed is enhanced as the weighted mean prediction for a given sample is determined by the nearby local models only. Our system also allows incrementally updating a specific local Gaussian Process in real time, which enhances the likelihood of adapting to run-time postures that are different from those in the database. Experimental results demonstrate that our system can generate high quality postures even under severe self-occlusion situations, which is beneficial for real-time applications such as motion-based gaming and sport training. Liuyang Zhou, Howard Leung, Hubert P. H. Shum |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2015 | High quality compatible triangulations for 2D shape morphingabstractWe propose a new method to compute compatible triangulations of two polygons in order to create a smooth geometric transformation between them. Compared with existing methods, our approach creates triangulations of better quality, that is, triangulations with fewer long thin triangles and Steiner points. This results in visually appealing morphing when transforming the shape from one to another. Our method consists of three stages. First, we compatibly decompose the target and source polygons into a set of sub-polygons, in which each source sub-polygon is triangulated. Second, we map the triangulation of a source sub-polygon onto the corresponding sub-polygon of the target polygon using linear transformation, thereby generating the compatible meshes between the source and the target. Third, we refine the compatible meshes, which can create better quality planar shape morphing with detailed textures. Experimental results show that our method can create compatible meshes of higher quality compared with existing methods, which facilitates smoother morphing process. The proposed algorithm is robust and computationally efficient. It can be applied to produce convincing transformations such as interactive 2D animation creation and special effects in movies. Howard Leung, Liuyang Zhou, Hubert P. H. Shum |
VRST | 2 |
| 2015 | EEG Activity During Movement Planning Encodes Upcoming Peak Speed and Acceleration and Improves the Accuracy in Predicting Hand KinematicsabstractThe relationship between movement kinematics and human brain activity is an important and fundamental question for the development of neural prosthesis. The peak velocity and the peak acceleration could best reflect the feedforward-type movement; thus, it is worthwhile to investigate them further. Most related studies focused on the correlation between kinematics and brain activity during the movement execution or imagery. However, human movement is the result of the motor planning phase as well as the execution phase and researchers have demonstrated that statistical correlations exist between EEG activity during the motor planning and the peak velocity and the peak acceleration using grand-average analysis. In this paper, we examined whether the correlations were concealed in trial-to-trial decoding from the low signal-to-noise ratio of EEG activity. The alpha and beta powers from the movement planning phase were combined with the alpha and beta powers from the movement execution phase to predict the peak tangential speed and acceleration. The results showed that EEG activity from the motor planning phase could also predict the peak speed and the peak acceleration with a reasonable accuracy. Furthermore, the decoding accuracy of the peak speed and the peak acceleration could both be improved by combining band powers from the motor planning phase with the band powers from the movement execution. Lingling Yang, Howard Leung, Markus Plank, Joseph Snider, Howard Poizner |
IEEE J. Biomed. Health Informatics | 2 |
| 2014 | A Tablet -based Chinese Composition Assessment System
Kat Leung, Barley Mak, Howard Leung |
ICCE | 3 |
| 2014 | Posture reconstruction using Kinect with a probabilistic modelabstractRecent work has shown that depth image based 3D posture estimation hardware such as Kinect has made interactive applications more popular. However, it is still challenging to accurately recognize postures from a single depth camera due to the inherently noisy data derived from depth images and self-occluding action performed by the user. While previous research has shown that data-driven methods can be used to reconstruct the correct postures, they usually require a large posture database, which greatly limit the usability for systems with constrained hardware such as game console. To solve this problem, we present a new probabilistic framework to enhance the accuracy of the postures live captured by Kinect. We adopt the Gaussian Process model as a prior to leverage position data obtained from Kinect and marker-based motion capture system. We also incorporate a temporal consistency term into the optimization framework to constrain the velocity variations between successive frames. To ensure that the reconstructed posture resembles the observed input data from Kinect when its tracking result is good, we embed joint reliability into the optimization framework. Experimental results demonstrate that our system can generate high quality postures even under severe self-occlusion situations, which is beneficial for real-time posture based applications such as motion-based gaming and sport training. Liuyang Zhou, Howard Leung, Hubert P. H. Shum |
VRST | 3 |
| 2014 | Human motion variation synthesis with multivariate Gaussian processesabstractABSTRACT Human motion variation synthesis is important for crowd simulation and interactive applications to enhance synthesis quality. In this paper, we propose a novel generative probabilistic model to synthesize variations of human motion. Our key idea is to model the conditional distribution of each joint via a multivariate Gaussian process model, namely semiparametric latent factor model (SLFM). SLFM can effectively model the correlations between degrees of freedom (DOFs) of joints rather than dealing with each DOF separately as implemented in existing methods. A detailed evaluation is performed to show that the proposed approach can effectively synthesize variations of different types of motions. Motions generated by our method show a richer variation compared with existing ones. Finally, our user study shows that the synthesized motion has a similar level of naturalness to captured human motions. Our method is best applied in computer games and animations to introduce motion variations. Copyright © 2014 John Wiley & Sons, Ltd. Liuyang Zhou, Lifeng Shang, Hubert P. H. Shum, Howard Leung |
Comput. Animat. Virtual Worlds | 4 |
| 2014 | Spatial temporal pyramid matching using temporal sparse representation for human motion retrieval
Liuyang Zhou, Zhiwu Lu 0001, Howard Leung, Lifeng Shang |
Vis. Comput. | 3 |
| 2013 | Human Motion Retrieval Based on Sparse Coding and Touchless InteractionsabstractTo search for a particular motion from a large database, a user-friendly and efficient retrieval mechanism is essential. In this paper, we propose a human motion retrieval system based on sparse coding and touch less interactions. Compared with existing methods that involve vector quantization, sparse coding leads to a more compact and discriminative representation. Motion comparison based on sparse representation greatly improves the effectiveness of motion retrieval in terms of accuracy and speed. With the recent advancement in human motion tracking hardware such as Kinect, our retrieval system allows the user to specify the query motion by performing it directly. Besides, the user interacts with the retrieval system interactively using gestures so no controller is required thus delivering a natural user interface. Liuyang Zhou, Howard Leung |
CAD/Graphics | 2 |
| 2013 | Synthesizing Two-character Interactions by Merging Captured Interaction Samples with their Spacetime RelationshipsabstractAbstract Existing synthesis methods for closely interacting virtual characters relied on user‐specified constraints such as the reaching positions and the distance between body parts. In this paper, we present a novel method for synthesizing new interacting motion by composing two existing interacting motion samples without the need to specify the constraints manually. Our method automatically detects the type of interactions contained in the inputs and determines a suitable timing for the interaction composition by analyzing the spacetime relationships of the input characters. To preserve the features of the inputs in the synthesized interaction, the two inputs will be aligned and normalized according to the relative distance and orientation of the characters from the inputs. With a linear optimization method, the output is the optimal solution to preserve the close interaction of two characters and the local details of individual character behavior. The output animations demonstrated that our method is able to create interactions of new styles that combine the characteristics of the original inputs. Jacky C. P. Chan, Jeff K. T. Tang, Howard Leung |
Comput. Graph. Forum | 3 |
| 2013 | Interactive partner control in close interactions for real-time applicationsabstractThis article presents a new framework for synthesizing motion of a virtual character in response to the actions performed by a user-controlled character in real time. In particular, the proposed method can handle scenes in which the characters are closely interacting with each other such as those in partner dancing and fighting. In such interactions, coordinating the virtual characters with the human player automatically is extremely difficult because the system has to predict the intention of the player character. In addition, the style variations from different users affect the accuracy in recognizing the movements of the player character when determining the responses of the virtual character. To solve these problems, our framework makes use of the spatial relationship-based representation of the body parts called interaction mesh, which has been proven effective for motion adaptation. The method is computationally efficient, enabling real-time character control for interactive applications. We demonstrate its effectiveness and versatility in synthesizing a wide variety of motions with close interactions. Edmond S. L. Ho, Jacky C. P. Chan, Taku Komura, Howard Leung |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2012 | Generalized Model-Based Human Motion Recognition with Body Partition Index MapsabstractAbstract Content‐based human motion analysis has captured extensive concerns of researchers from the domains of computer animation, human‐machine interaction, entertainment, etc. However, it is a non‐trivial task due to the spatial and temporal variations in the motion data. In this paper, we propose a generalized model (GM)‐based approach to model the variations and accurately recognize motion patterns. We partition the human character model into five parts, and extract the features of the submotions of each specific body part using clustering techniques. These features from the training trials in each class are combined to build the GM. We propose a new penalty based similarity measure for DTW to be used with the GMs for isolated motion recognition. On the other hand, from the GMs five body partition index maps are constructed and used for matching together with a flexible end point detection scheme during continuous motion recognition. In the experiments, we examine the effectiveness and efficiency of the approach in both isolated motion and continuous motion recognition. The results show that our proposed method has good performance compared with other state‐of‐the‐art methods in recognition accuracy and processing speed. Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046 |
Comput. Graph. Forum | 2 |
| 2012 | Interaction Retrieval by Spacetime Proximity GraphsabstractAbstract In this paper, we propose a new method to index and retrieve animation scenes in which multiple characters closely interact with one another. Such a technique can be an important tool for animators when they want to automatically extract the desired scene from a large database of animation sequence. Existing methods for single character movements do not scale well for multiple characters as they do not take into account the interaction of different body parts. In this paper, we propose a new distance function that computes the similarity of two‐character interations using the spatial relationship of the body parts. For each interaction, we produce a time‐varying graph structure based on the proximity of different joints, and compute the similarity of interactions by comparing the topology and Laplacian coordinates of the time‐varying graph. Experimental results show that the proposed method outperforms previous methods which are based on the kinematics of individual characters. The top retrieved samples are found similar in high level semantics while containing style variations. Jeff K. T. Tang, Jacky C. P. Chan, Howard Leung, Taku Komura |
Comput. Graph. Forum | 3 |
| 2012 | Retrieval of logically relevant 3D human motions by Adaptive Feature Selection with Graded Relevance Feedback
Jeff K. T. Tang, Howard Leung |
Pattern Recognit. Lett. | 2 |
| 2011 | Real-time mocap dance recognition for an interactive dancing gameabstractAbstract In this paper, we present an interactive dancing game based on motion capture technology. We address the problem of real‐time recognition of the user's live dance performance in order to determine the interactive motion to be rendered by a virtual dance partner. The real‐time recognition algorithm is based on a human body partition indexing scheme with flexible matching to determine the end of a move as well as to detect unwanted motion. We show that the system can recognize the live dance motions of users with good accuracy and render the interactive dance move of the virtual partner. Copyright © 2011 John Wiley & Sons, Ltd. Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046 |
Comput. Animat. Virtual Worlds | 2 |
| 2010 | Recognizing Dance Motions with Segmental SVDabstractIn this paper, a novel concept of segmental singular value decomposition (SegSVD) is proposed to represent a motion pattern with a hierarchical structure. The similarity measure based on the SegSVD representation is also proposed. SegSVD is capable of capturing the temporal information of the time series. It is effective in matching patterns in a time series in which the start and end points of the patterns are not known in advance. We evaluate the performance of our method on both isolated motion classification and continuous motion recognition for dance movements. Experiments show that our method outperforms existing work in terms of higher recognition accuracy. Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046 |
ICPR | 2 |
| 2010 | Automated Recognition of Sequential Patterns in Captured Motion Streams
Liqun Deng, Howard Leung, Naijie Gu, Yang Yang 0046 |
WAIM | 2 |
| 2009 | ACM 2009 workshop on ambient media computing (AMC'09) overviewabstractNo abstract available. Howard Leung, Cha Zhang, Qing Li 0001, Rynson W. H. Lau, Benjamin W. Wah, Abdulmotaleb El Saddik, K. Selçuk Candan, Irene Cheng 0001 |
ACM Multimedia | 1 |
| 2008 | Model-based analysis of Chinese calligraphy images
Tak-Sum Wong, Howard Leung, Horace Ho-Shing Ip |
Comput. Vis. Image Underst. | 2 |
| 2008 | Emulating human perception of motion similarityabstractAbstract Evaluating the similarity of motions is useful for motion retrieval, motion blending, and performance analysis of dancers and athletes. Euclidean distance between corresponding joints has been widely adopted in measuring similarity of postures and hence motions. However, such a measure does not necessarily conform to the human perception of motion similarity. In this paper, we propose a new similarity measure based on machine learning techniques. We make use of the results of questionnaires from subjects answering whether arbitrary pairs of motions appear similar or not. Using the relative distance between the joints as the basic features, we train the system to compute the similarity of arbitrary pair of motions. Experimental results show that our method outperforms methods based on Euclidean distance between corresponding joints. Our method is applicable to content‐based motion retrieval of human motion for large‐scale database systems. It is also applicable to e‐Learning systems which automatically evaluates the performance of dancers and athletes by comparing the subjects' motions with those by experts. Copyright © 2008 John Wiley & Sons, Ltd. Jeff K. T. Tang, Howard Leung, Taku Komura, Hubert P. H. Shum |
Comput. Animat. Virtual Worlds | 2 |
| 2006 | Teaching Chinese Handwriting by Automatic Feedback and Analysis for Incorrect Stroke Sequence and Stroke Production Errors
Jeff K. T. Tang, Howard Leung |
ICCE | 2 |
| 2006 | Fitting Ellipses to a Region with Application in Calligraphic Stroke ReconstructionabstractGiven a region, it is a challenge to find a set of primitive shapes such as rectangles, circles or ellipses to cover it. This is in fact a set-covering problem, which is known to be NP-hard. The focus of this paper is on fitting a set of ellipses onto an image region. This problem was first formulated by identifying a number of criteria required for the ellipse fitting. A solution is then proposed for automatically determining the set of ellipses that best fits onto an image region. The proposed ellipse fitting algorithm has also been applied to strokes forming characters of Chinese calligraphic artwork. The results show that our proposed algorithm generates ellipses fitting onto stroke regions and capturing the characteristics of the strokes during turning, tilting and back-trace. Tak-Sum Wong, Howard Leung, Horace Ho-Shing Ip |
ICIP | 2 |
| 2005 | A Feedback Controller for Biped Humanoids that Can Counteract Large Perturbations During GaitabstractIn this paper, we propose a new method for biped humanoids to compensate for large amounts of angular momentum induced by strong external perturbations applied to the body during gait motion. Such angular momentum can easily cause the humanoid to fall down onto the ground. We use an Angular Momentum inducing inverted Pendulum Model (AMPM), which is an enhanced version of the 3D linear inverted pendulum model to model the robot dynamics. Because the AMPM allows us to explicitly calculate the angular momentum generated by the ground reaction force, it is possible to calculate a counteracting motion that compensates for the angular momentum generated by external perturbations in real-time. Taku Komura, Howard Leung, Shunsuke Kudoh, James J. Kuffner |
ICRA | 2 |
| 2005 | Model-Based Analysis of Chinese Calligraphy ImagesabstractChinese fonts with smooth outlines and solid colouring have been produced for computer displays and printings for a long time. However, the aesthetic properties of the characters produced by calligraphers could not be simulated with these methods. Ip and Wong proposed a parameterised brush model that enables efficient generation of Chinese calligraphic writings such that the rendering is scalable in resolution and it allows high quality publishing. While this graphical model facilitates the synthesis of calligraphy writing given a set of writing parameters, for the inverse problem, a lot of user stroke manipulation is required to regenerate the model parameters given a calligraphic image. Consequently, an intelligent method is required to automatically determine the model parameters from images of Chinese calligraphy. This paper describes a methodology for automatically estimating the set of 3D geometric and dynamic writing parameters along a stroke trajectory from images of calligraphic writings. Tak-Sum Wong, Howard Leung, Horace Ho-Shing Ip |
IV | 2 |
| 2004 | Analysis of traditional Chinese seals and synthesis of personalized sealsabstractPeople have long used seals for various purposes. One practical use of seals is for authentication and a lot of research has been focused on this area. In ancient China, seals revealed one's identity and were commonly used by various classes of people. Traditional Chinese seals can also be considered as a form of art, for people to appreciate and to learn about the culture at different dynasties. Nowadays, at this digital age, we would like not only be able to represent the seals in digital format, but also we would like to use image processing techniques to help us better understand them so that although we are not experts we will still be able to produce good designs. In this paper, we propose two things: 1) an analysis method to better understand traditional Chinese seals so that we can describe them in a more semantic way rather than just representing each one as a bitmap image; and 2) a synthesis method to generate a new seal image given a handwritten character, while considering the information obtained from the analysis of the traditional Chinese seals Howard Leung |
ICME | 1 |
| 2004 | Animating reactive motions for biped locomotionabstractIn this paper, we propose a new method for simulating reactive motions for running or walking human figures. The goal is to generate realistic animations of how humans compensate for large external forces and maintain balance while running or walking. We simulate the reactive motions of adjusting the body configuration and altering footfall locations in response to sudden external disturbance forces on the body. With our proposed method, the user first imports captured motion data of a run or walk cycle to use as the primary motion. While executing the primary motion, an external force is applied to the body. The system automatically calculates a reactive motion for the center of mass and angular momentum around the center of mass using an enhanced version of the linear inverted pendulum model. Finally, the trajectories of the generalized coordinates that realize the precalculated trajectories of the center of mass, zero moment point, and angular momentum are obtained using constrained inverse kinematics. The advantage of our method is that it is possible to calculate reactive motions for bipeds that preserve dynamic balance during locomotion, which was difficult using previous techniques. We demonstrate our results on an application that allows a user to interactively apply external perturbations to a running or walking virtual human model. We expect this technique to be useful for human animations in interactive 3D systems such as games, virtual reality, and potentially even the control of actual biped robots. Taku Komura, Howard Leung, James J. Kuffner |
VRST | 2 |