EDBT 2026 Demo / reviewers in the wild / expert
Xiaohui Liang 0001
dblp:92/3528-1 · also Xiao-Hui Liang 0001
· DBLP profile ↗
68ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0001-6351-2538ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 1 first-author · 26 since 2021Artificial intelligence and machine learning · 21 · 13 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-temporal 3D Gaussian Splatting for Sparse-view Guided Scene UpdateabstractMaintaining consistent 3D scene representations over time is a significant challenge in computer vision. Updating 3D scenes from sparse-view observations is crucial for various real-world applications, including urban planning, disaster assessment, and historical site preservation, where dense scans are often unavailable or impractical. In this paper, we propose Cross-Temporal 3D Gaussian Splatting (Cross-Temporal 3DGS), a novel framework for efficiently reconstructing and updating 3D scenes across different time periods, using sparse images and previously captured scene priors. Our approach comprises three stages: 1) Cross-temporal camera alignment for estimating and aligning camera poses across different timestamps; 2) Interference-based confidence initialization to identify unchanged regions between timestamps, thereby guiding updates; and 3) Progressive cross-temporal optimization, which iteratively integrates historical prior information into the 3D scene to enhance reconstruction quality. Our method supports non-continuous capture, enabling not only updates using new sparse views to refine existing scenes, but also recovering past scenes from limited data with the help of current captures. Furthermore, we demonstrate the potential of this approach to achieve temporal changes using only sparse images, which can later be reconstructed into detailed 3D representations as needed. Experimental results show significant improvements over baseline methods in reconstruction quality and data efficiency, making this approach a promising solution for scene versioning, cross-temporal digital twins, and long-term spatial documentation. Zeyuan An, Yanghang Xiao, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001 |
AAAI | 5 |
| 2026 | Physics-based Hand-object Interaction via Control Force in Virtual RealityabstractPhysics-based hand-object interaction in VR/AR has been widely studied. Penetration-based dynamics, where the applied force is proportional to the degree of interpenetration, can effectively support hand-object grasping; however, they exhibit notable limitations when interacting with tiny objects and are restricted to a narrow range of interaction tasks. To address these issues, we introduce a control-based dynamics as an alternative. Specifically, we employ PD controllers to approximate the applied force based on the velocity and movements of the tracked hand. Unlike penetration-based methods, our approach relies solely on the hand’s movements for force computation, eliminating issues related to insufficient penetration distance and more accurately reflecting real-world physics. To evaluate the effectiveness of the proposed method, we conducted a series of user studies involving various interaction tasks and virtual objects, in comparison with state-of-the-art approaches. The results verify the effectiveness of our approach in hand-object interactions and demonstrate its significant positive sense of agency compared to prior works. Yue Ma 0035, Xiaohui Liang 0001 |
VR | 3 |
| 2026 | Uncertainty-aware calibrated 3D human motion forecasting with latent conformal prediction
Yue Ma 0035, Frederick W. B. Li, Xiaohui Liang 0001 |
Pattern Recognit. | 3 |
| 2026 | A comprehensive survey of action quality assessment: Method and benchmark
Kanglei Zhou, Ruizhi Cai, Hubert P. H. Shum, Xiaohui Liang 0001 |
Pattern Recognit. | 5 |
| 2026 | Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic ScenesabstractScenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world models, we introduce Dyn-HSI, the first cognitive architecture for dynamic human-scene interaction, which endows virtual humans with three humanoid components. (1) Vision (human eyes): we equip the virtual human with a Dynamic Scene-Aware Navigation, which continuously perceives changes in the surrounding environment and adaptively predicts the next waypoint. (2) Memory (human brain): we equip the virtual human with a Hierarchical Experience Memory, which stores and updates experiential data accumulated during training. This allows the model to leverage prior knowledge during inference for context-aware motion priming, thereby enhancing both motion quality and generalization. (3) Control (human body): we equip the virtual human with Human-Scene Interaction Diffusion Model, which generates high-fidelity interaction motions conditioned on multimodal inputs. To evaluate performance in dynamic scenes, we extend the existing static human-scene interaction datasets to construct a dynamic benchmark, Dyn-Scenes. We conduct extensive qualitative and quantitative experiments to validate Dyn-HSI, showing that our method consistently outperforms existing approaches and generates high-quality human-scene interaction motions in both static and dynamic settings. Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Uncertainty-aware Probabilistic 3D Human Motion Forecasting via Invertible Networksabstract3D human motion forecasting aims to enable autonomous applications. Estimating uncertainty for each prediction (i.e., confidence based on probability density or quantile) is essential for safety-critical contexts like human-robot collaboration to minimize risks. However, existing diverse motion fore-casting approaches struggle with uncertainty quantification due to implicit probabilistic representations hindering uncertainty modeling. We propose ProbHMI, which introduces invertible networks to parameterize poses in a disentangled latent space, enabling probabilistic dynamics modeling. A forecasting module then explicitly predicts future latent distributions, allowing effective uncertainty quantification. Evaluated on benchmarks, ProbHMI achieves strong performance for both deterministic and diverse prediction while validating uncertainty calibration, critical for risk-aware decision making. Yue Ma 0035, Kanglei Zhou, Fuyang Yu, Frederick W. B. Li, Xiaohui Liang 0001 |
ICRA | 5 |
| 2025 | Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | PHI: Bridging Domain Shift in Long-Term Action Quality Assessment via Progressive Hierarchical InstructionabstractLong-term Action Quality Assessment (AQA) aims to evaluate the quantitative performance of actions in long videos. However, existing methods face challenges due to domain shifts between the pre-trained large-scale action recognition backbones and the specific AQA task, thereby hindering their performance. This arises since fine-tuning resource-intensive backbones on small AQA datasets is impractical. We address this by identifying two levels of domain shift: task-level, regarding differences in task objectives, and feature-level, regarding differences in important features. For feature-level shifts, which are more detrimental, we propose Progressive Hierarchical Instruction (PHI) with two strategies. First, Gap Minimization Flow (GMF) leverages flow matching to progressively learn a fast flow path that reduces the domain gap between initial and desired features across shallow to deep layers. Additionally, a temporally-enhanced attention module captures long-range dependencies essential for AQA. Second, List-wise Contrastive Regularization (LCR) facilitates coarse-to-fine alignment by comprehensively comparing batch pairs to learn fine-grained cues while mitigating domain shift. Integrating these modules, PHI offers an effective solution. Experiments demonstrate that PHI achieves state-of-the-art performance on three representative long-term AQA datasets, proving its superiority in addressing the domain shift for long-term AQA. Kanglei Zhou, Hubert P. H. Shum, Frederick W. B. Li, Xingxing Zhang 0001, Xiaohui Liang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf InteractionabstractWe introduce MOST, a novel MOtion diffuSion model via Temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts. Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Adaptive Score Alignment Learning for Continual Perceptual Quality Assessment of 360-Degree Videos in Virtual RealityabstractVirtual Reality Video Quality Assessment (VR-VQA) aims to evaluate the perceptual quality of 360-degree videos, which is crucial for ensuring a distortion-free user experience. Traditional VR-VQA methods trained on static datasets with limited distortion diversity struggle to balance correlation and precision. This becomes particularly critical when generalizing to diverse VR content and continually adapting to dynamic and evolving video distribution variations. To address these challenges, we propose a novel approach for assessing the perceptual quality of VR videos, Adaptive Score Alignment Learning (ASAL). ASAL integrates correlation loss with error loss to enhance alignment with human subjective ratings and precision in predicting perceptual quality. In particular, ASAL can naturally adapt to continually changing distributions through a feature space smoothing process that enhances generalization to unseen content. To further improve continual adaptation to dynamic VR environments, we extend ASAL with adaptive memory replay as a novel Continual Learning (CL) framework. Unlike traditional CL models, ASAL utilizes key frame extraction and feature adaptation to address the unique challenges of non-stationary variations with both the computation and storage restrictions of VR devices. We establish a comprehensive benchmark for VR-VQA and its CL counterpart, introducing new data splits and evaluation metrics. Our experiments demonstrate that ASAL outperforms recent strong baseline models, achieving overall correlation gains of up to 4.78% in the static joint training setting and 12.19% in the dynamic CL setting on various datasets. This validates the effectiveness of ASAL in addressing the inherent challenges of VR-VQA. Our code is available at https://github.com/ZhouKanglei/ASAL_CVQA. Kanglei Zhou, Zikai Hao, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generationabstract3D shape generation from text is a fundamental task in 3D representation learning. The text-shape pairs exhibit a hierarchical structure, where a general text like “chair” covers all 3D shapes of the chair, while more detailed prompts refer to more specific shapes. Furthermore, both text and 3D shapes are inherently hierarchical structures. However, existing Text2Shape methods, such as SDFusion, do not exploit that. In this work, we propose HyperSD-Fusion, a dual-branch diffusion model that generates 3D shapes from a given text. Since hyperbolic space is suitable for handling hierarchical data, we propose to learn the hierarchical representations of text and 3D shapes in hyperbolic space. First, we introduce a hyperbolic text-image encoder to learn the sequential and multi-modal hierarchical features of text in hyperbolic space. In addition, we design a hyperbolic text-graph convolution module to learn the hierarchical features of text in hyperbolic space. In order to fully utilize these text features, we introduce a dual-branch structure to embed text features in 3D feature space. At last, to endow the generated 3D shapes with a hierarchical structure, we devise a hyperbolic hierarchical loss. Our method is the first to explore the hyperbolic hierarchical representation for text-to-shape generation. Experimental results on the existing text-to-shape paired dataset, Text2Shape, achieved state-of-the-art results. We release our implementation under HyperSDFusion.github.io. Zhiying Leng, Tolga Birdal, Xiaohui Liang 0001, Federico Tombari |
CVPR | 3 |
| 2024 | LightOctree: Lightweight 3D Spatially-Coherent Indoor Lighting EstimationabstractWe present a lightweight solution for estimating spatially-coherent indoor lighting from a single RGB image. Previous methods for estimating illumination using volumetric repre-sentations have overlooked the sparse distribution of light sources in space, necessitating substantial memory and computational resources for achieving high-quality results. We introduce a unified, voxel octree-based illumination estimation framework to produce 3D spatially-coherent lighting. Additionally, a differentiable voxel octree cone tracing ren-dering layer is proposed to eliminate regular volumetric representation throughout the entire process and ensure the retention of features across different frequency domains. This reduction significantly decreases spatial usage and required floating-point operations without substantially compromising precision. Experimental results demonstrate that our approach achieves high-quality coherent estimation with minimal cost compared to previous methods. Xuecan Wang, Shibang Xiao, Xiaohui Liang 0001 |
CVPR | 3 |
| 2024 | MAGR: Manifold-Aligned Graph Regularization for Continual Action Quality Assessment
Kanglei Zhou, Xingxing Zhang 0001, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
ECCV (11) | 7 |
| 2024 | CUS3D: Clip-Based Unsupervised 3D Segmentation via Object-Level DenoiseabstractTo ease the difficulty of acquiring annotation labels in 3D data, a common method is using unsupervised and open-vocabulary semantic segmentation, which leverage 2D CLIP semantic knowledge. In this paper, unlike previous research that ignores the "noise" raised during feature projection from 2D to 3D, we propose a novel distillation learning framework named CUS3D. In our approach, an object-level denosing projection module is designed to screen out the "noise" and ensure more accurate 3D feature. Based on the obtained features, a multimodal distillation learning module is designed to align the 3D feature with CLIP semantic feature space with object-centered constrains to achieve advanced unsupervised semantic segmentation. We conduct comprehensive experiments in both unsupervised and open-vocabulary segmentation, and the results consistently showcase the superiority of our model in achieving advanced unsupervised segmentation results and its effectiveness in open-vocabulary segmentation. Fuyang Yu, Runze Tian, Xiaohui Liang 0001 |
ICME | 5 |
| 2024 | CoFInAl: Enhancing Action Quality Assessment with Coarse-to-Fine Instruction Alignment
Kanglei Zhou, Ruizhi Cai, Xingxing Zhang 0001, Xiaohui Liang 0001 |
IJCAI | 6 |
| 2024 | Towards Cross-Modal Point Cloud Retrieval for Indoor Scenes
Fuyang Yu, Dongyuan Li, Peide Zhu, Xiaohui Liang 0001, Manabu Okumura |
MMM (4) | 5 |
| 2024 | Color Theme Evaluation through User Preference ModelingabstractColor composition (or color theme) is a key factor to determine how well a piece of art work or graphical design is perceived by humans. Despite a few color harmony models have been proposed, their results are often less satisfactory since they mostly neglect the variations of aesthetic cognition among individuals and treat the influence of all ratings equally as if they were all rated by the same anonymous user. To overcome this issue, in this article we propose a new color theme evaluation model by combining a back propagation neural network and a kernel probabilistic model to infer both the color theme rating and the user aesthetic preference. Our experiment results show that our model can predict more accurate and personalized color theme ratings than state of the art methods. Our work is also the first-of-its-kind effort to quantitatively evaluate the correlation between user aesthetic preferences and color harmonies of five-color themes, and study such a relation for users with different aesthetic cognition. Bailin Yang, Tianxiang Wei, Frederick W. B. Li, Xiaohui Liang 0001, Zhigang Deng 0001, Yili Fang |
ACM Trans. Appl. Percept. | 4 |
| 2024 | A Real-Time and Interactive Fluid Modeling System for Mixed RealityabstractWithin the realm of mixed reality, the capability to dynamically render environmental effects with high realism plays a crucial role in amplifying user engagement and interaction. Fluid dynamics, in particular, stand out as essential elements for crafting immersive virtual settings. This includes the simulation of phenomena like smoke, fire, and clouds, which are instrumental in enriching the virtual experience. This work showcases a cutting-edge system developed to produce dynamic and interactive fluid effects that mirror real captured data in real-time for mixed reality applications. This innovative system seamlessly incorporates fluid reconstruction alongside velocity estimation processes within the Unity engine environment. Our approach leverages a novel physics-based differentiable rendering technique, grounded in the principles of light transport in participating media, to simulate the intricate behaviors of fluid while ensuring high fidelity in visual appearance. To further enhance realism, we have expanded our framework to include the estimation of velocity fields, addressing the critical need for fluid motion simulation. The practical application of these techniques demonstrates the system's capacity to offer a robust platform for fluid modeling in mixed reality environments. Through extensive evaluations, we illustrate the effectiveness of our approach in various scenes, underscoring its potential to transform mixed reality content creation by providing developers with the tools to incorporate highly realistic and interactive fluid seamlessly. Yunchi Cen, Hanchen Deng, Yue Ma 0035, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Laplacian Projection Based Global Physical Prior Smoke ReconstructionabstractWe present a novel framework for reconstructing fluid dynamics in real-life scenarios. Our approach leverages sparse view images and incorporates physical priors across long series of frames, resulting in reconstructed fluids with enhanced physical consistency. Unlike previous methods, we utilize a differentiable fluid simulator (DFS) and a differentiable renderer (DR) to exploit global physical priors, reducing reconstruction errors without the need for manual regularization coefficients. We introduce divergence-free Laplacian eigenfunctions (div-free LE) as velocity bases, improving computational efficiency and memory usage. By employing gradient-related strategies, we achieve better convergence and superior results. Extensive experiments demonstrate the effectiveness of our method, showcasing improved reconstruction quality and computational efficiency compared to existing approaches. We validate our approach using both synthetic and real data, highlighting its practical potential. Shibang Xiao, Chao Tong 0001, Qifan Zhang 0003, Yunchi Cen, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Multi-Task Spatial-Temporal Graph Auto-Encoder for Hand Motion DenoisingabstractIn many human-computer interaction applications, fast and accurate hand tracking is necessary for an immersive experience. However, raw hand motion data can be flawed due to issues such as joint occlusions and high-frequency noise, hindering the interaction. Using only current motion for interaction can lead to lag, so predicting future movement is crucial for a faster response. Our solution is the Multi-task Spatial-Temporal Graph Auto-Encoder (Multi-STGAE), a model that accurately denoises and predicts hand motion by exploiting the inter-dependency of both tasks. The model ensures a stable and accurate prediction through denoising while maintaining motion dynamics to avoid over-smoothed motion and alleviate time delays through prediction. A gate mechanism is integrated to prevent negative transfer between tasks and further boost multi-task performance. Multi-STGAE also includes a spatial-temporal graph autoencoder block, which models hand structures and motion coherence through graph convolutional networks, reducing noise while preserving hand physiology. Additionally, we design a novel hand partition strategy and hand bone loss to improve natural hand motion generation. We validate the effectiveness of our proposed method by contributing two large-scale datasets with a data corruption algorithm based on two benchmark datasets. To evaluate the natural characteristics of the denoised and predicted hand motion, we propose two structural metrics. Experimental results show that our method outperforms the state-of-the-art, showcasing how the multi-task framework enables mutual benefits between denoising and prediction. Kanglei Zhou, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Dynamic Hyperbolic Attention Network for Fine Hand-object ReconstructionabstractReconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of meshes thanks to its exponentially-growing space distance, which amplifies the differences between the features based on similarity. In this work, we propose the first precise hand-object reconstruction method in hyperbolic space, namely Dynamic Hyperbolic Attention Network (DHANet), which leverages intrinsic properties of hyperbolic space to learn representative features. Our method that projects mesh and image features into a unified hyperbolic space includes two modules, i.e. dynamic hyperbolic graph convolution and image-attention hyperbolic graph convolution. With these two modules, our method learns mesh features with rich geometry-image multi-modal information and models better hand-object interaction. Our method provides a promising alternative for fine hand-object reconstruction in hyperbolic space. Extensive experiments on three public datasets demonstrate that our method outperforms most state-of-the-art methods. Zhiying Leng, Mahdi Saleh, Antonio Montanaro, Hao Yu 0010, Yin Wang 0005, Nassir Navab, Xiaohui Liang 0001, Federico Tombari |
ICCV | 8 |
| 2023 | Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion ModelabstractText-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text description. In this work, we propose a fine-grained method for generating high-quality, conditional human motion sequences supporting precise text description. Our approach consists of two key components: 1) a linguistics-structure assisted module that constructs accurate and complete language feature to fully utilize text information; and 2) a context-aware progressive reasoning module that learns neighborhood and overall semantic linguistics features from shallow and deep graph neural networks to achieve a multi-step inference. Experiments show that our approach outperforms text-driven motion generation methods on HumanML3D and KIT test sets and generates better visually confirmed motion to the text conditions. Yin Wang 0005, Zhiying Leng, Frederick W. B. Li, Xiaohui Liang 0001 |
ICCV | 5 |
| 2023 | PhyVR: Physics-based Multi-material and Free-hand Interaction in VRabstractThe realistic interaction with physical phenomena is a crucial aspect of human-computer interaction (HCI) in virtual reality (VR). However, the real-time performance of physical simulation, interactive computation, and rendering is the bottleneck of physics-based VR HCI. To address these challenges, we propose a novel physics-oriented framework for multi-material objects and free-hand interaction, termed PhyVR. This framework enables users to interact with diverse virtual phenomena dynamically. At the algorithm level, we develop a unified particle system to describe both the virtual multi-materials and the user’s avatar for the efficiency issue, optimize collision detection, and accelerate the HCI algorithms with a variable fine-coarse particle sampling scheme. At the rendering level, we introduce a hybrid particle-grid anisotropic algorithm for surface reconstruction, enabling real-time and visually convincing fluid rendering. Comprehensive experiments and user studies demonstrate that our framework effectively captures various physical interaction phenomena, providing an enhanced user experience and paving the way for expanding VR-related HCI applications. Hanchen Deng, Jin Li 0068, Yang Gao 0032, Xiaohui Liang 0001, Aimin Hao |
ISMAR | 4 |
| 2023 | A Mixed Reality Training System for Hand-Object Interaction in Simulated Microgravity EnvironmentsabstractAs human exploration of space continues to progress, the use of Mixed Reality (MR) for simulating microgravity environments and facilitating training in hand-object interaction holds immense practical significance. However, hand-object interaction in microgravity presents distinct challenges compared to terrestrial environments due to the absence of gravity. This results in heightened agility and inherent unpredictability of movements that traditional methods struggle to simulate accurately. To this end, we propose a novel MR-based hand-object interaction system in simulated microgravity environments, leveraging physics-based simulations to enhance the interaction between the user’s real hand and virtual objects. Specifically, we introduce a physics-based hand-object interaction model that combines impulse-based simulation with penetration contact dynamics. This accurately captures the intricacies of hand-object interaction in microgravity. By considering forces and impulses during contact, our model ensures realistic collision responses and enables effective object manipulation in the absence of gravity. The proposed system presents a cost-effective solution for users to simulate object manipulation in microgravity. It also holds promise for training space travelers, equipping them with greater immersion to better adapt to space missions. The system reliability and fidelity test verifies the superior effectiveness of our system compared to the state-of-the-art CLAP system. Kanglei Zhou, Yue Ma 0035, Zhiying Leng, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
ISMAR | 7 |
| 2023 | A Differential Diffusion Theory for Participating MediaabstractAbstract We present a novel approach to differentiable rendering for participating media, addressing the challenge of computing scene parameter derivatives. While existing methods focus on derivative computation within volumetric path tracing, they fail to significantly improve computational performance due to the expensive computation of multiply‐scattered light. To overcome this limitation, we propose a differential diffusion theory inspired by the classical diffusion equation. Our theory enables real‐time computation of arbitrary derivatives such as optical absorption, scattering coefficients, and anisotropic parameters of phase functions. By solving derivatives through the differential form of the diffusion equation, our approach achieves remarkable speed gains compared to Monte Carlo methods. This marks the first differentiable rendering framework to compute scene parameter derivatives based on diffusion approximation. Additionally, we derive the discrete form of diffusion equation derivatives, facilitating efficient numerical solutions. Our experimental results using synthetic and realistic images demonstrate the accurate and efficient estimation of arbitrary scene parameter derivatives. Our work represents a significant advancement in differentiable rendering for participating media, offering a practical and efficient solution to compute derivatives while addressing the limitations of existing approaches. Yunchi Cen, Frederick W. B. Li, Bailin Yang, Xiaohui Liang 0001 |
Comput. Graph. Forum | 5 |
| 2023 | Hierarchical Graph Convolutional Networks for Action Quality AssessmentabstractAction quality assessment (AQA) automatically evaluates how well humans perform actions in a given video, a technique widely used in fields such as rehabilitation medicine, athletic competitions, and specific skills assessment. However, existing works that uniformly divide the video sequence into small clips of equal length suffer from intra-clip confusion and inter-clip incoherence, hindering the further development of AQA. To address this issue, we propose a hierarchical graph convolutional network (GCN). First, semantic information confusion is corrected through clip refinement, generating the ‘shot’ as the basic action unit. We then construct a scene graph by combining several consecutive shots into meaningful scenes to capture local dynamics. These scenes can be viewed as different procedures of a given action, providing valuable assessment cues. The video-level representation is finally extracted via sequential action aggregation among scenes to regress the predicted score distribution, enhancing discriminative features and improving assessment performance. Experiments on the AQA-7, MTL-AQA, and JIGSAWS datasets demonstrate the superiority of the proposed hierarchical GCN over state-of-the-art methods. Kanglei Zhou, Yue Ma 0035, Hubert P. H. Shum, Xiaohui Liang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A Video-Based Augmented Reality System for Human-in-the-Loop Muscle Strength Assessment of Juvenile DermatomyositisabstractAs the most common idiopathic inflammatory myopathy in children, juvenile dermatomyositis (JDM) is characterized by skin rashes and muscle weakness. The childhood myositis assessment scale (CMAS) is commonly used to measure the degree of muscle involvement for diagnosis or rehabilitation monitoring. On the one hand, human diagnosis is not scalable and may be subject to personal bias. On the other hand, automatic action quality assessment (AQA) algorithms cannot guarantee 100% accuracy, making them not suitable for biomedical applications. As a solution, we propose a video-based augmented reality system for human-in-the-loop muscle strength assessment of children with JDM. We first propose an AQA algorithm for muscle strength assessment of JDM using contrastive regression trained by a JDM dataset. Our core insight is to visualize the AQA results as a virtual character facilitated by a 3D animation dataset, so that users can compare the real-world patient and the virtual character to understand and verify the AQA results. To allow effective comparisons, we propose a video-based augmented reality system. Given a feed, we adapt computer vision algorithms for scene understanding, evaluate the optimal way of augmenting the virtual character into the scene, and highlight important parts for effective human verification. The experimental results confirm the effectiveness of our AQA algorithm, and the results of the user study demonstrate that humans can more accurately and quickly assess the muscle strength of children using our system. Kanglei Zhou, Ruizhi Cai, Yue Ma 0035, Qingqing Tan, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 10 |
| 2023 | IAACS: Image Aesthetic Assessment Through Color Composition And Space FormationabstractJudging how an image is visually appealing is a complicated and subjective task. This highly motivates having a machine learning model to automatically evaluate image aesthetic by matching the aesthetics of general public. Although deep learning methods have been successfully learning good visual features from images, correctly assessing image aesthetic quality is still challenging for deep learning. To tackle this, we propose a novel multi-view convolutional neural network to assess image aesthetic by analyzing image color composition and space formation (IAACS). Specifically, from different views of an image, including its key color components with their contributions, the image space formation and the image itself, our network extracts their corresponding features through our proposed feature extraction module (FET) and the ImageNet weight-based classification model. By fusing the extracted features, our network produces an accurate prediction score distribution of image aesthetic. Experiment results have shown that we have achieved a superior performance. Bailin Yang, Changrui Zhu, Frederick W. B. Li, Tianxiang Wei, Xiaohui Liang 0001, Qingxu Wang |
Virtual Real. Intell. Hardw. | 5 |
| 2021 | STGAE: Spatial-Temporal Graph Auto-Encoder for Hand Motion DenoisingabstractHand object interaction in mixed reality (MR) relies on the accurate tracking and estimation of human hands, which provide users with a sense of immersion. However, raw captured hand motion data always contains errors such as joints occlusion, dislocation, high-frequency noise, and involuntary jitter. Denoising and obtaining the hand motion data consistent with the user’s intention are of the utmost importance to enhance the interactive experience in MR. To this end, we propose an end-to-end method for hand motion denoising using the spatial-temporal graph auto-encoder (STGAE). The spatial and temporal patterns are recognized simultaneously by constructing the consecutive hand joint sequence as a spatial-temporal graph. Considering the complexity of the articulated hand structure, a simple yet effective partition strategy is proposed to model the physic-connected and symmetry-connected relationships. Graph convolution is applied to extract structural constraints of the hand, and a self-attention mechanism is to adjust the graph topology dynamically. Combining graph convolution and temporal convolution, a fundamental graph encoder or decoder block is proposed. We finally establish the hourglass residual auto-encoder to learn a manifold projection operation and a corresponding inverse projection through stacking these blocks. In this work, the proposed framework has been successfully used in hand motion data denoising with preserving structural constraints between joints. Extensive quantitative and qualitative experiments show that the proposed method has achieved better performance than the state-of-the-art approaches. Kanglei Zhou, Zhiyuan Cheng 0004, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
ISMAR | 5 |
| 2021 | Stable Hand Pose Estimation under Tremor via Graph Neural NetworkabstractHand pose estimation, which predicts the spatial location of hand joints, is a fundamental task in VR/AR applications. Although existing methods can recover hand pose competently, the tremor issue occurring in hand motion has not been completely solved. Tremor is an involuntary motion accompanied by a desired gesture or hand motion, leading to hand pose that deviates from user's intentions. Considering the characteristic of tremor motion, we present a novel Graph Neural Network for stable 3D hand pose estimation. The input is depth images. The constraint adjacency matrix is devised in Graph Neural Network for dynamically adjusting the topology of a hand graph during message passing and aggregation. Firstly, since there are rich potential constraints among hand joints, we utilize the constraint adjacency matrix to mine the suitable topology, modeling spatial-temporal constraints of joints and outputting the precise tremor hand pose as the pre-estimation result. Then, for obtaining a stable hand pose, we provide a tremor compensation module based on the constraint adjacency matrix, which exploits the constraint between control points and tremor hand pose. Concretely, the control points represented the voluntary motion are employed as constraints to edit the tremor hand pose. Our extensive quantitative and qualitative experiments show that the proposed method has achieved decent performance for 3D tremor hand pose estimation. Zhiying Leng, Hubert P. H. Shum, Frederick W. B. Li, Xiaohui Liang 0001 |
VR | 5 |
| 2021 | Facial reshaping operator for controllable face beautification
Shanfeng Hu, Hubert P. H. Shum, Xiaohui Liang 0001, Frederick W. B. Li, Nauman Aslam |
Expert Syst. Appl. | 3 |
| 2021 | Cumulus cloud modeling from images based on VAE-GANabstractCumulus clouds are important elements in creating virtual outdoor scenes. Modeling cumulus clouds that have a specific shape is difficult owing to the fluid nature of the cloud. Image-based modeling is an efficient method to solve this problem. Because of the complexity of cloud shapes, the task of modeling the cloud from a single image remains in the development phase. In this study, a deep learning-based method was developed to address the problem of modeling 3D cumulus clouds from a single image. The method employs a three-dimensional autoencoder network that combines the variational autoencoder and the generative adversarial network. First, a 3D cloud shape is mapped into a unique hidden space using the proposed autoencoder. Then, the parameters of the decoder are fixed. A shape reconstruction network is proposed for use instead of the encoder part, and it is trained with rendered images. To train the presented models, we constructed a 3D cumulus dataset that included 200 3D cumulus models. These cumulus clouds were rendered under different lighting parameters. The qualitative experiments showed that the proposed autoencoder method can learn more structural details of 3D cumulus shapes than existing approaches. Furthermore, some modeling experiments on rendering images demonstrated the effectiveness of the reconstruction model. The proposed autoencoder network learns the latent space of 3D cumulus cloud shapes. The presented reconstruction architecture models a cloud from a single image. Experiments demonstrated the effectiveness of the two models. Yunchi Cen, Xiaohui Liang 0001 |
Virtual Real. Intell. Hardw. | 4 |
| 2020 | Perceptual Quality Assessment On DIBR Synthesized Videos With Composite DistortionsabstractDIBR (depth-image-based rendering) synthesized videos has become increasingly popular recently in supporting 3D-TV, Free-viewpoint video (FVV) and interactive graphic applications. Due to the synthesis process, it contains composite distortions, i.e., DCT-like distortions plus geometric distortions. It has led to new challenges for subjective and objective video quality assessment (VQA) community, where such composite distortions have not been taken into consideration. In this work, we construct a new database composed of DIBR synthesized videos, the reference view video is compressed by H.264 and HEVC encoders, which is subsequently synthesized to the virtual view with different DIBR algorithms. We also propose a novel no-reference (NR) synthesized video quality metric with multi-modal feature pooling and evaluate it against representative full reference (FR) and NR objective VQA models on the proposed database. The database and the proposed metric will be made available to the public. Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001 |
ICIP | 5 |
| 2020 | Small Components Parsing VIA Multi-Feature Fusion NetworkabstractPart parsing is a fundamental task towards fine image understanding in the multimedia and visual field. At present, the researchers working on part parsing focus on objects with large components, such as human, car. This paper centers on segmenting objects with small components. We call it small components parsing. In this paper, we propose a novel strategy for small components parsing, fusing multi-feature to utilize context information. We introduce Separable Spatial Pyramid module to embed spatial context information by fusing different scale spatial features. In decoding stage, attention-based feature fusion unit is drawn to utilize semantic context information in order to highlight details. Specifically, we design the Residual Upsampling manner to recover more details, considering spatial and channel characteristics. Experiments on RHD-PARSING and CamVid datasets demonstrate our method has achieved decent performance for small components parsing, taking hand parsing as an example, and has reached competitive results in scene parsing. Zhiying Leng, Xiaohui Liang 0001 |
ICME | 3 |
| 2020 | Cumuliform cloud formation control using parameter-predicting convolutional neural network
Yue Ma 0035, Frederick W. B. Li, Hubert P. H. Shum, Bailin Yang, Xiaohui Liang 0001 |
Graph. Model. | 8 |
| 2020 | Sparse metric-based mesh saliency
Shanfeng Hu, Xiaohui Liang 0001, Hubert P. H. Shum, Frederick W. B. Li, Nauman Aslam |
Neurocomputing | 2 |
| 2020 | Target-driven cloud evolution using position-based fluidsabstractAbstract To effectively control particle‐based cloud evolution without imposing strict position constraints, we propose a novel method integrating a control force field and a phase transition control into the position‐based fluids (PBF) framework. To produce realistic cloud simulation, we incorporate both fluid dynamics and thermodynamics to govern cloud particle movement. The fluid dynamics is simulated through our novel driving and damping force terms. As these terms are only formulated based on cloud particle density and position, they simplify the inputs and make our method free from artificial positional constraints. The thermodynamics is implemented by our phase transition control, which can effectively simulate cloud evolution between discrepant initial and target shapes, producing plausible results. Uniquely, our method can also support target shape change during cloud simulation. Experiment results have demonstrated our method surpasses existing methods. Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | CTS-LSTM: LSTM-based neural networks for correlatedtime series prediction
Huaiyu Wan, Shengnan Guo 0001, Xiaohui Liang 0001, Youfang Lin |
Knowl. Based Syst. | 4 |
| 2020 | A Unified Deep Metric Representation for Mesh Saliency Detection and Non-Rigid Shape MatchingabstractIn this paper, we propose a deep metric for unifying the representation of mesh saliency detection and non-rigid shape matching. While saliency detection and shape matching are two closely related and fundamental tasks in shape analysis, previous methods approach them separately and independently, failing to exploit their mutually beneficial underlying relationship. In view of the existing gap between saliency and matching, we propose to solve them together using a unified metric representation of surface meshes. We show that saliency and matching can be rigorously derived from our representation as the principal eigenvector and the smoothed Laplacian eigenvectors respectively. Learning the representation jointly allows matching to improve the deformation-invariance of saliency while allowing saliency to improve the feature localization of matching. To parameterize the representation from a mesh, we also propose a deep recurrent neural network (RNN) for effectively integrating multi-scale shape features and a soft-thresholding operator for adaptively enhancing the sparsity of saliency. Results show that by jointly learning from a pair of saliency and matching datasets, matching improves the accuracy of detected salient regions on meshes, which is especially obvious for small-scale saliency datasets, such as those having one to two meshes. At the same time, saliency improves the accuracy of shape matchings among meshes with reduced matching errors on surfaces. Shanfeng Hu, Hubert P. H. Shum, Nauman Aslam, Frederick W. B. Li, Xiaohui Liang 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | Deep Blind Synthesized Image Quality Assessment with Contextual Multi-Level Feature PoolingabstractBlind image quality metrics have achieved significant improvement on traditional 2D image dataset, yet still being insufficient for evaluating synthesized images generated from depth-image-based rendering. The geometric distortions in synthesized image are non-uniform, which is challenging for feature representation and pooling. To address this, we propose an end-to-end deep blind synthesized image quality metric SIQA-CFP. We particularly design a contextual multilevel feature pooling module to encode low- and high-level features, which are extracted by a deep pre-trained ResNet. Experimental results on IRCCyN/IVC DIBR dataset show that our method outperforms state-of-the-art synthesized image quality metrics. Our method also achieves competitive performance on traditional 2D image datasets like LIVE Challenge and TID2013. Bailin Yang, Frederick W. B. Li, Xiaohui Liang 0001 |
ICIP | 5 |
| 2019 | No-reference synthetic image quality assessment with convolutional neural network and local image saliencyabstractDepth-image-based rendering (DIBR) is widely used in 3DTV, free-viewpoint video, and interactive 3D graphics applications. Typically, synthetic images generated by DIBR-based systems incorporate various distortions, particularly geometric distortions induced by object dis-occlusion. Ensuring the quality of synthetic images is critical to maintaining adequate system service. However, traditional 2D image quality metrics are ineffective for evaluating synthetic images as they are not sensitive to geometric distortion. In this paper, we propose a novel no-reference image quality assessment method for synthetic images based on convolutional neural networks, introducing local image saliency as prediction weights. Due to the lack of existing training data, we construct a new DIBR synthetic image dataset as part of our contribution. Experiments were conducted on both the public benchmark IRCCyN/IVC DIBR image dataset and our own dataset. Results demonstrate that our proposed metric outperforms traditional 2D image quality metrics and state-of-the-art DIBR-related metrics. Xiaohui Liang 0001, Bailin Yang, Frederick W. B. Li |
Comput. Vis. Media | 2 |
| 2018 | A Deep Blind Image Quality Assessment with Visual Importance Based Patch Score
Zhengyi Lv, Xiaohui Liang 0001 |
ACCV (2) | 4 |
| 2018 | DOOBNet: Deep Object Occlusion Boundary Detection from an Image
Guoxia Wang, Frederick W. B. Li, Xiaohui Liang 0001 |
ACCV (6) | 4 |
| 2017 | Modeling Cumulus Cloud Scenes from High-resolution Satellite ImagesabstractAbstract We present a reconstruction framework, which fits physically‐based constraints to model large‐scale cloud scenes from satellite images. Applications include weather phenomena visualization, flight simulation, and weather spotter training. In our method, the cloud shape is assumed to be composed of a cloud top surface and a nearly flat cloud base surface. Based on this, an effective method of multi‐spectral data processing is developed to obtain relevant information for calculating the cloud base height and the cloud top height, including ground temperature, cloud top temperature and cloud shadow. A lapse rate model is proposed to formulate cloud shape as an implicit function of temperature lapse rate and cloud base temperature. After obtaining initial cloud shapes, we enrich the shapes by a fractal method and represent reconstructed clouds by a particle system. Experiment results demonstrate the capability of our method in generating physically sound large‐scale cloud scenes from high‐resolution satellite images. Xiaohui Liang 0001, Chunqiang Yuan, Frederick W. B. Li |
Comput. Graph. Forum | 2 |
| 2016 | A Cluster Sampling Method for Image Matting via Sparse Coding
Xiaoxue Feng, Xiaohui Liang 0001 |
ECCV (2) | 2 |
| 2016 | A strong bilayer appearance model for human pose estimation from a high freedom still imageabstractAppearance model is widely used for image description and demonstrates an impressive performance in object detection. However, most appearance models can not be applied to more freedom object in still image, especially when dealt with variant objects whose shapes are modified by warping, rotation, etc. In this article, a simple but effective method to build a regional rotation-invariant feature descriptor is proposed to catch discriminative information of the variant human pose, which has a superior advantage when targets are in arbitrary orientations and slightly warping. Moreover, a mixture spatial model with visible parameters is then presented to differentiate the body structure and estimate the visible accurate position of each joint. The experiment results indicate that the proposed descriptor give near state-of-the-art performance on both handwritten digit recognition database and two public human motion databases containing athletes or pedestrians under certain different variations. Songsong Ruan, Xiaohui Liang 0001 |
ICIP | 3 |
| 2016 | Derivation of 3D cloud animation from geostationary satellite images
Xiaohui Liang 0001, Chunqiang Yuan |
Multim. Tools Appl. | 1 |
| 2016 | Visual saliency guided textured model simplification
Bailin Yang, Frederick W. B. Li, Xun Wang 0007, Mingliang Xu 0001, Xiaohui Liang 0001, Zhaoyi Jiang, Yanhui Jiang |
Vis. Comput. | 5 |
| 2014 | Modelling Cumulus Cloud Shape from a Single ImageabstractAbstract Clouds are important components of the fascinating natural images. However, extracting cloud shapes from images remains a challenging task. This paper presents a calculation method for estimating the shape of a cumulus cloud from a single image suitable for flight simulations and games. The shape of the cloud is assumed to be symmetric. Based on this assumption, the intensities of pixels are correlated with the geometry of a cloud's surface via a simplified single scattering model. A propagation scheme is designed to derive the surface progressively, and mesh editing techniques are used to improve the surface. Finally, the cloud is represented by a particle system. The results show that the proposed method can generate realistic cumulus clouds that are similar to those found in the images in terms of the shape distribution. Chunqiang Yuan, Xiaohui Liang 0001, Shiyu Hao, Qinping Zhao |
Comput. Graph. Forum | 2 |
| 2014 | Recursive Templates Segmentation and Exemplars Matching for Human ParsingabstractMost previous studies need to learn a complex object model for parsing a specific object instance. This paper directly learns the general parsing patterns from the set of parsed objects and formalizes the parsing patterns as a series of parsing templates instead of learning the complex object model. Moreover, a novel hierarchical structure is presented to represent an object by using the parsing templates, which implicitly contains the multi-scale object parts and their relationships. For a single object, the parsing process is equivalent to establishing its hierarchical representation and determining the parsing template for each node. We combine the top-down decomposing scheme and the bottom-up composing scheme to infer the parsing process and formalize the inference as an energy minimization problem. The effect of our method is demonstrated by parsing the human body with aggressive pose variations. Compared with the state-of-the-art methods, the parsing results are more satisfying. Linjia Sun, Xiaohui Liang 0001, Qinping Zhao |
Comput. J. | 2 |
| 2014 | Automatic sub-category partitioning and parts localization for learning a robust object model
Linjia Sun, Xiaohui Liang 0001, Qinping Zhao |
Image Vis. Comput. | 2 |
| 2014 | Flexible editing of human motion by three-way decompositionabstractABSTRACT This paper proposes a new generative model for flexible editing of human motion. Different from previous work, three intuitive factors of motion, namely, content, identity and style, can be manipulated directly with the new model. With the new generative model, motion editing can be achieved in various aspects, including transferring an unknown style from an actor to another, synthesizing other styles for an unknown actor and generating a new motion with other content. Copyright © 2013 John Wiley & Sons, Ltd. Zhiying He, Xiaohui Liang 0001, Qinping Zhao |
Comput. Animat. Virtual Worlds | 2 |
| 2014 | ASEHM: a new transmission control mechanism for remote rendering system
Yajie Yan, Xiaohui Liang 0001, Ke Xie 0002, Qinping Zhao |
Multim. Tools Appl. | 2 |
| 2013 | Triangular Mesh Based Stroke Segmentation for Chinese CalligraphyabstractThis paper proposes a novel stroke extraction method for the Chinese character. In our method, a Chinese character is represented as a set of triangular mesh that is generated by using the canny contour detector and the constraint Delaunay triangulation (CDT). Based on the representation, the singular regions and the sub-strokes are firstly determined by the properties of triangular mesh. The point-to-boundary orientation distance (PBOD) of one triangular mesh is generated to discriminate whether it should be involved in singular regions. Then, the singular regions and sub-strokes are modeled with a graph. Two sub-strokes are connected if they are continuous in the singular region and recover the part contour of the stroke damaged by the singular region. Finally, this method is used to extract the stroke with variable width in Chinese characters, such as type of Kai in the tablet inscription. Experimental results show that the proposed method is feasible and effective. Xiaohui Liang 0001, Linjia Sun |
ICDAR | 2 |
| 2013 | Unsupervised image segmentation using global spatial constraint and multi-scale representation on multiple segmentation proposalsabstractThis paper presents a novel method for unsupervised image segmentation. The method determines the reasonable segments for final segmentation by exploiting both global and local context cues on multiple segmentation proposals. The proposal is obtained by using any existing segmentation algorithms, providing the diverse segment cues to guide segmentation. An iterative process is used to perform the cues integration and the image segmentation, including the segments modeling and the segments labeling. The former estimates the distribution of shared segments, while the latter labels each proposal into segments by minimizing an energy function. The final segmentation is produced when the consistent spatial layout is found in different proposals. Compared with the existing methods, the segmentation results are more satisfying on the Berkeley Segmentation Database. Linjia Sun, Xiaohui Liang 0001 |
ICIP | 2 |
| 2011 | Integrating Boundary Cue with Superpixel for Image SegmentationabstractThis paper researches image segmentation as a global optimization problem and proposes a new way, which is called superpixel status model, to integrate boundary and region cue. Superpixel status model is a label model which describes the joint distribution of boundary and region classification in a Bayesian framework. For organizing a boundary classifier, the contour of super pixel is decomposed into multiple line segments, and a robust line descriptor is presented to form line feature vector. Finally, an objective function is defined to assemble all super pixels statuses across the entire image for segmentation. Experiments and results show that the effectiveness of our approach. Linjia Sun, Xiaohui Liang 0001 |
ICIG | 2 |
| 2011 | An Adaptive Splitting and Transmission Control Method for Rendering Point Model on Mobile DevicesabstractThe physical characteristics of current mobile devices impose significant constraints on the processing of 3D graphics. The remote rendering framework is considered a better choice in this regard. However, limited battery life is a critical constraint when using this approach. Earlier methods based on this framework suffered from high transmission frequency. We present a software solution to this problem with a key element, an Adaptive Splitting and Error Handling Mechanism that indirectly reserves the electricity in mobile devices by reducing the transmission frequency. To achieve this goal, a geometric relation is maintained that tightly couples several consecutive Levels of Detail (LOD). Adaptive Splitting can then approximate the LOD from a much coarser split base under the guidance of the relation. Data transmission between the server and mobile device occurs only when out-ranged LOD is about to be displayed. Our remote rendering architecture, based on the above approach, trades splitting process for transmission, thereby alleviating the problem of frequent data transmission. Yajie Yan, Xiaohui Liang 0001, Ke Xie 0002, Qinping Zhao |
ISM | 2 |
| 2011 | Flexible editing of style, identity and content of human motionabstractMany prior works aim to provide style editing approaches which are intuitive and straightforward, such as [Min et al. 2010] etc. But they did not consider the content editing, like generate a running motion for an actor given his walking motion. For the purpose of style editing, most researches use linear model, such as PCA, ICA and multilinear model [Min et al. 2010]. However, the human motion is actually a highly nonlinear model [Elgammal and Lee 2004]. But the nonlinear model can not construct a direct mapping which makes the reconstruction of high-dimensional data complex. Besides, it can only manipulate data in the training database. In addition, when extracting the style of motion, most methods do not separate the content before learning the style. This confuses the content and style when doing the style learning, and therefore impairs the final result. Zhiying He, Xiaohui Liang 0001, Yiming Yue |
SI3D | 2 |
| 2011 | Light Space Cascaded Shadow Maps Algorithm for Real Time Rendering
Xiaohui Liang 0001, Shang Ma, Li-Xia Cen |
J. Comput. Sci. Technol. | 1 |
| 2010 | An adaptive splitting and transmission control method for rendering point model on mobile devicesabstractNo abstract available. Ke Xie 0002, Xiaohui Liang 0001, Zhiying He |
SI3D | 2 |
| 2009 | Multi-attributes controlled point-based rendering architecture for mobile devicesabstractTaking the characteristics of mobile devices and the demands for graphics rendering into account, this paper gives a set of definitions of attributes to express the performance of mobile terminals and algorithm performance. And then, according to the relationship of these attributes, a point-based rendering architecture for mobile devices is proposed. It includes pre-processing simplification, data organization and rendering. Instead of just simplifying the algorithm on PC, the entire process takes the multi-attributes of mobile devices and algorithm into account. Experiments shows that it can make full use of the limited resources to carry out a relatively realistic and real-time rendering. Zhiying He, Xiaohui Liang 0001 |
CAD/Graphics | 2 |
| 2009 | A novel simplification algorithm based on MLS and Splats for point modelsabstractThe simplification of point models is important in point-based processing technology because of the increasing of data complexity. Many researches focus on getting the subset from the initial point set, which can not represent the whole object properly. In this paper, we present a novel simplification algorithm based on Moving Least Square (MLS) and Splats for point models. The algorithm uses MLS to represent the point models and can get the minimum error point which is used to deputize its neighborhood. This approach can get new proper agent points instead of subset. Then we calculate the error based on the rendering results, which means considering the geometry of splat when calculating the error of simplification. Experiment results show that this algorithm is not only efficient, but also has good quality. Zhiying He, Xiaohui Liang 0001 |
CGI | 2 |
| 2009 | A point-based rendering approach for real-time interaction on mobile devices
Xiaohui Liang 0001, Qinping Zhao, Zhiying He, Ke Xie 0002 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2008 | A Registration Method Based on Nature Feature with KLT Tracking Algorithm for Wearable ComputersabstractKLT algorithm has been widely used in the registration process for natural features tracking in augmented reality (AR) systems. However, KLT is vulnerable by surrounding environments, and the feature points on screen borderlines or once be occluded may not be tracked persistently. Homographic matrices cannot be calculated accurately due to these disadvantages, which will result in registration failure. In this paper, KLT algorithm was improved based on the updating strategy of feature-points set, moreover, we applied RANSAC algorithm to filter mismatched feature points for calculating homographic matrices precisely. The new method was implemented in a prototype system which ran on a wearable computer, and the experiments results show that our system can not only track the feature points stably but also ensure the accuracy and the effectiveness of registration. Xiaohui Liang 0001, Zhi-Ying He, Guo-Liang Hua |
CW | 2 |
| 2008 | GPU-Based Feature-Preserving Distance Field ComputationabstractWe present an optimized algorithm to compute 3D distance fields using the bilinear interpolation capabilities of GPUs while preserving the features of the model. For a geometric model, our algorithm computes the Euclidean distance fields on each 2D slice of a 3D grid by applying linear decomposition to the non-linear distance function of each primitive and evaluating it using texture mapping hardware. We compute the bounds of the Voronoi region of each primitive on a 2D slice to reduce rasterization cost of the distance functions. Further more, culling techniques are incorporated to remove primitives that do not contribute to the distance field of a given slice. Our method is able to preserve the features of the model such as sharp edges and corners by detecting them and storing the associated information explicitly during the distance field computation. The experiment demonstrates that the algorithm is accurate and can compute 3D distance fields of complex models consisting of thousands of triangles while preserving the features efficiently. Sissi Xiaoxiao Wu, Xiaohui Liang 0001, Qidi Xu, Qinping Zhao |
CW | 2 |
| 2006 | A Graph Transformation System Model of Dynamic Reorganization in Multi-agent Systems
Xiaohui Liang 0001, Qinping Zhao |
IDEAL | 2 |
| 2006 | Adaptive Mechanisms of Organizational Structures in Multi-agent Systems
Xiaohui Liang 0001, Qinping Zhao |
PRIMA | 2 |
| 2005 | A Toolkit for Automatically Modeling and Simulating 3D Multi-articulation Entity in Distributed Virtual Environment
Xiaohui Liang 0001, Chuanpeng Wang, Yinghui Che, Yu Jiangying, Qu Na |
ICCSA (3) | 1 |