VLDB 2026 Research / reviewers in the wild / expert
Xia Hou
dblp:49/4722
· DBLP profile ↗
21ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0001-6006-4215ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IntentMotion: Learning Intent-Aware Human Motion from Language in 3D ScenesabstractGenerating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion, a novel framework that generates human motion in 3D scenes from natural language instructions by explicitly modeling intent. We first introduce the Intention-Guided Contact Field (IGCF). This differentiable voxel-based contact region representation explicitly aligns parsed language roles with spatial contact regions through a hierarchical attention mechanism. IGCF is jointly trained with a diffusion-based motion generator, allowing contact predictions to adapt dynamically through gradient feedback. To improve the controllability and physics-aware motion, we further propose an Intention-Aware Diffusion Model (IADM), which decouples the high-level semantic planning from the low-level contact refinement in a coarse-to-fine process. The optimized contact cues are utilized to guide the synthesis of a coarse trajectory, followed by refining detailed pose sequences under IGCF supervision. Experiments on the HUMANISE and LINGO datasets demonstrate that our IntentMotion outperforms recent baselines in contact accuracy, semantic alignment, and generalization to unseen scenes. Wenfeng Song, Shi Zheng, Xingliang Jin, Aimin Hao, Fei Hou 0001, Xia Hou, Shuai Li 0001 |
AAAI | 7 |
| 2026 | DynAvatar: Dynamic 3D Head Avatar Deformation With Expression Guided Gaussian SplattingabstractGenerating high-fidelity, expressive, and realistic 3D head avatars remains a fundamental challenge for immersive applications such as virtual reality, gaming, and telepresence. This task requires not only precise modeling of non-rigid facial deformations but also semantically controllable expression synthesis under diverse viewpoints and motion contexts. We present DynAvatar, a novel framework that integrates expression-guided deformation into the 3D Gaussian splatting pipeline to produce photorealistic and emotionally resonant head avatars. Our method introduces two key innovations: (1) an expression-guided Gaussian deformation module that tightly couples geometric displacement with high-level semantic cues, enabling fine-grained and anatomically meaningful facial animation; and (2) a spatial context embedding mechanism that encodes the canonical position of each Gaussian to preserve semantic coherence and spatial consistency during expression generation. Extensive experiments on both controlled and in-the-wild datasets demonstrate that DynAvatar significantly outperforms state-of-the-art methods in terms of visual realism, expression fidelity, and rendering quality. Wenfeng Song, Zhongyong Ye, Shuai Li 0001, Xia Hou, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible NetworksabstractAs virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avatars generation via disentangled invertible networks (CtrlAvatar), a real-time framework for generating lifelike and customizable avatars. CtrlAvatar uses disentangled invertible networks to separate the deformation process into implicit body geometry and explicit texture components. This approach eliminates the need for repeated occupancy reconstruction, enabling detailed and coherent animations. The body geometry component ensures anatomical accuracy, while the texture component allows for complex, artifact-free clothing customization. This architecture ensures smooth integration between body movements and surface details. By optimizing transformations with position-varying offsets from the avatar’s initial Linear Blend Skinning vertices, CtrlAvatar achieves flexible, natural deformations that adapt to various scenarios. Extensive experiments show that CtrlAvatar outperforms other methods in quality, diversity, controllability, and cost-efficiency, marking a significant advancement in avatar generation. Wenfeng Song, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Xia Hou |
AAAI | 6 |
| 2025 | ViMoGen: A Novel Motion Generator for Virtual Standard Patient
Xuehan Wang, Wenfeng Song, Shuai Li 0001, Xian'e Wang, Xia Hou |
ICXR | 6 |
| 2025 | AttriDiffuser: Adversarially enhanced diffusion model for text-to-facial attribute image synthesis
Wenfeng Song, Zhongyong Ye, Xia Hou, Shuai Li 0001, Aimin Hao |
Pattern Recognit. | 4 |
| 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style PreservationabstractIt is a challenging task to create realistic 3D avatars that accurately replicate individuals' speech and unique talking styles for speech-driven facial animation. Existing techniques have made remarkable progress but still struggle to achieve lifelike mimicry. This article proposes "TalkingStyle", a novel method to generate personalized talking avatars while retaining the talking style of the person. Our approach uses a set of audio and animation samples from an individual to create new facial animations that closely resemble their specific talking style, synchronized with speech. We disentangle the style codes from the motion patterns, allowing our method to associate a distinct identifier with each person. To manage each aspect effectively, we employ three separate encoders for style, speech, and motion, ensuring the preservation of the original style while maintaining consistent motion in our stylized talking avatars. Additionally, we propose a new style-conditioned transformer decoder, offering greater flexibility and control over the facial avatar styles. We comprehensively evaluate TalkingStyle through qualitative and quantitative assessments, as well as user studies demonstrating its superior realism and lip synchronization accuracy compared to current state-of-the-art methods. Wenfeng Song, Xuan Wang 0024, Shi Zheng, Shuai Li 0001, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Radiographic Reports Generation via Retrieval Enhanced Cross-modal FusionabstractAccurate radiographic reports are crucial for effective clinical decision-making and patient safety, as they directly influence diagnosis and treatment plans. Existing models for radiographic report generation often struggle with integrating medical image and textual report features and addressing data imbalances. To overcome these limitations, we propose an enhanced cross-modal aligned retrieval-driven network (EARnet). Our model incorporates two key innovations: Enhanced Cross-modal Alignment (ECA) module and Case-based Retrieval Augmenter (CRA) module. ECA ensures the effective integration of medical visual and textual data by aligning these modalities. CRA helps mitigate data imbalance by improving the representation of medical information and ensuring a more balanced and comprehensive coverage of both normal and abnormal cases. This dual-module approach significantly improves the coherence and accuracy of the generated medical reports. Evaluation results demonstrate that our EARnet model substantially outperforms existing methods in terms of report quality and accuracy across multiple metrics. Our codes and models are available at https://github.com/lyf616/EARnet. Xia Hou, Wenfeng Song, Wenzhe You, Shuai Li 0001 |
BIBM | 1 |
| 2024 | Arbitrary Motion Style Transfer with Multi-Condition Motion Latent Diffusion ModelabstractComputer animation's quest to bridge content and style has historically been a challenging venture, with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality, ensuring a harmonious fusion where the core narrative of the content is both preserved and elevated through stylistic enhancements. We propose a novel Multi-condition Motion Latent Diffusion Model (MCM-LDM) for Arbitrary Motion Style Transfer (AMST). Our MCM-LDM significantly emphasizes preserving trajectories, recognizing their fundamental role in defining the essence and fluidity of motion content. Our MCM-LDM's cornerstone lies in its ability first to disentangle and then intricately weave together motion's tripartite components: motion trajectory, motion content, and motion style. The critical insight of MCM-LDM is to embed multiple conditions with distinct priorities. The content channel serves as the primary flow, guiding the overall structure and movement, while the trajectory and style channels act as auxiliary components and synchronize with the primary one dynamically. This mechanism ensures that multi-conditions can seamlessly integrate into the main flow, enhancing the overall animation without overshadowing the core content. Empirical evaluations underscore the model's proficiency in achieving fluid and authentic motion style transfers, setting a new benchmark in the realm of computer animation. The source code and model are available at https://github.com/XingliangJin/MCM-LDM.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou, Hong Qin 0001 |
CVPR | 6 |
| 2024 | FusionCraft: Fusing Emotion and Identity in Cross-Modal 3D Facial Animation
Zhenyu Lv, Xuan Wang 0024, Wenfeng Song, Xia Hou |
ICIC (10) | 4 |
| 2024 | A review of automatic source code summarization
Xia Hou, Xiuming Qiao, Wenfeng Song |
Empir. Softw. Eng. | 2 |
| 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusionabstract3D Facial animations, crucial to augmented and mixed reality digital media, have evolved from mere aesthetic elements to potent storytelling media. Despite considerable progress in facial animation of neutral emotions, existing methods still struggle to capture the authenticity of emotions. This paper introduces a novel approach to capture fine facial expressions and generate facial animations using audio synchronization. Our method consists of two key components: First, the Local-to-global Latent Diffusion Model (LG-LDM) tailored for authentic facial expressions, which can integrate audio, time step, facial expressions, and other conditions towards possible encoding of emotionally rich yet latent features in response to possibly noisy raw audio signals. The core of LG-LDM is our carefully designed Facial Denoiser Model (FDM) for aligning the local-to-global animation feature with audio. Second, we redesign an Emotion-centric Vector Quantized-Variational AutoEncoder framework (EVQ-VAE) to finely decode the subtle differences under different emotions and reconstruct the final 3D facial geometry. Our work significantly contributes to the key challenges of emotionally realistic 3D facial animation for audio synchronization and enhances the immersive experience and emotional depth in augmented and mixed reality applications. We provide a reproducibility kit including our code, dataset, and detailed instructions for running the experiments. This kit is available at https://github.com/wangxuanx/Face-Diffusion-Model. Wenfeng Song, Xuan Wang 0024, Yiming Jiang 0018, Shuai Li 0001, Aimin Hao, Xia Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Tell Your Story: Text-Driven Face Video Synthesis with High Diversity via Adversarial LearningabstractFace synthesis is a rapidly growing area of research in computer vision. Text-driven face synthesis is particularly flexible, but challenges still exist in fusing the semantics of text and images, as well as generating diverse faces. To address these challenges, we propose a cross-modality adversarial learning framework to generate highly diverse face videos that correspond to given text descriptions. We encode text and images into a common latent space and align text and image features to control the synthesis of face attributes. We have designed a novel auto-encoder with a face identity discriminator that enlarges the margin between different individuals, increasing the variety of created faces while maintaining the semantic coherence of text and images. Our proposed method has been successfully tested on the recently released Multimodal VoxCeleb dataset. Our code is public available at https://github.com/sunmeng7/TYS.git. Xia Hou, Wenfeng Song |
ICIP | 1 |
| 2023 | Dual Temporal Transformers for Fine-Grained Dangerous Action RecognitionabstractRecognizing dangerous actions is a critical task in computer vision, especially for surveillance applications. While existing deep learning methods have been successful in confined environments, they struggle with the anomalous and salient variations of human postures in dangerous actions. Additionally, finer-grained dangerous actions require more discriminative cues, adding to the complexity of the task. To address these challenges, we propose a novel solution that models the intrinsic and invariant properties of dangerous actions at multiple temporal semantic levels. Concretely, we propose a Dual Temporal Transformers (DTT) to capture temporal interactions between distinct key points in the human body aggregation from shallow to deep layers, increasing the perception field from local to global, simultaneously. By doing so, our method avoids overfitting to unrelated or minor clues in videos and achieves a generalized representation of abnormal actions. We evaluate our approach on indoor and outdoor environments and found that DTT outperforms existing methods in terms of efficiency and accuracy. Our code and dataset are pubic available on https://github.com/AveryJohnsonJJ/DTT.git. Wenfeng Song, Xingliang Jin, Yang Gao 0032, Xia Hou |
ICIP | 5 |
| 2023 | Two-Stage Underwater Image Restoration Algorithm Based on Physical Model and Causal InterventionabstractUnderwater images often suffer from severe color degradation, haze and local blur, which are caused by the scattering and absorption effects of light in water. Firstly, to address the lack of annotations in underwater images, we propose a multi-degradation rate underwater image degradation model. Additionally, we use the generative adversarial network (GAN) to guide the underwater image degradation model and generate paired images that can be used for underwater image restoration (UIR) network training. Also, given the problems of poor restoration quality, serious loss of detail, and slow inference speed of existing deep learning algorithms, we construct an underwater image restoration network that can reference in real-time. Moreover, through causal interventions in the image generation process, spurious correlations between global features and detailed features are eliminated. As a result, the detail generation ability of the image is improved. Experiments on real underwater image datasets demonstrate that compared with existing methods, our proposed method more effectively solves the problem of underwater image degradation. Junyu Hao, Xia Hou, Yang Zhang 0099 |
IEEE Signal Process. Lett. | 3 |
| 2023 | An Intelligent Virtual Standard Patient for Medical Students Training Based on Oral Knowledge GraphabstractVirtual standard patient (VSP) is in high demand for medical students' diagnosis ability training in an efficient manner. Different from the traditional conversation system in medical dialogue generation, VSP needs a novel conversation paradigm to act as the patient instead of the doctor. However, existing conversation techniques still have limited ability in terms of generation of symptoms exhibited by patients with the personalized and knowledge-centered expressions. To alleviate these problems, we propose to construct a novel oral knowledge graph, which sufficiently provides medical clues of the certain disease. Accordingly, the VSP could accurately interact with the dentists for their underlying intention and express the symptoms characters in a natural style. To efficiently retrieve the related disease clues, the symptoms descriptions of the oral diseases are encoded into the oral knowledge graph, which could well organize the disease-centered symptom entities and speaking styles. Moreover, to transfer the common sense knowledge from existing large scale of medical knowledge graph to the specific oral knowledge graph, a coupled pre-trained Bert models is further designed to learn the related medical knowledge from coarse-level to fine-level hierarchically. Finally, a series of well-designed personalized templates are proposed to generate plausible and realistic answers in condition of the certain disease. We also conduct extensive user studies to demonstrate that the VSP satisfies the medical students' diagnosis practice requirement in terms of naturalness, realism, and topic relevance. Wenfeng Song, Xia Hou, Shuai Li 0001, Chenglizhao Chen, Danyang Gao, Xian'e Wang, Yuzhe Sun, Jianxia Hou, Aimin Hao |
IEEE Trans. Multim. | 2 |
| 2023 | FineStyle: Semantic-Aware Fine-Grained Motion Style Transfer with Dual Interactive-Flow FusionabstractWe present FineStyle, a novel framework for motion style transfer that generates expressive human animations with specific styles for virtual reality and vision fields. It incorporates semantic awareness, which improves motion representation and allows for precise and stylish animation generation. Existing methods for motion style transfer have all failed to consider the semantic meaning behind the motion, resulting in limited controls over the generated human animations. To improve, FineStyle introduces a new cross-modality fusion module called Dual Interactive-Flow Fusion (DIFF). As the first attempt, DIFF integrates motion style features and semantic flows, producing semantic-aware style codes for fine-grained motion style transfer. FineStyle uses an innovative two-stage semantic guidance approach that leverages semantic clues to enhance the discriminative power of both semantic and style features. At an early stage, a semantic-guided encoder introduces distinct semantic clues into the style flow. Then, at a fine stage, both flows are further fused interactively, selecting the matched and critical clues from both flows. Extensive experiments demonstrate that FineStyle outperforms state-of-the-art methods in visual quality and controllability. By considering the semantic meaning behind motion style patterns, FineStyle allows for more precise control over motion styles. Source code and model are available on https://github.com/XingliangJin/Fine-Style.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Person Re-Identification in Panoramic Views Based on Bayesian TransformersabstractThe panoramic view cameras offer more broad perspectives and continuous information for person re-identification (ReID). However, the panoramic-view videos suffer from objects distortion and bring more occlusion due to the fixed or moving capture points. This paper proposes a novel Bayesian Transformer Network (BTN) to adaptively capture the occlusion clues as Bayesian prior to guide the discriminative pedestrian-related feature extraction in the high-occlusion scenes. The Bayesian prior is built via a pre-trained CNN, which could recognize different occluded scenarios based on the severeness of noisy backgrounds. Moreover, to fully explore the occlusion prior, we propose to embed the semantic labels into a well-designed transformer network. By fostering the collaborative occlusion clues between the person and background, our method could achieve outstanding performance on both public benchmarks and panoramic view videos, which verifies the advantages of our BTN framework over existing methods. Wenfeng Song, Yang Gao 0032, Aimin Hao, Xia Hou |
ICIP | 7 |
| 2019 | Prediction for Student Academic Performance Using SMNaive Bayes Model
Baoting Jia, Ke Niu 0002, Xia Hou, Ning Li 0024, Xueping Peng, Peipei Gu, Ran Jia |
ADMA | 3 |
| 2015 | Texture segmentation using image decomposition and local self-similarity of different features
Xia Hou |
Multim. Tools Appl. | 2 |
| 2014 | Histogram modification using grey-level co-occurrence matrix for image contrast enhancementabstractHistogram modification is an important technique for contrast enhancement. Most changes of histogram are based on global or local region grey‐levels information. In this study, a novel grey‐level co‐occurrence matrix (GCOM)‐based histogram equalisation (COHE) method is proposed. A GCOM is a matrix or distribution of co‐occurring grey‐levels at a given offset, in which each row or column vector is actually a conditional histogram. The procedure of COHE has two steps. First, it is to equalise the modified conditional histograms, which are weighted sums of uniformly distributed histograms and the conditional histograms. An adjusting method of weight parameter is also presented in this study. Conditional histograms equalisations have the advantage of enlarging the difference between given grey‐levels and other spatially adjacent grey‐levels. Second, COHE algorithm finds mapping to obtain global enhance by weighting all the conditional translated grey‐levels with original image histogram. However, it could produce over‐enhanced unnatural looking images because of spikes of conditional histogram and original histogram. To deal with this, this study introduces methods of adjusting the conditional histogram and original histogram based on GCOM. Experimental results demonstrate that the proposed method can enhance the images effectively. Xia Hou |
IET Image Process. | 2 |
| 2007 | Approximation Property of Weighted Wavelet Neural Networks
Shou-Song Hu, Xia Hou |
ISNN (1) | 2 |