VLDB 2026 Research / reviewers in the wild / expert
Yuxin Hou
dblp:214/9400
· DBLP profile ↗
15ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0001-7266-9502ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DreamAnimate: Temporal Consistency and Detail Preservation for Character AnimationabstractCharacter animation aims to generate realistic, high-quality videos from a reference image and target frames. However, existing methods struggle to balance fine-grained detail preservation with temporal consistency. This limitation results in artifacts like flickering and unrealistic deformations, especially in facial and hand regions. To address these challenges, we propose DreamAnimate, a novel framework that synthesizes temporally consistent and detail-rich animations. DreamAnimate integrates three modules: the Progressive Motion Estimation module ensures accurate motion alignment and temporal stability by refining keypoint heatmaps, the Global Affine Transformation module generates dense motion flows to handle complex motions and occlusions, and the Character Animation Fusion module combines intermediate synthesis using a UNet architecture and an Animation Fusion Network to produce high-quality animations. Extensive experiments demonstrate that DreamAnimate outperforms state-of-the-art methods, achieving superior fidelity and effectively capturing intricate facial expressions and hand movements. Code and models will be released at https://github.com/cslltian/DreamAnimate in the near future. Lulu Tian, Hongxun Yao, Zhaopan Xu, Jiankun Zhu, Xi Chen 0110, Yuxin Hou |
ICME | 6 |
| 2024 | Hierarchical pose net: spatial hierarchical body tree driven multi-person pose estimation
Haoran Li 0012, Hongxun Yao, Yuxin Hou |
Multim. Tools Appl. | 3 |
| 2023 | Graph Convolutional GRU for Music-Oriented Dance Choreography GenerationabstractMusic-oriented dance choreography is a complex art form that requires consideration of various factors in music understanding and physical aesthetics. In addition, learning model to generate dance sequences automatically needs to overcome not only the challenges from choreography but also the complexity of human skeleton structure. In this paper, we propose a graph convolutional GRU (GraphGRU) model for music-oriented dance choreography by integrating the graph convolutional networks into the gated recurrent unit. The GraphGRU can model the spatial relationship among the human skeleton joints and the temporal relationship of the dancing sequences. Besides, we define multiple dance-driven kinematic graphs for GraphGRU and establish a path attention block to fuse the graph features. Moreover, we design a reverse-generation loss to ensure consistency between the generated dance and the music. Extensive experiments demonstrate that the proposed model GraphGRU can effectively solve the task of dance choreography. Yuxin Hou, Hongxun Yao, Haoran Li 0012 |
ICME | 1 |
| 2023 | Expansion of Visual Hints for Improved Generalization in Stereo MatchingabstractWe introduce visual hints expansion for guiding stereo matching to improve generalization. Our work is motivated by the robustness of Visual Inertial Odometry (VIO) in computer vision and robotics, where a sparse and unevenly distributed set of feature points characterizes a scene. To improve stereo matching, we propose to elevate 2D hints to 3D points. These sparse and unevenly distributed 3D visual hints are expanded using a 3D random geometric graph, which enhances the learning and inference process. We evaluate our proposal on multiple widely adopted benchmarks and show improved performance without access to additional sensors other than the image sequence. To highlight practical applicability and symbiosis with visual odometry, we demonstrate how our methods run on embedded hardware. Andrea Pilzer, Yuxin Hou, Niki Andreas Lopi, Arno Solin, Juho Kannala |
WACV | 2 |
| 2023 | Reachable Set Estimation for Memristive Complex-Valued Neural Networks With DisturbancesabstractThis brief focuses on reachable set estimation for memristive complex-valued neural networks (MCVNNs) with disturbances. Based on algebraic calculation and Gronwall-Bellman inequality, the states of MCVNNs with bounded input disturbances converge within a sphere. From this, the convergence speed is also obtained. In addition, an observer for MCVNNs is designed. Two illustrative simulations are also given to show the effectiveness of the obtained conclusions. Song Zhu, Yuxin Hou, Chunyu Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Gaussian Process Priors for View-Aware Inference
Yuxin Hou, Ari Heljakka, Arno Solin |
AAAI | 1 |
| 2021 | Learning Attribute-driven Disentangled Representations for Interactive Fashion RetrievalabstractInteractive retrieval for online fashion shopping provides the ability to change image retrieval results according to the user feedback. One common problem in interactive retrieval is that a specific user interaction (e.g., changing the color of a T-shirt) causes other aspects to change inadvertently (e.g., the retrieved item has a sleeve type different than the query). This is a consequence of existing methods learning visual representations that are semantically entangled in the embedding space, which limits the controllability of the retrieved results. We propose to leverage on the semantics of visual attributes to train convolutional networks that learn attribute-specific subspaces for each attribute to obtain disentangled representations. Thus operations, such as swapping out a particular attribute value for another, impact the attribute at hand and leave others untouched. We show that our model can be tailored to deal with different retrieval tasks while maintaining its disentanglement property. We obtain state-of-the-art performance on three interactive fashion retrieval tasks: attribute manipulation retrieval, conditional similarity retrieval, and outfit complementary item retrieval. Code and models are publicly available1. Yuxin Hou, Eleonora Vig, Michael Donoser, Loris Bazzani |
ICCV | 1 |
| 2021 | Novel View Synthesis via Depth-guided Skip ConnectionsabstractWe introduce a principled approach for synthesizing new views of a scene given a single source image. Previous methods for novel view synthesis can be divided into image-based rendering methods (e.g., flow prediction) or pixel generation methods. Flow predictions enable the target view to re-use pixels directly, but can easily lead to distorted results. Directly regressing pixels can produce structurally consistent results but generally suffer from the lack of low-level details. In this paper, we utilize an encoder-decoder architecture to regress pixels of a target view. In order to maintain details, we couple the decoder aligned feature maps with skip connections, where the alignment is guided by predicted depth map of the target view. Our experimental results show that our method does not suffer from distortions and successfully preserves texture details with aligned skip connections. Yuxin Hou, Arno Solin, Juho Kannala |
WACV | 1 |
| 2020 | Movement-induced Priors for Deep StereoabstractWe propose a method for fusing stereo disparity estimation with movement-induced prior information. Instead of independent inference frame-by-frame, we formulate the problem as a non-parametric learning task in terms of a temporal Gaussian process prior with a movement-driven kernel for inter-frame reasoning. We present a hierarchy of three Gaussian process kernels depending on the availability of motion information, where our main focus is on a new gyroscope-driven kernel for handheld devices with low-quality MEMS sensors, thus also relaxing the requirement of having full 6D camera poses available. We show how our method can be combined with two state-of-the-art deep stereo methods. The method either work in a plug-and-play fashion with pre-trained deep stereo networks, or further improved by jointly training the kernels together with encoder-decoder architectures, leading to consistent improvement. Yuxin Hou, Muhammad Kamran Janjua, Juho Kannala, Arno Solin |
ICPR | 1 |
| 2020 | Deep AutomodulatorsabstractWe introduce a new category of generative autoencoders called automodulators. These networks can faithfully reproduce individual real-world input images like regular autoencoders, but also generate a fused sample from an arbitrary combination of several such images, allowing instantaneous "style-mixing" and other new applications. An automodulator decouples the data flow of decoder operations from statistical properties thereof and uses the latent vector to modulate the former by the latter, with a principled approach for mutual disentanglement of decoder layers. Prior work has explored similar decoder architecture with GANs, but their focus has been on random sampling. A corresponding autoencoder could operate on real input images. For the first time, we show how to train such a general-purpose model with sharp outputs in high resolution, using novel training techniques, demonstrated on four image data sets. Besides style-mixing, we show state-of-the-art results in autoencoder comparison, and visual image quality nearly indistinguishable from state-of-the-art GANs. We expect the automodulator variants to become a useful building block for image applications and other data domains. Ari Heljakka, Yuxin Hou, Juho Kannala, Arno Solin |
NeurIPS | 2 |
| 2020 | Reachable set bounding for neural networks with mixed delays: Reciprocally convex approach
Song Zhu, Yongqiang Qi, Yuxin Hou |
Neural Networks | 4 |
| 2019 | Iterative Path Reconstruction for Large-Scale Inertial Navigation on Smartphones
Santiago Cortés Reina, Yuxin Hou, Juho Kannala, Arno Solin |
FUSION | 2 |
| 2019 | Multi-View Stereo by Temporal Nonparametric FusionabstractWe propose a novel idea for depth estimation from multi-view image-pose pairs, where the model has capability to leverage information from previous latent-space encodings of the scene. This model uses pairs of images and poses, which are passed through an encoder-decoder model for disparity estimation. The novelty lies in soft-constraining the bottleneck layer by a nonparametric Gaussian process prior. We propose a pose-kernel structure that encourages similar poses to have resembling latent spaces. The flexibility of the Gaussian process (GP) prior provides adapting memory for fusing information from nearby views. We train the encoder-decoder and the GP hyperparameters jointly end-to-end. In addition to a batch method, we derive a lightweight estimation scheme that circumvents standard pitfalls in scaling Gaussian process inference, and demonstrate how our scheme can run in real-time on smart devices. Yuxin Hou, Juho Kannala, Arno Solin |
ICCV | 1 |
| 2017 | Dancing like a superstar: Action guidance based on pose estimation and conditional pose alignmentabstractAction Guidance (AG) aims at scoring how accurate the action is and giving guidance to the learners on how to correct their actions according to the standard instructive videos. AG has plenty of real-world applications such as sports training, rehabilitation treatment, and dance teaching. However, the problem of assessing the accuracy of action has almost no effective solution. In this paper, we describe a two-stage framework for action guidance. Firstly, we estimate the poses in the test video and standard video using person detection method and convolutional pose machines (CPMs). As for action guidance, we propose a network to compute the essential differences between two poses under different sizes, views, and locations. Extensive experiments with real-world and synthetic datasets demonstrate the effectiveness of our framework. Yuxin Hou, Hongxun Yao, Haoran Li 0012, Xiaoshuai Sun |
ICIP | 1 |
| 2017 | Gated additive skip context connection for object detectionabstractContext information plays an important role in object detection. DeepID concatenates the global context for classification, while Yolo, SSD, and Crafting use local context information for detection. In this paper, we propose a straightforward method to plug the global context information into the Faster RCNN Framework, namely the Skip Context Connection (SCC). We use SCC to inject the global context into the object representation which skips the RoiPooling layer rather than drops it. Therefore, it can not only leverage the context information but also keep the location accuracy from the RCNN framework. We proposed three principles to construct the SCC blocks: effectiveness means fewer parameters, additivity means the features possess the same meaning, and selectable means soft gated addition. We also evaluate several different SCC blocks. The Gated Additive SCC(GA-SCC) which satisfy the three principles get the best performance. Our experiment results on PASCAL VOC 2007 show that GA-SCC can get the steady 1% improvement over the traditional RCNN method. Haoran Li 0012, Hongxun Yao, Yuxin Hou, Xiaoshuai Sun |
ICIP | 3 |