Hua Yu 0006

dblp:02/2407-6 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-7698-2173ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Grounding Programming Chatbot in Computational Thinking: Design and Evaluation of MazeMate
Chenyu Hou, Hua Yu 0006, Gaoxia Zhu, John Derek Anas, Jiao Liu 0006, Yew-Soon Ong
AIED (1)2
2025 Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesis
abstract
Human motion synthesis aims to generate plausible human motion sequences, which has raised widespread attention in computer animation. Recent score-based generative models (SGMs) have demonstrated impressive results on this task. However, their training process involves complex curvature trajectories, leading to unstable training process. In this paper, we propose a Deterministic-to-Stochastic Diverse Latent Feature Mapping (DSDFM) method for human motion synthesis. DSDFM consists of two stages. The first human motion reconstruction stage aims to learn the latent space distribution of human motions. The second diverse motion generation stage aims to build connections between the Gaussian distribution and the latent space distribution of human motions, thereby enhancing the diversity and accuracy of the generated human motions. This stage is achieved by the designed deterministic feature mapping procedure with DerODE and stochastic diverse output generation procedure with DivSDE. DSDFM is easy to train compared to previous SGMs-based methods and can enhance diversity without introducing additional training parameters. Through qualitative and quantitative experiments, DSDFM achieves state-of-the-art results surpassing the latest methods, validating its superiority in human motion synthesis.
Hua Yu 0006, Weiming Liu 0005, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008
CVPR1
2025 Distinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label Assignment
abstract
Nowadays, domain adaptation techniques have been widely investigated for knowledge sharing from labeled source domain to unlabeled target domain. However, target domain may include some data samples that belong to unknown categories in real-world scenarios. Moreover, the target domain cannot access the source data samples due to privacy-preserving restrictions. In this paper, we focus on the source-free open set domain adaptation problem which includes two main challenges, i.e., how to distinguish known and unknown target samples and how to exploit useful source information to provide trustworthy pseudo labels for known target samples. Existing approaches that directly apply conventional domain alignment methods could lead to sample mismatch and misclassification in this scenario. To overcome these issues, we propose a Distinguish Then Exploit model (DTE) with two components, i.e., weight barcode estimation and sparse label assignment. Weight barcode estimation first calculates the marginal probability of target samples via partially unbalanced optimal transport, then quantize barcode results to distinguish unknown target samples. Sparse label assignment utilizes sparse sample-label matching via proximal term to fully exploit useful source information. Our empirically study on several datasets shows that DTE outperforms the state-of-the-art models on tackling the source-free open set domain adaptation problem.
Weiming Liu 0005, Jun Dan, Fan Wang 0020, Xinting Liao, Junhao Dong 0001, Hua Yu 0006, Shunjie Dong, Lianyong Qi
CVPR6
2025 Solving Discrete (Semi) Unbalanced Optimal Transport with Equivalent Transformation Mechanism and KKT-Multiplier Regularization
abstract
Semi-Unbalanced Optimal Transport (SemiUOT) shows great promise in matching two probability measures by relaxing one of the marginal constraints. Previous solvers often incorporate an entropy regularization term, which can result in inaccurate matching solutions. To address this issue, we focus on determining the marginal probability distribution of SemiUOT with KL divergence using the proposed Equivalent Transformation Mechanism (ETM) approach. Furthermore, we extend the ETM-based method into exploiting the marginal probability distribution of Unbalanced Optimal Transport (UOT) with KL divergence for validating its generalization. Once the marginal probabilities of UOT/SemiUOT are determined, they can be transformed into a classical Optimal Transport (OT) problem. Moreover, we propose a KKT-Multiplier regularization term combined with Multiplier Regularized Optimal Transport (MROT) to achieve more accurate matching results. We conduct several numerical experiments to demonstrate the effectiveness of our proposed methods in addressing UOT/SemiUOT problems.
Weiming Liu 0005, Xinting Liao, Jun Dan, Fan Wang 0020, Hua Yu 0006, Junhao Dong 0001, Shunjie Dong, Lianyong Qi, Yew-Soon Ong
NeurIPS5
2025 GAM: A Generative Autoencoder for Diverse Human Motion Prediction
Jiapeng Bai, Hua Yu 0006, Yaqing Hou, Qiang Zhang 0008
PRICAI2
2025 A Spatio-Temporal Continuous Network for Stochastic 3D Human Motion Prediction
abstract
Stochastic Human Motion Prediction (HMP) has received increasing attention due to its wide applications. Despite the rapid progress in generative fields, existing methods often face challenges in learning continuous temporal dynamics and predicting stochastic motion sequences. They tend to overlook the flexibility inherent in complex human motions and are prone to mode collapse. To alleviate these issues, we propose a novel method called STCN, for stochastic and continuous human motion prediction, which consists of two stages. Specifically, in the first stage, we propose a spatio-temporal continuous network to generate smoother human motion sequences. In addition, the anchor set is innovatively introduced into the stochastic HMP task to prevent mode collapse, which refers to the potential human motion patterns. In the second stage, STCN endeavors to acquire the Gaussian mixture distribution (GMM) of observed motion sequences with the aid of the anchor set. It also focuses on the probability associated with each anchor, and employs the strategy of sampling multiple sequences from each anchor to alleviate intra-class differences in human motions. Experimental results on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy.
Hua Yu 0006, Yaqing Hou, Xu Gui 0001, Shanshan Feng 0001, Qiang Zhang 0008
IEEE Trans. Circuits Syst. Video Technol.1
2025 DivDiff: A Conditional Diffusion Model for Diverse Human Motion Prediction
abstract
Diverse human motion prediction (HMP) aims to predict multiple plausible future motions given an observed human motion sequence. It is a challenging task due to the diversity of potential human motions while ensuring an accurate description of future human motions. Current solutions are either low-diversity or limited in expressiveness. Recent denoising diffusion probabilistic models (DDPM) demonstrate promising performance in various generative tasks. However, introducing DDPM directly into diverse HMP incurs some issues. While DDPM can enhance the diversity of potential human motion patterns, the predicted human motions gradually become implausible over time due to significant noise disturbances in the forward process of DDPM. This phenomenon leads to the predicted human motions being unrealistic, seriously impacting the quality of predicted motions and restricting their practical applicability in real-world scenarios. To alleviate this, we propose a novel conditional diffusion-based generative model, called DivDiff, to predict more diverse and realistic human motions. Specifically, the DivDiff employs DDPM as our backbone and incorporates Discrete Cosine Transform (DCT) and Transformer mechanisms to encode the observed human motion sequence as a condition to instruct the reverse process of DDPM. More importantly, we design a diversified reinforcement sampling function (DRSF) to enforce human skeletal constraints on the predicted human motions. DRSF utilizes the acquired information from human skeletal as prior knowledge, thereby reducing significant disturbances introduced during the forward process. Extensive results received in the experiments on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy.
Hua Yu 0006, Yaqing Hou, Wenbin Pei, Yew-Soon Ong, Qiang Zhang 0008
IEEE Trans. Multim.1
2024 Similar Locality Based Transfer Evolutionary Optimization for Minimalistic Attacks
abstract
Deep neural networks are powerful and popular learning models; however, recent studies have shown that deep neural network-based policies are susceptible to deception by adversarial attacks. A minimalistic attack is a specialized form of adversarial attack that aims to accomplish successful attacks at the lowest possible cost. Recently, transfer optimization algorithms have been applied to deceive previously trained policies by acquiring knowledge from previously solved tasks. Experiments indicate that the transfer optimization algorithms perform well compared to traditional optimization algorithms. However, current transfer algorithms for addressing minimalistic attacks not only select a single source task for knowledge transfer but also tend to overly rely on identified appropriate source tasks. To address this issue, this paper introduces a similar locality based transfer evolutionary optimization algorithm. It can adaptively select multiple source tasks and extract valuable knowledge from these source tasks. Moreover, by leveraging the concept of similar locality, the algorithm alleviates its excessive dependence on familiar tasks, thereby providing fresh knowledge for the optimization of the target task. On this basis, the algorithm can mine more valuable knowledge from the large source task space to achieve a successful attack in a shorter period. The algorithm is tested on three Atari games-BeamRider, Qbert, and Seaquest-demonstrating its ability and potential to outperform other transfer optimization algorithms currently available in solving this problem.
Wenqiang Ma, Yaqing Hou, Hua Yu 0006, Xiangrong Tong, Zexuan Zhu 0001, Qiang Zhang 0008
CEC3
2024 Towards Efficient and Diverse Generative Model for Unconditional Human Motion Synthesis
abstract
Recent generative methods have revolutionized the way of human motion synthesis, such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Denoising Diffusion Probabilistic Models (DMs). These methods have gained significant attention in human motion fields. However, there are still challenges in unconditionally generating highly diverse human motions from a given distribution. To enhance the diversity of synthesized human motions, previous methods usually employ deep neural networks (DNNs) to train a transport map that transforms Gaussian noise distribution into real human motion distribution. According to Figalli's regularity theory, the optimal transport map computed by DNNs frequently exhibits discontinuities. This is due to the inherent limitation of DNNs in representing only continuous maps. Consequently, the generated human motions tend to heavily concentrate on densely populated regions of the data distribution, resulting in mode collapse or mode mixture. To address the issues, we propose an efficient method called MOOT for unconditional human motion synthesis. First, we utilize a reconstruction network based on GRU and transformer to map human motions to latent space. Next, we employ convex optimization to match the noise distribution with the latent space distribution of human motions through the Optimal Transport (OT) map. Then, we combine the extended OT map with the generator of reconstruction network to generate new human motions. Thereby overcoming the issues of mode collapse and mode mixture. MOOT generates a latent code distribution that is well-behaved and highly structured, providing a strong motion prior for various applications in the field of human motion. Through qualitative and quantitative experiments, MOOT achieves state-of-the-art results surpassing the latest methods, validating its superiority in unconditional human motion generation.
Hua Yu 0006, Weiming Liu 0005, Jiapeng Bai, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008
ACM Multimedia1
2024 TFGDA: Exploring Topology and Feature Alignment in Semi-supervised Graph Domain Adaptation through Robust Clustering
abstract
Semi-supervised graph domain adaptation, as a branch of graph transfer learning, aims to annotate unlabeled target graph nodes by utilizing transferable knowledge learned from a label-scarce source graph. However, most existing studies primarily concentrate on aligning feature distributions directly to extract domain-invariant features, while ignoring the utilization of the intrinsic structure information in graphs. Inspired by the significance of data structure information in enhancing models' generalization performance, this paper aims to investigate how to leverage the structure information to assist graph transfer learning. To this end, we propose an innovative framework called TFGDA. Specially, TFGDA employs a structure alignment strategy named STSA to encode graphs' topological structure information into the latent space, greatly facilitating the learning of transferable features. To achieve a stable alignment of feature distributions, we also introduce a SDA strategy to mitigate domain discrepancy on the sphere. Moreover, to address the overfitting issue caused by label scarcity, a simple but effective RNC strategy is devised to guide the discriminative clustering of unlabeled nodes. Experiments on various benchmarks demonstrate the superiority of TFGDA over SOTA methods.
Jun Dan, Weiming Liu 0005, Chunfeng Xie, Hua Yu 0006, Shunjie Dong, Yanchao Tan
NeurIPS4
2024 Human-Object Interaction detection via Global Context and Pairwise-level Fusion Features Integration
Haozhong Wang, Hua Yu 0006, Qiang Zhang 0008
Neural Networks2
2023 Toward Realistic 3D Human Motion Prediction With a Spatio-Temporal Cross- Transformer Approach
abstract
Human motion prediction intends to predict how humans move given a historical sequence of 3D human motions. Recent transformer-based methods have attracted increasing attentions and demonstrated their promising performance in 3D human motion prediction. However, existing methods generally decompose the input of human motion information into spatial and temporal branches in a separate way and seldom consider their inherent coherence between the two branches, hence often failing to register the dynamic spatio-temporal information during the training process. Motivated by these issues, we propose a spatio-temporal cross-transformer network (STCT) for 3D human motion predictions. Specifically, we investigate various types of interaction methods (i.e., Concatenation Interaction, Msg token interaction, and Cross-transformer) to capture the coherence of the spatial and temporal branches. According to the obtained results, the proposed cross-transformer interaction method shows its superiority over other methods. Meanwhile, considering that most existing works treat the human body as a set of 3D human joint positions, the predicted human joints are proportionally less appropriate to the realistic human body due to unreasonable bone length and non-plausible poses as time progresses. We further resort to the bone constraints of human mesh to produce more realistic human motions. By fitting a parametric body model (i.e., SMPL-X model) to the predicted human joints, a reconstruction loss function is proposed to remedy the unreasonable bone length and pose errors. Comprehensive experiments on AMASS and Human3.6M datasets have demonstrated that our method achieves superior performance over compared methods.
Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Wenbin Pei, Hong-Wei Ge, Xin Yang 0011, Qiang Zhang 0008, Mengjie Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 A Preliminary Study of Multi-task MAP-Elites with Knowledge Transfer for Robotic Arm Design
abstract
The structure design of robotic arms is of great importance on completing industrial tasks successfully. This is a typical multi-task optimization problem when considering different constraints as different tasks. However, mainstream methods for multi-task optimization such as evolutionary multitasking and Multi-task MAP-Elites algorithms tend to encounter problems such as high computational cost and slow convergence when solving large-scale robotic arm tasks. To this end, this paper proposes a new framework based on the MAP-Elites algorithms for solving large-scale robot arm design tasks, called Multi-task MAP-Elites with Knowledge Transfer (MMKT). Specifically, this paper designs the group-based knowledge transfer process for large-scale task optimization in which all tasks are classified into different groups according to their similarity to generate multiple knowledge transfer areas; and knowledge transfer strategies are designed to enhance the quality of solutions with low fitness value. We test the effectiveness of the MMKT framework in planar robotic arm experiments (2000, 5000, and 10,000 tasks; 10, 15-dimensional search space). The experimental results prove that the MMKT outperforms the MME, CMA-ES, and classical ES algorithms.
Hua Yu 0006, Han Linghu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008
CEC3
2022 Towards Efficient 3D Human Motion Prediction using Deformable Transformer-based Adversarial Network
abstract
Human motion prediction is a crucial step for achieving human-robot interactions. While recent transformer-based methods have shown great potentials in 3D human motion prediction, they still suffer from mode collapse to non-plausible poses and quadratically computational complexity with respect to the increasing length of input sequences. In this paper, we propose a novel spatio-temporal deformable transformer-based adversarial network (STDTA) for 3D human motion prediction. First, we design a spatio-temporal deformable transformer module to capture the correlations between human joints while reducing the computational costs. Second, we introduce the adversarial training mechanism and design fidelity and continuity discriminators to maintain smoothness and stability for the long-term prediction. Finally, extensive experiments on Human 3.6M and AMASS benchmarks demonstrate that the proposed STDTA achieves state-of-the-art performance.
Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Cai Kang, Qiang Zhang 0008
ICRA1
2021 Brain Tumor Segmentation based on Knowledge Distillation and Adversarial Training
abstract
3D MRI brain tumor segmentation is a reliable method for disease diagnosis and treatment plans in the future. Early on, the segmentation of brain tumors is mostly done manually. However, manual segmentation of 3D MRI brain tumor requires professional anatomical knowledge and may be inaccurate. In this paper, we propose a 3D MRI brain tumor segmentation architecture based on the encoder-decoder structure. Specially, we introduce knowledge distillation and adversarial training methods, which compresses and improves the accuracy and robustness of the model. Furthermore, we obtain soft targets by designing multiple teacher network training and then apply them to the student network. Finally, we evaluate our method on a challenging BraTS dataset. As a result, the performance of our proposed model is superior to state-of-the-art methods.
Yaqing Hou, Tianbo Li, Qiang Zhang 0008, Hua Yu 0006, Hong-Wei Ge
IJCNN4
2021 Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognition
abstract
Abstract In the study of human action recognition, two-stream networks have made excellent progress recently. However, there remain challenges in distinguishing similar human actions in videos. This paper proposes a novel local-aware spatio-temporal attention network with multi-stage feature fusion based on compact bilinear pooling for human action recognition. To elaborate, taking two-stream networks as our essential backbones, the spatial network first employs multiple spatial transformer networks in a parallel manner to locate the discriminative regions related to human actions. Then, we perform feature fusion between the local and global features to enhance the human action representation. Furthermore, the output of the spatial network and the temporal information are fused at a particular layer to learn the pixel-wise correspondences. After that, we bring together three outputs to generate the global descriptors of human actions. To verify the efficacy of the proposed approach, comparison experiments are conducted with the traditional hand-engineered IDT algorithms, the classical machine learning methods (i.e., SVM) and the state-of-the-art deep learning methods (i.e., spatio-temporal multiplier networks). According to the results, our approach is reported to obtain the best performance among existing works, with the accuracy of 95.3% and 72.9% on UCF101 and HMDB51, respectively. The experimental results thus demonstrate the superiority and significance of the proposed architecture in solving the task of human action recognition.
Yaqing Hou, Hua Yu 0006, Pengfei Wang 0013, Hong-Wei Ge, Jianxin Zhang 0001, Qiang Zhang 0008
Neural Comput. Appl.2