EDBT 2026 Demo / reviewers in the wild / expert
Qianhui Men
dblp:239/9182
· DBLP profile ↗
20ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-0059-5484ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Based Motion Tracking of Contact-Rich Interacting Characters
Xiaotang Zhang, Ziyi Chang, Qianhui Men, Hubert P. H. Shum |
Comput. Graph. Forum | 3 |
| 2026 | Physics-Based Motion Tracking of Contact-Rich Interacting CharactersabstractAbstract Motion tracking has been an important technique for imitating human‐like movement from large‐scale datasets in physics‐based motion synthesis. However, existing approaches focus on tracking either single character or a particular type of interaction, limiting their ability to handle contact‐rich interactions. Extending single‐character tracking approaches suffers from the instability due to the challenge of forces transferred through contacts. Contact‐rich interactions requires levels of control, which places much greater demands on model capacity. To this end, we propose a robust tracking method based on progressive neural network (PNN) where multiple experts are specialized in learning skills of various difficulties. Our method learns to assign training samples to experts automatically without requiring manually scheduling. Both qualitative and quantitative results show that our method delivers more stable motion tracking in densely interactive movements while enabling more efficient model training. Xiaotang Zhang, Ziyi Chang, Qianhui Men, Hubert P. H. Shum |
Comput. Graph. Forum | 3 |
| 2025 | Trustworthy and Practical AI for Healthcare: A Guided Deferral System with Large Language ModelsabstractLarge language models (LLMs) offer a valuable technology for various applications in healthcare. However, their tendency to hallucinate and the existing reliance on proprietary systems pose challenges in environments concerning critical decision-making and strict data privacy regulations, such as healthcare, where the trust in such systems is paramount. Through combining the strengths and discounting the weaknesses of humans and AI, the field of Human-AI Collaboration (HAIC) presents one front for tackling these challenges and hence improving trust. This paper presents a novel HAIC \textit{guided deferral} system that can simultaneously parse medical reports for disorder classification, and defer uncertain predictions with intelligent guidance to humans. We develop methodology which builds efficient, effective and open-source LLMs for this purpose, for the real-world deployment in healthcare. We conduct a pilot study which showcases the effectiveness of our proposed system in practice. Additionally, we highlight drawbacks of standard calibration metrics in imbalanced data scenarios commonly found in healthcare, and suggest a simple yet effective solution: the Imbalanced Expected Calibration Error. Joshua Strong, Qianhui Men, J. Alison Noble |
AAAI | 2 |
| 2025 | Motion In-Betweening for Densely Interacting CharactersabstractMotion in-betweening is the problem to synthesize movement between keyposes. Traditional research focused primarily on single characters. Extending them to densely interacting characters is highly challenging, as it demands precise spatial-temporal correspondence between the characters to maintain the interaction, while creating natural transitions towards predefined keyposes. In this research, we present a method for long-horizon interaction in-betweening that enables two characters to engage and respond to one another naturally. To effectively represent and synthesize interactions, we propose a novel solution called Cross-Space In-Betweening, which models the interactions of each character across different conditioning representation spaces. We further observe that the significantly increased constraints in interacting characters heavily limit the solution space, leading to degraded motion quality and diminished interaction over time. To enable long-horizon synthesis, we present two solutions to maintain long-term interaction and motion quality, thereby keeping synthesis in the stable region of the solution space. We first sustain interaction quality by identifying periodic interaction patterns through adversarial learning. We further maintain the motion quality by learning to refine the drifted latent space and prevent pose error accumulation. We demonstrate that our approach produces realistic, controllable, and long-horizon in-between motions of two characters with dynamic boxing and dancing actions across multiple keyposes, supported by extensive quantitative evaluations and user studies. Xiaotang Zhang, Ziyi Chang, Qianhui Men, Hubert P. H. Shum |
SIGGRAPH Asia | 3 |
| 2025 | Real-Time and Controllable Reactive Motion Synthesis via Intention GuidanceabstractAbstract We propose a real‐time method for reactive motion synthesis based on the known trajectory of input character, predicting instant reactions using only historical, user‐controlled motions. Our method handles the uncertainty of future movements by introducing an intention predictor, which forecasts key joint intentions to make pose prediction more deterministic from the historical interaction. The intention is later encoded into the latent space of its reactive motion, matched with a codebook which represents mappings between input and output. It samples from the categorical distribution for pose generation and strengthens model robustness through adversarial training. Unlike previous offline approaches, the system can recursively generate intentions and reactive motions using feedback from earlier steps, enabling real‐time, long‐term realistic interactive synthesis. Both quantitative and qualitative experiments show our approach outperforms other matching‐based motion synthesis approaches, delivering superior stability and generalisability. In our method, user can also actively influence the outcome by controlling the moving directions, creating a personalised interaction path that deviates from predefined trajectories. Xiaotang Zhang, Ziyi Chang, Qianhui Men, Hubert P. H. Shum |
Comput. Graph. Forum | 3 |
| 2025 | ScanAhead: Simplifying standard plane acquisition of fetal head ultrasoundabstractThe fetal standard plane acquisition task aims to detect an Ultrasound (US) image characterized by specified anatomical landmarks and appearance for assessing fetal growth. However, in practice, due to variability in human operator skill and possible fetal motion, it can be challenging for a human operator to acquire a satisfactory standard plane. To support a human operator with this task, this paper first describes an approach to automatically predict the fetal head standard plane from a video segment approaching the standard plane. A transformer-based image predictor is proposed to produce a high-quality standard plane by understanding diverse scales of head anatomy within the US video frame. Because of the visual gap between the video frames and standard plane image, the predictor is equipped with an offset adaptor that performs domain adaption to translate the off-plane structures to the anatomies that would usually appear in a standard plane view. To enhance the anatomical details of the predicted US image, the approach is extended by utilizing a second modality, US probe movement, that provides 3D location information. Quantitative and qualitative studies conducted on two different head biometry planes demonstrate that the proposed US image predictor produces clinically plausible standard planes with superior performance to comparative published methods. The results of dual-modality solution show an improved visualization with enhanced anatomical details of the predicted US image. Clinical evaluations are also conducted to demonstrate the consistency between the predicted echo textures and the expected echo patterns seen in a typical real standard plane, which indicates its clinical feasibility for improving the standard plane acquisition process. Qianhui Men, He Zhao 0002, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 1 |
| 2024 | MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video
Xiaoqing Guo, Qianhui Men, J. Alison Noble |
MICCAI (4) | 2 |
| 2024 | Pose-GuideNet: Automatic Scanning Guidance for Fetal Head Ultrasound from Pose Estimation
Qianhui Men, Xiaoqing Guo, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (4) | 1 |
| 2023 | Focalized contrastive view-invariant learning for self-supervised skeleton-based action recognition
Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
Neurocomputing | 1 |
| 2023 | Gaze-probe joint guidance with multi-task learning in obstetric ultrasound scanningabstractIn this work, we exploit multi-task learning to jointly predict the two decision-making processes of gaze movement and probe manipulation that an experienced sonographer would perform in routine obstetric scanning. A multimodal guidance framework, Multimodal-GuideNet, is proposed to detect the causal relationship between a real-world ultrasound video signal, synchronized gaze, and probe motion. The association between the multi-modality inputs is learned and shared through a modality-aware spatial graph that leverages useful cross-modal dependencies. By estimating the probability distribution of probe and gaze movements in real scans, the predicted guidance signals also allow inter- and intra-sonographer variations and avoid a fixed scanning path. We validate the new multi-modality approach on three types of obstetric scanning examinations, and the result consistently outperforms single-task learning under various guidance policies. To simulate sonographer's attention on multi-structure images, we also explore multi-step estimation in gaze guidance, and its visual results show that the prediction allows multiple gaze centers that are substantially aligned with underlying anatomical structures. Qianhui Men, Clare Teng, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 1 |
| 2022 | Geometric Features Informed Multi-person Human-Object Interaction Recognition in Videos
Tanqiu Qiao, Qianhui Men, Frederick W. B. Li, Yoshiki Kubotani, Shigeo Morishima, Hubert P. H. Shum |
ECCV (4) | 2 |
| 2022 | Multimodal-GuideNet: Gaze-Probe Bidirectional Guidance in Obstetric Ultrasound Scanning
Qianhui Men, Clare Teng, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (8) | 1 |
| 2022 | GAN-based reactive motion synthesis with class-aware discriminators for human-human interaction
Qianhui Men, Hubert P. H. Shum, Edmond S. L. Ho, Howard Leung |
Comput. Graph. | 1 |
| 2022 | Interaction Mix and Match: Synthesizing Close Interaction using Conditional Hierarchical GAN with Multi-Hot Class EmbeddingabstractAbstract Synthesizing multi‐character interactions is a challenging task due to the complex and varied interactions between the characters. In particular, precise spatiotemporal alignment between characters is required in generating close interactions such as dancing and fighting. Existing work in generating multi‐character interactions focuses on generating a single type of reactive motion for a given sequence which results in a lack of variety of the resultant motions. In this paper, we propose a novel way to create realistic human reactive motions which are not presented in the given dataset by mixing and matching different types of close interactions. We propose a Conditional Hierarchical Generative Adversarial Network with Multi‐Hot Class Embedding to generate the Mix and Match reactive motions of the follower from a given motion sequence of the leader. Experiments are conducted on both noisy (depth‐based) and high‐quality (MoCap‐based) interaction datasets. The quantitative and qualitative results show that our approach outperforms the state‐of‐the‐art methods on the given datasets. We also provide an augmented dataset with realistic reactive motions to stimulate future research in this area. Aman Goel, Qianhui Men, Edmond S. L. Ho |
Comput. Graph. Forum | 2 |
| 2021 | Semantics-STGCNN: A Semantics-guided Spatial-Temporal Graph Convolutional Network for Multi-class Trajectory PredictionabstractPredicting the movement trajectories of multiple classes of road users in real-world scenarios is a challenging task due to the diverse trajectory patterns. While recent works of pedestrian trajectory prediction successfully modelled the influence of surrounding neighbours based on the relative distances, they are ineffective on multi-class trajectory prediction. This is because they ignore the impact of the implicit correlations between different types of road users on the trajectory to be predicted—for example, a nearby pedestrian has a different level of influence from a nearby car. In this paper, we propose to introduce class information into a graph convolutional neural network to better predict the trajectory of an individual. We embed the class labels of the surrounding objects into the label adjacency matrix (LAM), which is combined with the velocity-based adjacency matrix (VAM) comprised of the objects’ velocity, thereby generating a semantics-guided graph adjacency (SAM). SAM effectively models semantic information with trainable parameters to automatically learn the embedded label features that will contribute to the fixed velocity-based trajectory. Such information of spatial and temporal dependencies is passed to a graph convolutional and temporal convolutional network to estimate the predicted trajectory distributions. We further propose new metrics, known as Average2Displacement Error (aADE) and Average Final Displacement Error (aFDE), that assess network accuracy more accurately. We call our framework Semantics-STGCNN. It consistently shows superior performance to the state-of-the-arts in existing and the newly proposed metrics. Ben A. Rainbow, Qianhui Men, Hubert P. H. Shum |
SMC | 2 |
| 2021 | A Quadruple Diffusion Convolutional Recurrent Network for Human Motion PredictionabstractRecurrent neural network (RNN) has become popular for human motion prediction thanks to its ability to capture temporal dependencies. However, it has limited capacity in modeling the complex spatial relationship in the human skeletal structure. In this work, we present a novel diffusion convolutional recurrent predictor for spatial and temporal movement forecasting, with multi-step random walks traversing bidirectionally along an adaptive graph to model interdependency among body joints. In the temporal domain, existing methods rely on a single forward predictor with the produced motion deflecting to the drift route, which leads to error accumulations over time. We propose to supplement the forward predictor with a forward discriminator to alleviate such motion drift in the long term under adversarial training. The solution is further enhanced by a backward predictor and a backward discriminator to effectively reduce the error, such that the system can also look into the past to improve the prediction at early frames. The two-way spatial diffusion convolutions and two-way temporal predictors together form a quadruple network. Furthermore, we train our framework by modeling the velocity from observed motion dynamics instead of static poses to predict future movements that effectively reduces the discontinuity problem at early prediction. Our method outperforms the state of the arts on both 3D and 2D datasets, including the Human3.6M, CMU Motion Capture and Penn Action datasets. The results also show that our method correctly predicts both high-dynamic and low-dynamic moving trends with less motion drift. Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | A Two-Stream Recurrent Network for Skeleton-based Human Interaction RecognitionabstractThis paper addresses the problem of recognizing human-human interaction from skeletal sequences. Existing methods are mainly designed to classify single human action. Many of them simply stack the movement features of two characters to deal with human interaction, while neglecting the abundant relationships between characters. In this paper, we propose a novel two-stream recurrent neural network by adopting the geometric features from both single actions and interactions to describe the spatial correlations with different discriminative abilities. The first stream is constructed under pairwise joint distance (PJD) in a fully-connected mesh to categorize the interactions with explicit distance patterns. To better distinguish similar interactions, in the second stream, we combine PJD with the spatial features from individual joint positions using graph convolutions to detect the implicit correlations among joints, where the joint connections in the graph are adaptive for flexible correlations. After spatial modeling, each stream is fed to a bi-directional LSTM to encode two-way temporal properties. To take advantage of the diverse discriminative power of the two streams, we come up with a late fusion algorithm to combine their output predictions concerning information entropy. Experimental results show that the proposed framework achieves state-of-the-art performance on 3D and comparable performance on 2D interaction datasets. Moreover, the late fusion results demonstrate the effectiveness of improving the recognition accuracy compared with single streams. Qianhui Men, Edmond S. L. Ho, Hubert P. H. Shum, Howard Leung |
ICPR | 1 |
| 2020 | Hierarchical Visual-aware Minimax Ranking Based on Co-purchase Data for Personalized RecommendationabstractPersonalized recommendation aims at ranking a set of items according to the learnt preferences of the user. Existing methods optimize the ranking function by considering an item that the user has not bought yet as a negative item and assuming that the user prefers the positive item that he has bought to the negative item. The strategy is to exclude irrelevant items from the dataset to narrow down the set of potential positive items to improve ranking accuracy. It conflicts with the goal of recommendation from the seller’s point of view, which aims to enlarge that set for each user. In this paper, we diminish this limitation by proposing a novel learning method called Hierarchical Visual-aware Minimax Ranking (H-VMMR), in which a new concept of predictive sampling is proposed to sample items in a close relationship with the positive items (e.g., substitutes, compliments). We set up the problem by maximizing the preference discrepancy between positive and negative items, as well as minimizing the gap between positive and predictive items based on visual features. We also build a hierarchical learning model based on co-purchase data to solve the data sparsity problem. Our method is able to enlarge the set of potential positive items as well as true negative items during ranking. The experimental results show that our H-VMMR outperforms the state-of-the-art learning methods. Xiaoya Chong, Qing Li 0001, Howard Leung, Qianhui Men, Xianjin Chao |
WWW | 4 |
| 2019 | Retrieval of spatial-temporal motion topics from 3D skeleton data
Qianhui Men, Howard Leung |
Vis. Comput. | 1 |
| 2019 | Self-feeding frequency estimation and eating action recognition from skeletal representation using Kinect
Qianhui Men, Howard Leung, Yang Yang 0046 |
World Wide Web | 1 |