EDBT 2026 Demo / reviewers in the wild / expert
Yixing Gao 0001
dblp:64/10194-1
· DBLP profile ↗
20ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-4475-2792ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective Robotic Cloth Grasping Through Suppressing False DiscoveriesabstractEnabling robots to grasp disorganized cloth for efficient storage is valuable in robot-assisted room organization. Diverse deformations of cloth and the stacking of multiple items limit grasping-pose estimation that relies on annotations. This necessitates segmenting each cloth item in an unsupervised manner before estimating the grasping position. However, existing segmentation methods primarily focus on improving metrics such as Intersection-over-Union and Pixel Accuracy, which cannot effectively measure the segmentation errors of the cloth area and thus lead to failure grasping position estimation. To address this challenge, we use False Discovery Rate (FDR) as a novel measure of segmentation errors and analyze its impact on grasping success. Our preliminary study reveals a negative correlation between segmentation FDR and grasping success rate, highlighting the need for more reliable segmentation in cluttered cloth scenarios. Therefore, we propose an unsupervised cloth segmentation network based on feature distance-weighted constraints, designed to reduce the false discovery rate in cloth area perception without requiring expensive pixel-level manual annotations. Additionally, to estimate the grasping position on the perceived cloth area, we introduce a strategy based on cloth surface wrinkle analysis, which operates without the need for annotations or training. By integrating the proposed segmentation network and grasping strategy, we develop a robotic system capable of sequentially grasping cluttered cloth from a table. Extensive real-world robotic experiments demonstrate the effectiveness of our approach, outperforming multiple baseline methods in segmentation FDR and grasping success rate. Xingyu Zhu 0014, Zhiwen Tu, Yan Wu 0002, Shan Luo 0001, Hechang Chen, Yixing Gao 0001 |
AAAI | 6 |
| 2026 | Video-based human pose estimation via feature decoupling and multi-hypothesis calibration
Runyang Feng, Tze Ho Elden Tse, Haoming Chen, Hyung Jin Chang, Haifeng Zhong, Yixing Gao 0001 |
Pattern Recognit. | 6 |
| 2026 | DualAdaptNet: Enhancing domain adaptation regression with error accumulation reduction and dual-head heterogeneous learning
Xin Wang 0035, Yuan Wu 0002, Xingyu Zhu 0014, Yi Chang 0001, Yixing Gao 0001 |
Pattern Recognit. | 5 |
| 2025 | Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using SuperquadricsabstractWith the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition. Tze Ho Elden Tse, Runyang Feng, Linfang Zheng, Yixing Gao 0001, Jihie Kim, Ales Leonardis, Hyung Jin Chang |
AAAI | 5 |
| 2025 | High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationabstractModeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods typically unify spatiotemporal learning within a single type of modeling structure (convolution or attention-based blocks), which inherently have difficulties in balancing global and local dynamic modeling and may bias the network to one of them, leading to suboptimal performance. Moreover, existing VHPE models suffer from quadratic complexity when capturing global dependencies, limiting their applicability especially for high-resolution sequences. Recently, the state space models (known as Mamba) have demonstrated significant potential in modeling long-range contexts with linear complexity; however, they are restricted to 1D sequential data. In this paper, we present a novel framework that extends Mamba from two aspects to separately learn global and local high-resolution spatiotemporal representations for VHPE. Specifically, we first propose a Global Spatiotemporal Mamba, which performs 6D selective space-time scan and spatial- and temporal-modulated scan merging to efficiently extract global representations from high-resolution sequences. We further introduce a windowed space-time scan-based Local Refinement Mamba to enhance the high-frequency details of localized keypoint motions. Extensive experiments on four benchmark datasets demonstrate that the proposed model outperforms state-of-the-art VHPE approaches while achieving better computational trade-offs. Runyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim, Yi Chang 0001, Yixing Gao 0001 |
ICCV | 6 |
| 2025 | AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationabstractThe challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additional multimodal fusion loss functions to establish fusion features. However, these approaches often overlook the challenge caused by feature discrepancies between modalities during the fusion process. To achieve precise segmentation, we propose an Attention-Driven Multimodal Discrepancy Alignment Network (AMDANet). AMDANet reallocates weights to reduce the saliency of discrepant features and utilizes low-weight features as cues to mitigate discrepancies between modalities, thereby achieving multimodal feature alignment. Furthermore, to simplify the feature alignment process, a semantic consistency inference mechanism is introduced to reveal the network’s inherent bias toward specific modalities, thereby compressing cross-modal feature discrepancies from the foundational level. Extensive experiments on the FMB, MFNet, and PST900 datasets demonstrate that AMDANet achieves mIoU improvements of 3.6%, 3.0%, and 1.6%, respectively, significantly outperforming state-of-the-art methods. The code is available at https://github.com/Zhonghaifeng6/AMDANet Haifeng Zhong, Fan Tang, Zhuo Chen 0028, Hyung Jin Chang, Yixing Gao 0001 |
ICCV | 5 |
| 2025 | DarkSeg: Infrared-Driven Semantic Segmentation for Garment Grasping Detection in Low-Light ConditionsabstractGarment grasping in low-light environments is a critical challenge for domestic intelligent robots, yet existing research has not sufficiently addressed this issue. In low-light conditions, the scarcity of visual features due to insufficient illumination causes different categories of garments to exhibit ambiguous feature similarities, thereby hindering the robot’s ability to detect the categories of different garments. Although traditional methods can compensate for visual deficiencies in low-light scenarios by applying preprocessing strategies that fuse infrared multimodal features, their complex computational processes incur significant computational overhead. To address this limitation, we propose a low-light garment detection model based on the student-teacher model. The innovation of DarkSeg lies in its replacement of complex multimodal feature fusion with an indirect feature alignment mechanism between the student and teacher models, thereby circumventing high computational demands. Through feature alignment, DarkSeg enables the student model to learn illumination-invariant structural representations from the infrared features provided by the teacher model, effectively correcting structural deficiencies in low-light environments. Furthermore, to evaluate DarkSeg’s feasibility for low-light clothing grasping, we propose a depth-perceptive grasping strategy and build a low-light multimodal garment detection dataset, DarkClothes. Extensive experiments deploying DarkSeg on a Baxter robot demonstrate that DarkSeg achieves a 22% improvement in the grasping success rate while reducing the model parameters by 99.08 million compared to traditional methods, validating the practical viability of DarkSeg for robotic garment grasping in low-light conditions. The code and dataset are available at https://github.com/Zhonghaifeng6/Darkseg Haifeng Zhong, Fan Tang, Hyung Jin Chang, Xingyu Zhu 0014, Yixing Gao 0001 |
IROS | 5 |
| 2025 | Generalizable Category-Level Topological Structure Learning for Clothing Recognition in Robotic GraspingabstractRecognizing various types of clothing is crucial for robotic clothing manipulation tasks, such as garment organization and robot-assisted dressing. Unlike rigid object recognition, clothing recognition remains a challenging task due to the diverse forms introduced by flexible deformations. However, existing classification models primarily focus on clothing color and texture while overlooking structural features, limiting their ability to distinguish between deformable clothing categories with similar color and texture. Moreover, due to the insufficient representation of structural features, these models heavily rely on manually annotated labels, making it difficult to accurately recognize unseen clothing items with new colors or textures. To address these challenges, we propose a novel topological structure representation and optimization strategy for category-level clothing structural feature learning. Additionally, we design a multi-clothing classification framework based on multiple mask generation to identify clothing regions within a scene. By leveraging our proposed structural feature learning strategy, our framework effectively generalizes to unseen clothing items. Finally, we introduce a fabric-specific grasping position estimation method and develop a corresponding robotic grasping system capable of selecting and grasping specified clothing items based on user instructions. Extensive real-world robotic experiments demonstrate the effectiveness of our system, and comprehensive comparisons with multiple baselines further validate the superiority of our approach. Xingyu Zhu 0014, Yan Wu 0002, Zhiwen Tu, Haifeng Zhong, Yixing Gao 0001 |
IROS | 5 |
| 2024 | Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition
Yunbing Jia, Xiaoyu Kong, Fan Tang, Yixing Gao 0001, Weiming Dong |
IJCAI | 4 |
| 2024 | JointLoc: A Real-time Visual Localization Framework for Planetary UAVs Based on Joint Relative and Absolute Pose EstimationabstractUnmanned aerial vehicles (UAVs) visual localization in planetary aims to estimate the absolute pose of the UAV in the world coordinate system through satellite maps and images captured by on-board cameras. However, since planetary scenes often lack significant landmarks and there are modal differences between satellite maps and UAV images, the accuracy and real-time performance of UAV positioning will be reduced. In order to accurately determine the position of the UAV in a planetary scene in the absence of the global navigation satellite system (GNSS), this paper proposes JointLoc, which estimates the real-time UAV position in the world coordinate system by adaptively fusing the absolute 2-degree-of-freedom (2-DoF) pose and the relative 6-degree-of-freedom (6-DoF) pose. Extensive comparative experiments were conducted on a proposed planetary UAV image cross-modal localization dataset, which contains three types of typical Martian topography generated via a simulation engine as well as real Martian UAV images from the Ingenuity helicopter. JointLoc achieved a root-mean-square error of 0.237m in the trajectories of up to 1,000m, compared to 0.594m and 0.557m for ORB-SLAM2 and ORB-SLAM3 respectively. The source code will be available at https://github.com/LuoXubo/JointLoc. Xubo Luo, Xue Wan, Yixing Gao 0001, Yaolin Tian, Wei Zhang 0260, Leizheng Shu |
IROS | 3 |
| 2024 | Row-Column Separated Attention Based Low-Light Image/Video EnhancementabstractAbstract U‐Net structure is widely used for low‐light image/video enhancement. The enhanced images result in areas with large local noise and loss of more details without proper guidance for global information. Attention mechanisms can better focus on and use global information. However, attention to images could significantly increase the number of parameters and computations. We propose a Row–Column Separated Attention module (RCSA) inserted after an improved U‐Net. The RCSA module's input is the mean and maximum of the row and column of the feature map, which utilizes global information to guide local information with fewer parameters. We propose two temporal loss functions to apply the method to low‐light video enhancement and maintain temporal consistency. Extensive experiments on the LOL, MIT Adobe FiveK image, and SDSD video datasets demonstrate the effectiveness of our approach. Chengqi Dong, Tuoshi Qi, Kexin Wu, Yixing Gao 0001, Fan Tang |
Comput. Graph. Forum | 5 |
| 2024 | GL-GNN: Graph learning via the network of graphs
Yixiang Shan, Jielong Yang, Yixing Gao 0001 |
Knowl. Based Syst. | 3 |
| 2023 | Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in VideoabstractTemporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excavate meaningful motion priors, their results are suboptimal, especially in complicated spatio-temporal interactions. On the other hand, the temporal difference has the ability to encode representative motion information which can potentially be valuable for pose estimation but has not been fully exploited. In this paper, we present a novel multi-frame human pose estimation framework, which employs temporal differences across frames to model dynamic contexts and engages mutual information objectively to facilitate useful motion information disentanglement. To be specific, we design a multi-stage Temporal Difference Encoder that performs incremental cascaded learning conditioned on multi-stage feature difference sequences to derive informative motion representation. We further propose a Representation Disentanglement module from the mutual information perspective, which can grasp discriminative task-relevant motion signals by explicitly defining useful and noisy constituents of the raw motion features and minimizing their mutual information. These place us to rank No.1 in the Crowd Pose Estimation in Complex Events Challenge on benchmark dataset HiEve, and achieve state-of-the-art performance on three benchmarks PoseTrack2017, PoseTrack2018, and PoseTrack21. Runyang Feng, Yixing Gao 0001, Xueqing Ma, Tze Ho Elden Tse, Hyung Jin Chang |
CVPR | 2 |
| 2023 | DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationabstractDenoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to multi-frame human pose estimation is non-trivial due to the presence of the additional temporal dimension in videos. More importantly, learning representations that focus on keypoint regions is crucial for accurate localization of human joints. Nevertheless, the adaptation of the diffusion-based methods remains unclear on how to achieve such objective. In this paper, we present DiffPose, a novel diffusion architecture that formulates video-based human pose estimation as a conditional heatmap generation problem. First, to better leverage temporal information, we propose SpatioTemporal Representation Learner which aggregates visual evidences across frames and uses the resulting features in each denoising step as a condition. In addition, we present a mechanism called Lookup-based Multi-Scale Feature Interaction that determines the correlations between local joints and global contexts across multiple scales. This mechanism generates delicate representations that focus on keypoint regions. Altogether, by extending diffusion models, we show two unique characteristics from DiffPose on pose estimation task: (i) the ability to combine multiple sets of pose estimates to improve prediction accuracy, particularly for challenging joints, and (ii) the ability to adjust the number of iterative steps for feature refinement without retraining the model. DiffPose sets new state-of-the-art results on three benchmarks: PoseTrack2017, PoseTrack2018, and PoseTrack21. Runyang Feng, Yixing Gao 0001, Tze Ho Elden Tse, Xueqing Ma, Hyung Jin Chang |
ICCV | 2 |
| 2023 | Clothes Grasping and Unfolding Based on RGB-D Semantic SegmentationabstractClothes grasping and unfolding is a core step in robotic-assisted dressing. Most existing works leverage depth images of clothes to train a deep learning-based model to recognize suitable grasping points. These methods often utilize physics engines to synthesize depth images to reduce the cost of real labeled data collection. However, the natural domain gap between synthetic and real images often leads to poor performance of these methods on real data. Furthermore, these approaches often struggle in scenarios where grasping points are occluded by the clothing item itself. To address the above challenges, we propose a novel Bi-directional Fractal Cross Fusion Network (BiFCNet) for semantic segmentation, enabling recognition of graspable regions in order to provide more possibilities for grasping. Instead of using depth images only, we also utilize RGB images with rich color features as input to our network in which the Fractal Cross Fusion (FCF) module fuses RGB and depth data by considering global complex features based on fractal geometry. To reduce the cost of real data collection, we further propose a data augmentation method based on an adversarial strategy, in which the color and geometric transformations simultaneously process RGB and depth data while maintaining the label correspondence. Finally, we present a pipeline for clothes grasping and unfolding from the perspective of semantic segmentation, through the addition of a strategy for grasp point selection from segmentation regions based on clothing flatness measures, while taking into account the grasping direction. We evaluate our BiFCNet on the public dataset NYUDv2 and obtained comparable performance to current state-of-the-art models. We also deploy our model on a Baxter robot, running extensive grasping and unfolding experiments as part of our ablation studies, achieving an 84% success rate. Xingyu Zhu 0014, Xin Wang 0035, Jonathan Freer, Hyung Jin Chang, Yixing Gao 0001 |
ICRA | 5 |
| 2023 | Knowing Before Seeing: Incorporating Post-retrieval Information into Pre-retrieval Query Intention Classification
Xueqing Ma, Xiaochi Wei, Yixing Gao 0001, Runyang Feng, Dawei Yin 0001, Yi Chang 0001 |
KSEM (2) | 3 |
| 2023 | Towards fidelity of graph data augmentation via equivariance
Bai Zhang, Yixing Gao 0001, Linbo Xie, Xiaofeng Cao 0002, Yixiang Shan, Jielong Yang |
Knowl. Based Syst. | 2 |
| 2022 | Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationabstractMulti-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art methods strive to incorporate additional visual evidences from neighboring frames (supporting frames) to facilitate the pose estimation of the current frame (key frame). One aspect that has been obviated so far, is the fact that current methods directly aggregate unaligned contexts across frames. The spatial-misalignment between pose features of the current frame and neighboring frames might lead to unsatisfactory results. More importantly, existing approaches build upon the straightforward pose estimation loss, which unfortunately cannot constrain the network to fully leverage useful information from neighboring frames. To tackle these problems, we present a novel hierarchical alignment framework, which leverages coarse-to-fine deformations to progressively update a neighboring frame to align with the current frame at the feature level. We further propose to explicitly supervise the knowledge extraction from neighboring frames, guaranteeing that useful complementary cues are extracted. To achieve this goal, we theoretically analyzed the mutual information between the frames and arrived at a loss that maximizes the task-relevant mutual information. These allow us to rank No.1 in the Multi-frame Person Pose Estimation Challenge on benchmark dataset PoseTrack2017, and obtain state-of-the-art performance on benchmarks Sub-JHMDB and Pose-Track2018. Our code is released at https://github.com/Pose-Group/FAMI-Pose, hoping that it will be useful to the community. Zhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 0002, Yixing Gao 0001, Yunjun Gao, Xiang Wang 0010 |
CVPR | 5 |
| 2016 | Iterative path optimisation for personalised dressing assistance using vision and force informationabstractWe propose an online iterative path optimisation method to enable a Baxter humanoid robot to assist human users to dress. The robot searches for the optimal personalised dressing path using vision and force sensor information: vision information is used to recognise the human pose and model the movement space of upper-body joints; force sensor information is used for the robot to detect external force resistance and to locally adjust its motion. We propose a new stochastic path optimisation method based on adaptive moment estimation. We first compare the proposed method with other path optimisation algorithms on synthetic data. Experimental results show that the performance of the method achieves the smallest error with fewer iterations and less computation time. We also evaluate real-world data by enabling the Baxter robot to assist real human users with their dressing. Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris |
IROS | 1 |
| 2015 | User modelling for personalised dressing assistance by humanoid robotsabstractAssistive robots can improve the well-being of disabled or frail human users by reducing the burden that activities of daily living impose on them. To enable personalised assistance, such robots benefit from building a user-specific model, so that the assistance is customised to the particular set of user abilities. In this paper, we present an end-to-end approach for home-environment assistive humanoid robots to provide personalised assistance through a dressing application for users who have upper-body movement limitations. We use randomised decision forests to estimate the upper-body pose of users captured by a top-view depth camera, and model the movement space of upper-body joints using Gaussian mixture models. The movement space of each upper-body joint consists of regions with different reaching capabilities. We propose a method which is based on real-time upper-body pose and user models to plan robot motions for assistive dressing. We validate each part of our approach and test the whole system, allowing a Baxter humanoid robot to assist human to wear a sleeveless jacket. Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris |
IROS | 1 |