Sisi Dai

dblp:11/2852 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-8262-8138ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation
abstract
Despite significant progress in text-driven 4D human-object interaction (HOI) generation with supervised methods, the scalability remains limited by the scarcity of large-scale 4D HOI datasets. To overcome this, recent approaches attempt zero-shot 4D HOI generation with pre-trained image diffusion models. However, interaction cues are minimally distilled during the generation process, restricting their applicability across diverse scenarios. In this paper, we propose AnchorHOI, a novel framework that thoroughly exploits hybrid priors by incorporating video diffusion models beyond image diffusion models, advancing 4D HOI generation. Nevertheless, directly optimizing high-dimensional 4D HOI with such priors remains challenging, particularly for human pose and compositional motion. To address this challenge, AnchorHOI introduces an anchor-based prior distillation strategy, which constructs interaction-aware anchors and then leverages them to guide generation in a tractable two-step process. Specifically, two tailored anchors are designed for 4D HOI generation: anchor Neural Radiance Fields (NeRFs) for expressive interaction composition, and anchor keypoints for realistic motion synthesis. Extensive experiments demonstrate that AnchorHOI outperforms previous methods with superior diversity and generalization.
Sisi Dai, Kai Xu 0004
AAAI1
2025 A cause fusion framework with information bottleneck for conversational causal emotion entailment
Xinxin Su, Zhen Huang 0006, Menglong Lu, Sisi Dai, Yong Dou
Neural Networks4
2024 InterFusion: Text-Driven Generation of 3D Human-Object Interaction
Sisi Dai, Haowen Sun 0001, Chongyang Ma, Hui Huang 0004, Kai Xu 0004, Ruizhen Hu
ECCV (48)1
2024 Synchronized Dual-arm Rearrangement via Cooperative mTSP
abstract
Synchronized dual-arm rearrangement is widely studied as a common scenario in industrial applications. It often faces scalability challenges due to the computational complexity of robotic arm rearrangement and the high-dimensional nature of dual-arm planning. To address these challenges, we formulated the problem as cooperative mTSP, a variant of mTSP where agents share cooperative costs, and utilized reinforcement learning for its solution. Our approach involved representing rearrangement tasks using a task state graph that captured spatial relationships and a cooperative cost matrix that provided details about action costs. Taking these representations as observations, we designed an attention-based network to effectively combine them and provide rational task scheduling. Furthermore, a cost predictor is also introduced to directly evaluate actions during both training and planning, significantly expediting the planning process. Our experimental results demonstrate that our approach outperforms existing methods in terms of both performance and planning efficiency.
Shishun Zhang, Sisi Dai, Hui Huang 0004, Ruizhen Hu, Kai Xu 0004
ICRA3
2024 DINA: Deformable INteraction Analogy
abstract
We introduce deformable interaction analogy (DINA) as a means to generate close interactions between two 3D objects. Given a single demo interaction between an anchor object (e.g. a hand) and a source object (e.g. a mug grasped by the hand), our goal is to generate many analogous 3D interactions between the same anchor object and various new target objects (e.g. a toy airplane), where the anchor object is allowed to be rigid or deformable. To this end, we optimize the pose or shape of the anchor object to adapt it to a new target object to mimic the demo. To facilitate the optimization, we advocate using interaction interface (ITF), defined by a set of points sampled on the anchor object, as a descriptive and robust interaction representation that is amenable to non-rigid deformation. We model similarity between interactions using ITF, while for interaction analogy, we transform the ITF, either rigidly or non-rigidly, to guide the feature matching to the reposing and deformation of the anchor object. Qualitative and quantitative experiments show that our ITF-guided deformable interaction analogy works surprisingly well even with simple distance features compared to variants of state-of-the-art methods that utilize more sophisticated interaction representations and feature learning from large datasets.
Sisi Dai, Kai Xu 0004, Hao (Richard) Zhang, Hui Huang 0004, Ruizhen Hu
Graph. Model.2
2023 NIFT: Neural Interaction Field and Template for Object Manipulation
abstract
We introduce NIFT, Neural Interaction Field and Template, a descriptive and robust interaction representation of object manipulations to facilitate imitation learning. Given a few object manipulation demos, NIFT guides the generation of the interaction imitation for a new object instance by matching the Neural Interaction Template (NIT) extracted from the demos in the target Neural Interaction Field (NIF) defined for the new object. Specifically, the NIF is a neural field that encodes the relationship between each spatial point and a given object, where the relative position is defined by a spherical distance function rather than occupancies or signed distances, which are commonly adopted by conventional neural fields but less informative. For a given demo interaction, the corresponding NIT is defined by a set of spatial points sampled in the demo NIF with associated neural features. To better capture the interaction, the points are sampled on the Interaction Bisector Surface (IBS), which consists of points that are equidistant to the two interacting objects and has been used extensively for interaction representation. With both point selection and pointwise features defined for better interaction encoding, NIT effectively guides the feature matching in the NIFs of the new object instances such that the relative poses are optimized to realize the manipulation while imitating the demo interactions. Experiments show that our NIFT solution outperforms state-of-the-art imitation learning methods for object manipulation and generalizes better to objects from new categories.
Juzhan Xu, Sisi Dai, Kai Xu 0004, Hao (Richard) Zhang, Hui Huang 0004, Ruizhen Hu
ICRA3
2022 Fusion Multiple Kernel K-means
abstract
Multiple kernel clustering aims to seek an appropriate combination of base kernels to mine inherent non-linear information for optimal clustering. Late fusion algorithms generate base partitions independently and integrate them in the following clustering procedure, improving the overall efficiency. However, the separate base partition generation leads to inadequate negotiation with the clustering procedure and a great loss of beneficial information in corresponding kernel matrices, which negatively affects the clustering performance. To address this issue, we propose a novel algorithm, termed as Fusion Multiple Kernel k-means (FMKKM), which unifies base partition learning and late fusion clustering into one single objective function, and adopts early fusion technique to capture more sufficient information in kernel matrices. Specifically, the early fusion helps base partitions keep more beneficial kernel details, and the base partitions learning further guides the generation of consensus partition in the late fusion stage, while the late fusion provides positive feedback on two former procedures. The close collaboration of three procedures results in a promising performance improvement. Subsequently, an alternate optimization method with promising convergence is developed to solve the resultant optimization problem. Comprehensive experimental results demonstrate that our proposed algorithm achieves state-of-the-art performance on multiple public datasets, validating its effectiveness. The code of this work is publicly available at https://github.com/ethan-yizhang/Fusion-Multiple-Kernel-K-means.
Yi Zhang 0104, Xinwang Liu 0002, Jiyuan Liu 0003, Sisi Dai, Changwang Zhang, Kai Xu 0004, En Zhu
AAAI4
2022 Sample Weighted Multiple Kernel K-means via Min-Max optimization
abstract
A representative multiple kernel clustering (MKC) algorithm, termed simple multiple kernel k-means (SMKKM), is recently proposed to optimally mine useful information from a set of pre-specified kernels to improve clustering performance. Different from existing min-min learning framework, it puts a novel min-max optimization manner, which attracts considerable attention in related community. Despite achieving encouraged success, we observe that SMKKM only focuses on combination coefficients among kernels and ignores the relationship among the importance of different samples. As a result, it does not sufficiently consider different contributions of each sample to clustering, and thus cannot effectively obtain the "ideal" similarity structure, leading to unsatisfying performance. To address this issue, this paper proposes a novel sample weighted multiple kernel k-means via min-max optimization (SWMKKM), which sufficiently considers the sum of relationship between one sample and the others to represent the sample weights. Such a weighting criterion helps clustering algorithm pay more attention to samples with more positive effects on clustering and avoids unreliable overestimation for samples with poor quality. Based on SMKKM, we adopt a reduced gradient algorithm with proved convergence to solve the resultant optimization problem. Comprehensive experiments on multiple benchmark datasets demonstrate that our proposed SWMKKM dramatically improves the state-of-the-art MKC algorithms, verifying the effectiveness of our proposed sample weighting criterion.
Yi Zhang 0104, Weixuan Liang, Xinwang Liu 0002, Sisi Dai, Siwei Wang 0001, En Zhu
ACM Multimedia4
2021 One-Stage Incomplete Multi-view Clustering via Late Fusion
abstract
As a representative of multi-view clustering (MVC), late fusion MVC (LF-MVC) algorithm has attracted intensive attention due to its superior clustering accuracy and high computational efficiency. One common assumption adopted by existing LF-MVC algorithms is that all views of each sample are available. However, it is widely observed that there are incomplete views for partial samples in practice. In this paper, we propose One-Stage Late Fusion Incomplete Multi-view Clustering (OS-LF-IMVC) to address this issue. Specifically, we propose to unify the imputation of incomplete views and the clustering task into a single optimization procedure, so that the learning of the consensus partition matrix can directly assist the final clustering task. To optimize the resultant optimization problem, we develop a five-step alternate strategy with theoretically proved convergence. Comprehensive experiments on multiple benchmark datasets are conducted to demonstrate the efficiency and effectiveness of the proposed OS-LF-IMVC algorithm.
Yi Zhang 0104, Xinwang Liu 0002, Siwei Wang 0001, Jiyuan Liu 0003, Sisi Dai, En Zhu
ACM Multimedia5
2021 Gaussian Mixture Model Clustering with Incomplete Data
abstract
Gaussian mixture model (GMM) clustering has been extensively studied due to its effectiveness and efficiency. Though demonstrating promising performance in various applications, it cannot effectively address the absent features among data, which is not uncommon in practical applications. In this article, different from existing approaches that first impute the absence and then perform GMM clustering tasks on the imputed data, we propose to integrate the imputation and GMM clustering into a unified learning procedure. Specifically, the missing data is filled by the result of GMM clustering, and the imputed data is then taken for GMM clustering. These two steps alternatively negotiate with each other to achieve optimum. By this way, the imputed data can best serve for GMM clustering. A two-step alternative algorithm with proved convergence is carefully designed to solve the resultant optimization problem. Extensive experiments have been conducted on eight UCI benchmark datasets, and the results have validated the effectiveness of the proposed algorithm.
Yi Zhang 0104, Miaomiao Li 0001, Siwei Wang 0001, Sisi Dai, Lei Luo 0002, En Zhu, Xinzhong Zhu, Chaoyun Yao
ACM Trans. Multim. Comput. Commun. Appl.4
2008 A Grid Trust Model Based On MADM Theory
abstract
In this paper, we propose a trust model for the open Grid market. Our trust model emphasizes the importance of both direct trust and indirect trust/reputation when evaluating the trustworthiness of a Grid service provider. Since many factors contribute to the direct and indirect trust, we propose a novel method based on multiple attribute decision making (MADM) theory to determine the objective weights of both direct and indirect trust. The simulation results demonstrate that our trust model reflects the trustworthiness of a service provider more accurately than the weighted feedback model and eBay trust model, and thus improves user satisfaction in the open Grid environment.
Yiyu Yu, Junhua Tang, Liming Hao, Sisi Dai
GLOBECOM4