VLDB 2026 Research / reviewers in the wild / expert
Haoye Dong
dblp:163/4089
· DBLP profile ↗
29ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0001-8169-8890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal deconfounding via multiplex spatial-temporal confounder disentanglement for next POI recommendation
Jie Li 0095, Zhengyang Wu 0001, Haoye Dong, Zetao Zheng, Mingrong Lin |
Inf. Process. Manag. | 3 |
| 2025 | Exploring a Tangible Interaction System for Behavior Management to Alleviate Children's Dental Anxiety in Waiting RoomsabstractOral health directly influences children's overall well-being, yet pervasive dental anxiety has become an unavoidable barrier to pediatric dental care.While behavior management proves more effective than environmental or equipment improvements in reducing pediatric dental anxiety while improving treatment understanding and oral health awareness, its time-intensive nature often conflicts with dentists' demanding workloads.Through formative research, we proposed a structured combination of behavior management techniques within dental waiting rooms, employing tangible interactions to create a comprehensive anxiety relief system for children aged 5-10 years.The system encompasses two core processes, dental caries treatment and caries prevention.We conducted the validation experiment and pilot study to refine the system.Then we deployed a user study involving 32 child-parent groups.The results demonstrated that the system effectively alleviates children's dental anxiety, prepares them for dental visits, promotes parent-child interaction, and supports children's participation. Weijia Lin, Xueyan Cai, Shichao Huang, Haoye Dong, Jiayu Yao, Jiayi Ma 0004, Shuyue Feng, Kecheng Jin |
IDC | 4 |
| 2025 | MV-SSM: Multi-View State Space Modeling for 3D Human Pose EstimationabstractWhile significant progress has been made in single-view 3D human pose estimation, multi-view 3D human pose estimation remains challenging, particularly in terms of generalizing to new camera configurations. Existing attention-based transformers often struggle to accurately model the spatial arrangement of keypoints, especially in occluded scenarios. Additionally, they tend to overfit specific camera arrangements and visual scenes from training data, resulting in substantial performance drops in new settings. In this study, we introduce a novel Multi-View State Space Modeling framework, named MV-SSM, for robustly estimating 3D human keypoints. We explicitly model the joint spatial sequence at two distinct levels: the feature level from multi-view images and the person keypoint level. We propose a Projective State Space (PSS) block to learn a generalized representation of joint spatial arrangements using state space modeling. Moreover, we modify Mamba’s traditional scanning into an effective Grid Token-guided Bidirectional Scanning (GTBS), which is integral to the PSS block. Multiple experiments demonstrate that MV-SSM achieves strong generalization, outperforming state-of-the-art methods: $ + {\mathbf{10}}.{\mathbf{8}}$ on AP25 $\left( { + 24\% {\text{ }} \uparrow } \right)$on the challenging three-camera setting in CMU Panoptic, $ + {\mathbf{7}}.{\mathbf{0}}$ on AP25 $\left( { + 13\% {\text{ }} \uparrow } \right)$on varying camera arrangements, and $ + {\mathbf{15}}.{\mathbf{3}}$ PCP $\left( { + 38\% {\text{ }} \uparrow } \right)$ on Campus A1 in cross-dataset evaluations. Project Website: https://aviralchharia.github.io/MV-SSM. Aviral Chharia, Wenbo Gou, Haoye Dong |
CVPR | 3 |
| 2025 | Learnable Infinite Taylor Gaussian for Dynamic View RenderingabstractCapturing the temporal evolution of Gaussian properties such as position, rotation, and scale is a challenging task due to the vast number of time-varying parameters and the limited photometric data available, which generally results in convergence issues, making it difficult to find an optimal solution. While feeding all inputs into an end-to- end neural network can effectively model complex temporal dynamics, this approach lacks explicit supervision and struggles to generate high-quality transformation fields. On the other hand, using time-conditioned polynomial functions to model Gaussian trajectories and orientations provides a more explicit and interpretable solution, but requires significant handcrafted effort and lacks generalizability across diverse scenes. To overcome these limitations, this paper introduces a novel approach based on a learnable infinite Taylor Formula to model the temporal evolution of Gaussians. This method offers both the flexibility of an implicit network-based approach and the interpretability of explicit polynomial functions, allowing for more robust and generalizable modeling of Gaussian dynamics across various dynamic scenes. Extensive experiments on dynamic novel view rendering tasks are conducted on public datasets, demonstrating that the proposed method achieves state-of-the-art performance in this domain. More information is available on our project page (https://ellisonking.github.io/TaylorGaussian). Haoye Dong, Junfeng Yao, Gim Hee Lee |
CVPR | 5 |
| 2025 | PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence Refinement
Haoye Dong, Gim Hee Lee |
ICCV | 1 |
| 2025 | CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian SplattingabstractRecent works in 3D multimodal learning have made remarkable progress. However, typically 3D multimodal models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point cloud cannot depict the texture information of 3D objects, resulting in inferior reconstruction capabilities. This limitation constrains the potential of point cloud-based 3D multimodal representation learning. In this paper, we present CLIP-GS, a novel multimodal representation learning framework grounded in 3DGS. We introduce the GS Tokenizer to generate serialized gaussian tokens, which are then processed through transformer layers pre-initialized with weights from point cloud models, resulting in the 3DGS embeddings. CLIP-GS leverages contrastive loss between 3DGS and the visual-text embeddings of CLIP, and we introduce an image voting loss to guide the directionality and convergence of gradient optimization. Furthermore, we develop an efficient way to generate triplets of 3DGS, images, and text, facilitating CLIP-GS in learning unified multimodal representations. Leveraging the well-aligned multimodal representations, CLIP-GS demonstrates versatility and outperforms point cloud-based models on various 3D tasks, including multimodal retrieval, zero-shot, and few-shot classification. Siyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie, Yinlong Qian, Yao Zhao 0001, Humphrey Shi, Yunchao Wei |
ICCV | 2 |
| 2024 | Generalizable Human Gaussians for Sparse View Synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang 0014, Francisco Vicente 0001, Albert Mosella-Montoro, Jianjin Xu, Shingo Takagi 0001, Daeil Kim, Aayush Prakash, Fernando De la Torre |
ECCV (78) | 4 |
| 2024 | DF-VTON: Dense Flow Guided Virtual Try-On NetworkabstractVirtual try-on system that transfers clothes onto the target person has attracted rapidly. Previous works use the affine or Thin Plate Spline (TPS) transformation for clothes warping and directly learn a composition mask to fuse the warped clothes and the person image, which usually causes rough shape and blurry details due to the poor warping mechanism and lack of human structure. In this paper, we propose a novel Dense Flow guided Virtual Try-On Network (DF-VTON), which contains a progressive warping network for clothes deformation, and a personalized fitting network for the fusion of the clothes and the person image. Specifically, given a target clothes image and a reference person image, the progressive warping network generates the dense flow in a progressive way and multi-scale views. The personalized fitting network aims to fuse warped clothes and person image seemly by using multi-scale composition masks. Extensive experiments on two challenging benchmarks demonstrate the superiority of our proposed DF-VTON over existing strong baselines with realistic high-resolution try-on results. Haoye Dong, Jun Liu 0075 |
ICASSP | 1 |
| 2024 | DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion Models
Zhenyu Xie, Haoye Dong, Zehua Ma, Xiaodan Liang |
ACM Multimedia | 2 |
| 2024 | Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mambaabstract3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, primarily due to inefficiently modeling spatial relations between joints. To address this problem, we propose a novel graph-guided Mamba framework, named Hamba, which bridges graph learning and state space modeling. Our core idea is to reformulate Mamba's scanning into graph-guided bidirectional scanning for 3D reconstruction using a few effective tokens. This enables us to efficiently learn the spatial relationships between joints for improving reconstruction performance. Specifically, we design a Graph-guided State Space (GSS) block that learns the graph-structured relations and spatial sequences of joints and uses 88.5\% fewer tokens than attention-based methods. Additionally, we integrate the state space features and the global features using a fusion module. By utilizing the GSS block and the fusion module, Hamba effectively leverages the graph-guided state space features and jointly considers global and local features to improve performance. Experiments on several benchmarks and in-the-wild tests demonstrate that Hamba significantly outperforms existing SOTAs, achieving the PA-MPVPE of 5.3mm and F@15mm of 0.992 on FreiHAND. At the time of this paper's acceptance, Hamba holds the top position, Rank 1, in two competition leaderboards on 3D hand reconstruction. Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente 0001, Fernando De la Torre |
NeurIPS | 1 |
| 2024 | E-Joint: Fabrication of Large-Scale Interactive Objects Assembled by 3D Printed Conductive Parts with Copper Plated JointsabstractThe advent of conductive thermoplastic filaments and multi-material 3D printing has made it feasible to create interactive 3D printed objects. Yet, challenges arise due to the volume constraints of desktop 3D printers and the high resistive characteristics of current conductive materials, making the fabrication of large-scale or highly conductive interactive objects can be daunting. We propose E-Joint, a novel fabrication pipeline for 3D printed objects utilizing mortise and tenon joint structures combined with a copper plating process. The segmented pieces and joint structures are customized in software along with integrated circuits. Then electroplate them for enhanced conductivity. We designed four distinct electrified joint structures in the experiment and evaluated the practical feasibility and effectiveness of fabricating pipes. By constructing three applications with those structures, we verified the usability of E-Joint in making large-scale interactive objects and showed the path to a more integrated future for manufacturing. Shuyue Feng, Haoye Dong, Shichao Huang, Xueyan Cai, Kecheng Jin, Fangtian Ying, Guanyun Wang |
UIST | 6 |
| 2024 | Physical-space Multi-body Mesh Detection Achieved by Local Alignment and Global Dense LearningabstractFrom monocular RGB images captured in the wild, detecting multi-body 3D meshes in physical sizes and locations is notoriously difficult due to the diverse visual ambiguity and lack of explicit depth measurement. Modern DNN approaches made numerous advances based on either two-stage Region-of-Interests(RoI)-Align or single-stage fixed Field-of-View (FoV) detector frameworks for two main subtasks: local pelvis-centered mesh regression and global body-to-camera translation regression. However, sub-meter-level physical-space monocular mesh detection is still out of reach by existing solutions. In this paper, we recognize two common drawbacks: (1) The local meshes are usually estimated without explicitly aligning body features under image-space scaling, occlusion, and truncation; (2) The global translations are estimated based on a weak-perspective assumption, which tricks the network into prioritizing image-space (front-view) mesh alignment and leads to inaccurate mesh depth. We introduce Physical-space Multi-body Mesh Detection (PMMD), in which (1) Locally, we preserve the body aspect ratio, align the body-to-RoI layout, and densely refine the person-wise RoI features for robustness; (2) Globally, we learn dense-depth-guided features to amend the body-wise local feature for physical depth estimation. With the cleaned local features and explicit local-global associations, PMMD achieves the best centimeter-level local mesh metrics and the first sub-meter-level global mesh metrics from monocular images in 3DPW and AGORA datasets. Haoye Dong, Tiange Xiang, Sravan Chittupalli, Jun Liu 0075 |
WACV | 1 |
| 2023 | GP-VTON: Towards General Purpose Virtual Try-On via Collaborative Local-Flow Global-Parsing LearningabstractImage-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for different garment parts, which fails to preserve the semantic information of different parts when receiving challenging inputs (e.g, intricate human poses, difficult garments). Moreover, most of them directly warp the input garment to align with the boundary of the preserved region, which usually requires texture squeezing to meet the boundary shape constraint and thus leads to texture distortion. The above inferior performance hinders existing methods from real-world applications. To address these problems and take a step towards real-world virtual try-on, we propose a General-Purpose Virtual Try-ON framework, named GP-VTON, by developing an innovative Local-Flow Global-Parsing (LFGP) warping module and a Dynamic Gradient Truncation (DGT) training strategy. Specifically, compared with the previous global warping mechanism, LFGP employs local flows to warp garments parts individually, and assembles the local warped results via the global garment parsing, resulting in reasonable warped parts and a semantic-correct intact garment even with challenging inputs. On the other hand, our DGT training strategy dynamically truncates the gradient in the overlap area and the warped garment is no more required to meet the boundary constraint, which effectively avoids the texture squeezing problem. Furthermore, our GP-VTON can be easily extended to multi-category scenario and jointly trained by using data from different garment categories. Extensive experiments on two high-resolution benchmarks demonstrate our superiority over the existing state-of-the-art methods.11Code is available at gp-vton. Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Xijin Zhang, Feida Zhu 0002, Xiaodan Liang |
CVPR | 5 |
| 2023 | Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from VideosabstractMulti-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person detection and tracking stages are performed in a multi-person setting, while temporal dynamics are only modeled for one person at a time. Consequently, their performance is severely limited by the lack of inter-person interactions in the spatial-temporal mesh recovery, as well as by detection and tracking defects. To address these challenges, we propose the Coordinate transFormer (Coord-Former) that directly models multi-person spatial-temporal relations and simultaneously performs multi-mesh recovery in an end-to-end manner Instead of partitioning the feature map into coarse-scale patch-wise tokens, CoordFormer leverages a novel Coordinate-Aware Attention to preserve pixel-level spatial-temporal coordinate information. Additionally, we propose a simple, yet effective Body Center Attention mechanism to fuse position information. Extensive experiments on the 3DPW dataset demonstrate that CoordFormer significantly improves the state-of-the-art, outperforming the previously best results by 4.2%, 8.8% and 4.7% according to the MPJPE, PAMPJPE, and PVE metrics, respectively, while being 40% faster than recent video-based approaches. The released code can be found at https://github.com/Li-Hao-yuan/CoordFormer. Haoye Dong, Hanchao Jia, Michael Kampffmeyer, Liang Lin 0004, Xiaodan Liang |
ICCV | 2 |
| 2023 | Human MotionFormer: Transferring Human Motions with Vision Transformers
Xintong Han, Chenbin Jin, Lihui Qian 0003, Huawei Wei, Haoye Dong, Yibing Song, Jia Xu 0011, Qifeng Chen 0001 |
ICLR | 8 |
| 2023 | XFormer: Fast and Accurate Monocular 3D Body CaptureabstractWe present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an image branch that makes prediction directly from the RGB image features. At the core of our method is a cross-modal transformer block that allows information flow across these two branches by modeling the attention between 2D keypoint coordinates and image spatial features. Our architecture is smartly designed, which enables us to train on various types of datasets including images with 2D/3D annotations, images with 3D pseudo labels, and motion capture datasets that do not have associated images. This effectively improves the accuracy and generalization ability of our system. Built on a lightweight backbone (MobileNetV3), our method runs blazing fast (over 30fps on a single CPU core) and still yields competitive accuracy. Furthermore, with a HRNet backbone, XFormer delivers state-of-the-art performance on Huamn3.6 and 3DPW datasets. Lihui Qian 0003, Xintong Han, Haoye Dong, Huawei Wei, Chengbin Jin |
IJCAI | 5 |
| 2021 | M3D-VTON: A Monocular-to-3D Virtual Try-On NetworkabstractVirtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garment templates, which hinders their applications in practical scenarios. 2D virtual try-on approaches provide a faster alternative to manipulate clothed humans, but lack the rich and realistic 3D representation. In this paper, we propose a novel Monocular-to-3D Virtual Try-On Network (M3D-VTON) that builds on the merits of both 2D and 3D approaches. By integrating 2D information efficiently and learning a mapping that lifts the 2D representation to 3D, we make the first attempt to reconstruct a 3D try-on mesh only taking the target clothing and a person image as inputs. The proposed M3D-VTON includes three modules: 1) The Monocular Prediction Module (MPM) that estimates an initial full-body depth map and accomplishes 2D clothes-person alignment through a novel two-stage warping procedure; 2) The Depth Refinement Module (DRM) that refines the initial body depth to produce more detailed pleat and face characteristics; 3) The Texture Fusion Module (TFM) that fuses the warped clothing with the non-target body part to refine the results. We also construct a high-quality synthesized Monocular-to-3D virtual try-on dataset, in which each person image is associated with a front and a back depth map. Extensive experiments demonstrate that the proposed M3D-VTON can manipulate and reconstruct the 3D human body wearing the given clothing with compelling details and is more efficient than other 3D approaches.1 Fuwei Zhao, Zhenyu Xie, Michael Kampffmeyer, Haoye Dong, Songfang Han, Tianxiang Zheng 0001, Tao Zhang 0042, Xiaodan Liang |
ICCV | 4 |
| 2021 | WAS-VTON: Warping Architecture Search for Virtual Try-on NetworkabstractDespite recent progress on image-based virtual try-on, current methods are constraint by shared warping networks and thus fail to synthesize natural try-on results when faced with clothing categories that require different warping operations. In this paper, we address this problem by finding clothing category-specific warping networks for the virtual try-on task via Neural Architecture Search (NAS). We introduce a NAS-Warping Module and elaborately design a bilevel hierarchical search space to identify the optimal network-level and operation-level flow estimation architecture. Given the network-level search space, containing different numbers of warping blocks, and the operation-level search space with different convolution operations, we jointly learn a combination of repeatable warping cells and convolution operations specifically for the clothing-person alignment. Moreover, a NAS-Fusion Module is proposed to synthesize more natural final try-on results, which is realized by leveraging particular skip connections to produce better-fused features that are required for seamlessly fusing the warped clothing and the unchanged person part. We adopt an efficient and stable one-shot searching strategy to search the above two modules. Extensive experiments demonstrate that our WAS-VTON significantly outperforms the previous fixed-architecture try-on methods with more natural warping results and virtual try-on results. Zhenyu Xie, Xujie Zhang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, Haonan Yan, Xiaodan Liang |
ACM Multimedia | 4 |
| 2021 | Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GANabstractImage-based virtual try-on is one of the most promising applications of human-centric image generation due to its tremendous real-world potential. Yet, as most try-on approaches fit in-shop garments onto a target person, they require the laborious and restrictive construction of a paired training dataset, severely limiting their scalability. While a few recent works attempt to transfer garments directly from one person to another, alleviating the need to collect paired datasets, their performance is impacted by the lack of paired (supervised) information. In particular, disentangling style and spatial information of the garment becomes a challenge, which existing methods either address by requiring auxiliary data or extensive online optimization procedures, thereby still inhibiting their scalability. To achieve a scalable virtual try-on system that can transfer arbitrary garments between a source and a target person in an unsupervised manner, we thus propose a texture-preserving end-to-end network, the PAtch-routed SpaTially-Adaptive GAN (PASTA-GAN), that facilitates real-world unpaired virtual try-on. Specifically, to disentangle the style and spatial information of each garment, PASTA-GAN consists of an innovative patch-routed disentanglement module for successfully retaining garment texture and shape characteristics. Guided by the source person's keypoints, the patch-routed disentanglement module first decouples garments into normalized patches, thus eliminating the inherent spatial information of the garment, and then reconstructs the normalized patches to the warped garment complying with the target person pose. Given the warped garment, PASTA-GAN further introduces novel spatially-adaptive residual blocks that guide the generator to synthesize more realistic garment details. Extensive comparisons with paired and unpaired approaches demonstrate the superiority of PASTA-GAN, highlighting its ability to generate high-quality try-on images when faced with a large variety of garments(e.g. vests, shirts, pants), taking a crucial step towards real-world scalable try-on. Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, Xiaodan Liang |
NeurIPS | 4 |
| 2021 | Image Comes Dancing With Collaborative Parsing-Flow Video SynthesisabstractTransferring human motion from a source to a target person poses great potential in computer vision and graphics applications. A crucial step is to manipulate sequential future motion while retaining the appearance characteristic. Previous work has either relied on crafted 3D human models or trained a separate model specifically for each target person, which is not scalable in practice. This work studies a more general setting, in which we aim to learn a single model to parsimoniously transfer motion from a source video to any target person given only one image of the person, named as Collaborative Parsing-Flow Network (CPF-Net). The paucity of information regarding the target person makes the task particularly challenging to faithfully preserve the appearance in varying designated poses. To address this issue, CPF-Net integrates the structured human parsing and appearance flow to guide the realistic foreground synthesis which is merged into the background by a spatio-temporal fusion module. In particular, CPF-Net decouples the problem into stages of human parsing sequence generation, foreground sequence generation and final video generation. The human parsing generation stage captures both the pose and the body structure of the target. The appearance flow is beneficial to keep details in synthesized frames. The integration of human parsing and appearance flow effectively guides the generation of video frames with realistic appearance. Finally, the dedicated designed fusion network ensure the temporal coherence. We further collect a large set of human dancing videos to push forward this research field. Both quantitative and qualitative results show our method substantially improves over previous approaches and is able to generate appealing and photo-realistic target videos given any input person image. All source code and dataset will be released at https://github.com/xiezhy6/CPF-Net. Zhenyu Xie, Xiaodan Liang, Yubei Xiao, Haoye Dong, Liang Lin 0004 |
IEEE Trans. Image Process. | 5 |
| 2020 | Fashion Editing With Adversarial Parsing LearningabstractInteractive fashion image manipulation, which enables users to edit images with sketches and color strokes, is an interesting research problem with great application value. Existing works often treat it as a general inpainting task and do not fully leverage the semantic structural information in fashion images. Moreover, they directly utilize conventional convolution and normalization layers to restore the incomplete image, which tends to wash away the sketch and color information. In this paper, we propose a novel Fashion Editing Generative Adversarial Network (FE-GAN), which is capable of manipulating fashion images by free-form sketches and sparse color strokes. FE-GAN consists of two modules: 1) a free-form parsing network that learns to control the human parsing generation by manipulating sketch and color; 2) a parsing-aware inpainting network that renders detailed textures with semantic guidance from the human parsing map. A new attention normalization layer is further applied at multiple scales in the decoder of the inpainting network to enhance the quality of the synthesized image. Extensive experiments on high-resolution fashion image datasets demonstrate that the proposed FE-GAN significantly outperforms the state-of-the-art methods on fashion image manipulation. Haoye Dong, Xiaodan Liang, Xujie Zhang, Xiaohui Shen, Zhenyu Xie, Jian Yin 0001 |
CVPR | 1 |
| 2019 | FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnabstractBeyond current image-based virtual try-on systems that have attracted increasing attention, we move a step forward to developing a video virtual try-on system that precisely transfers clothes onto the person and generates visually realistic videos conditioned on arbitrary poses. Besides the challenges in image-based virtual try-on (e.g., clothes fidelity, image synthesis), video virtual try-on further requires spatiotemporal consistency. Directly adopting existing image-based approaches often fails to generate coherent video with natural and realistic textures. In this work, we propose Flow-navigated Warping Generative Adversarial Network (FW-GAN), a novel framework that learns to synthesize the video of virtual try-on based on a person image, the desired clothes image, and a series of target poses. FW-GAN aims to synthesize the coherent and natural video while manipulating the pose and clothes. It consists of: (i) a flow-guided fusion module that warps the past frames to assist synthesis, which is also adopted in the discriminator to help enhance the coherence and quality of the synthesized video; (ii) a warping net that is designed to warp clothes image for the refinement of clothes textures; (iii) a parsing constraint loss that alleviates the problem caused by the misalignment of segmentation maps from images with different poses and various clothes. Experiments on our newly collected dataset show that FW-GAN can synthesize high-quality video of virtual try-on and significantly outperforms other methods both qualitatively and quantitatively. Haoye Dong, Xiaodan Liang, Xiaohui Shen, Bing-Cheng Chen, Jian Yin 0001 |
ICCV | 1 |
| 2019 | Towards Multi-Pose Guided Virtual Try-On NetworkabstractVirtual try-on systems under arbitrary human poses have significant application potential, yet also raise extensive challenges, such as self-occlusions, heavy misalignment among different poses, and complex clothes textures. Existing virtual try-on methods can only transfer clothes given a fixed human pose, and still show unsatisfactory performances, often failing to preserve person identity or texture details, and with limited pose diversity. This paper makes the first attempt towards a multi-pose guided virtual try-on system, which enables clothes to transfer onto a person with diverse poses. Given an input person image, a desired clothes image, and a desired pose, the proposed Multi-pose Guided Virtual Try-On Network (MG-VTON) generates a new person image after fitting the desired clothes into the person and manipulating the pose. MG-VTON is constructed with three stages: 1) a conditional human parsing network is proposed that matches both the desired pose and the desired clothes shape; 2) a deep Warping Generative Adversarial Network (Warp-GAN) that warps the desired clothes appearance into the synthesized human parsing map and alleviates the misalignment problem between the input human pose and the desired one; 3) a refinement render network recovers the texture details of clothes and removes artifacts, based on multi-pose composition masks. Extensive experiments on commonly-used datasets and our newly-collected largest virtual try-on benchmark demonstrate that our MG-VTON significantly outperforms all state-of-the-art methods both qualitatively and quantitatively, showing promising virtual try-on performances. Haoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang, Hanjiang Lai, Jia Zhu 0003, Zhiting Hu, Jian Yin 0001 |
ICCV | 1 |
| 2019 | Part-Preserving Pose Manipulation for Person Image SynthesisabstractManipulating person images under diverse poses, which transfers a person from one pose to another desired pose, is an interesting yet challenging task due to large non-rigid spatial deformation. Most existing works fail to preserve the fine-grained appearance consistency along with the pose changes due to the lack of explicit constraints and spatial modeling, leading to unrealistic results with severe artifacts. In this paper, we propose a novel Part-Preserving Generative Adversarial Network (PP-GAN) to achieve good manipulation quality by explicitly enforcing rich structure constraints over generative modeling. PP-GAN is proposed to decompose the challenging spatial transformation of the whole body into fine-grained part-level transformations, which are then integrated via human joint structure constraint. Given arbitrary poses, PP-GAN integrates human joint structure and region-level part cues as inputs to perform explicit generative modeling. Besides, we introduce a parsing-consistent loss to enforce semantic consistency among images with diverse poses, which guides the image synthesis from a semantic perspective. Extensive qualitative and quantitative evaluations on two benchmarks show that our PP-GAN significantly outperforms the state-of-the-art baselines in generating more realistic and plausible image synthesis results. PP-GAN successfully preserves part-level characteristics even for most challenging pose changes while prior works are easy to fail. Haoye Dong, Xiaodan Liang, Chenxing Zhou, Hanjiang Lai, Jia Zhu 0003, Jian Yin 0001 |
ICME | 1 |
| 2018 | Soft-Gated Warping-GAN for Pose-Guided Person Image SynthesisabstractDespite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most representative examples where the generation quality largely relies on the capability of identifying and modeling arbitrary transformations on different body parts. Current generative models are often built on local convolutions and overlook the key challenges (e.g. heavy occlusions, different views or dramatic appearance changes) when distinct geometric changes happen for each part, caused by arbitrary pose manipulations. This paper aims to resolve these challenges induced by geometric variability and spatial displacements via a new Soft-Gated Warping Generative Adversarial Network (Warping-GAN), which is composed of two stages: 1) it first synthesizes a target part segmentation map given a target pose, which depicts the region-level spatial layouts for guiding image synthesis with higher-level structure constraints; 2) the Warping-GAN equipped with a soft-gated warping-block learns feature-level mapping to render textures from the original image into the generated segmentation map. Warping-GAN is capable of controlling different transformation degrees given distinct target poses. Moreover, the proposed warping-block is light-weight and flexible enough to be injected into any networks. Human perceptual studies and quantitative evaluations demonstrate the superiority of our Warping-GAN that significantly outperforms all existing methods on two large datasets. Haoye Dong, Xiaodan Liang, Hanjiang Lai, Jia Zhu 0003, Jian Yin 0001 |
NeurIPS | 1 |
| 2018 | Deep Generative Models with Learnable Knowledge ConstraintsabstractThe broad set of deep generative models (DGMs) has achieved remarkable advances. However, it is often difficult to incorporate rich structured domain knowledge with the end-to-end DGMs. Posterior regularization (PR) offers a principled framework to impose structured constraints on probabilistic models, but has limited applicability to the diverse DGMs that can lack a Bayesian formulation or even explicit density evaluation. PR also requires constraints to be fully specified {\it a priori}, which is impractical or suboptimal for complex knowledge with learnable uncertain parts. In this paper, we establish mathematical correspondence between PR and reinforcement learning (RL), and, based on the connection, expand PR to learn constraints as the extrinsic reward in RL. The resulting algorithm is model-agnostic to apply to any DGMs, and is flexible to adapt arbitrary constraints with the model jointly. Experiments on human image generation and templated sentence generation show models with learned knowledge constraints by our algorithm greatly improve over base generative models. Zhiting Hu, Ruslan Salakhutdinov, Lianhui Qin, Xiaodan Liang, Haoye Dong, Eric P. Xing |
NeurIPS | 6 |
| 2016 | A novel feature selection strategy for friends recommendationabstractWith the social network being widely used, people would like to use the friends recommendation provided by a social websites. There are lots of methods to make the recommendation results more accurately and efficiently. By considering the feature selection strategy in the stage of data preprocessing, we propose a novel friend recommendation system using a classification model, which formulates the recommendation problem. We compare the performance of four classifiers, and draw a conclusion that our proposed method can get higher accuracy. Rui Ding 0007, Jia Zhu 0003, Yong Tang 0001, Xueqin Lin, Danyang Xiao, Haoye Dong |
CSCWD | 6 |
| 2015 | UBS: A Novel News Recommendation System Based on User Behavior SequenceabstractNews recommendation recently has attracted wide spread research attention because of the fast propagation of information on the Internet. Due to the large volume of information, a recommendation system which can provide the most important and useful information is required. Most of existing researches focus on providing recommendation based on news contents and predict the category of news only, which is inefficient if the news pool is very large or contains a lot of noisy data. In this study, we propose a novel news recommendation system called UBS, which recommends personalized news based on User Behavior Sequence (UBS) with high efficiency. We formulate the mining problem of user behavior sequence for Internet news reading, which can significantly enhance the performance of recommendation. Experimental validation was conducted using real datasets that obtained from news website. The results show that UBS can provide reasonable news recommendation compared to content-based recommendation as well as collaborative filtering. Haoye Dong, Jia Zhu 0003, Yong Tang 0001, Chuanhua Xu, Rui Ding 0007, Lingxiao Chen |
KSEM | 1 |
| 2015 | The Role of Physical Location in Our Online Social Networks
Jia Zhu 0003, Gabriel Pui Cheong Fung, Kam-Fai Wong, Binyang Li, Zhixu Li, Haoye Dong |
WAIM | 6 |