EDBT 2026 Demo / reviewers in the wild / expert
Dan Song 0006
dblp:33/4189-6
· DBLP profile ↗
74ranked-venue papers
24as first author
63since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 15 first-author · 41 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-task interactive attention for typhoon intensity and track forecasting
Yue Li 0042, Dan Song 0006, Anan Liu |
Expert Syst. Appl. | 4 |
| 2026 | Dual-path consistency constrained concept erasure for text-to-image diffusion models
Xiaoran Bai, Dan Song 0006, Shuangyan Yue |
Multim. Syst. | 2 |
| 2026 | Free-VTON: Cost-free acceleration and quality enhancement for diffusion-based virtual try-on
Dan Song 0006, Shuangyan Yue, Anan Liu |
Neural Networks | 1 |
| 2026 | DFS-Net: A Dense Focal Stack Image Generation Network From Misaligned Multi-Focus ImagesabstractDense focal stack images inherently encode depth cues and are crucial for various 3D vision applications. However, existing generation methods are susceptible to misalignment and introduce a domain gap between synthetic and real-world data due to off-axis aberrations. To address these challenges, we introduce DFS-Net, an aberration-aware dense focal stack image generation network. DFS-Net consists of two core modules: all-in-focus image synthesis and aberration-aware point spread function (PSF) generation. The all-in-focus image synthesis is achieved through a densely connected fusion network based on multi-scale focus migration and focus property detection. This fusion network can effectively fuse misaligned multi-focus images into an all-in-focus image. The aberration-aware PSF generation is realized through a multi-layer perceptron (MLP) network. Supervised by ray-tracing-based PSFs, the MLP network can generate spatially varying PSFs for arbitrary spatial positions and focus distances. By selecting a set of focus distances, the generated PSF maps are locally convolved with the all-in-focus image to produce an aberration-aware dense focal stack. We conduct extensive comparative experiments on all-in-focus image fusion and focal stack generation against state-of-the-art methods. The experimental results demonstrate that DFS-Net can synthesize all-in-focus images with high subjective and objective quality, as well as generate dense focal stacks that closely approximate ray-tracing results. In addition, we conduct comparative experiments on the depth-from-focus and salient object detection tasks using the generated focal stacks. The experimental results demonstrate that our DFS-Net can significantly enhance the performance of existing depth-from-focus and salient object detection models. The code and dataset will be publicly available at https://github.com/North-Li/DFS-Net. Zhilong Li, Pei An, You Yang 0002, Qiong Liu 0001, Dan Song 0006, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | RefreshReg: Receptive Field Reshaping and Multi-Layer Consistency Filtering for Point Cloud RegistrationabstractPoint cloud registration is a crucial task in the field of 3D processing research, which aims to align two or more point cloud scans into the same coordinate system. A significant factor limiting the performance of point cloud registration is the low proportion of inlier correspondences between two unaligned point clouds. It is particularly pronounced when the overlap between two scenes is low. Based on this observation, we propose a novel point cloud registration framework that enhances the proportion of correct correspondences via two aspects: extracting richer global geometric information for accurate identification of overlapping regions, and rejecting outliers based on spatial feature consistency. During the feature extraction phase, we first encode local geometry utilizing the Point Pair Features and then propose the Dual Graph Convolution module to reshape the receptive field, thereby expanding perception beyond small local areas. In the transformation estimation phase, we design a filtering module based on a multi-layer decoder. We extract point cloud features at different resolutions and select high-confidence point cloud pairs for registration based on the consistency of correspondences. We test the performance of our method on four datasets (3DMatch, ScanNet, KITTI, and MVP-RG). Compared with state-of-the-art approach NMCT, our method achieves improvements of 6% / 38% on KITTI / MVP-RG. Additionally, our filtering approach enhances the operational speed of RANSAC by more than 300%. Code is available at https://github.com/xiwanghuolight/RefreshReg. Dan Song 0006, Yue Zhang 0042, Weizhi Nie, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | MEF-GD: Multimodal Enhancement and Fusion Network for Garment DesignerabstractIn recent years, with advancements in generative models, an increasing number of garment design methods have been proposed. A generative model capable of generating garment images from text and sketches can provide designers with valuable visual references and creative inspiration to aid in the design process. Existing multimodal garment design methods face the challenge of lacking precise control over the generated results in relation to both sketches and text. In this paper, we propose Multimodal Enhancement and Fusion Network for Garment Design (MEF-GD). Our model inputs image conditions into Stable Diffusion based on ControlNet. On one hand, directly inputting image conditions can lead to feature forgetting, defined as the phenomenon in deep neural networks where previously learned feature representations are lost. To address this issue, we propose a multiple feature injection module to more effectively enhance image condition features. On the other hand, ControlNet fuses control features into Stable Diffusion through pointwise addition, which ignores the interaction between multimodal features and results in the fused features being biased towards the control features, overlooking Stable Diffusion features. To address this limitation, we introduce content-guided attention for more effective feature fusion and improve the expression of text features. Additionally, existing datasets often contain vague textual descriptions of garments. It is difficult to train the model on such a dataset to learn accurate alignment between generated image and the textual descriptions. To address this issue, we have designed a multimodal large model text optimization module to improve the quality and clarity of text generation. Compared to existing multimodal garment design methods, MEF-GD achieves more effective alignment with both textual and sketch-based inputs in generating garment images. Compared to MGD, MEF-GD achieves a decrease of 2.44 in FID and an increase of 0.83 in CLIP Score on Multi-VITON-HD dataset. The code will be available at https://github.com/fengyun691340/MEF-GD. Dan Song 0006, Jianhao Zeng, Hongshuo Tian, Bolun Zheng, Rongbao Kang, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | NDM: Boosting Dataset Distillation via Nested Difficulty Matching
Dongyang Zhang 0001, Hang Gou, Yue Zhang 0042, Dan Song 0006, Xiurui Xie |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prompt-Tuning-Guided Dual-Distribution Alignment for Unsupervised 2D-3D Cross-Modal Retrievalabstract2D-3D cross-modal retrieval (2D-3DCMR) aims at retrieving the most matching 3D models by leveraging 2D query images. However, the inherent modal discrepancy makes the 2D-3DCMR task still largely challenging. Besides, the scarcity of 3D labels in real-world applications severely hinders learning discriminative representations. To address these limitations, we present a prompt tuning guided dual distribution alignment (PTG-DDA) framework based on the CLIP model for the 2D-3DCMR task. Specifically, we design a learnable multi-view adaptive representation learning (MARL) module that adaptively integrates 3D features to merge complementary information and filter out redundant information across views, thereby improving the representation capability of 3D models. To mitigate the feature distribution shift between the 2D and 3D data, we design an attention-guided heterogeneous feature alignment (AHFA) module to guide the 2D and 3D inputs attend to feature banks by adopting the attention mechanism, thereby achieving heterogeneous feature alignment. Furthermore, to learn discriminative 3D features, we employ a multi-modal semantic prompt synergy (MSPS) module, which integrates class-related representations into learnable prompts to progressively learn the cross-modal synergy via a prompt synergy adapter, thereby achieving semantic feature alignment. Comprehensive experimental results on popular 2D-3DCMR benchmarks, i.e., MI3DOR and MI3DOR-2, demonstrate the superiority and effectiveness of PTG-DDA. Yaqian Zhou 0002, Ruiqiang Guo, Dan Song 0006, Jiayu Li 0004, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Separating Domain-Private Classes for Universal Unsupervised Cross-Domain 3D Model Retrieval
Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zan Gao 0002, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2025 | BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data TrainingabstractImage-based virtual try-on is an increasingly popular and important task to generate realistic try-on images of the specific person. Recent methods model virtual try-on as image mask-inpaint task, which requires masking the person image and results in significant loss of spatial information. Especially, for in-the-wild try-on scenarios with complex poses and occlusions, mask-based methods often introduce noticeable artifacts. Our research found that a mask-free approach can fully leverage spatial and lighting information from the original person image, enabling high-quality virtual try-on. Consequently, we propose a novel training paradigm for a mask-free try-on diffusion model. We ensure the model’s mask-free try-on capability by creating high-quality pseudo-data and further enhance its handling of complex spatial information through effective in-the-wild data augmentation. Besides, a try-on localization loss is designed to concentrate on try-on area while suppressing garment features in non-try-on areas, ensuring precise rendering of garments and preservation of fore/back-ground. In the end, we introduce BooW-VTON, the mask-free virtual try-on diffusion model, which delivers SOTA try-on quality without parsing cost. Extensive qualitative and quantitative experiments have demonstrated superior performance in wild scenarios with such a low-demand input. Xuanpu Zhang, Dan Song 0006, Pengxin Zhan, Jianhao Zeng, Weihua Luo, Anan Liu |
CVPR | 2 |
| 2025 | Image-Based Virtual Try-On: A Survey
Dan Song 0006, Xuanpu Zhang, Weizhi Nie, Ruofeng Tong 0001, Mohan Kankanhalli, Anan Liu |
Int. J. Comput. Vis. | 1 |
| 2025 | Adaptive CLIP for open-domain 3D model retrieval
Dan Song 0006, Zekai Qiang, Chumeng Zhang, Lanjun Wang, Qiong Liu 0001, You Yang 0002, Anan Liu |
Inf. Process. Manag. | 1 |
| 2025 | Generating counterfactual negative samples for image-text matching
Xinqi Su, Dan Song 0006, Wenhui Li 0001, Tongwei Ren, Anan Liu |
Inf. Process. Manag. | 2 |
| 2025 | S2Mix: Style and Semantic Mix for cross-domain 3D model retrieval
Xinwei Fu, Dan Song 0006 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Multi-view contrastive learning for unsupervised 3D model retrieval and classification
Wenhui Li 0001, Zhenghao Fang, Dan Song 0006, Weizhi Nie, Xuanya Li, Anan Liu |
Signal Process. Image Commun. | 3 |
| 2025 | Multi-Scale Spatial-Temporal Transformer for Meteorological Variable ForecastingabstractFrequent occurrences of marine extreme climate and weather events pose significant threats to human life and property, underscoring the practical significance of meteorological data forecasting methods. Notably, significant advancements in meteorological forecasting fields have been achieved by data-driven deep learning techniques, which leverage observed meteorological datasets and employ deep networks to capture complex patterns. However, challenges remain in accurately extracting local details and capturing spatial-temporal correlations when dealing with multiple meteorological forecasting tasks that exhibits diverse temporal and spatial scales. Hence, in this paper, we propose a Multi-Scale Spatial Temporal Transformer (MS-STT) framework to achieve efficient and accurate meteorological data forecasting. Specifically, to achieve more detailed and multi-scale representation of meteorological data, we design the regionally coherent encoding strategy and multi-scale feature aggregation for visual representation. To enhance the multi-scale ability in terms of learning spatial-temporal correlations, we propose a multi-scale spatial-temporal transformer network, which integrates a multi-scale spatial transformer to learn the spatial association between local patches and multi-scale regions and a temporal transformer to learn the temporal dynamic evolution properties. Extensive quantitative and qualitative experiments on three popular spatial temporal forecasting tasks validate the effectiveness of the proposed method. In particular, compared to the representative data-driven deep learning ENSO forecasting method Earthformer, our approach achieves a 3.7% performance improvement with only one-third of the parameters. Tianbao Li 0001, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Progressive Contrastive Label Optimization for Source-Free Universal 3D Model RetrievalabstractUnsupervised Cross-Domain 3D Model Retrieval (UCD3DMR) has emerged as an effective tool for managing 3D model data recently. However, existing UCD3DMR algorithms typically demand accessibility to source data and cross-domain label consistency, limiting their deployment in real-world industrial scenarios. Therefore, we relax the two demanding constraints and explore to address a newly challenging task, source-free universal 3D model retrieval (SFU3DMR). However, the inaccessibility to source data results in significant label noise in target pseudo-labels, while cross-domain label inconsistency introduces interference from target-private models, presenting tremendous challenges to model transfer. To address these challenges, we propose a novel SFU3DMR algorithm, Progressive Contrastive Label Optimization (PCLO). Specifically, we introduce the Neighbor-based Soft Label Optimization (NSLO) strategy, which refines target pseudo-labels based on the pseudo-label confidence of their nearest neighbors. Additionally, we design the Adaptive Hybrid Label Optimization (AHLO) strategy, which conducts positive label optimization to maximize label semantics for target-common models and executes negative label optimization to minimize label noise for target-private models. Experimental results confirm that the combined NSLO and AHLO strategies effectively refine the target pseudo-labels, and our PCLO achieves state-of-the-art performance for SFU3DMR on two well-established cross-domain benchmarks (MI3DOR and NTU/PSB). Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, You Yang 0002, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ClipMix for Domain GeneralizationabstractDomain Generalization (DG) is a growing field in machine learning that aims to train the model across multiple source domains, thereby enabling effective generalization to new, unseen target domains. Recent studies suggest that data augmentation, which enhances the diversity of the source domain, might be a promising solution to address this task. Current data augmentation methods use random fusion coefficients or local regional fusion, which cannot adaptively design the weights based on data, or preserve the integrity of original semantics. Inspired by the pre-trained model CLIP, which contains extensive multimodal knowledge, we propose ClipMix to address these limitations. Firstly, we use the CLIP model as the external knowledge to adaptively evaluate the alignment between images and their labels, using this alignment to assess the complexity of learning each image and guide adaptive augmentation. Secondly, we implement a label shift mechanism to dynamically assign soft labels to fused images, helping the model focus on hard-to-learn patterns and also gather domain-agnostic representation. Furthermore, we enhance the diversity of fused images at both the pixel and feature levels. Experimental results across sixteen domains from four databases verify the effectiveness of our method. Anan Liu, Hao-Chen Li, Wenhui Li 0001, Dan Song 0006, Hongshuo Tian, Lanjun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | MV-CLIP: Multi-View CLIP for Zero-Shot 3D Shape RecognitionabstractLarge-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models for 3D shapes, recent methods utilize language-image pre-training to realize zero-shot 3D shape recognition. However, due to the modality gap, pretrained language-image models are not confident enough in the generalization to 3D shape recognition. Consequently, this paper aims to improve the confidence with view selection and hierarchical prompts. Building on the well-established CLIP model, we introduce view selection in the vision side that minimizes entropy to identify the most informative views for 3D shape. On the textual side, hierarchical prompts combined of hand-crafted and GPT-generated prompts are proposed to refine predictions. The first layer prompts several classification candidates with traditional class-level descriptions, while the second layer refines the prediction based on function-level descriptions or further distinctions between the candidates. Extensive experiments demonstrate the effectiveness of the proposed modules for zero-shot 3D shape recognition. Remarkably, without the need for additional training, our proposed method achieves impressive zero-shot 3D classification accuracies of 84.44%, 91.51%, and 66.17% on ModelNet40, ModelNet10, and ShapeNet Core55, respectively. Furthermore, we will make the code publicly available to facilitate reproducibility and further research in this area. Dan Song 0006, Xinwei Fu, Weizhi Nie, Wenhui Li 0001, Lanjun Wang, You Yang 0002, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Better Fit: Accommodate Variations in Clothing Types for Virtual Try-OnabstractImage-based virtual try-on aims to transfer target in-shop clothing to a dressed model image, the objectives of which are totally taking off original clothing while preserving the contents outside of the try-on area, naturally wearing target clothing and correctly inpainting the gap between target clothing and original clothing. Tremendous efforts have been made to facilitate this popular research area, but cannot keep the type of target clothing with the try-on area affected by original clothing. In this paper, we focus on the unpaired virtual try-on situation where target clothing and original clothing on the model are different, i.e., the practical scenario. To break the correlation between the try-on area and the original clothing and make the model learn the correct information to inpaint, we propose an adaptive mask training paradigm that dynamically adjusts training masks. It not only improves the alignment and fit of clothing but also significantly enhances the fidelity of virtual try-on experience. Furthermore, we for the first time propose two metrics for unpaired try-on evaluation, the Semantic-Densepose-Ratio (SDR) and Skeleton-LPIPS (S-LPIPS), to evaluate the correctness of clothing type and the accuracy of clothing texture. For unpaired try-on validation, we construct a comprehensive cross-try-on benchmark (Cross-27) with distinctive clothing items and model physiques, covering a broad try-on scenarios. Experiments demonstrate the effectiveness of the proposed methods, contributing to the advancement of virtual try-on technology and offering new insights and tools for future research in the field. The code, model and benchmark will be publicly released. Dan Song 0006, Xuanpu Zhang, Jianhao Zeng, Pengxin Zhan, Weihua Luo, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Knowledge-Guided Prompt Learning for Tropical Cyclone Intensity EstimationabstractTropical cyclones (TCs) are one of the most destructive climatic phenomena, making accurate estimation of their intensity crucial for assessing disaster risks. However, due to the complex and indistinguishable structure of their satellite images, existing deep learning methods struggle to differentiate between images of varying intensity, which leads to unsatisfactory accuracy in intensity estimation. In this article, we propose a novel knowledge-guided prompt learning (KGPL) method, KGPL, for TC intensity estimation. KGPL utilizes a pretrained VLM to encode both historical intensity and current satellite image for estimating TC intensity. To minimize domain disparities and facilitate the transfer of a general large model to the TC domain, we introduce a cross-modal interactive prompt learning strategy. Specifically, we embed shared prompts in the text and vision encoders, aiming to learn domain knowledge and promote collaboration and interaction between these two modalities. Furthermore, we design a subregion contrastive learning strategy, which sets constraints on the intensity differences of convective activity in different subregions and guides the model to focus on learning strong convective areas such as the eye and eyewall. Extensive experiments show that our KGPL achieves a significant 41.1% reduction in root mean square error (RMSE) with only one-ninth of the trainable parameters compared to the state-of-the-art method, which validates the effectiveness of our method for TC intensity estimation. The code and data are available athttps://github.com/LiYue-TC/KGPL. Wenhui Li 0001, Yue Li 0042, Dan Song 0006, Jing Zhang 0038, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MMPFormer: A Memory-Aware Multiscale Predictive Transformer for Precipitation NowcastingabstractIn recent years, global climate change has intensified, leading to frequent occurrences of short-term extreme rain-fall events. How to accurately predict short-term precipitation within the next few hours has become a concern for people, especially for meteorologists. Traditional Numerical Weather Prediction (NWP) models rely on solving complex physical equations, resulting in significant computational costs and low temporal resolution, which limit their suitability for nowcasting tasks. Weather radar echo extrapolation, which forecasts future echoes from historical observations, has thus become the predominant approach for precipitation nowcasting. In the realm of deep learning-based methods, convolutional or recurrent neural network-based extrapolation pipelines inherently struggle in processing sequential data and capturing the multi-scale spatial features inherent in radar echo images. Additionally, current models often focus on enhancing dense prediction capabilities while neglecting to mine or utilize the evolutionary patterns of historical echoes, leading to diminished accuracy in long-term sequence forecasting. Moreover, high precipitation areas are critical for public safety, but current models often fail to forecast them accurately due to insufficient emphasis during learning. In this paper, we propose a memory-aware multi-scale predictive transformer (MMPFormer) for precipitation nowcasting. Specifically, we designed the Multi-Scale Spatial-Temporal Transformer (MS-STT), which incorporates multi-scale convolution techniques alongside self-attention mechanisms to effectively extract and model the spatiotemporal correlations within historical radar echoes. Additionally, we have designed a ConvGRU-Instructed Memory Mechanism (CIMM) to alleviate the degradation of radar echo details during extended extrapolation periods, enabling accurate forecasts over the next three hours. Furthermore, the Key Area Attention Module (KAAM), a plug-and-play module, has been introduced. It emphasizes high precipitation areas through attention mechanisms, mitigating the negative impacts of missing high precipitation by previous learning-based methods. Quantitative and qualitative experimental results on radar echo datasets demonstrate the superior performance of our method in modeling spatiotemporal dynamics and long-term extrapolation for precipitation forecasting. Dan Song 0006, Dehan Wang, Wenhui Li 0001, Lanjun Wang, Ryan Wen Liu, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Unified Multi-Modal Image Synthesis for Missing Modality ImputationabstractMulti-modal medical images provide complementary soft-tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scanning time, image corruption and various imaging protocols often result in incomplete multi-modal images, thus limiting the usage of multi-modal data for clinical purposes. To address this issue, in this paper, we propose a novel unified multi-modal image synthesis method for missing modality imputation. Our method overall takes a generative adversarial architecture, which aims to synthesize missing modalities from any combination of available ones with a single model. To this end, we specifically design a Commonality- and Discrepancy-Sensitive Encoder for the generator to exploit both modality-invariant and specific information contained in input modalities. The incorporation of both types of information facilitates the generation of images with consistent anatomy and realistic details of the desired distribution. Besides, we propose a Dynamic Feature Unification Module to integrate information from a varying number of available modalities, which enables the network to be robust to random missing modalities. The module performs both hard integration and soft integration, ensuring the effectiveness of feature combination while avoiding information loss. Verified on two public multi-modal magnetic resonance datasets, the proposed method is effective in handling various synthesis tasks and shows superior performance compared to previous methods. Yue Zhang 0042, Chengtao Peng, Qiuli Wang 0001, Dan Song 0006, Kaiyan Li 0005, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Source-Free Model Adaptation for Unsupervised 3D Object RetrievalabstractWith the explosive growth of 3D objects yet expensive annotation costs, unsupervised 3D object retrieval has become a popular but challenging research area. Existing labeled resources have been utilized to aid this task via transfer learning, which aligns the distribution of unlabeled data with the source one. However, the labeled resource are not always accessible due to the privacy disputes, limited computational capacity and other thorny restrictions. Therefore, we propose source-free model adaptation task for unsupervised 3D object management, which utilizes a pre-trained model to boost the performance with no access to source data and labels. Specifically, we compute representative prototypes to assume the source feature distribution, and design a bidirectional cumulative confidence-based adaptation strategy to adaptively align unlabeled samples towards prototypes. Subsequently, a dual-model distillation mechanism is proposed to generate source hypothesis for remedying the absence of ground-truth labels. The experiments on a cross-domain retrieval benchmark NTU-PSB (PSB-NTU) and a cross-modality retrieval benchmark MI3DOR also demonstrate the superiority of the proposed method even without access to raw data. Dan Song 0006, Yiyao Wu, Yuting Ling, Diqiong Jiang, Ruofeng Tong 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | CAT-DM: Controllable Accelerated Virtual Try-On with Diffusion ModelabstractGenerative Adversarial Networks (GANs) dominate the research field in image-based virtual try-on, but have not resolved problems such as unnatural deformation of garments and the blurry generation quality. While the generative quality of diffusion models is impressive, achieving controllability poses a significant challenge when applying it to virtual try-on and multiple denoising iterations limit its potential for real-time applications. In this paper, we propose Controllable Accelerated virtual Try-on with Diffusion Model (CAT-DM). To enhance the controllability, a basic diffusion-based virtual try-on network is designed, which utilizes ControlNet to introduce additional control conditions and improves the feature extraction of garment images. In terms of acceleration, CAT-DM initiates a reverse denoising process with an implicit distribution generated by a pre-trained GAN-based model. Compared with previous try-on methods based on diffusion models, CAT-DM not only retains the pattern and texture details of the in-shop garment but also reduces the sampling steps without compromising generation quality. Extensive experiments demonstrate the superiority of CAT-DM against both GAN-based and diffusion-based methods in producing more real-istic images and accurately reproducing garment patterns. Jianhao Zeng, Dan Song 0006, Weizhi Nie, Hongshuo Tian, Anan Liu |
CVPR | 2 |
| 2024 | FedDGP: Disentangling Global and Personal Models for Federated LearningabstractFederated learning (FL) aims to construct a global model by collaboratively training local models on data with diverse distributions, emphasizing the exchange of model parameters rather than raw data sharing. In the medical domain, achieving high performance in local models, strong generalization capabilities in the global model, and minimizing communication costs are all crucial. However, current federated learning methods struggle to concurrently optimize these three aspects. This paper introduces FedDGP, an innovative FL framework addressing these challenges. FedDGP comprises three key modules: Domain-guided Model Disentanglement (MD), Heterogeneous Aggregation (HA), and Reciprocal Iterative Training (RIT). MD divides a client model into two parts with different task objectives, enhancing both client and server model performance while avoiding issues like catastrophic forgetting. HA assigns reduced weight to underperforming models, limiting their impact. RIT minimizes communication costs by uploading only one client model’s parameters per round, facilitating local optimal model transfers. Extensive experiments on six domain classification datasets demonstrate FedDGP’s effectiveness, showcasing improved performance and reduced communication costs compared to existing approaches. Zhenhu Zhang, Dan Song 0006, Jiahua Dong 0001, Ruofeng Tong 0001 |
ICME | 3 |
| 2024 | Adaptive semantic transfer network for unsupervised 2D image-based 3D model retrieval
Dan Song 0006, Yuanxiang Yang, Wenhui Li 0001, Weizhi Nie, Xuanya Li, Anan Liu |
Comput. Vis. Image Underst. | 1 |
| 2024 | Structured serialization semantic transfer network for unsupervised cross-domain recognition and retrieval
Dan Song 0006, Yuanxiang Yang, Wenhui Li 0001, Xuanya Li, Min Liu 0008, Anan Liu |
Inf. Process. Manag. | 1 |
| 2024 | Hierarchical Transformer With Lightweight Attention for Radar-Based Precipitation NowcastingabstractThe U-net and Transformer have garnered significant attention in precipitation nowcasting due to their impressive capabilities in modeling sequential information. However, the performance is still constrained by the computational complexity of attention mechanism and the persistence of redundant information transmission between encoding and decoding stages. To address the above problems, we propose a novel hierarchical transformer with lightweight attention (HTLA) for precipitation nowcasting, which can integrate the Transformer and U-Net architectures to comprehensively explore the intrinsic characteristics of rainfall data with less complexity. Specifically, HTLA incorporates cross-channel self-attention with lightweight and dual feedforward module as fundamental components for encoding and decoding, efficiently fusing the advantages of Transformer and U-Net. A Gaussian pooling skip-connection strategy is proposed to adaptively weight information, effectively suppressing the redundant interference from the encoder to the decoder. The experimental results demonstrate the effectiveness and robustness of our HTLA, achieving improvements of 5.6% and 5.1% in terms of critical success index (CSI) and Heidke skill score (HSS) with only 3.6% parameters compared to the state-of-the-art method. The code is available athttps://github.com/precipitation-zy/HTLA. Wenhui Li 0001, Yue Li 0042, Dan Song 0006, Zhiqiang Wei 0002, Anan Liu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Text-Guided Knowledge Transfer for Remote Sensing Image-Text RetrievalabstractRemote sensing text-image retrieval aims to retrieve valuable information from diverse and complex remote sensing data, attracting significant attention. However, the performance is limited due to the complexity of scenes and their substantial content differences from natural domain images. To address these issues, we propose a simple but effective Text-guided Knowledge Transfer (TGKT) method for remote sensing image-text retrieval. TGKT utilizes CLIP to encode remote sensing data and transfer its rich semantic knowledge from natural to remote sensing domain. The textual information without significant domain differences is employed to bridge the semantic gap between these two domains, thereby enhancing image features. The extensive experimental results on both RSICD and RSITMD datasets demonstrate the effectiveness of our method. Anan Liu, Bo Yang 0055, Wenhui Li 0001, Dan Song 0006, Zhengya Sun, Tongwei Ren, Zhiqiang Wei 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Early prediction of sepsis using chatGPT-generated summaries and structured data
Qiang Li 0048, Hanbo Ma, Dan Song 0006, Yunpeng Bai, Keliang Xie |
Multim. Tools Appl. | 3 |
| 2024 | Fashion Customization: Image Generation Based on Editing ClueabstractFashion image generation attracts increasing attentions with wide applications in fashion design, virtual try-on, cosmetic industry, etc. Editing clues such as segmentation masks, keypoints and sketches are usually taken to guide the desired transformation of a reference image. However, spatial manipulation of the reference image remains a challenge, especially facing large-scale deformations and multiple editing requirements. In this paper, we propose a general model for multiple fashion editing tasks such as facial editing, pose transformation and clothes design based on user-defined editing instructions like semantic segmentation masks, keypoints, and sketches. With diverse editing requirements and various deformation scales, it is hard to learn the corresponding relationship between the editing clue and reference image with a uniform framework. Accordingly, we design a feature flow estimation network, which can adaptively adjust the feature flow according to the editing clue and the reference image, and generate a coarsely aligned image. Then we propose an image generative network to enrich the texture details of the transformed reference image. Experiments on three tasks verify the effectiveness of the proposed method and the adaptability to multiple tasks. The code and pretrained models will be available at https://github.com/zengjianhao/Fashion-Image-Generation-Based-on-Editing-Clue. Dan Song 0006, Jianhao Zeng, Min Liu 0008, Xuanya Li, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multimodal Adversarial Fusion for Typhoon Intensity ForecastingabstractTyphoon poses a significant threat to casualties and economic damage in coastal areas. Accurate forecasting of typhoon intensity is crucial for effective disaster management and mitigation strategies. Despite the certain progress has been made in integrating this task with deep learning techniques, it still faces limitations due to insufficient consideration of the intra-modal characteristics and the inter-modal discrepancy. In this paper, we propose a novel multi-modal adversarial fusion approach for typhoon intensity forecasting, which contains feature encoding and multi-modal fusion modules. Within the feature encoding module, we adopt a novel block attention mechanism to efficiently reveal the relationships of observational factors in multiple latent spaces, thereby obtaining more comprehensive representations feature of the observational data. Moreover, we introduce a complex convolutional operation to encode the components of wind speed within the reanalysis data, which can maintain their characteristics through the real and imaginary parts of complex numbers during the encoding procedure. Finally, we design the adversarial fusion scheme to effectively narrow the modality gap and facilitate a more comprehensive fusion for typhoon intensity forecasting. Comprehensive experiments are conducted on CMA-BST/ERA-Interim data and extensive results verify the effectiveness of our method for typhoon intensity forecasting. Wenhui Li 0001, Yue Li 0042, Yuanxiang Yang, Dan Song 0006, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multi-Task Spatial-Temporal Transformer for Multi-Variable Meteorological ForecastingabstractThis study delves into multi-variable meteorological spatial-temporal prediction, focusing on the simultaneous forecasting of key meteorological parameters such as temperature, wind speed, and atmospheric pressure. The core challenge of this task lies in identifying commonalities across different variables while capturing their unique features and the interactions among them. To address this, we propose a novel multi-task learning framework tailored for multi-variable meteorological forecasting. Our framework integrates a convolutional variable-specific visual representation module and a variable-interactive spatial-temporal inference module. The former extracts distinct variable information independently for each variable, while the latter employs a tri-level attention mechanism across space, time, and variables to uncover both commonalities and interactions among the variables. An adaptive multi-loss optimization strategy and a local information aggregation module are introduced to balance task optimization complexities and enhance representation stability. Comprehensive experiments across various meteorological prediction tasks confirm the effectiveness of our methods, showcasing superior performance over existing approaches. Tianbao Li 0001, Anan Liu, Dan Song 0006, Wenhui Li 0001, Jing Zhang 0038, Zhiqiang Wei 0002, Yuting Su 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Progressive Fourier Adversarial Domain Adaptation for Object Classification and RetrievalabstractDomain adaptation has been extensively explored as a means of transferring knowledge from the labeled source domain to the unlabeled target domain with disparate data distributions. However, the absence of target annotations and significant domain discrepancies pose a great challenge to transfer knowledge directly from source domain to target domain. To address this challenge, we propose a Progressive Fourier Adversarial Domain Adaptation (PFADA) framework, an effective and versatile framework which can generalize across multiple domain adaptation tasks. Firstly, we propose a Fourier-based style transfer strategy to generate a Fourier intermediate domain that incorporates source images with target domain-specific styles, while preserving the domain-invariant representations of the source data. Secondly, we introduce a progressive adversarial domain adaptation approach that utilizes the Fourier intermediate domain to facilitate the learning of domain-invariant representations. Finally, we present cross-domain semantic alignment and discriminative enhancement approach, which effectively guides the learning of discriminative cross-domain representations utilizing labeled source and intermediate domain data. Extensive experimental evaluations consistently validate the superior performance of the proposed method across diverse visual tasks, encompassing multiple domain adaptive image classification and retrieval scenarios. Tianbao Li 0001, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, Zhiqiang Wei 0002, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2024 | Cross-Modal Contrastive Learning with a Style-Mixed Bridge for Single Image 3D Shape RetrievalabstractImage-based 3D shape retrieval (IBSR) is a cross-modal matching task which searches similar shapes from a 3D repository using a natural image. Continuous attention has been paid to this topic, such as joint embedding, adversarial learning, and contrastive learning. Modality gap and diversity of instance similarities are two obstacles for accurate and fine-grained cross-modal matching. To overcome the two obstacles, we propose a style-mixed contrastive learning method (SC-IBSR). On one hand, we propose a style transition module to mix the styles of images and rendered shape views to form an intermediate style and inject it into image contents. The obtained style-mixed image features serve as a bridge for later contrastive learning in order to alleviate the modality gap. On the other hand, the proposed strategy of fine-grained consistency constraint aims at cross-domain contrast and considers the different importance of negative (positive) samples. Extensive experiments demonstrate the superiority of the style-mixed cross-modal contrastive learning on both the instance-level retrieval benchmark (i.e., Pix3D, Stanford Cars, and Comp Cars that annotate shapes to images), and the unsupervised category-level retrieval benchmark (i.e., MI3DOR-1 and MI3DOR-2 with unlabeled 3D shapes). Moreover, experiments are conducted on Office-31 dataset to validate the generalization capability of our method. Code and pretrained models will be available at https://github.com/honoria0204/SC-IBSR . Dan Song 0006, Shumeng Huo, Xinwei Fu, Chu-Meng Zhang, Wenhui Li 0001, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | StyleIPSB: Identity-Preserving Semantic Basis of StyleGAN for High Fidelity Face SwappingabstractRecent researches reveal that StyleGAN can generate highly realistic images, inspiring researchers to use pretrained StyleGAN to generate high-fidelity swapped faces. However, existing methods fail to meet the expectations in two essential aspects of high-fidelity face swapping. Their results are blurry without pore-level details and fail to preserve identity for challenging cases. To overcome the above artifacts, we innovatively construct a series of identity-preserving semantic bases of StyleGAN (called StyleIPSB) in respect of pose, expression, and illumination. Each basis of StyleIPSB controls one specific semantic attribute and disentangles with the others. The StyleIPSB constrains style code in the subspace of W+ space to preserve pore-level details and gives us a novel tool for high-fidelity face swapping, and we propose a three-stage framework for face swapping with StyleIPSB. Firstly, we transform the target facial images' attributes to the source image. We learn the mapping from 3D Morphable Model (3DMM) parameters, which capture the prominent semantic variance, to the coordinates of StyleIPSB that show higher identity-preserving and fidelity. Secondly, to transform detailed attributes which 3DMM does not capture, we learn the residual attribute between the reenacted face and the target face. Finally, the face is blended into the background of the target image. Extensive results and comparisons demonstrate that StyleIPSB can effectively preserve identity and pore-level details. The results of face swapping can achieve state-of-the-art performance. We will release our code at https://github.com/a686432/StyleIPSB Diqiong Jiang, Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001 |
CVPR | 2 |
| 2023 | Cross-domain Prototype Contrastive loss for Few-shot 2D Image-Based 3D Model Retrievalabstract2D image-based 3D model retrieval (IBMR) usually relies on abundant explicit supervision on 2D images, together with unlabeled 3D models to learn domain-aligned yet class-discriminative features for the retrieval task. However, collecting large-scale 2D labels is cost-effective and time-consuming. Therefore, we explore a challenging IBMR task, where only few-shot labeled 2D images are available while the rest of the 2D and 3D samples remain unlabeled. Limited annotation of 2D images further increases the difficulty of domain-aligned yet discriminative feature learning. Therefore, we propose cross-domain prototype contrastive loss (CPCL) for the few-shot IBMR task. Specifically, we capture semantic information to learn class-discriminative features in each domain by minimizing intra-domain prototype contrastive loss. Besides, we perform inter-domain transferable contrastive learning to align the features between instances and prototypes of the same class across domains. Comprehensive experiments on popular benchmarks, MI3DOR and MI3DOR-2, validate the superiority of CPCL. Yaqian Zhou 0002, Yu Liu 0004, Dan Song 0006, Jiayu Li 0004, Xuanya Li, Anan Liu |
ICME | 3 |
| 2023 | Chest X-ray Image Classification: A Causal Perspective
Weizhi Nie, Dan Song 0006, Yunpeng Bai, Keliang Xie, Anan Liu |
MICCAI (3) | 3 |
| 2023 | Towards Deconfounded Image-Text Matching with Causal InferenceabstractPrior image-text matching methods have shown remarkable performance on many benchmark datasets, but most of them overlook the bias in the dataset, which exists in intra-modal and inter-modal, and tend to learn the spurious correlations that extremely degrade the generalization ability of the model. Furthermore, these methods often incorporate biased external knowledge from large-scale datasets as prior knowledge into image-text matching model, which is inevitable to force model further learn biased associations. To address above limitations, this paper firstly utilizes Structural Causal Models (SCMs) to illustrate how intra- and inter-modal confounders damage the image-text matching. Then, we employ backdoor adjustment to propose an innovative Deconfounded Causal Inference Network (DCIN) for image-text matching task. DCIN (1) decomposes the intra- and inter-modal confounders and incorporates them into the encoding stage of visual and textual features, effectively eliminating the spurious correlations during image-text matching, and (2) uses causal inference to mitigate biases of external knowledge. Consequently, the model can learn causality instead of spurious correlations caused by dataset bias. Extensive experiments on two well-known benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superiority of our proposed method. Wenhui Li 0001, Xinqi Su, Dan Song 0006, Lanjun Wang, Kun Zhang 0040, Anan Liu |
ACM Multimedia | 3 |
| 2023 | Instrumental Variable Learning for Chest X-ray ClassificationabstractThe chest X-ray (CXR) is commonly employed to diagnose thoracic illnesses, but the challenge of achieving accurate automatic diagnosis through this method persists due to the complex relationship between pathology. In the past few years, numerous approaches based on deep learning have been proposed to address this issue but confounding factors such as image resolution or noise problems often damage model performance. In this paper, we focus on the chest X-ray classification task and proposed an interpretable instrumental variable (IV) learning framework, to eliminate the spurious association and obtain accurate causal representation. Specifically, we first construct a structural causal model (SCM) for our task and learn the confounders and the preliminary representations of IV, we then leverage electronic health record (EHR) as auxiliary information and we fuse the above feature with our transformer-based semantic fusion module, so the IV has the medical semantic. Meanwhile, the reliability of IV is further guaranteed via the constraints of mutual information between related causal variables. Finally, our approach's performance is demonstrated using the MIMIC-CXR, NIH ChestX-ray 14, and CheXpert datasets, and we achieve competitive results. Weizhi Nie, Dan Song 0006, Yunpeng Bai, Keliang Xie, Anan Liu |
SMC | 3 |
| 2023 | Domain-specific modeling and semantic alignment for image-based 3D model retrieval
Dan Song 0006, Xue-Jing Jiang, Yue Zhang 0042, Yun Zhang 0024 |
Comput. Graph. | 1 |
| 2023 | Hierarchical deep semantic alignment for cross-domain 3D model retrievalabstractWith the development of deep learning and the widespread application of 3D modeling technology, image-based cross-domain 3D model retrieval has attracted more and more researchers’ attention. Existing methods have achieved success by aligning the feature distributions from different domains. However, previous methods just statistically align the domain-level or class-level feature distributions, leaving sample discriminability a margin to be improved for retrieval. To address this issue, this paper proposes a Hierarchical Deep Semantic Alignment Network (HDSAN) for cross-domain 3D model retrieval, which combines the proposed sample-level semantic enhancement with global domain alignment and class semantic alignment. Concretely, we adopt adversarial domain adaptation at the domain level and dynamically align the class centers of two domains at the class level. To further improve sample discriminability, we design intra-domain and cross-domain triplet center alignment to enhance the semantic representation ability at the sample level. Experiments on two commonly-used cross-domain 3D model retrieval datasets MI3DOR-1 and MI3DOR-2 demonstrate the effectiveness of the proposed method. Dan Song 0006, Yuting Ling, Tianbao Li 0001, Xuanya Li |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Improving text-image cross-modal retrieval with contrastive loss
Chumeng Zhang, Junbo Guo, Guoqing Jin, Dan Song 0006, Anan Liu |
Multim. Syst. | 5 |
| 2023 | Focus on Hard Samples: Hierarchical Unbiased Constraints for Cross-Domain 3D Model RetrievalabstractCross-domain 3D model retrieval facilitates the management of explosively emerging unlabeled 3D models with conveniently available 2D images or RGB-D objects, which has attracted more and more attention. The modality gap between query samples (2D images or RGB-D objects) and 3D models makes the task challenging, and adversarial domain adaptation techniques have achieved success in narrowing such gaps. However, existing methods always pay excessive attention to the samples with high discriminability and transferability, whereas the hard samples with rich information are neglected. Accordingly, we propose hierarchical unbiased constraints to make full use of data at semantic level, sample level and feature level to improve the retrieval performance. At semantic level, we utilize maximum F-norm loss to constrain the semantic prediction results of target domain, which takes advantage of more hard samples to reduce ambiguous predictions and enhance discriminability. At sample level, we propose an adaptive triplet center loss to assign less confident samples with a farther negative class, which reliably compacts samples within the same class and expands the distance across different classes. At feature level, we perform SVD (singular value decomposition) for both source features and target features and suppress the relative value of the largest singular value, so that the information of other eigenvectors can be fully utilized to improve transferability. Experiments on two public datasets validate the superiority of the proposed method, and the ablation study analyzes different roles played by these hierarchical unbiased constraints. Tianbao Li 0001, Anan Liu, Dan Song 0006, Wenhui Li 0001, Xuanya Li, Yuting Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Self-supervised Image-based 3D Model RetrievalabstractImage-based 3D model retrieval aims at organizing unlabeled 3D models according to the relevance to the labeled 2D images. With easy accessibility of 2D images and wide applications of 3D models, image-based 3D model retrieval attracts more and more attentions. However, it is still a challenging problem due to the modality gap between 2D images and 3D models. In spite of the remarkable progress brought by domain adaptation techniques for this research topic, which usually propose to align the global distribution statistics of two domains, these methods are limited in learning discriminative features for target samples due to the lack of label information in target domain. In this article, besides utilizing the label information of 2D image domain and the adversarial domain alignment, we additionally incorporate self-supervision to address cross-domain 3D model retrieval problem. Specifically, we simultaneously optimize the adversarial adaptation for both domains based on visual features and the contrastive learning for unlabeled 3D model domain to help the feature extractor to learn discriminative feature representations. The contrastive learning is used to map view representations of the identical model nearby while view representations of different models far apart. To guarantee adequate and high-quality negative samples for contrastive learning, we design a memory bank to store and update representative view for each 3D model based on entropy minimization principle. Comprehensive experimental results on the public image-based 3D model retrieval datasets, i.e., MI3DOR and MI3DOR-2, have demonstrated the effectiveness of the proposed method. Dan Song 0006, Chu-Meng Zhang, Xiao-Qian Zhao, Weizhi Nie, Xuanya Li, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Cross-Domain 3D Model Retrieval Based On Contrastive Learning And Label PropagationabstractIn this work, we aim to tackle the task of unsupervised image based 3D model retrieval, where we seek to retrieve unlabeled 3D models that are most visually similar to the 2D query image. Due to the challenging modality gap between 2D images and 3D models, existing mainstream methods adopt domain-adversarial techniques to eliminate the gap, which cannot guarantee category-level alignment that is important for retrieval performance. Recent methods align the class centers of 2D images and 3D models to pay attention to the category-level alignment. However, there still exist two main issues: 1) the category-level alignment is too rough, and 2) the category prediction of unlabeled 3D models is not accurate. To overcome the first problem, we utilize contrastive learning for fine-grained category-level alignment across domains, which pulls both prototypes and samples with the same semantic information closer and pushes those with different semantic information apart. To provide reliable semantic prediction for contrastive learning and also address the second issue, we propose the consistent decision for pseudo labels of 3D models based on both the trained image classifier and label propagation. Experiments are carried out on MI3DOR and MI3DOR-2 datasets, and the results demonstrate the effectiveness of our proposed method. Dan Song 0006, Weizhi Nie, Xuanya Li, Anan Liu |
ACM Multimedia | 1 |
| 2022 | Improved Semantic Representation Learning by Multiple Clustering for Image-Based 3D Model RetrievalabstractUnder the heavy management on the increasing 3D models, the topic of image-based 3D model retrieval which organizes unlabeled 3D models based on abundant knowledge learned from labeled 2D images has drawn attention. However, prior methods are limited in aligning semantically at corresponding categories of two domains due to the lack of label information in the 3D domain. To this end, this paper proposes an improved semantic representation learning by multiple clustering approach, which improves the reliability of pseudo labels for 3D models, so as to achieve class-level semantic alignment. Specifically, this paper first extracts features for 2D images and 3D models. Then it clusters combining the 3D features with the semantic information from multiple clustering on 3D model features to obtain more reliable target pseudo labels. Extensive experiments have shown that the proposed method has achieved the gain of 3.0%-205.0% averagely for popular retrieval metrics on the benchmark of monocular image-based 3D object retrieval (MI3DOR), and 1.3%-69.7% on another advanced benchmark, MI3DOR-2. Jinghui Chu, Xiaoqian Zhao, Dan Song 0006, Wenhui Li 0001, Shenyuan Zhang, Xuanya Li, Anan Liu |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2022 | Gradual adaption with memory mechanism for image-based 3D model retrieval
Dan Song 0006, Yuting Ling, Tianbao Li 0001, Guoqing Jin, Junbo Guo, Xuanya Li |
Image Vis. Comput. | 1 |
| 2022 | Source-enhanced prototypical alignment for single image 3D model retrievalabstractAbstract Single image 3D model retrieval has attracted a lot of attentions with the convenience of organizing large‐scale unlabeled 3D models. Existing methods transfer the knowledge from well‐annotated 2D images (i.e., source domain) to unlabeled 3D models (i.e., target domain) to improve the discriminability of 3D models and align the feature distributions of 2D images and 3D models. However, during the alignment, the feature learning target of improving the discriminability of 3D models sometimes confuses the boundaries between 2D image categories, where prior methods ignore keeping the discriminability of 2D images. Motivated by this observation, we propose a source‐enhanced prototypical alignment framework to first remain the discriminability of 2D images and then guide the category‐level cross‐domain alignment with better image representations. Specifically, a novel separation and compactness loss is proposed for images to separate the samples from different categories and compact the samples within the same category. Then we perform prototypical alignment to make 2D image features assist in the discriminative feature learning for 3D models. We evaluate the proposed method on the commonly used cross‐domain 3D model retrieval benchmarks, namely MI3DOR and MI3DOR‐2, and the results demonstrate the effectiveness of the proposed method. Dan Song 0006, Chumeng Zhang, Xuanya Li, Ruofeng Tong 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2022 | Editorial paper for pattern recognition letters VSI on multi-view representation learning and multi-modal information representation
Dan Song 0006, Wenshu Zhang, Tongwei Ren, Xiaojun Chang |
Pattern Recognit. Lett. | 1 |
| 2022 | Monocular Image-Based 3-D Model Retrieval: A BenchmarkabstractMonocular image-based 3-D model retrieval aims to search for relevant 3-D models from a dataset given one RGB image captured in the real world, which can significantly benefit several applications, such as self-service checkout, online shopping, etc. To help advance this promising yet challenging research topic, we built a novel dataset and organized the first international contest for monocular image-based 3-D model retrieval. Moreover, we conduct a thorough analysis of the state-of-the-art methods. Existing methods can be classified into supervised and unsupervised methods. The supervised methods can be analyzed based on several important aspects, such as the strategies of domain adaptation, view fusion, loss function, and similarity measure. The unsupervised methods focus on solving this problem with unlabeled data and domain adaptation. Seven popular metrics are employed to evaluate the performance, and accordingly, we provide a thorough analysis and guidance for future work. To the best of our knowledge, this is the first benchmark for monocular image-based 3-D model retrieval, which aims to help related research in multiview feature learning, domain adaptation, and information retrieval. Dan Song 0006, Weizhi Nie, Wenhui Li 0001, Mohan Kankanhalli, Anan Liu |
IEEE Trans. Cybern. | 1 |
| 2022 | LR-GCN: Latent Relation-Aware Graph Convolutional Network for Conversational Emotion RecognitionabstractAs an intersection of artificial intelligence and human communication analysis, Emotion Recognition in Conversation (ERC) has attracted much research attention in recent years. Existing studies, however, are limited in adequately exploiting latent relations among the constituent utterances. In this paper, we address this issue by proposing a novel approach named Latent Relation-Aware Graph Convolutional Network (LR-GCN), where both speaker dependency of the interlocutors is leveraged and latent correlations among the utterances are captured for ERC. Specifically, we first establish a graph model to incorporate the context information and speaker dependency of the conversation. Afterward, the multi-head attention mechanism is introduced to explore the latent correlations among the utterances and generate a set of all-linked graphs. Here, aiming to simultaneously exploit the original modeled speaker dependency and the explored correlation information, we introduce a dense connection layer to capture more structural information of the generated graphs. Through a multi-branch graph network, we achieve a unified representation of each utterance for final prediction. Detailed evaluations on two benchmark datasets demonstrate LR-GCN outperforms the state-of-the-art approaches. Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Dan Song 0006, Weizhi Nie |
IEEE Trans. Multim. | 4 |
| 2021 | Toward Realistic Virtual Try-on Through Landmark Guided Shape MatchingabstractImage-based virtual try-on aims to synthesize the customer image with an in-shop clothes image to acquire seamless and natural try-on results, which have attracted increasing attentions. The main procedures of image-based virtual try-on usually consist of clothes image generation and try-on image synthesis, whereas prior arts cannot guarantee satisfying clothes results when facing large geometric changes and complex clothes patterns, which further deteriorates the afterwards try-on results. To address this issue, we propose a novel virtual try-on network based on landmark-guided shape matching (LM-VTON). Specifically, the clothes image generation progressively learns the warped clothes and refined clothes in an end-to-end manner, where we introduce a landmark-based constraint in Thin-Plate Spline (TPS) warping to inject finer deformation constraints around the clothes. The try-on process synthesizes the warped clothes with personal characteristics via a semantic indicator. Qualitative and quantitative experiments on two public datasets validate the superiority of the proposed method, especially for challenging cases such as large geometric changes and complex clothes patterns. Code will be available at https://github.com/lgqfhwy/LM-VTON. Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001 |
AAAI | 2 |
| 2021 | Multi-level similarity learning for image-text retrieval
Wenhui Li 0001, Yan Wang 0114, Dan Song 0006, Xuanya Li |
Inf. Process. Manag. | 4 |
| 2021 | Hierarchical multi-view context modelling for 3D object classification and retrieval
Anan Liu, Heyu Zhou, Weizhi Nie, Zhenguang Liu, Wu Liu 0005, Hongtao Xie 0001, Zhendong Mao 0001, Xuanya Li, Dan Song 0006 |
Inf. Sci. | 9 |
| 2021 | 3D shape recognition based on multi-modal information fusion
Qi Liang 0004, Mengmeng Xiao, Dan Song 0006 |
Multim. Tools Appl. | 3 |
| 2021 | FRSFN: A semantic fusion network for practical fashion retrieval
Anan Liu, Dan Song 0006, Wenhui Li 0001 |
Multim. Tools Appl. | 3 |
| 2021 | Multi-modal feature fusion based on multi-layers LSTM for video emotion recognition
Weizhi Nie, Dan Song 0006 |
Multim. Tools Appl. | 3 |
| 2021 | Coupled-dynamic learning for vision and language: Exploring Interaction between different tasks
Ning Xu 0003, Hongshuo Tian, Yanhui Wang 0001, Weizhi Nie, Dan Song 0006, Anan Liu, Wu Liu 0005 |
Pattern Recognit. | 5 |
| 2021 | DAN: Deep-Attention Network for 3D Shape RecognitionabstractDue to the wide applications in a rapidly increasing number of different fields, 3D shape recognition has become a hot topic in the computer vision field. Many approaches have been proposed in recent years. However, there remain huge challenges in two aspects: exploring the effective representation of 3D shapes and reducing the redundant complexity of 3D shapes. In this paper, we propose a novel deep-attention network (DAN) for 3D shape representation based on multiview information. More specifically, we introduce the attention mechanism to construct a deep multiattention network that has advantages in two aspects: 1) information selection, in which DAN utilizes the self-attention mechanism to update the feature vector of each view, effectively reducing the redundant information, and 2) information fusion, in which DAN applies attention mechanism that can save more effective information by considering the correlations among views. Meanwhile, deep network structure can fully consider the correlations to continuously fuse effective information. To validate the effectiveness of our proposed method, we conduct experiments on the public 3D shape datasets: ModelNet40, ModelNet10, and ShapeNetCore55. Experimental results and comparison with state-of-the-art methods demonstrate the superiority of our proposed method. Code is released on https://github.com/RiDang/DANN. Weizhi Nie, Yue Zhao 0042, Dan Song 0006, Yue Gao 0002 |
IEEE Trans. Image Process. | 3 |
| 2021 | Universal Cross-Domain 3D Model RetrievalabstractRecent advances in 3D modeling technologies such as 3D scanning, reconstruction and printing produce an explosive increasing of 3D models, consequently 3D model management becomes urgent to facilitate related applications such as CAD, VR/AR and autonomous driving. However, we usually lack the labels of the recently emerging 3D models and even have no prior knowledge toward the label set relationship between new datasets and existing labeled datasets, which makes the management challenging. In this paper, a universal cross-domain 3D model retrieval framework is proposed for utilizing the labeled 2D images or 3D models to manage unlabeled 3D models with no prior knowledge about label sets. Specifically, a sample-level weighting mechanism is adopted to automatically detect the samples from the common label set for both domains. Then, both the domain-level and class-level alignments are performed for domain adaptation. Finally, the adapted features are used for 3D model retrieval. We conduct experiments on the cross-domain 3D model retrieval dataset NTU-PSB (PSB-NTU) and image-based 3D model retrieval dataset MI3DOR, and the results validate the superiority and effectiveness of the proposed method. Dan Song 0006, Tianbao Li 0001, Wenhui Li 0001, Weizhi Nie, Wu Liu 0005, Anan Liu |
IEEE Trans. Multim. | 1 |
| 2021 | Joint Intermediate Domain Generation and Distribution Alignment for 2D Image-Based 3D Objects Retrievalabstract2D image-based 3D object retrieval provides a convenient way to manage 3D big data with easily accessed 2D images. It is also a challenging task due to the significant differences between 2D images and 3D objects. In this paper, we propose a 2D image-based 3D object retrieval method, which can reduce the distribution discrepancy between 2D images and 3D objects and learn invariant features between them. Specifically, we first construct an intermediate domain module based on maximum mean discrepancy (MMD) in an unsupervised way, which can reduce the 2D and 3D distribution discrepancy by marginal distribution constraint. Second, to further reduce conditional distribution discrepancy and learn invariant features, we use source domain labels as semantic information to dynamically guide distribution alignment. Moreover, in order to support the research in 3D object retrieval, we contribute a new dataset, MDI3D. We conducted extensive experiments on MDI3D and some popular datasets, such as MI3DOR and SHREC2013. The experimental results demonstrate the superiority of the proposed method by comparing with the state-of-the-art methods. Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Anan Liu, Jie Nie |
IEEE Trans. Multim. | 3 |
| 2020 | Consistent Domain Structure Learning and Domain Alignment for 2D Image-Based 3D Objects Retrievalabstract2D image-based 3D objects retrieval is a new topic for 3D objects retrieval which can be used to manage 3D data with 2D images. The goal is to search some related 3D objects when given a 2D image. The task is challenging due to the large domain gap between 2D images and 3D objects. Therefore, it is essential to consider domain adaptation problems to reduce domain discrepancy. However, most of the existing domain adaptation methods only utilize the semantic information from the source domain to predict labels in the target domain and neglect the intrinsic structure of the target domain. In this paper, we propose a domain alignment framework with consistent domain structure learning to reduce the large gap between 2D images and 3D objects. The domain structure learning module makes use of both the semantic information from the source domain and the intrinsic structure of the target domain, which provides more reliable predicted labels to the domain alignment module to better align the conditional distribution. We conducted experiments on two public datasets, MI3DOR and MI3DOR-2, and the experimental results demonstrate the proposed method outperforms the state-of-the-art methods. Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Weizhi Nie, Wenhui Li 0001, Anan Liu |
IJCAI | 3 |
| 2020 | Hierarchical Instance Feature Alignment for 2D Image-Based 3D Shape Retrievalabstract2D image-based 3D shape retrieval has become a hot research topic since its wide industrial applications and academic significance. However, existing view-based 3D shape retrieval methods are restricted by two settings, 1) learn the common-class features while neglecting the instance visual characteristics, 2) narrow the global domain variations while ignoring the local semantic variations in each category. To overcome these problems, we propose a novel hierarchical instance feature alignment (HIFA) method for this task. HIFA consists of two modules, cross-modal instance feature learning and hierarchical instance feature alignment. Specifically, we first use CNN to extract both 2D image and multi-view features. Then, we maximize the mutual information between the input data and the high-level feature to preserve as much as visual characteristics of an individual instance. To mix up the features in two domains, we enforce feature alignment considering both global domain and local semantic levels. By narrowing the global domain variations we impose the identical large norm restriction on both 2D and 3D feature-norm expectations to facilitate more transferable possibility. By narrowing the local variations we propose to minimize the distance between two centroids of the same class from different domains to obtain semantic consistency. Extensive experiments on two popular and novel datasets, MI3DOR and MI3DOR-2, validate the superiority of HIFA for 2D image-based 3D shape retrieval task. Heyu Zhou, Weizhi Nie, Wenhui Li 0001, Dan Song 0006, Anan Liu |
IJCAI | 4 |
| 2020 | Domain-Specific Alignment Network for Multi-Domain Image-Based 3D Object Retrievalabstract2D image-based 3D object retrieval is a very important task in computer vision and big data management. Conventional image-based 3D object retrieval usually assumes that the images are from one single domain. However, for real applications, 2D images may be from multiple domains (e.g., real image, sketch, and quick draw). It raises significant challenges for this task since these 2D images have a great domain gap with each other as well as a great modality gap with 3D objects. To address these issues, we propose an unsupervised Domain-Specific Alignment Network (DSAN) for multi-domain image-based 3D object retrieval. The proposed method aims to reduce domain discrepancy by domain-specific alignment network with multi-level moment matching, including first-order moment and second-order moment. Based on the observation that for any given sample, different domain classifiers should output the same label, we design a domain-specific classifier alignment module. To our knowledge, the proposed method is the first unsupervised work to align multiple-domain 2D images with 3D objects in an end-to-end manner. The multi-domain dataset MDI3D is utilized to advocate the research on this task, and the extensive experimental results demonstrate the superiority of the proposed method. Yuting Su 0001, Yuqian Li 0003, Dan Song 0006, Zhendong Mao 0001, Xuanya Li, Anan Liu |
ACM Multimedia | 3 |
| 2020 | Semantic Consistency Guided Instance Feature Alignment for 2D Image-Based 3D Shape Retrievalabstract2D image-based 3D shape retrieval (2D-to-3D) investigates the problem of matching the relevant 3D shapes from gallery dataset when given a query image. Recently, adversarial training and environmental style transfer learning have been successful applied to this task and achieved state-of-the-art performance. However, there still exist two problems. First, previous works only concentrate on the connection between the label and representation, where the unique visual characteristics of each instance are paid less attention. Second, the confused features or the transformed images can only cheat the discriminator but can not guarantee the semantic consistency. In another words, features of 2D desk may be mapped nearby the features of 3D chair. In this paper, we propose a novel semantic consistency guided instance feature alignment network (SC-IFA) to address these limitations. SC-IFA mainly consists of two parts, instance visual feature extraction and cross-domain instance feature adaptation. For the first module, unlike previous methods, which merely employ 2D CNN to extract the feature, we additionally maximize the mutual information between the input and feature to enhance the capability of feature representation for each instance. For the second module, we first introduce the margin disparity discrepancy model to mix up the cross-domain features in an adversarial training way. Then, we design two feature translators to transform the feature from one domain to another domain, and impose the translation loss and correlation loss on the transformed features to preserve the semantic consistency. Extensive experimental results on two benchmarks, MI3DOR and MI3DOR-2, verify SC-IFA is superior to the state-of-the-art methods. Heyu Zhou, Weizhi Nie, Dan Song 0006, Nian Hu, Xuanya Li, Anan Liu |
ACM Multimedia | 3 |
| 2020 | Joint deep feature learning and unsupervised visual domain adaptation for cross-domain 3D object retrieval
Wenhui Li 0001, Shu Xiang, Weizhi Nie, Dan Song 0006, Anan Liu, Xuanya Li |
Inf. Process. Manag. | 4 |
| 2020 | SP-VITON: shape-preserving image-based virtual try-on network
Dan Song 0006, Tianbao Li 0001, Zhendong Mao 0001, Anan Liu |
Multim. Tools Appl. | 1 |
| 2020 | Joint Heterogeneous Feature Learning and Distribution Alignment for 2D Image-Based 3D Object Retrievalabstract2D image-based 3D object retrieval is a novel but challenging task for 3D object retrieval. In this paper, we propose a 2D image-based 3D object retrieval method via joint heterogeneous feature learning and distribution alignment. Specifically, we propose to learn a mapping function in the Grassmann manifold to reduce the divergence of heterogeneous features of 2D images and 3D objects. We further employ the data distribution alignment method to adaptively integrate both marginal and conditional distributions. We embed both terms into the objective function to learn a domain-invariant classifier based on structural risk minimization. The output domain-invariant features of 2D images and 3D objects can be utilized for 2D image-based 3D object retrieval. Since there is lack of large-scale dataset for the evaluation of this task, we build two new datasets, MI3DOR and MI3DOR-2. We compare the proposed method against the representative methods for domain adaption and explore the influence of different components of the objective functions and key parameter. Comparison experiments show the superiority of this method. Yuting Su 0001, Yuqian Li 0003, Weizhi Nie, Dan Song 0006, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A shell space constrained approach for curve design on surface meshes
Dan Song 0006, Jin Huang 0001 |
Comput. Aided Des. | 2 |
| 2019 | Illumination-aware faster R-CNN for robust multispectral pedestrian detection
Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001 |
Pattern Recognit. | 2 |
| 2018 | Multispectral Pedestrian Detection via Simultaneous Detection and Segmentation
Dan Song 0006, Ruofeng Tong 0001, Min Tang 0001 |
BMVC | 2 |
| 2016 | 3D Body Shapes Estimation from Dressed-Human SilhouettesabstractAbstract Estimation of 3D body shapes from dressed‐human photos is an important but challenging problem in virtual fitting. We propose a novel automatic framework to efficiently estimate 3D body shapes under clothes. We construct a database of 3D naked and dressed body pairs, based on which we learn how to predict 3D positions of body landmarks (which further constrain a parametric human body model) automatically according to dressed‐human silhouettes. Critical vertices are selected on 3D registered human bodies as landmarks to represent body shapes, so as to avoid the time‐consuming vertices correspondences finding process for parametric body reconstruction. Our method can estimate 3D body shapes from dressed‐human silhouettes within 4 seconds, while the fastest method reported previously need 1 minute. In addition, our estimation error is within the size tolerance for clothing industry. We dress 6042 naked bodies with 3 sets of common clothes by physically based cloth simulation technique. To the best of our knowledge, We are the first to construct such a database containing 3D naked and dressed body pairs and our database may contribute to the areas of human body shapes estimation and cloth simulation. Dan Song 0006, Ruofeng Tong 0001, Jian Chang 0001, Xiaosong Yang, Min Tang 0001, Jian J. Zhang 0001 |
Comput. Graph. Forum | 1 |