EDBT 2026 Demo / reviewers in the wild / expert
Yiming Wu 0005
dblp:91/4697-5
· DBLP profile ↗
16ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-9866-669XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On-Device Diffusion Transformer Policy for Efficient Robot ManipulationabstractDiffusion Policies have significantly advanced robotic manipulation tasks via imitation learning, but their application on resource-constrained mobile platforms remains challenging due to computational inefficiency and extensive memory footprint. In this paper, we propose LightDP, a novel framework specifically designed to accelerate Diffusion Policies for real-time deployment on mobile devices. LightDP addresses the computational bottleneck through two core strategies: network compression of the denoising modules and reduction of the required sampling steps. We first conduct an extensive computational analysis on existing Diffusion Policy architectures, identifying the denoising network as the primary contributor to latency. To overcome performance degradation typically associated with conventional pruning methods, we introduce a unified pruning and retraining pipeline, optimizing the model's post-pruning recoverability explicitly. Furthermore, we combine pruning techniques with consistency distillation to effectively reduce sampling steps while maintaining action prediction accuracy. Experimental evaluations on the standard datasets, \ie, PushT, Robomimic, CALVIN, and LIBERO, demonstrate that LightDP achieves real-time action prediction on mobile devices with competitive performance, marking an important step toward practical deployment of diffusion-based policies in resource-limited environments. Extensive real-world experiments also show the proposed LightDP can achieve performance comparable to state-of-the-art Diffusion Policies. Yiming Wu 0005, Huan Wang 0001, Jianxin Pang, Dong Xu 0001 |
ICCV | 1 |
| 2025 | Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion ModelsabstractThe high computational cost and slow inference time are major obstacles to deploying Video Diffusion Models (VDMs). To overcome this, we introduce a new Video Diffusion Model Compression approach using individual content and motion dynamics preserved pruning and consistency loss. First, we empirically observe that deeper VDM layers are crucial for maintaining the quality of motion dynamics (e.g., coherence of the entire video), while shallower layers are more focused on individual content (e.g., individual frames). Therefore, we prune redundant blocks from the shallower layers while preserving more of the deeper layers, resulting in a lightweight VDM variant called VDMini. Moreover, we propose an Individual Content and Motion Dynamics (ICMD) Consistency Loss to gain comparable generation performance as larger VDM to VDMini. In particular, we first use the Individual Content Distillation (ICD) Loss to preserve the consistency in the features of each generated frame between the teacher and student models. Next, we introduce a Multi-frame Content Adversarial (MCA) Loss to enhance the motion dynamics across the generated video as a whole. This method significantly accelerates inference time while maintaining high-quality video generation. Extensive experiments demonstrate the effectiveness of our VDMini on two important video generation tasks, Text-to-Video (T2V) and Image-to-Video (I2V), where we respectively achieve an average 2.5 ×, 1.4 ×, and 1.25 × speed up for the I2V method SF-V, the T2V method T2V-Turbo-v2, and the T2V method HunyuanVideo, while maintaining the quality of the generated videos on several benchmarks including UCF101, VBench-T2V, and VBench-I2V. Yiming Wu 0005, Huan Wang 0014, Dong Xu 0001 |
ACM Multimedia | 1 |
| 2025 | SOEDiff: Efficient Distillation for Small Object EditingabstractIn this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion . These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA , which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation , which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID. Yiming Wu 0005, Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Ronghua Liang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Panoptic Scene Graph Generation with Semantics-Prototype LearningabstractPanoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates lead to biased predicate annotations in the dataset, i.e. different predicates for the same object pairs. Biased predicate annotations make PSG models struggle in constructing a clear decision plane among predicates, which greatly hinders the real application of PSG models. To address the intrinsic bias above, we propose a novel framework named ADTrans to adaptively transfer biased predicate annotations to informative and unified ones. To promise consistency and accuracy during the transfer process, we propose to observe the invariance degree of representations in each predicate class, and learn unbiased prototypes of predicates with different intensities. Meanwhile, we continuously measure the distribution changes between each presentation and its prototype, and constantly screen potentially biased data. Finally, with the unbiased predicate-prototype representation embedding space, biased annotations are easily identified. Experiments show that ADTrans significantly improves the performance of benchmark models, achieving a new state-of-the-art performance, and shows great generalization and effectiveness on multiple datasets. Our code is released at https://github.com/lili0415/PSG-biased-annotation. Li Li 0091, Wei Ji 0008, Yiming Wu 0005, Mengze Li 0001, You Qin, Lina Wei, Roger Zimmermann |
AAAI | 3 |
| 2024 | Progressive Classifier and Feature Extractor Adaptation for Unsupervised Domain Adaptation on Point Clouds
Zicheng Wang 0012, Zhen Zhao 0001, Yiming Wu 0005, Luping Zhou, Dong Xu 0001 |
ECCV (28) | 3 |
| 2024 | Mrtnet: Multi-Resolution Temporal Network for Video Sentence GroundingabstractVideo sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolution Temporal (MRT) module, a Query-aware Attention (QAM) module, and a predictor. The MRT module uses an encoder-decoder network and Transformers to predict start and end times. The QAM module fuses visual and text features. Both MRT and QAM modules are easily integrated into existing VSG models. We also employ a loss function for cross-modal feature supervision at multiple scales. Extensive experiments on two prevalent datasets have shown the effectiveness of MRTNet. Wei Ji 0008, You Qin, Long Chen 0016, Yinwei Wei, Yiming Wu 0005, Roger Zimmermann |
ICASSP | 5 |
| 2024 | Self-Distilled Dynamic Fusion Network for Language-Based Fashion RetrievalabstractIn the domain of language-based fashion image retrieval, pinpointing the desired fashion item using both a reference image and its accompanying textual description is an intriguing challenge. Existing approaches lean heavily on static fusion techniques, intertwining image and text. Despite their commendable advancements, these approaches are still limited by a deficiency in flexibility. In response, we propose a Self-distilled Dynamic Fusion Network to compose the multi-granularity features dynamically by considering the consistency of routing path and modality-specific information simultaneously. Two new modules are included in our proposed method: (1) Dynamic Fusion Network with Modality Specific Routers. The dynamic network enables a flexible determination of the routing for each reference image and modification text, taking into account their distinct semantics and distributions. (2) Self Path Distillation Loss. A stable path decision for queries benefits the optimization of feature extraction as well as routing, and we approach this by progressively refine the path decision with previous path information. Extensive experiments demonstrate the effectiveness of our proposed model compared to existing methods. Yiming Wu 0005, Hangfei Li, Yilong Zhang 0001, Ronghua Liang |
ICASSP | 1 |
| 2024 | RE-IDVIS: Person Re-Identification System based on Interactive Visualizationabstractpixel-based visual encoding attribute-based visual encoding image-based visual encoding Figure 1: The interface of the system.(A) the probe panel which allows users to select the person-of-interest as a probe and set up the visual parameter of the search space.(B) the ranking list composed of the pixel-based visual encoding which allows users to quickly retrieve strong negative samples.(C) the search space view which supports visual exploration and enables users to provide feedback on the samples.(D) the spatiotemporal view which summarizes the spatiotemporal information of the retrieval results.(E) a cluster sample at three different visualization scales. Guodao Sun, Pan Liang, Sujia Zhu, Yiming Wu 0005, Haoran Liang 0001, Ronghua Liang |
ICMR | 6 |
| 2024 | Towards Small Object Editing: A Benchmark Dataset and A Training-Free ApproachabstractA plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance, enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide SOEBench (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from MSCOCO[22] and OpenImage[18]. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical.We will release our dataset on our project page: https://soebench.github.io/ Qihe Pan, Zhen Zhao 0001, Zicheng Wang 0012, Sifan Long 0001, Yiming Wu 0005, Wei Ji 0008, Haoran Liang 0001, Ronghua Liang |
ACM Multimedia | 5 |
| 2023 | D3T-GAN: Data-Dependent Domain Transfer GANs for Image Generation with Limited DataabstractAs an important and challenging problem, image generation with limited data aims at generating realistic images through training a GAN model given few samples. A typical solution is to transfer a well-trained GAN model from a data-rich source domain to the data-deficient target domain. In this paper, we propose a novel self-supervised transfer scheme termed D 3 T-GAN, addressing the cross-domain GANs transfer in limited image generation. Specifically, we design two individual strategies to transfer knowledge between generators and discriminators, respectively. To transfer knowledge between generators, we conduct a data-dependent transformation, which projects target samples into the latent space of source generator and reconstructs them back. Then, we perform knowledge transfer from transformed samples to generated samples. To transfer knowledge between discriminators, we design a multi-level discriminant knowledge distillation from the source discriminator to the target discriminator on both the real and fake samples. Extensive experiments show that our method improves the quality of generated images and achieves the state-of-the-art FID scores on commonly used datasets. Xintian Wu, Yiming Wu 0005, Xi Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | MGH: Metadata Guided Hypergraph Modeling for Unsupervised Person Re-identificationabstractAs a challenging task, unsupervised person ReID aims to match the same identity with query images which does not require any labeled information. In general, most existing approaches focus on the visual cues only, leaving potentially valuable auxiliary metadata information (e.g., spatio-temporal context) unexplored. In the real world, such metadata is normally available alongside captured images, and thus plays an important role in separating several hard ReID matches. With this motivation in mind, we propose MGH, a novel unsupervised person ReID approach that uses meta information to construct a hypergraph for feature learning and label refinement. In principle, the hypergraph is composed of camera-topology-aware hyperedges, which can model the heterogeneous data correlations across cameras. Taking advantage of label propagation on the hypergraph, the proposed approach is able to effectively refine the ReID results, such as correcting the wrong labels or smoothing the noisy labels. Given the refined results, we further present a memory-based listwise loss to directly optimize the average precision in an approximate manner. Extensive experiments on three benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art. Yiming Wu 0005, Xintian Wu, Xi Li 0001 |
ACM Multimedia | 1 |
| 2021 | F³A-GAN: Facial Flow for Face Animation With Generative Adversarial NetworksabstractFormulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with 1D or 2D representation (e.g., action units, emotion codes, landmark), which often leads to low-quality results in some complicated scenarios such as continuous generation and large-pose transformation. To tackle this problem, the conditions are supposed to meet two requirements, i.e., motion information preserving and geometric continuity. To this end, we propose a novel representation based on a 3D geometric flow, termed facial flow, to represent the natural motion of the human face at any pose. Compared with other previous conditions, the proposed facial flow well controls the continuous changes to the face. After that, in order to utilize the facial flow for face editing, we build a synthesis framework generating continuous images with conditional facial flows. To fully take advantage of the motion information of facial flows, a hierarchical conditional framework is designed to combine the extracted multi-scale appearance features from images and motion features from flows in a hierarchical manner. The framework then decodes multiple fused features back to images progressively. Experimental results demonstrate the effectiveness of our method compared to other state-of-the-art methods. Xintian Wu, Qihang Zhang, Yiming Wu 0005, Lingyun Sun, Xi Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | BANet: Bidirectional Aggregation Network With Occlusion Handling for Panoptic SegmentationabstractPanoptic segmentation aims to perform instance segmentation for foreground instances and semantic segmentation for background stuff simultaneously. The typical top-down pipeline concentrates on two key issues: 1) how to effectively model the intrinsic interaction between semantic segmentation and instance segmentation, and 2) how to properly handle occlusion for panoptic segmentation. Intuitively, the complementarity between semantic segmentation and instance segmentation can be leveraged to improve the performance. Besides, we notice that using detection/mask scores is insufficient for resolving the occlusion problem. Motivated by these observations, we propose a novel deep panoptic segmentation scheme based on a bidirectional learning pipeline. Moreover, we introduce a plug-and-play occlusion handling algorithm to deal with the occlusion between different object instances. The experimental results on COCO panoptic benchmark validate the effectiveness of our proposed method. Codes will be released soon at https://github.com/Mooonside/BANet. Guangchen Lin, Omar El Farouk Bourahla, Yiming Wu 0005, Junyi Feng, Mingliang Xu 0001, Xi Li 0001 |
CVPR | 5 |
| 2020 | Context-Aware Deep Spatiotemporal Network for Hand Pose Estimation From Depth ImagesabstractAs a fundamental and challenging problem in computer vision, hand pose estimation aims to estimate the hand joint locations from depth images. Typically, the problems are modeled as learning a mapping function from images to hand joint coordinates in a data-driven manner. In this paper, we propose a context-aware deep spatiotemporal network, a novel method to jointly model the spatiotemporal properties for hand pose estimation. Our proposed network is able to learn the representations of the spatial information and the temporal structure from the image sequences. Moreover, by adopting the adaptive fusion method, the model is capable of dynamically weighting different predictions to lay emphasis on sufficient context. Our method is examined on two common benchmarks, the experimental results demonstrate that our proposed approach achieves the best or the second-best performance with the state-of-the-art methods and runs in 60 fps. Yiming Wu 0005, Wei Ji 0008, Xi Li 0001, Gang Wang 0012, Jianwei Yin, Fei Wu 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Adaptive Graph Representation Learning for Video Person Re-IdentificationabstractRecent years have witnessed the remarkable progress of applying deep learning models in video person re-identification (Re-ID). A key factor for video person Re-ID is to effectively construct discriminative and robust video feature representations for many complicated situations. Part-based approaches employ spatial and temporal attention to extract representative local features. While correlations between parts are ignored in the previous methods, to leverage the relations of different parts, we propose an innovative adaptive graph representation learning scheme for video person Re-ID, which enables the contextual interactions between relevant regional features. Specifically, we exploit the pose alignment connection and the feature affinity connection to construct an adaptive structure-aware adjacency graph, which models the intrinsic relations between graph nodes. We perform feature propagation on the adjacency graph to refine regional features iteratively, and the neighbor nodes' information is taken into account for part feature representation. To learn compact and discriminative representations, we further propose a novel temporal resolution-aware regularization, which enforces the consistency among different temporal resolutions for the same identities. We conduct extensive evaluations on four benchmarks, i.e. iLIDS-VID, PRID2011, MARS, and DukeMTMC-VideoReID, experimental results achieve the competitive performance which demonstrates the effectiveness of our proposed method. Code is available at https://github.com/weleen/AGRL.pytorch. Yiming Wu 0005, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Multi-Task Structure-Aware Context Modeling for Robust Keypoint-Based Object TrackingabstractIn the fields of computer vision and graphics, keypoint-based object tracking is a fundamental and challenging problem, which is typically formulated in a spatio-temporal context modeling framework. However, many existing keypoint trackers are incapable of effectively modeling and balancing the following three aspects in a simultaneous manner: temporal model coherence across frames, spatial model consistency within frames, and discriminative feature construction. To address this problem, we propose a robust keypoint tracker based on spatio-temporal multi-task structured output optimization driven by discriminative metric learning. Consequently, temporal model coherence is characterized by multi-task structured keypoint model learning over several adjacent frames; spatial model consistency is modeled by solving a geometric verification based structured learning problem; discriminative feature construction is enabled by metric learning to ensure the intra-class compactness and inter-class separability. To achieve the goal of effective object tracking, we jointly optimize the above three modules in a spatio-temporal multi-task learning scheme. Furthermore, we incorporate this joint learning scheme into both single-object and multi-object tracking scenarios, resulting in robust tracking results. Experiments over several challenging datasets have justified the effectiveness of our single-object and multi-object trackers against the state-of-the-art. Xi Li 0001, Wei Ji 0008, Yiming Wu 0005, Fei Wu 0001, Ming-Hsuan Yang 0001, Dacheng Tao, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |