Ming Li 0073

dblp:181/2821-73 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-7852-0159ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
abstract
The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent framework that seamlessly unifies personalized apparel design, shopping recommendation, virtual try-on, and systematic evaluation into a cohesive workflow. To this end, StyleTailor pioneers an iterative visual refinement paradigm driven by multi-level negative feedback, enabling adaptive and precise user alignment. Specifically, our framework features two core agents, i.e., Designer for personalized garment selection and Consultant for virtual try-on, whose outputs are progressively refined via hierarchical vision-language model feedback spanning individual items, complete outfits, and try-on efficacy. Counterexamples are aggregated into negative prompts, forming a closed-loop mechanism that enhances recommendation quality. To assess the performance, we introduce a comprehensive evaluation suite encompassing style consistency, visual quality, face similarity, and artistic appraisal. Extensive experiments demonstrate StyleTailor's superior performance in delivering personalized designs and recommendations, outperforming strong baselines without negative feedback and establishing a new benchmark for intelligent fashion systems.
Hongbo Ma, Fei Shen 0004, Xiaoce Wang, Jinkai Zheng, Liangqiong Qu, Ming Li 0073
AAAI8
2026 Restoring neural radiance fields performance under adverse weather conditions
Ying He 0006, Gan Chen, F. Richard Yu, Ming Li 0073, Fei Ma 0006, Guang Zhou
Eng. Appl. Artif. Intell.4
2026 VSE-MOT: Multi-object tracking in low-quality video scenes guided by visual semantic enhancement
Jun Du 0001, Weiwei Xing, Ming Li 0073, F. Richard Yu
Pattern Recognit.3
2026 Lifelong scene graph generation
Tao He 0007, Tongtong Wu, Dongyang Zhang 0001, Ming Li 0073, Yuan-Fang Li, F. Richard Yu
Pattern Recognit.5
2026 Understanding gait recognition through silhouette sequence disentanglement and fine-grained visualization
Shaoxiong Zhang 0001, Yixiu Liu, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Chenggang Yan 0001
Pattern Recognit.5
2025 EventGPT: Event Stream Understanding with Multimodal Large Language Models
abstract
Event cameras capture visual information as asynchronous pixel change streams, excelling in challenging lighting and high-dynamic scenarios. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we introduce EventGPT, the first MLLM for event stream understanding, pioneering the integration of large language models (LLMs) with event-based vision. To bridge the huge domain gap, we propose a three-stage optimization paradigm to progressively equip a pre-trained LLM with event understanding. Our EventGPT consists of an event encoder, a spatio-temporal aggregator, a linear projector, an event-language adapter, and an LLM. Firstly, GPT-generated RGB image-text pairs warm up the linear projector, following LLaVA, as the gap between natural images and language is smaller. Secondly, we construct N-ImageNet-Chat, a large synthetic dataset of event data and corresponding texts to enable the use of the spatio-temporal aggregator and to train the event-language adapter, thereby aligning event features more closely with the language space. Finally, we gather an instruction dataset, EventChat, which contains extensive real-world data to fine-tune the entire model, further enhancing its generalization ability. We construct a comprehensive benchmark, and experiments show that EventGPT surpasses previous state-of-the-art MLLMs in generation quality, descriptive accuracy, and reasoning capability. Code: EventGPT
Shaoyu Liu, Jianing Li 0001, Guanghui Zhao 0003, F. Richard Yu, Xiangyang Ji, Ming Li 0073
CVPR8
2025 DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models
abstract
Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas.
Ying He 0006, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
ICASSP4
2025 FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
abstract
Vision-language-action (VLA) models have significantly advanced robotic manipulation by enabling robots to interpret language instructions for task execution. However, training these models often relies on large-scale user-specific data, raising concerns about privacy and security, which in turn limits their broader adoption. To address this, we propose FedVLA, the first federated VLA learning framework, enabling distributed model training that preserves data privacy without compromising performance. Our framework integrates task-aware representation learning, adaptive expert selection, and expert-driven federated aggregation, enabling efficient and privacy-preserving training of VLA models. Specifically, we introduce an Instruction Oriented Scene-Parsing mechanism, which decomposes and enhances object-level features based on task instructions, improving contextual understanding. To effectively learn diverse task patterns, we design a Dual Gating Mixture-of-Experts (DGMoE) mechanism, where not only input tokens but also self-aware experts adaptively decide their activation. Finally, we propose an Expert-Driven Aggregation strategy at the federated server, where model aggregation is guided by activated experts, ensuring effective cross-client knowledge transfer.Extensive simulations and real-world robotic experiments demonstrate the effectiveness of our proposals. Notably, DGMoE significantly improves computational efficiency compared to its vanilla counterpart, while FedVLA achieves task success rates comparable to centralized training, effectively preserving data privacy.
Cui Miao, Tao Chang, Meihan Wu, Ming Li 0073, Xiaodong Wang 0002
ICCV6
2025 EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients
Meihan Wu, Tao Chang, Cui Miao, Jie Zhou 0001, Xiangyu Xu 0002, Ming Li 0073, Xiaodong Wang 0002
ICCV7
2025 OmniStyle: Attention-Optimized Global and Local Image Stylization with Diffusion Model Inversion
abstract
Recent interest in large-scale text-driven diffusion models has highlighted their ability to generate diverse images from textual prompts, with style transfer being a significant application. However, existing text-guided stylization methods are constrained by the reliance on manually created masks for localized transformations, which limits scalability and automation. In this work, a novel framework, OmniStyle, is introduced for text-driven global and localized image style transfer, leveraging diffusion model inversion and attention-based mask generation. Input images are mapped into latent noise representations using DDIM inversion, while cross-attention maps are utilized to automatically generate precise semantic masks, eliminating the need for manual annotations. Latent features are dynamically optimized with mask-guided constraints and attention manipulations, enabling fine-grained style transfer confined to target regions while maintaining the integrity of global content. Extensive experiments demonstrate the scalability and effectiveness of OmniStyle.
Jiarong Cheng, Xihang Qiu, Ming Li 0073, F. Richard Yu
ICME4
2025 LV-VTON: Long-Video Virtual Try-On via Enhanced Visual Autoregressive Modeling
abstract
Video virtual try-on (VTON) aims to dress a target person in a desired garment while preserving the motion and identity in the original video. Generating long-duration VTON videos exacerbates challenges in achieving temporal coherence and visual fidelity. Existing image-based methods struggle with temporal consistency due to their frame-by-frame manner, while video-based approaches often sacrifice details for global coherence, resulting in blurry results. To address these limitations, we propose the first VTON framework tailored for long-video generation, namely LV-VTON, based on UNet Diffusion Transformer (UDiT) and Visual Autoregressive Modeling (VAR). Our LV-VTON framework includes three key components, i.e., Garment Appearance Module, Temporal Consistency Module, and Extended Video Synthesis Module. Notably, we collected a diverse and high-quality dataset, namely LongTry, to advance long video virtual try-on research. Extensive experiments demonstrate that LV-VTON excels at synthesizing various long VTON videos, outperforming state-of-the-art methods in detail preservation, temporal consistency, and long-video synthesis.
Lulu Tian, Hongxun Yao, Ming Li 0073
ICME3
2025 L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning
abstract
Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads to incomplete historical information due to missing labels, and class imbalance, which results in the model bias toward majority classes. To address these challenges, we propose Label-Augmented Analytic Adaptation (L3A), an exemplar-free approach without storing past samples. L3A integrates two key modules. The pseudo-label (PL) module implements label augmentation by generating pseudo-labels for current phase samples, addressing the label absence problem. The weighted analytic classifier (WAC) derives a closed-form solution for neural networks. It introduces sample-specific weights to adaptively balance the class contribution and mitigate class imbalance. Experiments on MS-COCO and PASCAL VOC datasets demonstrate that L3A outperforms existing methods in MLCIL tasks. Our code is available at https://github.com/scut-zx/L3A.
Run He, Chen Jiao, Di Fang 0004, Ming Li 0073, Ziqian Zeng, Cen Chen 0002, Huiping Zhuang
ICML5
2025 Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction
abstract
Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https://github.com/Inter3D-ui/Inter3D.
Gan Chen, Ying He 0006, Mulin Yu, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
IJCAI7
2025 TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
abstract
Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving superior cross-scene generalization. However, while many methods focus on geometric consistency, they often neglect the potential of text-driven guidance to enhance semantic understanding, which is crucial for accurately reconstructing fine-grained details in complex scenes. To address this limitation, we propose TextSplat-the first text-driven Generalizable Gaussian Splatting framework. Specifically, our framework employs three parallel modules to obtain complementary representations: the Diffusion Prior Depth Estimator for accurate depth information, the Semantic Aware Segmentation Network for detailed semantic information, and the Multi-View Interaction Network for refined cross-view features. Then, in the Text-Guided Semantic Fusion Module, these representations are integrated via the text-guided and attention-based feature aggregation mechanism, resulting in enhanced 3D Gaussian parameters enriched with detailed semantic cues. Experimental results on various benchmark datasets demonstrate improved performance compared to existing methods across multiple evaluation metrics, validating the effectiveness of our framework. The code will be publicly available.
Zhicong Wu, Ping Nie, Zhixin Yan, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Liqiang Nie
ACM Multimedia8
2025 Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
abstract
The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-Sora, the first framework to embed graphical watermarks directly into the video generation process. Motivated by the observation that watermarking performance is closely tied to the visual similarity between the watermark and cover content, we introduce a hierarchical coarse-to-fine adaptive matching mechanism. Specifically, the watermark image is divided into patches, each assigned to the most visually similar video frame, and further localized to the optimal spatial region for seamless embedding. To enable spatiotemporal fusion of watermark patches across video frames, we develop a 3D wavelet transform-enhanced Mamba architecture with a novel scanning strategy, effectively modeling long-range dependencies during watermark embedding and retrieval. To the best of our knowledge, this is the first attempt to apply state space models to watermarking, opening new avenues for efficient and robust watermark protection. Extensive experiments demonstrate that Safe-Sora achieves state-of-the- art performance in terms of video quality, watermark fidelity, and robustness, which is largely attributed to our proposals. Code and additional supporting materials are provided in the supplementary.
Zihan Su, Xuerui Qiu, Tangyu Jiang, Junhao Zhuang, Chun Yuan 0003, Ming Li 0073, Shengfeng He, F. Richard Yu
NeurIPS7
2025 Correction: Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002
Int. J. Comput. Vis.1
2025 ColonNeRF: High-fidelity neural reconstruction of long colonoscopy
Yufei Shi 0003, Beijia Lu, Jia-Wei Liu, Ming Li 0073, Si Yong Yeo, Zheng Shou 0001
Neurocomputing4
2025 Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures
abstract
Existing multi-view classification and clustering methods typically improve task accuracy by leveraging and fusing information from different views. However, ensuring the reliability of multi-view integration and final decisions is crucial, particularly when dealing with noisy or corrupted data. Current methods often rely on Kullback-Leibler (KL) divergence to estimate uncertainty of network predictions, ignoring domain gaps between different modalities. To address this issue, KPHD-Net, based on Hölder divergence, is proposed for multi-view classification and clustering tasks. Generally, our KPHD-Net employs a variational Dirichlet distribution to represent class probability distributions, models evidences from different views, and then integrates it with Dempster-Shafer evidence theory (DST) to improve uncertainty estimation effects. Our theoretical analysis demonstrates that Proper Hölder divergence offers a more effective measure of distribution discrepancies, ensuring enhanced performance in multi-view learning. Moreover, Dempster-Shafer evidence theory, recognized for its superior performance in multi-view fusion tasks, is introduced and combined with the Kalman filter to provide future state estimations. This integration further enhances the reliability of the final fusion results. Extensive experiments show that the proposed KPHD-Net outperforms the current state-of-the-art methods in both classification and clustering tasks regarding accuracy, robustness, and reliability, with theoretical guarantees.
Zhipeng Xue 0001, Yan Zhang 0119, Ming Li 0073, Yue Liu 0005, F. Richard Yu
IEEE Trans. Image Process.3
2025 Uncertainty Quantification via Hölder Divergence for Multi-View Representation Learning
abstract
Evidence-based deep learning represents a burgeoning paradigm for uncertainty estimation, offering reliable predictions with negligible extra computational overheads. Existing methods usually adopt Kullback-Leibler divergence to estimate the uncertainty of network predictions, ignoring domain gaps among various modalities. To tackle this issue, this paper introduces a novel algorithm based on Hölder Divergence (HD) to enhance the reliability of multi-view learning by addressing inherent uncertainty challenges from incomplete or noisy data. Generally, our method extracts the representations of multiple modalities through parallel network branches, and then employs HD to estimate the prediction uncertainties. Through the Dempster-Shafer theory, integration of uncertainty from different modalities, thereby generating a comprehensive result that considers all available representations. Mathematically, HD proves to better measure the “distance” between real data distribution and predictive distribution of the model and improve the performances of multi-class recognition tasks. Specifically, our method surpasses the existing state-of-the-art counterparts on all evaluating benchmarks. We further conduct extensive experiments on different backbones to verify our superior robustness. It is demonstrated that our method successfully pushes the corresponding performance boundaries. Finally, we perform experiments on more challenging scenarios,i.e., learning with incomplete or noisy data, revealing that our method exhibits a high tolerance to such corrupted data.
Yan Zhang 0119, Ming Li 0073, Zhaoxia Liu, Ye Zhang 0017, F. Richard Yu
IEEE Trans. Multim.2
2024 Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002
Int. J. Comput. Vis.1
2024 Semi-Supervised Disease Classification Based on Limited Medical Image Data
abstract
Inrecent years, significant progress has been made in the field of learning from positive and unlabeled examples (PU learning), particularly in the context of advancing image and text classification tasks. However, applying PU learning to semi-supervised disease classification remains a formidable challenge, primarily due to the limited availability of labeled medical images. In the realm of medical image-aided diagnosis algorithms, numerous theoretical and practical obstacles persist. The research on PU learning for medical image-assisted diagnosis holds substantial importance, as it aims to reduce the time spent by professional experts in classifying images. Unlike natural images, medical images are typically accompanied by a scarcity of annotated data, while an abundance of unlabeled cases exists. Addressing these challenges, this paper introduces a novel generative model inspired by Hölder divergence, specifically designed for semi-supervised disease classification using positive and unlabeled medical image data. In this paper, we present a comprehensive formulation of the problem and establish its theoretical feasibility through rigorous mathematical analysis. To evaluate the effectiveness of our proposed approach, we conduct extensive experiments on five benchmark datasets commonly used in PU medical learning: BreastMNIST, PneumoniaMNIST, BloodMNIST, OCTMNIST, and AMD. The experimental results clearly demonstrate the superiority of our method over existing approaches based on KL divergence. Notably, our approach achieves state-of-the-art performance on all five disease classification benchmarks. By addressing the limitations imposed by limited labeled data and harnessing the untapped potential of unlabeled medical images, our novel generative model presents a promising direction for enhancing semi-supervised disease classification in the field of medical image analysis.
Yan Zhang 0119, Zhaoxia Liu, Ming Li 0073
IEEE J. Biomed. Health Informatics4
2024 DR-FER: Discriminative and Robust Representation Learning for Facial Expression Recognition
abstract
Learning discriminative and robust representations is important for facial expression recognition (FER) due to subtly different emotional faces and their subjective annotations. Previous works usually address one representation solely because these two goals seem to be contradictory for optimization. Their performances inevitably suffer from challenges from the other representation. In this article, by considering this problem from two novel perspectives, we demonstrate that discriminative and robust representations can be learned in a unified approach, i.e., DR-FER, and mutually benefit each other. Moreover, we make it with the supervision from only original annotations. Specifically, to learn discriminative representations, we propose performing masked image modeling (MIM) as an auxiliary task to force our network to discover expression-related facial areas. This is the first attempt to employ MIM to explore discriminative patterns in a self-supervised manner. To extract robust representations, we present a category-aware self-paced learning schedule to mine high-quality annotated (easy) expressions and incorrectly annotated (hard) counterparts. We further introduce a retrieval similarity-based relabeling strategy to correct hard expression annotations, exploiting them more effectively. By enhancing the discrimination ability of the FER classifier as a bridge, these two learning goals significantly strengthen each other. Extensive experiments on several popular benchmarks demonstrate the superior performance of our DR-FER. Moreover, thorough visualizations and extra experiments on manually annotation-corrupted datasets show that our approach successfully accomplishes learning both discriminative and robust representations simultaneously.
Ming Li 0073, Huazhu Fu, Shengfeng He, Hehe Fan, Jun Liu 0036, Jussi Keppo, Zheng Shou 0001
IEEE Trans. Multim.1
2023 STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition
abstract
Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Second, they are vulnerable to practical attacking scenarios where attackers probe for privacy from an entire video rather than individual frames. To address these issues, we propose a novel framework STPrivacy to perform video-level PPAR. For the first time, we introduce vision Transformers into PPAR by treating a video as a tubelet sequence, and accordingly design two complementary mechanisms, i.e., sparsification and anonymization, to remove privacy from a spatio-temporal perspective. In specific, our privacy sparsification mechanism applies adaptive token selection to abandon action-irrelevant tubelets. Then, our anonymization mechanism implicitly manipulates the remaining action-tubelets to erase privacy in the embedding space through adversarial learning. These mechanisms provide significant advantages in terms of privacy preservation for human eyes and action-privacy trade-off adjustment during deployment. We additionally contribute the first two large-scale PPAR benchmarks, VP-HMDB51 and VP-UCF101, to the community. Extensive evaluations on them, as well as two other tasks, validate the effectiveness and generalization capability of our framework.
Ming Li 0073, Xiangyu Xu 0002, Hehe Fan, Pan Zhou 0002, Jun Liu 0036, Jia-Wei Liu, Jiahe Li 0009, Jussi Keppo, Zheng Shou 0001, Shuicheng Yan
ICCV1
2023 FakePoI: A Large-Scale Fake Person of Interest Video Detection Benchmark and a Strong Baseline
abstract
Deepfake technique can synthesize realistic images, audios, and videos, facilitating the thriving of entertainment, education, healthcare, and other industries. However, its abuse may pose potential threats to personal privacy, social stability, and even national security. Therefore, the development of deepfake detection methods is attracting more and more attention. Existing works mainly focus on the detection of common videos for entertainment purposes. In contrast, fake videos maliciously synthesized for Person of Interest (PoI, i.e., who is in an authoritative position and has broadly public influences) are much more harmful to society because of celebrity endorsement. However, there is no particular benchmark for driving related research in the community. Motivated by this observation, we present the first large-scale benchmark dataset, named FakePoI, to enable the research on fake PoI detection. It contains numerous fake videos of important people from all walks of life, e.g., police chiefs, city mayors, famous artists, and well-known Internet bloggers. In summary, our FakePoI includes 11092 synthesized videos where only a few clips rather than the entire are fake. Previous fake detection algorithms deteriorate heavily or even fail on our FakePoI due to two main challenges. On the one hand, the rich diversity of our fake videos makes it pretty difficult to find universally applicable patterns for detection. On the other hand, the high credibility contributed by the presence of real frames easily confuses a common detector. To tackle these challenges, we present an amplifier framework, highlighting the feature gap between real and generated video frames. Specifically, we present a quadruplet loss to narrow the distance of all real PoIs and meanwhile push away each real and fake PoI in embedding space. We implement our framework and conduct extensive experiments on the proposed benchmark. The quantitative results demonstrate that our approach outperforms existing methods significantly, setting a strong baseline on FakePoI. The qualitative analysis also shows its superiority. We will release our dataset and code athttps://github.com/cslltian/deepfake-detectionto encourage future research on this valuable area.
Lulu Tian, Hongxun Yao, Ming Li 0073
IEEE Trans. Circuits Syst. Video Technol.3
2023 Exploiting Multi-View Part-Wise Correlation via an Efficient Transformer for Vehicle Re-Identification
abstract
Image-based vehicle re-identification (ReID) has witnessed much progress in recent years. However, most of existing works struggled to extract robust but discriminative features from a single image to represent one vehicle instance. We argue that images taken from distinct viewpoints,e.g.,front and back, have significantly different appearances and patterns for recognition. In order to identify each vehicle, these models have to capture consistent “ID codes” from totally different views, causing learning difficulties. Additionally, we claim that part-level correspondences among views,i.e.,various vehicle parts observed from the identical image and the same part visible from different viewpoints, contribute to instance-level feature learning as well. Motivated by these, we propose to extract comprehensive vehicle instance representations from multiple views through modelling part-wise correlations. To this end, we present our efficient transformer-based framework to exploit both inner- and inter-view correlations for vehicle ReID. In specific, we first adopt a convnet encoder to condense a series of patch embeddings from each view. Then our efficient transformer, consisting of a distillation token and a noise token in addition to a regular classification token, is constructed for enforcing these patch embeddings to interact with each other regardless of whether they are taken from identical or different views. We conduct extensive experiments on widely used vehicle ReID benchmarks, and our approach achieves the state-of-the-art performance, showing the effectiveness of our method.
Ming Li 0073, Jun Liu 0036, Xinming Huang 0001
IEEE Trans. Multim.1
2021 Self-supervised Geometric Features Discovery via Interpretable Attention for Vehicle Re-Identification and Beyond
abstract
To learn distinguishable patterns, most of recent works in vehicle re-identification (ReID) struggled to redevelop official benchmarks to provide various supervisions, which requires prohibitive human labors. In this paper, we seek to achieve the similar goal but do not involve more human efforts. To this end, we introduce a novel framework, which successfully encodes both geometric local features and global representations to distinguish vehicle instances, optimized only by the supervision from official ID labels. Specifically, given our insight that objects in ReID share similar geometric characteristics, we propose to borrow self-supervised representation learning to facilitate geometric features discovery. To condense these features, we introduce an interpretable attention module, with the core of local maxima aggregation instead of fully automatic learning, whose mechanism is completely understandable and whose response map is physically reasonable. To the best of our knowledge, we are the first that perform self-supervised learning to discover geometric features. We conduct comprehensive experiments on three most popular datasets for vehicle ReID, i.e., VeRi-776, CityFlow-ReID, and VehicleID. We report our state-of-the-art (SOTA) performances and promising visualization results. We also show the excel-lent scalability of our approach on other ReID related tasks, i.e., person ReID and multi-target multi-camera (MTMC) vehicle tracking.
Ming Li 0073, Xinming Huang 0001
ICCV1
2020 TreeRNN: Topology-Preserving Deep Graph Embedding and Learning
abstract
General graphs are difficult for learning due to their irregular structures. Existing works employ message passing along graph edges to extract local patterns using customized graph kernels, but few of them are effective for the integration of such local patterns into global features. In contrast, in this paper we study the methods to transfer the graphs into trees so that explicit orders are learned to direct the feature integration from local to global. To this end, we apply the breadth first search (BFS) to construct trees from the graphs, which adds direction to the graph edges from the center node to the peripheral nodes. In addition, we proposed a novel projection scheme that transfer the trees to image representations, which is suitable for conventional convolution neural networks (CNNs) and recurrent neural networks (RNNs). To best learn the patterns from the graph-tree-images, we propose TreeRNN, a 2D RNN architecture that recurrently integrates the image pixels by rows and columns to help classify the graph categories. We evaluate the proposed method on several graph classification datasets, and manage to demonstrate comparable accuracy with the state-of-the-art on MUTAG, PTC-MR and NCI1 datasets.
Yecheng Lyu, Ming Li 0073, Xinming Huang 0001, Ulkuhan Guler 0001, Patrick Schaumont
ICPR2