Minlong Lu

dblp:136/5085 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-9851-6480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Tracing Copied Pixels and Regularizing Patch Affinity in Copy Detection
Siwei Nie, Minlong Lu
ICCV3
2024 Self-supervised Video Copy Localization with Regional Token Representation
Minlong Lu, Siwei Nie
ECCV (12)1
2023 TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision
abstract
Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, and then detect and refine the boundaries of copied segments on similarity matrix under temporal constraints. In this paper, we propose TransVCL: an attention-enhanced video copy localization network, which is optimized directly from initial frame-level features and trained end-to-end with three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for similarity matrix generation, and a temporal alignment module for copied segments localization. In contrast to previous methods demanding the handcrafted similarity matrix, TransVCL incorporates long-range temporal information between feature sequence pair using self- and cross- attention layers. With the joint design and optimization of three components, the similarity matrix can be learned to present more discriminative copied patterns, leading to significant improvements over previous methods on segment-level labeled datasets (VCSL and VCDB). Besides the state-of-the-art performance in fully supervised setting, the attention architecture facilitates TransVCL to further exploit unlabeled or simply video-level labeled data. Additional experiments of supplementing video-level labeled datasets including SVD and FIVR reveal the high flexibility of TransVCL from full supervision to semi-supervision (with or without video-level annotation). Code is publicly available at https://github.com/transvcl/TransVCL.
Sifeng He, Minlong Lu, Chen Jiang 0006, Feng Qian 0006, Lei Yang 0061
AAAI3
2023 Entity-Graph Enhanced Cross-Modal Pretraining for Instance-Level Product Retrieval
abstract
Our goal in this research is to study a more realistic environment in which we can conduct weakly-supervised multi-modal instance-level product retrieval for fine-grained product categories. We first contribute the Product1M datasets and define two real practical instance-level retrieval tasks that enable evaluations on price comparison and personalized recommendations. For both instance-level tasks, accurately identifying the intended product target mentioned in visual-linguistic data and mitigating the impact of irrelevant content are quite challenging. To address this, we devise a more effective cross-modal pretraining model capable of adaptively incorporating key concept information from multi-modal data. This is accomplished by utilizing an entity graph, where nodes represented entities and edges denoted the similarity relations between them. Specifically, a novel Entity-Graph Enhanced Cross-Modal Pretraining (EGE-CMP) model is proposed for instance-level commodity retrieval, which explicitly injects entity knowledge in both node-based and subgraph-based ways into the multi-modal networks via a self-supervised hybrid-stream transformer. This could reduce the confusion between different object contents, thereby effectively guiding the network to focus on entities with real semantics. Experimental results sufficiently verify the efficacy and generalizability of our EGE-CMP, outperforming several SOTA cross-modal baselines like CLIP Radford et al. 2021, UNITER Chen et al. 2020 and CAPTURE Zhan et al. 2021.
Xunlin Zhan, Yunchao Wei, Xiaoyong Wei, Yaowei Wang 0001, Minlong Lu, Xiaochun Cao, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Caption-Aided Product Detection via Collaborative Pseudo-Label Harmonization
abstract
Product detection, which aims to localize products of interest in the advertising images, helps advance many potential E-commerce applications like product retrieval and recommendation. However, labeling a massive number of fine-grained product categories and accurate product boxes is costly and especially not practical since products are ever-changing on E-commerce websites. In this work, we step forward to train a fine-grained product detector solely supervised by the advertising captions, which are naturally available but often severely flawed and noisy. To reformulate the weakly supervised detection research into a real-world setting, we introduce a large-scale benchmark, namedCapProduct, where more than 80,000 product image-caption pairs are collected from E-commerce websites. The fine-grained nature of products and noisy captions inCapProductmake it intractable to excavate valid category labels to train a weakly supervised object detector. To tackle this challenge, we propose aCollaborativePseudo-LabelHarmonization (CoPLH) framework that harmonizes self-mined pseudo labels via modeling the global co-occurrence relationships of products. We construct a collaborative co-occurrence graph based on all training samples to improve the reliability of caption-predicted pseudo-labels as well as benefit the self-training procedure in a weakly supervised setting. Extensive experiments on theCapProductdataset demonstrate the effectiveness and the superiority of the proposed CoPLH over the state-of-the-art baselines.
Gengwei Zhang, Xunlin Zhan, Yi Ding 0026, Yunchao Wei, Minlong Lu, Xiaodan Liang
IEEE Trans. Multim.6
2022 M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal Pretraining
abstract
Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different modalities capture complementary semantic information, we contribute a large-scale multi-modal pretraining dataset M5Product. The dataset comprises 5 modalities (image, text, table, video, and audio), covers over 6,000 categories and 5,000 attributes, and is 500× larger than the largest publicly available dataset with a similar number of modalities. Furthermore, M5Product contains incomplete modality pairs and noise while also having a long-tailed distribution, resembling most real-world problems. We further propose Self-harmonized ContrAstive LEarning (SCALE), a novel pretraining framework that integrates the different modalities into a unified model through an adaptive feature fusion mechanism, where the importance of each modality is learned directly from the modality embeddings and impacts the inter-modality contrastive learning and masked tasks within a multi-modal transformer model. We evaluate the current multi-modal pre-training state-of-the-art approaches and benchmark their ability to learn from unlabeled data when faced with the large number of modalities in the M5Product dataset. We conduct extensive experiments on four downstream tasks and demonstrate the superiority of our SCALE model, providing insights into the importance of dataset scale and diversity. Dataset and codes are available at11https://xiaodongsuper.github.io/M5Product_dataset/.
Xunlin Zhan, Yangxin Wu, Yunchao Wei, Michael Kampffmeyer, Xiaoyong Wei, Minlong Lu, Yaowei Wang 0001, Xiaodan Liang
CVPR7
2022 Bridging the Gap between Reality and Ideality of Entity Matching: A Revisting and Benchmark Re-Constrcution
abstract
Entity matching (EM) is the most critical step for entity resolution (ER). While current deep learning-based methods achieve very impressive performance on standard EM benchmarks, their real-world application performance is much frustrating. In this paper, we highlight that such the gap between reality and ideality stems from the unreasonable benchmark construction process, which is inconsistent with the nature of entity matching and therefore leads to biased evaluations of current EM approaches. To this end, we build a new EM corpus and re-construct EM benchmarks to challenge critical assumptions implicit in the previous benchmark construction process by step-wisely changing the restricted entities, balanced labels, and single-modal records in previous benchmarks into open entities, imbalanced labels, and multi-modal records in an open environment. Experimental results demonstrate that the assumptions made in the previous benchmark construction process are not coincidental with the open environment, which conceal the main challenges of the task and therefore significantly overestimate the current progress of entity matching. The constructed benchmarks and code are publicly released at https://github.com/tshu-w/ember.
Tianshu Wang 0002, Cheng Fu 0003, Xianpei Han, Le Sun 0001, Feiyu Xiong, Minlong Lu, Xiuwen Zhu
IJCAI8
2021 Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal Pretraining
abstract
Nowadays, customer’s demands for E-commerce are more diversified, which introduces more complications to the product retrieval industry. Previous methods are either subject to single-modal input or perform supervised image-level product retrieval, thus fail to accommodate real-life scenarios where enormous weakly annotated multi-modal data are present. In this paper, we investigate a more realistic setting that aims to perform weakly-supervised multi-modal instance-level product retrieval among fine-grained product categories. To promote the study of this challenging task, we contribute Product1M, one of the largest multi-modal cosmetic datasets for real-world instance-level retrieval. Notably, Product1M contains over 1 million image-caption pairs and consists of two sample types, i.e., single-product and multi-product samples, which encompass a wide variety of cosmetics brands. In addition to the great diversity, Product1M enjoys several appealing characteristics including fine-grained categories, complex combinations, and fuzzy correspondence that well mimic the real-world scenes. Moreover, we propose a novel model named Cross-modal contrAstive Product Transformer for instance-level prodUct REtrieval (CAPTURE), that excels in capturing the potential synergy between multi-modal inputs via a hybrid-stream transformer in a self-supervised manner. CAPTURE generates discriminative instance features via masked multi-modal learning as well as cross-modal contrastive pretraining and it outperforms several SOTA cross-modal baselines. Extensive ablation studies well demonstrate the effectiveness and the generalization capacity of our model. Dataset and codes are available at https: //github.com/zhanxlin/Product1M.
Xunlin Zhan, Yangxin Wu, Yunchao Wei, Minlong Lu, Hang Xu 0004, Xiaodan Liang
ICCV5
2019 Deep Attention Network for Egocentric Action Recognition
abstract
Recognizing a camera wearer's actions from videos captured by an egocentric camera is a challenging task. In this paper, we employ a two-stream deep neural network composed of an appearance-based stream and a motion-based stream to recognize egocentric actions. Based on the insight that human action and gaze behavior are highly coordinated in object manipulation tasks, we propose a spatial attention network to predict human gaze in the form of attention map. The attention map helps each of the two streams to focus on the most relevant spatial region of the video frames to predict actions. To better model the temporal structure of the videos, a temporal network is proposed. The temporal network incorporates bi-directional long short-term memory to model the long-range dependencies to recognize egocentric actions. The experimental results demonstrate that our method is able to predict attention maps that are consistent with human attention and achieve competitive action recognition performance with the state-of-the-art methods on the GTEA Gaze and GTEA Gaze+ datasets.
Minlong Lu, Ze-Nian Li, Yueming Wang 0001, Gang Pan 0001
IEEE Trans. Image Process.1
2016 Learning Contextual Dependencies for Optical Flow with Recurrent Neural Networks
Minlong Lu, Zhiwei Deng, Ze-Nian Li
ACCV (4)1
2015 Accelerometer-Based Gait Recognition by Sparse Representation of Signature Points With Clusters
abstract
Gait, as a promising biometric for recognizing human identities, can be nonintrusively captured as a series of acceleration signals using wearable or portable smart devices. It can be used for access control. Most existing methods on accelerometer-based gait recognition require explicit step-cycle detection, suffering from cycle detection failures and intercycle phase misalignment. We propose a novel algorithm that avoids both the above two problems. It makes use of a type of salient points termed signature points (SPs), and has three components: 1) a multiscale SP extraction method, including the localization and SP descriptors; 2) a sparse representation scheme for encoding newly emerged SPs with known ones in terms of their descriptors, where the phase propinquity of the SPs in a cluster is leveraged to ensure the physical meaningfulness of the codes; and 3) a classifier for the sparse-code collections associated with the SPs of a series. Experimental results on our publicly available dataset of 175 subjects showed that our algorithm outperformed existing methods, even if the step cycles were perfectly detected for them. When the accelerometers at five different body locations were used together, it achieved the rank-1 accuracy of 95.8% for identification, and the equal error rate of 2.2% for verification.
Yuting Zhang 0001, Gang Pan 0001, Kui Jia, Minlong Lu, Yueming Wang 0001, Zhaohui Wu 0001
IEEE Trans. Cybern.4
2015 Continuous Depth Map Reconstruction From Light Fields
abstract
In this paper, we investigate how the recently emerged photography technology--the light field--can benefit depth map estimation, a challenging computer vision problem. A novel framework is proposed to reconstruct continuous depth maps from light field data. Unlike many traditional methods for the stereo matching problem, the proposed method does not need to quantize the depth range. By making use of the structure information amongst the densely sampled views in light field data, we can obtain dense and relatively reliable local estimations. Starting from initial estimations, we go on to propose an optimization method based on solving a sparse linear system iteratively with a conjugate gradient method. Two different affinity matrices for the linear system are employed to balance the efficiency and quality of the optimization. Then, a depth-assisted segmentation method is introduced so that different segments can employ different affinity matrices. Experiment results on both synthetic and real light fields demonstrate that our continuous results are more accurate, efficient, and able to preserve more details compared with discrete approaches.
Jianqiao Li, Minlong Lu, Ze-Nian Li
IEEE Trans. Image Process.2
2013 Generating fluent tubes in video synopsis
abstract
Video synopsis is one of the effective techniques to build a short video representation preserving the essential activities for a long video. Existing methods usually have the problem that a continuous activity (tube) from a single moving object is separated to a few small pieces. In this paper, two schemes are proposed to generate fluent tubes for video synopsis. The Gaussian mixture model and a texture method are combined to detect more compact foreground with shadow removed. The foreground constitutes a set of initial trajectories. A particle filter tracker is used to concatenate two trajectories if they belong to the same foreground activity, which generates more fluent tubes for video synopsis. Experimental results on 4 videos show that our method produces better accuracies and visual effects in video synopsis.
Minlong Lu, Yueming Wang 0001, Gang Pan 0001
ICASSP1