VLDB 2026 Research / reviewers in the wild / expert
Feng Zhu 0006
dblp:71/2791-6
· DBLP profile ↗
48ranked-venue papers
5as first author
37since 2021 · last 2025
0000-0003-4309-170XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 3 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated DataabstractYu Zhang, Ruijie Yu, Jidong Tian, Feng Zhu, Jiapeng Liu, Xiaokang Yang, Yaohui Jin, Yanyan Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruijie Yu, Jidong Tian, Feng Zhu 0006, Xiaokang Yang 0001, Yaohui Jin, Yanyan Xu 0002 |
ACL (1) | 4 |
| 2025 | Graph Pooling via Dropping Task-Irrelevant NodesabstractGraph neural networks (GNNs) face scalability challenges. While recent approaches have adopted pooling strategies inspired by convolutional neural networks (CNNs) to reduce graph size and improve efficiency, these methods often focus on local information and are optimized for single graph-level tasks. This limitation hinders their effectiveness in multi-task scenarios that require task-specific global information. We present DOTIN (Dropping Out Task-Irrelevant Nodes), an approach to graph size reductio. DOTIN utilizes K learnable virtual nodes to represent graph embeddings for K distinct graph-level tasks. By employing a transformer-based attention model, it adaptively removes up to 90% of low-attentiveness raw nodes without notable performance degradation. Our method achieves comparable accuracy to state-of-the-art techniques while offering substantial benefits in efficiency. Specifically, DOTIN accelerates Graph Attention Networks (GAT) by approximately 50% on graph-level tasks such as graph classification and graph edit distance (GED). Additionally, it reduces memory usage by about 60% on the D&D dataset. These results show DOTIN's potential to enhance the scalability and efficiency of deep GNNs across multiple graph-level tasks while maintaining high performance. Shaofeng Zhang, Feng Zhu 0006, Rui Zhao 0001, Xiaokang Yang 0001, Junchi Yan |
ICASSP | 3 |
| 2025 | Re-Aligning Language to Visual Objects with an Agentic WorkflowabstractLanguage-based object detection (LOD) aims to align visual objects with language expressions. A large amount of paired data is utilized to improve LOD model generalizations. During the training process, recent studies leverage vision-language models (VLMs) to automatically generate human-like expressions for visual objects, facilitating training data scaling up. In this process, we observe that VLM hallucinations bring inaccurate object descriptions (e.g., object name, color, and shape) to deteriorate VL alignment quality. To reduce VLM hallucinations, we propose an agentic workflow controlled by an LLM to re-align language to visual objects via adaptively adjusting image and text prompts. We name this workflow Real-LOD, which includes planning, tool use, and reflection steps. Given an image with detected objects and VLM raw language expressions, Real-LOD reasons its state automatically and arranges action based on our neural symbolic designs (i.e., planning). The action will adaptively adjust the image and text prompts and send them to VLMs for object re-description (i.e., tool use). Then, we use another LLM to analyze these refined expressions for feedback (i.e., reflection). These steps are conducted in a cyclic form to gradually improve language descriptions for re-aligning to visual objects. We construct a dataset that contains a tiny amount of 0.18M images with re-aligned language expression and train a prevalent LOD model to surpass existing LOD methods by around 50% on the standard benchmarks. Our Real-LOD workflow, with automatic VL refinement, reveals a potential to preserve data quality along with scaling up data quantity, which further improves LOD performance from a data-alignment perspective. Jiangyan Feng, Lijun Gong, Feng Zhu 0006, Rui Zhao 0001, Qibin Hou, Ming-Ming Cheng, Yibing Song |
ICLR | 5 |
| 2025 | Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-IdentificationabstractRecently, person re-identification (ReID) has witnessed fast development due to its broad practical applications and proposed various settings, e.g., traditional ReID, clothes-changing ReID, and visible-infrared ReID. However, current studies primarily focus on single specific tasks, which limits model applicability in real-world scenarios. This paper aims to address this issue by introducing a novel instruct-ReID task that unifies 6 existing ReID tasks in one model and retrieves images based on provided visual or textual instructions. Instruct-ReID is the first exploration of a general ReID setting, where 6 existing ReID tasks can be viewed as special cases by assigning different instructions. To facilitate research in this new instruct-ReID task, we propose a large-scale OmniReID++ benchmark equipped with diverse data and comprehensive evaluation methods, e.g., task-specific and task-free evaluation settings. In the task-specific evaluation setting, gallery sets are categorized according to specific ReID tasks. We propose a novel baseline model, IRM, with an adaptive triplet loss to handle various retrieval tasks within a unified framework. For task-free evaluation setting, where target person images are retrieved from task-agnostic gallery sets, we further propose a new method called IRM++ with novel memory bank-assisted learning. Extensive evaluations of IRM and IRM++ on OmniReID++ benchmark demonstrate the superiority of our proposed methods, achieving state-of-the-art performance on 10 test sets. Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu 0006, Yizhou Wang 0007, Lei Bai 0001, Qingsong Xie, Rui Zhao 0001, Donglian Qi, Wanli Ouyang, Shixiang Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Hulk: A Universal Knowledge Translator for Human-Centric TasksabstractHuman-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports analysis. There is a recent surge to develop human-centric foundation models that can benefit a broad range of human-centric perception tasks. While many human-centric foundation models have achieved success, they did not explore 3D and vision-language tasks for human-centric and required task-specific finetuning. These limitations restrict their application to more downstream tasks and situations. To tackle these problems, we present Hulk, the first multimodal human-centric generalist model, capable of addressing 2D vision, 3D vision, skeleton-based, and vision-language tasks without task-specific finetuning. The key to achieving this is condensing various task-specific heads into two general heads, one for discrete representations, e.g., languages, and the other for continuous representations, e.g., location coordinates. The outputs of two heads can be further stacked into four distinct input and output modalities. This uniform representation enables Hulk to treat diverse human-centric tasks as modality translation, integrating knowledge across a wide range of tasks. Comprehensive evaluations of Hulk on 12 benchmarks covering 8 human-centric tasks demonstrate the superiority of our proposed method, achieving state-of-the-art performance in 11 benchmarks. Yizhou Wang 0007, Weizhen He, Xun Guo 0001, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Jian Wu 0001, Tong He 0001, Wanli Ouyang, Shixiang Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | RelationLMM: Large Multimodal Model as Open and Versatile Visual Relationship GeneralistabstractVisual relationships are crucial for visual perception and reasoning, and cover tasks like Scene Graph Generation, Human-Object Interaction, and object affordance. Despite significant efforts, this field still suffers from the following limitations: specialists for a specific task without considering similar ones, strict and complex task formulations with limited flexibility, and underexploited reasoning with language and knowledge. To solve these limitations, we seek to build a new framework, one model for all tasks, over Large Multimodal Models (LMMs). LMMs offer the potential of unifying tasks, flexible forms, and reasoning with language. However, they fail to handle visual relationship tasks well. We find the obstacles include the conflicts between different tasks and insufficient instance-level information. We solve these problems by reforming the data for LMMs, rather than architectures, considering their strong language-in language-out capability. We propose to disassemble tasks into simple and common sub-tasks, verbally estimate instance confidence, and augment instance diversity, all without additional modules. These strategies help us build a visual relationship generalist, RelationLMM, with a simple architecture. Exhaustive experiments demonstrate RelationLMM is strong, generalizable and flexible to different tasks, with one model and one suite of weight. Chi Xie 0001, Shuang Liang 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | VisionTraj: A Noise-Robust Trajectory Recovery Framework Based on Large-Scale Camera NetworkabstractTrajectory recovery from snapshots captured by a city-wide multi-camera network facilitates urban mobility sensing and road network optimization. State-of-the-art solutions for such vision-based schemes typically rely on predefined rules or unsupervised iterative feedback, but they struggle with multiple challenges, such as the lack of open-source datasets for training the entire pipeline and the vulnerability to noise in visual inputs. In response to the dilemma, this paper proposes VisionTraj, the first learning-based model that reconstructs vehicle trajectories from snapshots recorded by road network cameras. Along with this, we present two well-designed vision-trajectory datasets that provide extensive trajectory data and corresponding visual snapshots, enabling the extraction of supervised vision-trajectory interactions. After the data creation, based on the results from the off-the-shelf multi-modal vehicle clustering, we first re-formulate the trajectory recovery problem as a generative task and introduce the canonical Transformer as the autoregressive backbone. Next, to identify clustering noise (i.e., false positives) based on the snapshots’ spatiotemporal dependencies, a graph convolutional neural network-based soft-denoising module is built upon the fine- and coarse-grained clusters. Additionally, we leverage strong semantic information extracted from the tracklet to provide detailed insights into the vehicle’s entry and exit behaviors during trajectory recovery. The denoising and tracklet components can also serve as plug-and-play modules to enhance baselines. Experimental results on the two hand-crafted datasets show that the proposed VisionTraj achieves a maximum improvement of +11.5% against the sub-best model. Furthermore, we explore potential downstream applications, and our model continues to outperform its peers. The code and data are available herehttps://github.com/bonaldli/VisionTraj Zhishuai Li, Ziyue Li 0002, Xiaoru Hu, Guoqing Du, Yunhao Nie, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Instruct-ReID: A Multi-Purpose Person Re-Identification Task with InstructionsabstractHuman intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a new instruct-ReID task that requires the model to retrieve images according to the given image or language instructions. Our instruct-ReID is a more general ReID setting, where existing 6 ReID tasks can be viewed as special cases by designing different instructions. We propose a large-scale OmniReID benchmark and an adaptive triplet loss as a baseline method to facilitate research in this new setting. Experimental results show that the proposed multi-purpose ReID model, trained on our OmniReID benchmark without finetuning, can improve +0.5%, +0.6%, +7.7% mAP on Market1501, MSMT17, CUHK03 for traditional ReID, +6.4%, +7.1%, +11.2% mAP on PRCC, VC-Clothes, LTCC for clothes-changing ReID, +11.7% mAP on COCAS+ real2 for clothes template based clothes-changing ReID when using only RGB images, +24.9% mAP on COCAS+ real2 for our newly defined language-instructed ReID, +4.3% on LLCM for visible-infrared ReID, +2.6% on CUHK-PEDES for text-to-image ReID. The datasets, the model, and code are available at https://github.com/hwz-zju/Instruct-ReID. Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang 0007, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Wanli Ouyang, Donglian Qi, Yunfeng Yan |
CVPR | 8 |
| 2024 | InstructDET: Diversifying Referring Object Detection with Generalized InstructionsabstractWe propose InstructDET, a data-centric method for referring object detection (ROD) that localizes target objects based on user instructions. While deriving from referring expressions (REC), the instructions we leverage are greatly diversified to encompass common user intentions related to object detection. For one image, we produce tremendous instructions that refer to every single object and different combinations of multiple objects. Each instruction and its corresponding object bounding boxes (bbxs) constitute one training data pair. In order to encompass common detection expressions, we involve emerging vision-language model (VLM) and large language model (LLM) to generate instructions guided by text prompts and object bbxs, as the generalizations of foundation models are effective to produce human-like expressions (e.g., describing object property, category, and relationship). We name our constructed dataset as InDET. It contains images, bbxs and generalized instructions that are from foundation models. Our InDET is developed from existing REC datasets and object detection datasets, with the expanding potential that any image with object bbxs can be incorporated through using our InstructDET method. By using our InDET dataset, we show that a conventional ROD model surpasses existing methods on standard REC datasets and our InDET test set. Our data-centric method InstructDET, with automatic data expansion by leveraging foundation models, directs a promising field that ROD can be greatly diversified to execute common object detection instructions. Ronghao Dang, Jiangyan Feng, Chongjian Ge, Lin Song 0002, Lijun Gong, Feng Zhu 0006, Rui Zhao 0001, Yibing Song |
ICLR | 9 |
| 2024 | Relation-Aware Distribution Representation Network for Person Clustering With Multiple ModalitiesabstractPerson clustering with multi-modal clues, including faces, bodies, and voices, is critical for various tasks, such as movie parsing and identity-based movie editing. Related methods such as multi-view clustering mainly project multi-modal features into a joint feature space. However, multi-modal clue features are usually rather weakly correlated due to the semantic gap from the modality-specific uniqueness. As a result, these methods are not suitable for person clustering. In this paper, we propose aRelation-AwareDistribution representation Network (RAD-Net) to generate adistribution representationfor multi-modal clues. The distribution representation of a clue is a vector consisting of the relation between this clue and all other clues from all modalities, thus beingmodality agnosticand good for person clustering. Accordingly, we introduce a graph-based method to construct distribution representation and employ a cyclic update policy to refine distribution representation progressively. Our method achieves substantial improvements of+6%and+8.2%in F-score on the Video Person-Clustering Dataset (VPCD) and VoxCeleb2 multi-view clustering dataset, respectively. Codes will be released athttps://github.com/bonaldli/RADNet. Kaijian Liu, Shixiang Tang, Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Structured Domain Adaptation With Online Relation Regularization for Unsupervised Person Re-IDabstractUnsupervised domain adaptation (UDA) aims at adapting the model trained on a labeled source-domain dataset to an unlabeled target-domain dataset. The task of UDA on open-set person reidentification (re-ID) is even more challenging as the identities (classes) do not have overlap between the two domains. One major research direction was based on domain translation, which, however, has fallen out of favor in recent years due to inferior performance compared with pseudo-label-based methods. We argue that domain translation has great potential on exploiting valuable source-domain data but the existing methods did not provide proper regularization on the translation process. Specifically, previous methods only focus on maintaining the identities of the translated images while ignoring the intersample relations during translation. To tackle the challenges, we propose an end-to-end structured domain adaptation framework with an online relation-consistency regularization term. During training, the person feature encoder is optimized to model intersample relations on-the-fly for supervising relation-consistency domain translation, which in turn improves the encoder with informative translated images. The encoder can be further improved with pseudo labels, where the source-to-target translated images with ground-truth identities and target-domain images with pseudo identities are jointly used for training. In the experiments, our proposed framework is shown to achieve state-of-the-art performance on multiple UDA tasks of person re-ID. With the synthetic→real translated images from our structured domain-translation network, we achieved second place in the Visual Domain Adaptation Challenge (VisDA) in 2020. Yixiao Ge, Feng Zhu 0006, Dapeng Chen, Rui Zhao 0001, Xiaogang Wang 0001, Hongsheng Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | UniHCP: A Unified Model for Human-Centric PerceptionsabstractHuman-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underlying semantic structure of the human body. However, few works have attempted to exploit such homogeneity and design a general-propose model for human-centric tasks. In this work, we revisit a broad range of human-centric tasks and unify them in a minimalist manner. We propose UniHCP, a Unified Model for Human-Centric Perceptions, which unifies a wide range of human-centric tasks in a simplified end-to-end manner with the plain vision transformer architecture. With large-scale joint training on 33 human-centric datasets, UniHCP can outperform strong baselines on several in-domain and downstream tasks by direct evaluation. When adapted to a specific task, UniHCP achieves new SOTAs on a wide range of human-centric tasks, e.g., 69.8 mIoU on CIHP for human parsing, 86.18 mA on PA100K for attribute prediction, 90.3 mAP on Market1501 for ReID, and 85.8 JI on CrowdHuman for pedestrian detection, performing better than specialized models tailored for each task. The code and pretrained model are available at https://github.com/OpenGVLab/UniHCP. Yuanzheng Ci, Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Fengwei Yu, Donglian Qi, Wanli Ouyang |
CVPR | 6 |
| 2023 | HumanBench: Towards General Human-Centric Perception with Projector Assisted PretrainingabstractHuman-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench. Shixiang Tang, Qingsong Xie, Meilin Chen, Yizhou Wang 0007, Yuanzheng Ci, Lei Bai 0001, Feng Zhu 0006, Haiyang Yang, Rui Zhao 0001, Wanli Ouyang |
CVPR | 8 |
| 2023 | CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingabstractOpen-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OVD methods rely on large-scale visual-language pre-trained models, such as CLIP, for recognizing novel objects. We identify the two core obstacles that need to be tackled when incorporating these models into detector training: (1) the distribution mismatch that happens when applying a VL-model trained on whole images to region recognition tasks; (2) the difficulty of localizing objects of unseen classes. To overcome these obstacles, we propose CORA, a DETR-style framework that adapts CLIP for Open-vocabulary detection by Region prompting and Anchor pre-matching. Region prompting mitigates the whole-to-region distribution gap by prompting the region features of the CLIP-based region classifier. Anchor pre-matching helps learning generalizable object localization by a class-aware matching mechanism. We evaluate CORA on the COCO OVD benchmark, where we achieve 41.7 AP50 on novel classes, which outperforms the previous SOTA by 2.4 AP50 even without resorting to extra training data. When extra training data is available, we train CORA+ on both ground-truth base-category annotations and additional pseudo bounding box labels computed by CORA. CORA+ achieves 43.1 AP50 on the COCO OVD benchmark and 28.1 box APr on the LVIS OVD benchmark. The code is available at https://github.com/tgxs002/CORA. Xiaoshi Wu, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001 |
CVPR | 2 |
| 2023 | Trust Your Partner's Friends: Hierarchical Cross-Modal Contrastive Pre-Training for Video-Text RetrievalabstractVideo-text retrieval has greatly benefited from the massive web video in recent years, while the performance is still limited to the weak supervision from the uncurated data. In this work, we propose to leverage the well-represented information of each original modality and exploit complementary information in two views of the same video, i.e., video clips and captions, by using one view to obtain positive samples with the neighboring samples of the other. Respecting the hierarchical organization of real-world data, we further design a hierarchical cross-modal pre-training method (HCP) to learn good representations in the common embedding space. We evaluate the pre-trained model on three downstream tasks, i.e. text-to-video retrieval, action step localization and video question answering and our method outperforms previous works under the same setting. Yuhan Xiang, Kaijian Liu, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Xianming Lin |
ICASSP | 5 |
| 2023 | Human Preference Score: Better Aligning Text-to-image Models with Human PreferenceabstractRecent years have witnessed a rapid growth of deep generative models, with text-to-image models gaining significant attention from the public. However, existing models often generate images that do not align well with human preferences, such as awkward combinations of limbs and facial expressions. To address this issue, we collect a dataset of human choices on generated images from the Stable Foundation Discord channel. Our experiments demonstrate that current evaluation metrics for generative models do not correlate well with human choices. Thus, we train a human preference classifier with the collected dataset and derive a Human Preference Score (HPS) based on the classifier. Using HPS, we propose a simple yet effective method to adapt Stable Diffusion to better align with human preferences. Our experiments show that HPS outperforms CLIP in predicting human choices and has good generalization capability toward images generated from other models. By tuning Stable Diffusion with the guidance of HPS, the adapted model is able to generate images that are more preferred by human users. The project page is available here: https://tgxs002.github.io/alignsd-web/. Xiaoshi Wu, Keqiang Sun, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001 |
ICCV | 3 |
| 2023 | Advancing Referring Expression Segmentation Beyond Single ImageabstractReferring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expression. However, in broader real-world scenarios, it is not always possible to determine if the described object exists in a specific image. Generally, a collection of images is available, some of which potentially contain the target objects. To this end, we propose a more realistic setting, named Group-wise Referring Expression Segmentation (GRES), which expands RES to a group of related images, allowing the described objects to exist in a subset of the input image group. To support this new setting, we introduce an elaborately compiled dataset named Grouped Referring Dataset (GRD), containing complete group-wise annotations of the target objects described by given expressions. Moreover, we also present a baseline method named Grouped Referring Segmenter (GRSer), which explicitly captures the language-vision and intra-group vision-vision interactions to achieve state-of-the-art results on the proposed GRES setting and related tasks, such as Co-Salient Object Detection and traditional RES. Our dataset and codes are publicly released in https://github.com/shikras/d-cube. Zhao Zhang 0018, Chi Xie 0001, Feng Zhu 0006, Rui Zhao 0001 |
ICCV | 4 |
| 2023 | Cycle-consistent Masked AutoEncoder for Unsupervised Domain Generalization
Haiyang Yang, Shixiang Tang, Feng Zhu 0006, Yizhou Wang 0007, Meilin Chen, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ICLR | 4 |
| 2023 | Contextual Image Masking Modeling via Synergized Contrasting without View Augmentation for Faster and Better Visual Pretraining
Shaofeng Zhang, Feng Zhu 0006, Rui Zhao 0001, Junchi Yan |
ICLR | 2 |
| 2023 | Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation Learning
Shaofeng Zhang, Feng Zhu 0006, Rui Zhao 0001, Junchi Yan |
ICLR | 2 |
| 2023 | Described Object Detection: Liberating Object Detection with Flexible ExpressionsabstractDetecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to flexible language expressions for OVD and overcoming the limitation of REC only grounding the pre-existing object. We establish the research foundation for DOD by constructing a *Description Detection Dataset* ($D^3$). This dataset features flexible language expressions, whether short category names or long descriptions, and annotating all described objects on all images without omission. By evaluating previous SOTA methods on $D^3$, we find some troublemakers that fail current REC, OVD, and bi-functional methods. REC methods struggle with confidence scores, rejecting negative instances, and multi-target scenarios, while OVD methods face constraints with long and complex descriptions. Recent bi-functional methods also do not work well on DOD due to their separated training procedures and inference strategies for REC and OVD tasks. Building upon the aforementioned findings, we propose a baseline that largely improves REC methods by reconstructing the training data and introducing a binary classification sub-task, outperforming existing methods. Data and code are available at https://github.com/shikras/d-cube and related works are tracked in https://github.com/Charles-Xie/awesome-described-object-detection. Chi Xie 0001, Zhao Zhang 0018, Feng Zhu 0006, Rui Zhao 0001, Shuang Liang 0001 |
NeurIPS | 4 |
| 2023 | COCAS+: Large-Scale Clothes-Changing Person Re-Identification With Clothes TemplatesabstractRecent years person re-identification (ReID) has been developed rapidly due to its broad practical applications. Most existing benchmarks assume that the same person wears the same clothes across captured images, while, in real-world scenarios, person may change his/her clothes frequently. Thus the Clothes-Changing person ReID (CC-ReID) problem is introduced and several related benchmarks are established. CC-ReID is a very difficult task as the main visual characteristics of a human body, clothes, are different between query and gallery, and clothes-irrelevant features are relatively weak. To promote the research and applications of person ReID in clothes-changing scenarios, in this paper, we introduce a new task called Clothes Template based Clothes-Changing person ReID (CTCC-ReID), where the query image is enhanced by a clothes template which shares similar visual patterns with the clothes of the target person image in the gallery. So, ReID methods are encouraged to jointly consider the original query image and the given clothes template for retrieval in the proposed CTCC-ReID setting. To facilitate research works on CTCC-ReID, we construct a novel large-scale ReID dataset named ClOthes ChAnging person Set Plus (COCAS+), which contains both realistic and synthetic clothes-changing person images with manually collected clothes templates. Furthermore, we propose a novel Dual-Attention Biometric-Clothes Transfusion Network (DualBCT-Net) for CTCC-ReID, which can effectively learn to extract biometric features from the original query person image and clothes features from the given clothes template and then fuse them through a Dual-Attention Fusion Module. Extensive experimental results show that the proposed CTCC-ReID setting and COCAS+ dataset can help greatly push the performance of clothes-changing ReID toward practical applications, and synthetic data is impressively effective for CTCC-ReID. What’s more, the proposed DualBCT-Net shows significant improvements over state-of-the-art methods on the CTCC-ReID task. COCAS+ and code of DualBCT-Net will be released inhttps://github.com/Chenhaobin/COCAS-plus. Shihua Li 0006, Shijie Yu, Zhiqun He, Feng Zhu 0006, Rui Zhao 0001, Jie Chen 0012, Yu Qiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | AutoMA: Towards Automatic Model Augmentation for Transferable Adversarial AttacksabstractRecent adversarial attack works attempt to improve the transferability by applying various differentiable transformations on input images. Considering the differentiable transformations and the original model together as a new model, these methods can be regarded as model augmentation that effectively derives an ensemble of models from the single original model. Despite their impressive performance, the model augmentation policies used in these methods are manually designed by experimental attempts, leaving the design of model augmentation policy an open question. In this paper, we propose an Automatic Model Augmentation (AutoMA) approach to find a strong model augmentation policy for transferable adversarial attacks. Specifically, we design a discrete search space that contains various diffierentiable transformations with different parameters and adopt reinforcement learning to search for the strong augmentation policy. The sampled augmentation policies together with the rewards they obtain during the searching process reveal several valuable observations for designing more powerful attacks using model augmentation policy:1) Augmentation transformations on color space are less effective; 2) The transformation type diversity matters; and 3) Using small distortion for geometric transformations while larger distortion for intensity transformations.Extensive experiments show that the augmentation policy found by AutoMA achieves superior performance than existing manually designed policies in a wide range of cases. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Multim. | 3 |
| 2022 | Revisiting the Transferability of Supervised Pretraining: an MLP PerspectiveabstractThe pretrain-finetune paradigm is a classical pipeline in visual learning. Recent progress on unsupervised pretraining methods shows superior transfer performance to their supervised counterparts. This paper revisits this phenomenon and sheds new light on understanding the transferability gap between unsupervised and supervised pretraining from a multilayer perceptron (MLP) perspective. While previous works [6], [8], [17] focus on the effectiveness of MLP on unsupervised image classification where pretraining and evaluation are conducted on the same dataset, we reveal that the MLP projector is also the key factor to better transferability of unsupervised pretraining methods than supervised pretraining methods. Based on this observation, we attempt to close the transferability gap between supervised and unsupervised pretraining by adding an MLP projector before the classifier in supervised pretraining. Our analysis indicates that the MLP projector can help retain intra-class variation of visual features, decrease the feature distribution distance between pretraining and evaluation datasets, and reduce feature redundancy. Extensive experiments on public benchmarks demonstrate that the added MLP projector significantly boosts the transferability of supervised pretraining, e.g. +7.2% top-1 accuracy on the concept generalization task, +5.8% top-1 accuracy for linear evaluation on 12 -domain classification tasks, and +0.8% AP on COCO object detection task, making supervised pretraining comparable or even better than unsupervised pretraining. Yizhou Wang 0007, Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Donglian Qi, Wanli Ouyang |
CVPR | 3 |
| 2022 | Feature Erasing and Diffusion Network for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) aims at matching occluded person images to holistic ones across different camera views. Target Pedestrians (TP) are often disturbed by Non-Pedestrian Occlusions (NPO) and Non-Target Pedestrians (NTP). Previous methods mainly focus on increasing the model's robustness against NPO while ignoring feature contamination from NTP. In this paper, we propose a novel Feature Erasing and Diffusion Network (FED) to simultaneously handle challenges from NPO and NTP. Specifically, aided by the NPO augmentation strategy that simulates NPO on holistic pedestrian images and gen-erates precise occlusion masks, NPO features are explicitly eliminated by our proposed Occlusion Erasing Module (OEM). Subsequently, we diffuse the pedestrian representations with other memorized features to synthesize the NTP characteristics in the feature space through the novel Feature Diffusion Module (FDM). With the guidance of the occlusion scores from OEM, the feature diffusion process is conducted on visible body parts, thereby improving the quality of the synthesized NTP characteristics. We can greatly improve the model's perception ability towards TP and alleviate the influence of NPO and NTP by jointly optimizing OEM and FDM. Furthermore, the proposed FDM works as an auxiliary module for training and will not be engaged in the inference phase, thus with high flexibility. Experiments on occluded and holistic person ReID benchmarks demonstrate the superiority of FED over state-of-the-art methods. Zhikang Wang, Feng Zhu 0006, Shixiang Tang, Rui Zhao 0001, Lihuo He, Jiangning Song |
CVPR | 2 |
| 2022 | Align Representations with Base: A New Approach to Self-Supervised LearningabstractExisting symmetric contrastive learning methods suffer from collapses (complete and dimensional) or quadratic complexity of objectives. Departure from these methods which maximize mutual information of two generated views, along either instance or feature dimension, the proposed paradigm introduces intermediate variables at the feature level, and maximizes the consistency between variables and representations of each view. Specifically, the proposed intermediate variables are the nearest group of base vectors to representations. Hence, we call the proposed method ARB (Align Representations with Base). Compared with other symmetric approaches, ARB 1) does not require negative pairs, which leads the complexity of the overall objective function is in linear order, 2) reduces feature redundancy, increasing the information density of training samples, 3) is more robust to output dimension size, which out-performs previous feature-wise arts over 28% Top-1 accuracy on ImageNet-100under low-dimension settings. Shaofeng Zhang, Lyn Qiu, Feng Zhu 0006, Junchi Yan, Rui Zhao 0001, Hongyang Li 0001, Xiaokang Yang 0001 |
CVPR | 3 |
| 2022 | Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification
Xulin Li, Yan Lu 0001, Bin Liu 0016, Guojun Yin, Qi Chu 0001, Jinyang Huang, Feng Zhu 0006, Rui Zhao 0001, Nenghai Yu |
ECCV (26) | 8 |
| 2022 | Unifying Visual Contrastive Learning for Object Recognition from a Graph Perspective
Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Chenyu Wang 0001, Wanli Ouyang |
ECCV (26) | 2 |
| 2022 | Relative Contrastive Loss for Unsupervised Representation Learning
Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ECCV (27) | 2 |
| 2022 | Domain Invariant Masked Autoencoders for Self-supervised Learning from Multi-domains
Haiyang Yang, Shixiang Tang, Meilin Chen, Yizhou Wang 0007, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Wanli Ouyang |
ECCV (31) | 5 |
| 2022 | Zero-CL: Instance and Feature decorrelation for negative-free symmetric contrastive learning
Shaofeng Zhang, Feng Zhu 0006, Junchi Yan, Rui Zhao 0001, Xiaokang Yang 0001 |
ICLR | 2 |
| 2022 | Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector LearningabstractUnsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to produce object priors, \emph{e.g.,} selective search, which separates the prior generation and detector learning and leads to sub-optimal solutions. In this work, we propose a novel object detection pretraining framework that could generate object priors and learn detectors jointly by generating accurate object priors from the model itself. Specifically, region priors are extracted by attention maps from the encoder, which highlights foregrounds. Instance priors are the selected high-quality output bounding boxes of the detection decoder. By assuming objects as instances in the foreground, we can generate object priors with both region and instance priors. Moreover, our object priors are jointly refined along with the detector optimization. With better object priors as supervision, the model could achieve better detection capability, which in turn promotes the object priors generation. Our method improves the competitive approaches by \textbf{+1.3 AP}, \textbf{+1.7 AP} in 1\% and 10\% COCO low-data regimes object detection. Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Feng Zhu 0006, Haiyang Yang, Lei Bai 0001, Rui Zhao 0001, Yunfeng Yan, Donglian Qi, Wanli Ouyang |
NeurIPS | 4 |
| 2021 | Temporal ROI Align for Video Object RecognitionabstractVideo object detection is challenging in the presence of appearance deterioration in certain video frames. Therefore, it is a natural choice to aggregate temporal information from other frames of the same video into the current frame. However, ROI Align, as one of the most core procedures of video detectors, still remains extracting features from a single-frame feature map for proposals, making the extracted ROI features lack temporal information from videos. In this work, considering the features of the same object instance are highly similar among frames in a video, a novel Temporal ROI Align operator is proposed to extract features from other frames feature maps for current frame proposals by utilizing feature similarity. The proposed Temporal ROI Align operator can extract temporal information from the entire video for proposals. We integrate it into single-frame video detectors and other state-of-the-art video detectors, and conduct quantitative experiments to demonstrate that the proposed Temporal ROI Align operator can consistently and significantly boost the performance. Besides, the proposed Temporal ROI Align can also be applied into video instance segmentation. Kai Chen 0026, Xinjiang Wang, Qi Chu 0001, Feng Zhu 0006, Dahua Lin, Nenghai Yu, Huamin Feng |
AAAI | 5 |
| 2021 | MixMix: All You Need for Data-Free Compression Are Feature and Data MixingabstractUser data confidentiality protection is becoming a rising challenge in the present deep learning research. Without access to data, conventional data-driven model compression faces a higher risk of performance degradation. Recently, some works propose to generate images from a specific pretrained model to serve as training data. However, the inversion process only utilizes biased feature statistics stored in one model and is from low-dimension to high-dimension. As a consequence, it inevitably encounters the difficulties of generalizability and inexact inversion, which leads to unsatisfactory performance. To address these problems, we propose MixMix based on two simple yet effective techniques: (1) Feature Mixing: utilizes various models to construct a universal feature space for generalized inversion; (2) Data Mixing: mixes the synthesized images and labels to generate exact label information. We prove the effectiveness of MixMix from both theoretical and empirical perspectives. Extensive experiments show that MixMix outperforms existing methods on the mainstream compression tasks, including quantization, knowledge distillation and pruning. Specifically, MixMix achieves up to 4% and 20% accuracy uplift on quantization and pruning, respectively, compared to existing data-free compression work. Yuhang Li 0001, Feng Zhu 0006, Ruihao Gong, Mingzhu Shen, Xin Dong 0009, Fengwei Yu, Shaoqing Lu, Shi Gu |
ICCV | 2 |
| 2021 | Progressive Correspondence Pruning by Consensus LearningabstractCorrespondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The pruning process is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further complicates the learning process for learning-based methods. To address this issue, we propose to progressively prune the correspondences via a local-to-global consensus learning procedure. We introduce a "pruning" block that lets us identify reliable candidates among the initial matches according to consensus scores estimated using local-to-global dynamic graphs. We then achieve progressive pruning by stacking multiple pruning blocks sequentially. Our method outperforms state-of-the-arts on robust line fitting, camera pose estimation and retrieval-based image localization benchmarks by significant margins and shows promising generalization ability to different datasets and detector/descriptor combinations. Chen Zhao 0025, Yixiao Ge, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001, Mathieu Salzmann |
ICCV | 3 |
| 2021 | Improving Facial Attribute Recognition by Group and Graph LearningabstractExploiting the relationships between attributes is a key challenge for improving multiple facial attribute recognition. In this work, we are concerned with two types of correlations that are spatial and non-spatial relationships. For the spatial correlation, we aggregate attributes with spatial similarity into a part-based group and then introduce a Group Attention Learning to generate the group attention and the part-based group feature. On the other hand, to discover the non-spatial relationship, we model a group-based Graph Correlation Learning to explore affinities of predefined part-based groups. We utilize such affinity information to control the communication between all groups and then refine the learned group features. Overall, we propose a unified network called Multi-scale Group and Graph Network. It incorporates these two newly proposed learning strategies and produces coarse-to-fine graph-based group features for improving facial attribute recognition. Comprehensive experiments demonstrate that our approach outperforms the state-of-the-art methods. Shuhang Gu, Feng Zhu 0006, Rui Zhao 0001 |
ICME | 3 |
| 2021 | Efficient Open-Set Adversarial Attacks on Deep Face RecognitionabstractDifferent from close-set classification task, deep face recognition models are often used in open-set scenarios, where the models need to handle arbitrary faces. Open-set adversarial attacks can identify the vulnerability of deep face recognition models. Compared to time-consuming iterative gradient-based methods, generator-based methods can produce adversarial examples with only one forward pass, which greatly improves attack efficiency. However, existing generator-based attack methods need to train an individual model for each target identity and can only generate a fixed perturbation pattern regardless of different attack intensity constraints, which is impractical and sub-optimal for open-set adversarial attacks. In this paper, we propose an efficient generator-based Single Model ARbitrary Target (SMART) approach for open-set adversarial attacks against deep face recognition models. Given an arbitrary source-target face image pair, SMART first generates an additive perturbation and then adds it to the source image to obtain the final adversarial face image. After the training with various source-target pairs randomly sampled on large scale face images, SMART could effectively learn inherent perturbation patterns for arbitrary source-target face images pairs. Besides, we also propose a novel Constraint-aware Adversarial Decoder (CAD) module, which makes SMART the first generator-based method that could produce adaptive adversarial patterns according to different constraints on attack intensity. Extensive experimental results in various settings demonstrate the effectiveness of the proposed method. Qi Chu 0001, Feng Zhu 0006, Rui Zhao 0001, Bin Liu 0016, Nenghai Yu |
ICME | 3 |
| 2020 | DASOT: A Unified Framework Integrating Data Association and Single Object Tracking for Online Multi-Object TrackingabstractIn this paper, we propose an online multi-object tracking (MOT) approach that integrates data association and single object tracking (SOT) with a unified convolutional network (ConvNet), named DASOTNet. The intuition behind integrating data association and SOT is that they can complement each other. Following Siamese network architecture, DASOTNet consists of the shared feature ConvNet, the data association branch and the SOT branch. Data association is treated as a special re-identification task and solved by learning discriminative features for different targets in the data association branch. To handle the problem that the computational cost of SOT grows intolerably as the number of tracked objects increases, we propose an efficient two-stage tracking method in the SOT branch, which utilizes the merits of correlation features and can simultaneously track all the existing targets within one forward propagation. With feature sharing and the interaction between them, data association branch and the SOT branch learn to better complement each other. Using a multi-task objective, the whole network can be trained end-to-end. Compared with state-of-the-art online MOT methods, our method is much faster while maintaining a comparable performance. Qi Chu 0001, Wanli Ouyang, Bin Liu 0016, Feng Zhu 0006, Nenghai Yu |
AAAI | 4 |
| 2020 | Towards Unified INT8 Training for Convolutional Neural NetworkabstractRecently low-bit (e.g., 8-bit) network quantization has been extensively studied to accelerate the inference. Besides inference, low-bit training with quantized gradients can further bring more considerable acceleration, since the backward process is often computation-intensive. Unfortunately, the inappropriate quantization of backward propagation usually makes the training unstable and even crash. There lacks a successful unified low-bit training framework that can support diverse networks on various tasks. In this paper, we give an attempt to build a unified 8-bit (INT8) training framework for common convolutional neural networks from the aspects of both accuracy and speed. First, we empirically find the four distinctive characteristics of gradients, which provide us insightful clues for gradient quantization. Then, we theoretically give an in-depth analysis of the convergence bound and derive two principles for stable INT8 training. Finally, we propose two universal techniques, including Direction Sensitive Gradient Clipping that reduces the direction deviation of gradients and Deviation Counteractive Learning Rate Scaling that avoids illegal gradient update along the wrong direction. The experiments show that our unified solution promises accurate and efficient INT8 training for a variety of networks and tasks, including MobileNetV2, InceptionV3 and object detection that prior studies have never succeeded. Moreover, it enjoys a strong flexibility to run on off-the-shelf hardware, and reduces the training time by 22% on Pascal GPU without too much optimization effort. We believe that this pioneering study will help lead the community towards a fully unified INT8 training for convolutional neural networks. Feng Zhu 0006, Ruihao Gong, Fengwei Yu, Xianglong Liu 0001, Zhelong Li, Xiuqi Yang |
CVPR | 1 |
| 2020 | Self-supervising Fine-Grained Region Similarities for Large-Scale Image Localization
Yixiao Ge, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001 |
ECCV (4) | 3 |
| 2020 | Self-paced Contrastive Learning with Hybrid Memory for Domain Adaptive Object Re-IDabstractDomain adaptive object re-ID aims to transfer the learned knowledge from the labeled source domain to the unlabeled target domain to tackle the open-class re-identification problems. Although state-of-the-art pseudo-label-based methods have achieved great success, they did not make full use of all valuable information because of the domain gap and unsatisfying clustering performance. To solve these problems, we propose a novel self-paced contrastive learning framework with hybrid memory. The hybrid memory dynamically generates source-domain class-level, target-domain cluster-level and un-clustered instance-level supervisory signals for learning feature representations. Different from the conventional contrastive learning strategy, the proposed framework jointly distinguishes source-domain classes, and target-domain clusters and un-clustered instances. Most importantly, the proposed self-paced method gradually creates more reliable clusters to refine the hybrid memory and learning targets, and is shown to be the key to our outstanding performance. Our method outperforms state-of-the-arts on multiple domain adaptation tasks of object re-ID and even boosts the performance on the source domain without any extra annotations. Our generalized version on unsupervised object re-ID surpasses state-of-the-art algorithms by considerable 16.7% and 7.9% on Market-1501 and MSMT17 benchmarks. Yixiao Ge, Feng Zhu 0006, Dapeng Chen, Rui Zhao 0001, Hongsheng Li 0001 |
NeurIPS | 2 |
| 2018 | Attention-Aware Compositional Network for Person Re-IdentificationabstractPerson re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was largely improved in accuracy. It is reasonable to use pose estimation results for handling pose variations and background clutters, and such attempts have obtained great improvement in ReID performance. However, we argue that the pose information was not well utilized and hasn't yet been fully exploited for person ReID. In this work, we introduce a novel framework called Attention-Aware Compositional Network (AACN) for person ReID. AACN consists of two main components: Pose-guided Part Attention (PPA) and Attention-aware Feature Composition (AFC). PPA is learned and applied to mask out undesirable background features in pedestrian feature maps. Furthermore, pose-guided visibility scores are estimated for body parts to deal with part occlusion in the proposed AFC module. Extensive experiments with ablation analysis show the effectiveness of our method, and state-of-the-art results are achieved on several public datasets, including Market-1501, CUHK03, CUHK01, SenseReID, CUHK03-NP and DukeMTMC-reID. Rui Zhao 0001, Feng Zhu 0006, Huaming Wang, Wanli Ouyang |
CVPR | 3 |
| 2018 | Crowd Tracking by Group Structure EvolutionabstractWe propose a new model-free approach for crowd tracking that integrates low-level keypoint tracking, midlevel patch tracking, and high-level group evolution in one unified framework. Instead of computing optical flows, tracking keypoints, or pedestrians, we propose to represent the crowd as a set of distinctive and stable midlevel patches. These patches are tracked together through occlusions, background clutter, and appearance variations, with spatial relations modeled by the proposed hierarchical tree structure. In the low level, keypoint tracking provides accurate local motions, which guides the detection of midlevel patches with stable internal motions, and also organizes patches into hierarchical groups with collective motions. In the high level, group evolution guides updating of the proposed hierarchical tree structure through merge and split events. The dynamically structured patches not only substantially improve their own tracking, but also act as assistant patches that can help track given targets more accurately in a crowd. Extensive experiments on both ours and publicly available data sets show that our proposed approach significantly outperforms current state-of-the-art trackers. Feng Zhu 0006, Xiaogang Wang 0001, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Learning Spatial Regularization with Image-Level Supervisions for Multi-label Image ClassificationabstractMulti-label image classification is a fundamental but challenging task in computer vision. Great progress has been achieved by exploiting semantic relations between labels in recent years. However, conventional approaches are unable to model the underlying spatial relations between labels in multi-label images, because spatial annotations of the labels are generally not provided. In this paper, we propose a unified deep neural network that exploits both semantic and spatial relations between labels with only image-level supervisions. Given a multi-label image, our proposed Spatial Regularization Network (SRN) generates attention maps for all labels and captures the underlying relations between them via learnable convolutions. By aggregating the regularized classification results with original results by a ResNet-101 network, the classification performance can be consistently improved. The whole deep neural network is trained end-to-end with only image-level annotations, thus requires no additional efforts on image annotations. Extensive evaluations on 3 public datasets with different types of labels show that our approach significantly outperforms state-of-the-arts and has strong generalization capability. Analysis of the learned SRN model demonstrates that it can effectively capture both semantic and spatial relations of labels for improving classification performance. Feng Zhu 0006, Hongsheng Li 0001, Wanli Ouyang, Nenghai Yu, Xiaogang Wang 0001 |
CVPR | 1 |
| 2016 | Consistent matching based on boosted salience channels for group re-identificationabstractAssociating groups of people across non-overlapping camera views is an important but unsolved problem. Compared with the similar person re-identification task, group re-identification introduces some new challenges, such as significant deformation in uncontrolled directions, great intra-group occlusions and so on. In this paper, we propose a novel patch matching based framework for group re-identification. Discriminative salience channels are learned to filter out highly unreliable and non-informative patch matches between two group images, while retain true matches undergoing appearance variations. The resulting candidate correspondences are further explored by the proposed consistent matching process, which prefers coherent matches in true group image pairs. The effectiveness of our approach is validated on two group re-identification datasets: ZeCSS and i-LIDS MCTS. It outperforms state-of-the-art methods on both datasets. Feng Zhu 0006, Qi Chu 0001, Nenghai Yu |
ICIP | 1 |
| 2016 | Multi-level visual tracking with hierarchical tree structural constraint
Jingjing Wang 0005, Nenghai Yu, Feng Zhu 0006, Liansheng Zhuang |
Neurocomputing | 3 |
| 2014 | Crowd Tracking with Dynamic Evolution of Group Structures
Feng Zhu 0006, Xiaogang Wang 0001, Nenghai Yu |
ECCV (6) | 1 |
| 2011 | Error Resilient Coding Based on Reversible Data Hiding and Redundant SliceabstractCompressed video streams are sensitive to errors and losses when transmitted over wireless error-prone channels. In this paper, we propose an Error Resilient (ER) scheme based on Reversible Data Hiding (RDH) and Redundant Slice (RDH-RS). Reversible data hiding is quite effective in unequal protection and can provide satisfactory protection to vital information, while redundant slice works well in protecting mass important data. Our scheme exploits both advantages. At the encoder side, we apply RS protection to odd frames, and embed vital data, the motion vectors (MV) of even frames, into the redundant slice of odd frames by bidirectional-RDH method, thus providing protection to even frames. If an MV of an odd frame cannot be correctly decoded at the decoder side, the redundant slice of odd frame will be utilized for error concealment. If the lost MV belong to an even frame's macro block (MB), then the MV will be retrieved from the redundant slice of prior odd frame and help restoration. As our data hiding method is reversible, no extra visual quality degradation will happen, and the computational burden is quite low. Experimental results demonstrate that RDH-RS method can provide much better protection to fragile compressed video when compared with the previous arts. Weiming Zhang 0001, Nenghai Yu, Feng Zhu 0006 |
ICIG | 4 |