EDBT 2026 Demo / reviewers in the wild / expert
Yali Li 0001
dblp:05/1013-1 · also Ya-Li Li 0001
· DBLP profile ↗
89ranked-venue papers
9as first author
55since 2021 · last 2026
0000-0002-6629-7228ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 5 first-author · 39 since 2021Artificial intelligence and machine learning · 49 · 5 first-author · 34 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-Background Decoupling for Domain-Generalizable Person Re-identification
Yali Li 0001, Shengjin Wang |
ICIC (12) | 2 |
| 2026 | HySpeFAS: A Hyperspectral Face Anti-Spoofing Dataset Based on Snapshot Compressive ImagingabstractFace anti-spoofing, which aims to prevent the attacks of widely-used face recognition systems, is highly related to personal privacy and property security. However, existing benchmarks on face anti-spoofing mainly focus on RGB images, further challenged by consistently developed 3D high-fidelity (HiFi) masks. To facilitate the research on multimodal face anti-spoofing, we construct the HyperSpectral Face Anti-Spoofing (HySpeFAS) dataset. We introduce the newly-developed snapshot spectral imaging (SSI) technology to capture real and spoof faces, as well as identify unknown HiFi masks. Specifically, hyperspectral images (HSIs) acquired by SSI sensor contain rich information about the chemical composition of the targets, which can be used to effectively distinguish live human skin and various spoof materials. The HySpeFAS dataset contains 22,368 multimodal images (i.e., RGB, SSI, HSI) of 17 live subjects and 60 spoof subjects. Moreover, extensive experiments with baseline deep learning models validate the special features of the SSI images and the potential of SSI in FAS. By publishing the dataset as well as the baseline models, we encourage the community to foster the algorithm study associated with hyperspectral images and the development of SSI-equipped intelligent systems. Shijie Rao, Yidong Huang, Xueqian Zhang, Ajian Liu 0001, Jun Wan 0001, Kaiyu Cui, Yali Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2026 | Unsupervised Temporal Correspondence Learning for Unified Video Object RemovalabstractVideo object removal aims at erasing a target object in the entire video and filling holes with plausible contents, given an object mask in the first frame as input. Existing solutions mostly break down the task into (supervised) mask tracking and (self-supervised) video completion, and then separately tackle them with tailored designs. In this paper, we introduce a new setup, coined as unified video object removal, where mask tracking and completion are addressed within a unified framework. Despite introducing more challenges, the setup is promising for future practical usage. We embrace the observation that these two sub-tasks have strong inherent connections in terms of pixel-level temporal correspondence. Making full use of the connections could be beneficial considering the complexity of both algorithm and deployment. We propose a single network linking the two sub-tasks by inferring temporal correspondences across multiple frames, i.e., correspondences between valid-valid (V-V) pixel pairs for mask tracking and correspondences between valid-hole (V-H) pixel pairs for video completion. Thanks to the unified setup, the network can be learned end-to-end in a totally unsupervised fashion without any annotations. We demonstrate that our method can generate visually pleasing results and perform favorably against existing separate solutions in realistic test cases. Zhongdao Wang, Jinglu Wang, Xiao Li 0030, Yali Li 0001, Yan Lu 0001, Shengjin Wang |
IEEE Trans. Image Process. | 4 |
| 2025 | LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual Groundingabstract3D Vision Grounding (3D-VG) seeks to unravel referential language and identify targets in 3D physical world. Prevailing methods align with the 2D-VG's pipeline to pinpoint the referred object in a categorical multi-modal reasoning manner. However, the geometric complexities of 3D scenes and the nuanced syntactic structures of language, exacerbates the \textbf{granularity inconsistency} of point cloud and text features, hindering the development of 3D-VG systems in complex scenarios. Towards this issue, we propose LIBA, a Language-Instructed multi-granularity Bridge Assistant tailored for 3D-VG task. LIBA tackles this issue as follows. (1) \textit{How to establish a multi-granularity 3D vision-text feature alignment in a unified model}? We advance a bilateral Dynamic Bridge Adapter (DBA) build multi-granularity interaction of 3D vision and language backnones during feature extraction. We further develop the Language-aware Cross-scale Object Modulation (LCOM) module to integrate multi-scale point cloud features modulated by language information. (2) After aligning multi-modal features, \textit{how to fully harness language model's knowledge to bolster vision concepts understanding}? A LLM-guided Hierarchical Query Selection (LLM-HQS) module incorporates world knowledge of Large Language Model~(LLM) to ground the target referral via an Attribute-then-Relation reasoning process. In this manner, our LIBA inherits reasoning prowess and world knowledge of LLM to bridge point clouds and texts at multiple granularities. Experiments on ScanRefer and Nr3D/Sr3D benchmarks substantiate the superiority of our LIBA, trumping state-of-the-arts by a considerable margin. Yali Li 0001, Eastman Z. Y. Wu, Shengjin Wang |
AAAI | 2 |
| 2025 | HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionabstractWhile flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored landscape. Current HSI methods naively rely on conditional Variational AutoEncoder (cVAE) and diffusion models. They are typically associated with limited modalities of control signals and task-specific frameworks design, leading to inflexible adaptation across various interaction scenarios and descriptive-unfaithful motions in diverse 3D physical environments. In this paper, we propose HSI-GPT, a General-Purpose Large Scene-Motion-Language Model that applies "next-token prediction" paradigm of Large Language Models to the HSI domain. HSI-GPT not only exhibits remarkable flexibility to accommodate diverse control signals (3D scenes, textual commands, key-frame poses, as well as scene affordances), but it seamlessly supports various HSI-related tasks (e.g., multi-modal controlled HSI generation, HSI understanding, and general motion completion in 3D scenes). First, HSI-GPT quantizes textual descriptions and human motions into discrete, LLM-interpretable tokens with multi-modal tokenizers. Inspired by multi-modal learning, we develop a recipe for aligning mixed-modality tokens into the shared embedding space of LLMs. These interaction tokens are then organized into unified instruction following prompts, allowing HSI-GPT to fine-tune on question-and-answer tasks. Extensive experiments and visualizations validate that our general-purpose HSI-GPT model delivers exceptional performance across multiple HSI-related tasks. Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2025 | Attention Augmented Structure-centric Bias Mitigation with Feature DisentanglementabstractImage classification models often rely on superficial visual features, such as textures or colors, leading to undesired bias. This can compromise the robustness and reliability of deep models, particularly their performance on out-of-distribution (o.o.d.) datasets. Existing approaches, focusing on data-centric aspects, typically predefine specific bias types to mitigate the impact of these superficial features. However, such data-centric methods may lack extensibility due to their focus on predefined biases. In this paper, we propose an attention augmented structure-centric bias mitigation method, considering network architecture can be flexibly manipulated to address a variety of visual features. This method captures global semantic representations by integrating the strengths of both self-attention and convolution, introducing a global receptive field to Convolutional Neural Networks. By incorporating feature disentanglement and augmentation, our concise network demonstrates improved performance as feature diversity increases in the latent space. Our method achieves state-of-the-art results on synthetic datasets (Colored MNIST and Corrupted CIFAR10) and shows impressive performance on challenging real-world datasets (ImageNet and BFFHQ), with improvements of about 2%-5% across different subsets. Xuege Hou, Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2025 | Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video UnderstandingabstractThe Large Vision-Language Model (LVLM) has achieved impressive performance in the field of visual-language understanding. However, its ability to understand longer videos is still limited due to the length and information diversity of multi-modal videos. Moreover, accurately matching detailed content within videos remains an open research problem. We design a new framework for LVLM inference, "Tower of Thoughts" (ToT), which extends the "Chain-of-Thought" (CoT) approach to the visual domain and constructs the high-dimensional semantics of the complete videos from the bottom up. Meanwhile, to achieve question-answering for video details within the constraints of the restricted context window, we propose a method of self-retrieval augmented generation (SRAG), which makes it possible to obtain details from long videos by storing and accessing video text as dense vectors in non-parametric memory. The solution of combining the ToT with SRAG enables our model to have cross-modal high-density semantic fusion and comprehensive and accurate generation capabilities, thereby achieving rationalized video answers. Experiments on public benchmarks demonstrate the effectiveness of our proposed method. In addition, we also conducted experiments on multi-modal long videos in the open world and achieved remarkable outcomes. These results provide new perspectives and technical routes for the future development of visual language models. Tong Yue, Mingrui Xiao, Dafeng Zhang, Yali Li 0001, Shengjin Wang |
ICASSP | 5 |
| 2025 | Dynamic Object Queries for Transformer-based Incremental Object DetectionabstractIncremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capacity and increasing knowledge. In this paper, we propose the Dynamic object Query-based DEtection TRansformer (DyQ-DETR), which incrementally expands the model representation ability to achieve stability-plasticity tradeoff. First, a new set of learnable object queries are fed into the decoder to represent new classes. Second, we propose the isolated bipartite matching for object queries in different phases, based on disentangled self-attention. Thanks to the separate supervision and computation over object queries, we further present the risk-balanced partial calibration for effective exemplar replay. Extensive experiments demonstrate that DyQ-DETR significantly surpasses the state-of-the-art methods, with limited parameter overhead. The code is available at https://github.com/THUzhangjic/DyQ-DETR. Jichuan Zhang, Wei Li 0110, Shuang Cheng, Yali Li 0001, Shengjin Wang |
ICASSP | 4 |
| 2025 | UPL-Net: Uncertainty-aware prompt learning network for semi-supervised action recognition
Shu Yang 0007, Yali Li 0001, Shengjin Wang |
Neurocomputing | 2 |
| 2025 | UniDetector: Towards Universal Object Detection With Heterogeneous SupervisionabstractIn this paper, we formally address universal object detection, which aims to detect every category in every scene. The dependence on human annotations, the limited visual information, and the novel categories in open world severely restrict the universality of detectors. We propose UniDetector, a universal object detector that recognizes enormous categories in the open world. The critical points for UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces in training through image-text alignment, which guarantees sufficient information for universal representations. 2) it involves heterogeneous supervision training, which alleviates the dependence on the limited fully-labeled images. 3) it generalizes to open world easily while keeping the balance between seen and unseen classes. 4) it further promotes generalizing to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7 k categories, the largest measurable size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot ability on large-vocabulary datasets - it surpasses supervised baselines by more than 5% without seeing any corresponding images. On 13 detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | G3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual GroundingabstractGrounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI, as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counterparts, challenges posed by the variability of 3D visual grounding remain relatively unsolved in existing studies: 1) the underlying geometric and complex spatial relationships in 3D scene. 2) the inherent complexity of 3D grounded language. 3) the inconsistencies between text and geometric features. To tackle these issues, we propose G3-LQ, a DEtection TRansformer-based model tailored for 3D visual grounding task. G3-LQ explicitly models Geometric-aware visual rep-resentations and Generates fine-Grained Language-guided object Queries in an overarching framework, which com-prises two dedicated modules. Specifically, the Position Adaptive Geometric Exploring (PAGE) unearths underlying information of 3D objects in the geometric details and spatial relationships perspectives. The Fine-grained Language-guided Query Selection (Flan-QS) delves into syntactic structure of texts and generates object queries that exhibit higher relevance towards fine-grained text features. Finally, a pioneering Poincaré Semantic Alignment (PSA) loss establishes semantic-geometry consistencies by modeling non-linear vision-text feature mappings and aligning them on a hyperbolic prototype-Poincaré ball. Extensive experiments verify the superiority of our G3-LQ method, trumping the state-of-the-arts by a considerable margin. Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2024 | Exploring Pose-Aware Human-Object Interaction via Hybrid LearningabstractHuman-Object Interaction (HOI) detection plays a crucial role in visual scene comprehension. In recent advancements, two-stage detectors have taken a prominent position. However, they are encumbered by two primary challenges. First, the misalignment between feature representation and relation reasoning gives rise to a deficiency in discrimi-native features crucial for interaction detection. Second, due to sparse annotation, the second-stage interaction head generates numerous candidatepairs, with only a small fraction receiving supervision. Towards these issues, we propose a hybrid learning method based on pose-aware HOI feature refinement. Specifically, we de-vise pose-aware feature refinement that encodes spatial fea-tures by considering human body pose characteristics. It can direct attention towards key regions, ultimately offering a wealth of fine-grained features imperative for HOI de-tection. Further, we introduce a hybrid learning method that combines HOI triplets with probabilistic soft labels supervision, which is regenerated from decoupled verb-object pairs. This method explores the implicit connections between the interactions, enhancing model generalization without requiring additional data. Our method establishes state-of-the-art performance on HICO-DET benchmark and excels notably in detecting rare HOIs. Eastman Z. Y. Wu, Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2024 | Risk-Aware Self-consistent Imitation Learning for Trajectory Planning in Autonomous Driving
Yixuan Fan, Yali Li 0001, Shengjin Wang |
ECCV (13) | 2 |
| 2024 | OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
Zhenyu Wang 0005, Yali Li 0001, Taichi Liu, Hengshuang Zhao, Shengjin Wang |
ECCV (47) | 2 |
| 2024 | Learning Generalizable Visual Representations via Self-Supervised Information BottleneckabstractNumerous approaches have recently emerged in the realm of self-supervised visual representation learning. While these methods have demonstrated empirical success, a theoretical foundation that understands and unifies these diverse techniques remains to be established. In this work, we draw inspiration from the principles underlying brain-based learning and propose a new method named self-supervised information bottleneck. Our method aims to maximize the mutual information between representations of views derived from the same image, while maintaining a minimal mutual information between the view and its corresponding representation at the same time. The brain-inspired method provides a unified information-theoretic perspective on various self-supervised approaches. This unified framework also empowers the model to learn generalizable visual representations for diverse downstream tasks and data distributions, achieving state-of-the-art performance across a wide variety of image and video tasks. Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2024 | Representation Distillation for Efficient Self-Supervised LearningabstractSiamese self-supervised learning has shown significant progress recently, which relies on Siamese networks with identical encoders in the two branches. However, due to this inherent design of Siamese networks, the overall model capacity is primarily constrained by the encoder of interest, resulting in the representation bottleneck problem during pre-training. To address this limitation, we propose a new Distill Your Own Latent (DYOL) method that can perform self-supervised learning between branches with different architectures. So a larger target network can be employed to provide stronger self-supervision. We first decouple the update process of the target network from the online network to prevent shortcut learning. Then we distill the representation directly from the target network into the online network by enforcing the view consistency between networks. Extensive experiments on various downstream tasks validate the effectiveness of our method. Importantly, the results demonstrate that strong target networks are efficient self-supervised distillers, which enable small online networks to attain similar results to large target networks (parameter efficiency) and achieve superior performance with a much smaller number of pretraining epochs and samples (time and data efficiency). Yali Li 0001, Shengjin Wang |
ICME | 2 |
| 2024 | A Dataset with Multi-Modal Information and Multi-Granularity Descriptions for Video CaptioningabstractVideo captioning aims to generate natural language descriptions automatically from videos. While datasets like MSVD and MSR-VTT have driven research in recent years, they predominantly focus on visual features and describe simple actions, ignoring audio, text, and other modal information. Which, however, is limited, because multi-modal information plays an important role in generating accurate captions. In this study, we introduce a dataset, News-11k, which includes over 150,000 captions with multi-modal information from more than 11,000 selected news video clips. We annotate multi-granularity captions from three perspectives: coarse-grained, medium-grained, and fine-grained captions. Due to the characteristics of news videos, generating accurate captions on our dataset requires multi-modal understanding ability. Therefore, we propose a baseline model for multi-modal video captioning. To address the challenge of multi-modal information fusion, we devise the concatenating modal embedding strategy. Experiments indicate that multi-modal information significantly enhances the understanding of the deeper semantics in videos. Data will be made available on https://github.com/David-Zeng-Zijian/News-11k. Mingrui Xiao, Shu Yang 0007, Yali Li 0001, Shengjin Wang |
ICME | 5 |
| 2024 | Transformer-Based Few-Shot Object Detection with Enhanced Fine-Tuning StabilityabstractFew-shot object detection (FSOD) aims at detecting unseen classes with limited annotated novel examples. while the fine-tuning paradigm has been proven effective for FSOD, most existing approaches are based on Faster R-CNN framework, neglecting exploration of advanced frameworks such as DETR. In this paper, we propose a transformer-based few-shot object detection method. To address the instability issue in few-shot fine-tuning of DETR, we propose a simple yet effective method to enhance the fine-tuning stability from two perspectives: 1) an enhanced initialization method for classification layer, hollow initialization. 2) Category-Level Dropout (CLD), to mitigate effect of missing annotations by controlling gradient backpropagation of classifier. Comprehensive experimental results on MS COCO show that our method significantly improves the baseline and achieves very promising performance. Zuyu Chen, Yali Li 0001, Shengjin Wang |
IJCNN | 2 |
| 2024 | SSINet: Snapshot Spectral Imaging Method with Matrix Fusion Module for Face Anti-spoofingabstractSnapshot spectral imaging technology based on the principle of compressed sensing can introduce additional spectral information and solve the problem of inadequate performance of existing optical sensors in dealing with high-fidelity masks. However, using compressed sensing for spectral image reconstruction requires a large amount of computational resources and has limited accuracy, making it the greatest obstacle for the practical application of snapshot spectral imaging technology. In this paper, we propose a new paradigm for Face Anti-Spoofing (FAS) based on snapshot spectral imaging technology. Our method does not require the reconstruction of compressive sensing spectral images, but instead directly treats the snapshot spectral imaging hardware as an optical encoder and designs the end-to-end Snapshot Spectral Imaging Network (SSINet) to decode the optical encoding results to obtain FAS results, which constitutes a software-hardware integrated FAS method. Experimental results show that our method not only surpasses previous state-of-the-art methods based on RGB images, but also performs the better than conventional spectral reconstruction methods in both efficiency and accuracy, with a nearly thousand-fold increase in FAS speed and an Average Classification Error Rate (ACER) decrease of 12.29%. Xueqian Zhang, Yidong Huang, Shijie Rao, Kaiyu Cui, Yali Li 0001 |
IJCNN | 5 |
| 2024 | One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object DetectionabstractThe current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however, such multi-domain joint training is highly challenging, because large domain gaps among point clouds from different datasets lead to the severe domain-interference problem. In this paper, we propose OneDet3D, a universal one-for-all model that addresses 3D detection across different domains, including diverse indoor and outdoor scenes, within the same framework and only one set of parameters. We propose the domain-aware partitioning in scatter and context, guided by a routing mechanism, to address the data interference issue, and further incorporate the text modality for a language-guided classification to unify the multi-dataset label spaces and mitigate the category interference issue. The fully sparse structure and anchor-free head further accommodate point clouds with significant scale disparities. Extensive experiments demonstrate the strong universal ability of OneDet3D to utilize only one trained model for addressing almost all 3D object detection tasks (Fig. 1). We will open-source the code for future research and applications. Zhenyu Wang 0005, Yali Li 0001, Hengshuang Zhao, Shengjin Wang |
NeurIPS | 2 |
| 2023 | Detecting Everything in the Open World: Towards Universal Object DetectionabstractIn this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We propose UniDetector, a universal object detector that has the ability to recognize enormous categories in the open world. The critical points for the universality of UniDetector are: 1) it leverages images of multiple sources and heterogeneous label spaces for training through the alignment of image and text spaces, which guarantees sufficient information for universal representations. 2) it generalizes to the open world easily while keeping the balance between seen and unseen classes, thanks to abundant information from both vision and language modalities. 3) it further promotes the generalization ability to novel categories through our proposed decoupling training manner and probability calibration. These contributions allow UniDetector to detect over 7k categories, the largest measurable category size so far, with only about 500 classes participating in training. Our UniDetector behaves the strong zero-shot generalization ability on largevocabulary datasets - it surpasses the traditional supervised baselines by more than 4% on average without seeing any corresponding images. On 13 public detection datasets with various scenes, UniDetector also achieves state-of-the-art performance with only a 3% amount of training data.11Codes are available at https://github.com/zhenyuw16/UniDetector. Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Ser-Nam Lim, Antonio Torralba 0001, Hengshuang Zhao, Shengjin Wang |
CVPR | 2 |
| 2023 | Divcon: Learning Concept Sequences for Semantically Diverse Image CaptioningabstractHuman generated image captions contain diverse semantic concepts, while this is still a difficult task for machines. The frequency distribution of semantic concepts in datasets is usually extremely imbalanced, leading to models repeatedly describe frequently occurring semantic concepts, resulting in a decline in the semantic diversity. In this paper, we propose a novel two-step method for diverse image captioning, generating descriptions with more diverse semantic concepts (Di-vCon). Firstly, we developed a concept sequence generator to auto-regressively generate concept sequences. This benefits the model by decoding sequences in a small searching space. Then a sentence generator takes as input the concept sequences and generates descriptions for each sequence. Experiments show that DivCon can generate captions containing diverse semantic concepts and pay more attention to the less occurring concepts. In the diverse image captioning task, Div-Con achieves the state-of-the-art results on MSCOCO dataset with oracle CIDEr and SPICE scores of 1.684 and 0.302. Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2023 | Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-identificationabstractThis paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the barriers of data and annotation, we propose to utilize large-scale unsupervised data for training. The key issue lies in how to mine identity information. To this end, we propose an Identity-seeking Self-supervised Representation learning (ISR) method. ISR constructs positive pairs from inter-frame images by modeling the instance association as a maximum-weight bipartite matching problem. A reliability-guided contrastive loss is further presented to suppress the adverse impact of noisy positive pairs, ensuring that reliable positive pairs dominate the learning process. The training cost of ISR scales approximately linearly with the data size, making it feasible to utilize large-scale data for training. The learned representation exhibits superior generalization ability. Without human annotation and fine-tuning, ISR achieves 87.0% Rank-1 on Market-1501 and 56.4% Rank-1 on MSMT17, outperforming the best supervised domain-generalizable method by 5.0% and 19.5%, respectively. In the pre-training→fine-tuning scenario, ISR achieves state-of-the-art performance, with 88.4% Rank-1 on MSMT17. The code is at https://github.com/dcp15/ISR_ICCV2023_Oral. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICCV | 3 |
| 2023 | RoICLIP: Text-Enhanced UAV-Based Video Object Detection
Yali Li 0001, Shengjin Wang |
ICIG (4) | 2 |
| 2023 | Feature Decoupling and Uncertainty Estimation for 3D Object DetectionabstractIn the real scene of 3D object detection, the point cloud collected for a single object is incomplete, resulting in misalignment of classification and regression features and uncertainty of object boundaries. Existing works pay little attention to the uncertainty caused by the incompleteness of point clouds. To address this issue, we propose a Feature Decoupling and Uncertainty Estimation single-stage 3D object detector named FDUE-Net. First, we design a Classification-Regression Attention Decoupling module to extract shared low-level and high-level features in a decoupling paradigm, generating the feature maps that are more helpful for classification or regression by layer-level and space-level attention modules. Furthermore, we propose an Uncertainty Estimation head (UE-head), which improves the quality of the predicted bounding boxes by modeling the regression values as general distributions, and uses the predicted distributions to correct the classification scores. Experiments on the KITTI dataset show that the proposed method achieves significant improvement, with 3D detection performance improved by 1.46% on the moderate set compared to the baseline, and competitive performance compared to the state-of-the-arts. Peiyuan Zhi, Kaiyue Zhou, Yali Li 0001, Shengjin Wang |
ICME | 3 |
| 2023 | Look Before You Drive: Boosting Trajectory Forecasting via Imagining FutureabstractPredicting the future trajectories of other agents in the scene fast and effectively is crucial for autonomous driving systems. We note that high-quality predictions require us to take into account the subjective initiative of the target agents, which is reflected by the fact that they themselves make decisions based on their own predictions about the future, just like our ego vehicle's prediction-planning system. However, this characteristic has been neglected in previous studies. We introduce Look Before You Drive (LBYD), a two-stage approach that explicitly incorporates both past observations and future estimates to make predictions. To get a preliminary estimate of the future, we propose a neat and effective baseline capable of making predictions for multiple agents simultaneously. We use only the most basic structures, mainly Transformer, to ensure sufficient inference speed and room for expansion. On this basis, we cooperatively train two networks to enable the coarse estimates to boost final forecasting. Our experiments demonstrate that LBYD can significantly surpass the baseline performance. Moreover, while state-of-the-art methods rely on considering heterogeneity and artificially designed inductive biases for attention modeling, LBYD performs on par with SOTA without them on both the Argoverse 1 and the large scale Argoverse 2 datasets, and can run at 67 FPS on an RTX 3090 GPU. Yixuan Fan, Yali Li 0001, Shengjin Wang |
IROS | 3 |
| 2023 | VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor ScenesabstractRobotic grasping faces new challenges in human-robot-interaction scenarios. We consider the task that the robot grasps a target object designated by human's language directives. The robot not only needs to locate a target based on vision-and-language information, but also needs to predict the reasonable grasp pose candidate at various views and postures. In this work, we propose a novel interactive grasp policy, named Visual-Lingual-Grasp (VL-Grasp), to grasp the target specified by human language. First, we build a new challenging visual grounding dataset to provide functional training data for robotic interactive perception in indoor environments. Second, we propose a 6- Dof interactive grasp policy combined with visual grounding and 6- Dof grasp pose detection to extend the universality of interactive grasping. Third, we design a grasp pose filter module to enhance the performance of the policy. Experiments demonstrate the effectiveness and extendibility of the VL-Grasp in real world. The VL-Grasp achieves a success rate of 72.5 % in different indoor scenes. The code and dataset is available at https://github.com/luyh20/VL-Grasp. Yuhao Lu, Yixuan Fan, Beixing Deng, Fangfu Liu, Yali Li 0001, Shengjin Wang |
IROS | 5 |
| 2023 | Uni3DETR: Unified 3D Detection TransformerabstractExisting point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, there is still a lack of a unified network architecture that can accommodate diverse scenes. In this paper, we propose Uni3DETR, a unified 3D detector that addresses indoor and outdoor 3D detection within the same framework. Specifically, we employ the detection transformer with point-voxel interaction for object prediction, which leverages voxel features and points for cross-attention and behaves resistant to the discrepancies from data. We then propose the mixture of query points, which sufficiently exploits global information for dense small-range indoor scenes and local information for large-range sparse outdoor ones. Furthermore, our proposed decoupled IoU provides an easy-to-optimize training target for localization by disentangling the $xy$ and $z$ space. Extensive experiments validate that Uni3DETR exhibits excellent performance consistently on both indoor and outdoor 3D detection. In contrast to previous specialized detectors, which may perform well on some particular datasets but suffer a substantial degradation on different scenes, Uni3DETR demonstrates the strong generalization ability under heterogeneous conditions (Fig. 1). Zhenyu Wang 0005, Yali Li 0001, Xi Chen 0119, Hengshuang Zhao, Shengjin Wang |
NeurIPS | 2 |
| 2023 | A prompt tuning method for few-shot action recognitionabstractVision-language pre-training models learn visual concepts from image-text or video-text pairs, which can be adopted for visual-textual tasks. In this paper, we adopt these concepts as prior knowledge to solve the unreliable problem of minimizing the loss of limited training samples in few-shot action recognition tasks. In particular, a two-stage framework of vision-language pre-training and prompt tuning is designed. In the pre-training stage, multi-modal encoding models are jointly trained on video-text pairs to learn the semantic correspondence between video and text. In the prompt tuning stage, a prompt module with instance-level bias is trained on a few video samples to utilize the pre-trained concepts for the classification task. The experimental results show that the proposed method is superior to the baseline and state-of-the-art few-shot action recognition methods on two public video benchmarks. Shu Yang 0007, Yali Li 0001, Shengjin Wang |
VCIP | 2 |
| 2023 | AdaZoom: Towards Scale-Aware Large Scene Object DetectionabstractDetection in large scenes is a challenging issue due to small objects and extreme scale variation. It is difficult for the deep-learning-based detector to extract features of small objects with only a few pixels. Most existing methods employ image pyramid and feature pyramid for multi-scale inference to alleviate this issue. However, they lack scale awareness to adapt to objects with different scales. In this paper, we propose a novel Adaptive Zoom (AdaZoom) network for scale-aware large scene object detection. There are three main contributions. First, an Adaptive Zoom network is proposed to actively focus and adaptively zoom the focused regions for high-performance object detection in large scenes. Second, to tackle the problem of missing annotations for focused regions, we train AdaZoom with the reward which measures the quality of generated regions, based on the paradigm of deep reinforcement learning. At last, we propose collaborative training to iteratively promote the joint performance of AdaZoom and the detector. To validate the effectiveness, we conduct extensive experiments on VisDrone2019, UAVDT and DOTA datasets. The experiments show AdaZoom brings consistent and significant improvement over different detection networks, achieving state-of-the-art performance on these datasets, especially outperforming the existing methods by AP of 4.64% on VisDrone2019. Jingtao Xu, Yali Li 0001, Shengjin Wang |
IEEE Trans. Multim. | 2 |
| 2022 | Delving into Probabilistic Uncertainty for Unsupervised Domain Adaptive Person Re-identificationabstractClustering-based unsupervised domain adaptive (UDA) person re-identification (ReID) reduces exhaustive annotations. However, owing to unsatisfactory feature embedding and imperfect clustering, pseudo labels for target domain data inherently contain an unknown proportion of wrong ones, which would mislead feature learning. In this paper, we propose an approach named probabilistic uncertainty guided progressive label refinery (P2LR) for domain adaptive person re-identification. First, we propose to model the labeling uncertainty with the probabilistic distance along with ideal single-peak distributions. A quantitative criterion is established to measure the uncertainty of pseudo labels and facilitate the network training. Second, we explore a progressive strategy for refining pseudo labels. With the uncertainty-guided alternative optimization, we balance between the exploration of target domain data and the negative effects of noisy labeling. On top of a strong baseline, we obtain significant improvements and achieve the state-of-the-art performance on four UDA ReID benchmarks. Specifically, our method outperforms the baseline by 6.5% mAP on the Duke2Market task, while surpassing the state-of-the-art method by 2.5% mAP on the Market2MSMT task. Code is available at: https://github.com/JeyesHan/P2LR. Yali Li 0001, Shengjin Wang |
AAAI | 2 |
| 2022 | Disentangling based Environment-Robust Feature Learning for Person ReID
Yali Li 0001, Shengjin Wang |
BMVC | 2 |
| 2022 | Polishing Network for Decoding of Higher-Quality Diverse Image Captions
Yali Li 0001, Shengjin Wang |
BMVC | 2 |
| 2022 | R(Det)2: Randomized Decision Routing for Object DetectionabstractIn the paradigm of object detection, the decision head is an important part, which affects detection performance significantly. Yet how to design a high-performance decision head remains to be an open issue. In this paper, we propose a novel approach to combine decision trees and deep neural networks in an end-to-end learning manner for object detection. First, we disentangle the decision choices and prediction values by plugging soft decision trees into neural networks. To facilitate effective learning, we propose randomized decision routing with node selective and associative losses, which can boost the feature representative learning and network decision simultaneously. Second, we develop the decision head for object detection with narrow branches to generate the routing probabilities and masks, for the purpose of obtaining divergent decisions from different nodes. We name this approach as the randomized decision routing for object detection, abbreviated as R(Det)2. Experiments on MS-COCO dataset demonstrate that R(Det)2 is effective to improve the detection performance. Equipped with existing detectors, it achieves 1.4 ~ 3.6% AP improvement. Yali Li 0001, Shengjin Wang |
CVPR | 1 |
| 2022 | OSKDet: Orientation-sensitive Keypoint Localization for Rotated Object DetectionabstractRotated object detection is a challenging issue in computer vision field. Inadequate rotated representation and the confusion of parametric regression have been the bottleneck for high performance rotated detection. In this paper, we propose an orientation-sensitive keypoint based rotated detector OSKDet. First, we adopt a set of keypoints to represent the target and predict the keypoint heatmap on ROI to get the rotated box. By proposing the orientation-sensitive heatmap, OSKDet could learn the shape and direction of rotated target implicitly and has stronger modeling capabilities for rotated representation, which improves the localization accuracy and acquires high quality detection results. Second, we explore a new unordered keypoint representation paradigm, which could avoid the confusion of keypoint regression caused by rule based ordering. Further-more, we propose a localization quality uncertainty module to better predict the classification score by the distribution uncertainty of keypoints heatmap. Experimental results on several public benchmarks show the state-of-the-art performance of OSKDet. Specifically, we achieve an AP of 80.91% on DOTA, 89.98% on HRSC2016, 97.27% on UCAS-AOD, and a F-measure of 92.18% on ICDAR2015, 81.43% on ICDAR2017, respectively. Dongchen Lu, Yali Li 0001, Shengjin Wang |
CVPR | 3 |
| 2022 | Noisy Boundaries: Lemon or Lemonade for Semi-supervised Instance Segmentation?abstractCurrent instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employed to boost the performance. We construct a framework for semi-supervised instance segmentation by assigning pixel-level pseudo labels. Under this framework, we point out that noisy boundaries associated with pseudo labels are double-edged. We propose to exploit and resist them in a unified manner simultaneously: 1) To combat the negative effects of noisy boundaries, we propose a noise-tolerant mask head by leveraging low-resolution features. 2) To enhance the positive impacts, we introduce a boundary-preserving map for learning detailed information within boundary-relevant regions. We evaluate our approach by extensive experiments. It behaves extraordinarily, outperforming the supervised baseline by a large margin, more than 6% on Cityscapes, 7% on COCO and 4.5% on BDD100k. On Cityscapes, our method achieves comparable performance by utilizing only 30% labeled images. Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2022 | Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval
Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ECCV (14) | 4 |
| 2022 | GraphCSPN: Geometry-Aware Depth Completion via Dynamic GCNs
Xiaofei Shao, Yali Li 0001, Shengjin Wang |
ECCV (33) | 4 |
| 2022 | Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-IdentificationabstractPerson re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularity Retrieval (PGR) method to tackle this issue. Specifically, (i) we define instance-level, part-level and pixel-level features for an image. PGR learns these features by a single feature extractor to capture hierarchical clues in the image. (ii) These features are inherently related but different in perceptual granularity, and they can provide complementary information. For each type of feature, we propose a corresponding similarity metric to achieve hierarchical feature alignment. (iii) In training, we learn the model end-to-end. In inference, a progressive retrieval strategy is introduced to efficiently aggregate the complementary information provided by these features. Extensive experiments on three bench-marks of both occluded and holistic-body re-ID tasks show the effectiveness of the proposed method. Especially, our method significantly outperforms state-of-the-art by 4.5% Rank-1 score on the challenging Occluded-Duke dataset. Zhaopeng Dou, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
ICASSP | 3 |
| 2022 | Few-Shot Object Detection with Local Correspondence RPN and Attentive HeadabstractExisting object detection methods rely heavily on a large number of annotated bounding boxes, which is expensive to collect. In this paper, we propose a novel few-shot object detection method named GCN-FSOD. Intending to find informal local correspondence to fully explore cues of novel classes, we propose the local correspondence region proposal network (lcRPN) and the attentive detection head for few-shot detection. Taking features from the support-query image pair as inputs, lcRPN generates region proposals by mining fine-grained local correspondence with the help of GCNs. Then the proposed attentive head performs precise detection. We conduct extensive experiments on the wildly adopted MS-COCO benchmark. The proposed GCN-FSOD brings significant performance gains and outperforms the state-of-the-art by a large margin (1.7% mAP for 10-shot). Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2022 | CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) has attracted more attention in computer vision, where only very few training examples are presented during model learning process. A commonly-overlooked issue in FSOD is that novel classes are usually classified as background clutters in the pre-training process. Another difficulty of FSOD is that the detection performance degrades especially under higher IoU thresholds since previous deep metric learning (DML) requires frozen region proposals without class-relevant box regression. In this work, we propose a Class-relevant Region Proposal Network (CRPN). The CRPN can derive network parameters for novel classes from pre-trained convolution kernels according to their feature similarity, which is used to eliminate the above mentioned adverse effects and improve the performance of few-shot object detection. The proposed CPRN is able to kill two birds with one stone and has two main contributions: (1) transfer a region proposal network pre-trained on base classes to novel classes; (2) perform class-dependent bounding-box regression which previous DML classifier lacks. For experimental testing, we achieve 12.7% AP75 in MS COCO dataset and 28.6% AP75 in ImageNet2015 dataset under the few-shot setting introduced by previous works, which exceeds the state-of-the-art by a certain margin. Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2022 | Hybrid Physical Metric For 6-DoF Grasp Pose Detectionabstract6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric usually generates several discrete levels of grasp confidence scores, which cannot finely distinguish millions of grasp poses and leads to inaccurate prediction results. In this paper, we propose a hybrid physical metric to solve this evaluation insufficiency. First, we define a novel metric is based on the force-closure metric, supplemented by the measurement of the object flatness, gravity and collision. Second, we leverage this hybrid physical metric to generate elaborate confidence scores. Third, to learn the new confidence scores effectively, we design a multi-resolution network called Flatness Gravity Collision GraspNet (FGC-GraspNet). FGC-GraspNet proposes a multi-resolution features learning architecture for multiple tasks and introduces a new joint loss function that enhances the average precision of the grasp detection. The network evaluation and adequate real robot experiments demonstrate the effectiveness of our hybrid physical metric and FGC-GraspNet. Our method achieves 90.5% success rate in real-world cluttered scenes. Our code is available at https://github.com/luyh20IFGC-GraspNet. Yuhao Lu, Beixing Deng, Zhenyu Wang 0005, Peiyuan Zhi, Yali Li 0001, Shengjin Wang |
ICRA | 5 |
| 2022 | Self-Supervised Learning via Maximum Entropy CodingabstractA mainstream type of current self-supervised learning methods pursues a general-purpose representation that can be well transferred to downstream tasks, typically by optimizing on a given pretext task such as instance discrimination. In this work, we argue that existing pretext tasks inevitably introduce biases into the learned representation, which in turn leads to biased transfer performance on various downstream tasks. To cope with this issue, we propose Maximum Entropy Coding (MEC), a more principled objective that explicitly optimizes on the structure of the representation, so that the learned representation is less biased and thus generalizes better to unseen downstream tasks. Inspired by the principle of maximum entropy in information theory, we hypothesize that a generalizable representation should be the one that admits the maximum entropy among all plausible representations. To make the objective end-to-end trainable, we propose to leverage the minimal coding length in lossy data coding as a computationally tractable surrogate for the entropy, and further derive a scalable reformulation of the objective that allows fast computation. Extensive experiments demonstrate that MEC learns a more generalizable representation than previous methods based on specific pretext tasks. It achieves state-of-the-art performance consistently on various downstream tasks, including not only ImageNet linear probe, but also semi-supervised classification, object detection, instance segmentation, and object tracking. Interestingly, we show that existing batch-wise and feature-wise self-supervised objectives could be seen equivalent to low-order approximations of MEC. Code and pre-trained models are available at https://github.com/xinliu20/MEC. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
NeurIPS | 3 |
| 2022 | Robust 3D face modeling and tracking from RGB-D images
Changwei Luo, Juyong Zhang, Changcun Bao, Yali Li 0001, Shengjin Wang |
Multim. Syst. | 4 |
| 2022 | BooDet: Gradient Boosting Object Detection With Additive Learning-Based Prediction AggregationabstractIn recent years, the community of object detection has witnessed remarkable progress with the development of deep neural networks. But the detection performance still suffers from the dilemma between complex networks and single-vector predictions. In this paper, we propose a novel approach to boost the object detection performance based on aggregating predictions. First, we propose a unified module with adjustable hyper-structure to generate multiple predictions from a single detection network. Second, we formulate the additive learning for aggregating predictions, which reduces the classification and regression losses by progressively adding the prediction values. Based on the gradient Boosting strategy, the optimization of the additional predictions is further modeled as weighted regression problems to fit the Newton-descent directions. By aggregating multiple predictions from a single network, we propose the BooDet approach which can Bootstrap the classification and bounding box regression for high-performance object Detection. In particular, we plug the BooDet into Cascade R-CNN for object detection. Extensive experiments show that the proposed approach is quite effective to improve object detection. We obtain a 1.3%~2.0% improvement over the strong baseline Cascade R-CNN on COCO val dataset. We achieve 56.5% AP on the COCO test-dev dataset with only bounding box annotations. Yali Li 0001, Shengjin Wang |
IEEE Trans. Image Process. | 1 |
| 2022 | Traffic Sign Recognition With Lightweight Two-Stage Model in Complex ScenesabstractTraffic sign recognition with high accuracy and real-time is an important part of the intelligent transportation system. In this article, based on large-scale traffic signs and the inherent conflict between location regression and classification of traffic signs, we propose a novel and flexible two-stage approach. It combines a lightweight superclass detector with a refinement classifier. The main contributions lie in three aspects: (1) We use locations and sizes of signs as prior knowledge to establish a probability distribution model. It can significantly decrease the search range of signs and improve the processing speed, as well as reducing false detection. (2) We propose a high-performance lightweight superclass detector. We introduce the Inception and Channel Attention, by generating multi-scale receptive fields and adaptively adjusting channel features. It alleviates the large scale variance challenge of objects and the interference of background information. Meanwhile, we present a merging Batch Normalization and multi-scale testing method to further improve detection performance. (3) We propose a refinement classifier based on similarity measure learning for the subclass classification. It increases the precision of discriminating similar subclasses and also improves the extensibility of our approach. Our two-stage approach is simple and effective, whose paradigm is different from others. Experiments on the Tsinghua-Tencent 100K dataset demonstrate the performance of our approach. Compared with the state-of-the-art methods, our method achieves competitive performance (92.16% mAP) with a lightweight detector ($6.49M $). The processing time is$0.150s $per frame, of which the speed is increased by 3 times compared with existing methods. Zhengshuai Wang, Yali Li 0001, Shengjin Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Sequence Model with Self-Adaptive Sliding Window for Efficient Spoken Document SegmentationabstractTranscripts generated by automatic speech recognition (ASR) systems for spoken documents lack structural annotations such as paragraphs, significantly reducing their readability. Automatically predicting paragraph segmentation for spoken documents may both improve readability and downstream NLP performance such as summarization and machine reading comprehension. We propose a sequence model with self-adaptive sliding window for accurate and efficient paragraph segmentation. We also propose an approach to exploit pho-netic information, which significantly improves robustness of spoken document segmentation to ASR errors. Evaluations are conducted on the English Wiki-727K document seg-mentation benchmark, a Chinese Wikipedia-based document segmentation dataset we created, and an in-house Chinese spoken document dataset. Our proposed model outperforms the state-of-the-art (SOTA) model based on the same BERT-Base, increasing segmentation F1 on the English benchmark by 4.2 points and on Chinese datasets by 4.3-10.1 points, while reducing inference time to less than 1/6 of inference time of the current SOTA. Qian Chen 0003, Yali Li 0001, Jiaqing Liu, Wen Wang 0001 |
ASRU | 3 |
| 2021 | A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance SegmentationabstractLearning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit FPN from further aggregating more discriminative features. In this work, we propose Attention Aggregation based Feature Pyramid Network (A2-FPN), to improve multi-scale feature learning through attention-guided feature aggregation. In feature extraction, it extracts discriminative features by collecting-distributing multi-level global context features, and mitigates the semantic information loss due to drastically reduced channels. In feature fusion, it aggregates complementary information from adjacent features to generate location-wise reassembly kernels for content-aware sampling, and employs channel-wise reweighting to enhance the semantic consistency before element-wise addition. A2-FPN shows consistent gains on different instance segmentation frameworks. By replacing FPN with A2-FPN in Mask R-CNN, our model boosts the performance by 2.1% and 1.6% mask AP when using ResNet-50 and ResNet-101 as backbone, respectively. Moreover, A2-FPN achieves an improvement of 2.0% and 1.4% mask AP when integrated into the strong baselines such as Cascade Mask R-CNN and Hybrid Task Cascade. Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 2 |
| 2021 | Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object DetectionabstractIn this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy labels, thus are deficient in learning different unlabeled knowledge well. To address this issue, we propose a data-uncertainty guided multi-phase learning method for semisupervised object detection. We comprehensively consider divergent types of unlabeled images according to their difficulty levels, utilize them in different phases, and ensemble models from different phases together to generate ultimate results. Image uncertainty guided easy data selection and region uncertainty guided RoI Re-weighting are involved in multi-phase learning and enable the detector to concentrate on more certain knowledge. Through extensive experiments on PASCAL VOC and MS COCO, we demonstrate that our method behaves extraordinarily compared to baseline approaches and outperforms them by a large margin, more than 3% on VOC and 2% on COCO. Zhenyu Wang 0005, Yali Li 0001, Lu Fang 0001, Shengjin Wang |
CVPR | 2 |
| 2021 | Disentangled Representation for Age-Invariant Face Recognition: A Mutual Information Minimization PerspectiveabstractGeneral face recognition has seen remarkable progress in recent years. However, large age gap still remains a big challenge due to significant alterations in facial appearance and bone structure. Disentanglement plays a key role in partitioning face representations into identity-dependent and age-dependent components for age-invariant face recognition (AIFR). In this paper we propose a multi-task learning framework based on mutual information minimization (MT-MIM), which casts the disentangled representation learning as an objective of information constraints. The method trains a disentanglement network to minimize mutual information between the identity component and age component of the face image from the same person, and reduce the effect of age variations during the identification process. For quantitative measure of the degree of disentanglement, we verify that mutual information can represent as metric. The resulting identity-dependent representations are used for age-invariant face recognition. We evaluate MT-MIM on popular public-domain face aging datasets (FG-NET, MORPH Album 2, CACD and AgeDB) and obtained significant improvements over previous state-of-the-art methods. Specifically, our method exceeds the baseline models by over 0.4% on MORPH Album 2, and over 0.7% on CACD subsets, which are impressive improvements at the high accuracy levels of above 99% and an average of 94%. Xuege Hou, Yali Li 0001, Shengjin Wang |
ICCV | 2 |
| 2021 | Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image CaptioningabstractHuman-oriented image captioning with both high diversity and accuracy is a challenging task in vision+language modeling. The reinforcement learning (RL) based frameworks promote the accuracy of image captioning, yet seriously hurt the diversity. In contrast, other methods based on variational auto-encoder (VAE) or generative adversarial network (GAN) can produce diverse yet less accurate captions. In this work, we devote our attention to promote the diversity of RL-based image captioning. To be specific, we devise a partial off-policy learning scheme to balance accuracy and diversity. First, we keep the model exposed to varied candidate captions by sampling from the initial state before RL launched. Second, a novel criterion named max-CIDEr is proposed to serve as the reward for promoting diversity. We combine the above-mentioned offpolicy strategy with the on-policy one to moderate the exploration effect, further balancing the diversity and accuracy for human-like image captioning. Experiments show that our method locates the closest to human performance in the diversity-accuracy space, and achieves the highest Pearson correlation as 0.337 with human performance. Jiahe Shi, Yali Li 0001, Shengjin Wang |
ICCV | 2 |
| 2021 | Combating Noise: Semi-supervised Learning by Region Uncertainty QuantificationabstractSemi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current methods are easily distracted by noisy regions generated by pseudo labels. To combat the noisy labeling, we propose noise-resistant semi-supervised learning by quantifying the region uncertainty. We first investigate the adverse effects brought by different forms of noise associated with pseudo labels. Then we propose to quantify the uncertainty of regions by identifying the noise-resistant properties of regions over different strengths. By importing the region uncertainty quantification and promoting multi-peak probability distribution output, we introduce uncertainty into training and further achieve noise-resistant learning. Experiments on both PASCAL VOC and MS COCO demonstrate the extraordinary performance of our method. Zhenyu Wang 0005, Yali Li 0001, Shengjin Wang |
NeurIPS | 2 |
| 2021 | Do Different Tracking Tasks Require Different Appearance Models?abstractTracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature has fragmented too, and now novel approaches proposed by the community are usually specialised to fit only one specific setup. To understand to what extent this specialisation is necessary, in this work we present UniTrack, a solution to address five different tasks within the same framework. UniTrack consists of a single and task-agnostic appearance model, which can be learned in a supervised or self-supervised fashion, and multiple ``heads'' that address individual tasks and do not require training. We show how most tracking tasks can be solved within this framework, and that the same appearance model can be successfully used to obtain results that are competitive against specialised methods for most of the tasks considered. The framework also allows us to analyse appearance models obtained with the most recent self-supervised methods, thus extending their evaluation and comparison to a larger variety of important problems. Zhongdao Wang, Hengshuang Zhao, Yali Li 0001, Shengjin Wang, Philip Torr 0001, Luca Bertinetto |
NeurIPS | 3 |
| 2021 | Learning Part-based Convolutional Features for Person Re-IdentificationabstractPart-level features offer fine granularity for pedestrian image description. In this article, we generally aim to learn discriminative part-informed feature for person re-identification. Our contribution is two-fold. First, we introduce a general part-level feature learning method, named Part-based Convolutional Baseline (PCB). Given an image input, it outputs a convolutional descriptor consisting of several part-level features. PCB is general in that it is able to accommodate several part partitioning strategies, including pose estimation, human parsing and uniform part partitioning. In experiment, we show that the learned descriptor has a significantly higher discriminative ability than the global descriptor. Second, based on PCB, we propose refined part pooling (RPP), which allows the parts to be more precisely located. Our idea is that pixels within a well-located part should be similar to each other while being dissimilar with pixels from other parts. We call it within-part consistency. When a pixel-wise feature vector in a part is more similar to some other part, it is then an outlier, indicating inappropriate partitioning. RPP re-assigns these outliers to the parts they are closest to, resulting in refined parts with enhanced within-part consistency. RPP requires no part labels and is trained in a weakly supervised manner. Experiment confirms that RPP allows PCB to gain another round of performance boost. For instance, on the Market-1501 dataset, we achieve (77.4+4.2) percent mAP and (92.3+1.5) percent rank-1 accuracy, a competitive performance with the state of the art. Yifan Sun 0003, Liang Zheng 0001, Yali Li 0001, Yi Yang 0001, Qi Tian 0001, Shengjin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Learning to Segment Video Object With Accurate BoundariesabstractVideo object segmentation has attracted considerable research interest these years. Top-performing video object segmentation methods mainly rely on fully convolutional neural networks which are specifically trained for predicting high-performance masks, resulting in a lack of preciseness in boundary details. This paper tackles the problem of predicting both mask-accurate and boundary-precise segmentation masks in videos. To solve this problem, we propose a simple and efficient network structure: the Mask-boundAry-Consistent Network (MAC-Net). TheMAC-Netis an end-to-end fully convolutional network, where both mask and boundaries are jointly optimized during training, enabling it to predict masks along with accurate boundaries. An inner-net boundary-computing module is incorporated in theMAC-Netfor producing spontaneously mask-consistent boundaries. We analyze the influence of parameter settings, network constructions of theMAC-Net, and compare with state-of-the-art algorithms on three widely-adopted datasets. Experimental results show that theMAC-Netachieves state-of-the-art performance, demonstrating the effectiveness of its mask-boundary-consistent network structure. We also propose that the boundary module inMAC-Nethas high compatibility, and can be easily adapted to other segmentation-related techniques. Jingchun Cheng, Yuhui Yuan, Yali Li 0001, Jingdong Wang 0001, Shengjin Wang |
IEEE Trans. Multim. | 3 |
| 2020 | Softmax Dissection: Towards Understanding Intra- and Inter-Class Objective for Embedding LearningabstractThe softmax loss and its variants are widely used as objectives for embedding learning applications like face recognition. However, the intra- and inter-class objectives in Softmax are entangled, therefore a well-optimized inter-class objective leads to relaxation on the intra-class objective, and vice versa. In this paper, we propose to dissect Softmax into independent intra- and inter-class objective (D-Softmax) with a clear understanding. It is straightforward to tune each part to the best state with D-Softmax as objective.Furthermore, we find the computation of the inter-class part is redundant and propose sampling-based variants of D-Softmax to reduce the computation cost. The face recognition experiments on regular-scale data show D-Softmax is favorably comparable to existing losses such as SphereFace and ArcFace. Experiments on massive-scale data show the fast variants significantly accelerates the training process (such as 64×) with only a minor sacrifice in performance, outperforming existing acceleration methods of Softmax in terms of both performance and efficiency. Lanqing He, Zhongdao Wang, Yali Li 0001, Shengjin Wang |
AAAI | 3 |
| 2020 | Video Super-Resolution With Temporal Group AttentionabstractVideo super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is divided into several groups, with each one corresponding to a kind of frame rate. These groups provide complementary information to recover missing details in the reference frame, which is further integrated with an attention module and a deep intra-group fusion module. In addition, a fast spatial alignment is proposed to handle videos with large motion. Extensive results demonstrate the capability of the proposed model in handling videos with various motion. It achieves favorable performance against state-of-the-art methods on several benchmark datasets. Takashi Isobe, Songjiang Li, Xu Jia 0012, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Yali Li 0001, Shengjin Wang, Qi Tian 0001 |
CVPR | 7 |
| 2020 | Towards Real-Time Multi-Object Tracking
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yali Li 0001, Shengjin Wang |
ECCV (11) | 4 |
| 2020 | CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions
Zhongdao Wang, Liang Zheng 0001, Yixuan Liu 0004, Yifan Sun 0003, Yali Li 0001, Shengjin Wang |
ECCV (11) | 6 |
| 2020 | CS-R-FCN: Cross-Supervised Learning for Large-Scale Object DetectionabstractGeneric object detection is one of the most fundamental problems in computer vision, yet it is difficult to provide all the bounding-box-level annotations aiming at large-scale object detection for thousands of categories. In this paper, we present a novel cross-supervised learning pipeline for large-scale object detection, denoted as CS-R-FCN. First, we propose to utilize the data flow of image-level annotated images in the fully-supervised two-stage object detection framework, leading to cross-supervised learning combining bounding-box-level annotated data and image-level annotated data. Second, we introduce a semantic aggregation strategy utilizing the relationships among the cross-supervised categories to reduce the unreasonable mutual inhibition effects during the feature learning. Experimental results show that the proposed CS-R-FCN improves the mAP by a large margin compared to previous related works. Yali Li 0001, Shengjin Wang |
ICASSP | 2 |
| 2020 | Intra-Clip Aggregation For Video Person Re-IdentificationabstractVideo-based person re-identification has drawn massive attention in recent years due to its extensive applications in video surveillance. While deep learning based methods have led to significant progress, these methods are limited by ineffectively using complementary information, which is blamed on necessary data augmentation in training process. Data augmentation has been widely used to mitigate the overfitting trap and improve the ability of network representation. However, the previous methods adopt image-based data augmentation scheme to individually process the input frames, which corrupts the complementary information between consecutive frames and causes performance degradation. In this paper, we propose a novel video-based data augmentation scheme, termed as Synchronous Data Augmentation, to address the challenge above. In order to represent discriminative clip-level features, we also propose a cascade integration module which hierarchically aggregates the intra-clip features with a linear-nonlinear combining projection. Extensive experiments on three benchmark datasets demonstrate that our framework outperforms the most recent state-of-the-art methods. We also perform cross-dataset validation to prove the generality of our method. Takashi Isobe, Yali Li 0001, Shengjin Wang |
ICIP | 4 |
| 2020 | Fianet: Video Object Detection Via Joint Feature-Level And Instance-Level AggregationabstractVideo object detection task is challenging due to the nonrigid and rigid appearance deformations in videos. Most of the typical competitive methods are to enhance per-frame features through aggregating lots of previous and future frames. But feature-level aggregation isn't robust to rigid deformations such as occlusion and rare postures. In this paper, we propose an online video object detection method with joint feature-level aggregation and instance-level aggregation network (FIANet). Besides feature-level aggregation, we design a spatial-temporal instance calibration module (STIC) to aggregate the instance as a whole, which can reduce the interference of local distorted and missed pixels. Joint featurelevel and instance-level aggregation can work collaboratively to overcome different deformations. Only using less previous frames, our method can achieve 81.6% mAP with relatively high speed on ImageNet VID, which is state-of-the-art compared with causal and non-causal methods. Zhengshuai Wang, Yali Li 0001, Shengjin Wang |
ICME | 2 |
| 2020 | Progressive Representation Adaptation for Weakly Supervised Object LocalizationabstractWe address the problem of weakly supervised object localization where only image-level annotations are available for training object detectors. Numerous methods have been proposed to tackle this problem through mining object proposals. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to undesirable local minimum solutions. In this paper, we propose to overcome these drawbacks by progressive representation adaptation with two main steps: 1) classification adaptation and 2) detection adaptation. In classification adaptation, we transfer a pre-trained network to a multi-label classification task for recognizing the presence of a certain object in an image. Through the classification adaptation step, the network learns discriminative representations that are specific to object categories of interest. In detection adaptation, we mine class-specific object proposals by exploiting two scoring strategies based on the adapted classification network. Class-specific proposal mining helps remove substantial noise from the background clutter and potential confusion from similar objects. We further refine these proposals using multiple instance learning and segmentation cues. Using these refined object bounding boxes, we fine-tune all the layer of the classification network and obtain a fully adapted detection network. We present detailed experimental validation on the PASCAL VOC and ILSVRC datasets. Experimental results demonstrate that our progressive representation adaptation algorithm performs favorably against the state-of-the-art methods. Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Adaptively Leverage Unlabeled Tracklets Based on Part Attention Model for Few-Example Re-IDabstractFew-example learning for video person re-ID is a challenging issue. Some studies use large unlabeled samples to mine more discriminative cues to overcome the visual scarcity. But how to develop a robust model to avoid overfitting and overcome noisy labels is still a remained problem. In this letter we focus on bridging between few labeled tracklets and numerous unlabeled tracklets. We leverage unlabeled tracklets by generating pseudo labels and adaptively joins them into the training set which consists of the labeled and unlabeled data. This work is distinguished by two key contributions. First, a novel model PAM (Part Attention Model) tailored for progressive learning is proposed. It has high accuracy and fast convergence. Second, we propose a novel sampling strategy ARS (Adaptively Relative Sampling). ARS adaptively filters out noisy labels and enlarges the training set. ARS-PAM achieves significant performance gains on four mainstream datasets. On the PRID2011 and iLIDS-VID dataset, ARS-PAM reaches 89.8%, 56.1% rank-1 accuracy, which exceeds the state-of-the-art by 5.5%, 17.5% respectively. Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2020 | Node-Adaptive Multi-Graph Fusion Using Extreme Value TheoryabstractThis letter considers the problem of grouping data by their underlying categories with inputs from multiple sources, known as the multi-view clustering problem. One of the most fundamental challenges lies in how to benefit from the complementary information in the multi-view data, so that clustering on such data consistently achieves higher accuracy than clustering on each single-view component. In this letter, to tackle the multi-view clustering problem, we propose a novel approach to fuse multiple affinity graphs computed in each single view to a unified affinity graph, so that single-view affinity-based clustering methods can be accordingly applied on it. The edges in the unified affinity graph between a node and its neighbors are computed as weighted average over the corresponding edges from multiple single graphs, and the weights here are adaptive to each node, estimated using the Extreme Value Theory (EVT). Experiments on two challenging multi-view clustering tasks show that, combined with existing off-the-shelf single-view clustering algorithms, the proposed graph fusion method brings consistently performance gain compared with naive graph fusion baselines. Zhongdao Wang, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 3 |
| 2020 | Unidirectional Representation-Based Efficient Dictionary LearningabstractDictionary learning (DL) has been widely studied for pattern classification. Most existing methods introduce multiple discriminative terms into objective functions for accuracy improvement, leading to complex learning frameworks and high computational burdens. This paper proposes a simple yet effective DL algorithm for classification, namely unidirectional representation dictionary learning (URDL). Unidirectional constraint is proposed to guide coefficient directions in the representation to be discriminative. Besides, direction-thresholding is proposed to exploit the direction property in the classification scheme. It suppresses the disturbance from undesired non-zero coefficients, and improves the representation discriminability. We adopt squared ℓ2-norm-based regularization for efficient coding, and systematically analyze the mechanism of the proposed method. Extensive experiments on five data sets are conducted, including object categorization, scene classification, face recognition, and fine-grained flower classification. The experimental results demonstrate that the proposed approach not only outperforms the state-of-the-art DL algorithms in terms of recognition accuracy significantly, but also exhibits a much higher computational efficiency. Xiudong Wang, Yali Li 0001, Shaodi You, Hongdong Li, Shengjin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | HAR-Net: Joint Learning of Hybrid Attention for Single-Stage Object DetectionabstractObject detection has been a challenging task in computer vision. Although significant progress has been made in object detection with deep neural networks, the attention mechanism has yet to be fully developed. In this paper, we propose a hybrid attention mechanism for single-stage object detection. First, we present the modules of spatial attention, channel attention and aligned attention for single-stage object detection. In particular, dilated convolution layers with symmetrically fixed rates are stacked to learn spatial attention. A channel attention mechanism with the cross-level group normalization and squeeze-and-excitation operation is proposed. Aligned attention is constructed with organized deformable filters. Second, the three types of attention are unified to construct the hybrid attention mechanism. We then plug the hybrid attention into Retina-Net and propose the efficient single-stage HAR-Net for object detection. The attention modules and the proposed HAR-Net are evaluated on the COCO detection dataset. The experiments demonstrate that hybrid attention can significantly improve the detection accuracy and that the HAR-Net can achieve a state-of-the-art 45.8% mAP, thus outperforming existing single-stage object detectors. Yali Li 0001, Shengjin Wang |
IEEE Trans. Image Process. | 1 |
| 2020 | Fast Pedestrian Detection With Attention-Enhanced Multi-Scale RPN and Soft-Cascaded Decision TreesabstractPedestrian detection has attracted more attention in the fields of computer vision and artificial intelligence. A variety of real-world applications involving pedestrian detection have been promoted, such as Advanced Driving Assistant System (ADAS). Although both two-stage and single-stage deeply learned object detectors have shown outstanding performance for general object detection, they are still facing the problem of poor accuracy in single-class detection senario because they are designed to distinguish objects from different categories rather than pay attention to various appearances of pedestrians. Previous leading pedestrian detectors F-DNN and F-DNN v2 fuse several neural networks like SSD, VGG16 and GoogLeNet to generate ROIs and supress false alarms with cascaded structure, resulting in low miss rate but high complexity. In this paper we propose a novel framework called Attention-Enhanced Multi-Scale Region Proposal Network (AEMS-RPN) for ROI generation, which also acts as first-stage classification. Inspired by the success of traditional pedestrian detectors, we use soft-cascaded decision trees instead of cascaded deep neural networks to achieve high accuracy and fast detection speed simultaneously. The decision tree classifier is used and enables us to combine features from different layers with various resolutions for classification and incorporate effective bootstrapping for mining hard negatives. We test our method on several pedestrian detection datasets and the experimental results certify the effectiveness of the proposed AEMS-RPN. Compared with the state-of-the-art, we obtain the competitive accuracy with near real-time efficiency. Yali Li 0001, Shengjin Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Perceive Where to Focus: Learning Visibility-Aware Part-Level Features for Partial Person Re-IdentificationabstractThis paper considers a realistic problem in person re-identification (re-ID) task, i.e., partial re-ID. Under partial re-ID scenario, the images may contain a partial observation of a pedestrian. If we directly compare a partial pedestrian image with a holistic one, the extreme spatial misalignment significantly compromises the discriminative ability of the learned representation. We propose a Visibility-aware Part Model (VPM) for partial re-ID, which learns to perceive the visibility of regions through self-supervision. The visibility awareness allows VPM to extract region-level features and compare two images with focus on their shared regions (which are visible on both images). VPM gains two-fold benefit toward higher accuracy for partial re-ID. On the one hand, compared with learning a global feature, VPM learns region-level features and thus benefits from fine-grained information. On the other hand, with visibility awareness, VPM is capable to estimate the shared regions between two images and thus suppresses the spatial misalignment. Experimental results confirm that our method significantly improves the learned feature representation and the achieved accuracy is on par with the state of the art. Yifan Sun 0003, Yali Li 0001, Chi Zhang 0026, Shengjin Wang, Jian Sun 0001 |
CVPR | 3 |
| 2019 | Linkage Based Face Clustering via Graph Convolution NetworkabstractIn this paper, we present an accurate and scalable approach to the face clustering task. We aim at grouping a set of faces by their potential identities. We formulate this task as a link prediction problem: a link exists between two faces if they are of the same identity. The key idea is that we find the local context in the feature space around an instance (face) contains rich information about the linkage relationship between this instance and its neighbors. By constructing sub-graphs around each instance as input data, which depict the local context, we utilize the graph convolution network (GCN) to perform reasoning and infer the likelihood of linkage between pairs in the sub-graphs. Experiments show that our method is more robust to the complex distribution of faces than conventional methods, yielding favorably comparable results to state-of-the-art methods on standard face clustering benchmarks, and is scalable to large datasets. Furthermore, we show that the proposed method does not need the number of clusters as prior, is aware of noises and outliers, and can be extended to a multi-view version for more accurate clustering accuracy. Zhongdao Wang, Liang Zheng 0001, Yali Li 0001, Shengjin Wang |
CVPR | 3 |
| 2019 | Intention Oriented Image Captions With Guiding ObjectsabstractAlthough existing image caption models can produce promising results using recurrent neural networks (RNNs), it is difficult to guarantee that an object we care about is contained in generated descriptions, for example in the case that the object is inconspicuous in the image. Problems become even harder when these objects did not appear in training stage. In this paper, we propose a novel approach for generating image captions with guiding objects (CGO). The CGO constrains the model to involve a human-concerned object when the object is in the image. CGO ensures that the object is in the generated description while maintaining fluency. Instead of generating the sequence from left to right, we start the description with a selected object and generate other parts of the sequence based on this object. To achieve this, we design a novel framework combining two LSTMs in opposite directions. We demonstrate the characteristics of our method on MSCOCO where we generate descriptions for each detected object in the images. With CGO, we can extend the ability of description to the objects being neglected in image caption labels and provide a set of more comprehensive and diverse descriptions for an image. CGO shows advantages when applied to the task of describing novel objects. We show experimental results on both MSCOCO and ImageNet datasets. Evaluations show that our method outperforms the state-of-the-art models in the task with average F1 75.8, leading to better descriptions in terms of both content accuracy and fluency. Yali Li 0001, Shengjin Wang |
CVPR | 2 |
| 2019 | Parallel-Structure-based Transfer Learning for Deep NIR-to-VIS Face Recognition
Yali Li 0001, Shengjin Wang |
ICIG (1) | 2 |
| 2019 | Cascade Attention: Multiple Feature Based Learning for Image CaptioningabstractMost recent researches in image captioning adopt attention mechanism based on encoder-decoder framework, where the attention module aligns input features for the decoder and boosts performance consequently. A common defect of traditional attention methods is that the inequality among different types of inputs is ignored, resulting in under-exploitation of certain informative features. In this paper, we propose a novel cascade attention module, which processes different types of input in a sequential manner. The cascade attention module enables inputs of higher priorities to affect the attention of other inputs so as to emphasize such inequality. We implement our model by introducing global feature of the image to the captioning process of R-CNN based frameworks, where such feature is rich of context information but takes few effects via traditional attention module. Experimental results demonstrate that our proposed method is able to exploit feature of different types, acquiring improvements on multiple automatic measurements. Jiahe Shi, Yali Li 0001, Shengjin Wang |
ICIP | 2 |
| 2019 | Secondary Information Aware Facial Expression RecognitionabstractFacial expression recognition (FER) is a key factor in human behavior analysis. Most algorithms deal with FER as a pure classification problem, assuming that expressions are exclusive to each other. In this letter, the problem of FER is tackled from a more detailed view: learning to discriminate expressions with consideration of the secondary information. We propose the Secondary Information aware Facial Expression Network (SIFE-Net) to explore the latent components without auxiliary labeling, and we propose a novel dynamic weighting strategy to teach the SIFE-Net. In contrast to traditional classifiers trained with one-hot labels, the proposed SIFE-Net takes advantage of secondary expression information and has more rational feature distributions. We carry out extensive experiments and analysis on three widely-used FER datasets, i.e. the CK+ dataset, the JAFFE dataset, and the RAF dataset. Experimental results show that the SIFE-Net achieves state-of-the-art performance on all three datasets, which demonstrates the effectiveness of our method. Ye Tian 0019, Jingchun Cheng, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 3 |
| 2019 | Open-World Person Re-Identification With Deep Hash Feature EmbeddingabstractMost existing person re-identification (re-id) methods are designed based on the artificial closed-set assumption that the probe and gallery identities are exactly overlapped with a small search pool. This leads to poor scalability in real-world applications where the task is often to re-id a small set of target people (i.e., watch-list) among a large search pool with unknown ID overlap, namely, an open-set deployment setting. In this paper, we firstly propose a new person re-id setting called Watch-List based Open-Set (WLOS) person re-id, which is characterised by the above open-set deployment and a watch-list available at the training stage. Then, we address such a under-studied WLOS problem by formulating a novel Task Dedicated Deep Hashing (TDDH) approach which learning a purpose-specific deep hash model particularly for the given target people in an efficient end-to-end manner. Extensive experiments on three large-scale re-id benchmarks are conducted to demonstrate the advantages and superiority of the TDDH over a wide range of the state-of-the-art hashing and re-id methods under the more realistic open-set setting. Yali Zhao, Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2019 | Unsupervised Deep Hashing With Adaptive Feature Learning for Image RetrievalabstractThe hashing method is widely used for large-scale image retrieval due to its low time and space complexity. However, the existing deep hashing methods are mainly designed for labeled datasets. Without supervised information, retrieval performance on unlabeled datasets is not guaranteed. In this letter, we propose a novel deep hashing approach for unsupervised image retrieval applications. The contributions are two-fold. First, the pseudolabels are generated using their global features aggregated from the pretrained network and employed as self-supervised information to optimize the objective function of training. Second, adaptive feature learning is used in this deep hashing framework to perform simultaneous hash function learning and global features learning in an unsupervised manner. The experimental results validated the effectiveness of the proposed method, obtaining state-of-the-art performances on several public datasets such as CIFAR-10, Holidays, and Oxford5k. Yali Li 0001, Shengjin Wang |
IEEE Signal Process. Lett. | 2 |
| 2018 | A Highly Accurate Feature Fusion Network For Vehicle Detection In Surveillance Scenarios
Yali Li 0001, Shengjin Wang |
BMVC | 2 |
| 2018 | Feature Learning for One-Shot Face RecognitionabstractOne-shot face recognition is a challenging open problem which requires recognizing novel identities from only one gallery face. One-shot classes are squeezed and neglected in the feature space for classification due to data imbalance. Moreover, training samples deficience is a major obstacle to intra-class clustering. In this paper, we propose a novel framework based on CNN of balancing regularizer and shifting center regeneration which regulates norms of weight vector into same scale and adjusts clustering center to deal with deficient training data. Comprehensive evaluations on MS-celeb-1M low-shot face dataset demonstrate that our methods improve one-shot face recognition notablely which achieve 88.78% coverage at precision=0.99 using restricted data without hybrid classifiers or multi-model. Moreover, experiments on LFW prove that CNN model trained with proposed methods can obtain more discriminative and compact feature representations. Since there are many identities that have only few training samples available online, our methods have great significance for improving data utilization and strengthening feature representation for face recognition. Lingxiao Wang 0009, Yali Li 0001, Shengjin Wang |
ICIP | 2 |
| 2018 | Attend and Align: Improving Deep Representations with Feature Alignment Layer for Person RetrievalabstractIn fine-grained recognition, object misalignment and background noise are two long-standing factors that influence the robustness of deep learning models. This paper mainly focuses on person re-identification (re-ID) and introduces a feature alignment layer (FAL) which alleviates the target misalignment and the background noise simultaneously. Through attention mechanism, FAL informs the underlying importance of each pixel on feature maps, i.e., whether the pixel is beneficial towards discriminating different persons. Then the discriminative regions relocate to the center and are stretched to fill the feature maps. Such an “attend and align” mechanism is specified into two steps: target position prediction and value assignment. In the first step, a pixel on feature maps learns to find a target position which is ID-discriminative. In the second step, the pixel is assigned with a new value using the context of the predicted position. Moreover, FAL can be easily plugged into a canonical Convolutional Neural Network (CNN) and learned in an end-to-end manner. In experiment, our method yields competitive results compared with the state-of-the-art approaches on three person re-ID datasets, Market-1501, DukeMTMC-reID and CUHK03. We also demonstrate that our method improves a competitive fine-grained recognition baseline on CUB-200-2011. Yifan Sun 0003, Yali Li 0001, Shengjin Wang |
ICPR | 3 |
| 2017 | A highly accurate facial region network for unconstrained face detectionabstractIn this paper, a new face detection method with very high accuracy is proposed. We introduce a novel facial region network to detect faces in unconstrained conditions. Firstly, a face proposal net is raised to generate possible face regions in the input image. Then, a novel weighted grid feature is applied to calculate features of face regions. Owing to that, faces with large pose variation and severe occlusion can be detected correctly. Furthermore, we use millions of general object data to pre-train the network to enhance the robustness of the extracted feature. Our method is evaluated on several public face detection datasets and achieves state-of-the-art performance on all of them. Specially, our method demonstrates a very high recall rate of 96.4% when false positives are 300 on the challenging FDDB benchmark, ranking first not only in the academic list but also in the commercial list which is much more competitive than the previous one. Han Shu, Dangdang Chen, Yali Li 0001, Shengjin Wang |
ICIP | 3 |
| 2016 | Weakly Supervised Object Localization with Progressive Domain AdaptationabstractWe address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to an undesirable local minimum. In this paper, we address this problem by progressive domain adaptation with two main steps: classification adaptation and detection adaptation. In classification adaptation, we transfer a pre-trained network to our multi-label classification task for recognizing the presence of a certain object in an image. In detection adaptation, we first use a mask-out strategy to collect class-specific object proposals and apply multiple instance learning to mine confident candidates. We then use these selected object proposals to fine-tune all the layers, resulting in a fully adapted detection network. We extensively evaluate the localization performance on the PASCAL VOC and ILSVRC datasets and demonstrate significant performance improvement over the state-of-the-art methods. Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2016 | Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signalabstractCommon Spatial Pattern(CSP) is a widely used algorithm in BCI application. However, it is sensitive to noise and artifact. In this paper, we propose a bagging regularized common spatial pattern (Bagging RCSP) approach for BCI with hybrid motor imagery and myoelectric signal. We divide the training samples into packets and choose training packets by Bagging to extract RCSP features. Furthermore, LDA is used to project the feature vector to lower space. In the end, a classification algorithm based on NNC is adopted. The Off-line experiment on BCI competition III attests Bagging RCSP versatile. The accuracy increases by 3%-5% in average than RCSP-A results. Furthermore, we designed and realized an online BCI system based on Bagging RCSP and evaluated through experiment involving four experimenters performing the BCI system of catching the apples. The results show the effectiveness of the proposed approach and the real time BCI system. Hongchuan Liu, Yali Li 0001, Hongma Liu, Shengjin Wang |
ICASSP | 2 |
| 2016 | A Boosting Approach to Exploit Instance Correlations for Multi-Instance ClassificationabstractWe propose a Boosting approach for multi-instance (MI) classification. Lp-norm is integrated to localize the witness instances and formulate the bag scores from classifier outputs. The contributions are twofold. First, a flexible and concise model for Boosting is proposed by the Lp-norm localization and exponential loss optimization. The scores for bag-level classification are directly fused from the instance feature space without probabilistic assumptions. Second, gradient and Newton descent optimizations are applied to derive the weak learners for Boosting. In particular, the instance correlations are exploited by fitting the weights and Newton updates for the weak learner construction. The final Boosted classifiers are the sums of iteratively chosen weak learners. Experiments demonstrate that the proposed Lp-norm-localized Boosting approach significantly improves the MI classification performance. Compared with the state of the art, the approach achieves the highest MI classification accuracy on 7/10 benchmark data sets. Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Selective parts for fine-grained recognitionabstractClassical visual bag-of-words approaches tackle fine-grained recognition using global features which discard spatial location of features. In this paper, we propose a novel part-based approach to distinguish fine-grained categories. This work is distinguished by two contributions. First, a fully automatic technique for selecting mid-level parts from large amounts of candidate regions without any part supervised information is presented. We call the selected parts by discriminative mining algorithm as selective parts. Second, a general effective evaluation criterion of quantifying part discriminability is built, which leads to joint selection process. For classification, feature ensembles are constructed based on global object and selective parts. Experimental results demonstrate the particular effectiveness of selective parts for fine-grained recognition on bird species on the Caltech UCSD Birds (CUB) dataset. Dong Li 0025, Yali Li 0001, Shengjin Wang |
ICIP | 2 |
| 2015 | A survey of recent advances in visual feature detection
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
Neurocomputing | 1 |
| 2015 | Feature representation for statistical-learning-based object detection: A review
Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
Pattern Recognit. | 1 |
| 2014 | Learning Cascaded Shared-Boost Classifiers for Part-Based Object DetectionabstractThis paper focuses on the problem of detecting a number of different class objects in images. We present a novel part-based model for object detection with cascaded classifiers. The coarse root and fine part classifiers are combined into the model. Different from the existing methods which learn root and part classifiers independently, we propose a shared-Boost algorithm to jointly train multiple classifiers. This paper is distinguished by two key contributions. The first is to introduce a new definition of shared features for similar pattern representation among multiple classifiers. Based on this, a shared-Boost algorithm which jointly learns multiple classifiers by reusing the shared feature information is proposed. The second contribution is a method for constructing a discriminatively trained part-based model, which fuses the outputs of cascaded shared-Boost classifiers as high-level features. The proposed shared-Boost-based part model is applied for both rigid and deformable object detection experiments. Compared with the state-of-the-art method, the proposed model can achieve higher or comparable performance. In particular, it can lift up the detection rates in low-resolution images. Also the proposed procedure provides a systematic framework for information reusing among multiple classifiers for part-based object detection. Yali Li 0001, Shengjin Wang, Qi Tian 0001, Xiaoqing Ding |
IEEE Trans. Image Process. | 1 |
| 2010 | Person-independent head pose estimation based on random forest regressionabstractIn this paper, a novel approach for person-independent head pose estimation in gray-level images is presented. There are two steps of the proposed method. In order to preserve similar patterns of faces under various poses, a novel multi-view face detector using tree-structured cascaded-Adaboost classifiers is applied. Furthermore, based on the cropped face images, randomized regression trees are learned and applied to estimate head pose precisely. Experiments show that our method achieves better pose estimation results in both horizontal and vertical orientations in comparison with the reported result with skin color information. Yali Li 0001, Shengjin Wang, Xiaoqing Ding |
ICIP | 1 |
| 2010 | Eye/eyes tracking based on a unified deformable template and particle filtering
Yali Li 0001, Shengjin Wang, Xiaoqing Ding |
Pattern Recognit. Lett. | 1 |